EconBase
← Back to paper

Welfare at Risk: Distributional impact of policy interventions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

99,804 characters · 18 sections · 106 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Welfare at Risk: Distributional impact of policy interventions

abstractThis paper proposes a framework for analyzing how the welfare effects of policy interventions are distributed across individuals when those effects are unobserved. Rather than focusing only on average outcomes, the approach uses readily available information on average welfare responses to uncover meaningful patterns in how gains and losses spread out across different populations. The framework is built around the concept of superquantiles and applies to a broad class of models with unobserved individual heterogeneity. It enables policymakers to identify which groups are most adversely affected by a policy and to evaluate trade-offs between efficiency and equity. We illustrate the approach in three widely studied economic settings: price changes and compensated variation, treatment allocation with self-selection, and the cost–benefit analysis of social programs. In this latter application, we show how standard tools from the marginal treatment effect and generalized Roy model literature are useful for implementing our bounds for both the overall population and for individuals who participate in the program.

{JEL classification: C35, C61, D90.} \\ {\bf Keywords:} Welfare, latent utility models, quantiles, superquantiles, compensated variation, treatment effect, generalized Roy models.

\thispagestyle{empty}

\textwidth 5.95in \textheight 600pt {1.1}

Introduction

Samuelson-Bergson welfare analysis of policy interventions typically focuses on average gains or losses, often overlooking their distributional consequences (e.g., bhattacharya2024nonparametric). Although mean welfare changes provide informative aggregate measures, they can mask substantial heterogeneity across individuals and groups. For instance, a policymaker may want to learn which subpopulations are most affected by a new intervention. Similarly, a policymaker may wish to know whether specific populations of interest are treated fairly in terms of utility gains. Addressing this limitation requires methods that capture distributional variation, such as quantile-based welfare analysis, heterogeneous-agent frameworks, or inequality measures such as the Atkinson index and the Gini coefficient.

A key requirement for assessing distributional heterogeneity in welfare analysis is knowledge of the distribution of individual welfare. Unfortunately, in many settings, this distribution is unknown to the analyst and cannot be inferred from available data. There are at least two reasons for this lack of identification. First, in many environments, the analyst observes only aggregate market data, so only average welfare changes can be identified (Berry1994,BerryLevinsohnPakes1995,Berry_Haile_ECMA2014). Second, even when individual (micro) data are available, outcomes are determined in a potential-outcomes framework in which each agent is observed in only one of two states—treated or untreated. This feature generates a fundamental identification problem: because the same individual cannot be observed in both states simultaneously, the individual welfare change, and hence its distribution, is unidentifiable (HECKMAN2007_Handbook_partI).

In this paper, we develop a distributional welfare framework that accounts for the heterogeneity in welfare generated by policy interventions. In particular, we derive quantile-based bounds that allow the analyst to learn about the distribution of individual welfare changes, even when this distribution is unobserved or unidentified.

Our approach builds on the perturbed utility model (PUM) of Allen_Rehbeck and McFadden_Fosgerau_2012, a latent choice model that accommodates general additive unobserved heterogeneity and nests several well-known frameworks, including the additive random utility model (ARUM) developed in mcfadden1972conditional, McFadden1978, mcf1.\footnote{Beyond ARUM, the PUM class also encompasses bundling models (Gentzkow_AER_2007; iaria2020identification; fox2017note), consumer choice models with latent constraints (agarwal2022demand), and demand systems with additive unobserved heterogeneity (brown2003strong; Matzkin_2007).} The ARUM framework is widely-used due to its tractability and its foundations in discrete choice theory, which facilitate counterfactual analysis (Berry1994; BerryLevinsohnPakes1995; Berry_Haile_ECMA2014). Furthermore, ARUMs play a central role in the treatment-effects literature, where self-selection is modeled as a discrete choice problem. This structure underpins the development of the marginal treatment effect (MTE), which makes explicit the roles of observables and unobservables in individuals’ decisions to participate in a treatment or social program (Bjorklund_Moffitt1987; Heckman_Vytlacil_ECMA_2005_MTE, HECKMAN2007_Handbook_partI; Heckman_Urzua_Vytlacil).

We consider environments in which the analyst observes only aggregate market data or estimates of conditional average welfare changes (such as average treatment effects). In such settings, the distribution of individual welfare changes is not identifiable, motivating the need for a framework that delivers informative bounds. As shown in Allen_Rehbeck, the generalizable PUM framework allows for the identification of conditional average treatment effects, a key requirement for our results.

Our framework relies on the concept of superquantiles, as developed in Rockafellar2000OptimizationOC, ROCKAFELLAR20021443 and Rockafellar_Royset. Superquantiles generalize quantiles by taking expectations over the distribution’s tail beyond a given quantile—for instance, the 20% superquantile reports the average welfare among the bottom 20%. They thus provide quantile-specific information while also summarizing average welfare across the distribution, yielding a richer view of distributional impacts.

Contributions

We make several contributions. First, Theorem (ref) shows that the superquantile of individual welfare changes—unobservable to the analyst—is bounded above by the superquantile of conditional average welfare changes, which is identified. This result highlights that information on conditional average welfare changes (or average treatment effects) provides valuable insights into the unobserved distribution of individual welfare changes. Economically, this allows the analyst to characterize, for any $\beta$-quantile, the average welfare gain or loss among the $(100\times \beta)\%$ worst affected by a policy change. More importantly, it enables the analyst to identify which subpopulations are most adversely affected and to assess the efficacy and fairness of alternative policy interventions. In addition, Theorem (ref) shows that, under additional assumptions regarding the distribution of individual welfare changes relative to conditional average welfare changes, the superquantile of individual welfare changes is also bounded below. Computationally, we can compute both bounds by solving linear programming problems.

Our second contribution applies the framework to three economic environments in which distributional welfare is central to policy analysis. The first application examines compensated variation (CV) arising from changes in goods prices. Within the PUM class, we derive bounds on the individual CV distributions across subpopulations and show that these bounds reveal heterogeneity in the unobserved CV distribution. More importantly, we demonstrate that conditional average CV — often far easier to obtain than individual-level data — can be used to extract meaningful information about the underlying individual CV distribution. From a data perspective, the analysis highlights that even aggregate market data can be informative about the distributional consequences of price changes, as captured by the CV.

In our second application, we study treatment allocation and welfare maximization when participation is endogenous due to self-selection. Building on Sasaki_Ura_2024, who focus on average welfare maximization, we show how our distributional framework can be combined with the MTE to evaluate the distributional welfare consequences of self-selection and unobserved heterogeneity across relevant populations and subgroups. In particular, we show that a planner with distributional objectives can identify which subpopulations are most adversely affected by a given treatment choice. A key insight of our analysis is that it disentangles the distinct roles of observable characteristics and unobserved heterogeneity in shaping welfare outcomes.

Moreover, our framework enables meaningful comparisons across policies based on their distributional welfare implications and the outcomes of worst-affected subpopulations. This perspective links our approach to regret-based criteria studied by Manski2004, Kiatagawa_Tetenov_2018, and Athey_Wagner_2021, while going beyond average-welfare comparisons by explicitly accounting for potential harm to vulnerable groups and the role of endogenous participation.

In our final application, we analyze social program evaluation in generalized Roy models with subjective participation costs (Heckman_Vytlacil_ECMA_2005_MTE). We extend the cost–benefit framework of Eisenhauer_Heckman_Vytlacil_2015 by incorporating explicit distributional considerations. Although the distributions of benefits, costs, and individual welfare are not directly observable, we show that our bounds are informative about each of these objects. In particular, the identification results in Eisenhauer_Heckman_Vytlacil_2015 can be leveraged to construct bounds that recover informative features of the unobserved benefit, cost, and welfare distributions.

Moreover, we derive bounds characterizing the distributions of costs, benefits, and welfare for agents who participate in the program, thereby providing distributional information for the treated (treatment on the treated). To the best of our knowledge, these results are novel within the context of generalized Roy models with participation costs.

The rest of the paper unfolds as follows. Section (ref) discusses the related literature. Section (ref) briefly presents the relevant aspects of the PUM framework and formalizes the notion of superquantiles, discussing its key properties. Section (ref) develops upper and lower bounds that allow us to learn about the distribution of individual-level welfare changes, which are often unidentified. Section (ref) examines the applications to compensated variation, welfare maximization, treatment allocation, and program evaluation with subjective participation costs. Section (ref) concludes.

Related literature

Welfare analysis in latent utility models has received significant attention. mcf1 introduced the general ARUM framework, demonstrating its aggregation properties and establishing the existence of a representative agent, making it well-suited for welfare analysis. Small_Rosen_1981 extended this by adapting traditional welfare economics methods to discrete choice models. Hanneman_1996 explored welfare changes in ARUM, emphasizing income effects, while Mcfadden_1996 discussed computational techniques to estimate mean CV changes. McFadden_Fosgerau_2012 further generalized demand and welfare analysis using the PUM framework. However, these studies primarily focus on average welfare and CV analysis, overlooking distributional aspects.

Much work has been done on non-parametric identification of latent utility models. In the context of the PUM model used in this paper, Allen_Rehbeck provides non-parametric identification methods for average welfare changes. Moreover, their results apply to aggregate data, a less demanding data requirement that our results also apply to. However, as with traditional welfare analysis, their results focus on average welfare measures. Therefore, our results are complementary to theirs, as we provide another tool for analyzing the distribution of welfare effects at the individual level that goes beyond their average analysis while leveraging their non-parametric identification results.

Bhattacharya_2015 studies welfare in non-additive random utility models (NARUMs) and identifies the marginal distributions of CV and EV using conditional choice probabilities. Although both their analysis and ours provide non-parametric tools for distributional welfare evaluation, the approaches differ in three main respects. First, their results are specific to NARUMs, whereas our framework applies to PUMs, allowing for richer demand patterns — including substitutability and complementarity. Second, Bhattacharya_2015 requires individual-level choice data, whereas our approach accommodates the standard setting in which the researcher observes only aggregate data (Berry_Haile_ECMA2014). We show that, by leveraging superquantiles, meaningful distributional welfare analysis remains feasible in this environment, making our results complementary to theirs. Third, our framework extends beyond demand to include treatment allocation, welfare maximization, and cost–benefit analysis in generalized Roy models. In the latter case, we show how our bounds deliver informative distributional insights for the subgroup that endogenously selects into treatment.

echenique2024utilitarian studies preference aggregation through Harsanyi’s utilitarian approach, establishing a foundation for distributional welfare measures based on quantiles of individual welfare effects. We differ in four key ways. First, rather than an axiomatic approach, we propose bounds that can be used when the researcher has aggregate data or access to a collection of estimates of conditional average welfare changes. Thus, we exploit available information to learn about the unknown distribution of individual welfare changes. Second, we leverage superquantiles to provide both quantile-specific welfare insights and the corresponding average welfare, thereby enhancing understanding of the welfare distribution. Third, whereas echenique2024utilitarian focuses on the ARUM class, our framework applies to PUMs and therefore to a broader class of models. Finally, they do not discuss either the problem of welfare maximization and treatment allocation, nor the cost-benefit analysis when agents have subjective participation costs.

Finally, our paper relates to the literature on distributional treatment effects. Qi_et_al_2023 and fan2025policylearningalphaexpectedwelfare also employ superquantiles for treatment allocation, but their analysis differs from ours in several ways. They study a binary-treatment environment without unobserved heterogeneity, so MTEs play no role - they restrict attention to linear regression models - and they do not address the problem of bounding the distribution of individual welfare changes or learning about welfare in settings such as compensated variation, welfare maximization with self-selection, or generalized Roy models with subjective participation costs.

The closest paper to ours is Kallus2023. Like us, he uses superquantiles to recover features of the individual welfare distribution from conditional average treatment effects. Nonetheless, the approaches diverge in essential respects. We work with the broader class of PUMs, covering a wide range of demand and behavioral models. We show that superquantiles can recover distributional welfare information even when the analyst observes only aggregate market data—an issue not treated in Kallus2023. We also provide bounds for welfare maximization and treatment choice under self-selection, and we develop a distributional analysis for generalized Roy models with subjective treatment costs, including the distribution of welfare among actual participants. These settings lie outside the scope of Kallus2023. Thus, while related, the two papers address distinct questions and deliver different types of distributional welfare insights.

Model and Distributional Framework

This section serves two purposes. First, we present the perturbed utility model (PUM), following Allen_Rehbeck and McFadden_Fosgerau_2012. The PUM is a broad class of latent utility models with additively separable unobserved heterogeneity, encompassing widely used formulations such as additive random utility models (ARUMs), bundling models, and matching models. Second, we develop a distributional framework for analyzing welfare in PUMs, extending the analysis beyond average effects. To do so, we briefly review standard quantiles and introduce the less familiar but central concept of superquantiles, which will form the basis of our distributional welfare results.

The PUM approach

Consider a decision-maker (DM) making a utility-maximizing choice among an unordered set of alternatives $j\in\mathcal{J}=\left\{1,\ldots, J\right\}$. The DM's optimal choice vector $Y$ satisfies:

equation[equation omitted — 114 chars of source]

where $B\subseteq\mathbb{R}^J$ denotes the DM's budget constraint and $X_j=(X_{j,1},\dots,X_{j,d_j})'$ denotes the observed characteristics of good $j$, which has $d_j$ observed characteristics. These good-specific regressors are collected into $X=(X_1',\dots,X_J')'$, and $X\subset \mathcal{X}$. We note that there can be common regressors across $X_j$ and $X_k$, $j\neq k$ and $j,k\in\mathcal{J}$. Therefore, $X$ can encode observed characteristics of both the goods and the DM. The vector $\vec{u}=(u_1,\dots,u_J)$ encodes how the desirability of a good varies with its regressors. Moreover, $\mathcal{R}$ is called the disturbance function and is a function of the DM's choice, $y$, and unobserved heterogeneity of the agent, $\varepsilon\in E$. In addition, the term $\mathcal{R}$ represents a regularizer term that smooths out choices.

To better see how PUM serves as an umbrella for a large class of models, consider a straightforward ARUM. In an ARUM environment, a DM must choose one of the choices $j\in\mathcal{J}$, whose utility is given by $$v_j=u_j(X_j)+\varepsilon_j$$ Setting $B=\Delta_J$, where $\Delta_J$ is the $J$-dimensional simplex defined as: $$\Delta_J=\Big\{y\in\mathbb{R}^J|\sum\limits_{j=1}^Jy_j=1,\quad y_j\geqslant0\forall j\in\mathcal{J}\Big\}$$ and disturbance function: $$\mathcal{R}(y,\varepsilon)=\sum\limits_{j=1}^Jy_j\varepsilon_j,$$ it follows that the ARUM framework corresponds to a particular instance of a PUM. Furthermore, in the ARUM case it is well-known that under the assumption that $\varepsilon$ is absolutely continuous, the DM's will choose one of the available options $j\in \mathcal{J}$, which corresponds to choosing one of the vertices of $\Delta_J$.

In general, we will focus on the case where $B=\Delta_J$ as above. However, most of our results apply to the general case of $B$ being a closed convex subset of $\mathbb{R}^J$. Following Allen_Rehbeck, we make use of the following assumption throughout the paper.

assumptionAssume the following: \begin{enumerate}[label=(\roman*)] • The random variables $Y$, $X$ and $\varepsilon$ satisfy ((ref)). • $B\subseteq\mathbb{R}^J$ is a nonempty, closed set. • $u_j:\mathbb{R}^{d_j}\rightarrow\mathbb{R}$ for $j\in\mathcal{J}$ where $d_j$ denotes the dimension of the space of covariates associated to choice $j\in \mathcal{J}$. • $\mathcal{R}:\mathbb{R}^J\times E\rightarrow \mathbb{R}\cup\{-\infty\}$, where $E$ denotes the support of $\varepsilon$, is an extended real-valued function. \end{enumerate}
comment\begin{definition}[Decomposability] The space $\mathcal{Z}$ is said to be decomposable if, whenever $T$ is a subset of $S$ of finite measure and $x_0: T \rightarrow X$ is a measurable function whose range is bounded, then for every $z \in \mathcal{Z}$ the function $$ z(s)=\left\{\begin{array}{lll} z_0(s) & \text { if } & s \in T, \\ z(s) & \text { if } & s \in S\setminus T, \end{array}\right. $$ also belongs to $\mathcal{Z}$. \end{definition} An important example of a decomposable space is $\mathcal{Z}=\mathcal{L}_X^p(S, \Sigma, \sigma)$, the space of all measurable functions $z: S \rightarrow X$. This measure satisfies the decomposability condition iff $$ \int_S\|z(s)\|^p \sigma(d s)<+\infty. $$ In the present context, the notion of decomposability allows one to provide a simpler proof that PUM aggregates.

Throughout the paper, we also make use of the following notation: $$U(y;x,\varepsilon)\triangleq\sum_{j=1}^Jy_ju_j(x_j)+\mathcal{R}(y,\varepsilon)$$ for all $y\in B$ and $X=x$. When $X$ is treated as a random variable, we use $U(y;X,\varepsilon)$.

Allen_Rehbeck show that the PUM aggregates. Based on their aggregation result, we can define the conditional average welfare $W(x)$ as: $$W(x)\triangleq \mathbb{E}\left[\max_{y\in B}U(y;x,\varepsilon)\mid X=x\right].$$

The previous expression can be interpreted as a generalization of the social surplus function introduced by mcf1 in the context of ARUMs. More importantly for our purposes, Allen_Rehbeck show that welfare changes, measured by $W(x^1)-W(x^0)$, are nonparametrically identified using only aggregate market data. We exploit this result as a key building block of our analysis.

Superquantiles

Instead of focusing on average welfare (or average welfare changes) across the entire distribution of $\varepsilon$, our focus will be on average welfare among a specific subset of the population, namely the $(100 \times \beta) \%$-worst affected. To formalize this notion, we use the concept of quantiles and the lesser-known concept of superquantiles introduced in Rockafellar2000OptimizationOC, ROCKAFELLAR20021443, and Rockafellar_Royset.

Let $Z$ be a random variable with a finite mean. The $\beta$-quantile corresponds to $F_Z^{-1}(\beta)=\inf \left\{\lambda: F_Z(\lambda) \geq \beta\right\},$ where $F_Z(z)=\mathbb{P}(Z \leq z)$. The $\beta$-superquantile Rockafellar2000OptimizationOC corresponds to the average among the $(100 \times \beta) \%$-lowest outcomes, which formally is defined as:

equation[equation omitted — 171 chars of source]

where $[t]_-\triangleq\min\{t,0\}$. The supremum is attained when $\lambda$ equals to the $\beta-$quantile, which corresponds to $F_Z^{-1}(\beta)=\inf \left\{\lambda: F_Z(\lambda) \geq \beta\right\}, \quad$ where $F_Z(z)=\mathbb{P}(Z \leq z)$.

Provided $F_Z\left(F_Z^{-1}(\beta)\right)=\beta$, i.e. assuming that $Z$ is continuous, then $\mathbb{S}_\beta(Z)=\mathbb{E}\left[Z \mid Z \leq F_Z^{-1}(\beta)\right]$. In the general case, where $Z$ may be discontinuous, we have the general inclusion:

equation[equation omitted — 181 chars of source]

Furthermore, $\mathbb{S}_\beta(Z)$ is continuous in $\beta$, concave, translation invariant, homogeneous, and monotone (Shapiro_et_al_2013). These properties make $\mathbb{S}_\beta(Z)$ the correct way to generalize the “average of the $(100 \times \beta) \%$-lowest values” when ambiguity occurs due to discontinuities. More importantly, as we shall see, the notion of superquantiles allows us to account for the distributional heterogeneity in welfare and in welfare changes.

figure[figure omitted — 732 chars of source]

To garner some intuition, Figure (ref) illustrates the superquantile of a normally distributed random variable, $Z$. Note that $f_Z(\cdot)$ denotes the probability density function of $Z$ while $F_Z(\ cdot) $ denotes the cumulative density function of $Z$. For a given quantile $\beta$, the lower superquantile $\mathbb{S}_\beta(Z)$ would (roughly) be the blue shaded area, normalized by multiplying by ${1\over\beta}.$ While this section has focused on lower superquantiles due to their relevance to the rest of the paper, it is straightforward to adapt (ref) to capture the right tail of the distribution associated with $Z$. The adaptation comes from the fact that the right superquantile can be defined as $\bar{\mathbb{S}}_{\alpha}(Z)=-\mathbb{S}_{1-\beta}(-Z),$ where $\alpha=1-\beta$. In particular, $\bar{\mathbb{S}}_{\alpha}(Z)$ can be written as: $$\overline{\mathbb{S}}_\alpha(Z)=\min_{\lambda\in \mathbb{R}}\left \{\lambda+{1\over 1-\alpha}\mathbb{E}(Z-\lambda)_+\right\}$$ where $(t)_{+}=\max\{t,0\}$.

commentUsing this notation, we can then characterize the probability that an agent with characteristics $X=x$ (and therefore deterministic utility $\vec{u}(x)$) will achieve indirect utility less than or equal to $z\in\mathbb{R}$. Formally, this characterization can be written as \begin{equation} \Psi(x,z)\triangleq \mathbb{P}(\{\varepsilon\in \mathbb{R}^{J}: \max_{y\in B}U(y, x,\varepsilon)\leq z\}). \end{equation} \textcolor{blue}{EM: I need to unify the notation.} The above characterization allows us to define a quantile and a superquantile. \begin{definition} The $\alpha$-quantile of the random variable $Z(x,\varepsilon)$ is given by: \begin{equation} q_\alpha(x)\triangleq \min \{z\in \mathbb{R} \mid \Psi(x,z) \geqslant \alpha\}, \alpha \in[0,1]. \end{equation} EM: these are not the proper definitions. Need to correct them. \end{definition} \begin{definition} For $\alpha\in [0,1)$ the right (upper tail) $\alpha$-superquantile is given by \begin{equation} \bar{\mathbb{S}}_{\alpha}(x)\triangleq\inf_{\lambda}\left\{\lambda+{1\over \beta}\mathbb{E}(Z(x,\varepsilon)-\lambda)_+\right \}. \end{equation} Similarly, the left (lower tail) $\beta$-superquantile is defined by \begin{equation} \mathbb{S}_{\beta}(x)\triangleq \sup_{\lambda}\left\{\lambda+{1\over \beta}\mathbb{E}(Z(x,\varepsilon)-\lambda)_{-}\right \}. \end{equation} where $q_{\alpha}(x)$ is the $\beta$-quantile in ((ref)) and expectations are taken over the marginal distribution of $\varepsilon$. \end{definition}

Using the notion of superquantile, we define the distributional welfare $W_\beta(x)$ as:

equation[equation omitted — 177 chars of source]

The term $W_\beta(x)$ measures the welfare for the $(100\times\beta)\%$ populations worst affected.\footnote{To make this intuition transparent, assume that the distribution of the random variable $M(y;x,\varepsilon)\triangleq \max_{y\in B}U(y;x,\varepsilon)$ is continuous. Then, we can use this assumption to get $W_\beta(x)=\mathbb{E}(M(y;x,\varepsilon)\mid X=x,M(y;x,\varepsilon)\leq F_{M(y;x,\varepsilon)}).$ Accordingly, $W_\beta(x)$ can be expressed as the average welfare for the $(100\times\beta)\%$ at the bottom of the welfare distribution. In the online Appendix, we discuss some basic properties of $W_\beta(x).$}

comment\begin{theorem} Let Assumption (ref) hold. For any given realization of observables $x\in\mathcal{X}$, we get the following: \begin{itemize} • The model aggregates: \begin{equation} \max_{Y\in\mathcal{Y}}\mathbb{S}_\beta\Big(U(Y;x,\varepsilon) \Big) =\mathbb{S}_\beta\Big(\max_{y\in B}U(y;x,\varepsilon)\Big). \end{equation} • For a solution $\bar{Y}$ of the above, it follows that: \begin{equation} \mathbb{S}_\beta(U(\bar{Y};x,\varepsilon))= \sum_{j=1}^J\bar{Y}_{\beta,j}u_j(x_j)+\bar{D}_\beta(y) \end{equation} where $$\bar{Y}_{\beta}\triangleq\mathbb{E}[\bar{Y}(\varepsilon)\mid U(\bar{Y}(\varepsilon);x,\varepsilon)\leqslant q_\beta(U(\bar{Y}(\varepsilon);x,\varepsilon))$$ and $$\bar{D}_\beta(y)\triangleq \sup_{Y\in\mathcal{Y}:\mathbb{E}({Y})=y}\mathbb{E}\left[D(Y(\varepsilon),\varepsilon)\mid U(Y(\varepsilon);x,\varepsilon)\leq q_\beta(U(Y(\varepsilon);x,\varepsilon))\right]$$ • For a given $x\in X$, let $M(x,\varepsilon)\triangleq \max_{y\in B}U(y,x,\varepsilon)$. Then, it follows that: \begin{equation} \mathbb{S}_\beta(M_x(\varepsilon ))=\mathbb{E}\left[M(x,\varepsilon)\mid M(x,\varepsilon)\leq q_\beta(M(x,\varepsilon))\right] \end{equation} where both the expectations and the quantile are being taken with respect to $\varepsilon$. \end{itemize} \end{theorem} {\em Proof.} Part (i) above follows from Theorem 4 in ruszczynski2006optimization. Part (ii) follows from the same cyclical proof used to prove Theorem 1 in Allen_Rehbeck ({\color{red} CL: I will add the actual proof here, pg. 26 of Allen Rehbeck}). \ $\square$ Part (i) of the theorem provides a distributional foundation to the PUM class. In particular, part (i) aggregates not only at the average level but also at the superquantile level. From an economic standpoint, this result extends the aggregation result beyond the average case and enables one to carry out distributional welfare analysis. Furthermore, part (ii) complements this finding by showing that a distributional representative agent is admitted. In subsequent sections we exploit these key features of Theorem (ref). From a mathematical standpoint, the key condition in proving Theorem (ref) is Assumption (ref)(v), that the space of random variables $\mathcal{Y}$ is decomposable. In particular, the equivalence in expression ((ref)) relies on this assumption. {\color{red} CL: I changed notation and removed $W_\beta(x)$. I did that for clarity, since we kept changing variable names and defining new terms. But if you want, I can change this back just let me know!} The following corollary formalizes the fact that Proposition (ref) is a direct corollary of the superquantile aggregation result. \textcolor{blue}{We now establish our aggregation theorem. We focus in the case of indirect utility for a representative agent.} \begin{theorem}Let Assumption (ref) hold. In addition, assume that the maximum of $U(y;x,\varepsilon)$ is attainable over $y\in B$ for all $\varepsilon\in E$. Then, for all $x\in \mathcal{X}$ and $\beta\in(0,1]$: \begin{equation} \max_{Y\in\mathcal{Y}}\mathbb{S}_\beta\Big(U(Y(\varepsilon);x,\varepsilon) \Big) =\mathbb{S}_\beta\Big(\max_{y\in B}U(y;x,\varepsilon)\Big). \end{equation} In addition suppose that the set $\arg\max_{Y\in\mathcal{Y}}\mathbb{S}_\beta(U(Y(\varepsilon);x,\varepsilon))$ is a singleton. Then the following holds for an optimal solution $\bar{Y}_x$: \begin{equation} \bar{Y}_x\in \arg\max_{Y\in\mathcal{Y}} \mathbb{S}_\beta(U(Y(\varepsilon);x,\varepsilon))\Leftrightarrow \bar{Y}_x(\varepsilon)\in\arg\max_{y\in B}U(y;x,\varepsilon) \end{equation} \end{theorem} {\em Proof.} All proofs are collected in Appendix (ref).\\ \ $\square$ Several points are worth emphasizing. Firstly, Theorem (ref) establishes that the PUM class aggregates at the level of the superquantile. From an economic perspective, this result implies that for any given $\beta \in (0,1]$, we can associate a representative agent to this segment of the distribution. This provides a tractable framework to account for welfare heterogeneity stemming from the unobservable $\varepsilon$. Secondly, Theorem (ref) connects to classical aggregation results. For instance, in the context of Gorman polar forms of indirect utility functions, aggregation typically proceeds by differentiating, applying Roy’s identity, and then taking expectations. In contrast, our result does not require differentiability or the existence of unique maximizers. More importantly, the aggregation here is distributional: it is achieved through superquantiles rather than expectations. When specialized to the ARUM class, Theorem (ref) non-trivially generalizes the aggregation result of mcf1, thereby enabling distributional welfare analysis even in the traditional ARUM framework. Third, from a technical standpoint, our aggregation result draws on the concept of interchangeability in stochastic optimization (see Shapiro_et_al_2013), which plays a central role in the proof. Finally, we point out that for $\beta=1$, Theorem (ref) generalizes Allen_Rehbeck. \begin{theorem} Let Assumption (ref) hold and let $x\in \supp(X)$. Suppose $Y$ is ($x,\varepsilon)$- measurable and both $\max\{D(Y(x,\varepsilon),\varepsilon)\}$ and $\max\{Y(x,\varepsilon)\}$ are finite, where the maximum is being taken with respect to $\varepsilon.$ Define the set $\hat{\varepsilon}_\alpha(\vec{u}(x))=\{\varepsilon|Z(\vec{u}(x),\varepsilon)\geqslant q_\alpha(\vec{u}(x))\}$. Then, for all $\alpha\in[0,1]$ \begin{enumerate}[label=(\roman*)] • It follows that $$\mathbb{E}\Big[Y(x,\varepsilon)|\varepsilon\in\hat{\varepsilon}_\alpha(\vec{u}(x))\Big]\in\arg\max_{y\in\bar{B}}\sum\limits_{j=1}^Jy_ju_j(X_j)+\tilde{D}(y)$$ where $\bar{B}$ is the convex hull of $B$ and $$\tilde{D}(y)=\sup_{\tilde{Y}\in\mathcal{Y}:\mathbb{E}\Big[\tilde{Y}(\varepsilon)|\varepsilon\in\hat{\varepsilon}_\alpha(\vec{u}(x))\Big]=y}\mathbb{E}\Big[D(\tilde{Y}(\varepsilon),\varepsilon)|\varepsilon\in\hat{\varepsilon}_\alpha(\vec{u}(x))\Big]$$ where $\mathcal{Y}$ is the set of $\varepsilon$-measurable functions that map to $B$. • Define the indirect utility function of the representative agent problem as \begin{align*} V_\alpha(\vec{v})=\max_{y\in\bar{B}}\sum\limits_{j=1}^Jy_jv_j+\tilde{D}(y) \end{align*} It then follows that \begin{align*} V_\alpha(\vec{u}(x))=\mathbb{E}\Big[\max_{y\in B}\sum\limits_{j=1}^Jy_ju_j(X_j)+D(y,\varepsilon)|\varepsilon\in\hat{\varepsilon}(\vec{u}(x))\Big]=\bar{\bar{\mathbb{S}}}_\alpha(\vec{u}(x)) \end{align*} \end{enumerate} \end{theorem} \subsection{Superquantiles and choice behavior} We now use the concept of superquantiles to provide a distributional version of the celebrated WDZ theorem presented in expression ((ref)). The following proposition presents this result. \begin{proposition} Let Assumption (ref) hold. In addition suppose that the set $\arg\max_{Y\in\mathcal{Y}}\mathbb{S}_\beta(U(Y;x,\varepsilon))$ is a singleton, and denote the solution as $\bar{Y}_x$. If the random variable $M(x,\varepsilon)\triangleq \max_{y\in B}U(y;x,\varepsilon)$ is continuous (in $\varepsilon$) with distribution $F_M$, then the following holds for all $\beta\in(0,1]$: \begin{equation} \mathbb{E}[\bar{Y}_x(\varepsilon)\mid X=x,M(x,\varepsilon)\leq F_M^{-1}(\beta)]=\nabla W_{\beta}(\vec{u}(x)) \end{equation} \end{proposition} The previous result characterizes conditional choice probabilities in terms of the covariates vector, $x$, and the quantile $F^{-1}_M(\beta)$. In particular, characterization (ref) accounts for heterogeneous choice behavior. This is a fundamental departure from the traditional WDZ which only characterizes average choice behavior. Moreover, as with the aggregation result discussed previously, the traditional WDZ theorem for PUMs is recovered from the above when $\beta=1$.

Bounds

In this section, we show how superquantiles allow us to move beyond purely average-based analysis. Our objective is to learn about the distribution of individual welfare changes without imposing strong structural assumptions or requiring detailed micro-level data. We focus on the empirically familiar setting in which only aggregate market data is available and demonstrate that the same information used to identify average welfare changes can, through superquantiles, yield informative insights into the distributional consequences of policy interventions.

Bounding individual welfare changes

We begin by defining the individual welfare change as the random variable

equation[equation omitted — 143 chars of source]

where $X \in \mathcal{X}$ denotes observable characteristics, and $t = (t_0, t_1)$ represents the values of a predetermined variable before and after an intervention. For example, $t_0$ and $t_1$ may correspond to prices observed before and after a policy intervention, such as the introduction of an ad valorem tax. Similarly, $t$ encapsulates information about whether an individual has been assigned to a social program or not, in which case $t_1=1$ and $t_0=0$, respectively.

It is worth noting that for a type $(X,\varepsilon)$, the realization of the random variable $\tau(X,t,\varepsilon)$ is unobservable to the analyst. More importantly, in many relevant economic settings, the distribution of $\tau(X,t,\varepsilon)$ is unknown, posing two key identification challenges when analyzing the distribution of individual welfare. The first arises in the typical situation in which the analyst observes only aggregate market data. Because individual decision-making is not available, the conventional approach in such settings is to work with the conditional average welfare change, defined as

eqnarray[eqnarray omitted — 312 chars of source]

where the expectations are taken with respect to $\varepsilon$.\footnote{We note that the second equality uses the assumption that $F_{\varepsilon\mid X,t_1}=F_{\varepsilon\mid X,t_0}=F_{\varepsilon\mid X}$ for all $X\in \mathcal{X}$.}

As shown in Allen_Rehbeck, $\tau(X,t)$ is non-parametrically identified for general distributions of $\varepsilon$ and flexible specifications of the deterministic utility component $u(X)$. However, such identification results are typically uninformative about the distributional implications of policy interventions.\footnote{See bhattacharya2024nonparametric for an excellent survey of approaches to achieve identification of welfare changes in demand models with aggregate and micro-data.}

The second issue arises even when the analyst has access to rich micro-level data. Even though the analyst may be able to observe individuals' decisions, identifying $\tau(X,t,\varepsilon)$ remains impossible due to the evaluation problem. To illustrate, suppose $t_1$ indicates that a worker participates in a training program aimed at increasing productivity, while $t_0$ represents the status quo of non-participation. In this environment, each worker is in only one of these states at a time, but never both. The treatment effect literature has traditionally addressed this challenge by also focusing on average treatment effects, which again does not address the distributional implications of interventions (HECKMAN2007_Handbook_partI).

Therefore, while the direct identification of $\tau(X,t,\varepsilon)$ remains infeasible in the two scenarios mentioned above, our goal is to gain some insight into the distribution of this random variable using the information contained in the conditional average welfare changes, $\tau(X,t)$. For this purpose, we define the $\beta$-superquantile associated with $\tau(X,t,\varepsilon)$ as follows:

equation[equation omitted — 186 chars of source]

Expression ((ref)) is concerned with the left tail of $\tau(X,t,\varepsilon)$, which depends on the joint distribution of $X$ and $\varepsilon$. In words, $\mathbb{S}_\beta(\tau(X,t,\varepsilon))$ quantifies the average change in utility among the bottom ($100\times\beta)\%$ of the population affected by an intervention that modifies $t_0$ to $t_1$. From an economic standpoint, expression ((ref)) can be interpreted as a welfare measure of how the bottom ($100\times\beta)\%$ of all individuals fare from this intervention.

While $\mathbb{S}_{\beta}(\tau(X,t,\varepsilon))$ is economically meaningful, fully identifying it first requires identifying the full distribution of $\tau(X,t,\varepsilon)$, which remains infeasible for aforementioned reasons. In fact, without parametric assumptions, the distribution of $\tau(X,t,\varepsilon)$, and in turn $\mathbb{S}_{\beta}(\tau(X,t,\varepsilon))$, is identifiable only with detailed, individual-level data in the absence of the missing outcome problem.\footnote{It is worth noting that the superquantile functional $\mathbb{S}_\beta(\tau(X,t,\varepsilon))$ is not additive. As a result $\mathbb{S}_\beta(\tau(X,t,\varepsilon))\neq \mathbb{S}_\beta\Big(\max_{y\in B}U(y;X,t_1,\varepsilon)\Big)-\mathbb{S}_\beta\Big(\max_{y\in B}U(y;X,t_0,\varepsilon)\Big)$. Consequently, differences in superquantiles of indirect utility cannot be used to infer the distribution of individual welfare changes.} However, the following theorem provides an upper bound for Eq. ((ref)) that does not rely on parametric assumptions nor detailed individual-level data.

theoremLet Assumption (ref) hold. Then: \begin{equation} \mathbb{S}_\beta(\tau(X,t,\varepsilon)) \leqslant \mathbb{S}_\beta(\tau(X,t)). \end{equation}

{\em Proof.} All proofs are collected in the Appendix (ref).

Theorem (ref) establishes that the superquantile of individual utility changes can be bounded above by the superquantile of conditional average welfare changes. The significance of bound ((ref)) lies in its empirical feasibility. To reiterate, the left-hand side cannot be estimated without detailed individual-level data and/or strong parametric assumptions. In contrast, the right-hand side can be identified non-parametrically from aggregate data alone. For example, identification can be achieved within the PUM framework using the results of Allen_Rehbeck, or under the class of RUMs following the nonparametric approach of Berry_Haile_ECMA2014. Similarly, bound (ref) is also relevant where the analyst faces the problem of missing outcomes, as discussed above, which persists even in the case of rich micro-level data.\footnote{In sections (ref) and (ref) we discuss this problem in further detail.}

From an economic perspective, Theorem (ref) is valuable because it offers a framework for evaluating the distributional welfare effects of policy interventions. To illustrate, define the total average welfare change as $\tau(t) = \mathbb{E}_{F_{X}}(\tau(X,t))$. Suppose a policy yields $\tau(t) > 0$, but $\mathbb{S}_\beta(\tau(X,t)) < 0$ for a substantial fraction of the population (e.g., $\beta = 0.2$). Comparing these two quantities provides a first-order assessment of whether a policy generates overall welfare gains while imposing losses on specific subgroups. This information might indicate that the policy change yields some large “winners" from the policy, masking harm that other subpopulations may be experiencing. Furthermore, $\mathbb{S}_\beta(\tau(X,t))$ enables us to identify which subpopulations are adversely affected. Assuming continuity, $\mathbb{S}_\beta(\tau(X,t))$ corresponds to the average welfare change among agents with $\tau(X,t) \leq F^{-1}_{\tau(X,t)}(\beta)$. Since $\tau(X,t)$ is identifiable, this approach allows us to identify which observable subgroups experience welfare losses under the policy. It also provides insights into how to improve a policy and minimize these harms. In summary, the key fact behind the bound

Theorem (ref) can be expressed in an equivalent form in terms of the average of quantiles. The following result formalizes this observation.

corollaryLet Assumption (ref) hold. In addition assume that $\tau(X,t,\varepsilon)$ and $\tau(X,t)$ have continuous distributions. Then the bound ((ref)) can be equivalently written as: \begin{equation} \int_0^\beta F^{-1}_{\tau(X,t,\varepsilon)}(\theta)d\theta\leq \int_0^\beta F^{-1}_{\tau(X,t)}(\theta)d\theta \end{equation}

It is worth remarking that $F^{-1}_{\tau(X,t,\varepsilon)}(\theta)$ corresponds to the $\theta$-quantile of the individual utility changes. In general, as we said earlier, this quantity is not identified. On the other hand, $F^{-1}_{\tau(X,t)}(\theta)$ gives us the $\theta$-quantile of the conditional average welfare changes, as measured by $\tau(X,t)$.

There is one final point worth mentioning regarding Theorem (ref). Notice, we did not exploit any structure of the PUM framework in the derivation of this bound. Therefore, the bound applies beyond the additive structure of the PUM class and is a general bound. However, the key requirement for this bound to be useful is the nonparametric identifiability of $\tau(X,t)$. The main advantage of using the PUM class is that it allows the identification of $\tau(X,t)$ with only aggregate market data. Therefore, while the PUM is not the only class of models for which Theorem (ref) applies, it is the most general class that we know of for which the bound can be non-parametrically estimated using only aggregate data. With rich micro-level data, nonparametric identification of $\tau(X,t)$ is again possible if utilities are additive (see, e.g., Bhattacharya_2015,bhattacharya2024nonparametric) or under certain conditions if utilities are non-additive (see e.g, imbens2009identification).

We now discuss how to lower bound $\mathbb{S}_\beta(\tau(X,t,\varepsilon))$. The following theorem provides a simple method to do so, assuming an additional condition.

theoremLet Assumption (ref) hold. In addition, suppose that $\tau(X,t)-\tau(X,t,\varepsilon)\leq \gamma$ for $X\in\mathcal{X}$ and $\gamma\in \mathbb{R}$. Then: \begin{equation} \mathbb{S}_\beta(\tau(X,t,\varepsilon))\geq \mathbb{S}_\beta(\tau(X,t))-\gamma. \end{equation}

The previous result shows that when the difference between $\tau(X,t,\varepsilon)$ and $\tau(X,t)$ is uniformly bounded by a constant $\gamma$, then $\mathbb{S}_\beta(\tau(X,t,\varepsilon)$ is lower bounded by $\mathbb{S}_\beta(\tau(X,t))-\gamma$, a quantity that is identified as long as $\gamma$ is known to the analyst (Manski1990).

It is worth emphasizing that Theorems (ref) and (ref) enable the analyst to learn about the distribution of individual welfare changes, even when only aggregate market data or average welfare measures are available. Theorem 1 does so without any knowledge of the distribution of unobservables nor any parametric assumptions; Theorem 2 requires an assumption on the range of treatment effects possible, conditional on observables. To fix ideas, consider the case in which the PUM reduces to a standard ARUM. In this setting, bhattacharya2024nonparametric argues that average welfare and observed choice probabilities are uninformative about the welfare change distribution. Our results overturn this conclusion: average welfare changes do contain information about the underlying individual welfare distribution. The source of this difference is methodological. We extract information from average welfare changes themselves, whereas bhattacharya2024nonparametric relies solely on choice probabilities, thereby discarding welfare-relevant variation.

comment\textcolor{blue}{Not sure if we need this litle paragraph and result} We close this section by noticing that combining Theorems (ref) and (ref) we get the following straightforward corollary. \begin{corollary} Let Assumption (ref) hold. In addition, suppose that $|\tau(X,t)-\tau(X,t,\varepsilon)|\leq \gamma$ for all $X\in\mathcal{X}$ and $\gamma\in \mathbb{R}$. Then: \begin{equation} \mathbb{S}_\beta(\tau(X,t,\varepsilon))\in[\mathbb{S}_\beta(\tau(X,t))- \gamma,\mathbb{S}_\beta(\tau(X,t))] \end{equation} \end{corollary}

Computing $\mathbb{S}_\beta(\tau(X,t))$

We conclude this section by discussing a simple approach to compute $\mathbb{S}_\beta(\tau(X,t))$. Suppose the analyst has access to a collection of K estimates $\hat{\tau}(X_1), \ldots, \hat{\tau}(X_K)$, where, for ease of exposition, we omit the dependence on $t$. Using these estimates, a direct approach to computing $\mathbb{S}_\beta(\tau(X,t))$ is to employ the variational representation (ref) together with a plug-in method, yielding the following program:

equation[equation omitted — 212 chars of source]

Problem (ref) is concave but, unfortunately, non-smooth in $\lambda$. However, we can avoid the analytical complications introduced by non-smoothness by observing that program (ref) can be equivalently expressed as a linear programming problem:

eqnarray[eqnarray omitted — 256 chars of source]

Program (ref) offers a straightforward way to compute $\hat{\lambda}$ and $\hat{\mathbb{S}}_\beta(\hat{\tau}(X,t))$. The problem involves a linear objective function with only $2K$ linear constraints, making it computationally simple and easy to implement with standard optimization software. The value of $\hat{\lambda}$ can be interpreted as an estimator of the $\beta$-quantile of the random variable $\tau(X,t)$, while the value of the objective function at the optimum, $V(\hat{\lambda})$, provides an estimate of $\mathbb{S}_\beta(\tau(X,t))$.

Because this approach is based on a set of pre-estimated values $\hat{\tau}(X_1), \ldots, \hat{\tau}(X_K)$, it effectively turns the estimation of $\mathbb{S}_\beta(\tau(X,t))$ into a linear program with estimated inputs. This feature introduces sampling variability that must be accounted for in inference. To handle this, one can apply the inferential methods developed in Shum_2022.

Applications

While the results in the preceding sections are general, this section illustrates how they apply to standard economic settings. We consider three applications. First, we examine the distributional properties of CV and discuss how information in conditional-average CV helps infer the distributional impacts of exogenous price changes. Second, we apply our framework to the analysis of distributional welfare in treatment allocation problems, demonstrating how our results account for unobserved heterogeneity and self-selection. Finally, we analyze treatment effects in a setting with endogenous participation and subjective treatment participation costs. In this context, we focus on the benefits, costs, and welfare implications of social programs.

Distributional CV analysis

We begin by discussing how to specialize the results from the previous section to the problem of compensated variation ($\operatorname{CV}$). In the ARUM literature, the problem of $\operatorname{CV}$ is well studied. Most studies of CV have two particular features. First, they focus on the average case. Second, they focus on the traditional ARUM. We aim to relax these two assumptions while fixing the constant marginal utility of income assumption.

We consider a setting with $J \geq 2$ goods. Let $X_j$ denote the observed non-price characteristics of good $j$, and let $\mathcal{X}_j$ represent the space of covariates associated with that good. The vector of covariates contains the qualities and prices of the different goods. The overall space of non-price covariates is $\mathcal{X} \triangleq \prod_{j=1}^J \mathcal{X}_j$. Let $p = (p_1, \ldots, p_J)\in \mathbb{R}_{+}^J$ denote the price vector, where $p_j$ is the price of good $j$ with $j = 1, \ldots, J$. For a given pair $(X,p)\in \mathcal{X}\times\mathbb{R}_{+}^J$, we assume that observable utility corresponds to the term $h_j(X_j)+\gamma(I-p_j)$ where $I>0$ is consumer's income and $\gamma$ denotes the marginal utility of income.

Our goal is to understand the distributional implications of an intervention that changes the prices of goods. In particular, let $p_0 = (p_{10}, \ldots, p_{J0})$ and $p_1 = (p_{11}, \ldots, p_{J1})$ denote the vectors of prices before and after the intervention, respectively. Note that in terms of our original notation $p_1=t_1$ and $p_0=t_0$ respectively. Accordingly, we use the notation $t=(p_0,p_1).$

Throughout the analysis, we assume that the non-price characteristics $X$ remain fixed and do not respond to the intervention.

Accordingly, at the individual level, we have the following condition that monetarily quantifies the impact of changing prices:

eqnarray[eqnarray omitted — 239 chars of source]

where $\operatorname{CV}$ is the CV necessary to keep the consumer with characteristics $X$ and unobserved tastes $\varepsilon$ with the same indirect utility level. Exploiting the quasi-linearity of the utility function, we obtain a simple analytical expression for the individual $\operatorname{CV}$, which we denote as $\operatorname{CV}(X,t, \varepsilon)$ to emphasize the dependence on $X$, $t$, and $\varepsilon$:

equation[equation omitted — 168 chars of source]

where $U(y;X,p_l,\varepsilon)\triangleq \sum_{j=1}^Jy_ju_j(X_j,p_{jl})+D(y,\varepsilon)$ and $u_j(X_j,p_{jl})=h_j(X_j)-\gamma p_{jl}$ for $l=0,1.$\\

Note that due to the quasilinear structure, income effects do not impact $\operatorname{CV}(X,t,\varepsilon)$. However, $\operatorname{CV}(X,t,\varepsilon)$ can capture rich patterns of complementarity and substitutability across goods, as well as general additive forms of unobserved heterogeneity. Furthermore, $\operatorname{CV}(X,t,\varepsilon)$ is not restricted to the case of discrete choice problems. In fact, when $\mathcal{R}(y,\varepsilon)=\sum_{j=1}^Jy_j\varepsilon_j$ and $B=\Delta_J$, expression ((ref)) yields the $\operatorname{CV}$ in ARUMs. In other words, (ref) applies when vector $y$ refers to continuous quantities, without imposing a particular structure on $B$, being the discrete choice model a particular case.

Taking conditional expectation with respect to $\varepsilon$ in ((ref)) we find:

eqnarray[eqnarray omitted — 171 chars of source]

The average $\operatorname{CV}(X)$ is commonly used when the analyst has access only to aggregate market data (see, for example, Berry_Haile_ECMA2014). However, this measure is uninformative about the potential distributional consequences of price changes. The following result shows how Theorems (ref) and (ref) in Section (ref) can be applied in the $\operatorname{CV}$ context to learn about the distributional consequences of price changes.

commentFrom Equation (ref) it follows that $CV(x,\varepsilon)$ is $(x,\varepsilon)$-measurable. Accordingly, we can use this structure to analyze the conditional (on observables) mean $CV$ which we define as $CV(x)\triangleq \mathbb{E}(CV(x,\varepsilon)\mid x)$. To see this, lets expand the definition of welfare to be a function of prices in addition to $x$. In particular, we define $W(\vec{u}(x,p^l))=\mathbb{E}[\max_{y\in \Delta_J}U(y;x,p^l,\varepsilon)\mid x]$ for $l=0,1.$ Then, the conditional average $CV$ corresponds to: \begin{equation} CV(x)={1\over\gamma}[W(\vec{u}(x,p^1))-W(\vec{u}(x,p^0))]. \end{equation} It is worth mentioning $CV(x)$ is the traditional formula commonly used in the ARUM framework. However, there is a significant difference given by the fact that expression (ref) is determined by the PUM class, allowing for complex patterns of both complementarity and substitutability. The ARUM rules out the former of these features. Thus, despite its mathematically similar expression, ((ref)) is a more general measure of CV. A second important implication is the fact that we can generalize expression (ref) by using Theorem (ref). Concretely, by combining (ref) with Theorem (ref) we are able to extend the traditional $CV$ average analysis. To do so, we define $CV_\beta(x)$ to be the conditional average $CV$ associated with the quantile $\beta\in (0,1]$. Formally, we have: \begin{equation} CV_\beta(x)={1\over \gamma}[W_\beta(\vec{u}(x,p^1))-W_\beta(\vec{u}(x,p^0))]. \end{equation} It is easy to see that (ref) is a simple extension of the average measure (ref). In fact, for $\beta=1$, we get that ((ref)) and ((ref)) coincide. Before moving forward, it is important to clarify a key subtlety in the above expression and its interpretation. Put simply, $CV_\beta(x)$ is the transfer necessary to equate utilities of a specified (by $\beta)$ segment of those facing $p^0$ to a specified segment of those facing $p^1$. If $CV_{\beta}(x)>0$, then those in the lowest $(100\times\beta)\%$ of utilities among those facing $p^0$ are worse off than those in the lowest $(100\times\beta)\%$ of utilities among those facing $p^1$; the opposite is true if $CV_{\beta}(x)<0$. This does not necessarily inform us about the distribution of CV across individuals, which will be specified and dealt with in the next subsection. Moreover, if we were to assume that $p^0$ and $p^1$ differ only in the $j$-th component, we have the following expression \begin{equation} CV_\beta(x)={1\over \gamma}\int_{p_j^0}^{p_j^1}{\partial W_\beta (\vec{u}(x,p))\over \partial u_j}dp_j \end{equation} where by Proposition (ref) we know that ${\partial W_\beta(\vec{u}(x,p))\over \partial u_j}$ corresponds to the conditional choice probability associated with the $\beta$-superquantile. Thus, expression (ref) tells us that $CV_\beta(x) $ represents the area below the demand curve associated with the $\beta$-superquantile. In other words, our framework can accommodate welfare heterogeneity generalizing the classic analysis of Small_Rosen_1981, mcf1, and Hanneman_1996. \subsection{Individual DCV and aggregate data} In the previous section, we discussed how our approach is helpful in extending the $CV$ analysis beyond the traditional average framework. As we show above, $CV_\beta$, as defined in (ref), measures the compensated variation associated with different quantiles of the welfare distribution. This measure allows for comparison between $\beta$-segments of individuals facing different prices. However, this is not necessarily informative about the distribution of individual level $CV$. Our goal in this section is to extend this analysis to the individual level by providing useful and informative bounds on individual $CV$ ($CV(x,\varepsilon)$) using conditional average $CV$ ($CV(x))$ measures.
commentBy using the concept of superquantiles allows us to quantify—and potentially mitigate—the adverse impact of an intervention on the worst-affected $(100 \times \beta)\%$ of the population, which formally corresponds to: \begin{equation} \mathbb{S}_\beta(CV(X,\varepsilon))={1\over\gamma}\mathbb{S}_\beta\left(\max_{y\in B}U(y;X,p^1,\varepsilon)-\max_{y\in B}U(y;X,p^0,\varepsilon)\right) \end{equation} where $\beta\in(0,1]$ and the superquantile is computed under the joint distribution of $(X,\varepsilon)\in(\mathcal{X}, E$).
commentIt is worth highlighting the difference in information captured by expression ((ref)) compared to what is captured by ((ref)). $CV_\beta(x)$ is a measure of the difference in superquantiles, whereas $\mathbb{S}_\beta(CV(x,\varepsilon)\mid X=x)$ is the superquantile of the individual differences given by $M(y;x,p^1,\varepsilon) - M(y;x,p^0,\varepsilon)$. In words, $CV_\beta(x)$ answers the question “how do the bottom $(100\times\beta)\%$ of those facing $p^0$ compare to the bottom $(100\times\beta)\%$ of those facing $p^1$?". In contrast, $\mathbb{S}_\beta(CV(x,\varepsilon)\mid X=x)$ is directly measuring the $(100\times\beta)\%$ of individual $CV(x,\varepsilon)$. This distinction can be further formalized by noting that, in general, superquantiles are non-additive. Specifically, under the strong assumption of comonotonicity between the random variables $M(y;x,p^1,\varepsilon)$ and $-M(y;x,p^0,\varepsilon)$, it can be shown that (see Dhaene_et_al_Risk_measures_review): $$ \mathbb{S}_\beta(CV(x,\varepsilon)) = \mathbb{S}_\beta(M(y;x,p^1,\varepsilon)) + \mathbb{S}_\beta(-M(y;x,p^0,\varepsilon))\neq CV_\beta(x). $$ The previous expression demonstrates that even under the admittedly strong comonotonicity assumption, $\mathbb{S}_\beta(CV(x,\varepsilon))$ cannot be expressed as the difference ${1\over \gamma}[W_\beta(\vec{u}(x,p^1))-W_\beta(\vec{u}(x,p^0))]$. A similar type of discrepancy has been noted in the econometric literature on quantile treatment effects (see, for instance, Imbens_Woodldridge_2009).
commentAn important observation comes from the fact that expression ((ref)) is fully pinned down by the distribution of $\varepsilon$. In terms of data, this condition requires that either we assume that the distribution of $\varepsilon$ is known or we are able to nonparametrically identify it, as discussed in papers such as matzkin1993nonparametric and chiappori2015nonparametric. However, in many relevant applications, the researcher only has access to aggregate data (e.g. Berry_Haile_ECMA2014) or identification of welfare changes is only possible at the average level (Allen_Rehbeck). In such instances where the distribution of $\varepsilon$ can not be reliably assumed nor identified, a natural question is how much we can learn about the distribution of individual $CV$ as specified in (ref).
propositionLet Assumption (ref) hold. Then the following statements hold: \begin{itemize} • $\operatorname{CV}(X,t,\varepsilon)$ and $\operatorname{CV}(X,t)$ satisfy the following inequality: \begin{equation} \mathbb{S}_\beta(\operatorname{CV}(X,t,\varepsilon))\leqslant \mathbb{S}_\beta(\operatorname{CV}(X,t)). \end{equation} • Suppose that $\operatorname{CV}(X,t)-\operatorname{CV}(X,t,\varepsilon)\leq \mu$ for $\mu\in \mathbb{R}_{++}$. Then it is also true that: \begin{equation} \mathbb{S}_\beta(\operatorname{CV}(X,t,\varepsilon))\geq \mathbb{S}_\beta(\operatorname{CV}(X,t))-\mu. \end{equation} \end{itemize}

Some remarks are in order. First, the left-hand side in (ref) corresponds to the $\beta$-superquantile of individual compensating variation across all realizations of $X$ and $\varepsilon$; the right-hand side is the $\beta$-superquantile of the average conditional compensating variation across all realizations of $X$ only. Therefore, Proposition (ref)$(i)$ informs us how the individual $\operatorname{CV}$ across an entire population can be bounded using the information contained in the $\operatorname{CV}(X,t)$s. In doing so, our analysis is refined to identify harmed (or worst-affected) subpopulations defined by specific covariates $X$, which is relevant for fairness and equity considerations in market interventions. This follows from interpreting $\mathbb{S}_\beta(\operatorname{CV}(X,t))$ as a summary of heterogeneity along realizations of the relevant regressors in $X$. Moreover, identifying the worst-affected subpopulations enables more targeted interventions.

A second important observation is that the bound ((ref)) can be identified using aggregate data. For instance, Berry_Haile_ECMA2014 and Allen_Rehbeck show that welfare averages can be identified in demand models with unobserved heterogeneity. Thus, bound ((ref)) is informative about the superquantile of the distribution of individual $\operatorname{CV}$ under minimal data and distributional assumptions. Proposition (ref)(ii) complements this finding by providing a lower bound ((ref)) to the unobserved $\operatorname{CV}$ distribution, which depends on observables where the parameter $\mu$ can represent an upper bound in the $\operatorname{CV}$ that the social planner can allocate (Manski1990).

At a higher level, Proposition (ref) is informative for learning about the distribution of individual $\operatorname{CV}$. As discussed earlier, our analysis offers a direct counterpoint to bhattacharya2024nonparametric, who argues that average welfare measures and conditional choice probabilities are uninformative about the distribution of individual welfare changes—even under additively separable unobserved heterogeneity, as in ARUMs. Proposition (ref) demonstrates that, within the PUM class, aggregate market data are sufficient to conduct meaningful distributional $\operatorname{CV}$ analysis.

Policy choice and welfare maximization

In this section, we apply our framework to the analysis of welfare in the allocation of social programs when participation decisions are endogenous due to self-selection. Following the approach in Sasaki_Ura_2024, our objective is to demonstrate how the MTE can be used to assess the distributional welfare consequences of self-selection and unobserved heterogeneity across different populations and groups of interest. Below, we present the model as described in Sasaki_Ura_2024.

We consider the following causal model:

eqnarray[eqnarray omitted — 127 chars of source]

where $V$ denotes an observed outcome variable, $D$ denotes an observed binary treatment variable, $Z$ denotes a vector of observed exogenous variables, $V_0$ and $V_1$ denote unobserved potential outcomes under no treatment and under treatment, respectively, and $\tilde{\varepsilon}$ denotes an unobserved factor (unobserved heterogeneity) of the treatment selection. Let $\mathcal{Z}$ denote the set of all observables covariates $Z.$

Equation ((ref)) models the outcome production using the potential-outcome framework, while equation ((ref)) models the treatment selection via a binary RUM (threshold-crossing) model. Note that ((ref)) is a particular instance of the PUM.

The function $\tilde{u}$ in the assignment model ((ref)) is nonparametric and is unknown to the econometrician. The model allows for endogeneity (unobserved confoundedness) in the sense that $(V_0, V_1)$ and $\tilde{\varepsilon}$ may be statistically dependent even when conditioned on $Z$. To achieve identification, we assume the vector $Z$ to contain excluded exogenous variables (i.e., excluded instruments) as well as included exogenous variables. Similar to Sasaki_Ura_2024, we use the following assumption:

assumptionEquations ((ref)) and ((ref)) hold, and the random vector $Z$ can be written as $\left(Z_0^{\prime}, X^{\prime}\right)^{\prime}$, where: \begin{itemize} • $\tilde{\varepsilon}$ and $Z_0$ are independent given $X$; • $\mathbb{E}\left[V_d \mid Z, \tilde{\varepsilon}\right]=\mathbb{E}\left[V_d \mid X, \tilde{\varepsilon}\right]$ and $\mathbb{E}\left[V_d^2\right]<\infty$; • $\tilde{\varepsilon}$ is continuously distributed with a convex support conditional on $X$ \end{itemize}

Part (i) concerns the treatment assignment model ((ref)) solely, and this is the only independence assumption to be imposed on the model, implying that we can allow for an arbitrary statistical dependence between the potential outcomes $(V_0, V_1)$ and $\tilde{\varepsilon}$, even conditional on $Z$. Part (ii) states the exclusion restriction of the random subvector $Z_0$ of $Z$, and bounded second moments of the potential outcomes $\left(V_0, V_1\right)$. Part (iii) rules out point masses and holes in the conditional distribution of $\tilde{\varepsilon}$ given $X$. For ease of exposition, we define $\mathcal{X}$ as the set of all observable covariates $X$.

Following the literature on the marginal treatment effect (MTE), it is standard to apply normalizing transformations ($\varepsilon \equiv F_{\tilde{\varepsilon} \mid X}(\tilde{\varepsilon})$ and $u(Z) \equiv F_{\tilde{\varepsilon} \mid X}(\tilde{u}(Z))$) in the threshold crossing model ((ref)). An important implication of Assumption (ref) is that $D=1\{u(Z)-\varepsilon \geq 0\}$ and $\varepsilon$ is distributed uniformly over $[0,1]$ conditional on $Z$. Accordingly, and without loss of generality, the threshold-crossing treatment selection model ((ref)) can be equivalently expressed as

equation[equation omitted — 146 chars of source]

As a consequence we will use ((ref)) in place of the original model ((ref)).

Treatment allocation and welfare

Now we study the case where the planner chooses a non-randomized policy/rule that maps $Z$ to treatment status. Formally we consider maps $\pi:\mathcal{Z}\mapsto \{0,1\}$;.The set of all Borel measurable functions from $\mathcal{Z}$ to $\{0,1\}$ is denoted by $\Pi$. For $\pi\in \Pi$, $V(\pi(Z))$ denotes the utility that an individual with covariates $Z$ derives from the policy $\pi$. In particular, we can express $V(\pi(Z))$ as

eqnarray[eqnarray omitted — 150 chars of source]

Expression ((ref)) represents the individual outcome of an individual with observables $Z$ when planner chooses $\pi\in \Pi.$ The role of the treatment assignment rule is to assign an individual with covariates $Z$ to the treatment, i.e., $D=1$, whenever $\pi(Z)=1$.

It is well known that the distribution of $V(\pi)$ is unidentified, a direct consequence of the missing-outcome problem. As a result, the policy-learning and welfare-maximization literature has largely focused on average welfare as the criterion for selecting an optimal rule $\pi \in \Pi$. Formally, we define the average welfare function $\mathcal{W}:\Pi\mapsto\mathbb{R}$

eqnarray*[eqnarray* omitted — 120 chars of source]

Let $\mathcal{W}(\pi, z)\triangleq\mathbb{E}\left[\mathcal{W}(\pi)\mid Z=z\right]=\mathbb{E}(V_0\mid Z=z)+\pi(Z)\mathbb{E}(V_1-V_0\mid Z=z)$ denote the conditional welfare associated with $\pi$ for a fixed realization $Z = z$.\footnote{ This definition makes explicit the fact that $\mathcal{W}(\cdot,\cdot)$ depends on the choice of $\pi$ and in the conditional value $z$. However, in our formal statements and proofs we work with the random variable $\mathcal{W}(\pi,Z)\triangleq\mathbb{E}\left[\mathcal{W}(\pi)\mid Z\right]$.}

Finally, we define the MTE as follows: $$\operatorname{MTE}( x, \bar{\varepsilon})=\mathbb{E}\left[V_1-V_0 \mid X=x, \varepsilon=\bar{\varepsilon}\right].$$

Under Assumption (ref), Sasaki_Ura_2024 shows that

equation[equation omitted — 199 chars of source]

Representation ((ref)) characterizes the average welfare associated with policy $\pi$. However, it is silent about the distributional welfare implications of implementing such a rule.

The following result shows that conditional average welfare contains meaningful information about the distributional consequences of policy $\pi$. Specifically, it establishes that the conditional average welfare is informative for bounding the distribution of individual welfare effects and, therefore, for assessing the potential harm faced by different subpopulations under the policy.

theoremLet Assumption (ref) hold. Then \begin{equation} \mathbb{S}_\beta(V(\pi))\leq \mathbb{S}_\beta(\mathcal{W}(\pi,Z))\quadfor $\beta\in(0,1]$, \end{equation} where $\mathcal{W}(\pi,Z)\triangleq\mathbb{E}(V_0|Z)+\pi(Z) \int_0^1 \operatorname{MTE}(X, \theta) d\theta$.

The previous result provides a tractable way to bound the distributional consequences of choosing a policy $\pi$. In particular, Theorem (ref) allows the analyst to characterize the $(100\times \beta)\%$ worst-off populations under policy $\pi$. A key feature of bound ((ref)) is that it highlights how unobserved heterogeneity $\varepsilon$ and the $\operatorname{MTE}$ shape the distributional impact of the chosen policy, even though the full distribution of $V(\pi)$ is not identified.

Another important implication of our framework is that it facilitates comparisons across policies regarding the potential harm that particular groups may experience. To formalize this idea, consider two policies $\pi$ and $\pi’ \in \Pi$. We are interested in the change in individual utility resulting from switching from $\pi$ to $\pi’$, given by

eqnarray*[eqnarray* omitted — 74 chars of source]

Because the distribution of $V(\pi’) - V(\pi)$ is not identified—due to the missing outcome problem—direct learning about its distribution is infeasible. However, the following result shows that the information contained in the MTE, together with the policies $\pi$ and $\pi’$, allows the analyst to bound the distributional consequences of moving from $\pi$ to $\pi’$.

propositionLet Assumption (ref) hold. Then \begin{equation} \mathbb{S}_\beta(V(\pi^\prime)-V(\pi))\leq \mathbb{S}_\beta\left((\pi^\prime(Z)-\pi(Z)) \cdot\int_0^1 \operatorname{MTE}(X, \theta) d\theta)\right)\quadfor $\beta\in(0,1]$. \end{equation}

The previous result provides a simple condition that allows the planner to compare two policies. In particular, Proposition (ref) provides a bound on the distribution of individual welfare changes associated with switching from policy $\pi$ to $\pi’$. These bounds enable the policymaker to assess the distributional consequences of alternative policies, with a particular focus on identifying the groups most affected. From an econometric standpoint, bound ((ref)) can be implemented using the results in byambadalai2022welfaregains.

It is worth noting that bound ((ref)) is closely related to regret-based criteria studied by Manski2004, Kiatagawa_Tetenov_2018, and Athey_Wagner_2021. To see this, let $\pi^\ast$ be the policy that maximizes welfare in the lower tail of the distribution. By applying the bound ((ref)) with $\pi^\prime = \pi^\ast$, Proposition (ref) allows us to bound the welfare loss from choosing an alternative policy $ \pi$ relative to $\pi^\ast$. In this sense, our framework provides distributional analogues of regret measures, informing policymakers about regret in terms of worst-case welfare outcomes for certain subpopulations of interest rather than average regret.

Finally, we note that Theorem (ref) and Proposition (ref) extend straightforwardly to settings with multi-valued treatments.

Treatment cost participation and welfare

In this section, we consider the welfare framework introduced by Eisenhauer_Heckman_Vytlacil_2015, which, in the context of treatment effects, examines the marginal benefits and marginal costs of policies. Their framework extends the modern treatment effect literature by providing a method to identify both the marginal benefits and the marginal costs of policy interventions. In particular, they incorporate agents’ subjective costs associated with participation in social programs, which allows them to analyze the benefits, costs, and surplus (benefits minus costs) of specific social programs in average terms. Our goal is to integrate their results with our framework to explore how policymakers can assess not only the average benefits, costs, and surpluses of different policies, but also their distributional implications.

Eisenhauer_Heckman_Vytlacil_2015's framework

We focus in the traditional binary case. As in Section (ref), we assume that are two potential outcomes $\left(V_0, V_1\right)$ and a choice indicator $D$, with $D=1$ if the agent selects into treatment so that $V_1$ is observed and $D=0$ if the agent does not select into treatment so that $V_0$ is observed. Similar to section (ref), we use the potential outcome equation to denote the value of $V$ as $V=DV_1+(1-D)V_0$. Following Eisenhauer_Heckman_Vytlacil_2015, we assume a separable structure in the outcomes where $\mathbb{E}\left(V_j \mid X\right)=\mu_j(X)$ and

equation[equation omitted — 88 chars of source]

As in previous sections, $X$ is a (random) vector of covariates observed by the analyst, while ($\nu_0, \nu_1$) are unobserved (to the analyst) heterogeneity terms. Combining $V=DV_1+(1-D)V_0$ with ((ref)) we get

equation[equation omitted — 126 chars of source]

Let $B=V_1-V_0$ denote the individual gross benefit of treatment, defined as the causal effect on $V$ of moving an otherwise identical individual from state 0 to state 1. Thus, $B$ measures the ceteris paribus change in the outcome induced by treatment.

Let $C$ denote the agent’s subjective treatment cost, defined as:

equation[equation omitted — 52 chars of source]

where $Z$ represents an observed random vector of cost shifters and $\nu_C$ is a random variable unobserved by the analyst. Under this specification, the conditional expectation in Eq. ((ref)) satisfies $\mathbb{E}[C \mid Z] = \mu_C(Z)$.

Individuals choose to participate in the treatment if the perceived benefit from participation is greater than the subjective cost:

equation[equation omitted — 125 chars of source]

where $\mathcal{W}$ is the individual welfare (surplus), that is, the net benefit, from treatment:

$$

aligned\mathcal{W} & \triangleq\left(V_1-V_0\right)-C \\ & =\mu_W(X, Z)-\varepsilon_\mathcal{W} ,

$$ with $\mu_\mathcal{W}(X, Z)=\left[\mu_1(X)-\mu_0(X)\right]-\mu_C(Z)$ and $\varepsilon_{\mathcal{W}}=\nu_C-\left(\nu_1-\nu_0\right)$.

Our distributional analysis of treatment costs and welfare does not impose functional-form restrictions on $\mu_0$, $\mu_1$, or $\mu_C$, nor does it require parametric assumptions on the distributions of $\nu_0$, $\nu_1$, or $\nu_C$. Instead, our results build on the identification framework of Eisenhauer_Heckman_Vytlacil_2015, which enables the analyst to recover the relevant conditional average objects needed for our bounds.

Similar to section (ref), it is easy to see that the choice model ((ref)) corresponds to a particular instance of the PUM. To formalize this, let $ P(X, Z)\triangleq \operatorname{Pr}(D=1 \mid X$, $Z)$ denote the probability of selecting into treatment given $(X,Z)$. Given the structure of this cross-threshold model, we note that $P(X, Z)=F_\varepsilon\left(\mu_\mathcal{W}(X, Z)\right)$, where $F_{\varepsilon_{\mathcal{W}}}(\cdot)$ denotes the distribution of $\varepsilon_{\mathcal{W}}$. For ease of exposition, we denote $P(X, Z)$ by $P$, suppressing the $(X, Z)$ argument. In addition, we use the fact that $U_\mathcal{W}=$ $F_{\varepsilon_\mathcal{W}}(\varepsilon_{\mathcal{W}})$ is a uniform random variable. In particular, different values of $\overline{\varepsilon}_\mathcal{W}$ denote different quantiles of $\varepsilon$. Given our previous assumptions, $F_{\varepsilon_\mathcal{W}}$ is strictly increasing, and $P(X, Z)$ is a continuous random variable conditional on $X$. Throughout this section, we assume the following.

assumptionThe following conditions are assumed to hold: \begin{itemize} • $( \nu_0, \nu_1, \nu_C)$ is independent of $(X,Z)$. • The distribution of $\mu_C(Z)$ conditional on $X$ is absolutely continuous with respect to Lebesgue measure. • The distribution of $\varepsilon_{\mathcal{W}}=\nu_C-\left(\nu_1-\nu_0\right)$ is absolutely continuous with respect to Lebesgue measure and has a cumulative distribution function that is strictly increasing. • The population means $\mathbb{E}(\left|V_1\right|), \mathbb{E}(\left|V_0\right|)$, and $\mathbb{E}(|C|)$ are finite. \end{itemize}

Parameters

Our goal is to learn about the distributions of benefits, costs, and welfare—objects that are unobserved by the analyst. To do so, we rely on several key parameters introduced by Eisenhauer_Heckman_Vytlacil_2015, which provide identification of the relevant conditional average objects that our analysis builds upon.

The first parameter is the conditional average treatment effect (ATE) benefit given by: $$ B^{\operatorname{ATE}}(x) \triangleq \mathbb{E}\left(Y_1-Y_0 \mid X=x\right)=\mu_1(x)-\mu_0(x) . $$

$B^{\operatorname{ATE}}(x)$ denotes the ATE for individuals with characteristics $X = x$: that is, the causal effect of assigning treatment randomly to all individuals of type $x$, under full compliance and abstracting from general equilibrium or spillover effects.

The second parameter of interest is the average treatment benefit for individuals who actually receive the treatment, commonly referred to as the benefit of treatment on the treated. $$

alignedB^{\mathrm{TT}}(x) & \triangleq \mathbb{E}\left(Y_1-Y_0 \mid X=x, D=1\right) \\ & =\mu_1(x)-\mu_0(x)+\mathbb{E}\left(\nu_1-\nu_0 \mid X=x, D=1\right).

$$

Heckman_Vytlacil_1999,Heckman_Vytlacil_ECMA_2005_MTE show that a uniform approach to $B^{\mathrm{ATE}}(x)$ and $B^{\mathrm{TT}}(x)$ is possible by using the MTE parameter, which is defined as:

$$

alignedB^{\mathrm{MTE}}\left(x, \overline{\varepsilon}_\mathcal{W}\right) & \triangleq \mathbb{E}\left(Y_1-Y_0 \mid X=x, U_\mathcal{W}=u_\mathcal{W}\right) \\ & =\mu_1(x)-\mu_0(x)+\mathbb{E}\left(\nu_1-\nu_0 \mid U_\mathcal{W}=u_\mathcal{W}\right) .

$$ The function $B^{\mathrm{MTE}}\left(x, u_\mathcal{W}\right)$ is the treatment effect parameter that conditions the unobserved desire to select into treatment.

Eisenhauer_Heckman_Vytlacil_2015 note that conventional treatment-effects analysis does not define, identify, or estimate any component of treatment costs. To fill this gap, they introduce three cost parameters: the average cost of treatment, the average cost of treatment for those who select into treatment, and the marginal cost of treatment.

$$

alignedC^{\mathrm{ATE}}(z) & =\mathbb{E}(C \mid Z=z)=\mu_C(z), \\ C^{\mathrm{TT}}(z) & = \mathbb{E}(C \mid Z=z, D=1),\\ &=\mu_C(z)+\mathbb{E}\left(\nu_C \mid Z=z, D=1\right),\\ C^{\mathrm{MTE}}\left(z, u_{\mathcal{W}}\right) & =\mathbb{E}\left(C \mid Z=z, U_\mathcal{W}=u_\mathcal{W}\right), \\ & =\mu_C(z)+\mathbb{E}\left(\nu_C\mid U_\mathcal{W}=u_\mathcal{W}\right).

$$

Finally, we define a set of welfare parameters. Recalling that $\mathcal{W}=B-C=\mu_\mathcal{W}(X, Z)-\varepsilon_{\mathcal{W}}$ we get: $$

aligned\mathcal{W}^{\mathrm{ATE}}(x, z) & =\mathbb{E}(\mathcal{W} \mid X=x, Z=z)=\mu_\mathcal{W}(x, z), \\ \mathcal{W}^{\mathrm{MTE}}\left(x, z, u_\mathcal{W}\right) & =\mathbb{E}\left(\mathcal{W} \mid X=x, Z=z, U_\mathcal{W}=u_\mathcal{W}\right) \\ & =\mu_\mathcal{W}(x, z)-\mathbb{E}\left(\nu_C \mid U_\mathcal{W}=u_\mathcal{W}\right)

$$ and $$

aligned\mathcal{W}^{\mathrm{TT}}(x, z) & =\mathbb{E}(\mathcal{W} \mid X=x, Z=z, D=1) \\ & =\mu_\mathcal{W}(x, z)-\mathbb{E}(\varepsilon \mid X=x, Z=z, D=1) .

$$

Eisenhauer_Heckman_Vytlacil_2015 show how to identify the previous parameters. Proposition (ref) below establishes that their average identification results help us to learn about the distributional aspects of the welfare distribution.

propositionLet Assumption (ref) hold. Then for $\beta\in[0,1)$ \begin{equation} \mathbb{S}_\beta(\mathcal{W})\leq \mathbb{S}_\beta(\mathcal{W}^{\operatorname{ATE}}(X,Z)) \end{equation} and \begin{equation} \mathbb{S}_\beta(\mathcal{W})\leq \mathbb{S}_\beta(\mathcal{W}^{\operatorname{MTE}}(X,Z, U_\mathcal{W})). \end{equation}

The result in Proposition (ref) shows that $\mathcal{W}^{\operatorname{ATE}}(X,Z)$ and $\mathcal{W}^{\operatorname{MTE}}(X,Z,U_{\mathcal{W}})$ constitute the best available approximations to the unobserved welfare effect $\mathcal{W}$. Consequently, they can be used to study the distributional behavior of $\mathcal{W}$ and to bound potential welfare losses for specific subpopulations. Intuitively, $\mathbb{S}_{\beta}(\mathcal{W})$ captures the welfare effect among the worst-off $(100\times\beta)\%$ of individuals, while $\mathbb{S}_{\beta}(\mathcal{W}^{\operatorname{ATE}}(X,Z))$ in bound ((ref)) captures the worst outcomes only across groups defined by $(X,Z)$. Similarly, in bound ((ref)), $\mathbb{S}_{\beta}(\mathcal{W}^{\operatorname{MTE}}(X,Z,U_{\mathcal{W}}))$ captures the worst outcomes only across groups defined by $\left(X, Z, U_{\mathcal{W}}\right)$. This distinction highlights the role of unobserved heterogeneity in assessing distributional impacts across subgroups. To the best of our knowledge, bounds in Proposition (ref) are new in the context of generalized Roy models with participation costs.

This distinction highlights the role of unobserved heterogeneity in assessing distributional impacts across subgroups. To the best of our knowledge, bounds in Proposition (ref) are new in the context of generalized Roy models with participation costs.

It is worth remarking that the bounds ((ref)) and ((ref)) explicitly incorporate the cost of treatment participation. Consequently, the result provides information about the potential welfare losses faced by different subpopulations as a function of observables $(X,Z)$, the cost $C$, and unobserved heterogeneity $U_{\mathcal{W}}$. Furthermore, our framework extends Eisenhauer_Heckman_Vytlacil_2015's results by providing lower bounds that identify which groups incur the highest costs. To do so, we focus on the right superquantile $\overline{\mathbb{S}}_\alpha(\cdot)$, which—as discussed earlier—captures the average outcome in the upper tail of the distribution and is therefore the appropriate object for assessing individuals who face the highest participation costs.

propositionLet Assumption (ref) hold. Then \begin{equation} \overline{\mathbb{S}}_\alpha(C)\geq \overline{\mathbb{S}}_\alpha(C^{\operatorname{ATE}}(Z))\quad for $\alpha\in[0,1)$ \end{equation} and \begin{equation} \overline{\mathbb{S}}_\alpha(C)\geq \overline{\mathbb{S}}_\alpha(C^{\operatorname{MTE}}(Z,U_\mathcal{W})) \end{equation}

The previous result provides information on the ATE and MTE costs of the $(100\times (1-\alpha ))\%$ worst affected. Expression ((ref)) provides a bound in terms of observables $Z$. Intuitively, this lower bound allows the analyst learn which particular subgroups are bearing a higher cost. Accordingly, this bound can guide the reduction of treatment costs in target subgroups, which can be welfare-improving. Similarly, bound ((ref)) complements the previous analysis by incorporating the role of unobservables.

Finally, we point out that Eisenhauer_Heckman_Vytlacil_2015's analysis is conducted in terms of average costs. Proposition (ref) shows that their results are informative about the distributional implications of treatment costs, considering both observables and unobservables.

Treatment on the treated

Our final result concerns the distributional properties of $\mathcal{W}^{\operatorname{TT}}$ and $C^{\operatorname{TT}}$.\footnote{For ease of exposition, we omit the analysis of $B^{\operatorname{TT}}$. However, all arguments extend directly to that case.} Because of the missing-outcome problem, neither distribution is observable to the analyst (HECKMAN2007_Handbook_partI). Our objective is to exploit the identification results in Eisenhauer_Heckman_Vytlacil_2015 for the parameters $\mathcal{W}^{\operatorname{TT}}(X,Z)$ and $C^{\operatorname{TT}}(Z)$ in order to recover informative bounds on the unknown distributions of $\mathcal{W}^{\operatorname{TT}}$ and $C^{\operatorname{TT}}$.

In doing so, we make use of the following notation:

eqnarray*[eqnarray* omitted — 233 chars of source]

In a similar way, for the cost $C$ we define :

eqnarray*[eqnarray* omitted — 222 chars of source]

Intuitively, $\mathbb{S}_\beta\!\left(\mathcal{W}^{\mathrm{TT}}\right)$ is the superquantile of the welfare distribution for the treated. It equals the average welfare among the $(100\times \beta)\% $ of treated individuals who experience the lowest welfare gains, including welfare losses. Similarly, $\overline{\mathbb{S}}_\alpha(C^{\mathrm{TT}})$ is the superquantile of the treatment-cost distribution for the treated. It corresponds to the average treatment cost for the $(100\times (1-\alpha))\%$ of treated individuals who incur the highest program costs.

The following is the main result of this section.

theoremLet Assumption (ref) hold. Then: \begin{itemize} • For $\beta\in(0,1]$ \begin{equation} \mathbb{S}_\beta(\mathcal{W}^{TT})\leq \mathbb{S}_\beta(\mathcal{W}^{TT}(X,Z)). \end{equation} • For $\alpha\in[0,1)$ \begin{equation} \overline{\mathbb{S}}_\alpha(C^{\operatorname{TT}})\geq \overline{\mathbb{S}}_\alpha(C^{\operatorname{TT}}(X,Z)) \end{equation} \end{itemize}

Some remarks are in order. First, part (i) bounds the welfare of the $(100\times \beta)\%$ worst-affected individuals in the population. To see the relevance of bound ((ref)), assume without loss of generality that $\mathcal{W}$ has a continuous distribution. In this case, $\mathbb{S}_\beta(\mathcal{W}^{\mathrm{TT}}) = \mathbb{E}\!\left[\mathcal{W}\mid \mathcal{W}\le F_{\mathcal{W}}^{-1}(\beta),\, D=1\right]$. However, because of the fundamental problem of missing outcomes, neither the distribution of $\mathcal{W}$ nor its conditional distribution (conditional on D=1) is identified. Thus $\mathbb{E}\!\left[\mathcal{W}\mid \mathcal{W}\le F_{\mathcal{W}}^{-1}(\beta), D=1\right]$ is not observed by the analyst. By contrast, using the identified function $\mathcal{W}^{\mathrm{TT}}(X,Z)$, we can construct the upper bound $\mathbb{E}\!\left[\mathcal{W}(X,Z)\mid \mathcal{W}(X,Z)\le F_{\mathcal{W}(X,Z)}^{-1}(\beta),\, D=1\right]$, which provides an informative measure of the potential gains or losses experienced by the bottom $(100\times \beta)\%$ of the welfare distribution among treated individuals. Thus, based on observables, the policymaker can identify which treated subgroups are likely to experience declines in welfare.

Part (ii) establishes a lower bound on the treatment cost among individuals who received the treatment. As with $\mathcal{W}^{\mathrm{TT}}$, the conditional distribution of $C^{\mathrm{TT}}$ is not identified. However, using the results in Eisenhauer_Heckman_Vytlacil_2015, the lower bound ((ref)) is identified and can be computed using the linear programming procedure described in Section (ref). From an economic perspective, the lower bound ((ref)) is informative along at least two dimensions. First, it enables the analyst to determine which subpopulations bear higher treatment costs among the treated. Second, it sheds light on the fairness of treatment costs across groups—for example, whether a particular intervention is regressive or progressive.

We close this section by noting that, to the best of our knowledge, Theorem (ref) is new in the cost-benefit analysis of treatment effects.

Conclusion

This paper develops a distributional framework for analyzing welfare heterogeneity across individuals. By combining the concept of superquantiles with observable-group average welfare effects, we show how to construct informative bounds on individual welfare changes without relying on strong structural assumptions or detailed micro-level data.

Although the framework is broadly applicable, much of the analysis is conducted within the class of perturbed utility models (PUMs), which offer a empirically tractable setting for nonparametrically estimating average welfare effects using standard data. We illustrate the usefulness of the framework in three economic environments: compensating variation from price changes, treatment decisions with endogenous participation, and social programs modeled through generalized Roy models with subjective participation costs. Across these applications, the core result established in Section (ref) extends naturally, even though the relevant welfare concepts and identification strategies differ.