EconBase
← Back to paper

Assessing Heterogeneity of Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,930 characters · 21 sections · 51 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Assessing Heterogeneity of Treatment Effects

abstractHeterogeneous treatment effects are of major interest in economics. For example, a poverty reduction measure would be best evaluated by its effects on those who would be poor in the absence of the treatment, or by the share among the poor who would increase their earnings because of the treatment. While these quantities are not identified, we derive nonparametrically sharp bounds using only the marginal distributions of the control and treated outcomes. Applications to microfinance and welfare reform demonstrate their utility even when the average treatment effects are not significant and when economic theory makes opposite predictions between heterogeneous individuals. JEL Codes: C01, I38, J21.

Introduction

Many questions in economics involve knowledge about heterogeneous treatment effects across different outcome levels. Consider, for example, the evaluation of microfinance as a poverty reduction measure. Suppose that the outcome of interest is wealth. If the average treatment effect (ATE) is positive, the access to microfinance increases individuals' wealth on average. However, if this is brought about by rich people being richer and poor people being poorer, the policymaker may want to think twice before rushing to implementation. In another example, consider evaluating a welfare reform on low\hypincome workers. The effect of the reform can be complex in that the anticipated income effect and price effect might point at different directions, even under simple and stylized economic theory. An important piece of information is how many workers choose to work more under the new welfare program.

This paper develops new methods to help answer these questions. To fix ideas, let $Y_{i1}$ be subject $i$'s treated outcome and $Y_{i0}$ subject $i$'s control outcome. Two quantities are of interest here. The first is the {\em subgroup treatment effect (STE)}, \[ \mathbb{E}[Y_{i1}-Y_{i0}\mid Y_{i0}<c], \] where $c$ is some number. In the first example, $Y$ corresponds to the wealth of individual $i$, so the STE represents the ATE of the subgroup defined by low wealth. The second quantity is the {\em subgroup proportion of winners (SPW)}, \[ P(Y_{i1}>Y_{i0}\mid Y_{i0}<c). \] In the second example, $Y$ corresponds to the earnings from work for individual $i$, so the SPW gives the share of low\hypearning individuals who earn more under the new welfare program. \footnote{If paid by hours, the hours worked are proportional to earnings. Then, we might regard higher earnings as more hours worked.}

Both quantities, however, require knowledge of individual treatment effects (ITEs) because of the conditioning on $Y_{i0}$. Since ITEs are not identified, neither quantity is point\hypidentified. However, the marginal distributions of $Y_{i0}$ and $Y_{i1}$ put restrictions on the maximum ranges in which they lie. Consequently, we seek bounds that are as tight as possible with no additional assumptions.

Our bounds provide useful complementary information to the widely\hypused quantile treatment effect (QTE). The QTE is often used to display the degree of heterogeneity immanent in the data but overlooked by the ATE bgh2006. However, the QTE does not come with a handy interpretation as the ITE unless one is willing to assume rank invariance; hence it calls for a delicate interpretation to say anything more than the mere existence of heterogeneity. Our bounds are available almost whenever QTEs are, and they characterize the furthest extent to which the data can speak about the heterogeneity in the forms of STEs and SPWs.

The bounds depend only on the marginal distributions of the treated outcome and the control outcome, so they are applicable widely beyond the randomized controlled trial (RCT) framework. \footnote{We do require, however, that the treatment be binary.} For example, there are existing estimators of the potential outcome distributions in the binary IV models with monotone compliance, selection\hypon\hypobservables models, panel data, and regression discontinuity designs. These estimators were mainly proposed as intermediate quantities to obtain the QTE, but instead of (or in addition to) calculating the QTE, we can feed them into our bounds to derive additional, interpretable insights on heterogeneity. The Stata program to compute the bounds from the RCT data (with monotone compliance) is available at an \href{https://github.com/jcao0/subgroup-treatment-effects}{author's website}.

This paper adds to the enormous literature on partial identification of heterogeneity measures in program evaluation. The literature is too vast for us to give a comprehensive review, but some papers close to ours include hsc1997, fp2010, and t2012. hsc1997 discussed bounds on the proportion of winners and the distribution of treatment effects and investigated how the assumption of rational choice by the participants helps reduce uncertainty. fp2010 gave bounds on the distribution of treatment effects in terms of second\hyporder stochastic dominance. t2012 derived bounds on positive and absolute treatment effects and the proportion of winners on the entire population. In contrast, we consider subgroups that are defined by the ranges of control outcome, giving a finer picture on the distribution of the treatment effects. jls2023 considered robust estimation of partially identified causal effects exploiting information of covariates. When we wish to incorporate covariates into estimation, their method provides a convenient way to produce reliable estimators of the bounds.

Our STE bounds are also related to the Lee bounds on the treatment effect under sample selection hm2000,l2009a. Suppose $Y_i$ is wage, $D_i$ is eligibility to a job training program, and $S_i$ the indicator of $i$'s employment status. Since the wage is only observed when employed, the outcome is missing when $S_i=0$. Denote by $S_{id}$ the potential indicator of employment when $D_i=d$. They concern bounding the effect of the job training program on wages for individuals who would be employed under both treatment statuses, $\mathbb{E}[Y_{i1}-Y_{i0}\mid S_{i0}=1,S_{i1}=1]$. Under the standard assumptions of the literature, this boils down to bounding $\mathbb{E}[Y_{i1}\mid S_{i0}=1,S_{i1}=0]$ when the distribution of $Y_{i1}$ given $S_{i1}=1$ and $P(S_{i0}=1\mid S_{i1}=1)$ are identified. We have an analogous situation: our objective is to bound $\mathbb{E}[Y_{i1}-Y_{i0}\mid Y_{i0}<c]$ when the marginal distribution of $Y_{i1}$ and $P(Y_{i0}<c)$ are identified.

The rest of the paper is organized as follows. (ref) motivates our bounds with several examples in development economics, health economics, and labor economics. (ref) lays out the formal statements of our results. (ref) presents an application of the STE to microfinance, drawn from tdj2015. (ref) gives an application of the SPW to the evaluation of a welfare reform, borrowing from bgh2006. (ref) concludes. All proofs are contained in (ref).

Motivating Examples

We discuss a few stylized examples to motivate each of the STE and SPW. In all examples, and throughout the paper, we maintain the potential outcomes notation: $Y_{i1}$ for the treated and $Y_{i0}$ for the control outcomes.

Subgroup Treatment Effects

exa[Microfinance for the poor] Microfinance allows poor households to borrow a small amount of money, with which they can invest in crops or livestock to escape from a poverty trap. The outcome of interest is therefore a measure of wealth or income, such as the value of livestock owned or the net revenues from crops. Let $Y_i$ be the value of livestock owned by $i$. On top of the ATE, $\mathbb{E}[Y_{i1}-Y_{i0}]$, we are interested specifically in the ATE for those who would have had low values of livestock in possession, \[ \mathbb{E}[Y_{i1}-Y_{i0}\mid Y_{i0}<Y_0^b], \] where $Y_0^b$ is the $b$th quantile of $Y_{i0}$. It is of interest to look at this value at various values of $b\in(0,1)$, and if it is positive for most low values of $b$, we might conclude that microfinance helps the poor. In (ref), we revisit this example using the RCT data from tdj2015. They gave estimates of the ATE \footnote{They considered access to microfinance as the treatment. If one sees borrowing from microfinance as the treatment, the ATE estimates can be understood as the intent\hypto\hyptreat (ITT) estimates.} and stated \begin{quote} Despite the large increase in borrowing, we find that for a large majority of socioeconomic outcomes the null of no impact cannot be rejected, although in several cases the point estimates are substantially large but imprecisely estimated. \end{quote} When we apply our bounds for the outcome of livestock value, we find that the STE for the poor is significantly positive, and the impreciseness of the estimates is most likely due to the large variation of the livestock values of the upper tail.
exa[Healthcare for the unhealthy] Many unhealthy individuals are not covered by health insurance, and it is of interest to know whether they would become healthier if they were covered. Therefore, the treatment of interest is the access to healthcare, and the outcome of interest is a measure of health such as the blood pressure. Let $Y_i$ be the blood pressure and $[c_1,c_2]$ be the ideal range thereof. One quantity of interest is the effect of access to healthcare on the blood pressure for those with otherwise low blood pressure, \[ \mathbb{E}[Y_{i1}-Y_{i0}\mid Y_{i0}<c_1]. \] If this is positive, availability of healthcare improves the health of those with low blood pressure. We may also be interested in the effect on the blood pressure for those with otherwise high blood pressure, \[ \mathbb{E}[Y_{i1}-Y_{i0}\mid Y_{i0}>c_2]. \] If this is negative, healthcare is beneficial for those as well. As the direction of benefits changes depending on where the control outcome is, such health benefits are difficult to capture by ATEs.

Subgroup Proportions of Winners

exa[Evaluation of welfare reform] When replacing an old welfare program with a new one, Connecticut conducted a randomized experiment to evaluate its effect. In essence, this reform replaced a mild indefinite support for women with a more generous but time\hyplimited support. bgh2006 observed that theory predicts heterogeneous effects regarding the sign and magnitude of the response of labor supply. Individual $i$'s income can be decomposed into earnings and government transfers. Let $Y_i$ be $i$'s earnings. As we will see in (ref), stylized theory predicts positive effects on earnings for low\hypearnings workers before the time limit, so if both leisure and consumption are normal goods for a good amount of workers, then we expect \[ P(Y_{i1}>Y_{i0}\mid Y_{i0}<c) \] to be positive for a low value of $c$. After presenting QTEs, bgh2006 stated \begin{quote} Without further assumptions, the possibility of rank reversals prevents us from being more specific about who the winners and losers are. \end{quote} In (ref), we apply our bounds on this example and find that the lower bound is significantly positive.
exa[Do-no-harm principle] The principle of “do no harm” is considered to be a primary ethical standard of the medical profession. Let $Y_i$ be a biomarker that measures the level of health. Assume for simplicity that the larger the better in the observed range. Then, the treatment is beneficial to $i$ when $Y_{i1}>Y_{i0}$ (a “winner” from the treatment), while it is harmful to $i$ when $Y_{i1}<Y_{i0}$ (a “loser” from the treatment). If there are losers among those with low levels of $Y_{i0}$, that is, $Y_{i0}<c$, that means the treatment exacerbates the conditions of those who are already bad. Thus, it is of interest to know the magnitude of \[ P(Y_{i1}<Y_{i0}\mid Y_{i0}<c) \] in deciding a suitable treatment.
exa[Persuasion effect] When the provision of information is the treatment, of interest is whether the information affects the decision of the information receiver. This is called the {\em persuasion effect}. dk2007 considered the persuasion effect of Fox News on voting for the Republican candidate in Presidential elections. To be precise, let $Y_i$ be a binary indicator of whether voter $i$ votes for the Republican candidate. The treatment is the availability of Fox News in $i$'s region. Persuasion by Fox News is defined as $Y_{i0}=0$ and $Y_{i1}=1$, that is, in the absence of Fox News, $i$ votes for the Democrat candidate, but with Fox News, for the Republican. This corresponds to the “winner” in our paper. The quantity of interest is the {\em persuasion rate}, \[ P(Y_{i1}>Y_{i0}\mid Y_{i0}=0). \] jl2023 showed that the persuasion rate is partially identified in empirically relevant setups and derived sharp bounds thereof. In fact, our bounds ((ref)) reduce to equation (3) of their paper when $Y$ is binary. Now, the results of this paper can be considered a natural extension of the persuasion effect to continuous decisions, e.g., whether the medical recommendation to sleep longer induces patients to sleep longer and how such effects differ across different levels of initial hours of sleep.

Main Results

Setting

We do not directly observe the potential outcomes $Y_{i0}$ and $Y_{i1}$. Instead, we observe the treatment indicator $D_i$ and the corresponding outcome $Y_i=D_i Y_{i1}+(1-D_i)Y_{i0}$. We let $F_0$ and $F_1$ denote the marginal cumulative distribution functions (cdfs) of $Y_{i0}$ and $Y_{i1}$ for the population of interest.

Depending on the application, $F_0$ and $F_1$ for the intended population may be identified through different mechanisms. Here, we list some examples. Thereafter, we take the identification of $F_0$ and $F_1$ as given.

itemize• {\em Case 1 (RCT):} The treatment indicator $D_i$ is exogenous. This arises when the treatment is randomized or when the assignment is randomized but the focus is on the ITT. In this case, $F_0$ and $F_1$ are trivially identified and correspond to the cdfs of outcomes for the entire population in the experiment. • {\em Case 2 (Imperfect compliance):} We observe a binary instrument $Z_i$ that satisfies monotonicity. This arises when assignment to the treatment is randomized. If one\hypsided compliance holds, $F_0$ and $F_1$ for the compliers are identified as a normalized difference of observable cdfs a2002. \footnote{The plug\hypin estimator is not guaranteed to be monotonic, but we can rearrange it to be cfg2009.} This has been extended to cases where $Z_i$ is valid conditional on covariates aai2003,fm2013,p2020. • {\em Case 3 (Difference\hypin\hypdifferences):} We observe cross\hypsectional pre\hyptreatment outcomes. \footnote{By “cross\hypsectional” we mean that individuals need not be tracked over time.} If the “change\hypin\hypchanges” assumption holds, $F_0$ and $F_1$ for the treated are identified as the composition of observable cdfs and quantiles ai2006. This has been generalized to synthetic control by g2023. He proposed to construct $F_0$ for the treated as a convex combination of those for the control, where the weights are determined as the minimizer of the 2\hypWasserstein distance in the pre\hyptreatment periods. • {\em Case 4 (Selection on observables):} We observe covariates $W_i$ such that $D_i$ is independent of $(Y_{i0},Y_{i1})$ conditional on $W_i$. If the common support assumption holds, $F_0$ and $F_1$ for the whole sample and for the treated are identified through propensity score weighting f2007. • {\em Case 5 (Regression discontinuity design):} We observe the running variable $R_i$ such that the propensity score $P(D_i=1\mid R_i)$ jumps discontinuously at a cutoff $R_i=r_0$. Under the assumption that there are no defiers at the cutoff, $F_0$ and $F_1$ for the compliers at the cutoff are identified, either for sharp or fuzzy regression discontinuity designs ffm2012.

In all cases, we may also observe exogenous covariates $X_i$ and wish to condition our analysis on the values of $X_i$. Then, we will be concerned of the conditional cdfs of $Y_{i0}$ and $Y_{i1}$ conditional on $X_i$. Unless $X_i$ is finitely supported, this calls for some modeling of the conditional distribution functions ch2006,aai2003. For simplicity, we hereafter suppress $X_i$ in our notation and understand implicitly that all results are allowed be conditioned on $X_i$.

From this point on, we take for granted that $F_0$ and $F_1$ for the population of interest are identified. Other than that, we do not make any more assumptions; notably, we do not require $Y_{i0}$ or $Y_{i1}$ to be continuous or discrete.

Definition of the Subgroup

Let $Q_0$ and $Q_1$ be the quantile functions corresponding to $F_0$ and $F_1$, that is, \[ Q_j(u)=\inf\{x\in\mathbb{R}:F_j(x)\geq u\} \] for $0<u<1$ and $j=0,1$. In plain words, $Q_j$ is the inverse of $F_j$ but assigns the smallest value when there are flat regions of $F_j$.

To define subgroups, we introduce the {\em rank} of individual $i$ in terms of the control outcome, to be denoted by $U_i\in[0,1]$. This rank is defined so that it satisfies \[ Y_{i0}=Q_0(U_i) \] for each individual in the population. If $Y_{i0}$ has a continuous distribution, there is a one\hypto\hypone relationship between $U_i$ and $Y_{i0}$ such that $U_i=F_0(Y_{i0})$. If $Y_{i0}$ has a discrete mass, there can be many distinct values of $U_i$ that correspond to a mass value of $Y_{i0}$ and we have $U_i\leq F_0(Y_{i0})$ in general. In this case, we may understand $U_i$ as representing the unobservable heterogeneity of otherwise observationally equivalent individuals in this mass. It is without loss of generality to assume that $U_i$ is distributed uniformly over $[0,1]$.

Consequently, the subgroup \( \{Y_{i0}<c\} \) can also be represented as \( \{U_i<F_0(c)\} \), but the subgroup \( \{U_i<b\} \) is not necessarily identical to \( \{Y_{i0}<Q_0(b)\} \). This means that conditioning on the values of $U_i$ is {\em finer} than conditioning on the values of $Y_{i0}$, that is, every conditioning on $Y_{i0}$ can be represented by some conditioning on $U_i$, but the converse is not true. For this generality, we focus on the subgroups of the form \[ \{a<U_i<b\} \] for various values of $0\leq a<b\leq 1$, rather than of the form $\{a<Y_{i0}<b\}$.

The introduction of this slight hassle will pay off when we plot the bounds against the rank. Conditioning on $U_i$ produces smooth curves that traverse over continuously increasing subpopulations no matter what the distribution of $Y_{i0}$ is. \footnote{For example, the Lipschitz property in (ref) is thanks to this formulation.}

Subgroup Treatment Effects

Statement

The following theorem gives the sharp lower and upper bounds on the STE for the subgroup defined by $\{a<U_i<b\}$.

thm[Bounds on subgroup treatment effects] For every $0\leq a<b\leq 1$, \begin{align} \mathbb{E}[Y_{i1}-Y_{i0}\mid a<U_i<b]&\geq\frac{1}{b-a}\int_a^b[Q_1(u-a)-Q_0(u)]du,\\ \mathbb{E}[Y_{i1}-Y_{i0}\mid a<U_i<b]&\leq\frac{1}{b-a}\int_a^b[Q_1(1+a-u)-Q_0(u)]du, \end{align} provided that the conditional expectation exists. For each of the bounds, there exists a joint distribution of $(Y_{i0},Y_{i1})$ that attains it.

These bounds are nonparametrically sharp, meaning that there exists a joint distribution of $(Y_{i0},Y_{i1})$, having marginal distributions $F_0$ and $F_1$, that attains each bound. To name a few intuitive cases, the lower bound ((ref)) for $a=0$ is attained when {\em rank invariance} holds, i.e., $Y_{i1}=Q_1(U_i)$; the upper bound ((ref)) for $a=0$ is attained when {\em complete rank reversal} holds, that is, $Y_{i1}=Q_1(1-U_i)$.

figure[figure omitted — 319 chars of source]

The intuition of the bounds is visualized in (ref). The left figure plots $Q_0$, and the red region describes our subpopulation of interest, $\{a<U_i<b\}$. The right figure plots $Q_1$. While we do not know where in $Q_1$ each individual in the red region corresponds to, we know that the lowest possible ATE these people can receive {\em collectively} is when they are matched with the blue region on $Q_1$. This gives ((ref)). Note that the exact correspondence between the points in red and those in blue can be left unspecified for this bound to be exact. Also, the highest possible ATE they can receive is when they are matched with the green region on $Q_1$. This gives ((ref)). If we let $a=b$, they reduce to the trivial bounds, \[ Q_1(0)-Q_0(a)\leq\mathbb{E}[Y_{i1}-Y_{i0}\mid U_i=a]\leq Q_1(1)-Q_0(a). \]

The above intuition suggests that the lower bound would be the most informative when $a=0$ and the upper bound when $b=1$. In fact, in the application in (ref), we fix $a=0$ (since we focus on the poor) and view the bounds as functions of $b$.

While the QTE cannot be interpreted as the ITE in the absence of rank invariance, ((ref)) states that the {\em integral} of the QTE can be interpreted as the lower bound for the STE with $a=0$ {\em without requiring} rank invariance. As stated earlier, this lower bound is exact when rank invariance holds, but it is also expected to be close when rank invariance approximately holds (see also (ref) in the next section). The idea that individuals do not move ranks too drastically by the treatment might appear reasonable in economic applications.

The natural estimator of the bounds is the plug\hypin estimator, where the quantile functions are replaced by the estimated ones. For example, in the ITT analysis in (ref), we plug in the empirical quantile functions of $Y_i\mid D_i=0$ and of $Y_i\mid D_i=1$ into $Q_0$ and $Q_1$.

If the estimated bounds are asymptotically jointly normal for fixed $a$ and $b$, we may construct an asymptotically valid confidence interval for the STE using the methods developed by im2004 and s2009. This kind of confidence interval does {\em not} cover the true bounds (an interval) but {\em does} cover the true STE (a single point) with a desired confidence level. This is also implemented in (ref).

Normal Distribution Example

figure[figure omitted — 954 chars of source]

We illustrate the STE using an example where the outcomes are normally distributed. Let $Y_{i0}\sim N(0,2)$ and $Y_{i1}\sim N(1/2,1)$, so their quantile functions look like (ref). If we specify their joint distribution to be normal, we can compute the unidentified population STE.

(ref) plots the STE, \[ \mathbb{E}[Y_{i1}-Y_{i0}\mid U_i<b], \] as a function of $b$ for various values of correlation between $Y_{i0}$ and $Y_{i1}$, namely for $\rho=-1$, $-0.5$, $0$, $0.5$, and $1$. The color represents the value of correlation as given by the color bar on the right.

For every value of $b$, the population $\{U<b\}$ receives the lowest STE when $Y_{i0}$ and $Y_{i1}$ are perfectly correlated, and the highest when they are perfectly negatively correlated. At any rate, we can infer that the lower group receives the STE higher than the ATE. The gray shaded area represents the region covered by our bounds.

(ref) plots the STE for $\mathbb{E}[Y_{i1}-Y_{i0}\mid U_i>a]$. On the contrary, the population $\{U>a\}$ receives the STE lower than the ATE. This STE is the highest when the outcomes are perfectly correlated and the lowest when perfectly negatively correlated.

In applications where rank migration is mild, we expect the true STEs to be close to the lower bound for $\{U<b\}$ and the upper bound for $\{U>a\}$.

Welfare Bounds

Sometimes the ATE is used for welfare analysis. We may also derive the bounds on possibly {\em non\hyputilitarian} welfare. In fact, (ref) is an immediate corollary of the following result.

thm[Bounds on subgroup welfare] Let $f:\mathbb{R}\to\mathbb{R}$ be a nondecreasing convex function and $g:\mathbb{R}\to\mathbb{R}$ a nonincreasing convex function. For every $0\leq a<b\leq 1$, \begin{align*} \mathbb{E}[f(Y_{i1}-Y_{i0})\mathbbm{1}\{a<U_i<b\}]&\geq\int_a^b f(Q_1(u-a)-Q_0(u))du,\\ \mathbb{E}[f(Y_{i1}-Y_{i0})\mathbbm{1}\{a<U_i<b\}]&\leq\int_a^b f(Q_1(1-u+a)-Q_0(u))du, \end{align*} and \begin{align*} \mathbb{E}[g(Y_{i1}-Y_{i0})\mathbbm{1}\{a<U_i<b\}]&\geq\int_a^b g(Q_1(1-b+u)-Q_0(u))du,\\ \mathbb{E}[g(Y_{i1}-Y_{i0})\mathbbm{1}\{a<U_i<b\}]&\leq\int_a^b g(Q_1(b-u)-Q_0(u))du, \end{align*} provided that the expectations exist. For each of the bounds, there exists a joint distribution of $(Y_0,Y_1)$ that attains it.

Non\hyputilitarian welfare might arise as a consequence of loss aversion. When the provider of the treatment may be held responsible for the {\em negative} effects of the treatment, such as doctors b2009 or policymakers nh2020, the provider might be interested in maximizing a non\hyputilitarian welfare. \footnote{On a related note, hsc1997 questioned utilitarian welfare in the context where individuals can choose treatment.} In (ref), we use this result to illustrate the policy that maximizes the worst\hypcase welfare.

While the claim is intuitive to understand, the proof without a distributional assumption on $Y$ calls for extention of classical optimal transport to conditioning on the subgroup. We achieve this by alternating the characterizations of the bound and of the optimal transport plan, as detailed for (ref) in (ref).

(ref) can also be used to bound the positive and negative STEs,

gather*[gather* omitted — 116 chars of source]

These bounds may help policymakers examine the sign of the treatment effects. Note that the sum of the bounds on positive and negative effects does not give the tightest bound on the net effect as given in (ref). Therefore, if we want the bounds on all of the positive, negative, and net effects, they need to be calculated for each case separately.

Subgroup Proportions of Winners

Now we turn to the SPW. We introduce the notation $[x]_+=\max\{x,0\}$ and $F(a-)=\lim_{x\nearrow a}F(x)$. The mirror image of the SPW is the {\em subgroup proportion of losers (SPL)}, \[ P(Y_{i1}<Y_{i0}\mid a<U_i<b). \]

Statement

The following theorem characterizes the sharp lower and upper bounds on the SPW and SPL for the subgroup of interest $\{a<U_i<b\}$.

thm[Bounds on subgroup proportions of winners and losers] For every $0\leq a<b\leq 1$, \begin{align} P(Y_{i1}>Y_{i0}\mid a<U_i<b)&\geq\frac{1}{b-a}\sup_{u\in(a,b)}[u-a-F_1(Q_0(u))]_+,\\ P(Y_{i1}>Y_{i0}\mid a<U_i<b)&\leq 1-\frac{1}{b-a}\sup_{u\in(a,b)}[b-u-1+F_1(Q_0(u))]_+,\notag \end{align} and \begin{align} P(Y_{i1}<Y_{i0}\mid a<U_i<b)&\geq\frac{1}{b-a}\sup_{u\in(a,b)}[b-u-1+F_1(Q_0(u)-)]_+,\\ P(Y_{i1}<Y_{i0}\mid a<U_i<b)&\leq1-\frac{1}{b-a}\sup_{u\in(a,b)}[u-a-F_1(Q_0(u)-)]_+.\notag \end{align} For each of the lower bounds, there exists a joint distribution of $(Y_0,Y_1)$ that attains it; for each of the upper bounds, there exists a joint distribution of $(Y_0,Y_1)$ that makes it arbitrarily tight.

Similarly as the STE, these bounds are the most informative when we let either $a=0$ or $b=1$. In practice, therefore, we recommend fixing $a=0$ and looking at the bounds as functions of $b$, or fixing $b=1$ and looking at the bounds as functions of $a$.

figure[figure omitted — 478 chars of source]

The analytical formulas of the SPW bounds are more complicated than those of the STE bounds, but we may nevertheless attempt to intuit them as in (ref). Since the upper bounds are complements of the lower bounds, it suffices to focus on the lower bounds. Assume for simplicity that both outcomes are continuous and let $a=0$. For an arbitrary value of outcome $y$, the portion of individuals whose $Y_0$ is below $y$ is given by $F_0(y)$, and the portion whose $Y_1$ is below $y$ by $F_1(y)$ ((ref)). If $F_0(y)>F_1(y)$, then at least $F_0(y)-F_1(y)$ portion of individuals in the subgroup $\{Y_0\leq y\}$ must move from $Y_0\leq y$ to $Y_1>y$. They are obviously winners. Since this is true for every value of $y$, we can take the supremum in the range of interest. Redefining $y=Q_0(u)$ gives ((ref)).

Next, suppose instead that $F_1(y)>F_0(y)$ as in (ref). Then, the difference $F_1(y)-F_0(y)$ gives the sure losers in the group $\{Y_0>y\}$. Meanwhile, this group contains $1-b$ portion of individuals who do not belong to the subgroup of interest. The worst scenario is that all of those individuals count toward the sure losers. Thus, the difference \[ [F_1(y)-F_0(y)]-(1-b) \] provides the lower bound of losers in the subgroup $\{U<b\}$. Redefining $y=Q_0(u)$ and taking the supremum in the range produce ((ref)). This intuition may smack as loose, but it turns out to be sharp. The proof in (ref) contains a construction of the joint distribution that attains this bound.

Note that the conditional probabilities in (ref) are obviously continuous in $(a,b)$ even when $Y_0$ and $Y_1$ are discrete since $U$ is continuously distributed, while the bounds appear susceptible to discontinuities. Somewhat surprisingly, the bounds are shown to be continuous, making them sharp regardless of the distributions of $Y$.

(ref) can be trivially extended to bound \[ P(Y_{i1}-Y_{i0}<c\mid a<U_i<b) \] for arbitrary $c$ by shifting the distribution of either $Y_{i0}$ or $Y_{i1}$. If we then set $a=0$ and $b=1$, the bounds reduce to the classical Makarov bounds m1981. In this sense, (ref) generalizes the Makarov bounds to conditioning on $U_i$.

Note that the bounds in (ref) are in the form of {\em intersection}, that is, the bounds are given by suprema. Simple plug\hypin estimators for the bounds of this kind can be severely biased in finite samples. clr2013 developed a procedure that gives an asymptotically median\hypbias\hypcorrected estimator of the bounds as well as an asymptotically valid confidence interval for the partially identified parameter. The idea is to approximate the supremand function by a Gaussian process and adjust for the precision of its estimator. The estimation procedure goes as follows.

enumerate• Estimate the covariance function of the supremand function. • Simulate the zero\hypmean Gaussian process with the corresponding covariance function and obtain the median of its supremum (over a shrinking set). • Subtract the median times the standard error function from the supremand function. • Maximize the precision\hypcorrected supremand function.

In the application in (ref), we use this method to construct the estimators of our bounds and the pointwise confidence intervals for the SPW. As before, the confidence interval thusly constructed does not cover the true bounds but covers the true SPW with a specified confidence level.

Finally, as the bounds are attained at somewhat quirky distributions, we may be interested in restricting the set of joint distributions we consider. In the most general case, we may numerically solve the optimal transport problem for the SPW or SPL with respect to the restricted set, but we may lose the closed\hypform expressions of the bounds. In (ref), we restrict the support of the joint distribution in a tractable way so that we maintain the closed\hypform expressions given by (ref) while still tightening the bounds.

Normal Distribution Example

figure[figure omitted — 909 chars of source]

We use the same example as (ref) and illustrate the SPW. The marginal distributions are maintained to be $Y_{i0}\sim N(0,2)$ and $Y_{i1}\sim N(1/2,1)$. If we impose joint normality, we can compute the unidentified population SPW.

If they are perfectly correlated, roughly the lower 69% of the population gains from the treatment, and the upper 31% loses. Therefore, the true SPW is one until $b=0.69$ and then decreases to 0.69 as we approach $b=1$ (the dark red line in (ref)). Similarly, if the outcomes are perfectly negatively correlated, the lower 57% gains from the treatment; this is the dark blue line in (ref). Other correlation values yield nondegenerate distributions, and there are both winners and losers at every rank.

The gray area gives the region between our bounds. For example, at $b=0.3$, the value of the lower bound is $0.80$. This means that among the lower 30% of the population in terms of $Y_{i0}$, at least 80% of them must be winners, $Y_{i1}>Y_{i0}$. This drops to 50% when we expand the subgroup to $b=0.5$. At $b=0.95$, the lower bound is $0.26$ and the upper bound is $0.96$. This means that at least 26% of $\{U_i<0.95\}$ are winners and at least 4% are losers; the remaining 70% may be a winner, a loser, or neither.

The losers are easier to identify on the right tail. In the right figure, at $a=0.9$, the value of the upper bound is $0.20$. Note that the flipside of a winner is a loser. \footnote{In this example, $Y_{i0}=Y_{i1}$ occurs with probability zero.} Therefore, we can interpret that among the upper 10% subpopulation, about 80% of them have to be losers.

Note that these bounds are pointwise tight, but there is usually no joint distribution that attains the bounds uniformly over $a$ or $b$.

Application 1: STE for Microfinance

We present an empirical application to illustrate how our methods can be used to draw information on the heterogeneous treatment effects in practice. We estimate the bounds on the effect of access to microfinance on the total value of livestock for the poor. We borrow the setup and data from tdj2015, who analyzed the RCT of microfinance conducted in rural Ethiopia from 2003 to 2006. \footnote{The replication files were downloaded at an author's website tdj2015dta.}

The brief outline of the experiment is as follows. Randomization was carried out at the Peasant Association (PA) level, a local administrative unit. Out of 133 PAs, 34 PAs were randomly assigned to the treatment and 33 PAs to the control. The remaining 66 PAs were assigned to a different combination of treatments, which we will not use. The implementation agencies sometimes did not comply with the experimental protocol, and the actual treatment coincided with the assignment in 78% of the cases. In this sense, our analysis is on the ITT---the effect of random assignment---following tdj2015. We will call this the ATE throughout this section.

The data we use are from the postintervention survey, but there was also a preintervention survey where individuals were not tracked between the two. Therefore, another possible direction, which we will not pursue, is to look at the effects on the treated using the framework of ai2006. We leave further details of the experiment to tdj2015.

tdj2015 noted that “most loans were initiated to fund crop cultivation or animal husbandry, with 80 percent of the 1,388 loans used for working capital or investment in these sectors . . .” In light of this, we choose the total value of livestock owned as the outcome of interest.

Do the Poor Benefit from Microfinance?

We are interested in knowing how the assignment to microfinance impacts those who are most in need, i.e., those who would attain low outcomes in the absence of the treatment. \footnote{acw2016 stated that “many researchers and policymakers are interested in estimating how treatments affect those most in need of help, that is, those who would attain unfavorable outcomes in the absence of the treatment.”}

We let $Y_i$ be individual $i$'s total value of livestock owned. It is in the local currency units, Birr, in their 2006 value. \footnote{In January 2006, 100 Birr was roughly worth 11.4 USD.} We define “those in need” as follows. Sort individuals based on their (unobservable) potential outcome under control, $Y_{i0}$. Then, take a sequence of cumulative subgroups whose $Y_{i0}$ belongs to the lower tail. Precisely, using the rank notation, we look at the subgroups $\{U_i<b\}$ as we vary $b$ from $0$ to $1$.

First, the ATE is estimated to be 226.50 Birr, but the 95% confidence interval contains 0, ranging from $-12.43$ to 469.69. However, this does not mean that the data have nothing to say about the effects on the low $Y_{i0}$. (ref) plots the bounds on the STE characterized by (ref), \[ \mathbb{E}[Y_{i1}-Y_{i0}\mid U_i<b], \] for various values of $0<b\leq 1$. The thick black line gives the lower bound of the STE for each $b$, and the dashed line the upper bound. As mentioned in (ref), since we set $a=0$ in (ref), the lower bound gives a meaningful picture but the upper bound is of little value. As the upper bound becomes unreasonably high, we show only the relevant range.

Note that the STE is point\hypidentified and equals the ATE when $b=1$; thus, the two bounds converge to the ATE at $b=1$. Between $b=0$ to $b=0.169$, the lower bound is equal to zero. This is because in both the treated and the control groups, we have about 17% of individuals who do not possess any livestock.

figure[figure omitted — 329 chars of source]

The shaded area presents the pointwise 95% confidence interval for each $b$, constructed with the method developed by im2004. Although hard to see it in the figure, at $b=1$, the lower end of the confidence interval swiftly crosses $0$, yielding an insignificant ATE estimate. The lower confidence interval is in fact positive from $b=0.271$ to $b=0.999$. Thus, unlike the ATE, the STE for those who possess a small to medium amount of livestock is estimated to be significantly positive.

One possible factor behind the insignificant ATE is that the large values of livestock of the “rich” were highly volatile. As is the case for the income or wealth distribution, we see those who have so much that a small multiplicative variation causes high variation in the average.

Who Should Be Treated? A Welfare Analysis

The heterogeneous effects across different levels of $Y_{i0}$ motivate optimizing assignment based on pre\hyptreatment $Y_i$. If we have panel data in which pre\hyptreatment outcome is observed and is associated with individuals in the post\hyptreatment period, we may simply include it as a covariate and proceed on to maximizing the empirical welfare as in kt2018. If we do not have such data, we might want to find a proxy for the pre\hyptreatment outcome. A natural proxy for the pre\hyptreatment outcome is the potential outcome $Y_{i0}$, especially in applications where rapid migration of ranks is considered uncommon. \footnote{Another proxy would be a fitted outcome predicted by covariates acw2016.}

The data we use consist of two cross\hypsectional surveys, so it is not an option to include the pre\hyptreatment outcome at the individual level. Moreover, in the application where motivation is drawn from “helping the poor escape the poverty trap,” uninterventional rank migration is expected to be limited, at least between the poor and the rich. \footnote{Note that rank migration {\em within} the subgroup does not affect the value of the STE or its bounds.} Thus, we consider an optimal treatment assignment based on $Y_{i0}$ that maximizes welfare to be a good policy\hyprelevant measure.

We consider an assignment rule in which individuals with values of livestock below some cutoff are granted access to microfinance. Let $b$ denote the cutoff of assignment in terms of the rank. Note that we can also state the equivalent assignment rule in terms of the value of livestock, in which case the cutoff is $Q_0(b)$.

We consider two welfare functions. In the first welfare model, we let the {\em social benefit} equal to the aggregate treatment effect, $\mathbb{E}[(Y_{i1}-Y_{i0})\mathbbm{1}\{U_i<b\}]$, which is normalized to the per\hypcapita level. The {\em social cost}, on the other hand, is set to 100 Birr per individual for the following reasoning. The average amount of outstanding loans from microfinance for a treated individual was 299 Birr, and we are assuming conservatively that one third of it goes sour. This is admittedly an assumption too simplistic for a policy analysis, and in reality the default rate would depend on the initial level of $Y$ as well as on many other factors. However, to concentrate on the key ideas of the paper, we will not seek to find more realistic versions of the social costs.

figure[figure omitted — 570 chars of source]

In sum, the first social welfare we consider is \[ \mathbb{E}[(Y_{i1}-Y_{i0})\mathbbm{1}\{U_i<b\}]-100\cdot\mathbb{E}[\mathbbm{1}\{U_i<b\}]. \] (ref) shows the social benefit bounds and social costs as functions of $b$. The solid blue line is the lower bound for the social benefits (the worst\hypcase social benefits), and the dashed blue line is the upper bound. The solid red line is the social costs. The shaded blue area is the pointwise 95% confidence interval for the social benefits. From $b=0.346$ to $b=1$, the worst\hypcase social benefits surpasses the social costs.

To determine the optimal assignment rule, we take the minimax approach: assign the treatment to maximize the worst\hypcase welfare. The worst\hypcase welfare corresponds to the lower bound of the welfare. Thus, the optimal assignment rule is characterized by the cutoff value $b$ that maximizes the difference between the social benefit lower bound and the social cost, which is $b=0.96$. This is equivalent to providing microfinance to those whose value of livestock is below 9,200 Birr.

Next, we consider a non\hyputilitarian welfare that exhibits loss aversion. Suppose that the policymaker weighs the loss 10% more than the benefits in individual $Y_i$, that is, for a function \[ h(x)=

casesx&x\geq 0,\\1.1 x&x<0,

\] define the second social welfare by \[ \mathbb{E}[h(Y_{i1}-Y_{i0})\mathbbm{1}\{U_i<b\}]. \] Here, we assume that the social costs are already built in to the loss\hypaversive welfare function of the policymaker.

figure[figure omitted — 557 chars of source]

(ref) shows the social welfare bounds as functions of $b$. The solid blue line is the lower bound (the worst\hypcase welfare) and the dashed blue line is the upper bound, as characterized by (ref). The shaded blue area is the pointwise 95% confidence interval for the social welfare. From $b=0.357$ to $b=0.976$, the worst\hypcase welfare is pointwise significantly positive.

The optimal assignment rule that maximizes the worst\hypcase welfare is found to be $b=0.928$, which is equivalent to treating everyone below 7,299 Birr in the total value of livestock owned.

Application 2: SPW for Welfare Reform

This section presents an empirical application showing how our bounds can be used to provide additional information in the context of testing economic theory. We estimate the bounds on the proportion of workers who earns more under the new welfare program than the counterfactual under the old program. The motivation and data are borrowed from bgh2006. \footnote{The replication files are accessible at the journal's website bgh2006dta.}

The summary of the background is as follows. Aid to Families with Dependent Children (AFDC) was a federal assistance program started in 1935 that required all states to provide financial support to children whose families had low or no income. During the 1990s, the federal government waived portions of the requirements and allowed states to make changes to various aspects of the program, such as expanding income disregards, increasing work requirements, and introducing time limits on benefits. In exchange for the waiver, states were required to conduct rigorous evaluations of the impacts of these changes. Connecticut introduced the Jobs First program as the AFDC waiver, and for the rigorous evaluation, conducted the random assignment study of Jobs First. The experiment took place between January 1996 and February 1997. We refer the reader to bgh2006 for more details.

bgh2006 stated that

quote. . . theory makes heterogeneous predictions concerning the sign and magnitude of the response of labor supply and welfare use to these reforms. . . . the vast majority of welfare reform studies rely on estimating mean impacts. Theory predicts that these mean impacts will average together positive and negative labor supply responses, possibly obscuring the extent of welfare reform's effects.

bgh2006 further used the QTE to demonstrate that the effects were heterogeneous across different levels of the potential outcome variables. In this section, we complement their analysis by identifying maximum levels of winners and losers given by our bounds.

figure[figure omitted — 559 chars of source]

(ref) illustrates the stylized budget constraints faced by women supported by AFDC and Jobs First. The horizontal axis is the time for leisure, which is a complement of working hours, and the vertical axis is the income. Assuming the standard utility theory, each program participant enjoys the income and leisure on the budget constraint. Under AFDC, the budget constraint is given by the black solid line in each figure of (ref). If an individual picks a point on segment A, it means that she receives benefits from AFDC. Jobs First replaced the AFDC benefits with the green solid line, meaningfully raising the budget constraint for individuals on A, B, and C. The Jobs First program also came with a time limit; after 21 months, the budget constraint is pushed down to just the original black line with no benefits. Before the time limit ((ref)), individuals on segments A, B, and C would move to somewhere on the green line under Jobs First; individuals on segment D may or may not move to a point on the green line. After the time limit ((ref)), individuals would receive zero cash transfers and become weakly worse off than AFDC. \footnote{They may still have received food stamps.}

While (ref) provides a helpful picture, the actual implementation was different from the stylized description. The most notable difference was an extension of the deadline. In a fair amount of cases, either an extension of 6 months or an indefinite exemption from the time limit was granted. Despite this, bgh2006 provided evidence that the time limit still mattered. However, it is important to keep in mind that there were beneficiaries that continued to receive benefits after the time limit.

How Did the New Program Affect Hours Worked?

We are interested in whether the new program increases or decreases earnings, which are roughly proportional to the hours worked.

figure[figure omitted — 879 chars of source]

There are three subpopulations of interest. We assume that each individual faces the same wage under the new and old programs, one of which is counterfactual. We do {\em not} assume that the wage is common across individuals. \footnote{However, we can only measure hours worked by earnings, so SPWs are conditional on earnings.} Under the new program, the following are the predictions of the stylized theory.

enumerate[i.] • The subpopulation on segment A in (ref) might work more or work less, depending on the magnitudes of the substitution effect (increases hours worked) and the income effect (decreases hours worked). This is shown in (ref). • The subpopulation on segment D in (ref) might work less or work as much. • The subpopulation on segment A in (ref) would work more if leisure is a normal good. This is shown in (ref).
figure[figure omitted — 792 chars of source]

First, we examine the sign of the program effects on subpopulation ((ref)). For this, we plot the estimated bounds on the SPW on earnings in the 4th post\hyptreatment quarter (Q4) for the subgroup defined by $\{U_i<b\}$ as a function of $b$ in (ref). The black line on the bottom that is zero until $b=0.47$ and takes a nonzero value thereafter is the estimated lower bound of the SPW \[ P(Y_{i1}>Y_{i0}\mid U_i<b), \] where $Y_i$ is $i$'s earnings. The black line that is almost identical to one is the upper bound. These are the median\hypbias\hypcorrected estimators given by clr2013. For this SPW, the upper bound is uninformative, while the lower bound identifies some winners who increased hours worked under the new program. Namely, the lower bound takes a maximum of $0.13$ at $b=0.55$. This means that, among the workers whose earnings are below the 55% quantile under the old program, at least 13% of them worked more under the new program. The gray area is the pointwise 95% confidence interval constructed by the method of clr2013.

These numbers are modest, but seem coherent with the QTE estimates of bgh2006. Considering that the theoretical prediction was ambiguous, the bounds successfully identified a small yet significant portion of individuals whom the new program encouraged to work more. We did not identify losers, who worked less because of the new program. The lack of evidence does not mean the lack of losers, but the data was not decisive enough to say anything about their existence or prevalence. In (ref), we will see that an additional assumption can sometimes tighten the bounds.

Second, we examine whether the possible negative effect on subpopulation ((ref)) can be identified in the data. (ref) plots the estimated bounds on the SPW for $\{U_i>a\}$ as a function of $a$. As before, the black line at the bottom is the lower bound, and the one at the top is the upper bound. The lower bound is zero between $a=0.055$ and $1$. The upper bound is slightly off $1$, but the confidence interval shows that it is not significantly away from $1$ at any value of $a$. In short, we do not see a statistically significant amount of losers on the right tail of the earnings distribution.

Third, after the time limit, the new program stopped providing benefits to some individuals, which was expected to encourage them to work more for subpopulation ((ref)). We examine this effect in (ref). It plots the bounds on the SPW on the earnings in the 12th post\hyptreatment quarter (Q12) for $\{U_i<b\}$ as a function of $b$. \footnote{By the 12th quarter, those who received a 6\hypmonth extension would have stopped receiving benefits.} The lower bound becomes positive after $b=0.42$ and takes the maximum of $0.097$ at $b=0.48$. Therefore, among those whose earnings are below the 48% quantile under the old program, at least 9.7% of them work more under the new program. The pointwise 95% confidence interval leaves zero at $b=0.44$. The fact that the bounds identify significant winners at lower values of $b$ compared to (ref) is consistent with the observation that if you receive low to zero earnings, the anticipated program effect on the hours worked would be large, due to both the positive substitution effect and the positive income effect ((ref)).

Using Economic Theory to Tighten the Bounds

figure[figure omitted — 644 chars of source]

Our bounds are nonparametrically tight, meaning that they characterize the maximum amounts of winners and losers without having to make additional assumptions. However, it is also true that the estimated bounds in (ref) were only modestly informative.

In this section, we investigate how we can tighten the bounds if we are open to making more assumptions. Recall that the theoretical prediction we examined came from a stylized utility theory. If we take the utility theory for granted, we see that those who do not work under the new program would not work under the old program if leisure is a normal good (see (ref)). \footnote{The leisure being a normal good is only needed for after the time limit. Also, the converse does not hold in that those who do not work under the old program may work nonzero hours as discussed for subpopulation ((ref)).}

This is in line with the data; in the 4th post\hyptreatment quarter, 47.5% do not work under the new program while 55.4% do not under the old program ((ref)). Thus, we may assume that those 47.5% in the treated go straight into the 55.4% in the control. The difference 7.9% are those who work under the new program but do not under the old. In light of this, we eliminate the 47.5% from the treated sample as well as the 47.5% off of the 55.4% from the control sample. This leaves us with a smaller sample, with which we compute the bounds as before. This is equivalent to imposing a block\hypdiagonal support restriction on the joint distribution of the potential outcomes ((ref)). \footnote{(ref) can be seen as a restriction on the support of the copula density.}

The new bounds estimated with the refined samples are presented in (ref). The horizontal axis now starts from $b=0.475$ since we have eliminated the 47.5% of non\hypworking individuals in both groups. Everyone represented in this figure works under the new program. Both the upper and lower bounds are one for the first 7.9%. Then, the lower bound decreases as we expand the subgroup. This sharp decline suggests that we hardly identify any more winners than the first 7.9%.

Meanwhile, the domain restriction identifies a small portion of significant losers on the upper tail ((ref)). The upper end of the pointwise 95% confidence interval is strictly below $1$ between $a=0$ and $a=0.90$. Within this range, the estimated upper bound takes the minimum of $0.879$ at $a=0.859$. This indicates that, among the subpopulation whose $Y_0$ is above \$4,500, at least 12.1% of them work fewer hours under the new program.

In the 12th quarter, 46.6% in the control did not work, and neither did 41.6% in the treated. Applying the same elimination procedure, the estimated bounds on the SPW for the subgroup $\{U_i<b\}$ are given in (ref). Like (ref), we do not identify more winners than the initial 5.0% who do not work under the old program.

figure[figure omitted — 755 chars of source]

Conclusion

Many empirical questions concern treatment effects at different levels of the outcome variables. We developed two bounds to investigate heterogeneous treatment effects of this type. These bounds can be computed with the estimated marginal distributions of the potential outcomes, and provide complementary information to the widely used QTE. They are sharp without requiring additional assumptions.

The first bounds are on the STE of the form $\mathbb{E}[Y_{i1}-Y_{i0}\mid Y_{i0}<c]$. In (ref), we applied them to assess the effect of microfinance on the poor, and found that, despite the ATE being insignificant, the STE for various values of $c$ was significantly positive. We also applied our bounds to the policy targeting problem. We considered two measures of welfare: one linear in the treatment effect with per\hypcapita fixed costs and the other a non\hyputilitarian welfare that embodied a loss\hypaversive preference. We illustrated how our bounds led to treatment assignment rules that maximized the worst\hypcase welfare for each welfare measure.

The second bounds are on the SPW of the form $P(Y_{i1}>Y_{i0}\mid Y_{i0}<c)$. In (ref), we investigated the heterogeneous effects of Connecticut's welfare reform on the earnings predicted by stylized labor supply theory. For a subgroup for whom the theoretical prediction was ambiguous, we detected a portion of individuals who increased hours worked because of the new welfare program. We also illustrated how economic theory could be used to tighten the bounds, and after tightening, we detected a small portion of individuals who decreased hours worked because of the new welfare program.

We conclude by noting that our results are applicable well beyond the RCT setups picked up in our empirical applications. The bounds are based solely on the marginal distributions of the potential outcomes, which most QTE estimation procedures estimate as byproducts. For example, there are existing QTE estimators for the RCT with imperfect compliance, difference\hypin\hypdifferences, selection\hypon\hypobservables, regression discontinuity designs, and instrumental variables regression models. Our bounds provide interpretable assessment of the heterogeneity of treatment effects that supplements the QTE. The Stata program for our bounds can be downloaded at an \href{https://github.com/jcao0/subgroup-treatment-effects}{author's website}.