EconBase
← Back to paper

Partial Identification and Inference for Conditional Distributions of Treatment Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,491 characters · 13 sections · 104 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Partial Identification and Inference for Conditional Distributions of Treatment Effects

abstractThis paper considers identification and inference for the distribution of treatment effects conditional on observable covariates. Since the conditional distribution of treatment effects is not point identified without strong assumptions, we obtain bounds on the conditional distribution of treatment effects by using the Makarov bounds. We also consider the case where the treatment is endogenous and propose two stochastic dominance assumptions to tighten the bounds. We develop a nonparametric framework to estimate the bounds and establish the asymptotic theory that is uniformly valid over the support of treatment effects. An empirical example illustrates the usefulness of the methods. \\ \\ Keywords: treatment effects, conditional distribution, heterogeneity, partial identification, uniform inference. \\ JEL Classification Numbers: C14, C21.

Introduction

This paper considers identification and estimation of the conditional distribution of treatment effects.\footnote{The distribution of treatment effects has received significant attention, and its importance becomes greater when the benefit from the treatment is non-transferrable (heckman2007econometric).} While the unconditional distribution of treatment effects has been considered extensively in the literature in both theoretical and applied econometrics (e.g., heckman1997making,fan2010sharp,fan2012confidence,fan2017partial,firpo2019partial,frandsen2021partial), we focus on the conditional distribution of treatment effects to take into account the heterogeneity caused by differences in the value of covariates.\footnote{For example, the effect of a mother's smoking behavior on her child's birth weight might differ across the mother's age (e.g., abrevaya2015estimating).} Such heterogeneous treatment effects, if they exist, may help policy makers develop more effective policies. We use the Makarov bounds, which depend on the marginal distributions of potential outcomes while the dependence structure between them is left unspecified, to partially identify the conditional distribution of treatment effects (e.g., makarov1982estimates, ruschendorf1982random, and williamson1990probabilistic). We then provide estimation and inference methods for the bounds.

We start with the case where the assumption of the conditional independence of the treatment and potential outcomes given covariates, which is called an unconfoundedness condition, is satisfied. However, there are many situations where such a conditional independence assumption fails to hold. We present identification results of the distribution of treatment effects when the treatment is endogenous without assuming the existence of an instrumental variable. We show that one can still obtain Makarov-type bounds on the distribution of treatment effects even with an endogenous treatment. When a treatment is endogenous, the bounds on the distribution of the treatment effect may be too wide to be informative.\footnote{This is also true for the case where we impose the unconfoundedness assumption. We can expect that, since we are interested in the distribution of treatment effects for some subgroup, the bounds on the conditional distribution may be narrower than those on the unconditional distribution. This is another motivation for considering the conditional distributions of treatment effects in this paper.} Motivated by the previous studies in the literature (e.g., manski1997monotone,manski2000monotone,blundell2007changes), we consider two kinds of stochastic dominance between potential outcomes in this paper to tighten the bounds. These stochastic dominance assumptions are useful in many empirical situations as they are consistent with many economic theories. The resulting bounds under the stochastic dominance assumptions are easy to compute.

We construct nonparametric kernel-type estimators of the bounds on the conditional distribution of treatment effects and establish asymptotic theory for them that is uniformly valid over the support of treatment effects for some fixed subgroup of the population. They are useful for comparing bounds between two subpopulations defined in terms of different values of observable characteristics.\footnote{One can perform statistical testing for global hypotheses, such as equality of or stochastic dominance between two bounds (e.g., barrett2003consistent,lee2009testing,seo2018tests), and this requires asymptotic theory uniformly valid over the support of treatment effects. } We adapt the results on uniform inference for value functions that were developed recently by firpo2021uniform and provide a set of conditions under which the estimated bounds are consistent for the true bounds. We propose a bootstrap procedure to mimic the asymptotic distributions of the estimated bounds and show its validity. The bootstrap scheme in this paper is based on the novel bootstrap procedure for Hadamard directionally differentiable functionals that was developed by fang2019inference.

The asymptotic theory developed in this paper relaxes some undesirable assumptions imposed in several existing studies in the literature. Specifically, the bounds on the distribution of treatment effects are defined as some functionals of marginal distributions of potential outcomes. These functionals involve the infimum and supremum over the supports of potential outcomes. The uniqueness of the infimum and supremum needs to be imposed to establish a pointwise asymptotic theory for those bounds, as in fan2010sharp,fan2012confidence.\footnote{fan2010sharp impose Assumptions 3 and 4 to guarantee the uniqueness of the maximizer and minimizer involved in the bounds on the distribution of treatment effects. fan2012confidence impose Assumptions A3 and A4 for the same purpose. } However, the uniqueness assumption may not hold for some data generating process (DGP), and the pointwise asymptotic results may not be valid in such cases (cf. milgrom2002envelope, firpo2021uniform). On the other hand, the inference results developed in this paper do not require such uniqueness assumptions.

One can use the novel approach developed by chernozhukov2013intersection that yields confidence bands uniformly valid over the joint support of an outcome variable and covariates.\footnote{chernozhukov2013intersection do not explicitly provide an inference result for nonparametric conditional distribution estimators that is uniformly valid over the joint support of an outcome variable and covariates. However, one can easily adapt their results in the literature to establish the strong Gaussian approximation of some nonparametric estimator of a conditional distribution function. This allows one to perform inference uniformly over the joint support of the outcome variable of interest and conditioning variables (e.g., chernozhukov2014gaussian).} Their inference methods for bounds defined by either supremum or infimum of some function are based on the strong Gaussian approximation using couplings. A key ingredient of their approach is the construction of argsup and/or arginf sets, and it is allowed to avoid imposing the uniqueness assumption on those sets. Our approach differs from theirs in that we rely on weak convergence of nonparametric estimators of the bounds to some fixed Gaussian process and the results on Hadamard directionally differentiable functionals. In addition, while our approach also requires to estimate argsup and arginf sets for the Hadamard directional derivatives, the construction of these sets is computationally easy in comparison to that of chernozhukov2013intersection. Therefore, the inference theory in this paper can be considered complementing the existing inference methods for the Makarov bounds.

The Monte Carlo simulation results show that the inference methods in this paper perform well in finite samples. We also provide an empirical example to illustrate the practical relevance of the methods proposed in this paper. We revisit the empirical question of the effect of 401(k) plans on net asset accumulations investigated by multiple papers in the literature (e.g., Abadie2003,Chernozhukov2004,Wuethrich2019,SantAnna2022). We confirm that the identifying power of the stochastic dominance assumptions is substantial from the empirical application. Furthermore, we find evidence on heterogeneity in the treatment effect that is consistent with the results in the literature.

This paper contributes to several strands of the literature on treatment effects. First, this paper mainly contributes to the literature on identification and estimation of the distribution of treatment effects by providing nonparametric estimation and inference methods for the conditional distribution of treatment effects. The literature is too vast to list all the papers relevant to this point. Identification of the unconditional distribution of treatment effects has been studied by, for example, heckman1997making, fan2010sharp,fan2012confidence, fan2017partial, vuong2017counterfactual, kim2018identifying, and firpo2019partial.

This paper is closely related to kim2018partial and frandsen2021partial. kim2018partial provides identification results under several distributional restrictions, together with an instrumental variable, that help tighten the bounds on the distribution of treatment effects, when the treatment is endogenous. The restrictions considered by kim2018partial are general and closely related to the stochastic dominance assumptions in this paper that were proposed independently of the results in kim2018partial. This paper differs from kim2018partial in that this paper focuses on estimation and inference for the conditional distribution of treatment effects with a motivation for treatment effect heterogeneity, whereas kim2018partial mainly considers identification of the unconditional distribution of treatment effects under some restrictions on the model. Moreover, it is much easier to compute and estimate the bounds proposed in this paper than those in kim2018partial. Therefore, the bounds in this paper have great applicability to empirical research.

In comparison with frandsen2021partial, we focus on estimation and inference for the bounds on the conditional distribution of treatment effects with a motivation for heterogeneity across subgroups. On the other hand, frandsen2021partial are mainly concerned with identification of the unconditional distribution of treatment effects under a novel restriction that is called the “stochastic increasingness” assumption. frandsen2021partial also discuss how to incorporate covariates and estimate the bounds on the unconditional distribution of treatment effects. However, they do not develop the asymptotic theory for the conditional distribution of treatment effects, especially when the bounds on the conditional distribution are the Makarov bounds.

This paper also contributes to the literature on heterogeneity in treatment effects across subpopulations by considering the conditional distribution of treatment effects (e.g., donald2012incorporating,abrevaya2015estimating,chang2015nonparametric,hsu2017consistent,shen2019estimation). Most existing studies focus on average/quantile treatment effects or the marginal distributions of potential outcomes. The results of this paper complement them by considering conditional distributions of treatment effects.

The rest of this paper is organized as follows. Section (ref) presents the model and identification results. Section (ref) provides nonparametric estimators of the Makarov bounds on the conditional distribution of treatment effects and the bootstrap procedure. Section (ref) develops the asymptotic theory for the estimated bounds on the conditional distribution of treatment effects. Section (ref) presents the Monte Carlo simulation study, and Section (ref) provides the empirical example considering the effect of 401(k) plan on net financial asset accumulations. We then conclude with Section (ref). All mathematical proofs, technical expressions, and additional results are presented in Appendix.

Before proceeding, we introduce some notation that will be used throughout this paper. Two random variables $A$ and $B$ that are independent of each other are denoted by $A\perp B$, and $\mathbb{E}[\cdot]$ is the expectation operator. For a matrix $A$, $A^{t}$ is the transpose of $A$. For a set $A$, $l^{\infty}(A)$ denotes the set of uniformly bounded functions on $A$. For a sequence of random variables $(Z_{n})$ and a random variable $Z$, we denote the weak convergence of $Z_{n}$ to $Z$ by $Z_{n}\Rightarrow Z$.\footnote{A formal definition of weak convergence can be found in, for example, kosorok2008.} For a set $A$, $int(A)$ is the interior of $A$.

Model and Identification

Identification Under the Unconfoundedness Condition

Let $(\Omega,\mathcal{S},\mathbf{P})$ be a probability space. Let $D$ be a binary variable that indicates whether a person gets the treatment or not, in other words, $D=1$ if the person gets the treatment, and $D=0$ if the person does not get the treatment. For each $d\in\{0,1\}$, let $Y_{d}$ denote the potential outcome when $D=d$, and the observed outcome is defined as $Y\equiv DY_{1}+(1-D)Y_{0}$. We assume that $Y_{d}$ is a continuous random variable for each $d\in\{0,1\}$. Let $X\in\mathbb{R}^{d_{x}}$ be a set of covariates and denote its support by $\mathcal{X}$. We can only observe $(Y,D,X^{t})^{t}$ from the data. We denote the conditional distribution of $Y_{d}$ on $X=x$ by $F_{d|X}(\cdot|x)$ for each $d\in\{0,1\}$, and let $p_{0}(x)\equiv\Pr(D=1|X=x)$ for a given $x\in\mathcal{X}$. $F_{X}(\cdot)$ denotes the distribution function of $X$.

We begin with the models under the conditional independence of the treatment variable and impose the following assumptions.

assumption$(Y_{1},Y_{0})\perp D|X$.
assumptionThere exist $\text{\ensuremath{\underline{p},\bar{p}\in(0,1)} such that \ensuremath{p_{0}(x)\in[\underline{p},\bar{p}]} uniformly in \ensuremath{x\in\mathcal{X}}}$.

Assumption (ref) is the unconfoundedness assumption, which means that the treatment $D$ is independent of the potential outcomes conditional on $X$. Assumption (ref) is an overlap condition that is commonly imposed in the relevant literature (e.g., imbens2004nonparametric). This assumption implies that we can observe individuals with $Y=Y_{1}$ and those with $Y=Y_{0}$ for any value of $x\in\mathcal{X}$. Under these assumptions, we have the following identification result:

lemmaLet $x\in\mathcal{X}$ be given. Suppose that Assumptions (ref) and (ref) are satisfied. For any measurable function $G$ such that $\mathbb{E}[|G(Y_{1})||X=x]<\infty$ and $\mathbb{E}[|G(Y_{0})||X=x]<\infty$, we have \begin{equation} \begin{aligned}\mathbb{E}[G(Y_{1})|X=x] & =\frac{\mathbb{E}[D\cdot G(Y)|X=x]}{\mathbb{E}[D|X=x]},\\ \mathbb{E}[G(Y_{0})|X=x] & =\frac{\mathbb{E}[(1-D)\cdot G(Y)|X=x]}{\mathbb{E}[(1-D)|X=x]}. \end{aligned} \end{equation}

Lemma (ref) implies that we can identify the conditional distributions of $Y_{1}$ and $Y_{0}$ on $X=x$ when we choose $G(Y)\equiv\mathbf{1}(Y\leq y)$ for a given $y\in\mathbb{R}$. This result is also closely related to the identification of the unconditional distributions of $Y_{1}$ and $Y_{0}$ that is considered by donald2014estimation.

The treatment effect $\Delta$ is defined as the difference between $Y_{1}$ and $Y_{0}$ (i.e., $\Delta\equiv Y_{1}-Y_{0}$). The parameter of interest is the conditional distribution of treatment effects given $X=x$ for a given value $x\in\mathcal{X}$, and this conditional distribution is denoted by $F_{\Delta|X}(\cdot|x)$.

The (unconditional) distribution of treatment effects has received a considerable amount of attention from the literature (e.g., heckman1997making, fan2010sharp,fan2012confidence, fan2017partial, firpo2019partial). The distribution function is useful in the context of program evaluation, for example, to determine the proportion of people who benefit from being treated, namely, $\Pr(\Delta\geq0)=1-F_{\Delta}(0)$. One can refer to abbring2007econometric for more discussion on the topic. The conditional distributions takes potential heterogeneity across subpopulations into account and thus can provide much richer information on treatment effects. Such heterogeneity, if it exists, may deliver different policy implications for different subpopulations.

It is worth noting that even in a randomized experiment where one can point identify the conditional distributions of $Y_{1}$ and $Y_{0}$ given $X$, the distribution of treatment effects is not point identified without additional structures of the model.\footnote{When the conditional distributions of $Y_{1}$ and $Y_{0}$ given $X$ are point identified, a sufficient condition for point identification of the joint distribution of $Y_{1}$ and $Y_{0}$ is the (conditional) rank invariance; that is, $F_{1|X}(Y_{1}|X)=F_{0|X}(Y_{0}|X)$ almost surely. } In particular, fan2010sharp provide sharp bounds on the distribution of treatment effects in randomized experiments that are based on makarov1982estimates, and firpo2019partial improve the bounds of fan2010sharp in the uniform sense. We do not focus on the uniform sharpness of the bounds but on pointwise sharp bounds of the conditional distribution of treatment effects and nonparametric estimation of the bounds. We recall the bounds on the conditional distribution of treatment effects $F_{\Delta|X}(\cdot|x)$:

proposition(Lemma 2.1 in fan2010sharp) Let $x\in\mathcal{X}$ be fixed. Then, for a given $\delta\in\mathbb{R}$, define \begin{eqnarray} F_{\Delta|X}^{L}(\delta|x) & = & \max\left(\sup_{y\in\mathbb{R}}\left\{ F_{1|X}(y|x)-F_{0|X}(y-\delta|x)\right\} ,0\right),\\ F_{\Delta|X}^{U}(\delta|x) & = & \min\left(\inf_{y\in\mathbb{R}}\left\{ F_{1|X}(y|x)-F_{0|X}(y-\delta|x)\right\} ,0\right)+1. \end{eqnarray} If Assumptions (ref) and (ref) are satisfied, we then have \begin{equation} F_{\Delta|X}(\delta|x)\in\left[F_{\Delta|X}^{L}(\delta|x),F_{\Delta|X}^{U}(\delta|x)\right]. \end{equation}

Note that this result is essentially identical to the identification result in fan2010sharp. Without additional structures on the model, these bounds are sharp in the pointwise sense for given $x\in\mathcal{X}$. It is also worth noting that one can derive bounds on the unconditional distribution of treatment effects by averaging the bounds on the conditional distribution. Specifically, if the potential outcomes are correlated with $X$, then the bounds on the unconditional distribution obtained from the conditional distributions of the potential outcomes are tighter than the those obtained from the unconditional distributions of the potential outcomes (cf. firpo2019partial).

Identification with an Endogenous Treatment

In this section, we discuss identification of the distribution of treatment effects with an endogenous binary treatment. When the treatment is endogenously determined, the conditional distributions of $Y_{1}$ and $Y_{0}$ given $X$ are generally partially identified without additional assumptions. The following theorem shows that even if the conditional distributions of $Y_{1}$ and $Y_{0}$ given $X$ are partially identified, we can still bound $F_{\Delta|X}(\delta)$.

theoremLet $x\in\mathcal{X}$ be fixed. Suppose that, for all $y\in\mathbb{R}$, the identified sets of $F_{1|X}(y|x)$ and $F_{0|X}(y|x)$ are given by $\left[F_{1|X}^{L}(y|x),F_{1|X}^{U}(y|x)\right]$, and $\left[F_{0|X}^{L}(y|x),F_{0|X}^{U}(y|x)\right]$, respectively. For a given $\delta\in Supp(\Delta|X=x)$, define \begin{eqnarray} F_{\Delta|X}^{e,L}(\delta|x) & = & \max\left(\sup_{y}\left\{ F_{1|X}^{L}(y|x)-F_{0|X}^{U}(y-\delta|x)\right\} ,0\right),\\ F_{\Delta|X}^{e,U}(\delta|x) & = & \min\left(\inf_{y}\left\{ F_{1|X}^{U}(y|x)-F_{0|X}^{L}(y-\delta|x)\right\} ,0\right)+1. \end{eqnarray} Then, \[ F_{\Delta|X}^{e,L}(\delta|x)\leq F_{\Delta|X}(\delta|x)\leq F_{\Delta|X}^{e,U}(\delta|x). \]

One example of the identified sets of $F_{1|X}$ and $F_{0|X}$ is Manski's bounds (manski1990nonparametric,manski1994selection) defined as follows:

equation[equation omitted — 336 chars of source]

The bounds presented in Theorem (ref) based on Manski's bounds are often too wide to be informative. Some model restrictions have been imposed to tighten bounds on the parameter of interest in the literature (e.g., manski1997monotone,manski2000monotone,blundell2007changes). Motivated by these studies, we propose two stochastic dominance assumptions that may help tighten the bounds and be consistent with many economic theories in the absence of instrumental variables. The bounds on the distribution of treatment effects in Theorem (ref) are based on the bounds on the conditional distributions of the potential outcomes, and therefore, we can tighten the bounds on $F_{\Delta|X}$ once we have tighter bounds on the conditional distributions of the potential outcomes than ((ref)).

Let $F_{d|d^{'}X}(y|x)\equiv\Pr\left(Y_{d}\leq y|D=d^{'},X=x\right)$ for given $d,d^{'}\in\{0,1\}$. The following assumption states that the potential outcome conditional on $D=1$ first-order stochastically dominates the potential outcome conditional on $D=0$.

assumptionLet $x\in\mathcal{X}$ be given. For all $d\in\{0,1\}$ and $y\in\mathbb{R}$, $F_{d|1,X}(y|x)\leq F_{d|0,X}(y|x)$.

Assumption (ref) is identical to the stochastic dominance assumption of blundell2007changes that is used to tighten the bounds on the wage distribution of the whole population in the presence of sample selection. okumura2014concave generalize this stochastic dominance assumption to the case where $D$ is not binary. Since Assumption (ref) implies that $E[Y_{d}|D=1,X=x]\geq E[Y_{d}|D=0,X=x]$, this assumption is a sufficient condition for the monotone treatment selection (MTS) assumption in manski2000monotone.

Assumption (ref) is consistent with some economic theories in empirical studies. As mentioned earlier, blundell2007changes utilize this assumption in conjunction with positive selection of labor force participation from the standard labor supply model that wage and probability of labor force participation are positively correlated. As another example, we can consider sorting models for return to education, where potential employees signal their ability to employers by using their education level. Considering return to college education (i.e., $Y$ is wage, and $D$ is college entrance), it is likely that the more capable people are, the more likely it is for them to complete a college education. As a result, we may anticipate that people with college degrees are more likely to have higher learning ability than those who did not complete a college education (e.g., bedard2001human). Many studies in the labor economics literature consider learning ability as an important factor affecting wage (lang1986human), and thus, we can assume that people with college degrees have higher wages than those without degrees. As a consequence, the sorting hypothesis is consistent with Assumption (ref).

The following theorem gives identification results for the conditional distributions of the potential outcomes under Assumption (ref).

theoremLet $x\in\mathcal{X}$ be fixed. Suppose that Assumption (ref) holds. For a given $y\in\mathbb{R}$, define \begin{eqnarray*} F_{1|X}^{L,FSD1}(y|x) & \equiv & \Pr(Y\leq y|D=1,X=x),\\ F_{1|X}^{U,FSD1}(y|x) & \equiv & \Pr(Y\leq y|D=1,X=x)\Pr(D=1|X=x)+\Pr(D=0|X=x),\\ F_{0|X}^{L,FSD1}(y|x) & \equiv & \Pr(Y\leq y|D=0,X=x)\Pr(D=0|X=x),\\ F_{0|X}^{U,FSD1}(y|x) & \equiv & \Pr(Y\leq y|D=0,X=x). \end{eqnarray*} Then, \begin{eqnarray} F_{1|X}(y|x) & \in & \left[F_{1|X}^{L,FSD1}(y|x),F_{1|X}^{U,FSD1}(y|x)\right],\\ F_{0|X}(y|x) & \in & \left[F_{0|X}^{L,FSD1}(y|x),F_{0|X}^{U,FSD1}(y|x)\right]. \end{eqnarray}

Comparing the bounds on the conditional distributions of the potential outcomes in Theorem (ref) with those in equation ((ref)), we can find that the upper bound on $F_{1|X}(y|x)$ and lower bound on $F_{0|X}(y|x)$ in Theorem (ref) are identical to their counterparts in ((ref)). Since Assumption (ref) designates only one direction of the monotonicity of the distribution functions, it is impossible to improve the lower bound on $F_{0|X}(y|x)$ and the upper bound on $F_{1|X}(y|x)$. Nevertheless, the bounds on the conditional distributions of the potential outcomes provided in Theorem (ref) are narrower than those in ((ref)).

We now introduce another stochastic dominance assumption. The following stochastic dominance assumption may be regarded as an assumption corresponding to the monotone treatment response (MTR) assumption in manski1997monotone.

assumptionLet $x\in\mathcal{X}$ be given. For all $d\in\{0,1\}$ and $y\in\mathbb{R}$, $F_{1|d,X}(y|x)\leq F_{0|d,X}(y|x)$.

Note that Assumption (ref) implies that $\mathbb{E}[Y_{1}|X=x]\geq\mathbb{E}[Y_{0}|X=x]$. To see this, recall that $F_{1|X}(y|x)=F_{1|1,X}(y|x)p_{0}(x)+F_{1|0,X}(y|x)(1-p_{0}(x))$. Under Assumption (ref), we have $F_{1|1,X}(y|x)\leq F_{0|1,X}(y|x)$ and $F_{1|0,X}(y|x)\leq F_{0|0,X}(y|x)$, and therefore, we have $F_{1|X}(y|x)\leq F_{0|X}(y|x)$ and $\mathbb{E}[Y_{1}|X=x]\geq\mathbb{E}[Y_{0}|X=x]$.

Assumption (ref) can be regarded as a distributional generalization of, but is weaker than, the MTR assumption of manski1997monotone. This assumption is also applicable to many empirical studies as it is compatible with some economic theories. Recalling the example of return to college education, Assumption (ref) may be consistent with the human capital theory in which education increases productivity through a human capital production function and return to education reflects the increased productivity (e.g., mincer1974schooling). For example, Assumption (ref) can be translated into that people who have completed their college degrees are likely to be paid higher wages than when they have not, and this argument can be justified by the human capital theory in the labor economics.

It is worth noting that Assumption (ref) is a special case of conditional negative quadrant dependence, which is a dependence concept considered by lehmann1966some. kim2018partial proposes a concept of conditional quadrant dependence, as well as a negative stochastic dominance, in the presence instrumental variables. He provides tighter bounds on the distribution of treatment effects than bounds using ((ref)) under these assumptions with the existence of an instrumental variable. A difference between kim2018partial and this paper is that this paper does not impose additional assumptions on the data generating process, such as monotonicity in structural functions for $Y_{1}$ and/or $Y_{0}$ and the availability of instrumental variables. Furthermore, this paper focuses on estimation of the bounds on conditional distribution of treatment effects, whereas kim2018partial focuses on identification of the unconditional distribution of treatment effects.

The following theorem shows that we can tighten the bounds on the conditional distributions of $Y_{1}$ and $Y_{0}$ given $X$ under Assumption (ref):

theoremLet $x\in\mathcal{X}$ be fixed. Suppose that Assumption (ref) holds. For a given $y\in\mathbb{R}$, define \begin{eqnarray*} F_{1|X}^{L,FSD2}(y|x) & = & \Pr(Y\leq y|D=1,X=x)\Pr(D=1|X=x),\\ F_{1|X}^{U,FSD2}(y|x) & = & \Pr(Y\leq y|X=x),\\ F_{0|X}^{L,FSD2}(y|x) & = & F_{1|X}^{U,FSD2}(y|x),\\ F_{0|X}^{U,FSD2}(y|x) & = & \Pr(Y\leq y|D=0,X=x)\Pr(D=0|X=x)+\Pr(D=1|X=x). \end{eqnarray*} Then, \begin{eqnarray} F_{1|X}(y|x) & \in & \left[F_{1|X}^{L,FSD2}(y|x),F_{1|X}^{U,FSD2}(y|x)\right],\\ F_{0|X}(y|x) & \in & \left[F_{0|X}^{L,FSD2}(y|x),F_{0|X}^{U,FSD2}(y|x)\right]. \end{eqnarray}

As Assumption (ref) is not enough to improve $F_{1|X}^{U}(y|x)$ and $F_{0|X}^{L}(y|x)$ in ((ref)), we can see that the lower bound on $F_{1|X}(y|x)$ and the upper bound on $F_{0|X}(y|x)$ in Theorem (ref) remain the same as those in ((ref)).

If both Assumptions (ref) and (ref) hold, one can further improve the bounds on the conditional distributions of the potential outcomes. This result is summarized in the following corollary:

corollaryLet $x\in\mathcal{X}$ be fixed. Suppose that Assumptions (ref) and (ref) hold. For a given $y\in\mathbb{R}$, \begin{eqnarray*} F_{1|X}(y|x) & \in & \left[F_{1|X}^{L,FSD1}(y|x),F_{1|X}^{U,FSD2}(y|x)\right],\\ F_{0|X}(y|x) & \in & \left[F_{0|X}^{L,FSD2}(y|x),F_{0|X}^{U,FSD1}(y|x)\right]. \end{eqnarray*}

It is straightforward to see that the identified sets for conditional distribution functions of $Y_{1}$ and $Y_{0}$ on $X=x$ presented in Corollary (ref) are connected and that the intersection of these sets is the boundary of each set, that is $F_{1|X}^{U,FSD2}(y|x)=F_{0|X}^{L,FSD2}(y|x)$ for all $y\in\mathbb{R}$.

It is worth mentioning that the stochastic dominance relationships in Assumptions (ref) and (ref) can be reversed, depending on the empirical context. For example, one can impose a condition that for all $d\in\{0,1\}$ and $y\in\mathbb{R}$, $F_{d|1,X}(y|x)\geq F_{d|0,X}(y|x)$ instead of Assumption (ref). Similarly, one can consider a stochastic dominance relationship that for all $d\in\{0,1\}$ and $y\in\mathbb{R}$, $F_{1|d,X}(y|x)\geq F_{0|d,X}(y|x)$ instead of Assumption (ref). These stochastic dominance relationships can be motivated by, for example, the empirical study of Angrist1990, who considers the long-term effect of veteran status on earnings. As pointed out by Angrist1990, it would be likely that men with fewer civilian opportunities served in the armed force, and this may yield a negative selection in the sense that the potential earnings of veterans are likely to be smaller than those of non-veterans (i.e., $F_{d|1,X}(y|x)\geq F_{d|0,X}(y|x)$). In addition, since military service may hinder human capital accumulation, the potential earnings they would have received when they served in the armed force are likely to be smaller than the potential earnings they would have received when they did not (i.e., $F_{1|d,X}(y|x)\geq F_{0|d,X}(y|x)$).

There are other ways to tighten the bounds on the distribution of treatment effects. For example, one can utilize support restrictions (e.g., kim2018identifying), consider a weaker condition than the conditional independence, such as $c$-independence proposed by masten2018identification, or impose some dependence structure as in frandsen2021partial. In addition, although we do not assume the availability of (monotone) instrumental variables in this paper, one can consider the monotone instrumental variable (MIV) assumption when there exists such a variable, as in manski2000monotone and kim2018partial. A related but stronger assumption on instrumental variables is the stochastically MIV (SMIV) assumption considered by mourifie2020sharp. The SMIV assumption imposes a stochastic dominance ordering on the joint distribution of $Y_{0}$ and $Y_{1}$ across values of an instrumental variable, and the conditional joint distribution of the potential outcomes given a value of the instrumental variable, say $z$, (first-order) stochastically dominates the conditional joint distribution given a value of the instrumental variable that is smaller than $z$. These restrictions help tighten the bounds on the conditional distributions of the potential outcomes and/or the bounds on $F_{\Delta|X}$, but estimation and (uniform) inference for such bounds is not trivial. Therefore, we leave it for future study.

Estimation of Bounds and Confidence Bands

Estimation and Bootstrap Procedure

We consider estimation of the bounds and construction of confidence bands for them in this section. To this end, we assume that the conditional distributions of the potential outcomes are identified as follows: for all $y\in\mathbb{R}$ and $x\in\mathcal{X}$,

equation[equation omitted — 187 chars of source]

We denote nonparametric estimators of $LB_{1|X}$, $UB_{1|X}$, $LB_{0|X}$, and $UB_{0|X}$ by $\hat{LB}_{1|X,n}$, $\hat{UB}_{1|X,n}$, $\hat{LB}_{0|X,n}$, and $\hat{UB}_{0|X,n}$, respectively. In this paper, we consider kernel-type estimators of the bounds. Let

align*[align* omitted — 377 chars of source]

Suppose that there exists a positive real sequence $\left(r_{n}\right)_{n}$ such that $r_{n}\rightarrow\infty$ and that \[ r_{n}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)-\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\Rightarrow\mathbb{G}\left(\cdot\right)\text{ in }\left(l^{\infty}(\mathcal{Y})\right)^{4}, \] where $\mathbb{G}(\cdot)$ is a four-dimensional centered Gaussian process. This weak convergence result can be derived under a set of mild regularity conditions when using standard nonparametric estimators. When we use kernel-type estimators, we have $r_{n}=\sqrt{nh_{n}^{d_{x}}}$, where $h_{n}$ is a bandwidth.

Define

align*[align* omitted — 221 chars of source]

and

align[align omitted — 608 chars of source]

where $F_{\Delta|X}^{L,0}$ and $F_{\Delta|X}^{U,0}$ are some fixed bounds that are true under the hypotheses $F_{\Delta|X}^{L}(\cdot|x)=F_{\Delta|X}^{L,0}(\cdot|x)$ and $F_{\Delta|X}^{U}(\cdot|x)=F_{\Delta|X}^{U,0}(\cdot|x)$. The functionals in ((ref)) and ((ref)) can be considered as the Kolmogorov-Smirnov (KS) type statistics, and they are useful to construct two-sided uniform confidence bands for the lower and upper bounds on the conditional distribution of treatment effects.\footnote{One can consider one-sided test statistics to construct one-sided uniform confidence bands. However, we do not examine such test statistics in this paper as the main purpose is to obtain uniform confidence sets for the identified region and the two-sided uniform confidence bands are enough to provide confidence sets for the identified region. }

Let $\mathbb{H}\equiv(h_{1},h_{2},h_{3},h_{4})^{t}$ be a four-dimensional vector-valued function and $f$ be a real-valued function. We denote the Hadamard directional derivatives of $\phi_{L}$ and $\phi_{U}$ at $f$ in direction $\mathbb{H}$ by $\phi_{L}^{'}(f;\mathbb{H})$ and $\phi_{U}^{'}(f;\mathbb{H})$, respectively. We also let $\widehat{\phi_{L}^{'}}\left(f;\mathbb{H},a_{n}\right)$ and $\widehat{\phi_{U}^{'}}\left(f;\mathbb{H},a_{n}\right)$ denote estimators of $\phi_{L}^{'}(f;\mathbb{H})$ and $\phi_{U}^{'}(f;\mathbb{H})$, respectively, where $(a_{n})$ is a positive real sequence decreasing to zero. We provide the forms of the Hadamard directional derivatives and their estimators in Appendix (ref). The main goal is to establish the limiting distributions of $r_{n}\left(\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}\right)-\phi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}\right)\right)$ and $r_{n}\left(\phi_{U}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}\right)-\phi_{U}\left(\mathbb{F}_{\mathbf{Y}|X}\right)\right)$. The limiting distributions allow one to construct confidence bands for the bounds on the conditional distribution of treatment effects.

As will be shown below, the limiting distributions are nonstandard. We propose to use a multiplier bootstrap to mimic the asymptotic distributions of $r_{n}\left(\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}\right)-\phi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}\right)\right)$ and $r_{n}\left(\phi_{U}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}\right)-\phi_{U}\left(\mathbb{F}_{\mathbf{Y}|X}\right)\right)$. Specifically, let \[ \hat{\Psi}_{i}((y_{1l},y_{1u},y_{0l},y_{0u})|x)\equiv\left(\hat{\psi}_{1l,i|X,n}(y_{1l}|x),\hat{\psi}_{1u,i|X,n}(y_{1u}|x),\hat{\psi}_{0l,i|X,n}(y_{0l}|x),\hat{\psi}_{0u,i|X,n}(y_{0u}|x)\right)^{t} \] be an estimated influence function for the $i$-th observation of $\hat{\mathbb{F}}_{\mathbf{Y}|X,n}((y_{1l},y_{1u},y_{0l},y_{0u})|x)$ and $\{B_{i}:i=1,2,...n\}$ be $n$ random draws from $B$ that is independent of the data and satisfies some moment conditions (cf. Assumption (ref)). Define

align*[align* omitted — 296 chars of source]

Then, one can implement the following bootstrap procedure to obtain confidence bands for $F_{\Delta|X}^{L}(\cdot|x)$ and $F_{\Delta|X}^{U}(\cdot|x)$.

\paragraph{Bootstrap procedure }

enumerate• Estimate the conditional distribution functions of the potential outcomes or bounds on them, and construct \begin{align*} \Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x) & =\hat{LB}_{1|X,n}(y|x)-\hat{UB}_{0|X,n}(y-\delta|x),\\ \Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x) & =\hat{UB}_{1|X,n}(y|x)-\hat{LB}_{0|X,n}(y-\delta|x). \end{align*} • Repeat the following procedure for $m$ times, where $m$ denotes the number of bootstrap iterations. \begin{enumerate} • Generate the bootstrap weights $\{B_{i}:i=1,2,...n\}$ from the random variable $B$ in Assumption (ref). • Using the estimated conditional distribution functions of the potential outcomes or bounds on them, compute $r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x)$. • For each iteration $b=1,2,...,m$, compute \begin{align*} & \widehat{\phi_{L}^{'}}^{(b)}\left(\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x);r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right),\\ & \widehat{\phi_{U}^{'}}^{(b)}\left(\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x)+1;r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right). \end{align*} \end{enumerate} • For a given significance level $\alpha\in(0,1)$, find the $(1-\frac{\alpha}{2})$-th quantiles of \begin{align*} & \left\{ \widehat{\phi_{L}^{'}}^{(b)}\left(\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x);r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right):b=1,2,...,m\right\} ,\\ & \left\{ \widehat{\phi_{U}^{'}}^{(b)}\left(\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x)+1;r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right):b=1,2,...m\right\} . \end{align*} We denote the quantiles by $c_{1-\frac{\alpha}{2}}^{L}$ and $c_{1-\frac{\alpha}{2}}^{U}$, respectively. • Let \begin{align*} \tilde{CI}\left(F_{\Delta|X}^{L}(\delta;x),1-\alpha\right) & \equiv\left[\hat{F}_{\Delta|X,n}^{L}(\delta|x)-c_{1-\frac{\alpha}{2}}^{L},\hat{F}_{\Delta|X,n}^{L}(\delta|x)+c_{1-\frac{\alpha}{2}}^{L}\right],\\ \tilde{CI}\left(F_{\Delta|X}^{U}(\delta;x),1-\alpha\right) & \equiv\left[\hat{F}_{\Delta|X,n}^{U}(\delta|x)-c_{1-\frac{\alpha}{2}}^{U},\hat{F}_{\Delta|X,n}^{U}(\delta|x)+c_{1-\frac{\alpha}{2}}^{U}\right]. \end{align*} The $(1-\alpha)\times100\%$ confidence bands of $F_{\Delta|X}^{L}(\cdot|x)$ and $F_{\Delta|X}^{U}(\cdot|x)$ can be constructed as \begin{align*} CI\left(F_{\Delta|X}^{L}(\delta;x),1-\alpha\right) & \equiv\left[\max\left\{ \hat{F}_{\Delta|X,n}^{L}(\delta|x)-c_{1-\frac{\alpha}{2}}^{L},0\right\} ,\min\left\{ \hat{F}_{\Delta|X,n}^{L}(\delta|x)+c_{1-\frac{\alpha}{2}}^{L},1\right\} \right],\\ CI\left(F_{\Delta|X}^{U}(\delta;x),1-\alpha\right) & \equiv\left[\max\left\{ \hat{F}_{\Delta|X,n}^{U}(\delta|x)-c_{1-\frac{\alpha}{2}}^{U},0\right\} ,\min\left\{ \hat{F}_{\Delta|X,n}^{U}(\delta|x)+c_{1-\frac{\alpha}{2}}^{U},1\right\} \right], \end{align*} respectively.

The difference between $\tilde{CI}$ and $CI$ in the above bootstrap procedure is that $CI$'s are confidence bands obtained to impose the logical bounds when necessary. chen2021shape show that one advantage of applying such operators to the original confidence bands is that the resulting confidence bands have greater coverage than the original confidence bands. In addition, this procedure is very simple and easy to implement in practice.

The above bootstrap procedure allows us to construct the pointwise confidence set of the identified set, where the identified set is $[F_{\Delta|X}^{L}(\delta|x),F_{\Delta|X}^{U}(\delta|x)]$ for given $\delta\in Supp(\Delta|X=x)$. Specifically, we can show that \[ \underset{n\rightarrow\infty}{\lim\inf}\Pr\left(F_{\Delta|X}(\delta|x)\in\left[\hat{F}_{\Delta|X,n}^{L}(\delta|x)-c_{1-\frac{\alpha}{2}}^{L},\hat{F}_{\Delta|X,n}^{U}(\delta|x)+c_{1-\frac{\alpha}{2}}^{U}\right]\right)\geq1-\alpha, \] and this can be used as a $(1-\alpha)\times100\%$ confidence set for the identified set. This confidence set does not require the uniqueness of the infimum and supremum in the Makarov bounds. As pointed out by firpo2021uniform, however, this confidence set is likely to be conservative.

Nonparametric Estimators of the Bounds on Conditional Distributions of Potential Outcomes

We propose to use kernel-type estimators of bounds on the conditional distributions of the potential outcomes in ((ref)). Let $K(\cdot):\mathbb{R}^{d_{x}}\rightarrow\mathbb{R}$ be a $d_{x}$-dimensional kernel function and $h_{n}$ be a bandwidth such that $h_{n}\rightarrow0$, $nh_{n}^{d_{x}}\rightarrow\infty$, and $nh_{n}^{d_{x}+4}\rightarrow0$ as $n\rightarrow\infty$. We start with the case where Assumptions (ref) and (ref) hold. In this case, the conditional distributions of the potential outcomes are point identified (i.e., $LB_{j|X}(y|x)=UB_{j|X}(y|x)=F_{j|X}(y|x)$ for each $j\in\{0,1\}$), and we can use

align*[align* omitted — 328 chars of source]

as estimators of $F_{1|X}(y_{1}|x)$ and $F_{0|X}(y_{0}|x)$, respectively. We then define

equation[equation omitted — 239 chars of source]

Note that the kernel function and bandwidth used to estimate $F_{1|X}$ can be different from those used to estimate $F_{0|X}$. In addition, we have $\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X,n})(y,\delta|x)=\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X,n})(y,\delta|x)=F_{1|X}(y|x)-F_{0|X}(y-\delta|x)$ and can use the bootstrap procedure with $r_{n}=\sqrt{nh_{n}^{d_{x}}}$ under a set of regularity conditions.

We now consider the case where the treatment is endogenous and Assumptions (ref) and (ref) hold. Recall that $F_{d|d^{'}X}(y|x)=\Pr(Y_{d}\leq y|D=d^{'},X=x)$ for given $d,d^{'}\in\{0,1\}$, and the conditional distribution of $Y$ given $X=x$ is denoted by $F_{Y|X}(\cdot|x)$. By corollary (ref), we can construct kernel estimators of these objects as

align*[align* omitted — 468 chars of source]

and define

equation[equation omitted — 233 chars of source]

Then, we have

align*[align* omitted — 233 chars of source]

We provide a description on how to estimate the influence functions in Appendix (ref).

Asymptotic Theory

Inference under the Unconfoundedness Assumption

We first develop the asymptotic theory for the kernel estimators in ((ref)).

assumption$\{W_{i}\equiv(Y_{i},D_{i},X_{i}^{t})^{t}:i=1,2,...,n\}$ is a random sample.
assumption(i) The support of $X$, $\mathcal{X}$, is a compact subset of $\mathbb{R}^{d_{x}}$; (ii) the distribution of $X$ admits its density $f_{X}(\cdot)$ on $\mathcal{X}$ such that $0<\inf_{x\in\mathcal{X}}f_{X}(x)<\sup_{x\in\mathcal{X}}f_{X}(x)<\infty$. The density function $f_{X}(\cdot)$ is twice continuously differentiable and $\sup_{x\in\mathcal{X}}|f_{X}^{(1)}(x)|$ and $\sup_{x\in\mathcal{X}}|f_{X}^{(2)}(x)|$ are bounded.
assumption(i) The propensity score function $p_{0}(x)$ is twice continuously differentiable and all derivatives are uniformly bounded over $\mathcal{X}$; (ii) For each $j\in\{0,1\}$, $F_{j|X}(y|x)$ is twice continuously differentiable with respect to $x$ and all derivatives are uniformly bounded.
assumptionThe kernel function $K(\cdot)$ is a product of a univariate bounded kernel function $k(\cdot):\mathbb{R}\rightarrow\mathbb{R}_{+}$ such that $\int k(u)du=1$, $\int uk(u)du=0$, and $\int u^{2}k(u)du\equiv k_{2}<\infty$. The support of the univariate kernel function is compact.
assumptionThe bandwidth $h_{n}$ satisfies the following conditions: (i) $h_{n}\rightarrow0$; (ii) $nh_{n}^{d_{x}}\rightarrow\infty$; (iii) $nh_{n}^{d_{x}+4}\rightarrow0$ as $n\rightarrow\infty$.

Assumption (ref) is an i.i.d assumption on the sample. This can be relaxed at a cost of more complicated proofs of the theoretical results. Assumption (ref) imposes some degree of smoothness on the distribution of $X$. This assumption is standard in the literature on kernel estimation. Assumption (ref) requires that the conditional distribution functions of $Y_{1}$ and $Y_{0}$ and the propensity score functions be smooth enough. This condition, together with Assumptions (ref) and (ref), allows to eliminate the bias of estimators of conditional distribution functions of $Y_{1}$ and $Y_{0}$. Assumption (ref) is also standard in the literature, and there are several kernel functions that satisfy this assumption (e.g., Epanechnikov, biweight, triweight kernels). Assumption (ref) restricts the rate of bandwidth. The last condition $nh_{n}^{d_{x}+4}\rightarrow0$ is required to eliminate the asymptotic bias of the kernel estimators. Note that when the dimension of $X$ is large, one can use a higher-order kernel to handle the bias term of kernel estimators.

We first establish the weak convergence of the kernel estimators under the unconfoundedness assumption. For given $x\in\mathcal{X}$, define $\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\equiv\left(F_{1|X}(\cdot|x),F_{1|X}(\cdot|x),F_{0|X}(\cdot|x),F_{0|X}(\cdot|x)\right)^{t}$ and $\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\equiv\left(\hat{F}_{1|X,n}(\cdot|x),\hat{F}_{1|X,n}(\cdot|x),\hat{F}_{0|X,n}(\cdot|x),\hat{F}_{0|X,n}(\cdot|x)\right)^{t}$.

theoremSuppose that Assumptions (ref) and (ref) hold. Let $x\in int(\mathcal{X})$ be given. If Assumptions (ref)--(ref) hold, then, \[ \sqrt{nh_{n}^{d_{x}}}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)-\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\Rightarrow\mathbb{G}_{x}\left(\cdot\right)\equiv\left(\mathbb{G}_{1,x}(\cdot),\mathbb{G}_{1,x}(\cdot),\mathbb{G}_{0,x}(\cdot),\mathbb{G}_{0,x}(\cdot)\right)^{t}\text{ in }\left(l^{\infty}(\mathcal{Y})\right)^{4}, \] where $\mathbb{G}_{1,x}(\cdot)$ and $\mathbb{G}_{0,x}(\cdot)$ are mean zero Gaussian processes with covariance kernels $H_{1,x}(s,t)$ and $H_{0,x}(s,t)$, respectively, whose the forms are given in Appendix (ref), and $\mathbb{G}_{x}((y_{1l},y_{1u},y_{0l},y_{0u}))=\left(\mathbb{G}_{1,x}(y_{1l}),\mathbb{G}_{1,x}(y_{1u}),\mathbb{G}_{0,x}(y_{0l}),\mathbb{G}_{0,x}(y_{0u})\right)^{t}$.

Theorem (ref) is useful not only for deriving the asymptotic distribution of the estimated bounds on $F_{\Delta|X}(\cdot|x)$, but also for conducting uniform inference for some policy-relevant parameter (e.g., quantile treatment effects).

We now establish the asymptotic theory for confidence bands for the identified set of $F_{\Delta|X}(\cdot|x)$. To this end, we consider the following two hypotheses: $F_{\Delta|X}^{L}(\delta|x)=F_{\Delta|X}^{L,0}(\delta|x)$ and $F_{\Delta|X}^{U}(\delta|x)=F_{\Delta|X}^{U,0}(\delta|x)$, where $F_{\Delta|X}^{L,0}(\delta|x)$ and $F_{\Delta|X}^{U,0}(\delta|x)$ are some fixed lower and upper bounds on the conditional distribution of treatment effects, respectively.

The main challenge to constructing uniform confidence bands, however, is that the mappings $\phi_{L}$ and $\phi_{U}$ are not Hadamard differentiable. To overcome this difficulty, we use Theorem 3.2 in firpo2021uniform that shows that these functionals are Hadamard directionally differentiable. Based on this result, we employ the inference method of fang2019inference to establish the asymptotic theory.

Recall that $\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)=\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)=F_{1|X}(y|x)-F_{0|X}(y-\delta|x)$. The following theorem establishes the limiting distributions of $\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)$ and $\phi_{U}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)$.

theoremSuppose that Assumptions (ref) and (ref) hold. Let $x\in int(\mathcal{X})$ be given. If Assumptions (ref)--(ref) hold, then, \begin{equation} \begin{aligned}\sqrt{nh_{n}^{d_{x}}}\left(\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)-\phi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\right) & \Rightarrow\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}(\cdot)\right)\ in l^{\infty}\left(\mathcal{Y}^{4}\right),\\ \sqrt{nh_{n}^{d_{x}}}\left(\phi_{U}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)-\phi_{U}\left(\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\right) & \Rightarrow\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}(\cdot)\right)\ in l^{\infty}\left(\mathcal{Y}^{4}\right), \end{aligned} \end{equation} where and $\mathbb{G}_{x}(\cdot)$ is the Gaussian process defined in Theorem (ref). The forms of $\phi_{L}^{'}$ and $\phi_{U}^{'}$ are provided in Appendix (ref).

It is worth pointing out that the development of weak convergence of the estimated bounds does not rely on some pointwise asymptotic theory when establishing the asymptotic theory for the estimated Makarov bounds. Instead, we first derive the weak convergence of the standard kernel estimators for a fixed conditioning value and consider the double-supremum as an operator to establish the weak convergence of the estimated bounds, as in firpo2021uniform. In doing so, we can avoid imposing the uniqueness of $argsup$ and $arginf$ that is needed for pointwise inference.

Theorem (ref) is a direct consequence of Theorem 2.1 of fang2019inference. The limiting distribution presented in Theorem (ref) is non-standard, and therefore we need to rely on some resampling method to mimic the limiting distribution and conduct inference. Since the functionals $\phi_{L}$ and $\phi_{U}$ are not Hadamard differentiable but only Hadamard directionally differentiable, the standard bootstrap fails (cf. Theorem 3.1 of fang2019inference). To resolve this problem, we utilize the result of bootstrap validity that was proposed by firpo2021uniform. The next assumption imposes conditions on the bootstrap weight $B$:

assumptionLet $B$ be a random variable that is independent of the data $\mathcal{W}$ such that $\mathbb{E}[B]=0$, $Var(B)=1$, and $\int_{0}^{\infty}\sqrt{\Pr(|B|>x)}dx<\infty$.

The last condition in Assumption (ref) is satisfied if $\mathbb{E}\left[|B|^{2+\epsilon}\right]<\infty$ for some $\epsilon>0$. One can use a standard normal random variable as the bootstrap weight $B$.

We employ the approach of fang2019inference to approximate the limiting distribution presented in Theorem (ref), which was also considered by firpo2021uniform. Note that it is required to consistently estimate the Hadamard directional derivatives in ((ref)) and that the Hadamard directional derivatives are defined in terms of the limit operator. To this end, we consider a tuning sequence $(a_{n})$ that satisfies the conditions in the following assumption:

assumptionLet $(a_{n})$ be a sequence of positive real numbers such that $a_{n}\downarrow0$ and $a_{n}\sqrt{nh_{n}^{d_{x}}}\rightarrow\infty$.

Since the limiting distributions in Theorem (ref) are nonstandard, we use a bootstrap to mimic the limiting distributions. The next theorem demonstrates that one can use the bootstrap procedure in Section (ref) to approximate the asymptotic distributions of $\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}(\cdot)\right)$ and $\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}(\cdot)\right)$. Let $\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x)$ denote the vector of simulated stochastic processes using estimated influence functions in ((ref)) in Appendix (ref).

theoremSuppose that Assumptions (ref) and (ref) hold. Let $x\in int(\mathcal{X})$ be given and Assumptions (ref)--(ref), (ref), and (ref) hold. Then, we have \begin{align*} \widehat{\phi_{L}^{'}}\left(\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x);\sqrt{nh_{n}^{d_{x}}}\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right) & \Rightarrow\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}(\cdot)\right),\\ \widehat{\phi_{U}^{'}}\left(\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x)+1;\sqrt{nh_{n}^{d_{x}}}\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right) & \Rightarrow\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}(\cdot)\right), \end{align*} in $l^{\infty}\left(\mathcal{Y}^{4}\right)$, conditional on data. The forms of $\widehat{\phi_{L}^{'}}$ and $\widehat{\phi_{U}^{'}}$ are provided in Appendix (ref).

Inference with an Endogenous Treatment

We now develop the asymptotic theory when the treatment is endogenous. Our asymptotic theory for an endogenous treatment focuses on the situation where Assumptions (ref) and (ref) hold, and thus, one can use the kernel estimators in Section (ref). To be concrete, define, for given $x\in\mathcal{X}$, $\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\equiv\left(F_{1|1X}(\cdot|x),F_{Y|X}(\cdot|x),F_{Y|X}(\cdot|x),F_{0|0X}(\cdot|x)\right)^{t}$ and $\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\equiv\left(\hat{F}_{1|1X,n}(\cdot|x),\hat{F}_{Y|X,n}(\cdot|x),\hat{F}_{Y|X,n}(\cdot|x),\hat{F}_{0|0X,n}(\cdot|x)\right)^{t}$, where each component of $\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)$ is given in ((ref)). Let $\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X,n})(y,\delta|x)\equiv F_{1|1X}(y|x)-F_{0|0X}(y-\delta|x)$ and $\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X,n})(y,\delta|x)\equiv F_{Y|X}(y|x)-F_{Y|X}(y-\delta|x)$.

assumption$F_{1|1,X}(y|x)$, $F_{0|0,X}(y|x)$, and $F_{Y|X}(y|x)$ are twice continuously differentiable with respect to $x$ and all derivatives are uniformly bounded.
theoremSuppose that Assumptions (ref) and (ref) hold. Let $x\in int(\mathcal{X})$ be given and Assumptions (ref), (ref), (ref), (ref), and (ref) hold. Then, \[ \sqrt{nh_{n}^{d_{x}}}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)-\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\Rightarrow\mathbb{G}_{x}^{e}(\cdot)\ \text{in }\left(l^{\infty}(\mathcal{Y})\right)^{4}, \] where $\mathbb{G}_{x}^{e}((y_{1l},y_{1u},y_{0l},y_{0u}))\equiv(\mathbb{G}_{1,x}^{e}(y_{1l}),\mathbb{G}_{Y,x}^{e}(y_{1u}),\mathbb{G}_{Y,x}^{e}(y_{0l}),\mathbb{G}_{0,x}^{e}(y_{0u}))^{t}$ and $\mathbb{G}_{1,x}^{e}$, $\mathbb{G}_{0,x}^{e}$, and $\mathbb{G}_{Y,x}^{e}$ are Gaussian processes with mean zero and covariance kernels being $H_{1,x}^{e}(\cdot,\cdot)$, $H_{0,x}^{e}(\cdot,\cdot)$, and $H_{Y,x}^{e}(\cdot,\cdot)$, respectively. The forms of the covariance kernels are given in Appendix (ref). In addition, \begin{equation} \begin{aligned}\sqrt{nh_{n}^{d_{x}}}\left(\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)-\phi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\right) & \Rightarrow\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}^{e}(\cdot)\right),\\ \sqrt{nh_{n}^{d_{x}}}\left(\phi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x))-\phi_{U}^{e}(\mathbb{F}_{\mathbf{Y}|X}(\cdot|x))\right) & \Rightarrow\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}^{e}(\cdot)\right), \end{aligned} \end{equation} in $l^{\infty}\left(\mathcal{Y}^{4}\right)$. The forms of $\phi_{L}^{'}$ and $\phi_{U}^{'}$ are provided in Appendix (ref).

We simulate the limiting distribution provided in Theorem (ref) in a similar way to the exogenous treatment case. Let $\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x)$ denote the vector of simulated stochastic processes using estimated influence functions in ((ref)) in Appendix (ref). The following theorem shows the validity of the bootstrap procedure:

theoremSuppose that conditions in Theorem (ref) are satisfied. In addition, if Assumptions (ref) and (ref) also hold, then, we have \begin{align*} \widehat{\phi_{L}^{'}}\left(\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x);\sqrt{nh_{n}^{d_{x}}}\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right) & \Rightarrow\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}^{e}(\cdot)\right),\\ \widehat{\phi_{U}^{'}}\left(\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x)+1;\sqrt{nh_{n}^{d_{x}}}\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right) & \Rightarrow\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}^{e}(\cdot)\right), \end{align*} in $l^{\infty}\left(\mathcal{Y}^{4}\right)$, conditional on data. The forms of $\widehat{\phi_{L}^{'}}$ and $\widehat{\phi_{U}^{'}}$ are provided in Appendix (ref).

We provide several extensions of the main results in this section in Appendix. We extend the results of abrevaya2015estimating who consider estimating conditional average treatment effects on a subset of covariates to the conditional distribution of treatment effects (Appendix (ref)). This result may be practically relevant when the number of covariates is large. We also discuss how to adopt the results to conduct global hypothesis testing (Appendix (ref)).

It is worth noting that firpo2019partial provide formal definitions of pointwise and uniform sharpness of bounds on distribution functions and show that the Makarov bounds are not uniformly sharp. The inference methods developed in this paper are based on pointwise sharp bounds on the conditional distribution of treatment effects and are valid uniformly over the support of treatment effects. Inference for uniformly sharp bounds on the distribution of treatment effects is left as future research.

Monte Carlo Simulation

We conduct a set of simulations to investigate the performance of the bootstrap in finite samples. To this end, the following DGP is considered:

align*[align* omitted — 163 chars of source]

where $X=2\tilde{X}-1$ with $\tilde{X}$ being an uniform a random variable on $[0,1]$, $U\equiv(U_{1},U_{0})^{t}\sim N\left(

pmatrix[pmatrix omitted — 19 chars of source]

,

pmatrix[pmatrix omitted — 27 chars of source]

\right)$, and $V\sim N(0,1)$. The conditional treatment effect on $X=x$ is defined as $(\mu_{1}-\mu_{0})+x(\beta_{1}-\beta_{0})+((\phi_{1}+X\gamma_{1})U_{1}-(\phi_{0}+X\gamma_{0})U_{0})$. The parameter values are set as follows: $\beta_{1}=\gamma_{1}=1$, $\beta_{0}=\gamma_{0}=0.9$, $\phi_{1}=\phi_{0}=1$, and $\alpha=1$. We consider various values for parameter vector $(\mu_{1},\mu_{0})^{t}$ to investigate the performance of a KS test statistic under the null and alternative hypotheses. Specifically, to investigate the performance of the KS test under null hypotheses, we consider the following hypothesis: \[ H_{0}:F_{\Delta|X}^{L}(\delta|x=0)=\left(2\cdot\Phi\left(\frac{\delta}{2}\right)-1\right)\cdot\mathbf{1}(\delta\geq0)\text{ for all }\delta, \] where $\Phi(\cdot)$ is the standard normal distribution function.

The lower bound in the null hypothesis is one derived by frank1987best, and the null hypothesis is true if and only if $\mu_{1}=\mu_{0}=0$. The KS statistic for this null is constructed as follows: \[ KS_{n}\equiv\sup_{\delta}\sqrt{nh_{n}}\Bigg|F_{\Delta|X}^{L}(\delta|x=0)-\left(2\cdot\Phi\left(\frac{\delta}{2}\right)-1\right)\cdot\mathbf{1}(\delta\geq0)\Bigg|. \] To investigate the performance of the KS test under some alternative hypotheses, we consider the cases where $(\mu_{1},\mu_{0})^{t}=(\mu,0)^{t}$ for $\mu\in\{-1,1\}$. The bandwidth is chosen to be $h_{n}=1.06\times s.d(X)\times n^{-1/6}$, where $s.d(X)$ denotes the (sample) standard deviation of $X$.

It is worth emphasizing that the asymptotic distribution of the KS test statistic is nonstandard and may not be uniformly valid with respect to the DGP. For this reason, it is important to investigate whether the KS statistic performs well with various choices for $(a_{n})$. Specifically, we consider several rates at which $a_{n}$ grows to the infinity: (i) $a_{n}=c\times\log\left(\log\left(nh_{n}\right)\right)/\sqrt{nh_{n}}$, (ii) $a_{n}=c\times\sqrt{\log\left(nh_{n}\right)}/\sqrt{nh_{n}}$, and (iii) $a_{n}=c\times\left(nh_{n}\right)^{1/6}/\sqrt{nh_{n}}$ for some $c>0$. These choices of $(a_{n})$ satisfy Assumption (ref). We consider various values of $c$ ranging from $0.1$ to $0.5$ to see whether the finite-sample performance of the KS test is sensitive to the choice of $c$.\footnote{firpo2021uniform suggest using $c=0.2$ with $a_{n}=c\times\log\left(\log\left(nh_{n}\right)\right)/\sqrt{nh_{n}}$ in our context. }

The sample size $n$ is set to be 500, and the number of bootstrap iterations is 500. The bootstrap weight $B$ is drawn from the standard normal distribution, and all simulation results are obtained from 500 iterations. The nominal level is set to be 0.05.

Table (ref) presents the simulation results for rejection probabilities. We find that the rejection probability under $H_{0}$ tends to decrease as $c$ increases, except for the case of $\sqrt{nh_{n}}a_{n}\propto\left(nh_{n}\right)^{1/6}$.\footnote{We also considered larger values of $c$ than $0.5$, but they are not reported here. The simulation results with those values of $c$ suggest that the resulting confidence bands would be too conservative (i.e., the rejection probabilities under the null hypothesis are relatively smaller than the nominal rate). As a result, we do not recommend using a too large value of $c$, based on our simulation results. It would be an interesting question how to choose $c$ or the rate of $(a_{n})$ in a data-dependent way, but this is far beyond the scope of this paper. We therefore leave this interesting and important question for future research.} The KS statistic performs well in finite samples for various values of $(a_{n})$ in the sense that the rejection probability under $H_{0}$ is close to the nominal probability and that the rejection probabilities under $H_{1}$ are large in general.

table[table omitted — 1,283 chars of source]

Empirical Application: The Effect of 401(k) Plans on Net Financial Assets

In this section, we provide an empirical example to illustrate the usefulness of the methods proposed in this paper. We revisit the empirical question on the effect of participation in 401(k) plans on net financial assets investigated by many studies in the literature (e.g., Abadie2003,Chernozhukov2004,Wuethrich2019,SantAnna2022). Our main goals in this empirical application are twofold. First, we empirically show that the stochastic dominance assumptions (Assumptions (ref) and (ref)) can provide considerable identifying power. Second, we complement the existing results that there is substantial heterogeneity in treatment effects across income groups by estimating the bounds on the conditional distribution of treatment effects on income levels without an instrumental variable. In doing so, we complement the empirical results on the effect of 401(k) plans on net financial assets documented in the literature by providing estimation results on the distribution of the treatment effect. Since our focus is on verifying the identifying power of the stochastic dominance assumptions and potential heterogeneity in the treatment effect across different subpopulations, we do not report confidence bands for clear illustration.

The U.S. introduced several tax-deferred retirement plans, including 401(k) plans and individual retirement accounts (IRAs), in the early 1980s. These retirement plans can be used as a way to accumulate individual assets. Many papers in the literature have considered how those tax-deferred retirement plans affect asset accumulation or savings. The main challenge with identifying and estimating the causal effect of 401(k) participation on assets is that participation in 401(k) plans is endogenously determined. Furthermore, the effect of 401(k) plans on net financial assets is heterogeneous across income levels, as shown by Chernozhukov2004. Motivated by the empirical results of Chernozhukov2004, we focus on the distribution of treatment effects conditional on an individual's income. Moreover, while it is common to use the eligibility for 401(k) as an instrumental variable to point identify some distributional effects (e.g., quantile treatment effects), our identification and estimation strategies do not rely on such an instrumental variable.

We use the data from SantAnna2022 for this empirical analysis. The original dataset contains 9,910 households from the 1991 Survey of Income and Program Participation. The dependent variable is the amount of net financial assets measured in ten thousand dollars. We exclude observations with a value of the dependent variable higher than the 0.99 sample quantile and lower than the 0.01 sample quantile of the net financial assets. This results in a sample of 9,712 households. The treatment variable is a binary variable indicating whether a household participates in 401(k) plans. To investigate the potential heterogeneity across income levels, we use the income variable as the covariate of interest. For stability of estimation, we standardize the original variable of income and consider the sample mean and various quantiles of income. Table (ref) reports the summary statistics of the data.

table[table omitted — 620 chars of source]

We first discuss the validity of Assumptions (ref) and (ref) in this empirical example. It is well known that the preference for saving is heterogeneous in that some people have a stronger preference for saving than others. This leads to the nonrandom selection into participation in 401(k) plans or other tax-deferred retirement plans (e.g., Chernozhukov2004). Based on this observation, it is likely that people who participate in 401(k) plans would have a stronger preference for saving than those who do not participate. Therefore, the net financial assets of people with a strong preference for saving tend to be larger than those of people with a weak preference for saving, suggesting that it may be plausible to impose Assumption (ref) on the model.

Assumption (ref) is consistent with the empirical example as most tax-deferred retirement plans, including 401(k) plans, by themselves increase the amount of assets that an individual possess. Therefore, regardless of whether an individual participates in 401(k) plans or not, the potential net financial assets that one would have had if she participated in 401(k) plans are likely to be larger than those she would have had if she did not participate in 401(k). As a result, it is plausible to impose Assumption (ref) on the model.

For estimation, we set $h_{n}=1.06\times n^{-1/(5+d_{x})}$ and $a_{n}=0.2\times\log\left(\log\left(nh_{n}^{d_{x}}\right)\right)/\sqrt{nh_{n}^{d_{x}}}$ with $d_{x}=1$. We denote the $\tau$-th (sample) quantile of income by $Q_{income}(\tau)$.

We note that when Assumptions (ref) and (ref) are not imposed, the estimated bounds are the logical ones, regardless of the conditioning value of the income.\footnote{By the logical bounds, we mean that the lower and upper bounds are equal to 0 and 1, respectively, for all $\delta\in Supp(\Delta|X=x)$. } On the other hand, the estimated bounds under Assumptions (ref) and (ref) are informative in the sense that they are not the logical bounds. This indicates that the stochastic dominance assumptions have considerable identifying power.

We find substantial heterogeneity in the distribution of treatment effects across different values of the income. Figure (ref) compares the estimated bounds on the conditional distribution of treatment effects at the 0.2 and 0.8 quantiles of the income. Specifically, the star-marked lines are the bounds on the conditional distribution of treatment effects given the 0.2 quantile of the income. The circle-marked lines are the bounds on the conditional distribution of treatment effects given the 0.8 quantile of income. We find that the lower bound conditional on the 0.2 quantile of the income is larger than the lower bound conditional on the 0.8 quantile of the income. Conversely, the upper bound conditional on the 0.8 quantile of the income is larger than the upper bound conditional on the 0.2 quantile of the income for all $\delta\in[-24,24]$ (i.e., from -\$240,000 to \$240,000). When considering the bounds on $\Pr\left(Y_{1}-Y_{0}\leq1|X=x\right)$ for $x\in\left\{ Q_{income}(0.2),Q_{income}(0.8)\right\} $, the estimation results show that $\Pr\left(Y_{1}-Y_{0}\leq1|X=Q_{income}(0.2)\right)\in\left[0.5804,1\right]$ and that $\Pr\left(Y_{1}-Y_{0}\leq1|X=Q_{income}(0.8)\right)\in\left[0.0638,1\right]$. These estimated bounds suggest that the proportion of individuals who experience a positive treatment effect larger than \$10,000 among those with the income being equal to the 0.2 sample quantile is at most 41.96%. The proportion among those with the income being equal to the 0.8 sample quantile is at most 93.62%. As a result, there may be a possibility that the proportion of individuals who experience a certain level of positive treatment effect is larger when considering a higher level of income. This argument is in part consistent with the finding of Chernozhukov2004 that the quantile treatment effects of 401(k) plans on net financial wealth tend to increase as the income level increases (see Figure 2 in Chernozhukov2004).

Figure (ref) compares the estimated bounds at a specific quantile level with those at the mean of income under Assumptions (ref) and (ref). When considering a low quantile level, e.g., $\tau\in\{0.1,0.2,0.3\}$, we find that the lower bound conditional on $X=Q_{income}(\tau)$ is larger than that conditional on the mean of income. The upper bound conditional on $X=Q_{income}(\tau)$ is smaller than that conditional on the mean of income over the potential support of the treatment effect. However, when considering a high quantile level, e.g., $\tau\in\{0.7,0.8,0.9\}$, the estimation results are the opposite. These estimation results indicate that the treatment effect of 401(k) plans on net financial assets is likely to be heterogeneous across income levels, which is consistent with the finding of Chernozhukov2004.

Conclusion

This paper considers identification and estimation of bounds on the conditional distribution of treatment effects. The conditional distribution may provide evidence on potential heterogeneity in treatment effects across subpopulations that are defined in terms of values of covariates, and therefore, is of practical importance in many empirical studies. We show that when the treatment is endogenously determined, one can tighten the bounds by imposing stochastic dominance assumptions. These assumptions are consistent with many economic theories and easy to interpret, and the resulting bounds on the distribution of treatment effects are easy to compute. We propose nonparametric estimators of the bounds and establish the uniform asymptotic theory based on the novel approach of fang2019inference and firpo2021uniform. The asymptotic theory in this paper is useful for constructing uniform confidence bands and conducting statistical tests for global hypotheses. We then provide an empirical application of the methodology proposed in this paper to illustrate its relevance to empirical research.

There are several interesting directions for future research. First, one can consider inference that is uniformly valid regardless of whether the (conditional) distributions of the potential outcomes are point identified or partially identified, as in, for example, Imbens2004, Stoye2009, and andrews2010inference. Second, one can consider testing global hypotheses, such as stochastic dominance. Although we cannot directly test stochastic dominance between two conditional distribution of treatment effects as they are not point identified, one can provide weak evidence by using similar arguments to those in firpo2021uniform. Third, it would be fruitful to develop inference methods that are uniformly valid in values of covariates. In work in progress, we consider a semiparametric approach to uniform inference over the support of treatment effects and covariates. It is expected to help resolve many interesting questions in empirical analysis that cannot be answered by the framework proposed in this paper. Lastly, it is worth considering inference in the presence of instrumental variables.