EconBase
← Back to paper

Partial Identification and Inference for Conditional Distributions of Treatment Effects

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

82,444 characters

Partial Identification and Inference for Conditional Distributions of Treatment Effects


\title{Partial Identification and Inference for \\
Conditional Distributions of Treatment Effects\thanks{This is a revised version of Chapter 3 of author's dissertation. I
am very grateful to Jason Abrevaya, Stephen Donald, Sukjin Han, Yu-Chin
Hsu, and participants in 2022 IAAE Annual Conference, 2022 KES Annual
Conference, and the seminar at Sungkyunkwan University for their valuable
comments and suggestions. Jiyong Jung provided excellent research
assistance. This paper was supported by the New Faculty Research Funding
provided by Sogang University. All errors are mine.}}
\author{Sungwon Lee\thanks{35 Baekbeom-ro Mapo-gu, Sogang University GN 710, Seoul 04107, South
Korea. }\\
Department of Economics\\
Sogang University\\
[email removed]\\
}
\date{This draft: \today }
\maketitle
\begin{abstract}
This paper considers identification and inference for the distribution
of treatment effects conditional on observable covariates. Since the
conditional distribution of treatment effects is not point identified
without strong assumptions, we obtain bounds on the conditional distribution
of treatment effects by using the Makarov bounds. We also consider
the case where the treatment is endogenous and propose two stochastic
dominance assumptions to tighten the bounds. We develop a nonparametric
framework to estimate the bounds and establish the asymptotic theory
that is uniformly valid over the support of treatment effects. An
empirical example illustrates the usefulness of the methods. \textit{\emph{ }}\\
\textit{\emph{}}\\
\emph{Keywords:} treatment effects, conditional distribution, heterogeneity,
partial identification, uniform inference. \\
\emph{JEL Classification Numbers:} C14, C21.
\end{abstract}
\newpage{}

\section{\label{sec:Introduction} Introduction }

This paper considers identification and estimation of the conditional
distribution of treatment effects.\footnote{The distribution of treatment effects has received significant attention,
and its importance becomes greater when the benefit from the treatment
is non-transferrable (\citet{heckman2007econometric}).} While the unconditional distribution of treatment effects has been
considered extensively in the literature in both theoretical and applied
econometrics (e.g., \citet{heckman1997making,fan2010sharp,fan2012confidence,fan2017partial,firpo2019partial,frandsen2021partial}),
we focus on the conditional distribution of treatment effects to take
into account the heterogeneity caused by differences in the value
of covariates.\footnote{For example, the effect of a mother's smoking behavior on her child's
birth weight might differ across the mother's age (e.g., \citet{abrevaya2015estimating}).} Such heterogeneous treatment effects, if they exist, may help policy
makers develop more effective policies. We use the \emph{Makarov bounds},
which depend on the marginal distributions of potential outcomes while
the dependence structure between them is left unspecified, to partially
identify the conditional distribution of treatment effects (e.g.,
\citet{makarov1982estimates}, \citet{ruschendorf1982random}, and
\citet{williamson1990probabilistic}). We then provide estimation
and inference methods for the bounds.

We start with the case where the assumption of the conditional independence
of the treatment and potential outcomes given covariates, which is
called an unconfoundedness condition, is satisfied. However, there
are many situations where such a conditional independence assumption
fails to hold. We present identification results of the distribution
of treatment effects when the treatment is endogenous \textit{without}
assuming the existence of an instrumental variable. We show that one
can still obtain Makarov-type bounds on the distribution of treatment
effects even with an endogenous treatment. When a treatment is endogenous,
the bounds on the distribution of the treatment effect may be too
wide to be informative.\footnote{This is also true for the case where we impose the unconfoundedness
assumption. We can expect that, since we are interested in the distribution
of treatment effects for some subgroup, the bounds on the conditional
distribution may be narrower than those on the unconditional distribution.
This is another motivation for considering the conditional distributions
of treatment effects in this paper.} Motivated by the previous studies in the literature (e.g., \citet{manski1997monotone,manski2000monotone,blundell2007changes}),
we consider two kinds of stochastic dominance between potential outcomes
in this paper to tighten the bounds. These stochastic dominance assumptions
are useful in many empirical situations as they are consistent with
many economic theories. The resulting bounds under the stochastic
dominance assumptions are easy to compute.

We construct nonparametric kernel-type estimators of the bounds on
the conditional distribution of treatment effects and establish asymptotic
theory for them that is uniformly valid over the support of treatment
effects for some fixed subgroup of the population. They are useful
for comparing bounds between two subpopulations defined in terms of
different values of observable characteristics.\footnote{One can perform statistical testing for global hypotheses, such as
equality of or stochastic dominance between two bounds (e.g., \citet{barrett2003consistent,lee2009testing,seo2018tests}),
and this requires asymptotic theory uniformly valid over the support
of treatment effects. } We adapt the results on uniform inference for value functions that
were developed recently by \citet{firpo2021uniform} and provide a
set of conditions under which the estimated bounds are consistent
for the true bounds. We propose a bootstrap procedure to mimic the
asymptotic distributions of the estimated bounds and show its validity.
The bootstrap scheme in this paper is based on the novel bootstrap
procedure for Hadamard directionally differentiable functionals that
was developed by \citet{fang2019inference}.

The asymptotic theory developed in this paper relaxes some undesirable
assumptions imposed in several existing studies in the literature. Specifically,
the bounds on the distribution of treatment effects are defined as
some functionals of marginal distributions of potential outcomes.
These functionals involve the infimum and supremum over the supports
of potential outcomes. The uniqueness of the infimum and supremum
needs to be imposed to establish a pointwise asymptotic theory for
those bounds, as in \citet{fan2010sharp,fan2012confidence}.\footnote{\citet{fan2010sharp} impose Assumptions 3 and 4 to guarantee the
uniqueness of the maximizer and minimizer involved in the bounds on
the distribution of treatment effects. \citet{fan2012confidence}
impose Assumptions A3 and A4 for the same purpose. } However, the uniqueness assumption may not hold for some data generating
process (DGP), and the pointwise asymptotic results may not be valid
in such cases (cf. \citet{milgrom2002envelope}, \citet{firpo2021uniform}).
On the other hand, the inference results developed in this paper do
not require such uniqueness assumptions.

One can use the novel approach developed by \citet{chernozhukov2013intersection}
that yields confidence bands uniformly valid over the joint support
of an outcome variable and covariates.\footnote{\citet{chernozhukov2013intersection} do not explicitly provide an
inference result for nonparametric conditional distribution estimators
that is uniformly valid over the joint support of an outcome variable
and covariates. However, one can easily adapt their results in the
literature to establish the strong Gaussian approximation of some
nonparametric estimator of a conditional distribution function. This
allows one to perform inference uniformly over the joint support of
the outcome variable of interest and conditioning variables (e.g.,
\citet{chernozhukov2014gaussian}).} Their inference methods for bounds defined by either supremum or
infimum of some function are based on the strong Gaussian approximation
using couplings. A key ingredient of their approach is the construction
of argsup and/or arginf sets, and it is allowed to avoid imposing
the uniqueness assumption on those sets. Our approach differs from
theirs in that we rely on weak convergence of nonparametric estimators
of the bounds to some fixed Gaussian process and the results on Hadamard
directionally differentiable functionals. In addition, while our approach
also requires to estimate argsup and arginf sets for the Hadamard
directional derivatives, the construction of these sets is computationally
easy in comparison to that of \citet{chernozhukov2013intersection}.
Therefore, the inference theory in this paper can be considered complementing
the existing inference methods for the Makarov bounds.

The Monte Carlo simulation results show that the inference methods
in this paper perform well in finite samples. We also provide an empirical
example to illustrate the practical relevance of the methods proposed
in this paper. We revisit the empirical question of the effect of
401(k) plans on net asset accumulations investigated by multiple papers
in the literature (e.g., \citet{Abadie2003,Chernozhukov2004,Wuethrich2019,SantAnna2022}).
We confirm that the identifying power of the stochastic dominance
assumptions is substantial from the empirical application. Furthermore,
we find evidence on heterogeneity in the treatment effect that is
consistent with the results in the literature.

This paper contributes to several strands of the literature on treatment
effects. First, this paper mainly contributes to the literature on
identification and estimation of the distribution of treatment effects
by providing nonparametric estimation and inference methods for the
conditional distribution of treatment effects. The literature is too
vast to list all the papers relevant to this point. Identification
of the unconditional distribution of treatment effects has been studied
by, for example, \citet{heckman1997making}, \citet{fan2010sharp,fan2012confidence},
\citet{fan2017partial}, \citet{vuong2017counterfactual}, \citet{kim2018identifying},
and \citet{firpo2019partial}.

This paper is closely related to \citet{kim2018partial} and \citet{frandsen2021partial}.
\citet{kim2018partial} provides identification results under several
distributional restrictions, together with an instrumental variable,
that help tighten the bounds on the distribution of treatment effects,
when the treatment is endogenous. The restrictions considered by \citet{kim2018partial}
are general and closely related to the stochastic dominance assumptions
in this paper that were proposed independently of the results in \citet{kim2018partial}.
This paper differs from \citet{kim2018partial} in that this paper
focuses on estimation and inference for the conditional distribution
of treatment effects with a motivation for treatment effect heterogeneity,
whereas \citet{kim2018partial} mainly considers identification of
the unconditional distribution of treatment effects under some restrictions
on the model. Moreover, it is much easier to compute and estimate
the bounds proposed in this paper than those in \citet{kim2018partial}.
Therefore, the bounds in this paper have great applicability to empirical
research.

In comparison with \citet{frandsen2021partial}, we focus on estimation
and inference for the bounds on the conditional distribution of treatment
effects with a motivation for heterogeneity across subgroups. On the
other hand, \citet{frandsen2021partial} are mainly concerned with
identification of the unconditional distribution of treatment effects
under a novel restriction that is called the ``stochastic increasingness''
assumption. \citet{frandsen2021partial} also discuss how to incorporate
covariates and estimate the bounds on the unconditional distribution
of treatment effects. However, they do not develop the asymptotic
theory for the conditional distribution of treatment effects, especially
when the bounds on the conditional distribution are the Makarov bounds.

This paper also contributes to the literature on heterogeneity in
treatment effects across subpopulations by considering the conditional
distribution of treatment effects (e.g., \citet{donald2012incorporating,abrevaya2015estimating,chang2015nonparametric,hsu2017consistent,shen2019estimation}).
Most existing studies focus on average/quantile treatment effects
or the marginal distributions of potential outcomes. The results of
this paper complement them by considering conditional distributions
of treatment effects.

The rest of this paper is organized as follows. Section \ref{sec:ID}
presents the model and identification results. Section \ref{sec:Estimation}
provides nonparametric estimators of the Makarov bounds on the conditional
distribution of treatment effects and the bootstrap procedure. Section
\ref{sec:Inference} develops the asymptotic theory for the estimated
bounds on the conditional distribution of treatment effects. Section
\ref{sec:MC} presents the Monte Carlo simulation study, and Section
\ref{sec:Empirical} provides the empirical example considering the
effect of 401(k) plan on net financial asset accumulations. We then
conclude with Section \ref{sec:Conclusion}. All mathematical proofs,
technical expressions, and additional results are presented in Appendix.

Before proceeding, we introduce some notation that will be used throughout
this paper. Two random variables $A$ and $B$ that are independent
of each other are denoted by $A\perp B$, and $\mathbb{E}[\cdot]$
is the expectation operator. For a matrix $A$, $A^{t}$ is the transpose
of $A$. For a set $A$, $l^{\infty}(A)$ denotes the set of uniformly
bounded functions on $A$. For a sequence of random variables $(Z_{n})$
and a random variable $Z$, we denote the weak convergence of $Z_{n}$
to $Z$ by $Z_{n}\Rightarrow Z$.\footnote{A formal definition of weak convergence can be found in, for example,
\citet[pp.107-108]{kosorok2008}.} For a set $A$, $int(A)$ is the interior of $A$.


\section{\label{sec:ID} Model and Identification}

\subsection{Identification Under the Unconfoundedness Condition}

Let $(\Omega,\mathcal{S},\mathbf{P})$ be a probability space. Let
$D$ be a binary variable that indicates whether a person gets the
treatment or not, in other words, $D=1$ if the person gets the treatment,
and $D=0$ if the person does not get the treatment. For each $d\in\{0,1\}$,
let $Y_{d}$ denote the potential outcome when $D=d$, and the observed
outcome is defined as $Y\equiv DY_{1}+(1-D)Y_{0}$. We assume that
$Y_{d}$ is a continuous random variable for each $d\in\{0,1\}$.
Let $X\in\mathbb{R}^{d_{x}}$ be a set of covariates and denote its
support by $\mathcal{X}$. We can only observe $(Y,D,X^{t})^{t}$
from the data. We denote the conditional distribution of $Y_{d}$
on $X=x$ by $F_{d|X}(\cdot|x)$ for each $d\in\{0,1\}$, and let
$p_{0}(x)\equiv\Pr(D=1|X=x)$ for a given $x\in\mathcal{X}$. $F_{X}(\cdot)$
denotes the distribution function of $X$.

We begin with the models under the conditional independence of the
treatment variable and impose the following assumptions.

\begin{assumption}\label{assu:unconfounded} $(Y_{1},Y_{0})\perp D|X$.
\end{assumption}

\begin{assumption}\label{assu:bound_propensity} There exist $\text{\ensuremath{\underline{p},\bar{p}\in(0,1)} such that \ensuremath{p_{0}(x)\in[\underline{p},\bar{p}]} uniformly in \ensuremath{x\in\mathcal{X}}}$.
\end{assumption}

Assumption \ref{assu:unconfounded} is the unconfoundedness assumption,
which means that the treatment $D$ is independent of the potential
outcomes conditional on $X$. Assumption \ref{assu:bound_propensity}
is an overlap condition that is commonly imposed in the relevant literature
(e.g., \citet{imbens2004nonparametric}). This assumption implies
that we can observe individuals with $Y=Y_{1}$ and those with $Y=Y_{0}$
for any value of $x\in\mathcal{X}$. Under these assumptions, we have
the following identification result:

\begin{lemma}\label{lem:id_unconfounded} Let $x\in\mathcal{X}$
be given. Suppose that Assumptions \ref{assu:unconfounded} and \ref{assu:bound_propensity}
are satisfied. For any measurable function $G$ such that $\mathbb{E}[|G(Y_{1})||X=x]<\infty$
and $\mathbb{E}[|G(Y_{0})||X=x]<\infty$, we have
\begin{equation}
\begin{aligned}\mathbb{E}[G(Y_{1})|X=x] & =\frac{\mathbb{E}[D\cdot G(Y)|X=x]}{\mathbb{E}[D|X=x]},\\
\mathbb{E}[G(Y_{0})|X=x] & =\frac{\mathbb{E}[(1-D)\cdot G(Y)|X=x]}{\mathbb{E}[(1-D)|X=x]}.
\end{aligned}
\label{eq:ID_Marginal_Unconfounded}
\end{equation}
 \end{lemma}

Lemma \ref{lem:id_unconfounded} implies that we can identify the
conditional distributions of $Y_{1}$ and $Y_{0}$ on $X=x$ when
we choose $G(Y)\equiv\mathbf{1}(Y\leq y)$ for a given $y\in\mathbb{R}$.
This result is also closely related to the identification of the unconditional
distributions of $Y_{1}$ and $Y_{0}$ that is considered by \citet{donald2014estimation}.

The treatment effect $\Delta$ is defined as the difference between
$Y_{1}$ and $Y_{0}$ (i.e., $\Delta\equiv Y_{1}-Y_{0}$). The parameter
of interest is the conditional distribution of treatment effects given
$X=x$ for a given value $x\in\mathcal{X}$, and this conditional
distribution is denoted by $F_{\Delta|X}(\cdot|x)$.

The (unconditional) distribution of treatment effects has received
a considerable amount of attention from the literature (e.g., \citet{heckman1997making},
\citet{fan2010sharp,fan2012confidence}, \citet{fan2017partial},
\citet{firpo2019partial}). The distribution function is useful in
the context of program evaluation, for example, to determine the proportion
of people who benefit from being treated, namely, $\Pr(\Delta\geq0)=1-F_{\Delta}(0)$.
One can refer to \citet{abbring2007econometric} for more discussion
on the topic. The conditional distributions takes potential heterogeneity
across subpopulations into account and thus can provide much richer
information on treatment effects. Such heterogeneity, if it exists,
may deliver different policy implications for different subpopulations.

It is worth noting that even in a randomized experiment where one
can point identify the conditional distributions of $Y_{1}$ and $Y_{0}$
given $X$, the distribution of treatment effects is not point identified
without additional structures of the model.\footnote{When the conditional distributions of $Y_{1}$ and $Y_{0}$ given
$X$ are point identified, a sufficient condition for point identification
of the joint distribution of $Y_{1}$ and $Y_{0}$ is the (conditional)
rank invariance; that is, $F_{1|X}(Y_{1}|X)=F_{0|X}(Y_{0}|X)$ almost
surely. } In particular, \citet{fan2010sharp} provide sharp bounds on the
distribution of treatment effects in randomized experiments that are
based on \citet{makarov1982estimates}, and \citet{firpo2019partial}
improve the bounds of \citet{fan2010sharp} in the uniform sense.
We do not focus on the uniform sharpness of the bounds but on pointwise
sharp bounds of the conditional distribution of treatment effects
and nonparametric estimation of the bounds. We recall the bounds on
the conditional distribution of treatment effects $F_{\Delta|X}(\cdot|x)$:

\begin{proposition}\label{prop:ID_dist_TE_Unconfounded} (Lemma 2.1
in \citet{fan2010sharp}) Let $x\in\mathcal{X}$ be fixed. Then, for
a given $\delta\in\mathbb{R}$, define
\begin{eqnarray}
F_{\Delta|X}^{L}(\delta|x) & = & \max\left(\sup_{y\in\mathbb{R}}\left\{ F_{1|X}(y|x)-F_{0|X}(y-\delta|x)\right\} ,0\right),\label{eq:ID_F_Delta_LB}\\
F_{\Delta|X}^{U}(\delta|x) & = & \min\left(\inf_{y\in\mathbb{R}}\left\{ F_{1|X}(y|x)-F_{0|X}(y-\delta|x)\right\} ,0\right)+1.\label{eq:ID_F_Delta_UB}
\end{eqnarray}
 If Assumptions \ref{assu:unconfounded} and \ref{assu:bound_propensity}
are satisfied, we then have
\begin{equation}
F_{\Delta|X}(\delta|x)\in\left[F_{\Delta|X}^{L}(\delta|x),F_{\Delta|X}^{U}(\delta|x)\right].\label{eq:ID_F_Delta}
\end{equation}
 \end{proposition}

Note that this result is essentially identical to the identification
result in \citet{fan2010sharp}. Without additional structures on
the model, these bounds are sharp in the pointwise sense for given
$x\in\mathcal{X}$. It is also worth noting that one can derive bounds
on the unconditional distribution of treatment effects by averaging
the bounds on the conditional distribution. Specifically, if the potential
outcomes are correlated with $X$, then the bounds on the unconditional
distribution obtained from the conditional distributions of the potential
outcomes are tighter than the those obtained from the unconditional
distributions of the potential outcomes (cf. \citet{firpo2019partial}).

\subsection{Identification with an Endogenous Treatment}

In this section, we discuss identification of the distribution of
treatment effects with an endogenous binary treatment. When the treatment
is endogenously determined, the conditional distributions of $Y_{1}$
and $Y_{0}$ given $X$ are generally partially identified without
additional assumptions. The following theorem shows that even if the
conditional distributions of $Y_{1}$ and $Y_{0}$ given $X$ are
partially identified, we can still bound $F_{\Delta|X}(\delta)$.

\begin{theorem}\label{thm:ID_Endo_FH_Bound} Let $x\in\mathcal{X}$
be fixed. Suppose that, for all $y\in\mathbb{R}$, the identified
sets of $F_{1|X}(y|x)$ and $F_{0|X}(y|x)$ are given by $\left[F_{1|X}^{L}(y|x),F_{1|X}^{U}(y|x)\right]$,
and $\left[F_{0|X}^{L}(y|x),F_{0|X}^{U}(y|x)\right]$, respectively. For a given $\delta\in Supp(\Delta|X=x)$, define
\begin{eqnarray}
F_{\Delta|X}^{e,L}(\delta|x) & = & \max\left(\sup_{y}\left\{ F_{1|X}^{L}(y|x)-F_{0|X}^{U}(y-\delta|x)\right\} ,0\right),\label{eq:trt1}\\
F_{\Delta|X}^{e,U}(\delta|x) & = & \min\left(\inf_{y}\left\{ F_{1|X}^{U}(y|x)-F_{0|X}^{L}(y-\delta|x)\right\} ,0\right)+1.\label{eq:trt2}
\end{eqnarray}
 Then,
\[
F_{\Delta|X}^{e,L}(\delta|x)\leq F_{\Delta|X}(\delta|x)\leq F_{\Delta|X}^{e,U}(\delta|x).
\]
 \end{theorem}

One example of the identified sets of $F_{1|X}$ and $F_{0|X}$ is
Manski's bounds (\citet{manski1990nonparametric,manski1994selection})
defined as follows:
\begin{equation}
\begin{aligned}F_{1|X}^{L}(y|x) & \equiv\Pr(Y\leq y|D=1,X=x)\Pr(D=1|X=x),\\
F_{1|X}^{U}(y|x) & \equiv\Pr(Y\leq y|D=1,X=x)\Pr(D=1|X=x)+\Pr(D=0|X=x),\\
F_{0|X}^{L}(y|x) & \equiv\Pr(Y\leq y|D=0,X=x)\Pr(D=0|X=x),\\
F_{0|X}^{U}(y|x) & \equiv\Pr(Y\leq y|D=0,X=x)\Pr(D=0|X=x)+\Pr(D=1|X=x).
\end{aligned}
\label{eq:manski_bound}
\end{equation}
 The bounds presented in Theorem \ref{thm:ID_Endo_FH_Bound} based
on Manski's bounds are often too wide to be informative. Some model
restrictions have been imposed to tighten bounds on the parameter
of interest in the literature (e.g., \citet{manski1997monotone,manski2000monotone,blundell2007changes}).
Motivated by these studies, we propose two stochastic dominance assumptions
that may help tighten the bounds and be consistent with many economic
theories in the absence of instrumental variables. The bounds on the
distribution of treatment effects in Theorem \ref{thm:ID_Endo_FH_Bound}
are based on the bounds on the conditional distributions of the potential
outcomes, and therefore, we can tighten the bounds on $F_{\Delta|X}$
once we have tighter bounds on the conditional distributions of the
potential outcomes than (\ref{eq:manski_bound}).

Let $F_{d|d^{'}X}(y|x)\equiv\Pr\left(Y_{d}\leq y|D=d^{'},X=x\right)$
for given $d,d^{'}\in\{0,1\}$. The following assumption states that
the potential outcome conditional on $D=1$ first-order stochastically
dominates the potential outcome conditional on $D=0$.

\begin{assumption}\label{assu:fsd1} Let $x\in\mathcal{X}$ be given.
For all $d\in\{0,1\}$ and $y\in\mathbb{R}$, $F_{d|1,X}(y|x)\leq F_{d|0,X}(y|x)$.
\end{assumption}

Assumption \ref{assu:fsd1} is identical to the stochastic dominance
assumption of \citet{blundell2007changes} that is used to tighten
the bounds on the wage distribution of the whole population in the
presence of sample selection. \citet{okumura2014concave} generalize
this stochastic dominance assumption to the case where $D$ is not
binary. Since Assumption \ref{assu:fsd1} implies that $E[Y_{d}|D=1,X=x]\geq E[Y_{d}|D=0,X=x]$,
this assumption is a sufficient condition for the monotone treatment
selection (MTS) assumption in \citet{manski2000monotone}.

Assumption \ref{assu:fsd1} is consistent with some economic theories
in empirical studies. As mentioned earlier, \citet{blundell2007changes}
utilize this assumption in conjunction with positive selection of
labor force participation from the standard labor supply model that
wage and probability of labor force participation are positively correlated.
As another example, we can consider sorting models for return to education,
where potential employees signal their ability to employers by using
their education level. Considering return to college education (i.e.,
$Y$ is wage, and $D$ is college entrance), it is likely that the
more capable people are, the more likely it is for them to complete
a college education. As a result, we may anticipate that people with
college degrees are more likely to have higher learning ability than
those who did not complete a college education (e.g., \citet{bedard2001human}).
Many studies in the labor economics literature consider learning ability
as an important factor affecting wage (\citet{lang1986human}), and
thus, we can assume that people with college degrees have higher wages
than those without degrees. As a consequence, the sorting hypothesis
is consistent with Assumption \ref{assu:fsd1}.

The following theorem gives identification results for the conditional
distributions of the potential outcomes under Assumption \ref{assu:fsd1}.

\begin{theorem}\label{thm:marginal_ID_FSD1} Let $x\in\mathcal{X}$
be fixed. Suppose that Assumption \ref{assu:fsd1} holds. For a given
$y\in\mathbb{R}$, define
\begin{eqnarray*}
F_{1|X}^{L,FSD1}(y|x) & \equiv & \Pr(Y\leq y|D=1,X=x),\\
F_{1|X}^{U,FSD1}(y|x) & \equiv & \Pr(Y\leq y|D=1,X=x)\Pr(D=1|X=x)+\Pr(D=0|X=x),\\
F_{0|X}^{L,FSD1}(y|x) & \equiv & \Pr(Y\leq y|D=0,X=x)\Pr(D=0|X=x),\\
F_{0|X}^{U,FSD1}(y|x) & \equiv & \Pr(Y\leq y|D=0,X=x).
\end{eqnarray*}
Then,
\begin{eqnarray}
F_{1|X}(y|x) & \in & \left[F_{1|X}^{L,FSD1}(y|x),F_{1|X}^{U,FSD1}(y|x)\right],\label{eq:fsd11}\\
F_{0|X}(y|x) & \in & \left[F_{0|X}^{L,FSD1}(y|x),F_{0|X}^{U,FSD1}(y|x)\right].\label{eq:fsd12}
\end{eqnarray}
 \end{theorem}

Comparing the bounds on the conditional distributions of the potential
outcomes in Theorem \ref{thm:marginal_ID_FSD1} with those in equation
(\ref{eq:manski_bound}), we can find that the upper bound on $F_{1|X}(y|x)$
and lower bound on $F_{0|X}(y|x)$ in Theorem \ref{thm:marginal_ID_FSD1}
are identical to their counterparts in (\ref{eq:manski_bound}). Since
Assumption \ref{assu:fsd1} designates only one direction of the monotonicity
of the distribution functions, it is impossible to improve the lower
bound on $F_{0|X}(y|x)$ and the upper bound on $F_{1|X}(y|x)$. Nevertheless,
the bounds on the conditional distributions of the potential outcomes
provided in Theorem \ref{thm:marginal_ID_FSD1} are narrower than
those in (\ref{eq:manski_bound}).

We now introduce another stochastic dominance assumption. The following
stochastic dominance assumption may be regarded as an assumption corresponding
to the monotone treatment response (MTR) assumption in \citet{manski1997monotone}.

\begin{assumption}\label{assu:fsd2} Let $x\in\mathcal{X}$ be given.
For all $d\in\{0,1\}$ and $y\in\mathbb{R}$, $F_{1|d,X}(y|x)\leq F_{0|d,X}(y|x)$.
\end{assumption}

Note that Assumption \ref{assu:fsd2} implies that $\mathbb{E}[Y_{1}|X=x]\geq\mathbb{E}[Y_{0}|X=x]$.
To see this, recall that $F_{1|X}(y|x)=F_{1|1,X}(y|x)p_{0}(x)+F_{1|0,X}(y|x)(1-p_{0}(x))$.
Under Assumption \ref{assu:fsd2}, we have $F_{1|1,X}(y|x)\leq F_{0|1,X}(y|x)$
and $F_{1|0,X}(y|x)\leq F_{0|0,X}(y|x)$, and therefore, we have $F_{1|X}(y|x)\leq F_{0|X}(y|x)$
and $\mathbb{E}[Y_{1}|X=x]\geq\mathbb{E}[Y_{0}|X=x]$.

Assumption \ref{assu:fsd2} can be regarded as a distributional generalization
of, but is weaker than, the MTR assumption of \citet{manski1997monotone}.
This assumption is also applicable to many empirical studies as it
is compatible with some economic theories. Recalling the example of
return to college education, Assumption \ref{assu:fsd2} may be consistent
with the human capital theory in which education increases productivity
through a human capital production function and return to education
reflects the increased productivity (e.g., \citet{mincer1974schooling}).
For example, Assumption \ref{assu:fsd2} can be translated into that
people who have completed their college degrees are likely to be paid
higher wages than when they have not, and this argument can be justified
by the human capital theory in the labor economics.

It is worth noting that Assumption \ref{assu:fsd2} is a special case
of conditional negative quadrant dependence, which is a dependence
concept considered by \citet{lehmann1966some}. \citet{kim2018partial}
proposes a concept of conditional quadrant dependence, as well as
a negative stochastic dominance, in the presence instrumental variables.
He provides tighter bounds on the distribution of treatment effects
than bounds using (\ref{eq:manski_bound}) under these assumptions
with the existence of an instrumental variable. A difference
between \citet{kim2018partial} and this paper is that this paper
does not impose additional assumptions on the data generating process,
such as monotonicity in structural functions for $Y_{1}$ and/or $Y_{0}$
and the availability of instrumental variables. Furthermore, this
paper focuses on estimation of the bounds on conditional distribution
of treatment effects, whereas \citet{kim2018partial} focuses on identification
of the unconditional distribution of treatment effects.

The following theorem shows that we can tighten the bounds on the
conditional distributions of $Y_{1}$ and $Y_{0}$ given $X$ under
Assumption \ref{assu:fsd2}:

\begin{theorem}\label{thm:marginal_ID_FSD2} Let $x\in\mathcal{X}$
be fixed. Suppose that Assumption \ref{assu:fsd2} holds. For a given
$y\in\mathbb{R}$, define
\begin{eqnarray*}
F_{1|X}^{L,FSD2}(y|x) & = & \Pr(Y\leq y|D=1,X=x)\Pr(D=1|X=x),\\
F_{1|X}^{U,FSD2}(y|x) & = & \Pr(Y\leq y|X=x),\\
F_{0|X}^{L,FSD2}(y|x) & = & F_{1|X}^{U,FSD2}(y|x),\\
F_{0|X}^{U,FSD2}(y|x) & = & \Pr(Y\leq y|D=0,X=x)\Pr(D=0|X=x)+\Pr(D=1|X=x).
\end{eqnarray*}
Then,
\begin{eqnarray}
F_{1|X}(y|x) & \in & \left[F_{1|X}^{L,FSD2}(y|x),F_{1|X}^{U,FSD2}(y|x)\right],\label{eq:fsd21}\\
F_{0|X}(y|x) & \in & \left[F_{0|X}^{L,FSD2}(y|x),F_{0|X}^{U,FSD2}(y|x)\right].\label{eq:fsd22}
\end{eqnarray}
\end{theorem}

As Assumption \ref{assu:fsd1} is not enough to improve $F_{1|X}^{U}(y|x)$
and $F_{0|X}^{L}(y|x)$ in (\ref{eq:manski_bound}), we can see that
the lower bound on $F_{1|X}(y|x)$ and the upper bound on $F_{0|X}(y|x)$
in Theorem \ref{thm:marginal_ID_FSD2} remain the same as those in
(\ref{eq:manski_bound}).

If both Assumptions \ref{assu:fsd1} and \ref{assu:fsd2} hold, one can further improve the bounds on the conditional
distributions of the potential outcomes. This result is summarized
in the following corollary:

\begin{corollary}\label{coro:marginal_ID_FSD12} Let $x\in\mathcal{X}$
be fixed. Suppose that Assumptions \ref{assu:fsd1} and \ref{assu:fsd2}
hold. For a given $y\in\mathbb{R}$,
\begin{eqnarray*}
F_{1|X}(y|x) & \in & \left[F_{1|X}^{L,FSD1}(y|x),F_{1|X}^{U,FSD2}(y|x)\right],\\
F_{0|X}(y|x) & \in & \left[F_{0|X}^{L,FSD2}(y|x),F_{0|X}^{U,FSD1}(y|x)\right].
\end{eqnarray*}
\end{corollary}

It is straightforward to see that the identified sets for conditional
distribution functions of $Y_{1}$ and $Y_{0}$ on $X=x$ presented
in Corollary \ref{coro:marginal_ID_FSD12} are connected and that
the intersection of these sets is the boundary of each set, that is
$F_{1|X}^{U,FSD2}(y|x)=F_{0|X}^{L,FSD2}(y|x)$ for all $y\in\mathbb{R}$.

It is worth mentioning that the stochastic dominance relationships
in Assumptions \ref{assu:fsd1} and \ref{assu:fsd2} can be reversed,
depending on the empirical context. For example, one can impose a
condition that for all $d\in\{0,1\}$ and $y\in\mathbb{R}$, $F_{d|1,X}(y|x)\geq F_{d|0,X}(y|x)$
instead of Assumption \ref{assu:fsd1}. Similarly, one can consider
a stochastic dominance relationship that for all $d\in\{0,1\}$ and
$y\in\mathbb{R}$, $F_{1|d,X}(y|x)\geq F_{0|d,X}(y|x)$ instead of
Assumption \ref{assu:fsd2}. These stochastic dominance relationships
can be motivated by, for example, the empirical study of \citet{Angrist1990},
who considers the long-term effect of veteran status on earnings.
As pointed out by \citet{Angrist1990}, it would be likely that men
with fewer civilian opportunities served in the armed force, and this
may yield a negative selection in the sense that the potential earnings
of veterans are likely to be smaller than those of non-veterans (i.e.,
$F_{d|1,X}(y|x)\geq F_{d|0,X}(y|x)$). In addition, since military
service may hinder human capital accumulation, the potential earnings
they would have received when they served in the armed force are likely
to be smaller than the potential earnings they would have received
when they did not (i.e., $F_{1|d,X}(y|x)\geq F_{0|d,X}(y|x)$).

There are other ways to tighten the bounds on the distribution of
treatment effects. For example, one can utilize support restrictions
(e.g., \citet{kim2018identifying}), consider a weaker condition than
the conditional independence, such as $c$-independence proposed by
\citet{masten2018identification}, or impose some dependence structure
as in \citet{frandsen2021partial}. In addition, although we do not
assume the availability of (monotone) instrumental variables in this
paper, one can consider the monotone instrumental variable (MIV) assumption
when there exists such a variable, as in \citet{manski2000monotone}
and \citet{kim2018partial}. A related but stronger assumption on
instrumental variables is the stochastically MIV (SMIV) assumption
considered by \citet{mourifie2020sharp}. The SMIV assumption imposes
a stochastic dominance ordering on the joint distribution of $Y_{0}$
and $Y_{1}$ across values of an instrumental variable, and the conditional
joint distribution of the potential outcomes given a value of the
instrumental variable, say $z$, (first-order) stochastically dominates
the conditional joint distribution given a value of the instrumental
variable that is smaller than $z$. These restrictions help tighten
the bounds on the conditional distributions of the potential outcomes
and/or the bounds on $F_{\Delta|X}$, but estimation and (uniform)
inference for such bounds is not trivial. Therefore, we leave it for
future study.

\section{\label{sec:Estimation} Estimation of Bounds and Confidence Bands}

\subsection{Estimation and Bootstrap Procedure }

We consider estimation of the bounds and construction of confidence
bands for them in this section. To this end, we assume that the conditional
distributions of the potential outcomes are identified as follows:
for all $y\in\mathbb{R}$ and $x\in\mathcal{X}$,
\begin{equation}
\begin{array}{cc}
F_{1|X}(y|x) & \in\left[LB_{1|X}(y|x),UB_{1|X}(y|x)\right],\\
F_{0|X}(y|x) & \in\left[LB_{0|X}(y|x),UB_{0|X}(y|x)\right].
\end{array}\label{eq:est_bound}
\end{equation}
 We denote nonparametric estimators of $LB_{1|X}$, $UB_{1|X}$, $LB_{0|X}$,
and $UB_{0|X}$ by $\hat{LB}_{1|X,n}$, $\hat{UB}_{1|X,n}$, $\hat{LB}_{0|X,n}$,
and $\hat{UB}_{0|X,n}$, respectively. In this paper, we consider
kernel-type estimators of the bounds. Let
\begin{align*}
\mathbb{F}_{\mathbf{Y}|X}((y_{1l},y_{1u},y_{0l},y_{0u})|x) & \equiv\left(LB_{1|X}(y_{1l}|x),UB_{1|X}(y_{1u}|x),LB_{0|X}(y_{0l}|x),UB_{0|X}(y_{0u}|x)\right)^{t},\\
\hat{\mathbb{F}}_{\mathbf{Y}|X,n}((y_{1l},y_{1u},y_{0l},y_{0u})|x) & \equiv\left(\hat{LB}_{1|X,n}(y_{1l}|x),\hat{UB}_{1|X,n}(y_{1u}|x),\hat{LB}_{0|X,n}(y_{0l}|x),\hat{UB}_{0|X,n}(y_{0u}|x)\right)^{t}.
\end{align*}
 Suppose that there exists a positive real sequence $\left(r_{n}\right)_{n}$
such that $r_{n}\rightarrow\infty$ and that
\[
r_{n}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)-\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\Rightarrow\mathbb{G}\left(\cdot\right)\text{ in }\left(l^{\infty}(\mathcal{Y})\right)^{4},
\]
 where $\mathbb{G}(\cdot)$ is a four-dimensional centered Gaussian
process. This weak convergence result can be derived under a set of
mild regularity conditions when using standard nonparametric estimators.
When we use kernel-type estimators, we have $r_{n}=\sqrt{nh_{n}^{d_{x}}}$,
where $h_{n}$ is a bandwidth.

Define
\begin{align*}
\Pi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}\right)(y,\delta|x) & \equiv LB_{1|X}(y|x)-UB_{0|X}(y-\delta|x),\\
\Pi_{U}\left(\mathbb{F}_{\mathbf{Y}|X}\right)(y,\delta|x) & \equiv UB_{1|X}(y|x)-LB_{0|X}(y-\delta|x),
\end{align*}
 and
\begin{align}
\phi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}\right) & \equiv\sup_{\delta\in\mathcal{D}}\Big|\sup_{y\in\mathcal{Y}}\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)-F_{\Delta|X}^{L,0}(\delta|x)\Big|=\sup_{\delta}\Big|F_{\Delta|X}^{L}(\delta|x)-F_{\Delta|X}^{L,0}(\delta|x)\Big|,\label{eq:lower_functional}\\
\phi_{U}\left(\mathbb{F}_{\mathbf{Y}|X}\right) & \sup_{\delta\in\mathcal{D}}\Big|\inf_{y\in\mathcal{Y}}\{\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1\}-F_{\Delta|X}^{U,0}(\delta|x)\Big|=\sup_{\delta}\Big|F_{\Delta|X}^{U}(\delta|x)-F_{\Delta|X}^{U,0}(\delta|x)\Big|,\label{eq:upper_functional}
\end{align}
 where $F_{\Delta|X}^{L,0}$ and $F_{\Delta|X}^{U,0}$ are some fixed
bounds that are true under the hypotheses $F_{\Delta|X}^{L}(\cdot|x)=F_{\Delta|X}^{L,0}(\cdot|x)$
and $F_{\Delta|X}^{U}(\cdot|x)=F_{\Delta|X}^{U,0}(\cdot|x)$. The
functionals in (\ref{eq:lower_functional}) and (\ref{eq:upper_functional})
can be considered as the Kolmogorov-Smirnov (KS) type statistics,
and they are useful to construct two-sided uniform confidence bands
for the lower and upper bounds on the conditional distribution of
treatment effects.\footnote{One can consider one-sided test statistics to construct one-sided
uniform confidence bands. However, we do not examine such test statistics
in this paper as the main purpose is to obtain uniform confidence
sets for the identified region and the two-sided uniform confidence
bands are enough to provide confidence sets for the identified region. }

Let $\mathbb{H}\equiv(h_{1},h_{2},h_{3},h_{4})^{t}$ be a four-dimensional
vector-valued function and $f$ be a real-valued function. We denote
the Hadamard directional derivatives of $\phi_{L}$ and $\phi_{U}$
at $f$ in direction $\mathbb{H}$ by $\phi_{L}^{'}(f;\mathbb{H})$
and $\phi_{U}^{'}(f;\mathbb{H})$, respectively. We also let $\widehat{\phi_{L}^{'}}\left(f;\mathbb{H},a_{n}\right)$
and $\widehat{\phi_{U}^{'}}\left(f;\mathbb{H},a_{n}\right)$ denote
estimators of $\phi_{L}^{'}(f;\mathbb{H})$ and $\phi_{U}^{'}(f;\mathbb{H})$,
respectively, where $(a_{n})$ is a positive real sequence decreasing
to zero. We provide the forms of the Hadamard directional derivatives
and their estimators in Appendix \ref{sec:HD-Derivatives}. The main
goal is to establish the limiting distributions of $r_{n}\left(\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}\right)-\phi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}\right)\right)$
and $r_{n}\left(\phi_{U}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}\right)-\phi_{U}\left(\mathbb{F}_{\mathbf{Y}|X}\right)\right)$.
The limiting distributions allow one to construct confidence bands
for the bounds on the conditional distribution of treatment effects.

As will be shown below, the limiting distributions are nonstandard.
We propose to use a multiplier bootstrap to mimic the asymptotic distributions
of $r_{n}\left(\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}\right)-\phi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}\right)\right)$
and $r_{n}\left(\phi_{U}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}\right)-\phi_{U}\left(\mathbb{F}_{\mathbf{Y}|X}\right)\right)$.
Specifically, let
\[
\hat{\Psi}_{i}((y_{1l},y_{1u},y_{0l},y_{0u})|x)\equiv\left(\hat{\psi}_{1l,i|X,n}(y_{1l}|x),\hat{\psi}_{1u,i|X,n}(y_{1u}|x),\hat{\psi}_{0l,i|X,n}(y_{0l}|x),\hat{\psi}_{0u,i|X,n}(y_{0u}|x)\right)^{t}
\]
 be an estimated influence function for the $i$-th observation of
$\hat{\mathbb{F}}_{\mathbf{Y}|X,n}((y_{1l},y_{1u},y_{0l},y_{0u})|x)$
and $\{B_{i}:i=1,2,...n\}$ be $n$ random draws from $B$ that is
independent of the data and satisfies some moment conditions (cf.
Assumption \ref{assu:boot_weight}). Define
\begin{align*}
 & \hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}((y_{1l},y_{1u},y_{0l},y_{0u})|x)\\
\equiv & \left(\sum_{i}B_{i}\hat{\psi}_{1l,i|X,n}(y_{1l}|x),\sum_{i}B_{i}\hat{\psi}_{1u,i|X,n}(y_{1u}|x),\sum_{i}B_{i}\hat{\psi}_{0l,i|X,n}(y_{0l}|x),\sum_{i}B_{i}\hat{\psi}_{0u,i|X,n}(y_{0u}|x)\right)^{t}.
\end{align*}


Then, one can implement the following bootstrap procedure to obtain
confidence bands for $F_{\Delta|X}^{L}(\cdot|x)$ and $F_{\Delta|X}^{U}(\cdot|x)$.

\paragraph{Bootstrap procedure }
\begin{enumerate}
\item Estimate the conditional distribution functions of the potential outcomes
or bounds on them, and construct
\begin{align*}
\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x) & =\hat{LB}_{1|X,n}(y|x)-\hat{UB}_{0|X,n}(y-\delta|x),\\
\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x) & =\hat{UB}_{1|X,n}(y|x)-\hat{LB}_{0|X,n}(y-\delta|x).
\end{align*}

\item Repeat the following procedure for $m$ times, where $m$ denotes
the number of bootstrap iterations.
\begin{enumerate}
\item Generate the bootstrap weights $\{B_{i}:i=1,2,...n\}$ from the random
variable $B$ in Assumption \ref{assu:boot_weight}.
\item Using the estimated conditional distribution functions of the potential
outcomes or bounds on them, compute $r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x)$.
\item For each iteration $b=1,2,...,m$, compute
\begin{align*}
 & \widehat{\phi_{L}^{'}}^{(b)}\left(\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x);r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right),\\
 & \widehat{\phi_{U}^{'}}^{(b)}\left(\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x)+1;r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right).
\end{align*}
\end{enumerate}
\item For a given significance level $\alpha\in(0,1)$, find the $(1-\frac{\alpha}{2})$-th
quantiles of
\begin{align*}
 & \left\{ \widehat{\phi_{L}^{'}}^{(b)}\left(\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x);r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right):b=1,2,...,m\right\} ,\\
 & \left\{ \widehat{\phi_{U}^{'}}^{(b)}\left(\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x)+1;r_{n}\hat{\mathbb{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right):b=1,2,...m\right\} .
\end{align*}
We denote the quantiles by $c_{1-\frac{\alpha}{2}}^{L}$ and $c_{1-\frac{\alpha}{2}}^{U}$,
respectively.
\item Let
\begin{align*}
\tilde{CI}\left(F_{\Delta|X}^{L}(\delta;x),1-\alpha\right) & \equiv\left[\hat{F}_{\Delta|X,n}^{L}(\delta|x)-c_{1-\frac{\alpha}{2}}^{L},\hat{F}_{\Delta|X,n}^{L}(\delta|x)+c_{1-\frac{\alpha}{2}}^{L}\right],\\
\tilde{CI}\left(F_{\Delta|X}^{U}(\delta;x),1-\alpha\right) & \equiv\left[\hat{F}_{\Delta|X,n}^{U}(\delta|x)-c_{1-\frac{\alpha}{2}}^{U},\hat{F}_{\Delta|X,n}^{U}(\delta|x)+c_{1-\frac{\alpha}{2}}^{U}\right].
\end{align*}
 The $(1-\alpha)\times100\%$ confidence bands of $F_{\Delta|X}^{L}(\cdot|x)$
and $F_{\Delta|X}^{U}(\cdot|x)$ can be constructed as
\begin{align*}
CI\left(F_{\Delta|X}^{L}(\delta;x),1-\alpha\right) & \equiv\left[\max\left\{ \hat{F}_{\Delta|X,n}^{L}(\delta|x)-c_{1-\frac{\alpha}{2}}^{L},0\right\} ,\min\left\{ \hat{F}_{\Delta|X,n}^{L}(\delta|x)+c_{1-\frac{\alpha}{2}}^{L},1\right\} \right],\\
CI\left(F_{\Delta|X}^{U}(\delta;x),1-\alpha\right) & \equiv\left[\max\left\{ \hat{F}_{\Delta|X,n}^{U}(\delta|x)-c_{1-\frac{\alpha}{2}}^{U},0\right\} ,\min\left\{ \hat{F}_{\Delta|X,n}^{U}(\delta|x)+c_{1-\frac{\alpha}{2}}^{U},1\right\} \right],
\end{align*}
 respectively.
\end{enumerate}
The difference between $\tilde{CI}$ and $CI$ in the above bootstrap
procedure is that $CI$'s are confidence bands obtained to impose
the logical bounds when necessary. \citet{chen2021shape} show that
one advantage of applying such operators to the original confidence
bands is that the resulting confidence bands have greater coverage
than the original confidence bands. In addition, this procedure is
very simple and easy to implement in practice.

The above bootstrap procedure allows us to construct the pointwise
confidence set of the identified set, where the identified set is
$[F_{\Delta|X}^{L}(\delta|x),F_{\Delta|X}^{U}(\delta|x)]$ for given
$\delta\in Supp(\Delta|X=x)$. Specifically, we can show that
\[
\underset{n\rightarrow\infty}{\lim\inf}\Pr\left(F_{\Delta|X}(\delta|x)\in\left[\hat{F}_{\Delta|X,n}^{L}(\delta|x)-c_{1-\frac{\alpha}{2}}^{L},\hat{F}_{\Delta|X,n}^{U}(\delta|x)+c_{1-\frac{\alpha}{2}}^{U}\right]\right)\geq1-\alpha,
\]
 and this can be used as a $(1-\alpha)\times100\%$ confidence set
for the identified set. This confidence set does not require the uniqueness
of the infimum and supremum in the Makarov bounds. As pointed out
by \citet{firpo2021uniform}, however, this confidence set is likely
to be conservative.

\subsection{Nonparametric Estimators of the Bounds on Conditional Distributions
of Potential Outcomes }

We propose to use kernel-type estimators of bounds on the conditional
distributions of the potential outcomes in (\ref{eq:est_bound}).
Let $K(\cdot):\mathbb{R}^{d_{x}}\rightarrow\mathbb{R}$ be a $d_{x}$-dimensional
kernel function and $h_{n}$ be a bandwidth such that $h_{n}\rightarrow0$,
$nh_{n}^{d_{x}}\rightarrow\infty$, and $nh_{n}^{d_{x}+4}\rightarrow0$
as $n\rightarrow\infty$. We start with the case where Assumptions
\ref{assu:unconfounded} and \ref{assu:bound_propensity} hold. In
this case, the conditional distributions of the potential outcomes
are point identified (i.e., $LB_{j|X}(y|x)=UB_{j|X}(y|x)=F_{j|X}(y|x)$
for each $j\in\{0,1\}$), and we can use
\begin{align*}
\hat{F}_{1|X,n}(y_{1}|x) & \equiv\frac{\sum_{i}^{n}\mathbf{1}(Y_{i}\leq y_{1})D_{i}K(\frac{X_{i}-x}{h_{n}})}{\sum_{i}^{n}D_{i}K(\frac{X_{i}-x}{h_{n}})},\\
\hat{F}_{0|X,n}(y_{0}|x) & \equiv\frac{\sum_{i}^{n}\mathbf{1}(Y_{i}\leq y_{0})(1-D_{i})K(\frac{X_{i}-x}{h_{n}})}{\sum_{i}^{n}(1-D_{i})K(\frac{X_{i}-x}{h_{n}})}
\end{align*}
 as estimators of $F_{1|X}(y_{1}|x)$ and $F_{0|X}(y_{0}|x)$, respectively.
We then define
\begin{equation}
\hat{\mathbb{F}}_{\mathbf{Y}|X,n}((y_{1l},y_{1u},y_{0l},y_{0u})|x)\equiv\left(\hat{F}_{1|X,n}(y_{1l}|x),\hat{F}_{1|X,n}(y_{1u}|x),\hat{F}_{0|X,n}(y_{0l}|x),\hat{F}_{0|X,n}(y_{0u}|x)\right)^{t}.\label{eq:kernel_unconfounded}
\end{equation}
 Note that the kernel function and bandwidth used to estimate $F_{1|X}$
can be different from those used to estimate $F_{0|X}$. In addition,
we have $\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X,n})(y,\delta|x)=\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X,n})(y,\delta|x)=F_{1|X}(y|x)-F_{0|X}(y-\delta|x)$
and can use the bootstrap procedure with $r_{n}=\sqrt{nh_{n}^{d_{x}}}$
under a set of regularity conditions.

We now consider the case where the treatment is endogenous and Assumptions \ref{assu:fsd1} and \ref{assu:fsd2} hold. Recall that
$F_{d|d^{'}X}(y|x)=\Pr(Y_{d}\leq y|D=d^{'},X=x)$ for given $d,d^{'}\in\{0,1\}$,
and the conditional distribution of $Y$ given $X=x$ is denoted by
$F_{Y|X}(\cdot|x)$. By corollary \ref{coro:marginal_ID_FSD12}, we
can construct kernel estimators of these objects as
\begin{align*}
\hat{F}_{1|1X,n}(y_{1}|x) & \equiv\frac{\sum_{i}^{n}\mathbf{1}(Y_{i}\leq y_{1})D_{i}K(\frac{X_{i}-x}{h_{n}})}{\sum_{i}^{n}D_{i}K(\frac{X_{i}-x}{h_{n}})},\\
\hat{F}_{0|0X,n}(y_{0}|x) & \equiv\frac{\sum_{i}^{n}\mathbf{1}(Y_{i}\leq y_{0})(1-D_{i})K(\frac{X_{i}-x}{h_{n}})}{\sum_{i}^{n}(1-D_{i})K(\frac{X_{i}-x}{h_{n}})},\\
\hat{F}_{Y|X,n}(y|x) & \equiv\frac{\sum_{i}^{n}\mathbf{1}(Y_{i}\leq y)K(\frac{X_{i}-x}{h_{n}})}{\sum_{i}^{n}K(\frac{X_{i}-x}{h_{n}})},
\end{align*}
 and define
\begin{equation}
\hat{\mathbb{F}}_{\mathbf{Y}|X,n}((y_{1l},y_{1u},y_{0l},y_{0u})|x)\equiv\left(\hat{F}_{1|1X,n}(y_{1l}|x),\hat{F}_{Y|X,n}(y_{1u}|x),\hat{F}_{Y|X,n}(y_{0l}|x),\hat{F}_{0|0X,n}(y_{0u}|x)\right)^{t}.\label{eq:kernel_endo}
\end{equation}
  Then, we have
\begin{align*}
\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x) & =\hat{F}_{1|1X,n}(y|x)-\hat{F}_{0|0X,n}(y-\delta|x),\\
\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x) & =\hat{F}_{Y|X,n}(y|x)-\hat{F}_{Y|X,n}(y-\delta|x).
\end{align*}
We provide a description on how to estimate the influence functions
in Appendix \ref{sec:Influence}.

\section{\label{sec:Inference} Asymptotic Theory }

\subsection{Inference under the Unconfoundedness Assumption }

We first develop the asymptotic theory for the kernel estimators in
(\ref{eq:kernel_unconfounded}).

\begin{assumption}\label{assu:iid} $\{W_{i}\equiv(Y_{i},D_{i},X_{i}^{t})^{t}:i=1,2,...,n\}$
is a random sample. \end{assumption}

\begin{assumption}\label{assu:dist_X} (i) The support of $X$, $\mathcal{X}$,
is a compact subset of $\mathbb{R}^{d_{x}}$; (ii) the distribution
of $X$ admits its density $f_{X}(\cdot)$ on $\mathcal{X}$ such
that $0<\inf_{x\in\mathcal{X}}f_{X}(x)<\sup_{x\in\mathcal{X}}f_{X}(x)<\infty$.
The density function $f_{X}(\cdot)$ is twice continuously differentiable
and $\sup_{x\in\mathcal{X}}|f_{X}^{(1)}(x)|$ and $\sup_{x\in\mathcal{X}}|f_{X}^{(2)}(x)|$
are bounded. \end{assumption}

\begin{assumption}\label{assu:smoothness} (i) The propensity score
function $p_{0}(x)$ is twice continuously differentiable and all
derivatives are uniformly bounded over $\mathcal{X}$; (ii) For each
$j\in\{0,1\}$, $F_{j|X}(y|x)$ is twice continuously differentiable
with respect to $x$ and all derivatives are uniformly bounded. \end{assumption}

\begin{assumption}\label{assu:kernel} The kernel function $K(\cdot)$
is a product of a univariate bounded kernel function $k(\cdot):\mathbb{R}\rightarrow\mathbb{R}_{+}$
such that $\int k(u)du=1$, $\int uk(u)du=0$, and $\int u^{2}k(u)du\equiv k_{2}<\infty$.
The support of the univariate kernel function is compact. \end{assumption}

\begin{assumption}\label{assu:bandwidth} The bandwidth $h_{n}$
satisfies the following conditions: (i) $h_{n}\rightarrow0$; (ii)
$nh_{n}^{d_{x}}\rightarrow\infty$; (iii) $nh_{n}^{d_{x}+4}\rightarrow0$
as $n\rightarrow\infty$. \end{assumption}

Assumption \ref{assu:iid} is an i.i.d assumption on the sample. This
can be relaxed at a cost of more complicated proofs of the theoretical
results. Assumption \ref{assu:dist_X} imposes some degree of smoothness
on the distribution of $X$. This assumption is standard in the literature
on kernel estimation. Assumption \ref{assu:smoothness} requires that
the conditional distribution functions of $Y_{1}$ and $Y_{0}$ and
the propensity score functions be smooth enough. This condition, together
with Assumptions \ref{assu:kernel} and \ref{assu:bandwidth}, allows
to eliminate the bias of estimators of conditional distribution functions
of $Y_{1}$ and $Y_{0}$. Assumption \ref{assu:kernel} is also standard
in the literature, and there are several kernel functions that satisfy
this assumption (e.g., Epanechnikov, biweight, triweight kernels).
Assumption \ref{assu:bandwidth} restricts the rate of bandwidth.
The last condition $nh_{n}^{d_{x}+4}\rightarrow0$ is required to
eliminate the asymptotic bias of the kernel estimators. Note that
when the dimension of $X$ is large, one can use a higher-order kernel
to handle the bias term of kernel estimators.

We first establish the weak convergence of the kernel estimators under
the unconfoundedness assumption. For given $x\in\mathcal{X}$, define
$\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\equiv\left(F_{1|X}(\cdot|x),F_{1|X}(\cdot|x),F_{0|X}(\cdot|x),F_{0|X}(\cdot|x)\right)^{t}$
and $\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\equiv\left(\hat{F}_{1|X,n}(\cdot|x),\hat{F}_{1|X,n}(\cdot|x),\hat{F}_{0|X,n}(\cdot|x),\hat{F}_{0|X,n}(\cdot|x)\right)^{t}$.

\begin{theorem}\label{thm:asymp_marginals} Suppose that Assumptions
\ref{assu:unconfounded} and \ref{assu:bound_propensity} hold. Let
$x\in int(\mathcal{X})$ be given. If Assumptions \ref{assu:iid}--\ref{assu:bandwidth}
hold, then,
\[
\sqrt{nh_{n}^{d_{x}}}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)-\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\Rightarrow\mathbb{G}_{x}\left(\cdot\right)\equiv\left(\mathbb{G}_{1,x}(\cdot),\mathbb{G}_{1,x}(\cdot),\mathbb{G}_{0,x}(\cdot),\mathbb{G}_{0,x}(\cdot)\right)^{t}\text{ in }\left(l^{\infty}(\mathcal{Y})\right)^{4},
\]
 where $\mathbb{G}_{1,x}(\cdot)$ and $\mathbb{G}_{0,x}(\cdot)$ are
mean zero Gaussian processes with covariance kernels $H_{1,x}(s,t)$
and $H_{0,x}(s,t)$, respectively, whose the forms are given in Appendix
\ref{sec:Proofs}, and $\mathbb{G}_{x}((y_{1l},y_{1u},y_{0l},y_{0u}))=\left(\mathbb{G}_{1,x}(y_{1l}),\mathbb{G}_{1,x}(y_{1u}),\mathbb{G}_{0,x}(y_{0l}),\mathbb{G}_{0,x}(y_{0u})\right)^{t}$.
\end{theorem}

Theorem \ref{thm:asymp_marginals} is useful not only for deriving
the asymptotic distribution of the estimated bounds on $F_{\Delta|X}(\cdot|x)$,
but also for conducting uniform inference for some policy-relevant
parameter (e.g., quantile treatment effects).

We now establish the asymptotic theory for confidence bands for the
identified set of $F_{\Delta|X}(\cdot|x)$. To this end, we consider
the following two hypotheses: $F_{\Delta|X}^{L}(\delta|x)=F_{\Delta|X}^{L,0}(\delta|x)$
and $F_{\Delta|X}^{U}(\delta|x)=F_{\Delta|X}^{U,0}(\delta|x)$, where
$F_{\Delta|X}^{L,0}(\delta|x)$ and $F_{\Delta|X}^{U,0}(\delta|x)$
are some fixed lower and upper bounds on the conditional distribution
of treatment effects, respectively.

The main challenge to constructing uniform confidence bands, however,
is that the mappings $\phi_{L}$ and $\phi_{U}$ are not Hadamard
differentiable. To overcome this difficulty, we use Theorem 3.2 in
\citet{firpo2021uniform} that shows that these functionals are Hadamard
directionally differentiable. Based on this result, we employ the
inference method of \citet{fang2019inference} to establish the asymptotic
theory.

Recall that $\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)=\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)=F_{1|X}(y|x)-F_{0|X}(y-\delta|x)$.
The following theorem establishes the limiting distributions of $\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)$
and $\phi_{U}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)$.

\begin{theorem}\label{thm:Asymptotic_Dist of Derivative} Suppose
that Assumptions \ref{assu:unconfounded} and \ref{assu:bound_propensity}
hold. Let $x\in int(\mathcal{X})$ be given. If Assumptions \ref{assu:iid}--\ref{assu:bandwidth}
hold, then,
\begin{equation}
\begin{aligned}\sqrt{nh_{n}^{d_{x}}}\left(\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)-\phi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\right) & \Rightarrow\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}(\cdot)\right)\ \text{in }l^{\infty}\left(\mathcal{Y}^{4}\right),\\
\sqrt{nh_{n}^{d_{x}}}\left(\phi_{U}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)-\phi_{U}\left(\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\right) & \Rightarrow\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}(\cdot)\right)\ \text{in }l^{\infty}\left(\mathcal{Y}^{4}\right),
\end{aligned}
\label{eq:te_asymptotic}
\end{equation}
 where and $\mathbb{G}_{x}(\cdot)$ is the Gaussian process defined
in Theorem \ref{thm:asymp_marginals}. The forms of $\phi_{L}^{'}$
and $\phi_{U}^{'}$ are provided in Appendix \ref{sec:HD-Derivatives}.
\end{theorem}

It is worth pointing out that the development of weak convergence
of the estimated bounds does not rely on some pointwise asymptotic
theory when establishing the asymptotic theory for the estimated Makarov
bounds. Instead, we first derive the weak convergence of the standard
kernel estimators for a fixed conditioning value and consider the
double-supremum as an operator to establish the weak convergence of
the estimated bounds, as in \citet{firpo2021uniform}. In doing so,
we can avoid imposing the uniqueness of $argsup$ and $arginf$ that
is needed for pointwise inference.

Theorem \ref{thm:Asymptotic_Dist of Derivative} is a direct consequence
of Theorem 2.1 of \citet{fang2019inference}. The limiting distribution
presented in Theorem \ref{thm:Asymptotic_Dist of Derivative} is non-standard,
and therefore we need to rely on some resampling method to mimic the
limiting distribution and conduct inference. Since the functionals
$\phi_{L}$ and $\phi_{U}$ are not Hadamard differentiable but only
Hadamard directionally differentiable, the standard bootstrap fails
(cf. Theorem 3.1 of \citet{fang2019inference}). To resolve this problem,
we utilize the result of bootstrap validity that was proposed by \citet{firpo2021uniform}.
The next assumption imposes conditions on the bootstrap weight $B$:

\begin{assumption}\label{assu:boot_weight} Let $B$ be a random
variable that is independent of the data $\mathcal{W}$ such that
$\mathbb{E}[B]=0$, $Var(B)=1$, and $\int_{0}^{\infty}\sqrt{\Pr(|B|>x)}dx<\infty$.
\end{assumption}

The last condition in Assumption \ref{assu:boot_weight} is satisfied
if $\mathbb{E}\left[|B|^{2+\epsilon}\right]<\infty$ for some $\epsilon>0$.
One can use a standard normal random variable as the bootstrap weight
$B$.

We employ the approach of \citet{fang2019inference} to approximate
the limiting distribution presented in Theorem \ref{thm:Asymptotic_Dist of Derivative},
which was also considered by \citet{firpo2021uniform}. Note that
it is required to consistently estimate the Hadamard directional derivatives
in (\ref{eq:phi_derivative}) and that the Hadamard directional derivatives
are defined in terms of the limit operator. To this end, we consider
a tuning sequence $(a_{n})$ that satisfies the conditions in the
following assumption:

\begin{assumption}\label{assu:est_der_tuning}Let $(a_{n})$ be a
sequence of positive real numbers such that $a_{n}\downarrow0$ and
$a_{n}\sqrt{nh_{n}^{d_{x}}}\rightarrow\infty$.\end{assumption}

Since the limiting distributions in Theorem \ref{thm:Asymptotic_Dist of Derivative}
are nonstandard, we use a bootstrap to mimic the limiting distributions.
The next theorem demonstrates that one can use the bootstrap procedure
in Section \ref{sec:Estimation} to approximate the asymptotic distributions
of $\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}(\cdot)\right)$
and $\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}(\cdot)\right)$.
Let $\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x)$ denote the vector
of simulated stochastic processes using estimated influence functions
in (\ref{eq:influence_exo}) in Appendix \ref{sec:Influence}.

\begin{theorem}\label{thm:bootstrap} Suppose that Assumptions \ref{assu:unconfounded}
and \ref{assu:bound_propensity} hold. Let $x\in int(\mathcal{X})$
be given and Assumptions \ref{assu:iid}--\ref{assu:bandwidth},
\ref{assu:boot_weight}, and \ref{assu:est_der_tuning} hold. Then,
we have
\begin{align*}
\widehat{\phi_{L}^{'}}\left(\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x);\sqrt{nh_{n}^{d_{x}}}\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right) & \Rightarrow\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}(\cdot)\right),\\
\widehat{\phi_{U}^{'}}\left(\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x)+1;\sqrt{nh_{n}^{d_{x}}}\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right) & \Rightarrow\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}(\cdot)\right),
\end{align*}
 in $l^{\infty}\left(\mathcal{Y}^{4}\right)$, conditional on data.
The forms of $\widehat{\phi_{L}^{'}}$ and $\widehat{\phi_{U}^{'}}$
are provided in Appendix \ref{sec:HD-Derivatives}. \end{theorem}

\subsection{Inference with an Endogenous Treatment}

We now develop the asymptotic theory when the treatment is endogenous.
Our asymptotic theory for an endogenous treatment focuses on the situation
where Assumptions \ref{assu:fsd1} and \ref{assu:fsd2} hold, and
thus, one can use the kernel estimators in Section \ref{sec:Estimation}.
To be concrete, define, for given $x\in\mathcal{X}$, $\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\equiv\left(F_{1|1X}(\cdot|x),F_{Y|X}(\cdot|x),F_{Y|X}(\cdot|x),F_{0|0X}(\cdot|x)\right)^{t}$
and $\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\equiv\left(\hat{F}_{1|1X,n}(\cdot|x),\hat{F}_{Y|X,n}(\cdot|x),\hat{F}_{Y|X,n}(\cdot|x),\hat{F}_{0|0X,n}(\cdot|x)\right)^{t}$,
where each component of $\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)$
is given in (\ref{eq:kernel_endo}). Let $\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X,n})(y,\delta|x)\equiv F_{1|1X}(y|x)-F_{0|0X}(y-\delta|x)$
and $\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X,n})(y,\delta|x)\equiv F_{Y|X}(y|x)-F_{Y|X}(y-\delta|x)$.

\begin{assumption}\label{assu:smooth_endo} $F_{1|1,X}(y|x)$, $F_{0|0,X}(y|x)$,
and $F_{Y|X}(y|x)$ are twice continuously differentiable with respect
to $x$ and all derivatives are uniformly bounded. \end{assumption}

\begin{theorem}\label{thm:endo_marginal_weak_limit} Suppose that
Assumptions \ref{assu:fsd1} and \ref{assu:fsd2} hold. Let $x\in int(\mathcal{X})$
be given and Assumptions \ref{assu:iid}, \ref{assu:dist_X}, \ref{assu:kernel},
\ref{assu:bandwidth}, and \ref{assu:smooth_endo} hold. Then,
\[
\sqrt{nh_{n}^{d_{x}}}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)-\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\Rightarrow\mathbb{G}_{x}^{e}(\cdot)\ \text{in }\left(l^{\infty}(\mathcal{Y})\right)^{4},
\]
 where $\mathbb{G}_{x}^{e}((y_{1l},y_{1u},y_{0l},y_{0u}))\equiv(\mathbb{G}_{1,x}^{e}(y_{1l}),\mathbb{G}_{Y,x}^{e}(y_{1u}),\mathbb{G}_{Y,x}^{e}(y_{0l}),\mathbb{G}_{0,x}^{e}(y_{0u}))^{t}$
and $\mathbb{G}_{1,x}^{e}$, $\mathbb{G}_{0,x}^{e}$, and $\mathbb{G}_{Y,x}^{e}$
are Gaussian processes with mean zero and covariance kernels being
$H_{1,x}^{e}(\cdot,\cdot)$, $H_{0,x}^{e}(\cdot,\cdot)$, and $H_{Y,x}^{e}(\cdot,\cdot)$,
respectively. The forms of the covariance kernels are given in Appendix
\ref{sec:Proofs}.

In addition,
\begin{equation}
\begin{aligned}\sqrt{nh_{n}^{d_{x}}}\left(\phi_{L}\left(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x)\right)-\phi_{L}\left(\mathbb{F}_{\mathbf{Y}|X}(\cdot|x)\right)\right) & \Rightarrow\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}^{e}(\cdot)\right),\\
\sqrt{nh_{n}^{d_{x}}}\left(\phi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n}(\cdot|x))-\phi_{U}^{e}(\mathbb{F}_{\mathbf{Y}|X}(\cdot|x))\right) & \Rightarrow\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}^{e}(\cdot)\right),
\end{aligned}
\label{eq:te_asymptotic_endo}
\end{equation}
 in $l^{\infty}\left(\mathcal{Y}^{4}\right)$. The forms of $\phi_{L}^{'}$
and $\phi_{U}^{'}$ are provided in Appendix \ref{sec:HD-Derivatives}.
\end{theorem}

We simulate the limiting distribution provided in Theorem \ref{thm:endo_marginal_weak_limit}
in a similar way to the exogenous treatment case. Let $\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x)$
denote the vector of simulated stochastic processes using estimated
influence functions in (\ref{eq:Influence_endo}) in Appendix \ref{sec:Influence}.
The following theorem shows the validity of the bootstrap procedure:

\begin{theorem}\label{thm:endo_bootstrap} Suppose that conditions
in Theorem \ref{thm:endo_marginal_weak_limit} are satisfied. In addition,
if Assumptions \ref{assu:boot_weight} and \ref{assu:est_der_tuning}
also hold, then, we have
\begin{align*}
\widehat{\phi_{L}^{'}}\left(\Pi_{L}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x);\sqrt{nh_{n}^{d_{x}}}\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right) & \Rightarrow\phi_{L}^{'}\left(\Pi_{L}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x);\mathbb{G}_{x}^{e}(\cdot)\right),\\
\widehat{\phi_{U}^{'}}\left(\Pi_{U}(\hat{\mathbb{F}}_{\mathbf{Y}|X,n})(y,\delta|x)+1;\sqrt{nh_{n}^{d_{x}}}\mathbb{\hat{F}}_{\mathbf{Y}|X,n}^{*}(\cdot|x),a_{n}\right) & \Rightarrow\phi_{U}^{'}\left(\Pi_{U}(\mathbb{F}_{\mathbf{Y}|X})(y,\delta|x)+1;\mathbb{G}_{x}^{e}(\cdot)\right),
\end{align*}
 in $l^{\infty}\left(\mathcal{Y}^{4}\right)$, conditional on data.
The forms of $\widehat{\phi_{L}^{'}}$ and $\widehat{\phi_{U}^{'}}$
are provided in Appendix \ref{sec:HD-Derivatives}. \end{theorem}

We provide several extensions of the main results in this section
in Appendix. We extend the results of \citet{abrevaya2015estimating}
who consider estimating conditional average treatment effects on a
subset of covariates to the conditional distribution of treatment
effects (Appendix \ref{sec:semipara}). This result may be practically
relevant when the number of covariates is large. We also discuss how
to adopt the results to conduct global hypothesis testing (Appendix
\ref{sec:hypothesis_testing}).

It is worth noting that \citet{firpo2019partial} provide formal definitions
of pointwise and uniform sharpness of bounds on distribution functions
and show that the Makarov bounds are not uniformly sharp. The inference
methods developed in this paper are based on pointwise sharp bounds
on the conditional distribution of treatment effects and are valid
uniformly over the support of treatment effects. Inference for uniformly
sharp bounds on the distribution of treatment effects is left as future
research.

\section{\label{sec:MC} Monte Carlo Simulation }

We conduct a set of simulations to investigate the performance of
the bootstrap in finite samples. To this end, the following DGP is
considered:
\begin{align*}
Y_{1} & =\mu_{1}+X\beta_{1}+(\phi_{1}+X\gamma_{1})U_{1},\\
Y_{0} & =\mu_{0}+X\beta_{0}+(\phi_{0}+X\gamma_{0})U_{0},\\
D & =\mathbf{1}(X\alpha\geq V),
\end{align*}
 where $X=2\tilde{X}-1$ with $\tilde{X}$ being an uniform a random
variable on $[0,1]$, $U\equiv(U_{1},U_{0})^{t}\sim N\left(\begin{pmatrix}0\\
0
\end{pmatrix},\begin{pmatrix}1 & 0\\
0 & 1
\end{pmatrix}\right)$, and $V\sim N(0,1)$. The conditional treatment effect on $X=x$
is defined as $(\mu_{1}-\mu_{0})+x(\beta_{1}-\beta_{0})+((\phi_{1}+X\gamma_{1})U_{1}-(\phi_{0}+X\gamma_{0})U_{0})$.
The parameter values are set as follows: $\beta_{1}=\gamma_{1}=1$,
$\beta_{0}=\gamma_{0}=0.9$, $\phi_{1}=\phi_{0}=1$, and $\alpha=1$.
We consider various values for parameter vector $(\mu_{1},\mu_{0})^{t}$
to investigate the performance of a KS test statistic under the null
and alternative hypotheses. Specifically, to investigate the performance
of the KS test under null hypotheses, we consider the following hypothesis:
\[
H_{0}:F_{\Delta|X}^{L}(\delta|x=0)=\left(2\cdot\Phi\left(\frac{\delta}{2}\right)-1\right)\cdot\mathbf{1}(\delta\geq0)\text{ for all }\delta,
\]
 where $\Phi(\cdot)$ is the standard normal distribution function.

The lower bound in the null hypothesis is one derived by \citet{frank1987best},
and the null hypothesis is true if and only if $\mu_{1}=\mu_{0}=0$.
The KS statistic for this null is constructed as follows:
\[
KS_{n}\equiv\sup_{\delta}\sqrt{nh_{n}}\Bigg|F_{\Delta|X}^{L}(\delta|x=0)-\left(2\cdot\Phi\left(\frac{\delta}{2}\right)-1\right)\cdot\mathbf{1}(\delta\geq0)\Bigg|.
\]
 To investigate the performance of the KS test under some alternative
hypotheses, we consider the cases where $(\mu_{1},\mu_{0})^{t}=(\mu,0)^{t}$
for $\mu\in\{-1,1\}$. The bandwidth is chosen to be $h_{n}=1.06\times s.d(X)\times n^{-1/6}$,
where $s.d(X)$ denotes the (sample) standard deviation of $X$.

It is worth emphasizing that the asymptotic distribution of the KS
test statistic is nonstandard and may not be uniformly valid with
respect to the DGP. For this reason, it is important to investigate
whether the KS statistic performs well with various choices for $(a_{n})$.
Specifically, we consider several rates at which $a_{n}$ grows to
the infinity: (i) $a_{n}=c\times\log\left(\log\left(nh_{n}\right)\right)/\sqrt{nh_{n}}$,
(ii) $a_{n}=c\times\sqrt{\log\left(nh_{n}\right)}/\sqrt{nh_{n}}$,
and (iii) $a_{n}=c\times\left(nh_{n}\right)^{1/6}/\sqrt{nh_{n}}$
for some $c>0$. These choices of $(a_{n})$ satisfy Assumption \ref{assu:est_der_tuning}.
We consider various values of $c$ ranging from $0.1$ to $0.5$ to
see whether the finite-sample performance of the KS test is sensitive
to the choice of $c$.\footnote{\citet{firpo2021uniform} suggest using $c=0.2$ with $a_{n}=c\times\log\left(\log\left(nh_{n}\right)\right)/\sqrt{nh_{n}}$
in our context. }

The sample size $n$ is set to be 500, and the number of bootstrap
iterations is 500. The bootstrap weight $B$ is drawn from the standard
normal distribution, and all simulation results are obtained from
500 iterations. The nominal level is set to be 0.05.

Table \ref{tab:sim_rp} presents the simulation results for rejection
probabilities. We find that the rejection probability under $H_{0}$
tends to decrease as $c$ increases, except for the case of $\sqrt{nh_{n}}a_{n}\propto\left(nh_{n}\right)^{1/6}$.\footnote{We also considered larger values of $c$ than $0.5$, but they are
not reported here. The simulation results with those values of $c$
suggest that the resulting confidence bands would be too conservative
(i.e., the rejection probabilities under the null hypothesis are relatively
smaller than the nominal rate). As a result, we do \textit{not }recommend
using a too large value of $c$, based on our simulation results.
It would be an interesting question how to choose $c$ or the rate
of $(a_{n})$ in a data-dependent way, but this is far beyond the
scope of this paper. We therefore leave this interesting and important
question for future research.} The KS statistic performs well in finite samples for various values
of $(a_{n})$ in the sense that the rejection probability under $H_{0}$
is close to the nominal probability and that the rejection probabilities
under $H_{1}$ are large in general.

\begin{table}[h]
\caption{\label{tab:sim_rp} Rejection Probabilities, $n=500$, $B=500$, $x=Q_{X}(0.5)$}

\bigskip{}

\begin{centering}
\begin{tabular}{cccccccccc}
\hline
$c$\textbackslash$a_{n}$ & \multicolumn{3}{c}{$c\cdot\log\left(\log\left(nh_{n}\right)\right)/\sqrt{nh_{n}}$} & \multicolumn{3}{c}{$c\sqrt{\log\left(nh_{n}\right)}/\sqrt{nh_{n}}$} & \multicolumn{3}{c}{$c\left(nh_{n}\right)^{1/6}/\sqrt{nh_{n}}$}\tabularnewline
\cline{2-10} \cline{3-10} \cline{4-10} \cline{5-10} \cline{6-10} \cline{7-10} \cline{8-10} \cline{9-10} \cline{10-10}
 & $\mu=0$ & $\mu=-1$ & $\mu_{1}=1$ & $\mu=0$ & $\mu=-1$ & $\mu_{1}=1$ & $\mu=0$ & $\mu=-1$ & $\mu_{1}=1$\tabularnewline
\hline
0.1 & 0.078 & 0.998 & 0.824 & 0.068 & 0.996 & 0.816 & 0.044 & 0.996 & 0.816\tabularnewline
0.2 & 0.066 & 0.986 & 0.806 & 0.060 & 0.992 & 0.782 & 0.064 & 0.992 & 0.782\tabularnewline
0.3 & 0.060 & 0.994 & 0.778 & 0.058 & 0.986 & 0.756 & 0.050 & 0.992 & 0.754\tabularnewline
0.4 & 0.058 & 0.994 & 0.734 & 0.044 & 0.992 & 0.748 & 0.056 & 0.988 & 0.746\tabularnewline
0.5 & 0.050 & 0.988 & 0.772 & 0.040 & 0.998 & 0.704 & 0.058 & 0.984 & 0.738\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\bigskip{}

Note: The nominal level is set to be 0.05. The case of $\mu=0$ is
where the null hypothesis is true.
\end{table}


\section{\label{sec:Empirical} Empirical Application: The Effect of 401(k)
Plans on Net Financial Assets}

In this section, we provide an empirical example to illustrate the
usefulness of the methods proposed in this paper. We revisit the empirical
question on the effect of participation in 401(k) plans on net financial
assets investigated by many studies in the literature (e.g., \citet{Abadie2003,Chernozhukov2004,Wuethrich2019,SantAnna2022}).
Our main goals in this empirical application are twofold. First, we
empirically show that the stochastic dominance assumptions (Assumptions
\ref{assu:fsd1} and \ref{assu:fsd2}) can provide considerable identifying
power. Second, we complement the existing results that there is substantial
heterogeneity in treatment effects across income groups by estimating
the bounds on the conditional distribution of treatment effects on
income levels without an instrumental variable. In doing so, we complement
the empirical results on the effect of 401(k) plans on net financial
assets documented in the literature by providing estimation results
on the distribution of the treatment effect. Since our focus is on
verifying the identifying power of the stochastic dominance assumptions
and potential heterogeneity in the treatment effect across different
subpopulations, we do not report confidence bands for clear illustration.

The U.S. introduced several tax-deferred retirement plans, including
401(k) plans and individual retirement accounts (IRAs), in the early
1980s. These retirement plans can be used as a way to accumulate individual
assets. Many papers in the literature have considered how those tax-deferred
retirement plans affect asset accumulation or savings. The main challenge
with identifying and estimating the causal effect of 401(k) participation
on assets is that participation in 401(k) plans is endogenously determined.
Furthermore, the effect of 401(k) plans on net financial assets is
heterogeneous across income levels, as shown by \citet{Chernozhukov2004}.
Motivated by the empirical results of \citet{Chernozhukov2004}, we
focus on the distribution of treatment effects conditional on an individual's
income. Moreover, while it is common to use the eligibility for 401(k)
as an instrumental variable to point identify some distributional
effects (e.g., quantile treatment effects), our identification and
estimation strategies do not rely on such an instrumental variable.

We use the data from \citet{SantAnna2022} for this empirical analysis.
The original dataset contains 9,910 households from the 1991 Survey
of Income and Program Participation. The dependent variable is the
amount of net financial assets measured in ten thousand dollars. We
exclude observations with a value of the dependent variable higher
than the 0.99 sample quantile and lower than the 0.01 sample quantile
of the net financial assets. This results in a sample of 9,712 households.
The treatment variable is a binary variable indicating whether a household
participates in 401(k) plans. To investigate the potential heterogeneity
across income levels, we use the income variable as the covariate
of interest. For stability of estimation, we standardize the original
variable of income and consider the sample mean and various quantiles
of income. Table \ref{tab:emp2_summary} reports the summary statistics
of the data.

\begin{table}[H]
\caption{\label{tab:emp2_summary} Summary Statistics of the Data}

\bigskip{}

\begin{centering}
\begin{tabular}{ccccccc}
\hline
Variables & Mean & Median & S.D. & Min & Max & Obs.\tabularnewline
\hline
Net financial assets & 1.4506 & 0.1499 & 3.1991 & -2.350 & 21.995 & 9,712\tabularnewline
Treatment & 0.26 & 0 & 0.4387 & 0 & 1 & 9,712\tabularnewline
Income & 3.661 & 3.120 & 2.379 & 0.003 & 19.299 & 9,712\tabularnewline
Age & 40.969 & 40 & 10.311 & 25 & 64 & 9,712\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\bigskip{}

Note: The net financial assets and income are measured in \$10,000.
\end{table}

We first discuss the validity of Assumptions \ref{assu:fsd1} and
\ref{assu:fsd2} in this empirical example. It is well known that
the preference for saving is heterogeneous in that some people have
a stronger preference for saving than others. This leads to the nonrandom
selection into participation in 401(k) plans or other tax-deferred
retirement plans (e.g., \citet{Chernozhukov2004}). Based on this
observation, it is likely that people who participate in 401(k) plans
would have a stronger preference for saving than those who do not
participate. Therefore, the net financial assets of people with a
strong preference for saving tend to be larger than those of people
with a weak preference for saving, suggesting that it may be plausible
to impose Assumption \ref{assu:fsd1} on the model.

Assumption \ref{assu:fsd2} is consistent with the empirical example
as most tax-deferred retirement plans, including 401(k) plans, by
themselves increase the amount of assets that an individual possess.
Therefore, regardless of whether an individual participates in 401(k)
plans or not, the potential net financial assets that one would have
had if she participated in 401(k) plans are likely to be larger than
those she would have had if she did not participate in 401(k). As
a result, it is plausible to impose Assumption \ref{assu:fsd2} on
the model.

For estimation, we set $h_{n}=1.06\times n^{-1/(5+d_{x})}$ and $a_{n}=0.2\times\log\left(\log\left(nh_{n}^{d_{x}}\right)\right)/\sqrt{nh_{n}^{d_{x}}}$
with $d_{x}=1$. We denote the $\tau$-th (sample) quantile of income
by $Q_{income}(\tau)$.

We note that when Assumptions \ref{assu:fsd1} and \ref{assu:fsd2}
are not imposed, the estimated bounds are the logical ones, regardless
of the conditioning value of the income.\footnote{By the logical bounds, we mean that the lower and upper bounds are
equal to 0 and 1, respectively, for all $\delta\in Supp(\Delta|X=x)$. } On the other hand, the estimated bounds under Assumptions \ref{assu:fsd1}
and \ref{assu:fsd2} are informative in the sense that they are not
the logical bounds. This indicates that the stochastic dominance assumptions
have considerable identifying power.

We find substantial heterogeneity in the distribution of treatment
effects across different values of the income. Figure \ref{fig:endo_treat}
compares the estimated bounds on the conditional distribution of treatment
effects at the 0.2 and 0.8 quantiles of the income. Specifically,
the star-marked lines are the bounds on the conditional distribution
of treatment effects given the 0.2 quantile of the income. The circle-marked
lines are the bounds on the conditional distribution of treatment
effects given the 0.8 quantile of income. We find that the lower bound
conditional on the 0.2 quantile of the income is larger than the lower
bound conditional on the 0.8 quantile of the income. Conversely, the
upper bound conditional on the 0.8 quantile of the income is larger
than the upper bound conditional on the 0.2 quantile of the income
for all $\delta\in[-24,24]$ (i.e., from -\$240,000 to \$240,000).
When considering the bounds on $\Pr\left(Y_{1}-Y_{0}\leq1|X=x\right)$
for $x\in\left\{ Q_{income}(0.2),Q_{income}(0.8)\right\} $, the estimation
results show that $\Pr\left(Y_{1}-Y_{0}\leq1|X=Q_{income}(0.2)\right)\in\left[0.5804,1\right]$
and that $\Pr\left(Y_{1}-Y_{0}\leq1|X=Q_{income}(0.8)\right)\in\left[0.0638,1\right]$.
These estimated bounds suggest that the proportion of individuals
who experience a positive treatment effect larger than \$10,000 among
those with the income being equal to the 0.2 sample quantile is at
most 41.96\%. The proportion among those with the income being equal
to the 0.8 sample quantile is at most 93.62\%. As a result, there
may be a possibility that the proportion of individuals who experience
a certain level of positive treatment effect is larger when considering
a higher level of income. This argument is in part consistent with
the finding of \citet{Chernozhukov2004} that the quantile treatment
effects of 401(k) plans on net financial wealth tend to increase as
the income level increases (see Figure 2 in \citet{Chernozhukov2004}).

Figure \ref{fig:comparison_all_quantiles} compares the estimated
bounds at a specific quantile level with those at the mean of income
under Assumptions \ref{assu:fsd1} and \ref{assu:fsd2}. When considering
a low quantile level, e.g., $\tau\in\{0.1,0.2,0.3\}$, we find that
the lower bound conditional on $X=Q_{income}(\tau)$ is larger than
that conditional on the mean of income. The upper bound conditional
on $X=Q_{income}(\tau)$ is smaller than that conditional on the mean
of income over the potential support of the treatment effect. However,
when considering a high quantile level, e.g., $\tau\in\{0.7,0.8,0.9\}$,
the estimation results are the opposite. These estimation results
indicate that the treatment effect of 401(k) plans on net financial
assets is likely to be heterogeneous across income levels, which is
consistent with the finding of \citet{Chernozhukov2004}.

\section{\label{sec:Conclusion} Conclusion }

This paper considers identification and estimation of bounds on the
conditional distribution of treatment effects. The conditional distribution
may provide evidence on potential heterogeneity in treatment effects
across subpopulations that are defined in terms of values of covariates,
and therefore, is of practical importance in many empirical studies.
We show that when the treatment is endogenously determined, one can
tighten the bounds by imposing stochastic dominance assumptions. These
assumptions are consistent with many economic theories and easy to
interpret, and the resulting bounds on the distribution of treatment
effects are easy to compute. We propose nonparametric estimators of
the bounds and establish the uniform asymptotic theory based on the
novel approach of \citet{fang2019inference} and \citet{firpo2021uniform}.
The asymptotic theory in this paper is useful for constructing uniform
confidence bands and conducting statistical tests for global hypotheses.
We then provide an empirical application of the methodology proposed
in this paper to illustrate its relevance to empirical research.

There are several interesting directions for future research. First,
one can consider inference that is uniformly valid regardless of whether
the (conditional) distributions of the potential outcomes are point
identified or partially identified, as in, for example, \citet{Imbens2004},
\citet{Stoye2009}, and \citet{andrews2010inference}. Second, one
can consider testing global hypotheses, such as stochastic dominance.
Although we cannot directly test stochastic dominance between two
conditional distribution of treatment effects as they are not point
identified, one can provide weak evidence by using similar arguments
to those in \citet{firpo2021uniform}. Third, it would be fruitful
to develop inference methods that are uniformly valid in values of
covariates. In work in progress, we consider a semiparametric approach
to uniform inference over the support of treatment effects and covariates.
It is expected to help resolve many interesting questions in empirical
analysis that cannot be answered by the framework proposed in this
paper. Lastly, it is worth considering inference in the presence of
instrumental variables.

\bibliographystyle{chicago}
\bibliography{dist_te}

\newpage{}