EconBase
← Back to paper

Stratifying on Treatment Status

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

24,623 characters

Stratifying on Treatment Status


\title{Stratifying on Treatment Status}
\author{Jinyong Hahn\thanks{Department of Economics, UCLA, Los Angeles, CA 90095, USA. Email: [email removed].}\\UCLA
\and John Ham\thanks{Department of Economics, NYU Abu Dhabi, Saadiyat Campus, P.O. Box 129188, Abu Dhabi, United Arab Emirates. Email: [email removed].}\\NYU Abu Dhabi
\and Geert Ridder\thanks{Department of Economics, USC, Los Angeles, CA 90089, USA. Email: [email removed].}\\USC
\and Shuyang Sheng\thanks {Shenzhen Finance Institute, School of Management and Economics, The Chinese University of Hong Kong, Shenzhen, 2001 Longxiang Boulevard, Longgang District, Shenzhen, China. Email: [email removed].}\\CUHK-Shenzhen}
\date{\today}
\maketitle

\begin{abstract}
We study the estimation of treatment effects using samples stratified by treatment status. Standard estimators of the average treatment effect and the local average treatment effect are inconsistent in this setting. We propose consistent estimators and characterize their asymptotic distributions.  \\
\noindent\textbf{JEL Codes}: C21, C83, C14. \\
\noindent\textbf{Keywords}: Stratification, Treatment Effects, Asymptotic Variance.

\end{abstract}

\newpage

\section{Introduction}

Sample surveys are usually not simple random samples. Often the population is
partitioned into strata, where the subpopulation in each stratum is sampled
with equal probability, but the inclusion probability is unequal between
strata. Stratified sampling is preferred for several reasons. First, it can
result in more precise estimates
\citep{Domencich_McFadden_1975}. Second, it may over-represent groups of
particular interest. Third, it may arise naturally from the combination of
independent data sets into a single sample \citep{Ridder_Moffitt_2007}.

Many studies select strata based on treatment status. For example,
\citet{Azoulay_2010} analyze the impact of collaborator quality on research
productivity. They compare researchers whose superstar collaborators died with
those whose superstar collaborators did not. Their dataset includes all researchers with a deceased superstar co-author and
a random sample of researchers with a living one. Similarly, \citet{Ham_Khan_2023} assess the effectiveness of a new teaching
approach, JAAGO, against government and non-governmental schools in urban
Bangladesh. The new approach is implemented in only two schools in Dhaka, so
they use the population of JAAGO students alongside a random sample of
students from other schools.

The classical literature showed in parametric settings that stratification on
endogenous variables leads to biased estimates,\footnote{Stratification on
endogenous variables was first discussed in discrete choice models under the
name choice-based sampling \citep{Manski_Lerman_1977,Manski_McFadden_1981,Hausman_Wise_1981}.} whereas
stratification on exogenous variables typically does not result in a bias. We
show that stratification on unconfounded treatment results in a biased estimate
of the average treatment effect (ATE). This may seem to contradict the results in
\cite{Heckman_Todd_2009}, who argued that \textquotedblleft matching and
selection procedures can still be applied when the propensity score is
estimated on unweighted choice based samples\textquotedblright.

The seeming contradiction is resolved by observing that although the conditional average treatment effect (CATE) given confounders $X$ can be identified (from the conditional distributions
of outcome $Y$ given $X$ and treatment status), the \textit{average} of the CATE over the distribution of $X$ (or that of the propensity
score) in a stratified sample results in inconsistent estimates of the ATE. There is a simple fix to the problem, and a reweighted method can identify the ATE, provided that the population
fraction treated is known or can be estimated. In contrast, the average
treatment effect on the treated (ATT) is identified by conventional methods
despite stratification, even if the population fraction treated is unknown.

In addition, we consider the case of an endogenous treatment with a valid
binary instrument. We show that the standard Wald ratio estimator is not
consistent for the local average treatment effect (LATE) when the sample is
stratified on the treatment status. We propose a reweighted method to recover
the LATE, which requires knowledge of the fraction treated in the population.

We contribute to the literature by developing \textit{constructive}
identification results that can form a basis for consistent
estimation. In particular, we present an analysis of the asymptotic distribution of various estimators under
stratified sampling, which is new to the
literature.\footnote{\citet{Correa_Tian_Bareinboim_2019} and \citet{Nabi2020}
provide generic identification results on the joint distribution of outcomes
and covariates for the treated and controls, but they do not have constructive
identification methods for estimating parameters of interest such as the ATE
or ATT. \citet{Hu2018} and \citet{Zhang2019} propose a reweighted method that is
the same as ours, but only for the ATE using covariate conditioning. We extend
this method to include other parameters of interest, such as the ATT and
LATE. \citet{Song2022} discuss the identification and estimation of ATE/ATT under unconfoundedness in the scenario where sampling is stratified based on the treatment and other discrete covariates. They focus on the case where the treatment probability in the population is known up to an interval. In contrast, we establish asymptotic results under the assumption that this probability is known. We also provide results for LATE, which is not covered in their paper.}

\section{Unconfounded Treatment\label{sect.exog}}

We start with the case where the treatment $D$ is independent of potential outcomes $(Y_{0},Y_{1})$ given $X$.
Let $Y\equiv DY_{1}+(1-D)Y_{0}$ denote the observed outcome. Denote the
fraction $D=1$ in the sample by $\pi^{\ast}$ and in the population by $\pi$. In a stratified sample, we have $\pi^{\ast}\neq\pi$.

For $d=0,1$, let $h(y|x,D=d)$ and $h^{\ast}(y|x,D=d)$ denote the conditional
densities of $Y$ given $X=x$ and $D=d$ in the population and sample,
respectively. Similarly, let $g(x|D=d)$ and $g^{\ast}(x|D=d)$ denote the
conditional densities of $X$ given $D=d$ in the population and sample. Because we sample randomly in the strata, the conditional distribution of $(Y,X)$ given $D$ is identical in the population and sample. Therefore for
$d=0,1$,
\begin{align}
h(y|x,D=d)  &  =h^{\ast}(y|x,D=d),\label{Ydensity}\\
g(x|D=d)  &  =g^{\ast}(x|D=d). \label{Xdensity}
\end{align}

Sampling objects that do not condition on the treatment status, however, may
differ from the population counterparts. To understand their relationship, let
$\pi(x)\equiv\Pr(D=1|X=x)$ denote the propensity score in the population, and
$\pi^{\ast}(x)\equiv\Pr^{\ast}(D=1|X=x)$ the propensity score in the
stratified sample. Moreover, let $g(x)$ and $g^{\ast}(x)$ denote the densities
of $X$ in the population and stratified sample, respectively. By Bayes'
theorem
\begin{equation}
g^{\ast}(x|D=1)=g^{\ast}(x)\frac{\pi^{\ast}(x)}{\pi^{\ast}},\qquad g^{\ast
}(x|D=0)=g^{\ast}(x)\frac{1-\pi^{\ast}(x)}{1-\pi^{\ast}}, \label{g*Bayes}
\end{equation}
and similar equations hold for the population counterparts. Therefore, we can
derive
\begin{equation}
\pi(x)=\frac{g(x|D=1)\pi}{g(x|D=1)\pi+g(x|D=0)(1-\pi)}=\frac{g^{\ast
}(x|D=1)\pi}{g^{\ast}(x|D=1)\pi+g^{\ast}(x|D=0)(1-\pi)}, \label{pscore}
\end{equation}
from which we obtain
\begin{equation}
\pi(x)=\frac{\pi^{\ast}(x)\frac{\pi}{\pi^{\ast}}}{\pi^{\ast}(x)\frac{\pi}
{\pi^{\ast}}+(1-\pi^{\ast}(x))\frac{1-\pi}{1-\pi^{\ast}}}. \label{pi}
\end{equation}
Because $\pi^{\ast}(x)$ and $\pi^{\ast}$ are identified from the sampling
distribution of $(Y,D,X)$, the population propensity score $\pi(x)$ is
identified \textit{if and only }$\pi$ is known.\ Switching the population and
sample objects in (\ref{pscore}), we can derive similarly
\begin{equation}
\pi^{\ast}(x)=\frac{\pi(x)\frac{\pi^{\ast}}{\pi}}{\pi(x)\frac{\pi^{\ast}}{\pi
}+(1-\pi(x))\frac{1-\pi^{\ast}}{1-\pi}}. \label{pi*}
\end{equation}
In addition, the population density $g(x)$ of $X$ satisfies
\[
g(x)=g(x|D=1)\pi+g(x|D=0)(1-\pi)=g^{\ast}(x|D=1)\pi+g^{\ast}(x|D=0)(1-\pi),
\]
so by (\ref{g*Bayes})
\begin{equation}
g(x)=g^{\ast}(x)\left(  \pi^{\ast}(x)\frac{\pi}{\pi^{\ast}}+(1-\pi^{\ast
}(x))\frac{1-\pi}{1-\pi^{\ast}}\right)  . \label{g*tog}
\end{equation}
Because $\pi^{\ast}(x)$, $g^{\ast}(x)$, and $\pi^{\ast}$ are identified from
the distribution of $(Y,D,X)$ in the stratified sample, the population density
$g(x)$ of $X$ can be identified \textit{if and only if} $\pi$ is known.

Equation (\ref{Ydensity}) implies that the conditional average treatment
effect (CATE) is identified by
\begin{equation}
\beta(X)\equiv E[Y|D=1,X]-E[Y|D=0,X]=E^{\ast}[Y|D=1,X]-E^{\ast}[Y|D=0,X]\equiv
\beta^{\ast}(X), \label{matching-works}
\end{equation}
where $E$ and $E^{\ast}$ denote the expectations taken with respect to
$h(y|X,D=d)$ and $h^{\ast}(y|X,D=d)$, which are equal. This is the sense in which conditioning
on $X$ \textquotedblleft works\textquotedblright\ \citep{Heckman_Todd_2009}.

\subsection{Average Treatment Effect (ATE)}

We now explore whether the success of conditioning translates into success in
identifying the ATE, $\beta\equiv E[Y_{1}-Y_{0}]$. By iterated expectations,
the ATE and its sample counterpart are
\begin{align*}
E[Y_{1}-Y_{0}]  &  =E[E[Y|D=1,X]-E[Y|D=0,X]],\\
E^{\ast}[Y_{1}-Y_{0}]  &  =E^{\ast}[E^{\ast}[Y|D=1,X]-E^{\ast}[Y|D=0,X]].
\end{align*}
Because $g^{\ast}(x)\neq g(x)$ by (\ref{g*tog}), the sample counterpart does
not identify the ATE even though $\beta(x)=\beta^*(x)$
\[
E[Y_{1}-Y_{0}]=\int\beta(x)g(x)dx\neq\int\beta^{\ast}(x)g^{\ast}(x)dx=E^{\ast
}[Y_{1}-Y_{0}].
\]
This result is intuitive and unsurprising. The problem is not the failure of
conditioning, but is due to
averaging the CATE over the wrong distribution of $X$.

The ATE cannot be identified by conditioning on the propensity score either.
Because the independence of the potential outcomes and $D$ given $X$ implies
independence given the propensity score, it follows by iterated expectations
that $E[Y_{1}-Y_{0}]=E[E[Y|D=1,\pi(X)]-E[Y|D=0,\pi(X)]]$. From (\ref{pi}) and
(\ref{pi*}), there is a 1-1 relation between $\pi(x)$ and $\pi^{\ast}(x)$. It
follows that the $\sigma$-algebras generated by these two random variables are
identical and thus conditioning on $\pi(X)$ and conditioning on $\pi^{\ast
}(X)$ are equivalent \citep{Heckman_Todd_2009}. Hence,
\begin{align*}
\delta(\pi(X))  &  \equiv E[Y|D=1,\pi(X)]-E[Y|D=0,\pi(X)]\\
&  =E^{\ast}[Y|D=1,\pi^{\ast}(X)]-E^{\ast}[Y|D=0,\pi^{\ast}(X)]\equiv
\delta^{\ast}(\pi^{\ast}(X)).
\end{align*}
Unfortunately, the averaging problem remains. We can see that
\[
E[Y_{1}-Y_{0}]=\int\delta(\pi(x))g(x)dx\neq\int\delta^{\ast}(\pi^{\ast
}(x))g^{\ast}(x)dx=E^{\ast}[Y_{1}-Y_{0}]
\]
because $g^{\ast}(x)\neq g(x)$ by (\ref{g*tog}). Conditioning only works
partially because it does not solve the problem of averaging. Therefore, we should be
careful interpreting \citet[p.S231]{Heckman_Todd_2009}'s observation that \textquotedblleft matching and selection procedures can
identify population treatment effects using misspecified estimates of
propensity scores fit on choice-based samples\textquotedblright\ when
estimating the ATE.

The problem remains if we use the inverse propensity score weighting (IPW).
IPW identifies the ATE by
\[
E[Y_{1}-Y_{0}]=E\left[  \frac{DY}{\pi(X)}-\frac{(1-D)Y}{1-\pi(X)}\right]  ,
\]
which is not identified by the sample counterpart
\[
E\left[  \frac{DY}{\pi(X)}-\frac{(1-D)Y}{1-\pi(X)}\right]  =E[\beta(X)]\neq
E^{\ast}[\beta^{\ast}(X)]=E^{\ast}\left[  \frac{DY}{\pi^{\ast}(X)}
-\frac{(1-D)Y}{1-\pi^{\ast}(X)}\right]
\]
again due to improper averaging, even though $\beta(x)=\beta^{\ast}(x)$.

If we know $\pi$, we can identify the ATE using reweighting based on
(\ref{g*tog})
\begin{align}
E[Y_{1}-Y_{0}]  &  =\int\beta^{\ast}(x)g(x)dx\nonumber\\
&  =\int\beta^{\ast}(x)\left(  \pi^{\ast}(x)\frac{\pi}{\pi^{\ast}}
+(1-\pi^{\ast}(x))\frac{1-\pi}{1-\pi^{\ast}}\right)  g^{\ast}(x)dx\nonumber\\
&  =E^{\ast}\left[  \beta^{\ast}(X)\left(  D\frac{\pi}{\pi^{\ast}}
+(1-D)\frac{1-\pi}{1-\pi^{\ast}}\right)  \right]  . \label{ATE}
\end{align}
If we condition on the propensity score, we can also identify the ATE through
reweighting
\begin{align}
E[Y_{1}-Y_{0}]  &  =\int\delta^{\ast}(\pi^{\ast}(x))g(x)dx\nonumber\\
&  =\int\delta^{\ast}(\pi^{\ast}(x))\left(  \pi^{\ast}(x)\frac{\pi}{\pi^{\ast
}}+(1-\pi^{\ast}(x))\frac{1-\pi}{1-\pi^{\ast}}\right)  g^{\ast}
(x)dx\nonumber\\
&  =E^{\ast}\left[  \delta^{\ast}(\pi^{\ast}(X))\left(  D\frac{\pi}{\pi^{\ast
}}+(1-D)\frac{1-\pi}{1-\pi^{\ast}}\right)  \right]  . \label{ATE-PS}
\end{align}
In addition, the ATE can be identified by a reweighted version of IPW
\begin{equation}
E^{\ast}\left[  \left(  \frac{DY}{\pi^{\ast}(X)}-\frac{(1-D)Y}{1-\pi^{\ast
}(X)}\right)  \left(  \pi^{\ast}(X)\frac{\pi}{\pi^{\ast}}+(1-\pi^{\ast
}(X))\frac{1-\pi}{1-\pi^{\ast}}\right)  \right]  . \label{IPW}
\end{equation}

In sum, conventional methods do not identify the ATE, but the problem can be
overcome by reweighting the observations according to the population
distribution of $X$.

\subsection{Average Treatment Effect on the Treated (ATT)}

Next we examine the ATT, $\gamma\equiv E[Y_{1}-Y_{0}|D=1]$. Note that by
(\ref{g*Bayes}) and (\ref{matching-works})
\begin{align}
E[Y_{1}-Y_{0}|D=1]  &  =\int\beta(x)g(x|D=1)dx\nonumber\\
&  =\int\beta^{\ast}(x)g^{\ast}(x|D=1)dx=E^{\ast}[Y_{1}-Y_{0}|D=1].
\label{ATT}
\end{align}
The ATT is identified despite stratification. We can also identify the ATT
conditioning on the propensity score because
\begin{align}
E[Y_{1}-Y_{0}|D=1]  &  =\int\delta(\pi(x))g(x|D=1)dx\nonumber\\
&  =\int\delta^{\ast}(\pi^{\ast}(x))g^{\ast}(x|D=1)dx=E^{\ast}[Y_{1}
-Y_{0}|D=1]. \label{ATT-PS}
\end{align}
Lastly, note that by iterated expectations
\[
E[Y_{1}-Y_{0}|D=1]=\frac{1}{\pi}E\left[  \left(  \frac{DY}{\pi(X)}
+\frac{(1-D)Y}{1-\pi(X)}\right)  \pi(X)\right]  ,
\]
whose sample counterpart identifies the ATT
\[
\frac{1}{\pi^{\ast}}E^{\ast}\left[  \left(  \frac{DY}{\pi^{\ast}(X)}
+\frac{(1-D)Y}{1-\pi^{\ast}(X)}\right)  \pi^{\ast}(X)\right]  =E^{\ast}
[\beta^{\ast}(X)|D=1]=E[\beta(X)|D=1]=E[Y_{1}-Y_{0}|D=1].
\]
Hence, the ATT can be identified using IPW as well.

In sum, the ATT is identified by conventional methods. Stratification has no
impact on the averaging because the distribution of $X$ given $D=1$ is the
same in the population and stratified sample.

\subsection{Asymptotic Distribution}

\subsubsection{ATE Estimators}

Assuming that the population fraction $\pi$ is known, equation (\ref{ATE}) suggests that we can estimate the ATE $\beta$ by a
semiparametric estimator $\hat{\beta}$ based on the moments
\begin{align}
E^{\ast}[Y-\beta_{1}^{\ast}(X)|D=1,X]  &  =0\nonumber\\
E^{\ast}[Y-\beta_{0}^{\ast}(X)|D=0,X]  &  =0\nonumber\\
E^{\ast}\left[  \left(  D\frac{\pi}{\pi^{\ast}}+(1-D)\frac{1-\pi}{1-\pi^{\ast
}}\right)  (\beta_{1}^{\ast}(X)-\beta_{0}^{\ast}(X))-\beta\right]   &  =0,
\label{ATE-estimator}
\end{align}
where $\beta_{d}^{\ast}(X)\equiv E^{\ast}[Y|D=d,X]$, $d=0,1$. Alternatively,
equation (\ref{ATE-PS}) suggests a semiparametric estimator using the
propensity score
\begin{align}
E^{\ast}[D-\pi^{\ast}(X)|X]  &  =0\nonumber\\
E^{\ast}[Y-\delta_{1}^{\ast}(\pi^{\ast}(X))|D=1,\pi^{\ast}(X)]  &
=0\nonumber\\
E^{\ast}[Y-\delta_{0}^{\ast}(\pi^{\ast}(X))|D=0,\pi^{\ast}(X)]  &
=0\nonumber\\
E^{\ast}\left[  \left(  D\frac{\pi}{\pi^{\ast}}+(1-D)\frac{1-\pi}{1-\pi^{\ast
}}\right)  (\delta_{1}^{\ast}(\pi^{\ast}(X))-\delta_{0}^{\ast}(\pi^{\ast
}(X)))-\beta\right]   &  =0, \label{ATE-PS-estimator}
\end{align}
where $\delta_{d}^{\ast}(\pi^{\ast}(X))\equiv E^{\ast}[Y|D=d,\pi^{\ast}(X)]$,
$d=0,1$.\footnote{The moment conditions (\ref{ATE-estimator}) and
(\ref{ATE-PS-estimator}) can be understood to be semiparametric
generalization of the parametric model considered by \citet{Nevo2003}, e.g.}
Applying \citet{Newey_1994}, we derive the influence functions of these ATE
estimators, presented in Propositions \ref{prop:ATE} and \ref{prop:ATE-PS}.\footnote{All the proofs are provided in the Supplementary Appendix, which is
available upon request.}  Propositions \ref{prop:ATE} and \ref{prop:ATE-PS} are both predicated on the assumption that the population fraction $\pi$ is known. The sample fraction $\pi^{\ast}$ may or may not be known, and the two propositions consider both cases.\footnote{The representations in (\ref{ATE-estimator}) and
(\ref{ATE-PS-estimator}) assume that $\pi^{\ast}$ is known. If $\pi^{\ast}$ is
unknown, we can add one more moment $E^{\ast}[D-\pi^{\ast}]=0$ to each system.}

\begin{proposition}
\label{prop:ATE}If $\pi^{\ast}$ is known, the influence function of the ATE
estimator based on (\ref{ATE-estimator}) is
\begin{align}
&  \left(  D\frac{\pi}{\pi^{\ast}}+(1-D) \frac{1-\pi}{1-\pi^{\ast}}\right)
(\beta_{1}(X)-\beta_{0}(X)) -\beta\nonumber\\
&  +\left(  \pi^{\ast}(X) \frac{\pi}{\pi^{\ast}}+(1-\pi^{\ast}(X)) \frac
{1-\pi}{1-\pi^{\ast}}\right)  \left(  \frac{D}{\pi^{\ast}(X) }(Y-\beta_{1}(X))
-\frac{1-D}{1-\pi^{\ast}( X)}(Y-\beta_{0}(X)) \right)  , \label{ATE-infl}
\end{align}
where $\beta_{d}(X)\equiv E[Y|D=d,X]$, $d=0,1$. If $\pi^{\ast}$ is estimated, the influence function is the sum of (\ref{ATE-infl}) and
\begin{equation}
E^{\ast}\left[  \left(  -\pi^{\ast}(X) \frac{\pi}{(\pi^{\ast})^{2}}+(
1-\pi^{\ast}(X)) \frac{1-\pi}{(1-\pi^{\ast})^{2}}\right)  (\beta_{1}
(X)-\beta_{0}(X)) \right]  (D-\pi^{\ast}). \label{ATE-infl-adjust}
\end{equation}

\end{proposition}

\begin{proposition}
\label{prop:ATE-PS}The influence function of the ATE estimator based on
(\ref{ATE-PS-estimator}) equals (\ref{ATE-infl}) if $\pi^{\ast}$ is known, and
equals the sum of (\ref{ATE-infl}) and (\ref{ATE-infl-adjust}) if $\pi^{\ast}$
is estimated.
\end{proposition}

The results mirror those in the non-stratified case, where various ATE estimators have the same asymptotic distribution. We can derive the asymptotic
variance of the ATE estimators using (\ref{ATE-infl}) if $\pi^{\ast}$ is
known. Denote $\epsilon_{d}\equiv Y-\beta_{d}(X)$ and $\sigma_{d}^{2}(X)
\equiv E^{\ast}[\epsilon_{d}^{2}|X]$, $d=0,1$, and $w(X)\equiv\pi^{\ast}(X)
\frac{\pi}{\pi^{\ast}}+(1-\pi^{\ast}(X)) \frac{1-\pi}{1-\pi^{\ast}}$. The
asymptotic variance of $\sqrt{n}(\hat{\beta}-\beta)$ is
\begin{equation}
E^{\ast}\left[  \left(  \left(  D\frac{\pi}{\pi^{\ast}}+(1-D) \frac{1-\pi
}{1-\pi^{\ast}}\right)  \beta(X) -\beta\right)  ^{2} +\frac{w(X)^{2}\sigma
_{1}^{2}(X)}{\pi^{\ast}(X)} +\frac{w(X)^{2}\sigma_{0}^{2}(X)}{1-\pi^{\ast}
(X)}\right]. \label{ATE-var}
\end{equation}

\subsubsection{ATT Estimators}

Equations (\ref{ATT}) and (\ref{ATT-PS}) serve as the basis for estimating the
ATT. Because $\beta^{\ast}(X)=E^{\ast}[Y-\beta_{0}^{\ast}(X)|D=1,X]$, where
$\beta_{0}^{\ast}(X)\equiv E^{\ast}[Y|D=0,X]$, it follows by iterated
expectations that $E^{\ast}[\beta^{\ast}(X)|D=1]=E^{\ast}[Y-\beta_{0}^{\ast
}(X)|D=1]=E^{\ast}[D(Y-\beta_{0}^{\ast}(X))]/\pi^{\ast}$. This suggests an
estimator of $\gamma$ of the form
\begin{equation}
\hat{\gamma}\equiv\frac{1}{n^{-1}\sum_{i=1}^{n}D_{i}}\cdot\frac{1}{n}
\sum_{i=1}^{n}D_{i}(Y_{i}-\hat{\beta}_{0}(X_{i})), \label{ATT-estimator}
\end{equation}
where $\hat{\beta}_{0}(\cdot)$ is a nonparametric estimator of $\beta
_{0}^{\ast}(\cdot)$. This estimator subtracts the predicted outcome for the
counterfactual from the outcome of a treated unit.
Although intuitive, we are not aware of any references that establish the
asymptotic distribution of such an estimator.

Define $\delta_{0}^{\ast}(\pi^{\ast}(X))\equiv E^{\ast}[Y|D=0,\pi^{\ast}(X)]$.
A similar estimator can be constructed based on the propensity score
\begin{equation}
\hat{\gamma}_{PS}\equiv\frac{1}{n^{-1}\sum_{i=1}^{n}D_{i}}\cdot\frac{1}{n}
\sum_{i=1}^{n}D_{i}(Y_{i}-\hat{\delta}_{0}(\hat{\pi}(X_{i}))),
\label{ATT-PS-estimator}
\end{equation}
where $\hat{\delta}_{0}(\cdot)$ and $\hat{\pi}(\cdot)$ are nonparametric
estimators of $\delta_{0}^{\ast}(\cdot)$ and $\pi^{\ast}(\cdot)$,
respectively. Propositions \ref{prop:ATT} and \ref{prop:ATT-PS} derive the
asymptotic variances of these ATT estimators.


\begin{proposition}
\label{prop:ATT}If $\pi^{\ast}$ is known, the asymptotic variance of $\sqrt
{n}(\hat{\gamma}-\gamma)$ is equal to
\begin{equation}
E^{\ast}\left[  \frac{\pi^{\ast}(X)(\beta(X)-\gamma)^{2}}{(\pi^{\ast})^{2}
}+\frac{\pi^{\ast}(X) \sigma_{1}^{2}(X)}{(\pi^{\ast})^{2}}+\frac{\pi^{\ast
}(X)^{2}\sigma_{0}^{2}(X) }{(\pi^{\ast})^{2}(1-\pi^{\ast}(X))}\right]  .
\label{ATT-var}
\end{equation}

\end{proposition}

\begin{proposition}
\label{prop:ATT-PS}If $\pi^{\ast}$ is known, the asymptotic variance of
$\sqrt{n}(\hat{\gamma}_{PS}-\gamma)$ is the same as in (\ref{ATT-var}).
\end{proposition}

The ATT estimators in (\ref{ATT-estimator}) and (\ref{ATT-PS-estimator}) have
the same asymptotic variance in (\ref{ATT-var}), which is equal to the
asymptotic variance bound based on the stratified sample.\footnote{Note that
(\ref{ATT-var}) is different from the asymptotic variance based on an
unstratified population. While the ATT can be consistently estimated, the
asymptotic variance is impacted by stratification.} See \cite{Hahn_1998}. So
both ATT estimators are efficient.\footnote{The fact that the ATE (or ATT)
estimators that condition on either the propensity score or the covariates
have the same asymptotic variance indicates that there is no efficiency gain
from using the propensity score in stratified samples. \citet{HR_2013} come to
the same conclusion in random sampling.}

\section{Endogenous Treatment}

In this section, we consider stratification on an endogenous treatment. We
have a binary instrument $Z$ and assume no covariates to focus on essentials.
\citet{AIR_1996} show that the local average treatment effect (LATE) is
identified by the Wald ratio $(E[Y|Z=1]-E[Y|Z=0])/(E[D|Z=1]-E[D|Z=0])$. If the
sample is stratified on the treatment status, we have
\begin{align*}
E[YZ|D=d]=  &  E^{\ast}[YZ|D=d]\\
E[Y(1-Z)|D=d]=  &  E^{\ast}[Y(1-Z)|D=d]\\
E[Z|D=d]=  &  E^{\ast}[Z|D=d],
\end{align*}
for $d=0,1$. However, because $\pi\neq\pi^{\ast}$, unconditional expectations
in the sample do not match those in the population. We can see that
\begin{align*}
E[Y|Z=1]-E[Y|Z=0]=  &  \frac{E[YZ]}{E[Z]}-\frac{E[Y(1-Z)]}{1-E[Z]}\\
\neq &  \frac{E^{\ast}[YZ]}{E^{\ast}[Z]}-\frac{E^{\ast}[Y(1-Z)]}{1-E^{\ast
}[Z]}=E^{\ast}[Y|Z=1]-E^{\ast}[Y|Z=0],
\end{align*}
and similarly $E[D|Z=1]-E[D|Z=0]\neq E^{\ast}[D|Z=1]-E^{\ast}[D|Z=0]$.
Therefore, the sample Wald ratio does not identify the population Wald ratio.

Observe that the population Wald ratio is the solution for $\beta$ in the
moment condition
\[
E\left\{
\begin{bmatrix}
1\\
Z
\end{bmatrix}
(Y-\alpha-\beta D)\right\}  =0.
\]
Proposition \ref{prop:weighting} suggests that we can identify the population
Wald ratio using reweighting.

\begin{proposition}
\label{prop:weighting}For any function $\varphi(Y,Z,D)$, we have
\[
E[\varphi(Y,Z,D)]=E^{\ast}\left[  \left(  D\frac{\pi}{\pi^{\ast}}
+(1-D)\frac{1-\pi}{1-\pi^{\ast}}\right)  \varphi(Y,Z,D)\right]  .
\]

\end{proposition}

Proposition \ref{prop:weighting} implies that the population Wald ratio is the
solution for $\beta$ in the reweighted sample moment condition
\[
E^{\ast}\left\{  \left(  D\frac{\pi}{\pi^{\ast}}+(1-D)\frac{1-\pi}{1-\pi
^{\ast}}\right)
\begin{bmatrix}
1\\
Z
\end{bmatrix}
(Y-\alpha-\beta D)\right\}  =0.
\]
This fix requires knowledge of $\pi$.\footnote{\citet{Froelich_2007} considers
a nonparametric estimator of LATE that allows for covariates.}

\begin{remark}
In the unconfounded case without $X$, i.e., under random assignment,
stratification on $D$ does not pose any issue, because without $X$ there is no
need for averaging. In the endogenous case, however, the bias is present even
without $X$. The intuition can be found in the literature on stratification on
endogenous variables (e.g, \citealp{Hausman_Wise_1981}).
\end{remark}

\section{Conclusion}

For unconfounded treatment, stratification on the treatment status does not
bias conditional treatment effects. It does bias the ATE, which can be
corrected through reweighting if we know the fraction treated in the
population. Stratification does not bias the ATT. In the case of endogenous
treatment with an instrument, stratification biases the LATE. We propose a
reweighted method to identify the LATE.

\bibliography{StraTreatment}

\newpage