EconBase
← Back to paper

Stratifying on Treatment Status

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

24,632 characters · 9 sections · 21 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Stratifying on Treatment Status

abstractWe study the estimation of treatment effects using samples stratified by treatment status. Standard estimators of the average treatment effect and the local average treatment effect are inconsistent in this setting. We propose consistent estimators and characterize their asymptotic distributions. \\ \noindentJEL Codes: C21, C83, C14. \\ \noindentKeywords: Stratification, Treatment Effects, Asymptotic Variance.

Introduction

Sample surveys are usually not simple random samples. Often the population is partitioned into strata, where the subpopulation in each stratum is sampled with equal probability, but the inclusion probability is unequal between strata. Stratified sampling is preferred for several reasons. First, it can result in more precise estimates Domencich_McFadden_1975. Second, it may over-represent groups of particular interest. Third, it may arise naturally from the combination of independent data sets into a single sample Ridder_Moffitt_2007.

Many studies select strata based on treatment status. For example, Azoulay_2010 analyze the impact of collaborator quality on research productivity. They compare researchers whose superstar collaborators died with those whose superstar collaborators did not. Their dataset includes all researchers with a deceased superstar co-author and a random sample of researchers with a living one. Similarly, Ham_Khan_2023 assess the effectiveness of a new teaching approach, JAAGO, against government and non-governmental schools in urban Bangladesh. The new approach is implemented in only two schools in Dhaka, so they use the population of JAAGO students alongside a random sample of students from other schools.

The classical literature showed in parametric settings that stratification on endogenous variables leads to biased estimates,\footnote{Stratification on endogenous variables was first discussed in discrete choice models under the name choice-based sampling Manski_Lerman_1977,Manski_McFadden_1981,Hausman_Wise_1981.} whereas stratification on exogenous variables typically does not result in a bias. We show that stratification on unconfounded treatment results in a biased estimate of the average treatment effect (ATE). This may seem to contradict the results in Heckman_Todd_2009, who argued that \textquotedblleft matching and selection procedures can still be applied when the propensity score is estimated on unweighted choice based samples\textquotedblright.

The seeming contradiction is resolved by observing that although the conditional average treatment effect (CATE) given confounders $X$ can be identified (from the conditional distributions of outcome $Y$ given $X$ and treatment status), the average of the CATE over the distribution of $X$ (or that of the propensity score) in a stratified sample results in inconsistent estimates of the ATE. There is a simple fix to the problem, and a reweighted method can identify the ATE, provided that the population fraction treated is known or can be estimated. In contrast, the average treatment effect on the treated (ATT) is identified by conventional methods despite stratification, even if the population fraction treated is unknown.

In addition, we consider the case of an endogenous treatment with a valid binary instrument. We show that the standard Wald ratio estimator is not consistent for the local average treatment effect (LATE) when the sample is stratified on the treatment status. We propose a reweighted method to recover the LATE, which requires knowledge of the fraction treated in the population.

We contribute to the literature by developing constructive identification results that can form a basis for consistent estimation. In particular, we present an analysis of the asymptotic distribution of various estimators under stratified sampling, which is new to the literature.\footnote{Correa_Tian_Bareinboim_2019 and Nabi2020 provide generic identification results on the joint distribution of outcomes and covariates for the treated and controls, but they do not have constructive identification methods for estimating parameters of interest such as the ATE or ATT. Hu2018 and Zhang2019 propose a reweighted method that is the same as ours, but only for the ATE using covariate conditioning. We extend this method to include other parameters of interest, such as the ATT and LATE. Song2022 discuss the identification and estimation of ATE/ATT under unconfoundedness in the scenario where sampling is stratified based on the treatment and other discrete covariates. They focus on the case where the treatment probability in the population is known up to an interval. In contrast, we establish asymptotic results under the assumption that this probability is known. We also provide results for LATE, which is not covered in their paper.}

Unconfounded Treatment

We start with the case where the treatment $D$ is independent of potential outcomes $(Y_{0},Y_{1})$ given $X$. Let $Y\equiv DY_{1}+(1-D)Y_{0}$ denote the observed outcome. Denote the fraction $D=1$ in the sample by $\pi^{\ast}$ and in the population by $\pi$. In a stratified sample, we have $\pi^{\ast}\neq\pi$.

For $d=0,1$, let $h(y|x,D=d)$ and $h^{\ast}(y|x,D=d)$ denote the conditional densities of $Y$ given $X=x$ and $D=d$ in the population and sample, respectively. Similarly, let $g(x|D=d)$ and $g^{\ast}(x|D=d)$ denote the conditional densities of $X$ given $D=d$ in the population and sample. Because we sample randomly in the strata, the conditional distribution of $(Y,X)$ given $D$ is identical in the population and sample. Therefore for $d=0,1$,

align[align omitted — 113 chars of source]

Sampling objects that do not condition on the treatment status, however, may differ from the population counterparts. To understand their relationship, let $\pi(x)\equiv\Pr(D=1|X=x)$ denote the propensity score in the population, and $\pi^{\ast}(x)\equiv\Pr^{\ast}(D=1|X=x)$ the propensity score in the stratified sample. Moreover, let $g(x)$ and $g^{\ast}(x)$ denote the densities of $X$ in the population and stratified sample, respectively. By Bayes' theorem

equation[equation omitted — 164 chars of source]

and similar equations hold for the population counterparts. Therefore, we can derive

equation[equation omitted — 156 chars of source]

from which we obtain

equation[equation omitted — 159 chars of source]

Because $\pi^{\ast}(x)$ and $\pi^{\ast}$ are identified from the sampling distribution of $(Y,D,X)$, the population propensity score $\pi(x)$ is identified if and only $\pi$ is known.\ Switching the population and sample objects in ((ref)), we can derive similarly

equation[equation omitted — 146 chars of source]

In addition, the population density $g(x)$ of $X$ satisfies \[ g(x)=g(x|D=1)\pi+g(x|D=0)(1-\pi)=g^{\ast}(x|D=1)\pi+g^{\ast}(x|D=0)(1-\pi), \] so by ((ref))

equation[equation omitted — 144 chars of source]

Because $\pi^{\ast}(x)$, $g^{\ast}(x)$, and $\pi^{\ast}$ are identified from the distribution of $(Y,D,X)$ in the stratified sample, the population density $g(x)$ of $X$ can be identified if and only if $\pi$ is known.

Equation ((ref)) implies that the conditional average treatment effect (CATE) is identified by

equation[equation omitted — 134 chars of source]

where $E$ and $E^{\ast}$ denote the expectations taken with respect to $h(y|X,D=d)$ and $h^{\ast}(y|X,D=d)$, which are equal. This is the sense in which conditioning on $X$ \textquotedblleft works\textquotedblright\ Heckman_Todd_2009.

Average Treatment Effect (ATE)

We now explore whether the success of conditioning translates into success in identifying the ATE, $\beta\equiv E[Y_{1}-Y_{0}]$. By iterated expectations, the ATE and its sample counterpart are

align*[align* omitted — 135 chars of source]

Because $g^{\ast}(x)\neq g(x)$ by ((ref)), the sample counterpart does not identify the ATE even though $\beta(x)=\beta^*(x)$ \[ E[Y_{1}-Y_{0}]=\int\beta(x)g(x)dx\neq\int\beta^{\ast}(x)g^{\ast}(x)dx=E^{\ast }[Y_{1}-Y_{0}]. \] This result is intuitive and unsurprising. The problem is not the failure of conditioning, but is due to averaging the CATE over the wrong distribution of $X$.

The ATE cannot be identified by conditioning on the propensity score either. Because the independence of the potential outcomes and $D$ given $X$ implies independence given the propensity score, it follows by iterated expectations that $E[Y_{1}-Y_{0}]=E[E[Y|D=1,\pi(X)]-E[Y|D=0,\pi(X)]]$. From ((ref)) and ((ref)), there is a 1-1 relation between $\pi(x)$ and $\pi^{\ast}(x)$. It follows that the $\sigma$-algebras generated by these two random variables are identical and thus conditioning on $\pi(X)$ and conditioning on $\pi^{\ast }(X)$ are equivalent Heckman_Todd_2009. Hence,

align*[align* omitted — 173 chars of source]

Unfortunately, the averaging problem remains. We can see that \[ E[Y_{1}-Y_{0}]=\int\delta(\pi(x))g(x)dx\neq\int\delta^{\ast}(\pi^{\ast }(x))g^{\ast}(x)dx=E^{\ast}[Y_{1}-Y_{0}] \] because $g^{\ast}(x)\neq g(x)$ by ((ref)). Conditioning only works partially because it does not solve the problem of averaging. Therefore, we should be careful interpreting Heckman_Todd_2009's observation that \textquotedblleft matching and selection procedures can identify population treatment effects using misspecified estimates of propensity scores fit on choice-based samples\textquotedblright\ when estimating the ATE.

The problem remains if we use the inverse propensity score weighting (IPW). IPW identifies the ATE by \[ E[Y_{1}-Y_{0}]=E\left[ \frac{DY}{\pi(X)}-\frac{(1-D)Y}{1-\pi(X)}\right] , \] which is not identified by the sample counterpart \[ E\left[ \frac{DY}{\pi(X)}-\frac{(1-D)Y}{1-\pi(X)}\right] =E[\beta(X)]\neq E^{\ast}[\beta^{\ast}(X)]=E^{\ast}\left[ \frac{DY}{\pi^{\ast}(X)} -\frac{(1-D)Y}{1-\pi^{\ast}(X)}\right] \] again due to improper averaging, even though $\beta(x)=\beta^{\ast}(x)$.

If we know $\pi$, we can identify the ATE using reweighting based on ((ref))

align[align omitted — 345 chars of source]

If we condition on the propensity score, we can also identify the ATE through reweighting

align[align omitted — 388 chars of source]

In addition, the ATE can be identified by a reweighted version of IPW

equation[equation omitted — 224 chars of source]

In sum, conventional methods do not identify the ATE, but the problem can be overcome by reweighting the observations according to the population distribution of $X$.

Average Treatment Effect on the Treated (ATT)

Next we examine the ATT, $\gamma\equiv E[Y_{1}-Y_{0}|D=1]$. Note that by ((ref)) and ((ref))

align[align omitted — 150 chars of source]

The ATT is identified despite stratification. We can also identify the ATT conditioning on the propensity score because

align[align omitted — 173 chars of source]

Lastly, note that by iterated expectations \[ E[Y_{1}-Y_{0}|D=1]=\frac{1}{\pi}E\left[ \left( \frac{DY}{\pi(X)} +\frac{(1-D)Y}{1-\pi(X)}\right) \pi(X)\right] , \] whose sample counterpart identifies the ATT \[ \frac{1}{\pi^{\ast}}E^{\ast}\left[ \left( \frac{DY}{\pi^{\ast}(X)} +\frac{(1-D)Y}{1-\pi^{\ast}(X)}\right) \pi^{\ast}(X)\right] =E^{\ast} [\beta^{\ast}(X)|D=1]=E[\beta(X)|D=1]=E[Y_{1}-Y_{0}|D=1]. \] Hence, the ATT can be identified using IPW as well.

In sum, the ATT is identified by conventional methods. Stratification has no impact on the averaging because the distribution of $X$ given $D=1$ is the same in the population and stratified sample.

Asymptotic Distribution

ATE Estimators

Assuming that the population fraction $\pi$ is known, equation ((ref)) suggests that we can estimate the ATE $\beta$ by a semiparametric estimator $\hat{\beta}$ based on the moments

align[align omitted — 299 chars of source]

where $\beta_{d}^{\ast}(X)\equiv E^{\ast}[Y|D=d,X]$, $d=0,1$. Alternatively, equation ((ref)) suggests a semiparametric estimator using the propensity score

align[align omitted — 423 chars of source]

where $\delta_{d}^{\ast}(\pi^{\ast}(X))\equiv E^{\ast}[Y|D=d,\pi^{\ast}(X)]$, $d=0,1$.\footnote{The moment conditions ((ref)) and ((ref)) can be understood to be semiparametric generalization of the parametric model considered by Nevo2003, e.g.} Applying Newey_1994, we derive the influence functions of these ATE estimators, presented in Propositions (ref) and (ref).\footnote{All the proofs are provided in the Supplementary Appendix, which is available upon request.} Propositions (ref) and (ref) are both predicated on the assumption that the population fraction $\pi$ is known. The sample fraction $\pi^{\ast}$ may or may not be known, and the two propositions consider both cases.\footnote{The representations in ((ref)) and ((ref)) assume that $\pi^{\ast}$ is known. If $\pi^{\ast}$ is unknown, we can add one more moment $E^{\ast}[D-\pi^{\ast}]=0$ to each system.}

propositionIf $\pi^{\ast}$ is known, the influence function of the ATE estimator based on ((ref)) is \begin{align} & \left( D\frac{\pi}{\pi^{\ast}}+(1-D) \frac{1-\pi}{1-\pi^{\ast}}\right) (\beta_{1}(X)-\beta_{0}(X)) -\beta\nonumber\\ & +\left( \pi^{\ast}(X) \frac{\pi}{\pi^{\ast}}+(1-\pi^{\ast}(X)) \frac {1-\pi}{1-\pi^{\ast}}\right) \left( \frac{D}{\pi^{\ast}(X) }(Y-\beta_{1}(X)) -\frac{1-D}{1-\pi^{\ast}( X)}(Y-\beta_{0}(X)) \right) , \end{align} where $\beta_{d}(X)\equiv E[Y|D=d,X]$, $d=0,1$. If $\pi^{\ast}$ is estimated, the influence function is the sum of ((ref)) and \begin{equation} E^{\ast}\left[ \left( -\pi^{\ast}(X) \frac{\pi}{(\pi^{\ast})^{2}}+( 1-\pi^{\ast}(X)) \frac{1-\pi}{(1-\pi^{\ast})^{2}}\right) (\beta_{1} (X)-\beta_{0}(X)) \right] (D-\pi^{\ast}). \end{equation}
propositionThe influence function of the ATE estimator based on ((ref)) equals ((ref)) if $\pi^{\ast}$ is known, and equals the sum of ((ref)) and ((ref)) if $\pi^{\ast}$ is estimated.

The results mirror those in the non-stratified case, where various ATE estimators have the same asymptotic distribution. We can derive the asymptotic variance of the ATE estimators using ((ref)) if $\pi^{\ast}$ is known. Denote $\epsilon_{d}\equiv Y-\beta_{d}(X)$ and $\sigma_{d}^{2}(X) \equiv E^{\ast}[\epsilon_{d}^{2}|X]$, $d=0,1$, and $w(X)\equiv\pi^{\ast}(X) \frac{\pi}{\pi^{\ast}}+(1-\pi^{\ast}(X)) \frac{1-\pi}{1-\pi^{\ast}}$. The asymptotic variance of $\sqrt{n}(\hat{\beta}-\beta)$ is

equation[equation omitted — 268 chars of source]

ATT Estimators

Equations ((ref)) and ((ref)) serve as the basis for estimating the ATT. Because $\beta^{\ast}(X)=E^{\ast}[Y-\beta_{0}^{\ast}(X)|D=1,X]$, where $\beta_{0}^{\ast}(X)\equiv E^{\ast}[Y|D=0,X]$, it follows by iterated expectations that $E^{\ast}[\beta^{\ast}(X)|D=1]=E^{\ast}[Y-\beta_{0}^{\ast }(X)|D=1]=E^{\ast}[D(Y-\beta_{0}^{\ast}(X))]/\pi^{\ast}$. This suggests an estimator of $\gamma$ of the form

equation[equation omitted — 158 chars of source]

where $\hat{\beta}_{0}(\cdot)$ is a nonparametric estimator of $\beta _{0}^{\ast}(\cdot)$. This estimator subtracts the predicted outcome for the counterfactual from the outcome of a treated unit. Although intuitive, we are not aware of any references that establish the asymptotic distribution of such an estimator.

Define $\delta_{0}^{\ast}(\pi^{\ast}(X))\equiv E^{\ast}[Y|D=0,\pi^{\ast}(X)]$. A similar estimator can be constructed based on the propensity score

equation[equation omitted — 178 chars of source]

where $\hat{\delta}_{0}(\cdot)$ and $\hat{\pi}(\cdot)$ are nonparametric estimators of $\delta_{0}^{\ast}(\cdot)$ and $\pi^{\ast}(\cdot)$, respectively. Propositions (ref) and (ref) derive the asymptotic variances of these ATT estimators.

propositionIf $\pi^{\ast}$ is known, the asymptotic variance of $\sqrt {n}(\hat{\gamma}-\gamma)$ is equal to \begin{equation} E^{\ast}\left[ \frac{\pi^{\ast}(X)(\beta(X)-\gamma)^{2}}{(\pi^{\ast})^{2} }+\frac{\pi^{\ast}(X) \sigma_{1}^{2}(X)}{(\pi^{\ast})^{2}}+\frac{\pi^{\ast }(X)^{2}\sigma_{0}^{2}(X) }{(\pi^{\ast})^{2}(1-\pi^{\ast}(X))}\right] . \end{equation}
propositionIf $\pi^{\ast}$ is known, the asymptotic variance of $\sqrt{n}(\hat{\gamma}_{PS}-\gamma)$ is the same as in ((ref)).

The ATT estimators in ((ref)) and ((ref)) have the same asymptotic variance in ((ref)), which is equal to the asymptotic variance bound based on the stratified sample.\footnote{Note that ((ref)) is different from the asymptotic variance based on an unstratified population. While the ATT can be consistently estimated, the asymptotic variance is impacted by stratification.} See Hahn_1998. So both ATT estimators are efficient.\footnote{The fact that the ATE (or ATT) estimators that condition on either the propensity score or the covariates have the same asymptotic variance indicates that there is no efficiency gain from using the propensity score in stratified samples. HR_2013 come to the same conclusion in random sampling.}

Endogenous Treatment

In this section, we consider stratification on an endogenous treatment. We have a binary instrument $Z$ and assume no covariates to focus on essentials. AIR_1996 show that the local average treatment effect (LATE) is identified by the Wald ratio $(E[Y|Z=1]-E[Y|Z=0])/(E[D|Z=1]-E[D|Z=0])$. If the sample is stratified on the treatment status, we have

align*[align* omitted — 120 chars of source]

for $d=0,1$. However, because $\pi\neq\pi^{\ast}$, unconditional expectations in the sample do not match those in the population. We can see that

align*[align* omitted — 196 chars of source]

and similarly $E[D|Z=1]-E[D|Z=0]\neq E^{\ast}[D|Z=1]-E^{\ast}[D|Z=0]$. Therefore, the sample Wald ratio does not identify the population Wald ratio.

Observe that the population Wald ratio is the solution for $\beta$ in the moment condition \[ E\left\{

bmatrix[bmatrix omitted — 20 chars of source]

(Y-\alpha-\beta D)\right\} =0. \] Proposition (ref) suggests that we can identify the population Wald ratio using reweighting.

propositionFor any function $\varphi(Y,Z,D)$, we have \[ E[\varphi(Y,Z,D)]=E^{\ast}\left[ \left( D\frac{\pi}{\pi^{\ast}} +(1-D)\frac{1-\pi}{1-\pi^{\ast}}\right) \varphi(Y,Z,D)\right] . \]

Proposition (ref) implies that the population Wald ratio is the solution for $\beta$ in the reweighted sample moment condition \[ E^{\ast}\left\{ \left( D\frac{\pi}{\pi^{\ast}}+(1-D)\frac{1-\pi}{1-\pi ^{\ast}}\right)

bmatrix[bmatrix omitted — 20 chars of source]

(Y-\alpha-\beta D)\right\} =0. \] This fix requires knowledge of $\pi$.\footnote{Froelich_2007 considers a nonparametric estimator of LATE that allows for covariates.}

remarkIn the unconfounded case without $X$, i.e., under random assignment, stratification on $D$ does not pose any issue, because without $X$ there is no need for averaging. In the endogenous case, however, the bias is present even without $X$. The intuition can be found in the literature on stratification on endogenous variables (e.g, Hausman_Wise_1981).

Conclusion

For unconfounded treatment, stratification on the treatment status does not bias conditional treatment effects. It does bias the ATE, which can be corrected through reweighting if we know the fraction treated in the population. Stratification does not bias the ATT. In the case of endogenous treatment with an instrument, stratification biases the LATE. We propose a reweighted method to identify the LATE.