Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
31,619 characters · 8 sections · 28 citation commands
Heterogeneous Coefficients, Control Variables, and Identification of Multiple Treatment Effects
Keywords: Treatment effect; Multiple treatments; Heterogeneous coefficients; Control variable; Identification; Conditional nonsingularity; Propensity score.\
Models that allow for multiple treatments are important for program evaluation and the estimation of treatment effects (Catt:2010,Imai vD:2004,Imbens:2000,Graham Pinto 2018,Lechner:2001). A general class is heterogeneous coefficients models where the outcome is a linear combination of dummy treatment variables and unobserved heterogeneity. These models allow for multiple treatment regimes, with each dummy variable representing a different kind of treatment. These models also feature multidimensional heterogeneity, with the dimension of unobserved heterogeneity being determined by the number of treatment regimes.
Endogeneity is often a problem in these models because we are interested in the effect of treatment variables on an outcome, and the treatment variables are correlated with heterogeneity. Control variables provide an important means of controlling for endogeneity with multidimensional heterogeneity. For treatment effects, a control variable is an observed variable that makes heterogeneity and treatment variables independent when it is conditioned on (RR:1983).
We use control variables to give necessary and sufficient conditions for identification of average treatment effects based on conditional nonsingularity of the second moment matrix of the vector of dummy treatment variables given the controls. This first main result is familiar in the binary treatment case, but its generalization to multiple treatments appears to be new. With mutually exclusive treatments we find that, provided the heterogeneous coefficients are mean independent from treatments given the controls, a simple identification condition is that the generalized propensity scores (Imbens:2000) be bounded away from zero and that their sum be bounded away from one, with probability one. This condition is the same as common support, that the support of treatment variables conditional on the controls is equal to the marginal support of the treatment variables. Thus our second main contribution is to show that, with mutually exclusive treatments, conditional mean independence and common support are jointly sufficient for identification, a substantial weakening of the standard assumption that conditional independence hold jointly with common support (e.g., Frolich:2004). We also extend our analysis to distributional and quantile treatment effects, as well as corresponding treatment effects on the treated. These results provide an important generalization of RR:1983's classical identification result for binary treatments.
Let $Y$ denote an outcome variable of interest, and $D$ a vector of dummy variables $D(t)$, $t\in\mathcal{T}\equiv\{1,\dots,T\}$, taking value one if treatment $t$ occurs and zero otherwise, and $\beta$ a structural disturbance vector of finite dimension. We consider a heterogeneous coefficients model of the form
This model is linear in the treatment dummy variables, with coefficients $\beta$ that are random and need not be independent of $D.$
When the potential outcome framework of Rubin:1974 is extended to mutually exclusive treatment regimes in the definition of $D$, linearity in model ((ref)) arises naturally. Denote the vector of potential outcomes by $\{Y(1),\ldots,Y(T)\}^{\textrm{T}}$, and the potential outcome in the absence of treatment by $Y(0)$. The observed outcome $Y$ and the vector of potential outcomes are related by \[ Y=Y(0)+\sum_{t=1}^{\textrm{T}}D(t)\{Y(t)-Y(0)\}, \] which is of the form ((ref)) upon setting $\beta=(\beta_{0},\beta_{1},\ldots,\beta_{T})^{\textrm{T}}$ with $\beta_{0}\equiv Y(0)$ and $\beta_{t}\equiv Y(t)-Y(0)$, $t\in\mathcal{T}$.
Mutually exclusive treatment regimes are important in a wide variety of nonexperimental settings, such as program or policy evaluation with a multivalued treatment (e.g., Ao et al:2021,Lechner:2002,Uysal:2015). Consider for example evaluation of active labor market programs with a treatment taking on $T\geq2$ values, according to different types or levels of program participation. Central objects of interest are average effects on earnings for each treatment value, \[ E(\beta_{t})=E\{Y(t)-Y(0)\},\quad t\in\mathcal{T}. \] Allowing for multivalued treatments thus permits to capture different average effects across program types or levels, going beyond the sole effect of program participation considered in binary treatment analysis. Another important example is policy or medical treatment evaluation with non-mutually exclusive policies or treatments, implemented both separately and jointly. In that case, a distinct treatment dummy variable is assigned to each policy or treatment and to each implemented policy mix (Becker Egger:2013,Tortu et al:2020) or combined therapy (Feng et al:2012,Nian et al:2019). Resulting treatment regimes in the definition of $D$ are then mutually exclusive, by construction.
A main motivation for defining the components of $D$ according to mutually exclusive treatment regimes is the general validity of model ((ref)) in that case. In contrast, when treatment regimes are non-mutually exclusive, model ((ref)) restricts the average effect of any combination $\mathcal{C}$ of $K\leq T$ treatments $t_{1},\ldots,t_{K}\in\mathcal{T}$ to be additive in the average effects of each component of $\mathcal{C}$, i.e., the average treatment effect of $\mathcal{C}$ is
In addition to average effects of each treatment, model ((ref)) is then able to capture average effects of potentially complex interventions by restricting the form the average effects of combined treatments can take. The additivity restriction has been used in evaluation of randomized medical experiments implementing combinations of a large number of treatments ( see Petropoulou et al:2021, for a literature review), but does not appear to be common in nonexperimental settings.
In general, heterogeneous coefficients $\beta$ need not be independent of $D$ because of confounding factors denoted $X$. Here we assume that these factors are observable and that there is sufficient independent variation in $D$ from $\beta$ once conditioning on $X$. In the empirical examples above, $X$ includes a variety of individual characteristics of program participants, such as age, gender, measures of cognitive and non-cognitive skills, as well as socio-economic characteristics. Formally, we assume that the vector $\beta$ is mean independent of the endogenous treatments $D$, conditional on an observable control variable $X$.
The RR:1983 treatment effects model is included as a special case where $D\in\{0,1\}$ is a treatment dummy variable that is equal to one if treatment occurs and equals zero without treatment, and \[ p(D)=(1,D)^{\textrm{T}}. \] In this case $\beta=(\beta_{0},\beta_{1})^{\textrm{T}}$ is two dimensional with $\beta_{0}$ giving the outcome without treatment, and $\beta_{1}$ being the treatment effect. Here the control variables in $X$ would be observable variables such that Assumption (ref) holds, i.e., the coefficients $(\beta_{0},\beta_{1})$ are mean independent of treatment conditional on controls; this is the unconfoundedness assumption of RR:1983.
A central object of interest in model ((ref)) is the average structural function given by $\mu(D)\equiv p(D)^{\textrm{T}}E(\beta)$; see Cham:1984, Blundell Powell 2003 and Wool:2005. This function is also referred to as the dose-response function in the statistics literature (e.g., Imbens:2000). When $D\in\{0,1\}$ is a dummy variable for treatment, $\mu(0)$ gives the average outcome if every unit remained untreated and $\mu(1)$ the average outcome if every unit were treated, with $\mu(1)-\mu(0)$ being the average treatment effect. In general, the average effect of some treatment $t\in\mathcal{T}$ is \[ \mu(e_{t})-\mu(0_{T}), \] with $e_{t}=(0,\ldots,0,1,0,\ldots,0)^{\textrm{T}}$ defined as a $T$-vector with all components equal to zero, except the $t$th, which is one, and $0_{T}$ a $T$-vector of zeros. Pairwise average treatment effect comparisons are formed as $\mu(e_{t})-\mu(e_{s})$, for any $s,t\in\mathcal{T}$, $s\neq t$. For non-mutually exclusive treatment regimes, the average effect of some combination $\mathcal{C}$ of treatments $t_{1},\ldots,t_{K}\in\mathcal{T}$, $K\leq T$, is formed as $\sum_{s\in\mathcal{C}}\left\{ \mu(e_{s})-\mu(0_{T})\right\} $, $\mathcal{C}=\{t_{1},\ldots,t_{K}\}$, and the corresponding relative average effect with respect to some treatment $t\in\mathcal{T}$ as $\sum_{s\in\mathcal{C}}\left\{ \mu(e_{s})-\mu(e_{t})\right\} $.
The conditional mean independence assumption and the form of the structural function $p(D)^{\textrm{T}}\beta$ in ((ref)) together imply that the control regression function of $Y$ given $(D,X)$, $E(Y\mid D,X)$, is a linear combination of the treatment variables:
The average structural function can thus be expressed as a known linear combination of $E\{q_{0}(X)\}$ from equation ((ref)). By iterated expectations,
We use the varying coefficient structure of the control regression function ((ref)) and the implied linear form of $\mu(D)$ to give conditions that are necessary as well as sufficient for identification. For non-mutually exclusive treatment regimes, Appendix (ref) gives an example of a model for which the implied average structural function is of the linear form ((ref)) while the average effect of some combination of treatments takes the additive form ((ref)).
Under the maintained Assumption (ref), a sufficient condition for identification of the average structural function is nonsingularity of the second moment matrix of the treatment dummies given the controls, \[ E\left\{p(D)p(D)^{\textrm{T}}\mid X\right\}, \] with probability one. Under the additional assumption that $E\{p(D)p(D)^{\textrm{T}}\}$ is nonsingular, this condition is also necessary.
Theorem (ref) states our first main result. The proofs of all formal results are given in Appendix (ref).
When $D\in\{0,1\}$ and $p(D)=(1,D)^{\textrm{T}}$, the identification condition becomes the standard condition for the treatment effect model \[ Y=\beta_{0}+\beta_{1}D,\quad E\left(\beta\mid D,X\right)=E\left(\beta\mid X\right),\quad\beta\equiv(\beta_{0},\beta_{1})^{\textrm{T}}. \] The identification condition is that the conditional second moment matrix of $(1,D)^{\textrm{T}}$ given $X$ is nonsingular with probability one, which is the same as
with probability one, where $P(X)$ is the propensity score. Here we can see that the identification condition is the same as $0<P(X)<1$ with probability one, which is the standard identification condition.
Because $p(D)$ includes an intercept, the identification condition is the same as nonsingularity of the variance matrix $\text{var}(D\mid X)$ with probability one. This result generalizes ((ref)).
For mutually exclusive treatment regimes, the two standard assumptions for identification are common support, i.e., $\Pr(D=0_{T}\mid X)>0$ and $\Pr\{D(t)=1\mid X\}>0$ with probability one for each $t\in\mathcal{T}$, and conditional independence, i.e.,
cf., for instance, Frolich:2004 for a review. By conditional probabilities adding up to unity, Theorem (ref) shows that common support is equivalent to conditional nonsingularity, and hence is necessary as well as sufficient for identification under Assumption (ref). It follows that, provided common support holds, identification only requires conditional mean independence, and assumption ((ref)) is not necessary.
For non-mutually exclusive treatment regimes, conditional independence assumption ((ref)) and common support are not jointly necessary either. In that case, the marginal support $\mathcal{D}$ of $D$ has cardinality $\widetilde{T}>T$. Suppose there exist both a subset $\widetilde{\mathcal{D}}\subset\mathcal{D}$ of cardinality $T$ such that $E\{\mathbb{1}(D\in\widetilde{\mathcal{D}})p(D)p(D)^{\textrm{T}}\mid X\}$ is nonsingular with probability one, and a value $\overline{d}\in\mathcal{D}\backslash\widetilde{\mathcal{D}}$ such that $\Pr(D=\overline{d}\mid X)=0$ with positive probability. Then common support does not hold but $E\{p(D)p(D)^{\textrm{T}}\mid X\}$ is nonsingular with probability one, and hence $\mu(D)$ is identified under Assumption (ref). Therefore, Theorems (ref) and (ref) together establish that conditional independence assumption ((ref)) and common support are not jointly necessary for identification, for either type of treatment regime.
Theorem (ref) also clarifies the role played by conditional nonsingularity in identification. For mutually exclusive treatment regimes $\Pr\{D(t)=1\mid X\}= \Pr(D=e_{t}\mid X)$, $t\in\mathcal{T}$, and hence conditional singularity on a set with positive probability means that at least one of the events $\{D=0_{T}\}$,$\{D=e_{1}\}$,$\ldots$,$\{D=e_{T}\}$, has probability zero conditional on $X$ on that set. Therefore, in this instance, failure of identification occurs because common support does not hold. Using a directed acyclic graph (Pearl2009), Figure (ref) illustrates this failure with the event $\{D=e_{1}\}$ having probability zero given $X=x$, for almost every $x$ in a set with positive probability. The absence of an arrow between this event and $Y$ reflects failure of identification.
Our identification results are useful for the analysis of other interesting objects. When Assumption (ref) is strengthened to conditional independence
conditional nonsingularity is also sufficient for identification of distributional and quantile treatment effects. Define the distribution structural function $G(y,d)$ and, when $Y$ is continuous, the quantile structural function $Q(\tau,d)$ by \[ G(y,d)\equiv\text{Pr}\{p(d)^{\textrm{T}}\beta\leq y\},\quad Q(\tau,d)\equiv\tau^{\text{th}}\text{ quantile of }p(d)^{\textrm{T}}\beta, \] where $d$ is fixed in these expressions. Distributional and quantile treatment effects are formed as $G(y,e_{t})-G(y,0_{T})$ and $Q(\tau,e_{t})-Q(\tau,0_{T})$, respectively, for each $t\in\mathcal{T}$, and pairwise distributional and quantile treatment comparisons as $G(y,e_{t})-G(y,e_{s})$ and $Q(\tau,e_{t})-Q(\tau,e_{s})$, respectively, for any $s,t\in\mathcal{T}$, $s\neq t$.
With mutually exclusive treatment regimes, by Theorem (ref) the conditional support of $D$ given $X$ coincides with the marginal support of $D$, and hence the conditional support of $X$ given $D$ coincides with the marginal support of $X$, with probability one. Therefore, by ImbensNewey 2009 conditional nonsingularity and conditional independence property ((ref)) together imply identification of $G(Y,D)$ and, when $Y$ is continuous, also of $Q(\tau,D)$, from \[ G(Y,D)=\int F_{Y\mid DX}(Y\mid D,X=x)F_{X}(dx),\quad Q(\tau,D)=G^{-1}(\tau,D), \] where $F_{Y\mid DX}(Y\mid D,X)$ and $F_{X}(X)$ are the cumulative distribution functions of $Y$ given $(D,X)$ and of $X$, respectively, and $\tau\mapsto G^{-1}(\tau,D)$ denotes the inverse function of $y\mapsto G(y,D)$.
With non-mutually exclusive treatment regimes, conditional nonsingularity need not coincide with common support. Identification without common support can nonetheless be achieved under additional restrictions imposed on model ((ref)). When $Y$ is continuous and letting $Q_{\beta_{t}\mid DX}(u\mid D,X)$ denote the conditional quantile function of $\beta_{t}$ given $(D,X)$, $u\in(0,1)$, an example of sufficient model restrictions is that unobserved heterogeneity components $\beta_{t}$ satisfy conditional independence property ((ref)) as well as the additional scalar heterogeneity restriction
where the unobservable $U$ is the same for each $\beta_{t}$. The control quantile regression function of $Y$ given $(D,X)$ then takes the linear form
Other objects of interest include treatment effects on the treated. For some specified treatment $s\in\mathcal{T}$, average effects are formed using the average structural function for the treated, $\mu(D,e_{s})\equiv p(D)^{\textrm{T}}E(\beta\mid D=e_{s})$. Distributional and, when $Y$ is continuous, quantile treatment effects are formed using the distribution and quantile structural functions for the treated,
respectively, where $d$ is fixed in these expressions. These structural objects are useful for decomposition and counterfactual analysis (e.g., Ao et al:2021). The average effect of treatment $t$ on units treated with treatment $s$ is $\mu(e_{t},e_{s})-\mu(0_{T},e_{s})$, and distributional and quantile effects of treatment $t$ on units treated with treatment $s$ are $G(y,e_{t},e_{s})-G(y,0_{T},e_{s})$ and $Q(\tau,e_{t},e_{s})-Q(\tau,0_{T},e_{s})$, respectively.
Let $\mathcal{X}(s)$ denote the conditional support of $X$ given $D=e_{s}$. The average structural function for the treated can be expressed as a linear combination of $E\{q_{0}(X)\mid D=e_{s}\}$. By conditional mean independence and iterated expectations, \[ p(D)^{\textrm{T}}E\{q_{0}(X)\mid D=e_{s}\}=p(D)^{\textrm{T}}E\{E(\beta\mid X,D=e_{s})\mid D=e_{s}\}=\mu(D,e_{s}), \] and hence $\mu(D,e_{s})$ is identified if $q_{0}(X)$ is identified on $\mathcal{X}(s)$. Thus, for average treatment effects on the treated, the identification condition becomes nonsingularity of $E\{p(D)p(D)^{\textrm{T}}\mid X=x\}$ for almost every $x$ in the set $\mathcal{X}(s)$. This conditional nonsingularity condition is also necessary for identification of $\mu(D,e_{s})$ under the additional condition that $E\{p(D)p(D)^{\textrm{T}}\}$ is nonsingular.
If conditional independence property ((ref)) holds, the distribution and, when $Y$ is continuous, quantile structural functions for the treated also are identified, from \[ G(Y,D,e_{s})=\int F_{Y| DX}(Y| D,X=x)F_{X| D}(dx| D=e_{s}),\,Q(\tau,D,e_{s})=G^{-1}(\tau,D,e_{s}), \] respectively, where $\tau\mapsto G^{-1}(\tau,D,e_{s})$ denotes the inverse function of $y\mapsto G(y,D,e_{s})$. Here identification only requires the support of $X$ conditional on $D$ to contain $\mathcal{X}(s)$ with probability one, and hence that the support of $D$ conditional on $X=x$ be the same as the marginal support of $D$ for almost every $x\in\mathcal{X}(s)$. With mutually exclusive treatment regimes, this support condition is equivalent to nonsingularity of $E\{p(D)p(D)^{\textrm{T}}\mid X=x\}$ for almost every $x\in\mathcal{X}(s)$, by Theorem (ref). Therefore, this conditional nonsingularity condition is sufficient for identification. With non-mutually exclusive treatment regimes, this condition is also sufficient for identification of $q_{U}(X)$ on $\mathcal{X}(s)$, and hence of $Q_{Y\mid DX}(U\mid D,X)$ and $F_{Y\mid DX}(Y\mid D,X)$ on $\mathcal{D}\times \mathcal{X}(s)$, when the outcome $Y$ is continuous and the scalar heterogeneity restriction ((ref)) holds. Thus, results analogous to Theorem (ref) hold for distribution and quantile treatment effects on the treated.
The heterogeneous coefficients formulation we propose for multiple treatment effects reveals the central role of the conditional nonsingularity condition for identification. Because this condition is in principle testable, establishing that it is also necessary demonstrates testability of identification (e.g., Breusch:1986). With mutually exclusive treatments, the formulation of the equivalent common support condition in Theorem (ref) thus relates testability of identification to the generalized propensity scores. This is a generalization of the relationship between testability of identification and the propensity score in the binary treatment case.
Conditions that are both necessary and sufficient are also important for the determination of minimal conditions for identification. In an unpublished 2004 working paper (cemmap CWP03/04), Wooldridge considers a restricted version of our model with $E(D\mid \beta,X)=E(D\mid X)$ and $E\{p(D)p(D)^{\textrm{T}}\mid \beta,X\}=E\{p(D)p(D)^{\textrm{T}}\mid X\}$, and shows that $q_{0}(X)$ is identified if $E\{p(D)p(D)^{\textrm{T}}\mid X\}$ is invertible. The additional conditional second moments assumption implies that his identification condition differs from ours. Thus his result and proof do not apply in our setting which only assumes conditional mean independence $E(\beta\mid D,X)=E(\beta\mid X)$, and our results show that conditional second moments independence is not necessary for identification in multiple treatment effect models. Graham Pinto 2018 consider a related approach in work independent of the first version of this paper NeweyStouli:2018 where we derived our identification result (Lemma (ref) in the Appendix). The conditional nonsingularity condition we propose is weaker than their identification condition, and we study necessity as well as sufficiency for identification of average treatment effects.
We analyze the role of conditional nonsingularity for identification of multiple treatment effects under the maintained conditional mean independence Assumption (ref). Although itself not testable in general, this assumption is substantially weaker than the standard conditional independence property ((ref)). In the general case of mutually exclusive treatment regimes, the relaxation of conditional independence afforded by our heterogeneous coefficients approach is the same as the conditional mean independence condition
because the formulation of Assumption (ref) in terms of potential outcomes, \[ E\{Y(0)| D,X\}=E\{Y(0)| X\},\;\; E\{Y(t)-Y(0)| D,X\}=E\{Y(t)-Y(0)| X\},\;\; t\in\mathcal{T}, \] reveals that Assumption (ref) is equivalent to ((ref)). Therefore, our results allow applied researchers to replace unconfoundedness requirement ((ref)) for identification by the weaker condition ((ref)) under which conditional nonsingularity is both necessary and sufficient for identification, thereby improving robustness of empirical studies in nonexperimental settings. In particular, conditional mean independence property ((ref)) allows for any higher conditional moment of $Y(t)$ to depend on both $D$ and $X$. Our identification results are thus of general interest for the vast treatment effects literature (e.g., Ath Imbens:2017 for a recent literature review) and complement existing results on identification of treatment effects.