Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,668 characters · 9 sections · 0 citation commands
Endogeneity in Weakly Separable Models without Monotonicity
{
We identify and estimate treatment effects when potential outcomes are weakly separable with a binary endogenous treatment. \citeasnoun{vytlacilyildiz} proposed an identification strategy that exploits the mean of observed outcomes, but their approach requires a monotonicity condition. In comparison, we exploit full information in the entire outcome distribution, instead of just its mean. As a result, our method does not require monotonicity and is also applicable to general settings with multiple indices. We provide examples where our approach can identify treatment effect parameters of interest whereas existing methods would fail. These include models where potential outcomes depend on multiple unobserved disturbance terms, such as a Roy model, a multinomial choice model, as well as a model with endogenous random coefficients. We establish consistency and asymptotic normality of our estimators.}
{ JEL Classification: C14, C31, C35 \newline Key Words Weak Separability, Treatment Effects, Monotonicity, Endogeneity}
\setcounter{equation}{0}
Consider a weakly separable model with a binary endogenous variable:
where $( v_1(X,D),v_2(X,D),...v_J(X,D) ) \equiv v(X,D) $ is a $J$-vector of unknown linear or nonlinear indices in the outcome equation (1.1) and $D$ is a binary endogenous variable defined by equation (1.2). Here $ X\in \mathbb{R}^{d_{x}}$ and $Z\in \mathbb{R}^{d_{z}}$ are vectors of observable exogenous variables, which may have overlapping elements. This paper is about the identification and estimation of the average treatment effect (ATE).\footnote{ The ATE is one of several parameters of interest in the program evaluation literature. In their seminal work, \citeasnoun{imbensangrist} focus on a different parameter of interest, the Local Average Treatment Effect (LATE) which is the ATE for the subset of the population referred to as compliers. However, as pointed out in \citeasnoun{HV05}, \citeasnoun{heckvyt-handbook}, \citeasnoun{CHV10}, and \citeasnoun{santos18}, there are settings where the LATE by itself is not always the focal point of policy studies.} Similar to conditions in \citeasnoun{vytlacilyildiz}, we require that there is some continuous element in $Z$ excluded from $X$ so that $\theta(Z)$ varies continuously conditional on $X$, and that we can vary $X$ after conditioning on $\theta(Z)$. \footnote{ This does not necessarily require a second exclusion restriction that there be an element in $X$ not in $Z$. As explained in \citeasnoun{vytlacilyildiz}, what is required is that the element in $Z$ not in $X$ be continuously distributed.} It is worth noting that the method we propose also applies directly in a more general setup where the error in the outcome equation is specific to $D$, i.e., with $\varepsilon$ replaced by $\varepsilon_D$, provided the marginal distribution of $\varepsilon_d$ is the same for $d\in\{0,1\}$.
The literature on program evaluation abounds in examples in which self-selected, endogenous treatment depends on exogenous instruments that have no direct impact on potential outcomes. For instance, \citeasnoun{vytlacilyildiz} provided an example where $D$ indicates an individual's enrollment in job training at date 0, and $Y$ is employment status at date 1. In this case, $Z$ may contain variables summarizing local labor market conditions at date 0. They noted these variables affect the opportunity cost of training, and consequently the training decision at date 0, but do not directly affect the individual's employment at date 1.
In the system of equations above, $U$ is the unobservable random variable normalized to follow the standard uniform distribution, and the error term $\varepsilon$ in the outcome equation is allowed to be a random vector with a known dimension. We assume $\left( X,Z\right) $ are independent of $ (\varepsilon ,U)$. Note that we allow $v(X,D)$ to be a vector of multiple indices, whereas existing methods, such as \citeasnoun{vytlacilyildiz}, can only be applied when it is a single index. \citeasnoun{HV05} and \citeasnoun{carneiroLee2009} maintained that $(\varepsilon,U)$ is independent of $Z$ conditional on $X$, and used local instrumental variables to estimate marginal treatment effect (MTE) for realized values of $U$ over the support of $\theta(Z)$ given $X$. In their case, integrating out MTE to ATE would require the support of $\theta(Z)$ given $X$ to be the complete unit interval, i.e. $[0,1]$. In comparison, we propose a method that relaxes such a support condition by exploiting more structure in the potential outcome equation and its full distribution. Moreover, while \citeasnoun{carneiroLee2009} use $E[1\{Y\leq y\}|X,D,U=p]$ for each fixed $y$ to recover the conditional distribution of potential outcomes at $y$, we use the full distribution of the observed outcomes to find the matching covariates. This helps us to recover ATE without the aforementioned full support condition. As in \citeasnoun{vytlacilyildiz}, our method requires that at least some covariates in the potential outcome equation are exogenous.
Since \citeasnoun{vytlacilyildiz}, important work has considered identification and estimation of similar models, but under alternative conditions, notably on the supports of $D$, and $\theta(Z)$ as well as the dimension of $\varepsilon$. \citeasnoun{neweyimbens} assume $D$ is continuous and monotone in the error term, $\theta(Z)$ is continuous with large support, and $\varepsilon$ is scalar. \citeasnoun{kasy14} allows $ \varepsilon$ to be a vector as we do here but imposes that both $D$ and $ \theta(Z)$ are continuously distributed and further assumes a monotone relationship between the two. \citeasnoun{dhault2014} and \citeasnoun{torgo2014} assume $D$ is continuous, $\varepsilon$ is scalar, and assume monotonicity in the error term in each of the two equations, though they allow for $\theta(Z)$ to be discrete. \citeasnoun{vuongxu} assume $D$ is discrete and $\theta(Z)$ is continuous, restricts $\varepsilon$ to be scalar and requires $Y$ in ((ref)) to be monotone in $ \varepsilon$. \citeasnoun{JunEtal2016} study an extension of \citeasnoun{vytlacilyildiz}, and also require the same monotonicity condition as in the latter. \citeasnoun{fengjun} shows how to identify nonseparable triangular models where the endogenous variable is discrete but restricted to have larger support than the instrument variable. \footnote{ These papers focus on point identification. For partial identification of a model with a binary outcome, see \citeasnoun{shaikhvytlacil11} and \citeasnoun{mourifie15}. \citeasnoun{santos18} explore partial identification of relevant treatment effect parameters in models without structure imposed on the outcome equation. In such settings, attaining point identification requires support conditions on the propensity score function that are stronger than imposed here.}
As in the conventional framework, two potential outcomes $Y_{1}$ and $Y_{0}$ satisfy
We only observe $\left( Y,D,X,Z\right) $, where $Y=DY_{1}+(1-D)Y_{0}$. In this model, we do not impose parametric distribution on the error terms $(\varepsilon,U)$ or a linear index structure. Similarly, \citeasnoun{vytlacilyildiz} also do not require such restrictions, but assume that $v(X,D) \in \mathbb{R}$ is a single index, and
We do not impose any such monotonicity structure. That is because our approach exploits the full information in the distribution of the outcome variable, instead of just its mean. Indeed, when the outcome distribution function is more informative than the mean, our method is applicable to more general settings; in particular, not only do we not rely on such a monotonicity assumption, but we also allow for multiple indices.
In Section (ref) we present the identification argument and discuss its required conditions. In Sections (ref) and (ref) we provide some examples in which such a monotonicity condition fails, but the average effect of the binary endogenous variable is still identified. To reiterate, we allow the potential outcomes to be weakly separable in multiple indices, that is, $ v(X,D)=(v_{1}(X,D),v_{2}(X,D),...,v_{J}(X,D)) \in \mathbb{R}^J$. We consider the identification and estimation of the average treatment effect of $D$ on $Y$, $E(Y_{1}|X\in A)$, $E(Y_{0}|X\in A)$, and $ E(Y_{1}-Y_{0}|X\in A)$, for some set $A$, without the aforementioned monotonicity. Indeed, for the case with multiple indices $v(X,D) \in \mathbb{ R}^J$, the monotonicity condition is no longer well defined.
\setcounter{equation}{0}
Our identification strategy is based on the notion of matching covariates. Let $Supp(V)$ denote the support of a generic random vector $V$. Consider the identification of $E(Y_{1}|X=x)$ for some $x\in Supp(X)$. Under the assumption that $ (\varepsilon ,U)\bot (X,Z)$,
where $P(z)\equiv E(D|Z=z) = \theta(z)$ because $U$ is normalized to follow a standard uniform distribution. The only term not directly identifiable on the right-hand side of ((ref)) is:\footnote{ If the support of $P(Z)$ given $x$ covers the full closed interval $[0,1]$, then $E(Y_1|X=x)$ is directly identified as $E(Y|D=1,X=x,Z=z)$ at $z$ s.t. $P(z)=1$. However this means point identification hinges on the event “$P(Z)=1$ given $x$”. }
The main idea behind our approach would at first seem to be similar to that of \citeasnoun{vytlacilyildiz} , which is to use exogenous variations of covariates in the outcome equation to find some $(\tilde{x},\tilde{z})\in Supp(X,Z)$ such that
so that
However, unlike \citeasnoun{vytlacilyildiz}, we utilize the full distribution of $Y$ (rather than its first moment only) while searching for such pairs of $(x, \tilde{x})$ in ((ref)). This allows us to relax the single-index and monotonicity conditions.
To fix ideas, let there be continuous components in $Z$ that are excluded from $X$. For any $p$ on the support of $P(Z)$ given $X=x$, and for any $y$, define
where
with $v(x,d)$ being a realized index at $X=x$ and the expectation in the definition of $F_{g|u}$ is with respect to the distribution of $\varepsilon $ given $U=u$. The last equality in ((ref)) holds because of independence between $(\varepsilon ,U)$ and $(X,Z)$. Assume that for all $y,x$ and $d \in \{ 0,1 \}$, the function $F_{g|u}(y;v(x,d))$ is continuous in $u$ over $[0,1]$. This implies $h^*_1$ is differentiable in $p$. By construction, $h_{1}^{\ast }(x,y,p)$ is directly identified from the joint distribution of $(D,Y,X,Z)$ in the data-generating process. Furthermore, for any pair $p_{1}>p_{2}$, define:
Likewise, define
and let
Let $\mathcal{P}_{x}$ denote the support of $P(Z)$ given $X=x$; let $ Int( \mathcal{P}_{x}\cap \mathcal{P}_{\tilde{x}} )$ denote the interior of the intersection of $\mathcal{P}_{x}$ and $\mathcal{P}_{\tilde{x}}$. Similar to \citeasnoun{HV05} and \citeasnoun{carneiroLee2009}, we use continuous, exogenous variation in $Z$ to establish here that for any $ x \neq \tilde{x}$ with $ Int( \mathcal{P}_{x}\cap \mathcal{P}_{\tilde{x}} ) \neq \emptyset$ and any $y$,
if and only if
Sufficiency of ((ref)) is immediate from the definition of $h_{1}$ and $h_{0}$. To see its necessity, note that for all $p>p^{\prime }$ on $Int(\mathcal{P}_{x}\cap \mathcal{P}_{ \tilde{x}})$,
{and}
Thus ((ref)) and ((ref)) are equivalent.
Next, we collect identifying assumptions as follows:
ASSUMPTION A-1: The distribution of $U$ is absolutely continuous with respect to Lebesgue measure, and is normalized to standard uniform, and $\theta(z)$ in ((ref)) is continuous in $z$.
ASSUMPTION A-2: The random vectors $(U,\varepsilon )$ and $(X,Z)$ are independent. There exists at least one continuously distributed component in $Z$ that is excluded from $X$.
ASSUMPTION A-3: Both $g(v(X,1),\varepsilon )$ and $g(v(X,0),\varepsilon )$ have finite first moments conditional on $U=u$ for all $u \in [0,1]$.
ASSUMPTION A-4: For all realized values of $y,x$ and $d\in\{0,1\}$, $E[1\{g(v(x,d),\epsilon) \leq y\}|U=u] $ is continuous in $u$ over $[0,1]$.
ASSUMPTION A-5: For any $x\neq \tilde{x}$ such that $Int(\mathcal{P}_{x}\cap \mathcal{P}_{\tilde{x}})$ is nonempty, $F_{g|p}(y;v(x,1))=F_{g|p}(y;v(\tilde{x},0))$ for all $y$ and $p\in Int(\mathcal{P}_{x}\cap \mathcal{P}_{\tilde{x}})$ if and only if $v(x,1)=v(\tilde{x},0)$.
Assumptions A-1 and A-3 are common regularity conditions in the literature. Assumption A-2 consists of a standard condition of instrument exogeneity. It allows $P(Z)$ to have continuous, exogenous variation conditional on $X$. The continuity of instrument is also common in the literature on treatment effects. Examples include \citeasnoun{heckmanvytlacil} and \citeasnoun{vytlacilyildiz}. As noted above, Assumption A-4 ensures the identifiable functions $h^*_1(x,y,p),h^*_0(x,y,p)$ are both differentiable in the last argument $p$.
It is important to note that Assumption A-5 allows the support of $P(Z)$ to be a strict subset of the interval $(0,1)$. This is because our method does not use an identification-at-infinity argument, which would require $P(Z)$ to have full support [0,1] given $X$ in order to identify $E(Y_0|X)$ directly. (See Footnote 4 above.) We also note that Assumption A-5 relaxes two limitations of Assumption 4 in \citeasnoun{vytlacilyildiz}. Specifically, to identify pairs $(x,\tilde{x})$ with $v(x,1)=v(\tilde{x},0) $, \citeasnoun{vytlacilyildiz} rely on two assumptions that $v(X,D) \in \mathbb{R}$ is a single index and that $E\left[ g(v(x,1),\varepsilon )|U=u\right] =E\left[ g(v(\tilde{x},0),\varepsilon )|U=u\right] $ if and only if $v(x,1)=v(\tilde{x},0) $. The latter holds under a maintained assumption that $E\left[ g(v(x,d),\varepsilon )|U=p\right] $ is a strictly monotonic function of $ v(x,d)$. In comparison, we construct an identification strategy without the single index and monotonicity restrictions by matching conditional distributions $F_{g|p}(\cdot ;v(x,1))$ and $\ F_{g|p}(\cdot ;v(\tilde{x},0))$.
The role of Assumption A-5 in our method can be illustrated by drawing an analogy with a standard identifying condition in nonlinear regression: $ Y = f(X,\theta ^*) + \varepsilon$. In this case, identification of $\theta^*$ requires: “$f(x,\theta^* ) = f(x,\theta)$ for all $x$ if and only if $\theta^* =\theta$ ”. To see how this is related to Assumption A-5 in our setting, let $ \theta_1, \theta_0$ be shorthand for $v(x,1), v(\tilde{x},0)$ respectively, and let $m(y,p,\theta) \equiv F_{g|p}(y,\theta )$. Then our method requires: “$m(y,p,\theta _{1})=m(y,p,\theta _{0})$ for all $(y,p)$ if and only if $\theta _{1}=\theta _{0}$ ”.
It is worth emphasizing that Assumption A-5 only presents our identification condition in the weakest form, for the sake of generality. Later in Section (ref), we exploit the structures embedded in specific examples to show how Assumption A-5 can be satisfied under intuitive, mild primitive conditions.
Next, we specify conditions under which such pairs of covariates exist on the support. Define $\mathcal S \equiv \{(x,\tilde{x}):v(x,1)=v(\tilde{x},0)\} $ and $ \mathcal T \equiv \{(x,\tilde{x}): \exists z, \tilde{z} \text{ with } (x,z),(\tilde{x},\tilde{z})\in Supp(X,Z) \text{ and } P(z)=P(\tilde{z}) \in Int(\mathcal{P}_x\cap \mathcal{P}_{\tilde{x}})\}$. Let $ \mathcal X^1 \equiv \{x:\exists \tilde{x} \text{ with }(x,\tilde{x})\in \mathcal S \cap \mathcal T\}$ and $ \mathcal X^0 \equiv \{x:\exists \tilde{x} \text{ with }(\tilde{x},x)\in \mathcal S \cap \mathcal T\}$.
ASSUMPTION A-6: $\Pr (X\in \mathcal X^1)>0$ \ and $\Pr (X\in \mathcal X^0)>0$.
This condition is similar to Assumption A-4 in \citeasnoun{vytlacilyildiz}. By definition, for each $x\in\mathcal X^1$, we can find $\tilde{x}\in Supp(X)$ such that there exists $(z,\tilde{z}) $ with $ (x,z),(\tilde{x},\tilde{z}) \in Supp(X,Z)$, $v(x,1)=v(\tilde{x},0)$, and $P(z)=P(\tilde{z})$. This requires variation in $ x $ while holding $P(Z)$ fixed. As noted earlier (Footnote 2), this does not necessarily require a “second exclusion restriction” that there is an element in $X$ that is excluded from $Z$. In general, there exist multiple such values of $\tilde{x}$ that can be matched with this $x$. Hence for each $x\in\mathcal X^1$, we define the set of such matched values as
and define $ \mathcal{P}_0^*(x) \equiv \bigcup_{\tilde{x} \in \lambda_0(x)}(\mathcal{P}_{\tilde{x}}\cap \mathcal{P}_x)$. Note that by definition of $\mathcal{X}^1$, the set $\mathcal{P}_x \cap \mathcal{P}_{\tilde{x}}$ must be non-empty for $x \in \mathcal{X}^1$ and $ \tilde{x} \in \lambda_0(x)$ under our maintained assumptions. Likewise, by symmetry, for each $x\in\mathcal X^0$, we define
and let $ \mathcal{P}_1^*(x) \equiv \bigcup_{\tilde{x} \in \lambda_1(x)}(\mathcal{P}_{\tilde{x}}\cap \mathcal{P}_x)$.
The theorem below shows how to use such matched values to identify the conditional mean of potential outcomes.
It is worth mentioning that the conditions for this theorem do not directly restrict the support of potential outcomes. As we show in the next section, our method applies in important applications where the potential outcome is either discrete (e.g., determined by multinomial choices), or multi-dimensional with both discrete and continuous components (e.g., determined in a Roy model).
\setcounter{equation}{0}
In this section, we present several examples in which the potential outcomes depend on multi-dimensional indices with an endogenous treatment. More importantly, we show how the specific structure embedded in each application naturally leads to transparent, primitive conditions that imply Assumption A-5, thus corroborating the wide applicability of this general approach.
In the first and third examples, the monotonicity condition in \citeasnoun{vytlacilyildiz} does not hold; in the second example, the identification requires a generalization of the monotonicity condition into an invertibility condition in higher dimensions.
Example 1. (Heteroskaedastic shocks in Uncensored or Censored outcomes) First, consider a triangular system where a continuous uncensored outcome is determined by double indices $v(X,D)\equiv (v_{1}(X,D),v_{2}(X,D))$:\footnote{ \citeasnoun{abrevaya2019estimation} adopted a different approach to identify ATE in the same model with uncensored, continuous outcome and multiplicative, heteroskaedastic shocks above, using a binary instrument $Z$. They showed the conditional mean of the observed outcome $Y$, when scaled by the conditional covariance of $Y1\{D=d\}$ and $Z$, is a weighted sum of the conditional means of potential outcomes: $v_1(x,0)$ and $v_1(x,1)$. Thus, using the means of the scaled $Y$ conditional on $Z = 0,1$ respectively, one can construct a linear system that identifies $v_1(x,0)$ and $v_1(x,1)$, provided the proper rank condition holds. Their method leverages this particular specification of multiplicative endogeneity, but does not generalize to the case of censored outcomes or in the other examples we consider later. }
The selection equation determining the actual treatment is the same as (1.2). In this case the concept of monotonicity in $v\in \mathbb{R}^{2}$ is not well-defined, so the procedure proposed in \citeasnoun{vytlacilyildiz} is not suitable here.\footnote{ For this particular design, the approach proposed in \citeasnoun{vuongxu} should be valid. But it will not be for a slightly modified model, such as $ Y=v_1(X,D)+(e_2+v_2(X,D)*e_1)$, whereas ours will be.} Nevertheless, we can apply the method in Section (ref)\ to identify the average treatment effect by using the distribution of outcomes to find pairs of $x$ and $\tilde{x}$ such that $v(x,1)=v(\tilde{x},0)$.
To see how Assumption A-5 holds, assume the range of $v_2(\cdot)$ is positive. Note that
for $d=0,1$. Suppose the distribution of $\varepsilon $ conditional on $u$ is increasing over $\mathbb{R}$. Then for all $y$ and $x \in \mathcal{X}^1$ and matched values $\tilde{x} \in \lambda_0(x)$,
Differentiating with respect to $y$ yields $v_{2}(x,1)=v_{2}(\tilde{x},0)$, which implies $v_{1}(x,1)=v_{1}(\tilde{x},0)$.
Next, consider a similar model with continuous censored outcomes, where
In this case, our identification argument for the uncensored case above applies almost immediately. The only adjustment needed is to confine the values of $y$ used for constructing matched pairs $(x,\tilde{x})$ over the uncensored segment, i.e. $y > 0$. In contrast, other methods for identifying ATE in the uncensored model which rely strictly on the linear structure with multiplicative heterogeneity, such as \citeasnoun{abrevaya2019estimation}, no longer applies.
Example 2. (Multinomial potential outcome) Consider a triangular system where the outcome is multinomial. The multinomial response model has a long and rich history in both applied and theoretical econometrics. Recent examples in the semiparametric literature include \citeasnoun{lee95}, \citeasnoun{pakesporter2014}, \citeasnoun{ahnetal2018}, \citeasnoun{shietal2018}, and \citeasnoun{KOT2019}. None of those papers allow for dummy endogenous variables or potential outcomes.
In this example, let the observed outcome be
where
In this case, the index $v\equiv (v_{j})_{j\leq J}$ and the errors $ \varepsilon \equiv (\varepsilon _{j})_{j\leq J}$ are both $J$-dimensional. The selection equation that determines $D$ is the same as (1.2). In this case, we can replace $1\{Y\leq y\}$ by $1\{Y=y\}$ in the definition of $ h_{1},h_{0},h_{1}^{\ast },h_{0}^{\ast }$ and $F_{g|u}(\cdot ;v)$. Then for $ d=0,1$ and $j\leq J$,\
By \citeasnoun{ruud2000} and \citeasnoun{ahnetal2018}, the mapping from $ v\in \mathbb{R}^{J}$ to $(F_{g|u}(j;v):j\leq J)\in \mathbb{R}^{J}$ is smooth and invertible provided that $\varepsilon \in \mathbb{R}^{J}$ has non-negative density everywhere. This implies Assumption A-5.
Example 3. (Potential outcome from a Roy model) Consider a treatment effect model with an endogenous binary treatment $D$ and with the potential outcome determined by a latent Roy model. The Roy model has also been studied extensively from both applied and theoretical perspectives. See, for example, the literature survey in \citeasnoun{heckvyt-handbook} and the seminal paper in \citeasnoun{heckmanhonoreb}.
Here the observed outcome consists of two pieces:\ \ a continuous measure $ Y=DY_{1}+(1-D)Y_{0}$ and a discrete indicator $W=DW_{1}+(1-D)W_{0}$ for $ d=0,1$. These potential outcomes are given by
where $a$ and $b$ index potential outcomes realized in different sectors, with
The binary endogenous treatment $D$ is determined as in equation (1.2). For example, $D\in \{1,0\}$ indicates whether an individual participates in a professional training program, $W_{d}\in \{a,b\}$ indicates the potential sector in which the individual is employed, $ y^*_{j,d}$ is the potential wage from sector $j$ under treatment $D=d$, and $Y_{d}\in \mathbb{R}$ is the potential wage if the treatment status is $D=d$.\footnote{ Note this differs from the classical approach that formulates treatment effects using Roy models (e.g. \citeasnoun{heckmanUrzuaVytlacil2006}) in that the potential outcome $Y_d$ itself is determined by a latent Roy model.}
As before, we maintain that $(X,Z)\bot (\varepsilon ,U)$. The parameter of interest is
By the independence condition that $(X,Z)\bot (\varepsilon ,U)$ and an application of the law of total probability, this conditional probability can be expressed in terms of directly identifiable quantities and the following counterfactual quantity
Again, we seek to identify this counterfactual quantity by finding $\tilde{x}$ such that there exists $\tilde{z} $ with $ (\tilde{x},\tilde{z}) \in Supp(X,Z)$, $P(z) = P(\tilde{z}) $, and
This would allow us to recover the counterfactual conditional probability in ((ref)) as
To find such a pair of $(x,\tilde{x})$, define $h_{d,W}(x,p,p^{\prime }),h_{d,W}^{\ast }(x,p)$ by replacing $1\{Y\leq y\}$ with $1\{W=a\}$ in the definition of $h_{d},h_{d}^{\ast }$ in Section (ref). Similarly, define $h_{d,Y}(x,y,p,p^{\prime }),h_{d,Y}^{\ast }(x,y,p)$ by replacing $1\{Y\leq y\}$ with $1\{Y\leq y,W=a\}$ in the definition of $ h_{d},h_{d}^{\ast }$ in Section (ref). Then
and $h_{d,W}(x,p_{1},p_{2})$ and $h_{d,Y}(x,y,p_{1},p_{2})$ are both identified over their respective domains by construction.
Assume $(\varepsilon _{a},\varepsilon _{b})$ is continuously distributed with positive density over $\mathbb{R}^{2}$ conditional on all $u$, and the distribution of $ (\varepsilon_0,\varepsilon_1) $ given $U=u$ is continuous in $u $ over $[0,1]$. Then the statement
holds true if and only if ((ref)) holds. To see this, first note that matching $ h_{1,W}(x,p,p^{\prime })=h_{0,W}(\tilde{x},p,p^{\prime })$ requires
while matching $h_{1,Y}(x,y,p,p^{\prime })=h_{0,Y}(\tilde{x},y,p,p^{\prime }) $ at the same time requires
Thus requiring ((ref)) and ((ref)) to hold jointly is equivalent to ((ref)).
It is worth mentioning that the identification strategy above also applies in a more general setup where the error term in the potential outcome is specific to the treatment, i.e., with $\varepsilon_j$ replaced by $\varepsilon_{j,d}$ in $y^*_{j,d}$. In such cases, the argument above remains valid under a “rank similarity” condition (that the marginal distribution of $\varepsilon_{j,d}$ is the same for $d\in\{0,1\}$). The rank similarity condition has been used for identifying treatment effects in instrumental quantile regression models such as \citeasnoun{chernozhukovhansenjoe}.
\setcounter{equation}{0}
The identification strategy we use requires finding matched pairs for $x$ in $\mathcal{X}^1$ and $\mathcal{X}^0$. In some cases, with the outcome being continuous, we can construct similar arguments for identifying a counterfactual quantity in a treatment effect model by matching different elements on the support of continuous outcomes. To the best of our knowledge, this approach has not been explored in the literature on the effects of endogenous treatments. The following example illustrates this point.
Example 4. (Potential outcome with random coefficients) Random coefficient models are prominent in both the theoretical and applied econometrics literature. They permit a flexible way to allow for conditional heteroscedasticity and unobserved heterogeneity. For a survey and recent developments, see \citeasnoun{hsiaochapter}, \citeasnoun{hoderlein2010analyzing}, \citeasnoun{arellano2012identifying}, and \citeasnoun{masten2018random}.
We consider a treatment effect model where the potential outcome is determined through random coefficients:
and the binary endogenous treatment $D$ is determined as in equation (1.2). The random intercepts $\alpha _{d}\in \mathbb{R}$ and the random vectors of coefficients $\beta _{d}$ are given by
where for any $x \in Supp(X)$ and $d \in \{0,1\}$, $(\bar{\alpha}_{d}(x),\bar{\beta} _{d}(x))\in \mathbb{R}^{K+1}$ is a vector of constant parameters while $\eta _{d}\in \mathbb{R}$ and $\varepsilon _{d}\in \mathbb{R}^{K}$ are unobservable noises.
As before, suppose some elements of $Z$ in the treatment equation are excluded from $X$. We allow the vector of unobservable terms $(\varepsilon_{1},\varepsilon _{0},\eta _{0},\eta _{1},U)$ to be arbitrarily correlated, and assume:
Aslo, normalize the marginal distribution of $U$ to standard uniform, so that $\theta (Z)$ is directly identified as $P(Z)\equiv E(D|Z)$.
Our goal is to identify the distribution of potential outcomes $ Y_{d}$ given $X=x$ for $d=0,1$. From this result, we can recover other parameters of interest such as average treatment effects, quantile treatment effects, etc. As a preliminary step, we start by pinpointing a counterfactual item that is crucial for this identification question.
Let $G_{P|x}$ denote the distribution of $P\equiv P(Z)$ given $X=x$ (recall that its support is denoted as $\mathcal{P}_x$). This conditional distribution is directly identifiable from the data-generating process. By construction,
where
The first term on the right-hand side of ((ref)) is identified as
The second term on the right-hand side of ((ref)), denoted by $\phi_0(x,y,p)$ , is counterfactual and can be written as
Hence identification of the conditional distribution of $Y_1$ amounts to identification of $\phi_0(\cdot) $.
As noted at the beginning of this section, we will identify the counterfactual $\phi_0(\cdot)$ by finding matched elements on the support of the observed outcome $Y$. This takes two steps. First, we show that if for a given pair $(x,y)$ one can find $t(x,y)$ such that
then one can use $t(x,y)$ to identify $\phi_0(x,y,p)$ for any $p\in \mathcal{P}_x$. Specifically, for any $p$ on the support of $P$ given $X=x$, define
where the second equality uses ((ref)). Likewise, under ((ref)) we have:
Assume\footnote{ Such a distributional equality condition has been used to motivate the rank similarity condition imposed frequently in the econometrics literature -- see, for example, \citeasnoun{chernozhukovhansen}, \citeasnoun{vytlacilyildiz}, \citeasnoun{chenstackhan}, \citeasnoun{franlef}, \citeasnoun{dongshen}.}
Under ((ref)), we have
Then by definition of $t(x,y)$ in ((ref)),
Thus the counterfactual $\phi _{0}(x,y,p)$ would be identified as $h_{0}^{\ast }(x,t(x,y),p)$.
The second step is to show that for each pair $(x,y)$ we can indeed uniquely recover $t(x,y) $ using quantities that are identifiable in the data-generating process. To do so, we define two auxiliary functions as follows: for $p_{1}>p_{2}$ on the support of $P$ given $X=x$, let
and
Suppose $\eta _{d}+x^{\prime }\varepsilon _{d}$ is continuously distributed over $\mathbb{R}$ for all values of $x$ conditional on all $u \in [0,1]$. Then for any fixed pair $(x,y)$ and $p_{1}>p_{2}$,
if and only if
\setcounter{equation}{0}
In this section, we outline estimation procedures from a random sample of the observed variables that are motivated by our identification results. We first describe an estimation procedure for the parameter $E[Y_1]$ in the first three examples. Recall $\mathcal{P}_{x}$ denotes the support of $P(Z) \equiv P $ given $X=x$. Let $f_P(.|x)$ denote the density of $P(Z)$ given $X=x$, and define
For simplicity, assume
Define a measure of distance between $h_{1}(x_{1},\cdot )$ and $ h_{0}(x_{0},\cdot )$ as follows:
where $w(y)$ is a chosen weight function.
Consider the case when $h_{0}(x,y,p_{1},p_{2})$, $h_{1}(x,y,p_{1},p_{2})$ and $P(z)$ are known. For any given $x_{i}$, let $\tilde x_{i}$ be such that
which, under Assumption A-5 in Section (ref), is equivalent to
Let $P_i$ be shorthand for $P(Z_i)$ and define
This conditional expectation equals $E[ Y | D=0, v( X , 0 ) = v( X_i , 1), P=P_i] $, which in turn equals $E( Y_1 | D=0, X=X_i, P=P_i) $.
The parameter of interest $\Delta \equiv E[Y_{1}]$ can be written as $ \Delta = E[D_iY_i+(1-D_i)Y^*_i]$. Therefore, we estimate $\Delta$ by its sample analog, after replacing $Y^*_i$ with its Nadaraya-Watson estimates. That is, we estimate $\Delta$ by
where $\hat{Y}_i$ is a kernel regression estimator of $Y^*_i$, using knowledge of $h_0,h_1$ and $P(Z)$. Likewise, for estimating the conditional mean $E(Y_1|X \in A)$ where $A$ is a generic subset of the support of $X$, we use a weighted version
Limiting distribution theory for each of these estimators follows from identical arguments in \citeasnoun{vytlacilyildiz}. Here we formally state the theorem for the first estimator:
Next, we describe an estimation procedure for the distributional treatment effect in Example 4, where potential outcomes depend on random coefficients. In this case, the parameter of interest is, for a given value $y \in \mathbb{R}$,
For fixed values of $y$ and $p_{1}>p_{2}$, we propose to estimate $t(x,y)$ as
and then average over values of $p_{1},p_{2}$:
An infeasible estimator for $\Delta_2(y)$, which assumes $ t(x,y)$ is known, would be
In practice, for feasible estimation, one needs to replace $t(x,y)$ with its estimator $\hat{\tau}(x,y)$.
We conclude this section with some discussion about the computational aspects. The computational costs for finding the suitable pairs $(x,\tilde{x})$ are modest in comparison with typical semi- or nonparametric methods in the literature. To see this, consider the case with $J=2$, where $v=(v_{1},v_{2})\in \mathbb{R}^{2}$ consists of two index functions $v_{j}(x,d):\mathbb{R}^{K}\times \{0,1\}\rightarrow R$ for $j=1,2$. The actual dimensions that matter in implementation are: (i) $K$ in the estimation of $h_{1}$ and $h_{0}$ in the first stage, and (ii) the dimension of the indexes to be matched in the second stage: $J=2$. Therefore, the dimensionality that causes difficulty in estimation is $\max \left\{J,K\right\} $, not $J\times K$. In our view, this is no more difficult than implementing estimators that are based on matching or pairwise comparison, such as Blundell and Powell (2003).
\setcounter{equation}{0}
This section presents simulation evidence for the performance of the estimator in Section (ref), for both the Average Treatment Effect and the Distributional Treatment Effect.
We report results for both our estimator and that in \citeasnoun{vytlacilyildiz}, for several designs where the potential outcomes are real-valued and continuous. These include designs where the monotonicity condition fails, and designs where the disturbance terms in the outcome equation are multi-dimensional.
Throughout all designs, we model the treatment or dummy endogenous variable as
where $Z,U$ are independent standard normal. We experiment with the following designs for the outcome.
We note that the monotonicity condition holds in Design 1 but fails in the other two designs. For each of these designs, we report results for estimating $E[Y_1]$, i.e., the mean potential outcome under treatment $D=1$. The two estimators reported in the simulation study are our estimator proposed in Section (ref) and the one proposed in \citeasnoun{vytlacilyildiz}. The summary statistics, scaled by the true parameter value, Mean Bias, Median Bias, Root Mean Squared Error, (RMSE), and Median Absolute Deviation (MAD) are evaluated for sample sizes of $n = 100, 200, 400$ for 401 replications.
Results for each of these designs are reported in Tables 1 to 3 respectively. In implementing our estimator, we assume the propensity score function is known, and conduct next stage estimation using a nonparametric kernel estimator with normal kernel function, and a bandwidth of $n^{-1/5}$. This rate reflects “undersmoothing" as there are two regressors, the propensity score and the regressor $X$. For the estimator in \citeasnoun{vytlacilyildiz}, which involves the derivative of conditional expectation functions as well, we also report results for an infeasible version of their estimator, assuming such functions, as well as the propensity scores, are known. To implement the second stage of our estimator, in calculating the distance $\|h_1(x_i,\cdot)-h_0(x_0,\cdot)\|$ we used an evenly spaced grid of values for $y$, and selected $n/50$ grid points, with $n $ denoting the sample size.
The results indicate the desirable properties of our estimator, generally agreeing with Theorem (ref). In all designs, our estimator has small values for bias and RMSE, with the value of RMSE decreasing as the sample size grows. In contrast, the procedure based on \citeasnoun{vytlacilyildiz} only performs well in Design 1, with the sizes of bias and RMSE comparable to those using our method. As in our estimator, these values decrease as the sample size grows, which is expected, as the monotonicity condition they require is satisfied in this design. In this case, their approach has smaller standard errors largely due to the relatively simpler structure of the infeasible version of the estimator, but their biases persist even when the sample size increases.
For Designs 2 and 3, where the monotonicity condition is violated, the estimator proposed in \citeasnoun{vytlacilyildiz} does not perform well. Table 2 shows that in Design 2 both the bias and RMSE of their estimator are generally decreasing slowly with the sample size. Results for their estimator are better in Design 3 in Table 3, but the bias hardly converges with the sample size, and is much larger compared to our estimator.
We also report estimator performance in samples simulated from a model where potential outcomes are determined by random coefficients and dummy endogenous variables. It is important to note that for this design, the estimator in \citeasnoun{vytlacilyildiz} does not apply. This is because different values of $x$ lead to different distributions of the composite error $\eta_d + x^{\prime }\epsilon_d$. Our contribution in Section (ref) is to propose a new approach based on matching different values of the observed outcome $y$, rather than the exogenous covariates $x$. Based on the counterfactual framework discussed in Section (ref), here the treatment variable $D$ is modeled as the same way as in the first three designs, with the regressor $X$ being standard normal. For both $Y_0,Y_1$, the intercepts were modeled as constants (0 and 1, respectively) and the additive error terms were each standard normal. For the random slopes, the means were 1 and 2 respectively, and the additive error terms were also standard normal, independent of all other disturbance terms and each other. Here we use the procedure in Section (ref) to estimate the parameter $\Delta_2=P(Y_1<y)$, where in the simulation, we set $y=1$.
Results for this design with random coefficients are reported in Table 4. The same four summary statistics are reported for sample sizes $n \in \{100,200,400\}$, based on 401 replications. The estimator proposed in Section (ref) performs well. The bias and RMSE are much smaller for a bigger sample with $n=400$ than for smaller samples with $n = 100$ and $n=200$, indicating convergence at the parametric rate.
We also report the performance of our estimator in a sample drawn from Example 2, where the observed outcomes are determined in a multinomial choice model: \[ Y_{i}(D_{i})=\arg \max_{j\in \{0,1,2\}}Y_{i,j}^{\ast }(D_{i})\text{,} \] where the potential outcomes are $Y_{i,j}^{\ast }(d)\equiv v_{j}(X_{i,j},d)+\varepsilon _{i,j}$ with \[ v_{j}(X_{i,j},d)=\alpha _{j}(d)+X_{i,j}\beta _{j}(d) \] for $j=1,2$, and $v_0(X_{i,j},d) =0$ by way of normalization. The exogenous covariates $X_{i,j}$ are drawn independently from standard normal, and the intercepts and slope coefficients are:
For each individual $i$, the binary treatment $D_{i}$ is determined as before: \[ D_{i}=1\{U_{i} < Z_{i}\}\text{,} \] where the instrument $Z_{i}$ is independently drawn from a standard uniform distribution. The marginal distribution of the selection error $U_{i}$ is standard uniform. Conditional on $U_{i}=u$, the outcome errors $\varepsilon_{i,j}$ are independent across $j=0,1,2$ and are distributed as type-1 extreme value with unit variance and means $(0,\delta u,2\delta u)$ respectively, where $\delta$ is a parameter to be specified in the data-generating process. We adopt this specification as it allows for substantial dependence between $U_{i}$ and $\varepsilon _{i}\equiv (\varepsilon_{i,j})_{j=0,1,2}$.
Our simulation study uses sample sizes $n\in \{250,500,1000,2000\}$. For each sample size $n$, we generate $S=400$ independent samples from the data-generating process above. Throughout this section, we focus on estimating a conditional distribution of potential outcomes $\Pr\{Y_d=j|X \in \omega\}$ for $d\in \{0,1\}$ and $ j \in \{1,2\}$, where $\omega \equiv\{x:x_{j}\in \lbrack -1,1]$ for $j=1,2\}$ is a subset of the support of covariates.
As a benchmark, we first implement the infeasible version of the estimator in Section (ref), where knowledge of $h_{1}^{\ast }(\cdot )$ and $h_{0}^{\ast }(\cdot )$ are used for finding $(x,\tilde{x})$ with $v(x,1)=v(\tilde{x},0)$ for estimating $\Pr\{Y_1=j|X \in \omega\}$ (or $v(x,0)=v(\tilde{x},1)$ for estimating $\Pr\{Y_0=j|X \in \omega\}$). In this case, $F_{g|u}$ has a known close form, and we use numerical integration via mid-point approximation to calculate $h_{1}^{\ast }$ and $h_{0}^{\ast }$. Table 5 shows the mean bias and mean squared error (M.S.E.) for the infeasible estimator calculated from the $S$ simulated samples. We report these measures for $Pr\{Y_d = j | X \in \omega \}$ for $d=0,1$ and $j=1,2$. In the last two columns of the table, we report the M.S.E. for the full vector summarizing the probability mass function of $Y_d$, i.e., $[\Pr\{Y_d=1|X \in \omega \}, \Pr\{Y_d=2|X \in \omega \}]$.
The M.S.E. in Table 5 diminishes at a root-n rate that is proportional to the sample size. While in most cases the mean bias decreases with the sample size, the root-n rate of convergence appears to be substantially driven by the diminishing variance of the estimator. The performance does not vary substantively across different designs with the parameter values that affect the strength of correlation between the structural error $\varepsilon _{i,j}$ and the selection error $U_{i}$, i.e., the parameter values $\delta=(1/4,1/3,1/2)$.
We then construct a feasible version of the estimator by using $h_{1}^{\ast }$ and $h_{0}^{\ast }$ with their corresponding kernel estimates. Specifically, we use bivariate Gaussian kernels with bandwidths $1.06\hat{\sigma}^{-1/7}$, where $\hat{\sigma}$ denotes the sample standard deviation of components in $(X_{i},P_{i})$. Table 6 reports the mean bias and M.S.E. of this feasible estimator in the same data-generating process as in Table 5. The estimation errors in Table 6 are overall larger than those reported in Table 5 but demonstrate similar patterns of convergence in most cases, even though the convergence appears to be slower for the MSE of $F_{Y_{0}|X\in \omega}$ and $F_{Y_{1}|X\in \omega}$ when $\delta $ is large. The difference in the magnitude of estimation errors across Table 5 and Table 6 is attributable to the estimation error in $h_{1}^{\ast }$ and $h_{2}^{\ast }$. All in all, we conclude our estimator has decent finite-sample performance in these designs.
{\scalefont{0.75}
}
In this paper, we consider identification and estimation of weakly separable models with endogenous binary treatment. Existing approaches are based on a monotonicity condition, which is violated in models with multiple unobserved idiosyncratic shocks. Such models arise in many important empirical settings, including cases where potential outcomes are determined by Roy models, multinomial choice models, or random coefficients with dummy endogenous variables. We establish new identification results for these models which are constructive and conducive to estimation procedures. A simulation study indicates adequate finite sample performance of our method.
This paper leaves several open questions for future research. For example, it may not be feasible to locate pairs of $(x,\tilde{x})$ that satisfy the matching criterion perfectly due to limited, say, discrete, support of $X$ or $Z$. In this case, it remains an open question how or whether the partial identification approach proposed in \citeasnoun{shaikhvytlacil11} can be applied in the current setting where multiple indices in potential outcomes defy any notion of monotonicity. Besides, our method requires the selection of the number and location of cutoff points, so a data-driven method for selecting these would be useful. Furthermore, the relative efficiency of our proposed estimation approach needs to be explored, perhaps by deriving efficiency bounds for these new classes of models.