Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
56,944 characters · 10 sections · 56 citation commands
Nonparametric Bounds on Treatment Effects with Imperfect Instruments
\nonstopmode
The use of an instrumental variable (IV) is a popular solution to deal with endogeneity in social sciences. However, this approach may yield misleading conclusions when the instrument is invalid. A valid instrument must be uncorrelated with the unobservables in the model.\footnote{A stronger version of this condition is that the instrument is statistically (or mean) independent of the unobservables.} This requirement is often difficult to justify, and may call into question empirical findings. For this reason, Nevo2012 derived bounds on the parameters of interest (e.g., the average treatment effect) in parametric models under weaker conditions. They first assume that the sign of correlation between the imperfect IV (IIV)\footnote{An instrument that is potentially correlated with the unobservables.} and the unobserved latent variables is the same as that of the correlation between the endogenous variable and the latent variables. Second, they add the assumption that the correlation between the IIV and the latent variables is less than the correlation between the endogenous variable and the latent variables to tighten the bounds on the parameters of interest.
In this paper, we derive nonparametric bounds on the average treatment effect with an imperfect IV under the above assumptions when the outcome variable has bounded support. We introduce the concept of binarized MTS-MIV, which is implied by the monotone treatment selection (MTS) and monotone IV (MIV) assumptions developed by Manski2000, Manski2009. We show that the correlation between a binarized MTS-MIV and the unobserved latent variables has the same sign as the correlation between the endogenous variable and the latent variables. Hence, we link the Nevo2012 same direction of correlation assumption to the Manski2000 monotone treatment selection and monotone IV assumptions. We believe this result is new in the literature. Furthermore, we show how additional restrictions such as the less endogenous instrument, and the monotone treatment response can help tighten the bounds. As in Nevo2012, the bounds take the form of intersection bounds and can be implemented using the inferential methods developed by CLR2013 or AS2013. We illustrate our methodology using the National Longitudinal Survey of Young Men (NLSYM) data to estimate returns to schooling.
There is an increasing interest in the identification of causal effects with imperfect instrumental variables. Recently, Masten2020 developed a methodology that allows researchers to consider continuous relaxations of the IV models when they are refuted by the data. Their approach is data driven as it exploits the extent of falsification of the model to construct the identified set for the parameter of interest. Although their method helps salvage some invalid IVs, their identifying assumptions may seem difficult to interpret. Our approach, as well as that of Nevo2012 and Manski2000, Manski2009, is not data driven and has clearer interpretations of the identifying assumptions. This paper also relates to the work of Kedagni2018b, who consider weaker version of the mean independence assumption for the IV model. They derived nonparametric bounds on the average treatment effect under unconditional moment restrictions for the IV. Several other papers have also studied identification of model parameters when the IV is invalid using a framework different from ours; see Hotz1997, Conley2012, among others.
The remainder of the paper is organized as follows. Section (ref) presents the model, the assumptions and their link with the literature. In Section (ref), we derive our main identification results. We discuss inference and implementation in Section (ref). Section (ref) presents an empirical illustration of our proposed methodology, while Section (ref) concludes. Proofs and additional results are relegated to the appendix.
Consider the following potential outcome model (POM)
where $Y$ is the outcome variable taking values in $\mathcal Y \subset \mathbb R$, $D$ is a discrete endogenous treatment variable taking values in $\mathcal D=\{1,2,\ldots,T\}$, $Y_{d}$ is the potential outcome that would have been observed if the treatment $D$ had externally been set to $d$. Let $Z\in \mathcal Z \subseteq \mathbb R$ be an imperfect IV in the sense that it may be correlated with the potential outcome $Y_d$. In what follows, we assume that the random variable $Y_d$ is integrable, i.e., $\text{E}[Y_d]<\infty$. The objects of interest in this paper are the potential outcome means $\theta_d \equiv \text{E}[Y_d]$, for all $d \in \mathcal D$, and some treatment effects $ATE(d,d')\equiv \theta_d-\theta_{d'}$, for $d,d' \in \mathcal D$. We allow for heterogeneous treatment effects, so that $ATE(d,d')$ may vary across $(d,d')$. The methodology that we develop in this paper can also be used to identify other commonly used parameters of interest such as the average treatment effect on the treated $ATT(d,d')\equiv\text{E}[Y_d-Y_{d'}\vert D=d]$, and the average treatment effect on the untreated $ATU(d,d')\equiv\text{E}[Y_d-Y_{d'}\vert D=d']$. But, for the sake of clarity of the exposition, we focus our attention on the $ATE$.
We observe a random sample of the vector $(Y,D,Z)$. For simplicity, we drop exogenous covariates from the analysis. For example, $Y$ could be earnings, $D$ years of schooling, and $Z$ parental education. In this example, $Y_d$ is the potential earnings for an individual with $d$ years of schooling. We now state our main identifying assumptions:
Assumption BoS states that the support of the counterfactual outcome is the same as that of the factual. It is standard and similar to the usual bounded outcome assumption considered in Manski1990, Manski1994, and many other papers. Like in Kedagni2018b, it allows the support of the potential outcome $Y_d$ to vary across all treatment levels $d$.
Assumption SDC is equivalent to Assumption 3 in Nevo2012. It states that the correlation between the imperfect instrument $Z$ and the potential outcome $Y_d$ has weakly the same sign as the correlation between the endogenous treatment $D$ and the potential outcome. For example, it is documented that parental education is not a valid instrument; see Kedagni20, Mourifie2020, among many others. However, one could assume that parental education has the same sign of correlation with the potential earnings as does the individual's education. Note that if either the treatment $D$ or the instrument $Z$ is exogenous, this assumption holds. If Assumption BoS holds from the definition of the outcome variable (e.g., market share lies between 0 and 1), Assumption SDC has a testable implication. Indeed, when BoS holds and the bounds derived in Proposition (ref) for $\theta_d$ under BoS and SDC are empty, then SDC is rejected.
Assumption SDC can be seen as a weaker version of the concepts of monotone IV (MIV: $\text{E}\left[Y_d \vert Z=z\right]$ is monotone in $z$ for all $d$) and monotone treatment selection (MTS: $\text{E}\left[Y_d \vert D=\ell\right]$ is monotone in $\ell$ for all $d$) developed by Manski2000, Manski2009. To show this result, we introduce the concept of binarized MTS-MIV that is intermediate between SDC and MTS-MIV. We use the following notation.
In words, we say that $Z$ is a binarized MTS-MIV for $D$ if all binarized treatments $\mathbbm{1}\{D\geq j\}$ satisfy the MTS restriction, and all binarized instruments $\mathbbm{1}\{Z\geq z\}$ satisfy the MIV restriction.
Lemma (ref) shows that MTS-MIV is a sufficient condition for binarized MTS-MIV, while Lemma (ref) shows that binarized MTS-MIV is a sufficient condition for SDC.
Another sufficient condition for binarized MTS-MIV is the joint positive quadrant dependence between the potential outcome $Y_d$ and the instrument $Z$, and between the potential outcome $Y_d$ and the treatment $D$. The concept of positive quadrant dependence has been considered by BSV2012 in a different framework. Two random variables $\varepsilon$ and $\nu$ are positive quadrant dependent (PQD) if
As pointed out by BSV2012, the PQD assumption implies that
From this implication, we conclude that if $Y_d$ and $Z$ are PQD, then the distribution $Y_d$ conditional on $\{Z \geq z\}$ first-order stochastically dominates that of $Y_d$ conditional on $\{Z < z\}$ for all $z$. Therefore, $\text{E}[Y_d \vert Z \geq z] \geq \text{E}[Y_d \vert Z < z]$, i.e., $h_d^+(z)-h_d^-(z) \geq 0$ for all $z$. Similarly, if $Y_d$ and $D$ are PQD, then $g_d^+(j)-g_d^-(j) \geq 0$ for all $j$. Hence, binarized MTS-MIV holds.
Note that joint positive quadrant dependence between $Y_d$ and $Z$, and between $Y_d$ and $D$ does not imply joint MTS-MIV, and the converse does not hold either. However, joint positive regression dependence (PRD) between $Y_d$ and $Z$ (i.e., $\text{P}(Y_d > y \vert Z=z)$ is nondecreasing in $z$ for all $y$), and between $Y_d$ and $D$ implies both joint MTS-MIV, and joint positive quadrant dependence between $Y_d$ and $Z$, and between $Y_d$ and $D$. See the proof in the appendix. Figure (ref) below summarizes the relationship between the different concepts.
Assumption LEI is the same as Assumption 4 in Nevo2012, which they refer to as the “instrument less endogenous than treatment” assumption. It states that the imperfect instrument $Z$ is less correlated with the potential outcome than is the endogenous treatment $D$. In this paper, we use the shorthand “less endogenous instrument” to call this assumption. In the context of our empirical example, it is reasonable to assume that parental education is less correlated with the individual's potential wage than is the individual's own education.
Assumption MTR states that the potential outcome weakly increases with the level of the treatment. It was introduced by Manski1997, and considered in Manski2000, Manski2009, among many others. For instance, in the returns to schooling example, it implies that the wage that a worker earns weakly increases as a function of the worker's years of schooling. We show how this assumption can help tighten the bounds derived under Assumptions BoS, SDC and LEI.
Now that we have discussed the model and our identifying assumptions, we are going to present our main identification results.
In this section, we derive under the different assumptions discussed in the previous section bounds on the potential outcome expectation $\theta_d$, for each $d \in \mathcal D$: $LB_d \leq \theta_d \leq UB_d$. Bounds on $ATE(d,d')$ are then obtained as: $LB_d-UB_{d'} \leq ATE(d,d') \leq UB_d-LB_{d'}$.
\setcounter{equation}{0}
Assumption SDC is equivalent to $\text{E}\left[Y_d\tilde{D}\right]\text{E}\left[Y_d\tilde{Z}\right] \geq 0$, where $\tilde{D}\equiv D-\text{E}[D]$ and $\tilde{Z}\equiv Z-\text{E}[Z]$, which in turn is equivalent to: either
or
We first derive bounds on the potential outcome mean $\theta_d$ using inequalities ((ref)). Similarly, we can derive the bounds implied by inequalities ((ref)).
Inequality ((ref)) implies that, for all $(\lambda,\gamma)\in \mathbb R^2_+\setminus \left\{(0,0)\right\}$, we have
By factorizing $(\lambda + \gamma)$ in the above inequality, we have
where $\beta=\frac{\lambda}{\lambda + \gamma}$. We can normalize $\lambda + \gamma$ to lie within the interval $[0,1]$.\footnote{If $\lambda + \gamma >1$, we can multiply each side of the inequality by $\frac{1}{\lambda + \gamma+1}$, and have $\frac{\lambda + \gamma}{\lambda + \gamma+1} \in [0,1]$.} By setting $\alpha=\lambda + \gamma$, this last inequality becomes
Hence, Inequality ((ref)) implies that, for any $(\alpha,\beta) \in [0,1]^2$, we have
Intuitively, $\beta$ measures how tight is the constraint $\text{E}\left[Y_d\tilde{D}\right] \geq 0$ relatively to the constraint $ \text{E}\left[Y_d\tilde{Z}\right] \geq 0$, while $\alpha$ measures the extent to which the mixture of the constraints is binding. For example, when the constraint $\text{E}\left[Y_d\tilde{Z}\right] \geq 0$ is binding while the constraint $\text{E}\left[Y_d\tilde{D}\right] \geq 0$ is slack, then $\beta=0$, and the instrument $Z$ satisfies the zero covariance assumption, while the treatment variable $D$ is endogenous. In such a scenario, we expect the mixed constraint to be binding as well, that is $\alpha=1$. If instead, the treatment variable $D$ satisfies the zero covariance assumption, and the instrument $Z$ is endogenous, then we expect $\beta=1$ and $\alpha=1$.
The latter inequalities are respectively equivalent to:\footnote{In the case where the $ATT/ATU$ is our parameter of interest, we would bound $\theta_{d|d'}\equiv \text{E}[Y_d\vert D=d']=\frac{\text{E}[Y_d\mathbbm{1}\{D=d'\}]}{\text{E}[\mathbbm{1}\{D=d'\}]}$. In such a case, we will equivalently write these inequalities as:
and use the same technique we develop in this paper.}
where $\delta^+_{S} \equiv 1+ \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)$ and $\delta^-_{S} \equiv 1- \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)$.
Furthermore, using the identity $\mathbbm{1}\left\{D=d\right\} + \mathbbm{1}\left\{D\neq d\right\}=1$, we rewrite them as
respectively, given that $Y=Y_d$ when $D=d$.
Now, using Assumption BoS, we can bound the counterfactuals $\delta^+_{S} Y_d \mathbbm{1}\left\{D \neq d\right\}$ and $\delta^-_{S} Y_d \mathbbm{1}\left\{D \neq d\right\}$ as follows:
Therefore, using inequalities ((ref)) and ((ref)), it follows that
for any $(\alpha, \beta) \in [0,1]^2$, where we define the function $\underline{f}_d$ and $\overline{f}_d$ as
We can then take the supremum and the infimum of the lower and upper bounds over $(\alpha, \beta)$, respectively, to obtain the following bounds for $\theta_d$:
Similarly, using inequalities ((ref)), we derive the following bounds for $\theta_d$:
All these results are summarized in the following proposition.
Proposition (ref) provides two-sided bounds on the potential outcome means, and then on the average treatment effects, which mainly relies on the bounded outcome assumption. Nevo2012 obtain two-sided bounds if $Cov(D,Z) <0$, and one-sided bounds if $Cov(D,Z) > 0$. We relax the parametric linear assumption at the expense of the bounded support assumption. In light of the following statement from the fourth paragraph of Section VI in Nevo2012 “\ldots However, with a nonparametric functional, it is doubtful that our assumptions on the correlations of endogenous regressors and imperfect instruments with econometric errors would prove anywhere near as fruitful,” we believe that the result of Proposition (ref) makes a positive contribution to the literature. However, we do not have a proof that the derived bounds are sharp at this point. We believe that this is an important theoretical question that can be investigated in future research.
The bounds derived in Proposition (ref) are no wider than the usual Manski worst-case bounds without instrument, as the latter bounds are a special case of ours where $\alpha=0$. Furthermore, we expect the Manski bounds derived under the strict IV exogeneity condition, $\text{E}\left[Y_d\vert Z\right]=\text{E}\left[Y_d\right]$, to be weakly narrower than the bounds $I_{SDC}^d$. We provide a heuristic proof for this conjecture in the appendix. Intuitively, we expect the optimal value of $\beta$ to be equal to 0 under the strict IV exogeneity assumption, since the treatment variable $D$ is endogenous and the intrument $Z$ satisfies the zero covariance assumption. Using this restriction, we show that the lower (upper) bounds of $I_{SDC1}^d$ and $I_{SDC2}^d$ are each less (greater) than the Manski lower (upper) bound.
It is possible that the bounds in Proposition (ref) be empty. Indeed, if the supremum and the infimum in the bounds' expressions are attained at different values of $(\alpha, \beta)$, the lower bounds of $I_{SDC1}^d$ and $I_{SDC2}^d$ could be bigger than their respective upper bounds. When that happens for both bounds $I_{SDC1}^d$ and $I_{SDC2}^d$, we say that the model (Assumptions BoS and SDC) is rejected in the data. When only $I_{SDC1}^d$ ($I_{SDC2}^d$) is empty, then the assumption that the instrument $Z$ and the treatment $D$ are positively (negatively) correlated with the potential outcome $Y_d$ is rejected.
In this subsection, we combine Assumptions SDC and LEI in order to get tighter bounds on the parameter $\theta_d$. Assumption LEI is equivalent to $\vert \frac{\text{E}\left[Y_d\tilde{D}\right]}{\sigma_D}\vert \geq \vert\frac{\text{E}\left[Y_d\tilde{Z}\right]}{\sigma_Z}\vert$. Hence, Assumptions LEI and SDC imply that either of the followings is always true:
or
Differently, we can rewrite these inequalities as either
or
Applying a similar reasoning as in the previous subsection to ((ref)), for any $(\alpha, \beta, \gamma) \in [0, 1]^3$ such that $1-\beta-\gamma \geq 0$, we have
Since $\beta$ and $\gamma$ belong to a 2-simplex, we can parametrize $\gamma=(1-\beta)\mu$, where $\mu \in [0,1]$. Therefore, the above inequalities are equivalent to
where $\delta^+_{L} \equiv 1+ \alpha\big((1-\beta)\mu (\tilde{D}\sigma_Z - \tilde{Z}\sigma_D) + \beta \tilde{D} + (1-\beta)(1-\mu) \tilde{Z} \big)$ and $\delta^-_{L} \equiv 1- \alpha\big((1-\beta)\mu (\tilde{D}\sigma_Z - \tilde{Z}\sigma_D) + \beta \tilde{D} + (1-\beta)(1-\mu) \tilde{Z} \big)$. Hence, we obtain the following bounds for $\theta_d$ from ((ref)):
Likewise, we derive bounds for $\theta_d$ from ((ref)) as
and the following proposition holds.
The following corollary shows that $I_{LEI}^d$ is weakly tighter than $I_{SDC}^d$.
The proof of Corollary (ref) follows from the fact that the bounds $I_{SDC}^d$ are a special case of the bounds $I_{LEI}^d$ where $\mu=0$.
In this subsection, we derive bounds on the potential outcome mean $\theta_d$ under Assumptions BoS, SDC, LEI, and MTR . Under Assumption MTR, we have
where the inequality holds as $Y_d \mathbbm{1}\left\{D=j\right\} \leq Y_j \mathbbm{1}\left\{D=j\right\}=Y \mathbbm{1}\left\{D=j\right\}$ for all $j>d$ for each $d$. This result shows that under the MTR assumption, the counterfactual random variable $Y_d \mathbbm{1}\left\{D>d\right\}$ is bounded from above by the observed variable $Y \mathbbm{1}\left\{D>d\right\}$. As we can see, this assumption considerably shrinks the upper bound on $Y_d \mathbbm{1}\left\{D>d\right\}$, which would be $\overline{y}_d \mathbbm{1}\left\{D>d\right\}$ otherwise under Assumption BoS. Similarly, we also have
for each $d$. Without the MTR assumption, the lower bound on $Y_d \mathbbm{1}\left\{D<d\right\}$ would be $\underline{y}_d \mathbbm{1}\left\{D<d\right\}$ under Assumption BoS. Combining these results together, the following inequalities hold under Assumptions BoS and MTR:
Thus, for any $\delta \in \mathbb{R}$, we have the following bounds
Moreover, as we have
inequalities ((ref)) imply
for any $\delta \in \mathbb{R}$.
Now, recall that inequalities ((ref)) implied by Assumptions SDC and LEI yield
for any $(\alpha, \beta, \mu) \in [0,1]^3$, and thus we have
for any $(\alpha, \beta,\mu) \in [0,1]^3$, where we define the functions $\underline{m}_d$ and $\overline{m}_d$ as
Likewise, inequalities ((ref)) under the MTR assumption yield the following implications for $\theta_d$:
for any $(\alpha, \beta,\mu) \in [0,1]^3$. Therefore, we conclude that Assumptions BoS, SDC, LEI, and MTR together imply
where
Note that if we are not willing to impose Assumption LEI, the bounds can be obtained from $I_{MTR1}^d \cup I_{MTR2}^d$ using $\delta^+_{S}$ and $\delta^-_{S}$ instead of $\delta^+_{L}$ and $\delta^-_{L}$.
An important implication of the MTR assumption is that it weakly signs the ATE, i.e., $ATE(d,d')\geq 0$ for all $d > d'$. More precisely, bounds for $ATE(d,d')$, where $d>d'$, can be characterized as follows:
\setcounter{equation}{0} We want to construct confidence bounds for the set $I_{SDC}^d = I_{SDC1}^d \cup I_{SDC2}^d $. This is a special case of an intersection-union test as described in Berger1982. We are going to construct confidence regions for the sets $I_{SDC1}^d$ and $I_{SDC2}^d$ using the intersection bounds framework of CLR2013 or AS2013, and then take the union of the two confidence regions. Berger1996 showed that the union of the confidence regions has at least the same coverage rate as each confidence region. Identified sets with a similar structure have been considered in Chesher2020 and Machado2019. As we explain in Section (ref), the sets $I_{SDC1}^d$ and $I_{SDC2}^d$ could each be empty, and may not overlap. However, in our empirical illustration below, they are nonempty and overlap in all cases.
We now explain how to rewrite the intersection bounds $ I_{SDC1}^d$ in such a way that it can be easily implemented using the CKLRstata or AKS2017 Stata packages. Suppose that we draw two independent random variables $U_1$ and $U_2$ from the uniform distribution over $[0,1]$, independently of the data $(Y,D,Z)$. Then, we have
since $U_1$ and $U_2$ are independent of $(Y,D,Z)$. Therefore,
Hence, these bounds take the form of conditional moment inequalities, which can be implemented using existing inferential methods like CKLRstata or AKS2017. Similar results apply to the bounds $I_{LEI}^d$ and $I_{MTR}^d$, where there are three conditioning variables $U_1$, $U_2$, and $U_3$ instead of two. A similar technique has been proposed in Kedagni2018b where they construct confidence sets for the potential outcome means under the IV zero-covariance assumption.
Using the CKLRstata Stata package, we obtain confidence sets that asymptotically cover either the true parameter $\theta_d$ or the bounds for $\theta_d$ with pre-specified probability, through the clr3bound command.\footnote{The clr2bound command could be used to obtain confidence sets for the bounds for $\theta_d$ only.}
\setcounter{equation}{0} In this application, we use a data set drawn from the NLSYM. This data includes 3,010 young men who were ages 24-34 in 1976. It is the same data used in Card1995. In our analysis, the outcome variable is log hourly wage in cents $(lwage)$, and the treatment variable is education $(educ)$ grouped in 4 categories: less than high school $(educ< 12\ years)$, high school $(12 \leq educ < 16)$, college degree $(16 \leq educ < 18)$, and graduate $(educ \geq 18)$.\footnote{As discussed in Andresen2021, this discretization may induce some identification issues, and the results may be sensitive to it. For this reason, our empirical results should be seen as illustrative.}
Our imperfect IV is parental education. Since the work of Willis79, parental education has been used as an IV. However, an individual's ability can be dependent on her parents' ability, which is correlated with parental education. For this reason, parents' education will not be a valid instrument. This fact is documented in Kedagni20, who provided evidence that even after controlling for a measure of ability, parental education is not a good instrument. This result is contrary to the Lemke03 idea that controlling for some measure of child ability could make parental education a valid IV. Nonetheless, it is reasonable to assume that parental education has the same sign of correlation with the individual's potential wage as the correlation between the person's potential wage and her own education.\footnote{The MIV and possibly MTS assumptions also seem reasonable in this empirical example. However, our goal in this section is to show how our derived bounds can be implemented in a real-world application.} It is also likely that parental education be less endogenous than is the person's own education. Finally, as in Manski2000, we use the monotone treatment response assumption to tighten the bounds on the average returns to education.
In theory, the outcome variable $lwage$ is unbounded. For practical reasons, we follow Ginther2000 to trim the log wage. The outcome variable that we use is defined as $Y=\tau$-quantile of $lwage$ if $lwage$ is less than or equal to its $\tau$-quantile, $Y=(1-\tau)$-quantile of $lwage$ if $lwage$ is greater than or equal to its $(1-\tau)$-quantile, and $Y=lwage$ otherwise. In our empirical illustration, we set $\tau=0.05$. We construct 95% two-sided confidence sets on potential average log wages and their bounds using the clr3bound command of CKLRstata in the Stata software. We estimate the conditional expectations using the parametric method, which is the default option for this Stata command. See the appendix for more details on the implementation.
We present the results with mother's education as an IIV. The results for father's education are in the appendix. Table (ref) displays the 95% confidence sets for the potential wage means and average returns to schooling under the SDC assumption, while Table (ref) shows the confidence sets under both the SDC and LEI assumptions. Each table shows the confidence set for bounds on $\theta_d$ in the first two columns, the confidence set for the parameter $\theta_d$ in the third and forth columns, and the point estimates of the bounds in the fifth and sixth columns. The SDC+LEI bounds in Table (ref) are generally narrower than the SDC bounds in Table (ref). This suggests that Assumption LEI provides some extra identifying power to the SDC assumption. However, the bounds seem wide and less informative in both cases. For example, the confidence set for $ATE(2,1) \equiv \theta_2-\theta_1$ under SDC+LEI is $[-0.98,1.02]$, implying that the average return to college degree compared to high school education varies between $-98\%$ and $102\%$, which is not very informative for an individual making a college decision. The corresponding point estimates of the bounds seem tighter $[-0.53, 0.69]$.
Furthermore, the confidence regions for the bounds and the parameters considerably shrink and become more informative when we add the MTR assumption (see Table (ref)). We assume that the lower bound of $\theta_d$ is equal to the upper bound of $\theta_{d-1}$ whenever the former is less than the latter, because $Y_{d-1}$ cannot exceed $Y_d$ for each $d=1, 2, 3$ under the MTR assumption. Individuals with less than high school education could earn up to 94% less than high school graduates ($ATE(0,1)$). Moreover, the confidence regions for $ATE(2, 1)$ and $ATE(3,1)$ suggest that college graduates could earn up to 55% more than high school graduates, while individuals with a graduate degree earn between 41% and 65% higher wages than high school graduates (which approximately represents an annual return between 6.8% and 10.8%). Table II in Card2001 shows that the point estimates of the annual return to schooling in the US vary roughly between 5% and 13%. As we can see, our set estimates for the annual return are consistent with the existing range in the literature.
If we impose the linear structure of Nevo2012 in this empirical exercise, we obtain the following point estimate bounds for $\theta$ under SDC and MTR: $[0, 0.16] \cup [0.24, +\infty)$. Indeed, the estimated correlation between $D$ and $Z$ is $0.3824>0$, and under SDC, the point estimate bounds for $\theta$ are $(-\infty, 0.16] \cup [0.24, +\infty)$. If one is willing to further assume that $Cov(D,U) \geq 0$, then the bounds for $\theta$ reduce to $[0, 0.16]$, which imply an annual return between 0 and 16%. This range is consistent with the literature, but remains wide.
In this paper, we derive nonparametric bounds on the average treatment effect when an imperfect instrument is available. We extend Nevo2012's (Nevo2012) identification results to nonparametric models. We first assume that the sign of correlation between the imperfect instrument and the unobserved latent variables is the same as the correlation between the endogenous variable and the latent variables. We show that the MTS-MIV restrictions introduced by Manski2000, Manski2009, jointly imply this assumption. Second, we show how the assumption that the imperfect instrument is less endogenous than the treatment variable can help tighten the bounds. We also use the monotone treatment response assumption to get tighter bounds. The identified set takes the form of intersection bounds, which can be implemented using CLR2013's (CLR2013) inferential method. Finally, we illustrate our methodology using the National Longitudinal Survey of Young Men data to estimate returns to schooling.
The authors are grateful to the Editor Petra Todd, and three anonymous referees for valuable suggestions and comments. They also thank Santiago Acerenza, Otavio Bartalotti, Helle Bunzel, Ismael Mourifi\'e, Vitor Possebom, and participants at the Iowa State econometrics workshop for helpful comments. All errors are ours.