EconBase
← Back to paper

Nonparametric Bounds on Treatment Effects with Imperfect Instruments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

56,944 characters · 10 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric Bounds on Treatment Effects with Imperfect Instruments

\nonstopmode

abstractThis paper extends the identification results in Nevo2012 to nonparametric models. We derive nonparametric bounds on the average treatment effect when an imperfect instrument is available. As in Nevo2012, we assume that the correlation between the imperfect instrument and the unobserved latent variables has the same sign as the correlation between the endogenous variable and the latent variables. We show that the monotone treatment selection and monotone instrumental variable restrictions, introduced by Manski2000, Manski2009, jointly imply this assumption. Moreover, we show how the monotone treatment response assumption can help tighten the bounds. The identified set can be written in the form of intersection bounds, which is more conducive to inference. We illustrate our methodology using the National Longitudinal Survey of Young Men data to estimate returns to schooling.

Introduction

The use of an instrumental variable (IV) is a popular solution to deal with endogeneity in social sciences. However, this approach may yield misleading conclusions when the instrument is invalid. A valid instrument must be uncorrelated with the unobservables in the model.\footnote{A stronger version of this condition is that the instrument is statistically (or mean) independent of the unobservables.} This requirement is often difficult to justify, and may call into question empirical findings. For this reason, Nevo2012 derived bounds on the parameters of interest (e.g., the average treatment effect) in parametric models under weaker conditions. They first assume that the sign of correlation between the imperfect IV (IIV)\footnote{An instrument that is potentially correlated with the unobservables.} and the unobserved latent variables is the same as that of the correlation between the endogenous variable and the latent variables. Second, they add the assumption that the correlation between the IIV and the latent variables is less than the correlation between the endogenous variable and the latent variables to tighten the bounds on the parameters of interest.

In this paper, we derive nonparametric bounds on the average treatment effect with an imperfect IV under the above assumptions when the outcome variable has bounded support. We introduce the concept of binarized MTS-MIV, which is implied by the monotone treatment selection (MTS) and monotone IV (MIV) assumptions developed by Manski2000, Manski2009. We show that the correlation between a binarized MTS-MIV and the unobserved latent variables has the same sign as the correlation between the endogenous variable and the latent variables. Hence, we link the Nevo2012 same direction of correlation assumption to the Manski2000 monotone treatment selection and monotone IV assumptions. We believe this result is new in the literature. Furthermore, we show how additional restrictions such as the less endogenous instrument, and the monotone treatment response can help tighten the bounds. As in Nevo2012, the bounds take the form of intersection bounds and can be implemented using the inferential methods developed by CLR2013 or AS2013. We illustrate our methodology using the National Longitudinal Survey of Young Men (NLSYM) data to estimate returns to schooling.

There is an increasing interest in the identification of causal effects with imperfect instrumental variables. Recently, Masten2020 developed a methodology that allows researchers to consider continuous relaxations of the IV models when they are refuted by the data. Their approach is data driven as it exploits the extent of falsification of the model to construct the identified set for the parameter of interest. Although their method helps salvage some invalid IVs, their identifying assumptions may seem difficult to interpret. Our approach, as well as that of Nevo2012 and Manski2000, Manski2009, is not data driven and has clearer interpretations of the identifying assumptions. This paper also relates to the work of Kedagni2018b, who consider weaker version of the mean independence assumption for the IV model. They derived nonparametric bounds on the average treatment effect under unconditional moment restrictions for the IV. Several other papers have also studied identification of model parameters when the IV is invalid using a framework different from ours; see Hotz1997, Conley2012, among others.

The remainder of the paper is organized as follows. Section (ref) presents the model, the assumptions and their link with the literature. In Section (ref), we derive our main identification results. We discuss inference and implementation in Section (ref). Section (ref) presents an empirical illustration of our proposed methodology, while Section (ref) concludes. Proofs and additional results are relegated to the appendix.

Analytical Framework

Consider the following potential outcome model (POM)

eqnarray[eqnarray omitted — 78 chars of source]

where $Y$ is the outcome variable taking values in $\mathcal Y \subset \mathbb R$, $D$ is a discrete endogenous treatment variable taking values in $\mathcal D=\{1,2,\ldots,T\}$, $Y_{d}$ is the potential outcome that would have been observed if the treatment $D$ had externally been set to $d$. Let $Z\in \mathcal Z \subseteq \mathbb R$ be an imperfect IV in the sense that it may be correlated with the potential outcome $Y_d$. In what follows, we assume that the random variable $Y_d$ is integrable, i.e., $\text{E}[Y_d]<\infty$. The objects of interest in this paper are the potential outcome means $\theta_d \equiv \text{E}[Y_d]$, for all $d \in \mathcal D$, and some treatment effects $ATE(d,d')\equiv \theta_d-\theta_{d'}$, for $d,d' \in \mathcal D$. We allow for heterogeneous treatment effects, so that $ATE(d,d')$ may vary across $(d,d')$. The methodology that we develop in this paper can also be used to identify other commonly used parameters of interest such as the average treatment effect on the treated $ATT(d,d')\equiv\text{E}[Y_d-Y_{d'}\vert D=d]$, and the average treatment effect on the untreated $ATU(d,d')\equiv\text{E}[Y_d-Y_{d'}\vert D=d']$. But, for the sake of clarity of the exposition, we focus our attention on the $ATE$.

We observe a random sample of the vector $(Y,D,Z)$. For simplicity, we drop exogenous covariates from the analysis. For example, $Y$ could be earnings, $D$ years of schooling, and $Z$ parental education. In this example, $Y_d$ is the potential earnings for an individual with $d$ years of schooling. We now state our main identifying assumptions:

assumption[Bounded support (BoS)] \begin{eqnarray*} Supp(Y_d\vert D\neq d)= Supp(Y_d \vert D = d) =\left[y_d,\overline{y}_d\right]\ for each \ d \in \mathcal{D}. \end{eqnarray*}

Assumption BoS states that the support of the counterfactual outcome is the same as that of the factual. It is standard and similar to the usual bounded outcome assumption considered in Manski1990, Manski1994, and many other papers. Like in Kedagni2018b, it allows the support of the potential outcome $Y_d$ to vary across all treatment levels $d$.

assumption[Same direction of correlation (SDC)] \begin{eqnarray*} Cov\left(Y_d,D\right) Cov\left(Y_d,Z\right)\geq 0\ for each \ d \in \mathcal{D}. \end{eqnarray*}

Assumption SDC is equivalent to Assumption 3 in Nevo2012. It states that the correlation between the imperfect instrument $Z$ and the potential outcome $Y_d$ has weakly the same sign as the correlation between the endogenous treatment $D$ and the potential outcome. For example, it is documented that parental education is not a valid instrument; see Kedagni20, Mourifie2020, among many others. However, one could assume that parental education has the same sign of correlation with the potential earnings as does the individual's education. Note that if either the treatment $D$ or the instrument $Z$ is exogenous, this assumption holds. If Assumption BoS holds from the definition of the outcome variable (e.g., market share lies between 0 and 1), Assumption SDC has a testable implication. Indeed, when BoS holds and the bounds derived in Proposition (ref) for $\theta_d$ under BoS and SDC are empty, then SDC is rejected.

Assumption SDC can be seen as a weaker version of the concepts of monotone IV (MIV: $\text{E}\left[Y_d \vert Z=z\right]$ is monotone in $z$ for all $d$) and monotone treatment selection (MTS: $\text{E}\left[Y_d \vert D=\ell\right]$ is monotone in $\ell$ for all $d$) developed by Manski2000, Manski2009. To show this result, we introduce the concept of binarized MTS-MIV that is intermediate between SDC and MTS-MIV. We use the following notation.

notationDenote $g_d^+(j)=\text{E}[Y_d|D\geq j]$, $g_d^-(j)=\text{E}[Y_d|D < j]$, $h_d^+(z)=\text{E}[Y_d|Z\geq z]$, $h_d^-(z)=\text{E}[Y_d|Z < z]$. $\rho_{UV}$ denotes the coefficient of correlation between two random variables $U$ and $V$.
definitionThe variable $Z$ is a binarized MTS-MIV for $D$ if for each $d \in \mathcal D$, \begin{eqnarray} \left(g_d^+(j)-g_d^-(j)\right)\left(h_d^+(z)-h_d^-(z)\right)\geq 0\ for all j,\ z. \end{eqnarray}

In words, we say that $Z$ is a binarized MTS-MIV for $D$ if all binarized treatments $\mathbbm{1}\{D\geq j\}$ satisfy the MTS restriction, and all binarized instruments $\mathbbm{1}\{Z\geq z\}$ satisfy the MIV restriction.

remarkIf $Z$ is a binarized MTS-MIV for $D$ then the functions $g_d^+$ and $g_d^-$ do not cross, nor do the functions $h_d^+$ and $h_d^-$ for all $d$; that is, either [$g_d^+(j) \geq g_d^-(j)$ for all $j$ and $h_d^+(z) \geq h_d^-(z)$ for all $z$] or [$g_d^+(j) \leq g_d^-(j)$ for all $j$ and $h_d^+(z) \leq h_d^-(z)$ for all $z$]. Moreover, if $g_d^+ \geq g_d^-$ for some $d$ then $h_d^+ \geq h_d^-$, and vice versa.

Lemma (ref) shows that MTS-MIV is a sufficient condition for binarized MTS-MIV, while Lemma (ref) shows that binarized MTS-MIV is a sufficient condition for SDC.

lemmaMTS-MIV in the same direction for $D$ and $Z$ implies that $Z$ is a binarized MTS-MIV for $D$.
lemmaIf $Z$ is a binarized MTS-MIV for $D$, then Assumption SDC holds.
remarkFrom Lemmas (ref) and (ref), we conclude that MTS-MIV in the same direction implies Assumption SDC. Moreover, when both the treatment $D$ and the imperfect instrument $Z$ are binary, MTS-MIV in the same direction, binarized MTS-MIV and Assumption SDC are equivalent. However, Example (ref) in the appendix shows a case where binarized MTS-MIV holds, but the joint MTS-MIV fails. If MTS and MIV hold in the opposite directions, then SDC will not hold. Instead, the “opposite directions of correlation” assumption $Cov(Y_d,D) Cov(Y_d,Z) \leq 0$ holds. The identification strategy developed in this paper can easily be adapted to this case. In general, if the directions of the MTS and MIV assumptions are unknown, SDC is not weaker than MTS-MIV.

Another sufficient condition for binarized MTS-MIV is the joint positive quadrant dependence between the potential outcome $Y_d$ and the instrument $Z$, and between the potential outcome $Y_d$ and the treatment $D$. The concept of positive quadrant dependence has been considered by BSV2012 in a different framework. Two random variables $\varepsilon$ and $\nu$ are positive quadrant dependent (PQD) if

eqnarray*[eqnarray* omitted — 128 chars of source]

As pointed out by BSV2012, the PQD assumption implies that

eqnarray*[eqnarray* omitted — 147 chars of source]

From this implication, we conclude that if $Y_d$ and $Z$ are PQD, then the distribution $Y_d$ conditional on $\{Z \geq z\}$ first-order stochastically dominates that of $Y_d$ conditional on $\{Z < z\}$ for all $z$. Therefore, $\text{E}[Y_d \vert Z \geq z] \geq \text{E}[Y_d \vert Z < z]$, i.e., $h_d^+(z)-h_d^-(z) \geq 0$ for all $z$. Similarly, if $Y_d$ and $D$ are PQD, then $g_d^+(j)-g_d^-(j) \geq 0$ for all $j$. Hence, binarized MTS-MIV holds.

Note that joint positive quadrant dependence between $Y_d$ and $Z$, and between $Y_d$ and $D$ does not imply joint MTS-MIV, and the converse does not hold either. However, joint positive regression dependence (PRD) between $Y_d$ and $Z$ (i.e., $\text{P}(Y_d > y \vert Z=z)$ is nondecreasing in $z$ for all $y$), and between $Y_d$ and $D$ implies both joint MTS-MIV, and joint positive quadrant dependence between $Y_d$ and $Z$, and between $Y_d$ and $D$. See the proof in the appendix. Figure (ref) below summarizes the relationship between the different concepts.

figure[figure omitted — 336 chars of source]
assumption[Less endogenous instrument (LEI)] \begin{eqnarray*} \mid\rho_{Y_d D}\mid \geq \mid \rho_{Y_d Z}\mid for each \ d \in \mathcal D. \end{eqnarray*}

Assumption LEI is the same as Assumption 4 in Nevo2012, which they refer to as the “instrument less endogenous than treatment” assumption. It states that the imperfect instrument $Z$ is less correlated with the potential outcome than is the endogenous treatment $D$. In this paper, we use the shorthand “less endogenous instrument” to call this assumption. In the context of our empirical example, it is reasonable to assume that parental education is less correlated with the individual's potential wage than is the individual's own education.

assumption[Monotone treatment response (MTR)] \begin{eqnarray*} Y_d \geq Y_{d'}\ for all \ d>d'. \end{eqnarray*}

Assumption MTR states that the potential outcome weakly increases with the level of the treatment. It was introduced by Manski1997, and considered in Manski2000, Manski2009, among many others. For instance, in the returns to schooling example, it implies that the wage that a worker earns weakly increases as a function of the worker's years of schooling. We show how this assumption can help tighten the bounds derived under Assumptions BoS, SDC and LEI.

Now that we have discussed the model and our identifying assumptions, we are going to present our main identification results.

Identification results

In this section, we derive under the different assumptions discussed in the previous section bounds on the potential outcome expectation $\theta_d$, for each $d \in \mathcal D$: $LB_d \leq \theta_d \leq UB_d$. Bounds on $ATE(d,d')$ are then obtained as: $LB_d-UB_{d'} \leq ATE(d,d') \leq UB_d-LB_{d'}$.

\setcounter{equation}{0}

Identification under the same direction of correlation assumption

Assumption SDC is equivalent to $\text{E}\left[Y_d\tilde{D}\right]\text{E}\left[Y_d\tilde{Z}\right] \geq 0$, where $\tilde{D}\equiv D-\text{E}[D]$ and $\tilde{Z}\equiv Z-\text{E}[Z]$, which in turn is equivalent to: either

eqnarray[eqnarray omitted — 142 chars of source]

or

eqnarray[eqnarray omitted — 141 chars of source]

We first derive bounds on the potential outcome mean $\theta_d$ using inequalities ((ref)). Similarly, we can derive the bounds implied by inequalities ((ref)).

Inequality ((ref)) implies that, for all $(\lambda,\gamma)\in \mathbb R^2_+\setminus \left\{(0,0)\right\}$, we have

eqnarray*[eqnarray* omitted — 95 chars of source]

By factorizing $(\lambda + \gamma)$ in the above inequality, we have

eqnarray*[eqnarray* omitted — 124 chars of source]

where $\beta=\frac{\lambda}{\lambda + \gamma}$. We can normalize $\lambda + \gamma$ to lie within the interval $[0,1]$.\footnote{If $\lambda + \gamma >1$, we can multiply each side of the inequality by $\frac{1}{\lambda + \gamma+1}$, and have $\frac{\lambda + \gamma}{\lambda + \gamma+1} \in [0,1]$.} By setting $\alpha=\lambda + \gamma$, this last inequality becomes

eqnarray*[eqnarray* omitted — 113 chars of source]

Hence, Inequality ((ref)) implies that, for any $(\alpha,\beta) \in [0,1]^2$, we have

eqnarray*[eqnarray* omitted — 224 chars of source]

Intuitively, $\beta$ measures how tight is the constraint $\text{E}\left[Y_d\tilde{D}\right] \geq 0$ relatively to the constraint $ \text{E}\left[Y_d\tilde{Z}\right] \geq 0$, while $\alpha$ measures the extent to which the mixture of the constraints is binding. For example, when the constraint $\text{E}\left[Y_d\tilde{Z}\right] \geq 0$ is binding while the constraint $\text{E}\left[Y_d\tilde{D}\right] \geq 0$ is slack, then $\beta=0$, and the instrument $Z$ satisfies the zero covariance assumption, while the treatment variable $D$ is endogenous. In such a scenario, we expect the mixed constraint to be binding as well, that is $\alpha=1$. If instead, the treatment variable $D$ satisfies the zero covariance assumption, and the instrument $Z$ is endogenous, then we expect $\beta=1$ and $\alpha=1$.

The latter inequalities are respectively equivalent to:\footnote{In the case where the $ATT/ATU$ is our parameter of interest, we would bound $\theta_{d|d'}\equiv \text{E}[Y_d\vert D=d']=\frac{\text{E}[Y_d\mathbbm{1}\{D=d'\}]}{\text{E}[\mathbbm{1}\{D=d'\}]}$. In such a case, we will equivalently write these inequalities as:

eqnarray*[eqnarray* omitted — 342 chars of source]

and use the same technique we develop in this paper.}

eqnarray*[eqnarray* omitted — 192 chars of source]

where $\delta^+_{S} \equiv 1+ \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)$ and $\delta^-_{S} \equiv 1- \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)$.

Furthermore, using the identity $\mathbbm{1}\left\{D=d\right\} + \mathbbm{1}\left\{D\neq d\right\}=1$, we rewrite them as

eqnarray[eqnarray omitted — 349 chars of source]

respectively, given that $Y=Y_d$ when $D=d$.

Now, using Assumption BoS, we can bound the counterfactuals $\delta^+_{S} Y_d \mathbbm{1}\left\{D \neq d\right\}$ and $\delta^-_{S} Y_d \mathbbm{1}\left\{D \neq d\right\}$ as follows:

eqnarray*[eqnarray* omitted — 382 chars of source]

Therefore, using inequalities ((ref)) and ((ref)), it follows that

eqnarray*[eqnarray* omitted — 206 chars of source]

for any $(\alpha, \beta) \in [0,1]^2$, where we define the function $\underline{f}_d$ and $\overline{f}_d$ as

eqnarray*[eqnarray* omitted — 425 chars of source]

We can then take the supremum and the infimum of the lower and upper bounds over $(\alpha, \beta)$, respectively, to obtain the following bounds for $\theta_d$:

equation*[equation* omitted — 266 chars of source]

Similarly, using inequalities ((ref)), we derive the following bounds for $\theta_d$:

eqnarray*[eqnarray* omitted — 262 chars of source]

All these results are summarized in the following proposition.

propositionUnder Assumptions BoS and SDC, nonparametric bounds for the parameter $\theta_d$ are given by: \begin{equation*} I_{SDC}^d \equiv I_{SDC1}^d \cup I_{SDC2}^d . \end{equation*}

Proposition (ref) provides two-sided bounds on the potential outcome means, and then on the average treatment effects, which mainly relies on the bounded outcome assumption. Nevo2012 obtain two-sided bounds if $Cov(D,Z) <0$, and one-sided bounds if $Cov(D,Z) > 0$. We relax the parametric linear assumption at the expense of the bounded support assumption. In light of the following statement from the fourth paragraph of Section VI in Nevo2012\ldots However, with a nonparametric functional, it is doubtful that our assumptions on the correlations of endogenous regressors and imperfect instruments with econometric errors would prove anywhere near as fruitful,” we believe that the result of Proposition (ref) makes a positive contribution to the literature. However, we do not have a proof that the derived bounds are sharp at this point. We believe that this is an important theoretical question that can be investigated in future research.

The bounds derived in Proposition (ref) are no wider than the usual Manski worst-case bounds without instrument, as the latter bounds are a special case of ours where $\alpha=0$. Furthermore, we expect the Manski bounds derived under the strict IV exogeneity condition, $\text{E}\left[Y_d\vert Z\right]=\text{E}\left[Y_d\right]$, to be weakly narrower than the bounds $I_{SDC}^d$. We provide a heuristic proof for this conjecture in the appendix. Intuitively, we expect the optimal value of $\beta$ to be equal to 0 under the strict IV exogeneity assumption, since the treatment variable $D$ is endogenous and the intrument $Z$ satisfies the zero covariance assumption. Using this restriction, we show that the lower (upper) bounds of $I_{SDC1}^d$ and $I_{SDC2}^d$ are each less (greater) than the Manski lower (upper) bound.

It is possible that the bounds in Proposition (ref) be empty. Indeed, if the supremum and the infimum in the bounds' expressions are attained at different values of $(\alpha, \beta)$, the lower bounds of $I_{SDC1}^d$ and $I_{SDC2}^d$ could be bigger than their respective upper bounds. When that happens for both bounds $I_{SDC1}^d$ and $I_{SDC2}^d$, we say that the model (Assumptions BoS and SDC) is rejected in the data. When only $I_{SDC1}^d$ ($I_{SDC2}^d$) is empty, then the assumption that the instrument $Z$ and the treatment $D$ are positively (negatively) correlated with the potential outcome $Y_d$ is rejected.

Adding the less endogenous instrument assumption

In this subsection, we combine Assumptions SDC and LEI in order to get tighter bounds on the parameter $\theta_d$. Assumption LEI is equivalent to $\vert \frac{\text{E}\left[Y_d\tilde{D}\right]}{\sigma_D}\vert \geq \vert\frac{\text{E}\left[Y_d\tilde{Z}\right]}{\sigma_Z}\vert$. Hence, Assumptions LEI and SDC imply that either of the followings is always true:

equation*[equation* omitted — 258 chars of source]

or

equation*[equation* omitted — 256 chars of source]

Differently, we can rewrite these inequalities as either

equation[equation omitted — 248 chars of source]

or

equation[equation omitted — 248 chars of source]

Applying a similar reasoning as in the previous subsection to ((ref)), for any $(\alpha, \beta, \gamma) \in [0, 1]^3$ such that $1-\beta-\gamma \geq 0$, we have

eqnarray*[eqnarray* omitted — 361 chars of source]

Since $\beta$ and $\gamma$ belong to a 2-simplex, we can parametrize $\gamma=(1-\beta)\mu$, where $\mu \in [0,1]$. Therefore, the above inequalities are equivalent to

eqnarray*[eqnarray* omitted — 192 chars of source]

where $\delta^+_{L} \equiv 1+ \alpha\big((1-\beta)\mu (\tilde{D}\sigma_Z - \tilde{Z}\sigma_D) + \beta \tilde{D} + (1-\beta)(1-\mu) \tilde{Z} \big)$ and $\delta^-_{L} \equiv 1- \alpha\big((1-\beta)\mu (\tilde{D}\sigma_Z - \tilde{Z}\sigma_D) + \beta \tilde{D} + (1-\beta)(1-\mu) \tilde{Z} \big)$. Hence, we obtain the following bounds for $\theta_d$ from ((ref)):

equation*[equation* omitted — 276 chars of source]

Likewise, we derive bounds for $\theta_d$ from ((ref)) as

equation*[equation* omitted — 276 chars of source]

and the following proposition holds.

propositionUnder Assumptions BoS, SDC and LEI, nonparametric bounds for the parameter $\theta_d$ are given by: \begin{equation*} I_{LEI}^d \equiv I_{LEI1}^d \cup I_{LEI2}^d. \end{equation*}

The following corollary shows that $I_{LEI}^d$ is weakly tighter than $I_{SDC}^d$.

corollaryUnder Assumptions BoS, SDC and LEI, we have $$I_{LEI}^d \subseteq I_{SDC}^d.$$

The proof of Corollary (ref) follows from the fact that the bounds $I_{SDC}^d$ are a special case of the bounds $I_{LEI}^d$ where $\mu=0$.

Adding the monotone treatment response assumption

In this subsection, we derive bounds on the potential outcome mean $\theta_d$ under Assumptions BoS, SDC, LEI, and MTR . Under Assumption MTR, we have

eqnarray*[eqnarray* omitted — 340 chars of source]

where the inequality holds as $Y_d \mathbbm{1}\left\{D=j\right\} \leq Y_j \mathbbm{1}\left\{D=j\right\}=Y \mathbbm{1}\left\{D=j\right\}$ for all $j>d$ for each $d$. This result shows that under the MTR assumption, the counterfactual random variable $Y_d \mathbbm{1}\left\{D>d\right\}$ is bounded from above by the observed variable $Y \mathbbm{1}\left\{D>d\right\}$. As we can see, this assumption considerably shrinks the upper bound on $Y_d \mathbbm{1}\left\{D>d\right\}$, which would be $\overline{y}_d \mathbbm{1}\left\{D>d\right\}$ otherwise under Assumption BoS. Similarly, we also have

equation*[equation* omitted — 104 chars of source]

for each $d$. Without the MTR assumption, the lower bound on $Y_d \mathbbm{1}\left\{D<d\right\}$ would be $\underline{y}_d \mathbbm{1}\left\{D<d\right\}$ under Assumption BoS. Combining these results together, the following inequalities hold under Assumptions BoS and MTR:

eqnarray*[eqnarray* omitted — 362 chars of source]

Thus, for any $\delta \in \mathbb{R}$, we have the following bounds

eqnarray[eqnarray omitted — 518 chars of source]

Moreover, as we have

eqnarray*[eqnarray* omitted — 148 chars of source]

inequalities ((ref)) imply

eqnarray*[eqnarray* omitted — 461 chars of source]

for any $\delta \in \mathbb{R}$.

Now, recall that inequalities ((ref)) implied by Assumptions SDC and LEI yield

eqnarray*[eqnarray* omitted — 308 chars of source]

for any $(\alpha, \beta, \mu) \in [0,1]^3$, and thus we have

align*[align* omitted — 175 chars of source]

for any $(\alpha, \beta,\mu) \in [0,1]^3$, where we define the functions $\underline{m}_d$ and $\overline{m}_d$ as

eqnarray*[eqnarray* omitted — 509 chars of source]

Likewise, inequalities ((ref)) under the MTR assumption yield the following implications for $\theta_d$:

align*[align* omitted — 175 chars of source]

for any $(\alpha, \beta,\mu) \in [0,1]^3$. Therefore, we conclude that Assumptions BoS, SDC, LEI, and MTR together imply

equation*[equation* omitted — 79 chars of source]

where

align*[align* omitted — 533 chars of source]

Note that if we are not willing to impose Assumption LEI, the bounds can be obtained from $I_{MTR1}^d \cup I_{MTR2}^d$ using $\delta^+_{S}$ and $\delta^-_{S}$ instead of $\delta^+_{L}$ and $\delta^-_{L}$.

An important implication of the MTR assumption is that it weakly signs the ATE, i.e., $ATE(d,d')\geq 0$ for all $d > d'$. More precisely, bounds for $ATE(d,d')$, where $d>d'$, can be characterized as follows:

eqnarray*[eqnarray* omitted — 177 chars of source]

Inference

\setcounter{equation}{0} We want to construct confidence bounds for the set $I_{SDC}^d = I_{SDC1}^d \cup I_{SDC2}^d $. This is a special case of an intersection-union test as described in Berger1982. We are going to construct confidence regions for the sets $I_{SDC1}^d$ and $I_{SDC2}^d$ using the intersection bounds framework of CLR2013 or AS2013, and then take the union of the two confidence regions. Berger1996 showed that the union of the confidence regions has at least the same coverage rate as each confidence region. Identified sets with a similar structure have been considered in Chesher2020 and Machado2019. As we explain in Section (ref), the sets $I_{SDC1}^d$ and $I_{SDC2}^d$ could each be empty, and may not overlap. However, in our empirical illustration below, they are nonempty and overlap in all cases.

We now explain how to rewrite the intersection bounds $ I_{SDC1}^d$ in such a way that it can be easily implemented using the CKLRstata or AKS2017 Stata packages. Suppose that we draw two independent random variables $U_1$ and $U_2$ from the uniform distribution over $[0,1]$, independently of the data $(Y,D,Z)$. Then, we have

eqnarray*[eqnarray* omitted — 186 chars of source]

since $U_1$ and $U_2$ are independent of $(Y,D,Z)$. Therefore,

eqnarray*[eqnarray* omitted — 365 chars of source]

Hence, these bounds take the form of conditional moment inequalities, which can be implemented using existing inferential methods like CKLRstata or AKS2017. Similar results apply to the bounds $I_{LEI}^d$ and $I_{MTR}^d$, where there are three conditioning variables $U_1$, $U_2$, and $U_3$ instead of two. A similar technique has been proposed in Kedagni2018b where they construct confidence sets for the potential outcome means under the IV zero-covariance assumption.

Using the CKLRstata Stata package, we obtain confidence sets that asymptotically cover either the true parameter $\theta_d$ or the bounds for $\theta_d$ with pre-specified probability, through the clr3bound command.\footnote{The clr2bound command could be used to obtain confidence sets for the bounds for $\theta_d$ only.}

Empirical illustration

\setcounter{equation}{0} In this application, we use a data set drawn from the NLSYM. This data includes 3,010 young men who were ages 24-34 in 1976. It is the same data used in Card1995. In our analysis, the outcome variable is log hourly wage in cents $(lwage)$, and the treatment variable is education $(educ)$ grouped in 4 categories: less than high school $(educ< 12\ years)$, high school $(12 \leq educ < 16)$, college degree $(16 \leq educ < 18)$, and graduate $(educ \geq 18)$.\footnote{As discussed in Andresen2021, this discretization may induce some identification issues, and the results may be sensitive to it. For this reason, our empirical results should be seen as illustrative.}

Our imperfect IV is parental education. Since the work of Willis79, parental education has been used as an IV. However, an individual's ability can be dependent on her parents' ability, which is correlated with parental education. For this reason, parents' education will not be a valid instrument. This fact is documented in Kedagni20, who provided evidence that even after controlling for a measure of ability, parental education is not a good instrument. This result is contrary to the Lemke03 idea that controlling for some measure of child ability could make parental education a valid IV. Nonetheless, it is reasonable to assume that parental education has the same sign of correlation with the individual's potential wage as the correlation between the person's potential wage and her own education.\footnote{The MIV and possibly MTS assumptions also seem reasonable in this empirical example. However, our goal in this section is to show how our derived bounds can be implemented in a real-world application.} It is also likely that parental education be less endogenous than is the person's own education. Finally, as in Manski2000, we use the monotone treatment response assumption to tighten the bounds on the average returns to education.

In theory, the outcome variable $lwage$ is unbounded. For practical reasons, we follow Ginther2000 to trim the log wage. The outcome variable that we use is defined as $Y=\tau$-quantile of $lwage$ if $lwage$ is less than or equal to its $\tau$-quantile, $Y=(1-\tau)$-quantile of $lwage$ if $lwage$ is greater than or equal to its $(1-\tau)$-quantile, and $Y=lwage$ otherwise. In our empirical illustration, we set $\tau=0.05$. We construct 95% two-sided confidence sets on potential average log wages and their bounds using the clr3bound command of CKLRstata in the Stata software. We estimate the conditional expectations using the parametric method, which is the default option for this Stata command. See the appendix for more details on the implementation.

We present the results with mother's education as an IIV. The results for father's education are in the appendix. Table (ref) displays the 95% confidence sets for the potential wage means and average returns to schooling under the SDC assumption, while Table (ref) shows the confidence sets under both the SDC and LEI assumptions. Each table shows the confidence set for bounds on $\theta_d$ in the first two columns, the confidence set for the parameter $\theta_d$ in the third and forth columns, and the point estimates of the bounds in the fifth and sixth columns. The SDC+LEI bounds in Table (ref) are generally narrower than the SDC bounds in Table (ref). This suggests that Assumption LEI provides some extra identifying power to the SDC assumption. However, the bounds seem wide and less informative in both cases. For example, the confidence set for $ATE(2,1) \equiv \theta_2-\theta_1$ under SDC+LEI is $[-0.98,1.02]$, implying that the average return to college degree compared to high school education varies between $-98\%$ and $102\%$, which is not very informative for an individual making a college decision. The corresponding point estimates of the bounds seem tighter $[-0.53, 0.69]$.

table[table omitted — 1,216 chars of source]
table[table omitted — 1,224 chars of source]

Furthermore, the confidence regions for the bounds and the parameters considerably shrink and become more informative when we add the MTR assumption (see Table (ref)). We assume that the lower bound of $\theta_d$ is equal to the upper bound of $\theta_{d-1}$ whenever the former is less than the latter, because $Y_{d-1}$ cannot exceed $Y_d$ for each $d=1, 2, 3$ under the MTR assumption. Individuals with less than high school education could earn up to 94% less than high school graduates ($ATE(0,1)$). Moreover, the confidence regions for $ATE(2, 1)$ and $ATE(3,1)$ suggest that college graduates could earn up to 55% more than high school graduates, while individuals with a graduate degree earn between 41% and 65% higher wages than high school graduates (which approximately represents an annual return between 6.8% and 10.8%). Table II in Card2001 shows that the point estimates of the annual return to schooling in the US vary roughly between 5% and 13%. As we can see, our set estimates for the annual return are consistent with the existing range in the literature.

If we impose the linear structure of Nevo2012 in this empirical exercise, we obtain the following point estimate bounds for $\theta$ under SDC and MTR: $[0, 0.16] \cup [0.24, +\infty)$. Indeed, the estimated correlation between $D$ and $Z$ is $0.3824>0$, and under SDC, the point estimate bounds for $\theta$ are $(-\infty, 0.16] \cup [0.24, +\infty)$. If one is willing to further assume that $Cov(D,U) \geq 0$, then the bounds for $\theta$ reduce to $[0, 0.16]$, which imply an annual return between 0 and 16%. This range is consistent with the literature, but remains wide.

table[table omitted — 1,230 chars of source]

Conclusion

In this paper, we derive nonparametric bounds on the average treatment effect when an imperfect instrument is available. We extend Nevo2012's (Nevo2012) identification results to nonparametric models. We first assume that the sign of correlation between the imperfect instrument and the unobserved latent variables is the same as the correlation between the endogenous variable and the latent variables. We show that the MTS-MIV restrictions introduced by Manski2000, Manski2009, jointly imply this assumption. Second, we show how the assumption that the imperfect instrument is less endogenous than the treatment variable can help tighten the bounds. We also use the monotone treatment response assumption to get tighter bounds. The identified set takes the form of intersection bounds, which can be implemented using CLR2013's (CLR2013) inferential method. Finally, we illustrate our methodology using the National Longitudinal Survey of Young Men data to estimate returns to schooling.

Acknowledgements

The authors are grateful to the Editor Petra Todd, and three anonymous referees for valuable suggestions and comments. They also thank Santiago Acerenza, Otavio Bartalotti, Helle Bunzel, Ismael Mourifi\'e, Vitor Possebom, and participants at the Iowa State econometrics workshop for helpful comments. All errors are ours.

thebibliography\bibitem[\citeauthoryear{Andresen and Huber}{Andresen and Huber}{2021}]{Andresen2021} Andresen, M. E. and M. Huber (2021). \newblock Instrument-based estimation with binarized treatments: Issues and tests for the exclusion restriction. \newblock {\em The Econometrics Journal (forthcoming)\/}. \bibitem[\citeauthoryear{Andrews, Kim, and Shi}{Andrews et al.}{2017}]{AKS2017} Andrews, D. W. K., W. Kim, and X. Shi (2017). \newblock Stata commands for testing conditional moment inequalities/equalities. \newblock {\em Stata Journal\/} {\em 17\/}(1), 56--72. \bibitem[\citeauthoryear{Andrews and Shi}{Andrews and Shi}{2013}]{AS2013} Andrews, D. W. K. and X. Shi (2013). \newblock Inference based on conditional moment inequalities. \newblock {\em Econometrica\/} {\em 81}, 609--666. \bibitem[\citeauthoryear{Berger}{Berger}{1982}]{Berger1982} Berger, R. L. (1982). \newblock Multiparameter hypothesis testing and acceptance sampling. \newblock {\em Technometrics\/} {\em 24}, 295--300. \bibitem[\citeauthoryear{Berger and Hsu}{Berger and Hsu}{1996}]{Berger1996} Berger, R. L. and J. C. Hsu (1996). \newblock Bioequivalence trials, intersection-union tests and equivalence confidence sets. \newblock {\em Statistical Science\/} {\em 11\/}(4), 283--319. \bibitem[\citeauthoryear{Bhattacharya, Shaikh, and Vytlacil}{Bhattacharya et al.}{2012}]{BSV2012} Bhattacharya, J., A. Shaikh, and E. Vytlacil (2012). \newblock Treatment effect bounds: An application to swan-ganz catheterization. \newblock {\em Journal of Econometrics\/} {\em 168\/}(2), 223--243. \bibitem[\citeauthoryear{Card}{Card}{1995}]{Card1995} Card, D. (1995). \newblock Using geographic variation in college proximity to estimate the return to schooling. \newblock In L. N. Christofides, E. K. Grant, and R. Swidinsky (Eds.), {\em Aspects of Labour Market Behaviour: Essays in Honour of John Vanderkamp}, pp.\ 201--222. Toronto, Canada: University of Toronto Press. \bibitem[\citeauthoryear{Card}{Card}{2001}]{Card2001} Card, D. (2001). \newblock Estimating the return to schooling: Progress on some persistent econometric problems. \newblock {\em Econometrica\/} {\em 69}, 1127--1160. \bibitem[\citeauthoryear{Chernozhukov, Kim, Lee, and Rosen}{Chernozhukov et al.}{2015}]{CKLRstata} Chernozhukov, V., W. Kim, S. Lee, and A. M. Rosen (2015). \newblock Implementing intersection bounds in stata. \newblock {\em Stata Journal\/} {\em 15\/}(1), 21--44. \bibitem[\citeauthoryear{Chernozhukov, Lee, and Rosen}{Chernozhukov et al.}{2013}]{CLR2013} Chernozhukov, V., S. Lee, and A. M. Rosen (2013). \newblock Intersection bounds: Estimation and inference. \newblock {\em Econometrica\/} {\em 81\/}(2), 667--737. \bibitem[\citeauthoryear{Chesher and Rosen}{Chesher and Rosen}{2020}]{Chesher2020} Chesher, A. and A. Rosen (2020). \newblock Generalized instrumental variable models, methods, and applications. \newblock In J. J. H. S. N. Durlauf, L. P. Hansen and R. L. Matzkin (Eds.), {\em Handbook of Econometrics}, Volume 7A, pp.\ 1--110. Amsterdam: North-Holland: Elsevier. \bibitem[\citeauthoryear{Conley, Hansen, and Rossi}{Conley et al.}{2012}]{Conley2012} Conley, T. G., C. B. Hansen, and P. E. Rossi (2012). \newblock Plausibly exogenous. \newblock {\em Review of Economics and Statistics\/} {\em 94\/}(1), 260--272. \bibitem[\citeauthoryear{Ginther}{Ginther}{2000}]{Ginther2000} Ginther, D. K. (2000). \newblock Alternative estimates of the effect of schooling on earnings. \newblock {\em Review of Economics and Statistics\/} {\em 82\/}(1), 103--116. \bibitem[\citeauthoryear{Hotz, Mullin, and Sanders}{Hotz et al.}{1997}]{Hotz1997} Hotz, V. J., C. Mullin, and S. Sanders (1997). \newblock Bounding causal effects using data from a contaminated natural experiment: Analyzing the effects of teenage childbearing. \newblock {\em Review of Economic Studies\/} {\em 64\/}(4), 575--603. \bibitem[\citeauthoryear{K\'edagni, Li, and Mourifi\'e}{K\'edagni et al.}{2018}]{Kedagni2018b} K\'edagni, D., L. Li, and I. Mourifi\'e (2018). \newblock Bounding average returns to schooling using unconditional moment restrictions. \newblock Working Paper 18022, Department of Economics, Iowa State University. \bibitem[\citeauthoryear{K\'edagni and Mourifi\'e}{K\'edagni and Mourifi\'e}{2020}]{Kedagni20} K\'edagni, D. and I. Mourifi\'e (2020). \newblock Generalized instrumental inequalities: Testing the instrumental variable independence assumption. \newblock {\em Biometrika\/} {\em 107\/}(3), 661--675. \bibitem[\citeauthoryear{Lemke and Rischall}{Lemke and Rischall}{2003}]{Lemke03} Lemke, R. J. and I. C. Rischall (2003). \newblock Skill, parental income, and iv estimation of the returns to schooling. \newblock {\em Applied Economics Letters\/} {\em 10\/}(5), 281–--286. \bibitem[\citeauthoryear{Machado, Shaikh, and Vytlacil}{Machado et al.}{2019}]{Machado2019} Machado, C., A. Shaikh, and E. Vytlacil (2019). \newblock Instrumental variables and the sign of the average treatment effect. \newblock {\em Journal of Econometrics\/} {\em 212}, 522--555. \bibitem[\citeauthoryear{Manski}{Manski}{1990}]{Manski1990} Manski, C. F. (1990). \newblock Nonparametric bounds on treatment effects. \newblock {\em American Economic Reviews, Papers and Proceedings of the Hundred and Second Annual Meeting of the American Economic Association\/} {\em 80\/}(2), 319--323. \bibitem[\citeauthoryear{Manski}{Manski}{1994}]{Manski1994} Manski, C. F. (1994). \newblock The selection problem. \newblock {\em in Advances Economics, Sixth World Congress, C. Sims (ed.), Cambridge University Press\/} {\em 1}, 143--170. \bibitem[\citeauthoryear{Manski}{Manski}{1997}]{Manski1997} Manski, C. F. (1997). \newblock Monotone treatment response. \newblock {\em Econometrica\/} {\em 65\/}(6), 1311--1334. \bibitem[\citeauthoryear{Manski and Pepper}{Manski and Pepper}{2000}]{Manski2000} Manski, C. F. and J. Pepper (2000). \newblock Monotone instrumental variables: With an application to the returns to schooling. \newblock {\em Econometrica\/} {\em 68}, 997--1010. \bibitem[\citeauthoryear{Manski and Pepper}{Manski and Pepper}{2009}]{Manski2009} Manski, C. F. and J. Pepper (2009). \newblock More on monotone instrumental variables. \newblock {\em Econometrics Journal\/} {\em 12}, S200--S216. \bibitem[\citeauthoryear{Masten and Poirier}{Masten and Poirier}{2020}]{Masten2020} Masten, M. A. and A. Poirier (2020). \newblock Salvaging falsified instrumental variable models. \newblock {\em Econometrica (forthcoming)\/}. \bibitem[\citeauthoryear{Mourifi\'e, Henry, and M\'eango}{Mourifi\'e et al.}{2020}]{Mourifie2020} Mourifi\'e, I., M. Henry, and R. M\'eango (2020). \newblock Sharp bounds and testability of a roy model of stem major choices. \newblock {\em Journal of Political Economy\/} {\em 8\/}(128), 3220--3283. \bibitem[\citeauthoryear{Nevo and Rosen}{Nevo and Rosen}{2012}]{Nevo2012} Nevo, A. and A. Rosen (2012). \newblock Identification with imperfect instruments. \newblock {\em The Review of Economics and Statistics\/} {\em 94\/}(3), 659--671. \bibitem[\citeauthoryear{Willis and Rosen}{Willis and Rosen}{1979}]{Willis79} Willis, R. and S. Rosen (1979). \newblock Education and self-selection. \newblock {\em Journal of Political Economy\/} {\em 87\/}(5), Pt2:S7--36.