EconBase
← Back to paper

Nonparametric Bounds on Treatment Effects with Imperfect Instruments

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

56,944 characters

Nonparametric Bounds on Treatment Effects with Imperfect Instruments


\nonstopmode

  \begin{abstract}
   This paper extends the identification results in \cite{Nevo2012} to nonparametric models. We derive nonparametric bounds on the average treatment effect when an imperfect instrument is available. As in \cite{Nevo2012}, we assume that the correlation between the imperfect instrument and the unobserved latent variables has the same sign as the correlation between the endogenous variable and the latent variables. We show that the monotone treatment selection and monotone instrumental variable restrictions, introduced by \cite{Manski2000, Manski2009}, jointly imply this assumption.
Moreover, we show how the monotone treatment response assumption can help tighten the bounds. The identified set can be written in the form of intersection bounds, which is more conducive to inference. We illustrate our methodology using the National Longitudinal Survey of Young Men data to estimate returns to schooling.

  \keywords{Imperfect instrumental variables, nonparametric bounds, average treatment effect, monotone treatment response.}

  \end{abstract}


\section{Introduction}

   The use of an instrumental variable (IV) is a popular solution to deal with endogeneity in social sciences. However, this approach may yield misleading conclusions when the instrument is invalid. A valid instrument must be uncorrelated with the unobservables in the model.\footnote{A stronger version of this condition is that the instrument is statistically (or mean) independent of the unobservables.} This requirement is often difficult to justify, and may call into question empirical findings. For this reason, \cite{Nevo2012} derived bounds on the parameters of interest (e.g., the average treatment effect) in parametric models under weaker conditions. They first assume that the sign of correlation between the imperfect IV (IIV)\footnote{An instrument that is potentially correlated with the unobservables.} and the unobserved latent variables is the same as that of the correlation between the endogenous variable and the latent variables. Second, they add the assumption that the correlation between the IIV and the latent variables is less than the correlation between the endogenous variable and the latent variables to tighten the bounds on the parameters of interest.

In this paper, we derive nonparametric bounds on the average treatment effect with an imperfect IV under the above assumptions when the outcome variable has bounded support. We introduce the concept of binarized MTS-MIV, which is implied by the monotone treatment selection (MTS) and monotone IV (MIV) assumptions developed by \cite{Manski2000, Manski2009}. We show that the correlation between a binarized MTS-MIV and the unobserved latent variables has the same sign as the correlation between the endogenous variable and the latent variables. Hence, we link the \cite{Nevo2012} same direction of correlation assumption to the \cite{Manski2000} monotone treatment selection and monotone IV assumptions.
We believe this result is new in the literature. Furthermore, we show how additional restrictions such as the less endogenous instrument, and the monotone treatment response can help tighten the bounds. As in \cite{Nevo2012}, the bounds take the form of intersection bounds and can be implemented using the inferential methods developed by \cite{CLR2013} or \cite{AS2013}. We illustrate our methodology using the National Longitudinal Survey of Young Men (NLSYM) data to estimate returns to schooling.

There is an increasing interest in the identification of causal effects with imperfect instrumental variables. Recently, \cite{Masten2020} developed a methodology that allows researchers to consider continuous relaxations of the IV models when they are refuted by the data. Their approach is data driven as it exploits the extent of falsification of the model to construct the identified set for the parameter of interest. Although their method helps salvage some invalid IVs, their identifying assumptions may seem difficult to interpret. Our approach, as well as that of \cite{Nevo2012} and \cite{Manski2000, Manski2009}, is not data driven and has clearer interpretations of the identifying assumptions. This paper also relates to the work of \cite{Kedagni2018b}, who consider weaker version of the mean independence assumption for the IV model. They derived nonparametric bounds on the average treatment effect under unconditional moment restrictions for the IV. Several other papers have also studied identification of model parameters when the IV is invalid using a framework different from ours; see \cite{Hotz1997}, \cite{Conley2012}, among others.

The remainder of the paper is organized as follows. Section \ref{anaF} presents the model, the assumptions and their link with the literature. In Section \ref{ident}, we derive our main identification results. We discuss inference and implementation in Section \ref{inference}. Section \ref{empirical} presents an empirical illustration of our proposed methodology, while Section \ref{conclusion} concludes. Proofs and additional results are relegated to the appendix.


\section{Analytical Framework}\label{anaF}
Consider the following potential outcome model (POM)
\begin{eqnarray}\label{eq:pom}
Y=\sum_{d=1}^T Y_d\mathbbm{1}\left\{D=d\right\},
\end{eqnarray}
where $Y$ is the outcome variable taking values in $\mathcal Y \subset \mathbb R$, $D$ is a discrete endogenous treatment variable taking values in $\mathcal D=\{1,2,\ldots,T\}$, $Y_{d}$ is the potential outcome that would have been observed if the treatment $D$ had externally been set to $d$. Let $Z\in \mathcal Z \subseteq \mathbb R$ be an imperfect IV in the sense that it may be correlated with the potential outcome $Y_d$. In what follows, we assume that the random variable $Y_d$ is integrable, i.e., $\text{E}[Y_d]<\infty$. The objects of interest in this paper are the potential outcome means $\theta_d \equiv \text{E}[Y_d]$, for all $d \in \mathcal D$, and some treatment effects $ATE(d,d')\equiv \theta_d-\theta_{d'}$, for $d,d' \in \mathcal D$. We allow for heterogeneous treatment effects, so that $ATE(d,d')$ may vary across $(d,d')$. The methodology that we develop in this paper can also be used to identify other commonly used parameters of interest such as the average treatment effect on the treated $ATT(d,d')\equiv\text{E}[Y_d-Y_{d'}\vert D=d]$, and the average treatment effect on the untreated $ATU(d,d')\equiv\text{E}[Y_d-Y_{d'}\vert D=d']$. But, for the sake of clarity of the exposition, we focus our attention on the $ATE$.

We observe a random sample of the vector $(Y,D,Z)$. For simplicity, we drop exogenous covariates from the analysis.
For example, $Y$ could be earnings, $D$ years of schooling, and $Z$ parental education. In this example, $Y_d$ is the potential earnings for an individual with $d$ years of schooling.
We now state our main identifying assumptions:
\begin{assumption}[Bounded support (BoS)]\label{BS}
\begin{eqnarray*}
Supp(Y_d\vert D\neq d)= Supp(Y_d \vert D = d) =\left[\underline{y}_d,\overline{y}_d\right]\ \text{ for each }\ d \in \mathcal{D}.
\end{eqnarray*}
\end{assumption}
Assumption BoS states that the support of the counterfactual outcome is the same as that of the factual. It is standard and similar to the usual bounded outcome assumption considered in \cite{Manski1990, Manski1994}, and many other papers. Like in \cite{Kedagni2018b}, it allows the support of the potential outcome $Y_d$ to vary across all treatment levels $d$.

\begin{assumption}[Same direction of correlation (SDC)]\label{SDC}
\begin{eqnarray*}
Cov\left(Y_d,D\right) Cov\left(Y_d,Z\right)\geq 0\  \text{ for each }\ d \in \mathcal{D}.
\end{eqnarray*}
\end{assumption}

Assumption SDC is equivalent to Assumption 3 in \cite{Nevo2012}.  It states that the correlation between the imperfect instrument $Z$ and the potential outcome $Y_d$ has weakly the same sign as the correlation between the endogenous treatment $D$ and the potential outcome. For example, it is documented that parental education is not a valid instrument; see \cite{Kedagni20}, \cite{Mourifie2020}, among many others. However, one could assume that parental education has the same sign of correlation with the potential earnings as does the individual's education. Note that if either the treatment $D$ or the instrument $Z$ is exogenous, this assumption holds. If Assumption BoS holds from the definition of the outcome variable (e.g., market share lies between 0 and 1),
Assumption SDC has a testable implication. Indeed, when BoS holds and the bounds derived in Proposition \ref{thm1} for $\theta_d$ under BoS and SDC are empty, then SDC is rejected.

Assumption SDC can be seen as a weaker version of the concepts of monotone IV (MIV: $\text{E}\left[Y_d \vert Z=z\right]$ is monotone in $z$ for all $d$) and monotone treatment selection (MTS: $\text{E}\left[Y_d \vert D=\ell\right]$ is monotone in $\ell$ for all $d$) developed by \cite{Manski2000, Manski2009}. To show this result, we introduce the concept of binarized MTS-MIV that is intermediate between SDC and MTS-MIV.
We use the following notation.
\begin{notation}
Denote $g_d^+(j)=\text{E}[Y_d|D\geq j]$, $g_d^-(j)=\text{E}[Y_d|D < j]$, $h_d^+(z)=\text{E}[Y_d|Z\geq z]$, $h_d^-(z)=\text{E}[Y_d|Z < z]$.
$\rho_{UV}$ denotes the coefficient of correlation between two random variables $U$ and $V$.
\end{notation}
\begin{definition}
The variable $Z$ is a \textit{binarized MTS-MIV} for $D$ if for each $d \in \mathcal D$,
\begin{eqnarray}
\left(g_d^+(j)-g_d^-(j)\right)\left(h_d^+(z)-h_d^-(z)\right)\geq 0\ \text{ for all } j,\ z. \label{eq:BMTSMIV}
\end{eqnarray}
\end{definition}
In words, we say that $Z$ is a binarized MTS-MIV for $D$ if all binarized treatments $\mathbbm{1}\{D\geq j\}$ satisfy the MTS restriction, and all binarized instruments $\mathbbm{1}\{Z\geq z\}$ satisfy the MIV restriction.
\begin{remark}
\textnormal{
If $Z$ is a binarized MTS-MIV for $D$ then the functions $g_d^+$ and $g_d^-$ do not cross, nor do the functions $h_d^+$ and $h_d^-$ for all $d$; that is, either [$g_d^+(j) \geq g_d^-(j)$ for all $j$ and $h_d^+(z) \geq h_d^-(z)$ for all $z$] or [$g_d^+(j) \leq g_d^-(j)$ for all $j$ and $h_d^+(z) \leq h_d^-(z)$ for all $z$]. Moreover, if $g_d^+ \geq g_d^-$ for some $d$ then $h_d^+ \geq h_d^-$, and vice versa.
}
\end{remark}
Lemma \ref{lem1} shows that MTS-MIV is a sufficient condition for binarized MTS-MIV, while Lemma \ref{lem2} shows that binarized MTS-MIV is a sufficient condition for SDC.
\begin{lemma}\label{lem1}
MTS-MIV in the same direction for $D$ and $Z$ implies that $Z$ is a binarized MTS-MIV for $D$.
\end{lemma}

\begin{lemma}\label{lem2}
If $Z$ is a binarized MTS-MIV for $D$, then Assumption SDC holds.
\end{lemma}
\begin{remark}
\textnormal{
From Lemmas \ref{lem1} and \ref{lem2}, we conclude that MTS-MIV in the same direction implies Assumption SDC. Moreover, when both the treatment $D$ and the imperfect instrument $Z$ are binary, MTS-MIV in the same direction,  binarized MTS-MIV and Assumption SDC are equivalent. However, Example \ref{ex.0921} in the appendix shows a case where binarized MTS-MIV holds, but the joint MTS-MIV fails.}

\textnormal{
If MTS and MIV hold in the opposite directions, then SDC will not hold. Instead, the ``opposite directions of correlation''  assumption $Cov(Y_d,D) Cov(Y_d,Z) \leq 0$ holds. The identification strategy developed in this paper can easily be adapted to this case. In general, if the directions of the MTS and MIV assumptions are unknown, SDC is not weaker than MTS-MIV.
}
\end{remark}

Another sufficient condition for binarized MTS-MIV is the joint positive quadrant dependence between the potential outcome $Y_d$ and the instrument $Z$, and between the potential outcome $Y_d$ and the treatment $D$.
The concept of positive quadrant dependence has been considered by \cite{BSV2012} in a different framework. Two random variables $\varepsilon$ and $\nu$ are positive quadrant dependent (PQD) if
\begin{eqnarray*}
\text{P}(\varepsilon \leq t_0 \vert \nu < t_1) \geq \text{P}(\varepsilon \leq t_0)\ \text{ for all }\ t_0, t_1.
\end{eqnarray*}
As pointed out by \cite{BSV2012}, the PQD assumption implies that
\begin{eqnarray*}
\text{P}(\varepsilon \leq t_0 \vert \nu < t_1) \geq \text{P}(\varepsilon \leq t_0 \vert \nu \geq t_1)\ \text{ for all }\ t_0, t_1.
\end{eqnarray*}
From this implication, we conclude that if $Y_d$ and $Z$ are PQD, then the distribution $Y_d$ conditional on $\{Z \geq z\}$ first-order stochastically dominates that of $Y_d$ conditional on $\{Z < z\}$ for all $z$. Therefore, $\text{E}[Y_d \vert Z \geq z] \geq \text{E}[Y_d \vert Z < z]$, i.e., $h_d^+(z)-h_d^-(z) \geq 0$ for all~$z$. Similarly, if $Y_d$ and $D$ are PQD, then $g_d^+(j)-g_d^-(j) \geq 0$ for all $j$. Hence, binarized MTS-MIV holds.

Note that joint positive quadrant dependence between $Y_d$ and $Z$,  and between $Y_d$ and $D$ does not imply joint MTS-MIV, and the converse does not hold either. However, joint positive regression dependence (PRD) between $Y_d$ and $Z$ (i.e., $\text{P}(Y_d > y \vert Z=z)$ is nondecreasing in $z$ for all $y$), and between $Y_d$ and $D$ implies both joint MTS-MIV, and joint positive quadrant dependence between $Y_d$ and $Z$,  and between $Y_d$ and $D$. See the proof in the appendix. Figure \ref{fig:imp} below summarizes the relationship between the different concepts.

\begin{figure}[!htbp]
        \centering
        \begin{tikzcd}[arrows=Rightarrow]
		\textrm{Joint PRD} \arrow[r] \arrow[d] & \textrm{Joint PQD} \arrow[d] \arrow[dr]\\
		\textrm{Joint MTS-MIV} \arrow[r] & \textrm{Binarized MTS-MIV} \arrow[r] & \textrm{SDC}
		\end{tikzcd}
        \caption{Implications Between Assumptions}\label{fig:imp}
\end{figure}

\begin{assumption}[Less endogenous instrument (LEI)]\label{LEI}
\begin{eqnarray*}
\mid\rho_{Y_d D}\mid \geq \mid \rho_{Y_d Z}\mid \text{ for each }\ d \in \mathcal D.
\end{eqnarray*}
\end{assumption}

Assumption LEI is the same as Assumption 4 in \cite{Nevo2012}, which they refer to as the ``instrument less endogenous than treatment'' assumption. It states that the imperfect instrument $Z$ is less correlated with the potential outcome than is the endogenous treatment $D$. In this paper, we use the shorthand ``less endogenous instrument'' to call this assumption. In the context of our empirical example, it is reasonable to assume that parental education is less correlated with the individual's potential wage than is the individual's own education.

\begin{assumption}[Monotone treatment response (MTR)]\label{MTR}
\begin{eqnarray*}
Y_d \geq Y_{d'}\ \text{ for all }\ d>d'.
\end{eqnarray*}
\end{assumption}
Assumption MTR states that the potential outcome weakly increases with the level of the treatment. It was introduced by \cite{Manski1997}, and considered in \cite{Manski2000, Manski2009}, among many others. For instance, in the returns to schooling example, it implies that the wage that a worker earns weakly increases as a function of the worker's years of schooling. We show how this assumption can help tighten the bounds derived under Assumptions BoS, SDC and LEI.

Now that we have discussed the model and our identifying assumptions, we are going to present our main identification results.

\section{Identification results}\label{ident}
In this section, we derive under the different assumptions discussed in the previous section bounds on the potential outcome expectation $\theta_d$, for each $d \in \mathcal D$: $LB_d \leq \theta_d \leq UB_d$. Bounds on $ATE(d,d')$ are then obtained as: $LB_d-UB_{d'} \leq ATE(d,d') \leq UB_d-LB_{d'}$.

\setcounter{equation}{0}
\subsection{Identification under the same direction of correlation assumption}
Assumption SDC is equivalent to $\text{E}\left[Y_d\tilde{D}\right]\text{E}\left[Y_d\tilde{Z}\right] \geq 0$, where $\tilde{D}\equiv D-\text{E}[D]$ and $\tilde{Z}\equiv Z-\text{E}[Z]$, which in turn is equivalent to: either
\begin{eqnarray}
    \text{E}\left[Y_d\tilde{D}\right] \geq 0  \ \text{  and  } \ \text{E}\left[Y_d\tilde{Z}\right] \geq 0,
    \label{eq.SDC1}
\end{eqnarray}
 or
\begin{eqnarray}
    \text{E}\left[Y_d\tilde{D}\right] \leq 0 \ \text{  and  } \ \text{E}\left[Y_d\tilde{Z}\right] \leq 0.
    \label{eq.SDC2}
\end{eqnarray}
We first derive bounds on the potential outcome mean $\theta_d$ using inequalities (\ref{eq.SDC1}). Similarly, we can derive the bounds implied by inequalities (\ref{eq.SDC2}).

Inequality (\ref{eq.SDC1}) implies that, for all $(\lambda,\gamma)\in \mathbb R^2_+\setminus \left\{(0,0)\right\}$, we have
\begin{eqnarray*}
\text{E}\left[Y_d \left(\lambda\tilde{D}+\gamma\tilde{Z}\right)\right] \geq 0.
\end{eqnarray*}
By factorizing $(\lambda + \gamma)$ in the above inequality, we have
\begin{eqnarray*}
    \text{E}\left[Y_d (\lambda + \gamma) \left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)\right] \geq 0,
\end{eqnarray*}
where $\beta=\frac{\lambda}{\lambda + \gamma}$. We can normalize $\lambda + \gamma$ to lie within the interval $[0,1]$.\footnote{If $\lambda + \gamma >1$, we can multiply each side of the inequality by $\frac{1}{\lambda + \gamma+1}$, and have $\frac{\lambda + \gamma}{\lambda + \gamma+1} \in [0,1]$.} By setting $\alpha=\lambda + \gamma$, this last inequality becomes
\begin{eqnarray*}
    \text{E}\left[Y_d \alpha \left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right) \right] \geq 0.
\end{eqnarray*}

Hence, Inequality (\ref{eq.SDC1}) implies that, for any $(\alpha,\beta) \in [0,1]^2$, we have
\begin{eqnarray*}
    \text{E}\left[Y_d \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)\right] \geq 0\ \text{  and  } \
    \text{E}\left[-Y_d \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)\right] \leq 0.
\end{eqnarray*}
Intuitively, $\beta$ measures how tight is the constraint $\text{E}\left[Y_d\tilde{D}\right] \geq 0$ relatively to the constraint $ \text{E}\left[Y_d\tilde{Z}\right] \geq 0$, while $\alpha$ measures the extent to which the mixture of the constraints is binding. For example, when the constraint $\text{E}\left[Y_d\tilde{Z}\right] \geq 0$ is binding while the constraint $\text{E}\left[Y_d\tilde{D}\right] \geq 0$ is slack, then $\beta=0$, and the instrument $Z$ satisfies the zero covariance assumption, while the treatment variable $D$ is endogenous. In such a scenario, we expect the mixed constraint to be binding as well, that is $\alpha=1$. If instead, the treatment variable $D$ satisfies the zero covariance assumption, and the instrument $Z$ is endogenous, then  we expect $\beta=1$ and $\alpha=1$.

The latter inequalities are respectively equivalent to:\footnote{In the case where the $ATT/ATU$ is our parameter of interest, we would bound $\theta_{d|d'}\equiv \text{E}[Y_d\vert D=d']=\frac{\text{E}[Y_d\mathbbm{1}\{D=d'\}]}{\text{E}[\mathbbm{1}\{D=d'\}]}$. In such a case, we will equivalently write these inequalities as: \begin{eqnarray*}
    \text{E}\left[Y_d\left(\mathbbm{1}\{D=d'\}+ \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)\right)\right] &\geq& \text{E}[Y_d \mathbbm{1}\{D=d'\}] \\
    \text{E}\left[Y_d\left(\mathbbm{1}\{D=d'\}- \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)\right)\right] &\leq& \text{E}[Y_d \mathbbm{1}\{D=d'\}],
\end{eqnarray*}
and use the same technique we develop in this paper.}
\begin{eqnarray*}
    \text{E}\left[\delta^+_{S} Y_d \right] \geq \text{E}[Y_d ] \equiv \theta_d \ \text{  and  } \
    \text{E}\left[\delta^-_{S} Y_d\right] \leq \text{E}[Y_d] \equiv \theta_d,
\end{eqnarray*}
where $\delta^+_{S} \equiv 1+ \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)$ and $\delta^-_{S} \equiv 1- \alpha\left(\beta \tilde{D} + (1-\beta) \tilde{Z} \right)$.

Furthermore, using the identity $\mathbbm{1}\left\{D=d\right\} + \mathbbm{1}\left\{D\neq d\right\}=1$, we rewrite them as
\begin{eqnarray}
    \text{E}\left[\delta^+_{S} Y \mathbbm{1}\left\{D=d\right\} + \delta^+_{S} Y_d \mathbbm{1}\left\{D \neq d\right\} \right] &\geq& \theta_d,
    \label{eq.SDC11}
    \\
    \text{E}\left[\delta^-_{S} Y \mathbbm{1}\left\{D=d\right\} + \delta^-_{S} Y_d \mathbbm{1}\left\{D \neq d\right\} \right] &\leq&  \theta_d,
    \label{eq.SDC12}
\end{eqnarray}
respectively, given that $Y=Y_d$ when $D=d$.

Now, using Assumption BoS, we can bound the counterfactuals $\delta^+_{S} Y_d \mathbbm{1}\left\{D \neq d\right\}$ and $\delta^-_{S} Y_d \mathbbm{1}\left\{D \neq d\right\}$ as follows:
\begin{eqnarray*}
    \delta^+_{S} Y_d \mathbbm{1}\left\{D \neq d\right\}
    &\leq& \max\left\{ \delta^+_{S} \underline{y}_d, \delta^+_{S} \overline{y}_d \right\}\mathbbm{1}\left\{D \neq d\right\}
    \\
    \delta^-_{S} Y_d\mathbbm{1}\left\{D \neq d\right\}
    &\geq& \min\left\{ \delta^-_{S} \underline{y}_d,\delta^-_{S} \overline{y}_d \right\}\mathbbm{1}\left\{D \neq d\right\}.
\end{eqnarray*}
Therefore, using inequalities (\ref{eq.SDC11}) and (\ref{eq.SDC12}), it follows that
\begin{eqnarray*}
    \text{E} \Big[ \overline{f}_d\left( Y, D, \delta^+_{S} \right) \Big] \geq \theta_d \text{  and  }
    \text{E} \Big[ \underline{f}_d\left( Y, D, \delta^-_{S} \right) \Big] \leq \theta_d
\end{eqnarray*}
for any $(\alpha, \beta) \in [0,1]^2$, where we define the function $\underline{f}_d$ and $\overline{f}_d$ as
\begin{eqnarray*}
    \underline{f}_d\left(Y, D, \delta \right) &\equiv&
    \mathbbm{1}\left\{D=d\right\}\delta Y + \mathbbm{1}\left\{D\neq d\right\} \min \left\{ \delta \underline{y}_d, \delta \overline{y}_d \right\}
    \\
    \overline{f}_d\left(Y, D, \delta \right) &\equiv&
    \mathbbm{1}\left\{D=d\right\}\delta Y + \mathbbm{1}\left\{D\neq d\right\} \max \left\{ \delta \underline{y}_d, \delta \overline{y}_d \right\}.
\end{eqnarray*}
We can then take the supremum and the infimum of the lower and upper bounds over $(\alpha, \beta)$, respectively, to obtain the following bounds for $\theta_d$:
\begin{equation*}
  I_{SDC1}^d \equiv \left[ \sup_{(\alpha, \beta) \in [0,1]^2} \text{E} \Big[ \underline{f}_d\left( Y, D, \delta^-_{S} \right) \Big],
    \inf_{(\alpha, \beta) \in [0,1]^2} \text{E} \Big[ \overline{f}_d\left( Y, D, \delta^+_{S} \right) \Big] \right].
\end{equation*}
Similarly, using inequalities (\ref{eq.SDC2}), we derive the following bounds for $\theta_d$:
\begin{eqnarray*}
I_{SDC2}^d &\equiv& \left[  \sup_{(\alpha, \beta) \in [0,1]^2} \text{E} \Big[ \underline{f}_d\left( Y, D,\delta^+_{S} \right) \Big], \inf_{(\alpha, \beta) \in [0,1]^2} \text{E} \Big[ \overline{f}_d\left( Y, D, \delta^-_{S} \right) \Big] \right].
\end{eqnarray*}
All these results are summarized in the following proposition.
\begin{proposition}\label{thm1}
Under Assumptions BoS and SDC, nonparametric bounds for the parameter $\theta_d$ are given by:
\begin{equation*}
I_{SDC}^d \equiv I_{SDC1}^d  \cup  I_{SDC2}^d .
\end{equation*}
\end{proposition}
Proposition \ref{thm1} provides two-sided bounds on the potential outcome means, and then on the average treatment effects, which mainly relies on the bounded outcome assumption. \cite{Nevo2012} obtain two-sided bounds if $Cov(D,Z) <0$, and one-sided bounds if $Cov(D,Z) > 0$. We relax the parametric linear assumption at the expense of the bounded support assumption. In light of the following statement from the fourth paragraph of Section VI in \cite{Nevo2012} ``\textit{\ldots However, with a nonparametric functional, it is doubtful that our assumptions on the correlations of endogenous regressors and imperfect instruments with econometric errors would prove anywhere near as fruitful},'' we believe that the result of Proposition \ref{thm1} makes a positive contribution to the literature. However, we do not have a proof that the derived bounds are sharp at this point. We believe that this is an important theoretical question that can be investigated in future research.

The bounds derived in Proposition \ref{thm1} are no wider than the usual Manski worst-case bounds without instrument, as the latter bounds are a special case of ours where $\alpha=0$. Furthermore, we expect the Manski bounds derived under the strict IV exogeneity condition, $\text{E}\left[Y_d\vert Z\right]=\text{E}\left[Y_d\right]$, to be weakly narrower than the bounds $I_{SDC}^d$. We provide a heuristic proof for this conjecture in the appendix. Intuitively, we expect the optimal value of $\beta$ to be equal to 0 under the strict IV exogeneity assumption, since the treatment variable $D$ is endogenous and the intrument $Z$ satisfies the zero covariance assumption. Using this restriction, we show that the lower (upper) bounds of $I_{SDC1}^d$ and $I_{SDC2}^d$ are each less (greater) than the Manski lower (upper) bound.

It is possible that the bounds in Proposition \ref{thm1} be empty. Indeed, if the supremum and the infimum in the bounds' expressions are attained at different values of $(\alpha, \beta)$, the lower bounds of $I_{SDC1}^d$ and $I_{SDC2}^d$ could be bigger than their respective upper bounds. When that happens for both bounds $I_{SDC1}^d$ and $I_{SDC2}^d$, we say that the model (Assumptions BoS and SDC) is rejected in the data. When only $I_{SDC1}^d$ ($I_{SDC2}^d$) is empty, then the assumption that the instrument $Z$ and the treatment $D$ are positively (negatively) correlated with the potential outcome $Y_d$ is rejected.

\subsection{Adding the less endogenous instrument assumption}
In this subsection, we combine Assumptions SDC and LEI in order to get tighter bounds on the parameter $\theta_d$.
Assumption LEI is equivalent to $\vert \frac{\text{E}\left[Y_d\tilde{D}\right]}{\sigma_D}\vert \geq \vert\frac{\text{E}\left[Y_d\tilde{Z}\right]}{\sigma_Z}\vert$.
Hence, Assumptions LEI and SDC imply that either of the followings is always true:
\begin{equation*}
    \frac{\text{E}\left[Y_d\tilde{D}\right]}{\sigma_D}
    \geq \frac{\text{E}\left[Y_d\tilde{Z}\right]}{\sigma_Z},
     \qquad \text{E}\left[Y_d\tilde{D}\right] \geq 0
    \qquad \textrm{and} \qquad \text{E}\left[Y_d\tilde{Z}\right] \geq 0,
\end{equation*}
or
\begin{equation*}
  \frac{\text{E}\left[Y_d\tilde{D}\right]}{\sigma_D}
    \leq \frac{\text{E}\left[Y_d\tilde{Z}\right]}{\sigma_Z},
     \qquad \text{E}\left[Y_d\tilde{D}\right] \leq 0
    \qquad \textrm{and} \qquad \text{E}\left[Y_d\tilde{Z}\right] \leq 0.
\end{equation*}
Differently, we can rewrite these inequalities as either
\begin{equation}
    \text{E}\left[Y_d \left(\tilde{D}\sigma_Z - \tilde{Z}\sigma_D\right)\right] \geq 0,
    \qquad \text{E}\left[Y_d\tilde{D}\right] \geq 0
    \qquad \textrm{and} \qquad \text{E}\left[Y_d\tilde{Z}\right] \geq 0,
    \label{eq.LEI1}
\end{equation}
or
\begin{equation}
    \text{E}\left[Y_d \left(\tilde{D}\sigma_Z - \tilde{Z}\sigma_D\right)\right] \leq 0,
    \qquad \text{E}\left[Y_d\tilde{D}\right] \leq 0
    \qquad \textrm{and} \qquad \text{E}\left[Y_d\tilde{Z}\right] \leq 0.
    \label{eq.LEI2}
\end{equation}
Applying a similar reasoning as in the previous subsection to (\ref{eq.LEI1}), for any $(\alpha, \beta, \gamma) \in [0, 1]^3$ such that $1-\beta-\gamma \geq 0$, we have
\begin{eqnarray*}
    &\text{E}&\left[Y_d \alpha\left(\gamma \big(\tilde{D}\sigma_Z - \tilde{Z}\sigma_D\big) + \beta \tilde{D} + (1-\beta-\gamma) \tilde{Z} \right)\right] \geq 0\\
    \text{  and  } \
    &\text{E}&\left[-Y_d \alpha\left(\gamma \big(\tilde{D}\sigma_Z - \tilde{Z}\sigma_D\big) + \beta \tilde{D} + (1-\beta-\gamma) \tilde{Z} \right)\right] \leq 0.
\end{eqnarray*}
Since $\beta$ and $\gamma$ belong to a 2-simplex, we can parametrize $\gamma=(1-\beta)\mu$, where $\mu \in [0,1]$. Therefore, the above inequalities are equivalent to
\begin{eqnarray*}
    \text{E}\left[\delta^+_{L} Y_d \right] \geq \text{E}[Y_d ] \equiv \theta_d \ \text{  and  } \
    \text{E}\left[\delta^-_{L} Y_d\right] \leq \text{E}[Y_d] \equiv \theta_d,
\end{eqnarray*}
where $\delta^+_{L} \equiv 1+ \alpha\big((1-\beta)\mu (\tilde{D}\sigma_Z - \tilde{Z}\sigma_D) + \beta \tilde{D} + (1-\beta)(1-\mu) \tilde{Z} \big)$ and $\delta^-_{L} \equiv 1- \alpha\big((1-\beta)\mu  (\tilde{D}\sigma_Z - \tilde{Z}\sigma_D) + \beta \tilde{D} + (1-\beta)(1-\mu) \tilde{Z} \big)$.
Hence, we obtain the following bounds for $\theta_d$ from (\ref{eq.LEI1}):
\begin{equation*}
  I_{LEI1}^d \equiv \left[ \sup_{(\alpha, \beta, \mu) \in [0,1]^3} \text{E} \Big[ \underline{f}_d\left( Y, D, \delta^-_{L} \right) \Big],
    \inf_{(\alpha, \beta, \mu) \in [0,1]^3} \text{E} \Big[ \overline{f}_d\left( Y, D, \delta^+_{L} \right) \Big] \right].
\end{equation*}
Likewise, we derive bounds for $\theta_d$ from (\ref{eq.LEI2}) as
\begin{equation*}
  I_{LEI2}^d \equiv \left[ \sup_{(\alpha, \beta, \mu) \in [0,1]^3} \text{E} \Big[ \underline{f}_d\left( Y, D, \delta^+_{L} \right) \Big],
    \inf_{(\alpha, \beta, \mu) \in [0,1]^3} \text{E} \Big[ \overline{f}_d\left( Y, D, \delta^-_{L} \right) \Big] \right],
\end{equation*}
and the following proposition holds.
\begin{proposition}\label{thm2}
Under Assumptions BoS, SDC and LEI, nonparametric bounds for the parameter $\theta_d$ are given by:
\begin{equation*}
I_{LEI}^d \equiv I_{LEI1}^d \cup I_{LEI2}^d.
\end{equation*}
\end{proposition}

The following corollary shows that $I_{LEI}^d$ is weakly tighter than $I_{SDC}^d$.
\begin{corollary}\label{cor1}
Under Assumptions BoS, SDC and LEI, we have
$$I_{LEI}^d \subseteq I_{SDC}^d.$$
\end{corollary}
The proof of Corollary \ref{cor1} follows from the fact that the bounds $I_{SDC}^d$ are a special case of the bounds $I_{LEI}^d$ where $\mu=0$.

\subsection{Adding the monotone treatment response assumption} In this subsection, we derive bounds on the potential outcome mean $\theta_d$ under Assumptions BoS, SDC, LEI, and MTR .
Under Assumption MTR, we have
\begin{eqnarray*}
    Y_d \mathbbm{1}\left\{D>d\right\} &=& Y_d \sum_{j=d+1}^{T} \mathbbm{1}\left\{D=j\right\}
    \, = \, \sum_{j=d+1}^{T} Y_d  \mathbbm{1}\left\{D=j\right\}
    \\
    &\leq& \sum_{j=d+1}^{T} Y  \mathbbm{1}\left\{D=j\right\}\, = \, Y \sum_{j=d+1}^{T}  \mathbbm{1}\left\{D=j\right\} \, = \, Y  \mathbbm{1}\left\{D>d\right\},
\end{eqnarray*}
where the inequality holds as
    $Y_d  \mathbbm{1}\left\{D=j\right\} \leq  Y_j  \mathbbm{1}\left\{D=j\right\}=Y \mathbbm{1}\left\{D=j\right\}$ for all $j>d$ for each $d$.
This result shows that under the MTR assumption, the counterfactual random variable $Y_d \mathbbm{1}\left\{D>d\right\}$ is bounded from above by the observed variable $Y  \mathbbm{1}\left\{D>d\right\}$. As we can see, this assumption considerably shrinks the upper bound on $Y_d \mathbbm{1}\left\{D>d\right\}$, which would be $\overline{y}_d \mathbbm{1}\left\{D>d\right\}$ otherwise under Assumption BoS.
Similarly, we also have
\begin{equation*}
    Y_d \mathbbm{1}\left\{D<d\right\} \quad \geq \quad Y  \mathbbm{1}\left\{D<d\right\}
\end{equation*}
for each $d$.
Without the MTR assumption, the lower bound on $Y_d \mathbbm{1}\left\{D<d\right\}$ would be $\underline{y}_d \mathbbm{1}\left\{D<d\right\}$ under Assumption BoS.
Combining these results together, the following inequalities hold under Assumptions BoS and MTR:
\begin{eqnarray*}
\begin{array}{lcl}
    \underline{y}_d  \mathbbm{1}\left\{D>d\right\} \quad \leq \quad &Y_d \mathbbm{1}\left\{D>d\right\}& \quad \leq \quad Y \mathbbm{1}\left\{D>d\right\},
    \\
    Y  \mathbbm{1}\left\{D<d\right\} \quad \leq \quad &Y_d \mathbbm{1}\left\{D<d\right\}& \quad \leq \quad \overline{y}_d \mathbbm{1}\left\{D<d\right\}.
 \end{array}
\end{eqnarray*}
Thus, for any $\delta \in \mathbb{R}$, we have the following bounds
\begin{eqnarray}
\begin{array}{lcl}
    \min\{\delta \underline{y}_d, \delta Y \} \mathbbm{1}\left\{D>d\right\}
    \quad \leq  &\delta Y_d  \mathbbm{1}\left\{D>d\right\}& \leq \quad
    \max\{\delta \underline{y}_d, \delta Y \} \mathbbm{1}\left\{D>d\right\},
    \\
    \min\{\delta\overline{y}_d, \delta Y \} \mathbbm{1}\left\{D<d\right\}
    \quad \leq &\delta Y_d \mathbbm{1}\left\{D<d\right\}& \leq \quad
    \max\{\delta\overline{y}_d, \delta Y \} \mathbbm{1}\left\{D<d\right\}.
    \end{array}
    \label{eq.mon}
\end{eqnarray}
Moreover, as we have
\begin{eqnarray*}
    Y_d \delta \mathbbm{1}\left\{D>d\right\}+Y_d \delta \mathbbm{1}\left\{D<d\right\}=Y_d \delta \mathbbm{1}\left\{D\neq d\right\},
\end{eqnarray*}
inequalities (\ref{eq.mon}) imply
\begin{eqnarray*}
    \delta Y_d\mathbbm{1}\left\{D\neq d\right\} &\leq& \max\left\{\delta\underline{y}_d, \delta Y\right\} \mathbbm{1}\left\{D>d\right\} + \max\left\{\delta\overline{y}_d, \delta Y\right\} \mathbbm{1}\left\{D<d\right\},
    \\
    \delta Y_d\mathbbm{1}\left\{D\neq d\right\} &\leq& \min\left\{\delta\underline{y}_d, \delta Y\right\} \mathbbm{1}\left\{D>d\right\} + \min\left\{\delta\overline{y}_d, \delta Y\right\} \mathbbm{1}\left\{D<d\right\},
\end{eqnarray*}
for any $\delta \in \mathbb{R}$.

Now, recall that inequalities (\ref{eq.LEI1}) implied by Assumptions SDC and LEI yield
\begin{eqnarray*}
    \text{E}\left[\delta^+_{L} Y \mathbbm{1}\left\{D=d\right\} + \delta^+_{L} Y_d \mathbbm{1}\left\{D \neq d\right\} \right] &\geq& \theta_d,
    \\
    \text{E}\left[\delta^-_{L} Y \mathbbm{1}\left\{D=d\right\} + \delta^-_{L} Y_d \mathbbm{1}\left\{D \neq d\right\} \right] &\leq&  \theta_d,
\end{eqnarray*}
for any $(\alpha, \beta, \mu) \in [0,1]^3$, and thus we have
\begin{align*}
    \text{E} \Big[ \underline{m}_d\left( Y, D, \delta^-_{L} \right) \Big] \leq \theta_d \leq \text{E} \Big[ \overline{m}_d\left( Y, D, \delta^+_{L} \right) \Big]
\end{align*}
for any $(\alpha, \beta,\mu) \in [0,1]^3$, where we define the functions $\underline{m}_d$ and $\overline{m}_d$ as
\begin{eqnarray*}
    \underline{m}_d\left(Y, D, \delta \right) &\equiv&
    \mathbbm{1}\left\{D=d\right\}\delta Y + \mathbbm{1}\left\{D>d\right\} \min\{\delta\underline{y}_d, \delta Y \}  + \mathbbm{1}\left\{D<d\right\} \min\{\delta\overline{y}_d, \delta Y \},
    \\
    \overline{m}_d\left(Y, D, \delta \right) &\equiv&
    \mathbbm{1}\left\{D=d\right\}\delta Y + \mathbbm{1}\left\{D>d\right\}\max\{\delta\underline{y}_d, \delta Y \}  + \mathbbm{1}\left\{D<d\right\}\max\{\delta\overline{y}_d, \delta Y \} .
\end{eqnarray*}
Likewise, inequalities (\ref{eq.LEI2}) under the MTR assumption yield the following implications for $\theta_d$:
\begin{align*}
    \text{E} \Big[ \underline{m}_d\left( Y, D, \delta^+_{L} \right) \Big] \leq \theta_d \leq \text{E} \Big[ \overline{m}_d\left( Y, D, \delta^-_{L} \right) \Big]
\end{align*}
for any $(\alpha, \beta,\mu) \in [0,1]^3$.
Therefore, we conclude that Assumptions BoS, SDC, LEI, and MTR together imply
\begin{equation*}
    \theta_d \in  I_{MTR1}^d \cup I_{MTR2}^d \equiv I_{MTR}^d,
\end{equation*}
where
\begin{align*}
    I_{MTR1}^d &\equiv \left[ \sup_{(\alpha, \beta, \mu) \in [0,1]^3} \text{E} \Big[ \underline{m}_d\left( Y, D, \delta^-_{L} \right) \Big], \inf_{(\alpha, \beta, \mu) \in [0,1]^3} \text{E} \Big[ \overline{m}_d\left( Y, D, \delta^+_{L} \right) \Big] \right],\\
    I_{MTR2}^d &\equiv \left[ \sup_{(\alpha, \beta, \mu) \in [0,1]^3} \text{E} \Big[ \underline{m}_d\left( Y, D, \delta^+_{L} \right) \Big], \inf_{(\alpha, \beta, \mu) \in [0,1]^3} \text{E} \Big[ \overline{m}_d\left( Y, D, \delta^-_{L} \right) \Big] \right].
\end{align*}

Note that if we are not willing to impose Assumption LEI, the bounds can be obtained from $I_{MTR1}^d \cup I_{MTR2}^d$ using $\delta^+_{S}$ and $\delta^-_{S}$ instead of $\delta^+_{L}$ and $\delta^-_{L}$.

An important implication of the MTR assumption is that it weakly signs the ATE, i.e., $ATE(d,d')\geq 0$ for all $d > d'$. More precisely, bounds for $ATE(d,d')$, where $d>d'$, can be characterized as follows:
\begin{eqnarray*}
I_{MTR}^{ATE(d,d')}\equiv \left\{\theta^d-\theta^{d'}:\ \theta^d-\theta^{d'} \geq 0,\ \theta^d \in I_{MTR}^d \text{ and }  \theta^{d'} \in I_{MTR}^{d'}\right\}.
\end{eqnarray*}


\section{Inference}\label{inference}
\setcounter{equation}{0}
We want to construct confidence bounds for the set $I_{SDC}^d = I_{SDC1}^d  \cup  I_{SDC2}^d $.
This is a special case of an intersection-union test as described in \cite{Berger1982}. We are going to construct confidence regions for the sets $I_{SDC1}^d$ and $I_{SDC2}^d$ using the intersection bounds framework of \cite{CLR2013} or \cite{AS2013}, and then take the union of the two confidence regions. \cite{Berger1996} showed that the union of the confidence regions has at least the same coverage rate as each confidence region. Identified sets with a similar structure have been considered in \cite{Chesher2020} and \cite{Machado2019}. As we explain in Section \ref{ident}, the sets $I_{SDC1}^d$ and $I_{SDC2}^d$ could each be empty, and may not overlap. However, in our empirical illustration below, they are nonempty and overlap in all cases.

We now explain how to rewrite the intersection bounds $ I_{SDC1}^d$ in such a way that it can be easily implemented using the \cite{CKLRstata} or \cite{AKS2017} Stata packages. Suppose that we draw two independent random variables $U_1$ and $U_2$ from the uniform distribution over $[0,1]$, independently of the data $(Y,D,Z)$. Then, we have
\begin{eqnarray*}
\text{E} \Big[ \underline{f}_d\big( Y, D, \delta^-_{S} \big) \Big\vert  U_1=\alpha, U_2=\beta \Big] = \text{E} \Big[ \underline{f}_d\big( Y, D, \delta^-_{S} \big) \Big],
\end{eqnarray*}
since $U_1$ and $U_2$ are independent of $(Y,D,Z)$. Therefore,
\begin{eqnarray*}
    I_{SDC1}^d &=& \Bigg[ \sup_{(\alpha, \beta) \in [0,1]^2} \text{E} \Big[ \underline{f}_d\big( Y, D, \delta^-_{S} \big) \Big\vert U_1=\alpha, U_2 = \beta \Big],\\
    && \qquad \qquad \qquad \qquad \qquad \inf_{(\alpha, \beta) \in [0,1]^2} \text{E} \Big[ \overline{f}_d\big( Y, D,\delta^+_{S} \big) \Big\vert U_1=\alpha, U_2 = \beta \Big] \Bigg].
\end{eqnarray*}

Hence, these bounds take the form of conditional moment inequalities, which can be implemented using existing inferential methods like \cite{CKLRstata} or \cite{AKS2017}. Similar results apply to the bounds $I_{LEI}^d$ and $I_{MTR}^d$, where there are three conditioning variables $U_1$, $U_2$, and $U_3$ instead of two. A similar technique has been proposed in \cite{Kedagni2018b} where they construct confidence sets for the potential outcome means under the IV zero-covariance assumption.

Using the \cite{CKLRstata} Stata package, we obtain confidence sets that asymptotically cover either the true parameter $\theta_d$ or the bounds for $\theta_d$ with pre-specified probability, through the \texttt{clr3bound} command.\footnote{The \texttt{clr2bound} command could be used to obtain confidence sets for the bounds for $\theta_d$ only.}

\section{Empirical illustration}\label{empirical}
\setcounter{equation}{0}
In this application, we use a data set drawn from the NLSYM. This data includes 3,010 young men who were ages 24-34 in 1976. It is the same data used in \cite{Card1995}. In our analysis, the outcome variable is log hourly wage in cents $(lwage)$, and the treatment variable is education $(educ)$ grouped in 4 categories: less than high school $(educ< 12\ years)$, high school $(12 \leq educ < 16)$, college degree $(16 \leq educ < 18)$, and graduate $(educ \geq 18)$.\footnote{As discussed in \cite{Andresen2021}, this discretization may induce some identification issues, and the results may be sensitive to it. For this reason, our empirical results should be seen as illustrative.}

Our imperfect IV is parental education. Since the work of \cite{Willis79}, parental education has been used as an IV. However, an individual's ability can be dependent on her parents' ability, which is correlated with parental education. For this reason, parents' education will not be a valid instrument. This fact is documented in \cite{Kedagni20}, who provided evidence that even after controlling for a measure of ability, parental education is not a good instrument. This result is contrary to the \cite{Lemke03} idea that controlling for some measure of child ability could make parental education a valid~IV. Nonetheless, it is reasonable to assume that parental education has the same sign of correlation with the individual's potential wage as the correlation between the person's potential wage and her own education.\footnote{The MIV and possibly MTS assumptions also seem reasonable in this empirical example. However, our goal in this section is to show how our derived bounds can be implemented in a real-world application.} It is also likely that parental education be less endogenous than is the person's own education. Finally, as in \cite{Manski2000}, we use the monotone treatment response assumption to tighten the bounds on the average returns to education.

In theory, the outcome variable $lwage$ is unbounded. For practical reasons, we follow \cite{Ginther2000} to trim the log wage. The outcome variable that we use is defined as $Y=\tau$-quantile of $lwage$ if $lwage$ is less than or equal to its $\tau$-quantile, $Y=(1-\tau)$-quantile of $lwage$ if $lwage$ is greater than or equal to its $(1-\tau)$-quantile, and $Y=lwage$ otherwise. In our empirical illustration, we set $\tau=0.05$. We construct 95\% two-sided confidence sets on potential average log wages and their bounds using the \texttt{clr3bound} command of \cite{CKLRstata} in the Stata software. We estimate the conditional expectations using the parametric method, which is the default option for this Stata command. See the appendix for more details on the implementation.

We present the results with mother's education as an IIV. The results for father's education are in the appendix. Table \ref{Table:cs1} displays the 95\% confidence sets for the potential wage means and average returns to schooling under the SDC assumption, while Table \ref{Table:cs2} shows the confidence sets under both the SDC and LEI assumptions. Each table shows the confidence set for bounds on $\theta_d$ in the first two columns, the confidence set for the parameter $\theta_d$ in the third and forth columns, and the point estimates of the bounds in the fifth and sixth columns. The SDC+LEI bounds in Table \ref{Table:cs2} are generally narrower than the SDC bounds in Table \ref{Table:cs1}. This suggests that Assumption LEI provides some extra identifying power to the SDC assumption. However, the bounds seem wide and less informative in both cases. For example, the confidence set for $ATE(2,1) \equiv \theta_2-\theta_1$ under SDC+LEI is $[-0.98,1.02]$, implying that the average return to college degree compared to high school education varies between $-98\%$ and $102\%$, which is not very informative for an individual making a college decision. The corresponding point estimates of the bounds seem tighter $[-0.53, 0.69]$.

\begin{table}[ht]
\caption{95\% Confidence regions for bounds and parameters under SDC}
\begin{center}
\begin{tabular}{l|cccccc} \label{Table:cs1}
           &   \multicolumn{2}{c}{For Bounds} & \multicolumn{2}{c}{For Parameters} & \multicolumn{2}{c}{Point Estimates} \\
Parameters &    Cf. LB & Cf. UB  &  Cf. LB & Cf. UB & $\,\,\,$ LB $\,\,\,$ & UB\\ \hline \hline
$\theta_0$ ($<$ high) &   	 5.51 &      6.91 &      5.53 &      6.89 &      5.65 &      6.70 \\
$\theta_1$ (high) &   		 5.86 &      6.64 &      5.88 &      6.62 &      5.97 &      6.45 \\
$\theta_2$ (college) &   	 5.62 &      6.94 &      5.64 &      6.92 &      5.77 &      6.76 \\
$\theta_3$ (graduate) &   	 5.51 &      7.01 &      5.53 &      6.99 &      5.64 &      6.85 \\
$\theta_0-\theta_1$   & 	 -1.13 &      1.05 &     -1.09 &      1.01 &     -0.81 &      0.73 \\
$\theta_2-\theta_1$ &   	 -1.02 &      1.08 &     -0.98 &      1.04 &     -0.68 &      0.79 \\
$\theta_3-\theta_1$ &  		 -1.13 &      1.15 &     -1.09 &      1.11 &     -0.81 &      0.88 \\
\hline\hline
\end{tabular}
\end{center}
\footnotesize \textbf{Note:} Cf. LB (UB): lower (upper) bound of the 95\% confidence region; LB (UB): point estimate for the lower (upper) bound.
\end{table}

\begin{table}[ht]
\caption{95\% Confidence regions for bounds and parameters under SDC and LEI}
\begin{center}
\begin{tabular}{l|cccccc} \label{Table:cs2}
           &   \multicolumn{2}{c}{For Bounds} & \multicolumn{2}{c}{For Parameters} & \multicolumn{2}{c}{Point Estimates} \\
Parameters &    Cf. LB & Cf. UB  &  Cf. LB & Cf. UB & $\,\,\,$ LB $\,\,\,$ & UB\\ \hline \hline
$\theta_0$ ($<$ high) &   	 5.52 &      6.91 &      5.54 &      6.89 &      5.70 &      6.64 \\
$\theta_1$ (high) &   		 5.86 &      6.64 &      5.88 &      6.62 &      6.01 &      6.36 \\
$\theta_2$ (college) &   	 5.62 &      6.92 &      5.64 &      6.90 &      5.83 &      6.70 \\
$\theta_3$ (graduate) &   	 5.51 &      7.00 &      5.53 &      6.98 &      5.69 &      6.78 \\
$\theta_0-\theta_1$   & 	 -1.12 &      1.05 &     -1.08 &      1.01 &     -0.66 &      0.63 \\
$\theta_2-\theta_1$ &   	 -1.02 &      1.06 &     -0.98 &      1.02 &     -0.53 &      0.69 \\
$\theta_3-\theta_1$ &  		 -1.13 &      1.14 &     -1.09 &      1.10 &     -0.67 &      0.77 \\
\hline\hline
\end{tabular}
\end{center}
\footnotesize \textbf{Note:} Cf. LB (UB): lower (upper) bound of the 95\% confidence region; LB (UB): point estimate for the lower (upper) bound.
\end{table}
Furthermore, the confidence regions for the bounds and the parameters considerably shrink and become more informative when we add the MTR assumption (see Table \ref{Table:cs3}). We assume that the lower bound of $\theta_d$ is equal to the upper bound of $\theta_{d-1}$ whenever the former is less than the latter, because $Y_{d-1}$ cannot exceed $Y_d$ for each $d=1, 2, 3$ under the MTR assumption. Individuals with less than high school education could earn up to 94\% less than high school graduates ($ATE(0,1)$). Moreover, the confidence regions for $ATE(2, 1)$ and $ATE(3,1)$ suggest that college graduates could earn up to 55\% more than high school graduates, while individuals with a graduate degree earn between 41\% and 65\% higher wages than high school graduates (which approximately represents an annual return between 6.8\% and 10.8\%). Table II in \cite{Card2001} shows that the point estimates of the annual return to schooling in the US vary roughly between 5\% and 13\%. As we can see, our set estimates for the annual return are consistent with the existing range in the literature.

If we impose the linear structure of \cite{Nevo2012} in this empirical exercise, we obtain the following point estimate bounds for $\theta$ under SDC and MTR: $[0, 0.16] \cup [0.24, +\infty)$. Indeed, the estimated correlation between $D$ and $Z$ is $0.3824>0$, and under SDC, the point estimate bounds for $\theta$ are $(-\infty, 0.16] \cup [0.24, +\infty)$. If one is willing to further assume that $Cov(D,U) \geq 0$, then the bounds for $\theta$ reduce to $[0, 0.16]$, which imply an annual return between 0 and 16\%. This range is consistent with the literature, but remains wide.

\begin{table}[ht]
\caption{95\% Confidence regions for bounds and parameters under SDC, LEI, and MTR}
\begin{center}
\begin{tabular}{l|cccccc} \label{Table:cs3}
           &   \multicolumn{2}{c}{For Bounds} & \multicolumn{2}{c}{For Parameters} & \multicolumn{2}{c}{Point Estimates} \\
Parameters &    Cf. LB & Cf. UB  &  Cf. LB & Cf. UB & $\,\,\,$ LB $\,\,\,$ & UB\\ \hline \hline
$\theta_0$ ($<$ high) &   	 5.52 &      6.35 &      5.54 &      6.33 &      5.70 &      6.08 \\
$\theta_1$ (high) &   		 6.35 &      6.50 &      6.33 &      6.48 &      6.08 &      6.26 \\
$\theta_2$ (college) &   	 6.50 &      6.91 &      6.48 &      6.89 &      6.32 &      6.66 \\
$\theta_3$ (graduate) &   	 6.91 &      7.00 &      6.89 &      6.98 &      6.66 &      6.78 \\
$\theta_0-\theta_1$   & 	 -0.98 &      0.00 &     -0.94 &      0.00 &     -0.56 &      0.00 \\
$\theta_2-\theta_1$ &   	  0.00 &      0.55 &      0.00 &      0.55 &      0.06 &      0.58 \\
$\theta_3-\theta_1$ &  		  0.41 &      0.65 &      0.41 &      0.65 &      0.40 &      0.70 \\
\hline\hline
\end{tabular}
\end{center}
\footnotesize \textbf{Note:} Cf. LB (UB): lower (upper) bound of the 95\% confidence region; LB (UB): point estimate for the lower (upper) bound.
\end{table}

\section{Conclusion}\label{conclusion}
In this paper, we derive nonparametric bounds on the average treatment effect when an imperfect instrument is available. We extend \citeauthor{Nevo2012}'s (\citeyear{Nevo2012}) identification results to nonparametric models.
We first assume that the sign of correlation between the imperfect instrument and the unobserved latent variables is the same as the correlation between the endogenous variable and the latent variables.
We show that the MTS-MIV restrictions introduced by \cite{Manski2000, Manski2009}, jointly imply this assumption.
Second, we show how the assumption that the imperfect instrument is less endogenous than the treatment variable can help tighten the bounds.
We also use the monotone treatment response assumption to get tighter bounds. The identified set takes the form of intersection bounds, which can be implemented using \citeauthor{CLR2013}'s (\citeyear{CLR2013}) inferential method. Finally, we illustrate our methodology using the National Longitudinal Survey of Young Men data to estimate returns to schooling.

\section*{Acknowledgements}
The authors are grateful to the Editor Petra Todd, and three anonymous referees for valuable suggestions and comments. They also thank Santiago Acerenza, Otavio Bartalotti, Helle Bunzel, Ismael Mourifi\'e, Vitor Possebom, and participants at the Iowa State econometrics workshop for helpful comments. All errors are ours.

\bibliographystyle{chicago}
\begin{thebibliography}{}

\bibitem[\protect\citeauthoryear{Andresen and Huber}{Andresen and
  Huber}{2021}]{Andresen2021}
Andresen, M.~E. and M.~Huber (2021).
\newblock Instrument-based estimation with binarized treatments: Issues and
  tests for the exclusion restriction.
\newblock {\em The Econometrics Journal (forthcoming)\/}.

\bibitem[\protect\citeauthoryear{Andrews, Kim, and Shi}{Andrews
  et~al.}{2017}]{AKS2017}
Andrews, D. W.~K., W.~Kim, and X.~Shi (2017).
\newblock Stata commands for testing conditional moment
  inequalities/equalities.
\newblock {\em Stata Journal\/}~{\em 17\/}(1), 56--72.

\bibitem[\protect\citeauthoryear{Andrews and Shi}{Andrews and
  Shi}{2013}]{AS2013}
Andrews, D. W.~K. and X.~Shi (2013).
\newblock Inference based on conditional moment inequalities.
\newblock {\em Econometrica\/}~{\em 81}, 609--666.

\bibitem[\protect\citeauthoryear{Berger}{Berger}{1982}]{Berger1982}
Berger, R.~L. (1982).
\newblock Multiparameter hypothesis testing and acceptance sampling.
\newblock {\em Technometrics\/}~{\em 24}, 295--300.

\bibitem[\protect\citeauthoryear{Berger and Hsu}{Berger and
  Hsu}{1996}]{Berger1996}
Berger, R.~L. and J.~C. Hsu (1996).
\newblock Bioequivalence trials, intersection-union tests and equivalence
  confidence sets.
\newblock {\em Statistical Science\/}~{\em 11\/}(4), 283--319.

\bibitem[\protect\citeauthoryear{Bhattacharya, Shaikh, and
  Vytlacil}{Bhattacharya et~al.}{2012}]{BSV2012}
Bhattacharya, J., A.~Shaikh, and E.~Vytlacil (2012).
\newblock Treatment effect bounds: An application to swan-ganz catheterization.
\newblock {\em Journal of Econometrics\/}~{\em 168\/}(2), 223--243.

\bibitem[\protect\citeauthoryear{Card}{Card}{1995}]{Card1995}
Card, D. (1995).
\newblock Using geographic variation in college proximity to estimate the
  return to schooling.
\newblock In L.~N. Christofides, E.~K. Grant, and R.~Swidinsky (Eds.), {\em
  Aspects of Labour Market Behaviour: Essays in Honour of John Vanderkamp},
  pp.\  201--222. Toronto, Canada: University of Toronto Press.

\bibitem[\protect\citeauthoryear{Card}{Card}{2001}]{Card2001}
Card, D. (2001).
\newblock Estimating the return to schooling: Progress on some persistent
  econometric problems.
\newblock {\em Econometrica\/}~{\em 69}, 1127--1160.

\bibitem[\protect\citeauthoryear{Chernozhukov, Kim, Lee, and
  Rosen}{Chernozhukov et~al.}{2015}]{CKLRstata}
Chernozhukov, V., W.~Kim, S.~Lee, and A.~M. Rosen (2015).
\newblock Implementing intersection bounds in stata.
\newblock {\em Stata Journal\/}~{\em 15\/}(1), 21--44.

\bibitem[\protect\citeauthoryear{Chernozhukov, Lee, and Rosen}{Chernozhukov
  et~al.}{2013}]{CLR2013}
Chernozhukov, V., S.~Lee, and A.~M. Rosen (2013).
\newblock Intersection bounds: Estimation and inference.
\newblock {\em Econometrica\/}~{\em 81\/}(2), 667--737.

\bibitem[\protect\citeauthoryear{Chesher and Rosen}{Chesher and
  Rosen}{2020}]{Chesher2020}
Chesher, A. and A.~Rosen (2020).
\newblock Generalized instrumental variable models, methods, and applications.
\newblock In J.~J.~H. S.~N.~Durlauf, L. P.~Hansen and R.~L. Matzkin (Eds.),
  {\em Handbook of Econometrics}, Volume~7A, pp.\  1--110. Amsterdam:
  North-Holland: Elsevier.

\bibitem[\protect\citeauthoryear{Conley, Hansen, and Rossi}{Conley
  et~al.}{2012}]{Conley2012}
Conley, T.~G., C.~B. Hansen, and P.~E. Rossi (2012).
\newblock Plausibly exogenous.
\newblock {\em Review of Economics and Statistics\/}~{\em 94\/}(1), 260--272.

\bibitem[\protect\citeauthoryear{Ginther}{Ginther}{2000}]{Ginther2000}
Ginther, D.~K. (2000).
\newblock Alternative estimates of the effect of schooling on earnings.
\newblock {\em Review of Economics and Statistics\/}~{\em 82\/}(1), 103--116.

\bibitem[\protect\citeauthoryear{Hotz, Mullin, and Sanders}{Hotz
  et~al.}{1997}]{Hotz1997}
Hotz, V.~J., C.~Mullin, and S.~Sanders (1997).
\newblock Bounding causal effects using data from a contaminated natural
  experiment: Analyzing the effects of teenage childbearing.
\newblock {\em Review of Economic Studies\/}~{\em 64\/}(4), 575--603.

\bibitem[\protect\citeauthoryear{K\'edagni, Li, and Mourifi\'e}{K\'edagni
  et~al.}{2018}]{Kedagni2018b}
K\'edagni, D., L.~Li, and I.~Mourifi\'e (2018).
\newblock Bounding average returns to schooling using unconditional moment
  restrictions.
\newblock Working Paper 18022, Department of Economics, Iowa State University.

\bibitem[\protect\citeauthoryear{K\'edagni and Mourifi\'e}{K\'edagni and
  Mourifi\'e}{2020}]{Kedagni20}
K\'edagni, D. and I.~Mourifi\'e (2020).
\newblock Generalized instrumental inequalities: Testing the instrumental
  variable independence assumption.
\newblock {\em Biometrika\/}~{\em 107\/}(3), 661--675.

\bibitem[\protect\citeauthoryear{Lemke and Rischall}{Lemke and
  Rischall}{2003}]{Lemke03}
Lemke, R.~J. and I.~C. Rischall (2003).
\newblock Skill, parental income, and iv estimation of the returns to
  schooling.
\newblock {\em Applied Economics Letters\/}~{\em 10\/}(5), 281–--286.

\bibitem[\protect\citeauthoryear{Machado, Shaikh, and Vytlacil}{Machado
  et~al.}{2019}]{Machado2019}
Machado, C., A.~Shaikh, and E.~Vytlacil (2019).
\newblock Instrumental variables and the sign of the average treatment effect.
\newblock {\em Journal of Econometrics\/}~{\em 212}, 522--555.

\bibitem[\protect\citeauthoryear{Manski}{Manski}{1990}]{Manski1990}
Manski, C.~F. (1990).
\newblock Nonparametric bounds on treatment effects.
\newblock {\em American Economic Reviews, Papers and Proceedings of the Hundred
  and Second Annual Meeting of the American Economic Association\/}~{\em
  80\/}(2), 319--323.

\bibitem[\protect\citeauthoryear{Manski}{Manski}{1994}]{Manski1994}
Manski, C.~F. (1994).
\newblock The selection problem.
\newblock {\em in Advances Economics, Sixth World Congress, C. Sims (ed.),
  Cambridge University Press\/}~{\em 1}, 143--170.

\bibitem[\protect\citeauthoryear{Manski}{Manski}{1997}]{Manski1997}
Manski, C.~F. (1997).
\newblock Monotone treatment response.
\newblock {\em Econometrica\/}~{\em 65\/}(6), 1311--1334.

\bibitem[\protect\citeauthoryear{Manski and Pepper}{Manski and
  Pepper}{2000}]{Manski2000}
Manski, C.~F. and J.~Pepper (2000).
\newblock Monotone instrumental variables: With an application to the returns
  to schooling.
\newblock {\em Econometrica\/}~{\em 68}, 997--1010.

\bibitem[\protect\citeauthoryear{Manski and Pepper}{Manski and
  Pepper}{2009}]{Manski2009}
Manski, C.~F. and J.~Pepper (2009).
\newblock More on monotone instrumental variables.
\newblock {\em Econometrics Journal\/}~{\em 12}, S200--S216.

\bibitem[\protect\citeauthoryear{Masten and Poirier}{Masten and
  Poirier}{2020}]{Masten2020}
Masten, M.~A. and A.~Poirier (2020).
\newblock Salvaging falsified instrumental variable models.
\newblock {\em Econometrica (forthcoming)\/}.

\bibitem[\protect\citeauthoryear{Mourifi\'e, Henry, and M\'eango}{Mourifi\'e
  et~al.}{2020}]{Mourifie2020}
Mourifi\'e, I., M.~Henry, and R.~M\'eango (2020).
\newblock Sharp bounds and testability of a roy model of stem major choices.
\newblock {\em Journal of Political Economy\/}~{\em 8\/}(128), 3220--3283.

\bibitem[\protect\citeauthoryear{Nevo and Rosen}{Nevo and
  Rosen}{2012}]{Nevo2012}
Nevo, A. and A.~Rosen (2012).
\newblock Identification with imperfect instruments.
\newblock {\em The Review of Economics and Statistics\/}~{\em 94\/}(3),
  659--671.

\bibitem[\protect\citeauthoryear{Willis and Rosen}{Willis and
  Rosen}{1979}]{Willis79}
Willis, R. and S.~Rosen (1979).
\newblock Education and self-selection.
\newblock {\em Journal of Political Economy\/}~{\em 87\/}(5), Pt2:S7--36.

\end{thebibliography}