The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
46,485 characters
Phase transition of the monotonicity assumption in learning local average treatment effects
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\title{Phase transition of the monotonicity assumption in learning local
average treatment effects\thanks{This version: March 24, 2021}}
\author{Yinchu Zhu\thanks{Email: [email removed]. We thank Tymon S{\l}oczy{\'n}ski,
Kaspar W{\"u}thrich and the participants of MIT Econometrics Lunch
for very helpful discussions and comments. We thank Krisztian Gado
for excellent research assistance. All the errors are mine. }\\
\\
Department of Economics\\
Brandeis University}
\maketitle
\begin{abstract}
We consider the setting in which a strong binary instrument is available
for a binary treatment. The traditional LATE approach assumes the
monotonicity condition stating that there are no defiers (or compliers).
Since this condition is not always obvious, we investigate the sensitivity
and testability of this condition. In particular, we focus on the
question: does a slight violation of monotonicity lead to a small
problem or a big problem?
We find a phase transition for the monotonicity condition. On one
of the boundary of the phase transition, it is easy to learn the sign
of LATE and on the other side of the boundary, it is impossible to
learn the sign of LATE. Unfortunately, the impossible side of the
phase transition includes data-generating processes under which the
proportion of defiers tends to zero. This boundary of phase transition
is explicitly characterized in the case of binary outcomes. Outside
a special case, it is impossible to test whether the data-generating
process is on the nice side of the boundary. However, in the special
case that the non-compliance is almost one-sided, such a test is possible.
We also provide simple alternatives to monotonicity.
\begin{comment}
We derive a lower bound for the magnitude of LATE without assuming
monotonicity. This lower bound is larger than zero if and only if
the usual IV regression coefficient is non-zero. However, monotonicity
is crucial for identifying the sign of LATE. Using local asymptotic
analysis, we show that in many situations, we can easily determine
that the magnitude of LATE is non-zero but it is impossible to consistently
estimate the sign of LATE even if the monotonicity condition is slightly
violated (proportion of defiers goes to zero). We also look at the
adaptivity issue; we show that if a confidence set for the sign of
LATE is valid without monotonicity, then this confidence set must
be uninformative even when monotonicity holds. Therefore, it is impossible
to use specification tests to construct a procedure that is widely
valid and is also accurate under monotonicity.
\end{comment}
\begin{comment}
Let $\mu_{1}=E(Y_{i}(1)-Y_{i}(0)|\text{complier})$, $\mu_{2}=E(Y_{i}(1)-Y_{i}(0)|\text{defier})$
and $\beta$ be the IV estimator (LATE). Without assuming monotonicity,
we show that $\max\{|\mu_{1}|,|\mu_{2}|\}\geq|\beta\rho|$, where
$\rho=cov(D_{i},Z_{i})/(ED_{i}(1)+ED_{i}(0))$. In other words, when
the IV is strong ($cov(D_{i},Z_{i})\neq0$), the magnitude of the
treatment effect of at least one group (compliers or defiers) is at
least $|\beta\rho|$. We can easily estimate $\beta\rho$ and thus
obtain a lower bound for the magnitude of the treatment effect on
some subpopulation (either complier or defier). We also show that
monotonicity is a delicate condition; a slight violation can lead
to big uncertainty on the sign of the treatment effects.
\end{comment}
\end{abstract}
\section{Introduction}
Instrumental variables (IV) regressions have been widely used to study
treatment effects in economics and other disciplines. One important
conceptual framework that justifies the causal interpretation of IV
regressions is the local average treatment (LATE). In this paper,
we consider a simple setting with no covariates and discuss the sensitivity
of the key monotonicity assumption in the LATE framework.
We observe iid data $\{(Y_{i},D_{i},Z_{i})\}_{i=1}^{n}$, where $Y_{i}=Y_{i}(1)D_{i}+Y_{i}(0)(1-D_{i})$,
$D_{i}=D_{i}(1)Z_{i}+D_{i}(0)(1-Z_{i})$ and $Z_{i},D_{i}(1),D_{i}(0)\in\{0,1\}$.
The treatment effect is $Y_{i}(1)-Y_{i}(0)$. (Notice that this assumes
that $Z_{i}$ does not directly affect the potential outcomes: $Y_{i}(z,d)=Y_{i}(d)$
for $z,d\in\{0,1\}$.) Throughout the paper, we maintain the assumption
that the IV is strong ($|cov(D_{i},Z_{i})|\geq C$ for a constant
$C>0$) and the following exogeneity condition
\begin{assumption}
\label{assu: IV exogeneity}$Z_{i}$ is independent of $(Y_{i}(1),Y_{i}(0),D_{i}(1),D_{i}(0))$.
\end{assumption}
The typical IV regression exploits the moment condition $EZ_{i}(Y_{i}-D_{i}\beta)=EZ_{i}E(Y_{i}-D_{i}\beta)$
(i.e., $cov(Z_{i},Y_{i}-D_{i}\beta)=0$). This means\footnote{Under Assumption \ref{assu: IV exogeneity}, $\beta$ can be written
in other ways. For example, $\beta=\frac{E(Y_{i}\mid Z_{i}=1)-E(Y_{i}\mid Z_{i}=0)}{E(D_{i}\mid Z_{i}=1)-E(D_{i}\mid Z_{i}=0)}$.} that
\[
\beta=\frac{E(Y_{i}Z_{i})-E(Y_{i})E(Z_{i})}{E(D_{i}Z_{i})-E(D_{i})E(Z_{i})}.
\]
To describe the causal interpretation of $\beta$, we categorize the
population into four types depending on the value of $(D_{i}(1),D_{i}(0))\in\{0,1\}\times\{0,1\}$.
We introduce their definitions and their probability:
\begin{equation}
\begin{cases}
P(D_{i}(1)=D_{i}(0)=1)=a & \text{always\ taker}\\
P(D_{i}(1)=1,D_{i}(0)=0)=b & \text{complier}\\
P(D_{i}(1)=0,D_{i}(0)=1)=c & \text{defier}\\
P(D_{i}(1)=0,D_{i}(0)=0)=1-a-b-c & \text{never taker}.
\end{cases}\label{eq: four types}
\end{equation}
As shown in \citet{angrist1996identification},
\begin{equation}
\beta=\frac{\mu_{1}b-\mu_{2}c}{b-c},\label{eq: IV beta}
\end{equation}
where $\mu_{1}$ and $\mu_{2}$ are the local average treatment effects
(LATE) for compliers and defiers, respectively:
\[
\begin{cases}
\mu_{1}=E(Y_{i}(1)-Y_{i}(0)\mid D_{i}(1)=1,D_{i}(0)=0)\\
\mu_{2}=E(Y_{i}(1)-Y_{i}(0)\mid D_{i}(1)=0,D_{i}(0)=1).
\end{cases}
\]
Since (\ref{eq: IV beta}) is in general not a convex combination
of $\mu_{1}$ and $\mu_{2}$, $\beta$ typically does not have a causal
interpretation without further assumptions. The classical assumption
that makes $\beta$ causally interpretable is the following monotonicity
condition.
\begin{quotation}
\textbf{Monotonicity condition:} either $b=0$ or $c=0$.
\end{quotation}
Clearly, under the monotonicity condition, (\ref{eq: IV beta}) implies
that $\beta=\mu_{1}$ or $\beta=\mu_{2}$, which can be summarized
as $\beta=E(Y_{i}(1)-Y_{i}(0)\mid D_{i}(1)\neq D_{i}(0))$. Hence,
$\beta$ is interpreted as the average treatment effect on the sub-population
for which $D_{i}(1)\neq D_{i}(0)$.
As pointed out by \citet{imbens2014instrumental}, perhaps the strongest
justification of the monotonicity condition is when the instrument
provides an incentive to choose the treatment or when the treatment
is simply not an option without $Z_{i}=1$. Outside these situations,
the validity of the monotonicity condition is not always obvious.
In this paper, we try to answer the following questions
\begin{itemize}
\item Suppose that the data can easily reject $H_{0}:\ \beta=0$. If the
monotonicity is slightly violated ($b$ is far from zero but $c$
is close to zero), would this create a big problem or small problem
for learning LATE?
\item In applications with almost one-sided non-compliance ($P(D_{i}=1\mid Z_{i}=0)\approx0$),
should we worry about the interpretation of $\beta$?
\item What are other options for learning LATE without the monotonicity
condition?
\end{itemize}
\subsection{Background of the problem and summary of main results}
Let us explain why (some of) these questions might be quite subtle
and difficult although they seem to have an obvious answer at the
first glance.
The majority of the paper focuses on the seemingly simple question
of learning the sign of LATE (so we can answer the basic question
of whether the treatment is beneficial or harmful). In particular,
whether we can conclude that $\mu_{1}$ and $\beta$ have the same
sign when monotonicity is slightly violated ($c\approx0$). By rearranging
(\ref{eq: IV beta}), we have
\[
\mu_{1}=\frac{c\mu_{2}+(b-c)\beta}{b}.
\]
Suppose that $\beta<0$. It is easy to see that $\mu_{1}<0$ (and
thus has the same sign as $\beta$) if and only if $\mu_{2}<-(\beta/c)(b-c)$.
Throughout the paper, we assume that $P(|Y_{i}|\leq M)=1$ for a constant
$M>0$. Then the question of learning the sign of $\mu_{1}$ would
seem straight-forward. If $c\rightarrow0$ and $|\beta|$ and $b$
are bounded below by a positive constant, then the threshold $-(\beta/c)(b-c)$
tends to infinity. Since $\mu_{2}$ is bounded (due to the boundedness
of $Y_{i}$), the condition of $\mu_{2}<-(\beta/c)(b-c)$ is asymptotically
satisfied. Hence, the conclusion would be that no matter how slowly
$c$ goes to zero, it is asymptotically valid to conclude that $\mu_{1}$
and $\beta$ have the same sign.
One subtly is whether modeling $|\beta|$ as a quantity bounded below
by a positive constant is an asymptotic framework that is empirically
relevant. In many empirical studies, if we throw away half of the
data and run the IV regression, we often do not find a statistically
significant $\beta$ anymore. Then it might be too strong to assume
that $|\beta|$ is of a much larger order of magnitude compared to
the estimation noise. Moreover, statistically significance of $\beta$
does not mean that $|\beta|$ is bounded below by a positive constant;
statistical significance is asymptotically guaranteed even if $|\beta|\rightarrow0$
and $\sqrt{n}|\beta|\rightarrow\infty$. To provide robust results
that are empirically relevant, we shall allow $|\beta|\rightarrow0$.
In fact, the use of drifting sequences is the standard practice for
establishing robust analysis in many areas of econometrics.\footnote{Examples include weak instruments (e.g., \citet{staiger1997instrumental}),
local-to-unit-root process (e.g., \citet{Stock1991}), estimation
on the boundary (e.g., \citet{andrews1999estimation}), model selection
(e.g., \citet{Leeb2005}), moment inequalities (e.g., \citet{andrews2009validity})
and time series forecasting (e.g., \citet{hirano2017forecasting})
among others. }
When $|\beta|$ is allowed to tend to zero, the situation is less
straight-forward. When $|b|,c\rightarrow0$,\footnote{Under strong IV condition (say $cov(D_{i},Z_{i})>0$), $b\gtrsim cov(D_{i},Z_{i})$,
which is bounded below by a positive constant.} the threshold of $-(\beta/c)(b-c)$ may or may not be tending to
infinity, depending on the ratio $|\beta|/c$. This paper tries to
find out how worried we should be about $c\rightarrow0$ (but $c\neq0$)
in this case. It turns out that the answer depends on whether $P(D_{i}=1\mid Z_{i}=0)$
is close to zero or not. We now explain our findings. Let us try to
construct a confidence set for the sign of $\mu_{1}$, i.e., a mapping
from the data to a subset of $\{-1,0,1\}$.
The case with $P(D_{i}=1\mid Z_{i}=0)$ being far away from zero is
common, e.g., $P(D_{i}=1\mid Z_{i}=0)>30\%$ in \citet{angrist1998children}.
The question in this case is whether a slight violation of monotonicity
is a big deal. From the discussion above, it is obvious that it is
not a big deal if $|\beta|/c\rightarrow\infty$. The natural way to
proceed is to construct a test or a data-dependent check. If the test
or data check suggests that monotonicity might be a problem, then
use $\{-1,0,1\}$ as the confidence set; if the test suggests otherwise,
then use the more informative set $\{-1\}$ (because $\beta<0$).
This overall procedure has an answer in every situation, regardless
of whether violation of monotonicity is a big deal. However, we show
that if this procedure is robust (i.e., valid with or without monotonicity),
then it must be uninformative (contains both $-1$ and $1$) under
monotonicity ($c=0$). Notice that this is true no matter how sophisticated
the test is. Therefore, although small enough violation of monotonicity
($|\beta|/c\rightarrow\infty$) does not cause a problem, we cannot
really check whether potential violation of monotonicity is small
enough. As a result, if $\beta$ is statistically significant and
the violation of monotonicity tends to zero, this violation may or
may not cause a problem, and we show that no data-dependent procedure
is smart enough to find out (even after imposing constraints such
as $\mu_{1}$ and $\mu_{2}$ having the same sign and both have magnitude
at least $|\beta|$ plus strong distributional restrictions such as
Bernoulli).
Another common case is $P(D_{i}=1\mid Z_{i}=0)\approx0$. This is
typical when the non-compliance is almost one-sided, e.g., $P(D_{i}=1\mid Z_{i}=0)<2\%$
in the example of Job Training Partnership Act (JPTA). Of course,
$c\rightarrow0$ would still cause a problem if $|\beta|/c\rightarrow0$.
However, since $P(D_{i}=1\mid Z_{i}=0)=a+c\geq c$, we can at least
carve out a ``safe'' region based on the data. For example, if $|\beta|/P(D_{i}=1\mid Z_{i}=0)\rightarrow\infty$,
then $|\beta|/c\rightarrow\infty$ and thus violation of monotonicity
does not cause a problem. Notice that $|\beta|/P(D_{i}=1\mid Z_{i}=0)\rightarrow\infty$
is testable since both $|\beta|$ and $P(D_{i}=1\mid Z_{i}=0)$ can
be learned from the data. In the case of binary outcomes, we provide
a precise characterization of the ``safe'' region. It turns out
that this ``safe'' region also highlights a sharp contrast. If the
data-generating process is in the ``safe'' region, $\mu_{1}$ and
$\beta$ have the same sign; otherwise, the impossibility result from
before holds.
We refer to this sharp contrast as a phase transition. On side of
the boundary, learning the sign of $\mu_{1}$ is trivial, whereas
it is impossible on the other side of the boundary. There is little
or nothing in the middle. This is the case no matter whether $P(D_{i}=1\mid Z_{i}=0)$
is far away from or close to zero. The difference is that in the former
case, it is impossible to find out on which side of the phase-transition
bound the data-generating process is; the testability is possible
in the latter case. In the former case, we still provide a precise
characterization of the phase transition at least for binary outcomes
because it is useful for robustness checks. For example, suppose that
the boundary of the phase transition is $0.5\%$ of defiers. Although
it is impossible to check which side of the boundary the data-generating
process is, it is still important to know that a mere $1\%$ of defiers
would put the data-generating process on the ``dangerous'' side
of the phase-transition boundary.
\begin{comment}
We find that there is a phase transition for the monotonicity condition.
If the degree to which monotonicity is violated exceeds a boundary,
it is impossible to learn the sign of LATE; otherwise, if the violation
of monotonicity is within the boundary, the sign of LATE is the sign
of $\beta$. One observation is that this boundary depends on $|\beta|$
so we should not assess the problem of monotonicity by the value of
$c$ alone. What makes this finding interesting is that this boundary
can go to zero. As a result, even if the proportion of defiers tends
to zero ($c\rightarrow0$), it does not necessarily mean that we can
reliably identify the sign of LATE. Moreover, we show that for the
purpose of learning the sign of LATE, it is impossible to test whether
the violation of monotonicity is beyond or within the phase-transition
boundary. Therefore, a subtle and undetectable violation of the monotonicity
condition could change the fundamental understanding of empirical
results: does the treatment benefit or hurt?
The situation is more tractable under almost one-sided non-compliance
$P(D_{i}=1\mid Z_{i}=0)\approx0$. In the case of binary outcomes,
the boundary of phase transition can be characterized by a simple
inequality that can be checked using the data.
\end{comment}
We also outline other ways of learning LATE. We show that the magnitude
of $\mu_{1}$ and $\mu_{2}$ is bounded below by $|\beta|\cdot\gamma$,
where $\gamma$ can be consistently estimated and satisfies $\gamma\asymp|cov(D_{i},Z_{i})|$.
We also show that imposing $|\mu_{1}|\geq|\mu_{2}|$ is enough to
identify the sign of LATE. These results do not rely on monotonicity
at all.
\begin{comment}
We provide more discussions in Section \ref{sec: conclusion}.
\end{comment}
\subsection{Related literature}
The literature of IV regressions has a long history dating back to
at least \citet{wright1928tariff}. An excellent review on this vast
literature can be found in \citet{imbens2014instrumental}. The framework
of LATE was started by the seminal work of \citet{imbens1994identification},
\citet{angrist1996identification} and \citet{abadie2003semiparametric}.
Since then the LATE-type idea has also been explored in the study
of quantile treatment effects, e.g., \citet{abadie2002instrumental}
and \citet{wuthrich2020comparison}. The framework of LATE fueled
many empirical work ever since the early influential studies including
\citet{angrist1991draft} and \citet{angrist1998children}. The nature
of monotonicity condition has been discussed for decades, e.g., \citet{robins1989analysis},
\citet{balke1995counterfactuals}, \citet{vytlacil2002independence}
and \citet{heckman2005structural}. Since there is not always an obvious
justification for the monotonicity condition, various specification
tests and alternatives have been proposed, see \citet{huber2015testing},
\citet{kitagawa2015test}, \citet{mourifie2017testing}, \citet{de2017tolerating}
and \citet{2011.06695} among many others. Another interesting approach
focuses on the partial identification of average treatment effects
or other quantities under various restrictions, see \citet{balke1997bounds},
\citet{manski2003partial}, \citet{swanson2018partial} and \citet{machado2019instrumental}
among many others.
\section{\label{sec: sign LATE}Learning the sign of LATE}
In the rest of the paper, we use the following notation. For $x\in\mathbb{R}$,
let
\[
{\rm sign}(x)=\begin{cases}
1 & \text{if }x>0\\
0 & \text{if }x=0\\
-1 & \text{if }x<0.
\end{cases}
\]
In Table \ref{tab: empirical}, we consider two empirical studies.
In the JPTA study, the treatment $D_{i}$ is job training and $Z_{i}$
is the indicator of the randomized offer of training and the treat.
In the example of \citet{angrist1998children}, we consider case with
$D_{i}$ being the indicator of being more than 2 children and $Z_{i}$
being the indicator of same sex in the first two children.
\begin{table}
\caption{\label{tab: empirical}Some estimates in two empirical studies}
\begin{centering}
\begin{tabular}{ccccc}
& & & & \tabularnewline
& $P(Z_{i}=1)$ & $P(D_{i}=1\mid Z_{i}=1)$ & $P(D_{i}=1\mid Z_{i}=0)$ & \tabularnewline
\cline{1-4} \cline{2-4} \cline{3-4} \cline{4-4}
JPTA & 0.6662 & 0.6228 & 0.0112 & \tabularnewline
\citet{angrist1998children} & 0.5048 & 0.4105 & 0.3557 & \tabularnewline
& & & & \tabularnewline
\end{tabular}
\par\end{centering}
\end{table}
We use these two studies to illustrate the two cases. In \citet{angrist1998children},
$P(D_{i}=1\mid Z_{i}=0)$ is not close to zero. Although the interpretation
of the result is clear under the monotonicity condition ($c=0$),
what if we have a small proportion of defiers ($c\approx0$)? We consider
this setting in Section \ref{subsec: small violation of mono}. In
the JPTA study, $P(D_{i}=1\mid Z_{i}=0)$ is close to zero, but does
this mean that we do not need to worry? We provide analysis for this
setting in Section \ref{subsec: small violation of one-sided noncomp}.
We state most of the theoretical results for $\beta<0$, but results
for $\beta>0$ can be obtained analogously.
\begin{comment}
For much of this section, we focus on local analysis ($|\beta|\rightarrow0$).
This is motivated by finite-sample observations. In many studies,
we often cannot reject $H_{0}:\ \beta=0$ at 5\% significance level
if we cut the sample size by half. To analyze such problems in an
asymptotic framework, it seems unrealistic to assume that $|\beta|$
is bounded away from zero and $n$ grows to infinity. A more reasonable
framework is to consider the case of $|\beta|\rightarrow0$. It is
worth noting that the theoretical results do not rule out any specific
rate at which $|\beta|$ tends to zero. Perhaps the most interesting
case is $|\beta|\gtrsim n^{-1/2}$ so that we can detect $\beta$
from the data but the question is how worried we should be even if
$c$ also tends to zero.
\end{comment}
\subsection{\label{subsec: small violation of mono}Small violation of monotonicity:
$P(D_{i}=1\mid Z_{i}=0)\gg0$}
For simplicity, we assume that the distribution of $Z_{i}\in\{0,1\}$
is known. We first introduce notations for the distribution of $(Y_{i}(1),Y_{i}(0),D_{i}(1),D_{i}(0))$.
We specify the distribution of $(D_{i}(1),D_{i}(0))$ and then the
conditional distribution of $(Y_{i}(1),Y_{i}(0))\mid(D_{i}(1),D_{i}(0))$.
The former is straight-forward; we simply use the same notation $a,b,c$
as in (\ref{eq: four types}). Define the conditional distribution
\[
H(y_{1},y_{0},d_{1},d_{0})=P\left(Y_{i}(1)\leq y_{1}\ and\ Y_{i}(0)\leq y_{0}\mid D_{i}(1)=d_{1},D_{i}(0)=d_{0}\right).
\]
Let $\theta=(a,b,c,H)$. Then the distribution of $(Z_{i},Y_{i}(1),Y_{i}(0),D_{i}(1),D_{i}(0))$
is indexed by $\theta$. Let $P_{\theta}$ and $E_{\theta}$ denote
the distribution and expectation under $\theta$, respectively. The
following quantities can be written as a function of $\theta$:
\begin{itemize}
\item LATE for compliers: $\mu_{1}(\theta)=E_{\theta}(Y_{i}(1)-Y_{i}(0)\mid D_{i}(1)=1,D_{i}(0)=0)$
\item LATE for defiers: $\mu_{2}(\theta)=E_{\theta}(Y_{i}(1)-Y_{i}(0)\mid D_{i}(1)=0,D_{i}(0)=1)$
\end{itemize}
From the data $W=\{(Y_{i},D_{i},Z_{i})\}_{i=1}^{n}$, we can identify
the following quantities:
\begin{itemize}
\item $k_{1}=a+b=E(D_{i}\mid Z_{i}=1)$
\item $k_{2}=a+c=E(D_{i}\mid Z_{i}=0)$
\item $\beta=[\mu_{1}(\theta)b-\mu_{2}(\theta)c]/(b-c)$.
\end{itemize}
Assuming that these three quantities are known, consider the following
parameter space:
\begin{multline*}
\Theta(\eta)=\biggl\{\theta=(a,b,c,H):\ a,b,c\in[0,1],\ a+b+c\in[0,1],\ a+b=k_{1},\ a+c=k_{2},\\
\max_{d,z\in\{0,1\}}P_{\theta}(|Y_{i}|\geq M\mid D_{i}=d,Z_{i}=z)=0,\ \frac{\mu_{1}(\theta)b-\mu_{2}(\theta)c}{b-c}=\beta,\\
\ |\mu_{1}(\theta)|\geq|\beta|,\ {\rm sign}(\mu_{1}(\theta))={\rm sign}(\mu_{2}(\theta)),\ 0\leq c\leq\eta\biggr\},
\end{multline*}
where $M>0$ is a constant.
Clearly, $\Theta(\eta)$ assumes a lot of structures that are typically
unavailable in practice. In particular, it assumes that $P(D_{i}=1\mid Z_{i}=1)$,
$P(D_{i}=1\mid Z_{i}=0)$ and the population IV regression coefficient
$\beta$ are known. Moreover, it assumes that the LATE for the compliers
and defiers has the same sign and that the magnitude of LATE for compliers
is not too small. The only difficulty is that $c$ (proportion of
defiers) might not be exactly zero and is allowed to be between $0$
and a small tolerance level $\eta$. The point of this subsection
is to show that even under these additional assumptions, allowing
for a small $\eta$ makes it impossible to learn the sign of LATE.
To make this point, we show that many data generating processes with
no defiers and ${\rm sign}(\mu_{1})=\beta$ are observationally equivalent
to those with a small proportion of defiers and ${\rm sign}(\mu_{1})\neq\beta$.
To formally state this, we define the following subset
\[
\Theta_{*}=\left\{ \theta=(a,b,c,H)\in\Theta(\eta):\ c=0,\ Q_{1,\theta}(1-\varepsilon_{1})-Q_{2,\theta}(\varepsilon_{1})>\varepsilon_{2}\right\} ,
\]
where $\varepsilon_{1},\varepsilon_{2}>0$ are constants and $Q_{1,\theta}$
and $Q_{2,\theta}$ are the quantile functions of $Y_{i}\mid(D_{i}=1,Z_{i}=0)$
and $Y_{i}\mid(D_{i}=0,Z_{i}=1)$ under $P_{\theta}$, respectively;
in other words, for any $\varepsilon\in(0,1)$,
\[
Q_{1,\theta}(\varepsilon)=\inf\left\{ t\in\mathbb{R}:\ P_{\theta}\left(Y_{i}\leq t\mid D_{i}=1,Z_{i}=0\right)\geq\varepsilon\right\}
\]
and
\[
Q_{2,\theta}(\varepsilon)=\inf\left\{ t\in\mathbb{R}:\ P_{\theta}\left(Y_{i}\leq t\mid D_{i}=0,Z_{i}=1\right)\geq\varepsilon\right\} .
\]
\begin{rem}
\label{rem: overlap}The condition of $Q_{1,\theta}(1-\varepsilon_{1})-Q_{2,\theta}(\varepsilon_{1})>\varepsilon_{2}$
is not very restrictive. When $Y_{i}$ is binary in $\{0,1\}$, $Q_{1,\theta}(1-\varepsilon_{1})-Q_{2,\theta}(\varepsilon_{1})>\varepsilon_{2}$
holds if
\[
E_{\theta}(Y_{i}\mid D_{i}=1,Z_{i}=0),E_{\theta}(Y_{i}\mid D_{i}=0,Z_{i}=1)\in(\varepsilon_{1},1-\varepsilon_{1}).
\]
This seems to be reasonable since it might be a bit unrealistic to
expect extreme situations with $E_{\theta}(Y_{i}\mid D_{i}=1,Z_{i}=0)\rightarrow0$
or $E_{\theta}(Y_{i}\mid D_{i}=0,Z_{i}=1)\rightarrow1$.
\end{rem}
We now state the key observation.
\begin{thm}
\label{thm: key imps}Let $M,\varepsilon_{2}>0$ and $\varepsilon_{1}\in(0,1)$.
Assume that $\beta<0$, $0<\eta<\varepsilon_{1}\min\{k_{2},1-k_{1},k_{1}-k_{2}\}$
and $3|\beta|/\eta<\varepsilon_{2}/(k_{1}-k_{2})$. Then for any $\theta\in\Theta_{*}$,
there exists $\tilde{\theta}\in\Theta(\eta)$ such that \\
(1) $P_{\theta}$ and $P_{\tilde{\theta}}$ imply the same distribution for
the observed data $(Y_{i},D_{i},Z_{i})$\\
(2) $\mu_{1}(\theta)=\beta<0$ and $\mu_{1}(\tilde{\theta}),\mu_{2}(\tilde{\theta})>-\beta>0$.
\end{thm}
Theorem \ref{thm: key imps} provides the key insight on why lack
of monotonicity creates difficult issues. For a small tolerance level
$\eta$, as long as $|\beta|/\eta$ is not too large, a data-generating
process with no defiers would look exactly like another data-generating
process with a small proportion of defiers such that LATE has different
signs under the two data-generating processes.
This sheds light on one of the most common problems in IV regressions.
If we reject $H_{0}:\ \beta=0$ and have some arguments against the
presence of defiers (e.g., $Z_{i}$ provides more information and
thus encourages $D_{i}=1$), can we reliably say that $\mu_{1}$ is
non-zero and has the same sign as $\beta$? By Theorem \ref{thm: key imps},
we see that a small proportion of defiers might be enough to invalidate
the result. In the asymptotic framework, rejecting $H_{0}:\ \beta=0$
in large samples is almost guaranteed when $|\beta|\gg n^{-1/2}$.
However, even if $\eta\rightarrow0$ (the proportion of defiers is
small), the observed data is indistinguishable from a distribution
with $\mu_{1}\neq{\rm sign}(\beta)$ when $\eta\gg|\beta|$. Therefore,
a slight violation of monotonicity creates a problem if $\eta\gg|\beta|$.
The natural question is whether or not we could check $\eta\gg|\beta|$
in the data. Unfortunately, the answer is no. We now show this using
an adaptivity argument based on Theorem \ref{thm: key imps}. We define
a confidence set of $\mu_{1}$ to be any measurable function mapping
the observed data $W_{n}$ to a subset of $\{-1,0,1\}$ with a guarantee
on the coverage probability.
\begin{cor}
\label{cor: adaptivity}Let $M,\varepsilon_{1},\varepsilon_{2},k_{1},k_{2}>0$
be any fixed constants such that $k_{1}-k_{2}>0$. Assume that $\beta<0$,
$\eta\rightarrow0$ and $|\beta|\rightarrow0$ such that $|\beta|\ll\eta$.
Let $CS(W_{n})$ be a confidence set for ${\rm sign}(\mu_{1}(\theta))$
with validity over $\Theta(\eta)$, i.e.,
\[
\liminf_{n\rightarrow\infty}\inf_{\theta\in\Theta(\eta)}P_{\theta}\left({\rm sign}(\mu_{1}(\theta))\in CS(W_{n})\right)\geq1-\alpha,
\]
where $\alpha\in(0,1)$. Then
\[
\liminf_{n\rightarrow\infty}\inf_{\theta\in\Theta_{*}}P_{\theta}\left(\{-1,1\}\subset CS(W_{n})\right)\geq1-2\alpha.
\]
\end{cor}
The assumption of $k_{2}=P(D_{i}=1\mid Z_{i}=0)$ being fixed models
the situation of $P(D_{i}=1\mid Z_{i}=0)$ being far from zero. The
strong IV condition corresponds to the requirement of $k_{1}-k_{2}$
being fixed and positive. There are some important implications of
Corollary \ref{cor: adaptivity}. The following discussions are for
$\eta\rightarrow0$. The same argument obvious holds if $\eta=k_{2}$.
First, when $n^{-1/2}\ll|\beta|\ll\eta\ll1$, the data allows us to
distinguish $|\beta|$ from zero, but it is still impossible to consistently
estimate the sign of $\mu_{1}$ over $\Theta(\eta)$. To see this,
consider an argument by contradiction. Suppose that there exists a
consistent estimator, a function $\sigma$ that maps $W_{n}$ to $\{-1,1,0\}$
and $\inf_{\theta\in\Theta(\eta)}P_{\theta}(\mu_{1}(\theta)=\sigma(W_{n}))\geq1-o(1)$.
In other words, $\sigma(W_{n})$ is only one value in $\{-1,0,1\}$.
Then we can set $CS(W_{n})=\{\sigma(W_{n})\}$ and the assumption
of Corollary \ref{cor: adaptivity} holds with an arbitrary $\alpha$,
say $\alpha=0.05$. The conclusion of Corollary \ref{cor: adaptivity}
says that with asymptotic probability at least 90\%, $CS(W_{n})=\{\sigma(W_{n})\}$
contains at least two elements, which is impossible since by construction
$\{\sigma(W_{n})\}$ is always a singleton. Hence, no consistent estimator
for ${\rm sign}(\mu_{1}(\theta))$ exists on $\Theta(\eta)$ when $n^{-1/2}\ll|\beta|\ll\eta\ll1$.
Second, clever specification tests (for monotonicity or $\eta\gg|\beta|$)
or other data-dependent procedures might not be able to address the
instability arising from a slight violation of the monotonicity condition.
One common purpose of specification tests is to allow us to handle
the problem based on the result of the tests. For example, when the
test tells us the monotonicity fails, we use a cautious set, say $\{-1,0,1\}$,
as the confidence set for ${\rm sign}(\mu_{1})$; when the test tells us
that the monotonicity holds, we use $\{{\rm sign}(\beta)\}$ as the confidence
set for ${\rm sign}(\mu_{1})$. Then by Corollary \ref{cor: adaptivity},
if this confidence set has uniform validity\footnote{One might wonder whether the requirement of uniform validity is too
stringent. It turns out that a similar result holds even if we replace
uniform validity with pointwise validity.} over $\Theta(\eta)$, the confidence set must be uninformative for
${\rm sign}(\mu_{1})$ on the nice set $\Theta_{*}$. If this confidence
set does not have uniform validity over $\Theta(\eta)$, then one
might question why we want to use a specification test in the first
place. Therefore, for the purpose of learning ${\rm sign}(\mu_{1})$, even
if we know that $|\mu_{1}(\theta)|\gg n^{-1/2}$, the monotonicity
condition is not really testable even when the alternative is only
a slight violation of monotonicity ($\eta\rightarrow0$).
Third, Corollary \ref{cor: adaptivity} implies a severe lack of adaptivity.
It states that it is impossible to be valid over the bigger set $\Theta(\eta)$
while maintaining efficiency on the nice set $\Theta_{*}$. Hence,
requiring validity over $\Theta(\eta)$ necessarily causes loss of
efficiency on $\Theta_{*}$. Notice that the loss of efficiency is
not on some points in $\Theta_{*}$. The efficiency loss occurs at
\textbf{every} point in $\Theta_{*}$; note that the second inequality
in Corollary \ref{cor: adaptivity} has $\inf_{\theta\in\Theta_{*}}$
rather than $\sup_{\theta\in\Theta_{*}}$. Therefore, the trade-off
of robustness and efficiency is quite stark.
The condition of $\eta\gg|\beta|$ turns out to define the boundary
of a ``phase transition''. We have seen that if we allow for $\eta\gg|\beta|$,
it is impossible to actually learn ${\rm sign}(\mu_{1})$. On the other
hand, we can show that if $\eta\ll|\beta|$, learning the sign of
LATE is trivial: ${\rm sign}(\mu_{1})={\rm sign}(\beta)$. To see this, notice
that $|\mu_{2}(\theta)|\leq2M$ (since $P_{\theta}(|Y_{i}|\leq M)=1$).
Since $\beta=(\mu_{1}(\theta)b-\mu_{2}(\theta)c)/(b-c)$, it follows
that
\[
\mu_{1}(\theta)=\lambda\mu_{2}(\theta)+(1-\lambda)\beta\leq2M\lambda+(1-\lambda)\beta,
\]
with $\lambda=c/b$. Since $\lambda=c/(k_{1}-k_{2}+c)$ and $c\leq\eta\ll|\beta|$,
we have that $\lambda=o(|\beta|)$. This means that
\[
\mu_{1}(\theta)\leq o(|\beta|)+(1-o(|\beta|))\beta=\beta(1+o(1)).
\]
By $\beta<0$, we have ${\rm sign}(\mu_{1}(\theta))={\rm sign}(\beta)$ asymptotically.
We now summarize these results.
\begin{thm}[Phase transition]
\label{thm: phase transition}Let $M,\varepsilon_{1},\varepsilon_{2},k_{1},k_{2}>0$
be any fixed constants such that $k_{1}-k_{2}>0$. Assume that $\beta<0$,
$\eta\rightarrow0$ and $|\beta|\rightarrow0$. \\
(1) If $\eta\gg|\beta|$, then for any $CS(W_{n})$ satisfying
\[
\liminf_{n\rightarrow\infty}\inf_{\theta\in\Theta(\eta)}P_{\theta}\left({\rm sign}(\mu_{1}(\theta))\in CS(W_{n})\right)\geq1-\alpha,
\]
with $\alpha\in(0,1)$, we have
\[
\liminf_{n\rightarrow\infty}\inf_{\theta\in\Theta_{*}}P_{\theta}\left(\{-1,1\}\subset CS(W_{n})\right)\geq1-2\alpha.
\]
(2) If $\eta\ll|\beta|$, then
\[
\liminf_{n\rightarrow\infty}\inf_{\theta\in\Theta(\eta)}P_{\theta}\left({\rm sign}(\mu_{1}(\theta))={\rm sign}(\beta)\right)=1.
\]
\end{thm}
By Theorem \ref{thm: phase transition}, the magnitude of $|\beta|$
serves as the boundary (in rate) of phase transition. A slight violation
may or may not be a huge problem depending on the order of magnitude
of the violation $\eta$. If $|\beta|\ll\eta$, even imposing the
extra condition of ${\rm sign}(\mu_{1}(\theta))={\rm sign}(\mu_{2}(\theta))$
does not help with learning the sign of $\mu_{1}$. In contrast, if
$|\beta|\gg\eta$, we can easily learn ${\rm sign}(\mu_{1}(\theta))$ without
assuming ${\rm sign}(\mu_{1}(\theta))={\rm sign}(\mu_{2}(\theta))$; the proof
of the second part of Theorem \ref{thm: phase transition} does not
rely on this condition.
Since $\Theta(\eta)$ allows for a large class of distributions, it
is difficult to say much more than the rate. However, when the outcome
variable is binary, we can precisely determine the boundary for the
phase transition.
\subsubsection{\label{subsec: small violation of mono binary}Exact boundary of
the phase transition for binary outcomes}
We define the counterparts of $\Theta(\eta)$ and $\Theta_{*}$ for
the binary outcomes. Let
\begin{multline}
\Theta_{binary}(\eta)=\biggl\{\theta=(a,b,c,H):\ a,b,c\in[0,1],\ a+b+c\in[0,1],\ a+b=k_{1},\ a+c=k_{2},\\
P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0),P_{\theta}(Y_{i}=D_{i}=0\mid Z_{i}=1)\geq\varepsilon,\\
P_{\theta}(Y_{i}\in\{0,1\})=1,\ 0\leq c\leq\eta\biggr\},\label{eq: binary para space}
\end{multline}
where $\varepsilon>0$ is a constant. The requirement that $P_{\theta}(Y_{i}=D_{i}\mid Z_{i})$
be bounded away from zero and one is mild in many applications.
\begin{thm}
\label{thm: phase trans binary easy}Let $k_{1},k_{2},\varepsilon\in(0,1)$
be given constants such that $k_{1}-k_{2}>0$. Suppose that $\eta\in[0,k_{2}]$
and $\beta$ satisfy $\beta<0$, $|\beta|\rightarrow0$ and $\eta\rightarrow0$.
\\
(1) If $\eta\geq|\beta|(k_{1}-k_{2})$, then there does not exist
any estimator of $\mu_{1}(\theta)$ that is consistent uniformly over
$\Theta_{binary}(\eta)$.\\
(2) If $\eta<|\beta|(k_{1}-k_{2})$, then ${\rm sign}(\mu_{1}(\theta))={\rm sign}(\beta)$
for any $\theta\in\Theta_{binary}(\eta)$.
\end{thm}
When the outcome variable is not binary, we can dichotomize it to
binary variables. We can define the new outcome variable $\tilde{Y}_{i}(1)=\mathbf{1}\{Y_{i}(1)\geq y\}$
and $\tilde{Y}_{i}(0)=\mathbf{1}\{Y_{i}(0)\geq y\}$, where $y$ is given. Then
the treatment effect is how much the treatment changes the probability
of $Y_{i}\geq y$.
For example, in \citet{angrist1998children}, one outcome variable
of interest $Y_{i}$ is the number of weeks a person worked in a year.
We can set $y=1$ and ask how the treatment changes the probability
of a person working for at least one week. Once we do this, we can
estimate $|\beta|(k_{1}-k_{2})$, the boundary of phase transition.
The results are in Table \ref{tab: AK98}. We see that any tolerance
level of $c$ above $0.52\%$ can cause a serious problem for the
question of whether or not the LATE is negative. Hence, even if the
proportion of defiers is known to be at most 1\%, it might not be
obvious that we can safely conclude a negative LATE.
\begin{table}
\caption{\label{tab: AK98}Estimating the phase-transition boundary using data
in \citet{angrist1998children}}
\begin{centering}
\begin{tabular}{crc}
& & \tabularnewline
& Point estimate & \tabularnewline
\cline{1-2} \cline{2-2}
$\beta$ & -0.0950 & \tabularnewline
$|\beta|(k_{1}-k_{2})$ & 0.0052 & \tabularnewline
& & \tabularnewline
\end{tabular}
\par\end{centering}
{\small{}The outcome variable is whether or not a person worked for
at least one week in the year.}{\small\par}
\end{table}
\subsection{\label{subsec: small violation of one-sided noncomp}What about $P(D_{i}=1\mid Z_{i}=0)\approx0$?}
In many studies, the absence of defiers is justified by the one-sidedness
of non-compliance. For example, $Z_{i}\in\{0,1\}$ is the randomly
assigned treatment and $D_{i}$ is the actually treatment status.
When the compliance is not perfect (i.e., $P(Z_{i}=D_{i})<1$), the
non-compliance is often one-sided: $P(D_{i}=1\mid Z_{i}=1)<1$ but
$P(D_{i}=1\mid Z_{i}=0)=0$. However, we discuss a small violation
to this ideal case $P(D_{i}=1\mid Z_{i}=0)\approx0$, see JPTA in
Table \ref{tab: empirical} as an example.
Here, we provide a discussion for the case of $k_{2}=P(D_{i}=1\mid Z_{i}=0)\rightarrow0$
in the case of binary outcomes. This is different from Theorem \ref{thm: phase trans binary easy},
which assumes that $k_{2}$ is bounded away from zero. Moreover, when
$k_{2}\rightarrow0$, the natural choice of $\eta$ is $\eta=k_{2}$.
To analyze this case, we consider the following parameter space
\[
\Theta_{binary,*}=\biggl\{\theta=(a,b,c,H)\in\Theta_{binary}(k_{2}):\ P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)<|\beta|(k_{1}-k_{2})\biggr\},
\]
where $\Theta_{binary}(\cdot)$ is defined in (\ref{eq: binary para space}).
\begin{thm}
\label{thm: phase trans binary hard}Let $k_{1},\varepsilon\in(0,1)$
be given constants. Suppose that $k_{2},|\beta|\rightarrow0$ and
$\beta<0$.\\
(1) there does not exist any estimator of $\mu_{1}(\theta)$ that
is consistent over $\Theta_{binary}(k_{2})\backslash\Theta_{binary,*}$.\\
(2) ${\rm sign}(\mu_{1}(\theta))={\rm sign}(\beta)$ for any $\theta\in\Theta_{binary,*}$.
\end{thm}
We notice that $P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)<|\beta|(k_{1}-k_{2})$
(the boundary in Theorem \ref{thm: phase trans binary hard}) is not
the same as applying $\eta=k_{2}$ to Theorem \ref{thm: phase trans binary easy}.
Applying $\eta=k_{2}$ to $\eta<|\beta|(k_{1}-k_{2})$ in Theorem
\ref{thm: phase trans binary easy} leads to $k_{2}<|\beta|(k_{1}-k_{2})$.
However, $P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)\leq k_{2}$. To see
this, observe that $P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)=E_{\theta}(Y_{i}D_{i}\mid Z_{i}=0)\leq E_{\theta}(D_{i}\mid Z_{i}=0)=k_{2}$.
What this means in practice is that Theorem \ref{thm: phase trans binary hard}
makes it easier to be on the ``nice'' side of the boundary; instead
of requiring $|\beta|(k_{1}-k_{2})$ to be above $k_{2}$, we require
it to be above $P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)$.
A more important implication of Theorem \ref{thm: phase trans binary hard}
is that it is possible to check whether or not we are on the nice
side of the phase-transition boundary. The set $\Theta_{binary,*}$
is defined by the testable condition
\[
P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)<|\beta|(k_{1}-k_{2}).
\]
We apply this to the JPTA study. We set the outcome variable to be
$\mathbf{1}\{\text{income}<\$50000\}$. This means that we study the effect
of treatment on the probability of earning less than \$50000.\footnote{We choose ``less than'' instead of ``more than'' to get a negative
$\beta$. The interpretation is intuitively the same. Receiving treatment
makes it less likely to earn less than \$50000 so it makes it more
likely to earn at least \$50000.} The results are in Table \ref{tab: JPTA}. Based on only the point
estimates, the data-generating process is in $\Theta_{binary,*}$,
which is the ``nice'' side of the phase-transition boundary.
\begin{table}
\caption{\label{tab: JPTA}Estimating the phase-transition boundary using data
in JPTA}
\begin{centering}
\begin{tabular}{crc}
& & \tabularnewline
& Point estimate & \tabularnewline
\cline{1-2} \cline{2-2}
$\beta$ & -0.0363 & \tabularnewline
$|\beta|(k_{1}-k_{2})$ & 0.0222 & \tabularnewline
$P_{\theta}(Y_{i}=1,D_{i}=1\mid Z_{i}=0)$ & 0.0157 & \tabularnewline
\end{tabular}
\par\end{centering}
{\small{}The outcome variable is whether or not a person earns less
than \$50000 a year.}{\small\par}
\end{table}
\section{\label{sec: no monotonicity}Learning LATE without monotonicity:
an alternative}
There are already alternatives to the monotonicity condition in the
literature. We add the following discussions.
\subsection{Learning the magnitude without any additional assumption}
We now outline a lower bound for the magnitude of LATE.
\begin{thm}
\label{thm: key}Let Assumption \ref{assu: IV exogeneity} hold. Assume
$cov(D_{i},Z_{i})\neq0$. Then
\[
\max\{|\mu_{1}|,|\mu_{2}|\}\geq|\beta|\cdot\gamma,
\]
where
\[
\gamma=\frac{\left|E(D_{i}\mid Z_{i}=1)-E(D_{i}\mid Z_{i}=0)\right|}{E(D_{i}\mid Z_{i}=1)+E(D_{i}\mid Z_{i}=0)}.
\]
\end{thm}
Notice that $\gamma\gtrsim|cov(D_{i},Z_{i})|$. Therefore, in the
case of strong instruments, the lower bound $|\beta|\cdot\gamma$
is not too small compared to $|\beta|$. From the proof, we can see
that the lower bound is also tight in that the equality can hold (because
the minimum in the proof can be achieved).
It is worth noting that the lower bound in Theorem \ref{thm: key}
can be related to intent-to-treat (ITT) effects. We notice that
\[
|\beta|\cdot\gamma=\frac{\left|E(Y_{i}\mid Z_{i}=1)-E(Y_{i}\mid Z_{i}=0)\right|}{E(D_{i}\mid Z_{i}=1)+E(D_{i}\mid Z_{i}=0)}=\frac{|ITT|}{E(D_{i}\mid Z_{i}=1)+E(D_{i}\mid Z_{i}=0)}.
\]
Therefore, the lower bound satisfies $|\beta|\cdot\gamma\geq|ITT|/2$.
Therefore, whenever we find that $\beta\neq0$ or $ITT\neq0$, it
means that the treatment effect is not zero and we can use $|\beta|\cdot\gamma$
as a lower bound.
\subsection{Learning the sign: $|\mu_{1}|\protect\geq|\mu_{2}|$}
Learning the sign of LATE requires extra restrictions. This is inevitable;
otherwise, monotonicity would be testable in Section \ref{subsec: small violation of mono binary}.
It turns out that simple restrictions such as $|\mu_{1}|\geq|\mu_{2}|$
would suffice.
\begin{thm}
\label{thm: new ID sign}Let Assumption \ref{assu: IV exogeneity}
hold. Suppose that $\beta\neq0$ and $cov(D_{i},Z_{i})>0$. If $|\mu_{1}|\geq|\mu_{2}|$,
then ${\rm sign}(\mu_{1})={\rm sign}(\beta)$.
\end{thm}
We should notice that Theorem \ref{thm: new ID sign} imposes more
than $|\mu_{1}|\geq|\mu_{2}|$. The assumption of $cov(D_{i},Z_{i})>0$
is not without loss of generality since the condition of $|\mu_{2}|\geq|\mu_{1}|$
is not enough to identify the sign of LATE. Therefore, the result
should be viewed in the context of the empirical application.
\begin{comment}
Conclusion and implications for practitioners
In Sections \ref{subsec: small violation of mono} and \ref{subsec: small violation of one-sided noncomp},
we have considered two cases: $P(D_{i}=1\mid Z_{i}=0)$ is large and
$P(D_{i}=1\mid Z_{i}=0)$ is close to zero. In both cases, there is
a phase transition. When the proportion of defiers is below a threshold,
we can easily learn the sign of LATE; otherwise, it is impossible
to consistently estimate the sign of LATE. In both cases, the boundary
of phase transition can be accurately stated for binary outcomes.
However, there is a crucial difference. In the former case, it is
impossible to test whether or not we are on the nice side of the phase-transition
boundary. Therefore, small and undetectable violations of the monotonicity
can lead to serious trouble in learning the sign of LATE, thereby
making the result hinge on untestable conditions. In contrast, in
the latter case ($P(D_{i}=1\mid Z_{i}=0)\approx0$), it is possible
to check whether the data is on the impossible side or the easy side
of the phase-transition boundary.
For practitioners, the results in Section \ref{sec: sign LATE} provide
useful tools for understanding and interpreting empirical results.
When $P(D_{i}=1\mid Z_{i}=0)$ is not close to zero, we should take
extra care. One simple check is as follows: convert the outcome to
a binary variable and estimate the phase-transition boundary $|\beta|(k_{1}-k_{2})$.
For example, if the presence of 0.1\% defiers is enough to flip the
sign of LATE, it is useful to know. Unfortunately, if $|\beta|(k_{1}-k_{2})$
is high, say 20\%, the data still cannot tell us that there are less
than 20\% of defiers due to the untestability result in Section \ref{subsec: small violation of mono}.
Hopefully, we can rely on economic theory for this.
When $P(D_{i}=1\mid Z_{i}=0)$ is close to zero, we can rely more
on the data. We can check the explicitly written boundary of phase
transition in Section \ref{subsec: small violation of one-sided noncomp}
(after possibly dichotomizing $Y_{i}$).
Alternatively, we can choose to completely avoid monotonicity. Section
\ref{sec: no monotonicity} provides some simple options, but their
appropriateness should be evaluated in the context of the empirical
study.
\end{comment}
\bibliographystyle{apalike}
\bibliography{SC_biblio}