Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
46,492 characters · 10 sections · 36 citation commands
Phase transition of the monotonicity assumption in learning local average treatment effects
\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long
Instrumental variables (IV) regressions have been widely used to study treatment effects in economics and other disciplines. One important conceptual framework that justifies the causal interpretation of IV regressions is the local average treatment (LATE). In this paper, we consider a simple setting with no covariates and discuss the sensitivity of the key monotonicity assumption in the LATE framework.
We observe iid data $\{(Y_{i},D_{i},Z_{i})\}_{i=1}^{n}$, where $Y_{i}=Y_{i}(1)D_{i}+Y_{i}(0)(1-D_{i})$, $D_{i}=D_{i}(1)Z_{i}+D_{i}(0)(1-Z_{i})$ and $Z_{i},D_{i}(1),D_{i}(0)\in\{0,1\}$. The treatment effect is $Y_{i}(1)-Y_{i}(0)$. (Notice that this assumes that $Z_{i}$ does not directly affect the potential outcomes: $Y_{i}(z,d)=Y_{i}(d)$ for $z,d\in\{0,1\}$.) Throughout the paper, we maintain the assumption that the IV is strong ($|cov(D_{i},Z_{i})|\geq C$ for a constant $C>0$) and the following exogeneity condition
The typical IV regression exploits the moment condition $EZ_{i}(Y_{i}-D_{i}\beta)=EZ_{i}E(Y_{i}-D_{i}\beta)$ (i.e., $cov(Z_{i},Y_{i}-D_{i}\beta)=0$). This means\footnote{Under Assumption (ref), $\beta$ can be written in other ways. For example, $\beta=\frac{E(Y_{i}\mid Z_{i}=1)-E(Y_{i}\mid Z_{i}=0)}{E(D_{i}\mid Z_{i}=1)-E(D_{i}\mid Z_{i}=0)}$.} that \[ \beta=\frac{E(Y_{i}Z_{i})-E(Y_{i})E(Z_{i})}{E(D_{i}Z_{i})-E(D_{i})E(Z_{i})}. \]
To describe the causal interpretation of $\beta$, we categorize the population into four types depending on the value of $(D_{i}(1),D_{i}(0))\in\{0,1\}\times\{0,1\}$. We introduce their definitions and their probability:
As shown in angrist1996identification,
where $\mu_{1}$ and $\mu_{2}$ are the local average treatment effects (LATE) for compliers and defiers, respectively: \[
\]
Since ((ref)) is in general not a convex combination of $\mu_{1}$ and $\mu_{2}$, $\beta$ typically does not have a causal interpretation without further assumptions. The classical assumption that makes $\beta$ causally interpretable is the following monotonicity condition.
Clearly, under the monotonicity condition, ((ref)) implies that $\beta=\mu_{1}$ or $\beta=\mu_{2}$, which can be summarized as $\beta=E(Y_{i}(1)-Y_{i}(0)\mid D_{i}(1)\neq D_{i}(0))$. Hence, $\beta$ is interpreted as the average treatment effect on the sub-population for which $D_{i}(1)\neq D_{i}(0)$.
As pointed out by imbens2014instrumental, perhaps the strongest justification of the monotonicity condition is when the instrument provides an incentive to choose the treatment or when the treatment is simply not an option without $Z_{i}=1$. Outside these situations, the validity of the monotonicity condition is not always obvious. In this paper, we try to answer the following questions
Let us explain why (some of) these questions might be quite subtle and difficult although they seem to have an obvious answer at the first glance.
The majority of the paper focuses on the seemingly simple question of learning the sign of LATE (so we can answer the basic question of whether the treatment is beneficial or harmful). In particular, whether we can conclude that $\mu_{1}$ and $\beta$ have the same sign when monotonicity is slightly violated ($c\approx0$). By rearranging ((ref)), we have \[ \mu_{1}=\frac{c\mu_{2}+(b-c)\beta}{b}. \]
Suppose that $\beta<0$. It is easy to see that $\mu_{1}<0$ (and thus has the same sign as $\beta$) if and only if $\mu_{2}<-(\beta/c)(b-c)$. Throughout the paper, we assume that $P(|Y_{i}|\leq M)=1$ for a constant $M>0$. Then the question of learning the sign of $\mu_{1}$ would seem straight-forward. If $c\rightarrow0$ and $|\beta|$ and $b$ are bounded below by a positive constant, then the threshold $-(\beta/c)(b-c)$ tends to infinity. Since $\mu_{2}$ is bounded (due to the boundedness of $Y_{i}$), the condition of $\mu_{2}<-(\beta/c)(b-c)$ is asymptotically satisfied. Hence, the conclusion would be that no matter how slowly $c$ goes to zero, it is asymptotically valid to conclude that $\mu_{1}$ and $\beta$ have the same sign.
One subtly is whether modeling $|\beta|$ as a quantity bounded below by a positive constant is an asymptotic framework that is empirically relevant. In many empirical studies, if we throw away half of the data and run the IV regression, we often do not find a statistically significant $\beta$ anymore. Then it might be too strong to assume that $|\beta|$ is of a much larger order of magnitude compared to the estimation noise. Moreover, statistically significance of $\beta$ does not mean that $|\beta|$ is bounded below by a positive constant; statistical significance is asymptotically guaranteed even if $|\beta|\rightarrow0$ and $\sqrt{n}|\beta|\rightarrow\infty$. To provide robust results that are empirically relevant, we shall allow $|\beta|\rightarrow0$. In fact, the use of drifting sequences is the standard practice for establishing robust analysis in many areas of econometrics.\footnote{Examples include weak instruments (e.g., staiger1997instrumental), local-to-unit-root process (e.g., Stock1991), estimation on the boundary (e.g., andrews1999estimation), model selection (e.g., Leeb2005), moment inequalities (e.g., andrews2009validity) and time series forecasting (e.g., hirano2017forecasting) among others. }
When $|\beta|$ is allowed to tend to zero, the situation is less straight-forward. When $|b|,c\rightarrow0$,\footnote{Under strong IV condition (say $cov(D_{i},Z_{i})>0$), $b\gtrsim cov(D_{i},Z_{i})$, which is bounded below by a positive constant.} the threshold of $-(\beta/c)(b-c)$ may or may not be tending to infinity, depending on the ratio $|\beta|/c$. This paper tries to find out how worried we should be about $c\rightarrow0$ (but $c\neq0$) in this case. It turns out that the answer depends on whether $P(D_{i}=1\mid Z_{i}=0)$ is close to zero or not. We now explain our findings. Let us try to construct a confidence set for the sign of $\mu_{1}$, i.e., a mapping from the data to a subset of $\{-1,0,1\}$.
The case with $P(D_{i}=1\mid Z_{i}=0)$ being far away from zero is common, e.g., $P(D_{i}=1\mid Z_{i}=0)>30\%$ in angrist1998children. The question in this case is whether a slight violation of monotonicity is a big deal. From the discussion above, it is obvious that it is not a big deal if $|\beta|/c\rightarrow\infty$. The natural way to proceed is to construct a test or a data-dependent check. If the test or data check suggests that monotonicity might be a problem, then use $\{-1,0,1\}$ as the confidence set; if the test suggests otherwise, then use the more informative set $\{-1\}$ (because $\beta<0$). This overall procedure has an answer in every situation, regardless of whether violation of monotonicity is a big deal. However, we show that if this procedure is robust (i.e., valid with or without monotonicity), then it must be uninformative (contains both $-1$ and $1$) under monotonicity ($c=0$). Notice that this is true no matter how sophisticated the test is. Therefore, although small enough violation of monotonicity ($|\beta|/c\rightarrow\infty$) does not cause a problem, we cannot really check whether potential violation of monotonicity is small enough. As a result, if $\beta$ is statistically significant and the violation of monotonicity tends to zero, this violation may or may not cause a problem, and we show that no data-dependent procedure is smart enough to find out (even after imposing constraints such as $\mu_{1}$ and $\mu_{2}$ having the same sign and both have magnitude at least $|\beta|$ plus strong distributional restrictions such as Bernoulli).
Another common case is $P(D_{i}=1\mid Z_{i}=0)\approx0$. This is typical when the non-compliance is almost one-sided, e.g., $P(D_{i}=1\mid Z_{i}=0)<2\%$ in the example of Job Training Partnership Act (JPTA). Of course, $c\rightarrow0$ would still cause a problem if $|\beta|/c\rightarrow0$. However, since $P(D_{i}=1\mid Z_{i}=0)=a+c\geq c$, we can at least carve out a “safe” region based on the data. For example, if $|\beta|/P(D_{i}=1\mid Z_{i}=0)\rightarrow\infty$, then $|\beta|/c\rightarrow\infty$ and thus violation of monotonicity does not cause a problem. Notice that $|\beta|/P(D_{i}=1\mid Z_{i}=0)\rightarrow\infty$ is testable since both $|\beta|$ and $P(D_{i}=1\mid Z_{i}=0)$ can be learned from the data. In the case of binary outcomes, we provide a precise characterization of the “safe” region. It turns out that this “safe” region also highlights a sharp contrast. If the data-generating process is in the “safe” region, $\mu_{1}$ and $\beta$ have the same sign; otherwise, the impossibility result from before holds.
We refer to this sharp contrast as a phase transition. On side of the boundary, learning the sign of $\mu_{1}$ is trivial, whereas it is impossible on the other side of the boundary. There is little or nothing in the middle. This is the case no matter whether $P(D_{i}=1\mid Z_{i}=0)$ is far away from or close to zero. The difference is that in the former case, it is impossible to find out on which side of the phase-transition bound the data-generating process is; the testability is possible in the latter case. In the former case, we still provide a precise characterization of the phase transition at least for binary outcomes because it is useful for robustness checks. For example, suppose that the boundary of the phase transition is $0.5\%$ of defiers. Although it is impossible to check which side of the boundary the data-generating process is, it is still important to know that a mere $1\%$ of defiers would put the data-generating process on the “dangerous” side of the phase-transition boundary.
We also outline other ways of learning LATE. We show that the magnitude of $\mu_{1}$ and $\mu_{2}$ is bounded below by $|\beta|\cdot\gamma$, where $\gamma$ can be consistently estimated and satisfies $\gamma\asymp|cov(D_{i},Z_{i})|$. We also show that imposing $|\mu_{1}|\geq|\mu_{2}|$ is enough to identify the sign of LATE. These results do not rely on monotonicity at all.
The literature of IV regressions has a long history dating back to at least wright1928tariff. An excellent review on this vast literature can be found in imbens2014instrumental. The framework of LATE was started by the seminal work of imbens1994identification, angrist1996identification and abadie2003semiparametric. Since then the LATE-type idea has also been explored in the study of quantile treatment effects, e.g., abadie2002instrumental and wuthrich2020comparison. The framework of LATE fueled many empirical work ever since the early influential studies including angrist1991draft and angrist1998children. The nature of monotonicity condition has been discussed for decades, e.g., robins1989analysis, balke1995counterfactuals, vytlacil2002independence and heckman2005structural. Since there is not always an obvious justification for the monotonicity condition, various specification tests and alternatives have been proposed, see huber2015testing, kitagawa2015test, mourifie2017testing, de2017tolerating and 2011.06695 among many others. Another interesting approach focuses on the partial identification of average treatment effects or other quantities under various restrictions, see balke1997bounds, manski2003partial, swanson2018partial and machado2019instrumental among many others.
In the rest of the paper, we use the following notation. For $x\in\mathbb{R}$, let \[ {\rm sign}(x)=
\]
In Table (ref), we consider two empirical studies. In the JPTA study, the treatment $D_{i}$ is job training and $Z_{i}$ is the indicator of the randomized offer of training and the treat. In the example of angrist1998children, we consider case with $D_{i}$ being the indicator of being more than 2 children and $Z_{i}$ being the indicator of same sex in the first two children.
We use these two studies to illustrate the two cases. In angrist1998children, $P(D_{i}=1\mid Z_{i}=0)$ is not close to zero. Although the interpretation of the result is clear under the monotonicity condition ($c=0$), what if we have a small proportion of defiers ($c\approx0$)? We consider this setting in Section (ref). In the JPTA study, $P(D_{i}=1\mid Z_{i}=0)$ is close to zero, but does this mean that we do not need to worry? We provide analysis for this setting in Section (ref). We state most of the theoretical results for $\beta<0$, but results for $\beta>0$ can be obtained analogously.
For simplicity, we assume that the distribution of $Z_{i}\in\{0,1\}$ is known. We first introduce notations for the distribution of $(Y_{i}(1),Y_{i}(0),D_{i}(1),D_{i}(0))$. We specify the distribution of $(D_{i}(1),D_{i}(0))$ and then the conditional distribution of $(Y_{i}(1),Y_{i}(0))\mid(D_{i}(1),D_{i}(0))$. The former is straight-forward; we simply use the same notation $a,b,c$ as in ((ref)). Define the conditional distribution \[ H(y_{1},y_{0},d_{1},d_{0})=P\left(Y_{i}(1)\leq y_{1}\ and\ Y_{i}(0)\leq y_{0}\mid D_{i}(1)=d_{1},D_{i}(0)=d_{0}\right). \]
Let $\theta=(a,b,c,H)$. Then the distribution of $(Z_{i},Y_{i}(1),Y_{i}(0),D_{i}(1),D_{i}(0))$ is indexed by $\theta$. Let $P_{\theta}$ and $E_{\theta}$ denote the distribution and expectation under $\theta$, respectively. The following quantities can be written as a function of $\theta$:
From the data $W=\{(Y_{i},D_{i},Z_{i})\}_{i=1}^{n}$, we can identify the following quantities:
Assuming that these three quantities are known, consider the following parameter space:
where $M>0$ is a constant.
Clearly, $\Theta(\eta)$ assumes a lot of structures that are typically unavailable in practice. In particular, it assumes that $P(D_{i}=1\mid Z_{i}=1)$, $P(D_{i}=1\mid Z_{i}=0)$ and the population IV regression coefficient $\beta$ are known. Moreover, it assumes that the LATE for the compliers and defiers has the same sign and that the magnitude of LATE for compliers is not too small. The only difficulty is that $c$ (proportion of defiers) might not be exactly zero and is allowed to be between $0$ and a small tolerance level $\eta$. The point of this subsection is to show that even under these additional assumptions, allowing for a small $\eta$ makes it impossible to learn the sign of LATE. To make this point, we show that many data generating processes with no defiers and ${\rm sign}(\mu_{1})=\beta$ are observationally equivalent to those with a small proportion of defiers and ${\rm sign}(\mu_{1})\neq\beta$.
To formally state this, we define the following subset \[ \Theta_{*}=\left\{ \theta=(a,b,c,H)\in\Theta(\eta):\ c=0,\ Q_{1,\theta}(1-\varepsilon_{1})-Q_{2,\theta}(\varepsilon_{1})>\varepsilon_{2}\right\} , \] where $\varepsilon_{1},\varepsilon_{2}>0$ are constants and $Q_{1,\theta}$ and $Q_{2,\theta}$ are the quantile functions of $Y_{i}\mid(D_{i}=1,Z_{i}=0)$ and $Y_{i}\mid(D_{i}=0,Z_{i}=1)$ under $P_{\theta}$, respectively; in other words, for any $\varepsilon\in(0,1)$, \[ Q_{1,\theta}(\varepsilon)=\inf\left\{ t\in\mathbb{R}:\ P_{\theta}\left(Y_{i}\leq t\mid D_{i}=1,Z_{i}=0\right)\geq\varepsilon\right\} \] and \[ Q_{2,\theta}(\varepsilon)=\inf\left\{ t\in\mathbb{R}:\ P_{\theta}\left(Y_{i}\leq t\mid D_{i}=0,Z_{i}=1\right)\geq\varepsilon\right\} . \]
We now state the key observation.
Theorem (ref) provides the key insight on why lack of monotonicity creates difficult issues. For a small tolerance level $\eta$, as long as $|\beta|/\eta$ is not too large, a data-generating process with no defiers would look exactly like another data-generating process with a small proportion of defiers such that LATE has different signs under the two data-generating processes.
This sheds light on one of the most common problems in IV regressions. If we reject $H_{0}:\ \beta=0$ and have some arguments against the presence of defiers (e.g., $Z_{i}$ provides more information and thus encourages $D_{i}=1$), can we reliably say that $\mu_{1}$ is non-zero and has the same sign as $\beta$? By Theorem (ref), we see that a small proportion of defiers might be enough to invalidate the result. In the asymptotic framework, rejecting $H_{0}:\ \beta=0$ in large samples is almost guaranteed when $|\beta|\gg n^{-1/2}$. However, even if $\eta\rightarrow0$ (the proportion of defiers is small), the observed data is indistinguishable from a distribution with $\mu_{1}\neq{\rm sign}(\beta)$ when $\eta\gg|\beta|$. Therefore, a slight violation of monotonicity creates a problem if $\eta\gg|\beta|$.
The natural question is whether or not we could check $\eta\gg|\beta|$ in the data. Unfortunately, the answer is no. We now show this using an adaptivity argument based on Theorem (ref). We define a confidence set of $\mu_{1}$ to be any measurable function mapping the observed data $W_{n}$ to a subset of $\{-1,0,1\}$ with a guarantee on the coverage probability.
The assumption of $k_{2}=P(D_{i}=1\mid Z_{i}=0)$ being fixed models the situation of $P(D_{i}=1\mid Z_{i}=0)$ being far from zero. The strong IV condition corresponds to the requirement of $k_{1}-k_{2}$ being fixed and positive. There are some important implications of Corollary (ref). The following discussions are for $\eta\rightarrow0$. The same argument obvious holds if $\eta=k_{2}$.
First, when $n^{-1/2}\ll|\beta|\ll\eta\ll1$, the data allows us to distinguish $|\beta|$ from zero, but it is still impossible to consistently estimate the sign of $\mu_{1}$ over $\Theta(\eta)$. To see this, consider an argument by contradiction. Suppose that there exists a consistent estimator, a function $\sigma$ that maps $W_{n}$ to $\{-1,1,0\}$ and $\inf_{\theta\in\Theta(\eta)}P_{\theta}(\mu_{1}(\theta)=\sigma(W_{n}))\geq1-o(1)$. In other words, $\sigma(W_{n})$ is only one value in $\{-1,0,1\}$. Then we can set $CS(W_{n})=\{\sigma(W_{n})\}$ and the assumption of Corollary (ref) holds with an arbitrary $\alpha$, say $\alpha=0.05$. The conclusion of Corollary (ref) says that with asymptotic probability at least 90%, $CS(W_{n})=\{\sigma(W_{n})\}$ contains at least two elements, which is impossible since by construction $\{\sigma(W_{n})\}$ is always a singleton. Hence, no consistent estimator for ${\rm sign}(\mu_{1}(\theta))$ exists on $\Theta(\eta)$ when $n^{-1/2}\ll|\beta|\ll\eta\ll1$.
Second, clever specification tests (for monotonicity or $\eta\gg|\beta|$) or other data-dependent procedures might not be able to address the instability arising from a slight violation of the monotonicity condition. One common purpose of specification tests is to allow us to handle the problem based on the result of the tests. For example, when the test tells us the monotonicity fails, we use a cautious set, say $\{-1,0,1\}$, as the confidence set for ${\rm sign}(\mu_{1})$; when the test tells us that the monotonicity holds, we use $\{{\rm sign}(\beta)\}$ as the confidence set for ${\rm sign}(\mu_{1})$. Then by Corollary (ref), if this confidence set has uniform validity\footnote{One might wonder whether the requirement of uniform validity is too stringent. It turns out that a similar result holds even if we replace uniform validity with pointwise validity.} over $\Theta(\eta)$, the confidence set must be uninformative for ${\rm sign}(\mu_{1})$ on the nice set $\Theta_{*}$. If this confidence set does not have uniform validity over $\Theta(\eta)$, then one might question why we want to use a specification test in the first place. Therefore, for the purpose of learning ${\rm sign}(\mu_{1})$, even if we know that $|\mu_{1}(\theta)|\gg n^{-1/2}$, the monotonicity condition is not really testable even when the alternative is only a slight violation of monotonicity ($\eta\rightarrow0$).
Third, Corollary (ref) implies a severe lack of adaptivity. It states that it is impossible to be valid over the bigger set $\Theta(\eta)$ while maintaining efficiency on the nice set $\Theta_{*}$. Hence, requiring validity over $\Theta(\eta)$ necessarily causes loss of efficiency on $\Theta_{*}$. Notice that the loss of efficiency is not on some points in $\Theta_{*}$. The efficiency loss occurs at every point in $\Theta_{*}$; note that the second inequality in Corollary (ref) has $\inf_{\theta\in\Theta_{*}}$ rather than $\sup_{\theta\in\Theta_{*}}$. Therefore, the trade-off of robustness and efficiency is quite stark.
The condition of $\eta\gg|\beta|$ turns out to define the boundary of a “phase transition”. We have seen that if we allow for $\eta\gg|\beta|$, it is impossible to actually learn ${\rm sign}(\mu_{1})$. On the other hand, we can show that if $\eta\ll|\beta|$, learning the sign of LATE is trivial: ${\rm sign}(\mu_{1})={\rm sign}(\beta)$. To see this, notice that $|\mu_{2}(\theta)|\leq2M$ (since $P_{\theta}(|Y_{i}|\leq M)=1$). Since $\beta=(\mu_{1}(\theta)b-\mu_{2}(\theta)c)/(b-c)$, it follows that \[ \mu_{1}(\theta)=\lambda\mu_{2}(\theta)+(1-\lambda)\beta\leq2M\lambda+(1-\lambda)\beta, \] with $\lambda=c/b$. Since $\lambda=c/(k_{1}-k_{2}+c)$ and $c\leq\eta\ll|\beta|$, we have that $\lambda=o(|\beta|)$. This means that \[ \mu_{1}(\theta)\leq o(|\beta|)+(1-o(|\beta|))\beta=\beta(1+o(1)). \]
By $\beta<0$, we have ${\rm sign}(\mu_{1}(\theta))={\rm sign}(\beta)$ asymptotically. We now summarize these results.
By Theorem (ref), the magnitude of $|\beta|$ serves as the boundary (in rate) of phase transition. A slight violation may or may not be a huge problem depending on the order of magnitude of the violation $\eta$. If $|\beta|\ll\eta$, even imposing the extra condition of ${\rm sign}(\mu_{1}(\theta))={\rm sign}(\mu_{2}(\theta))$ does not help with learning the sign of $\mu_{1}$. In contrast, if $|\beta|\gg\eta$, we can easily learn ${\rm sign}(\mu_{1}(\theta))$ without assuming ${\rm sign}(\mu_{1}(\theta))={\rm sign}(\mu_{2}(\theta))$; the proof of the second part of Theorem (ref) does not rely on this condition.
Since $\Theta(\eta)$ allows for a large class of distributions, it is difficult to say much more than the rate. However, when the outcome variable is binary, we can precisely determine the boundary for the phase transition.
We define the counterparts of $\Theta(\eta)$ and $\Theta_{*}$ for the binary outcomes. Let
where $\varepsilon>0$ is a constant. The requirement that $P_{\theta}(Y_{i}=D_{i}\mid Z_{i})$ be bounded away from zero and one is mild in many applications.
When the outcome variable is not binary, we can dichotomize it to binary variables. We can define the new outcome variable $\tilde{Y}_{i}(1)=\mathbf{1}\{Y_{i}(1)\geq y\}$ and $\tilde{Y}_{i}(0)=\mathbf{1}\{Y_{i}(0)\geq y\}$, where $y$ is given. Then the treatment effect is how much the treatment changes the probability of $Y_{i}\geq y$.
For example, in angrist1998children, one outcome variable of interest $Y_{i}$ is the number of weeks a person worked in a year. We can set $y=1$ and ask how the treatment changes the probability of a person working for at least one week. Once we do this, we can estimate $|\beta|(k_{1}-k_{2})$, the boundary of phase transition. The results are in Table (ref). We see that any tolerance level of $c$ above $0.52\%$ can cause a serious problem for the question of whether or not the LATE is negative. Hence, even if the proportion of defiers is known to be at most 1%, it might not be obvious that we can safely conclude a negative LATE.
In many studies, the absence of defiers is justified by the one-sidedness of non-compliance. For example, $Z_{i}\in\{0,1\}$ is the randomly assigned treatment and $D_{i}$ is the actually treatment status. When the compliance is not perfect (i.e., $P(Z_{i}=D_{i})<1$), the non-compliance is often one-sided: $P(D_{i}=1\mid Z_{i}=1)<1$ but $P(D_{i}=1\mid Z_{i}=0)=0$. However, we discuss a small violation to this ideal case $P(D_{i}=1\mid Z_{i}=0)\approx0$, see JPTA in Table (ref) as an example.
Here, we provide a discussion for the case of $k_{2}=P(D_{i}=1\mid Z_{i}=0)\rightarrow0$ in the case of binary outcomes. This is different from Theorem (ref), which assumes that $k_{2}$ is bounded away from zero. Moreover, when $k_{2}\rightarrow0$, the natural choice of $\eta$ is $\eta=k_{2}$. To analyze this case, we consider the following parameter space \[ \Theta_{binary,*}=\biggl\{\theta=(a,b,c,H)\in\Theta_{binary}(k_{2}):\ P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)<|\beta|(k_{1}-k_{2})\biggr\}, \] where $\Theta_{binary}(\cdot)$ is defined in ((ref)).
We notice that $P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)<|\beta|(k_{1}-k_{2})$ (the boundary in Theorem (ref)) is not the same as applying $\eta=k_{2}$ to Theorem (ref). Applying $\eta=k_{2}$ to $\eta<|\beta|(k_{1}-k_{2})$ in Theorem (ref) leads to $k_{2}<|\beta|(k_{1}-k_{2})$. However, $P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)\leq k_{2}$. To see this, observe that $P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)=E_{\theta}(Y_{i}D_{i}\mid Z_{i}=0)\leq E_{\theta}(D_{i}\mid Z_{i}=0)=k_{2}$. What this means in practice is that Theorem (ref) makes it easier to be on the “nice” side of the boundary; instead of requiring $|\beta|(k_{1}-k_{2})$ to be above $k_{2}$, we require it to be above $P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)$.
A more important implication of Theorem (ref) is that it is possible to check whether or not we are on the nice side of the phase-transition boundary. The set $\Theta_{binary,*}$ is defined by the testable condition \[ P_{\theta}(Y_{i}=D_{i}=1\mid Z_{i}=0)<|\beta|(k_{1}-k_{2}). \]
We apply this to the JPTA study. We set the outcome variable to be $\mathbf{1}\{\text{income}<\$50000\}$. This means that we study the effect of treatment on the probability of earning less than \$50000.\footnote{We choose “less than” instead of “more than” to get a negative $\beta$. The interpretation is intuitively the same. Receiving treatment makes it less likely to earn less than \$50000 so it makes it more likely to earn at least \$50000.} The results are in Table (ref). Based on only the point estimates, the data-generating process is in $\Theta_{binary,*}$, which is the “nice” side of the phase-transition boundary.
There are already alternatives to the monotonicity condition in the literature. We add the following discussions.
We now outline a lower bound for the magnitude of LATE.
Notice that $\gamma\gtrsim|cov(D_{i},Z_{i})|$. Therefore, in the case of strong instruments, the lower bound $|\beta|\cdot\gamma$ is not too small compared to $|\beta|$. From the proof, we can see that the lower bound is also tight in that the equality can hold (because the minimum in the proof can be achieved).
It is worth noting that the lower bound in Theorem (ref) can be related to intent-to-treat (ITT) effects. We notice that \[ |\beta|\cdot\gamma=\frac{\left|E(Y_{i}\mid Z_{i}=1)-E(Y_{i}\mid Z_{i}=0)\right|}{E(D_{i}\mid Z_{i}=1)+E(D_{i}\mid Z_{i}=0)}=\frac{|ITT|}{E(D_{i}\mid Z_{i}=1)+E(D_{i}\mid Z_{i}=0)}. \]
Therefore, the lower bound satisfies $|\beta|\cdot\gamma\geq|ITT|/2$. Therefore, whenever we find that $\beta\neq0$ or $ITT\neq0$, it means that the treatment effect is not zero and we can use $|\beta|\cdot\gamma$ as a lower bound.
Learning the sign of LATE requires extra restrictions. This is inevitable; otherwise, monotonicity would be testable in Section (ref). It turns out that simple restrictions such as $|\mu_{1}|\geq|\mu_{2}|$ would suffice.
We should notice that Theorem (ref) imposes more than $|\mu_{1}|\geq|\mu_{2}|$. The assumption of $cov(D_{i},Z_{i})>0$ is not without loss of generality since the condition of $|\mu_{2}|\geq|\mu_{1}|$ is not enough to identify the sign of LATE. Therefore, the result should be viewed in the context of the empirical application.