EconBase
← Back to paper

Robust Identification in Randomized Experiments with Noncompliance

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

77,231 characters · 13 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust Identification in Randomized Experiments with Noncompliance

\onehalfspacing

abstractInstrument variable (IV) methods are widely used in empirical research to identify causal effects of a policy. In the local average treatment effect (LATE) framework, the IV estimand identifies the LATE under three main assumptions: random assignment, exclusion restriction, and monotonicity. However, these assumptions are often questionable in many applications, leading some researchers to doubt the causal interpretation of the IV estimand. This paper considers a robust identification of causal parameters in a randomized experiment setting with noncompliance where the standard LATE assumptions could be violated. We discuss identification under two sets of weaker assumptions: random assignment and exclusion restriction (without monotonicity), and random assignment and monotonicity (without exclusion restriction). We derive sharp bounds on some causal parameters under these two sets of relaxed LATE assumptions. Finally, we apply our method to revisit the random information experiment conducted in bursztyn2020misperceived and find that the standard LATE assumptions are jointly incompatible in this application. We then estimate the robust identified sets under the two sets of relaxed assumptions.

{ Keywords: Instrumental variable, monotonicity, exclusion restriction, local average controlled direct effect.

JEL subject classification: C14, C31, C35, C36.}

Introduction

Researchers are often concerned about the credibility of their assumptions and the conclusions they reach manski2003partial, manski2011policy. When the stated assumptions are refuted by the data, analysts usually resort to weaker versions of the assumptions. There may exist multiple ways of relaxing a set of assumptions when they are rejected by the data. Some researchers may choose to relax the assumptions in a continuous way Masten2021SalvagingModels, while others may consider discrete relaxations by dropping some of the assumptions li2024discordant. In this paper, we consider a robust identification of causal parameters in a randomized experiment setting with noncompliance where the standard local average treatment effect (LATE) assumptions could be violated. We follow li2024discordant's (li2024discordant) approach to propose a misspecification robust bound for a vector-valued parameter of various causal parameters.

Three main assumptions (random assignment, exclusion restriction, and monotonicity) are commonly used in the LATE framework. When these assumptions are jointly rejected, researchers could choose to relax some of the assumptions if they have some additional information on which assumptions may be causing the rejection. For example, Kedagni2023IdentifyingTypes proposes an identification strategy that allows for violations of the random assignment assumption, deChaisemartin2017ToleratingMonotonicity, Huber2017SharpNoncompliance, small2017instrumental, Noack2021SensitivityAssumption, dahl2023never, among others have proposed different ways of relaxing the monotonicity assumption, while cai2008bounds and flores2013partial allow for violations of the exclusion restriction assumption. But these relaxations may appear arbitrary in some circumstances unless the analyst has an unambiguous reason to convince the scientific community that she knows exactly which assumptions are the source of the rejection in the data.

The current paper aims at developing a unifying approach that would provide the maximal combination of the three assumptions that are compatible with the data. We consider a randomized experiment setting where the random assignment assumption is maintained. We discuss identification under two sets of weaker assumptions: random assignment and exclusion restriction (without monotonicity), and random assignment and monotonicity (without exclusion restriction).

We first introduce two causal parameters: the local average treatment-controlled direct effect (LATCDE), and the local average instrument-controlled direct effect (LAICDE). We define a twenty-dimensional real-valued vector parameter of interest that contains the eight LATCDEs, the eight LAICDEs, and the four types' probabilities. Second, we derive sharp bounds for the twenty parameters under respectively (i) random assignment, exclusion restriction, and monotonicity; (ii) random assignment, and exclusion restriction; and (iii) random assignment, and monotonicity. Specifically, under the random assignment and monotonicity assumptions, we derive sharp bounds on the local average treatment-controlled direct effects for the always-takers and never-takers, respectively, and the total average controlled direct effect for the compliers. Additionally, we show that the intent-to-treat effect can be expressed as a convex weighted average of these three effects. Third, we propose a misspecification robust bound for the vector parameter of interest under the three possible sets of assumptions. Finally, we apply our method to analyze the randomized information experiment conducted by bursztyn2020misperceived. We find that the LATE assumptions are jointly rejected in this setting. However, when either the monotonicity or exclusion restriction is dropped, the remaining two assumptions cannot be rejected. We provide estimates of identified sets for our parameters of interest, which are robust to violations of monotonicity or exclusion assumptions.

The remainder of the paper is organized as follows. Section (ref) presents the model, the assumptions, and the parameters of interest. In Sections (ref), (ref), and (ref), we derive bounds on the parameters under respectively (i) random assignment, exclusion restriction, and monotonicity; (ii) random assignment, and exclusion restriction; and (iii) random assignment, and monotonicity. Section (ref) presents the misspecification robust bound, Section (ref) shows the empirical illustration, and finally Section (ref) concludes. The proofs of the main results are relagated to the appendix from the supplementary material.

Analytical framework

We consider a potential outcome model with a binary treatment and a binary instrumental variable. Let $D \in \{0, 1\}$ be the observed treatment, where $D = 1$ indicates receiving the treatment and $D = 0$ not receiving it. Let $Y$ denote the observed outcome, and $Z$ denote a binary instrumental variable, which takes values in $\mathcal{Z} = \{0, 1\}$. Following the convention, we use $D_0$ and $D_1$ to represent the potential treatment status, where $D_z$ denotes the potential treatment if the instrument $Z$ is externally set to $z$. Also, $Y_{dz}$ represents the potential outcome if the treatment and instrument are externally set to $d$ and $z$, respectively. Then, we have the following potential outcome model,

equation[equation omitted — 228 chars of source]

where $(Y, D, Z)$ is the observed data, $\left(Y_{00}, Y_{01}, Y_{10}, Y_{11}, D_0, D_1\right)$ is a vector of latent variables.

The vector $\left(D_0, D_1\right)$ represents the four unobserved groups commonly known in the treatment effect literature as the types in the language of Angrist1996IdentificationVariables. We have $$

array[array omitted — 234 chars of source]

$$ Let $T$ denote the random type of an individual with support $\{a, c, n, d f\}$. Let $p_t$ denote the proportion of type $t$ in the population. Thus, we have the corresponding probabilities for these four types (strata/groups) of the population: $p_n$, $p_{df}$, $p_c$, and $p_a$. We define our parameters of interest:

eqnarray*[eqnarray* omitted — 274 chars of source]

for $z \in\{0,1\}$, $d \in\{0,1\},$ and $t\in\{a,c,df,n\}$. We have eight $\theta$-parameters, and eight $\delta$-parameters. Combined with the four types' probabilities $p_t$, we define a $20 \times 1$ column vector $\Gamma$ of our parameters of interest, $$\Gamma\equiv [\theta_{0a}, \theta_{1a}, \theta_{0c}, \theta_{1c}, \theta_{0df}, \theta_{1df}, \theta_{0n}, \theta_{1n}, \delta_{0a}, \delta_{1a}, \delta_{0c}, \delta_{1c}, \delta_{0df}, \delta_{1df}, \delta_{0n}, \delta_{1n}, p_a, p_c, p_{df}, p_n]'.$$ Elements of the vector $\Gamma$ can be viewed as building block parameters for many commonly used causal parameters. For instance, we can define some commonly used parameters as functionals of $\Gamma$:

eqnarray*[eqnarray* omitted — 379 chars of source]

In this paper, our goal is to provide a robust identified set for these parameters under the following set of assumptions considered in the original work by imbens1994identification.

assumption[Random assignment, RA] The instrument variable $Z$ is independent of potential outcomes and potential treatments: \begin{equation*} Z \perp \!\!\! \perp (Y_{11}, Y_{10}, Y_{01}, Y_{00}, D_1, D_0). \end{equation*}

Assumption (ref) requires that the instrument $Z$ be independent of all potential outcomes and potential treatments. This assumption will likely hold in randomized experiments. We may assume that it holds conditional on covariates in stratified randomized experiments. This assumption is maintained throughout the paper.

assumption[Exclusion restriction, ER] $Y_{dz}=Y_{dz'}=Y_d$ for all $d$, $z$, and $z'$.

Assumption (ref) imposes that the instrument $Z$ does not have a direct effect on the potential outcome. It is allowed to affect the outcome $Y$ only through the treatment $D$. We allow for potential violations of this assumption in this paper.

assumption[Monotonicity, MON] Either $D_1 \geq D_0$ or $D_0 \geq D_1$.

Assumption (ref), also often referred to as uniformity (see heckman2018unordered, heckman2018unordered), rules out the existence of both compliers and defiers in the population. This assumption is questionable in many empirical settings, especially the judge IV design literature. To illustrate, suppose there are two judges who are randomly assigned to cases/defendants. A jury will make a decision on whether or not to convict a defendant. A researcher is interested in studying the effect of the jury's decision $D$ on an outcome $Y$ (recidivism/employment). The judges assignment variable $Z$ can be considered as an IV. There is no a priori reason to assume monotonicity in the jury's decision with respect to the judges. Hence, all four types are likely present in this setting. Therefore, we allow for complete violations of this monotonicity assumption in this paper.

Fiorini2021ScrutinizingDesigns extensively discussed situations in which the monotonicity assumption might be questionable. For illustration purposes, we consider the measurement of the returns to schooling as our motivating example.

example[Returns to schooling] Suppose that the college decision is based on whether the opportunity cost $V_1$ is less than a threshold $Q_1(Z)$ and the cost of attending college $V_2$ (e.g., psychological cost of effort, tuition) is less than a different threshold $Q_2(Z)$, where $Z$ denotes an IV, say college proximity. This two-dimensional decision model is called the double hurdle model in Lee2018IdentifyingTreatments. Let $Y$ be the outcome of interest (wage), $D$ be the treatment variable (college attendance), and $Z$ be the instrument (college proximity) so that the college choice model is given as below: \begin{equation} D= \mathbbm{1}\left\{V_1 < Q_1(Z), V_2 < Q_2(Z) \right \}, \end{equation} In this framework, the monotonicity assumption will likely fail if for example the instrument lowers $Q_1$ while at the same time it increases $Q_2$, that is, $Q_1(0) < Q_1(1)$, and $Q_2(1) < Q_2(0)$. As a consequence, we will have all four types of individuals in the population: compliers $(c)$, defiers $(df)$, never-takers $(n)$, and always-takers $(a)$ as illustrated in Figure (ref).\footnote{Definition of the four types in the population: $c=\{V_1 < Q_1(1),V_2 < Q_2(1)\}\cap\{V_1 \geq Q_1(0)\cup V_2 \geq Q_2(0)\},$ $df=\{V_1 < Q_1(0),V_2 < Q_2(0)\}\cap\{V_1 \geq Q_1(1)\cup V_2 \geq Q_2(1)\},$ $a=\{V_1 < Q_1(0),V_2 < Q_2(0)\}\cap\{V_1 < Q_1(1), V_2 < Q_2(1)\},$ $n=\{V_1 \geq Q_1(0)\cup V_2 \geq Q_2(0)\}\cap\{V_1 \geq Q_1(1)\cup V_2 \geq Q_2(1)\}$.} \begin{figure} \begin{tikzpicture}[x=0.75pt,y=0.75pt,yscale=-1,xscale=1] \draw (172,267.25) -- (474.5,267.25)(202.25,22.45) -- (202.25,294.45) (467.5,262.25) -- (474.5,267.25) -- (467.5,272.25) (197.25,29.45) -- (202.25,22.45) -- (207.25,29.45) ; \draw (202,61) -- (427.5,60.45) ; \draw (427.5,60.45) -- (426.5,267.45) ; \draw [line width=2.25] [dash pattern={on 2.53pt off 3.02pt}] (203,104) -- (253.5,103.45) ; \draw [line width=2.25] [dash pattern={on 2.53pt off 3.02pt}] (253.5,103.45) -- (252.5,266.45) ; \draw [line width=2.25] [dash pattern={on 2.53pt off 3.02pt}] (202.5,217.45) -- (364.5,217.45) ; \draw [line width=2.25] [dash pattern={on 2.53pt off 3.02pt}] (364.5,217.45) -- (364.5,267.45) ; \draw (164,28) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle V_{2}$}; \draw (486,272) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle V_{1}$}; \draw (324,141) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle n$}; \draw (217,150) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle \textcolor[rgb]{0.82,0.01,0.11}{df}$}; \draw (303,232) node [anchor=north west][inner sep=0.75pt] [align=left] {\textcolor[rgb]{0.29,0.56,0.89}{$\displaystyle c$}}; \draw (224,232) node [anchor=north west][inner sep=0.75pt] [align=left] {\textcolor[rgb]{0.49,0.83,0.13}{$\displaystyle a$}}; \draw (187,272) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle 0$}; \draw (181,51) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle 1$}; \draw (422,278) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle 1$}; \draw (156,96) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle Q_{2}( 0)$}; \draw (157,210) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle Q_{2}( 1)$}; \draw (228,276) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle Q_{1}( 0)$}; \draw (343,276) node [anchor=north west][inner sep=0.75pt] [align=left] {$\displaystyle Q_{1}( 1)$}; \end{tikzpicture} \caption{Visualization of types (double hurdle model)} \end{figure}

In addition, growing up near a college could have a direct effect on someone's ability or skills necessary to complete some tasks that they would not have done otherwise. These skills would positively affect people's wages, which would make the college proximity instrument violate the exclusion restriction assumption.

We consider three different combinations (or menus) of assumptions.

eqnarray*[eqnarray* omitted — 107 chars of source]

We denote $A$ the set of all three combinations of assumptions: $A\equiv \{A_1,A_2,A_3\}.$ We assume that we are in a randomized experiment setting where Assumption (ref) holds. For this reason, we do not consider the menu $\{ER,MON\}$ as an option. In general, this combination alone would not yield informative bounds without additional assumptions Kedagni2023IdentifyingTypes.

Identification under random assignment, monotonicity, and exclusion restriction

Let $\mathcal{Y}$ denote the support of the outcome $Y$. We first note that under Assumption (ref), $Y_{dz}=Y_{d}$ for all $d$ and $z$. Under Assumptions (ref) and (ref), we have for any Borel set $A\subseteq \mathcal{Y}$,

eqnarray[eqnarray omitted — 580 chars of source]

\equationautorefname (ref) shows that the distribution of $Y$ in the treatment group for individuals assigned to this group is a mixture of the distribution of $Y_1$ for the compliers and the always-takers. And \equationautorefname (ref) shows that the distribution of $Y$ in the treatment group for individuals assigned to the control group is a mixture of the distribution of $Y_1$ for the defiers and the always-takers. A similar decomposition holds for the control group, as shown in Equations (ref)-(ref).

Let us take the difference between Equations (ref) and (ref). We have

eqnarray[eqnarray omitted — 243 chars of source]

Similarly, the following holds from Equations (ref) and (ref).

eqnarray[eqnarray omitted — 245 chars of source]

For $A=\mathcal{Y}$, Equation (ref) implies

equation[equation omitted — 104 chars of source]

Under Assumption (ref), either $p_{df}=0$ or $p_c=0$. Without loss of generality, assume that $p_{df}=0$. Then, Equation (ref) implies that $p_c=\mathbb{P}(D=1 \vert Z=1)-\mathbb{P}(D=1 \vert Z=0)$. As a result, $p_a$ and $p_n$ are also identified: $p_a=\mathbb P(D=1\vert Z=0)$, and $p_n=\mathbb P(D=0\vert Z=1)$. Under Assumption (ref), $\delta_{dt}=0$ for all $d$ and $t$, and $\theta_{0t}=\theta_{1t}$ for all $t$. Under Assumptions (ref)-(ref), $\theta_{zc}$ is identified if $\mathbb{P}(D=1 \vert Z=1)-\mathbb{P}(D=1 \vert Z=0)>0$ imbens1994identification: $$\theta_{1c}=\theta_{0c}=\frac{\mathbb E[Y\vert Z=1]-\mathbb E[Y\vert Z=0]}{\mathbb E[D\vert Z=1]-\mathbb E[D\vert Z=0]},$$ and from Equations (ref)-(ref), the following testable implications must hold:

eqnarray[eqnarray omitted — 227 chars of source]

for all Borel set $A$ Balke1997BoundsCompliance, Heckman2005StructuralEvaluation, Kitagawa2015AValidity, Mourifie2017TestingAssumptions. The remaining $\theta$-parameters are not identified and lie in the real line $\mathbb R$. The identified set for $\Gamma$ is

eqnarray*[eqnarray* omitted — 375 chars of source]

where

eqnarray*[eqnarray* omitted — 965 chars of source]

Let $\tilde{Z}\equiv 1-Z$, and let $\tilde{T}$ and $\tilde{\Gamma}$ be respectively the corresponding $T$ and $\Gamma$ defined based on the instrument $\tilde{Z}$. We have $\tilde{a}=n$, $\tilde{c}=df$, $\tilde{df}=c$, $\tilde{n}=a$, $\tilde{\theta}_{\tilde{z}\tilde{a}}=\theta_{(1-z)n}$, $\tilde{\theta}_{\tilde{z}\tilde{n}}=\theta_{(1-z)a}$, $\tilde{\theta}_{\tilde{z}\tilde{c}}=\theta_{(1-z)df}$, $\tilde{\theta}_{\tilde{z}\tilde{df}}=\theta_{(1-z)c}$, $\tilde{\delta}_{d\tilde{a}}=\delta_{dn}$, $\tilde{\delta}_{d\tilde{n}}=\delta_{da}$, $\tilde{\delta}_{d\tilde{c}}=\delta_{d df}$, and $\tilde{\delta}_{d\tilde{df}}=\delta_{dc}$. Then $\Theta_I^2(A_1)=\tilde{\Theta}_I^1(A_1),$ where $\tilde{\Theta}_I^1(A_1)$ is the corresponding $\Theta_I^1(A_1)$ defined for $\tilde{\Gamma}$ based on $(Y,D,\tilde{Z})$ and $\tilde{T}$.

If $\mathbb E[D\mid Z=1]-\mathbb E[D\mid Z=0]=0$, then $p_c-p_{df}=0,$ and under Assumption (ref), $p_c=p_{df}=0$. This implies $p_a=\mathbb P(D=1\mid Z=1)=\mathbb P(D=1\mid Z=0)=\mathbb P(D=1),$ and $p_n=\mathbb P(D=0)$. Furthermore, inequalities (ref)-(ref) must hold with equality, i.e., $Z \perp \!\!\! \perp (Y,D)$. Hence,

eqnarray*[eqnarray* omitted — 824 chars of source]

Identification under random assignment and exclusion restriction

Building on the above equations (ref)-(ref), we can derive bounds on the probability of being a defier as follows:

align*[align* omitted — 249 chars of source]

Kedagni2020GeneralizedAssumption and Kitagawa2021TheIndependence show that the following restriction is a sharp testable implication of Assumptions (ref) and (ref):

eqnarray[eqnarray omitted — 125 chars of source]

In the case when the instrument $Z$ is binary, we can prove that the violation of testable implication in Equation (ref) is equivalent to the identified bounds for the proportion of defiers being empty. Therefore, the following proposition holds.

proposition[Sharp bounds for $p_{df}$] Under Assumptions (ref) and (ref), the identified set $\Theta_{I}(p_{df})$ for the probability of defiers is given by \begin{equation*} \begin{aligned} \Theta_{I}(p_{df}) = & \bigg[\max _{s \in\{0,1\}}\left\{\sup _A\{\mathbb{P}(Y \in A, D=s \vert Z=1-s)-\mathbb{P}(Y \in A, D=s \vert Z=s)\}\right\}, \\ & \qquad \qquad \qquad \min \{\mathbb{E}[D \vert Z=0], \mathbb{E}[1-D \vert Z=1]\}\bigg]. \end{aligned} \end{equation*} Moreover, the identified set $\Theta_{I}(p_{df})$ is empty if and only if inequality (ref) is violated.
proofThe detailed proof is provided in Appendix (ref).

The bounds in $\Theta_I(p_{df})$ were first derived by richardson2010analysis for binary outcomes and recently generalized to any outcomes by Noack2021SensitivityAssumption. This is an intermediate result for our main results in this paper. If the lower bound of $\Theta_I(p_{df})$ is greater than zero, the standard LATE assumptions are jointly rejected (or incompatible), aligning with Balke1997BoundsCompliance, imbens1997estimating, Heckman2005StructuralEvaluation, Kitagawa2015AValidity, and Mourifie2017TestingAssumptions. If instead inequalities (ref)-(ref) hold, then the testable implications developed for the LATE assumptions hold, and inequality (ref) becomes redundant, since the LATE testable implications are sharp Kitagawa2015AValidity, Mourifie2017TestingAssumptions. In such a context, the lower bound of $\Theta_I(p_{df})$ is 0, as the supremum in the lower bound is achieved when $A=\emptyset$. Huber2017SharpNoncompliance derive bounds on $p_{df}$ under a weaker version of Assumption (ref), $\mathbb E[Y_{dz}\mid T, Z]=E[Y_{dz}\mid T]$ and $Z \perp \!\!\! \perp T$, which implies the Manski1990NonparametricEffects mean independence assumption between the potential outcomes and the instrument, $\mathbb E[Y_{dz} \mid Z]=\mathbb E[Y_{dz}]$. This latter mean independence assumption together with the exclusion restriction assumption (ref) imply similar conditions to (ref): $\sup_z \mathbb E[YD+(1-D)\inf \mathcal Y \mid Z=z] \leq \inf_z \mathbb E[YD+(1-D)\sup \mathcal Y \mid Z=z]$ and $\sup_z \mathbb E[Y(1-D)+D\inf \mathcal Y\mid Z=z] \leq \inf_z \mathbb E[Y(1-D)+D\sup \mathcal Y\mid Z=z]$. These conditions could be rejected if the support $\mathcal Y$ is bounded. For example, when the outcome is binary (i.e., $\mathcal Y=\{0,1\}$), the testable implication for mean independence and Assumption (ref) is identical to inequality (ref).

Kitagawa2021TheIndependence focuses on the identified region of marginal distributions of potential outcomes, $Y_1$ and $Y_0$, under Assumptions (ref) and (ref), in the same context of a binary treatment and a binary instrument. This paper has a result that the identified region of marginal distributions of potential outcomes is non-empty if and only if inequality (ref) holds. We focus on the identified set of type proportions and local average treatment effects here and also prove in Proposition (ref) that the identified set of the defier proportion is non-empty if and only if inequality (ref) is satisfied. Both of these results indicate that the non-emptyness of two different parameters (marginal potential outcome distributions and defiers probability) under Assumptions (ref) and (ref) leads to the same testable inequality (ref).

kwon2024testing consider a more general framework with a binary instrument and a multivalued discrete treatment. They derive bounds on the proportion of always-takers. Their bounds on the fraction of always-takers can be used to derive bounds on the defiers probability in our framework, since we can express the defiers probability as a function of the always-takers probability.

While our ultimate goal is to derive the identified set for $\Gamma$ under the bundle of assumptions $A_2$, we are first to going derive bounds on the LATE for compliers and defiers, respectively, as these are causal parameters that receive a particular attention in the causal inference literature. Before we move on, we introduce some additional notation. Denote $\mu_{d t} \equiv \mathbb{E}\left[Y_d \vert T=t\right], d \in\{0,1\}$ and $t \in$ $\{a, c, d f, n\}$. Under Assumptions (ref) and (ref), We have $p_a=\mathbb{E}[D \vert Z=0]-p_{d f}$, and $p_n= \mathbb{E}[1-D \vert Z=1]-p_{d f}.$ Lastly, we have $p_c=1-p_{df}-p_n-p_a=\mathbb{E}[D \vert Z=1] - \mathbb{E}[D \vert Z=0] +p_{df}$, where $p_{df} \in \Theta_I(p_{df}).$

To proceed, we are going to derive sharp bounds for $\mathbb P(Y_1 \in A \vert T=a)$ using Equations (ref)-(ref). For simplicity, suppose $p_{df}$ is an interior point of $\Theta_I(p_{df})$ such that $0 < p_{df} < \min\{\mathbb E[D\vert Z=0], \mathbb E[1-D \vert Z=1]\}$. Equation (ref) implies

eqnarray*[eqnarray* omitted — 137 chars of source]

Since $\mathbb P(Y_1 \in A \vert T=c) \in [0,1],$ and $\mathbb P(Y_1 \in A \vert T=a) \in [0,1]$, we have

eqnarray*[eqnarray* omitted — 201 chars of source]

Similarly, Equation (ref) implies

eqnarray*[eqnarray* omitted — 204 chars of source]

Hence, Equations (ref)-(ref) together imply

eqnarray[eqnarray omitted — 370 chars of source]

A similar reasoning holds for (ref)-(ref) and helps partially identify $\mathbb P(Y_0 \in A \vert T=n)$. The identified sets for $F_{dt}\equiv F_{Y_d\vert T=t}$, $d\in\{0,1\}$, $t\in \{c,df\}$ are given in Proposition (ref).

propositionFor a given $p_{df}$ interior point of $\Theta_I(p_{df})$, pointwise sharp bounds for the distributions $F_{dt}$ are given below: \begin{eqnarray*} F_{1a}^{LB}(y)&\equiv& \max\left\{F_{1a}^{LB_1}(y),F_{1a}^{LB_0}(y)\right\} \leq F_{1a}(y) \leq \min\left\{F_{1a}^{UB_1}(y),F_{1a}^{UB_0}(y)\right\}\equiv F_{1a}^{UB}(y),\\ F_{0n}^{LB}(y) &\equiv& \max\left\{F_{0n}^{LB_1}(y),F_{0n}^{LB_0}(y)\right\} \leq F_{0n}(y) \leq \min\left\{F_{0n}^{UB_1}(y),F_{0n}^{UB_0}(y)\right\}\equiv F_{0n}^{UB}(y),\\ F_{1c}(y) &=& \frac{\mathbb P(Y \leq y, D=1 \vert Z=1)-p_a F_{1a}(y)}{p_c },\ F_{0c}(y) = \frac{\mathbb P(Y \leq y, D=0 \vert Z=0)-p_n F_{0n}(y)}{p_c },\\ F_{1df}(y) &=& \frac{\mathbb P(Y \leq y, D=1 \vert Z=0)-p_a F_{1a}(y)}{p_{df} },\ F_{0df}(y) = \frac{\mathbb P(Y \leq y, D=0 \vert Z=1)-p_n F_{0n}(y)}{p_{df}}, \end{eqnarray*} where $F_{1a}^{\ell}(y)$, $F_{0n}^{\ell}(y)$, $\ell \in \{LB_0, LB_1, UB_0, UB_1\}$ are defined in Appendix (ref), and $p_a=\mathbb{E}[D \vert Z=0]-p_{d f}$, $p_n= \mathbb{E}[1-D \vert Z=1]-p_{d f},$ $p_c=1-p_{df}-p_n-p_a=\mathbb{E}[D \vert Z=1] - \mathbb{E}[D \vert Z=0] +p_{df}$.

While Proposition (ref) shows pointwise sharp bounds for the distributions, we characterize their functional sharp bounds (in the terminology of mourifie2020sharp) in Appendix (ref). In the following, we explain the intuition behind the reason why the bounds for $F_{1a}(y)$ are not functionally sharp. Suppose that we are interested in bounding $F_{1a}(y')-F_{1a}(y)$ where $y < y'$. One can take the difference of the bounds in Proposition (ref) and obtain $$F_{1a}^{LB}(y')-F_{1a}^{UB}(y) \leq F_{1a}(y')-F_{1a}(y) \leq F_{1a}^{UB}(y')-F_{1a}^{LB}(y).$$ But this approach will not lead to sharp bounds for $F_{1a}(y')-F_{1a}(y)$. Instead, we are going to replace the Borel set $A$ by $(y,y']$ in Equation (ref). Doing this will yield tighter bounds than the pointwise difference of the bounds in Proposition (ref). Pointwise sharp bounds for $F_{1a}(y)$ yield sharp bounds on location parameters such as the mean and the quantiles, but they will not deliver sharp bounds on spread parameters like the interquantile range. Functional sharp bounds for $F_{1a}(y)$ yield sharp bounds for interquantile range parameters.

From the results in Proposition (ref), we derive bounds on the mean potential outcomes for types. Let $\mu_F$ denote the expected value of a given cumulative distribution function (cdf) $F$, and $\mu_{dt}\equiv \mathbb E[Y_d \vert T=t]$. Corollary (ref) provides sharp bounds on $\mu_{dt}$.

corollaryFor a given $p_{df}$ interior point of $\Theta_I(p_{df})$, sharp bounds on $\mu_{1a}$ and $\mu_{0n}$ are given by: \begin{eqnarray*} && \mu_{F_{1a}^{UB}}(p_a) \leq \mu_{1a}(p_a) \leq \mu_{F_{1a}^{LB}}(p_a),\\ && \mu_{F_{0n}^{UB}}(p_n) \leq \mu_{0n}(p_n) \leq \mu_{F_{0n}^{LB}}(p_n), \end{eqnarray*} where $p_a=\mathbb{E}[D \vert Z=0]-p_{d f}$, $p_n= \mathbb{E}[1-D \vert Z=1]-p_{d f}.$

The sharp bounds in Corollary (ref) can be difficult to compute in practice. Tractable valid outer sets are then proposed:

eqnarray*[eqnarray* omitted — 309 chars of source]
remark$\max\left\{\mu_{F_{1a}^{UB_1}}, \mu_{F_{1a}^{UB_0}}\right\}=\mu_{F_{1a}^{UB}}$ if $F_{1a}^{UB_1}$ first-order stochastically dominates $F_{1a}^{UB_0}$ or vice versa. Similarly, $\min\left\{\mu_{F_{1a}^{LB_1}}, \mu_{F_{1a}^{LB_0}}\right\}=\mu_{F_{1a}^{LB}}$ if $F_{1a}^{LB_1}$ first-order stochastically dominates $F_{1a}^{LB_0}$ or vice versa. However, in general, $\max\left\{\mu_{F_{1a}^{UB_1}}, \mu_{F_{1a}^{UB_0}}\right\} \leq \mu_{F_{1a}^{UB}}$, and $\min\left\{\mu_{F_{1a}^{LB_1}}, \mu_{F_{1a}^{LB_0}}\right\}\geq \mu_{F_{1a}^{LB}}$.

The bounds in Corollary (ref) provides sharp bounds on $\mu_{1a}(p_a)$ and $\mu_{0n}(p_n)$ for discrete, continuous, or mixed outcomes. Closed-form expressions for the bounds can be obtained for continuous outcomes with strictly increasing cdfs. These are called Lee2009TrainingEffects's (Lee2009TrainingEffects) bounds. Lemma (ref) displays the bounds.

lemmaThe following holds for continuous outcomes with strictly increasing cdfs. \begin{eqnarray*} \mu_{F_{1a}^{LB_z}} &=& \mathbb E\left[Y\vert D=1, Z=z, Y> F^{-1}_{Y\vert D=1, Z=z}\left(1-\frac{p_a}{\mathbb E[D\vert Z=z]}\right)\right],\\ \mu_{F_{1a}^{UB_z}} &=& \mathbb E\left[Y\vert D=1, Z=z, Y< F^{-1}_{Y\vert D=1, Z=z}\left(\frac{p_a}{\mathbb E[D\vert Z=z]}\right)\right],\\ \mu_{F_{0n}^{LB_z}} &=& \mathbb E\left[Y\vert D=0, Z=z, Y> F^{-1}_{Y\vert D=0, Z=z}\left(1-\frac{p_{n}}{\mathbb E[1-D\vert Z=z]}\right)\right],\\ \mu_{F_{0n}^{UB_z}} &=& \mathbb E\left[Y\vert D=0, Z=z, Y< F^{-1}_{Y\vert D=0, Z=z}\left(\frac{p_{n}}{\mathbb E[1-D\vert Z=z]}\right)\right]. \end{eqnarray*}

Let $\mathring{\Theta}_I(p_{df})$ denote the interior of $\Theta_I(p_{df})$. The following theorem proposes sharp bounds on the LATEs for compliers $\theta_{0c}=\theta_{1c}$ and defiers $\theta_{0df}=\theta_{1df}$.

theoremSuppose inequality (ref) holds, and $p_{df}$ is an interior point of $\Theta_I(p_{df})$. Under Assumptions (ref) and (ref), sharp bounds for $\mu_{1 c}$ and $\mu_{1 d f}$ are as follows: \begin{eqnarray*} \begin{aligned} & \mu_{1 c}(p_a) = \frac{\mathbb{E}[Y D \vert Z=1]-p_a \mu_{1 a}(p_a)}{\mathbb{E}[D \vert Z=1]-p_a},\ \ \ \mu_{1 d f}(p_a) = \frac{\mathbb{E}[Y D \vert Z=0]-p_a \mu_{1 a}(p_a)}{\mathbb{E}[D \vert Z=0]-p_a}, \end{aligned} \end{eqnarray*} where $p_a=\mathbb{E}[D \vert Z=0]-p_{d f}$, $p_{df} \in \mathring{\Theta}_I(p_{df}),$ and $\mu_{1a}(p_a) \in \left[\mu_{F_{1a}^{UB}}(p_a),\mu_{F_{1a}^{LB}}(p_a)\right]$. Similarly, sharp bounds for $\mu_{0 c}$ and $\mu_{0 d f}$ are given by: \begin{eqnarray*} \begin{aligned} & \mu_{0 c}(p_n) = \frac{\mathbb{E}[Y(1-D) \vert Z=0]-p_n \mu_{0 n}(p_n)}{\mathbb{E}[1-D \vert Z=0]-p_n},\ \ \ \mu_{0 d f}(p_n) = \frac{\mathbb{E}[Y(1-D) \vert Z=1]-p_n \mu_{0 n}(p_n)}{\mathbb{E}[1-D \vert Z=1]-p_n}, \end{aligned} \end{eqnarray*} where $p_n=$ $\mathbb{E}[1-D \vert Z=1]-p_{d f},$ $p_{df} \in \mathring{\Theta}_I(p_{df})$, $\mu_{0n}(p_n)\in \left[\mu_{F_{0n}^{UB}}(p_{n}),\mu_{F_{0n}^{LB}}(p_{n})\right].$ Sharp bounds for $\theta_{0c}$, $\theta_{1c}$, $\theta_{0df}$, $\theta_{1df}$ are given by: \begin{eqnarray*} \theta_{0c}(p_c)=\theta_{1c}(p_c)&=& \mu_{1c}(p_c)-\mu_{0c}(p_c),\ \ \ \theta_{0df}(p_{df})=\theta_{1df}(p_{df})= \mu_{1df}(p_{df})-\mu_{0df}(p_{df}), \end{eqnarray*} where $p_c=1-p_{df}-p_n-p_a=\mathbb{E}[D \vert Z=1] - \mathbb{E}[D \vert Z=0] +p_{df}$, and $p_{df} \in \mathring{\Theta}_I(p_{df})$.
proofSee detailed proofs in Appendix (ref).

Theorem (ref) provides sharp bounds on the LATEs for compliers and defiers under random assignment and exclusion restriction, without requiring any kind of monotonicity assumption. The bounds differ from those in Huber2017SharpNoncompliance in two ways. First, they are derived under a stronger assumption (random assignment instead of mean independence assumption). Second, as we previously discuss, these bounds take into account the testable implication (ref), while the Huber2017SharpNoncompliance bounds do not take into account the mean independence version of this inequality. As explained in li2024discordant, these kinds of outer bounds could be misleading.

Proposition (ref) and Theorem (ref) differ from the results in Kitagawa2021TheIndependence. While Kitagawa2021TheIndependence studies the identified sets of potential outcome distributions and derives sharp bounds for the average treatment effect, our focus is on the identified set of the vector parameter $\Gamma$, as it is a building block for many commonly used parameters, including the average and local average treatment effects. The proposed approach in Theorem (ref) also differs from other existing papers deChaisemartin2017ToleratingMonotonicity, Noack2021SensitivityAssumption, dahl2023never as it does not assume any weaker version of the monotonicity assumption (ref).

The identified set for $\Gamma$ is

eqnarray*[eqnarray* omitted — 482 chars of source]

where

eqnarray*[eqnarray* omitted — 1,339 chars of source]

which arises when $p_{df}>0$ is in the interior of the identified set for the proportion of defiers, $\Theta_I(p_{df})$. When $p_{df}$ is equal to 0, the identified set for $\Gamma$ coincides with $\Theta_I(A_1)$, which corresponds to the setting where all assumptions in $A_1$ (RA, ER, MON) hold. When $p_{df}$ is equal to the upper bound of $\Theta_I(p_{df})$, the identified set for $\Gamma$ under $A_2$ is $\Theta_I^3(A_2)$, which is formally defined in Appendix (ref).

Numerical Illustration

In this section, we illustrate how informative our proposed bounds can be in the context of a double hurdle model for a given data generating process. This is an example where our bounds correctly identify the sign of the LATEs for compliers and defiers separately while the IV estimand yields the wrong sign. Consider the following data-generating process\footnote{Note that this a version of the double hurdle model considered in the introduction, since we can equivalently write $D$ as follows:\\ $D = \mathbbm{1}\{V_1\leq 2 Z, -V_2< -Z\}=\mathbbm{1}\{\Phi(V_1)\leq \Phi(2 Z), \Phi(-V_2) < \Phi(-Z)\}=\mathbbm{1}\{\tilde{V}_1\leq Q_1(Z), \tilde{V}_2 < Q_2(Z)\}$, where $\tilde{V_1}\equiv \Phi(V_1)$, $\tilde{V}_2\equiv \Phi(-V_2)$, $Q_1(Z)\equiv \Phi(2Z)$, $Q_2(Z) \equiv \Phi(-Z)$.}

equation*[equation* omitted — 157 chars of source]

where $\beta=5 \Phi(2V_1+V_2), U=\frac{1}{2} (V_1+V_2), (V_1, V_2, \varepsilon)^{\prime} \sim N(0, I)$, and $\Phi(\cdot)$ is the standard normal cdf.

The potential treatments are then defined as

equation*[equation* omitted — 150 chars of source]

In Table (ref), we can compute the following some important parameters from this DGP.

table[table omitted — 768 chars of source]

In this DGP, the true causal effects for compliers and defiers are $L A T E_c \equiv$ $\mathbb{E}\left[Y_1-Y_0 \vert T=c\right]=4.93$ and $L A T E_{d f} \equiv \mathbb{E}\left[Y_1-Y_0 \vert T=d f\right]=1.23$. Our lower bounds for compliers and defiers from Figure (ref) are strictly larger than zero, showing our bound can clearly identify the positive sign and reasonable regions of LATE for two subgroups of people. However, when we add the IV estimand to Figure (ref), we can see that the IV estimand is worse than our bounds under this DGP in two dimensions: magnitude and sign. The IV estimand is lower than the lower bounds. Also, compared to the positive true LATE for defiers ($LATE_{df}= 1.23$), the sign of the IV estimand ($LATE_{IV}= -1.69$) is negative, which is misleading to policymakers. Indeed, when there exist both compliers and defiers in the population, the IV estimand is a weighted average of the LATEs for compliers and defiers, with negative weights.\footnote{IV estimand $=\frac{p_c LATE_c -p_{df} LATE_{df}}{p_c-p_{df}}$.} These negative weights lead to sign reversal of the IV estimand in this example.

figure[figure omitted — 172 chars of source]

Identification under random assignment and monotonicity

In this section, the potential outcome $Y_{dz}$ is allowed to vary with $z$ in such a way that we may have $\theta_{0t}\neq \theta_{1t}$ and $\delta_{dt}\neq 0$. This section demonstrates how researchers can use a randomly assigned instrument that may violate the exclusion assumption (ref). We maintain the monotonicity assumption (ref). To proceed, as in the previous section, we write the identified probabilities in each of the observed four subgroups $\{(D=d,Z=z): z, d \in \{0,1\}\}$ as a mixture of the potential outcome distributions for the unobserved types. More precisely, under Assumption (ref), Equations (ref)-(ref) become

eqnarray[eqnarray omitted — 597 chars of source]

Suppose first that $\mathbb{E}[D \vert Z=1] - \mathbb{E}[D \vert Z=0]> 0$. Then, under Assumption (ref), there are no defiers, i.e., $p_{df}=0$. As a result, $p_c=\mathbb{E}[D \vert Z=1] - \mathbb{E}[D \vert Z=0]$, $p_a=\mathbb E[D\vert Z=0]$, and $p_n=\mathbb E[1-D\vert Z=1]$. Equation (ref) implies

eqnarray*[eqnarray* omitted — 143 chars of source]

Since $\mathbb P(Y_{11} \in A \vert T=c) \in [0,1],$ and $\mathbb P(Y_{11} \in A \vert T=a) \in [0,1]$, we have

eqnarray*[eqnarray* omitted — 201 chars of source]

Equation (ref) implies $\mathbb P(Y_{10}\in A \vert T=a)=\mathbb P(Y \in A \vert D=1, Z=0)$, while Equation (ref) implies $\mathbb P(Y_{01}\in A \vert T=n)=\mathbb P(Y \in A \vert D=0, Z=1)$. Finally, Equation (ref) implies

eqnarray*[eqnarray* omitted — 201 chars of source]

and

eqnarray*[eqnarray* omitted — 143 chars of source]

The above bounding approach builds on Horowitz1995IdentificationData. We can then suitably take the expectations of the bounds and obtain the so-called Lee2009TrainingEffects bounds. Let $\mu_{dzt}\equiv \mathbb E[Y_{dz} \vert T=t]$, $F_{dzt}\equiv F_{Y_{dz}\vert T=t}$. The following proposition (ref) derives sharp bounds on the distributions $F_{dzt}$.

propositionSuppose Assumptions (ref) and (ref) hold, and $\mathbb{E}[D \vert Z=1] - \mathbb{E}[D \vert Z=0]> 0$. Then, \begin{eqnarray*} && p_c=\mathbb{E}[D \vert Z=1] - \mathbb{E}[D \vert Z=0],\ p_a=\mathbb E[D\vert Z=0],\ p_n=\mathbb E[1-D\vert Z=1], p_{df}=0,\\ && F_{11a}^{LB}(y)\equiv \max\left\{\frac{\mathbb P(Y\leq y, D=1 \vert Z=1)- p_c}{p_a},0\right\} \leq F_{11a}(y)\\ && \qquad \qquad \qquad \qquad \qquad \qquad \qquad \qquad \qquad \qquad \qquad \leq \min\left\{\frac{\mathbb P(Y\leq y, D=1 \vert Z=1)}{p_a},1\right\} \equiv F_{11a}^{UB}(y),\\ &&F_{00n}^{LB}(y) \equiv \max\left\{\frac{\mathbb P(Y\leq y, D=0 \vert Z=0)- p_c}{p_n},0\right\} \leq F_{00n}(y)\\ && \qquad \qquad \qquad \qquad \qquad \qquad \qquad \qquad \qquad \qquad \qquad \leq \min\left\{\frac{\mathbb P(Y\leq y, D=0 \vert Z=0)}{p_n},1\right\} \equiv F_{00n}^{UB}(y),\\ && F_{11c}(y) = \frac{\mathbb P(Y\leq y, D=1 \vert Z=1)- p_a F_{11a}(y)}{p_c},\\ && F_{00c}(y) = \frac{\mathbb P(Y\leq y, D=0 \vert Z=0)- p_n F_{00n}(y)}{p_c},\\ && F_{10a(y)}=\mathbb P(Y \leq y \vert D=1, Z=0),\ F_{01n(y)}=\mathbb P(Y \leq y \vert D=0, Z=1). \end{eqnarray*} These above bounds are pointwise sharp.

Proposition (ref) provides important results that help derive sharp bounds on the local average treatment-controlled direct effects for the always-takers and never-takers. A direct implication of this proposition is Corollary (ref) in Appendix (ref) , which derives sharp bounds on the parameters $\mu_{dzt}$. Building on the results in Corollary (ref), Corollary (ref) below provides sharp bounds on the parameters $\delta_{1a}$, $\delta_{0n}$, and $\theta_{1c}+\delta_{0c}=\theta_{0c}+\delta_{1c}$. To the best of our knowledge, this is the first paper to formally derive sharp bounds on these parameters under the random assignment and monotonicity assumptions. Note however that $\delta_{1a}$ and $\delta_{0n}$ coincide with the local net average treatment effect (LNATE) for the always-takers and never-takers, respectively, defined in flores2013partial. In general, $LNATE_t\equiv \mathbb E[Y_{D_0 1}-Y_{D_0 0}\vert T=t]$ is different from our LATCDE parameter $\delta_{dt}$.

corollarySuppose Assumptions (ref) and (ref) hold, and $\mathbb{E}[D \vert Z=1] - \mathbb{E}[D \vert Z=0]> 0$. Then, \begin{eqnarray*} && p_c=\mathbb{E}[D \vert Z=1] - \mathbb{E}[D \vert Z=0],\ p_a=\mathbb E[D\vert Z=0],\ p_n=\mathbb E[1-D\vert Z=1], p_{df}=0,\\ && \mu_{F_{11a}^{UB}}-\mathbb E[Y \vert D=1, Z=0] \leq \delta_{1a} \leq \mu_{F_{11a}^{LB}}-\mathbb E[Y \vert D=1, Z=0],\\ &&\mathbb E[Y \vert D=0, Z=1]-\mu_{F_{00n}^{LB}} \leq \delta_{0n} \leq \mathbb E[Y \vert D=0, Z=1]-\mu_{F_{00n}^{UB}},\\ && \delta_{0c}+\theta_{1c}=\delta_{1c}+\theta_{0c} = \frac{\mathbb E[Y D \vert Z=1]- p_a \mu_{11a}}{p_c}-\frac{\mathbb E[Y (1-D) \vert Z=0]- p_n \mu_{00n}}{p_c},\\ && \mu_{F_{11a}^{UB}} \leq \mu_{11a} \leq \mu_{F_{11a}^{LB}},\ \mu_{F_{00n}^{UB}} \leq \mu_{00n} \leq \mu_{F_{00n}^{LB}}. \end{eqnarray*} These bounds are sharp.

Proposition (ref) and Corollaries (ref)-(ref) are intermediate results that are key to deriving the identified set for the vector of parameters $\Gamma$. The identified set for $\Gamma$ is

eqnarray*[eqnarray* omitted — 375 chars of source]

where

eqnarray*[eqnarray* omitted — 935 chars of source]

$\Theta_I^2(A_3)=\tilde{\Theta}_I^1(A_3),$ where $\tilde{\Theta}_I^1(A_3)$ is the corresponding $\Theta_I^1(A_3)$ defined for $\tilde{\Gamma}$ based on $(Y,D,\tilde{Z})$ and $\tilde{T}$, and

eqnarray*[eqnarray* omitted — 698 chars of source]
remarkWe can express $\mathbb{E}[Y \mid Z = 1]$ as \begin{equation*} \begin{aligned} \mathbb{E}[Y \mid Z = 1] =& \mathbb{E}[Y_{11} D_1 + Y_{01} (1 - D_1) \mid Z = 1] \\ =& \mathbb{E}[Y_{11} D_1 + Y_{01} (1 - D_1)] \\ =& \mathbb{E}[Y_{11} \mid T = a] p_a + \mathbb{E}[Y_{11} \mid T = c] p_c + \mathbb{E}[Y_{01} \mid T = n] p_n, \end{aligned} \end{equation*} where the second equality holds under Assumption (ref), and we can apply the law of iterated expectation to obtain the third equality. Using the similar reasoning, we can also write $\mathbb{E}\left[Y \mid Z = 0\right]$ as \begin{equation*} \begin{aligned} \mathbb{E}[Y \mid Z = 0] =& \mathbb{E}[Y_{10} D_0 + Y_{00} (1 - D_0) \mid Z = 0] \\ =& \mathbb{E}[Y_{10} D_0 + Y_{00} (1 - D_0)] \\ =& \mathbb{E}[Y_{10} \mid T = a] p_a + \mathbb{E}[Y_{00} \mid T = c] p_c + \mathbb{E}[Y_{00} \mid T = n] p_n. \end{aligned} \end{equation*} As is customary, define the intent-to-treat (ITT) parameter as $\operatorname{ITT} \equiv \mathbb{E}[Y \mid Z = 1] - \mathbb{E}[Y \mid Z = 0]$. Then \begin{equation*} \begin{aligned} \operatorname{ITT} =& \mathbb{E}[Y \mid Z = 1] - \mathbb{E}[Y \mid Z = 0] \\ =& \mathbb{E}[Y_{11} - Y_{10} \mid T = a] p_a + \mathbb{E}[Y_{01} - Y_{00} \mid T = n] p_n + \mathbb{E}[Y_{11} - Y_{00} \mid T = c] p_c \\ =& \delta_{1a} p_a + \delta_{0n} p_n + \left[ \delta_{1c}+\theta_{0c} \right] p_c. \end{aligned} \end{equation*} Since $p_a, p_n, p_c \geq 0$, and $p_a + p_n + p_c = 1$ under Assumption (ref), we have shown that $\operatorname{ITT}$ can be expressed as a convex combination of $\delta_{1a}$, $\delta_{0n}$, and $\delta_{1c}+\theta_{0c}$.
remark(Testable implication) In Corollary (ref), we point identify $\mathbb{E}[Y_{10} \mid T = a]$ and $\mathbb{E}[Y_{01} \mid T = n]$, and partially identify $\mathbb{E}[Y_{11} \mid T = a]$ and $\mathbb{E}[Y_{00} \mid T = n]$. Those identification results can be exploited as a testable implication for the exclusion restriction. Since the exclusion restriction states that $Y_{11} = Y_{10}$ and $Y_{01} = Y_{00}$, it implies that $\mathbb{E}[Y_{11} \mid T = a] = \mathbb{E}[Y_{10} \mid T = a]$ and $\mathbb{E}[Y_{01} \mid T = n] = \mathbb{E}[Y_{00} \mid T = n]$. Therefore, if the point identified $\mathbb{E}[Y_{10} \mid T = a]$ is not in the identified set of $\mathbb{E}[Y_{11} \mid T = a]$, or the point identified $\mathbb{E}[Y_{01} \mid T = n]$ is not contained in the identified set of $\mathbb{E}[Y_{00} \mid T = n]$, we have enough evidence to reject the exclusion restriction assumption (ref). However, this testable implication is not sharp for testing the exclusion restriction assumption (ref). Kitagawa2021TheIndependence showed that inequality (ref) is the sharp testable implication for the random assignment and exclusion restriction assumptions (ref) and (ref) together. This result is confirmed in the sharpness proof of Proposition (ref) in this paper. In a randomized experiment setting like the one considered in this paper, the testable inequality (ref) can be interpreted as a testable implication for only the exclusion restriction assumption.
remarkNotice that unlike $\Theta_I(A_1)$ and $\Theta_I(A_2)$ which could be empty, $\Theta_I(A_3)$ is never empty. This implies that Assumptions (ref) and (ref) do not have any testable implication.

Misspecification robust bounds

In this section, we follow li2024discordant to derive a misspecification robust bound for the model and the various assumptions considered in this paper. This is a discrete way of salvaging the model when it is refuted by the data. This relaxation is different from the continuous relaxation approach introduced in Masten2021SalvagingModels. We find this discrete relaxation easier to interpret than the continuous one. Before we proceed to show the misspecification robust bound for the collection of assumptions $A$, we discuss when each of the assumptions $A_1$, $A_2$, and $A_3$ is data-consistent. According to li2024discordant, an assumption $A_k$ is data-consistent if its identified set is nonempty. Therefore, $A_1$ is data-consistent if $\Theta_I^1(A_1)$ is nonempty when $\mathbb E[D\mid Z=1]-\mathbb E[D\mid Z=0]>0$, $\Theta_I^2(A_1)$ is nonempty when $\mathbb E[D\mid Z=1]-\mathbb E[D\mid Z=0]<0$, and $\Theta_I^3(A_1)$ is nonempty when $\mathbb E[D\mid Z=1]-\mathbb E[D\mid Z=0]=0$. Assumption $A_2$ is data-consistent if inequality (ref) holds. Assumption $A_3$ is always data-consistent as its identified set is never empty.

definition[li2024discordant, li2024discordant] An assumption $A_k\in A$ is a minimum data-consistent relaxation of $A$ if $\Theta_I(A_k)\neq \emptyset$ and $\Theta_I(A_k\cup A_{k'})=\emptyset$ for any $A_{k'} \neq A_k: A_{k'} \in A$.

The misspecification robust bound $\Theta_I^*(A)$ is defined as the union of the identified sets of all minimum data-consistent relaxations of $A$. Note here that $A_2 \cup A_3=A_1$. Then $\Theta_I(A)=\Theta_I(A_1\cup A_2\cup A_3)=\Theta_I(A_1).$ If $\Theta(A)\neq \emptyset$, there exists a unique minimum data-consistent relaxation by definition, it is $A_1$. But, if $\Theta_I(A_1)=\emptyset$ and $\Theta_I(A_2)\neq \emptyset$, we have two minimum data-consistent relaxations, $A_2$ and $A_3$. Finally, if $\Theta(A_2)=\emptyset$ (i.e., inequality (ref) is violated), then $A_3$ will be the only minimum data-consistent relaxation. Hence, we have

eqnarray*[eqnarray* omitted — 517 chars of source]

In the next section, we illustrate the theoretical results with a real-world application.

Empirical illustration

In this section, we illustrate our methodology by revisiting bursztyn2020misperceived, which is also re-examined by kwon2024testing in a mechanism framework. bursztyn2020misperceived study how misperceived social norms affect women's labor force participation in Saudi Arabia. They find that while many men privately support female employment, they incorrectly believe that their peers do not share this view. To study the effects of correcting these misperceptions, the authors design a randomized experiment in which men are randomly informed that a majority of their peers actually support female employment. The results show that women whose husbands receive the information are significantly more likely to register for a job-search service. In addition, the study finds persistent effects on women's employment behavior, as women from treated households are more likely to be employed several months after the intervention.

To align this study with our framework, we define the women long-term labor supply outcome\footnote{bursztyn2020misperceived construct an index by pooling six measures of women's long-term labor supply outcomes, which we use as the outcome in our analysis.} as the outcome $Y$, the decision to sign up for a job-search service as the binary treatment variable $D$, and the information regarding social norms as a randomized experiment $Z$. Given that the information experiments in this study are randomly assigned, we maintain Assumption (ref) throughout the analysis. The dataset applied here comes from the online appendix of bursztyn2020misperceived. Our sample is restricted to observations with non-missing data on women's long-term labor supply outcomes, resulting in a final sample size of 381 observations. Some summary statistics are provided in Table (ref). We observe that households assigned to the treatment group ($Z=1$) are more likely to enroll in job-search services compared to those in the control group ($Z=0$). Additionally, treated households, on average, have better long-term female labor force outcomes than the control households.

The results presented in this section are entirely based on estimation. We provide a brief discussion of the inference methods and present the corresponding inference results in Appendix (ref).

table[table omitted — 500 chars of source]

Identified set under $A_1$

We estimate the difference in probabilities, $\mathbb{P}(D=1 \mid Z=1)-\mathbb{P}(D=1 \mid Z=0)$, and obtain a result of $\hat{\mathbb{P}}(D=1 \mid Z=1)-\hat{\mathbb{P}}(D=1 \mid Z=0) = 0.0929 > 0$. Given this positive estimate, we want to test the inequalities (ref)-(ref), which are the sharp testable implications of the joint assumptions in $A_1 = \{RA, ER, MON\}$. We apply the testing method proposed in Mourifie2017TestingAssumptions and reject the validity of inequalities (ref)-(ref) at the 1% level. This rejection implies that the identified set $\Theta_I(A_1)$ is empty, suggesting that the assumptions in $A_1$ are not consistent with the dataset used in bursztyn2020misperceived.

The application in kwon2024testing differs from ours. First, we use a continuous outcome variable constructed in bursztyn2020misperceived, which is an index summarizing six measures of female labor force participation, including an indicator for whether the wife applies for jobs outside the home. In contrast, kwon2024testing use only this binary indicator as their outcome variable. Second, in our notation, the primary objective of kwon2024testing is to test the validity of the exclusion restriction assumption (ref). For their implementation, they assume monotonicity (Assumption (ref)) and find that the set of assumptions $A_1$ is rejected. So, the rejection of $A_1$ implies that the random assignment $Z$ violates the exclusion restriction if monotonicity holds. They also estimate the average direct effects of the random assignment on the outcome of always-takers and never-takers. Finally, they perform a sensitivity analysis by gradually relaxing the monotonicity assumption, assuming a certain proportion of defiers to exist. In contrast, our method constructs misspecification robust bounds, as discussed in Section (ref), by applying the minimum data-consistent relaxation of the LATE assumptions. Unlike kwon2024testing, we do not impose any prior restrictions on the proportion of defiers. Instead, we discretely relax the LATE assumptions in a manner that ensures they remain consistent with the observed data.

Identified set under $A_2$

Next, we relax Assumption (ref) and consider identification under the assumption set $A_2 = \{RA, ER\}$. Based on Proposition (ref), we estimate the identified set for the proportion of defiers as

equation*[equation* omitted — 67 chars of source]

Since this identified set is nonempty, Proposition (ref) implies that we fail to reject the inequality (ref). Therefore, the combination of assumptions in $A_2$ cannot be rejected. Table (ref) reports the estimated identified sets for the proportions of types under the assumption set $A_2$.

table[table omitted — 465 chars of source]

Building on the analysis in Section (ref), we derive informative identified sets for the local average instrument-controlled direct effects for compliers, $\theta_{1c} = \theta_{0c}$, and for defiers, $\theta_{1df} = \theta_{0df}$. Figure (ref) presents the estimated bounds for these identified sets.

figure[figure omitted — 174 chars of source]

Figure (ref) shows the heterogeneous effects of job-search services across different household types. Specifically, we observe positive treatment effects for compliers across all values of $p_c$ within the identified set. In complier households, individuals enroll in job-search services only when they are informed about the social norms. Upon learning that more women are employed than they initially thought, they may perceive an increased likelihood of getting a job, as a larger female workforce indicates greater opportunities in the labor market, which motivates them to comply. These characteristics of complier households align with the increase effects on long-term female labor force participation following their enrollment in job-search services.

In contrast, we observe negative treatment effects for defiers across most values of $p_{df}$ within the identified set. Unlike compliers, women in defier households may regard increased female workforce as a signal of intensified competition. This may discourage them from seeking jobs, particularly if their relative advantage has diminished. Consequently, they choose not to enroll in job-search services once they are informed about the social norms. These characteristics of defier households are consistent with the decreased effects on long-term female labor force participation following enrollment in job-search services, as they may face greater competition in the female labor market.

Notably, the IV estimate $(2.68)$ falls outside our estimated bounds. It overestimates the effect of job-search services for compliers relative to our identified set.

Identified set under $A_3$

Finally, we relax Assumption (ref) while maintaining Assumption (ref) and study identification under the assumption set $A_3 = \{RA, MON\}$. As discussed in Section (ref), the identified set under $A_3$ is always non-empty, implying that the joint assumptions in $A_3$ can never be rejected.

With Assumption (ref) imposed, defiers cannot exist and all proportions of types can be point identified in this framework. Table (ref) presents the estimations of type proportions under $A_3$. The results show that never-takers constitute the largest proportion of the population, accounting for 70.68% of all observations. This finding suggests that the majority of households in the dataset would not sign up job-search services for women, regardless of whether or not they receive the information about correct social norms.

table[table omitted — 415 chars of source]

According to the identified set $\Theta_I(A_3)$, we have informative identified sets for the local average treatment-controlled direct effects for always-takers and never-takers, $\delta_{1a}$ and $\delta_{0n}$, and the total effects for compliers, $\delta_{1c} + \theta_{0c} = \delta_{0c} + \theta_{1c}$. Table (ref) presents the estimated results for these identified sets.

table[table omitted — 483 chars of source]

The results in Table (ref) indicate that the information experiment has positive direct effects on long-term female labor force outcomes for both always-takers and never-takers, who together account for nearly 90% of the population. This finding implies that, even though the information experiment does not affect job-search service enrollment decisions in these households, exposure to correct social norm information directly benefits women's long-term labor force participation.

Misspecification robust bound

Since $\widehat{\Theta}_I(A_2)$ and $\widehat{\Theta}_I(A_3)$ are nonempty, the estimated misspecification robust bound in this empirical example is $\widehat{\Theta}^*_I(A)=\widehat{\Theta}_I(A_2) \cup \widehat{\Theta}_I(A_3)$. We explained earlier in Section (ref) how this result is achieved.

Conclusion

In this paper, we propose a robust identification approach in a randomized experiment setting with imperfect compliance where the LATE assumptions could potentially be rejected by the data. We propose bounds that are never empty and robust to detectable violations of the LATE assumptions. The bounds are also robust to violations of the monotonicity or exclusion restriction assumptions. We introduce two causal parameters: the local average treatment-controlled direct effect, and the local average instrument-controlled direct effect. We apply our methods to revisit bursztyn2020misperceived and study the effect of randomly providing accurate social norm information on the decision to enroll women in job-search services, which may affect women's labor force participation in Saudi Arabia. We show that the LATE assumptions are jointly rejected in this context. However, the assumptions cannot be rejected when either the monotonicity or exclusion restriction is relaxed. Under the assumptions of random assignment and the exclusion restriction, we find that job-search service enrollment has a positive effect on long-term female labor force outcomes for compliers, whereas the effect is negative for defiers. Under the assumptions of random assignment and monotonicity, our results suggest that the information experiment itself has positive direct effects on long-term female labor force outcomes for both always-takers and never-takers, who together account for more than 90% of the population.

The approach developed in this paper works well for a binary instrument. An interesting area for future research would consist of extending the proposed method to a multivalued discrete or continuous instrument in the context of marginal treatment effects introduced in Heckman2001Policy-RelevantEffects, Heckman2005StructuralEvaluation. One challenge is that the number of causal parameters would be infinite in this setting. Recently, sun2022pairwise proposed a pairwise validity approach for multivalued discrete instruments in the LATE framework. A combination of their method with the one developed in this paper could be an interesting area for future research. Another consideration for future research could be how to incorporate exogenous covariates in the analysis. While the approach still works conditional on covariates, it may suffer from the curse of dimensionality as the number of parameters grows with the cardinality of the support of the covariates. Semiparametric assumptions could also enter the set $A$ of assumptions considered in the paper.

\onehalfspacing