EconBase
← Back to paper

Sensitivity Analysis for Instrumental Variables Under Joint Relaxations of Monotonicity and Independence

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

67,266 characters · 16 sections · 50 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sensitivity Analysis for Instrumental Variables Under Joint Relaxations of Monotonicity and Independence

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi

\if01 {

center[center omitted — 35 chars of source]

} \fi

abstractIn this paper I develop a breakdown frontier approach to assess the sensitivity of Local Average Treatment Effects (LATE) estimates to violations of monotonicity and independence of the instrument. I parametrize violations of independence using the concept of $c$-dependence from mastenpoirier2018 and allow for the share of defiers to be greater than zero but smaller than the share of compliers. I derive identified sets for the LATE and the Average Treatment Effect (ATE) in which the bounds are functions of these two sensitivity parameters. Using these bounds, I derive the breakdown frontier for the LATE, which is the weakest set of assumptions such that a conclusion regarding the LATE holds. I derive consistent sample analogue estimators for the breakdown frontiers and provide a valid bootstrap procedure for inference. Monte Carlo simulations show the desirable finite-sample properties of the estimators and an empirical application shows that the conclusions regarding the effect of family size on female labor force participation from angev are highly sensitive to violations of independence and monotonicity.

{\it Keywords:} Partial identification, heterogeneous treatment effects, selection on unobservables.

\spacingset{1.8}

Introduction

Instrumental variables (IV) techniques are among the most widely used empirical tools in social sciences. In the canonical IV setting, the causal effect of a binary treatment is identified by exploiting variations in a binary instrument in the form of the wald estimand. Point identification is achieved if the instrument satisfies a set of assumptions. For instance, the instrumental variable must be independent from potential treatments and potential outcomes. Also, the instrument must affect treatment uptake in the same direction for all individuals, which is usually referred to as the monotonicity assumption.

If the instrument satisfies the independence and the monotonicity assumption, along additional assumptions, then the Wald estimand identifies the average effect of the treatment for compliers, the subpopulation of individuals whose treatment status mimics its assignment, which is called the Local Average Treatment Effect, or simply LATE imbensangrist.

In the recent years, applied researchers have grown increasingly skeptical of IV methods cinelli. The identifying assumptions are often unverifiable, and although in certain cases some assumptions are readily justified (for instance, the independence assumption in experimental studies with imperfect compliance), in most cases they are defended by appealing to context-specific knowledge.

In this paper I study what can be learned about treatment effects in IV settings under relaxations of independence and monotonicity and develop a breakdown frontier approach for assessing the sensitivity of IV estimates. I focus in the case where the outcome is binary. I begin by deriving bounds for potential treatments and potential outcomes under a bounded dependence assumption called c-dependence mastenpoirier2018, which bounds the distance between the probability of being assigned to treatment given observed covariates and unobserved potential quantities and the probability of assignment given just the observed covariates.

I then use the bounds for potential quantities to derive identified sets for the causal effect of assignment (which I will refer to as the Intention-to-Treat, or simply ITT) and the LATE, under a share of defiers which is greater than zero, but always smaller than the share of compliers. I derive the conditions under which these identified sets are sharp. In the cases where these conditions do not hold, the identified sets still provide a valid outer region for the parameters of interest.

One can argue that once the identifying assumptions for IV settings are violated, the LATE is no longer an interesting causal parameter. Thus, I also derive the identified set for the Average Treatment Effect (ATE) and show how the bounds under violations of independence and monotonicity are connected to well known bounds for the ATE using IVs in the causal inference literature Balke01091997,Chen1015-7483R1.

I use the bounds of the ITT and the LATE to construct breakdown frontiers for conclusions regarding causal effects. The breakdown frontier in this setting provides the largest combination of violations of independence and monotonicity under which a particular conclusion holds. For instance, suppose a researcher finds a positive point estimate for the LATE, but is skeptical towards the identifying assumptions. To provide evidence of robustness of the qualitative takeaways of its findings (for instance, that the effect is indeed positive), the researcher can use the breakdown frontier to show the combinations of violations under which one can still conclude that the LATE is greater than zero.

I propose nonparametric estimators for the bounds of causal effects and breakdown values, and derive their asymptotic properties using convergence results for Hadamard directional differentiable functions fangsantos. Standard inference methods such as the nonparametric bootstrap are not consistent for the breakdown frontiers. I show, however, that valid uniform confidence bands can be estimated using the boostrap procedure for Hadamard directional differentiable functions in fangsantos and the numerical estimator for the Hadamard derivative from HONG2018379. Monte Carlo simulations show the desirable finite sample properties of the estimators and inference procedures.

For the empirical application, I revisit angev, which studies the effects of family size on female employment using same-sex siblings as the instrument. The estimated breakdown frontier for the LATE shows that the qualitative takeaway from this study only holds under very small violations of the identifying assumptions. Therefore, the breakdown frontier approach suggests that the conclusions of the study are highly sensitive to violations of independence and monotonicity.

Related Literature: This paper relates broadly to three strands of the causal inference literature. First, it is connected to the literature on partial identification and sensitivity analysis in IV settings. Most papers in this literature focus on partial identification and sensitivity analysis under violations of independence and the exclusion restrictionconley,wang18,mastenporirer21,cinelli. There also papers that focus on identification and sensitivity analysis under violations of monotonicity tolerating,noack2026sensitivity. In this paper, I consider both relaxations of independence and monotonicity.

Second, this paper relates to the literature on the identification of breakdown values, introduced by horwitzmanski. My approach to inference follows closely the one introduced in mastenpoirier2020 as it also uses $c$-dependence to parametrize violations of independence. While most of the work in this literature focuses on missing data settings klinesantos and selection on observables mastenpoirier2020, this is one of the first papers studying inference for breakdown values in settings with non-compliance. In that sense, it is closely related to the work of noack2026sensitivity, but under a different parametrization for violations of monotonicity. A desirable feature of the breakdown analysis in this paper is that the violations of independence and monotonicity are measured in the same unit, which makes the interpretation of the tradeoffs of violations displayed by the breakdown frontier particularly easy.

Finally, this paper is related to the literature on IV settings with binary outcomes, which dates back to the seminal work of heckman78. While most prominent work on this literature focuses on the identification of the average structural functions vytlacilyildiz,shaikhvytlacil or partial identification of Average Treatment Effects Balke01091997,Chen1015-7483R1,MACHADO2019522, this paper considers both the partial identification of the LATE and the ATE.

Outline of the paper: The rest of the paper is organized as follows: Section 2 describes the framework and target parameters in the setting. Section 3 provides the partial identification of potential treatments and outcomes, and in Section 4 I derive the identified sets for the ITT and the LATE show the identification of the breakdown frontiers. I also derive the identified sets for the ATE. Section 5 introduces the estimators and their asymptotic properties, as well as the bootstrap procedure used for inference. Section 6 presents the Monte Carlo simulation studies. Section 7 presents the empirical application and Section 8 concludes. Appendix A contains the main proofs from the results in the paper, and Appendix B contains auxiliary lemmas.

General Framework

Setup

Let $Z\in\left\{ 0,1\right\}$ denote a binary variable that indicates whether an individual was assigned to treatment ($Z=1$) or control ($Z=0$). In this setting, non-compliance is allowed, which means that not all individuals assigned to treatment will actually take the treatment and not all individuals assigned to control will remain untreated. Rather than determining treatment status, the assignment represents an encouragement (or discouragement) towards treatment.

Let $D\in\left\{ 0,1\right\}$ denote the actual treatment status. Define the potential treatment associated to assignment $z$ as $D(z)$. We observe the treatment status

equation*[equation* omitted — 38 chars of source]

Let $Y\in\left\{ 0,1\right\}$ denote the observed binary outcome. The potential outcome associated to assignment $z$ is defined as $Y(D(z),z)$. At first, I allow potential outcomes depend arbitrarily on treatment and assignment. Observed and potential outcomes are related by

equation*[equation* omitted — 48 chars of source]

Let $X\in\mathcal{S}(X)$ be a vector of observed covariates and $p_{z|x}=\mathbb{P}\left ( Z=z|X=x \right )$ be the observed propensity score for assignment. I maintain the following assumption regarding the joint distribution of $(D(z),Y(D(z),z),Z,X)$ throughout the paper:

Assumption 1: For each $z,z'\in\left\{ 0,1\right\}$ and $x\in\mathcal{S}(X)$:

enumerate$\mathbb{P}\left ( D(z)=1|Z=z',X=x \right )\in \left ( 0,1 \right )$$\mathbb{P}\left ( Y(D(z),z)=1|Z=z',X=x \right )\in \left ( 0,1 \right )$$\mathbb{P}\left ( Y(D(z),z)=1,D(z)=1|Z=z',X=x \right )\in \left ( 0,1 \right )$$p_{z|x}>0$

Assumptions 1.1 to 1.3 state that the support of potential quantities does not depend on the assignment. Assumption 1.4 states that all individuals can be assigned to treatment and control with probability greater than zero, and is usually referred to as the common support, or overlap assumption.

I also maintain the standard exclusion restriction assumption from IV settings, which imposes that assignment does not affect potential outcomes directly.

Assumption 2: For $(D(z),z)\in\left\{0,1\right\}^{2}$, $Y(D(z),z)=Y(D(z))$.

In order to identify treatment effects with instrumental variables, it is standard to assume that the instrument is independent of potential outcomes and potential treatments conditional on $X=x$. The goal of this identification analysis is to study what can be said about treatment effects when standard IV assumptions fail to hold. To do this, I replace these standard assumptions by a bounded dependence assumption, called c-dependence mastenpoirier2018:

Definition ($c$-dependence): Let $z\in\left\{0,1\right\}$ and $x\in\mathcal{S}(X)$. Let $c$ be a scalar between 0 and 1. $Z$ is conditionally $c$-dependent with $Y(D(z)),D(z)$ given $X=x$ if

equation*[equation* omitted — 171 chars of source]

where $\mathcal{S}_{x}((Y(D(z)),D(z)))$ is the support of $(Y(D(z)),D(z))$ conditional on $X=x$. Conditional $c$-dependence provides a parametrization of violations of independence which has a straightforward interpretation. The sensitivity parameter $c$ can be interpreted as the difference between the unobserved assignment probability and the observed propensity score in terms of probability units. When $c=0$, independence holds, and potential probabilities $\mathbb{P}\left(Y(D(z))=y|X=x\right)$ and $\mathbb{P}\left(D(z)=d|X=x\right)$ are point identified. Throughout this paper, $c$-dependence is assume to hold.

Assumption 3: $Z$ is $c$-dependent with $(Y(D(1)),D(1))$ given $X$ and $(Y(D(0)),D(0))$ given $X$.

Without further assumptions, individuals can be partitioned into four groups regarding how they respond to assignment: always-takers ($at$), never-takers ($nt$), compliers ($co$) and defiers ($def$). Let $\pi_{g|x}$ denote the proportion of individuals from group $g\in\left\{at,nt,co,def\right\}$ with covariates equal to $x$. The fundamental behavioral assumption in IV settings is the monotonicity assumption, which imposes that for all $x\in\mathcal{S}(X)$, $\pi_{def|x}=0$.

I relax this assumption to allow for the presence of defiers, but I restrict the proportion of defiers to be smaller than the share of compliers:

Assumption 4: For all $x\in\mathcal{S}(X)$, $\pi_{co|x}>\pi_{def|x}$.

Assumption 4 is analogous to Assumption 5 from tolerating. If this assumption holds, then the the estimate for the first-stage is positive\footnote{See noack2026sensitivity for a partial ID framework where defiance is allowed which relaxes this assumption.}. Thus, the sensitivity parameter $\pi_{def|x}$ can be seen as a measure of deviation between the proportion of compliers $\pi_{co|x}$ and the first-stage estimand $\mathbb{E}\left[D|Z=1,X=x\right]-\mathbb{E}\left[D|Z=0,X=x\right]$ in terms of probability units.

If assumption 3 holds with $c=0$, then potential quantities are identified. If Assumption 4 further holds with $\pi_{def|x}=0$ for all $x$, then we go back to the standard IV setting with binary outcomes, where the LATE is point identified by the Wald estimand, and the ATE is partially identified within the bounds provided by Balke01091997.

Target Parameters

In this paper, I focus on the partial Identification of the Local Average Treatment Effect for compliers, which is usually the target parameter in IV settings and the Average Treatment Effect (ATE), which is typically the causal parameter that researchers would ideally like to identify. Define $LATE_{g}=\mathbb{E}\left[Y(1)-Y(0)|g\right]$ as the local average treatment effect for group $g$, with $g\in\left\{at,nt,co,def\right\}$. We are thus, interested in the partial identification of $LATE_{co}$ and the identification of a breakdown frontier which can be used to assess the robustness of results from studies that employ IV methods.

Researchers often report a point estimate of the paramerer $LATE_{co}$ because they assume the parameter is point identified in their instrumental variable setting. However, it is often argued that $LATE_{co}$ is not necessarily a relevant parameter huber2017jae. Moreover, once the instrument is not assumed to be independent from conditional quantities, nor it is assumed to be monotonic, then potential quantities for the sub-population of compliers are no longer point identified. Therefore, I also focus on the partial identification of the ATE.

The partial identification approach is built using the following steps. First, I derive the bounds for conditional potential joint probabilities $\mathbb{P}\left(Y(D(z))=y,D(z)=d|X=x\right)$. Then, I derive the bounds for the marginal probabilities of potential outcomes and potential treatments, $\mathbb{P}\left(Y(D(z))=y|X=x\right)$ and $\mathbb{P}\left(D(z)=d|X=x\right)$. Then, I derive bounds for the parameter $LATE_{co|x}$, which is the LATE for compliers conditional on $X=x$, and a breakdown frontier, and finally bounds for the conditional ATE. Unconditional quantities are partially identified by integrating the conditional bounds over the distribution of covariates.

Partial Identification of Potential Probabilities

I begin with the identified set for the joint probability of potential quantities. I begin with the joint probability of potential quantities. Under Assumptions 1 and 2, the results from Proposition 5 of mastenpoirier2018 can be readily adapted to the IV setting and the modified conditional $c$-dependence assumption. Let $p_{y,d|z,x}=\mathbb{P}\left(Y=y,D=d|Z=z,X=x\right)$:

propositionSuppose Assumptions 1-3 hold. Then the sharp identified set for $\mathbb{P}\left(Y(D(z))=y,D(z)=d|X=x\right)$ is { \begin{equation*} \mathbb{P}\left ( Y(D(z))=y,D(z)=d|X=x \right )\in\left [\mathbb{P}_{c}^{l}\left ( Y(D(z))=y,D(z)=d|X=x \right ),\mathbb{P}_{c}^{u}\left ( Y(D(z))=y,D(z)=d|X=x \right ) \right ] \end{equation*}} where \begin{align*} &\mathbb{P}_{c}^{l}\left ( Y(D(z))=y,D(z)=d|X=x \right )=\&\max\left\{\frac{p_{y,d|z,x}p_{z|x}}{p_{z|x}+c},\frac{p_{y,d|z,x}p_{z|x}-c}{p_{z|x}-c}\mathbf{1}\left ( p_{z|x}>c \right ),p_{y,d|z,x}p_{z|x}\right\},\&\mathbb{P}_{c}^{u}\left ( Y(D(z))=y,D(z)=d|X=x \right )\&=\min\left\{\frac{p_{y,d|z,x}p_{z|x}}{p_{z|x}-c}\mathbf{1}\left ( p_{z|x}>c \right )+\mathbf{1}\left ( p_{z|x}\leq c \right ),\frac{p_{y,d|z,x}p_{z|x}+c}{p_{z|x}+c},p_{y,d|z,x}p_{z|x}+(1-p_{z|x})\right\} \end{align*}

The notation introduced in Proposition 1 for the bounds on the joint probability of potential quantities illustrates the fact the bounds are functions of the sensitivity parameter $c$. When $c=0$ the joint probability is point identified by the conditional probability $\mathbb{P}\left(Y=y,D=d|Z=z,X=x\right)$,and as $c$ increases the identified set becomes larger until we reach the worst-case identified set.

The bounds from Proposition 1 can be combined using the Law of Total Probabilities to obtain bounds for the marginal probabilities of potential quantities. I begin with the bounds for potential outcomes:

propositionSuppose Assumptions 1-3 hold. Then the identified set for $\mathbb{P}\left(Y(D(z))=y|X=x\right)$ is \begin{equation*} \mathbb{P}\left ( Y(D(z))=y|X=x \right )\in\left [\mathbb{P}_{c}^{l}\left ( Y(D(z))=y|X=x \right ),\mathbb{P}_{c}^{u}\left ( Y(D(z))=y|X=x \right ) \right ] \end{equation*} where \begin{align*} &\mathbb{P}_{c}^{l}\left ( Y(D(z))=y|X=x \right )\&=\max\left\{ \mathbb{P}^{l}_{c}\left(Y(D(z))=y,D(z)=1|X=x\right)+\mathbb{P}^{l}_{c}\left(Y(D(z))=y,D(z)=0|X=x\right),p_{y|z,x}p_{z|x}\right\},\&\mathbb{P}_{c}^{u}\left ( Y(D(z))=y|X=x \right )\&=\min\left\{\mathbb{P}^{u}_{c}\left(Y(D(z))=y,D(z)=1|X=x\right)+\mathbb{P}^{u}_{c}\left(Y(D(z))=y,D(z)=0|X=x\right),p_{y|z,x}p_{z|x}+(1-p_{z|x})\right\} \end{align*} Moreover, if $p_{y,d|z,x}<1/2$ and $c\in\left ( 0,\min\left\{ p_{z|x}(1-2p_{y,d|z,x}),\frac{p_{z|x}(1-p_{z|x})(1-p_{y,d|z,x})}{p_{y,d|z,x}p_{z|x}+1-p_{z|x}}\right\} \right )$, then the identified set is sharp.

Proposition 2 shows that the bounds for potential outcomes can be obtained by combining the bounds for join potential quantities, and the additional conditions under which the identified set is sharp. Essentially, the additional conditions imply that the upper bound for joint potential probabilities simplifies to $\mathbb{P}^{u}_{c}\left(Y(D(z))=y,D(z)=d|X=x\right)=\frac{p_{y,d|z,x}p_{z|x}}{p_{z|x}-c}$, and that the lower bound simplifies to $\mathbb{P}^{l}_{c}\left(Y(D(z))=y,D(z)=d|X=x\right)=\frac{p_{y,d|z,x}p_{z|x}}{p_{z|x}+c}$. If this conditions do not hold, the identified set still provides a valid outer region.

The results from Proposition 2 can be easily adapted to obtain bounds for potential treatments:

propositionSuppose Assumptions 1-3 hold. Then the identified set for $\mathbb{P}\left(D(z)=d|X=x\right)$ is \begin{equation*} \mathbb{P}\left ( D(z)=d|X=x \right )\in\left [\mathbb{P}_{c}^{l}\left ( D(z)=d|X=x \right ),\mathbb{P}_{c}^{u}\left ( D(z)=d|X=x \right ) \right ] \end{equation*} where \begin{align*} &\mathbb{P}_{c}^{l}\left ( D(z)=d|X=x \right )\&=\max\left\{ \mathbb{P}^{l}_{c}\left(Y(D(z))=1,D(z)=d|X=x\right)+\mathbb{P}^{l}_{c}\left(Y(D(z))=0,D(z)=d|X=x\right),p_{d|z,x}p_{z|x}\right\},\&\mathbb{P}_{c}^{u}\left ( D(z)=d|X=x \right )\&=\min\left\{\mathbb{P}^{u}_{c}\left(Y(D(z))=1,D(z)=d|X=x\right)+\mathbb{P}^{u}_{c}\left(Y(D(z))=0,D(z)=d|X=x\right),p_{d|z,x}p_{z|x}+(1-p_{z|x})\right\} \end{align*} Moreover, if $p_{y,d|z,x}<1/2$ and $c\in\left ( 0,\min\left\{ p_{z|x}(1-2p_{y,d|z,x}),\frac{p_{z|x}(1-p_{z|x})(1-p_{y,d|z,x})}{p_{y,d|z,x}p_{z|x}+1-p_{z|x}}\right\} \right )$, then the identified set is sharp.

Bounds for unconditional potential probabilities are obtained by integrating the bounds of conditional probabilities over the distribution of covariates. The bounds derived in this section are the building blocks for the partial identification of treatment effects which is presented in the next section.

Partial Identification of Treatment Effects

Partial Identification of $LATE_{co}$

I begin deriving bounds for the average treatment effects of the sub-population of compliers. In the standard IV setting with covariates, the parameter $LATE_{co|x}$ is partially identified by the conditional Wald estimand:

{

equation*[equation* omitted — 269 chars of source]

}

The numerator of the Wald estimand, usually referred to as the reduced form estimand, identifies the causal effect of assignment, $\mathbb{E}\left[Y(D(1))-Y(D(0))|X=x\right]$, which is equal to the treatment effect for compliers multiplied by the share of compliers in the standard IV setting. This parameter is often called the Intention-to-Treat effect (I will refer to its conditional as $ITT_{x}$ and its unconditional version as $ITT$). The ITT is rarely the parameter of interest in IV settings, but it carries important information regarding the LATE for compliers. For instance, the paramater $LATE_{co}$ has the same sign as the ITT if monotonicity holds, or if the share of compliers is greater than the share of defiers. The next proposition provides the bounds for the ITT as functions of the sensitivity parameters $c$ and $\pi_{def|x}$, as well as the conditions under which these bounds are sharp.

propositionSuppose Assumptions 1-4 hold. Then $ITT_{x}\in\left[ITT^{l}(c,\pi_{def|x}),ITT^{u}(c,\pi_{def|x})\right]$, where \begin{align*} &ITT^{u}(c,\pi_{def|x})=\min\left\{\mathbb{P}_{c}^{u}\left(Y(D(1))=1|X=x\right)-\mathbb{P}_{c}^{l}\left(Y(D(0))=1|X=x\right)+\pi_{def|x},1\right\},\&ITT^{l}(c,\pi_{def|x})=\max\left\{\mathbb{P}_{c}^{l}\left(Y(D(1))=1|X=x\right)-\mathbb{P}_{c}^{u}\left(Y(D(0))=1|X=x\right)-\pi_{def|x},-1\right\} \end{align*} Moreover, if \begin{align*} &\&(i)\max\left\{0, \mathbb{P}\left ( D(0)=1|X=x \right )-\mathbb{P}\left ( D(1)=1|X=x \right ) \right\}\&\leq \pi_{def|x}\&\leq \min\left\{ \mathbb{P}\left ( D(0)=1|X=x \right ),\mathbb{P}\left ( D(1)=0|X=x \right )\right\}\&(ii)\ p_{y,d|z,x}<1/2\&(iii)\ c\in\left ( 0,\min\left\{ p_{z|x}(1-2p_{y,d|z,x}),\frac{p_{z|x}(1-p_{z|x})(1-p_{y,d|z,x})}{p_{y,d|z,x}p_{z|x}+1-p_{z|x}}\right\} \right )\&(iv)\ \mathbb{P}^{u}_{c}\left(Y(D(0))=1|X=x\right)-\mathbb{P}^{l}_{c}\left(Y(D(1))=1|X=x\right)-1\&<\pi_{def|x}\&<1-\left(\mathbb{P}^{u}_{c}\left(Y(D(1))=1|X=x\right)-\mathbb{P}^{l}_{c}\left(Y(D(0))=1|X=x\right)\right) \end{align*} for all $\left(y,d\right)\in\left\{0,1\right\}^{2}$ and $x\in\mathcal{S}(X)$, then the identified set is sharp.

Proposition 4 provides bounds for the conditional ITT. The bounds for the unconditional ITT are obtained by integrating the conditional bounds over the distribution of covariates. Under additional assumptions, the bound is sharp. These assumptions restrict the share of complier to lie within the Fréchet-feasible interval (i), the values which joint potential probabilities can take (ii), the values which the sensitivity parameter $c$ can take (iii) and the share of defiers to lie in an interval in which the truncations of the bounds are not active (iv). If these assumptions fail to hold, the bounds still provide a valid outer region for the ITT.

Proposition 4 provides bounds for the ITT. Once bounds for the share of compliers $\pi_{co|x}$ are obtained, one can derive the identified set for $LATE_{co|x}$:

propositionSuppose Assumptions 1-4 hold. Then the identified set for $LATE_{co|x}$ is $\left[LATE^{u}_{co|x}(c,\pi_{def|x}),LATE^{l}_{co|x}(c,\pi_{def|x})\right]$, where \begin{align*} &LATE_{co|x}^{UB}(c,\pi_{def|x})=\min\left\{\frac{ITT_{x}^{u}(c,\pi_{def|x})}{\pi_{co|x}^{l}(c,\pi_{def|x})},1\right\},\&LATE_{co|x}^{LB}(c,\pi_{def|x})=\max\left\{\frac{ITT^{l}_{x}(c,\pi_{def|x})}{\pi_{co|x}^{u}(c,\pi_{def|x})},-1\right\} \end{align*} with \begin{align*} &\pi_{co|x}^{u}(c,\pi_{def|x})=\min\left\{\mathbb{P}_{c}^{u}\left(D(1)=1|X=x\right)-\mathbb{P}^{l}\left(D(0)=1|X=x\right)+\pi_{def|x},1\right\},\&\pi_{co|x}^{l}(c,\pi_{def|x})=\max\left\{ \mathbb{P}_{c}^{l}\left(D(1)=1|X=x\right)-\mathbb{P}^{u}\left(D(0)=1|X=x\right)+\pi_{def|x},0\right\} \end{align*} Moreover, if \begin{align*} &\&(i)\max\left\{0, \mathbb{P}\left ( D(0)=1|X=x \right )-\mathbb{P}\left ( D(1)=1|X=x \right ) \right\}\&\leq \pi_{def|x}\&\leq \min\left\{ \mathbb{P}\left ( D(0)=1|X=x \right ),\mathbb{P}\left ( D(1)=0|X=x \right )\right\}\&(ii)\ c=0\&(iii)\ \pi_{def|x}<1-\left(\mathbb{P}\left(D=1|Z=1,X=x\right)-\mathbb{P}\left(D=1|Z=0,X=x\right)\right) \end{align*} for all $\left(y,d\right)\in\left\{0,1\right\}^{2}$ and $x\in\mathcal{S}(X)$, then the identified set is sharp.

Proposition 5 provides the bounds for the LATE of compliers. In general, the bounds will not be sharp, since under violations of independence (Assumption 4 holds with $c>0$), the upper bound of the conditional ITT and the lower bound of the conditional share of compliers (and vice-versa) cannot be attained simultaneously while satisfying Assumptions 1-4. Nevertheless, in the case where independence holds and the share of compliers is such that it satisfies the Frechet inequalities and the bounds for the share of compliers are not the worst-case bounds, the identified set is sharp.

Breakdown Frontier

In this section I provide a breakdown frontier approach to assess the robustness of ITT and LATE estimates to violations of independence and monotonicity. I focus on the breakdown frontiers for the conclusions that $ITT\geq \mu_{1}$ and $LATE_{co}\geq \mu_{2}$. Choosing $\mu_{1}$ and $\mu_{2}$ equal to 0, for instance, provides us the breakdown analysis for the conclusion that the treatment has a positive effect.

When deriving the breakdown frontier, it is important to consider the same share of defiers across all values of covariates ($\pi_{def|x}=\pi_{def}$) in order to go obtain unconditional bounds. First, consider all the values of $c$ and $\pi_{def}$ under which the conclusion holds. This sets are called the robust regions and are defined for the ITT and the LATE, respectively, as

align*[align* omitted — 262 chars of source]

Robust regions are simply combinations of $\left(c,\pi_{def}\right)$ which respectively deliver identified sets for the ITT and $LATE_{co}$ that contain the values $\mu_{1}$ and $\mu_{2}$. The breakdown frontiers are the sets $\left(c,\pi_{def}\right)$ in the boundary of the robust region for given conclusions. The breakdown frontiers are

align*[align* omitted — 256 chars of source]

Note that, in the cases where the bounds for the ITT and the LATE are not sharp, the breakdown frontiers operate as a conservative sufficient-robustness frontier rather than the exact combination of breakdown values.

Solving for $\pi_{def}$ in the equations $ITT^{l}(c,\pi_{def})=\mu_{1}$ and $LATE_{c}^{l}(c,\pi_{def})=\mu_{2}$ yields

align*[align* omitted — 479 chars of source]

Therefore, we obtain the following analytical expressions for the breakdown frontiers:

align*[align* omitted — 199 chars of source]

The frontiers provide the largest relaxations $c$ and $\pi_{def}$ under which predetermined conclusions regarding the ITT and the LATE hold. The shape of the frontier allows us to analyze the trade-off between the two types of relaxations considered when drawing conclusions regarding the target parameters. A desirable feature of this approach is that the sensitivity parameters $c$ and $\pi_{def}$ are measured in the same unit. Although this is not necessary, it certainly can be helpful.

Note that when we are interested in assessing the conclusion regarding the sign of treatment effect we can always use the breakdown frontiers for the ITT, as $bf_{LATE}^{}(c,0)=bf_{ITT}^{}(c,0)$. Next, I provide a simple numerical illustration of the bounds of the treatment effects and the breakdown frontier approach.

Numerical Illustration

I consider a simple DGP with a single covariate $X$ where $x\in\left\{0,1\right\}$. The instrument is assigned according to a Bernoulli distribution with parameter $p=0.6$. The covariate is distributed according to a Bernoulli distribution with parameter $p=0.5$. See Appendix C for the entire characterization of the DGP. Potential outcomes and treatments are defined in a way such that for $x\in\left\{0,1\right\}$, we have

align*[align* omitted — 254 chars of source]

Therefore, in the absence of violations of the identifying assumptions in IV settings, the ITT is equal to 0.25 and the LATE is equal to 0.5 under this DGP. To analyze the sensitivity to violations of independence and monotonicity, Figures 1 shows the identified sets for the LATE under different shares of defiers.

figure[figure omitted — 502 chars of source]

The plot on the left of Figure 1 shows the identified set for the LATE under the monotonicity assumption ($\pi_{def|x}=0$ for all $x$). When $c=0$, the identified set collapses to 0.5, which is the value which would be point identified in the absence of any violation. As $c$ increases, the set becomes less informative. The vertical dotted line marks the largest violation $c$ under which the identified set does not contain 0. That is, in the absence of defiers, we can conclude that the LATE is positive under violations of independence for all sensitivity parameters $c\leq0.15$.

The plot on the right shows the identified set when defiance is allowed (I set $\pi_{def|x}=0.1$ for all $x$). Note that in this case, the LATE is no longer point identified when $c=0$. The vertical dotted line is moved to the left, and shows that we can conclude that the LATE is positive for violation parameters $c\leq 0.1$.

The plots with the identified sets under different shares of defiers illustrate the tradeoff between the magnitude of the violations when assessing the robustness of a given conclusion regarding the LATE. If we want to conclude that the LATE is positive, we can allow for smaller deviations from independence as we allow for larger shares of defiers.

The breakdown frontier format captures the tradeoffs between these violations. Figure 2 shows the breakdown frontiers for two conclusions regarding the LATE.

figure[figure omitted — 431 chars of source]

The plot on the left of Figure 2 provides the breakdown frontier for the conclusion that $LATE_{co}^{l}(c,\pi_{def})\geq 0$. The area painted in blue represents the robust region for the conclusion that the LATE is positive, and the black line denotes the breakdown frontier. The breakdown frontier shows that if we are willing to assume independence, then the share of defiers can be as great as 0.25 and the conclusion that the LATE is positive still holds. If we are willing to assume monotonicity, then the observed and unobserved propensity scores can differ by up to 0.15 probability units and the conclusion still holds.

The plot on the right shows the robust region and the breakdown frontier for the conclusion that $LATE^{l}(c,\pi_{def})\geq 0.25$, which is half of the value that is point identified under the standard assumptions. Note that the robust region is smaller that the one for the conclusion that the LATE is positive, and smaller violations of monotonicity are admitted in order for the conclusion to hold. If we are willing to assume that independence holds, then we can allow for a share of defiers no greater than 0.1. If we are willing to assume monotonicity, then the observed and unobserved propensity scores can differ by up to 0.075 probability units and the conclusion still holds.

Partial Identification of the ATE

Researchers usually report the LATE in IV settings because that is the causal parameter that is point identified under the standard IV assumptions imbensangrist. However, whether or not the LATE is a relevant parameter depends on the empirical context huber2017jae,Chen1015-7483R1. Researchers are typically interested in the Average Treatment Effect (ATE), which is the most general average causal parameter. Moreover, once the standard IV assumptions are violated and potential quantities are no longer point identified for the group of compliers, it might be of interest to analyze what can be learned about the ATE.

The ATE is a parameter that is not point identified in standard IV settings, as the quantities $\mathbb{P}\left(Y(0)=1|at,X=x\right)$ and $\mathbb{P}\left(Y(1)=1|nt,X=x\right)$ cannot be point identified from the data without further assumptions. If violations of monotonicity are allowed, further potential outcomes cannot be point identified. If violations of independence is also allowed, then none of the potential quantities are identified. The next proposition shows what are the bounds for the ATE under violations of monotonicity and independence.

propositionSuppose Assumptions 1-4 hold. Then, $ATE_{x}\in\left[ATE^{l}_{x}(c,\pi_{def|x}),ATE^{u}_{x}(c,\pi_{def|x})\right]$, where \begin{align*} &ATE_{x}^{u}(c,\pi_{def|x})=\mathbb{P}^{u}\left(Y(1)=1|X=x\right)-\mathbb{P}^{l}\left(Y(0)=1|X=x\right),\&ATE^{l}_{x}(c,\pi_{def|x})=\mathbb{P}^{l}\left(Y(1)=1|X=x\right)-\mathbb{P}^{u}\left(Y(0)=1|X=x\right) \end{align*} with \begin{equation*} \begin{aligned} \mathbb{P}^{u}\left ( Y(1)=1|X=x \right ) = \min\Bigg\{& \mathbb{P}^{u}_{c}\left(Y(D(1))=1,D(1)=1|X=x\right) +\mathbb{P}^{u}_{c}\left(Y(D(0))=1,D(0)=1|X=x\right) \\ &+\mathbb{P}^{u}_{c}\left(D(1)=0|X=x\right)-\pi_{def|x},\&\mathbb{P}\left (Y=1|D=1,X=x \right )\mathbb{P}\left (D=1|X=x \right )+(1-\mathbb{P}\left (D=1|X=x \right )) \Bigg\}, \end{aligned} \end{equation*} \begin{equation*} \begin{aligned} \mathbb{P}^{l}\left ( Y(1)=1|X=x \right ) \\= \max\Bigg\{& \mathbb{P}^{l}_{c}\left(Y(D(1))=1,D(1)=1|X=x\right) +\mathbb{P}^{l}_{c}\left(Y(D(0))=1,D(0)=1|X=x\right) \\ &-\min\left\{\mathbb{P}^{u}_{c}\left(Y(D(1))=1,D(1)=1|X=x\right),\mathbb{P}^{u}_{c}\left(Y(D(0))=1,D(0)=1|X=x\right)\right\},\&\mathbb{P}\left (Y=1|D=1,X=x \right )\mathbb{P}\left (D=1|X=x \right ) \Bigg\}, \end{aligned} \end{equation*} \begin{equation*} \begin{aligned} \mathbb{P}^{u}\left ( Y(0)=1|X=x \right ) = \min\Bigg\{& \mathbb{P}^{u}_{c}\left(Y(D(1))=1,D(1)=0|X=x\right) +\mathbb{P}^{u}_{c}\left(Y(D(0))=1,D(0)=0|X=x\right) \\ &+\mathbb{P}^{u}_{c}\left(D(0)=1|X=x\right)-\pi_{def|x},\&\mathbb{P}\left (Y=1|D=0,X=x \right )\mathbb{P}\left (D=0|X=x \right )+(1-\mathbb{P}\left (D=0|X=x \right )) \Bigg\}, \end{aligned} \end{equation*} \begin{equation*} \begin{aligned} \mathbb{P}^{l}\left ( Y(0)=1|X=x \right ) \\=\max\Bigg\{& \mathbb{P}^{l}_{c}\left(Y(D(1))=1,D(1)=0|X=x\right) +\mathbb{P}^{l}_{c}\left(Y(D(0))=1,D(0)=0|X=x\right) \\ &-\min\left\{\mathbb{P}^{u}_{c}\left(Y(D(1))=1,D(1)=0|X=x\right),\mathbb{P}^{u}_{c}\left(Y(D(0))=1,D(0)=0|X=x\right)\right\},\&\mathbb{P}\left (Y=1|D=0,X=x \right )\mathbb{P}\left (D=0|X=x \right ) \Bigg\} \end{aligned} \end{equation*} Moreover, if \begin{align*} &\&(i)\max\left\{0, \mathbb{P}\left ( D(0)=1|X=x \right )-\mathbb{P}\left ( D(1)=1|X=x \right ) \right\}\&\leq \pi_{def|x}\&\leq \min\left\{ \mathbb{P}\left ( D(0)=1|X=x \right ),\mathbb{P}\left ( D(1)=0|X=x \right )\right\}\&(ii)\ p_{y,d|z,x}<1/2\&(iii)\ c\in\left ( 0,\min\left\{ p_{z|x}(1-2p_{y,d|z,x}),\frac{p_{z|x}(1-p_{z|x})(1-p_{y,d|z,x})}{p_{y,d|z,x}p_{z|x}+1-p_{z|x}}\right\} \right )\&(iv)\ \mathbb{P}^{u}_{c}\left(Y(D(0))=1|X=x\right)-\mathbb{P}^{l}_{c}\left(Y(D(1))=1|X=x\right)-1\&<\pi_{def|x}\&<1-\left(\mathbb{P}^{u}_{c}\left(Y(D(1))=1|X=x\right)-\mathbb{P}^{l}_{c}\left(Y(D(0))=1|X=x\right)\right) \end{align*} for all $\left(y,d\right)\in\left\{0,1\right\}^{2}$ and $x\in\mathcal{S}(X)$, then the identified set is sharp.

The bounds in Proposition 6 can be directly connected to the existing bounds for the ATE in the IV literature. The next corollary shows that the bounds from proposition 6 are equivalent to the bounds from Balke01091997 and Chen1015-7483R1 in the absence of violations.

corollarySuppose Assumptions 1-4 hold. Furthermore, suppose that Assumption 3 holds with $c=0$ and Assumption 4 holds with $\pi_{def|x}=0$. Then, the bounds for $ATE_{x}$ become \begin{align*} &ATE^{u}_{x}=\mathbb{P}\left(Y=1,D=1|Z=1,X=x\right)-\mathbb{P}\left(Y=1,D=0|Z=0,X=x\right)+\mathbb{P}\left(D=0|Z=1,X=x\right),\&ATE^{l}_{x}=\mathbb{P}\left(Y=1,D=1|Z=1,X=x\right)-\mathbb{P}\left(Y=1,D=0|Z=0,X=x\right)-\mathbb{P}\left(D=1|Z=0,X=x\right) \end{align*}

Estimation and Inference

In this section, I study estimation and inference of the bounds for the LATE and the breakdown frontiers for the LATE and ITT defined in Section 4.1. The bounds and the breakdown frontier are known functionals of conditional probabilities of treatments and outcomes given assignments and covariates, and the conditional probabilities of assignments given covariates. Hence, I propose nonparametric sample analogue estimators for the bounds and the breakdown frontier.

First I assume a random sample of data is available for the researcher:

Assumption 5: The random variables $\left\{Y_{i},D_{i},Z_{i},X_{i}\right\}_{i=1}^{N}$ are independently and identically distributed according to the distribution of $\left(Y,D,Z,X\right)$.

Furthermore, assume that the support of the vector of covariates is discrete:

Assumption 6: The support of $X$ is discrete and finite. Let $\mathcal{S}(X)=\left\{x_{1},...,x_{K}\right\}$.

Next, I invoke an assumption which is an important regularity condition for the derivation of the asymptotic properties of the estimator.

Assumption 7: For all $x\in\mathcal{S}(X)$, we have $c<\min\left\{p_{1|x},p_{0|x}\right\}$.

Assumption 7 is necessary for the proposed bounds to be sharp, but is also key for asymptotics. The asymptotic results are obtained using a delta method for directionally differentiable functionals. Under assumption 7, the indicator functions inside the min and max operators that determines the bounds disappear, and therefore, there are no Dirac delta functions in the analytical expression.

I begin with the asymptotic properties of the bounds for the LATE and its breakdown frontier.

LATE and Breakdown Frontier

The parameters of interest defined in Section 4.1 are functionals of the parameters $p_{y,d|z,x}=\mathbb{P}\left(Y=y,D=d|Z=z,X=x\right)$, $p_{z|x}=\mathbb{P}\left(Z=z|X=x\right)$ and $q_{x}=\mathbb{P}\left(X=x\right)$. Let

align*[align* omitted — 450 chars of source]

denote the sample analog estimators of these probabilities. In Lemma 1 of Appendix B, I show that the estimators of these quantities converge uniformly to a Gaussian process at a $\sqrt{N}$-rate.

The bounds in Propositions 1-6 are functionals evaluated at $p_{y,d|z,x}$, $p_{z|x}$ and $q_{x}$. The bounds are estimated by these functionals evaluated at the sample analogue estimators. If these functionals are Hadamard directional differentiable, then $\sqrt{N}$-convergence in distribution of the sample analogue estimators will carry over to the functionals by the delta method.

I use the functional delta method for Hadamard directionally differentiable mappings fangsantos to show convergence in distribution of the estimators. Convergence is usually to a non-Gaussian limiting process. Thus, analytical asymptotic bands are challenging to obtain. I follow mastenpoirier2020 and propose a bootstrap procedure to obtain asymptotically valid uniform confidence bands for the breakdown frontier and the estimators for the bounds.

Consider the bounds from Proposition 1. Under Assumptions 1-7, we estimate them by

align*[align* omitted — 555 chars of source]

The estimators perform poorly when $c$ is close to $p_{z|x}$. Assumption 7 ensures that $c$ is bounded away from $p_{z|x}$. The estimators for the bounds of potential treatments are analogous. In Lemmas 3 and 4 from Appendix B I show that these estimators converge in distribution to a nonstandard distribution.

For the main results in this section I establish convergence uniformly over $c\in\mathcal{C}$, where $\mathcal{C}$ is a finite grid $\mathcal{C}\subset\left[0,\min\left\{p_{1|x},p_{0|x}\right\}\right]$ for all $x\in\mathcal{S}(X)$. Therefore, the asymptotic results are valid for values of $c$ which satisfy Assumption 7.

Next, consider the bounds for the conditional ITT introduced in Proposition 4. We estimate them by

align*[align* omitted — 382 chars of source]

The unconditional bounds are estimated by integrating over the empirical distribution of the covariates $X$. Let

equation*[equation* omitted — 217 chars of source]

In Lemma 5 of Appendix B, I show that these estimators for the ITT bounds converge weakly to a Gaussian element.

Now, consider the estimation for the breakdown frontier for the conclusion that the ITT is above a certain threshold $\mu_{1}$. Although the ITT is not the usual parameter of interest in IV settings, the breakdown frontier for the conclusion that the ITT is greater than zero coincides with the breakdown frontier for the conclusion that the LATE is greater than zero, so it is interesting to analyze its asymptotic properties.

Denote the breakdown frontier for the conclusion that $ITT\geq\mu_{1}$ by

equation*[equation* omitted — 122 chars of source]

where

equation*[equation* omitted — 211 chars of source]

I show that the estimator for the breakdown frontier of the ITT converges in distribution.

theoremSuppose Assumptions 1-7 hold and that $c\in\mathcal{C}$ for some finite grid $\mathcal{C}\subset \overline{C}\in\left(0,\min\left\{p_{1|x},p_{0|x}\right\}\right)$. Let $\mathcal{M}\subset\left[-1,1\right]$ be a finite grid of points. Then, \begin{equation*} \sqrt{N}\left(\widehat{BF}_{ITT}(c,\mu)-BF_{ITT}(c,\mu)\right)\xrightarrow[d]Z_{BF_{ITT}}(c,\mu) \end{equation*} a tight random element of $l^{\infty}\left(\mathcal{C}\times\mathcal{M}\right)$.

Now, consider the bounds for the conditional LATE introduced in Proposition 5. They are obtained by combining the bounds for the ITT with the bounds for the share of compliers. We estimate the bounds for the share of compliers by

align*[align* omitted — 375 chars of source]

The unconditional bounds are estimated by integrating over the empirical distribution of the covariates $X$. Let

equation*[equation* omitted — 239 chars of source]

In Lemma 6 of Appendix B, I show that these estimators converge weakly to a Gaussian element. The estimators for the bounds of the LATE are obtained by combining the bounds of the ITT and the share of compliers:

align*[align* omitted — 278 chars of source]

In Lemma 7 of Appendix B, I show that these estimators converge weakly to a Gaussian element. The estimator for the breakdown frontier for the conclusion that $LATE\geq \mu_{2}$ is

equation*[equation* omitted — 124 chars of source]

where

align*[align* omitted — 346 chars of source]

I show that the estimator for the breakdown frontier of the LATE converges in distribution.

theoremSuppose Assumptions 1-7 hold and that $c\in\mathcal{C}$ for some finite grid $\mathcal{C}\subset \overline{C}\in\left(0,\min\left\{p_{1|x},p_{0|x}\right\}\right)$. Let $\mathcal{M}\subset\left[-1,1\right]$ be a finite grid of points. Then, \begin{equation*} \sqrt{N}\left(\widehat{BF}_{LATE}(c,\mu)-BF_{LATE}(c,\mu)\right)\xrightarrow[d]Z_{BF_{LATE}}(c,\mu) \end{equation*} a tight random element of $l^{\infty}\left(\mathcal{C}\times\mathcal{M}\right)$.

The results in this section essentially follow from the $\sqrt{N}$-convergence rate of the sample analogue estimators to a Gaussian process and by sequential applications of the Delta Method for Hadamard directionally differentiable functions.

Bootstrap Inference

The limiting processes of the estimators presented in this Section are non-Gaussian, so relying on analytical estimates of quantiles of functionals of these processes would be challenging. In order to overcome these challenges I use the bootstrap procedure from mastenpoirier2020. The bootstrap procedure is subsequently used to construct uniform confidence bands for the breakdown frontiers.

Let $W_{i}=\left(Y_{i},D_{i},Z_{i},X_{i}\right)$ and $W^{N}=\left\{W_{1},W_{2},...,W_{N}\right\}$. Let $\theta_{0}$ denote a parameter of interest and $\widehat{\theta}$ be an estimator of $\theta_{0}$ based on $W^{N}$. Define $\textbf{A}_{N}^{*}=\sqrt{N}\left(\widehat{\theta}^{*}-\widehat{\theta}\right)$, where $\widehat{\theta}^{*}$ is a draw from the nonparametric bootstrap distribution of $\widehat{\theta}$.

I focus on

equation*[equation* omitted — 215 chars of source]

Let $\textbf{Z}_{1}$ denote the limiting distribution of $\sqrt{N}\left(\widehat{\theta}-\theta_{0}\right)$, which is defined in Lemma 1 of Appendix B. It is well known that $\textbf{A}_{N}^{*}$ converges weakly to $\textbf{Z}_{1}$. The parameters of interest are functionals $\phi$ of $\theta_{0}$. For Hadamard differentiable functions, the nonparametric bootstrap is valid fangsantos. However, when parameters are only Hadamard directionally differentiable, which is the case for the bounds of the ITT and LATE, and the breakdown frontiers, the nonparametric bootstrap is not consistent.

To construct a consistent bootstrap distribution, I use the bootstrap procedure from fangsantos, which relies on a consistent estimator $\widehat{\phi^{'}}_{\theta_{0}}$ of the Hadamard derivative at $\theta_{0}$. These estimates can be obtained by using the numerical derivative estimator proposed by HONG2018379, which is

equation*[equation* omitted — 278 chars of source]

and is computed across the bootstrap estimates $\widehat{\theta}^{*}$ Under the constraints $\varepsilon_{N}\rightarrow0$ and $\sqrt{N}\varepsilon_{N}\rightarrow\infty$ and additional regularity conditions, this numerical derivative bootstrap procedure is consistent (Li and Hong, 2018).

I use this bootstrap procedure construct uniform confidence bands for the breakdown frontiers. I focus on one-sided lower uniform confidence bands. I am looking for a lower bound function $\widehat{LB}(c)$ such that

align*[align* omitted — 149 chars of source]

I consider bands of the form

equation*[equation* omitted — 105 chars of source]

where $\widehat{z}_{1-\alpha}$ is a scalar and $\sigma(.)$ is a known function. Note that under Assumptions 1-7, the estimators for the breakdown frontiers can be written as $\widehat{BF}(c,\mu)=\left[\phi(\widehat{\theta})\right](c)$, where $\phi$ is Hadamard directionally differentiable. If we further assume that $\varepsilon_{N}\rightarrow0$ and $\sqrt{N}\varepsilon_{N}\rightarrow\infty$, then the conditions in Proposition 2 from mastenpoirier2020 hold, and the estimator

equation*[equation* omitted — 301 chars of source]

is consistent for $z_{1-\alpha}$, the $1-\alpha$ quantile of the cdf of

equation*[equation* omitted — 83 chars of source]

Note that this holds for the estimators of both breakdown frontiers. It follows that the proposed lower bands are valid uniformly on the grid $\mathcal{C}$. In the next section, I study the finite-sample properties of the estimation and inference procedures for breakdown frontiers.

Monte Carlo Simulations

In this section I study the finite sample performance of the estimation and inference procedures proposed in Section 5. I consider the same DGP from the numerical illustration in Section 4.2.1, which implies a joint distribution for $\left(Y,D,Z,X\right)$ from which I draw independently.

I consider two sample sizes, $N=1000$ and $N=2000$. For each sample size, I conduct 500 Monte Carlo simulations. For each exercise, I compute the estimated breakdown frontier and a 95% lower bootstrap uniform confidence band. In all simulations, I set $\varepsilon_{N}=2/\sqrt{N}$, which is the choice of $\varepsilon_{N}$ which shows the best finite-sample coverage in MastenPoirier2020_supplement. I estimate the breakdown frontier over a finite grid of points $c$. In the simulation, I use 100 values of $c$ equally spaced between 0 and 0.15 both in all simulations and bootstrap procedures.

figure[figure omitted — 525 chars of source]

Figure 3 shows the shows the sampling distribution of the breakdown frontier estimator for the conclusion that the LATE is greater than zero. The first thing that shows out is that, as implied by the consistency result in Section 5, the distribution of the estimator becomes tighter around the true frontier as the sample size increases. Second, the sampling distribution looks fairly symmetric around the true frontier. This contrasts with the findings of MastenPoirier2020_supplement, which find that the estimator for the breakdown frontier of Distributional Treatment Effects is biased downwards. The difference might arise due to several factors: we consider different target parameters and different sensitivity parameters for the relaxation of the identifying assumptions, which inevitably leads to different functional forms for the breakdown frontiers. Nevertheless, the fact that the estimator for the breakdown frontier of the LATE is symmetric around the true frontier is a desirable feature which is does not hold generally for breakdown approach settings.

figure[figure omitted — 479 chars of source]

Figure 4 shows the true breakdown frontier as the solid line, and the sample mean of breakdown estimates across the Monte Carlo simulations with $N=1.000$ as the dashed line. The two lines are pretty much overlapped, which shows that the finite-sample bias of the estimator for the frontier is very small across all considered values of $c$. The dotted line below represents the sample mean of the lower confidence band with nominal coverage $1-\alpha=0.95$.

Overall the results of the Monte Carlo exercise show desirable finite-sample properties of the estimator for the breakdown frontier. A pervasive concern when conducting inference procedures in IV settings is the so-called weak instrument problem. Although there are several bootstrap procedures that improve inference in settings with weak instruments where the standard assumptions hold, it is unclear how to improve the bootstrap for nondifferentiable functions. I leave this analysis for future work.

Empirical Application

In this section, I use the estimators from Section 5 to perform the breakdown analysis for the results regarding family size and female employment in angev, using data from the US Census Public Use Microsamples married mothers aged 21–35 in 1980 with at least 2 children and oldest child less than 18.

In this setting, the dependent variable is and indicator for women who did not work for pay in 1979. Treatment is an indicator for women having three or more children, and the instrument is an indicator for women whose first two children have the same sex. The authors control for age, age at the first birth, race and sex of the first child as covariates.

Two concerns regarding the assumptions that lead to point identification of the LATE in this setting arise. The first, regards violations of monotonicity. The assumption holds if all parents in the sample have weak preferences towards mixed-sibling compositions. Although there is evidence that more families with two same-sex siblings have a higher probability of third birth than families with two siblings with mixed composition, this does not guarantee that there are no families which prefer same-sex siblings over mixed compositions. The second concern comes from the independence assumption. Genetic conditions which determine fertility outcomes can be correlated to economic outcomes farb, which would lead to violations of independence. Under the light of this concerns, the angev setting seems to be well suited for the breakdown analysis approach.

To begin the sensitivity analysis, I use selection on observables to to calibrate the beliefs regarding the amount of selection on unobservables. I take the approach from altonji and mastenpoirier2018. I partition the vector of covariates $X$ as $(X_{k},X_{-k})$, where $X_{k}$ is the $k$-th component and $X_{-k}$ is a vector with remaining components. The measures used to calibrate the beliefs regarding deviations from independence are

equation*[equation* omitted — 171 chars of source]

In the data, the largest value obtained form $\overline{c}_{k}$ is associated to to the indicator for women whose first child is a man, which was estimated to be $\overline{c}_{1st\ sex}=0.011$.

Using this result as a reference for the breakdown analysis, a robust result would have a breakdown frontier which admits values of $c$ above $\overline{c}_{1st\ sex}$.

To calibrate the beliefs regarding violations of monotonicity, I follow tolerating, which uses a survey from Peru in which women were asked about their ideal sex composition for their children. In the survey, 1.8% of the respondents had three children or more and declared that ideal sex sibship composition would have been two boys and no girl, or no boy and two girls. Thus, one can argue that these women seem to have been induced to having a third child because their first two children were a boy and a girl. Using this result as a reference, a robust result would have a breakdown frontier which admits a share of defiers greater than 0.018.

figure[figure omitted — 396 chars of source]

I estimate the breakdown frontier using 50 values of $c$, equally spaced between 0 and 0.1. For the construction of the lower confidence bands, I draw 999 boostrap samples from the data and set the tuning parameter to $\varepsilon_{N}=\frac{2}{\sqrt{N}}$. The implementation algorithm is similar to the one in Section 4 of mastenpoirier2020, although it is much simpler since the considered outcome is binary.

Figure 5 shows the estimated breakdown frontier for the conclusion that the the effects of family size on employment is negative. The solid line is the estimated breakdown frontier, and the dashed line is the lower confidence band at the $1-\alpha=0.95$ level.

One can think of this the breakdown frontier as the frontier for the conclusion that the qualitative takeaways from angev hold. The plot shows that when independence holds ($c=0$) the maximum share of defiers uner which the qualitative takeaways hold is 0.008, which lies below the baseline share of defiers implied the by the Peruvian survey. When monotonicity holds ($def=0$) the largest admissible difference between the observable and unobservable propensity scores is around 0.004 probability units, which lies below the baseline violation of 0.011 implied by the calibration based on selection on observables.

Overall, the results from this breakdown analysis suggest that the conclusion that effect of family size on employment is negative is not robust to violations of independence or monotonicity of the same-sex siblings instrument. The results align with the findings of noack2026sensitivity which shows that small violations of monotonicity lead to uninformative results in this setting, and also add to the discussion that small deviations from independence also lead to uninformative results.

Conclusion

In this paper, I provide a breakdown frontier approach to sensitivity analysis in Instrumental Variables settings. I study the partial identification of the LATE under parametrizations of violations of independence and monotonicity. The bounds for the LATE are used to derived breakdown frontiers, the weakest set of assumptions such that a particular conclusion of interest holds. Also, I derive identified set for the ATE under violations of independence and monotonicity given the fact that when the population of compliers is not point-identified, the LATE is no longer such a relevant parameter.

I propose sample analogue estimators and uniform confidence bands for the breakdown frontiers. Monte Carlo simulations show that the estimator exhibits desirable finite-sample properties.

Finally, I use the proposed breakdown frontier approach to revisit the results from angev, and find that the conclusions regarding the effect of family size on unemployment are highly sensitive to violations of independence and monotonicity.