EconBase
← Back to paper

Breakdown Analysis for Instrumental Variables with Binary Outcomes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

55,011 characters · 14 sections · 30 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Breakdown Analysis for Instrumental Variables with Binary Outcomes

titlepage\begin{abstract} This paper studies the partial identification of treatment effects in Instrumental Variables (IV) settings with binary outcomes under violations of independence. I derive the identified sets for the treatment parameters of interest in the setting, as well as breakdown values for conclusions regarding the true treatment effects. I derive $\sqrt{N}$-consistent nonparametric estimators for the bounds of treatment effects and for breakdown values. These results can be used to assess the robustness of empirical conclusions obtained under the assumption that the instrument is independent from potential quantities, which is a pervasive concern in studies that use IV methods with observational data. In the empirical application, I show that the conclusions regarding the effects of family size on female unemployment using same-sex siblings as the instrument are highly sensitive to violations of independence. \\ \noindentKeywords: Partial Identification, Sensitivity Analysis, Instrumental Variables. \\ \noindentJEL Codes: C01, C13, C21.\\ \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

\doublespacing

Introduction

Instrumental variables (IV) techniques are among the most widely used empirical tools in social sciences. In the canonical IV setting, the causal effect of a binary treatment is identified by exploiting variations in a binary instrument in the form of the wald estimand. Point identification is achieved if the instrument satisfies a set of assumptions. For instance, the instrumental variable must be independent from potential treatments and potential outcomes.

Although in certain cases the independence assumption is readily justified (experimental studies), it is often unverifiable and must be defended by appealing to context specific knowledge, specially in observational studies. In this paper I study what can be learned about treatment effects in IV settings under weaker versions of the instrument independence assumption.

I focus in the case where the outcome is binary. I derive bounds for the first-stage and reduced form parameters, as well as bounds for the Local Average Treatment Effect (LATE) under a bounded dependence assumption called c-dependence mastenpoirier2018, which bounds the distance between the probability of receiving the instrument given observed covariates and unobserved potential quantities and the probability of being treated given just the observed covariates.

The first-stage parameter, the share of compliers, is partially identified as function of the difference between the probability of assignment given covariates and the probability of assignment given covariates and potential treatments. The reduced form parameter, the intention-to-treat effect, is partially identified as function of the difference between the probability of assignment given covariates and the probability of assignment given covariates and potential outcomes. The LATE is partially identified as a function of both sensitivity parameters.

I identify breakdown values for the first-stage and reduced form, as well as the breakdown frontier for the LATE. Breakdown values are the violations of the independence assumptions under which a particular conclusion no longer holds. For instance, one could be interested to learn under which violations of independence the conclusion that the treatment effect has a particular sign holds. If a researcher is concerned about the external validity of the study, the breakdown analysis of the first-stage is useful to understand under which violations of independence the share of compliers is above a certain share of the study population.

I propose nonparametric estimators for the bounds of causal effects and breakdown values, and derive their asymptotic properties using convergence results for Hadamard directional differentiable functions fangsantos.

The identified sets for the LATE are not sharp. I derive sharp bounds for the LATE under a joint $c$-dependence assumption for potential outcomes and potential treatments. The bounds of the set can be used to derive the breakdown point for conclusions regarding the LATE.

Monte Carlo simulations show the desirable finite sample properties of the proposed estimators.

For the empirical application, I revisit angev, which studies the effects of family size on female employment using same-sex siblings as the instrument, and estimate the identified sets for the share of compliers, the ITT and the LATE under different relaxations of independence. The estimated breakdown values for the LATE and the ITT are not statistically different from zero, which suggests that the conclusions of the study are highly sensitive to violations of independece.

Related Literature: This paper relates broadly to three strands of the causal inference literature. First, ir is connected to the literature on partial identification and sensitivity analysis in IV settings. While most results on the literature focus on partial identification under violations of the exclusion restriction conley,wang18,mastenporirer21,cinelli, this paper focuses solely on violations of independence. In that sense, it is similar to mastenkline and jonashesh, but differs from the former by allowing two-sided noncompliance and from the latter by choosing a different sensitivity parameter.

Second, this paper relates to the literature on the identification of breakdown values, introduced by horwitzmanski. My approach to inference follows closely the one introduced in mastenpoirier2020. While most of the work in this literature focuses on missing data settings klinesantos and selection on observables mastenpoirier2020, this is the first paper studying inference for breakdown values in settings with non-compliance.

Finally, this paper is related to the literature on IV settings with binary outcomes, which dates back to the seminal work of heckman78. While most prominent work on this literature focuses on the identification of the average structural functions vytlacilyildiz,shaikhvytlacil or partial identification of Average Treatment Effects MACHADO2019522, this paper is more closely related to chesher as it builds on the LATE framework for identification and sensitivity analysis.

Outline of the paper: The rest of the paper is organized as follows: Section 2 describes the framework and target parameters in the setting. Section 3 provides the partial identification results for the case of binary outcomes and in Section 4 I show the identification of the breakdown values. In Section 5 I perform a numerical illustration of the method. Section 6 introduces the estimators and their asymptotic properties. Section 7 presents the partial identification results for the case of joint $c$-dependence. Section 8 presents the Monte Carlo simulations, Section 9 presents the empirical application and Section 10 concludes.

Framework

Let $Z\in\left\{ 0,1\right\}$ denote a binary variable that indicates whether an individual was assigned to treatment ($Z=1$) or control ($Z=0$). In the setting, non-compliance is allowed, which means that not all individuals assigned to treatment will actually take the treatment and not all individuals assigned to control will remain untreated. Rather than determining treatment status, the assignment represents an encouragement towards treatment.

Let $D\in\left\{ 0,1\right\}$ denote the actual treatment status. Define the potential treatment associated to assignment $z$ as $D(z)$. We observe the treatment status

equation*[equation* omitted — 38 chars of source]

Let $Y\in\left\{ 0,1\right\}$ denote the observed binary outcome. The potential outcome associated to assignment $z$ is defined as $Y(D(z),z)$. At first, I allow potential outcomes depend arbitrarily on treatment and assignment. Observed and potential outcomes are related by

equation*[equation* omitted — 48 chars of source]

Let $X\in\mathcal{S}(X)$ be a vector of observed covariates and $p_{z|x}=\mathbb{P}\left ( Z=z|X=x \right )$ be the observed propensity score for assignment. I maintain the following assumption regarding the joint distribution of $(D(z),Y(D(z),z),Z,X)$ throughout the paper:

Assumption 1: For each $z,z'\in\left\{ 0,1\right\}$ and $x\in\mathcal{S}(X)$:

enumerate$\mathbb{P}\left ( D(z)=1|Z=z',X=x \right )\in \left ( 0,1 \right )$$\mathbb{P}\left ( Y(D(z),z)=1|Z=z',X=x \right )\in \left ( 0,1 \right )$$p_{z|x}>0$

Assumptions 1.1 and 1.2 state that the support of potential quantities does not depend on the assignment. Assumption 1.3 states that all individuals can be assigned to treatment and control with probability greater than zero, and is usually referred to as the common support, or overlap assumption.

The fundamental behavioral assumption for identification in IV settings restricts how individuals respond to assignment, and is formalized below:

Assumption 2: For all $x\in\mathcal{S}(X)$, we have $D(1)\geq D(0)$ conditional on $X=x$.

Assumption 2 is referred to as the monotonicity condition imbensangrist and it states that there are no individuals that would take treatment if assigned to control and would not take treatment in the presence of the encouragement. Under assumption 2 individuals can be divided into three groups regarding their response to assignment: Always-takers (individuals that take treatment regardless of the assignment), Compliers (individuals that mimick their assignment) and Never-takers (individuals that don't take treatment regardless of their assignment).

Another necessary assumption for identification is the exclusion restriction:

Assumption 3: For all $x\in\mathcal{S}(X)$ and $z\in\left\{ 0,1\right\}$, $Y(D(z),z)=Y(D(z ))$.

The exclusion restriction states that the outcome is only affected directly by the actual uptake of the treatment. Hence, assignment only affects outcomes to the extent that it affects the choice of treatment. Under the exclusion restriction, the observed outcome relates to potential outcomes simply by $Y=ZY(D(1))+(1-Z)Y(D(0))$.

Point identification in IV settings usually relies on two additional assumptions, which are stated below:

Assumption 4: For all $x\in\mathcal{S}(X)$, $\mathbb{E}\left [ D|Z=1,X=x \right ]\neq\mathbb{E}\left [D|Z=0,X=x \right ]$

Assumption 4 is a technical assumption often referred to as relevance.

Assumption 5: For all $x\in\mathcal{S}(X)$, $\left (Y(D(z)),D(z) \right )\perp Z|X=x$.

Assumption 5 states that the assignment of treatment is independent of the potential quantities. Although it is usually justified in experimental settings, it is hard to justify and verify in observational settings.

Under these five assumptions, it is well known that the average treatment effect for compliers (LATE) conditional on $X_{i}=x$ is identified by the conditional Wald estimand:

align*[align* omitted — 374 chars of source]

The unconditional LATE is identified by integrating $\tau_{Y}(x)$ and $\tau_{D}(x)$ over the distribution of covariates. In this paper, I study the partial identification of the LATE in settings where the independence assumption is violated. The approach consists in finding bounds for the first-stage $(\tau_{D}(x))$ and the reduced form $(\tau_{Y}(x))$ as functions of the magnitude of the dependence of treatment assignment on potential quantities.

The partial identification results are used to conduct sensitivity analysis and identifying breakdown frontiers, that is, the boundary between the set of assumptions which lead to a specific conclusion and those which do not. For instance, one might be interested in the values of dependence which change the conclusion that the LATE is greater than zero.

Partial Identification with Binary Outcomes

I consider the partial identification as a function of violations of independence in a setting where the outcome $Y$ is binary. In that case, the conditional LATE can be interpreted as the increase in probability of observing $Y=1$ due to the treatment for compliers with covariates $X=x$:

equation*[equation* omitted — 124 chars of source]

I begin with the partial identification of the share of compliers.

First-Stage

I begin with the partial identification of $\tau_{D}(x)$. For that purpose, write $\tau_{D}(x)=\tau_{D(1)}(x)-\tau_{D(0)}(x)$, where $\tau_{D(z)}(x)=\mathbb{E}\left [D(z)|X=x \right ]$. The parameter $\tau_{D}(x)$ can be interpreted as the share of compliers with $X=x$: $\mathbb{P}\left(D(1)>D(0)|X=x\right)$. In the case where Assumptions 1-5 and independence hold, $\tau_{D}(x)$ is identified by

equation*[equation* omitted — 106 chars of source]

which is usually referred to as the first-stage in the Wald estimand.

We replace the independence assumption by a bounded dependence assumption, called c-dependence mastenpoirier2018:

Definition: Let $x\in\mathcal{S}(X)$. Let $c_{1}$ be a scalar between 0 and 1. We say that $Z$ is $c_1$-dependent with potential treatment $D(z)$ given $X=x$ if

equation*[equation* omitted — 153 chars of source]

\

Under $c_{1}$-dependence, the unobserved propensity score is allowed to deviate $c_1$ probability units away from the observed propensity score $p_{1|x}$. For $c_{1}=0$, the assignment is independent from potential treatments (Assumption 5 holds). Throughout the paper, I assume $c_1$-dependence:

Assumption 5A: $Z$ is $c_1$-dependent with $D(1)$ given $X$ and with $D(0)$ given $X$.

Let $p_{D|z,x}=\mathbb{P}\left ( D=1|Z=z,X=x \right )$. Lemma 1 provides sharp identified sets for potential treatments and the share of compliers:

lemmaSuppose Assumptions 1-3 and 5A hold. Then, the sharp identified set for potential treatment associated to assignment $z$ is $\tau_{D(z)}(x)\in\left[\tau^{LB}_{D(z)}(c_{1},x),\tau^{UB}_{D(z)}(c_{1},x)\right]$ is, where \begin{equation*} \tau^{LB}_{D(z)}(c_{1},x)=\max\left\{ \frac{p_{D|z,x}p_{z|x}}{p_{z|x}+c_{1}},\frac{p_{D|z,x}p_{z|x}-c_{1}}{p_{z|x}-c_{1}},p_{D|z,x}p_{z|x}\right\} \end{equation*} and \begin{equation*} \tau_{D(z)}^{UB}(c_{1},x)=\min\left\{ \frac{p_{D|z,x}p_{z|x}}{p_{z|x}-c_{1}}\mathbf{1}\left ( p_{z|x}>c_{1} \right )+\mathbf{1}\left ( p_{z|x}\leq c_{1} \right ),\frac{p_{D|z,x}p_{z|x}+c_{1}}{p_{z|x}+c_{1}},p_{D|z,x}p_{z|x}+(1-p_{z|x})\right\} \end{equation*} Consequently, the sharp identified set for the share of compliers is $\tau_{D}(x)\left[\tau^{LB}_{D}(c_{1},x),\tau^{UB}_{D}(c_{1},x)\right]$, where \begin{equation*} \tau_{D}^{LB}(c_{1},x)=\max\left ( 0,\tau^{LB}_{D(1)}(c_{1},x)-\tau^{UB}_{D(0)}(c_{1},x) \right ) \end{equation*} and \begin{equation*} \tau_{D}^{UB}(c_{1},x)=\tau^{UB}_{D(1)}(c_{1},x)-\tau_{D(0)}^{LB}(c_{1},x) \end{equation*}

Lemma 1 follows directly from Proposition 5 of mastenpoirier2018. The upper bound for the first-stage is identified by the difference between the upper bound of $\tau_{D(1)}(c_{1},x)$ and the lower bound of $\tau_{D(0)}(c_{1},x)$. These are both quantities between zero and one, and the monotonicity assumptions ensures that the difference between these quantities is positive.

The lower bound is identified by the difference beteween the lower bound of $\tau_{D(1)}(c_{1},x)$ and the upper bound of $\tau_{D(0)}(c_{1},x)$. There is no guarantee that these difference is greater than zero. Since the first-stage identifies a share between zero and one, the lower bound for is restricted to be at least as great as zero.

The unconditional bounds for the first-stage are obtained by integrating the conditional bounds over the distribution of covariates:

align*[align* omitted — 241 chars of source]

Reduced Form

Now, I focus on the identification of $\tau_{Y}(x)$, which I write as $\tau_{Y}(x)=\tau_{Y(D(1))}(x)-\tau_{Y(D(0))}(x)$, where $\tau_{Y(D(z))}(x)=\mathbb{E}\left [ Y(D(z))|X=x \right ]$. The parameter $\tau_{Y}(x)$ captures the effect of assignment on potential outcomes, which is often referred to in the literature as the Intention-to-Treat effect (ITT). In the case where Assumptions 1-5 and independence hold, $\tau_{Y}(x)$ is identified by

equation*[equation* omitted — 106 chars of source]

which is referred to as the reduced form. As in the first-stage, I relax the independence assumption replace it by a c-dependence assumption of $Z_{i}$ with the potential outcomes.

Definition: Let $x\in\mathcal{S}(X)$. Let $c_{2}$ be a scalar between 0 and 1. We say that $Z$ is $c_2$-dependent with potential outcome $Y(D(z))$ given $X=x$ if

equation*[equation* omitted — 156 chars of source]

For $c_{2}=0$, the assignment is independent from potential outcomes (Assumption 5 holds). Throughout the paper, I assume $c_2$-dependence:

Assumption 5B: $Z$ is $c_2$-dependent with $Y(D(1))$ given $X$ and with $Y(D(0))$ given $X$.

Let $p_{Y|z,x}=\mathbb{P}\left ( Y=1|Z=z,X=x \right )$. I use the results from mastenpoirier2018 to derive sharp identified sets for potential outcomes and the ITT:

lemmaSuppose Assumptions 1-3 and 5B hold. Then, the sharp identified set for potential outcome associated to assignment $z$ is $\tau_{Y(D(z))}(x)\in\left[\tau^{LB}_{Y(D(z))}(c_{2},x),\tau^{UB}_{Y(D(z))}(c_{2},x)\right]$ is, where \begin{equation*} \tau^{LB}_{Y(D(z))}(c_{2},x)=\max\left\{ \frac{p_{Y|z,x}p_{z|x}}{p_{z|x}+c_{2}},\frac{p_{Y|z,x}p_{z|x}-c_{2}}{p_{z|x}-c_{2}},p_{Y|z,x}p_{z|x}\right\} \end{equation*} and \begin{equation*} \tau_{Y(D(z))}^{UB}(c_{2},x)=\min\left\{ \frac{p_{Y|z,x}p_{z|x}}{p_{z|x}-c_{2}}\mathbf{1}\left ( p_{z|x}>c_{2} \right )+\mathbf{1}\left ( p_{z|x}\leq c_{2} \right ),\frac{p_{Y|z,x}p_{z|x}+c_{2}}{p_{z|x}+c_{2}},p_{Y|z,x}p_{z|x}+(1-p_{z|x})\right\} \end{equation*} Consequently, the sharp identified set for the share of compliers is $\tau_{Y}(x)\left[\tau^{LB}_{Y}(c_{2},x),\tau^{UB}_{Y}(c_{2},x)\right]$, where \begin{equation*} \tau_{Y}^{LB}(c_{2},x)=\max\left ( 0,\tau^{LB}_{Y(D(1))}(c_{2},x)-\tau^{UB}_{Y(D(0))}(c_{2},x) \right ) \end{equation*} and \begin{equation*} \tau_{Y}^{UB}(c_{2},x)=\tau^{UB}_{Y(D(1))}(c_{2},x)-\tau_{Y(D(0))}^{LB}(c_{2},x) \end{equation*}

The bounds are similar to those obtained for the first-stage. Note that the ITT is not bounded by definition between zero and one, and thus, there is no need to constraint the lower bound to be at least as great as zero.

The unconditional bounds for the ITT are identified by integrating the conditional bounds over the distribution of covariates:

align*[align* omitted — 230 chars of source]

Local Average Treatment Effect

The Local Average Treatment Effect (LATE), the average treatment effect for the subgroup of compliers, under the standard IV assumptions, is point-identified by the ratio of the reduced form and the first-stage. Replacing Assumption 5 with Assumptions 5A and 5B, we obtain the following bounds for the LATE. Putting the pieces from Sections 3.1 and 3.2 together, we find that $\tau(x)\in\left[\tau^{LB}(c_{1},c_{2},x),\tau^{UB}(c_{1},c_{2},x)\right]$, with

align*[align* omitted — 217 chars of source]

In the case of a binary outcome, treatment effects are not greater than $1$ nor smaller than $-1$. Hence, the lower bound is the greatest value between the ratio of the lower bound of the ITT and the upper bound of the first-stage, and $-1$. The upper bound is the smallest value between the ratio of the upper bound of the ITT and the lower bound of the first-stage, and $1$. The unconditional bounds for the LATE are identified by integrating the conditional LATE bounds over the distribution of covariates:

equation*[equation* omitted — 147 chars of source]

and

equation*[equation* omitted — 146 chars of source]

The bounds for the LATE are functions of violations of instrument independence with respect to both potential treatments and potential outcomes. In that sense, this result is similar to the partial identification result presented in Section 5.3 of jonashesh, which partially identifies the LATE as a function of the finite-population covariance between the assignment probabilities and the potential outcomes and treatments. In the next section, I show how researchers can identify the set of violations of independence under which a conclusion is invalidated.

Breakdown Analysis

The fundamental interest here is to understand under which violations of independence a particular conclusion still hold. In this section I define breakdown points for conclusions regarding the firs-stage and the reduced form separately, and a breakdown frontier for conclusions regarding the LATE.

I begin by defining the breakdown point for conclusions of the first-stage. That is, what is the largest value of $c_1$ under which we can conclude that $\mathbb{P}\left(D(1)>D(0)\right)\geq\mu$? First, define the robust region for the conclusion as the set of values of $c_{1}$ under which the conclusion holds:

equation*[equation* omitted — 107 chars of source]

The robust region for the first-stage is the set of values of $c_{1}$ for which the identified set for the share of compliers is above $\mu$. The breakdown point is the value of $c_{1}$ on the boundary of the robust region. Formally, the breakdown point is $c_{1}^{*}$ such that $\tau^{LB}_{D}(c_{1}^{*})=\mu$. The breakdown point is identified by solving the expression $\int\tau^{LB}_{D(1)}(c_{1},x)dF_{X}(x)-\int\tau^{UB}_{D(0)}(c_{1},x)dF_{X}(x)=\mu$ for $c_{1}$. Let $bp_{FS}(\mu)$ denote the solutions of the expression above. The analytical expression for the breakdown point is

equation*[equation* omitted — 87 chars of source]

That is, the breakdown point is the smallest value of $c_{1}$ which solves the breakdown equation, as long as it is bounded in the unit interval. If that is not case, then the breakdown point depends on the worst-case bounds. If the worst-case bounds lie within the robust region, then the breakdown point is 1. Otherwise, it is equal to 0.

Similarly, when it comes to the reduced form, the breakdown point is the largest value of $c_{2}$ under which one can conclude that $\mathbb{E}\left[Y(D(1))-Y(D(0))\right]\geq\mu$. It is defined implicitly by $c_{2}^{*}$ such that $\tau^{LB}_{Y}(c_{2}^{*})=\mu$. Analogously to the first-stage, the breakdown point for the ITT is identified by solving the expression $\int\tau^{LB}_{Y(D(1))}(c_{2},x)dF_{X}(x)-\int\tau^{UB}_{Y(D(0))}(c_{2},x)dF_{X}(x)=\mu$ for $c_{2}$. Let $bp_{RF}(\mu)$ denote the solutions to the expression above, we obtain the following expression for the breakdown point:

equation*[equation* omitted — 87 chars of source]

When it comes to the LATE, the parameter is partially identified as function of both sensitivity parameters $c_{1},c_{2}$. In that case, the robust region of identification is the set of values of $c_{1}$ and $c_{2}$ under which the conclusion holds. It is defined as

equation*[equation* omitted — 131 chars of source]

The breakdown frontier is the set of values of $c_{1}$ and $c_{2}$ on the boundary of the robust region. Specifically, the breakdown frontier is defined as

equation*[equation* omitted — 128 chars of source]

Consider the case where $\mu\neq0$. The breakdown frontier can be identified as a function of $c_{1}$ by solving the following equation for $c_{2}$:

equation*[equation* omitted — 143 chars of source]

Let $bf(c_{1},\mu)$ denote the solutions, we obtain the following expression for the breakdown frontier as a function of $c_{1}$:

equation*[equation* omitted — 92 chars of source]

Note that in the case where $\mu=0$, the breakdown frontier collapses to the breakdown point for the ITT. That is, $BF(c_{1},0)=bp_{RF}(0)$.

The shape of the breakdown frontier provides insights on the tradeoffs between these two types of relaxations of independence when researchers are drawing conclusions. The relaxations $c_1$ and $c_2$ are measured in the same unit, which helps the interpretation of the breakdown analysis.

Numerical Illustration

I illustrate the breakdown approach with a simple numerical illustration. Let $X$ be a binary covariate that follows a Bernoulli distribution with parameter $p_{x}=0.5$ Let $Z$ denote the instrument which also follows a Bernoulli distribution with parameter $p_{z}$. Let $p_{z|x}=0.5$, for the sake of simplicity. Selection into treatment follows a threshold-crossing model as in vytla:

equation*[equation* omitted — 72 chars of source]

where $V$ has a standard uniform distribution. The parameter $\pi_{z}$ is the share of compliers, and is set to be equal to 0.5.

The binary outcome also follows a threshold model, as in Heckman (1978):

equation*[equation* omitted — 76 chars of source]

where $U$ is also uniformly distributed. The random variables $U$ and $V$ are linearly correlated as in olsen, which generates the selection problem. The parameter $\beta_{d}$ is the LATE in this DGP, which is set to be equal to 0.5. Hence, it follows that the ITT is equal to 0.25.

figure[figure omitted — 563 chars of source]

Figure 1 (a) shows the identified set for the first-stage as a function of $c_{1}$. The breakdown point for the conclusion that the share of compliers is greater than zero is $c_{1}^{*}=0.25$. Figure 1 (b) show the identified set of the ITT as a function of $c_{2}$. The breakdown point for the conclusion that the ITT, and hence the LATE, is greater than zero is $c_{1}^{ *}=0.1$.

Figure 2 shows the breakdown frontier for the LATE. I specify the breakdown frontier for the conclusion that the LATE is greater than 0.25, which is half of its true value. The blue area represents the robust region, that is, this is the set of values for violations of independence under which the conclusion still holds.

figure[figure omitted — 276 chars of source]

Estimation and Inference

In this section I study estimation and inference on the identified sets and breakdown values defined above. The estimands for the bounds of assignment effects, the LATE and the breakdown values are functionals of the conditional cdf of outcomes and treatment given assignment and covariates, the probability of assignment given covariates, and the marginal distribution of the covariates. I propose nonparametric sample analog estimators and derive asymptotic distributional results using a delta method for directionally differentiable functionals. First, I assume we observe a random sample of data.

Assumption 6: The random variables $\left\{ \left ( Y_{i},D_{i},Z_{i},X_{i} \right )\right\}_{i=1}^{n}$ are independently and identically distributed according to the distribution of $\left(Y,D,Z,X\right)$.

Furthermore, I assume the support of covariates is discrete.

Assumption 7: The support of X is discrete and finite. Let $\mathcal{S}(X)=\left\{ x_{1},...,x_{K}\right\}$ up to a finite $K$.

All target parameters are functionals of the underlying parameters $p_{Y|z,x}$, $p_{D|z,x}$, $p_{z|x}$ and $q_{x}=\mathbb{P}\left(X=x\right)$. Their parametric estimators are, respectively,

equation*[equation* omitted — 219 chars of source]
equation*[equation* omitted — 219 chars of source]
equation*[equation* omitted — 174 chars of source]

and

equation*[equation* omitted — 96 chars of source]

These quantities converge uniformly to a Gaussian process at a $\sqrt{N}$-rate; see Lemma B1 in Appendix B. Next, consider the bounds for potential treatments and outcomes. I estimate these bounds by a plug-in estimator of the quantities above. First, I introduce an additional assumption:

Assumption 8: For all $x\in\mathcal{S}(X)$, we have $\max\left\{ c_{1},c_{2}\right\}<\min\left\{ p_{1|x},p_{0|x}\right\}$ and $\tau^{LB}_{D}(c_{1},x)>0$.

Assumption 8 is a technical assumption which simplifies the bounds for potential treatments to

align*[align* omitted — 326 chars of source]

and the bounds for potential outcomes to

align*[align* omitted — 332 chars of source]

This simplification is particularly important to guarantee Hadamard directional differentiablity of the estimators. Also, it modifies that standard relevance assumption to the partially identified case, ensuring that the upper bound for the LATE is not subject to the weak instrument problem.

The bounds for potential quantities are estimated by replacing the population quantities in the expressions above by its sample analogues. In Lemma B2 of Appendix B I show that the estimators for potential quantities converge in distribution to a nonstandard distribution given by a continuous transformation of Gaussian processes. This result is the building block for deriving the asymptotic properties of the estimators for bounds and breakdown values.

First, consider the bounds for the first-stage. The plug-in estimators for the lower and the upper bound, respectively are

align*[align* omitted — 232 chars of source]

The unconditional bounds are estimated by integrating the estimates of conditional bounds over the empirical distribution of covariates:

equation*[equation* omitted — 113 chars of source]

and

equation*[equation* omitted — 114 chars of source]

Now consider the bounds for the reduced form. The plug-in estimators are

align*[align* omitted — 243 chars of source]

The unconditional bounds are obtained by integrating the estimates over the empirical distribution of covariates:

equation*[equation* omitted — 113 chars of source]

and

equation*[equation* omitted — 113 chars of source]

The asymptotic distribution of the estimator for these bounds is formalized in the proposition below:

propositionAssume Assumptions 1-3, 5A, 5B and 6-8 hold. Then, \begin{equation*} \sqrt{N}\begin{pmatrix} \widehat{\tau}_{D}^{LB}(c_{1})-\tau^{LB}_{D}(c_{1}) \\\widehat{\tau}_{D}^{UB}(c_{1})-\tau_{D}^{UB}(c_{1}) \end{pmatrix}\overset{d}{\rightarrow}Z_{FS}(d,z,x,c_{1}) \end{equation*} and \begin{equation*} \sqrt{N}\begin{pmatrix} \widehat{\tau}_{Y}^{LB}(c_{2})-\tau^{LB}_{Y}(c_{2}) \\\widehat{\tau}_{Y}^{UB}(c_{2})-\tau_{Y}^{UB}(c_{2}) \end{pmatrix}\overset{d}{\rightarrow}Z_{RF}(y,z,x,c_{2}) \end{equation*} where $\textbf{Z}_{FS}(d,z,x,c_{1})$ and $\textbf{Z}_{RF}(y,z,x,c_{2})$ are Gaussian Elements defined in the Section 1 of Appendix A.

Now consider the estimation for the breakdown point for the claim that $\tau_{D}\geq \mu$. We focus on that case where $\tau^{LB}_{D}(0)>\mu$, which implies that $c_{1}^{*}>0$. Define the estimator for the breakdown point $c_{1}^{*}$ as

equation*[equation* omitted — 128 chars of source]

and it is obtained by replacing the population quantities from $bf_{FS}(\mu)$ with the sample analogues defined in this section. The estimator for the breakdown frontier for the first-stage is thus $\widehat{c}_{1}^{*}=\min\left\{ \max\left\{ \widehat{b}f_{FS}(\mu),0\right\},1\right\}$

Similarly, when it comes to the estimation for the breakdown point for the claim that $\tau_{Y}\geq \mu$, define the estimator for the breakdown point $c_{2}^{*}$ as

equation*[equation* omitted — 128 chars of source]

which is obtained by replacing the population quantities from $bf_{RF}(\mu)$ with the sample analogues. The estimator for the breakdown frontier for the first-stage is thus $\widehat{c}^{*}_{2}=\min\left\{ \max\left\{ \widehat{b}f_{FS}(\mu),0\right\},1\right\}$.

I now provide a formal result regarding the asymptotic distribution of $\widehat{c}_{1}^{*}$ and $\widehat{c}_{2}^{*}$:

theoremAssume Assumptions 1-3, 5A, 5B and 6-8 hold. Furthermore, assume that $c_{1},c_{2}\in\left [ 0,\overline{C} \right ]$. Then, \begin{equation*} \sqrt{N}\left ( \widehat{c}_{1}^{*}-c_{1}^{*} \right )\overset{d}{\rightarrow}Z_{FS}^{BP} \end{equation*} and \begin{equation*} \sqrt{N}\left ( \widehat{c}_{2}^{*}-c_{2}^{*} \right )\overset{d}{\rightarrow}Z_{RF}^{BP} \end{equation*} where $\textbf{Z}_{FS}^{BP}$ and $\textbf{Z}_{RF}^{BP}$ are Gaussian random variables defined in Section 2 of Appendix A.

Finally, consider the estimators for the bounds and breakdown frontier of the LATE. The estimators for the lower and upper bound are respectively

equation*[equation* omitted — 146 chars of source]

and

equation*[equation* omitted — 145 chars of source]

The next lemma formalizes the asymptotic distribution of the bounds:

propositionSuppose Assumptions 1-3, 5A, 5B and 6-8 hold. Then, \begin{equation*} \sqrt{N}\begin{pmatrix} \widehat{\tau}^{LB}(c_{1},c_{2})-\tau^{LB}(c_{1},c_{2}) \\ \widehat{\tau}^{UB}(c_{1},c_{2})-\tau^{UB}(c_{1},c_{2}) \end{pmatrix}\overset{d}{\rightarrow}Z_{\tau}(y,d,z,x,c_{1},c_{2}) \end{equation*}

Denote the estimated breakdown frontier for the conclusion that $\tau\geq\mu$ by

equation*[equation* omitted — 112 chars of source]

\

I show that the estimated breakdown frontier converges in distribution.

theoremLet Assumptions 1-3, 5A, 5B and 6-8 hold. Furthermore, let $\mathcal{C}\subset \left [ 0,\overline{C} \right ]$ and $\mathcal{M}\subset \left [ -1,1 \right ]$ be finite grids of points. Then, \begin{equation*} \sqrt{N}\left ( \widehat{BF}(c_{1},\mu) -BF(c_{1},\mu)\right )\overset{d}{\rightarrow}Z_{BF}(c_{1},\mu), \end{equation*} a tight random element of $l^{\infty}\left ( \mathcal{C}\times\mathcal{M} \right )$.

Since the limiting process is non-Gaussian, inference on the breakdown values will not be based on standard errors. The processes’ distribution is characterized fully by the expressions in Appendices B1 and B2, but obtaining analytical estimates of functionals of these processes is challenging. In the next subsection I give describe a bootstrap procedure that can be used to construct confidence intervals for the breakdown points and confidence bands for the breakdown frontier.

The breakdown points and frontier estimators can be obtained using standard root finding algorithms, such as Matlab's fzero and R's findZeros. The solutions provide the estimates.

Bootstrap Inference

Before describing the procedure, I introduce some notation. Let $\mathcal{F}_{i}=\left(Y_{i},D_{i},Z_{i},X_{i}\right)$ and $F_{1:N}=\left\{\mathcal{F}_{1},...,\mathcal{F}_{N}\right\}$. Let $\widehat{\theta}$ denote the estimator of a parameter $\theta_{0}$ based on $\mathcal{F}_{1:N}$. Let $\textbf{A}^{*}_{1:N}\equiv\sqrt{N}\left(\widehat{\theta}^{*}-\widehat{\theta}\right)$ where $\widehat{\theta}^{*}$ is a draw from the nonparametric bootstrap distribution of $\widehat{\theta}$. Suppose $\textbf{A}$ is the tight limiting process of $\sqrt{N}\left(\widehat{\theta}-\theta_{0}\right)$. Bootstrap consistency is given by weak convergence in probability conditional on $\mathcal{F}_{1:N}$, that is,

equation*[equation* omitted — 173 chars of source]

where $BL_1$ denotes the set of Lipschitz functions into $\mathbb{R}$ with Lipschitz constant no greater than 1. For $\theta_{0}$ and $\widehat{\theta}$, I focus on the parameters introduced in Section 3 and their sample analogue estimators, which are plugged-in in the bounds estimators.

Let $\textbf{Z}_{1:N}=\sqrt{N}\left(\widehat{\theta}^{*}-\widehat{\theta}\right)$. Theorem 3.6.1 of van der Vaart and Wellner (1996) implies the bootstrap consistency of $\textbf{Z}_{1:N}$. Since the parameters of interest are Hadamard differentiable functionals of $\theta_{0}$, it follows from fangsantos that the nonparametric bootstrap can be used to do inference on the identified sets and breakdown values.

For the breakdown points of the first-stage and the reduced form, the bootstrap procedure can be used to construct one-sided confidence intervals as in klinesantos. For the breakdown frontier of the LATE, the bootstrap can be used to construct uniform one-sided confidence bands as in mastenpoirier2020.

Partial Identification under joint $c$-dependence

The results from Sections 3 and 4 provide the breakdown analysis framework for IV settings with binary outcomes in the case where the assumption of instrument independence is replaced by a bounded dependence assumption, that consider the probability of assignment conditional on potential treatments and potential outcomes separately.

Relaxing the independence assumption in terms of $c_1$ and $c_{2}$-dependence allows the researcher to conduct breakdown analysis for the share of compliers and the ITT separately, while also allowing for the possibility of assessing tradeoffs between these assumptions in the breakdown frontier for the LATE.

Despite the several desirable features of this approach, it does not provide a sharp identified set for the LATE. In this section, I derive sharp bounds for the LATE under a joint $c$-dependence assumption, which I define below:

Definition: Let $x\in\mathcal{S}(X)$. Let $c$ be a scalar between 0 and 1. We say that $Z$ is joint $c$-dependent with $\left(Y(D(z)),D(z)\right)$ given $X=x$ if

equation*[equation* omitted — 166 chars of source]

As in Section 3, the sensitivity parameter captures deviations from the independence assumption in terms of the distance in probability units between the probability of assignment given covariates and the probability of assignment given covariates and potential quantities. Note that $c=0$ is the case where independence holds, and the target parameters in the setting are point identified. Throughout this section, I assume joint $c$-dependence:

Assumption 9: $Z$ is joint $c$-dependent with $\left(Y(D(1)),D(1)\right)$ given $X$ and $\left(Y(D(0)),D(0)\right)$ given $X$. Next, I derive the sharp identified set for potential quantities and the LATE:

theoremFor a random variable $Q\in\left\{Y,D\right\}$, denote its potential value associated to assignment $z$ as $Q(z)$. Suppose Assumptions 1-3 and 6-9 hold. The sharp identified set for potential quantities is $\tau_{Q(z)}(x)\in\left [ \tau^{LB}_{Q(z)}(c,x),\tau^{UB}_{Q(z)}(c,x) \right ]$, where \begin{align*} \tau^{LB}_{Q(z)}(c,x)=\max \left\{ \frac{p_{Q|z,x}p_{z|x}}{p_{z|x}+c},\frac{p_{Q|z,x}p_{z|x}-2c}{p_{z|x}-c},p_{Q|z,x}p_{z|x}\right\} \end{align*} and \begin{align*} \tau^{UB}_{Q(z)}(c,x)=\min \left\{ \frac{p_{Q|z,x}p_{z|x}}{p_{z|x}-c},\frac{p_{Q|z,x}p_{z|x}+2c}{p_{z|x}+c},p_{Q|z,x}p_{z|x}+(1-p_{z|x})\right\} \end{align*} Consequently, the sharp identified set for the LATE is $\tau(x)\in\left [ \tau^{LB}(c,x),\tau^{UB}(c,x)\right ]$, where \begin{equation*} \tau^{LB}(c,x)=\frac{\tau^{LB}_{Y(D(1))}(c,x)-\tau^{UB}_{Y(D(0))}(c,x)}{\tau^{UB}_{D(1)}(c,x)-\tau^{LB}_{D(0)}(c,x)} \end{equation*} and \begin{equation*} \tau^{UB}(c,x)=\frac{\tau^{UB}_{Y(D(1))}(c,x)-\tau^{LB}_{Y(D(0))}(c,x)}{\tau^{LB}_{D(1)}(c,x)-\tau^{UB}_{D(0)}(c,x)} \end{equation*}

Theorem 3 provides the partial identification results for the LATE as a function of the sensitivity parameter $c$. As in Section 4, the bounds for the LATE can be used to identify the breakdown point to a particular conclusion under joint $c$-dependence. Estimation and inference procedures for this case are similar to the ones presented in Section 6. Once again, the bounds for the unconditional LATE are identified by integrating the conditional bounds over the distribution of covariates.

When it comes to estimation, the nonparametric estimators for conditional probabilities defined in Section 6 can be used as plug-ins to build an estimator for the bounds of the LATE under joint $c$-dependence. The proposition below shows that the estimators of the bounds converge to Gaussian elements:

propositionSuppose Assumptions 1-3 and 6-9 hold. Then, \begin{equation*} \sqrt{N}\begin{pmatrix} \widehat{\tau}^{LB}(c)-\tau^{LB}(c) \\ \widehat{\tau}^{UB}(c)-\tau^{UB}(c) \end{pmatrix}\overset{d}{\rightarrow}Z_{j,\tau}(y,d,z,x,c) \end{equation*}

Now, turn to the breakdown point for the conclusion that the LATE is equal or greater than $\mu$, which is denoted by $c^{*}$. Let

equation*[equation* omitted — 111 chars of source]

denote the estimated breakdown point. The next result formally present a result about the asymptotic distribution of $\widehat{c^{*}}$:

theoremSuppose Assumptions 1-3 and 6-9 hold. Furthermore, assume that $c\in\left [ 0,\overline{C} \right ]$. Then, \begin{equation*} \sqrt{N}\left ( \widehat{c}^{*}-c^{*} \right )\overset{d}{\rightarrow}Z_{j}^{BP} \end{equation*} where $\textbf{Z}^{BP}_{j}$ is a Gaussian random variable defined in Section 7 of Appendix A.

When it comes to inference, the same Bootstrap procedures introduced in Section 6.1 can be used in the case of joint $c$-dependence to perform valid inference over the bounds for the LATE and the breakdown point.

Monte Carlo Simulations

In this section I perform Monte Carlo exercises to study the properties of the estimators for the breakdown points under the DGP described in Section 5. I consider a sample size $n$ equal to 1000 and conduct 1000 Monte Carlo simulations to study the performance of the estimators for breakdown points regarding different conclusions.

To focus the nondegenerate case, I only consider claims that yield breakdown points greater than zero.

I analyze the performance of the estimators for the breakdown points of conclusions regarding the share of compliers and the ITT under separate $c_{1}$ and $c_{2}$-dependence, and conclusions regarding the LATE under joint $c$-dependence. Table 1 displays the results. The estimators are analyzed in terms of their average and median bias, and the 95 % coverage of the one-sided confidence interval from klinesantos.

Overall, the estimators exhibit desirable finite-sample properties, expressed in terms of close to zero finite-sample bias and empirical coverage of the confidence interval being close to the target 95 %. The performance of the estimators is stable across different values of $\mu$.

table[table omitted — 896 chars of source]

Empirical Application: Family Size and Employment

In this Section, I use the estimators from Sections 6 and 7 to perform the breakdown analysis for the results regarding family size and female employment in angev, using data from the US Census Public Use Microsamples married mothers aged 21–35 in 1980 with at least 2 children and oldest child less than 18.

In this setting, the dependent variable is and indicator for women who worked for pay in 1979. Treatment is an indicator for women having three or more children, and the instrument is an indicator for women whose first two children have the same sex. The authors control for age, age at the first birth, race and sex of the first two children as covariates.

To begin the sensitivity analysis, I use selection on observables to to calibrate the beliefs regarding the amount of selection on unobservables. I take the approach from altonji and mastenpoirier2018. I partition the vector of covariates $X$ as $(X_{k},X_{-k})$, where $X_{k}$ is the $k$-th component and $X_{-k}$ is a vector with remaining components. The measures used to calibrate the beliefs regarding deviations from independence are

equation*[equation* omitted — 171 chars of source]

In the data, the largest value obtained form $\overline{c}_{k}$ is associated to to the indicator for women whose second child is a man, which was estimated to be $\overline{c}_{2nd\ sex}=0.011$.

figure[figure omitted — 532 chars of source]

Picture 3 shows the identified sets for the share of compliers and the ITT in the application. The share of compliers in the case where $c_{1}=0$ is approximately 0.060 and the estimated breakdown point for the conclusion that the share of compliers is greater than zero is $\widehat{c}_{1}^{*}=0.037$. The identified set for the ITT is displayed on the left. If point identification holds, then the ITT is equal to -0.008. However, the identified set is uninformative for most of the values of $c_{2}$. The estimated breakdown point for the conclusion that treatment effects are negative is $\widehat{c}_{2}^{*}=0.004$. However, it is not statistically different from zero.

figure[figure omitted — 346 chars of source]

Figure 4 shows the identified set for the LATE under joint $c$-dependence. In the case of full independence of the instrument ($c=0)$, the LATE is equal to -0.132. The estimated breakdown point for the conclusion that the LATE is negative is $\widehat{c}^{*}=0.004$, and it is not statistically different from zero.

Overall, the results from the breakdown analysis suggest that the conclusions regarding the effects of family size on unemployment using same-sex siblings as the instrument are not robust to violations of independence of the instrument.

Conclusion

In this paper I discuss the partial identification of treatment effects in IV settings with binary outcomes under violations of independence. I derive identified sets for the first-stage, the reduced form and the LATE. Building on this result, I identify breakdown values for conclusions regarding these parameters. I derive the asymptotic properties for the estimators of the bounds and the breakdown values. I also derive sharp bounds for the LATE under a joint $c$-dependence assumption.

Monte Carlo simulations show the desirable properties of the estimators for the breakdown points and the bootstrap procedure for constructing one-sided confidence intervals. This is still a work in progress. In the empirical application I study the effects of family size on female employment and find that weak conclusions about the share of compliers to the same-sex sibling instrument and treatment effects are highly sensitive to relaxations of random assignment of the instrument.