EconBase
← Back to paper

Sensitivity of LATE Estimates to Violations of the Monotonicity Assumption

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

108,866 characters · 34 sections · 49 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\selectlanguage{english}

titlepage\thispagestyle{empty} \begin{verbatim} \end{verbatim} \begin{center} { Sensitivity of LATE Estimates to}\\ {Violations of the Monotonicity Assumption} \end{center} \begin{flushleft} \begin{center} Claudia Noack\footnote[$\star$]{This version: June 10, 2021. Department of Economics, University of Mannheim; E-mail: [email removed]. Website: claudianoack.github.io. I am grateful to my advisor Christoph Rothe for invaluable support on this project. I am thankful for insightful discussions with Tim Armstrong, Matthew Masten, Yoshiyasu Rai, and Ed Vytlacil. Furthermore, I thank Michael Dobrew, Jasmin Fliegner, Martin Huber, Paul Goldsmith-Pinkham, Sacha Kapoor, Lukas Laffers, Tomasz Olma, Vitor Possebom, Alexandre Poirier, Jonathan Roth, Pedro S'Antanna, Konrad Stahl, Matthias Stelter, Jörg Stoye, Philipp Wangner, and seminar and conference participants at the University of Mannheim, University of Heidelberg, University of Bonn, University of Oxford, Yale University, ESEM 2019, ESWC 2020, IAAE 2019, and 1st International Ph.D. Conference at Erasmus University Rotterdam for helpful comments and suggestions. I gratefully acknowledge financial support by the European Research Council (ERC) through grant SH1-77202.} \\ \end{center} \end{flushleft} \begin{abstract} In this paper, we develop a method to assess the sensitivity of local average treatment effect estimates to potential violations of the monotonicity assumption of Imbens and Angris (1994). We parameterize the degree to which monotonicity is violated using two sensitivity parameters: the first one determines the share of defiers in the population, and the second one measures differences in the distributions of outcomes between compliers and defiers. For each pair of values of these sensitivity parameters, we derive sharp bounds on the outcome distributions of compliers in the first-order stochastic dominance sense. We identify the robust region that is the set of all values of sensitivity parameters for which a given empirical conclusion, e.g. that the local average treatment effect is positive, is valid. Researchers can assess the credibility of their conclusion by evaluating whether all the plausible sensitivity parameters lie in the robust region. We obtain confidence sets for the robust region through a bootstrap procedure and illustrate the sensitivity analysis in an empirical application. We also extend this framework to analyze treatment effects of the entire population. \end{abstract} \setcounter{page}{0}

\pagenumbering{arabic} \onehalfspacing

Introduction

The local average treatment effect framework (LATE) is used for instrumental variable analysis in setups of heterogeneous treatment effects imbens1994identification. We consider settings of a binary instrumental variable and a binary treatment variable. The Wald estimand then equals the treatment effect of compliers, individuals for which the instrument influences the treatment status, given the well-known classical LATE assumptions: monotonicity, independence, and relevance. Monotonicity states that the effect of the instrument on the treatment decision is monotone across all units. In the canonical example, in which the instrument encourages units to take up the treatment, monotonicity rules out the existence of defiers, i.e., units that receive the treatment only if the instrument discourages them. Researchers might question the validity of this assumption in empirical applications. In these settings, the local treatment effect estimates might be biased and might lead the researchers to draw incorrect conclusions about the true treatment effect.

As an example of a setup in which monotonicity could plausibly be violated, consider the study of angristevens1998, who analyze the effect of having a third child on the labor market outcomes of mothers. As the decision to have a third child is endogenous, the authors use a dummy for whether the first two children are of the same sex as an instrument. The underlying reasoning is that some parents would only decide to have a third child if their first two children were of the same sex; these parents are compliers. The monotonicity assumption seems questionable in this setting as parents, who have a strong preference for one specific sex, might act as a defier in this setup. Consider, for example parents who want to have at least two boys and their first child is a boy. Contrary to the incentive given by instrument, they have two children if their second child is a boy, and three children if their second child is a girl. As the monotonicity assumption might be questionable in this example, one can question the validity of empirical conclusions drawn from the classical LATE analysis.\footnote{The other LATE assumptions seem to be plausible here. As the sex of a child is determined by nature and as only the number of and not the sex of the child arguably influences the labor market outcome of mothers, the independence assumption seems to be satisfied. The relevance assumption is testable.}

In this paper, we provide a framework to evaluate the sensitivity of treatment effect estimates to a potential violation of the monotonicity assumption. As noted in AngristImbensRubin1996, a violation of the monotonicity assumption always has two dimensions: The first dimension is the heterogeneous effect of the instrumental variable on the treatment variable, the presence of defiers. The second dimension is the heterogeneous effect of the treatment variable on the outcome variable, the outcome heterogeneity between defiers and compliers. We derive the degree to which monotonicity is violated by parameterizing these two dimensions.

We parameterize the existence of defiers by their population size and the outcome heterogeneity by the Kolmogorov-Smirnov norm, which bounds the difference of the cumulative distribution functions of compliers and defiers. For each of these two sensitivity parameters, we identify sharp bounds of the outcome distribution of compliers in a first-order stochastic dominance sense. These bounds also imply sharp bounds on various treatment effects, e.g., the average treatment effect or quantile treatment effects of compliers.

Our analysis proceeds in two steps. In a first step, we identify the sensitivity region. The sensitivity region defines the set of sensitivity parameters for which a data generating process exists, that is consistent with our model assumptions and implies both the observed probabilities and the sensitivity parameters. Since sensitivity parameters lying in the complement of the sensitivity region are not compatible with our model, we do not analyze them further. For the derivation of the sensitivity region, we also derive sharp bounds of the population size of defiers. In a second step, we identify the robust region, which is the set of sensitivity parameters that imply treatment effects that are consistent with a particular empirical conclusion; for instance, the treatment effect of compliers has a specific sign or a particular order of magnitude.\footnote{See masten2020 for a detailed exposition of this approach.} Parameters lying in the complement of the robust region, the nonrobust region, imply treatment effects that are not, or may not be, consistent with the given empirical conclusion. The robust region and the nonrobust region are separated from each other by the breakdown frontier, following the terminology of masten2020. For each population size of defiers, the breakdown frontier identifies the weakest assumption about outcome heterogeneity, which is necessary to be imposed to imply treatment effects being consistent with the particular empirical conclusion under consideration.

This framework can be used in the following ways. First, by evaluating the size of the sensitivity region, one can determine the plausibility of the model. If this set is empty, the model is refuted, which implies that even if one would allow for an arbitrary violation of the monotonicity assumption, at least one of the model assumptions has to be violated. Second, researchers can analyze the sensitivity of their estimates with respect to the degree to which the monotonicity assumption is violated by varying the sensitivity parameters within the sensitivity region. Third, by evaluating the plausibility of the parameters within the robust region, researchers can assess the sign or the order of magnitude of the treatment effect. While being transparent about the imposed assumptions, they might still arrive at a particular empirical conclusion of interest in a credible way. Fourth, one can assess to which degree monotonicity has to be violated to overturn a particular empirical conclusion. Within our framework, researchers can use their economic insights about the analyzed situation to judge the severity of a violation monotonicity.

While the main focus of this paper lies on treatment effects of compliers, we also show how this framework can be exploited to analyze treatment effects of the entire population. Under further support assumptions of the outcome variable and for given sensitivity parameters, the average treatment effect of the entire population is partially identified, which complements known results in the literature kitagawa2021identification, Balke97, machado2019instrumental, kamat2018identifying. Since the analytic expressions of the sensitivity and robust regions are rather complicated and difficult to interpret, we provide simplified analytical expressions of these regions for a binary outcome.

To construct confidence sets for both the sensitivity and the robust region, we show that both regions are determined through mappings of some underlying parameters. These mappings are not Hadamard-differentiable, and inference methods relying on standard Delta method arguments are therefore not applicable. We show how to construct smooth mappings that bound the parameters of interest. This construction leads to mappings for which standard Delta method arguments are applicable, and we use the nonparametric bootstrap to construct valid confidence sets for the parameters of interest. With a binary outcome variable, the mappings resulting in the sensitivity and robust region are considerably simpler. Therefore, we can use a generalized Delta method to show asymptotic distributional results and apply a bootstrap procedure to construct asymptotically valid confidence sets.

We show in a Monte Carlo study that our proposed inference method has good finite sample properties. We further apply our method to the setup studied by angristevens1998 introduced above. We show that relatively strong assumptions on either the population size of the defiers or the outcome heterogeneity have to be imposed to preserve the sign of the estimated treatment effect. This result demonstrates that the monotonicity assumption is key in the local treatment effect framework.

The remainder of this paper is structured as follows: A literature review follows, and Section (ref) illustrates the setup in a simplified setting. Section (ref) introduces the sensitivity parameters and Section (ref) derives sharp bounds on the distribution functions of compliers. The main sensitivity analysis is presented in Section (ref). Section (ref) discusses extensions and Section (ref) derives estimation and inference results. Section (ref) contains a simulation study and Section (ref) an empirical example. Section (ref) concludes. All proofs and additional materials are deferred to the appendix.

Literature

This paper relates to several strands of the literature. First, this paper contributes to the growing strand of the literature, which considers sensitivity analysis in various applications. These applications include, among many others, violations of parametric assumptions, violations of moment conditions, and multiple examples within the treatment effect literature armstrong2021sensitivity, mukhin2018sensitivity, christensen2019counterfactual, kitamura2013robustness, bonhomme2018minimizing, bonhomme2019posterior, andrews2017measuring, andrews2020informativeness, andrews2020model, andrews2020transparency, rambachan2020honest, conley2012plausibly, imbens2003, chen2015sens. This paper is closely related to the literature about breakdown points of HorowitzManski1995, imbens2003, KleinSantos2013, stoye2005, stoye2010partial, and especially closely related to masten2020, masten2021salvaging. These papers consider several assumptions in the treatment effect literature, but not the monotonicity assumption.

Second, it is related to the local average treatment effect framework literature, which is formally introduced in imbens1994identification and further in vytlacil2002independence. Several papers consider violations of the monotonicity assumption through different types of assumptions. Balke97, machado2019instrumental, hubermellace2010, manski1990, huber2015testing, huber2015testmonotonicity consider a binary and kitagawa2021identification a continuous outcome variable and partially identify the average treatment effect. small2017instrumental, manskipepper2000,dahl2017s, hubermellace2010 propose alternative assumptions on the data generating process, which are strictly weaker than monotonicity and obtain bounds on various treatment effects. richardson2010analysis consider a binary outcome variable and derive bounds on outcome distributions for a given population size of always takers. We consider not only a binary outcome variable and we introduce a second parameter to also bound outcome heterogeneity. We then consider these sensitivity parameters within the framework of a breakdown frontier.

de2017tolerating shows that in the presence of defiers, under certain assumptions, the Wald estimand still identifies a convex combination of causal treatment effects of only a subpopulation of compliers. In a policy context, the treatment effect of compliers might be of particular interest because the treatment status of compliers is most likely to change with a small policy change. However, the same reasoning does not apply to the subpopulation of compliers. Klein2010 evaluates the sensitivity of the treatment effect of compliers to random departures from monotonicity. FioriniStevens2014 give examples of analyzing the sensitivity of the monotonicity, and Huber2014 considers a violation of monotonicity in a specific example. They do not provide sharp identification results of the treatment effect of compliers in the presence of defiers, nor do they derive the robust region. A violation of the monotonicity assumption with a non-binary instrumental variable is considered, and alternative assumptions and testing procedures are proposed in mogstad2019identification, frandsen2019judging, norris2020effects. This paper contributes to this literature by presenting an effective tool to analyze the severity of a potential violation of the monotonicity assumption. It thus gives applied researchers a new tool to evaluate the robustness of their estimates to a violation of the monotonicity assumption, and their estimates may thereby gain credibility.

Our proposed inference procedure builds on seminal work about Delta methods for non-differentiable mappings by Shapiro1991, fansantos2016, dumbgen1993nondifferentiable, Hong2016, and it further exploits ideas of smoothing population parameters by masten2020, chernozhukov2010quantile, haile2003.

Setup

Model of the Local Average Treatment Effect

We observe the distribution of the random variables $(Y, D, Z)$, where $Y$ is the outcome of interest; $D$ is the actual treatment status, with $D=1$ if the person is treated and $D=0$ otherwise; and $Z$ is the instrument, with $Z=1$ if the person is assigned to treatment and $Z=0$ otherwise. We assume that each unit has potential outcomes $Y_0$ in the absence and $Y_1$ in the presence of treatment, and potential treatment status $D_1$ when assigned to treatment and $D_0$ when not assigned to treatment. The observed and potential outcomes are related by $Y=D Y_1 + (1-D)Y_0$, and observed and potential treatment status by $D=Z D_1 + (1-Z)D_0$.

Based on the effect of the instrument on the treatment status, we distinguish four different groups: compliers that are only treated if they are assigned to treatment (CO); defiers that are only treated if they are not assigned to treatment (DF); always takers that are independently of the instrument always treated (AT), and never takers that are never treated (NT). We denote the population sizes of the respective group by $\pi_{\textnormal{AT}}$, $\pi_{\textnormal{NT}}$, $\pi_{\textnormal{CO}}$, and $\pi_{\textnormal{DF}}$. We denote by $Y_d^{T}$ the potential outcome variable of group $T \in \{AT, NT, CO, DF\}$ under treatment status $d$. To simplify the notation, we write $Y_d^{dT}$ for the potential outcome variable of always takers if $d=1$ and otherwise of never takers, and similarly $\pi_{dT}$ for the respective population size. We denote the outcome distribution of a variable $Y$ by $F_Y$, its density function, if it exists, by $f_Y$, and its support by $\mathbb{Y}$.\footnote{Throughout the paper, we implicitly assume that all necessary moments of all random variables for the parameter of interest exist; for instance, if we consider the local average treatment effect, we assume $Y_d^{T}$ has first moments for all $d\in\{0,1\}$ and $T \in \{C, DF, AT, NT\}$.}

The key parameters of interest in this analysis are treatment effects of compliers. We denote the average treatment effect of compliers by\footnote{Similarly, the average treatment effect of defiers is denoted by $\Delta_{DF}=\mathbb{E}[Y_1-Y_0 |D_0=1, \; D_1=0]$.} $$\Delta_{CO}=\mathbb{E}[Y_1-Y_0 | D_0=0, \; D_1=1]. $$

Throughout the paper, we assume that $\mathbb P(D=1|Z=1) \geq \mathbb P (D=1|Z=0)$ without loss of generality, and we impose the following identifying assumptions.

assumptionThe instrument satisfies $(Y_{1}, Y_{0}, D_1, D_0) \perp Z$ (Independence), and $\mathbb P(D=1|Z=1)>\mathbb P(D=1|Z=0)$ (Relevance).

We refer to AngristImbensRubin1996 for an extensive discussion of these assumptions.

Illustration of the Sensitivity Analysis

In this section, we illustrate the sensitivity analysis in a very simplified framework, where we introduce the sensitivity parameters, the sensitivity and the robust region. In contrast to our main sensitivity analysis in Section (ref)-(ref), we do not consider any sharp identification results in this illustration.

Sensitivity Parameter Space

In the presence of defiers, AngristImbensRubin1996 show that the average treatment effect of compliers is not point identified. The Wald estimand, $\beta^{IV} =\textrm{Cov}(Y,Z) / \textrm{Cov}(D,Z)$, equals a weighted difference of the average treatment effect of compliers and defiers:

align[align omitted — 204 chars of source]

Three parameters in equation (ref) are in general not identified: the population size of defiers $\pi_{\textnormal{DF}}$, the treatment effect of compliers $\Delta_{CO}$ and of defiers $\Delta_{DF}$.\footnote{Clearly, if either $\pi_{\textnormal{DF}}=0$, implying the absence of defiers, or $\Delta_{CO}= \Delta_{DF}$, implying that compliers and defiers have the same average treatment effect, the treatment effect $\Delta_{CO}$ is still point identified.} To bound the average treatment effect of compliers, we introduce two sensitivity parameters. The first one determines the population size of defiers, and the second one outcome heterogeneity between compliers and defiers. These two parameters measure the degree to which monotonicity is violated and represent the two dimensions of heterogeneity: (i) heterogeneous effects of the instrument on the treatment status and (ii) heterogeneous effects of the treatment on the outcome.

The heterogeneous impact of the instrument on the treatment status, is parameterized, in the most simplest ways, by the population size of defiers

align[align omitted — 91 chars of source]

A larger sensitivity parameter $\pi_{\textnormal{DF}}$ implies a more severe violation of monotonicity. It is clear that, for a given population size of defiers, $\pi_{\textnormal{DF}}$, the population sizes of the other groups are point identified. In our analysis, these population sizes are, therefore, functions of the sensitivity parameter $\pi_{\textnormal{DF}}$, but we leave this dependence implicit.\footnote{It follows from the definitions of the groups and our assumptions that ${\pi_{\textnormal{AT}}=\mathbb{P}(D=1|Z=0)-\pi_{\textnormal{DF}}}$, $\pi_{\textnormal{NT}}=\mathbb{P}(D=0|Z=1)-\pi_{\textnormal{DF}} $ and $\pi_{\textnormal{CO}}=\mathbb{P}(D=1|Z=1)-\mathbb{P}(D=1|Z=0)+\pi_{\textnormal{DF}}$.}

We parameterize the second dimension of heterogeneity by the sensitivity parameter $\delta_a$ which equals the absolute differences in treatment effects of both groups

equation*[equation* omitted — 57 chars of source]

A larger sensitivity parameter $\delta_a$ implies a more severe violation of monotonicity.

Sensitivity Region and Robust Region

The sensitivity region is the set of sensitivity parameters which do not violate our model assumptions. For instance, a sensitivity parameter ${\pi_{\textnormal{DF}}\geq0.5}$ would violate our model assumptions as the relevance assumption implies that ${\pi_{\textnormal{CO}} > \pi_{\textnormal{DF}}}$. Therefore, such a sensitivity parameter does not lie within our sensitivity region, which is identified without imposing any additional assumptions. In this illustrative example, we simplify the derivation and say that the sensitivity region is trivially given by $$\textnormal{SR}_a= [0,0.5) \times \mathbb{R}_+.$$ In our main sensitivity analysis, this set, however, is nontrivial and can be empty. In this case, the model is rejected, implying that even though the monotonicity assumption may be violated, at least one of the other model assumptions has to be violated as well.

Even though the treatment effect of compliers is generally not point identified if $\pi_{\textnormal{DF}}>0$, using (ref), it is partially identified for any given pair of sensitivity parameters $(\pi_{\textnormal{DF}}, \delta_a)$ by $$ \Delta_{CO} \in \left[ \beta^{IV}- \frac{\pi_{\textnormal{DF}}}{\pi_{\textnormal{CO}}-\pi_{\textnormal{DF}}} \delta_a, \; \beta^{IV}+\frac{\pi_{\textnormal{DF}}}{\pi_{\textnormal{CO}} - \pi_{\textnormal{DF}}} \delta_a \right].$$ In a typical sensitivity analysis, researchers now consider different values of the sensitivity parameters to evaluate the identified sets of the parameter of interest and to evaluate the robustness of the LATE estimates to a potential violation of monotonicity. However, in many empirical applications, the interest does not lie in the precise treatment effect but in its sign or in its order of magnitude. It is, therefore, natural to start with the empirical conclusion of interest and to ask which sensitivity parameters imply treatment effects that are consistent with this conclusion. This approach is formalized by the breakdown frontier KleinSantos2013, masten2020.

We now consider the empirical conclusion that $\Delta{CO} \geq \mu$, and we assume that ${\beta^{IV} \geq \mu}$.\footnote{If $\beta^{IV} \leq \mu$, the robust region for the conclusion that $\Delta{CO} \geq \mu$ is empty.} Under our model assumptions and for a given value of the population size of defiers $\pi_{\textnormal{DF}}$, the breakdown point determines the largest value of outcome heterogeneity $\delta_a$ that implies treatment effects that are consistent with our empirical conclusion of interest. Specifically, for any $\pi_{\textnormal{DF}} \in [0,0.5]$, the breakdown point is given by $$BP_a(\pi_{\textnormal{DF}}) =\frac{\pi_{\textnormal{CO}}-\pi_{\textnormal{DF}}}{\pi_{\textnormal{DF}}} (\beta^{IV} - \mu).$$

The breakdown frontier (BF) is the set of all breakdown points and the robust region (RR) is the set of all sensitivity parameters that are consistent with the empirical conclusion of interest. They are respectively given by

align*[align* omitted — 284 chars of source]
figure[figure omitted — 1,251 chars of source]

The nonrobust region is the complement of the robust region within the sensitivity region. It contains sensitivity parameters that may or may not be consistent with the empirical conclusion. Due to the functional form of the breakdown frontier, the nonrobust region is a convex set in this example. An illustrative example of this setup is shown in Figure (ref).

In this simple example, neither the sensitivity region nor the robust regions are sharp. For example, if the outcome is binary than the difference between compliers and defiers treatment effects is clearly bounded. Similarly, the robust region might also be substantially reduced by taking into account the actually observed outcomes.\footnote{To give a concrete example, assume that all treated units have a realized outcome of 1 and all nontreated units have a realized outcome of 0. Then it is clear, that the treatment effect of compliers is point identified to be one.} This reasoning means that even though a parameter pair may lie within the sensitivity region, it might not imply a well-defined data generating process that is consistent with the model assumptions and the observed probabilities. Similarly, even though a parameter pair may lie within the nonrobust region, it might be robust. Empirical conclusions that can be drawn from this analysis might, therefore, not be very informative. Consequently, we improve upon this framework in the remainder of this paper.

Sensitivity Parameters

In this section, we introduce two sensitivity parameters that are interpretable and imply bounds on the outcome distributions of compliers so that the parameter of interest is partially identified. They allow us to consider a trade-off between the strength of the imposed assumption and the size of the identified set. To derive the sensitivity parameters, we consider the following function:

gather*[gather* omitted — 122 chars of source]

for $d \in \{0,1\}$. In the absence of defiers, $G_d(y)$ is the cumulative distribution function of compliers under treatment status $d$. In the presence of defiers, it holds analogously to (ref) that

align[align omitted — 188 chars of source]

The outcome distributions of compliers are thus identified up to the population size of defiers and the heterogeneity between the outcome distributions of compliers and defiers. We introduce two sensitivity parameters to parameterize these two dimensions. First, the presence of defiers is parameterized by the population size of defiers $\pi_{\textnormal{DF}}$ (ref). Second, outcome heterogeneity is represented by $\delta$, which bounds the maximal difference between cumulative distribution functions of the outcome of compliers and defiers by the Kolmogorov-Smirnov (KS) norm

equation*[equation* omitted — 127 chars of source]

where $\delta \in [0,1]$. Without a restriction on $\delta$, the outcome distributions can be arbitrarily different. If $\delta=0$, the outcome distributions are restricted the most as both distribution functions coincide. A larger value of the parameter $\delta$ implies a more severe violation of monotonicity.

There are clearly many different possibilities for how heterogeneity between distribution functions can be specified. In this paper, we choose the Kolmogorov-Smirnov norm, as it leads to tractable analytical solutions of the bounds on the compliers outcome distribution. More importantly, this parameterization is simple enough to be interpretable in an empirical conclusion.

A similar parameterization is chosen in KleinSantos2013 in a different context.\footnote{Since the parameterization of $\delta$ is weak on the tails of the distributions, the bounds on the tails are likely to be uninformative. Imposing a weighted KS assumption, that penalizes deviations at the tails of the two distributions more, would overcome this issue but would also lead to less tractable results.}

Partial Identification of Distribution Functions

Since our main sensitivity analysis exploits bounds on parameters defined by the distribution function $F_{Y_d^{CO}}$ for $d \in \{0, 1\}$, we bound this distribution function for a fixed given sensitivity parameter pair $(\pi_{\textnormal{DF}}, \delta)$ in this section. We illustrate the derivation of the bounds in the subsequent sections, and the main result is stated in Section (ref).

Preliminaries

Identification Strategy

Our goal is to obtain sharp lower and upper bounds of the distribution function $F_{Y_d^{CO}}$ in a first-order stochastic dominance sense. That is, we derive analytical characterizations of the distribution functions $\underline{F}_{Y_d^{CO}}$ and $ \overline{F}_{Y_d^{CO}}$ that are feasible candidates for $F_{Y_d^{CO}}$, in the sense that they are compatible with the imposed sensitivity parameters, our assumptions, and the population distributions of observable probabilities. They are further such that $ \underline{F}_{Y_d^{CO}}(y) \leq F_{Y_d^{CO}}(y) \leq \overline{F}_{Y_d^{CO}}(y), $ for all $y \in \mathbb Y$. The identification strategy for deriving such sharp bounds $\underline{F}_{Y_d^{CO}}$ and $\overline{F}_{Y_d^{CO}}$ is based on the premise that any candidate distribution function of $F_{Y_d^{CO}}$ then also implies distribution functions of $F_{Y_d^{dT}}$ and $F_{Y_d^{DF}}$. Our candidate function $F_{Y_d^{CO}}$ is therefore feasible, only if the implied functions of $F_{Y_d^{dT}}$ and $F_{Y_d^{DF}}$ are indeed distribution functions.

The explicit analytical characterization of these sharp bounds illustrates the effect of the sensitivity parameters on the bounds, and more importantly, it implies sharp bounds on a variety of treatment effects of interest, e.g., the average treatment effect of compliers stoye2010partial.\footnote{The explicit characterization also allows the inference procedure to be based on $\underline{F}_{Y_d^{CO}}$ and $\overline{F}_{Y_d^{CO}}$.}

Notation

We here collect the notation used in the following subsections. Let $d,s \in \{0,1\}$ and $y \in \mathbb Y$. Let the differences in population sizes of compliers and defiers be denoted by $\pi_{\Delta}=\pi_{\textnormal{CO}}-\pi_{\textnormal{DF}}$. Let $Q_{ds}(y) \equiv \mathbb{P}(Y\leq y, D=d|Z=s)$ be the observed joint distribution of $Y$ and $D$. We further let, for $\mathscr{B}$ denoting the Borel $\sigma$-algebra,

align*[align* omitted — 151 chars of source]

and $G_d^+=\frac{1}{\pi_{\textnormal{CO}}} \widetilde G_d^+(y)$. Our sensitivity analysis is based on the following observed underlying parameters

align[align omitted — 133 chars of source]

Preliminary Bounds

To illustrate the identification argument, we first derive preliminaries bounds on the distribution function $F_{Y_d^{CO}}$, which are not necessarily sharp in general. Based on the law of total probability and our assumptions, the probability function $Q_{dd}$ is a weighted average of the distribution functions $F_{Y_d^{CO}}$ and $F_{Y_d^{dT}}$, specifically $Q_{dd}(y)=\pi_{\textnormal{CO}} F_{Y_d^{CO}}(y)+ \pi_{\textnormal{d}} F_{Y_d^{dT}}(y) $. Any feasible distribution function of $F_{Y_d^{CO}}$ has to imply a function $F_{Y_d^{dT}}$ that is a distribution function. Exploiting this argument and using our sensitivity parameter $\pi_{\textnormal{DF}}$, it follows that

align[align omitted — 217 chars of source]

These bounds correspond to the extreme scenarios where compliers have the highest or the lowest outcomes compared to always and never takers.

Using the same argument for defiers and the definition of $G_d(y)$ in (ref), it further follows that

align[align omitted — 237 chars of source]

We now consider the second sensitivity parameter $\delta$. Based on the definition of $G_d(y)$ in (ref), we conclude that any feasible candidate of $F_{Y_d^{CO}}$ also has to satisfy that

align[align omitted — 213 chars of source]

Since the function $G_d$ is not necessarily increasing in $y$ for all $y \in \mathbb Y$, bounds on the distribution function $F_{Y_d^{CO}}$ based on (ref) and (ref) have to take this into account. We therefore directly consider bounds on $F_{Y_d^{CO}}$ that employ this information. To be precise, for the lower bound, we consider equation (ref) and (ref), where we replace $G_d$ by its smallest, nondecreasing upper envelope; vice versa, for the upper bound, where we replace $G_d$ by its greatest, nondecreasing lower envelope.\footnote{We give an illustration of this derivation in Appendix (ref).} Following this reasoning and taking (ref)-(ref) into account, the lower bound is given by

align[align omitted — 361 chars of source]

and the upper bound by

align[align omitted — 350 chars of source]

Any value outside of these bounds is clearly incompatible with the distribution of $(Y,D,Z)$ and our assumptions. To illustrate the effect of our sensitivity parameters, we consider the width of these bounds for any fixed $y \in \mathbb Y$ as a function of $(\pi_{\textnormal{DF}}, \delta)$, that is\footnote{This comparison is helpful as the qualitative size of the width of the bounds on the distribution functions is related to the width of the identified set of many parameters of interest, e.g., the LATE.} $ \overline{H}_{Y_d^{CO}}(y,\pi_{\textnormal{DF}},\delta) - \underline{H}_{Y_d^{CO}}(y,\pi_{\textnormal{DF}},\delta). $ The width is weakly increasing in the sensitivity parameter $\delta$, which implies that a larger violation of monotonicity leads to a larger identified set. However, the effect of the sensitivity parameter $\pi_{\textnormal{DF}}$ on this width can be both negative and positive depending on the specific underlying parameters $\theta$. For example, we note that $F_{Y_d^{CO}}$ is point identified either if $\pi_{\textnormal{DF}}=0$ or $\pi_d=0$, which denotes the absence of always or never takers. Heuristically speaking, the parameter $\pi_{\textnormal{DF}}$, therefore, trades off the identification power gained from the non-existence of defiers and the non-existence of always or never takers.

The functions $\underline{H}_{Y_d^{CO}}$ and $\overline{H}_{Y_d^{CO}}$ clearly bound $F_{Y_d^{CO}}$ in a first-order stochastic dominance sense. However, since they do not imply that the implied functions of $F_{Y_d^{dT}}$ and $F_{Y_d^{DF}}$ are nondecreasing, they are not necessarily a feasible candidate of $F_{Y_d^{CO}}$. To give an intuition for this result and for the sake of argument, we now assume that all outcome variables are continuously distributed. We consider $\underline{H}_{Y_d^{CO}}$, and we assume that the bound on the outcome heterogeneity $\delta$ determines the bound, i.e., $ \underline{H}_{Y_d^{CO}}(y)= G_d(y) - \frac{\pi_{\textnormal{DF}}}{\pi_{\Delta}} \delta. $ This bound does not necessarily imply that the always takers have a positive density. Specifically, the density of the lower bound is $ g_d(y) = (q_{dd}-q_{d(1-d)}(y))/\pi_{\Delta}, $ whereas to guarantee that the density function $f_{Y_d^{dT}}$ does not take any negative value, any feasible candidate of $f_{Y_d^{CO}}$ has to satisfy that

align[align omitted — 105 chars of source]

for all $y \in \mathbb Y$.\footnote{To be precise, one can assume that $q_{d(1-d)}(y)=0$ and as $\pi_\Delta =\pi_{\textnormal{CO}}- \pi_{\textnormal{DF}} \leq \pi_{\textnormal{CO}}$ the claim follows.} A similar restriction as (ref) can be derived for defiers such that any feasible candidate of the density $f_{Y_d^{CO}}(y)$ has to also satisfy that, for all $y \in \mathbb Y$,

align[align omitted — 125 chars of source]

Based on this argument, we construct our final bounds, $\underline{F}_{Y_d^{CO}}$ and $\overline{F}_{Y_d^{CO}}$. Specifically, the distribution function $\underline{F}_{Y_d^{CO}}$ is dominated by $\underline{H}_{Y_d^{CO}}$ in a first-order stochastic dominance sense, and the distribution function $\overline{F}_{Y_d^{CO}}$ dominates $\overline{H}_{Y_d^{CO}}$ in a first-order stochastic dominance sense, and they both carefully take into account the reasoning of (ref) and (ref). In Appendix (ref), we show that these distribution functions both bound the distribution function $F_{Y_d^{CO}}$ and are feasible candidates.

Identification Result

We first provide the analytical expressions of the bounds in the following. The lower bound of the distribution functions $F_{Y_d^{CO}}$ is given by

align[align omitted — 494 chars of source]

and similarly the upper bound by

align[align omitted — 475 chars of source]

Based on the derivation above, Theorem (ref) summarizes the result.

theoremSuppose that Assumption (ref) holds, and the data generating process is compatible with the sensitivity parameters $(\pi_{\textnormal{DF}}, \delta)$. Then, it holds that \begin{align*} F_{Y_d^{CO}}(y,\pi_{DF},\delta) \leq F_{Y^{CO}_d}(y) \leq \overline{F}_{Y_d^{CO}}(y,\pi_{DF},\delta), \end{align*} for $d \in \{0,1\}$ and for all $y \in \mathbb Y$. Moreover, there exist data generating processes which are compatible with the above assumptions such that the outcome distribution of compliers equals either $\underline{F}_{Y_d^{CO}}(y,\pi_{\textnormal{DF}},\delta)$, $\overline{F}_{Y_d^{CO}}(y,\pi_{\textnormal{DF}},\delta)$, or any convex combination of these bounds.

Theorem (ref) shows not only that the proposed bounds are valid but also that without imposing further assumptions, the bounds cannot be tightened in a first-order stochastic dominance sense.\footnote{As the derived bounds are rather complicated, we propose simpler bounds for each of our sensitivity parameters in Appendix (ref). These bounds are possibly conservative. We explain how to evaluate in an empirical setting whether they are close to the sharp bounds derived in this section.}

remarkTheorem (ref) does clearly not imply that all distribution functions that are bounded by the distribution functions $\underline{F}_{Y_d^{CO}}$ and $\overline{F}_{Y_d^{CO}}$ are feasible candidates of the distribution function of $F_{Y^{CO}_d}$. The reason for that is that these functions do not necessarily imply nondecreasing distribution functions of the other groups. Since we are not interested in the distributions functions themselves but in parameters defined through the bounds, this result is sufficient to derive sharp bounds on the sensitivity and robust region for empirical conclusions about these parameters.
remarkThe parameter of interest is often not only the average treatment effect but also, e.g., quantile and distribution treatment effects. As Theorem (ref) bounds the entire outcome distribution functions of compliers, in a first-order stochastic dominance sense, these treatment effects are identified as well and are sharp for many relevant parameters. We present them in Appendix (ref).
remarkIn empirical applications, researchers also often have access to pre-intervention covariates. In Appendix (ref), we show how these covariates can be exploited to reduce the size of the identified set of the distribution function $F_{Y_d^{CO}}$. These covariates can then be used to tighten the sensitivity and to enlarge the robust regions.

Sensitivity Analysis

We present our main sensitivity analysis in this section.

Sensitivity Region

We derive the sensitivity region, which is the set of sensitivity parameter pairs for which a feasible candidate of the distribution function $F_{Y_d^{CO}}$ exists. Sensitivity parameters that lie in the complement of this set refute the model, and we, therefore, do not consider them further.\footnote{masten2021salvaging denote the complement of the sensitivity region the falsification region.}

Population Size of Defiers

We show that the population size of defiers is partially identified. We denote an upper bound by

align[align omitted — 113 chars of source]

The first element of the minimum represents the sum of the population size of always takers and defiers, whereas the second one of never takers and defiers. The population size of defiers is clearly smaller than both of these quantities.

The lower bound on the population size of defiers is denoted by

align[align omitted — 184 chars of source]

The supremum is taken over the differences in population distributions of defiers and compliers, which bounds the population size of defiers from below. The Proposition (ref) shows that these bounds are sharp.\footnote{richardson2010analysis present sharp bounds on $\pi_{\textnormal{DF}}$ for a binary outcome variable.}

propositionSuppose Assumption (ref) holds. Then the population size of defiers, $\pi_{\textnormal{DF}}$, is bounded by $[\underline{\pi}_{DF}, \overline{\pi}_{DF}]$. Moreover, there exist data generating processes which are compatible with the above assumptions such that the population size of defiers equals any value within these bounds. Thus, the bounds are sharp.

If the lower bound on population size of defiers is greater than zero, $\underline{\pi}_{DF} > 0$, at least one of the classical LATE assumptions, including monotonicity, is violated.

This reasoning is align with the result of Kitagawa2015, who shows that $\underline{\pi}_{DF} > 0$ is sufficient and necessary such that the LATE assumptions are valid.

However, if the above inequalities contradict, i.e., $\underline{\pi}_{DF}>\overline{\pi}_{DF}$, the sensitivity region is empty. This implies that even if one allows for a violation of monotonicity, our model assumptions must be violated as well.

Outcome Heterogeneity

We now consider the sensitivity parameter $\delta$. Based on Theorem (ref), we can bound the sensitivity parameter $\delta$ from below and from above for a given value of the sensitivity parameter $\pi_{\textnormal{DF}}$.

A given pair fo sensitivity parameters $(\pi_{\textnormal{DF}}, \delta)$ is refuted if the implied lower and upper bounds, $\underline{F}_{Y_d^{CO}}$ and $\overline{F}_{Y_d^{CO}}$, intersect, so that there does not exists a feasible candidate of the distribution function $F_{Y_d^{CO}}$ which is compatible with these sensitivity parameters. The domain of the sensitivity parameter $\delta$ is bounded from below by

align[align omitted — 252 chars of source]

The feasible set of the sensitivity parameter $\delta$ is further bounded from above. The bounds $\underline{F}_{Y_d^{CO}}$ and $\overline{F}_{Y_d^{CO}}$ imply bounds on the distribution function of $F_{Y_d^{DF}}$, where the largest value of the Kolmogorov-Smirnov norm between the distributions of $F_{Y_d^{CO}}$ and $F_{Y_d^{DF}}$ is achieved when $\delta=1$. It follows that there does not exists a feasible candidate function of $F_{Y_d^{CO}}$ such that the implied outcome heterogeneity parameter exceeds this value. We denote the upper bounds by

align[align omitted — 427 chars of source]

By the reasoning of Theorem (ref), these bounds are sharp, and any convex combination of these bounds is feasible as well. It follows that our sensitivity region is given by

align[align omitted — 260 chars of source]

Robust Region

We now derive the robust region for the empirical conclusion that $\Delta_{CO}\geq \mu$.\footnote{In Appendix (ref), we also consider other treatment effects than the average treatment effect of compliers. Sensitivity and robust regions for empirical conclusions about these parameters can then also be derived based on the reasoning of this section.} To simplify the presentation, we assume in the following that the sensitivity region is nonempty and that $\Delta_{CO}(\underline{\pi}_{DF}, \underline{\delta}(\underline{\pi}_{DF}))\geq \mu$.\footnote{If $\Delta_{CO}(\underline{\pi}_{DF}, \underline{\delta}(\underline{\pi}_{DF}))< \mu$, the robust region is empty.}

By first-order stochastic dominance of the distribution functions $\underline{F}_{Y_d^{CO}}$ and $\overline{F}_{Y_d^{CO}}$, we can construct sharp bounds on many treatment effect parameters, that depend on these bounds stoye2010partial. Specifically, let

align[align omitted — 516 chars of source]
corollarySuppose that Assumption (ref) holds, and the data generating process is compatible with the sensitivity parameters $(\pi_{\textnormal{DF}}, \delta)$. Then, the average treatment effect of compliers, $\Delta_{CO}$, is bounded by $ [\underline{\Delta}_{CO}(\pi_{\textnormal{DF}}, \delta) , \overline{\Delta}_{CO}(\pi_{\textnormal{DF}}, \delta) ]$. Moreover, there exist data generating processes which are compatible with the above assumptions such that the average treatment effect of compliers equals any value within these bounds. Thus, the bounds are sharp.

For a given sensitivity parameter $\pi_{\textnormal{CO}}$, we now consider the breakdown point given by $$BP(\pi_{\textnormal{DF}} )=\sup \{\delta: (\pi_{\textnormal{DF}}, \delta) \in \textnormal{SR} \text{ and }\underline{ \Delta}_{CO}(\pi_{\textnormal{DF}},\delta) \geq\mu\}. $$ For a given sensitivity parameter $\pi_{\textnormal{DF}}$, it identifies the weakest assumption on outcome heterogeneity between compliers and defiers such that the empirical conclusion holds. The breakdown point, as a function of the sensitivity parameter $\pi_{\textnormal{DF}}$, is not necessarily decreasing in the population size of defiers as the bounds on the outcome distribution of compliers can become tighter if the value of $\pi_{\textnormal{DF}}$ increases (see the discussion in Section (ref)). The breakdown frontier of the average treatment effect is the boundary of the robust region and given by the set of all breakdown points

align[align omitted — 134 chars of source]

The robust region of the empirical conclusion that $\Delta_{CO}\geq \mu$ is characterized by

align[align omitted — 147 chars of source]

The nonrobust region, that is the complement of the robust region within the sensitivity region, contains pairs of sensitivity parameters which only may not imply treatment effects being consistent with the empirical conclusion. Figure (ref) illustrates one example of the sensitivity and robust region.\footnote{We refer to a discussion on how these sets can be used in an empirical setting to Section (ref) and (ref)}

figure[figure omitted — 1,500 chars of source]

Extensions

In this section, we show how our framework can be exploited to draw empirical conclusions about other population parameters, and how it simplifies if the outcome variable is binary.

Treatment Effects for other Populations

To show how empirical questions about treatment effects of the entire population can be analyzed, we exploit that the proof of Theorem (ref) presents sharp bounds on all groups in a first-order stochastic dominance sense. For $d \in \{0,1\}$, let the lower bound be denoted by

align*[align* omitted — 177 chars of source]

and the upper bound by

align*[align* omitted — 182 chars of source]
propositionSuppose the instrument satisfies Assumption (ref), and the data generating process is compatible with the sensitivity parameters $(\pi_{\textnormal{DF}}, \delta)$. Then, it holds that $$ \underline F_{Y_d} (y, \pi_{\textnormal{DF}}, \delta) \leq F_{Y_d} (y, \pi_{\textnormal{DF}}, \delta) \leq \overline F_{Y_d} (y, \pi_{\textnormal{DF}}, \delta)$$ for $d \in \{0,1\}$ and for all $y \in \mathbb Y$. Moreover, there exist data generating processes which are compatible with the above assumptions such that the potential outcome distributions equal either $\overline{F}_{Y_d}(y,\pi_{\textnormal{DF}},\delta)$, $\underline{F}_{Y_d}(y,\pi_{\textnormal{DF}},\delta)$, or any convex combination of these bounds.

As the data do not contain any information about the distribution functions $F_{Y_0^{AT}}$ and $F_{Y_1^{NT}}$, the bounds $ \underline F_{Y_d}$ and $\overline F_{Y_d}$ are such that their respective probability mass is shifted to the extreme of the support $\mathbb Y$. To interpret these bounds, for any $y$ in the interior of $\mathbb Y$, we consider the difference $$\overline F_{Y_d} (y, \pi_{\textnormal{DF}}, \delta) - \underline F_{Y_d} (y, \pi_{\textnormal{DF}}, \delta) = \pi_d.$$ The sensitivity parameter $\delta$ does not influence the distribution of the outcome of the entire population, as it only influences how the observed outcome probability mass is distributed between the groups. However, the size of the bounds decreases with the population size of defiers, $\pi_{\textnormal{DF}}$, as the population size of always and never takers $\pi_{\textnormal{d}}$ decreases with $\pi_{\textnormal{DF}}$. This reasoning is intuitive as if $\pi_{\textnormal{d}}$ decreases, the observed outcome probability mass represent more of the population under consideration. This result aligns with results of the literature kitagawa2021identification,kamat2018identifying, who shows that imposing monotonicity (e.g., $\pi_{\textnormal{DF}}=0$) does not imply a smaller identified set of the average treatment effect of the entire population if the LATE assumptions are not violated.

Based on the bounds presented in Proposition (ref), we can now derive a sensitivity analysis similar to the one presented in Section (ref). However, to derive informative results about the average treatment effect of the entire population, we would have to impose that the outcome is bounded as otherwise the average treatment effect is not identified in general.

The sensitivity analysis of this paper is based on the premise that the treatment effect of compliers is the object of interest. However, if the parameter of interest is the treatment effect of the entire population, one might then be willing to impose assumptions not only on outcome heterogeneity between compliers and defiers but also between other groups. To be precise, we can replace the sensitivity parameter $\delta $ by $\delta_p\in [0,1]$ such that $$ \max_d \sup_y \{ | F_{Y_d^{T}} (y) -F_{Y_d^{T'}} (y)| \} | \leq \delta_p \quad \forall \; T,T' \in \{ AT, NT, CO, DF\}. $$ Using similar arguments as in the proof of Theorem (ref), one can then derive sharp bounds on the outcome distribution functions of the entire population and then conduct a sensitivity analysis similar to the one described in Section (ref). Empirical conclusions drawn on this parameterization might be substantially more informative.

Binary Outcome Variable

In many empirical applications, the outcome of interest is binary. The results of Section (ref) and (ref) are still valid in this case, but we show in this section that the bounds substantially simplify so that they are easier applicable. Let $P_d^{T}= \mathbb P (Y_d^T=1)$ denote the probability that the random variable $Y_d^T$ equals one, and let the conditional joint probability of the outcome and the treatment status be given by $P_{ds}=\mathbb P(Y=1, D=d|Z=s)$. We denote the underlying parameters by $\theta_b=(P_{11},P_{10},P_{01}, P_{00}, P_0, P_1) \in [0,1]^6$.

Following the same arguments as above, the sensitivity and robust region depend on the marginal outcome distributions of the compliers. The presence of defiers is also bounded by $\pi_{\textnormal{DF}} $, and the parameter of outcome heterogeneity simplifies to $$ \delta_b= \max_{d \in \{0,1\}} |P_d^{CO}-P_d^{DF}|. $$ The outcome probabilities of compliers are bounded from below by

align[align omitted — 338 chars of source]

and from above by

align[align omitted — 332 chars of source]
corollarySuppose Assumption (ref) holds, and the data generating process is compatible with the sensitivity parameters $(\pi_{\textnormal{DF}}, \delta)$. The outcome probabilities of compliers are bounded by $ \underline P_d^{CO} \leq P_d^{CO} \leq \overline P_d^{CO}$. Moreover, there exist data generating processes which are compatible with the above assumptions such that the population size of defiers equals any value within these bounds. Thus, the bounds are sharp.

The interpretation of the width of these bounds follows the same reasoning as in Section (ref). The lower bound of the population size of defiers simplifies to

align*[align* omitted — 167 chars of source]

The upper bound on $\pi_{\textnormal{DF}}$ cannot be simplified further and is given by (ref). The lower bound on outcome heterogeneity is given by

align*[align* omitted — 114 chars of source]

The lower bound on the sensitivity parameter $\delta$ decreases with the population size of defiers. The upper bound on the sensitivity parameter $\delta$ is given by the maximal difference between the outcome probabilities of compliers and defiers

gather*[gather* omitted — 290 chars of source]

The sensitivity parameter space is given by $$ \textnormal{SR}_b=\{ (\pi_{\textnormal{DF}} , \delta_b) \in [\underline{\pi}_{DF}, \overline{\pi}_{DF}] \times [0,1]: \underline{\delta}_b(\pi_{\textnormal{DF}}) \leq \delta_b \leq \overline{\delta}_b(\pi_{\textnormal{DF}}) \}, $$ and the robust region for the claim $\Delta_{CO}\geq \mu$, if $\underline P_1^{CO}(\underline{\pi}_{DF} , \underline{\delta}_b) - \overline P_0^{CO}(\underline{\pi}_{DF} , \underline{\delta}_b) \geq \mu$, is given by

align*[align* omitted — 204 chars of source]

Using the simple algebra structure of the bounds of the outcome probabilities, a closed-form expression for both the robust and the sensitivity region can be derived. As this expression is rather lengthy without providing much intuition, we state it in Appendix (ref).

Estimation and Inference

Even though the contribution of this paper is the derivation of the sensitivity and robust region for a particular empirical conclusion, we consider some methods for estimation and inference of these two regions. While the technical details are deferred to Appendix (ref), in this section, we sketch the main issues of conducting inference in this setting and our proposed solutions. To simplify the exposition, we consider the setting of a continuous and a binary outcome variable, but our method is not restricted to these distributions.

Throughout this section, we assume that we have access to the data $\{(Y^z_{i},D^z_{i})\}_{i=1}^{n_z}$ for $z\in \{0,1\}$ that are independent and identically distributed according to the distribution of $(Y,D)$ conditionally on $Z=z$ with support $\mathbb{Y}\times \{0,1\}$. We denote this distribution by $(Y^z, D^z)$ and we let $n=n_0+n_1$, where $n_0/n$ converges to a nonzero constant as $n\rightarrow \infty$.\footnote{We discuss this assumption in Assumption (ref).}

Estimation

To construct estimators of the sensitivity and robust region for a particular empirical conclusion, we note that the identification argument of these regions are constructive. It follows from Section (ref) that the boundaries of both regions are identified by the following mapping,\footnote{The signs of the components of the mapping $\phi(\theta, \pi_{\textnormal{DF}})$ simplify the subsequent analysis.}

align[align omitted — 245 chars of source]

which is evaluated at the sensitivity parameter $\pi_{\textnormal{DF}} \in \left[0,0.5\right)$ and the underlying parameters $\theta$, that is defined in (ref). Estimating the sensitivity and robust region is then equivalent to estimating this mapping. To do so, we consider estimates of the underlying parameters $\theta$ that are simply obtained by replacing unknown population quantities by their corresponding nonparametric sample counterparts and by standard nonparametric kernel methods. We denote the estimates of $\theta$ by $\widehat \theta$. Point estimates of the mapping $\phi(\theta, \pi_{\textnormal{DF}})$ can then be derived by simple plug-in methods. We defer a detailed description to Appendix (ref).

Goal of Inference

We propose to construct confidence sets for the sensitivity and robust region such that the confidence set for the sensitivity region is an outer confidence set and for the robust region is an inner confidence set.\footnote{Considering inner confidence set for the robust region follows from masten2020.} These confidence sets should therefore jointly satisfy with probability approaching the confidence level, $1-\alpha$, that (i) any sensitivity parameter pair of the sensitivity region lies within the confidence set for the sensitivity region and (ii) not any single parameter pair of the nonrobust region lies within the confidence set for the robust region.\footnote{To give one more interpretation of the confidence sets and using the language of hypothesis testing, a sensitivity parameter pair, $(\pi_{\textnormal{DF}}, \delta)$, does not lie in the sensitivity region only if we can reject such a hypothesis with confidence level $1-\alpha$. Contrary, $(\pi_{\textnormal{DF}}, \delta)$ lies in the robust region, only if we can reject that it is nonrobust with confidence level $1-\alpha$. The confidence sets are constructed so that the hypothesis tests are valid uniformly in the sensitivity parameter space. } Let $\widehat \textnormal{SR}_{L}$ and $\widehat{RR}_{L}$ denote two sets of the sensitivity parameters. They satisfy the described condition if

align[align omitted — 247 chars of source]

Based on the definition of the mapping $\phi (\theta, \pi_{\textnormal{DF}}) $, it therefore suffices to construct a lower confidence band for each component of the estimator $\phi (\widehat \theta, \pi_{\textnormal{DF}}) $ as a function of $\pi_{\textnormal{DF}}$ that are jointly valid.\footnote{Throughout this section, we consider confidence sets that are uniformly valid in the sensitivity parameter space, but not necessarily in the distribution of the underlying parameters $\theta$} That is, we need to find a function that is componentwise a uniformly lower bound $\phi_L (\widehat \theta, \pi_{\textnormal{DF}})$ in $\pi_{\textnormal{DF}}$ of $\phi ( \theta, \pi_{\textnormal{DF}}) $ so that\footnote{We verify this equivalence in Appendix (ref).}

align[align omitted — 291 chars of source]

where $e_l$ is the l-th unit vector.\footnote{Conservative confidence sets for only the average treatment effect of compliers for specific values of $(\pi_{\textnormal{DF}}, \delta)$ directly follow from the presented procedure. To obtain nonconservative confidence sets, one can follow the literature on partially identified parameters imbens2004confidence}

Inference for a Continuous Outcome Variable

We analyze the distribution of $\phi(\widehat \theta, \pi_{\textnormal{DF}})$ in order to construct confidence sets for the mapping $\phi(\theta, \pi_{\textnormal{DF}})$.\footnote{We want to emphasize that this procedure is valid for a fixed distribution. In particular, we do not consider settings of weak instruments or data generating processes which are such that the robust region becomes empty.} Under regularity assumptions presented in Appendix (ref), the estimators of the underlying parameters $\widehat \theta$ converge in $\sqrt{n}$ to a tight Gaussian process. Since the mapping $\phi$ is not Hadamard-differentiable, as it depends on minimum, maximum, supremum, and infimum of random functions, standard Delta method arguments do not apply in this setup fansantos2016. We propose a method to construct confidence sets that are asymptotically conservative but valid in the sense of (ref). It is based on ideas of population smoothing that have been suggested by, e.g., haile2003, chernozhukov2010quantile, masten2020.

In contrast to considering the mapping $\phi$, which identifies the sensitivity and robust region, we construct a smooth mapping, $\phi_\kappa$, which yields valid bounds of both regions. The smoothed mapping $\phi_\kappa$ is indexed by a fixed smoothing parameter $\kappa\in \mathbb N$. The mapping $\phi_\kappa$ is differentiable such that the standard functional Delta method can be applied to $\phi_\kappa$ and we can study its asymptotic distribution by standard methods. The mapping $\phi_\kappa$ is further such that it yields an outer set of the sensitivity region and an inner set of the robust region. This reasoning implies that confidence sets of the smooth mappings $\phi_\kappa$, which are valid in the sense of (ref), are also valid for the mapping $\phi$.

In finite samples, the choice of the smoothing parameter $\kappa$ comprises the trade-off of constructing conservative confidence sets and better finite sample approximations of the underlying distributions. Suppose the smoothing parameter $\kappa$ is small. In that case, the smoothed sensitivity and robust region are very similar to the original regions. However, the finite-sample distribution of $\phi_\kappa(\widehat \theta \pi_{\textnormal{DF}})$ might not be well-approximated by its asymptotic distribution. Vice versa, suppose the smoothing parameter $\kappa$ is large. The finite-sample distribution of $\phi_\kappa(\widehat \theta \pi_{\textnormal{DF}})$ might be well-approximated by its asymptotic distribution. However, the smoothed sensitivity and robust region are conservative to the original regions.

In Appendix (ref), we show how the smoothed mappings can be constructed. It then follows that plug-in estimators of the smoothed mappings converge in $\sqrt{n}$ to a Gaussian process by standard functional Delta method arguments. The covariance structure of this process is, in general, rather complicated and tedious to estimate. We, therefore, apply the nonparametric bootstrap to simulate its distribution. Consistency of this bootstrap procedure follows from arguments of fansantos2016. In Appendix (ref), we show how to construct the confidence sets based on the described procedure and that they achieve the outlined goal (ref).

Inference for a Binary Outcome Variable

Following the discusssion about a bianry outcome model in Section (ref), the mapping yielding the sensitivity and the robust region for a particular conclusion for a binary outcome variable is given by\footnote{where its precise definition follows from Section (ref) and Appendix (ref).} $$\phi_b(\theta_b, \pi_{\textnormal{DF}})= (\underline{\pi}_{\text{DF,b}}, -\overline{\pi}_{\text{DF,b}}, - \overline{\delta}_b( \pi_{\textnormal{DF}}), BP_b(\pi_{\textnormal{DF}})),$$ The interpretation of $\phi_b$ follows the one for a continuously distributed outcome variable, and in principle, we could apply the same inference procedure as described above. However, the mapping $\phi_b(\theta_b, \pi_{\textnormal{DF}})$ is substantially simpler than the mapping $\phi(\theta, \pi_{\textnormal{DF}})$ so that, in this section, we can apply more classical inference procedure to obtain confidence sets in the sense of (ref); in particular, we follow ideas of masten2020 and the literature about moment inequalities andrews2010inference.

Under standard sampling assumptions, it follows that the estimators of the underlying parameters are jointly $\sqrt{n}$ normally distributed (see Appendix (ref)). The mapping $\phi_b(\theta_b, \pi_{\textnormal{DF}})$ is clearly not Hadamard-differentiable, as it consist of minimum and maximum of random functions. Standard Delta method arguments are therefore not applicable here as well. Valid confidence sets could be obtained by projection arguments, which, however, are known to be conservative in general.

We show instead that the mapping $\phi_b$ is Hadamard directionally differentiable in the direction of $\theta$ when evaluated at finitely many $\{ \pi_{\textnormal{DF}}^k\}_{k=1}^K$. Using generalized Delta method arguments, the estimator of the mapping $\phi_b$ converges to some tight random process, which is a continuous transformation of a Gaussian process, indexed at the finite set $\{\pi_{\textnormal{DF}}^k\}_{k=1}^K$. As this limiting distribution is rather complicated, we do not construct our inference procedure directly on its limiting distribution, but one can choose various modified bootstrap methods to simulate this distribution, e.g., subsampling or numerical-Delta method dumbgen1993nondifferentiable, Hong2016. In this paper, we follow a bootstrap method which relies on ideas based on the moment inequality literature andrews2010inference, bugni2010bootstrap and we explain the procedure in detail in Appendix (ref). Based on this bootstrap procedure, we can construct valid lower confidence sets for $\phi_b$ indexed at the finite set of sensitivity parameters $\{\pi_{\textnormal{DF}}^k\}_{k=1}^K$. Using these confidence sets and exploiting the functional form of $\phi_b$, we then obtain lower confidence sets for the estimator of the mapping $\phi_b$, which are uniformly valid in $\pi_{\textnormal{DF}}$. We state these arguments precisely in Appendix (ref) and show that these confidence sets are asymptotically valid in the sense of our goal of inference (ref).

Simulations

Setup

We study the finite sample performance of the proposed estimators of the sensitivity and robust regions through a Monte Carlo study. We consider different data generating processes with varying degrees of violations of monotonicity, implying different sizes and shapes of both the sensitivity and robust regions. Specifically, we consider the following population sizes $(\pi_{\textnormal{CO}}, \pi_{\textnormal{DF}}) \in \{ (0.35, 0.05) , (0.25,0.15) \}$, where $\pi_{\textnormal{DF}}=\pi_{\textnormal{AT}} =0.3.$ We set $\mathbb{P}(Z=1)=0.5$ and we generate the outcome by

align*[align* omitted — 184 chars of source]

where $\Delta_{CO} \in \{0.2,0.1\}$, and $\mathcal{B}( 1,p)$ denotes the Bernoulli distribution with parameter $p$. The sensitivity region is nonempty as the data generating process satisfies our model assumptions. We consider the empirical conclusion of a positive treatment effect of compliers, so that the robust region is nonempty in each of the data generating processes as the Wald estimand is positive. The bootstrap procedure requires to choose the tuning parameter $\eta$, which is explained in Appendix (ref). We consider different values of $\eta$ given by $\{0.2, 0.5,1,1.5,2\}/\sqrt{N}$. The results are based on 10,000 Monte Carlo draws.

Simulation Results

Table (ref) shows the results of the simulated coverage rates at which the confidence sets cover the population sensitivity and nonrobust region for the different data generating processes and choices of tuning parameters. Our considered choice of tuning parameters implies that the simulated coverage of our confidence sets is close to the nominal one in most data generating processes. These results illustrate that the confidence method performs reasonably well in finite samples.

table[table omitted — 1,033 chars of source]

Empirical Application

To illustrate our proposed framework, we apply this sensitivity analysis to data from angristevens1998, who analyze the effect of having a third child on the labor market outcomes of mothers. It is shown that even small violations of the monotonicity assumption may have a large impact on the robustness of the estimated treatment effects such that even the sign of the treatment effects may be indeterminate. The same-sex instrument in angristevens1998 arguably satisfies Assumption (ref): The independence assumption seems to be plausible by the following reasoning: The sex of a child is determined by nature, and only the number of and not the sex of the child arguably influences the labor market outcome. The relevance assumption is testable. However, monotonicity might be violated. We apply the proposed sensitivity analysis to evaluate the robustness of the estimated treatment effects to a potential violation of monotonicity in this setting. For simplicity, we focus on two outcome variables: the labor market participation of mothers and their annual wage.\footnote{The annual wage is a continuously distributed running variable with a point mass at zero.} The binary decision to treat represents the extensive margin and the continuous outcome variable a mix of extensive and intensive variables. We use the same data as angristevens1998.\footnote{Data are taken from the website Joshua D. Angrist website www.economics.mit.edu/faculty/angrist from 1980. The sample is restricted to women at the age of 20-36, having at least two children, being white, and having their first child at the age of 19-25.} The sample size is 211,983. The point estimated difference of the population sizes of compliers and defiers is given by 0.06.

Sensitivity Analysis for Binary Outcome Variable

We consider the labor market participation of mothers as the outcome variable. The Wald estimate is given by $-0.13$. Figure (ref) illustrates the 95% confidence set for the sensitivity and the robust region for the claim that the treatment effect of compliers is negative. The formal definition of these confidence sets is given in Section (ref).

figure[figure omitted — 11,063 chars of source]

In this example, a (conservative) 95% confidence set for $\pi_{\textnormal{DF}}$ is given by $[0, 0.37]$. Following the literature, one can therefore not conclude that monotonicity is violated in this example small2017instrumental. The sensitivity parameter pairs below the red line represent the robust region, which is the estimated set of sensitivity parameters implying a negative treatment effect. This figure shows that concerns about the validity of the monotonicity assumption have to be taken seriously. Since $BP(0.37)$ is almost zero, the hypothesis that the treatment effect is negative cannot be rejected without imposing any assumptions on the data generating process, If the population size of defiers increases, the breakdown frontier is relatively steeply declining, and thus the robust region is rather small. This implies that relatively strong assumptions on the outcome distributions of compliers and defiers have to be imposed to conclude that the treatment effect is negative in the presence of defiers. In contrast, if the population size of defiers is small, it is not necessary to impose strong assumptions about heterogeneity in the outcome variables to imply a negative effect.

This example shows that without imposing any assumptions on the data generating process, only non-informative conclusions can be drawn in this example, which is the case as the population size of defiers is not much restricted and is arguably implausible high.\footnote{ To interpret these numbers, we note that the upper bound is a rather conservative estimate. If roughly 37% of the population were a defier, then approximately 43% of the population would have been a complier. This reasoning implies that roughly 90% of the population would base their decision to have a third child on the sex composition of the first two children.} One, therefore, might be willing to impose further assumptions to arrive at more interesting results, and we show how one could plausibly proceed. These assumptions should only serve as an example, and obviously, they have to be always adapted to the analyzed situation. We adopt the approach of de2017tolerating. One of the most essential inherently unknown quantities of interest is the population size of defiers. Imposing a smaller upper bound of this quantity based on economic reasoning allows us to derive sharper results. Based on a survey conducted in the US, de2017tolerating states that it seems reasonable that 5% of defiers is a conservative upper bound of the population size of defiers in this setting. If one is willing to impose this assumption, one would still have to assume that the differences in the Kolmogorov-Smirnov norm are less than 0.05, which is a quite strong assumption. Therefore, we would conclude that the treatment effect is not robust to a potential violation in this specific example.

Sensitivity Analysis for Continuous Outcome Variable

We now consider the annual log income of the mother. This variable has a point mass at zero, representing all women who do not work but is otherwise continuously distributed. The Wald estimate is given by $-1.23$. Figure (ref) shows the corresponding 95% confidence sets for the robust and sensitivity region. If the monotonicity assumption were not violated, this estimate would imply that women who get a third child have an annual log wage reduced by $1.23.$

Figure (ref) shows 95 % confidence sets for both the sensitivity and robust region. The same line of interpretation applies as in the case of a binary outcome variable. One can see that without imposing any assumption about the population size of defiers, the empirical conclusion of a negative treatment effect is not robust to a potential violation of monotonicity. However, applying the same reasoning as above and imposing a maximal population size of 5% as an upper bound of the population size of defiers, one can see that the empirical conclusion is now robust to a potential violation of monotonicity.

figure[figure omitted — 7,699 chars of source]

To conclude, this sensitivity analysis is of interest, as one can identify the sign and the order of magnitude of the treatment effects by imposing further assumptions. These imposed assumptions are substantially weaker than the monotonicity assumption so that the estimates gain credibility.

Conclusion

The local average treatment effect framework is popular to evaluate heterogeneous treatment effects in settings of endogenous treatment decisions and instrumental variables. In some empirical settings, one might doubt the validity of one of its key identifying assumptions, the monotonicity assumption. Conducting a sensitivity analysis of the estimates in these settings improves the reliability of the results. This paper, therefore, proposes a new framework, which allows researchers to assess the robustness of the treatment effect estimates to a potential violation of monotonicity. It parameterizes a violation of monotonicity by two parameters, the presence of defiers and heterogeneity of defiers and compliers. The former parameter is represented by the population size of defiers and the latter by the Kolmogorov-Smirnov norm bounding the outcome distributions of both groups. Based on these two parameters, we derive sharp identified sets for the average treatment effect of compliers and for any other group under further mild support assumption on the outcome variable. These identification results allow us to identify the sensitivity parameters that imply conclusions of treatment effect being consistent with the empirical conclusion. The empirical example of angristevens1998 same-sex instruments underlines the importance of the validity of the monotonicity assumptions as small violations of monotonicity may already lead to uninformative results.