EconBase
← Back to paper

Inference on Breakdown Frontiers

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

126,489 characters · 21 sections · 94 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference on Breakdown Frontiers

abstractGiven a set of baseline assumptions, a breakdown frontier is the boundary between the set of assumptions which lead to a specific conclusion and those which do not. In a potential outcomes model with a binary treatment, we consider two conclusions: First, that ATE is at least a specific value (e.g., nonnegative) and second that the proportion of units who benefit from treatment is at least a specific value (e.g., at least 50%). For these conclusions, we derive the breakdown frontier for two kinds of assumptions: one which indexes relaxations of the baseline random assignment of treatment assumption, and one which indexes relaxations of the baseline rank invariance assumption. These classes of assumptions nest both the point identifying assumptions of random assignment and rank invariance and the opposite end of no constraints on treatment selection or the dependence structure between potential outcomes. This frontier provides a quantitative measure of robustness of conclusions to relaxations of the baseline point identifying assumptions. We derive $\sqrt{N}$-consistent sample analog estimators for these frontiers. We then provide two asymptotically valid bootstrap procedures for constructing lower uniform confidence bands for the breakdown frontier. As a measure of robustness, estimated breakdown frontiers and their corresponding confidence bands can be presented alongside traditional point estimates and confidence intervals obtained under point identifying assumptions. We illustrate this approach in an empirical application to the effect of child soldiering on wages. We find that sufficiently weak conclusions are robust to simultaneous failures of rank invariance and random assignment, while some stronger conclusions are fairly robust to failures of rank invariance but not necessarily to relaxations of random assignment.

JEL classification: C14; C18; C21; C25; C51

Keywords: Nonparametric Identification, Partial Identification, Sensitivity Analysis, Selection on Unobservables, Rank Invariance, Treatment Effects, Directional Differentiability

\onehalfspacing

\onehalfspacing

Introduction

Traditional empirical analysis combines the observed data with a set of assumptions to draw conclusions about a parameter of interest. Breakdown frontier analysis reverses this ordering. It begins with a fixed conclusion and a set of baseline assumptions and asks, `What are the weakest assumptions needed to draw that conclusion, given the observed data?' For example, consider the impact of a binary treatment on some outcome variable. The traditional approach might assume random assignment, point identify the average treatment effect (ATE), and then report the obtained value. The breakdown frontier approach instead begins with a conclusion about ATE, like `ATE is positive', and reports the weakest assumption---relative to random assignment---on the relationship between treatment assignment and potential outcomes needed to obtain this conclusion, when such an assumption exists. When more than one kind of assumption is considered, this approach leads to a curve, representing the weakest combinations of assumptions which lead to the desired conclusion. This curve is the breakdown frontier.

At the population level, the difference between the traditional approach and the breakdown frontier approach is a matter of perspective: an answer to one question is an answer to the other. This relationship has long been present in the literature initiated by Manski on partial identification (for example, see Manski2007 or section 3 of Manski2013). In finite samples, however, which approach one chooses has important implications for how one does statistical inference. Specifically, the traditional approach estimates the parameter or its identified set. Here we instead estimate the breakdown frontier. The traditional approach then performs inference on the parameter or its identified set. Here we instead perform inference on the breakdown frontier. Thus the breakdown frontier approach puts the weakest assumptions necessary to draw a conclusion at the center of attention. Consequently, by construction, this approach avoids the non-tight bounds critique of partial identification methods (for example, see section 7.2 of HoRosen2016). One distinction is that the traditional approach may require inference on a partially identified parameter. The breakdown frontier approach, however, only requires inference on a point identified object.

The breakdown frontier we study generalizes the concept of an “identification breakdown point” introduced by HorowitzManski1995, a one dimensional breakdown frontier.\footnote{The identification breakdown point is distinct from the breakdown point introduced earlier by Hampel1968,Hampel1971 in the robust statistics literature that began with Huber1964; also see DonohoHuber1983. HorowitzManski1995 give a detailed comparison of the two concepts. Throughout this paper we use the term “breakdown” in the same sense as Horowitz and Manski's identification breakdown point.} Their breakdown point was further studied and generalized by Stoye2005,Stoye2010. Our emphasis on inference on the breakdown frontier follows KlineSantos2013, who proposed doing inference on a breakdown point. Finally, our focus on multi-dimensional frontiers builds on the graphical sensitivity analysis of Imbens2003 and the multi-dimensional sensitivity analysis of ManskiPepper2018. We discuss these papers and others in detail in appendix (ref).

The breakdown frontier approach

The breakdown frontier approach requires six main steps: (a) specify a parameter of interest, (b) specify a set of baseline assumptions, (c) define a class of assumptions indexed by a sensitivity parameter which deliver a nested sequence of identified sets, with the baseline assumptions obtained at one extreme and the no assumptions bounds obtained at the other, (d) characterize identified sets for the parameter of interest as a function of the sensitivity parameter, (e) use those identified sets to define the breakdown frontier for a conclusion of interest, and (f) develop estimation and inference procedures for that frontier based on its characterization.

In principle this analysis can be done for a general class of models, for example, by using the general identification analysis in ChesherRosen2015 or Torgovitsky2015 for step (d), and then applying general tools for nonparametric estimation and inference for step (f). While such a general analysis is an important next step for future work, in this paper we focus on just one important and widely used model: the potential outcomes model with a binary treatment. By focusing on a single concrete model we can clearly illustrate how to do the six main steps (a)--(f) required for a breakdown frontier analysis in any model. While the mathematical analysis will differ from model to model, the general ideas and approach do not.

In the rest of this section and this paper, we use the binary treatment potential outcomes model to explain and illustrate the breakdown frontier approach. Our main parameter of interest is the proportion of units who benefit from treatment. Under random assignment of treatment and rank invariance, this parameter is point identified. One may be concerned, however, that these two assumptions are too strong. We relax rank invariance by supposing that there are two types of units in the population: one type for which rank invariance holds and another type for which it may not. The proportion $t$ of the second type measures the relaxation of rank invariance. We relax random assignment using a propensity score distance $c \geq 0$ as in our previous work, MastenPoirier2017. We give more details on both of these relaxations in section (ref). We derive the identified set for $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ as a function of $(c,t)$. For a specific conclusion, such as $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq 0.5$, this identification result defines a breakdown frontier.

figure[figure omitted — 314 chars of source]

Figure (ref) illustrates this breakdown frontier. The horizontal axis measures $c$, the relaxation of the random assignment assumption. The vertical axis measures $t$, the relaxation of rank invariance. The origin represents the baseline point identifying assumptions of random assignment and rank invariance. Points along the vertical axis represent random assignment paired with various relaxations of rank invariance. Points along the horizontal axis represent rank invariance paired with various relaxations of random assignment. Points in the interior of the box represent relaxations of both assumptions. The points in the lower left region are pairs of assumptions $(c,t)$ such that the data allow us to draw our desired conclusion: $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq 0.5$. We call this set the robust region. Specifically, no matter what value of $(c,t)$ we choose in this region, the identified set for $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ always lies completely above 0.5. The points in the upper right region are pairs of assumptions that do not allow us to draw this conclusion. For these pairs $(c,t)$ the identified set for $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ contains elements smaller than 0.5. The boundary between these two regions is precisely the breakdown frontier. The area under the breakdown frontier---the robust region---is a quantitative measure of robustness.

Figure (ref) illustrates how the breakdown frontier changes as our conclusion of interest changes. Specifically, consider the conclusion that \[ \ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p} \] for five different values for $\underline{p}$. The figure shows the corresponding breakdown frontiers. As $\underline{p}$ increases towards one, we are making a stronger claim about the true parameter and hence the set of assumptions for which the conclusion holds shrinks. For strong enough claims, the claim may be refuted even with the strongest assumptions possible. Conversely, as $\underline{p}$ decreases towards zero, we are making progressively weaker claims about the true parameter, and hence the set of assumptions for which the conclusion holds grows larger.

figure[figure omitted — 1,657 chars of source]

Under the strongest assumptions of $(c,t) = (0,0)$, the parameter $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ is point identified. Let $p_{0,0}$ denote its value. The value $p_{0,0}$ is often strictly less than 1. In this case, any $\underline{p} \in (p_{0,0},1]$ yields a degenerate breakdown frontier: This conclusion is refuted under the point identifying assumptions. Even if $p_{0,0}<\underline{p}$, the conclusion $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$ may still be correct. This follows since, for strictly positive values of $c$ and $t$, the identified sets for $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ do contain values larger than $p_{0,0}$. But they also contain values smaller than $p_{0,0}$. Hence there do not exist any assumptions for which we can draw the desired conclusion.

The breakdown frontier is similar to the classical Pareto frontier from microeconomic theory: It shows the trade-offs between different assumptions in drawing a specific conclusion. For example, consider figure (ref). If we are at the top left point, where the breakdown frontier intersects the vertical axis, then any relaxation of random assignment requires strengthening the rank invariance assumption in order to still be able to draw our desired conclusion. The breakdown frontier specifies the precise marginal rate of substitution between the two assumptions. If we can interpret the units of each relaxation separately, we can interpret trade-offs between them. We discuss these individual interpretations in section (ref). It can sometimes be helpful to measure the two different relaxations in common units. This is not necessary, however. For example, in labor supply models agents trade off leisure hours for consumption, although time and quantities of goods are measured in fundamentally different units. Many common rates outside of economics, like kilometers per hour or beats per minute, also do not have common units.

We also study breakdown frontiers and derive asymptotic distributional results for the average treatment effect $\text{ATE} = \ensuremath{\mathbb{E}}(Y_1 - Y_0)$ and quantile treatment effects $\text{QTE}(\tau) = Q_{Y_1}(\tau) - Q_{Y_0}(\tau)$ for $\tau \in (0,1)$. Because the breakdown frontier analysis for these two parameters is nearly identical, we focus on ATE for brevity. Since ATE does not rely on the dependence structure between potential outcomes, assumptions regarding rank invariance do not affect the identified set. Hence, for a specific conclusion like $\text{ATE} \geq 0$, the breakdown frontiers are vertical lines. This was not the case for the breakdown frontier for claims about $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$, where the assumptions on independence of treatment assignment and on rank invariance interact.

Figure (ref) illustrates how the breakdown frontier for ATE changes as our conclusion of interest changes. Specifically, consider the conclusion that \[ \text{ATE} \geq \mu \] for five different values for $\mu$. The figure shows the corresponding breakdown frontiers. As $\mu$ increases, we are making a stronger claim about the true parameter and hence the set of assumptions for which the conclusion holds shrinks. For strong enough claims, the claim may be refuted even with the strongest assumptions possible. This happens when $\ensuremath{\mathbb{E}}(Y \mid X=1) - \ensuremath{\mathbb{E}}(Y \mid X=0) < \mu$. Conversely, as $\mu$ decreases, we are making progressively weaker claims about the true parameter, and hence the set of assumptions for which the conclusion holds grows larger.

Breakdown frontiers can be defined for many simultaneous claims. For example, figure (ref) illustrates breakdown frontiers for the joint claim that $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$ and $\text{ATE} \geq 0$, for five different values of $\underline{p}$. Compare these plots to those of figures (ref) and (ref). For small $\underline{p}$, many pairs $(c,t)$ lead to the conclusion $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$. But many of these pairs do not also lead to the conclusion $\ensuremath{\mathbb{E}}(Y_1 - Y_0) \geq 0$, and hence those pairs are removed. As $\underline{p}$ increases, however, the constraint that we also want to conclude $\text{ATE} \geq 0$ eventually does not bind. Similarly, if we look at the breakdown frontier solely for our claim about ATE, many points $(c,t)$ are included which are ruled out when we also wish to make our claim about $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$. This is especially true as $\underline{p}$ gets larger.

Although we focus on one-sided claims like $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$, one may also be interested in the simultaneous claim that $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$ and $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \leq \overline{p}$, where $\underline{p}<\overline{p}$. Similarly to the joint one-sided claims about $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ and ATE discussed above, the breakdown frontier for this two-sided claim on $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ can be obtained by taking the minimum of the two breakdown frontier functions for each one-sided claim separately. Equivalently, the area under the breakdown frontier for the two-sided claim is just the intersection of the areas under the breakdown frontiers for the one-sided claims.

Inference

As mentioned above, although identified sets are used in its definition, the breakdown frontier is always point identified. Hence for estimation and inference we use mostly standard nonparametric techniques. We first propose a consistent sample-analog estimator of the breakdown frontier. We then do inference via two asymptotically valid bootstrap procedures. Our main approach, discussed in section (ref), relies on important work by Dumbgen1993, FangSantos2014, and HongLi2015. We discuss an alternative approach in supplemental appendix D. In both approaches we construct one-sided lower uniform confidence bands for the breakdown frontier. A one-sided band is appropriate because our inferential goal is to determine the set of assumptions for which we can still draw our conclusion. If $\alpha$ denotes the nominal size, then approximately $100(1-\alpha)$% of the time, all elements $(c,t)$ which are below the confidence band lead to population level identified sets for the parameters of interest which allow us to conclude that our conclusions of interest hold; we also discuss a testing interpretation in supplemental appendix A. We examine the finite sample performance of our main approach in several Monte Carlo simulations.

An empirical illustration of the breakdown frontier approach

Finally, we use our results to examine the role of assumptions in determining the effects of child soldiering on wages, as studied in BlattmanAnnan2010. We illustrate how our results can be used as a sensitivity analysis within a larger empirical study. Specifically, we begin by first estimating and doing inference on parameters under point identifying assumptions. Next, we estimate breakdown frontiers for several claims about these parameters. We present our one-sided confidence bands as well. We then use these breakdown frontier point estimates and confidence bands to judge the sensitivity of our conclusion to the point identifying assumptions: A large inner confidence set for the robust region which includes many different points $(c,t)$ suggests robust results while a small inner confidence set close to the origin suggests non-robust results.

Motivation for the parameter $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$

Our empirical illustration uses data from BlattmanAnnan2010, which is part of a large literature on the impact of compulsory military service on wages. As CardCardoso2012 note,

quoteAngrist1990 showed that Vietnam-era draftees had lower earnings than non-draftees, a finding he attributed to the low value of military experience in the civilian labor market. Subsequent research in the United States and other countries, however, has uncovered a surprisingly mixed pattern of impacts.”

On one hand, enlistees learn basic skills and receive occupational training in the military. On the other hand, they forgo civilian schooling and work experience, and may experience debilitating psychological trauma. Which of these two explanations is correct? Most likely, both. By using models where the treatment effects $Y_1 - Y_0$ are heterogeneous, we allow some people to have overall positive effects, and thus satisfy the first explanation, and some people to have overall negative effects, and thus satisfy the second explanation. The parameter $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ tells us precisely which proportion of the population primarily satisfies the first versus the second explanation. Thus it gives researchers a more nuanced understanding of treatment effects, by quantitatively measuring the prevalence of two opposing explanations of the impact of treatment on outcomes. We suspect this parameter can be particularly helpful in literatures which find a “mixed pattern of impacts,” like the work on compulsory military service and wages discussed above.

That said, one may be concerned that the rank invariance assumption used to point identify this parameter is too strong. As in HeckmanSmithClements1997, this concern motivates our sensitivity analysis. For further motivation of this parameter in various settings, see BedoyaEtAl2017 and Mullahy2018.

Model and identification

To illustrate the breakdown frontier approach, we study the standard potential outcomes model with a binary treatment. We focus on two key assumptions: random assignment and rank invariance. We discuss how to relax these two assumptions and derive identified sets for various parameters under these relaxations. While there are many different ways to relax these assumptions, our goal is to illustrate the breakdown frontier methodology and hence we focus on just one kind of relaxation for each assumption.

Setup

Let $Y_1$ and $Y_0$ denote the unobserved potential outcomes. Let $X \in \{ 0,1\}$ be an observed binary treatment. Let $W\in \operatorname*{supp}(W)$ be a vector of observed covariates. This vector may contain discrete covariates, continuous covariates, or both. We observe the scalar outcome variable

equation[equation omitted — 68 chars of source]

Let $p_{x|w} = \ensuremath{\mathbb{P}}(X=x \mid W=w)$ denote the observed propensity score. We maintain the following assumption on the joint distribution of $(Y_1,Y_0,X,W)$ throughout.

\setcounter{partialIndepSection}{1}

partialIndepAssumpFor each $x,x' \in \{0,1\}$ and $w \in \operatorname*{supp}(W)$: \begin{enumerate} • $Y_x \mid X=x', W=w$ has a strictly increasing and continuous distribution function on its support, $\operatorname*{supp}(Y_x \mid X=x', W=w)$. • $\operatorname*{supp}(Y_x \mid X=x',W=w) = \operatorname*{supp}(Y_x \mid W=w) = [\underline{y}_x(w),\overline{y}_x(w)]$ where $-\infty \leq \underline{y}_x(w) < \overline{y}_x(w) \leq \infty$. • $p_{x|w} > 0$. \end{enumerate}

Via A(ref).(ref), we restrict attention to continuously distributed potential outcomes. A(ref).(ref) states that the support of $Y_x \mid X=x',W=w$ does not depend on $x'$, and is a possibly infinite closed interval. This assumption implies that the endpoints $\underline{y}_x(w)$ and $\overline{y}_x(w)$ are point identified. We maintain A(ref).(ref) for simplicity, but it can be relaxed using similar derivations as in MastenPoirier2016. A(ref).(ref) is an overlap assumption.

Define the conditional rank random variables $U_0 = F_{Y_0 \mid W}(Y_0 \mid W)$ and $U_1 = F_{Y_1 \mid W}(Y_1 \mid W)$. Since $F_{Y_1 \mid W}(\cdot \mid w)$ and $F_{Y_0 \mid W}(\cdot \mid w)$ are strictly increasing (by A(ref).(ref)), $U_0 \mid W$ and $U_1 \mid W$ are uniformly distributed on $[0,1]$. The value of unit $i$'s conditional rank random variable $U_x$ tells us where unit $i$ lies in the distribution of $Y_x \mid W$.

Identifying assumptions

It is well known that the joint conditional distribution of potential outcomes $(Y_1,Y_0) \mid W$ is point identified under two assumptions:

enumerate• Conditional random assignment of treatment: $X \protect\mathpalette{\protect\independenT}{\perp} Y_1 \mid W$ and $X \protect\mathpalette{\protect\independenT}{\perp} Y_0 \mid W$. • Conditional rank invariance: Conditional on $W$, $U_1 = U_0$ almost surely.

Note that the joint conditional independence assumption $X \protect\mathpalette{\protect\independenT}{\perp} (Y_1,Y_0) \mid W$ provides no additional identifying power beyond the marginal conditional independence assumption stated above. Any functional of $F_{Y_1,Y_0 \mid W}$ is point identified under these random assignment and rank invariance assumptions. The goal of our identification analysis is to study what can be said about such functionals when one or both of these point-identifying assumptions fails. To do this we define two classes of assumptions: one which indexes the relaxation of random assignment of treatment, and one which indexes the relaxation of rank invariance. These classes of assumptions nest both the point identifying assumptions of random assignment and rank invariance and the opposite end of no constraints on treatment selection or the dependence structure between potential outcomes.

A key feature of the relaxations we use is that they are `orthogonal' in the sense that we can relax each of the two assumptions separately: The amount by which we relax one assumption does not constrain the amount by which we can relax the other assumption. This feature is important since a key goal of our analysis is to quantify the trade-off between relaxations of these two assumptions.

We begin with our measure of distance from conditional independence.

definitionLet $c$ be a scalar between 0 and 1. Say $X$ is conditionally $c$-dependent with $Y_x$ given $W$ if \begin{equation} \sup_{y \in \operatorname*{supp}(Y_x \mid W=w)} | \ensuremath{\mathbb{P}}(X=x \mid Y_x = y, W=w) - \ensuremath{\mathbb{P}}(X=x\mid W=w) | \leq c \end{equation} for $x \in \{ 0,1\}$ and $w\in \operatorname*{supp}(W)$.

For $c=0$, conditional $c$-dependence implies $X \protect\mathpalette{\protect\independenT}{\perp} Y_1 \mid W$ and $X \protect\mathpalette{\protect\independenT}{\perp} Y_0 \mid W$. For $c > 0$, however, it allows for some deviations from conditional independence. Specifically, it allows the unobserved treatment assignment probability $\ensuremath{\mathbb{P}}(X=1 \mid Y_x = y, W=w)$ to be at most $c$ probability units away from the observed propensity score $p_{1|w}$. We discuss one way to interpret the magnitude of $c$ on page (ref). We give further discussion in our previous paper MastenPoirier2017.

Our second class of assumptions constrains the dependence structure between $Y_1$ and $Y_0$, conditional on $W$. By Sklar's Theorem (Sklar1959), write \[ F_{Y_1,Y_0 \mid W}(y_1,y_0 \mid w) = C(F_{Y_1 \mid W}(y_1 \mid w),F_{Y_0 \mid W}(y_0 \mid w) \mid w) \] where $C(\cdot,\cdot \mid w)$ is a unique conditional copula function. See Nelsen2006 for an overview of copulas and FanPatton2014 for a survey of their use in econometrics. Restrictions on $C$ constrain the dependence between potential outcomes. For example, if \[ C(u_1,u_0 \mid w) = \min\{u_1,u_0\} \equiv M(u_1,u_0), \] then $U_1 = U_0$ almost surely, conditional on $W$. Thus conditional rank invariance holds. In this case the potential outcomes $Y_1$ and $Y_0$ are sometimes called conditionally comonotonic and $M$ is called the comonotonicity copula. At the opposite extreme, when $C$ is an arbitrary copula, the dependence between $Y_1 \mid W$ and $Y_0 \mid W$ is constrained only by the Fr\'echet-Hoeffding bounds, which state that \[ \max\{u_1 + u_0 -1, 0\} \leq C(u_1,u_0 \mid w) \leq M(u_1,u_0). \] We next define a class of assumptions which includes both conditional rank invariance and no assumptions on the dependence structure as special cases.

definitionThe potential outcomes $(Y_1,Y_0)$ satisfy (1$-t)$-percent conditional rank invariance given $W$ if for all $w \in \operatorname*{supp}(W)$ their conditional copula $C$ satisfies \begin{equation} C(u_1,u_0 \mid w) = (1-t) M(u_1,u_0) + t H(u_1,u_0 \mid w) \end{equation} where $M(u_1,u_0) = \min\{ u_1,u_0 \}$ and $H$ is some conditional copula.

This assumption says that within each covariate cell the population is a mixture of two parts: In one part, rank invariance holds. This part contains $100 \cdot t \,$% of the overall population in that cell. In the second part, rank invariance fails in an arbitrary, unknown way. Hence, for this part, the dependence structure is unconstrained beyond the Fr\'echet-Hoeffding bounds. This part contains $100(1-t)$% of the overall population in that cell. Thus for $t=0$ the usual conditional rank invariance assumption holds, while for $t=1$ no assumptions are made about the dependence structure. For $t \in (0,1)$ we obtain a kind of partial conditional rank invariance. Note that by exercise 2.3 on page 14 of Nelsen2006, a mixture of copulas like that in equation (ref) is also a copula.

To see this mixture interpretation formally, let $T$ follow a Bernoulli distribution with $\ensuremath{\mathbb{P}}(T=1 \mid W=w) = t$, where $T\protect\mathpalette{\protect\independenT}{\perp} Y_1 \mid W$ and $T \protect\mathpalette{\protect\independenT}{\perp} Y_0 \mid W$, but $T$ is not independent of $(Y_1,Y_0) \mid W$ jointly. This implies that $T$ has an effect on the dependence structure of $(Y_1,Y_0) \mid W$ but not on their conditional marginal distributions. Suppose that individuals for whom $T_i = 1$ have an arbitrary dependence structure, while those with $T_i=0$ have conditionally rank invariant potential outcomes. Then by the law of total probability,

align*[align* omitted — 483 chars of source]

Our approach to relaxing rank invariance is an example of a more general approach. In this approach we take a weak assumption and a stronger assumption and use them to define a continuous class of assumptions by considering the population as a mixture of two subpopulations. The weak assumption holds in one subpopulation while the stronger assumption holds in the other subpopulation. The mixing proportion $t$ continuously spans the two distinct assumptions we began with. This approach was used earlier by HorowitzManski1995 in their analysis of the contaminated sampling model. While this general approach may not always be the most natural way to relax an assumption, it is always available and hence can be used to facilitate breakdown frontier analyses.

Throughout the rest of this section we impose both conditional $c$-dependence and (1$-t$)-percent conditional rank invariance given $W$.

partialIndepAssump\begin{enumerate} • $X$ is conditionally $c$-dependent with the potential outcomes $Y_x$ given $W$, where for all $w \in \operatorname*{supp}(W)$ we have $c < \min\{ p_{1|w}, p_{0|w} \}$. • $(Y_1,Y_0)$ satisfies (1$-t$)-percent conditional rank invariance given $W$, where $t \in [0,1]$. \end{enumerate}

For brevity we focus on the case $c < \min \{ p_{1|w}, p_{0|w} \}$ throughout this paper. This allows us to explain the key ideas while keeping the notation and derivations relatively simple. All of our results, however, can be relaxed to the general case where $c \in [0,1]$.

Partial identification under relaxations of independence and rank invariance

We next study identification under the deviations from full independence and rank invariance defined above. We begin by briefly recalling results from MastenPoirier2017 on identification of the conditional quantile treatment effect $\text{CQTE}(\tau \mid w) =Q_{Y_1 \mid W}(\tau \mid w) - Q_{Y_0}(\tau \mid w)$, the conditional average treatment effect $\text{CATE}(w) = \ensuremath{\mathbb{E}}(Y_1 - Y_0 \mid W=w)$, and the conditional marginal cdfs of potential outcomes under $c$-dependence. We then derive new identification results for the distribution of treatment effects (DTE), $F_{Y_1 - Y_0}(z)$, and its related parameter $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$.

In MastenPoirier2017, we showed that A(ref) and A(ref).(ref) imply that the identified set for $\text{CQTE}(\tau \mid w)$ is\footnote{In that paper we also extended the bounds in equations (ref)--(ref) to the $c \geq \min \{ p_{1|w}, p_{0|w} \}$ case. As noted earlier, here we focus on the case $c < \min \{ p_{1|w}, p_{0|w} \}$ case for brevity.}

align[align omitted — 345 chars of source]

where

align[align omitted — 336 chars of source]

We further showed that, assuming $\ensuremath{\mathbb{E}}(| Y | \mid X=x, W=w) < \infty$ for $x \in \{0,1\}$ and all $w \in \operatorname*{supp}(W)$, the identified set for $\text{CATE}(w)$ is

equation[equation omitted — 256 chars of source]

and that

equation[equation omitted — 213 chars of source]

and

equation[equation omitted — 224 chars of source]

are functionally sharp bounds on the cdf $F_{Y_x \mid W}$. We use these conditional cdf bounds in our DTE bounds. Bounds on the corresponding unconditional parameters, like $\text{ATE} = \ensuremath{\mathbb{E}}[\text{CATE}(W)]$, can be obtained by integrating the conditional bounds over the marginal distribution of $W$. These results are unchanged if we further impose A(ref).(ref). That is, assumptions on rank invariance have no identifying power for functionals of the marginal distributions of potential outcomes.

We next derive the identified set for the distribution of treatment effects (DTE), the cdf \[ \text{DTE}(z) = \ensuremath{\mathbb{P}}(Y_1 - Y_0 \leq z). \] To do this, we first derive the identified set for the conditional distribution of treatment effects (CDTE), the cdf \[ \text{CDTE}(z \mid w) = \ensuremath{\mathbb{P}}(Y_1 - Y_0 \leq z \mid W=w). \] By the law of iterated expectations, \[ \text{DTE}(z) = \ensuremath{\mathbb{E}}[ \text{CDTE}(z \mid W)]. \] Thus we will obtain the identified set for the DTE by averaging bounds for the CDTE. While the ATE only depends on the conditional marginal distributions of potential outcomes, the CDTE depends on the joint distribution of $(Y_1,Y_0) \mid W$. Consequently, as we'll see below, the identified set for the CDTE depends on the value of $t$.

For any $z \in \ensuremath{\mathbb{R}}$ define $\mathcal{Y}_z(w) = [\underline{y}_1(w), \overline{y}_1(w)] \cap [\underline{y}_0(w) + z, \overline{y}_0(w) + z]$. Note that $\operatorname*{supp}(Y_1 - Y_0 \mid W=w) \subseteq [\underline{y}_1(w) - \overline{y}_0(w), \overline{y}_1(w) - \underline{y}_0(w)]$. Let $z$ be an element of $[\underline{y}_1(w) - \overline{y}_0(w), \overline{y}_1(w) - \underline{y}_0(w)]$ such that $\mathcal{Y}_z(w)$ is nonempty. If $z$ is such that $\mathcal{Y}_z(w)$ is empty, then the CDTE is either 0 or 1 depending solely on the relative location of the two supports, which is point identified by A(ref).(ref). In this case, define $\overline{\text{CDTE}}(z,c,t \mid w)$ and $\underline{\text{CDTE}}(z,c,t \mid w)$ to equal this point identified value. If $z > \overline{y}_1(w) - \underline{y}_0(w)$, define these CDTE bounds to equal 1. If $z < \underline{y}_1(w) - \overline{y}_0(w)$, define these CDTE bounds to equal 0.

If $\mathcal{Y}_z(w)$ is nonempty, define

multline*[multline* omitted — 342 chars of source]
multline*[multline* omitted — 329 chars of source]

where $U \sim \text{Unif}[0,1]$. The following result shows that (a) these are sharp bounds on the CDTE, and (b) the integral of these bounds over the marginal distribution of $W$ yields sharp bounds on the DTE, defined as $\ensuremath{\mathbb{P}}(Y_1 - Y_0 \leq z)$.

theorem[DTE bounds] Suppose the joint distribution of $(Y,X,W)$ is known. Suppose A(ref) and A(ref) hold. Let $z \in \ensuremath{\mathbb{R}}$. Then the identified set for $\ensuremath{\mathbb{P}}(Y_1 - Y_0 \leq z \mid W=w)$ is \[ [ \underline{\text{CDTE}}(z,c,t \mid w), \overline{\text{CDTE}}(z,c,t \mid w) ]. \] Moreover, the identified set for $\ensuremath{\mathbb{P}}(Y_1 - Y_0 \leq z)$ is \begin{multline*} [ DTE(z,c,t), \overline{DTE}(z,c,t) ] \\ = \left[ \int_{\operatorname*{supp}(W)} CDTE(z,c,t \mid w) \; dF_W(w), \int_{\operatorname*{supp}(W)}\overline{CDTE}(z,c,t \mid w) \; dF_W(w) \right]. \end{multline*}

The bound functions $\underline{\text{DTE}}(z,\cdot,\cdot)$ and $\overline{\text{DTE}}(z,\cdot,\cdot)$ are continuous and monotonic in both arguments. When both conditional random assignment ($c=0$) and conditional rank invariance ($t=0$) hold, these bounds collapse to a single point and we obtain point identification. If we impose conditional random assignment ($c=0$) but allow arbitrary dependence between $Y_1$ and $Y_0$ ($t=1$) then we obtain a conditional version of the well known Makarov Makarov1982 bounds. For example, see equation (2) of FanPark2010. DTE bounds have been studied extensively by Fan and coauthors; see the introduction of FanGuerreZhu2017 for a recent and comprehensive discussion of this literature.

Theorem (ref) immediately implies that the identified set for $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z)$ is \[ \ensuremath{\mathbb{P}}(Y_1 - Y_0 > z) \in [1 - \overline{\text{DTE}}(z,c,t), 1-\underline{\text{DTE}}(z,c,t)]. \] In particular, setting $z=0$ yields the proportion who benefit from treatment, $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$. Thus theorem (ref) allows us to study the sensitivity of this parameter to relaxations of full conditional independence and conditional rank invariance.

Finally, notice that all of the bounds and identified sets discussed in this section are analytically tractable and depend on just three functions identified from the population---the conditional cdf $F_{Y \mid X,W}$, the propensity scores $p_{x|w}$, and the marginal distribution of covariates $F_W$. This suggests a plug-in estimation approach which we study in section (ref).

Breakdown frontiers

We now formally define the breakdown frontier, which generalizes the scalar breakdown point to multiple assumptions or dimensions. We also define the robust region, the area below the breakdown frontier. These objects can be defined for different conclusions about different parameters in various models. For concreteness, however, we focus on just a few conclusions about $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z)$ and ATE in the potential outcomes model discussed above.

We begin with the conclusion that $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z) \geq \underline{p}$ for a fixed $\underline{p} \in [0,1]$ and $z \in \ensuremath{\mathbb{R}}$. For example, if $z = 0$ and $\underline{p} = 0.5$, then this conclusion states that at least 50% of people have higher outcomes with treatment than without. If we impose conditional random assignment and conditional rank invariance, then $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z)$ is point identified and hence we can directly check whether this conclusion holds. But the breakdown frontier approach asks: Relative to these baseline assumptions, what are the weakest assumptions that allow us to draw this conclusion, given the observed distribution of $(Y,X,W)$? Specifically, since larger values of $c$ and $t$ correspond to weaker assumptions, what are the largest values of $c$ and $t$ such that we can still definitively conclude that $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z) \geq \underline{p}$?

We answer this question in two steps. First we gather all values of $c$ and $t$ such that the conclusion holds. We call this set the robust region. Since the lower bound of the identified set for $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z)$ is $1 - \overline{\text{DTE}}(z,c,t)$ (by theorem (ref)), the robust region for the conclusion that $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z) \geq \underline{p}$ is

align*[align* omitted — 203 chars of source]

The robust region is simply the set of all $(c,t)$ which deliver an identified set for $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z)$ which lies on or above $\underline{p}$. See pages 60--61 of Stoye2005 for similar definitions in the scalar assumption case in a different model. Since $\overline{\text{DTE}}(z,c,t)$ is increasing in $c$ and $t$, the robust region will be empty if $\overline{\text{DTE}}(z,0,0) > 1-\underline{p}$, and non-empty if $\overline{\text{DTE}}(z,0,0) \leq 1-\underline{p}$. That is, if the conclusion of interest does not hold under the point identifying assumptions, it certainly will not hold under weaker assumptions. From here on we only consider the first case, where the conclusion of interest holds under the point identifying assumptions. That is, we suppose $\overline{\text{DTE}}(z,0,0) \leq 1-\underline{p}$ so that $\text{RR}(z,\underline{p}) \neq \emptyset$.

The breakdown frontier is the set of points $(c,t)$ on the boundary of the robust region. Specifically, for the conclusion that $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$, this frontier is the set \[ \text{BF}(\underline{p}) = \{ (c,t) \in[0,1]^2 : \overline{\text{DTE}}(0,c,t) = 1-\underline{p} \}. \] Solving for $t$ in the equation $\overline{\text{DTE}}(0,c,t) = 1-\underline{p}$ yields

equation[equation omitted — 95 chars of source]

where \[ \text{num} = 1 - \underline{p} - \int_{\operatorname*{supp}(W)}\ensuremath{\mathbb{P}}(\underline{Q}^c_{Y_1 \mid W}(U \mid w) - \overline{Q}^c_{Y_0 \mid W}(U \mid w) \leq 0) \; dF_W(w) \] and

multline*[multline* omitted — 352 chars of source]

Thus we obtain the following analytical expression for the breakdown frontier as a function of $c$: \[ \text{BF}(c,\underline{p}) = \min\{ \max \{ \text{bf}(c,\underline{p}), 0 \}, 1 \}. \] This frontier provides the largest relaxations $c$ and $t$ which still allow us to conclude that $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$. It thus provides a quantitative measure of robustness of this conclusion to relaxations of the baseline point identifying assumptions of conditional random assignment and conditional rank invariance. Moreover, the shape of this frontier allows us to understand the trade-off between these two types of relaxations in drawing our conclusion. We illustrate this trade-off between assumptions in our empirical illustration of section (ref).

We next consider breakdown frontiers for ATE. Consider the conclusion that $\text{ATE} \geq \mu$ for some $\mu \in \ensuremath{\mathbb{R}}$. Analogously to above, the robust region for this conclusion is \[ \text{RR}_\textsc{ate}(\mu) = \{ (c,t) \in [0,1]^2 : \underline{\text{ATE}}(c) \geq \mu \} \] and the breakdown frontier is \[ \text{BF}_\textsc{ate}(\mu) = \{ (c,t) \in [0,1]^2 : \underline{\text{ATE}}(c) = \mu \}. \] These sets are nonempty if $\underline{\text{ATE}}(0) \geq \mu$; that is, if our conclusion holds under the point identifying assumptions. As we mentioned earlier, rank invariance has no identifying power for ATE, and hence the breakdown frontier is a vertical line at the point \[ c^* = \inf\{ c \in [0,1] : \underline{\text{ATE}}(c) \leq \mu \}. \] This point $c^*$ is a breakdown point for the conclusion that $\text{ATE} \geq \mu$. Note that continuity of $\underline{\text{ATE}}(\cdot)$ implies $\underline{\text{ATE}}(c^*) = \mu$. Thus we've seen two kinds of breakdown frontiers so far: The first had nontrivial curvature, which indicates a trade-off between the two assumptions. The second was vertical in one direction, indicating a lack of identifying power of that assumption.

We can also derive robust regions and breakdown frontiers for more complicated joint conclusions. For example, suppose we are interested in concluding that both $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$ and $\text{ATE} \geq \mu$ hold. Then the robust region for this joint conclusion is just the intersection of the two individual robust regions: \[ \text{RR}(0,\underline{p}) \cap \text{RR}_\textsc{ate}(\mu). \] The breakdown frontier for the joint conclusion is just the boundary of this intersected region. Viewing these frontiers as functions mapping $c$ to $t$, the breakdown frontier for this joint conclusion can be computed as the minimum of the two individual frontier functions. For example, see figure (ref) on page (ref).

Above we focused on one-sided conclusions about the parameters of interest. Another natural joint conclusion is the two-sided conclusion that $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z) \geq \underline{p}$ and $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z) \leq \overline{p}$, for $0 \leq \underline{p} < \overline{p} \leq 1$. No new issues arise here: the robust region for this joint conclusion is still the intersection of the two separate robust regions. Keep in mind, though, that whether we look at a one-sided or a two-sided conclusion is unrelated to the fact that we use lower confidence bands in section (ref).

Finally, the bootstrap procedures we propose in section (ref) can also be used to do inference on these joint breakdown frontiers. For simplicity, though, in that section we focus on the case where we are only interested in the conclusion $\ensuremath{\mathbb{P}}(Y_1 - Y_0 > z) \geq \underline{p}$.

Estimation and inference

In this section we study estimation and inference on the breakdown frontier defined above. The breakdown frontier is a known functional of the conditional cdf of outcomes given treatment and covariates, the probability of treatment given covariates, and the marginal distribution of the covariates. Hence we propose nonparametric sample analog plug-in estimators of the breakdown frontier. We derive $\sqrt{N}$-consistency and asymptotic distributional results using a delta method for directionally differentiable functionals. We then use a bootstrap procedure to construct asymptotically valid lower confidence bands for the breakdown frontier. We conclude by discussing selection of the tuning parameter for this bootstrap procedure.

Although we focus on inference on the breakdown frontier, one might also be interested in doing inference directly on the parameters of interest. If we fix $c$ and $t$ a priori then we obtain identified sets for ATE, QTE, and the DTE from section (ref). Our asymptotic results below may be used as inputs to traditional inference on partially identified parameters. See CanayShaikh2017 for a survey of this literature.

To establish our main asymptotic results, we present a sequence of results. Each result focuses on a different component of the overall breakdown frontier: (1) the bounds on the marginal distributions of potential outcomes conditional on $W$, (2) the CQTE bounds, (3) the CATE and ATE bounds, (4) breakdown points for ATE, (5) the CDTE under conditional rank invariance but without full conditional independence, (6) the DTE without either conditional rank invariance or full conditional independence, and finally (7) the breakdown frontier itself.

We first suppose we observe a random sample of data.

partialIndepAssumpThe random variables $\{ (Y_i,X_i,W_i) \}_{i=1}^N$ are independently and identically distributed according to the distribution of $(Y,X,W)$.

We assume the support of $W$ is discrete. We sketch an approach to handling continuous covariates in appendix (ref). Note that $W$ may still be a vector.

partialIndepAssumpThe support of $W$ is discrete and finite. Let $\operatorname*{supp}(W) = \{w_1,\ldots, w_K\}$.

All parameters of interest are defined as functionals of the underlying parameters $F_{Y \mid X,W}(y \mid x,w)$, $p_{x|w} = \ensuremath{\mathbb{P}}(X=x \mid W=w)$, and $q_w = \ensuremath{\mathbb{P}}(W=w)$. Let \[ \widehat{F}_{Y \mid X,W}(y \mid x,w) = \frac{\frac{1}{N} \sum_{i=1}^N \ensuremath{\mathbbm{1}}(Y_i \leq y) \ensuremath{\mathbbm{1}}(X_i = x ,W_i = w) }{ \frac{1}{N} \sum_{i=1}^N \ensuremath{\mathbbm{1}}(X_i = x, W_i = w)}, \] \[ \widehat{p}_{x|w} = \frac{\frac{1}{N} \sum_{i=1}^N \ensuremath{\mathbbm{1}}(X_i = x, W_i = w)}{\frac{1}{N} \sum_{i=1}^N \ensuremath{\mathbbm{1}}(W_i = w)}, \] and \[ \widehat{q}_w = \frac{1}{N} \sum_{i=1}^N \ensuremath{\mathbbm{1}}(W_i = w) \] denote the sample analog estimators of these three quantities, which converge uniformly to a Gaussian process at a $\sqrt{N}$-rate; see lemma (ref) in appendix (ref).

Next consider the bounds (ref) and (ref) on the marginal distributions of potential outcomes. These population bounds are a functional $\phi_1$ evaluated at $(F_{Y \mid X,W}(\cdot \mid \cdot, \cdot),p_{(\cdot|\cdot)},q_{(\cdot)})$ where $p_{(\cdot|\cdot)}$ denotes the probability $p_{x|w}$ as a function of $(x,w)\in\{0,1\}\times \operatorname*{supp}(W)$, and $q_{(\cdot)}$ denotes $q_w$ as a function of $w \in \operatorname*{supp}(W)$. We estimate these bounds by a plug-in estimator $\phi_1(\widehat{F}_{Y \mid X,W}(\cdot \mid \cdot, \cdot),\widehat{p}_{(\cdot|\cdot)}, \widehat{q}_{(\cdot)})$. If this functional is differentiable in an appropriate sense, $\sqrt{N}$-convergence in distribution of its arguments will carry over to the functional by the delta method. The type of differentiability we require is Hadamard directional differentiability, first defined by Shapiro1990 and Dumbgen1993, and further studied in FangSantos2014.

definitionLet $\mathbb{D}$ and $\mathbb{E}$ be Banach spaces with norms $\|\cdot\|_{\mathbb{D}}$ and $\|\cdot\|_{\mathbb{E}}$. Let $\mathbb{D}_\phi \subseteq \mathbb{D}$ and $\mathbb{D}_0 \subseteq \mathbb{D}$. The map $\phi:\mathbb{D}_\phi \rightarrow\mathbb{E}$ is Hadamard directionally differentiable at $\theta\in \mathbb{D}_\phi$ tangentially to $\mathbb{D}_0$ if there is a continuous map $\phi'_\theta:\mathbb{D}_0 \rightarrow \mathbb{E}$ such that \begin{align*} \lim_{n\rightarrow\infty} \left\| \frac{\phi(\theta + t_nh_n) - \phi(\theta)}{t_n} - \phi'_\theta(h)\right\|_{\mathbb{E}} &= 0 \end{align*} for all sequences $\{h_n\}\subset \mathbb{D}$ and $\{t_n\} \in \ensuremath{\mathbb{R}}_+$ such that $t_n \searrow 0$, $\|h_n - h\|_{\mathbb{D}}\rightarrow 0$, $h\in\mathbb{D}_0$ as $n\rightarrow\infty$, and $\theta + t_nh_n \in \mathbb{D}_\phi$ for all $n$.

If we further have that $\phi_\theta'$ is linear, then we say $\phi$ is Hadamard differentiable (proposition 2.1 of FangSantos2014). Not every Hadamard directional derivative $\phi_\theta'$ must be linear, however.

We use the functional delta method for Hadamard directionally differentiable mappings (e.g., theorem 2.1 in FangSantos2014) to show convergence in distribution of our estimators. Such convergence is usually to a non-Gaussian limiting process. We do not use this distribution to do inference since obtaining analytical asymptotic confidence bands would be challenging. Instead, we use a bootstrap procedure to obtain asymptotically valid uniform confidence bands for our breakdown frontier and associated estimators.

Returning to our population bounds (ref) and (ref), we estimate these by

align[align omitted — 540 chars of source]

Note that these estimators may not perform well when $c$ is close to $p_{x|w}$. In our analysis we assume $c$ is bounded away from $p_{x|w}$.

In addition to assumption A(ref), we make the following regularity assumptions.

partialIndepAssump\begin{enumerate} • For each $x \in \{0,1 \}$ and $w\in\operatorname*{supp}(W)$, $-\infty < \underline{y}_x(w) < \overline{y}_x(w) < +\infty$. • For each $x\in\{0,1\}$ and $w\in\operatorname*{supp}(W)$, $F_{Y \mid X,W}(y \mid x,w)$ is continuously differentiable everywhere with density $f_{Y \mid X,W}(y \mid x,w)$ uniformly continuous in $y$, uniformly bounded from above, and uniformly bounded away from zero on $\operatorname*{supp}(Y \mid X=x,W=w)$. \end{enumerate}

A(ref).(ref) combined with our earlier assumption A(ref).(ref) constrain the potential outcomes to have compact support. This compact support assumption is not used to analyze our cdf bounds estimators (ref), but we use it later to obtain estimates of the corresponding quantile function bounds uniformly over their arguments $u \in (0,1)$, which we then use to estimate the bounds on $\ensuremath{\mathbb{P}}(Q_{Y_1 \mid W}(U \mid w) - Q_{Y_0 \mid W}(U \mid w) \leq z)$. This is a well known issue when estimating quantile processes; for example, see Vaart2000 lemma 21.4(ii). A(ref).(ref) requires the density of $Y \mid X,W$ to be bounded away from zero uniformly. This ensures that conditional quantiles of $Y \mid X,W$ are uniquely defined. It also implies that the limiting distribution of the estimated quantile bounds will be well-behaved. Uniform continuity of the density implies that the derivatives of the conditional quantile function with respect to $\tau$ are uniformly continuous.

For some of our main results in this section (lemmas (ref), (ref), (ref) and theorem (ref)) we establish convergence uniformly over $c \in \mathcal{C}$ for some finite grid $\mathcal{C} =\{c_1,c_2,\ldots,c_J\} \subset [0, \min \{p_{1|w}, p_{0|w}\} )$. We discuss the choice of these grid points on page (ref). We constrain this grid to be below $\min \{p_{1|w}, p_{0|w}\}$ solely for simplicity, as all our results can be extended to grids $\mathcal{C}\subset [0,1]$ by combining our present bound estimates with estimates based on the $c \geq \min \{ p_{1|w}, p_{0|w} \}$ case given in MastenPoirier2017. Weak convergence does not hold uniformly over an interval of $c$ since some of the functionals below are not Hadamard directionally differentiable when their codomain is a set of functions on that interval. To resolve this issue, we propose two ways of conducting inference on the breakdown frontier uniformly over intervals of $c$. The first is to use the fixed grid and monotonicity of the breakdown frontier to construct a uniform band. The second is to smooth the population breakdown frontier such that it is Hadamard differentiable when viewed as a function of $c$. We use the first approach in this section and the second approach in supplemental appendix D.

The next result establishes convergence in distribution of the cdf bound estimators. Here and below we use the following notation: For an arbitrary set $\mathcal{A}$ and a Banach space $\mathcal{B}$, $\ell^\infty(\mathcal{A},\mathcal{B})$ denotes the set of all maps $z: \mathcal{A} \rightarrow \mathcal{B}$ with finite sup-norm $\| z \| = \sup_{a \in \mathcal{A}} \| z(a) \|_{\mathcal{B}}$, equipped with this norm. For example, see VaartWellner1996.

lemmaSuppose A(ref), A(ref), and A(ref) hold. Let $\mathcal{Y} \subset \ensuremath{\mathbb{R}}$ be a finite grid of points. Then \begin{align*} \sqrt{N} \begin{pmatrix} \widehat{\overline{F}}^c_{Y_x \mid W}(y \mid w) - \overline{F}^c_{Y_x \mid W}(y \mid w) \\ \widehat{F}^c_{Y_x \mid W}(y \mid w) - F^c_{Y_x \mid W}(y \mid w) \end{pmatrix} &\rightsquigarrow \mathbf{Z}_2(y,x,w,c), \end{align*} a tight random element of $\ell^\infty(\mathcal{Y} \times \{0,1\} \times \operatorname*{supp}(W)\times \mathcal{C},\ensuremath{\mathbb{R}}^2)$.

$\mathbf{Z}_2$ is not Gaussian itself, but it is a continuous transformation of Gaussian processes. For given $(x,c,w)$, the limit will be Gaussian at all values of $y$ except for \[ y \in\left\{ Q_{Y \mid X,W}\left(\frac{p_{x|w} - c}{2p_{x|w}} \mid x,w\right), \ Q_{Y \mid X,W}\left(\frac{p_{x|w} + c}{2p_{x|w}} \mid x,w\right)\right\}. \] A characterization of this limiting process is given in the proof in appendix (ref).

Next consider the conditional quantile bounds (ref), which we estimate by

align*[align* omitted — 323 chars of source]

The next result establishes uniform convergence in distribution of these quantile bounds estimators. For the following results, let $\overline{C} \in (0,\min\{p_{1|w},p_{0|w}\})$ for all $w \in \operatorname*{supp}(W)$.

lemmaSuppose A(ref), A(ref), A(ref), and A(ref) hold. Then \begin{align*} \sqrt{N} \begin{pmatrix} \widehat{\overline{Q}}^c_{Y_x \mid W}(\tau \mid w) - \overline{Q}^c_{Y_x \mid W}(\tau \mid w) \\ \widehat{Q}^c_{Y_x \mid W}(\tau \mid w) - Q^c_{Y_x \mid W}(\tau \mid w) \end{pmatrix} &\rightsquigarrow \mathbf{Z}_3(\tau,x,w,c), \end{align*} a mean-zero Gaussian process in $\ell^\infty((0,1) \times \{0,1\} \times \operatorname*{supp}(W) \times [0,\overline{C}],\ensuremath{\mathbb{R}}^2)$ with continuous paths.

This result is uniform in $c$ on an interval, in $x\in\{0,1\}$, $w\in\operatorname*{supp}(W)$, and in $\tau\in(0,1)$. This result directly implies convergence over $c\in\mathcal{C}$ as well. Unlike the distribution of the cdf bounds estimators, this process is Gaussian. This follows by Hadamard differentiability of the mapping between $\theta_0 \equiv (F_{Y \mid X,W}(\cdot \mid \cdot,\cdot), p_{(\cdot|\cdot)}, q_{(\cdot)})$ and the conditional quantile bounds. By applying the functional delta method, we can show asymptotic normality of smooth functionals of these bounds. A first set of functionals are the CQTE bounds of equation (ref), which are a linear combination of the quantile bounds. Let \[ \widehat{\underline{\text{CQTE}}}(\tau,c \mid w) = \widehat{\underline{Q}}^c_{Y_1 \mid W}(\tau \mid w) - \widehat{\overline{Q}}^c_{Y_0 \mid W}(\tau \mid w) \] and \[ \widehat{\overline{\text{CQTE}}}(\tau,c \mid w) = \widehat{\overline{Q}}^c_{Y_1 \mid W}(\tau \mid w) - \widehat{\underline{Q}}^c_{Y_0 \mid W}(\tau \mid w). \] Then, \[ \sqrt{N}

pmatrix[pmatrix omitted — 225 chars of source]

\rightsquigarrow

pmatrix[pmatrix omitted — 177 chars of source]

, \] where the superscript $\mathbf{Z}^{(j)}$ denotes the $j$th component of the vector $\mathbf{Z}$.

A second set of functionals are the CATE bounds from equation (ref). These bounds are smooth linear functionals of the CQTE bounds. Therefore the joint asymptotic distribution of these bounds can be established by the continuous mapping theorem. Let \[ \widehat{\underline{\text{CATE}}}(c \mid w) = \int_0^1 \widehat{\underline{\text{CQTE}}}(u,c \mid w) \; du \qquad \text{and} \qquad \widehat{\overline{\text{CATE}}}(c \mid w) = \int_0^1 \widehat{\overline{\text{CQTE}}}(u,c \mid w) \; du. \] Then, by the linearity of the integral operator, these estimated CATE bounds converge to their population counterpart at a $\sqrt{N}$-rate and therefore

align*[align* omitted — 574 chars of source]

a mean-zero Gaussian process in $\ell^\infty(\operatorname*{supp}(W) \times [0,\overline{C}],\ensuremath{\mathbb{R}}^2)$ with continuous paths.

We estimate the unconditional $\text{ATE}$ bounds by integrating over the empirical distribution of the covariates $W$: Let \[ \widehat{\underline{\text{ATE}}}(c) = \frac{1}{N} \sum_{i=1}^N \widehat{\underline{\text{CATE}}}(c \mid W_i) \qquad \text{and} \qquad \widehat{\overline{\text{ATE}}}(c) = \frac{1}{N} \sum_{i=1}^N \widehat{\overline{\text{CATE}}}(c \mid W_i). \] The following decomposition implies that the estimated ATE upper bound converges weakly to a Gaussian element:

align*[align* omitted — 528 chars of source]

A similar result holds for the estimated ATE lower bound.

Next consider estimation of the breakdown point for the claim that $\text{ATE}\geq \mu$ where $\mu \in \ensuremath{\mathbb{R}}$. To focus on the nondegenerate case, suppose the population value of ATE obtained under full independence is greater than $\mu$, $\underline{\text{ATE}}(0) > \mu$ which implies $c^* > 0$. Let \[ \widehat{c}^* = \inf\{c\in[0,1]: \widehat{\underline{\text{ATE}}}(c) \leq \mu\} \] be the estimated breakdown point. This is the estimated smallest relaxation of independence such that we cannot conclude that the ATE is strictly greater than $\mu$. By the properties of the quantile bounds as a function of $c$, the function $\underline{\text{ATE}}(c)$ is non-increasing and differentiable in $c$. We now present a result about the asymptotic distribution of $\widehat{c}^*$.

propositionSuppose A(ref), A(ref), A(ref), and A(ref) hold. Assume $c^*\in(0,\overline{C}]$. Then $\sqrt{N}(\widehat{c}^* - c^*) \rightsquigarrow \mathbf{Z}_{4}$, a Gaussian random variable.

The assumption that $c^*\in(0,\overline{C}]$ can again be relaxed to the general case where $c^*\in(0,1]$ but we maintain the stronger assumption for brevity.

Under conditional rank invariance, we can also establish asymptotic normality of bounds for $\ensuremath{\mathbb{P}}(Q_{Y_1 \mid W}(U \mid w) - Q_{Y_0 \mid W}(U \mid w) \leq z)$ for a fixed $z \in \ensuremath{\mathbb{R}}$. These bounds are given by \[ (\underline{P}(c \mid w),\overline{P}(c \mid w)) \equiv (\underline{\text{CDTE}}(z,c,0 \mid w),\overline{\text{CDTE}}(z,c,0 \mid w)). \] We keep $z$ implicit in the notation for these bounds. Estimates for these quantities are provided by

align*[align* omitted — 388 chars of source]

Asymptotic normality can be established using the Hadamard directional differentiability of the mapping from the differences in quantile bounds to the bounds $(\underline{P}(c \mid w),\overline{P}(c \mid w))$. This mapping is called the pre-rearrangement operator. ChernozhukovFernandez-ValGalichon2010 showed that this operator was Hadamard differentiable when the quantile functions are continuously differentiable for all $u \in (0,1)$. In our case, the underlying quantile functions are continuously differentiable on $(0,1/2)\cup(1/2,1)$, and continuous but not differentiable at $u=1/2$. At this value, the left and right derivatives exist and are finite, but are generally different from one another. We extend the result of ChernozhukovFernandez-ValGalichon2010 to the case where the quantile function has a point of non-differentiability by showing Hadamard directional differentiability of this mapping.

To do so, we make additional assumptions on the behavior of these quantile functions.

partialIndepAssumpFor each $c \in \mathcal{C}$ and $w\in\operatorname*{supp}(W)$, \begin{enumerate} • The number of elements in each of the sets \begin{align*} \mathcal{U}^*_1(c \mid w) = \{u\in(0,1) &: \partial_u^- (\overline{Q}^c_{Y_1 \mid W}(u \mid w) - Q^c_{Y_0 \mid W}(u \mid w)) = 0 \\ &\quad or \partial_u^+ (\overline{Q}^c_{Y_1 \mid W}(u \mid w) - Q^c_{Y_0 \mid W}(u \mid w)) = 0\} \\[0.5em] \mathcal{U}^*_2(c \mid w) = \{u\in(0,1) &: \partial_u^- (Q^c_{Y_1 \mid W}(u \mid w) - \overline{Q}^c_{Y_0 \mid W}(u \mid w)) = 0 \\ &\quad or \partial_u^+ (Q^c_{Y_1 \mid W}(u \mid w) - \overline{Q}^c_{Y_0 \mid W}(u \mid w)) = 0\} \end{align*} is finite. • The following hold. \begin{enumerate} • For any $u\in\mathcal{U}^*_1(c \mid w)$, $\overline{Q}^c_{Y_1 \mid W}(u \mid w) - \underline{Q}^c_{Y_0 \mid W}(u \mid w) \neq z$. • For any $u\in\mathcal{U}^*_2(c \mid w)$, $\underline{Q}^c_{Y_1 \mid W}(u \mid w) - \overline{Q}^c_{Y_0 \mid W}(u \mid w) \neq z$. \end{enumerate} \end{enumerate}

These assumptions imply that the respective function's derivatives change signs a finite number of times, and therefore they cross the horizontal line at $z$ a finite number of times. These functions are continuously differentiable in $u$ everywhere on $(0,1/2)\cup(1/2,1)$, and are directionally differentiable at $1/2$. The second assumption rules out the functions being flat when exactly valued at $z$. Failure of the second condition in this assumption implies that convergence will hold uniformly over any compact subset that excludes these values, which typically form a measure-zero set. Therefore this assumption can be satisfied by considering convergence for values of $c$ which exclude those where the second part of assumption A(ref) fails. Without knowing a priori at which values this assumption may fail, selecting grid points randomly from a continuous distribution ensures that these values are selected with probability zero.

An alternative approach to inference if the second condition fails for some values of $c$ is to smooth the population function using methods described in supplemental appendix D. Like in ChernozhukovFernandez-ValGalichon2010, we require a tuning parameters to control the level of smoothing. We show that $\sqrt{N}$-convergence holds for all parameter values when introducing any amount of fixed smoothing.

Finally, note that A(ref) is refutable, since it is expressed as a function of identified quantities, namely the CQTE bounds for all $u \in (0,1)$.

With this additional assumption we can show $\sqrt{N}$-convergence of the bounds uniformly in $\operatorname*{supp}(W)\times \mathcal{C}$.

lemmaSuppose A(ref), A(ref), A(ref), A(ref), and A(ref) hold. Then \[ \sqrt{N} \begin{pmatrix} \widehat{\overline{P}}(c \mid w) - \overline{P}(c \mid w) \\ \widehat{\underline{P}}(c \mid w) - \underline{P}(c \mid w) \end{pmatrix} \rightsquigarrow \mathbf{Z}_{5}(w,c), \] a tight random element in $\ell^\infty(\operatorname*{supp}(W)\times\mathcal{C},\ensuremath{\mathbb{R}}^2)$.

If conditional random assignment holds ($c=0$) in addition to conditional rank invariance ($t=0$), then the CDTE is point identified and lemma (ref) gives the asymptotic distribution of the sample analog CDTE estimator (in this case the upper and lower bound functions are equal). This can be considered an estimator of the CDTE in one of the models of Matzkin2003.

We now establish the limiting distribution of the CDTE bounds uniformly in $(c,t,w)$. Let

align*[align* omitted — 528 chars of source]

We estimate the unconditional DTE bounds by integrating the estimated CDTE bounds over the empirical distribution of the covariates: Let \[ \widehat{\underline{\text{DTE}}}(z,c,t) = \frac{1}{N}\sum_{i=1}^N \widehat{\underline{\text{CDTE}}}(z,c,t \mid W_i) \qquad \text{and} \qquad \widehat{\overline{\text{DTE}}}(z,c,t) = \frac{1}{N}\sum_{i=1}^N \widehat{\overline{\text{CDTE}}}(z,c,t \mid W_i). \]

We have shown in lemma (ref) that the terms $\underline{P}(c \mid w)$ and $\overline{P}(c \mid w)$ are estimated at a $\sqrt{N}$-rate by the Hadamard directional differentiability of the mapping linking empirical cdfs and these terms. We now show that the second components of the CDTE bounds are a Hadamard directionally differentiable functional as well, leading to the $\sqrt{N}$ joint convergence of the DTE bounds to a tight, random element uniformly in $c$ and $t$.

lemmaFix $z \in \ensuremath{\mathbb{R}}$. Suppose A(ref), A(ref), A(ref), A(ref), and A(ref) hold. Then \begin{equation} \sqrt{N} \begin{pmatrix} \widehat{\overline{DTE}}(z,c,t) - \overline{DTE}(z,c,t) \\ \widehat{DTE}(z,c,t) - DTE(z,c,t) \end{pmatrix} \rightsquigarrow \mathbf{Z}_6(c,t), \end{equation} a tight random element of $\ell^\infty(\mathcal{C}\times[0,1],\ensuremath{\mathbb{R}}^2)$ with continuous paths.

Having established the convergence in distribution of the DTE, we can now show that the breakdown frontier also converges in distribution uniformly over its arguments. Denote the estimated breakdown frontier for the conclusion that $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$ by

equation[equation omitted — 142 chars of source]

where

equation[equation omitted — 139 chars of source]

with

align*[align* omitted — 376 chars of source]

By combining our previous lemmas, we can show that the estimated breakdown frontier converges in distribution.

theoremSuppose A(ref), A(ref), A(ref), A(ref), and A(ref) hold. Let $\mathcal{P}\subset [0,1]$ be a finite grid of points. Then \[ \sqrt{N}(\widehat{\text{BF}}(c,\underline{p}) - \text{BF}(c,\underline{p})) \rightsquigarrow \mathbf{Z}_7(c,\underline{p}), \] a tight random element of $\ell^{\infty}(\mathcal{C}\times\mathcal{P})$.

This result essentially follows from the convergence of the preliminary estimators established in lemma (ref) in appendix (ref) and by showing that the breakdown frontier is a composition of a number of Hadamard differentiable and Hadamard directionally differentiable mappings, implying convergence in distribution of the estimated breakdown frontier.

Breakdown frontiers for more complex conclusions can typically be constructed from breakdown frontiers for simpler conclusions. For example, consider the breakdown frontier for the joint conclusion that $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$ and $\text{ATE} \geq \mu$. Then the breakdown frontier for this joint conclusion is the minimum of the two individual frontier functions. Alternatively, consider the conclusion that $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$ or $\text{ATE} \geq \mu$, or both, hold. Then the breakdown frontier for this joint conclusion is the maximum of the two individual frontier functions. Since the minimum and maximum operators are Hadamard directionally differentiable, these joint breakdown frontiers will also converge in distribution.

Since the limiting process is non-Gaussian, inference on the breakdown frontier is not based on standard errors as with Gaussian limiting theory. Our processes' distribution is characterized fully by the expressions in appendix (ref), but obtaining analytical estimates of quantiles of functionals of these processes would be challenging. In the next subsection we give details on the bootstrap procedure we use to construct confidence bands for the breakdown frontier.

Bootstrap inference

As mentioned earlier, we use a bootstrap procedure to do inference on the breakdown frontier rather than directly using its limiting process. In this subsection we discuss how to use the bootstrap to approximate this limiting process. In the next subsection we discuss its application to constructing uniform confidence bands.

First we define some general notation. Let $Z_i = (Y_i,X_i,W_i)$ and $Z^N = \{Z_1,\ldots,Z_N\}$. Let $\theta_0$ denote some parameter of interest and let $\widehat{\theta}$ be an estimator of $\theta_0$ based on the data $Z^N$. Let $\mathbf{A}_N^*$ denote $\sqrt{N}(\widehat{\theta}^* - \widehat{\theta})$ where $\widehat{\theta}^*$ is a draw from the nonparametric bootstrap distribution of $\widehat{\theta}$. Suppose $\ensuremath{\mathbf{A}}$ is the tight limiting process of $\sqrt{N} ( \widehat{\theta} - \theta_0)$. Denote bootstrap consistency by $\mathbf{A}_N^* \overset{P}{\rightsquigarrow} \mathbf{A}$ where $\overset{P}{\rightsquigarrow}$ denotes weak convergence in probability, conditional on the data $Z^N$. Weak convergence in probability conditional on $Z^N$ is defined as \[ \sup_{h\in \text{BL}_1} \left| \ensuremath{\mathbb{E}} [h(\mathbf{A}_N^*) \mid Z^N] - \ensuremath{\mathbb{E}} [h(\mathbf{A})] \right| = o_p(1) \] where $\text{BL}_1$ denotes the set of Lipschitz functions into $\ensuremath{\mathbb{R}}$ with Lipschitz constant no greater than 1.

We focus on the following specific choices of $\theta_0$ and $\widehat{\theta}$: \[ \theta_0 =

pmatrix[pmatrix omitted — 95 chars of source]

\qquad and \qquad \widehat{\theta} =

pmatrix[pmatrix omitted — 125 chars of source]

. \] For these choices, let $\mathbf{Z}_N^* = \sqrt{N}(\widehat{\theta}^* - \widehat{\theta})$. Let $\mathbf{Z}_1$ denote the limiting distribution of $\sqrt{N}(\widehat{\theta} - \theta_0)$; see lemma (ref) in appendix (ref). Theorem 3.6.1 of VaartWellner1996 implies that $\mathbf{Z}_N^* \overset{P}{\rightsquigarrow} \mathbf{Z}_1$. Our parameters of interest are all functionals $\phi$ of $\theta_0$. For Hadamard differentiable functionals $\phi$, the nonparametric bootstrap is consistent. For example, see theorem 3.1 of FangSantos2014. They further show that $\phi$ is Hadamard differentiable if and only if \[ \sqrt{N}(\phi(\widehat{\theta}^*) - \phi(\widehat{\theta})) \overset{P}{\rightsquigarrow} \phi'_{\theta_0}(\mathbf{Z}_1) \] where $\phi_{\theta_0}'$ denotes the Hadamard derivative at $\theta_0$. This implies that the nonparametric bootstrap can be used to do inference on the QTE and ATE bounds since they are Hadamard differentiable functionals of $\theta_0$. A second implication is that the nonparametric bootstrap is not consistent for the DTE or for the breakdown frontier for claims about the DTE since they are Hadamard directionally differentiable mappings of $\theta_0$, but they are not ordinary Hadamard differentiable.

In such cases, FangSantos2014 show that a different bootstrap procedure is consistent. Specifically, let $\widehat{\phi}_{\theta_0}'$ be a consistent estimator of $\phi'_{\theta_0}$. Then their results imply that \[ \widehat{\phi}_{\theta_0}'(\mathbf{Z}_N^*) \overset{P}{\rightsquigarrow} \phi'_{\theta_0}(\mathbf{Z}_1). \] Analytical consistent estimates of $\phi'_{\theta_0}$ are often difficult to obtain, so Dumbgen1993 and HongLi2015 propose using a numerical derivative estimate of $\phi'_{\theta_0}$. Their estimate of the limiting distribution of $\sqrt{N} ( \phi( \widehat{\theta}) - \phi(\theta_0) )$ is given by the distribution of

equation[equation omitted — 269 chars of source]

across the bootstrap estimates $\widehat{\theta}^*$. Under the rate constraints $\varepsilon_N \rightarrow 0$ and $\sqrt{N} \varepsilon_N \rightarrow \infty$, and some measurability conditions stated in their appendix, HongLi2015 show \[ \widehat{\phi}_{\theta_0}'(\sqrt{N}(\widehat{\theta}^* - \widehat{\theta})) \overset{P}{\rightsquigarrow} \phi'_{\theta_0}(\mathbf{Z}_1). \] where the left hand side is defined in equation (ref).

This bootstrap procedure requires evaluating $\phi$ at two values, which is computationally simple. It also requires selecting the tuning parameter $\varepsilon_N$, which we discuss later. Note that the standard, or naive, bootstrap is a special case of this numerical delta method bootstrap where $\varepsilon_N = N^{-1/2}$.

Uniform confidence bands for the breakdown frontier

In this subsection we combine all of our asymptotic results thus far to construct uniform confidence bands for the breakdown frontier. As in section (ref) we use the function $\text{BF}(\cdot,\underline{p})$ to characterize this frontier. We specifically construct one-sided lower uniform confidence bands. That is, we will construct a lower band function $\widehat{\text{LB}}(c)$ such that \[ \lim_{N \rightarrow \infty} \ensuremath{\mathbb{P}} \left( \widehat{\text{LB}}(c) \leq \text{BF}(c,\underline{p}) \text{ for all $c \in [0,1]$} \right) = 1-\alpha. \] We use a one-sided lower uniform confidence band because this gives us an inner confidence set for the robust region. Specifically, define the set \[ \text{RR}_L = \{ (c,t) \in [0,1]^2 : t \leq \widehat{\text{LB}}(c) \}. \] Then validity of the confidence band $\widehat{\text{LB}}$ implies \[ \lim_{N \rightarrow \infty} \ensuremath{\mathbb{P}} \left( \text{RR}_L \subseteq \text{RR}(0,\underline{p}) \right) = 1-\alpha. \] Thus the area underneath our confidence band, $\text{RR}_L$, is interpreted as follows: Across repeated samples, approximately $100(1-\alpha)$% of the time, every pair $(c,t) \in \text{RR}_L$ leads to a population level identified set for the parameter $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ which lies weakly above $\underline{p}$. Put differently, approximately $100(1-\alpha)$% of the time, every pair $(c,t) \in \text{RR}_L$ still lets us draw the conclusion we want at the population level. Hence the size of this set $\text{RR}_L$ is a finite sample measure of robustness of our conclusion to failure of the point identifying assumptions. We discuss an alternative testing-based interpretation on page 5 in supplemental appendix A.

One might be interested in constructing one-sided upper confidence bands if the goal was to do inference on the set of assumptions for which we cannot come to the conclusion of interest. This might be useful in situations where two opposing sides are debating a conclusion. But since our focus is on trying to determine when we can come to the desired conclusion, rather than looking for when we cannot, we only describe the one-sided lower confidence band case.

When studying inference on scalar breakdown points, KlineSantos2013 constructed one-sided lower confidence intervals. Unlike for breakdown frontiers, uniformity over different points in the assumption space is not a concern for inference on breakdown points. See supplemental appendix A for more discussion.

We consider bands of the form \[ \widehat{\text{LB}}(c) = \widehat{\text{BF}}(c,\underline{p}) - \widehat{k}(c) \] for some function $\widehat{k}(\cdot) \geq 0$. This band is an asymptotically valid lower uniform confidence band of level $1-\alpha$ if \[ \lim_{N \rightarrow \infty} \ensuremath{\mathbb{P}} \left( \widehat{\text{BF}}(c,\underline{p}) - \widehat{k}(c) \leq \text{BF}(c,\underline{p}) \text{ for all $c \in [0,1]$} \right) = 1-\alpha, \] or, equivalently, if \[ \lim_{N \rightarrow \infty} \ensuremath{\mathbb{P}} \left( \sup_{c\in[0,1]} \sqrt{N} \Big( \widehat{\text{BF}}(c,\underline{p}) - \text{BF}(c,\underline{p}) - \widehat{k}(c) \Big) \leq 0\right) = 1-\alpha. \]

In our theoretical analysis, we consider $\widehat{k}(c) = \widehat{z}_{1-\alpha} \sigma(c)$ for a scalar $\widehat{z}_{1-\alpha}$ and a function $\sigma$. We focus on known $\sigma$ for simplicity. We start by deriving a uniform band over a grid $\mathcal{C}$, then extend it over an interval using monotonicity of the breakdown frontier. As discussed earlier, we only derive uniformity of the band over $c \in [0,\overline{C}]$ rather than over $c \in [0,1]$, but this is also for brevity and can be relaxed. The choice of $\sigma$ affects the shape of the confidence band, and there are many possible choices of the function $\sigma$ which yield valid level $1-\alpha$ uniform confidence bands. See FreybergerRai2016 for a detailed analysis. A simple choice of $\sigma$ is the constant function: $\sigma(c) = 1$, which delivers an equal width uniform band. Alternatively, as we do below, one could choose $\sigma(c)$ to construct a minimum width confidence band (equivalently, maximum area of $\text{RR}_L$).

propositionSuppose A(ref), A(ref), A(ref), and A(ref) hold. Define $\phi : \ell^\infty(\ensuremath{\mathbb{R}} \times \{0,1\},\ensuremath{\mathbb{R}}^2) \rightarrow \ell^\infty(\mathcal{C})$ such that \[ \widehat{\text{BF}}(c,\underline{p}) = [\phi(\widehat{\theta})](c). \] Then $\phi$ is Hadamard directionally differentiable. Suppose that $\varepsilon_N \rightarrow 0$ and $\sqrt{N} \varepsilon_N \rightarrow\infty$. Let $\widehat{\theta}^*$ denote a draw from the nonparametric bootstrap distribution of $\widehat{\theta}$. Then \begin{align} \left[ \widehat{\phi}_{\theta_0}'(\sqrt{N}(\widehat{\theta}^* - \widehat{\theta})) \right] \overset{P}{\rightsquigarrow} \left[ \phi'_{\theta_0}(\mathbf{Z}_1) \right] &\equiv \mathbf{Z}_7. \end{align} For a given function $\sigma(\cdot)$ such that $\inf_{c\in\mathcal{C}}\sigma(c)>0$, define \begin{equation} \widehat{z}_{1-\alpha} = \inf\left\{ z \in \ensuremath{\mathbb{R}} : \ensuremath{\mathbb{P}}\left(\sup_{c\in\mathcal{C}}\frac{ \left[ \widehat{\phi}_{\theta_0}'(\sqrt{N}(\widehat{\theta}^* - \widehat{\theta})) \right] (c,p)}{\sigma(c)} \leq z \mid Z^N\right)\geq 1-\alpha\right\}. \end{equation} Finally, suppose also that the cdf of \[ \sup_{c\in\mathcal{C}}\frac{[\phi_{\theta_0}'(\mathbf{Z}_1) ](c,\underline{p})}{\sigma(c)} = \sup_{c\in \mathcal{C}}\frac{\mathbf{Z}_7(c,\underline{p})}{\sigma(c)} \] is continuous and strictly increasing at its $1-\alpha$ quantile, denoted $z_{1-\alpha}$. Then $\widehat{z}_{1-\alpha} = z_{1-\alpha} + o_p(1)$.

This proposition is a variation of corollary 3.2 in FangSantos2015workingPaper. As a consequence of this result, the lower $1-\alpha$ band $\widehat{\text{LB}}(c) = \widehat{\text{BF}}(c,\underline{p}) - \widehat{z}_{1-\alpha}\sigma(c)$ is valid uniformly on the grid $\mathcal{C}$. To extend the uniformity to all of $[0,\overline{C}]$ we propose the following lower confidence band: \[ \widetilde{LB}(c) =

cases\widehat{LB}(c_1) & if $c\in [0,c_1]$ \\ \quad \vdots \\ \widehat{LB}(c_j) & if $c \in (c_{j-1},c_j]$, for $j=2,\ldots,J$ \\ \quad \vdots \\ \quad 0 & if $c \in (c_J, \overline{C}]$.

\] This band is a step function which interpolates between grid points using the least monotone interpolation. The following result shows its validity.

corollaryLet the assumptions of proposition (ref) hold. Then, $\widetilde{\text{LB}}(c)$ is a uniform lower $1-\alpha$ band for $\text{BF}(c,\underline{p})$ over $c\in[0,\overline{C}]$.

Corollary (ref) shows that, for any fixed $J \geq 1$, the interpolated lower confidence band preserves exact the $1-\alpha$ coverage on the grid points. This follows by monotonicity of the breakdown frontier; see lemma (ref) in appendix (ref). That said, this interpolated band might not be taut, in the sense that there may exist other lower bands with $1-\alpha$ coverage that are weakly larger than $\widetilde{\text{LB}}(c)$ for all $c$ and strictly larger at some values of $c$. See FreybergerRai2016 for further discussion of taut confidence bands.

Proposition (ref) can be extended to estimated functions $\sigma$, although we leave the details for future work. We use an estimated $\sigma$ in our application, as described next. When both $z_{1-\alpha}$ and $\sigma$ are estimated, we work directly with $\widehat{k}(c) = \widehat{z}_{1-\alpha} \widehat{\sigma}(c)$. We choose $\widehat{k}(c)$ to minimize an approximation to the area between the confidence band and the estimated function; equivalently, to maximize the area of $\text{RR}_L$. Specifically, we let $\widehat{k}(c_1),\ldots,\widehat{k}(c_J)$ solve \[ \min_{k(c_1),\ldots,k(c_J) \geq 0} \; \sum_{j=2}^J k(c_j)(c_{j} - c_{j-1}) \] subject to \[ \ensuremath{\mathbb{P}}\left(\sup_{c\in\{c_1,\ldots,c_J\}} \sqrt{N} \Big(\widehat{\text{BF}}(c,\underline{p}) - \text{BF}(c,\underline{p}) - k(c) \Big) \leq 0 \right) = 1-\alpha, \] where we approximate the left-hand side probability via the numerical delta method bootstrap. The criterion function here is just a right Riemann sum over the grid points. This optimization is not computationally costly: It is only performed once per value of $\underline{p}$ and $\varepsilon_N$. Moreover, in our empirical illustration it takes an average of 15 seconds per run on a mid-2013 MacBook Air.

Choosing the grid points $\mathcal{C}$

Here we suggest three approaches for choosing the number and location of the grid points $\mathcal{C} = \{ c_1,\ldots,c_J \}$. First, one can let $\mathcal{C} = \{ \overline{c}_1,\ldots, \overline{c}_K \}$ where $K$ is the number of observed covariates and, for each $k \in \{1,\ldots,K\}$, $\overline{c}_k$ is the maximal deviation between the observed propensity score and the “leave out variable $k$” propensity score, which we define and discuss in our empirical illustration on page (ref). There we argue that these are natural points to consider.

Second, researchers can choose equally spaced grid points, for a fixed $J$. Third, researchers can randomly select grid points from a continuous distribution, for a fixed $J$. As mentioned on page (ref), this random selection ensures that assumption A(ref) holds with probability one. Both of these approaches are commonly used in the literature. Moreover, both of these approaches can be used in combination with the first approach. In this case, researchers may want to begin with $\{ \overline{c}_1,\ldots, \overline{c}_K \}$ and then add a multiple $m$ of $K$ additional points, so that the total number of points $J$ is $K + mK$ for some positive integer $m$.

Finally, we emphasize that for our asymptotic approximations to be valid, the key restriction is that the grid $\mathcal{C}$ is not too dense around the points $c$ of nondifferentiability. Picking a sufficiently sparse fixed finite grid is one way to ensure this. Thus, although we state our formal results for a fixed grid, one could instead let $J \rightarrow \infty$ sufficiently slowly as $N \rightarrow \infty$. Alternatively, one could pre-estimate the points of nondifferentiability and then let $\mathcal{C}$ equal the entire domain, minus sufficiently large intervals around these points. The length of these removed intervals then shrinks asymptotically. HorowitzLee2012,HorowitzLee2017 discuss approaches like these in different settings.

Bootstrap selection of $\varepsilon_N$

While Dumbgen1993 and HongLi2015 provide rate constraints on $\varepsilon_N$, they do not recommend a procedure for picking $\varepsilon_N$ in practice. In this section, we suggest a heuristic bootstrap method for picking $\varepsilon_N$. We use this method for our empirical illustration in section (ref); we also present the full range of bands considered. Since the question of choosing $\varepsilon_N$ goes beyond the purpose of the present paper, we defer a formal analysis of this method to future research. For discussions of bootstrap selection of tuning parameters in other problems, see Taylor1989, LegerRomano1990, Marron1992, and CaoCuevasManteiga1994.

Fix a $\underline{p}$. Let $\text{CP}_N(\varepsilon ; F_{Y,X,W})$ denote the finite sample coverage probability of our confidence band as described above, for a fixed $\varepsilon$. This statistic depends on the unknown distribution of the data, $F_{Y,X,W}$. The bootstrap replaces $F_{Y,X,W}$ with an estimator $\widehat{F}_{Y,X,W}$. We pick a grid $\{ \varepsilon_1,\ldots,\varepsilon_K \}$ of $\varepsilon$'s and let $\widehat{\varepsilon}_N$ solve \[ \min_{k =1,\ldots, K} | \text{CP}_N(\varepsilon_k ; \widehat{F}_{Y,X,W}) - (1-\alpha) |. \] We compute $\text{CP}_N$ by simulation. In our empirical illustration, we take $B=500$ draws. We use the same grid of $\varepsilon$'s as in our Monte Carlo simulations in supplemental appendix C. Larger grids and larger values of $B$ can be chosen subject to computational constraints. We furthermore must choose an estimator $\widehat{F}_{Y,X,W}$. The nonparametric bootstrap uses the empirical distribution. We use the smoothed bootstrap (DeAngelisYoung1992, PolanskySchucany1997). Specifically, we estimate the distribution of $(X,W)$ by its empirical distribution. We then let $\widehat{F}_{Y \mid X,W}$ be a kernel smoothed cdf estimate of the conditional cdf of $Y|X,W$. We use the standard logistic cdf kernel and the method proposed by Hansen2004 to choose the smoothing bandwidths. We divide these bandwidths in half since this visually appears to better capture the shape of the conditional empirical cdfs, and since smaller order bandwidths are recommended for the smoothed bootstrap (section 4 of DeAngelisYoung1992).

Bootstrap consistency requires sufficient smoothness of the functional of interest in the underlying cdf. It may be that the lack of smoothness that requires us to use the methods of FangSantos2014 and HongLi2015 in the first place also cause the naive bootstrap to be inconsistent for approximating the distribution of $\text{CP}_N(\varepsilon; F_{Y,X,W})$. As mentioned earlier, formally investigating this issue is beyond the scope of this paper. Our goal here is merely to suggest a simple first-pass approach at choosing $\varepsilon_N$.

Empirical illustration: The effects of child soldiering

In this section we use our results to examine the impact of assumptions in determining the effects of child soldiering on wages. We first briefly discuss the background and then we present our analysis.

Background

We use data from phase 1 of SWAY, the Survey of War Affected Youth in northern Uganda, conducted by principal researchers Jeannie Annan and Chris Blattman (see AnnanBlattmanHorton2006). As BlattmanAnnan2010 discuss on page 882, a primary goal of this survey was to understand the effects of a twenty year war in Uganda, where “an unpopular rebel group has forcibly recruited tens of thousands of youth”. In that paper, they use this data to examine the impacts of abduction on educational, labor market, psychosocial, and health outcomes. In our illustration, we focus solely on the impact of abduction on wages.

Blattman and Annan note that self-selection into the military is a common problem in the literature studying the effects of military service on outcomes. They argue that forced recruitment in Uganda led to random assignment of military service in their data. They first provide qualitative evidence for this, based on interviews with former rebels who led raiding parties. After murdering and mutilating civilians, the rebels had no public support, making abduction the only means of recruitment. Youths were generally taken during nighttime raids on rural households. According to the former rebel leaders, “targets were generally unplanned and arbitrary; they raided whatever homesteads they encountered, regardless of wealth or other traits.”

This qualitative evidence is supported by their survey data, where Blattman and Annan show that most pre-treatment covariates are balanced across the abducted and nonabducted groups (see their table 2). Only two covariates are not balanced: year of birth and prewar household size. They say this is unsurprising because

quote“a youth's probability of ever being abducted depended on how many years of the conflict he fell within the [rebel group's] target age range. Moreover, abduction levels varied over the course of the war, so youth of some ages were more vulnerable to abduction than others. The significance of household size, meanwhile, is driven by households greater than 25 in number. We believe that rebel raiders, who traveled in small bands, were less likely to raid large, difficult-to-control households.” (Page 887)

Hence they use a selection-on-observables identification strategy, conditioning on these two variables.

While their evidence supporting the full conditional independence assumption is compelling, this assumption is still nonrefutable. Hence they apply the methods of Imbens2003 to analyze the sensitivity to this assumption. In this analysis they only consider one outcome variable, years of education. Likewise, as in Imbens2003, they only look at one parameter, the constant treatment effect in a fully parametric model.

We complement their results by applying the breakdown frontier methods we develop in this paper. We focus on the log-wage outcome variable. We look at both the average treatment effect and $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$, which was not studied in BlattmanAnnan2010.

Analysis

The original phase 1 SWAY data has 1216 males born between 1975 and 1991. Of these, wage data is available for 504 observations. 56 of these earned zero wages; we drop these and only look at people who earned positive wages. This leaves us with our main sample of 448 observations. In addition to this outcome variable, we let our treatment variable be an indicator that the person was not abducted. We include the two covariates discussed above, age when surveyed and household size in 1996. Additional covariates can be included, but we focus on just these two for simplicity.

Table (ref) shows summary statistics for these four variables. 36% of our sample were not abducted. Age ranges from 14 years old to 30 years old, with a median of 22 years old. Household size ranges from 2 people to 28, with a median of 8 people. Wages range from as low as 36 shillings to as high as about 83,300 shillings, with a median of 1,400 shillings.

table[table omitted — 1,484 chars of source]

Age has 17 support points and household size has 21 support points. Hence there are 357 total covariate cells. Including the treatment variable, this yields 714 total cells, compared to our sample size of 448 observations. Since we focus on unconditional parameters, having small or zero observations per cell is not a problem in principle. However, in the finite sample we have, to ensure that our estimates of the cdf bounds $\overline{F}_{Y_x \mid W}^c(y \mid w)$ and $\underline{F}_{Y_x \mid W}^c(y \mid w)$ are reasonably smooth in $y$, we collapse our covariates as follows. We replace age with a binary indicator of whether one is above or below the median age. Likewise, we replace household size with a binary indicator of whether one lived in a household with above or below median household size. This reduces the number of covariate cells to 4, giving 8 total cells including the treatment variable. This yields approximately 55 observations per cell. While this crude approach suffices for our illustration, in more extensive empirical analyses one may want to use more sophisticated methods. For example, we could use discrete kernel smoothing, as discussed in LiRacine2008, who also provide additional references. We also consider alternative coarsenings in supplemental appendix E.

Table (ref) shows unconditional comparisons of means of the outcome and the original covariates across the treatment and control groups. Wages for people who were not abducted are 702 shillings larger on average. People who were not abducted are also about 1.4 years younger than those who were abducted. People who were not abducted also had a slightly larger household size than those who were abducted. Only the difference in ages is statistically significant at the usual levels, but as in tables 2 and 3 of BlattmanAnnan2010 the standard errors can be decreased by including additional controls. These extra covariates are not essential for illustrating our breakdown frontier methods, however.

table[table omitted — 1,491 chars of source]

The point estimates in table (ref) are unconditional on the two covariates. Next we consider the conditional independence assumption, with age and household size in 1996 as our covariates. Under this assumption, our estimate of ATE is 890 [726.13] shillings when the outcome variable is level of wages, and is 0.21 [0.11] when the outcome variable is log wage.\footnote{Using their full set of control variables, BlattmanAnnan2010 estimate ATE to be 0.33 [0.15] when the outcome is log wage. See column 1 of their table 3.} To check the robustness of these point estimates to failure of conditional independence, we estimate the breakdown point $c^*$ for the conclusion $\text{ATE} \geq 0$, where we use log-wages as our outcome variable. We measure relaxations of conditional independence by our conditional $c$-dependence distance. The estimated breakdown point is $\widehat{c}^* = 0.041$. Based on this point estimate, for all $x \in \{0,1\}$ and $w \in \operatorname*{supp}(W)$ we can allow the conditional propensity scores $\ensuremath{\mathbb{P}}(X=x \mid Y_x = y, W=w)$ to vary $\pm 4$ percentage points around the observed propensity scores $\ensuremath{\mathbb{P}}(X=x \mid W=w)$ without changing our conclusion.

Is this a big or small amount of variation? Well, as a baseline, the upper bound on $c$ is about 0.73. This is an estimate of \[ \max_{w \in \operatorname*{supp}(W)} \max \{ \ensuremath{\mathbb{P}}(X=1 \mid W=w), \ensuremath{\mathbb{P}}(X=0 \mid W=w) \}. \] Any $c \geq 0.73$ would lead to the no assumptions identified set for ATE. In this sense, $0.041$ is quite small, which would suggest that our results are quite fragile. Next we examine variation in the observed propensity scores as we suggested in MastenPoirier2017. Specifically, we consider the difference between the “full” propensity score and the “leave out variable $k$” propensity score which omits variable $k$: Define \[ \overline{c}_\texttt{age} = \sup_{s=0,1} \sup_{a=0,1} | \widehat{\ensuremath{\mathbb{P}}}(X=1 \mid \texttt{age}=a,\texttt{hhSize}=s) - \widehat{\ensuremath{\mathbb{P}}}(X=1 \mid \texttt{hhSize} = s) | \] and \[ \overline{c}_\texttt{hhSize} = \sup_{a=0,1} \sup_{s=0,1} | \widehat{\ensuremath{\mathbb{P}}}(X=1 \mid \texttt{age}=a,\texttt{hhSize}=s) - \widehat{\ensuremath{\mathbb{P}}}(X=1 \mid \texttt{age} = a) |. \] Using these numbers as a reference, a robust result would have a breakdown point above one or both of the $\overline{c}$'s. In the data, we obtain $\overline{c}_\texttt{age} = 0.0625$ and $\overline{c}_\texttt{hhSize} = 0.0403$. The estimated breakdown point $\widehat{c}^* = 0.041$ is below $\overline{c}_\texttt{age}$ and approximately equal to $\overline{c}_\texttt{hhSize}$. This latter result suggests that perhaps our conclusion could be considered somewhat robust. Accounting for sampling uncertainty in the breakdown point, however, shows that the true breakdown point may be less than $\overline{c}_\texttt{hhSize}$. Overall, this suggests that our conclusion that $\text{ATE} \geq 0$ is not robust to relaxations of full conditional independence.

This argument for judging the plausibility of specific values of $c$ relies on using variation in the observed propensity score to ground our beliefs about reasonable variation in the unobserved propensity scores. The general question here is how one should quantitatively distinguish `large' and `small' relaxations of an assumption. This is an old and ongoing question in the sensitivity analysis literature, and much work remains to be done. For discussions on this point for different measures of deviations or relaxations from independence in various settings, see RotnitzkyRobinsScharfstein1998, Robins1999, Imbens2003, AltonjiElderTaber2005,AltonjiElderTaber2008, and Oster2016.

Next consider the parameter $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$. Since we define treatment as not being abducted, this parameter measures the proportion of people who earn higher wages when they are not abducted, compared to when they are abducted. For this parameter, we must make both the full conditional independence assumption and the conditional rank invariance assumption to obtain point identification. Under these assumptions, our point estimate is $0.67$ with a one-sided lower 95% CI of $[0.48, 1]$.

Is this point estimate robust to failures of full conditional independence and conditional rank invariance? We examine this question by estimating breakdown frontiers and corresponding confidence bands for the conclusion that $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq \underline{p}$. We do this for $\underline{p} = 0.1$, $0.25$, $0.5$ as in our Monte Carlo simulations in supplemental appendix C. We do not consider $\underline{p} = 0.75$ or $0.9$ since these values are larger than our point estimate under the baseline assumptions; they yield empty estimated robust regions. Besides picking a grid of $\underline{p}$'s a priori, one could let $\underline{p} = \widehat{p}_{0,0} / 2$, half the value of the parameter estimated under the baseline point identifying assumptions. In our application this is $0.34$; we omit this choice for brevity. Imbens2003 suggests a similar choice of cutoff in his approach. We use the same eight ratios of $\varepsilon_N / \varepsilon_N^\text{naive}$ as in our Monte Carlo simulations in supplemental appendix C.

figure[figure omitted — 670 chars of source]

Figure (ref) shows the results. As in our earlier plots, the horizontal axis plots $c$, relaxations of full conditional independence, while the vertical axis plots $t$, relaxations of conditional rank invariance. As mentioned earlier, the natural upper bound for $c$ is about $0.73$. Since all of the breakdown frontiers intersect the horizontal axis at much smaller values, we have cut off the part of the overall assumption space with $c \geq 0.2$. Remember that, for the following analysis, it's valid to examine various $(c,t)$ combinations since we use uniform confidence bands.

First consider the left plot, $\underline{p} = 0.1$. Since this is the weakest conclusion of the five we consider, the estimated breakdown frontier and the corresponding robust region are the largest among the three plots. If we impose full conditional independence, then our estimated frontier suggests that we can completely relax conditional rank invariance and still conclude that at least 10% of people benefit from not being forced into military service. Even accounting for sampling uncertainty, we can still draw this conclusion. Moreover, looking at all choices of $\varepsilon_N$---not just our selected one---the lowest the vertical intercept ever gets is about 61%. Next suppose we relax full conditional independence. Recall that the maximal relaxation between the observed propensity score and the “leave out variable $k$” propensity scores gave $\overline{c}_\texttt{age} = 0.0625$ and $\overline{c}_\texttt{hhSize} = 0.0403$. Both of these numbers are substantially smaller than the horizontal intercept of our selected confidence band. Hence, if we impose full conditional rank invariance, our conclusion that $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq 0.1$ is robust to relaxations of full conditional independence. Suppose instead that we think selection on unobservables is at most the largest $\overline{c}$ value, about $0.06$. Then for $c$'s in the range $[0,0.06]$, and accounting for sampling uncertainty, we can still conclude $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq 0.1$ so long as at least 30% of the population satisfies rank invariance. Thus we can relax full independence within this range without paying too high a cost in terms of requiring stronger rank invariance assumptions.

If we are willing to restrict selection on unobservables to be smaller than the largest $\overline{c}$ value, then we can allow for larger relaxations of conditional rank invariance. To quantify this trade-off, we can compute the difference between the values of the estimated breakdown frontier at two different points. As a starting point we recommend computing \[ \widehat{\text{BF}} (\overline{c}_{(K)}, \underline{p}) - \widehat{\text{BF}}(\overline{c}_{(K-1)}, \underline{p}) , \] where $K$ denotes the number of observed regressors, $\overline{c}_{(K)}$ denotes the largest value of $\overline{c}$, and $\overline{c}_{(K-1)}$ denotes the second largest value. One could also divide this difference by $\overline{c}_{(K)} - \overline{c}_{(K-1)}$ to get a secant line. In our empirical application, $\overline{c}_{(K)} = \overline{c}_\texttt{age}$ and $\overline{c}_{(K-1)} = \overline{c}_\texttt{hhSize}$. Thus \[ \widehat{\text{BF}}(\overline{c}_\texttt{age}, 0.1) - \widehat{\text{BF}}(\overline{c}_\texttt{hhSize}, 0.1) = -6\%. \] The corresponding secant line has slope about $-3$. Thus if we assume selection on unobservables is at most as large as the second largest amount of variation in “leave out variable $k$” propensity scores, we can allow for an additional 6% of the population to violate conditional rank invariance. Put differently, around these values of $c$, allowing latent conditional propensity scores to vary an extra 1 percentage point requires us to impose that an additional 3% of the population must satisfy conditional rank invariance. This rate of substitution generally increases as $c$ gets larger. Our ability to quantify this kind of trade-off between assumptions is a primary goal of our breakdown frontier analysis.

Overall, our results from this top left plot suggest that the conclusion that at least 10% of people benefit from not being forced into military service is robust to relaxations of full conditional independence up to twice the size we see between the observed and leave out variable $k$ propensity scores, depending on how much conditional rank invariance failure we allow. For relaxations of full conditional independence up to the largest value of $\overline{c}$, we can allow up to 70% of the population to deviate from conditional rank invariance, accounting for sampling uncertainty.

Next consider the middle plot, $\underline{p} = 0.25$. Since this is a stronger conclusion than the previous one, all the frontiers are shifted towards the origin. Consequently, by construction, this conclusion is not as robust as the other one. Our qualitative conclusions, however, as similar to those obtained for $\underline{p} = 0.1$. If we impose full conditional independence we can allow conditional rank invariance to fail for about 70% of the population. Conversely, if we impose full conditional rank invariance, we can allow the latent conditional propensity scores to vary by about 10 percentage points---well beyond the largest observed variation $\overline{c}$. For $\underline{p} = 0.25$, we have \[ \widehat{\text{BF}}(\overline{c}_\texttt{age}, 0.25) - \widehat{\text{BF}}(\overline{c}_\texttt{hhSize}, 0.25) = -17\%. \] Hence the slope around our observed maximal $\overline{c}$'s is much larger for $\underline{p}=0.25$ as compared to $\underline{p}=0.1$. An important caveat to our conclusions for both $\underline{p} = 0.1$ and $\underline{p} = 0.25$ is that there is substantial variation in confidence bands as $\varepsilon_N$ changes. This point underscores the need for future work on the choice of $\varepsilon_N$.

Next consider the right plot, $\underline{p} = 0.5$. Here we consider the conclusion that at least half of people benefit from not being forced into military service. If we impose full conditional independence, and accounting for sampling uncertainty, then we can allow conditional rank invariance to fail for about 25% of the population. This is quite large, but it relies on full conditional independence holding exactly. If we also relax conditional independence to $c = 0.03$ then we need conditional rank invariance to hold for everyone if we still want to conclude that at least 50% of people benefit from not being forced into military service. $0.03$ is smaller than both $\overline{c}_\texttt{age}$ and $\overline{c}_\texttt{hhSize}$. Hence we might not be comfortable with such small values of $c$. This suggests the data do not definitively support the conclusion $\ensuremath{\mathbb{P}}(Y_1 > Y_0) \geq 0.5$, even though our point estimate under the baseline assumptions is $0.67$.

In this section we used our breakdown frontier methods to study the robustness of conclusions about ATE and $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ to failures of conditional independence and conditional rank invariance. We first considered the conclusion that the average treatment effect of not being abducted on log wages is nonnegative. Our point estimates suggest that this conclusion is robust to deviations in unobserved latent propensity scores up to the same value as $\overline{c}_\texttt{age}$, which is also about two-thirds as large as $\overline{c}_\texttt{hhSize}$; this robustness does not hold up when accounting for sampling uncertainty, however. We then considered the conclusion that at least $\underline{p}$% of people earn higher wages when they are not abducted. This conclusion is robust to large simultaneous relaxations of conditional rank invariance and conditional independence for $\underline{p} =$ 10%. For $\underline{p} =$ 25%, This conclusion continues to be robust to reasonable relaxations, although after accounting for the variation in confidence bands over $\varepsilon_N$, this conclusion appears to be more sensitive to conditional independence than to conditional rank invariance. This robustness to rank invariance matches the findings of HeckmanSmithClements1997, who imposed full independence and studied deviations from rank invariance. In their table 5B they found that, in their empirical application, one could generally conclude that $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ was at least 50%, regardless of the assumption on rank invariance. In our empirical application our results are not quite as robust to rank invariance failures, which could be because we use a different measure of relaxation of rank invariance, and also because of differences in the empirical applications.

Conclusion

Summary

In this paper we advocated the breakdown frontier approach to sensitivity analysis. Given a set of baseline assumptions, this approach defines the population breakdown frontier as the weakest set of assumptions such that a specific conclusion of interest holds. Sample analog estimates and lower uniform confidence bands allow researchers to do inference on this frontier. The area under the confidence band is a quantitative, finite sample measure of the robustness of a conclusion to relaxations of point-identifying assumptions. To examine this robustness, empirical researchers can present these estimated breakdown frontiers and their accompanying confidence bands along with traditional point estimates and confidence intervals obtained under point identifying assumptions. We illustrated this general approach in the context of a treatment effects model, where the robustness of conclusions about ATE and $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ to relaxations of random assignment and rank invariance are examined. We applied these results in an empirical study of the effect of child soldiering on wages. We found that weak conclusions about $\ensuremath{\mathbb{P}}(Y_1 > Y_0)$ are fairly robust to failures of both rank invariance and random assignment, but stronger conclusions are more sensitive to relaxations of random assignment.

Breakdown frontier analysis for other models and other relaxations

As discussed in section (ref), breakdown frontier analysis can in principle be done in most models. In that section we outlined the six main steps required for any breakdown frontier analysis. In this paper we illustrated this general approach by studying a single important and widely used model: the potential outcomes model with a binary treatment. In future work it would be helpful to perform breakdown frontier analyses in other models. In particular, it may be possible to do breakdown frontier analyses in a large class of models by using the general identification analysis in ChesherRosen2015 or Torgovitsky2015.

A key conceptual step in any breakdown frontier analysis is deciding how to define the indexed classes of assumptions such that the magnitude of the relaxation can be reasonably interpreted. This is not easy, and will generally depend on the model, the specific kind of assumption being relaxed, and the empirical context. Moreover, this choice may affect our findings: A conclusion can be robust with respect to one measure of relaxation but not another. Thus one goal of future research is to explore this space of assumption relaxations, to understand their substantive interpretations, and to chart their implications for the robustness of empirical findings. In MastenPoirier2016 we have already compared three different measures of relaxation of the random assignment assumption, including the one used here. We further studied quantile independence, a common relaxation of random assignment, in MastenPoirier2018QI. In the present paper, we also used a general method for spanning two discrete assumptions by defining a (1$-t$)-percent relaxation, as we did with rank invariance. But much work still remains to be done.

\singlespacing