EconBase
← Back to paper

Assessing Sensitivity to IV Exclusion and Exogeneity without First Stage Monotonicity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

117,170 characters · 21 sections · 80 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Assessing Sensitivity to IV Exclusion and Exogeneity without First Stage Monotonicity

abstractExclusion and exogeneity are core assumptions in instrumental variable (IV) analyses, but their empirical validity is often debated. This paper develops new sensitivity analyses for these assumptions. Our results accommodate arbitrary heterogeneity in treatment effects and do not impose any monotonicity requirements on the first stage. Specifically, we derive identified sets for the marginal distributions of potential outcomes and their functionals, like average treatment effects, under a broad class of nonparametric relaxations of the exclusion and exogeneity assumptions. These identified sets are characterized as solutions to linear programs and have desirable theoretical properties. We explain how to estimate these solutions using computationally tractable methods even when the linear program is infinite-dimensional. We illustrate these methods with an empirical application to peer effects in movie viewership, using weather as a potentially imperfect instrument.

JEL classification: C14, C18, C21, C26, C51

Keywords: Instrumental Variables, Sensitivity Analysis, Nonparametric Identification, Partial Identification

\allowdisplaybreaks

\onehalfspacing

Introduction

Instrumental variable (IV) analyses typically rely on two core assumptions: Instrument exclusion and instrument exogeneity. Exclusion holds when the instrument has no direct effect on the outcome, while exogeneity holds when the instrument is randomly assigned. Since the work of ImbensAngrist1994, a third assumption is also often imposed: First stage monotonicity. In the simplest setting where the treatment and instrument are binary, monotonicity holds if the instrument's effect on treatment is always of the same sign.

All three assumptions can be hard to justify in some empirical settings. Instruments may have direct effects on outcomes or may not be randomly assigned. Monotonicity can also fail. This occurs, for example, in leniency designs (also called `judge IV' designs) where monotonicity implies that a judge is stricter or more lenient in the face of any possible case; see FrandsenLefgrenLeslie2023Judging for details. In designs with many treatment and instrument values, there is no single monotonicity assumption to choose from, and it may be difficult to find one suitable for one's empirical setting.

To address these concerns, we study identification of treatment effects in a setting where no monotonicity conditions are imposed whatsoever, but where exogeneity and exclusion are assumed to at least partially hold in some sense. Specifically, we introduce a unifying class of continuous relaxations of instrument exclusion and exogeneity that nests several prominent approaches in the literature. In particular, it includes as special cases the marginal sensitivity model (MSM) of Tan2006, $c$-dependence by MastenPoirier2018, and supremum distance approaches, as in Manski1983 and KlineSantos2013. All these approaches were developed as sensitivity models for unconfoundedness (or selection on observables), but we develop modified versions suitable for IV sensitivity analysis. In each of those cases, the sensitivity model is indexed by a scalar, unit-free sensitivity parameter that is easy to interpret.

When the outcome variable is discrete, we show that the identified sets for the conditional probabilities of the potential outcomes given the instruments can be characterized as the intersection of two convex sets, parameterized by the relaxation of instrument exclusion and exogeneity. Using this result, we then show that the identified set for a class of linear functionals of the densities of potential outcomes is the solution to a linear program that can be computed efficiently. This class of functionals includes the standard treatment parameters such as the ATE, the average effect of treatment on the treated (ATT), and quantile treatment effects (QTE).

We show that these identified sets exhibit many desirable properties, including continuity and monotonicity with respect to the sensitivity parameter. As is well known (e.g., BalkePearl1997), IV models have testable implications, which can fail in practice. In this case, we characterize the smallest deviations from the baseline model that prevent the model from being refuted.

We then extend our results to the case where outcomes are continuous. This case is more delicate, as the distribution of outcomes is now characterized by a density function, an infinite-dimensional object. We show that the identified set for densities and its functionals (like ATE, ATT, or QTE) can also be characterized via a linear program, albeit an infinite-dimensional one. As in the discrete case, we show that identified sets derived from this linear program have desirable properties by analyzing them as infinite-dimensional-valued correspondences, since each sensitivity parameter is now associated with a set of infinite-dimensional-valued density function. As these linear programs cannot be solved in practice, we propose a tractable approach for approximating the problem with a finite-dimensional one.

Using these computational results, we show how applied researchers can produce sensitivity plots that show the sensitivity (or robustness) of their parameter of interest to exclusion and exogeneity violations. These plots can be used, for example, to determine how strong exclusion or exogeneity violations can be before the data is consistent with a zero treatment effect.

To illustrate our approach, we revisit Gilchrist and Sands' GilchristSands2016 study of peer effects in movie viewership, using weather as an instrument for opening-weekend viewership. While extremely popular in empirical practice, weather instruments have come under increasing scrutiny in recent years (e.g., Sarsons2015, GallenRaymond2023, Mellon2025). In this application, social learning and dynamic behavior could lead to violations of instrument exclusion, and use our results to assess the robustness of conclusions to relaxations of this assumption. Using both discretized and continuous outcomes, we confirm that under the baseline of instrument exclusion, there is a positive peer effect on viewership, but show that this conclusion is sensitive to relatively small relaxations of the exogeneity assumption.

The rest of the paper is organized as follows. We first provide an overview of the related literature. Section (ref) then develops the framework for binary outcomes, introduces the relaxation class, and derives sharp identified sets, falsification frontiers, and falsification adaptive sets in the discrete setting. Section (ref) extends the analysis to continuous outcomes, establishes the corresponding identification and continuity results, and details a sieve-based computational strategy. Section (ref) presents our empirical application.

Related Literature

Research on the sensitivity of IV results to violations of exclusion and exogeneity go back to at least Fisher1961. More recent developments were proposed by BoundJaegerBaker1995, Small2007 and ConleyHansenRossi2012. All of these methods assume a linear outcome equation, motivated by treatment effect homogeneity, which we do not assume. These papers bound the direct effect of the instrument on the outcome, which can be done by bounding the coefficient on the instrument under the assumption that the potential outcomes depend linearly on it. Various approaches have been proposed to bound this direct effect, including NunnWantchekon2011, ConleyHansenRossi2012, Kraay2012, KippersluisRietveld2017, KippersluisRietveld2018, and MastenPoirier2021. Also see AltonjiElderTaber2005a, Ashley2009, and AshleyParmeter2015 for alternative approaches.

Our paper contributes to the literature on sensitivity analysis in instrumental variable models with heterogeneous treatment effects, which is much sparser than that for homogeneous treatment effects. Specifically, few papers consider continuous relaxations of the baseline instrumental variable assumptions while still allowing for heterogeneous treatment effects.

Early work by Manski1990 characterizes sharp bounds on average treatment effects under two sets of assumptions: (i) instrument exclusion and exogeneity hold (formulated as mean independence in his general analysis) or (ii) instrument exclusion and exogeneity fail arbitrarily. Our continuously parameterized sensitivity model spans these two sets of assumptions, allowing users to calibrate the degree of exclusion and exogeneity violations from “no-violations” (i.e., full exclusion and exogeneity) to “no assumptions”.

HotzMullinSanders1997 used a mixture model to allow for relaxations of the baseline assumptions. They focus on the average effect of treatment on the treated, whereas our sensitivity analysis allows for a broader set of parameters of interest. Ramsahai2012 studied a heterogeneous treatment effect model with a binary outcome, binary treatment, and a binary instrument. He defines a continuous relaxation of the instrument exogeneity assumption and then shows how to numerically compute identified sets for a single value of this relaxation. On pages 842-–843, he notes that “it is not obvious how the methods described in [his] paper can be extended to compute bounds” as a function of his relaxation. In our analysis, we allow all variables to be nonbinary, and even continuous for the outcome variable, and we allow for multiple instruments. We also consider a large set of target parameters and derive theoretical and computational properties for the sensitivity plots, which map the sensitivity parameters into this range of target parameters. Also see Huber2014 and MachadoShaikhVytlacil2019 for related, but different, approaches.

In fully discrete cases, identified sets for causal parameters and counterfactual distributions can often be obtained via linear programming. This observation goes back at least to BalkePearl1997 (and related work by Pearl1995) and is emphasized in more recent reviews of discrete partial identification methods. For example, see the literature review in Torgovitsky2019. Linear programming has been used in several papers to do sensitivity analysis. One paper is Ramsahai2012, which we already discussed above. Laffers2019 considers continuous relaxations of instrument exogeneity. He then computes identified sets for ATE for several values of this relaxation. In Laffers2018, he applies this approach to various additional forms of continuous relaxations. Duarte2024 also uses linear programming to bound parameters under exclusion and monotonicity violations. These papers all require all variables to be discrete. A key contribution of our paper is that our results allow for continuous outcome variables.

Our paper also contributes to assessments of IV model falsification. BalkePearl1997 characterize when Manski's bounds are empty, and hence when the model is falsified.\footnote{BalkePearl1997 assume the instrument is independent of the potential outcomes jointly, whereas Manski1990 only assumed the instrument is independent of each potential outcome separately. (Here we suppose outcomes are binary, so that mean independence is equivalent to statistical independence.) This difference does not affect whether the identified set is empty, given any fixed distribution of observables. Hence, it does not change the testable implications of the model. When the identified set is nonempty, however, this difference can affect its size. See the second paragraph of section 3 in SwansonHernanMillerRobinsRichardson2018 for further discussion.} Kitagawa2021 generalizes this characterization to allow for continuous outcomes, still requiring the treatment and instrument to be binary. As Kitagawa2021 notes, his extension is an adaptation of Corollary 2.2.1 in Manski's Manski2003 analysis of missing data. BeresteanuMolchanovMolinari2012 further generalizes this characterization to allow for continuous instruments and discrete treatment, for discrete or continuous outcomes. KedagniMourifie2020 provides an alternative characterization when instruments and outcomes are continuous, treatment is binary, and under the stronger assumption that the instrument is independent of the potential outcomes jointly; also see Proposition 2.5 of BeresteanuMolchanovMolinari2012 for a result under this stronger independence assumption.

Finally, a large literature on the testable implications of instrument exclusion and exogeneity, combined with other assumptions, has developed. Most notably, many papers have studied the testable implications of the monotonicity assumption of ImbensAngrist1994. FloresChen2018 give a comprehensive review. Also see FrandsenLefgrenLeslie2023Judging for discussions of monotonicity in the judge IV framework. In this paper, we focus on instrument exclusion and exogeneity only.

Sensitivity Analysis with Binary Outcomes

We begin by considering analyses with a binary outcome. For further simplicity, we also assume that the treatment and instrument are binary. The results below generalize to a setting with multiple treatment values and multiple discrete instruments, but we focus on the binary case, which allows us to explain the main ideas and results while keeping the notation simple. See Section (ref) for their generalization. The case where the outcome variable is continuously distributed presents additional technical challenges and is analyzed in Section (ref).

Model, Parameters of Interest, and Assumptions

Let $X \in \{ 0, 1 \}$ denote the observed binary treatment variable and $Z \in \{0,1 \}$ denote an observed instrument. As mentioned above, we consider multiple treatments and discrete instruments later. Let $\{Y(x,z)\}_{x,z \in \{0,1\}}$ denote potential outcomes for both treatment and instrument values. The observed outcome is denoted by

equation[equation omitted — 61 chars of source]

We assume the joint distribution of $(Y,X,Z)$ is known in this identification analysis. Our analysis could be done conditional on a vector $W$ of covariates, but we omit them for simplicity. Let $p_Z \coloneqq \ensuremath{\mathbb{P}}(Z=1)$ and $\pi(x \mid z) \coloneqq \ensuremath{\mathbb{P}}(X = x \mid Z = z)$. We maintain the following assumption to rule out trivial cases. \begin{het1Assump} Let $p_Z \in (0,1)$ and $\pi(x \mid z) \in (0,1)$ for all $x,z \in \{0,1\}$. \end{het1Assump}

We define $p_Y(x,z) \coloneqq \ensuremath{\mathbb{P}}(Y(x,z) = 1\mid Z = z)$, conditional probabilities of the potential outcomes given the instrument. Let $\textbf{p}_Y(x) \coloneqq (p_Y(x,0),p_Y(x,1)) \in [0,1]^2$ and $\textbf{p}_{Y} \coloneqq (\textbf{p}_Y(0),\textbf{p}_Y(1)) \in [0,1]^4$ be collections of these conditional probabilities. We are interested in functionals of these conditional probabilities, denoted by $\Gamma: [0,1]^4 \rightarrow \ensuremath{\mathbb{R}}$, which include various treatment effect parameters. In this section, we focus our attention on averages of treatment effects such as the average treatment effect (ATE) and the average treatment effect on the treated (ATT). They can be viewed as functionals of $\textbf{p}_Y$ as follows:

align*[align* omitted — 201 chars of source]

where

align[align omitted — 327 chars of source]

These parameters are well-defined even in the absence of exclusion or exogeneity assumptions about the instruments. Additional parameters could be of interest, such as the local average treatment effect (LATE). The LATE is defined in terms of potential treatments, which could be incorporated into our framework but are not required by it.

Before introducing additional assumptions, we first characterize the identified set for $\textbf{p}_{Y}$ when no assumptions are made about the joint distribution of $(Y,X,Z)$ beyond the regularity assumption (ref). To do so, define

align[align omitted — 245 chars of source]

which depends on the joint distribution of $(Y,X,Z)$. With this notation, we obtain the following result which is adapted from Manski1990.

proposition[Manski1990] Suppose Assumption (ref) holds. Then the identified set for $\textbf{p}_Y$ is $\mathcal{H}_0 \times \mathcal{H}_1$.

This result shows that the identified set for the conditional probabilities $\textbf{p}_Y$ is a Cartesian product of intervals, i.e., a hyperrectangle. As these bounds are sharp, they can be used to obtain sharp bounds on any functional of $\textbf{p}_Y$. For example, the functional $\Gamma_{\text{ATE}}$ is linear and the set $\mathcal{H}_0 \times \mathcal{H}_1$ is a Cartesian product of intervals, so appropriately evaluating $\Gamma_{\text{ATE}}$ at the lower/upper bounds of the intervals in (ref) will yield sharp bounds for it. The same approach can be used to obtain sharp bounds of the ATT, for example. For any linear functional, this is equivalent to a linear program, which is easy to solve analytically given the discrete supports of $X$ and $Z$. Figure (ref) illustrates this identified set and the optimization of the ATE over the identified set for $\textbf{p}_Y(x)$.

figure[figure omitted — 409 chars of source]

These bounds can be considerably tightened by assuming exogeneity or exclusion, as we formally define below.

Baseline Assumptions

We now introduce the assumptions we will study in this model. We compare these assumptions to the four assumptions usually imposed in a large segment of the literature, including the traditional Local Average Treatment Effect (LATE) framework: exogeneity, exclusion, monotonicity, and relevance. For brevity, we do not include covariates in this discussion, although all the upcoming assumptions can be stated conditional on a covariate vector $W$.

First, we formally define the exogeneity and exclusion assumptions we consider.

definition[Exogeneity] The instrument is exogenous if $Z \mathbin{ \mathpalette{\@indep}{} } Y(x,z)$ holds for each $(x,z) \in \operatorname*{supp}(X) \times \operatorname*{supp}(Z)$.

Exogeneity holds when the instrument is randomly assigned, or as good as randomly assigned, with respect to the potential outcomes. We do not require that the instrument be independent of potential treatment values, although this assumption can be incorporated into the framework. As mentioned earlier, we could consider relaxing the conditional exogeneity assumption $Y(x) \mathbin{ \mathpalette{\@indep}{} } Z \mid W$, at the cost of additional notation.

Next, we consider an exclusion assumption that is weaker than the most commonly used version.

definition[Weak Exclusion] The instrument is weakly excluded if $Y(x,z) \mid \{Z = z''\} \overset{d}{=} Y(x,z') \mid \{Z = z''\}$ for all $x \in \operatorname*{supp}(X)$ and $z,z',z'' \in \operatorname*{supp}(Z)$.

The standard exclusion assumption is that $Y(x,z) = Y(x,z')$ with probability 1 for any possible treatment value $x$ and any possible instrument values $z$ and $z'$, whereas weak exclusion only requires the (conditional) distributions of these potential outcomes to be identical. This has also been called stochastic exclusion; see, for example, SwansonHernanMillerRobinsRichardson2018. Although we do not study the LATE here, the arguments used to obtain a causal interpretation for the Wald estimand are not impacted if exclusion is replaced by weak exclusion. In particular, the Wald estimand equals the LATE when the treatment and instrument are binary under weak exclusion, provided appropriate exogeneity, relevance, and monotonicity conditions hold.

We will assume that the instrument is exogenous or weakly excluded, without requiring it to satisfy both. \begin{het1Assump} The instrument $Z$ is exogenous or weakly excluded. \end{het1Assump} Under this assumption, $p_Y(x,z)$ can be interpreted in one of two ways. Under exogeneity, this probability equals the unconditional probability $\mathbb{P}(Y(x,z) = 1)$, while under weak exclusion it denotes the conditional probability $\mathbb{P}(Y(x) = 1 \mid Z = z)$. If both hold, then $p_Y(x,z)$ does not depend on $z$, meaning that $\mathbb{P}(Y(x,1) = 1 \mid Z = 1) = \mathbb{P}(Y(x,0) = 1 \mid Z = 0)$, but Assumption (ref) allows the dependence of $p_Y(x,z)$ on $z$ to be nontrivial. We formally show that under Assumption (ref), $p_Y(x,z)$ not depending on $z$ implies the exogeneity and weak exclusion of the instrument.

lemma[Condition for Exogeneity and Weak Exclusion] Suppose assumptions (ref) and (ref) hold. Then $p_Y(x,1) = p_Y(x,0)$ for all $x \in \{0,1\}$ if and only if $Z$ is exogenous and weakly excluded.

Thus, we can view failures of exogeneity or exclusion as mathematically equivalent to the probabilities $p_Y(x,z)$ being nonconstant in $z$.

To simplify our exposition going forward, we let \[ Y(x) \coloneqq Y(x,Z). \] Note that $p_Y(x,z) = \mathbb{P}(Y(x) = 1\mid Z=z)$, and that the ATE and ATT functionals $\Gamma_{\text{ATE}}$ and $\Gamma_{\text{ATT}}$ are defined as functionals of $Y(x)$, independently of whether weak exclusion or exogeneity holds. Also, the bounds of Proposition (ref) do not change when Assumption (ref) is imposed. With these definitions, we can see that the instrument is exogenous and weakly excluded if and only if \[ Y(x) \mathbin{ \mathpalette{\@indep}{} } Z \] for $x \in\{0,1\}$. Hence, we will consider relaxations of exogeneity or weak exclusion as relaxations of an independence assumption, as they are mathematically equivalent here.

To finish our comparison to the standard IV assumption, we note that we allow for positive masses of both compliers, i.e., units for whom $X(1) > X(0)$, and defiers, i.e., units for whom $X(0)>X(1)$. Again, this assumption could be added to our framework at the cost of additional notation, but we focus on the case where no restrictions are imposed on these potential treatments. We also do not require that $\ensuremath{\mathbb{P}}(X=1 \mid Z=1) \neq \ensuremath{\mathbb{P}}(X=1\mid Z=0)$, the usual relevance assumption assumed in the LATE framework.

We will maintain assumptions (ref) and (ref), and consider a continuum of exogeneity and exclusion assumptions that range from no assumptions to full exogeneity and exclusion. Thus, they will include the no-assumptions case and the case where $Z$ is excluded and exogenous as special cases.

Sensitivity Models for the Exogeneity or Exclusion Assumptions

We now consider a menu of assumptions that can be interpreted as relaxations of the exogeneity or exclusion assumption. The results of Proposition (ref) show one extreme: bounds under no dependence assumptions. We briefly consider bounds under the other extreme, where exogeneity and weak exclusion exactly hold. Manski1990 derived the identified set for $\mathbb{P}(Y(x) = 1)$ for $x \in \{0,1\}$ as well as the identified set for the ATE under this assumption.\footnote{Manski's Manski1990 analysis considered a general case which does not require outcomes, treatment, or instruments to be binary. In this general setting, he used a mean independence assumption. When outcomes are binary, mean independence of $Y(x)$ from $Z$ is equivalent to statistical independence of $Y(x)$ and $Z$.}

Under this assumption, $p_Y(x,0) = p_Y(x,1)$ for $x \in \{0,1\}$ by Lemma (ref). This restricts the probabilities $\textbf{p}_Y$ to lie in the set

align*[align* omitted — 132 chars of source]

Therefore, the identified set for $\textbf{p}_Y$ is given by the set of probabilities as restricted by the observed distribution of $(Y,X,Z)$, namely $\mathcal{H}_0 \times \mathcal{H}_1$, intersected with the set of probabilities satisfying exclusion and exogeneity, given by $\mathcal{A}_\text{indep}$. Thus, the identified set for $\textbf{p}_Y$ is

align[align omitted — 103 chars of source]

which can also be written as

align*[align* omitted — 274 chars of source]

These bounds take the form of intersections. Pearl1995 and BalkePearl1997 showed that this identified set can be empty, and hence that this model is falsifiable. This set is empty if and only if the two sets in (ref) are disjoint. An empty identified set corresponds to a falsification of the original model, or of an exogeneity or exclusion assumption, when other model assumptions are maintained. Figure (ref) shows the identified set both when the model is falsified (left panel), and when it is not (right panel).

figure[figure omitted — 431 chars of source]

Full exogeneity or exclusion of the instrument may be a strong assumption in contexts where we do not believe that $Z$ is assigned randomly, or if we cannot rule out a direct effect of the instrument on the potential outcomes. In these cases, relaxing exogeneity or exclusion is appropriate. The no-assumption bounds of Manski remain valid, but partial validity of the instrument will yield intermediate bounds that are potentially significantly narrower than those in Proposition (ref).

We will consider relaxations of exogeneity and exclusion by characterizing sets of conditional probabilities $\textbf{p}_{Y}$. A large literature on sensitivity analysis has proposed various approaches for relaxing assumptions, often independence or conditional independence assumptions. We will focus on three examples, which are special cases of a unifying class of relaxations from independence we define in Section (ref).

Marginal Sensitivity Model: Tan (2006)

The Marginal Sensitivity Model (MSM) of Tan2006 consists of a class of relaxations of an independence assumption between a potential outcome and a binary treatment. It is generalized to multivariate treatments in ZhaoSmallBhattacharya2019 and BasitLatifWahed2023. We consider a version of the MSM that constrains the dependence of the potential outcomes on the instruments, rather than the treatment.

definitionLet $\Lambda \in [1,+\infty]$ be a known sensitivity parameter. The distribution of $(\{Y(x)\}_{x \in \operatorname*{supp}(X)},Z)$ satisfies the Marginal Sensitivity Model with parameter $\Lambda$ if \begin{align} \frac{\ensuremath{\mathbb{P}}(Z=z)}{\ensuremath{\mathbb{P}}(Z=z')} \Big/ \frac{\ensuremath{\mathbb{P}}(Z=z \mid Y(x) = y)}{\ensuremath{\mathbb{P}}(Z=z' \mid Y(x) = y)} \in \left[\Lambda^{-1},\Lambda\right] \end{align} for all $x \in \operatorname*{supp}(X)$, $y \in \operatorname*{supp}(Y(x))$, and $z, z' \in \operatorname*{supp}(Z)$.\footnote{We let $\Lambda^{-1} = 0$ when $\Lambda = +\infty$.}

With a binary instrument, this restriction places a bound on the odds ratio between the conditional odds of the instrument $\ensuremath{\mathbb{P}}(Z=z \mid Y(x) = y)/\ensuremath{\mathbb{P}}(Z=1-z \mid Y(x) = y)$ and its unconditional counterpart, $\ensuremath{\mathbb{P}}(Z=z)/\ensuremath{\mathbb{P}}(Z=1-z)$, for $z = 0,1$. In the binary outcome and instrument setting, equation (ref) can be rearranged as

align[align omitted — 218 chars of source]

for $x,y \in \{0,1\}$. When $\Lambda = 1$, this ratio is 1 and $p_Y(x,1) = p_Y(x,0)$ for $x = 0,1$. By Lemma (ref), this means that weak exclusion and exogeneity hold when $\Lambda = 1$ under Assumption (ref). When $\Lambda = +\infty$, these inequalities do not impose any restrictions on $\textbf{p}_Y$. Intermediate values of $\Lambda$ yield intermediate levels of restrictions on $\textbf{p}_Y$. Note that we can choose different $\Lambda$ values for $x = 0,1$, but we omit this generalization for brevity.

We note that equation (ref) can be written as four linear constraints on $\textbf{p}_Y$ by varying $y$ and $x$ over their support. This will be useful for casting this sensitivity analysis exercise as a linear program, as linear programming is a reliably fast and scalable computation method whose implementation is standard. Define

align*[align* omitted — 249 chars of source]

where $\lambda \coloneqq 1- \Lambda^{-1}$. The set of conditional probabilities satisfying the marginal sensitivity model with sensitivity parameter $\lambda$ is

align[align omitted — 183 chars of source]

where the weak inequality in (ref) is component-wise. We reparametrized the sensitivity parameter $\Lambda$ as $\lambda = 1 - \Lambda^{-1}$ to standardize its scale to $[0,1]$. Here $\Lambda = 1$, or full exogeneity and exclusion, maps into $\lambda = 0$ while $\Lambda = +\infty$, or no assumptions, maps into $\lambda = 1$.

$c$-dependence

Introduced in MastenPoirier2018, $c$-dependence imposes a bound on the maximum difference between the conditional probability of receiving a binary treatment $\ensuremath{\mathbb{P}}(X=1\mid Y(x))$ and its unconditional probability $\ensuremath{\mathbb{P}}(X=1)$. This was proposed in a setting where the unconfoundedness of treatment $X$ is relaxed. We adapt this sensitivity model to the case where exogeneity or exclusion of an instrument is relaxed. Here is the formal definition of this sensitivity model.

definitionLet $c\in[0,1]$ be a known sensitivity parameter. The distribution of $(\{Y(x)\}_{x \in \operatorname*{supp}(X)},Z)$ satisfies $c$-dependence if \begin{equation} |\mathbb{P}(Z=z \mid Y(x) = y) - \ensuremath{\mathbb{P}}(Z=z) | \leq c \end{equation} for all $y \in \operatorname*{supp}(Y(x))$, $x \in \operatorname*{supp}(X)$, and $z \in \operatorname*{supp}(Z)$.

When $Z$ is binary, it suffices to impose this inequality for $z=1$ only. When $c=0$, $c$-dependence is equivalent to imposing full exclusion and exogeneity. Values of $c$ exceeding $\max\{ p_Z,1-p_Z \}$ do not constrain the stochastic relationship between $Z$ and $Y(x)$. $c \in (0,1)$ partially constrains the stochastic relationship between $Z$ and $Y(x)$. MastenPoirier2023 give additional discussion of how to interpret $c$-dependence.

We can again rewrite the above restriction into a system of four linear restrictions on $\textbf{p}_Y$. Let

align*[align* omitted — 268 chars of source]

where $k_z(c) \coloneqq \frac{\ensuremath{\mathbb{P}}(Z=z)\max\{\ensuremath{\mathbb{P}}(Z=1-z) - c,0\}}{\ensuremath{\mathbb{P}}(Z = 1-z)\min\{\ensuremath{\mathbb{P}}(Z=z)+c,1\}}$ for $z = 0,1$ and $c \in [0,1]$. We can show that the set of conditional probabilities consistent with $c$-dependence with sensitivity parameter $c$ is

align[align omitted — 189 chars of source]

This set depends only on $c$ and $p_Z$.

Kolmogorov-Smirnov Distance

Consider a sensitivity model bounding a metric between the distributions of $Y(x) \mid \{Z=0\}$ and $Y(x) \mid \{Z=1\}$. This type of restriction was used in KlineSantos2013 to relax a missingness at random assumption. It was also considered for estimation in Manski1983.

definitionLet $K \in [0,1]$ be a known sensitivity parameter. The distribution of $(\{Y(x)\}_{x \in \operatorname*{supp}(X)},Z)$ satisfies the Kolmogorov-Smirnov (KS) model if \begin{align} |\ensuremath{\mathbb{P}}(Y(x) \leq y\mid Z=z) - \ensuremath{\mathbb{P}}(Y(x) \leq y\mid Z=z')| \leq K \end{align} for all $x \in \operatorname*{supp}(X)$, $y \in \ensuremath{\mathbb{R}}$, and $z,z' \in \operatorname*{supp}(Z)$.

When outcomes and instruments are binary, this sensitivity model is equivalent to bounding the magnitude of the difference between $\mathbb{P}(Y(x) = 1 \mid Z = 1)$ and $\mathbb{P}(Y(x) = 1 \mid Z = 0)$ by $K$. This assumption directly bounds the maximum deviation between the potential outcomes distribution given the instrument's two values. As in the previous two definitions, this class of restrictions encompasses independence ($K = 0$), no assumptions ($K = 1$), and intermediate cases ($K \in (0,1)$).

The set of conditional probabilities satisfying the Kolmogorov-Smirnov restrictions is characterized by the two linear inequalities

align[align omitted — 159 chars of source]

where

align[align omitted — 199 chars of source]

A Unifying Sensitivity Model

We now consider a general sensitivity model that encompasses the previous three sensitivity models as special cases. We will derive our main theoretical results under this sensitivity model. We assume that $X$ and $Z$ are binary for ease of notation and discuss the generalization to discrete $X$ and $Z$ in Section (ref). In what follows, $\theta \in [0,1]$ is a sensitivity parameter that indexes relaxations of exogeneity or weak exclusion of the instrument.

\begin{het1Assump}[General Sensitivity Model] For a known sensitivity parameter $\theta \in [0,1]$, let

align*[align* omitted — 82 chars of source]

where, for $x \in \{0,1\}$, $\mathcal{A}_x$ satisfies

enumerate• (Spanning) $\mathcal{A}_x(0) = \{a \in [0,1]^2: a_0 = a_1\}$ and $\mathcal{A}_x(1) = [0,1]^2$; • (Monotonicity) $\mathcal{A}_x(\theta) \subseteq \mathcal{A}_x(\theta')$ when $\theta \leq \theta'$; • (Linearity of Constraints) $\mathcal{A}_x(\theta)$ is a closed convex polytope for each $\theta \in [0,1]$; • (Continuity) The correspondence $\mathcal{A}_x: [0,1] \rightrightarrows [0,1]^2$ is continuous.

\end{het1Assump}

The first part of this assumption implies that setting $\theta = 0$ imposes exogeneity and weak exclusion of the instrument, while setting $\theta = 1$ implies no restrictions on the dependence between $Z$ and the potential outcomes. The second part assumes these restrictions are monotonic in $\theta$, meaning that increasing $\theta$ yields a (weakly) larger set of conditional probabilities. These two parts combined yield that $\{\mathcal{A}_x(\theta):\theta \in [0,1]\}$ monotonically connects no assumptions to exogeneity and weak exclusion. The third restriction says that these sets are characterized by finitely many weak linear inequalities. This is crucial in obtaining a linear programming formulation for the bounds of various causal objects, such as the ATE. The last part of this assumption assumes the continuity of the correspondence between the sensitivity parameter $\theta$ and the set of restricted conditional probabilities. Recall that a correspondence is continuous if it is both upper and lower hemicontinuous (uhc and lhc) at all points of its domain. See Border1985 for a compendium of results related to continuity of correspondences we make use of in our proofs. This assumption will yield continuity in the sensitivity parameter of the causal bounds obtained from linear programming.

This high-level assumption has useful properties, and all three previously considered relaxations are special cases of it. This is formalized in this proposition.

propositionSuppose Assumption (ref) holds. Relabeling $\lambda$, $c$, and $K$ as $\theta \in [0,1]$, the sets $\mathcal{A}_\text{MSM}(\lambda)$, $\mathcal{A}_{c\text{-dep}}(c)$, and $\mathcal{A}_\text{KS}(K)$ satisfy Assumption (ref).

Under this general relaxation, we will derive identified sets for various parameters of interest. We use these identified sets to characterize sharp bounds on causal objects using linear programming. We can also use them to determine what values of $\theta$ correspond to falsified models.

Before continuing our discussion, we present the identified set for conditional outcome probabilities under this general restriction.

theoremSuppose assumptions (ref), (ref), and (ref) hold. Then: \begin{enumerate} • The identified set for $\textbf{p}_Y$ is \begin{align} \Pi(\theta) \coloneqq \Pi_0(\theta) \times \Pi_1(\theta), \end{align} where $\Pi_x(\theta) \coloneqq \mathcal{H}_x \cap \mathcal{A}_x(\theta)$; • There exists $\underline{\theta}\in [0,1]$ such that $\Pi(\theta)$ is non-empty for $\theta \geq \underline{\theta}$ and empty for $\theta < \underline{\theta}$; • For all $\theta \in [\underline{\theta},1]$, $\Pi(\theta)$ is a closed convex polytope in $[0,1]^4$; • For $x \in \{0,1\}$, suppose $\text{int}(\mathcal{H}_x \cap \mathcal{A}_x(\theta)) \neq \emptyset$ for all $\theta > \underline{\theta}$. The correspondence $\Pi:[\underline{\theta},1] \rightrightarrows [0,1]^4$ defined in equation (ref) is continuous. \end{enumerate}

This theorem has several implications. First, the identified set for the set of probabilities $\textbf{p}_Y$ is a Cartesian product of two sets. Each of these two sets is characterized as the intersection between a set containing all vectors $\textbf{p}_Y(x)$ consistent with the distribution of observables $F_{Y,X,Z}$, and the set of vectors consistent with a sensitivity model indexed by $\theta$.

The second implication is that the sensitivity model is falsified for an open, but potentially empty, subset of $[0,1]$. The minimum value at which the model is not falsified, called the falsification point by MastenPoirier2021, is $\underline{\theta}$ and is identified since it is a property of the sets $\Pi(\theta)$ for $\theta \in [0,1]$, all of which are known from the distribution of $(Y,X,Z)$. Moreover, the set of values for which the identified set is non-empty is closed, and always contains $\theta = 1$.

Third, this set is a closed, convex polytope, meaning it is defined by finitely many linear inequalities. This ensures that optimizing linear functions, such as $\Gamma_\text{ATE}$ and $\Gamma_\text{ATT}$ from (ref) and (ref), can be performed using linear programming. This will be the key computational tool for implementing these methods.

Fourth, and finally, the mapping from $\theta$ into the identified set is continuous as a correspondence. This allows us to show the continuity in $\theta$ of extrema of continuous functionals of $\textbf{p}_Y$ over the identified set, again including the ATE and ATT.

We now illustrate the identified set for a sensitivity model corresponding to $c$-dependence. The shaded boxes in Figure (ref) show examples of the no-assumption bounds $\mathcal{H}_0 \times \mathcal{H}_1$.

figure[figure omitted — 712 chars of source]

The set $\mathcal{A}_x(\theta)$ is a parallelogram imposing the $c$-dependence constraint. The identified set for $\textbf{p}_Y(x)$ is given by the intersection of the parallelogram and shaded box.

While the no-assumption bounds are never empty, the bounds under exogeneity and weak exclusion ($\theta=0$) can be empty, and hence the baseline statistical independence assumption can be falsified. This happens when, for some $x \in \{0,1\}$, the no assumption bounds $\mathcal{H}_x$ have an empty intersection with the statistical independence constraint set $\mathcal{A}_x(\theta)$. Graphically, this happens when the box defined by the no assumption bounds does not intersect the 45-degree line. This is shown in the first plot Figure (ref). The falsification point is simply the smallest value of $\theta$ such that the parallelogram defined by $\mathcal{A}_x(\theta)$ has a nonempty intersection with the no assumption bounds $\mathcal{H}_x$ for each $x \in \{0,1\}$. This intersection is illustrated in the second plot of Figure (ref). Increasing the sensitivity parameter increases the size of this intersection (see the third plot of Figure (ref)), until the intersection equals $\mathcal{H}_x$, the no assumption bounds, which can be seen in the fourth plot of Figure (ref).

We next show how to use Theorem (ref) to get identified sets for counterfactual probabilities $\ensuremath{\mathbb{P}}(Y(x)=1)$ and for the ATE. By the law of total probability, \[ \ensuremath{\mathbb{P}}(Y(x)=1) = p_Y(x,0) (1-p_Z) + p_Y(x,1) p_Z. \] The weight $p_Z$ is identified, while the identified set for $\textbf{p}_Y(x)$ is given by $\Pi_x(\theta)$. Thus, we can simply minimize and maximize the above convex combination over this set to obtain the identified set for $\ensuremath{\mathbb{P}}(Y(x)=1)$. Hence we define

align*[align* omitted — 251 chars of source]

These are both finite-dimensional linear programs and hence can be computed easily given estimates of the joint distribution of $(Y,X,Z)$. Figure (ref) illustrates the minimization/maximization of a linear functional over the identified set $\mathcal{A}_x(\theta) \cap \mathcal{H}_x$.

figure[figure omitted — 232 chars of source]

The following corollary lists properties of these bounds.

corollarySuppose assumptions (ref), (ref), and (ref) hold. Then, for $x \in \{0,1\}$: \begin{enumerate} • The identified set for $(\ensuremath{\mathbb{P}}(Y(0)=1), \ensuremath{\mathbb{P}}(Y(1)=1))$ is $I_0(\theta) \times I_1(\theta)$ where $I_x \coloneqq [\underline{P}_x(\theta), \overline{P}_x(\theta)]$ when $\theta\in [\underline{\theta},1]$, and the empty set when $\theta < \underline{\theta}$; • The functions $\underline{P}_x(\theta)$ and $\overline{P}_x(\theta)$ are continuous and monotonic over $\theta \in [\underline{\theta},1]$; • Let $\theta \in [\underline{\theta},1]$. The identified set for ATE is $[\underline{\text{ATE}}(\theta),\overline{\text{ATE}}(\theta)]$ where \[ \underline{\text{ATE}}(\theta) \coloneqq \underline{P}_1(\theta) - \overline{P}_0(\theta) \qquad \text{and} \qquad \overline{\text{ATE}}(\theta) \coloneqq \overline{P}_1(\theta) - \underline{P}_0(\theta). \] \end{enumerate}

This discussion implies that ATE will typically be partially identified at the falsification point. That is: The falsification adaptive set for ATE, $[\underline{\text{ATE}}(\underline\theta), \overline{\text{ATE}}(\underline\theta)]$, will generally be an interval with a nonempty interior.

Generalization to Non-Binary Discrete Variables

The previous results illustrate that common sensitivity models yield identified sets for parameters of interest with desirable properties. However, these were illustrated only for cases where $Y$, $X$, and $Z$ were all binary. In practice, many empirical settings have multiple instruments, and treatments or outcomes may be multivalued as well. In this section, we sketch a generalization of the previous results to cases where the support of $Y$, $X$, may be discrete instead of binary, where there may be multiple instruments, and where each instrument may possess finite support rather than being binary.

Let $X$ be a discrete treatment, let $Z$ be a vector of instruments, where each instrument is discrete, and let $Y(x,z)$ be discrete as well. We let $Y = Y(X,Z)$ be the realized, observed outcome. We suppose all their supports are finite.

Let $p_Y(y \mid z ; \ x) \coloneqq \ensuremath{\mathbb{P}}(Y(x) = y\mid Z = z)$, $\textbf{p}_{Y}(x,z) \coloneqq \{p_Y(y \mid z ; \ x)\}_{y \in \operatorname*{supp}(Y(x))}$, $\textbf{p}_{Y}(x) \coloneqq \{\textbf{p}_{Y}(x,z)\}_{z \in \operatorname*{supp}(Z)}$, and $\textbf{p}_{Y} \coloneqq \{\textbf{p}_{Y}(x)\}_{x \in \operatorname*{supp}(X)}$. The vector $\textbf{p}_{Y}$ contains the full distribution of $Y(x) \mid \{Z = z\}$ for all $(x,z) \in \operatorname*{supp}(X,Z)$. We define $s_Z \coloneqq |\operatorname*{supp}(Z)|$, $s_{Y(x)} \coloneqq |\operatorname*{supp}(Y(x))|$, and $\operatorname*{supp}(Y(x)) \coloneqq \{y_1,\ldots,y_{s_{Y(x)}}\}$.

To avoid trivial cases, we make the following assumption. \begin{het1Assump} For all $(x, z) \in \operatorname*{supp}(X) \times\operatorname*{supp}(Z)$, $\mathbb{P}(Z = z) \in (0,1)$ and $\ensuremath{\mathbb{P}}(X=x \mid Z =z) \in (0,1)$. \end{het1Assump}

Let $\Delta_{S}$ denote the simplex of dimension $S$:

align*[align* omitted — 144 chars of source]

For $K \in \mathbb{N}$, let $\Delta_S^K$ denote the $K$-fold cartesian product of $\Delta_S$.

The no-assumption identified set for $\textbf{p}_Y$ is given by

align*[align* omitted — 110 chars of source]

where

align*[align* omitted — 297 chars of source]

The set $\mathcal{H}_{x,z}$ is the identified set for $\textbf{p}_Y(x,z)$ under no assumptions. We also note the similar structure of $\mathcal{H}_x = \prod_{z \in \operatorname*{supp}(Z)} \mathcal{H}_{x,z}$ and of the rectangles defined in (ref) for the binary case.

The three sensitivity models we investigated earlier can be defined independently of the supports of the potential outcomes, treatments, or instruments, so they can be used when these variables are non-binary. We can also embed these sensitivity models in a general sensitivity model similar to the one in Assumption (ref). The following assumption simplifies to (ref) when all variables are binary.

\begin{het1Assump}[General Sensitivity Model] Suppose Assumption (ref) holds. For a known sensitivity parameter $\theta \in [0,1]$, let

align*[align* omitted — 92 chars of source]

where, for $x \in \operatorname*{supp}(X)$, $\mathcal{A}(\theta; x)$ satisfies

enumerate• (Spanning) $\mathcal{A}(0; x) = \{ (a_1,\ldots,a_{s_Z}) \in \Delta_{s_Y(x) - 1}^{s_Z} : a_1 = \cdots = a_{s_Z}\}$ and $\mathcal{A}(1; x) = \Delta_{s_Y(x) - 1}^{s_Z}$; • (Monotonicity) $\mathcal{A}(\theta; x) \subseteq \mathcal{A}(\theta'; x)$ when $\theta \leq \theta'$; • (Linearity of Constraints) $\mathcal{A}(\theta; x)$ is a closed convex polytope for each $\theta \in [0,1]$; • (Continuity) The correspondence $\mathcal{A}(\cdot; x): [0,1] \rightrightarrows \Delta_{s_Y(x) - 1}^{s_Z}$ is continuous.

\end{het1Assump} This assumption is similar to its counterpart with binary variables, except for parts 1 and 4, which have been modified to allow $Y(x)$ to be nonbinary. The restriction in part 1 states that $\mathbb{P}(Y(x) = y \mid Z = z)$ is constant in $z \in \operatorname*{supp}(Z)$ for each $y \in \operatorname*{supp}(Y(x))$, and is stated as equality constraints for on the components of $\mathcal{A}_x(0)$.

As in the binary case, all these assumptions can be written as linear inequalities in the components of vector $\textbf{p}_Y$. Therefore, the bounds on various causal objects can be obtained by solving linear programs. We expect similar results to Theorem (ref) and Corollary (ref) to hold in this setting, so that the bounds enjoy the same monotonicity and continuity property.

Identification with Continuous Outcomes

We now consider cases where the outcome variable is continuously distributed. In this case, we may view the corresponding problem as an infinite dimensional program, whose theoretical properties are harder to analyze. Nevertheless, in this section we show that the previous sensitivity models can be used with continuous outcomes, and we obtain theoretical properties of the corresponding sensitivity analyses for the exogeneity/exclusion of an instrument. To keep other aspects of the problem relatively simple, we consider the case where the treatment and instrument are both binary, although this can be naturally generalized as in Section (ref). In this section, we show that the analytical results we derived under binary outcomes generalize to continuous outcomes. This leads us to a relatively simple and feasible approach for computing identified sets under relaxations of instrument exogeneity with continuous outcomes.

We begin by assuming that outcomes are continuously distributed.

\begin{het1Assump} Suppose that $\operatorname*{supp}(Z) = \operatorname*{supp}(X) = \{0,1\}$. For any $x, x',z \in \{0,1\}$ the distribution of $Y(x) \mid \{X=x', Z=z\}$ is continuous with respect to the Lebesgue measure and is supported on a compact interval $\mathcal{Y}_x \coloneqq \operatorname*{supp}(Y(x))$, which is independent of $x'$ and $z$. \end{het1Assump}

Assumption (ref) supposes that, conditional on the treatment and instruments, potential outcomes are continuously distributed. It implies that, conditional on the treatment and instruments, observed outcomes are also continuously distributed. We can allow for discrete instruments as in Section (ref), but we only consider a binary instrument to simplify the notation.

This assumption also states that the conditional support of $Y(x)$ given $(X,Z) = (x',z)$ does not depend on $(x',z)$, which is made for convenience. Our results would remain valid without this restriction, but notation in the proofs would have to be heavier.

Let $f_{Y}(y \mid z; \ x) \coloneqq f_{Y(x) \mid Z}(y \mid z)$ denote the conditional density of $Y(x)$ given $Z = z$. We also let $\textbf{f}_{Y}(y ; \ x) \coloneqq (f_{Y}(y \mid 0 ; \ x),f_{Y}(y \mid 1 ; \ x))$ and $\textbf{f}_{Y}(y) \coloneqq (\textbf{f}_{Y}(y ; \ 0), \textbf{f}_{Y}(y ; \ 1))$ denote collections of these densities across instrument and treatment values. We assume that the potential outcomes' densities belong to a convex class of densities that is compact with respect to the supremum norm. \begin{het1Assump} For $x,z\in\{0,1\}$, let

align*[align* omitted — 179 chars of source]

where $\mathcal{F}_x(\mathcal{Y}_x)$ is a convex set of bounded functions supported on $\mathcal{Y}_x$ that is compact with respect to the norm $\|f\|_\infty \coloneqq \sup_{y \in \mathcal{Y}_x}|f(y)|$. \end{het1Assump}

Examples of compact sets $\mathcal{F}_x(\mathcal{Y}_x)$ include the set of bounded Lipschitz functions:

align*[align* omitted — 240 chars of source]

where $\mathcal{C}_0(A)$ denotes the set of continuous functions on domain $A$, and $M<\infty$ is a constant. See FreybergerMasten2019 for alternative compact sets of functions and associated discussion.

We start by deriving the no-assumptions bounds for this set of conditional densities.

propositionSuppose assumptions (ref), (ref), and (ref) hold. The identified set for $\textbf{f}_{Y}$ is \begin{align*} \mathcal{H} \coloneqq \prod_{x = 0,1} \mathcal{H}_{x} \end{align*} where $\mathcal{H}_{x} \coloneqq \prod_{z = 0,1} \mathcal{H}_{x,z}$ and $\mathcal{H}_{x,z} \coloneqq \{f(\cdot) \in \mathcal{F}_{\text{den},x}: f \geq f_{Y|X,Z}(\cdot\mid x,z)\pi(x\mid z)\}$.

We next consider the baseline case where the instruments are exogenous and excluded. In this case, the instrument's validity implies that the densities $\textbf{f}_Y$ must lie in

align*[align* omitted — 192 chars of source]

since this set imposes that $f_{Y(x)|Z}(\cdot\mid 0) = f_{Y(x)|Z}(\cdot\mid 1)$ for $x = 0,1$. Thus, the identified set for $\textbf{f}_Y$ under independence is given by

align*[align* omitted — 57 chars of source]

This is precisely the setting studied in Kitagawa2021, and he provides a characterization of this set in his Proposition 3.1, which we include without proof.

proposition(Proposition 3.1 in Kitagawa2021) Suppose assumptions (ref), (ref), and (ref) hold. Suppose $Z$ is exogenous and weakly excluded. Then the identified set for $(f_{Y(0)},f_{Y(1)})$ is \[ \left\{ f_0 \in \mathcal{F}_{\text{den},0}: f_0(\cdot) \geq \max_{z=0,1} \; f_{Y|X,Z}(\cdot\mid 0 , z) \pi(0\mid z)\right\} \times \left\{ f_1 \in \mathcal{F}_{\text{den},1}: f_1(\cdot) \geq \max_{z=0,1} \; f_{Y|X,Z}(\cdot \mid 1, z)\pi(1\mid z)\right\}. \] Consequently, the model is refuted if \[ \int_{\mathcal{Y}_x} \max_{z=0,1} \; f_{Y,X|Z}(y,x \mid z) \; dy > 1 \] for some $x \in \{0,1 \}$.

The previous two results establish the identification region for conditional densities of $Y(x)$ given $Z$ under no-assumptions, and under the full validity of the instrument, which correspond to the ends of a spectrum of assumptions about the dependence between $Y(x)$ and $Z$. We now consider sensitivity models that consider intermediate assumptions on the instrument's validity. We again consider the following three restrictions, which are adapted from Section (ref).

Marginal Sensitivity Model

Consider the Marginal Sensitivity Model of definition (ref). When the outcome is continuously distributed, Bayes' rule allows us to rewrite equation (ref) as a density ratio:

align*[align* omitted — 110 chars of source]

for $(x,z,z') \in \operatorname*{supp}(X) \times \operatorname*{supp}(Z)^2$. As in the previous sections, we reparametrize $\Lambda$ as $\lambda = 1 - \Lambda^{-1} \in [0,1]$. The set of densities satisfying this restriction can be viewed as a set of functions satisfying linear inequality constraints. Specifically, we can write the set of restricted densities as

align[align omitted — 142 chars of source]

where $\mathcal{A}_\text{MSM}(\lambda;x) \coloneqq \{\textbf{f} \in \mathcal{F}_{\text{den},x}^2: A_\text{MSM}(\lambda)\textbf{f} \leq \textbf{0}\}$ and

align*[align* omitted — 114 chars of source]

Inequalities involving functions $\textbf{f}$ are meant to hold across all $y \in \ensuremath{\mathbb{R}}$.

$c$-dependence

As defined in equation (ref), $c$-dependence is collection of inequalities across values of $y$. Again using Bayes' Rule, we can rewrite these inequalities using conditional densities of $Y(x)$ given the instrument:

align*[align* omitted — 209 chars of source]

These are densities restricted by linear inequalities that depend on the observed variables only through the marginal distribution of the instrument. The set of densities as restricted by $c$-dependence is given by

align[align omitted — 145 chars of source]

where $\mathcal{A}_\text{$c$-dep}(c;x) \coloneqq \{\textbf{f} \in \mathcal{F}_{\text{den},x}^2: A_\text{$c$-dep}(c)\textbf{f} \leq \textbf{0}\}$ and

align[align omitted — 134 chars of source]

We can see that setting $c = 0$ implies that $k_0(c) = k_1(c) = 1$, which mechanically imposes that the conditional densities $f_{Y(x)|Z}(\cdot\mid 0)$ and $f_{Y(x)|Z}(\cdot\mid 1)$ are equal. As a result, we can verify that $c=0$ implies independence of potential outcomes and the instrument, as it does when the outcome is discrete.

Supremum Distance

Using the Kolmogorov-Smirnov as a starting point, we consider a sensitivity model that bounds the supremum distance between densities rather than distribution functions. Hence, we assume that

align*[align* omitted — 99 chars of source]

for $x \in \operatorname*{supp}(X)$, for some known $K$ satisfying $K \in [0,1]$.\footnote{We let $K/(1-K) = +\infty$ when $K = 1$.} The sensitivity parameter $K$ bounds the difference between density functions, and we used the strictly increasing mapping $a\mapsto 1/(1-a)$ to span the continuum between independence and no restrictions, as $K=0$ maps to exact equality of densities, and $K=1$ does not impose any restrictions on the dependence of the distribution of $Y(x)$ in $Z$. An alternate mapping from $[0,1]$ to $[0,+\infty]$ could be used instead.

The set of densities as restricted by this sup distance is given by

align[align omitted — 128 chars of source]

where $\mathcal{A}_\text{KS}(K;x) \coloneqq \{\textbf{f} \in \mathcal{F}_{\text{den},x}^2: A_\text{KS}\textbf{f} \leq (K/(1-K), K/(1-K))^\top\}$, where $A_\text{KS}$ is defined as in equation (ref).

A Unifying Sensitivity Model with Continuous Outcomes

As in Section (ref), all these relaxations can be viewed as special cases of a unifying class of relaxations encoding various types of departures from independence.

\begin{het1Assump}[General Sensitivity Model with Continuous Outcomes] For a known sensitivity parameter $\theta \in [0,1]$, suppose

align*[align* omitted — 82 chars of source]

where, for $x \in \{0,1\}$, $\mathcal{A}_x$ satisfies

enumerate• (Spanning) $\mathcal{A}_x(0) = \{(f_0,f_1) \in \mathcal{F}_{\text{den},x}^2: f_0 = f_1\}$ and $\mathcal{A}_x(1) = \mathcal{F}_{\text{den},x}^2$; • (Monotonicity) $\mathcal{A}_x(\theta) \subseteq \mathcal{A}_x(\theta')$ when $\theta \leq \theta'$; • (Linearity of Constraints) The set $\mathcal{A}_x(\theta)$ is a closed convex subset of $\mathcal{F}_{\text{den},x}^2$ characterized by finitely many componentwise weak linear inequalities in the densities for each $\theta \in [0,1]$; • (Continuity) The correspondence $\mathcal{A}_x: [0,1] \rightrightarrows \mathcal{F}_{\text{den},x}^2$ is continuous with respect to the sup-norm.

\end{het1Assump} The constraint set $\mathcal{A}_0(\theta) \times \mathcal{A}_1(\theta)$ is a convex set of functions defined by linear inequalities that weakly expands as $\theta$ increases. It nests the identified set under the baseline independence assumption ($\theta=0$) and the identified set under no assumptions on the dependence between potential outcomes and instruments ($\theta=1$). The third requirement is that the constraint set is of the form $\mathcal{A}_x(\theta) = \{\textbf{f} = (f_0,f_1) \in \mathcal{F}_{\text{den},x}^2: A(\theta)\textbf{f} \leq a(\theta)\}$ where $A(\theta)$ is a finite dimensional matrix. It involves finitely many componentwise weak inequalities, even though the inequality $A(\theta)\textbf{f} \leq a(\theta)$ hold for infinitely many values on the support $\mathcal{Y}_x$. As in the binary outcome case, this relaxation encompasses the previous three restrictions.

propositionSuppose assumptions (ref), (ref), and (ref) hold. Relabeling $(\lambda,K,c)$ as $\theta \in [0,1]$, the restrictions from equations (ref), (ref), and (ref) all satisfy Assumption (ref).

We now state our main result about theoretical properties of the identified set for densities of the potential outcomes.

theoremSuppose assumptions (ref), (ref), (ref), and (ref) hold, and suppose that $\mathcal{F}_{\text{den},x}^2 \cap \mathcal{H}_x \neq \emptyset$ for $x \in \{0,1\}$. Then, \begin{enumerate} • The identified set for $\textbf{f}_Y$ is \begin{align} \Pi(\theta) \coloneqq \Pi_0(\theta) \times \Pi_1(\theta), \end{align} where \begin{align*} \Pi_x(\theta) &\coloneqq \mathcal{H}_x \cap \mathcal{A}_x(\theta); \end{align*} • There exists $\underline{\theta}\in [0,1]$ such that $\Pi(\theta)$ is non-empty for $\theta \geq \underline{\theta}$ and empty for $\theta < \underline{\theta}$; • Suppose $\text{int}(\mathcal{H}_x \cap \mathcal{A}_x(\theta)) \neq \emptyset$ for all $\theta > \underline{\theta}$. Then, the correspondence $\Pi:[\underline{\theta},1] \rightrightarrows \mathcal{F}_{\text{den},0}^2 \times \mathcal{F}_{\text{den},1}^2$ defined by $\Pi(\theta)$ in equation (ref) is continuous. \end{enumerate}

This theorem establishes the main theoretical properties of the identified sets for densities, including their continuity as an infinite-dimensional correspondence. This continuity will carry over to functionals of these densities, in particular to linear or continuous functionals.

In particular, consider the class of linear mappings, for which the sharp bounds can be obtained as the solution to a linear program. Let \[ \Gamma(\textbf{f}) \coloneqq \int_{\mathcal{Y}_0} \omega_0(y)'\textbf{f}_Y(y ; \ 0)dy + \int_{\mathcal{Y}_1} \omega_1(y)'\textbf{f}_Y(y ; \ 1)dy \] where, for $x = 0,1$, $\omega_x$ is a known weight function that maps $\ensuremath{\mathbb{R}}$ to $\ensuremath{\mathbb{R}}^2$. The $\Gamma$ mapping is used to characterize a functional of the conditional densities of $Y(x)\mid Z$.

For example, with $\omega_1(y) = -\omega_0(y) = (y (1-p_Z), y p_Z)$, we have that

align*[align* omitted — 389 chars of source]

the average treatment effect. Letting $\omega_x(y) = \ensuremath{\mathbbm{1}}(y \leq a)$ and $\omega_{1-x}(y) = 0$ yields $\Gamma(\textbf{f}_Y) = \mathbb{P}(Y(x) \leq a)$, the cumulative distribution function evaluated at $a \in \ensuremath{\mathbb{R}}$.. This choice can be used to obtain bounds on quantiles of $Y(x)$ or on the quantile treatment effect $\text{QTE}(\tau) \coloneqq Q_{Y(1)}(\tau) - Q_{Y(0)}(\tau)$ for a quantile index $\tau \in (0,1)$.

The proposition below shows that bounds on these functionals are continuous and monotonic. This result uses the Maximum Theorem Berge1959 applied to an infinite-dimensional correspondence. Let \[ \overline{\Gamma}(\theta) \coloneqq \sup_{\textbf{f}_1 \in \Pi_1(\theta)} \int_{\mathcal{Y}_1} \omega_1(y_1)'\textbf{f}_1(y_1) \; dy_1 + \sup_{\textbf{f}_0 \in \Pi_0(\theta)}\int_{\mathcal{Y}_0} \omega_0(y_0)'\textbf{f}_0(y_0) \; dy_0 \] and \[ \underline{\Gamma}(\theta) \coloneqq \inf_{\textbf{f}_1 \in \Pi_1(\theta)} \int_{\mathcal{Y}_1} \omega_1(y_1)'\textbf{f}_1(y_1) \; dy_1 + \inf_{\textbf{f}_0 \in \Pi_0(\theta)}\int_{\mathcal{Y}_0} \omega_0(y_0)'\textbf{f}_0(y_0) \; dy_0 \] denote the lower and upper bounds of the functional over the sets $\Pi_x(\theta)$, $x=0,1$.

corollarySuppose the assumptions of Theorem (ref) hold. Let $\|\omega_x(\cdot)\|_\infty < \infty$. Then, \begin{enumerate} • Let $\theta\in [\underline{\theta},1]$. The identified set for $(f_{Y(0)}, f_{Y(1)})$ is $I_0(\theta) \times I_1(\theta)$ where $I_x(\theta) \coloneqq \{f_0 (1-p_Z) + f_1 p_Z: (f_0,f_1) \in \Pi_x(\theta)\}$ when $\theta\in [\underline{\theta},1]$, and the empty set when $\theta < \underline{\theta}$; • The functions $\underline{\Gamma}(\theta)$ and $\overline{\Gamma}(\theta)$ are continuous and monotonic over $\theta \in [\underline{\theta},1]$. • Let $\theta \in [\underline{\theta},1]$. The identified set for $\Gamma(\textbf{f}_{Y})$ is $[\underline{\Gamma}(\theta), \overline{\Gamma}(\theta)]$. \end{enumerate}

Therefore, as in the discrete case, bounds a can be obtained in the continuous case through infinite-dimensional linear programming. To make this approach feasible, we show in the next section how to convert an infinite-dimensional linear program into a feasible, finite-dimensional linear program that can be directly implemented.

Computation

The identified set $\Pi(\theta)$ is an infinite-dimensional set of continuous densities. If we restrict attention to the class of linear functionals described in Corollary (ref), the corresponding identified set is an interval (or the empty set). However, Corollary (ref) characterizes this interval by optimization over the infinite-dimensional spaces $\Pi(\theta)$, which is generally not feasible to compute directly. In this section, we discuss one approach to computing these identified sets by approximating the infinite-dimensional space of densities with a finite-dimensional sieve space and the constraint sets with a finite set of constraints. Similar approximations of identified sets have been used, for example, in MogstadSantosTorgovitsky2018. Alternatively, the computational approach developed in ChristensenConnault2023 could be adapted to our setting. Unlike the sieve-based approach we consider below, the dimension of their optimization problem does not depend on the precision of the density approximation. We leave the application of their approach to our problem to future work.

For simplicity, let $\mathcal{Y}_x = [0,1]$ for $x \in \{0,1\}$. This restriction can be relaxed by linearly transforming the outcome variable so that it has support on the unit interval. We also assume that $\mathcal{F} \coloneqq \mathcal{F}_0 = \mathcal{F}_1$, and therefore $\mathcal{F}_{\text{den}} \coloneqq \mathcal{F}_{\text{den},0} = \mathcal{F}_{\text{den}, 1}$. We also impose assumptions (ref) and (ref).

We will approximate $\mathcal{F}_{\text{den}}$ by the convex sieve space $\mathcal{F}_M$, defined by

align*[align* omitted — 101 chars of source]

where $\textbf{b}^M \coloneqq \{b_0^M, b_1^M, \ldots,b_{M}^M\}$ are the $M$-degree Bernstein basis polynomials scaled by $M + 1$. That is, $$ b^M_m(y) \coloneqq (M+1) \binom{M}{m} y^m (1-y)^{M-m} $$ for $m \in \{0, \ldots, M\}$.

Since $\mathcal{F}_M$ is increasing in $M$ and $\bigcup_{M:M > 0}\mathcal{F}_M$ is dense in $\mathcal{F}_{\text{den}}$, $\mathcal{F}_M$ is a sieve space for $\mathcal{F}_{\text{den}}$. We denote the Bernstein polynomial approximation to function $f$ at $y \in [0,1]$ as

align*[align* omitted — 100 chars of source]

We also define approximate constraint sets, which are characterized by a finite number of linear equality or inequality constraints. First, we approximate $\mathcal{H}_x$ by the sets,

align*[align* omitted — 198 chars of source]

where $\mathcal{F}_M^{s_Z}$ is the $s_Z$-fold Cartesian product of $\mathcal{F}_M$. In the proposition below, we show that replacing $f_{Y \mid X, Z}$ by its Bernstein approximation is sufficient to characterize this set by a finite number of linear constraints.

Next, we approximate $\mathcal{A}(\theta)$ using a finite set of inequalities. Each model in Section (ref) uses linear inequalities: $\mathcal{A}(\theta) = \{\mathbf{f} \in \mathcal{F}^{s_Z}_{\text{den}} : A(\theta)\mathbf{f} \le a(\theta)\}$. We use a grid of $N$ points in $[0,1]$ (for example, $y_n = n/(N + 1)$ for $n = 1, ..., N$), and then define $\mathcal{A}^{M,N}(\theta)$ as all $\textbf{f} \in \mathcal{F}_M^{s_Z}$ such that $A(\theta) \textbf{f}(y_n) \le a(\theta)$ for each grid point.

The approximate identified set for $\textbf{f}_Y$ is $\Pi_0^{M,N}(\theta) \times \Pi_1^{M,N}(\theta)$, where $\Pi_x^{M,N}(\theta)$ is the intersection of $\mathcal{A}^{M,N}(\theta)$ and $\mathcal{H}_x^M$. The next proposition gives a more convenient representation of this set for computation. Here, $\bar{\Delta}_s^r \coloneqq \left\{

bmatrix[bmatrix omitted — 33 chars of source]

^{\top} : a_j \in \Delta_s for j = 1, \ldots, r\right\}$. $\operatorname{vec}(W)$ is the vectorization of matrix $W$, and $\otimes$ is the Kronecker product.

propositionFor $\mathcal{A}(\theta) = \{\mathbf{f} \in \mathcal{F}_{\text{den}}^{s_Z} : A(\theta) \mathbf{f} \le a(\theta)\}$, $N,M \in \mathbb{N}$, and $\{y_1, \ldots, y_N\} \subset [0,1]$, the approximate constraint sets $\mathcal{A}^{M,N}(\theta)$ and $\mathcal{H}_x^M$ can be represented as \begin{align*} \mathcal{A}^{M,N}(\theta) = \{ W b^M : W \in \mathcal{W}^{M,N}(\theta) \} \end{align*} and \begin{align*} \mathcal{H}_x^M = \{ W b^M : W \in \mathcal{W}_x^M \} \end{align*} where \begin{align*} \mathcal{W}^{M,N}(\theta) &\coloneqq \left\{ W \in \bar{\Delta}_{M}^{s_Z} : \left( (B^{M,N})^{\top} \otimes A(\theta) \right) \operatorname{vec}(W) \le \iota_{N} \otimes a(\theta) \right\} \\ \mathcal{W}_x^{M} &\coloneqq \left\{ D_x \Xi^M_x + D_{1-x} W : W \in \bar{\Delta}_{M}^{s_Z} \right\}. \end{align*} In $\mathcal{W}^{M,N}(\theta)$, we define $B^{M,N} \coloneqq \begin{bmatrix} \textbf{b}^M(y_1) & \ldots & \textbf{b}^M(y_{N}) \end{bmatrix}$ and $\iota_{N}$ to be the $N$-dimensional vector of ones. In $\mathcal{W}_x^M$, we define $\textbf{D}_x \coloneqq \operatorname{diag}(\pi(x \mid z_1), \ldots, \pi(x \mid z_{s_Z}))$ and $\Xi^M_x$ to be the $s_Z \times (M + 1)$ matrix with elements $f_{Y \mid X, Z}\left( \frac{m-1}{M} \mid x, z_j\right)$ in the $(j, m)$-th position.

This proposition shows that the approximate identified set, $\prod_{x \in \{0, 1\}} \left(\mathcal{A}(\theta)^{M,N} \cap \mathcal{H}_x^M\right)$ can be characterized by a finite number of linear constraints. Following Corollary 2, we use this result to characterize the approximate identified set of a functional of $\textbf{f}_Y$ as the solution to a finite linear program.

Approximating the functional $\Gamma(\textbf{f}_Y) = \int_{\mathcal{Y}_0} \omega_0(y)^{\top} f_0(y) dy + \int_{\mathcal{Y}_1} \omega_1(y)^{\top} f_1(y) dy$ with a Riemann sum with $L$ points, we can characterize $\underline{\Gamma}^{M,N,L}(\theta)$ as the solution to the linear program,

align[align omitted — 790 chars of source]

The linear inequalities (ref) and (ref) correspond to the constraints that $W_1, W_0 \in \mathcal{W}^{M,N}(\theta)$, and the equality constraints (ref) and (ref) together with the simplex constraints on $W_{x,1-x}$ for $x \in \{0, 1\}$ correspond to the constraints that $W_1 \in \mathcal{W}_1^M$ and $ W_0 \in \mathcal{W}_0^M$ respectively. The optimization program is therefore a linear program in the $s_Z \times (M + 1)$ weight matrices $W_1, W_{1,0}, W_0, W_{0,1}$, which can be solved using standard software.

$\overline{\Gamma}^{M,N,L}(\theta)$ is the solution to the corresponding maximization problem, which is also a linear program.

Since $\Pi^{M,N}(\theta)$ is closed, bounded, and convex, the approximate identified set is

align*[align* omitted — 161 chars of source]

Although we omit a full analysis, we expect that $\underline{\Gamma}^{M,N,L}(\theta)$ and $\overline{\Gamma}^{M,N,L}(\theta)$ will converge to $\underline{\Gamma}(\theta)$ and $\overline{\Gamma}(\theta)$ respectively as $M, N, L \rightarrow \infty$ under suitable regularity conditions.

Empirical Application

Here we revisit the empirical study of peer effects in consumer demand by GilchristSands2016. Specifically, they study whether movie viewership is affected by peer viewership choices. They provide evidence that movie viewership can have “momentum" from one weekend to the next. They argue that this is partly because if a movie does well on its opening weekend, it motivates people to see it in subsequent weekends, so they can discuss it with their peers or attend it as a social event.

Identifying this effect is a challenging empirical problem: an apparent peer effect on consumer demand could simply reflect a common understanding of the movie’s unobserved quality. To address this, the authors use a classic instrumental variables approach, using weather as an instrument for opening weekend viewership. They argue that outdoor activities are a substitute for going to the movies, so days with especially nice weather provide a plausibly negative, exogenous shock to viewership.

While its inherent randomness makes weather an appealing instrument, recent literature has cast doubt on its validity as an instrument in many contexts (e.g., Mellon2025). For this application, we highlight three potential violations of the exclusion assumption: (1) social learning about movie quality, (2) dynamic consumer behavior, and (3) dynamic behavior by movie studios.

GilchristSands2016 acknowledge that social learning is an important alternative explanation for the observed momentum in movie viewership. The concern is that consumers may be uncertain about a movie's quality and rely on their peers to learn about it. When viewership is high, there is a higher probability that a consumer has friends who have seen the movie and can share their opinion of it. For more reluctant consumers, they may wait until they have good information about the film's quality before seeing it. This is a similar but distinct mechanism from the social incentive that the authors are interested in.

One approach would be to redefine the “peer effect” to include this learning effect; however, GilchristSands2016 are clear that they are interested in the direct social incentive to see the movie. Instead, they explore whether there are learning effects by testing an implication from a model of social learning in Young2009. This auxiliary model introduces several additional strong behavioral and distributional assumptions, and the results are not decisive. They conclude that “Although our estimates do not rule out some role for learning, taken together the results suggest that the observed momentum is driven in part by a preference for shared experience, and not only by learning.” GilchristSands2016.

Dynamic behavior could also lead to violations of exclusion. When a consumer skips seeing a particular movie one weekend to enjoy the weather, she may simply plan to see the movie on a future weekend. However, the set of available movies in that future weekend is often different, possibly leading them to make a different choice about what movie to see altogether. Similarly, movie studios may respond to first-weekend viewership by adjusting their advertising strategy, which could affect subsequent viewership.

Finally, we note an additional challenge to the exogeneity condition which GilchristSands2016 address directly in their main specifications. Movie studios may strategically time movie release dates based on seasonal weather patterns, inducing a correlation between weather shocks and unobserved movie quality. To address this problem, the authors condition on several calendar controls, including the week of the year, the year, and holiday indicators. Since movie studios have to release movies based on their expectations of the weather far in advance rather than short-term forecasts, they argue that this strategic behavior should be captured by these calendar controls. In our analyses, we follow their approach of controlling for these time-of-year variables. However, this could still be insufficient if movie studios use more accurate long-term weather forecasts than the average weather for that week of the year.

These potential violations of the exclusion and exogeneity assumptions motivate the importance of assessing sensitivity in this application.

Data and Definitions

We use the dataset assembled by GilchristSands2016 for our analysis. Viewership data on daily ticket sales is obtained from the Internet Movie Database (IMDb) for all movies released between 2002 and 2013. The sample is restricted to movies that were in theaters for at least six weeks, and uses only data on ticket sales on Friday, Saturday, and Sunday.

The instruments are measures of the weather on each weekend. These data come from Weather Underground and consist of (1) the daily maximum temperature, (2) inches of rain, and (3) inches of snow in \(1,941\) weather stations across the country. In order to create national aggregate measures, weather station-level data is weighted by $\omega_j = \frac{n_j}{\sum_j n_j}$ for each weather station $j$ where $n_j$ is the number of movie theaters for which $j$ is the closest weather station.\footnote{To do this, they first assign each zip code to the closest weather station, and obtain the number of movie theaters in each zip code from the U.S. Census Zip Code Business Patterns data.} For any weather station-level weather measure, $Z_{tj}$, the aggregate instrument is, $Z_{t} = \sum_j \omega_j Z_{tj}$.

We define the potential outcome, $Y_{i}(x)$, to be the viewership of movie $i$ in the second weekend of its release with or without a negative shock to viewership in the opening weekend, $x$. The treatment is binary, with $x = 1$ when opening-weekend viewership is below its 25th percentile. This specification of the treatment is motivated by the observation in GilchristSands2016 that good weather tends to suppress viewership.

We want to ask whether such a negative shock to initial viewership increases the probability of low viewership in subsequent weekends through peer effects. We begin by defining low viewership in the second weekend analogously to the treatment. Specifically, we consider the summary outcome $\mathbf{1}(Y_i(x) \le \underbar{y})$, where $\underbar{y}$ is the 25th percentile of viewership in the second weekend across all movies. The natural parameter of interest is the average treatment effect (ATE), $\mathbb{E}(\mathbf{1}(Y_i(1) \le \underbar{y}) - \mathbf{1}(Y_i(0) \le \underbar{y}))$. This is the effect of a negative shock to opening weekend viewership on the probability of low viewership in the second weekend. Moving beyond this coarse measure of low viewership in the second weekend, we also consider quantile treatment effects across the distribution of viewership in that weekend. That is for, a range of quantiles $\tau \in (0,1)$, we consider the parameter $\text{QTE}(\tau) = Q_{Y_i(1)}(\tau) - Q_{Y_i(0)}(\tau)$ where $Q_{Y_i(x)}(\tau)$ is the $\tau$th quantile of the distribution of $Y_i(x)$.

To minimize endogeneity between movie quality and opening weekend weather, we follow the approach of Gilchrist and Sands (2016) and residualize all variables (viewership in the first and second weekends and the weather instrument) using a set of week-of-year dummies. We use their preferred weather instrument, the share of theaters with a daily high temperature between 75 and 80 degrees Fahrenheit, which we discretize into quintiles. Finally, we condition on this same weather variable on the second weekend. This helps control for potential serial correlation in weather across weekends, which is not captured by the week-of-year dummies.

Sensitivity Analysis

We begin with the discretized outcome. Under the baseline of exogeneity, we find that a negative shock on viewership in the initial weekend increases the probability of low viewership in the second weekend. The estimated identified set for the ATE is $[0.04, 0.87]$. This result, which bounds the ATE above zero, is qualitatively consistent with the conclusion of Gilchrist and Sands (2016), who find a positive effect of opening weekend viewership on subsequent weekend viewership using a 2SLS estimator. While the lower bound is small, this means that peer effects increase the probability of low viewership in the second weekend by at least $4.4\%$, which is a quantitatively important effect size.

We find, however, that this conclusion is sensitive to relatively small violations of the exogeneity assumption. In Table (ref) we present the estimated ATE bounds for different levels of $c$-dependence. The interval between the lower and upper lines is the identified set for the ATE at each level of $c$-dependence. Even at low levels of $c$-dependence, the identified set for the ATE includes zero. The lowest level of $c$-dependence at which the identified set for the ATE includes $0$ -- the breakdown point -- is $0.015$, or when the latent propensity score is allowed to be 1.5 percentage points away from the observed propensity score.

table[table omitted — 476 chars of source]

To explore the distributional effects of a negative shock to opening weekend viewership, we now turn to the quantile treatment effects (QTE) for the continuous outcome $Y_i(x)$. In Table (ref), we report the identified set of the QTE across several quantiles and different levels of $c$-dependence. Consistent with the results of the discretized outcome, we find that under the baseline assumption of exogeneity, the identified set for the QTE at the $10$th and $25$th percentile is negative and bounded away from zero. A negative shock to opening weekend viewership causes the 25th percentile of viewership in the second weekend to decrease by at least $0.39$ million tickets. These results, however, hold only for the bottom half of the distribution of potential outcomes. At the $50$th, $75$th, and $90$th percentiles, the identified set is very wide and includes zero.

table[table omitted — 698 chars of source]

To see why the identified set for the QTE is much less informative for higher quantiles, it is useful to examine the identified sets for the potential outcome CDFs directly. Figure (ref) shows the upper and lower bounds on the CDF for $(Y_i(1), Y_i(0))$ at different levels of $c$-dependence. The first panel shows the bounds under exogeneity, while the second shows a $c$-dependence level of $0.1$. There is an asymmetry in the bounds of the distributions of potential outcomes, with much tighter bounds for the potential outcome with $x = 0$ in which viewership in the opening weekend is above the $25$th percentile. This is because there is a much larger mass of observations with $X = 0$ than with $X = 1$. In addition, the data is largely uninformative about the top half of the distribution of $Y_i(1)$. This reflects the fact that nearly all of the observed mass of $Y$ conditional on $X = 1$ is in the lower half of the support of $Y$. Since we make no monotonicity assumption or other shape restriction, the bounds on the CDF of $Y_i(1)$ have no other restriction except for the lower bound from the mass below $0$.

figure[figure omitted — 23,716 chars of source]

Conclusion

We introduced a new, computationally tractable approach for conducting sensitivity to the instrument exclusion and exogeneity assumptions. Our approach does not impose any kind of monotonicity assumption in the first stage, and allows for arbitrarily heterogeneous treatment effects. We did this by developing a unifying sensitivity model which nests several well known approaches to continuously parameterizing relaxations of statistical independence assumptions from the literature. We showed that, under those relaxations, identified sets for parameters like ATE and QTE are solutions to linear programs. Our approach can be used when the outcome is discrete or continuous, and when there are one or multiple discretely supported instruments.

We illustrated the practical value of our results in an empirical study of peer effects in movie viewership. There our sensitivity analysis shows that although ATE is positive under full exclusion and exogeneity (meaning peer effects are present), that conclusion is highly sensitive to minor relaxations of the exclusion and exogeneity assumptions. Overall, our results allow researchers to transparently study and report the robustness of their instrumental variable conclusions to violations of exclusion or exogeneity.