EconBase
← Back to paper

Assessing Sensitivity to IV Exclusion and Exogeneity without First Stage Monotonicity

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

117,170 characters

Assessing Sensitivity to IV Exclusion and Exogeneity without First Stage Monotonicity


\maketitle

\vspace{-2em}

\begin{abstract}
Exclusion and exogeneity are core assumptions in instrumental variable (IV) analyses, but their empirical validity is often debated. This paper develops new sensitivity analyses for these assumptions. Our results accommodate arbitrary heterogeneity in treatment effects and do not impose any monotonicity requirements on the first stage. Specifically, we derive identified sets for the marginal distributions of potential outcomes and their functionals, like average treatment effects, under a broad class of nonparametric relaxations of the exclusion and exogeneity assumptions. These identified sets are characterized as solutions to linear programs and have desirable theoretical properties. We explain how to estimate these solutions using computationally tractable methods even when the linear program is infinite-dimensional. We illustrate these methods with an empirical application to peer effects in movie viewership, using weather as a potentially imperfect instrument.
\end{abstract}

\bigskip
\small
\noindent \textbf{JEL classification:}
C14, C18, C21, C26, C51


\bigskip
\noindent \textbf{Keywords:}
Instrumental Variables, Sensitivity Analysis, Nonparametric Identification, Partial Identification

\allowdisplaybreaks

\onehalfspacing

\newpage

\section{Introduction}\label{sec:intro}

Instrumental variable (IV) analyses typically rely on two core assumptions: Instrument \textit{exclusion} and instrument \textit{exogeneity}. Exclusion holds when the instrument has no direct effect on the outcome, while exogeneity holds when the instrument is randomly assigned. Since the work of \cite{ImbensAngrist1994}, a third assumption is also often imposed: First stage \emph{monotonicity}. In the simplest setting where the treatment and instrument are binary, monotonicity holds if the instrument's effect on treatment is always of the same sign.

All three assumptions can be hard to justify in some empirical settings. Instruments may have direct effects on outcomes or may not be randomly assigned. Monotonicity can also fail. This occurs, for example, in leniency designs (also called `judge IV' designs) where monotonicity implies that a judge is stricter or more lenient in the face of any possible case; see \cite{FrandsenLefgrenLeslie2023Judging} for details. In designs with many treatment and instrument values, there is no single monotonicity assumption to choose from, and it may be difficult to find one suitable for one's empirical setting.

To address these concerns, we study identification of treatment effects in a setting where no monotonicity conditions are imposed whatsoever, but where exogeneity and exclusion are assumed to at least partially hold in some sense. Specifically, we introduce a unifying class of continuous relaxations of instrument exclusion and exogeneity that nests several prominent approaches in the literature. In particular, it includes as special cases the marginal sensitivity model (MSM) of \cite{Tan2006}, $c$-dependence by \cite{MastenPoirier2018}, and supremum distance approaches, as in \cite{Manski1983} and \cite{KlineSantos2013}. All these approaches were developed as sensitivity models for unconfoundedness (or selection on observables), but we develop modified versions suitable for IV sensitivity analysis. In each of those cases, the sensitivity model is indexed by a scalar, unit-free sensitivity parameter that is easy to interpret.


When the outcome variable is discrete, we show that the identified sets for the conditional probabilities of the potential outcomes given the instruments can be characterized as the intersection of two convex sets, parameterized by the relaxation of instrument exclusion and exogeneity.   Using this result, we then show that the identified set for a class of linear functionals of the densities of potential outcomes is the solution to a linear program that can be computed efficiently. This class of functionals includes the standard treatment parameters such as the ATE, the average effect of treatment on the treated (ATT), and quantile treatment effects (QTE).

We show that these identified sets exhibit many desirable properties, including continuity and monotonicity with respect to the sensitivity parameter. As is well known (e.g., \citealt{BalkePearl1997}), IV models have testable implications, which can fail in practice. In this case, we characterize the smallest deviations from the baseline model that prevent the model from being refuted.

We then extend our results to the case where outcomes are continuous. This case is more delicate, as the distribution of outcomes is now characterized by a density function, an infinite-dimensional object. We show that the identified set for densities and its functionals (like ATE, ATT, or QTE) can also be characterized via a linear program, albeit an infinite-dimensional one. As in the discrete case, we show that identified sets derived from this linear program have desirable properties by analyzing them as infinite-dimensional-valued correspondences, since each sensitivity parameter is now associated with a set of infinite-dimensional-valued density function. As these linear programs cannot be solved in practice, we propose a tractable approach for approximating the problem with a finite-dimensional one.

Using these computational results, we show how applied researchers can produce sensitivity plots that show the sensitivity (or robustness) of their parameter of interest to exclusion and exogeneity violations. These plots can be used, for example, to determine how strong exclusion or exogeneity violations can be before the data is consistent with a zero treatment effect.

To illustrate our approach, we revisit Gilchrist and Sands' \citeyearpar{GilchristSands2016} study of peer effects in movie viewership, using weather as an instrument for opening-weekend viewership. While extremely popular in empirical practice, weather instruments have come under increasing scrutiny in recent years (e.g., \citealt{Sarsons2015}, \citealt{GallenRaymond2023}, \citealt{Mellon2025}). In this application, social learning and dynamic behavior could lead to violations of instrument exclusion, and use our results to assess the robustness of conclusions to relaxations of this assumption. Using both discretized and continuous outcomes, we confirm that under the baseline of instrument exclusion, there is a positive peer effect on viewership, but show that this conclusion is sensitive to relatively small relaxations of the exogeneity assumption.

The rest of the paper is organized as follows. We first provide an overview of the related literature. Section \ref{sec:binary_outcome} then develops the framework for binary outcomes, introduces the relaxation class, and derives sharp identified sets, falsification frontiers, and falsification adaptive sets in the discrete setting. Section \ref{sec:cont_Y} extends the analysis to continuous outcomes, establishes the corresponding identification and continuity results, and details a sieve-based computational strategy. Section \ref{subsec:empirical} presents our empirical application.



\subsubsection*{Related Literature}\label{subsec:litreview}

Research on the sensitivity of IV results to violations of exclusion and exogeneity go back to at least \cite{Fisher1961}. More recent developments were proposed by \cite{BoundJaegerBaker1995}, \cite{Small2007} and  \cite{ConleyHansenRossi2012}. All of these methods assume a linear outcome equation, motivated by treatment effect homogeneity, which we do not assume. These papers bound the direct effect of the instrument on the outcome, which can be done by bounding the coefficient on the instrument under the assumption that the potential outcomes depend linearly on it. Various approaches have been proposed to bound this direct effect, including \cite{NunnWantchekon2011}, \cite{ConleyHansenRossi2012}, \cite{Kraay2012},  \cite{KippersluisRietveld2017}, \cite{KippersluisRietveld2018}, and \cite{MastenPoirier2021}.  Also see \cite{AltonjiElderTaber2005a},  \cite{Ashley2009}, and \cite{AshleyParmeter2015} for alternative approaches.

Our paper contributes to the literature on sensitivity analysis in instrumental variable models with heterogeneous treatment effects, which is much sparser than that for homogeneous treatment effects. Specifically, few papers consider continuous relaxations of the baseline instrumental variable assumptions while still allowing for heterogeneous treatment effects.

Early work by \citet{Manski1990} characterizes sharp bounds on average treatment effects under two sets of assumptions: (i) instrument exclusion and exogeneity hold (formulated as mean independence in his general analysis) or (ii) instrument exclusion and exogeneity fail arbitrarily. Our continuously parameterized sensitivity model spans these two sets of assumptions, allowing users to calibrate the degree of exclusion and exogeneity violations from ``no-violations'' (i.e., full exclusion and exogeneity) to ``no assumptions''.

\cite{HotzMullinSanders1997} used a mixture model to allow for relaxations of the baseline assumptions. They focus on the average effect of treatment on the treated, whereas our sensitivity analysis allows for a broader set of parameters of interest. \cite{Ramsahai2012} studied a heterogeneous treatment effect model with a binary outcome, binary treatment, and a binary instrument. He defines a continuous relaxation of the instrument exogeneity assumption and then shows how to numerically compute identified sets for a single value of this relaxation. On pages 842-–843, he notes that ``it is not obvious how the methods described in [his] paper can be extended to compute bounds'' as a function of his relaxation. In our analysis, we allow all variables to be nonbinary, and even continuous for the outcome variable, and we allow for multiple instruments. We also consider a large set of target parameters and derive theoretical and computational properties for the sensitivity plots, which map the sensitivity parameters into this range of target parameters. Also see \cite{Huber2014} and \cite{MachadoShaikhVytlacil2019} for related, but different, approaches.

In fully discrete cases, identified sets for causal parameters and counterfactual distributions can often be obtained via linear programming. This observation goes back at least to \citet{BalkePearl1997} (and related work by \citealt{Pearl1995}) and is emphasized in more recent reviews of discrete partial identification methods. For example, see the literature review in \cite{Torgovitsky2019}. Linear programming has been used in several papers to do sensitivity analysis. One paper is \cite{Ramsahai2012}, which we already discussed above. \citet[section 4]{Laffers2019} considers continuous relaxations of instrument exogeneity. He then computes identified sets for ATE for several values of this relaxation. In \cite{Laffers2018}, he applies this approach to various additional forms of continuous relaxations. \cite{Duarte2024} also uses linear programming to bound parameters under exclusion and monotonicity violations. These papers all require all variables to be discrete. A key contribution of our paper is that our results allow for continuous outcome variables.

Our paper also contributes to assessments of IV model falsification. \cite{BalkePearl1997} characterize when Manski's bounds are empty, and hence when the model is falsified.\footnote{\cite{BalkePearl1997} assume the instrument is independent of the potential outcomes jointly, whereas \cite{Manski1990} only assumed the instrument is independent of each potential outcome separately. (Here we suppose outcomes are binary, so that mean independence is equivalent to statistical independence.) This difference does not affect whether the identified set is empty, given any fixed distribution of observables. Hence, it does not change the testable implications of the model. When the identified set is nonempty, however, this difference \emph{can} affect its size. See the second paragraph of section 3 in \cite{SwansonHernanMillerRobinsRichardson2018} for further discussion.} \citet[Proposition 3.1]{Kitagawa2021} generalizes this characterization to allow for continuous outcomes, still requiring the treatment and instrument to be binary. As \cite{Kitagawa2021} notes, his extension is an adaptation of Corollary 2.2.1 in Manski's \citeyearpar{Manski2003} analysis of missing data. \citet[Proposition 2.4]{BeresteanuMolchanovMolinari2012} further generalizes this characterization to allow for continuous instruments and discrete treatment, for discrete or continuous outcomes. \citet[Proposition 1]{KedagniMourifie2020} provides an alternative characterization when instruments and outcomes are continuous, treatment is binary, and under the stronger assumption that the instrument is independent of the potential outcomes jointly; also see Proposition 2.5 of \cite{BeresteanuMolchanovMolinari2012} for a result under this stronger independence assumption.


Finally, a large literature on the testable implications of instrument exclusion and exogeneity, combined with other assumptions, has developed. Most notably, many papers have studied the testable implications of the monotonicity assumption of \cite{ImbensAngrist1994}. \cite{FloresChen2018} give a comprehensive review. Also see \cite{FrandsenLefgrenLeslie2023Judging} for discussions of monotonicity in the judge IV framework. In this paper, we focus on instrument exclusion and exogeneity only.


\section{Sensitivity Analysis with Binary Outcomes}\label{sec:binary_outcome}

We begin by considering analyses with a binary outcome. For further simplicity, we also assume that the treatment and instrument are binary. The results below generalize to a setting with multiple treatment values and multiple discrete instruments, but we focus on the binary case, which allows us to explain the main ideas and results while keeping the notation simple. See Section \ref{subsec:discrete_generalization} for their generalization. The case where the outcome variable is continuously distributed presents additional technical challenges and is analyzed in Section \ref{sec:cont_Y}.

\subsection{Model, Parameters of Interest, and Assumptions}
Let $X \in \{ 0, 1 \}$ denote the observed binary treatment variable and $Z \in \{0,1 \}$ denote an observed instrument. As mentioned above, we consider multiple treatments and discrete instruments later. Let $\{Y(x,z)\}_{x,z \in \{0,1\}}$ denote potential outcomes for both treatment and instrument values. The observed outcome is denoted by
\begin{equation}\label{eq:generalOutcomeEquation}
 Y = Y(X,Z).
\end{equation}
We assume the joint distribution of $(Y,X,Z)$ is known in this identification analysis. Our analysis could be done conditional on a vector $W$ of covariates, but we omit them for simplicity.  Let $p_Z \coloneqq \ensuremath{\mathbb{P}}(Z=1)$ and $\pi(x \mid z) \coloneqq \ensuremath{\mathbb{P}}(X = x \mid Z = z)$. We maintain the following assumption to rule out trivial cases.
\begin{het1Assump}\label{assn:NonTrivialinstrument}
 Let $p_Z \in (0,1)$ and $\pi(x \mid z) \in (0,1)$ for all $x,z \in \{0,1\}$.
\end{het1Assump}

We define $p_Y(x,z) \coloneqq \ensuremath{\mathbb{P}}(Y(x,z) = 1\mid Z = z)$, conditional probabilities of the potential outcomes given the instrument. Let $\textbf{p}_Y(x) \coloneqq (p_Y(x,0),p_Y(x,1)) \in [0,1]^2$ and $\textbf{p}_{Y} \coloneqq (\textbf{p}_Y(0),\textbf{p}_Y(1)) \in [0,1]^4$ be collections of these conditional probabilities. We are interested in functionals of these conditional probabilities, denoted by $\Gamma: [0,1]^4 \rightarrow \ensuremath{\mathbb{R}}$, which include various treatment effect parameters. In this section, we focus our attention on averages of treatment effects such as the average treatment effect (ATE) and the average treatment effect on the treated (ATT). They can be viewed as functionals of $\textbf{p}_Y$ as follows:
\begin{align*}
 \text{ATE} &\coloneqq \mathbb{E}[Y(1,Z) - Y(0,Z)] = \Gamma_\text{ATE}(\textbf{p}_{Y})\\
 \text{ATT} &\coloneqq \mathbb{E}[Y(1,Z) - Y(0,Z) \mid X = 1] = \Gamma_\text{ATT}(\textbf{p}_{Y}),
\end{align*}
where
\begin{align}
 \Gamma_{\text{ATE}}(\textbf{p}_{Y}) &\coloneqq  p_Y(1,1)p_Z + p_Y(1,0)(1-p_Z) - p_Y(0,1)p_Z - p_Y(0,0)(1-p_Z)\label{eq:ATE_functional}\\
 \Gamma_{\text{ATT}}(\textbf{p}_{Y}) &\coloneqq \mathbb{E}[Y \mid X = 1] - \frac{p_Y(0,1)p_Z + p_Y(0,0)(1-p_Z) - \mathbb{E}[Y \mid X = 0](1-p_Z)}{p_Z}.\label{eq:ATT_functional}
\end{align}
These parameters are well-defined even in the absence of exclusion or exogeneity assumptions about the instruments. Additional parameters could be of interest, such as the local average treatment effect (LATE). The LATE is defined in terms of potential treatments, which could be incorporated into our framework but are not required by it.

Before introducing additional assumptions, we first characterize the identified set  for $\textbf{p}_{Y}$ when no assumptions are made about the joint distribution of $(Y,X,Z)$ beyond the regularity assumption \ref{assn:NonTrivialinstrument}. To do so, define
\begin{align}
\label{eq:boxSet}
 \mathcal{H}_x
 &\coloneqq [\mathbb{P}(Y=1,X=x \mid Z=0), \ \mathbb{P}(Y=1,X=x \mid Z=0) + \pi(1-x \mid 0)]\notag \\
 &\qquad \times [\mathbb{P}(Y=1,X=x \mid Z=1), \ \mathbb{P}(Y=1,X=x \mid Z=1) + \pi(1-x \mid 1)],
\end{align}
which depends on the joint distribution of $(Y,X,Z)$. With this notation, we obtain the following result which is adapted from \cite{Manski1990}.
\begin{proposition}[\citealt{Manski1990}]\label{prop:noassn_discrete}
 Suppose Assumption \ref{assn:NonTrivialinstrument} holds. Then the identified set for $\textbf{p}_Y$ is $\mathcal{H}_0 \times \mathcal{H}_1$.
\end{proposition}
This result shows that the identified set for the conditional probabilities $\textbf{p}_Y$ is a Cartesian product of intervals, i.e., a hyperrectangle. As these bounds are sharp, they can be used to obtain sharp bounds on any functional of $\textbf{p}_Y$. For example, the functional $\Gamma_{\text{ATE}}$ is linear and the set $\mathcal{H}_0 \times \mathcal{H}_1$ is a Cartesian product of intervals, so appropriately evaluating $\Gamma_{\text{ATE}}$ at the lower/upper bounds of the intervals in \eqref{eq:boxSet} will yield sharp bounds for it. The same approach can be used to obtain sharp bounds of the ATT, for example. For any linear functional, this is equivalent to a linear program, which is easy to solve analytically given the discrete supports of $X$ and $Z$. Figure \ref{fig:FFbinaryYsingleIV} illustrates this identified set and the optimization of the ATE over the identified set for $\textbf{p}_Y(x)$.


\begin{figure}[t]
\centering
\includegraphics[width=75mm]{images/newpics/pic1.pdf}
\hspace{5mm}
\includegraphics[width=75mm]{images/newpics/pic2.pdf}
\footnotesize
\caption{Left: Example identified set for $\textbf{p}_Y(x) = (p_Y(x,0),p_Y(x,1))$ under no exogeneity assumptions. Right: corresponding linear program minimizing/maximizing the ATE $= p_Y(x,0)(1-p_Z) + p_Y(x,1)p_Z$.}
\label{fig:FFbinaryYsingleIV}
\end{figure}

These bounds can be considerably tightened by assuming exogeneity or exclusion, as we formally define below.

\subsubsection*{Baseline Assumptions}

We now introduce the assumptions we will study in this model. We compare these assumptions to the four assumptions usually imposed in a large segment of the literature, including the traditional Local Average Treatment Effect (LATE) framework: exogeneity, exclusion, monotonicity, and relevance. For brevity, we do not include covariates in this discussion, although all the upcoming assumptions can be stated conditional on a covariate vector $W$.

First, we formally define the exogeneity and exclusion assumptions we consider.
\begin{definition}[Exogeneity] \label{def:exogeneity}
 The instrument is exogenous if $Z 
  \mathbin{
    \mathpalette{\@indep}{}
  }
 Y(x,z)$ holds for each $(x,z) \in \operatorname*{supp}(X) \times \operatorname*{supp}(Z)$.
\end{definition}

Exogeneity holds when the instrument is randomly assigned, or as good as randomly assigned, with respect to the potential outcomes. We do not require that the instrument be independent of potential treatment values, although this assumption can be incorporated into the framework. As mentioned earlier, we could consider relaxing the conditional exogeneity assumption $Y(x) 
  \mathbin{
    \mathpalette{\@indep}{}
  }
 Z \mid W$, at the cost of additional notation.

Next, we consider an exclusion assumption that is weaker than the most commonly used version.
\begin{definition}[Weak Exclusion]\label{def:exclusion}
 The instrument is weakly excluded if $Y(x,z) \mid \{Z = z''\} \overset{d}{=} Y(x,z') \mid \{Z = z''\}$ for all $x \in \operatorname*{supp}(X)$ and $z,z',z'' \in \operatorname*{supp}(Z)$.
\end{definition}

The standard exclusion assumption is that $Y(x,z) = Y(x,z')$ with probability 1 for any possible treatment value $x$ and any possible instrument values $z$ and $z'$, whereas weak exclusion only requires the (conditional) distributions of these potential outcomes to be identical. This has also been called \textit{stochastic exclusion}; see, for example, \cite{SwansonHernanMillerRobinsRichardson2018}.  Although we do not study the LATE here, the arguments used to obtain a causal interpretation for the Wald estimand are not impacted if exclusion is replaced by weak exclusion. In particular, the Wald estimand equals the LATE when the treatment and instrument are binary under weak exclusion, provided appropriate exogeneity, relevance, and monotonicity conditions hold.

We will assume that the instrument is exogenous or weakly excluded, without requiring it to satisfy both.
\begin{het1Assump}\label{assn:exog_or_excl}
 The instrument $Z$ is exogenous or weakly excluded.
\end{het1Assump}
Under this assumption, $p_Y(x,z)$ can be interpreted in one of two ways. Under exogeneity, this probability equals the unconditional probability $\mathbb{P}(Y(x,z) = 1)$, while under weak exclusion it denotes the conditional probability $\mathbb{P}(Y(x) = 1 \mid Z = z)$. If both hold, then $p_Y(x,z)$ does not depend on $z$, meaning that $\mathbb{P}(Y(x,1) = 1 \mid Z = 1) = \mathbb{P}(Y(x,0) = 1 \mid Z = 0)$, but Assumption \ref{assn:exog_or_excl} allows the dependence of $p_Y(x,z)$ on $z$ to be nontrivial. We formally show that under Assumption \ref{assn:exog_or_excl}, $p_Y(x,z)$ not depending on $z$ implies the exogeneity \textbf{and} weak exclusion of the instrument.

\begin{lemma}[Condition for Exogeneity and Weak Exclusion]\label{lemma:equivalence}
 Suppose assumptions \ref{assn:NonTrivialinstrument} and \ref{assn:exog_or_excl} hold. Then $p_Y(x,1) = p_Y(x,0)$ for all $x \in \{0,1\}$ if and only if $Z$ is exogenous and weakly excluded.
\end{lemma}

Thus, we can view failures of exogeneity or exclusion as mathematically equivalent to the probabilities $p_Y(x,z)$ being nonconstant in $z$.

To simplify our exposition going forward, we let
\[
	Y(x) \coloneqq Y(x,Z).
\]
Note that $p_Y(x,z) = \mathbb{P}(Y(x) = 1\mid Z=z)$, and that the ATE and ATT functionals $\Gamma_{\text{ATE}}$ and $\Gamma_{\text{ATT}}$ are defined as functionals of $Y(x)$, independently of whether weak exclusion or exogeneity holds.  Also, the bounds of Proposition \ref{prop:noassn_discrete} do not change when Assumption \ref{assn:exog_or_excl} is imposed.  With these definitions, we can see that the instrument is exogenous and weakly excluded if and only if
\[
	Y(x) 
  \mathbin{
    \mathpalette{\@indep}{}
  }
 Z
\]
for $x \in\{0,1\}$. Hence, we will consider relaxations of exogeneity or weak exclusion as relaxations of an independence assumption, as they are mathematically equivalent here.


To finish our comparison to the standard IV assumption, we note that we allow for positive masses of both compliers, i.e., units for whom $X(1) > X(0)$, and defiers, i.e., units for whom $X(0)>X(1)$. Again, this assumption could be added to our framework at the cost of additional notation, but we focus on the case where no restrictions are imposed on these potential treatments. We also do not require that $\ensuremath{\mathbb{P}}(X=1 \mid Z=1) \neq \ensuremath{\mathbb{P}}(X=1\mid Z=0)$, the usual relevance assumption assumed in the LATE framework.

We will maintain assumptions \ref{assn:NonTrivialinstrument} and \ref{assn:exog_or_excl}, and consider a continuum of exogeneity and exclusion assumptions that range from no assumptions to full exogeneity and exclusion. Thus, they will include the no-assumptions case and the case where $Z$ is excluded and exogenous as special cases.


\subsection{Sensitivity Models for the Exogeneity or Exclusion Assumptions}\label{subsec:partial_indep_assn}

We now consider a menu of assumptions that can be interpreted as relaxations of the exogeneity or exclusion assumption. The results of Proposition \ref{prop:noassn_discrete} show one extreme: bounds under no dependence assumptions. We briefly consider bounds under the other extreme, where exogeneity and weak exclusion exactly hold.  \cite{Manski1990} derived the identified set for $\mathbb{P}(Y(x) = 1)$ for $x \in \{0,1\}$ as well as the identified set for the ATE under this assumption.\footnote{Manski's \citeyearpar{Manski1990} analysis considered a general case which does not require outcomes, treatment, or instruments to be binary. In this general setting, he used a mean independence assumption. When outcomes are binary, mean independence of $Y(x)$ from $Z$ is equivalent to statistical independence of $Y(x)$ and $Z$.}

Under this assumption, $p_Y(x,0) = p_Y(x,1)$ for $x \in \{0,1\}$ by Lemma \ref{lemma:equivalence}. This restricts the probabilities $\textbf{p}_Y$ to lie in the set
\begin{align*}
 \mathcal{A}_\text{indep} &\coloneqq \{(p_{00},p_{01},p_{10},p_{11}) \in [0,1]^4 : p_{00} = p_{01}, p_{10} = p_{11}\}.
\end{align*}
Therefore, the identified set for $\textbf{p}_Y$ is given by the set of probabilities as restricted by the observed distribution of $(Y,X,Z)$, namely $\mathcal{H}_0 \times \mathcal{H}_1$, intersected with the set of probabilities satisfying exclusion and exogeneity, given by $\mathcal{A}_\text{indep}$. Thus, the identified set for $\textbf{p}_Y$ is
\begin{align}\label{eq:IDset_indep}
 (\mathcal{H}_0 \times \mathcal{H}_1) \cap \mathcal{A}_\text{indep},
\end{align}
which can also be written as
\begin{align*}
 &\bigcap_{z=0,1}[\ensuremath{\mathbb{P}}(Y=1,X=0 \mid Z=z), \ \ensuremath{\mathbb{P}}(Y=1,X=0 \mid Z=z) + \pi(1 \mid z)]\\
 &\qquad \times \bigcap_{z=0,1}[\ensuremath{\mathbb{P}}(Y=1,X=1 \mid Z=z), \ \ensuremath{\mathbb{P}}(Y=1,X=1 \mid Z=z) + \pi(0 \mid z)].
\end{align*}
These bounds take the form of intersections. \cite{Pearl1995} and \cite{BalkePearl1997} showed that this identified set can be empty, and hence that this model is falsifiable. This set is empty if and only if the two sets in \eqref{eq:IDset_indep} are disjoint. An empty identified set corresponds to a falsification of the original model, or of an exogeneity or exclusion assumption, when other model assumptions are maintained. Figure \ref{fig:IDset_indep} shows the identified set both when the model is falsified (left panel), and when it is not (right panel).

\begin{figure}[t]
\centering
\includegraphics[width=77mm]{images/newpics/pic4.pdf}
\hspace{5mm}
\includegraphics[width=77mm]{images/newpics/pic3.pdf}
\footnotesize
\caption{Left: Empty identified set under exogeneity and exclusion. Right: Nonempty identified set under exogeneity and exclusion. The upper and lower bounds for $p_Y(x,1) = p_Y(x,0)$ are denoted by $\overline{p}_Y(x)$ and $\underline{p}_Y(x)$}
\label{fig:IDset_indep}
\end{figure}

Full exogeneity or exclusion of the instrument may be a strong assumption in contexts where we do not believe that $Z$ is assigned randomly, or if we cannot rule out a direct effect of the instrument on the potential outcomes. In these cases, relaxing exogeneity or exclusion is appropriate. The no-assumption bounds of Manski remain valid, but partial validity of the instrument will yield intermediate bounds that are potentially significantly narrower than those in Proposition \ref{prop:noassn_discrete}.

We will consider relaxations of exogeneity and exclusion by characterizing sets of conditional probabilities $\textbf{p}_{Y}$. A large literature on sensitivity analysis has proposed various approaches for relaxing assumptions, often independence or conditional independence assumptions. We will focus on three examples, which are special cases of a unifying class of relaxations from independence we define in Section \ref{subsec:unifying_sensitivity_model_discrete}.

\subsubsection{Marginal Sensitivity Model: Tan (2006)}
The \textit{Marginal Sensitivity Model} (MSM) of \cite{Tan2006} consists of a class of relaxations of an independence assumption between a potential outcome and a binary treatment. It is generalized to multivariate treatments in \cite{ZhaoSmallBhattacharya2019} and \cite{BasitLatifWahed2023}. We consider a version of the MSM that constrains the dependence of the potential outcomes on the instruments, rather than the treatment.

\begin{definition}\label{def:MSModel}
Let $\Lambda \in [1,+\infty]$ be a known sensitivity parameter. The distribution of $(\{Y(x)\}_{x \in \operatorname*{supp}(X)},Z)$ satisfies the Marginal Sensitivity Model with parameter $\Lambda$ if
\begin{align}\label{eq:MSModel}
 \frac{\ensuremath{\mathbb{P}}(Z=z)}{\ensuremath{\mathbb{P}}(Z=z')} \Big/ \frac{\ensuremath{\mathbb{P}}(Z=z \mid Y(x) = y)}{\ensuremath{\mathbb{P}}(Z=z' \mid Y(x) = y)} \in \left[\Lambda^{-1},\Lambda\right]
\end{align}
for all  $x \in \operatorname*{supp}(X)$, $y \in \operatorname*{supp}(Y(x))$, and $z, z' \in \operatorname*{supp}(Z)$.\footnote{We let $\Lambda^{-1} = 0$ when $\Lambda = +\infty$.}
\end{definition}


With a binary instrument, this restriction places a bound on the odds ratio between the conditional odds of the instrument $\ensuremath{\mathbb{P}}(Z=z \mid Y(x) = y)/\ensuremath{\mathbb{P}}(Z=1-z \mid Y(x) = y)$ and its unconditional counterpart, $\ensuremath{\mathbb{P}}(Z=z)/\ensuremath{\mathbb{P}}(Z=1-z)$, for $z = 0,1$. In the binary outcome and instrument setting, equation \eqref{eq:MSModel} can be rearranged as
\begin{align}\label{eq:MSM_constraints}
 \frac{\mathbb{P}(Y(x) = y \mid Z = 1)}{\mathbb{P}(Y(x) = y \mid Z = 0)} = \frac{p_Y(x,1)^y(1-p_Y(x,1))^{1-y}}{p_Y(x,0)^y(1-p_Y(x,0))^{1-y}}  \in \left[\Lambda^{-1},\Lambda\right]
\end{align}
for $x,y \in \{0,1\}$. When $\Lambda = 1$, this ratio is 1 and $p_Y(x,1) = p_Y(x,0)$ for $x = 0,1$. By Lemma \ref{lemma:equivalence}, this means that weak exclusion \textit{and} exogeneity hold when $\Lambda = 1$ under Assumption \ref{assn:exog_or_excl}. When $\Lambda = +\infty$, these inequalities do not impose any restrictions on $\textbf{p}_Y$. Intermediate values of $\Lambda$ yield intermediate levels of restrictions on $\textbf{p}_Y$. Note that we can choose different $\Lambda$ values for $x = 0,1$, but we omit this generalization for brevity.

We note that equation \eqref{eq:MSM_constraints} can be written as four linear constraints on $\textbf{p}_Y$ by varying $y$ and $x$ over their support. This will be useful for casting this sensitivity analysis exercise as a linear program, as linear programming is a reliably fast and scalable computation method whose implementation is standard. Define
\begin{align*}
 A_\text{MSM}(\lambda) &\coloneqq \begin{pmatrix}
 1-\lambda & -1\\
 -1 & 1-\lambda\\
 \lambda - 1& 1\\
 1 & \lambda -1
 \end{pmatrix}, \qquad a_\text{MSM}(\lambda) \coloneqq \begin{pmatrix}
 0\\
 0\\
 \lambda\\
 \lambda
 \end{pmatrix}
\end{align*}
where $\lambda \coloneqq 1- \Lambda^{-1}$. The set of conditional probabilities satisfying the marginal sensitivity model with sensitivity parameter $\lambda$ is
\begin{align}\label{eq:MSM_discrete_correspondence}
 \mathcal{A}_\text{MSM}(\lambda) &\coloneqq \{\textbf{p} \in [0,1]^2: A_\text{MSM}(\lambda)\textbf{p} \leq a_\text{MSM}(\lambda)\}^2
\end{align}
where the weak inequality in \eqref{eq:MSM_discrete_correspondence} is component-wise. We reparametrized the sensitivity parameter $\Lambda$ as $\lambda = 1 - \Lambda^{-1}$ to standardize its scale to $[0,1]$. Here $\Lambda = 1$, or full exogeneity and exclusion, maps into $\lambda = 0$ while $\Lambda = +\infty$, or no assumptions, maps into $\lambda = 1$.




\subsubsection{$c$-dependence}
Introduced in \cite{MastenPoirier2018}, $c$-dependence imposes a bound on the maximum difference between the conditional probability of receiving a binary treatment $\ensuremath{\mathbb{P}}(X=1\mid Y(x))$ and its unconditional probability $\ensuremath{\mathbb{P}}(X=1)$. This was proposed in a setting where the unconfoundedness of treatment $X$ is relaxed. We adapt this sensitivity model to the case where exogeneity or exclusion of an instrument is relaxed. Here is the formal definition of this sensitivity model.
\begin{definition}\label{def:c-dep}
Let $c\in[0,1]$ be a known sensitivity parameter. The distribution of $(\{Y(x)\}_{x \in \operatorname*{supp}(X)},Z)$ satisfies \emph{$c$-dependence} if
\begin{equation}\label{eq:c-indep1}
 |\mathbb{P}(Z=z \mid Y(x) = y) - \ensuremath{\mathbb{P}}(Z=z) | \leq c
\end{equation}
for all $y \in \operatorname*{supp}(Y(x))$, $x \in \operatorname*{supp}(X)$, and $z \in \operatorname*{supp}(Z)$.
\end{definition}
When $Z$ is binary, it suffices to impose this inequality for $z=1$ only. When $c=0$, $c$-dependence is equivalent to imposing full exclusion and exogeneity. Values of $c$ exceeding $\max\{ p_Z,1-p_Z \}$ do not constrain the stochastic relationship between $Z$ and $Y(x)$. $c \in (0,1)$ partially constrains the stochastic relationship between $Z$ and $Y(x)$. \cite{MastenPoirier2023} give additional discussion of how to interpret $c$-dependence.


We can again rewrite the above restriction into a system of four linear restrictions on $\textbf{p}_Y$. Let
 \begin{align*}
 A_\text{$c$-dep}(c) &\coloneqq \begin{pmatrix}
  k_0(c) & -1\\
  -1 & k_1(c)\\
  -k_0(c) & 1\\
  1 & -k_1(c)
 \end{pmatrix} \qquad \text{ and } \qquad a_\text{$c$-dep}(c) \coloneqq \begin{pmatrix}
  0 \\
  0 \\
  1- k_0(c)\\
  1 - k_1(c)
 \end{pmatrix},
\end{align*}
where $k_z(c) \coloneqq \frac{\ensuremath{\mathbb{P}}(Z=z)\max\{\ensuremath{\mathbb{P}}(Z=1-z) - c,0\}}{\ensuremath{\mathbb{P}}(Z = 1-z)\min\{\ensuremath{\mathbb{P}}(Z=z)+c,1\}}$ for $z = 0,1$ and $c \in [0,1]$. We can show that the set of conditional probabilities consistent with $c$-dependence with sensitivity parameter $c$ is
\begin{align}\label{eq:cdep_discrete_correspondence}
 \mathcal{A}_\text{$c$-dep}(c) &\coloneqq \{\textbf{p} \in [0,1]^2: \textbf{A}_\text{$c$-dep}(c) \textbf{p} \leq a_\text{$c$-dep}(c)\}^2.
\end{align}
This set depends only on $c$ and $p_Z$.

\subsubsection{Kolmogorov-Smirnov Distance}
Consider a sensitivity model bounding a metric between the distributions of $Y(x) \mid \{Z=0\}$ and $Y(x) \mid \{Z=1\}$. This type of restriction was used in \cite{KlineSantos2013} to relax a missingness at random assumption. It was also considered for estimation in \cite{Manski1983}.

\begin{definition}\label{def:KSmodel}
Let $K \in [0,1]$ be a known sensitivity parameter. The distribution of $(\{Y(x)\}_{x \in \operatorname*{supp}(X)},Z)$ satisfies the Kolmogorov-Smirnov (KS) model if
\begin{align}\label{eq:KSdist}
 |\ensuremath{\mathbb{P}}(Y(x) \leq y\mid Z=z) - \ensuremath{\mathbb{P}}(Y(x) \leq y\mid Z=z')| \leq K
\end{align}
for all $x \in \operatorname*{supp}(X)$, $y \in \ensuremath{\mathbb{R}}$, and $z,z' \in \operatorname*{supp}(Z)$.
\end{definition}
When outcomes and instruments are binary, this sensitivity model is equivalent to bounding the magnitude of the difference between $\mathbb{P}(Y(x) = 1 \mid  Z = 1)$ and $\mathbb{P}(Y(x) = 1 \mid  Z = 0)$ by $K$. This assumption directly bounds the maximum deviation between the potential outcomes distribution given the instrument's two values. As in the previous two definitions, this class of restrictions encompasses independence ($K = 0$), no assumptions ($K = 1$), and intermediate cases ($K \in (0,1)$).

The set of conditional probabilities satisfying the Kolmogorov-Smirnov restrictions is characterized by the two linear inequalities
\begin{align}\label{eq:KS_discrete_correspondence}
 \mathcal{A}_\text{KS}(K) \coloneqq \{\textbf{p} \in [0,1]^2 : A_\text{KS} \textbf{p} \leq a_\text{KS}(K)\}^2
\end{align}
where
\begin{align}\label{eq:KS_Matrix}
 A_\text{KS} &\coloneqq \begin{pmatrix}
 1 & -1\\
 -1 & 1\\
 \end{pmatrix} \qquad \text{ and } \qquad a_\text{KS}(K) \coloneqq \begin{pmatrix}
 K\\
 K
 \end{pmatrix}.
\end{align}


\subsection{A Unifying Sensitivity Model}\label{subsec:unifying_sensitivity_model_discrete}

We now consider a general sensitivity model that encompasses the previous three sensitivity models as special cases. We will derive our main theoretical results under this sensitivity model. We assume that $X$ and $Z$ are binary for ease of notation and discuss the generalization to discrete $X$ and $Z$ in Section \ref{subsec:discrete_generalization}. In what follows, $\theta \in [0,1]$ is a sensitivity parameter that indexes relaxations of exogeneity or weak exclusion of the instrument.

\begin{het1Assump}[General Sensitivity Model]\label{assn:ind_relax}
 For a known sensitivity parameter $\theta \in [0,1]$, let
\begin{align*}
 \textbf{p}_Y \in \mathcal{A}_0(\theta) \times \mathcal{A}_1(\theta)
\end{align*}
where, for $x \in \{0,1\}$, $\mathcal{A}_x$ satisfies
\begin{enumerate}
 \item(Spanning) $\mathcal{A}_x(0) = \{a \in [0,1]^2: a_0 = a_1\}$ and $\mathcal{A}_x(1) = [0,1]^2$;
 \item(Monotonicity) $\mathcal{A}_x(\theta) \subseteq \mathcal{A}_x(\theta')$ when $\theta \leq \theta'$;
 \item(Linearity of Constraints)  $\mathcal{A}_x(\theta)$ is a closed convex polytope for each $\theta \in [0,1]$;
 \item(Continuity) The correspondence $\mathcal{A}_x: [0,1] \rightrightarrows [0,1]^2$ is continuous.
\end{enumerate}
\end{het1Assump}

The first part of this assumption implies that setting $\theta = 0$ imposes exogeneity and weak exclusion of the instrument, while setting $\theta = 1$ implies no restrictions on the dependence between $Z$ and the potential outcomes. The second part assumes these restrictions are monotonic in $\theta$, meaning that increasing $\theta$ yields a (weakly) larger set of conditional probabilities. These two parts combined yield that $\{\mathcal{A}_x(\theta):\theta \in [0,1]\}$ monotonically connects no assumptions to exogeneity and weak exclusion. The third restriction says that these sets are characterized by finitely many weak linear inequalities. This is crucial in obtaining a linear programming formulation for the bounds of various causal objects, such as the ATE.  The last part of this assumption assumes the continuity of the correspondence between the sensitivity parameter $\theta$ and the set of restricted conditional probabilities. Recall that a correspondence is continuous if it is both upper and lower hemicontinuous (uhc and lhc) at all points of its domain. See \cite{Border1985} for a compendium of results related to continuity of correspondences we make use of in our proofs. This assumption will yield continuity in the sensitivity parameter of the causal bounds obtained from linear programming.

This high-level assumption has useful properties, and all three previously considered relaxations are special cases of it. This is formalized in this proposition.
\begin{proposition}\label{prop:high-level_relax_discrete}
 Suppose Assumption \ref{assn:NonTrivialinstrument} holds. Relabeling $\lambda$, $c$, and $K$ as $\theta \in [0,1]$, the sets $\mathcal{A}_\text{MSM}(\lambda)$, $\mathcal{A}_{c\text{-dep}}(c)$, and $\mathcal{A}_\text{KS}(K)$ satisfy Assumption \ref{assn:ind_relax}.
\end{proposition}

Under this general relaxation, we will derive identified sets for various parameters of interest. We use these identified sets to characterize sharp bounds on causal objects using linear programming. We can also use them to determine what values of $\theta$ correspond to falsified models.

Before continuing our discussion, we present the identified set for conditional outcome probabilities under this general restriction.
\begin{theorem}\label{thm:prob_IDset}
 Suppose assumptions \ref{assn:NonTrivialinstrument}, \ref{assn:exog_or_excl}, and \ref{assn:ind_relax} hold. Then:
 \begin{enumerate}
  \item The identified set for $\textbf{p}_Y$ is
\begin{align}\label{eq:IDset_corr}
 \Pi(\theta) \coloneqq \Pi_0(\theta) \times \Pi_1(\theta),
\end{align}
where $\Pi_x(\theta) \coloneqq \mathcal{H}_x \cap \mathcal{A}_x(\theta)$;
  \item There exists $\underline{\theta}\in [0,1]$ such that $\Pi(\theta)$ is non-empty for $\theta \geq \underline{\theta}$ and empty for $\theta < \underline{\theta}$;
  \item For all $\theta \in [\underline{\theta},1]$, $\Pi(\theta)$ is a closed convex polytope in $[0,1]^4$;
  \item For $x \in \{0,1\}$, suppose $\text{int}(\mathcal{H}_x \cap \mathcal{A}_x(\theta)) \neq \emptyset$ for all $\theta > \underline{\theta}$. The correspondence $\Pi:[\underline{\theta},1] \rightrightarrows [0,1]^4$ defined in equation \eqref{eq:IDset_corr} is continuous.
 \end{enumerate}
\end{theorem}
This theorem has several implications. First, the identified set for the set of probabilities $\textbf{p}_Y$ is a Cartesian product of two sets. Each of these two sets is characterized as the intersection between a set containing all vectors $\textbf{p}_Y(x)$ consistent with the distribution of observables $F_{Y,X,Z}$, and the set of vectors consistent with a sensitivity model indexed by $\theta$.

The second implication is that the sensitivity model is falsified for an open, but potentially empty, subset of $[0,1]$. The minimum value at which the model is not falsified, called the \textit{falsification point} by \cite{MastenPoirier2021}, is $\underline{\theta}$ and is identified since it is a property of the sets $\Pi(\theta)$ for $\theta \in [0,1]$, all of which are known from the distribution of $(Y,X,Z)$. Moreover, the set of values for which the identified set is non-empty is closed, and always contains $\theta = 1$.

Third, this set is a closed, convex polytope, meaning it is defined by finitely many linear inequalities. This ensures that optimizing linear functions, such as $\Gamma_\text{ATE}$ and $\Gamma_\text{ATT}$ from \eqref{eq:ATE_functional} and \eqref{eq:ATT_functional}, can be performed using linear programming. This will be the key computational tool for implementing these methods.

Fourth, and finally, the mapping from $\theta$ into the identified set is continuous as a correspondence. This allows us to show the continuity in $\theta$ of extrema of continuous functionals of $\textbf{p}_Y$ over the identified set, again including the ATE and ATT.

We now illustrate the identified set for a sensitivity model corresponding to $c$-dependence. The shaded boxes in Figure \ref{fig:FFbinaryYsingleIV} show examples of the no-assumption bounds $\mathcal{H}_0 \times \mathcal{H}_1$.

\begin{figure}[t]
\centering
\includegraphics[width=77mm]{images/newpics/pic5.pdf}
\hspace{5mm}
\includegraphics[width=77mm]{images/newpics/pic6.pdf}\\
\includegraphics[width=77mm]{images/newpics/pic7.pdf}
\hspace{5mm}
\includegraphics[width=77mm]{images/newpics/pic8.pdf}
\footnotesize
\caption{Example identified sets for different sensitivity parameter values. Relaxation is $c$-dependence. Top Left: $\theta = 0$. Top Right: $\theta = \underline{\theta}_x$, the falsification point. Bottom left: $\theta > \underline{\theta}_x$. Bottom right: $\theta$ sufficiently large that the identified set equals the no-assumption identified set from Proposition \ref{prop:noassn_discrete}. }
\label{fig:IDset_sensparam}
\end{figure}

The set $\mathcal{A}_x(\theta)$ is a parallelogram imposing the $c$-dependence constraint. The identified set for $\textbf{p}_Y(x)$ is given by the intersection of the parallelogram and shaded box.


While the no-assumption bounds are never empty, the bounds under exogeneity and weak exclusion ($\theta=0$) can be empty, and hence the baseline statistical independence assumption can be falsified. This happens when, for some $x \in \{0,1\}$, the no assumption bounds $\mathcal{H}_x$ have an empty intersection with the statistical independence constraint set $\mathcal{A}_x(\theta)$. Graphically, this happens when the box defined by the no assumption bounds does not intersect the 45-degree line. This is shown in the first plot Figure \ref{fig:IDset_sensparam}. The falsification point is simply the smallest value of $\theta$ such that the parallelogram defined by $\mathcal{A}_x(\theta)$ has a nonempty intersection with the no assumption bounds $\mathcal{H}_x$ for each $x \in \{0,1\}$. This intersection is illustrated in the second plot of Figure \ref{fig:IDset_sensparam}. Increasing the sensitivity parameter increases the size of this intersection (see the third plot of Figure \ref{fig:IDset_sensparam}), until the intersection equals $\mathcal{H}_x$, the no assumption bounds, which can be seen in the fourth plot of Figure \ref{fig:IDset_sensparam}.

We next show how to use Theorem \ref{thm:prob_IDset} to get identified sets for  counterfactual probabilities $\ensuremath{\mathbb{P}}(Y(x)=1)$ and for the ATE. By the law of total probability,
\[
 \ensuremath{\mathbb{P}}(Y(x)=1) = p_Y(x,0) (1-p_Z) + p_Y(x,1) p_Z.
\]
The weight $p_Z$ is identified, while the identified set for $\textbf{p}_Y(x)$ is given by $\Pi_x(\theta)$. Thus, we can simply minimize and maximize the above convex combination over this set to obtain the identified set for $\ensuremath{\mathbb{P}}(Y(x)=1)$. Hence we define
\begin{align*}
	\overline{P}_x(\theta) &\coloneqq \sup_{(a_0,a_1) \in \Pi_x(\theta)} \big( a_0 (1 - p_Z) + a_1 p_Z \big) \qquad \text{and} \qquad \underline{P}_x(\theta) &\coloneqq \inf_{(a_0,a_1) \in \Pi_x(\theta)} \big( a_0 (1 - p_Z) + a_1 p_Z \big).
\end{align*}
These are both finite-dimensional linear programs and hence can be computed easily given estimates of the joint distribution of $(Y,X,Z)$. Figure \ref{fig:IDset_sensparam_optima} illustrates the minimization/maximization of a linear functional over the identified set $\mathcal{A}_x(\theta) \cap \mathcal{H}_x$.

\begin{figure}[t]
\centering
\includegraphics[width=105mm]{images/newpics/pic9.pdf}
\footnotesize
\caption{Example identified set and minimization/maximization of the ATE under a sensitivity model.}
\label{fig:IDset_sensparam_optima}
\end{figure}

The following corollary lists properties of these bounds.
\begin{corollary}\label{corr:hetTrtBinATEbounds}
Suppose assumptions \ref{assn:NonTrivialinstrument}, \ref{assn:exog_or_excl}, and \ref{assn:ind_relax} hold. Then, for $x \in \{0,1\}$:
\begin{enumerate}
\item The identified set for $(\ensuremath{\mathbb{P}}(Y(0)=1), \ensuremath{\mathbb{P}}(Y(1)=1))$ is $I_0(\theta) \times I_1(\theta)$ where $I_x \coloneqq [\underline{P}_x(\theta), \overline{P}_x(\theta)]$ when $\theta\in [\underline{\theta},1]$, and the empty set when $\theta < \underline{\theta}$;

\item The functions $\underline{P}_x(\theta)$ and $\overline{P}_x(\theta)$ are continuous and monotonic over $\theta \in [\underline{\theta},1]$;

\item Let $\theta \in [\underline{\theta},1]$. The identified set for ATE is $[\underline{\text{ATE}}(\theta),\overline{\text{ATE}}(\theta)]$ where
\[
 \underline{\text{ATE}}(\theta) \coloneqq \underline{P}_1(\theta) - \overline{P}_0(\theta)
 \qquad \text{and} \qquad
 \overline{\text{ATE}}(\theta) \coloneqq \overline{P}_1(\theta) - \underline{P}_0(\theta).
\]
\end{enumerate}
\end{corollary}


 This discussion implies that ATE will typically be partially identified at the falsification point. That is: The falsification adaptive set for ATE, $[\underline{\text{ATE}}(\underline\theta), \overline{\text{ATE}}(\underline\theta)]$, will generally be an interval with a nonempty interior.




\subsection{Generalization to Non-Binary Discrete Variables}\label{subsec:discrete_generalization}
The previous results illustrate that common sensitivity models yield identified sets for parameters of interest with desirable properties. However, these were illustrated only for cases where $Y$, $X$, and $Z$ were all binary. In practice, many empirical settings have multiple instruments, and treatments or outcomes may be multivalued as well. In this section, we sketch a generalization of the previous results to cases where the support of $Y$, $X$, may be discrete instead of binary, where there may be multiple instruments, and where each instrument may possess finite support rather than being binary.

Let $X$ be a discrete treatment, let $Z$ be a vector of instruments, where each instrument is discrete, and let $Y(x,z)$ be discrete as well. We let $Y = Y(X,Z)$ be the realized, observed outcome. We suppose all their supports are finite.

Let $p_Y(y \mid z ; \ x) \coloneqq \ensuremath{\mathbb{P}}(Y(x) = y\mid Z = z)$, $\textbf{p}_{Y}(x,z) \coloneqq \{p_Y(y \mid z ; \ x)\}_{y \in \operatorname*{supp}(Y(x))}$, $\textbf{p}_{Y}(x) \coloneqq \{\textbf{p}_{Y}(x,z)\}_{z \in \operatorname*{supp}(Z)}$, and $\textbf{p}_{Y} \coloneqq \{\textbf{p}_{Y}(x)\}_{x \in \operatorname*{supp}(X)}$. The vector $\textbf{p}_{Y}$ contains the full distribution of $Y(x) \mid \{Z = z\}$ for all $(x,z) \in \operatorname*{supp}(X,Z)$. We define $s_Z \coloneqq |\operatorname*{supp}(Z)|$, $s_{Y(x)} \coloneqq |\operatorname*{supp}(Y(x))|$, and $\operatorname*{supp}(Y(x)) \coloneqq \{y_1,\ldots,y_{s_{Y(x)}}\}$.

To avoid trivial cases, we make the following assumption.
\begin{het1Assump}\label{assn:NonTrivialinstrument_disc}
For all $(x, z) \in \operatorname*{supp}(X) \times\operatorname*{supp}(Z)$, $\mathbb{P}(Z = z) \in (0,1)$ and $\ensuremath{\mathbb{P}}(X=x \mid Z =z) \in (0,1)$.
\end{het1Assump}

Let $\Delta_{S}$ denote the simplex of dimension $S$:
\begin{align*}
 \Delta_{S} &\coloneqq \left\{\textbf{p} \in \ensuremath{\mathbb{R}}^{S+1}: \textbf{p} \geq 0, \sum_{s = 1}^{S+1} p_s = 1\right\}.
\end{align*}
For $K \in \mathbb{N}$, let $\Delta_S^K$ denote the $K$-fold cartesian product of $\Delta_S$.

The no-assumption identified set for $\textbf{p}_Y$ is given by
\begin{align*}
 \prod_{x \in \operatorname*{supp}(X)} \prod_{z \in \operatorname*{supp}(Z)}  \mathcal{H}_{x,z},
\end{align*}
where
\begin{align*}
 \mathcal{H}_{x,z} &= \{ (p_1,\ldots,p_{s_{Y(x)}}) \in \Delta_{s_{Y}(x) - 1}: p_s \in [\ensuremath{\mathbb{P}}(Y=y_s,X=x\mid Z = z),\ensuremath{\mathbb{P}}(Y=y_s,X=x\mid Z = z) \\
 &\qquad + \ensuremath{\mathbb{P}}(X \neq x\mid Z = z)], \text{for all } s \in \{1,\ldots,s_{Y(x)}\}\}.
\end{align*}
The set $\mathcal{H}_{x,z}$ is the identified set for $\textbf{p}_Y(x,z)$ under no assumptions. We also note the similar structure of $\mathcal{H}_x = \prod_{z \in \operatorname*{supp}(Z)}  \mathcal{H}_{x,z}$  and of the rectangles defined in \eqref{eq:boxSet} for the binary case.


The three sensitivity models we investigated earlier can be defined independently of the supports of the potential outcomes, treatments, or instruments, so they can be used when these variables are non-binary. We can also embed these sensitivity models in a general sensitivity model similar to the one in Assumption \ref{assn:ind_relax}. The following assumption simplifies to \ref{assn:ind_relax} when all variables are binary.

\begin{het1Assump}[General Sensitivity Model]\label{assn:ind_relax_discrete}
 Suppose Assumption \ref{assn:NonTrivialinstrument_disc} holds. For a known sensitivity parameter $\theta \in [0,1]$, let
\begin{align*}
 \textbf{p}_Y \in \prod_{x \in \operatorname*{supp}(X)} \mathcal{A}(\theta; x)
\end{align*}
where, for $x \in \operatorname*{supp}(X)$, $\mathcal{A}(\theta; x)$ satisfies
\begin{enumerate}
 \item(Spanning) $\mathcal{A}(0; x) = \{ (a_1,\ldots,a_{s_Z}) \in \Delta_{s_Y(x) - 1}^{s_Z} : a_1 = \cdots = a_{s_Z}\}$ and $\mathcal{A}(1; x) = \Delta_{s_Y(x) - 1}^{s_Z}$;
 \item(Monotonicity) $\mathcal{A}(\theta; x) \subseteq \mathcal{A}(\theta'; x)$ when $\theta \leq \theta'$;
 \item(Linearity of Constraints)  $\mathcal{A}(\theta; x)$ is a closed convex polytope for each $\theta \in [0,1]$;
 \item(Continuity) The correspondence $\mathcal{A}(\cdot; x): [0,1] \rightrightarrows \Delta_{s_Y(x) - 1}^{s_Z}$ is continuous.
\end{enumerate}
\end{het1Assump}
This assumption is similar to its counterpart with binary variables, except for parts 1 and 4, which have been modified to allow $Y(x)$ to be nonbinary. The restriction in part 1 states that $\mathbb{P}(Y(x) = y \mid  Z = z)$ is constant in $z \in \operatorname*{supp}(Z)$ for each $y \in \operatorname*{supp}(Y(x))$, and is stated as equality constraints for on the components of $\mathcal{A}_x(0)$.

As in the binary case, all these assumptions can be written as linear inequalities in the components of vector $\textbf{p}_Y$. Therefore, the bounds on various causal objects can be obtained by solving linear programs. We expect similar results to Theorem \ref{thm:prob_IDset} and Corollary \ref{corr:hetTrtBinATEbounds} to hold in this setting, so that the bounds enjoy the same monotonicity and continuity property.


\section{Identification with Continuous Outcomes}\label{sec:cont_Y}

We now consider cases where the outcome variable is continuously distributed. In this case, we may view the corresponding problem as an infinite dimensional program, whose theoretical properties are harder to analyze. Nevertheless, in this section we show that the previous sensitivity models can be used with continuous outcomes, and we obtain theoretical properties of the corresponding sensitivity analyses for the exogeneity/exclusion of an instrument. To keep other aspects of the problem relatively simple, we consider the case where the treatment and instrument are both binary, although this can be naturally generalized as in Section \ref{subsec:discrete_generalization}. In this section, we show that the analytical results we derived under binary outcomes generalize to continuous outcomes. This leads us to a relatively simple and feasible approach for computing identified sets under relaxations of instrument exogeneity with continuous outcomes.

We begin by assuming that outcomes are continuously distributed.


\begin{het1Assump}
\label{assn:contsupport}
Suppose that $\operatorname*{supp}(Z) = \operatorname*{supp}(X) = \{0,1\}$. For any $x, x',z \in \{0,1\}$ the distribution of $Y(x) \mid \{X=x', Z=z\}$ is continuous with respect to the Lebesgue measure and is supported on a compact interval $\mathcal{Y}_x \coloneqq \operatorname*{supp}(Y(x))$, which is independent of $x'$ and $z$.
\end{het1Assump}

Assumption \ref{assn:contsupport} supposes that, conditional on the treatment and instruments, potential outcomes are continuously distributed. It implies that, conditional on the treatment and instruments, observed outcomes are also continuously distributed. We can allow for discrete instruments as in Section \ref{sec:binary_outcome}, but we only consider a binary instrument to simplify the notation.

This assumption also states that the conditional support of $Y(x)$ given $(X,Z) = (x',z)$ does not depend on $(x',z)$, which is made for convenience. Our results would remain valid without this restriction, but notation in the proofs would have to be heavier.

Let $f_{Y}(y \mid z; \ x) \coloneqq f_{Y(x) \mid Z}(y \mid z)$ denote the conditional density of $Y(x)$ given $Z = z$. We also let $\textbf{f}_{Y}(y ; \ x) \coloneqq (f_{Y}(y \mid 0 ; \ x),f_{Y}(y \mid 1 ; \ x))$ and  $\textbf{f}_{Y}(y) \coloneqq (\textbf{f}_{Y}(y ; \ 0), \textbf{f}_{Y}(y ; \ 1))$ denote collections of these densities across instrument and treatment values.  We assume that the potential outcomes' densities belong to a convex class of densities that is compact with respect to the supremum norm.
\begin{het1Assump}\label{assn:compactdensities}
For $x,z\in\{0,1\}$, let
\begin{align*}
	f_{Y}(\cdot \mid z ; \ x) &\in \left\{f \in \mathcal{F}_x(\mathcal{Y}_x): \int_{\mathcal{Y}_x}f(y) \, dy = 1, f \geq 0\right\} \eqqcolon \mathcal{F}_{\text{den},x},
\end{align*}
where $\mathcal{F}_x(\mathcal{Y}_x)$ is a convex set of bounded functions supported on $\mathcal{Y}_x$ that is compact with respect to the norm $\|f\|_\infty \coloneqq \sup_{y \in \mathcal{Y}_x}|f(y)|$.
\end{het1Assump}

Examples of compact sets $\mathcal{F}_x(\mathcal{Y}_x)$  include the set of bounded Lipschitz functions:
\begin{align*}
	\mathcal{C}_{0,\infty,1,1}(\mathcal{Y}_x) \coloneqq \left\{f \in \mathcal{C}_0(\mathcal{Y}_x): \sup_{y \in \mathcal{Y}_x}|f(y)| + \sup_{y,y' \in \text{int}(\mathcal{Y}_x), y\neq y'} \frac{|f(y') - f(y)|}{|y' - y|} <M \right\}
\end{align*}
where $\mathcal{C}_0(A)$ denotes the set of continuous functions on domain $A$, and $M<\infty$ is a constant. See \cite{FreybergerMasten2019} for alternative compact sets of functions and associated discussion.

We start by deriving the no-assumptions bounds for this set of conditional densities.
\begin{proposition}\label{prop:cont_noassbounds}
Suppose assumptions \ref{assn:NonTrivialinstrument}, \ref{assn:contsupport}, and \ref{assn:compactdensities} hold. The identified set for $\textbf{f}_{Y}$ is
\begin{align*}
	\mathcal{H} \coloneqq \prod_{x = 0,1} \mathcal{H}_{x}
\end{align*} where $\mathcal{H}_{x} \coloneqq \prod_{z = 0,1} \mathcal{H}_{x,z}$ and $\mathcal{H}_{x,z} \coloneqq \{f(\cdot) \in \mathcal{F}_{\text{den},x}: f \geq f_{Y|X,Z}(\cdot\mid x,z)\pi(x\mid z)\}$.
\end{proposition}

We next consider the baseline case where the instruments are exogenous and excluded. In this case, the instrument's validity implies that the densities $\textbf{f}_Y$ must lie in
\begin{align*}
	\mathcal{A}_\text{indep} \coloneqq \{(f_{00},f_{01},f_{10},f_{11}) \in \mathcal{F}_0(\mathcal{Y}_0)^2 \times \mathcal{F}_1(\mathcal{Y}_1)^2 : f_{00} = f_{01}, f_{10} = f_{11}\},
\end{align*}
since this set imposes that $f_{Y(x)|Z}(\cdot\mid 0) = f_{Y(x)|Z}(\cdot\mid 1)$ for $x = 0,1$. Thus, the identified set for $\textbf{f}_Y$ under independence is given by
\begin{align*}
	\mathcal{H} \cap \mathcal{A}_\text{indep}.
\end{align*}
This is precisely the setting studied in \cite{Kitagawa2021}, and he provides a characterization of this set in his Proposition 3.1, which we include without proof.
\begin{proposition}\label{prop:kitagawa3point1} (Proposition 3.1 in \cite{Kitagawa2021}) Suppose assumptions \ref{assn:NonTrivialinstrument}, \ref{assn:contsupport}, and \ref{assn:compactdensities} hold. Suppose $Z$ is exogenous and weakly excluded. Then the identified set for $(f_{Y(0)},f_{Y(1)})$ is
\[
	 \left\{ f_0 \in \mathcal{F}_{\text{den},0}: f_0(\cdot) \geq \max_{z=0,1} \; f_{Y|X,Z}(\cdot\mid  0 , z) \pi(0\mid z)\right\} \times \left\{ f_1 \in \mathcal{F}_{\text{den},1}: f_1(\cdot) \geq \max_{z=0,1} \; f_{Y|X,Z}(\cdot \mid 1, z)\pi(1\mid z)\right\}.
\]
Consequently, the model is refuted if
\[
	\int_{\mathcal{Y}_x} \max_{z=0,1} \; f_{Y,X|Z}(y,x \mid z) \; dy > 1
\]
for some $x \in \{0,1 \}$.
\end{proposition}

The previous two results establish the identification region for conditional densities of $Y(x)$ given $Z$ under no-assumptions, and under the full validity of the instrument, which correspond to the ends of a spectrum of assumptions about the dependence between $Y(x)$ and $Z$. We now consider sensitivity models that consider intermediate assumptions on the instrument's validity. We again consider the following three restrictions, which are adapted from Section \ref{subsec:partial_indep_assn}.

\subsubsection*{Marginal Sensitivity Model}
Consider the Marginal Sensitivity Model of definition \ref{def:MSModel}. When the outcome is continuously distributed, Bayes' rule allows us to rewrite equation \eqref{eq:MSModel} as a density ratio:
\begin{align*}
	 \frac{f_{Y}(\cdot\mid z; \ x)}{f_{Y}(\cdot\mid z'; \ x)} \in \left[\Lambda^{-1},\Lambda\right]
\end{align*}
for $(x,z,z') \in \operatorname*{supp}(X) \times \operatorname*{supp}(Z)^2$. As in the previous sections, we reparametrize $\Lambda$ as $\lambda = 1 - \Lambda^{-1} \in [0,1]$. The set of densities satisfying this restriction can be viewed as a set of functions satisfying linear inequality constraints. Specifically, we can write the set of restricted densities as
\begin{align}\label{eq:MSM_cont}
	\mathcal{A}_\text{MSM}(\lambda) &= \mathcal{A}_\text{MSM}(\lambda;0) \times \mathcal{A}_\text{MSM}(\lambda;1)
\end{align}
where $\mathcal{A}_\text{MSM}(\lambda;x) \coloneqq \{\textbf{f} \in \mathcal{F}_{\text{den},x}^2: A_\text{MSM}(\lambda)\textbf{f} \leq \textbf{0}\}$ and
\begin{align*}
	A_\text{MSM}(\lambda) \coloneqq \begin{pmatrix}
		-1 & 1-\lambda\\
		1-\lambda & -1
	\end{pmatrix}.
\end{align*}
Inequalities involving functions $\textbf{f}$ are meant to hold across all $y \in \ensuremath{\mathbb{R}}$.


\subsubsection*{$c$-dependence}
As defined in equation \eqref{eq:c-indep1}, $c$-dependence is collection of inequalities across values of $y$. Again using Bayes' Rule, we can rewrite these inequalities using conditional densities of $Y(x)$ given the instrument:
\begin{align*}
	-\min\{p_Z + c,1\}(1-p_Z)f_{Y}(y\mid 0;x) + (1 - \min\{p_Z + c,1\})p_Z f_{Y}(y\mid 1;x) &\leq 0\\
	\max\{p_Z - c,0\}(1-p_Z)f_{Y}(y\mid 0;x) + (\max\{p_Z - c,0\} - 1)p_Z f_{Y}(y\mid 1;x) &\leq 0.
\end{align*}

These are densities restricted by linear inequalities that depend on the observed variables only through the marginal distribution of the instrument. The set of densities as restricted by $c$-dependence is given by
\begin{align}\label{eq:cdep_cont}
	\mathcal{A}_\text{$c$-dep}(c) &\coloneqq \mathcal{A}_\text{$c$-dep}(c;0) \times \mathcal{A}_\text{$c$-dep}(c;1)
\end{align}
where $\mathcal{A}_\text{$c$-dep}(c;x) \coloneqq \{\textbf{f} \in \mathcal{F}_{\text{den},x}^2: A_\text{$c$-dep}(c)\textbf{f} \leq \textbf{0}\}$ and
\begin{align}\label{eq:cdep_cont_matrix}
		A_\text{$c$-dep}(c) &\coloneqq \begin{pmatrix}
		-1 & k_1(c)\\
		k_0(c) & -1
	\end{pmatrix}.
\end{align}
We can see that setting $c = 0$ implies that $k_0(c) = k_1(c) = 1$, which mechanically imposes that the conditional densities $f_{Y(x)|Z}(\cdot\mid 0)$ and $f_{Y(x)|Z}(\cdot\mid 1)$ are equal. As a result, we can verify that $c=0$ implies independence of potential outcomes and the instrument, as it does when the outcome is discrete.

\subsubsection*{Supremum Distance}
Using the Kolmogorov-Smirnov as a starting point, we consider a sensitivity model that bounds the supremum distance between densities rather than distribution functions. Hence, we assume that
\begin{align*}
	\sup_{y \in \ensuremath{\mathbb{R}}}|f_Y(y\mid 0;x) - f_{Y}(y\mid 1;x)| \leq K/(1-K)
\end{align*}
for $x \in \operatorname*{supp}(X)$, for some known $K$ satisfying $K \in [0,1]$.\footnote{We let $K/(1-K) = +\infty$ when $K = 1$.} The sensitivity parameter $K$ bounds the difference between density functions, and we used the strictly increasing mapping $a\mapsto 1/(1-a)$ to span the continuum between independence and no restrictions, as $K=0$ maps to exact equality of densities, and $K=1$ does not impose any restrictions on the dependence of the distribution of $Y(x)$ in $Z$. An alternate mapping from $[0,1]$ to $[0,+\infty]$ could be used instead.

The set of densities as restricted by this sup distance is given by
\begin{align}\label{eq:KS_cont}
	\mathcal{A}_\text{KS}(K) &\coloneqq \mathcal{A}_\text{KS}(K;0) \times \mathcal{A}_\text{KS}(K;1)
\end{align}
where $\mathcal{A}_\text{KS}(K;x) \coloneqq \{\textbf{f} \in \mathcal{F}_{\text{den},x}^2: A_\text{KS}\textbf{f} \leq (K/(1-K), K/(1-K))^\top\}$, where $A_\text{KS}$ is defined as in equation \eqref{eq:KS_Matrix}.

\subsection{A Unifying Sensitivity Model with Continuous Outcomes}

As in Section \ref{sec:binary_outcome}, all these relaxations can be viewed as special cases of a unifying class of relaxations encoding various types of departures from independence.

\begin{het1Assump}[General Sensitivity Model with Continuous Outcomes]\label{assn:ind_relax_cont}
	For a known sensitivity parameter $\theta \in [0,1]$, suppose
\begin{align*}
	\textbf{f}_Y \in \mathcal{A}_0(\theta) \times \mathcal{A}_1(\theta)
\end{align*}
where, for $x \in \{0,1\}$, $\mathcal{A}_x$ satisfies
\begin{enumerate}
	\item(Spanning) $\mathcal{A}_x(0) = \{(f_0,f_1) \in \mathcal{F}_{\text{den},x}^2: f_0 = f_1\}$ and $\mathcal{A}_x(1) = \mathcal{F}_{\text{den},x}^2$;
	\item(Monotonicity) $\mathcal{A}_x(\theta) \subseteq \mathcal{A}_x(\theta')$ when $\theta \leq \theta'$;
	\item(Linearity of Constraints)  The set $\mathcal{A}_x(\theta)$ is a closed convex subset of $\mathcal{F}_{\text{den},x}^2$ characterized by finitely many componentwise weak linear inequalities in the densities for each $\theta \in [0,1]$;
	\item(Continuity) The correspondence $\mathcal{A}_x: [0,1] \rightrightarrows \mathcal{F}_{\text{den},x}^2$ is continuous with respect to the sup-norm.
\end{enumerate}
\end{het1Assump}
The constraint set $\mathcal{A}_0(\theta) \times \mathcal{A}_1(\theta)$ is a convex set of functions defined by linear inequalities that weakly expands as $\theta$ increases. It nests the identified set under the baseline independence assumption ($\theta=0$) and the identified set under no assumptions on the dependence between potential outcomes and instruments ($\theta=1$). The third requirement is that the constraint set is of the form $\mathcal{A}_x(\theta) = \{\textbf{f} = (f_0,f_1) \in \mathcal{F}_{\text{den},x}^2: A(\theta)\textbf{f} \leq a(\theta)\}$ where $A(\theta)$ is a finite dimensional matrix. It involves finitely many componentwise weak inequalities, even though the inequality $A(\theta)\textbf{f} \leq a(\theta)$ hold for infinitely many values on the support $\mathcal{Y}_x$.   As in the binary outcome case, this relaxation encompasses the previous three restrictions.

\begin{proposition}\label{prop:high-level_relax_cont}
	Suppose assumptions \ref{assn:NonTrivialinstrument}, \ref{assn:contsupport}, and \ref{assn:compactdensities} hold. Relabeling $(\lambda,K,c)$ as $\theta \in [0,1]$, the restrictions from equations \eqref{eq:MSM_cont}, \eqref{eq:cdep_cont}, and \eqref{eq:KS_cont} all satisfy Assumption \ref{assn:ind_relax_cont}.
\end{proposition}

We now state our main result about theoretical properties of the identified set for densities of the potential outcomes.
\begin{theorem}\label{thm:cont_prob_IDset}
	Suppose assumptions \ref{assn:NonTrivialinstrument}, \ref{assn:contsupport}, \ref{assn:compactdensities}, and \ref{assn:ind_relax_cont} hold, and suppose that $\mathcal{F}_{\text{den},x}^2 \cap \mathcal{H}_x \neq \emptyset$ for $x \in \{0,1\}$. Then,
	\begin{enumerate}
		\item The identified set for $\textbf{f}_Y$ is
\begin{align}\label{eq:cont_IDset_corr}
	\Pi(\theta) \coloneqq \Pi_0(\theta) \times \Pi_1(\theta),
\end{align}
where
\begin{align*}
	\Pi_x(\theta) &\coloneqq \mathcal{H}_x \cap \mathcal{A}_x(\theta);
\end{align*}
		\item There exists $\underline{\theta}\in [0,1]$ such that $\Pi(\theta)$ is non-empty for $\theta \geq \underline{\theta}$ and empty for $\theta < \underline{\theta}$;
		\item Suppose $\text{int}(\mathcal{H}_x \cap \mathcal{A}_x(\theta)) \neq \emptyset$ for all $\theta > \underline{\theta}$. Then, the correspondence $\Pi:[\underline{\theta},1] \rightrightarrows \mathcal{F}_{\text{den},0}^2 \times \mathcal{F}_{\text{den},1}^2$ defined by $\Pi(\theta)$ in equation \eqref{eq:cont_IDset_corr} is continuous.
	\end{enumerate}
\end{theorem}

This theorem establishes the main theoretical properties of the identified sets for densities, including their continuity as an infinite-dimensional correspondence. This continuity will carry over to functionals of these densities, in particular to linear or continuous functionals.

In particular, consider the class of linear mappings, for which the sharp bounds can be obtained as the solution to a linear program.
Let
\[
	\Gamma(\textbf{f})
	\coloneqq \int_{\mathcal{Y}_0} \omega_0(y)'\textbf{f}_Y(y ; \ 0)dy + \int_{\mathcal{Y}_1} \omega_1(y)'\textbf{f}_Y(y ; \ 1)dy
\]
where, for $x = 0,1$, $\omega_x$ is a known weight function that maps $\ensuremath{\mathbb{R}}$ to $\ensuremath{\mathbb{R}}^2$. The $\Gamma$ mapping is used to characterize a functional of the conditional densities of $Y(x)\mid Z$.

For example, with $\omega_1(y) = -\omega_0(y) = (y (1-p_Z), y p_Z)$, we have that
\begin{align*}
	\Gamma(\textbf{f}_Y) &= \int_{\mathcal{Y}_1} y \left(p_Zf_{Y}(y\mid 1; \ 1) + (1-p_Z)f_{Y}(y\mid 0 ; \ 1)\right) dy\\
	&\qquad - \int_{\mathcal{Y}_0} y \left(p_Z f_{Y}(y\mid 1; \ 0) + (1-p_Z)f_{Y}(y\mid 0 ; \ 0)\right) dy\\
	&= \int_{\mathcal{Y}_1} y f_{Y(1)}(y) dy - \int_{\mathcal{Y}_0} y f_{Y(0)}(y) dy\\
	&= \ensuremath{\mathbb{E}}[Y(1)] - \ensuremath{\mathbb{E}}[Y(0)],
\end{align*}
the average treatment effect. Letting $\omega_x(y) = \ensuremath{\mathbbm{1}}(y \leq a)$ and $\omega_{1-x}(y) = 0$ yields $\Gamma(\textbf{f}_Y) = \mathbb{P}(Y(x) \leq a)$, the cumulative distribution function evaluated at $a \in \ensuremath{\mathbb{R}}$.. This choice can be used to obtain bounds on quantiles of $Y(x)$ or on the quantile treatment effect $\text{QTE}(\tau) \coloneqq Q_{Y(1)}(\tau) - Q_{Y(0)}(\tau)$ for a quantile index $\tau \in (0,1)$.

The proposition below shows that bounds on these functionals are continuous and monotonic. This result uses the Maximum Theorem \citep{Berge1959} applied to an infinite-dimensional correspondence.
Let
\[
	\overline{\Gamma}(\theta)
	\coloneqq \sup_{\textbf{f}_1 \in \Pi_1(\theta)} \int_{\mathcal{Y}_1} \omega_1(y_1)'\textbf{f}_1(y_1) \; dy_1
	+ \sup_{\textbf{f}_0 \in \Pi_0(\theta)}\int_{\mathcal{Y}_0} \omega_0(y_0)'\textbf{f}_0(y_0) \; dy_0
\]
and
\[
	\underline{\Gamma}(\theta)
	\coloneqq \inf_{\textbf{f}_1 \in \Pi_1(\theta)} \int_{\mathcal{Y}_1} \omega_1(y_1)'\textbf{f}_1(y_1) \; dy_1
	+ \inf_{\textbf{f}_0 \in \Pi_0(\theta)}\int_{\mathcal{Y}_0} \omega_0(y_0)'\textbf{f}_0(y_0) \; dy_0
\]
denote the lower and upper bounds of the functional over the sets $\Pi_x(\theta)$, $x=0,1$.
\begin{corollary}\label{corr:cont_functionalbounds}
Suppose the assumptions of Theorem \ref{thm:cont_prob_IDset} hold. Let $\|\omega_x(\cdot)\|_\infty < \infty$. Then,
\begin{enumerate}
\item Let $\theta\in [\underline{\theta},1]$. The identified set for $(f_{Y(0)}, f_{Y(1)})$ is $I_0(\theta) \times I_1(\theta)$ where $I_x(\theta) \coloneqq \{f_0 (1-p_Z) + f_1 p_Z: (f_0,f_1) \in \Pi_x(\theta)\}$ when $\theta\in [\underline{\theta},1]$, and the empty set when $\theta < \underline{\theta}$;

\item The functions $\underline{\Gamma}(\theta)$ and $\overline{\Gamma}(\theta)$ are continuous and monotonic over $\theta \in [\underline{\theta},1]$.

\item Let $\theta \in [\underline{\theta},1]$. The identified set for $\Gamma(\textbf{f}_{Y})$ is $[\underline{\Gamma}(\theta), \overline{\Gamma}(\theta)]$.
\end{enumerate}
\end{corollary}

Therefore, as in the discrete case, bounds a can be obtained in the continuous case through infinite-dimensional linear programming. To make this approach feasible, we show in the next section how to convert an infinite-dimensional linear program into a feasible, finite-dimensional linear program that can be directly implemented.

\subsection{Computation}\label{sec:comp}

The identified set $\Pi(\theta)$ is an infinite-dimensional set of continuous densities. If we restrict attention to the class of linear functionals described in Corollary \ref{corr:cont_functionalbounds}, the corresponding identified set is an interval (or the empty set). However, Corollary \ref{corr:cont_functionalbounds} characterizes this interval by optimization over the infinite-dimensional spaces $\Pi(\theta)$, which is generally not feasible to compute directly. In this section, we discuss one approach to computing these identified sets by approximating the infinite-dimensional space of densities with a finite-dimensional sieve space and the constraint sets with a finite set of constraints. Similar approximations of identified sets have been used, for example, in \cite{MogstadSantosTorgovitsky2018}. Alternatively, the computational approach developed in \cite{ChristensenConnault2023} could be adapted to our setting. Unlike the sieve-based approach we consider below, the dimension of their optimization problem does not depend on the precision of the density approximation. We leave the application of their approach to our problem to future work.

For simplicity, let $\mathcal{Y}_x = [0,1]$ for $x \in \{0,1\}$. This restriction can be relaxed by linearly transforming the outcome variable so that it has support on the unit interval. We also assume that $\mathcal{F} \coloneqq \mathcal{F}_0 = \mathcal{F}_1$, and therefore $\mathcal{F}_{\text{den}} \coloneqq \mathcal{F}_{\text{den},0} = \mathcal{F}_{\text{den}, 1}$. We also impose assumptions \ref{assn:contsupport} and \ref{assn:compactdensities}.

We will approximate $\mathcal{F}_{\text{den}}$ by the convex sieve space $\mathcal{F}_M$, defined by
\begin{align*}
 \mathcal{F}_{M} &\coloneqq \left\{f^M = w^{\top}\textbf{b}^M: w \in \Delta_M \right\},
\end{align*}
where $\textbf{b}^M \coloneqq \{b_0^M, b_1^M, \ldots,b_{M}^M\}$ are the $M$-degree Bernstein basis polynomials scaled by $M + 1$. That is,
$$
b^M_m(y) \coloneqq (M+1) \binom{M}{m} y^m (1-y)^{M-m}
$$
for $m \in \{0, \ldots, M\}$.

Since $\mathcal{F}_M$ is increasing in $M$ and $\bigcup_{M:M > 0}\mathcal{F}_M$ is dense in $\mathcal{F}_{\text{den}}$, $\mathcal{F}_M$ is a sieve space for $\mathcal{F}_{\text{den}}$. We denote the Bernstein polynomial approximation to function $f$ at $y \in [0,1]$ as
\begin{align*}
 (B_M f)(y) &\coloneqq \frac{1}{M + 1}\sum_{m=0}^M f\left(\frac{m}{M}\right) b_m^M(y).
\end{align*}

We also define approximate constraint sets, which are characterized by a finite number of linear equality or inequality constraints. First, we approximate $\mathcal{H}_x$ by the sets,
\begin{align*}
 \mathcal{H}_x^M \coloneqq \{(f_1,\ldots,f_{s_Z}) \in \mathcal{F}_M^{s_Z} : f_j(\cdot) \geq \pi(x \mid z_j) (B_M f_{Y \mid X, Z})( \cdot \mid x, z_j) \text{ for } j = 1, \ldots, s_Z\},
\end{align*}
where $\mathcal{F}_M^{s_Z}$ is the $s_Z$-fold Cartesian product of $\mathcal{F}_M$. In the proposition below, we show that replacing $f_{Y \mid X, Z}$ by its Bernstein approximation is sufficient to characterize this set by a finite number of linear constraints.

Next, we approximate $\mathcal{A}(\theta)$ using a finite set of inequalities. Each model in Section \ref{sec:cont_Y} uses linear inequalities: $\mathcal{A}(\theta) = \{\mathbf{f} \in \mathcal{F}^{s_Z}_{\text{den}} : A(\theta)\mathbf{f} \le a(\theta)\}$. We use a grid of $N$ points in $[0,1]$ (for example, $y_n = n/(N + 1)$ for $n = 1, ..., N$), and then define $\mathcal{A}^{M,N}(\theta)$ as all $\textbf{f} \in \mathcal{F}_M^{s_Z}$ such that $A(\theta) \textbf{f}(y_n) \le a(\theta)$ for each grid point.

The approximate identified set for $\textbf{f}_Y$ is $\Pi_0^{M,N}(\theta) \times \Pi_1^{M,N}(\theta)$, where $\Pi_x^{M,N}(\theta)$ is the intersection of $\mathcal{A}^{M,N}(\theta)$ and $\mathcal{H}_x^M$. The next proposition gives a more convenient representation of this set for computation. Here, $\bar{\Delta}_s^r \coloneqq \left\{\begin{bmatrix} a_1 & \cdots & a_r \end{bmatrix}^{\top} : a_j \in \Delta_s \text{ for } j = 1, \ldots, r\right\}$. $\operatorname{vec}(W)$ is the vectorization of matrix $W$, and $\otimes$ is the Kronecker product.

\begin{proposition}\label{prop:approxIDset}
 For $\mathcal{A}(\theta) = \{\mathbf{f} \in \mathcal{F}_{\text{den}}^{s_Z} : A(\theta) \mathbf{f} \le a(\theta)\}$, $N,M \in \mathbb{N}$, and $\{y_1, \ldots, y_N\} \subset [0,1]$, the approximate constraint sets $\mathcal{A}^{M,N}(\theta)$ and $\mathcal{H}_x^M$ can be represented as
\begin{align*}
 \mathcal{A}^{M,N}(\theta) = \{ W \textbf{b}^M : W \in \mathcal{W}^{M,N}(\theta) \}
\end{align*}
and
\begin{align*}
 \mathcal{H}_x^M = \{ W \textbf{b}^M : W \in \mathcal{W}_x^M \}
\end{align*}
where
\begin{align*}
 \mathcal{W}^{M,N}(\theta) &\coloneqq \left\{ W \in \bar{\Delta}_{M}^{s_Z} : \left( (B^{M,N})^{\top} \otimes A(\theta) \right) \operatorname{vec}(W) \le \iota_{N} \otimes a(\theta) \right\} \\
 \mathcal{W}_x^{M} &\coloneqq \left\{ \textbf{D}_x \Xi^M_x + \textbf{D}_{1-x} W : W \in \bar{\Delta}_{M}^{s_Z} \right\}.
\end{align*}
In $\mathcal{W}^{M,N}(\theta)$, we define $B^{M,N} \coloneqq \begin{bmatrix} \textbf{b}^M(y_1) & \ldots & \textbf{b}^M(y_{N}) \end{bmatrix}$ and $\iota_{N}$ to be the $N$-dimensional vector of ones. In $\mathcal{W}_x^M$, we define $\textbf{D}_x \coloneqq \operatorname{diag}(\pi(x \mid z_1), \ldots, \pi(x \mid z_{s_Z}))$ and $\Xi^M_x$ to be the $s_Z \times (M + 1)$ matrix with elements $f_{Y \mid X, Z}\left( \frac{m-1}{M} \mid x, z_j\right)$ in the $(j, m)$-th position.
\end{proposition}


This proposition shows that the approximate identified set, $\prod_{x \in \{0, 1\}} \left(\mathcal{A}(\theta)^{M,N} \cap \mathcal{H}_x^M\right)$ can be characterized by a finite number of linear constraints. Following Corollary 2, we use this result to characterize the approximate identified set of a functional of $\textbf{f}_Y$ as the solution to a finite linear program.

Approximating the functional $\Gamma(\textbf{f}_Y) = \int_{\mathcal{Y}_0} \omega_0(y)^{\top} f_0(y) dy + \int_{\mathcal{Y}_1} \omega_1(y)^{\top} f_1(y) dy$ with a Riemann sum with $L$ points, we can characterize $\underline{\Gamma}^{M,N,L}(\theta)$ as the solution to the linear program,
\begin{align}
\operatorname*{minimize}_{\substack{W_1, W_{1,0},\\ W_0, W_{0,1}}} \quad &\frac{1}{L} \sum_{n=0}^{L-1} \left( \omega_1\left(\frac{n}{L}\right)^{\top} W_1 + \omega_0\left(\frac{n}{L}\right)^{\top} W_0 \right) \textbf{b}^M\left(\frac{n}{L}\right) \notag \\
\text{subject to} \quad
&W_{0}, W_{0,1}, W_{1}, W_{1,0} \in \bar{\Delta}_{M}^{s_Z} \notag \\
&W_1 - \textbf{D}_1 \Xi^M_1 - \textbf{D}_0 W_{1,0} = 0 \label{eq:LPconstraints1}\\
&W_0 - \textbf{D}_0 \Xi^M_0 - \textbf{D}_1 W_{0,1} = 0 \label{eq:LPconstraints2}\\
&\left((B^{M,N})^{\top} \otimes A(\theta)\right) \operatorname{vec}(W_1) \le \iota_{N} \otimes a(\theta) \label{eq:LPconstraints3}\\
&\left((B^{M,N})^{\top} \otimes A(\theta)\right) \operatorname{vec}(W_0) \le \iota_{N} \otimes a(\theta) \label{eq:LPconstraints4}
\end{align}
The linear inequalities \eqref{eq:LPconstraints3} and \eqref{eq:LPconstraints4} correspond to the constraints that $W_1, W_0 \in \mathcal{W}^{M,N}(\theta)$, and the equality constraints \eqref{eq:LPconstraints1} and \eqref{eq:LPconstraints2} together with the simplex constraints on $W_{x,1-x}$ for $x \in \{0, 1\}$ correspond to the constraints that $W_1 \in \mathcal{W}_1^M$ and $ W_0 \in \mathcal{W}_0^M$ respectively. The optimization program is therefore a linear program in the $s_Z \times (M + 1)$ weight matrices $W_1, W_{1,0}, W_0, W_{0,1}$, which can be solved using standard software.

$\overline{\Gamma}^{M,N,L}(\theta)$ is the solution to the corresponding maximization problem, which is also a linear program.

Since $\Pi^{M,N}(\theta)$ is closed, bounded, and convex, the approximate identified set is
\begin{align*}
	[\underline{\Gamma}^{M,N,L}(\theta), \overline{\Gamma}^{M,N,L}(\theta)] = \{ \Gamma^{M,N,L}(\textbf{f}_Y) : \textbf{f}_Y \in \Pi^{M,N}(\theta) \}.
\end{align*}

Although we omit a full analysis, we expect that $\underline{\Gamma}^{M,N,L}(\theta)$ and $\overline{\Gamma}^{M,N,L}(\theta)$ will converge to $\underline{\Gamma}(\theta)$ and $\overline{\Gamma}(\theta)$ respectively as $M, N, L \rightarrow \infty$ under suitable regularity conditions.



\section{Empirical Application}\label{subsec:empirical}

Here we revisit the empirical study of peer effects in consumer demand by \cite{GilchristSands2016}. Specifically, they study whether movie viewership is affected by peer viewership choices. They provide evidence that movie viewership can have ``momentum" from one weekend to the next. They argue that this is partly because if a movie does well on its opening weekend, it motivates people to see it in subsequent weekends, so they can discuss it with their peers or attend it as a social event.

Identifying this effect is a challenging empirical problem: an apparent peer effect on consumer demand could simply reflect a common understanding of the movie’s unobserved quality. To address this, the authors use a classic instrumental variables approach, using weather as an instrument for opening weekend viewership. They argue that outdoor activities are a substitute for going to the movies, so days with especially nice weather provide a plausibly negative, exogenous shock to viewership.

While its inherent randomness makes weather an appealing instrument, recent literature has cast doubt on its validity as an instrument in many contexts (e.g., \citealt{Mellon2025}). For this application, we highlight three potential violations of the exclusion assumption: (1)  social learning about movie quality, (2) dynamic consumer behavior, and (3) dynamic behavior by movie studios.

\cite{GilchristSands2016} acknowledge that social learning is an important alternative explanation for the observed momentum in movie viewership. The concern is that consumers may be uncertain about a movie's quality and rely on their peers to learn about it. When viewership is high, there is a higher probability that a consumer has friends who have seen the movie and can share their opinion of it. For more reluctant consumers, they may wait until they have good information about the film's quality before seeing it. This is a similar but distinct mechanism from the social incentive that the authors are interested in.

One approach would be to redefine the ``peer effect'' to include this learning effect; however, \cite{GilchristSands2016} are clear that they are interested in the direct social incentive to see the movie. Instead, they explore whether there are learning effects by testing an implication from a model of social learning in \cite{Young2009}. This auxiliary model introduces several additional strong behavioral and distributional assumptions, and the results are not decisive. They conclude that ``Although our estimates do not rule out some role for learning, taken together the results suggest that the observed momentum is driven in part by a preference for shared experience, and not only by learning.'' \citep[][p.1342]{GilchristSands2016}.

Dynamic behavior could also lead to violations of exclusion. When a consumer skips seeing a particular movie one weekend to enjoy the weather, she may simply plan to see the movie on a future weekend. However, the set of available movies in that future weekend is often different, possibly leading them to make a different choice about what movie to see altogether. Similarly, movie studios may respond to first-weekend viewership by adjusting their advertising strategy, which could affect subsequent viewership.

Finally, we note an additional challenge to the exogeneity condition which \cite{GilchristSands2016} address directly in their main specifications. Movie studios may strategically time movie release dates based on seasonal weather patterns, inducing a correlation between weather shocks and unobserved movie quality. To address this problem, the authors condition on several calendar controls, including the week of the year, the year, and holiday indicators. Since movie studios have to release movies based on their expectations of the weather far in advance rather than short-term forecasts, they argue that this strategic behavior should be captured by these calendar controls. In our analyses, we follow their approach of controlling for these time-of-year variables. However, this could still be insufficient if movie studios use more accurate long-term weather forecasts than the average weather for that week of the year.

These potential violations of the exclusion and exogeneity assumptions motivate the importance of assessing sensitivity in this application.

\subsection{Data and Definitions}\label{data-and-definitions}

We use the dataset assembled by \cite{GilchristSands2016} for our analysis. Viewership data on daily ticket sales is obtained from the Internet Movie Database (IMDb) for all movies released between 2002 and 2013. The sample is restricted to movies that were in theaters for at least six weeks, and uses only data on ticket sales on Friday, Saturday, and Sunday.

The instruments are measures of the weather on each weekend. These data come from Weather Underground and consist of (1) the daily maximum temperature, (2) inches of rain, and (3) inches of snow in \(1,941\) weather stations across the country. In order to create national aggregate measures, weather station-level data is weighted by $\omega_j = \frac{n_j}{\sum_j n_j}$ for each weather station $j$ where $n_j$ is the number of movie theaters for which $j$ is the closest weather station.\footnote{To do this, they first assign each zip code to the closest weather station, and obtain the number of movie theaters in each zip code from the U.S. Census Zip Code Business Patterns data.} For any weather station-level weather measure, $Z_{tj}$, the aggregate instrument is, $Z_{t} = \sum_j \omega_j Z_{tj}$.

We define the potential outcome, $Y_{i}(x)$, to be the viewership of movie $i$ in the second weekend of its release with or without a negative shock to viewership in the opening weekend, $x$. The treatment is binary, with $x = 1$ when opening-weekend viewership is below its 25th percentile. This specification of the treatment is motivated by the observation in \cite{GilchristSands2016} that good weather tends to suppress viewership.

We want to ask whether such a negative shock to initial viewership increases the probability of low viewership in subsequent weekends through peer effects. We begin by defining low viewership in the second weekend analogously to the treatment. Specifically, we consider the summary outcome $\mathbf{1}(Y_i(x) \le \underbar{y})$, where $\underbar{y}$ is the 25th percentile of viewership in the second weekend across all movies. The natural parameter of interest is the average treatment effect (ATE), $\mathbb{E}(\mathbf{1}(Y_i(1) \le \underbar{y}) - \mathbf{1}(Y_i(0) \le \underbar{y}))$. This is the effect of a negative shock to opening weekend viewership on the probability of low viewership in the second weekend. Moving beyond this coarse measure of low viewership in the second weekend, we also consider quantile treatment effects across the distribution of viewership in that weekend. That is for, a range of quantiles $\tau \in (0,1)$, we consider the parameter $\text{QTE}(\tau) = Q_{Y_i(1)}(\tau) - Q_{Y_i(0)}(\tau)$ where $Q_{Y_i(x)}(\tau)$ is the $\tau$th quantile of the distribution of $Y_i(x)$.

To minimize endogeneity between movie quality and opening weekend weather, we follow the approach of Gilchrist and Sands (2016) and residualize all variables (viewership in the first and second weekends and the weather instrument) using a set of week-of-year dummies. We use their preferred weather instrument, the share of theaters with a daily high temperature between 75 and 80 degrees Fahrenheit, which we discretize into quintiles. Finally, we condition on this same weather variable on the second weekend. This helps control for potential serial correlation in weather across weekends, which is not captured by the week-of-year dummies.

\subsection{Sensitivity Analysis}\label{sensitivity-analysis}

We begin with the discretized outcome. Under the baseline of exogeneity, we find that a negative shock on viewership in the initial weekend increases the probability of low viewership in the second weekend. The estimated identified set for the ATE is $[0.04, 0.87]$. This result, which bounds the ATE above zero, is qualitatively consistent with the conclusion of Gilchrist and Sands (2016), who find a positive effect of opening weekend viewership on subsequent weekend viewership using a 2SLS estimator. While the lower bound is small, this means that peer effects increase the probability of low viewership in the second weekend by at least $4.4\%$, which is a quantitatively important effect size.

We find, however, that this conclusion is sensitive to relatively small violations of the exogeneity assumption. In Table \ref{tab:ate_bounds} we present the estimated ATE bounds for different levels of $c$-dependence. The interval between the lower and upper lines is the identified set for the ATE at each level of $c$-dependence. Even at low levels of $c$-dependence, the identified set for the ATE includes zero. The lowest level of $c$-dependence at which the identified set for the ATE includes $0$ -- the \textit{breakdown point} -- is $0.015$, or when the latent propensity score is allowed to be 1.5 percentage points away from the observed propensity score.

\begin{table}[tbh]
\begin{center}
\begin{tabular}{cc}
\hline
$c$ & Estimated ATE bounds \\
\hline
0.000 & [0.038, 0.872] \\
0.010 & [0.012, 0.880] \\
0.015 & [0.000, 0.883] \\
0.025 & [-0.024, 0.889] \\
0.050 & [-0.055, 0.901] \\
0.100 & [-0.071, 0.917] \\
0.200 & [-0.071, 0.928] \\
0.500 & [-0.071, 0.929] \\
1.000 & [-0.071, 0.929] \\
\hline
\end{tabular}

\caption{Estimated ATE bounds for different values of sensitivity parameter $c$.}
\label{tab:ate_bounds}
\end{center}
\end{table}


To explore the distributional effects of a negative shock to opening weekend viewership, we now turn to the quantile treatment effects (QTE) for the continuous outcome $Y_i(x)$. In Table \ref{tab:qte_bounds}, we report the identified set of the QTE across several quantiles and different levels of $c$-dependence. Consistent with the results of the discretized outcome, we find that under the baseline assumption of exogeneity, the identified set for the QTE at the $10$th and $25$th percentile is negative and bounded away from zero. A negative shock to opening weekend viewership causes the 25th percentile of viewership in the second weekend to decrease by at least $0.39$ million tickets. These results, however, hold only for the bottom half of the distribution of potential outcomes. At the $50$th, $75$th, and $90$th percentiles, the identified set is very wide and includes zero.

\begin{table}[tbh]
\begin{center}
\begin{tabular}{lccccc}
\toprule
Percentile           & 10\% & 25\% & 50\% & 75\% & 90\% \\
\midrule
$c$ = 0.00   & [-2.94, -0.60] & [-3.45, -0.39] & [-4.02, 8.89] & [-3.11, 7.99] & [-5.57, 6.39] \\
$c$ = 0.02   & [-2.97, 1.90] & [-3.50, -0.13] & [-4.15, 9.02] & [-3.90, 8.11] & [-9.92, 6.57] \\
$c$ = 0.10   & [-3.06, 2.29] & [-3.58, 0.90] & [-4.29, 9.15] & [-7.16, 8.32] & [-10.05, 7.08] \\
\bottomrule
\end{tabular}

\caption{QTE bounds under $c$-dependence. This table shows the estimated identified set for $Q_{Y_i(1)}(\tau) - Q_{Y_i(0)}(\tau)$ for a range of values of $\tau$ (columns) and levels of $c$-dependence (rows).}
\label{tab:qte_bounds}
\end{center}
\end{table}

To see why the identified set for the QTE is much less informative for higher quantiles, it is useful to examine the identified sets for the potential outcome CDFs directly. Figure \ref{fig:cdf_bounds} shows the upper and lower bounds on the CDF for $(Y_i(1), Y_i(0))$ at different levels of $c$-dependence. The first panel shows the bounds under exogeneity, while the second shows a $c$-dependence level of $0.1$. There is an asymmetry in the bounds of the distributions of potential outcomes, with much tighter bounds for the potential outcome with $x = 0$ in which viewership in the opening weekend is above the $25$th percentile. This is because there is a much larger mass of observations with $X = 0$ than with $X = 1$. In addition, the data is largely uninformative about the top half of the distribution of $Y_i(1)$. This reflects the fact that nearly all of the observed mass of $Y$ conditional on $X = 1$ is in the lower half of the support of $Y$. Since we make no monotonicity assumption or other shape restriction, the bounds on the CDF of $Y_i(1)$ have no other restriction except for the lower bound from the mass below $0$.

\begin{figure}[tbh]
\centering
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (505.89,252.94);
\begin{scope}
\path[clip] ( 33.59, 27.33) rectangle (234.10,231.58);
\definecolor{drawColor}{gray}{0.92}

\path[draw=drawColor,line width= 0.3pt,line join=round] ( 33.59, 55.96) --
	(234.10, 55.96);

\path[draw=drawColor,line width= 0.3pt,line join=round] ( 33.59, 94.64) --
	(234.10, 94.64);

\path[draw=drawColor,line width= 0.3pt,line join=round] ( 33.59,133.32) --
	(234.10,133.32);

\path[draw=drawColor,line width= 0.3pt,line join=round] ( 33.59,172.01) --
	(234.10,172.01);

\path[draw=drawColor,line width= 0.3pt,line join=round] ( 33.59,210.69) --
	(234.10,210.69);

\path[draw=drawColor,line width= 0.3pt,line join=round] ( 33.59,230.03) --
	(234.10,230.03);

\path[draw=drawColor,line width= 0.3pt,line join=round] ( 60.67, 27.33) --
	( 60.67,231.58);

\path[draw=drawColor,line width= 0.3pt,line join=round] (132.73, 27.33) --
	(132.73,231.58);

\path[draw=drawColor,line width= 0.3pt,line join=round] (204.80, 27.33) --
	(204.80,231.58);

\path[draw=drawColor,line width= 0.6pt,line join=round] ( 33.59, 36.61) --
	(234.10, 36.61);

\path[draw=drawColor,line width= 0.6pt,line join=round] ( 33.59, 75.30) --
	(234.10, 75.30);

\path[draw=drawColor,line width= 0.6pt,line join=round] ( 33.59,113.98) --
	(234.10,113.98);

\path[draw=drawColor,line width= 0.6pt,line join=round] ( 33.59,152.67) --
	(234.10,152.67);

\path[draw=drawColor,line width= 0.6pt,line join=round] ( 33.59,191.35) --
	(234.10,191.35);

\path[draw=drawColor,line width= 0.6pt,line join=round] ( 96.70, 27.33) --
	( 96.70,231.58);

\path[draw=drawColor,line width= 0.6pt,line join=round] (168.76, 27.33) --
	(168.76,231.58);
\definecolor{drawColor}{RGB}{0,0,0}

\path[draw=drawColor,line width= 0.6pt,line join=round] ( 42.70, 36.70) --
	( 44.56, 36.90) --
	( 46.42, 37.38) --
	( 48.28, 38.13) --
	( 50.14, 38.66) --
	( 52.00, 39.12) --
	( 53.86, 39.54) --
	( 55.72, 40.03) --
	( 57.58, 40.67) --
	( 59.44, 41.53) --
	( 61.30, 42.64) --
	( 63.16, 44.01) --
	( 65.02, 45.66) --
	( 66.88, 47.58) --
	( 68.74, 49.80) --
	( 70.60, 52.37) --
	( 72.46, 55.31) --
	( 74.32, 58.57) --
	( 76.18, 62.09) --
	( 78.04, 65.76) --
	( 79.90, 69.55) --
	( 81.76, 73.42) --
	( 83.62, 77.28) --
	( 85.48, 81.02) --
	( 87.34, 84.45) --
	( 89.20, 87.37) --
	( 91.06, 89.76) --
	( 92.92, 91.59) --
	( 94.78, 92.93) --
	( 96.64, 93.90) --
	( 98.50, 94.56) --
	(100.36, 95.02) --
	(102.22, 95.32) --
	(104.08, 95.54) --
	(105.94, 95.69) --
	(107.80, 95.80) --
	(109.66, 95.90) --
	(111.52, 95.98) --
	(113.38, 96.05) --
	(115.24, 96.11) --
	(117.10, 96.16) --
	(118.96, 96.20) --
	(120.82, 96.22) --
	(122.68, 96.24) --
	(124.54, 96.25) --
	(126.40, 96.25) --
	(128.26, 96.26) --
	(130.12, 96.26) --
	(131.98, 96.26) --
	(133.84, 96.26) --
	(135.70, 96.26) --
	(137.56, 96.26) --
	(139.42, 96.26) --
	(141.28, 96.26) --
	(143.14, 96.26) --
	(145.00, 96.26) --
	(146.86, 96.26) --
	(148.72, 96.26) --
	(150.58, 96.26) --
	(152.44, 96.26) --
	(154.30, 96.26) --
	(156.16, 96.26) --
	(158.02, 96.26) --
	(159.88, 96.26) --
	(161.74, 96.26) --
	(163.60, 96.26) --
	(165.46, 96.26) --
	(167.33, 96.26) --
	(169.19, 96.26) --
	(171.05, 96.26) --
	(172.91, 96.26) --
	(174.77, 96.26) --
	(176.63, 96.26) --
	(178.49, 96.26) --
	(180.35, 96.26) --
	(182.21, 96.26) --
	(184.07, 96.26) --
	(185.93, 96.26) --
	(187.79, 96.26) --
	(189.65, 96.26) --
	(191.51, 96.26) --
	(193.37, 96.26) --
	(195.23, 96.26) --
	(197.09, 96.26) --
	(198.95, 96.26) --
	(200.81, 96.26) --
	(202.67, 96.26) --
	(204.53, 96.26) --
	(206.39, 96.26) --
	(208.25, 96.26) --
	(210.11, 96.26) --
	(211.97, 96.26) --
	(213.83, 96.26) --
	(215.69, 96.26) --
	(217.55, 96.26) --
	(219.41, 96.26) --
	(221.27, 96.26) --
	(223.13, 96.26) --
	(224.99, 96.26);

\path[draw=drawColor,line width= 0.6pt,line join=round] ( 42.70, 79.54) --
	( 44.56,114.79) --
	( 46.42,128.97) --
	( 48.28,132.09) --
	( 50.14,133.65) --
	( 52.00,134.27) --
	( 53.86,134.71) --
	( 55.72,135.19) --
	( 57.58,135.84) --
	( 59.44,136.70) --
	( 61.30,137.82) --
	( 63.16,139.22) --
	( 65.02,140.88) --
	( 66.88,142.82) --
	( 68.74,145.05) --
	( 70.60,147.64) --
	( 72.46,150.60) --
	( 74.32,153.88) --
	( 76.18,157.42) --
	( 78.04,161.12) --
	( 79.90,164.96) --
	( 81.76,168.87) --
	( 83.62,172.76) --
	( 85.48,176.47) --
	( 87.34,179.82) --
	( 89.20,182.71) --
	( 91.06,185.05) --
	( 92.92,186.85) --
	( 94.78,188.18) --
	( 96.64,189.12) --
	( 98.50,189.77) --
	(100.36,190.21) --
	(102.22,190.51) --
	(104.08,190.71) --
	(105.94,190.85) --
	(107.80,190.96) --
	(109.66,191.05) --
	(111.52,191.13) --
	(113.38,191.20) --
	(115.24,191.25) --
	(117.10,191.30) --
	(118.96,191.34) --
	(120.82,191.36) --
	(122.68,191.38) --
	(124.54,191.39) --
	(126.40,191.39) --
	(128.26,191.39) --
	(130.12,191.40) --
	(131.98,191.40) --
	(133.84,191.40) --
	(135.70,191.40) --
	(137.56,191.40) --
	(139.42,191.40) --
	(141.28,191.40) --
	(143.14,191.40) --
	(145.00,191.40) --
	(146.86,191.40) --
	(148.72,191.40) --
	(150.58,191.40) --
	(152.44,191.40) --
	(154.30,191.40) --
	(156.16,191.40) --
	(158.02,191.40) --
	(159.88,191.40) --
	(161.74,191.40) --
	(163.60,191.40) --
	(165.46,191.40) --
	(167.33,191.40) --
	(169.19,191.40) --
	(171.05,191.40) --
	(172.91,191.40) --
	(174.77,191.40) --
	(176.63,191.40) --
	(178.49,191.40) --
	(180.35,191.40) --
	(182.21,191.40) --
	(184.07,191.40) --
	(185.93,191.40) --
	(187.79,191.40) --
	(189.65,191.40) --
	(191.51,191.40) --
	(193.37,191.40) --
	(195.23,191.40) --
	(197.09,191.40) --
	(198.95,191.40) --
	(200.81,191.40) --
	(202.67,191.40) --
	(204.53,191.40) --
	(206.39,191.40) --
	(208.25,191.40) --
	(210.11,191.40) --
	(211.97,191.40) --
	(213.83,191.40) --
	(215.69,191.40) --
	(217.55,191.40) --
	(219.41,191.40) --
	(221.27,191.40) --
	(223.13,191.40) --
	(224.99,191.40);

\path[draw=drawColor,line width= 0.6pt,dash pattern=on 2pt off 2pt ,line join=round] ( 42.70, 36.61) --
	( 44.56, 36.61) --
	( 46.42, 36.61) --
	( 48.28, 36.61) --
	( 50.14, 36.61) --
	( 52.00, 36.61) --
	( 53.86, 36.61) --
	( 55.72, 36.62) --
	( 57.58, 36.62) --
	( 59.44, 36.63) --
	( 61.30, 36.64) --
	( 63.16, 36.67) --
	( 65.02, 36.74) --
	( 66.88, 36.85) --
	( 68.74, 37.05) --
	( 70.60, 37.38) --
	( 72.46, 37.92) --
	( 74.32, 38.73) --
	( 76.18, 39.90) --
	( 78.04, 41.54) --
	( 79.90, 43.74) --
	( 81.76, 46.59) --
	( 83.62, 50.18) --
	( 85.48, 54.58) --
	( 87.34, 59.83) --
	( 89.20, 65.87) --
	( 91.06, 72.56) --
	( 92.92, 79.76) --
	( 94.78, 87.26) --
	( 96.64, 94.86) --
	( 98.50,102.32) --
	(100.36,109.45) --
	(102.22,116.13) --
	(104.08,122.29) --
	(105.94,127.91) --
	(107.80,133.01) --
	(109.66,137.58) --
	(111.52,141.66) --
	(113.38,145.25) --
	(115.24,148.40) --
	(117.10,151.17) --
	(118.96,153.60) --
	(120.82,155.73) --
	(122.68,157.62) --
	(124.54,159.29) --
	(126.40,160.80) --
	(128.26,162.18) --
	(130.12,163.44) --
	(131.98,164.61) --
	(133.84,165.69) --
	(135.70,166.70) --
	(137.56,167.65) --
	(139.42,168.54) --
	(141.28,169.40) --
	(143.14,170.23) --
	(145.00,171.02) --
	(146.86,171.78) --
	(148.72,172.46) --
	(150.58,173.07) --
	(152.44,173.60) --
	(154.30,174.06) --
	(156.16,174.47) --
	(158.02,174.84) --
	(159.88,175.18) --
	(161.74,175.51) --
	(163.60,175.83) --
	(165.46,176.15) --
	(167.33,176.46) --
	(169.19,176.77) --
	(171.05,177.08) --
	(172.91,177.38) --
	(174.77,177.67) --
	(176.63,177.95) --
	(178.49,178.23) --
	(180.35,178.50) --
	(182.21,178.77) --
	(184.07,179.02) --
	(185.93,179.25) --
	(187.79,179.46) --
	(189.65,179.66) --
	(191.51,179.84) --
	(193.37,180.02) --
	(195.23,180.18) --
	(197.09,180.34) --
	(198.95,180.48) --
	(200.81,180.62) --
	(202.67,180.75) --
	(204.53,180.88) --
	(206.39,181.03) --
	(208.25,181.17) --
	(210.11,181.32) --
	(211.97,181.44) --
	(213.83,181.52) --
	(215.69,181.58) --
	(217.55,181.61) --
	(219.41,181.65) --
	(221.27,181.69) --
	(223.13,181.74) --
	(224.99,181.80);

\path[draw=drawColor,line width= 0.6pt,dash pattern=on 2pt off 2pt ,line join=round] ( 42.70, 40.84) --
	( 44.56, 44.28) --
	( 46.42, 45.64) --
	( 48.28, 45.89) --
	( 50.14, 46.00) --
	( 52.00, 46.02) --
	( 53.86, 46.02) --
	( 55.72, 46.02) --
	( 57.58, 46.03) --
	( 59.44, 46.04) --
	( 61.30, 46.05) --
	( 63.16, 46.08) --
	( 65.02, 46.15) --
	( 66.88, 46.26) --
	( 68.74, 46.47) --
	( 70.60, 46.80) --
	( 72.46, 47.34) --
	( 74.32, 48.16) --
	( 76.18, 49.36) --
	( 78.04, 51.02) --
	( 79.90, 53.23) --
	( 81.76, 56.09) --
	( 83.62, 59.71) --
	( 85.48, 64.14) --
	( 87.34, 69.38) --
	( 89.20, 75.41) --
	( 91.06, 82.14) --
	( 92.92, 89.38) --
	( 94.78, 96.93) --
	( 96.64,104.54) --
	( 98.50,112.00) --
	(100.36,119.14) --
	(102.22,125.83) --
	(104.08,132.01) --
	(105.94,137.63) --
	(107.80,142.70) --
	(109.66,147.23) --
	(111.52,151.24) --
	(113.38,154.79) --
	(115.24,157.93) --
	(117.10,160.70) --
	(118.96,163.13) --
	(120.82,165.26) --
	(122.68,167.14) --
	(124.54,168.81) --
	(126.40,170.31) --
	(128.26,171.69) --
	(130.12,172.96) --
	(131.98,174.14) --
	(133.84,175.24) --
	(135.70,176.26) --
	(137.56,177.20) --
	(139.42,178.08) --
	(141.28,178.93) --
	(143.14,179.74) --
	(145.00,180.52) --
	(146.86,181.26) --
	(148.72,181.93) --
	(150.58,182.53) --
	(152.44,183.06) --
	(154.30,183.51) --
	(156.16,183.91) --
	(158.02,184.28) --
	(159.88,184.62) --
	(161.74,184.95) --
	(163.60,185.27) --
	(165.46,185.59) --
	(167.33,185.90) --
	(169.19,186.22) --
	(171.05,186.53) --
	(172.91,186.84) --
	(174.77,187.13) --
	(176.63,187.41) --
	(178.49,187.68) --
	(180.35,187.96) --
	(182.21,188.23) --
	(184.07,188.48) --
	(185.93,188.71) --
	(187.79,188.91) --
	(189.65,189.10) --
	(191.51,189.27) --
	(193.37,189.44) --
	(195.23,189.60) --
	(197.09,189.76) --
	(198.95,189.90) --
	(200.81,190.04) --
	(202.67,190.17) --
	(204.53,190.31) --
	(206.39,190.45) --
	(208.25,190.61) --
	(210.11,190.76) --
	(211.97,190.88) --
	(213.83,190.95) --
	(215.69,190.99) --
	(217.55,191.03) --
	(219.41,191.06) --
	(221.27,191.10) --
	(223.13,191.15) --
	(224.99,191.22);
\end{scope}
\begin{scope}
\path[clip] (239.60, 27.33) rectangle (440.11,231.58);
\definecolor{drawColor}{gray}{0.92}

\path[draw=drawColor,line width= 0.3pt,line join=round] (239.60, 55.96) --
	(440.11, 55.96);

\path[draw=drawColor,line width= 0.3pt,line join=round] (239.60, 94.64) --
	(440.11, 94.64);

\path[draw=drawColor,line width= 0.3pt,line join=round] (239.60,133.32) --
	(440.11,133.32);

\path[draw=drawColor,line width= 0.3pt,line join=round] (239.60,172.01) --
	(440.11,172.01);

\path[draw=drawColor,line width= 0.3pt,line join=round] (239.60,210.69) --
	(440.11,210.69);

\path[draw=drawColor,line width= 0.3pt,line join=round] (239.60,230.03) --
	(440.11,230.03);

\path[draw=drawColor,line width= 0.3pt,line join=round] (266.68, 27.33) --
	(266.68,231.58);

\path[draw=drawColor,line width= 0.3pt,line join=round] (338.74, 27.33) --
	(338.74,231.58);

\path[draw=drawColor,line width= 0.3pt,line join=round] (410.81, 27.33) --
	(410.81,231.58);

\path[draw=drawColor,line width= 0.6pt,line join=round] (239.60, 36.61) --
	(440.11, 36.61);

\path[draw=drawColor,line width= 0.6pt,line join=round] (239.60, 75.30) --
	(440.11, 75.30);

\path[draw=drawColor,line width= 0.6pt,line join=round] (239.60,113.98) --
	(440.11,113.98);

\path[draw=drawColor,line width= 0.6pt,line join=round] (239.60,152.67) --
	(440.11,152.67);

\path[draw=drawColor,line width= 0.6pt,line join=round] (239.60,191.35) --
	(440.11,191.35);

\path[draw=drawColor,line width= 0.6pt,line join=round] (302.71, 27.33) --
	(302.71,231.58);

\path[draw=drawColor,line width= 0.6pt,line join=round] (374.77, 27.33) --
	(374.77,231.58);
\definecolor{drawColor}{RGB}{0,0,0}

\path[draw=drawColor,line width= 0.6pt,line join=round] (248.71, 36.68) --
	(250.57, 36.84) --
	(252.43, 37.08) --
	(254.29, 37.59) --
	(256.15, 37.93) --
	(258.01, 38.21) --
	(259.87, 38.48) --
	(261.73, 38.79) --
	(263.59, 39.20) --
	(265.45, 39.77) --
	(267.31, 40.50) --
	(269.17, 41.40) --
	(271.03, 42.46) --
	(272.89, 43.72) --
	(274.75, 45.18) --
	(276.61, 46.89) --
	(278.47, 48.91) --
	(280.33, 51.24) --
	(282.19, 53.85) --
	(284.05, 56.69) --
	(285.91, 59.66) --
	(287.77, 62.68) --
	(289.63, 65.63) --
	(291.49, 68.38) --
	(293.35, 70.83) --
	(295.21, 72.92) --
	(297.07, 74.60) --
	(298.93, 75.91) --
	(300.79, 76.90) --
	(302.65, 77.59) --
	(304.51, 78.05) --
	(306.37, 78.36) --
	(308.23, 78.57) --
	(310.09, 78.72) --
	(311.95, 78.82) --
	(313.81, 78.89) --
	(315.67, 78.96) --
	(317.53, 79.02) --
	(319.39, 79.07) --
	(321.25, 79.11) --
	(323.11, 79.15) --
	(324.97, 79.18) --
	(326.83, 79.19) --
	(328.69, 79.20) --
	(330.55, 79.21) --
	(332.41, 79.21) --
	(334.28, 79.22) --
	(336.14, 79.22) --
	(338.00, 79.22) --
	(339.86, 79.22) --
	(341.72, 79.22) --
	(343.58, 79.22) --
	(345.44, 79.22) --
	(347.30, 79.22) --
	(349.16, 79.22) --
	(351.02, 79.22) --
	(352.88, 79.22) --
	(354.74, 79.22) --
	(356.60, 79.22) --
	(358.46, 79.22) --
	(360.32, 79.22) --
	(362.18, 79.22) --
	(364.04, 79.22) --
	(365.90, 79.22) --
	(367.76, 79.22) --
	(369.62, 79.22) --
	(371.48, 79.22) --
	(373.34, 79.22) --
	(375.20, 79.22) --
	(377.06, 79.22) --
	(378.92, 79.22) --
	(380.78, 79.22) --
	(382.64, 79.22) --
	(384.50, 79.22) --
	(386.36, 79.22) --
	(388.22, 79.22) --
	(390.08, 79.22) --
	(391.94, 79.22) --
	(393.80, 79.22) --
	(395.66, 79.22) --
	(397.52, 79.22) --
	(399.38, 79.22) --
	(401.24, 79.22) --
	(403.10, 79.22) --
	(404.96, 79.22) --
	(406.82, 79.22) --
	(408.68, 79.22) --
	(410.54, 79.22) --
	(412.40, 79.22) --
	(414.26, 79.22) --
	(416.12, 79.22) --
	(417.98, 79.22) --
	(419.84, 79.22) --
	(421.70, 79.22) --
	(423.56, 79.22) --
	(425.42, 79.22) --
	(427.28, 79.22) --
	(429.14, 79.22) --
	(431.00, 79.22);

\path[draw=drawColor,line width= 0.6pt,line join=round] (248.71, 87.26) --
	(250.57,128.54) --
	(252.43,145.19) --
	(254.29,148.48) --
	(256.15,150.01) --
	(258.01,150.43) --
	(259.87,150.71) --
	(261.73,151.03) --
	(263.59,151.46) --
	(265.45,152.03) --
	(267.31,152.78) --
	(269.17,153.70) --
	(271.03,154.76) --
	(272.89,156.03) --
	(274.75,157.53) --
	(276.61,159.30) --
	(278.47,161.33) --
	(280.33,163.63) --
	(282.19,166.19) --
	(284.05,168.99) --
	(285.91,171.97) --
	(287.77,174.98) --
	(289.63,177.90) --
	(291.49,180.63) --
	(293.35,183.08) --
	(295.21,185.18) --
	(297.07,186.88) --
	(298.93,188.20) --
	(300.79,189.17) --
	(302.65,189.86) --
	(304.51,190.33) --
	(306.37,190.64) --
	(308.23,190.84) --
	(310.09,190.98) --
	(311.95,191.07) --
	(313.81,191.14) --
	(315.67,191.20) --
	(317.53,191.25) --
	(319.39,191.30) --
	(321.25,191.34) --
	(323.11,191.37) --
	(324.97,191.39) --
	(326.83,191.41) --
	(328.69,191.42) --
	(330.55,191.43) --
	(332.41,191.43) --
	(334.28,191.43) --
	(336.14,191.43) --
	(338.00,191.43) --
	(339.86,191.43) --
	(341.72,191.43) --
	(343.58,191.43) --
	(345.44,191.43) --
	(347.30,191.43) --
	(349.16,191.43) --
	(351.02,191.43) --
	(352.88,191.43) --
	(354.74,191.43) --
	(356.60,191.43) --
	(358.46,191.43) --
	(360.32,191.43) --
	(362.18,191.43) --
	(364.04,191.43) --
	(365.90,191.43) --
	(367.76,191.43) --
	(369.62,191.43) --
	(371.48,191.43) --
	(373.34,191.43) --
	(375.20,191.43) --
	(377.06,191.43) --
	(378.92,191.43) --
	(380.78,191.43) --
	(382.64,191.43) --
	(384.50,191.43) --
	(386.36,191.43) --
	(388.22,191.43) --
	(390.08,191.43) --
	(391.94,191.43) --
	(393.80,191.43) --
	(395.66,191.43) --
	(397.52,191.43) --
	(399.38,191.43) --
	(401.24,191.43) --
	(403.10,191.43) --
	(404.96,191.43) --
	(406.82,191.43) --
	(408.68,191.43) --
	(410.54,191.43) --
	(412.40,191.43) --
	(414.26,191.43) --
	(416.12,191.43) --
	(417.98,191.43) --
	(419.84,191.43) --
	(421.70,191.43) --
	(423.56,191.43) --
	(425.42,191.43) --
	(427.28,191.43) --
	(429.14,191.43) --
	(431.00,191.43);

\path[draw=drawColor,line width= 0.6pt,dash pattern=on 2pt off 2pt ,line join=round] (248.71, 36.61) --
	(250.57, 36.61) --
	(252.43, 36.61) --
	(254.29, 36.61) --
	(256.15, 36.61) --
	(258.01, 36.61) --
	(259.87, 36.61) --
	(261.73, 36.62) --
	(263.59, 36.62) --
	(265.45, 36.62) --
	(267.31, 36.63) --
	(269.17, 36.65) --
	(271.03, 36.69) --
	(272.89, 36.77) --
	(274.75, 36.90) --
	(276.61, 37.12) --
	(278.47, 37.47) --
	(280.33, 38.00) --
	(282.19, 38.80) --
	(284.05, 39.95) --
	(285.91, 41.57) --
	(287.77, 43.79) --
	(289.63, 46.72) --
	(291.49, 50.42) --
	(293.35, 54.90) --
	(295.21, 60.09) --
	(297.07, 65.90) --
	(298.93, 72.17) --
	(300.79, 78.70) --
	(302.65, 85.32) --
	(304.51, 91.84) --
	(306.37, 98.10) --
	(308.23,103.98) --
	(310.09,109.38) --
	(311.95,114.28) --
	(313.81,118.67) --
	(315.67,122.57) --
	(317.53,126.01) --
	(319.39,129.03) --
	(321.25,131.69) --
	(323.11,134.01) --
	(324.97,136.04) --
	(326.83,137.81) --
	(328.69,139.36) --
	(330.55,140.73) --
	(332.41,141.93) --
	(334.28,143.01) --
	(336.14,143.99) --
	(338.00,144.87) --
	(339.86,145.69) --
	(341.72,146.43) --
	(343.58,147.12) --
	(345.44,147.76) --
	(347.30,148.35) --
	(349.16,148.91) --
	(351.02,149.45) --
	(352.88,149.95) --
	(354.74,150.42) --
	(356.60,150.83) --
	(358.46,151.20) --
	(360.32,151.51) --
	(362.18,151.79) --
	(364.04,152.04) --
	(365.90,152.28) --
	(367.76,152.51) --
	(369.62,152.74) --
	(371.48,152.96) --
	(373.34,153.18) --
	(375.20,153.39) --
	(377.06,153.60) --
	(378.92,153.80) --
	(380.78,153.98) --
	(382.64,154.16) --
	(384.50,154.34) --
	(386.36,154.52) --
	(388.22,154.70) --
	(390.08,154.87) --
	(391.94,155.03) --
	(393.80,155.17) --
	(395.66,155.30) --
	(397.52,155.42) --
	(399.38,155.54) --
	(401.24,155.65) --
	(403.10,155.75) --
	(404.96,155.84) --
	(406.82,155.93) --
	(408.68,156.01) --
	(410.54,156.09) --
	(412.40,156.19) --
	(414.26,156.28) --
	(416.12,156.37) --
	(417.98,156.45) --
	(419.84,156.51) --
	(421.70,156.55) --
	(423.56,156.58) --
	(425.42,156.61) --
	(427.28,156.64) --
	(429.14,156.68) --
	(431.00,156.72);

\path[draw=drawColor,line width= 0.6pt,dash pattern=on 2pt off 2pt ,line join=round] (248.71, 52.15) --
	(250.57, 64.78) --
	(252.43, 69.80) --
	(254.29, 70.71) --
	(256.15, 71.12) --
	(258.01, 71.18) --
	(259.87, 71.19) --
	(261.73, 71.19) --
	(263.59, 71.20) --
	(265.45, 71.20) --
	(267.31, 71.21) --
	(269.17, 71.23) --
	(271.03, 71.28) --
	(272.89, 71.36) --
	(274.75, 71.50) --
	(276.61, 71.73) --
	(278.47, 72.11) --
	(280.33, 72.68) --
	(282.19, 73.48) --
	(284.05, 74.59) --
	(285.91, 76.19) --
	(287.77, 78.39) --
	(289.63, 81.31) --
	(291.49, 85.00) --
	(293.35, 89.47) --
	(295.21, 94.67) --
	(297.07,100.48) --
	(298.93,106.74) --
	(300.79,113.28) --
	(302.65,119.90) --
	(304.51,126.42) --
	(306.37,132.68) --
	(308.23,138.55) --
	(310.09,143.96) --
	(311.95,148.86) --
	(313.81,153.25) --
	(315.67,157.14) --
	(317.53,160.59) --
	(319.39,163.61) --
	(321.25,166.26) --
	(323.11,168.59) --
	(324.97,170.62) --
	(326.83,172.39) --
	(328.69,173.94) --
	(330.55,175.31) --
	(332.41,176.52) --
	(334.28,177.60) --
	(336.14,178.58) --
	(338.00,179.46) --
	(339.86,180.27) --
	(341.72,181.02) --
	(343.58,181.71) --
	(345.44,182.36) --
	(347.30,182.97) --
	(349.16,183.55) --
	(351.02,184.10) --
	(352.88,184.60) --
	(354.74,185.06) --
	(356.60,185.45) --
	(358.46,185.80) --
	(360.32,186.11) --
	(362.18,186.39) --
	(364.04,186.64) --
	(365.90,186.88) --
	(367.76,187.11) --
	(369.62,187.34) --
	(371.48,187.55) --
	(373.34,187.77) --
	(375.20,187.99) --
	(377.06,188.20) --
	(378.92,188.40) --
	(380.78,188.59) --
	(382.64,188.77) --
	(384.50,188.96) --
	(386.36,189.16) --
	(388.22,189.36) --
	(390.08,189.52) --
	(391.94,189.66) --
	(393.80,189.78) --
	(395.66,189.90) --
	(397.52,190.01) --
	(399.38,190.12) --
	(401.24,190.23) --
	(403.10,190.33) --
	(404.96,190.43) --
	(406.82,190.51) --
	(408.68,190.59) --
	(410.54,190.68) --
	(412.40,190.79) --
	(414.26,190.90) --
	(416.12,190.99) --
	(417.98,191.05) --
	(419.84,191.09) --
	(421.70,191.13) --
	(423.56,191.16) --
	(425.42,191.19) --
	(427.28,191.22) --
	(429.14,191.26) --
	(431.00,191.31);
\end{scope}
\begin{scope}
\path[clip] ( 33.59,231.58) rectangle (234.10,247.44);
\definecolor{drawColor}{gray}{0.10}

\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale=  0.80] at (133.84,236.76) {c = 0};
\end{scope}
\begin{scope}
\path[clip] (239.60,231.58) rectangle (440.11,247.44);
\definecolor{drawColor}{gray}{0.10}

\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale=  0.80] at (339.86,236.76) {c = 0.1};
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{gray}{0.30}

\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale=  0.70] at ( 96.70, 17.56) {0};

\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale=  0.70] at (168.76, 17.56) {5};
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{gray}{0.30}

\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale=  0.70] at (302.71, 17.56) {0};

\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale=  0.70] at (374.77, 17.56) {5};
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{gray}{0.30}

\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale=  0.70] at ( 28.64, 34.20) {0.00};

\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale=  0.70] at ( 28.64, 72.89) {0.25};

\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale=  0.70] at ( 28.64,111.57) {0.50};

\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale=  0.70] at ( 28.64,150.25) {0.75};

\node[text=drawColor,anchor=base east,inner sep=0pt, outer sep=0pt, scale=  0.70] at ( 28.64,188.94) {1.00};
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{RGB}{0,0,0}

\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale=  0.90] at (236.85,  7.25) {Residual Viewership on Second Weekend};
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{RGB}{0,0,0}

\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale=  0.90] at ( 11.70,129.45) {Cumulative Distribution};
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{RGB}{0,0,0}

\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale=  0.80] at (456.61,152.54) {Initial};

\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale=  0.80] at (456.61,143.90) {Residual};

\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale=  0.80] at (456.61,135.26) {Viewership};
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{RGB}{0,0,0}

\path[draw=drawColor,line width= 0.6pt,line join=round] (458.06,121.76) -- (469.62,121.76);
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{RGB}{0,0,0}

\path[draw=drawColor,line width= 0.6pt,dash pattern=on 2pt off 2pt ,line join=round] (458.06,107.31) -- (469.62,107.31);
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{RGB}{0,0,0}

\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale=  0.70] at (476.56,119.35) {Low};
\end{scope}
\begin{scope}
\path[clip] (  0.00,  0.00) rectangle (505.89,252.94);
\definecolor{drawColor}{RGB}{0,0,0}

\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale=  0.70] at (476.56,104.89) {High};
\end{scope}
\end{tikzpicture}

\caption{Distributional bounds. These plots show the bounds on the cumulative distribution function of $Y_i(x)$ at two values of $c$-dependence.}
\label{fig:cdf_bounds}
\end{figure}

\section{Conclusion}

We introduced a new, computationally tractable approach for conducting sensitivity to the instrument exclusion and exogeneity assumptions. Our approach does not impose any kind of monotonicity assumption in the first stage, and allows for arbitrarily heterogeneous treatment effects. We did this by developing a unifying sensitivity model which nests several well known approaches to continuously parameterizing relaxations of statistical independence assumptions from the literature. We showed that, under those relaxations, identified sets for parameters like ATE and QTE are solutions to linear programs. Our approach can be used when the outcome is discrete or continuous, and when there are one or multiple discretely supported instruments.

We illustrated the practical value of our results in an empirical study of peer effects in movie viewership. There our sensitivity analysis shows that although ATE is positive under full exclusion and exogeneity (meaning peer effects are present), that conclusion is highly sensitive to minor relaxations of the exclusion and exogeneity assumptions. Overall, our results allow researchers to transparently study and report the robustness of their instrumental variable conclusions to violations of exclusion or exogeneity.



\bibliographystyle{econometrica}
\bibliography{FF_paper}