EconBase
← Back to paper

On the falsification of instrumental variable models for heterogeneous treatment effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

63,330 characters · 11 sections · 80 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the falsification of instrumental variable models for heterogeneous treatment effects

abstractIn this paper I derive a set of testable implications for econometric models defined by three assumptions: (i) the existence of strictly exogenous discrete instruments, (ii) restrictions on how the instruments affect adoption of a finite number of treatment types (such as monotonicity), and (iii) the assumption that the instruments only affect outcomes through their effect on treatment adoption (i.e. an exclusion restriction). The testable implications aggregate (via integration) an otherwise potentially infinite set of inequalities that must hold for every measurable subset of the outcome's support. For binary instruments the testable implications are sharp. Furthermore, I propose an implementation that links restrictions on latent response types to a generalization of first-order stochastic dominance and random utility models, allowing to distinguish violations of the exclusion restriction from violations of monotonicity-type assumptions. The testable implications extend naturally to the many instruments case.

JEL classification: C12, C21, C25, C26, C31, C36.

\onehalfspacing

Introduction

Consider a potential outcomes model defined by the following two equations:

align[align omitted — 82 chars of source]

Suppose that the joint distribution of $(Z_i,X_i,Y_i)$ over the support $\mathcal{Z}\times \mathcal{X}\times \mathcal{Y}$ is observed, and that both $\mathcal{X}$ and $\mathcal{Z}$ are finite while $\mathcal{Y}$ is unrestricted. The goal of this paper is to investigate the joint falsification of three assumptions: (i) that $Z$ takes values across the population as if it was randomly assigned, (ii) restrictions on the effect of $Z$ on $X$ (typically called instrument-response restrictions), and (iii) an exclusion restriction on $Z$, that is, the assumption that:

align[align omitted — 57 chars of source]

These three assumptions play an important role in modern econometrics. First, they are the cornerstone of a range instrumental variables (IV) methods for heterogeneous treatment effects in the spirit of angristImbens1994late. Furthermore, as noted by kwon2024testing, the three assumptions combined are equivalent to the sharp null of full mediation in the context of randomized control trials, that is, the hypothesis that a set of mediators $\mathcal{X}$ can fully account for the effect on $Y$ of an exogenous set of treatments $\mathcal{Z}$.

In this paper I propose a novel set of observable implications of the aforementioned assumptions. The observable implications are sharp when the instrument is binary, and extend naturally to the multiple instruments case. The core idea is to transform (via integration) a potentially infinite set of inequalities imposed by the exclusion restriction, into a finite set of linear restrictions on conditional treatment and response type probabilities. The resulting restrictions are characterized by the integral of the point-wise minimum of multiple conditional sub-densities.

In the many instruments case the observable implications have two advantages: (i) they use information from the join distribution of $(Z_i,X_i,Y_i)$ that previous approaches tailored to test IV validity do not incorporate, and (ii) they do not rely on monotonicity or support restrictions, allowing for general instrument-response restrictions.

Moreover, for instrument-response restrictions that rule out certain behaviors or latent types, the characterization can be further simplified to a finite set of inequalities with a generalized first order stochastic dominance (FOSD) interpretation. This set of inequalities can be used to distinguish violations of the exclusion restriction from violations of instrument-response restrictions, and to link the latter to non-parametric assumptions about the potential preferences of units in the population. Thus, as a byproduct and in the absence of an exclusion restriction, these inequalities can also be used to falsify the assumption that some observed treatment choice probabilities are consistent with a set of pre-specified non-parametric restrictions on how an exogenous binary variable affects the preferences and choices of a population of rational decision makers.

To obtain the generalized FOSD inequalities I show that finding a distribution over potential treatments and outcomes that satisfies the sharp instrument validity conditions amounts to solving a well known discrete transportation problem. This technique, and its equivalence to the maximum flow/minimum cut problems, was first used to characterize incomplete econometric models by ekeland2010optimal and galichon2011set.\footnote{My approach differs in two ways: first, instead of considering transportation problems from the space of observed variables to the domain of latent unobservables, my transportation problem is defined between observed treatment probabilities conditional on different instrument values. Second, to incorporate the exclusion restriction I extend their minimum-cut characterization method to a broader class of capacity constraints.}

My optimal transport approach is related to contemporaneous and independent work by kaido2025testing who study the falsification of shape constraints and exclusion restrictions using graph theory. In general, neither of the two approaches is nested by the other. The main advantage of the approach from kaido2025testing is its generality since it is sharp for the multiple instruments case and applies to other econometric models beyond IV. The main advantage of my approach is that it does not rely on partitioning the support of the outcome but instead incorporates the sharp observable implications of the exclusion restriction as a capacity constraint on the transportation problem which yields a finite linear system of inequalities, even when the outcome has infinite support.

The observable implications: Let $\tilde{Z}\subset\mathcal{Z}$ be a subset of instrument values, and let $S(x_l,\tilde{Z})$ be the group \footnote{The notation $S(x_l,\tilde{Z})$ comes from an acronym for sufficient takers, i.e., units such that $Z_i\in\tilde{Z}$ is a sufficient condition for them to realize treatment $x_l$} of units that would realize treatment $x_l$ whenever $Z_i\in\tilde{Z}$. Suppose that all units in $S(x_l,\tilde{Z})$ realize outcomes in some measurable set $B_Y\subset\mathcal{Y}$ when the treatment is set to $x_l$. Then, the units in $S(x_l,\tilde{Z})$ must be a subset of all the units that realize outcomes $Y_i\in B_y$ when $X_i=x_l$. Consequently, the mass of group $S(x_l,\tilde{Z})$ is bounded above by the minimum over $z_k\in\tilde{Z}$, of the conditional probability of observing $(Y_i\in B_Y,X=x_l)$ given $Z=z_k$. In other words, the mass of group $S(x_l,\tilde{Z})$ is bounded by the minimum over instrument values in $\tilde{Z}$ of the integrals of the sub-densities $\phi_{z_0,x_l}:=\mathbb{P}[X_i=x_l,Z=z_k]f_{y|Z_i=z_0,X_i=x_l}(y)$ over the set $B_Y$.

A core result of this paper is to note that if the previous bound holds for every measurable subset of $\mathcal{Y}$, then, the instrument-wise minimum can be replaced by a point-wise minimum. That is, the probability that units in $S(x_l,\tilde{Z})$ realize outcomes in $B_Y$ is also bounded by the integral over $B_Y$ of the point-wise minimum $\psi(y)=\underset{z_k\in\tilde{Z}}{\min}\{\phi_{z_k,x_k}(y)\}$. Since this bound must hold for every measurable $B_Y$, setting $B_Y=\mathcal{Y}$ yields an upper bound on the probability that a unit belongs to $S(x_l,\tilde{Z})$. The result is a finite set of observable necessary conditions for instrument validity.

Other related work: This paper contributes to the literature on testing the validity of instrumental variables models. The first test for instrument validity for heterogeneous treatment effects was proposed by balke1997bounds and subsequently refined by kitagawa2015test, huber2015testing, and mourifie2017testing. These refinements, however, only apply when both the instrument and treatment are binary, and treatment adoption is weakly monotone in the instrument. kedagni2020generalized proposed a joint test for the exclusion restriction and random assignment --without monotonicity or any other instrument-response restriction-- but their results are only sharp for a binary outcome. Subsequent work has derived sharp testable implications for IV designs on a case-by-case basis bai2024sharp, or proposed general tests that are not sharp sun2023instrument. More recently, kwon2024testing provided the first sharp characterization for general binary instruments and kaido2025testing for multi-valued instruments.

The observable implications I propose apply to a range of IV methods for heterogeneous treatment effects. angristImbens1994late were the first to use non-parametric restrictions on instrument-response heterogeneity to identify causal parameters. Subsequent work has extended their approach to a wide range of settings. Major extensions include: re-interpretations of monotonicity (heckman2018unordered, navjeevan2022ordered); alternative restrictions on how the instrument affects treatment take-up (richardson2010analysis, de2017tolerating, dahl2023never, goff2024does); and generalizations to settings with multiple instruments and treatmens (lee2020treatment, goff2020vector, mogstad2021causal, van2023limited, kazemi2024instrumental, fusejima2024identification, bai2024inference). These methodological advances have been applied in diverse empirical contexts, including education and labor market studies (kline2016evaluating, kirkeboen2016field, pinto2021beyond, mountjoy2022community).

The FOSD characterization I offer and its relation to non-parametric random utility models contributes to the literature that investigates the empirical content of response type restrictions (goff2024does and bai2024identifying) and their microeconomic foundations (vytlacil2002independence, heckman2018unordered, navjeevan2022ordered).

All the testable implications I provide can be represented by the requirement that a linear system of inequalities admits at least one solution. Inference on the feasibility of linear systems and linear programs is an active area of research (bai2022testing, fang2023inference, cho2024simple, goff2025inference) tightly related (and in some cases equivalent) to inference methods for models characterized by moment inequalities such as cox2023simple or andrews2023inference.

The remainder of this paper is organized as follows: In section (ref) I introduce the class of models and restrictions considered. To illustrate the intuition behind the observable implications and compare my approach to the existing literature, in section (ref) I consider the binary instrument case. In section (ref) I specialize the characterization from section (ref) to instrument-response-type restrictions and link them to FOSD and random utility models using optimal transport. In section (ref) I extend the observable implications to the multiple instruments case. Section (ref) concludes

Instrument validity: Notation and framework

The IV model with discrete instruments and treatments

The population consists of units indexed by $i\in\mathcal{I}$. I assume that the joint distribution of the outcome $Y_i$, treatment $X_i$, and instrument $Z_i$ is observed, that is, the joint distribution of $(Y,X,Z)$ is identified. I impose no restrictions on the cardinality or dimension of the support of the outcome, which I denote as $\mathcal{Y}$. The treatment and instrumental variables have finite supports $\mathcal{X}=\{x_0,x_1,...,x_{L-1}\}$ and $\mathcal{Z}=\{z_0,z_1,...,z_{K-1}\}$ respectively.

To formalize the structure of $\mathcal{Y}$, I assume it is equipped with a $\sigma-$algebra $\mathcal{B}_Y$ and a measure $\mu_Y$ so that $(\mathcal{Y},\mathcal{B}_Y,\mu_Y)$ forms a measure space. Analogous constructions hold for $\mathcal{X}$ and $\mathcal{Z}$, which, being finite, are naturally endowed with power set $\sigma$-algebras and counting measures.

I assume that all other relevant covariates are exogenous. Since they play no role in the characterizations I provide, I assume that all the identification arguments are conditional on such covariates and omit them from the formal proofs and results.

Unless otherwise noted, all the results will be presented and discussed using the IV notation and terminology. All results would also hold for the sharp null of full mediation after appropriately relabeling the values of $X$ as mediators and the values of $Z$ as treatments.

I adopt the potential outcomes framework. $X_i(z_k)$ denotes the treatment that unit $i$ would realize if assigned to instrument value $z_{k}$, and $Y_i(x_{l},z_{k})$ denotes the outcome under treatment $x_l$ and instrument $z_k$. The observed values of $Y_i$ and $X_i$ correspond almost surely to the potential outcomes and treatments evaluated at the realized values of $Z_i$ and $X_i$.

align[align omitted — 99 chars of source]

For $Z$ to be a valid instrument, two assumptions are typically invoked, instrument independence and exclusion:

assumption[Instrument validity]\phantom{a} \begin{enumerate} • Instrument independence: \begin{align} \mathbb{P}[Y_i(x_{l'},z_{k'})\in B_Y,X_i(z_{k”})=x_{l}|Z_i=z_k]=\mathbb{P}[Y_i(x_{l'},z_{k'})\in B_Y,X_i(z_{k”})=x_{l}] . \end{align} For every $\phantom{a}x_l,x_{l'}\in\mathcal{X},\phantom{a}z_k,z_{k'},z_{k''}\in\mathcal{Z}$ and $B_Y\in\mathcal{B}_Y$. • Exclusion restriction: \begin{align} Y_i(x_{l},z_k)=Y_i(x_{l}) a.s. \end{align} For every $\phantom{a}x_l\in\mathcal{X}$ and $ z_k\in\mathcal{Z}$ \end{enumerate}

Put in words, the instrument must be independent from potential outcomes and treatments as if it was randomly assigned, and it must have no effect on outcomes after accounting for treatment.

Assumption (ref) implies that the observed joint distribution of $(X,Y,Z)$ is fully determined by the observed distribution of $Z$, and the unobserved distribution of potential outcomes and potential treatments. This motivates the following definitions:

definition\phantom{a} \begin{enumerate} • An instrument-response type $t^Z\in\mathcal{T}_Z$ is a function $t^Z:\mathcal{Z}\longrightarrow \mathcal{X}$ that maps instrument values to treatment values. • A treatment-response type $t^X\in\mathcal{T}_X$ is a function $t^X:\mathcal{X}\longrightarrow \mathcal{Y}$ that maps treatment values to outcome values. • A response type $t$ is a pair $t=(t_z,t_x)\in\mathcal{T}=\mathcal{T}_Z\times\mathcal{T}_X$. \end{enumerate}

Thus, units are fully characterized by their response type $t\in\mathcal{T}$. That is, their potential treatment and outcome realizations over the support of $X$ and $Z$. An important feature of this definition of response types is that by construction they satisfy the exclusion restriction. The reason is that for any $x_l\in\mathcal{X}$, a response type $t=(t^Z,t^X)$ assigns a unique outcome value $t^X(x_l)$ when $X=x_l$ regardless of the value that the instrument takes.

Note that since $\mathcal{X}$ is finite, any instrument-response type $t^Z$ can be represented as a $|\mathcal{Z}|$-dimensional vector $(t^Z(z_1),...,t^Z(z_k),...,t^Z(z_{K}))$. Similarly, any treatment-response type $t^X$ can be represented as a $|\mathcal{X}|$-dimensional vector $\big(t^X(x_1),...,t^X(x_{L})\big)$. Therefore, in a slight abuse of notation, I will refer to both instrument-response and treatment-response types in their vector form so that $t^Z=\big(t^Z(z_1),...,t^Z(z_{K})\big)$ and $t^X=\big(t^X(x_1),...,t^X(x_l),...,t^X(x_{L})\big)$. Consistent with this notation, for $1\leq k \leq K$ and $1\leq l\leq L$, I denote the $k-th$ element of $t^Z$ and the $l-th$ element of $t^X$ as $(t^Z)_k$ and $(t^X)_{l}$ respectively. In other words, $(t^Z)_k$ denotes the treatment value that instrument-response type $t^Z$ realizes when $Z=z_k$, and $(t^X)_{l}$ denotes the outcome that treatment-response type $t^X$ realizes when $X=x_{l}$.

Furthermore, any response type $t=(t^Z,t^X)$ can be represented as the concatenation of the vector representations of $t^Z$ and $t^X$. This representation results in a $|\mathcal{Z}|+|\mathcal{X}|$ dimensional vector that belongs to the set $\mathcal{T}:=\mathcal{T}_Z\times\mathcal{T}_X=\mathcal{X}^{|\mathcal{Z}|}\times\mathcal{Y}^{|\mathcal{X}|}$.

$\mathcal{T}$ is the Cartesian product of a finite number of finite sets, each identical to $\mathcal{X}$, and a finite number of sets identical to $\mathcal{Y}$. Hence, the product measure $\mu_T:=\mu_X\otimes\mu_X\otimes...\mu_X\otimes\mu_Y\otimes\mu_Y\otimes...\otimes\mu_Y$ with respect to the product $\sigma$-algebra $\mathcal{B}_T:=\mathcal{B}_X\otimes\mathcal{B}_X\otimes...\mathcal{B}_X\otimes\mathcal{B}_Y\otimes\mathcal{B}_Y\otimes...\otimes\mathcal{B}_Y$ over $\mathcal{T}$ is well defined. Having equipped $\mathcal{T}$ with a measure I can formally define a distribution over response and instrument-response types.

definition\phantom{a} \begin{enumerate} • A distribution over response types is a probability measure $g:\mathcal{B}_T\longrightarrow [0,1]$ over the measure space $(\mathcal{T},\mathcal{B}_T)$. • A distribution over instrument-response types is a function $p:\mathcal{T}_Z\longrightarrow[0,1]$ such that $p(t^Z)\geq 0$ for all $t^Z\in\mathcal{T}_Z$, and $\sum\limits_{t^Z\in\mathcal{T}_Z}p(t^Z)=1$. \end{enumerate}
remark\phantom{a}Any distribution over response types $g$ defines a unique distribution over instrument-response types $f_g$ according to the following equation: \begin{align} \phantom{aaaaaaaaaa}p_g(t'^Z):=&g\Big(\big\{t=(t^Z,t^X)\in\mathcal{T}:t^Z=t'^Z,t^X\in\mathcal{Y }^{L}\big\}\Big), & \forall\phantom{a} t'^Z\in\mathcal{T}_Z,\\ \phantom{aaaaaaaaaa}=&g\Big(\big\{t=(t^Z,t^X)\in\mathcal{T}:t^Z=t'^Z\big\}\Big), & \forall \phantom{a}t'^Z\in\mathcal{T}_Z. \end{align}

If assumption (ref) holds, then an unobserved true distribution over response types $g_0$ must exist that generates the observed distribution of $(Y,X,Z)$.

definition[Consistency]\phantom{a} A distribution over response types $g\in\Delta(\mathcal{T})$ is consistent with observed conditional outcome and treatment probabilities if: \begin{align} \mathbb{P}[y_i\in B_Y,X_i=x_{l}|Z_i=z_k]=g_0\Big(\big\{(t^Z,t^X)\in\mathcal{T}:(t^Z)_k=x_l,(t^X)_l\in B_Y\big\}\Big),\\ \forall x_l\in\mathcal{X}, z_k\in\mathcal{Z}, and B_Y\in\mathcal{B}_Y \nonumber. \end{align}

In the previous equation, the observed conditional probability on the left hand side must correspond to the mass that $g_0$ assigns to response types whose potential treatment is equal to $x_l$ when the instrument takes the value $z_k$ and their potential outcome is contained in $B_Y$ when treatment is equal to $x_l$.

Let $\Delta(\mathcal{T})$ be the set of all probability measures over $(\mathcal{T},\mathcal{B}_T)$ (i.e. distributions over response types), assumption (ref) is falsified if no $g\in\Delta(\mathcal{T})$ exists that satisfies equation $\ref{Equation:Consistency}$.

Restrictions on instrument-response heterogeneity

To facilitate the identification of causal parameters using the IV method, assumptions that restrict instrument-response heterogeneity are often invoked. The most prominent example is monotonicity in the binary instrument and treatment case which, as angristImbens1994late showed, allows identification of a local average treatment effect. Subsequent work has explored both relaxations and extensions of this assumption to settings with multiple instruments and multiple, ordered and unordered, treatments (heckman2018unordered,goff2020vector, mogstad2021causal, van2023limited, lee2020treatment, bai2024sharp), these generalizations typically rely on ruling out particular instrument-response types.

Alternatively, researchers may impose inequality restrictions on the prevalence or relative frequency of certain instrument-response types—for example, requiring that the proportion of compliers exceeds that of defiers as de2017tolerating, or bounding the share of defiers for sensitivity analysis as in huber2014sensitivity, noack2021sensitivity or yap2025sensitivity). The framework of this paper accomodates all the aforementioned restrictions on their own and combined.

To introduce the class of instrument-response restrictions I consider, note that the total number of possible instrument-response types is $|\mathcal{T}_Z|:=|\mathcal{X}|^{|\mathcal{Z}|}=(L)^{K}$. Therefore, a distribution $p$ over $\mathcal{T}_Z$ can be represented as a vector $\boldsymbol{p}$ in $\mathbb{R}_+^{|\mathcal{T}_Z|}$, where each coordinate corresponds to the probability that $p\in\Delta(\mathcal{T}_Z)$ assigns to a different $t\in\mathcal{T}_Z$.

To formally define $\boldsymbol{p}$ is suffices to note that the set of response types is finite and thus can be well ordered. Then, the $n-th$ coordinate of $\boldsymbol{p}$ represents the probability that the distribution $p\in\Delta(\mathcal{T}_Z)$ assigns to the $n-th$ response type. The specific ordering of response types is not relevant as long as it is well defined.

The class of restrictions on heterogeneity I consider take the form $A_Z\boldsymbol{p}\leq \boldsymbol{r}$ where $A_Z$ is an $M$ by $|\mathcal{T}_Z|$ matrix with known coefficients for some fixed $M\geq 1$, and $\boldsymbol{r}$ is an $M$ dimensional vector with entries that are either known or point identified from the data. I use the subscript $Z$ to emphasize that the matrix $A_Z$ encodes restrictions that pertain to treatment assignment as a function of the instrument.

assumption\phantom{a} Let $M\geq$ be a fixed integer, $A_Z$ a known fixed matrix $A_Z\in\mathbb{R}^{M\times|\mathcal{T}_Z|}$, and $\boldsymbol{r}_Z\in\mathbb{R}^{M}$ a point identified vector that depends on the joint distribution of $(Y,X,Z)$. A distribution over response types $g_0\in\Delta(\mathcal{T})$ exists such that if $\boldsymbol{p}_{g_0}$ denotes the vector representation of $p_{g_0}$, the following system of inequalities holds: \begin{align} A_Z\boldsymbol{p}_{g_0}\leq \boldsymbol{r}_Z. \end{align}

Note that strict equality restrictions can be incorporated by representing them as pairs of inequality constraints. In particular, if the $k$-th entry of $\boldsymbol{p}$ corresponds to the instrument-response type $t^Z\in\mathcal{T}_Z$, the restriction that $t^{Z}$ occurs with zero probability can be imposed by including the canonical basis vectors $e_k$ and $-e_k$ as rows of $A_Z$, and setting the corresponding entries of $\boldsymbol{r}_Z$ to zero. \footnote{$e_k$ denotes the vector that takes the value 0 in every entry except the $k$-th which is equal to 1.} Thus, assumption (ref) can accommodate restrictions that rule-out specific response types. In appendix (ref) I provide examples of models and restrictions encompassed by assumption (ref).

To conclude, note that if no distribution over instrument-response types $p\in\Delta(\mathcal{T}_Z)$ exists such that $\boldsymbol{p}$ satisfies equation (ref) and is consistent with observed conditional treatment probabilities, then, the restrictions on instrument-response heterogeneity imposed by assumption (ref) are falsified.

proposition[Sharp observable implications of assumption (ref)]\phantom{a}A distribution over instrument-response types $p$ is consistent with assumption (ref) and observed conditional treatment probabilitites $\mathbb{P}[X_i=x_l|Z_i=x_l]$, if and only if it satisfies the following system of linear equalities and inequalities: \begin{align} A_{Z}\boldsymbol{p}&\leq \boldsymbol{r}_Z,\\ A_{X}\boldsymbol{p}&=\boldsymbol{r}_X. \end{align} Where $A_Z$ and $\boldsymbol{r}_Z$ are defined as in assumption (ref) and the system $A_X\boldsymbol{p}=\boldsymbol{r}_X$ stacks the following set of equations: \begin{align} \phantom{aaaaa}\mathbb{P}[X_i=x_{l}|Z_i=z_k]&=\sum\limits_{t^Z\in\mathcal{T}_Z:t^Z(z_k)=x_l}p(t^Z), \phantom{aaaaa}\forall z_{l}\in\mathcal{Z},x_k\in\mathcal{X}. \end{align}

The proof of proposition (ref) is straightforward and follows from the previous discussion.

The binary instrument case

Assumption (ref) and the observed conditional treatment probabilities restrict the set of admissible instrument-response probabilitites. Now I show that, for binary instruments, the restrictions implied by the observed conditional distribution of the outcome combined with the exclusion restriction can be expressed as a finite set of linear restrictions on instrument-response probabilities too. This insight reduces a potentially infinite set of inequalities to a finite one.

To this end, I first introduce some additional notation. Define the function $\Psi_{x_l}:\mathcal{Y}\longrightarrow \mathbb{R}_{+}$ as follows:

align[align omitted — 128 chars of source]

Where $\boldsymbol{P}_{x_l|z_k}:=\mathbb{P}[X=x_l|Z=z_k]$ and $f_{Y|X=x_{l},Z=z_{k}}$ denote the distribution of $Y$ conditional on $X=X_k$ and $Z=Z_k$.

Now define the quantity:

align[align omitted — 72 chars of source]

Finally, for any $x_l\in \mathcal{X}$, define the set of always-$x_l$ takers as units such that $X_i(z_0)=X_i(z_1)=x_l$ and note that a distribution over response types $p\in\Delta(\mathcal{T}_Z)$ assigns probability $p\left(AT_{x_l}\right)$ to always-$x_l$ takers where $p\left(AT_{x_l}\right)$ is defined as follows:

align[align omitted — 122 chars of source]

Where $\boldsymbol{I}_{\{t(z_0)=t(z_1)\}}$ is an indicator that response type $t^Z$ realizes treatment $x_l$ at both values of $Z$.

Finally, Let $t^Z_j$ denote the $j$-th instrument-response type according to the fixed pre-specified order over instrument-response types. Let $A_{Y}$ be a $L\times|\mathcal{T}_Z|$ matrix with entires $(A_{Y})_{l,j}=\boldsymbol{I}_{\{t^Z_j(z_0)=t^Z_j(z_1)=x_l\}}$ and let $\boldsymbol{\Psi}$ denote the vector that stacks the parameters $\Psi_{x_0},...,\Psi_{x_K}$ and has $\Psi_{x_l}$ in its $l$-th coordinate.

The following theorem characterizes the sharp observable implications of binary instruments for heterogeneous treatment effects:

theorem[Sharp observable implications of binary instruments]\phantom{a}\\ The following statements are equivalent: \begin{enumerate} • A distribution over response types $g\in\Delta(\mathcal{T})$ that satisfies, assumptions (ref) (instrument validity), and (ref) (instrument-response restriction) and definition (ref) (consistency with observed probabilities) exists. • A distribution over instrument-response types $p\in\Delta(\mathcal{T}_Z)$ exists such that $p$ satisfies: \begin{align} A_{Z}\boldsymbol{p}&\leq \boldsymbol{r}_Z,\\ A_{X}\boldsymbol{p}&=\boldsymbol{r}_X,\\ A_{Y}\boldsymbol{p}&\leq \boldsymbol{\Psi}. \end{align} \end{enumerate} Where $A_Z,A_X,r_Z$ and $r_X$ are finite dimensional matrices as defined in assumption (ref) and proposition (ref).

The formal proof of theorem (ref) is given in appendix (ref). Here, I provide some intuition on the result and compare it to existing falsification procedures.

To see why the inequality $p(AT_{x_l})\leq \Psi_{x_l}$ must hold for every $x_l\in\mathcal{X}$ note that always-$x_l$ takers are a potentially strict subset of all the units that could potentially realize treatment $x_{l}$ both when $Z=z_0$ and when $Z=z_1$. Therefore, if the exclusion restriction holds, the group of always-$x_l$ takers that realize outcomes in any measurable set $B_Y\subset\mathcal{B}_Y$ must also be a (potentially strict) subset of the units that realize outcomes in $B_Y$ when $X=x_l$ and $Z=Z_k$ for $z_k\in\{z_0,z_1\}$. This observation is illustrated in figure (ref).

To see the sufficiency of the restrictions, suppose that some $p\in\mathcal{T}_Z$ is consistent with observed conditional treatment probabilities and satisfies $p(AT_{x_l})\leq\Psi_{x_l}$ for every $x_\in\mathcal{X}$. To construct a distribution over response types consistent with the observed joint distribution of $(Y,X,Z)$, the unobserved marginal distribution of potential outcomes for always-$x_l$ takers when $X=x_l$ can be set to $\big[\frac{p(AT_{x_l})}{\Psi_{x_l}}\big]\psi_{x_l}(y)$ so that always-$x_l$ takers account for a fraction $\frac{p(AT_{x_l})}{\Psi_{x_l}}$ of the mass represented by $\Psi_{x_l}$. This construction guarantees that the unconditional sub-density associated to the outcomes of always-$x_l$-takers is dominated by $\boldsymbol{P}_{x_l|z_0}\left[f_{Y|X=x_{l},Z=z_{0}}(y)\right]$ and $\boldsymbol{P}_{x_l|z_1}\left[f_{Y|X=x_{l},Z=z_{1}}(y)\right]$.

\FloatBarrier

figure[figure omitted — 2,764 chars of source]

It remains to distribute the fraction $1-\frac{p(AT_{x_l})}{\Psi_{x_l}}$ of the area $\Psi_{x_l}$ as well as any mass strictly above at least one sub-density across response types that realize $x_l$ at just one instrument value (i.e. $x_l$-compliers and $x_l$-defiers). Assigning potential outcomes to these units is simple because --since they realize treatment $x_l$ at just one instrument value-- the exclusion restriction does not bind their potential outcomes. Figure (ref) provides a visual illustration of this procedure. The formal construction is described in appendix (ref).

figure[figure omitted — 2,048 chars of source]

Comparison with kitagawa2015test and kwon2024testing: A benchmark case is when assumption (ref) imposes that treatment $x_l$ can only increase (or decrease) as the instrument changes. This case corresponds to the classical monotone binary instrument and treatment model from angristImbens1994late,kitagawa2015test and mourifie2017testing. It also corresponds to maximal elements in the ordered or unordered case as in sun2023instrument.

Suppose that addoption of $x_l$ weakly increases as $Z$ changes from $z_0$ to $z_1$. Note that in this special case, the function $\boldsymbol{P}_{x_l|z_1}\left[f_{Y|X=x_{l},Z=z_{1}}(y)\right]$ weakly dominates the function $\boldsymbol{P}_{x_l|z_0}\left[f_{Y|X=x_{l},Z=z_{0}}(y)\right]$, therefore, $\psi_{x_l}(y)$ is equal to $\boldsymbol{P}_{x_l|z_0}\left[f_{Y|X=x_{l},Z=z_{0}}(y)\right]$ almost everywhere. Furthermore, the fraction of always $x_l$ takers is identified and coincides with $\Psi_{x_l}$ and $\boldsymbol{P}_{x_l|z_0}$. This particular case is illustrated in figure (ref).

Figure (ref) illustrates why, while seemingly simpler, my approach is equivalent to kitagawa2015test in the binary monotone case. The integral of $\psi_{x_l}(y)$ over $\mathcal{Y}$ constrains the frequency of always takers, this statement is true even if monotonicity does not hold. The test of kitagawa2015test takes the function $\boldsymbol{P}_{x_l|z_0}\left[f_{Y|X=x_{l},Z=z_{0}}(y)\right]$ and asks if, indeed, it coincides with $\psi_{x_l}$ almost everywhere. Instead, I first compute the integral of $\psi_{x_l,\mathcal{Z}}$ without invoking monotonicity. Then, I note that if violations of monotonicity occur with positive probability, $\psi_{x_l,\mathcal{Z}}$ will be strictly smaller than $\boldsymbol{P}_{x_l|z_0}\left[f_{Y|X=x_{l},Z=z_{0}}(y)\right]$ in some positive measure set. Thus, the two integrals are equal if and only if random assignment, exclusion, and monotonicity are not violated.

For the non-monotone case, kwon2024testing note that the population frequency of $x_l$-defiers (units that realize $X_i=x_l$ only when $Z_i=z_0$) is bounded below by the supremum over measurable sets $B_Y\in\mathcal{B}_Y$ of the difference $\mathbb{P}[Y_i\in B_Y,X_i=x_l|Z_i=z_0]-\mathbb{P}[Y_i\in B_Y,X_i=x_l|Z_i=z_0]$. Note that this supremum is attained in the set $\left\{y\in\mathcal{Y}: \boldsymbol{P}_{x_l|z_0}\left[f_{Y|X=x_{l},Z=z_{0}}(y)\right]\geq \boldsymbol{P}_{x_l|z_1}\left[f_{Y|X=x_{l},Z=z_{1}}(y)\right]\right\}$ and its value coincides with the quantity $\mathbb{P}[X_i=x_l,Z_i=z_0]-\Psi_{x_l}$. Since at $Z=z_0$ only $x_l$-defiers and $x_l$-always takers can realize $x_l$, the upper bound on always takers I derive implies and is implied by the lower bound on compliers from kwon2024testing. The difference is that instead of taking a supremum over a potentially infinite collection of measurable sets, I obtain an equivalent bound by integrating the point-wise minimum of two sub-densities.

Generalized FOSD and response type restrictions

Sharp observable implications of response type restrictions for binary instruments

In this section I specialize theorem (ref) to instrument-response-type restrictions, i.e. restrictions that rule out from the population specific instrument-response types. The most common example of a restriction in this class is the no-defiers assumption from angristImbens1994late. The specialized result is a set of sharp inequalities that only depend on point identified parameters.

The resulting inequalities can be interpreted as a generalization of FOSD. This interpretation allows me to link response type restrictions to non-parametric random utility models. Specifically, I show that any instrument-response type restriction can be derived from the assumption that, as the instrument changes from $z_0$ to $z_1$, it affects certain pairwise comparisons in one direction and therefore prevents specific preference reversals over treatments.

The refined inequalities come from representing the instrument validity conditions as a flow problem. The observed conditional treatment probability $\mathbb{P}[X_i=x_l|Z_i=z_0]$ represents the share of the population that realizes treatment $x_l$ at instrument value $z_0$. In the absence of instrument-response type restrictions, these units could realize any other treatment at the instrument value $z_1$. But instrument-response type restrictions rule out certain potential treatments for units that realize treatment $x_l$ when $Z_i=z_k$. That is, flows from $x_k$ to certain treatment values $\tilde{X}_{x_k}\subset \mathcal{X}$ are not allowed. \footnote{This restriction is reminiscent of the minimal monotonicity idea discussed by navjeevan2022ordered, but it need not apply to every pair of treatment values, just to a subset.}

The refined observable implications come from taking the flow idea literally. Finding a valid distribution over response types is equivalent to finding a valid set of flows that, for every $x_l\in\mathcal{X}$, takes the share $\mathbb{P}[X_i=x_l|Z_i=z_0]$ of the population and distributes it across potential counterfactual treatments $X_i(z_1)$ in a way that matches the observed treatment probabilities conditional on $Z_i=z_1$ without recurring to the flows that the instrument-response type restriction forbids.

To formally define the transportation problem I require additional notation and definitions. I begin by formally defining response type restrictions.

assumption[Instrument-response type restriction]\phantom{a}\\ Let $R\subset\mathcal{T}_Z=\mathcal{X}\times \mathcal{X}$ be a fixed and known collection of instrument-response types. A distribution over response types $g_0\in\Delta(\mathcal{T})$ exists such that $p_{g_0}(t^Z)=0$ for all $t^Z\in R$.

Put in words, a response type restriction simply imposes that some instrument-response types occur with zero probability.

Now define the sets $\mathcal{X}_k$ for $k=0,1$ as copies of $\mathcal{X}$ but with elements $x_l^k\in\mathcal{X}_k$ labeled with the superscript $k$ to make explicit that $x^0_l$ refers to treatment $x_l\in\mathcal{X}$ realized when $Z=0$ and $x^1_l$ to treatment $x_l\in\mathcal{X}$ realized when $Z=1$.

Next, define the directed graph with vertex set $\mathcal{X}_0\cup \mathcal{X}_0$ and let the edges be $\mathcal{E}_R\subset\mathcal{X}_{1}\times\mathcal{X}_{2}$, where the ordered pair $(x_l^0,x_{l'}^1)\in\mathcal{X}_0\times\mathcal{X}_1$ belongs to $\mathcal{E}_R$ if and only if the instrument-response type $t^Z=(x_l,x_{l'})$ does not belong to $R$. In order words, an edge $(x_l^0,x_{l'}^1)$ belongs to the graph if the response type that realizes $x_l$ at $Z=0$ and $x_{l'}$ at $Z=1$ is not ruled out by assumption (ref).

For instance, in the case of the no defiers assumption, the set $R$ would consist of the singleton $\{(x_1,x_0)\}$. Therefore, the edges $(x_0^0,x_0^1), (x_0^0,x_1^1)$ and $(x_1^0,x_1^1)$ --corresponding to never takers, compliers, and always takers-- would belong to $\mathcal{E}_R$, in contrast, the edge $(x_1^0,x_0^1 )$ would not. This particular case is illustrated in figure (ref).

\FloatBarrier

figure[figure omitted — 669 chars of source]

\FloatBarrier

The transportation problem is as follows: every node $x_l^0\in\mathcal{X}_0$ starts with a mass $\mathbb{P}[X_i=x_l|Z_i=z_0]$ of units. Each unit can travel from its origin node $x_l^0$ through the admissible edges $(x_l^0,x_{l'}^1)\in\mathcal{E}_R$ in order to reach some destination node $x_{l'}^1\in\mathcal{X}_1$ subject to the constraint that each destination node $x_{l'}$ cannot receive more than $\mathbb{P}[X_i=x_{l'}|Z_i=z_1]$ units. If it is possible to empty all origin nodes $x_l^0$ so that all destination nodes end at full capacity, then, the transportation problem is solved.

If the transportation problem is solved, a distribution over response types exists that is consistent with both the observed joint distribution of $(X,Y)$ and the imposed response type restriction. In appendix (ref) I provide the formal definition of the transportation problem and prove the equivalence between the existence of a feasible solution and a valid distribution over response types. The central argument, however, is simple: There is a one to one mapping between response types in $R$ and edges in $\mathcal{E}_R$. From here, it is easy to see that the value of the flow that travels through edge $e$ according to a solution to the transportation problem corresponds to the probability that a valid distribution over instrument-response types assigns to the unique instrument-response type associated to $e$.

Two well known combinatorial problems are equivalent to the transportation problem described: the maximum flow and minimum cut problems. \footnote{The equivalence is also proven in appendix (ref). The entire reasoning would also apply to more general restrictions such as those considered in assumption (ref). The only difference would be that more general restrictions cannot necessarily be stated as capacity constraints on the edges of the graph. This impossibility would complicate using the equivalence with minimum-cut problem to obtain observable inequalities that do not depend on unobservables.} The minimum cut formulation is particularly useful to derive the sharp observable implications of incomplete models in terms of inequalities that only depend on observable quantities as first noted by ekeland2010optimal and galichon2011set due to its equivalence to plausibility constraints in the context of random set theory (beresteanu2012partial).

Before stating the observable implications, I make explicit the one to one correspondence between response type restrictions and binary relations that exists in the binary instrument case, and introduce a generalization of the lower contour to arbitrary binary relations.

definition[Binary relation associated to $R\subset\mathcal{T}_Z$]\phantom{a}\\ Let $R\subset\mathcal{T}_Z$ be a collection of response types. Define the binary relation $\geq^R\subset\mathcal{X}\times\mathcal{X}$ as follows: \begin{align} x_l\geq^R x_{l'}\Longleftrightarrow (x_l,x_{l'})\in R \end{align}

That is, $x_l$ is related to $x_{l'}$ according to $\geq^{R}$ if switches from $x_l$ to $x_{l'}$ when the instrument goes from $z_0$ to $z_1$ are ruled out by the instrument-response type restriction associated to the set $R$.

definition[Common lower contour of $\geq^R$]\phantom{a}\\ Let $\geq^R$ be a binary relation over $\mathcal{X}$. For any $S\subset\mathcal{X}$, the common lower contour of $S$ according to $\geq^R$ is: \begin{align} L_{\geq^R}(S)=\{x_k\in\mathcal{X}:s \geq^R x_{k},\forall s\in S\} \end{align}

In the context of instrument-response types, $L_{\geq^R}(S)$ is the (potentially empty) subset of elements of $\mathcal{X}$ such that no flows from an element in $S$ to any element in $L_{\geq^R}(S)$ are allowed. Thus, its complement $L_{\geq^R}(S)^C$ can be interpreted as the set of treatments that units that start with treatment values in $S$ are weakly incentivized to adopt as the instrument changes from $z_0$ to $z_1$. Note that for total orders (or their linear extensions), the set $L_{\geq^R}(S)$ coincides with the standard definition of the upper contour of $S$ and the set $L_{\geq^R}(S)^C$ with the upper contour.

Now I state the sharp observable implications of response type restrictions for binary instruments:

theorem\phantom{a}\\ A distribution over instrument-response types $p\in\Delta(\mathcal{T}_Z)$ that satisfies assumption (ref) (instrument validity), assumption (ref) (instrument-response type restriction), and definition (ref) (consistency) exists, if and only if: \begin{enumerate} • The following inequality holds for every $S\subset\mathcal{X}$. \begin{align} \mathbb{P}[X_i\in S |Z_i=z_0]\leq \mathbb{P}[X_i\in L_{\geq^{R}}(S)^C|Z_i=z_1] \end{align} • For any $S\subset\mathcal{X}$ and any partition $(\Lambda_S,\Lambda'_S)$ of $L_{\geq^R}(S)^C$ satisfying (i) $\emptyset\neq\Lambda'_S\subset S$, (ii) $L_{\geq^R}(x_{l'})^C\setminus\{x_{l'}\}\subset \Lambda_S$ for every $x_{l'}\in \Lambda_S'$, and (iii) $L_{\geq^R}(S\setminus \Lambda_S')^C\subset \Lambda_S$: \begin{align} \mathbb{P}[X_i\in S |Z=z_0]\leq \mathbb{P}[X_i\in \Lambda_S|Z=z_1]+\sum\limits_{x_{l'}\in \Lambda'_S} \Psi_{x_{l'}} \end{align} Where $\Psi_{x_l}$ is defined as in the previous section. \end{enumerate}

The proof of theorem (ref) is given in appendix (ref). One interpretation of the inequalities from part one is that mutually exclusive events cannot occur in the population with probability greater than one, this condition is implicitly verified by the minimum cut problem. A generalization of a logically equivalent idea beyond the binary instrument case was recently developed by kaido2025testing.

Another way to interpret the inequalities from part one of theorem (ref) is generalized FOSD. The response type restriction associated to $R$ implicitly assumes that, if $x_l\geq^R x_{l'}$, then treatment $x_l$ is incentivized by the instrument relative to $x_{l'}$. Thus, the relative likelihood of observing $x_{l}$ relative to $x_{l'}$ must increase. As such, this set of inequalities can be independently used as a sharp test for the validity of response type restrictions or, as formalized in the next sub-section, to test non parametric hypotheses about how an exogenous variable affects the preferences of a population of decision makers.

The second part of theorem (ref) notes that the exclusion restriction and the observed conditional distribution of the outcome bind the prevalence of always-$x_l$ takers, and, in some cases, these constraints on always takers, tighten the FOSD inequalities by further limiting the maximum of units that could potentially realize treatments in certain upper contours.

The fact that the exclusion restriction tightens the generalized FOSD inequalities by replacing some conditional treatment probabilities $\mathbb{P}[X_i=x_{l'}|Z=z_1]$ with the weakly smaller quantities $\Psi_{x_{l'}}$ allows to distinguish violations from the exclusion restriction from violations of instrument-response restrictions. In particular, one of the following three cases must occur:

enumerate• All inequalities are satisfied. Hence, neither the exclusion restriction nor the response-type restriction is falsified. • Only inequalities from part 2 are violated. Hence, the response type restriction is not falsified on its own, but only when combined with the exclusion restriction. • At least one inequality from part 1 is violated. Hence, the response type restriction is falsified

Additionally, every inequality is linked to exactly one set $S$. Thus, if the inequalities associated with $S$ are not satisfied, it must be that the assumed restriction on how the instrument affects the potential treatments (or preferences over treatments) of units that realize treatments in $S$ when $Z_i=z_0$ is falsified, while restrictions associated to other treatment switches might not.

One final advantage of theorem (ref) relative to theorem (ref) is that it yields a linear system of inequalities with no nuisance parameters.

To illustrate its usefulness, I use theorem (ref) to characterize the binary monotone instrument IV model from angrist1995two.

corollary[Sharp observable implications of binary monotone instruments] Suppose that $\mathcal{X}=\{x_0,x_1,...,x_{L}\}$, with $x_{l}<x_{l'}$ for every $l<l'$. A distribution over response types exists that satisfies assumptions (ref) and: \begin{align} X_i(z_0)\leq X_i(z_1) \phantom{aaa} a.s. \end{align} If and only if: \begin{align} \mathbb{P}[X\geq x_l|Z=z_0]\leq \Psi_{x_l} + \mathbb{P}[X> x_{l+1}|Z=z_1] \phantom{aaa}\forall x_l\in\mathcal{X} \end{align}

Note that in the ordered monotone case, the set $L_{\geq}(x_l)^{C}$ corresponds to the standard upper contour of $x_l$ (i.e. $\{x_{l},x_{l+1},...,x_L\}$), thus, the sharp inequalities reduce to FOSD but tightened by the exclusion restriction in the sense that the probability $\mathbb{P}[X= x_{l}|Z=z_1]$ associated to the minimal element of every $L_{\geq}(x_l)^{C}$ (i.e. $x_l$) is replaced by the weakly smaller quantity $\Psi_{x_l}$. \footnote{As a consequence of this result, it is easy to see that the conditions from sun2023instrument do not incorporate the effect that the exclusion restriction has on upper contours associated to non minimal/maximal elements. This illustrates why the observable implications from sun2023instrument are not sharp for monotone instruments, and why they are nested by the ones proposed in this paper.}

A submonotone interpretation for response type restrictions

If treatment is assumed to be chosen by rational decision makers, the connection between binary relations, response type restrictions, and first order stochastic dominance extends beyond realized treatments to utility maximization. This relation does not depend on functional or distributional assumptions --such as additive separability. Instead, it can be directly expressed in terms of latent potential preferences.

assumption[$\geq^{*}$-Submonotonicity] Let $\geq^{*}$ be an irreflexive \footnote{The extension to reflexive binary relations is conceptually simple. The reasoning remains the same but requires special treatment of response type restrictions that rule out “always takers” of one particular treatment. In the interest of brevity this case is addressed in appendix (ref).} binary relation over $\mathcal{X}$. Suppose that: \begin{enumerate} • Units are endowed with strict preference profiles over $\mathcal{X}$ that depend on $Z_i$ denoted $\succ_i^{(z_k)}$ for $k=0,1$ where the function $\succ_i^{(z_k)}$ maps $\mathcal{Z}$ to strict preference profiles over $\mathcal{X}$. In other words, potential preferences. • Units are rational in the sense that $X_i(z_k)=x_l$ implies $x_l \succ_i^{(z)} x_{l'}$ almost surely for all $l,l'\in\{0,1,...,L\}$ such that $l\neq l'$. • There are no preference reversals against $\geq^*$ when the instrument changes from $z_0$ to $z_1$. That is, if $x_l\phantom{a}\geq^*x_{l'}$ and $x_l\succ_i^{(z_0)} x_{l'}$, then, $x_l\succ_i^{(z_1)} x_{l'}$ almost surely. \end{enumerate}

The first two points of assumption (ref) simply posit that response types arise from latent utility maximization. The third point captures the the notion that if $x_l\geq^*x_{l'}$, then alternative $x_l$ is being weakly incentivized or promoted (not necesarilly relative to all other alternatives but relative to $x_{l'}$) by the instrument.

I call assumption (ref) submonotonicity because it encodes a notion weaker than monotonicity; it rules out two-way flows between certain pairs of treatments, but not necessarily between all. In this sense, any submonotonicity assumption can always be extended to a full monotonicity condition with respect to a total order (or the linear extension of a total order) by incorporating additional restrictions. For instance, both ordered and unordered monotonicity correspond to special cases of submonotonicity where the binary relation is transitive. The notion of preventing “two-way flows” from navjeevan2022ordered corresponds to assuming that the binary relation in question is asymmetric.

In Appendix (ref) I explore in depth the interpretation of different assumptions within the submonotonicity class and provide multiple motivating examples. These examples include extensions of unordered monotonicity (heckman2018unordered); assumptions that draw from both ordered and unordered monotonicity (rose2024recoding); specific cases of targeting and encouragement designs as well as some possible relaxations and generalizations (lee2020treatment, bai2022testing); and designs in which the instrument encourages variety and experimentation, or its effect depends on status-quo ($Z=z_0$) choices.

theoremLet $\geq^{*}$ be an irreflexive binary relation over $\mathcal{X}$. Then: \begin{enumerate} • If $\geq^{*}$-Submonotonicity holds, then, $p_{g_0}(t^Z)=0$ for every $t^Z\not\in R$. • Suppose that $R\subset\mathcal{T}$ is such that $\geq^{R}=\geq^*$. Then, $R$ is the smallest cardinality subset of $\mathcal{T}_Z$ such that if $p_{g_0}(t)=0$ for all $t\not\in R$, then, $\geq^{*}$-Submonotonicity cannot be falsified. \end{enumerate}

The assumption that $\geq^R$ is irreflexive is not necessary, but it substantially simplifies exposition. The proof of theorem (ref) is straightforward and can be found in appendix (ref) where the general case is covered.

Theorem (ref) shows that $\geq^R$-submonotonicity is sufficient to guarantee that instrument-response types outside of $R$ occur with zero probability. Moreover, $R$ is the minimal (in the sense of set inclusion) restriction over instrument-response types, such that $\geq^R$-submonotonicity is not falsified. Thus, the set of inequalities from (ref) comprises the sharp observable implications of $\geq^{R}$-submonotonicity.

The multiple instruments case

Sufficient takers: A finite set of necessary conditions

The bounds on always takers from the previous sections can be naturally extended to the multiple instruments case. In this more general setting, the exclusion restriction imposes upper bounds on the prevalence of multiple groups of instrument-response types in addition to always takers. The following definition formalizes such instrument-response type groups.

definition[Sufficient-$x_l$-takers] For any treatment value $x_l\in\mathcal{X}$ and any non-empty subset of instrument values $\tilde{Z}\subset\{z_0,z_1,...,z_{K_{Z}-1}\}$ define the set of response types that realize $x_l$ whenever $Z$ takes any value in $\tilde{Z}$ as: \begin{align} \tilde{S}(x_l,\tilde{Z}):=\{t^Z:(t^Z)_k=x_l \phantom{a} \Longrightarrow z_k\in\tilde{Z}, \} \end{align}

For any subset $\tilde{Z}\subset\mathcal{Z}$ with $|\tilde{Z}|\geq 2$, define the function $\psi_{x_l\tilde{Z}}(y)$ and the constant $\Psi_{x_l,\tilde{Z}}$ as follows:

align[align omitted — 234 chars of source]

As in the binary instrument case, the function \( \psi_{x_l, \tilde{Z}}(y) \) represents the point-wise minimum of the sub-distributions \( \phi_{Y \mid x_l, z_{\tilde{k}}} \) over all \( z_k \in \tilde{Z} \). The constant \( \Psi_{x_l, \tilde{Z}} \) measures the area that lies below all these subdensities (the overlap).

propositionSuppose assumption (ref) holds, then: \begin{align} \mathbb{P}[Y_i\in B_Y,X_i(z_k)=x_l\phantom{a}\forall z_k\in\tilde{Z}] \leq \int\limits_{B_Y}\psi_{x_l,\tilde{Z}}(y)d\mu_Y \phantom{aaa} \forall B_Y\in\mathcal{B}_Y \end{align}

Therefore, for any collection of instrument values $\tilde{Z'}$, setting $B_Y=\mathcal{Y}$ yields an upper bound on the probability that a unit belongs to the sufficient-$x_l$-takers group $\tilde{S}(x_l,\tilde{Z})$. In terms of $\boldsymbol{p}\in\Delta(\mathcal{T}_Z)$:

align[align omitted — 197 chars of source]

Once more, the previous equation comprises a finite system of ineequalities that are linear in $\boldsymbol{p}$. Therefore, it can be expressed in matrix form as $A_\Psi\boldsymbol{p}\leq \boldsymbol{\Psi}$ where $\boldsymbol{\Psi}$ stacks all the constants $\Psi_{x_l,\tilde{Z}}$ across treatment values and subsets of the support of the instrument.

corollaryA valid distribution over instrument response types must satisfy: \begin{align} A_\Psi\boldsymbol{p}&\leq \boldsymbol{\Psi}\\ A_{Z}\boldsymbol{p}&\leq \boldsymbol{r}_Z\\ A_{X}\boldsymbol{p}&=\boldsymbol{r}_X \end{align}

Conclusion

This paper proposes a set of observable implications for IV models. For binary instruments, the observable implications are sharp. For binary instruments with instrument-response-type restrictions I provide a characterization with a generalized FOSD interpretation and use this interpretation to link instrument-reponse-type restrictions to non-parametric assumptions on the preferences of units in the population. The testable implications generalize naturally to the multiple instruments case.

\singlespacing

\onehalfspacing