EconBase
← Back to paper

When does IV identification not restrict outcomes?

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,879 characters · 27 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

When does IV identification not restrict outcomes?

abstractMany identification results in instrumental variables (IV) models hold without requiring any restrictions on the distribution of potential outcomes, or how those outcomes are correlated with selection behavior. This enables IV models to allow for arbitrary heterogeneity in treatment effects and the possibility of selection on gains in the outcome. I provide a necessary and sufficient condition for treatment effects to be point identified in a manner that does not restrict outcomes, when the instruments take a finite number of values. The condition generalizes the well-known LATE monotonicity assumption, and unifies a wide variety of other known IV identification results. The result also yields a brute-force approach to reveal all selection models that allow for point identification of treatment effects without restricting outcomes, and then enumerate all of the identified parameters within each such selection model. The search uncovers new selection models that yield identification, provides impossibility results for others, and offers opportunities to relax assumptions on selection used in existing literature. An application considers the identification of complementarities between two cross-randomized treatments, obtaining a necessary and sufficient condition on selection for local average complementarities among compliers to be identified in a manner that does not restrict outcomes. I use this result to revisit two empirical settings, one in which the data are incompatible with this restriction on selection, and another in which the data are compatible with the restriction.

Introduction

To leverage instrumental variables (IV) with heterogeneous treatment effects, researchers often make assumptions about selection behavior, such as the “monotonicity” assumption of the seminal local average treatment effects (LATE) model Imbens2018. Given this monotonicity assumption, the average treatment effect among compliers is point identified, with no restrictions imposed on the distribution of potential outcomes beyond them being independent of the instrument.

This paper shows that a similar result holds broadly across IV models. If restrictions on selection behavior are sufficient to establish a particular generalization of the monotonicity assumption, then corresponding local average treatment effects are identified. Strikingly, when the treatments and instruments have finite support, this identification result also has a converse: for any treatment effect that conditions on selection behavior to be point identified without further restrictions on the distribution of potential outcomes, the selection model must permit the generalization of monotonicity to hold.

Together this yields a necessary and sufficient condition for when identification of a given treatment effect or counterfactual mean is possible without ad-hoc assumptions regarding outcomes such as treatment effect homogeneity. I say that a parameter is identified in an outcome-nonrestrictive way when it is point identified without any restrictions on the distribution of potential outcomes, beyond statistical independence between those outcomes and the instruments. This paper shows that quite generally, the price of outcome-nonrestrictive identification is the need to make assumptions about selection into treatment. This trade-off may be quite appealing, particularly in “design-based” studies in which the researcher has contextual knowledge about a factor that affects treatment uptake in a given setting carddesignbased, but may be reluctant to make assumptions about the very causal effect of interest (e.g. that it is homogeneous across individuals).

Outcome-nonrestrictive identification results also have the practical benefit of paving the way for analysis to be repeated across multiple outcome variables, with maintained assumptions about selection in a given (natural) experiment aiding identification for each outcome. An outcome-nonrestrictive identification result is not fully indifferent to which variable one uses as an outcome: one must make the standard independence and exclusion restrictions for each one. But in settings where treatment is as-good-as-randomly assigned and there are limited opportunities for the assignment to affect anything except via treatment (e.g. in certain experimental settings), such assumptions may be quite natural without much further justification specific to each outcome.

When specialized to the case of a binary treatment, the main result of this paper can be stated succinctly as follows. Let $D_i(z)$ denote the counterfactual treatment of individual $i$ when the available instruments take value $z$, for each $z$ in a set $\mathcal{Z}$.

result*[(in the case of binary treatments)] Suppose $|\mathcal{Z}|$ is finite, and consider any subgroup of the population defined by their counterfactual selection behavior $D_i(\cdot)$. Then the average treatment effect among this subgroup is identified in an outcome-nonrestrictive way if and only if there exists a function $\alpha: \mathcal{Z} \rightarrow \mathbbm{R}$ such that \begin{equation} \left\{\sum_{z} \alpha(z) \cdot D_i(z)\right\} \in \{0,1\} \quad for all i, \end{equation} where the subgroup is composed of the individuals $i$ for whom the term in brackets is equal to one. The “if” direction of the above further assumes that $\sum_{z} \alpha(z)=0$, and this restriction is needed for the “only if” direction if there are always-takers or never-takers.

The above result is formalized in Theorem (ref) (which establishes sufficiency) and Theorem (ref) (which establishes necessity) of this paper, where each result is provided more generally for any finite set of treatments that are not necessarily binary, and is established for means of each potential outcome alone in addition to holding for treatment effects.\footnote{To see how Eq. (ref) nests the LATE identification result of Imbens2018 in the case of a single binary instrument, take the function $\alpha(z) = (-1)^{z+1}$, so that $\alpha(0)=-1$ and $\alpha(1)=1$. The quantity $\sum_{z} \alpha(z) \cdot D_i(z) = D_i(1)-D_i(0)$ is equal to either $0$ or $1$ for all individuals $i$, and is equal to unity only for the compliers.}

When generalized beyond the binary-treatment case, the above result delivers a simple geometric characterization of outcome-nonrestrictive identification in IV models. Units in the population can be partitioned by which “response type” $G_i \in \mathcal{G}$ they belong to, and the conditioning event that defines a target parameter by functions $c: \mathcal{G}\rightarrow \{0,1\}$ that indicates which response types are considered by that parameter. In models with a finite selection model $\mathcal{G}$, a local counterfactual mean $\mathbbm{E}[Y_i(t)|c(G_i)=1]$ is point-identified in an outcome-nonrestrictive manner if and only if a vector representation of $c$ belongs to a particular linear subspace of $\mathbbm{R}^{|\mathcal{G}|}$, where the subspace depends on what restrictions are assumed about selection behavior through the choice of $\mathcal{G}$. Intuitively, “linearity” arises throughout from the law of iterated expectations over the latent types $g$. While a counterfactual mean $\mathbbm{E}[Y_i(t)|c(G_i)=1]$ can be expressed as a convex linear combination of $\mathbbm{E}[Y_i(t)|G_i=g]$ over the $g$ for which $c(g)=1$, the observable distribution of $Y$ conditioned treatment realization $t$ similarly amounts to a mixture over conditional distributions of $Y_i(t)$ that condition on the various response types.\footnote{A similar structure is exploited in previous work that has used linear programming approaches to identification under LATE monotonicity (e.g. MTS). These authors generally focus on partial identification under specific monotonicity assumptions (see also Mogstad2020b,kamat2023identification for related results). I characterize point identification for arbitrary selection models.}

The perspective of outcome-nonrestrictive identification turns out to unify a wide variety of existing IV identification results in the literature. Theorem (ref) provides a simple and common proof of point identification for settings including: i) the original LATE model Imbens2018; ii) the marginal treatment effect (Heckman2001, Heckman2005) and its generalization to multivalued treatments Lee2018a; iii) unordered monotonicity Heckman2018; iv) vector monotonicity with multiple instruments goff2024vector; v) restrictions on choice and/or knowledge of second-best options kirkeboenleuvenmogstad; vi) interaction effects between two treatments blackwell2017; and vii) recent notions of monotonicity that are only required to hold between particular pairs of instrument values sun2024pairwise,cclate, sigstad2024marginal. I contrast the above results with other identification results from the IV literature that weaken assumptions about selection while leveraging additional assumptions about outcomes and are therefore not outcome-nonrestrictive imbensangristrubin,Kolesar2013,toleratingdefiance,comey2023supercompliers.

In the other direction, the necessary part of my result (Theorem (ref)) motivates a comprehensive search over all outcome-nonrestrictive identification results, which is feasible in settings in which the instruments and treatments have small support.\footnote{Further, I provide code that enables quick enumeration of identified parameters given a user-provided selection model.} For a given finite set of support points of the instrument and treatment variables, the set of possible selection models is finite and can in principle be enumerated by brute force. Within each such selection model, I show that the set of possible functions $\alpha(\cdot)$ can also be enumerated. We can therefore search systematically over the opportunities for IV identification that place no modeling restrictions on outcomes and have not yet been revealed in the literature. Section (ref) proposes two algorithms that implement this insight, which I apply to uncover all outcome-nonrestrictive identification results for binary or ternary instruments and treatments. In many cases the selection models underlying these results can be empirically motivated, and in some cases they offer opportunities to relax assumptions made in existing work. Section (ref) presents an extended application of this search to the identification of interaction effects in cross-randomized experiments. This setting illustrates how the computational approach can distill a unified and economically meaningful model of selection that is necessary (and not just sufficient) for identification.

In establishing sufficient and necessary conditions for identification in IV models, the perspective of this paper is related to recent results by \citet*{navjeevan2023identification} (NPS). NPS begin with an IV setup that is similar to that of this paper, but consider the identification of unconditional moments of latent heterogeneity in general rather than focusing on moments of potential outcomes that condition on selection. The conditions for identification of “conditional” and “unconditional” versions of such moments are not equivalent in general, but my Theorem (ref) establishes that they are when the former is outcome-nonrestrictive. In Appendix (ref) I show that while my Theorem (ref) can be obtained as a corollary to results found in NPS, the same is not true of Theorem (ref). Without Theorem (ref), we would lack a guarantee that the search for outcome-nonrestrictive identified local average treatment effects executed in this paper is exhaustive.

The structure of the paper is as follows. Section (ref) begins by formalizing a definition of the notion of “outcome-nonrestrictive” identification in IV models. Section (ref) introduces the idea of binary combinations, a name I give to instances in which an analog of (ref) holds for general unordered discrete treatments. I show there that whether the instruments are discrete or continuous, binary combinations are sufficient for outcome-nonrestrictive identification of local counterfactual means. Further, collections of binary combinations across different treatment values (but isolating the same response types) are sufficient to identify local average treatment effects. I call such collections of binary combinations binary collections. Appendix (ref) details how the notion of binary collections nests a broad range of identification results for treatment effects from the literature.

Section (ref) then specializes to the case of discrete and finite instruments, and shows that in such settings binary combinations and binary collections are the only cases in which local counterfactual means or local average treatment effects can be identified in an outcome-nonrestrictive way. Building on an algebraic characterization of binary combinations, Section (ref) presents new results on the identification of various local average treatment effect parameters after enumerating all binary collections by brute-force, with some detailed examples examined in Appendix (ref). A full catalog of such identification results is presented in Appendix (ref), for settings with small instrument/treatment support.

Section (ref) turns to a specific application of the results to the identification of interaction effects in order to assess the “complementarity” between two treatments in cross-randomized designs. I find that the average interaction effect between the treatments can be identified among a complier subgroup without restricting outcomes if and only if the selection model allows just five response types, composed of the individuals who decide “separately” whether to select into each of the two treatments. I use this result to revisit two empirical applications: one in which the data are incompatible with observable implications of this selection model, and another in which the data are. In this application, I proceed to estimating the extent of complementarity, yielding new evidence on the interaction between pharmacotherapy and livelihoods assistance in combatting depression.

Defining outcome-nonrestrictive IV identification

Setup and notation

Let treatment $t$ take values in a finite set $\mathcal{T}$. Denote potential outcomes as $Y_i(t,z)$ and potential treatments as $T_i(z)$, where $Z_i$ are instruments with support $\mathcal{Z}$. I'll refer to $Z_i$ as “the instruments”, since in general it can be a vector of instrumental variables. The index $i$ corresponds to observational units, i.e. “individuals”.

IV model assumptions

Let $D^{[t]}_i(z) = \mathbbm{1}(T_i(z) = t)$ be an indicator for $i$ taking treatment $t$ when the instruments are equal to $z$. I throughout impose the exclusion restriction that $Z_i$ only affects $Y_i$ through $T_i = T_i(Z_i)$, so that $Y_i(t,z)=Y_i(t)$ for all $i$ and $z \in \mathcal{T}$. The observed outcome $Y_i$ is then:

equation[equation omitted — 108 chars of source]

Let $G_i: \mathcal{Z} \rightarrow \mathcal{T}$ be $i$'s “response type”, i.e. the function that yields individual $i$'s counterfactual treatment value for each possible instrument value $z$. Let $\mathcal{G}$ be the set of all admissible response types. Let us denote the full set of conceivable functions from $\mathcal{Z}$ to $\mathcal{T}$ as $\mathcal{T}^{\mathcal{Z}}$. Any $\mathcal{G} \subset \mathcal{T}^{\mathcal{Z}}$ reflects a restriction on the response types, which I refer to as a selection model or choice model.

I will also assume throughout that the instruments are exogenous in the sense that

equation[equation omitted — 81 chars of source]

where $\tilde{Y} = \{Y_i(t)\}_{t \in \mathcal{T}}$ is a vector of potential outcomes across all treatments $t$. Eq. (ref) says that potential outcomes and potential treatments are jointly independent of the instruments. In applications, researchers often defend a conditional version of (ref), i.e. $\{Z_i \perp\!\!\!\!\perp (\tilde{Y}_i,G_i)\} | X_i$, where $X_i$ are observed covariates unaffected by treatment. Since my focus in this paper is on identification and not estimation of treatment effects, I suppress throughout conditioning on any such covariates for ease of exposition, and consider them in Appendix (ref).

Intuitively, the notion of outcome-nonrestrictive identification amounts to identification of a causal parameter that holds without restrictions on the distribution of $(\tilde{Y}_i,G_i)$, apart from (ref) and that $supp\{G_i\} \subseteq \mathcal{G}$, where $supp\{G_i\}$ is the support of the response types $G_i$. However, defining the notion of outcome-nonrestrictive identification in a formal way requires accounting for possible restrictions on observables implied by a given selection model $\mathcal{G}$. The remainder of this Section (subsections (ref) and (ref) below) develops some notation to give this formal definition, before turning to the first main result in Section (ref). Note that neither the definition of outcome-nonrestrictive identification---not my results concerning it---restrict the support of $Y_i$ (e.g. that it be discrete or bounded).

Observable restrictions implied by the model

Let $\mathcal{P}$ denote the joint distribution of the model fundamentals $(G_i,\tilde{Y}_i,Z_i)$. Given (ref), we can decompose $\mathcal{P}$ as $$\mathcal{P} = \mathcal{P}_{latent} \times \mathcal{P}_Z,$$ where $\mathcal{P}_Z$ denotes the distribution of the instruments $Z_i$ and $\mathcal{P}_{latent}$ denotes the distribution of the latent (counterfactual) variables of the model $\tilde{Y}$ and $G$.\footnote{By $\mathcal{P} = \mathcal{P}_{latent} \times \mathcal{P}_Z$, I mean that for any Borel set $\mathcal{B}_{L}$ of values for $(G_i,\tilde{Y}_i)$ and $\mathcal{B}_{Z}$ of values for $\mathcal{P}_Z$ we have $\mathcal{P}(\mathcal{B}_{L} \times \mathcal{B}_{Z}) = \mathcal{P}_{latent}(\mathcal{B}_{L}) \cdot \mathcal{P}_{Z}(\mathcal{B}_{Z})$, where $\mathcal{B}_{L} \times \mathcal{B}_{Z}$ is the Cartesian product of $\mathcal{B}_{L}$ and $\mathcal{B}_{Z}$.}

A generic causal parameter of interest $\theta$ is a functional $\theta(\mathcal{P})$ of the distribution $\mathcal{P}$ of model variables. Let $\mathcal{P}_{obs}$ denote the distribution of observable variables $(Y_i,T_i,Z_i)$. Note that $\mathcal{P}_Z$ is a marginalization of $\mathcal{P}_{obs}$ over $Y_i$ and $T_i$. I make use of the following notational convention: for a sub-vector $W_0$ of a random vector $W$, let $\mathcal{P}_{W_0}(\mathcal{P}_{W})$ be the distribution of $W_0$ that arises after marginalizing distribution $\mathcal{P}_W$ over the components of $W$ not included in $W_0$. In this notation, for example, $\mathcal{P}_Z = \mathcal{P}_Z(\mathcal{P}_{obs})$.

Define $\mathscr{P}_{latent}(\mathcal{G})$ to be the set of $\mathcal{P}_{latent}$ compatible with a given selection model $\mathcal{G}$ and admitting of finite moments:

equation[equation omitted — 207 chars of source]

where we let $\mathscr{P}_{\tilde{Y}G}$ denote the set of all distributions over $(\tilde{Y}_i,G_i)$, such that $\mathbbm{E}[Y_i(t)|G_i=g]$ exists and is finite for each $t \in \mathcal{T}$ and $g \in \mathcal{G}$. Employing a similar notation, we let $\mathscr{P}_Z$ be the set of distributions over instrument values that embed any maintained support restrictions (e.g. that $Z_i$ is binary with $P(Z_i=1) \in (0,1)$).

Note that for any $\mathcal{P} = \mathcal{P}_{latent} \times \mathcal{P}_Z$, Eq. (ref) and $T_i=T_i(Z_i)$ imply a distribution of observables. Let $\phi$ denote this map so that $ \mathcal{P}_{obs} = \phi(\mathcal{P})$. The set of possible distributions of observables given a selection model $\mathcal{G}$ is $$\mathscr{P}_{obs}(\mathcal{G}) := \{\phi(\mathcal{P}_{latent} \times \mathcal{P}_Z): \mathcal{P}_{latent} \in \mathscr{P}_{latent}(\mathcal{G}), \mathcal{P}_Z \in \mathscr{P}_Z\}$$ All together, we can think of the basic IV model as the set of distributions $M=\{\mathcal{P}_{latent} \times \mathcal{P}_{Z}: \mathcal{P}_{latent} \in \mathscr{P}_{latent}(\mathcal{G}), \mathcal{P}_{Z} \in \mathscr{P}_{Z}\}$. In this notation note that $\mathscr{P}_{obs}(\mathcal{G})=\phi(M)$.

In general, $\mathscr{P}_{obs}(\mathcal{G})$ is a strict subset of the set of all joint distributions of $(Y_i,T_i,Z_i)$, i.e. restrictions on $\mathcal{G}$ coupled with Eq. ((ref)) imply testable implications on $\mathcal{P}_{obs}$. These testable implications have been studied in the case of the classic LATE model (see e.g. kitagawatesting,mourifiewan,kedagniandmourifie, see also jiang2023testing). Such restrictions are discussed further in Section (ref).

Outcome nonrestrictive IV identification

In defining outcome-nonrestrictive identification, I focus on parameters $\theta$ that take the form of a conditional counterfactual mean $\mu_c^t := \mathbbm{E}[Y_i(t)|c(G_i)=1]$, a conditional treatment effect $\Delta^{t,t'}_c := \mathbbm{E}[Y_i(t')-Y_i(t)|c(G_i)=1]=\mu_c^{t'}-\mu_c^t$, or a probability $P(c(G_i)=1)$. In all three cases, the target parameter is defined from of a function $c: \mathcal{G} \rightarrow \{0,1\}$ that represents inclusion in some collection of response types. For example, in the LATE model, the LATE is a conditional treatment effect $\Delta^{t,t'}_c$ with $t'=1,t=0$ and $c(g) = \mathbbm{1}(g = \textrm{complier})$.

Given such a function $c(\cdot)$, it will be useful to denote the subset of $\mathscr{P}_{latent}(\mathcal{G})$ for which $P(c(G_i)=1)>0$ given the distribution $\mathcal{P}_G$ of $G_i$ as:

equation[equation omitted — 232 chars of source]

Similarly, let $\mathscr{P}_{obs,c}(\mathcal{G}) := \{\phi(\mathcal{P}_{latent} \times \mathcal{P}_Z): \mathcal{P}_{latent} \in \mathscr{P}_{latent,c}(\mathcal{G}), \mathcal{P}_Z \in \mathscr{P}_Z\}$. $\mathscr{P}_{obs,c}(\mathcal{G})$ consist of the distributions of observables that respect selection model $\mathcal{G}$ and put positive probability on the groups $g \in \mathcal{G}$ such that $c(g)=1$. These sets are used in defining outcome-nonrestrictive identification as a simple guarantee that the target parameters $\mu^t_c$ and $\Delta^{t,t'}_c$ are well-defined. But causal parameters that condition on a probability-zero event---such as the marginal treatment effect---can also be accommodated in this framework, as limiting cases of a sequence of parameters for $c_j$ satisfying $P(c_j(G_i)=1)>0$ (see Appendix (ref)).

We are now ready to give a definition of outcome-nonrestrictive identification, where the target parameter $\theta$ is expressed as a function $\theta=\theta(\mathcal{P})$ of the data generating process $\mathcal{P}$:

definitionGiven a choice model $\mathcal{G}$, we say that parameter $\theta$ with conditioning function $c$ is outcome-nonrestrictive identified under $\mathcal{G}$ if the set $$\{\theta(\mathcal{P}): \phi(\mathcal{P}) = \mathcal{P}_{obs} \textrm{ and } \mathcal{P} = (\mathcal{P}_{latent} \times \mathcal{P}_{Z}) \textrm{ for some } \mathcal{P}_{latent} \in \mathscr{P}_{latent,c}(\mathcal{G}) \textit{ and } \mathcal{P}_{Z} \in \mathscr{P}_{Z}\}$$ is a singleton for all $\mathcal{P}_{obs} \in \mathscr{P}_{obs,c}(\mathcal{G})$.

Point identification in general says there is a unique value $\theta(\mathcal{P})$ compatible with Eq. (ref) and $\phi(\mathcal{P})$, for all $\mathcal{P}$ in some set defined by the model. The key requirement that identification be outcome-nonrestrictive is that this model is broad enough to include all of $\mathscr{P}_{latent,c}(\mathcal{G})$.\footnote{Definition (ref) represents a case of point identification as defined in idzoo (see also matzkin2007), where the known information ($\phi$ in Lewbel's notation) is the distribution $\mathcal{P}_{obs}$, the model value ($m \in M$ in Lewbel's notation) is $\mathcal{P} = \mathcal{P}_{latent} \times \mathcal{P}_Z$, and the model $M$ is the Cartesian product of $\mathscr{P}_{latent,c}(\mathcal{G})$ and $\mathscr{P}_Z$.} The set $\mathscr{P}_{latent,c}(\mathcal{G})$ allows what Heckman2004 call essential heterogeneity. The only restrictions on outcomes amount to IV independence (imposed by taking the product measure $\mathcal{P}=\mathcal{P}_{latent} \times \mathcal{P}_Z$), exclusion (implicit in the notation $Y_i(t)$), and finite group-specific means of $\tilde{Y}_i$ (imposed through $\mathcal{P}_{\tilde{Y}G}$ in (ref)).\footnote{$\mathscr{P}_{latent,c}(\mathcal{G})$ does restrict the marginal distributions of $G_i$ and $Z_i$: through $\mathcal{G}$, $P(c(G_i)=1)>0$, and $\mathscr{P}_Z$.} Thus $\mathscr{P}_{latent,c}(\mathcal{G})$ is compatible with any marginal distribution $\mathcal{P}_{\tilde{Y}}$ of $\tilde{Y} = \{Y_i(t)\}_{t \in \mathcal{T}}$ or selection-type conditioned distributions $\mathcal{P}_{\tilde{Y}|G=g}$ across various $g \in \mathcal{G}$ whatsoever (provided that they have finite means), so there is no assumption that e.g. treatment effects are homogeneous across units, or are unrelated to counterfactual selection behavior $G_i$.

Note that identification of $\mu_c^t$ and $\mu_c^{t'}$ immediately implies identification of $\Delta^{t,t'}_c=\mu_c^{t'}-\mu_c^{t}$. With outcome-nonrestrictive identification, this implication in fact goes the other way as well: outcome nonrestrictive identification of $\Delta^{t,t'}_c$ requires outcome-nonrestrictive identification of each constituent part $\mu_c^t$, $\mu_c^{t'}$ . The intuition is that absent assumptions about the joint distribution of potential outcomes, data from individuals with $T_i=t'$ provide no information about $Y_i(t')$ for a different treatment $t'\ne t$, and vice-versa. Thus, we have:

proposition$\Delta^{t,t'}_c$ for $t'\ne t$ is outcome-nonrestrictive identified if and only if $\mu_c^t$ and $\mu_c^{t'}$ are.

See Appendix (ref) for proofs. Although researchers are typically more interested in treatment effect parameters like $\Delta^{t,t'}_c$ than they are in counterfactual means, we can---informed by Proposition (ref)---begin our analysis of outcome-nonrestrictive identification with the simpler counterfactual means, before later considering treatment effects.\footnote{The proof in Section (ref) extends Proposition (ref) to cover treatment effect parameters that involve more than two separate treatment states, which is useful for the study of interaction effects in Section (ref).}

A generic outcome-nonrestrictive identification result

Identifying counterfactual means through “binary combinations”

We begin with a very simple sufficient condition for outcome-nonrestrictive identification of counterfactual means of potential outcomes taking the form $\mu_c^t = \mathbbm{E}[Y_i(t)|c(G_i)=1]$. Consider any finite collection of distinct instrument values $z_k \in \mathcal{Z}$ for $k=1 \dots K$ and corresponding coefficients $\alpha_k$. The following quantity is then identified:

align*[align* omitted — 280 chars of source]

where $D_i^{[t]} = D_i^{[t]}(Z_i) = \mathbbm{1}(T_i=t)$ and the second equality follows from independence ((ref)).

Suppose that the $z_k$ and $\alpha_k$ could be chosen in such a way as to guarantee that the linear combination $\sum_{k}\alpha_{k}\cdot D^{[t]}_i(z_k)$ (in parentheses above) could only take values of 0 or 1 for any given $i$. In such a case, the above simplifies to $$ \sum_{k=1}^K \alpha_k \cdot \mathbbm{E}\left[Y_i\cdot D^{[t]}_i|Z_i=z_k\right] = P\left(\sum_{k}\alpha_{k}\cdot D^{[t]}_i(z_k)=1\right)\cdot \mathbbm{E}\left[Y_i(t)\left|\sum_{k}\alpha_{k}\cdot D^{[t]}_i(z_k)=1\right.\right]$$ Meanwhile

equation[equation omitted — 160 chars of source]

Therefore, provided that $P\left(\sum_{k}\alpha_{k}\cdot D^{[t]}_i(z_k)=1\right)>0:$

equation[equation omitted — 276 chars of source]

Eq. ((ref)) represents a generalization of the “Wald ratio” form common among IV estimands, and turns out to nest a surprising variety of point identification results from the IV literature, both for counterfactual means as well as treatment effects.

To make our way from the former to the latter, let us establish some terminology to refer to situations in which result (ref) can be applied.\\

Definition. Given selection model $\mathcal{G}$, a binary combination is a treatment value $t \in \mathcal{T}$ and a function $\alpha: \mathcal{Z} \rightarrow \mathbb{R}$ of finite support $\mathcal{Z}_K=\{z_k\}_{k=1}^K$ such that $\sum_{k=1}^K\alpha(z_k)\cdot D^{[t]}_i(z_k) \in \{0,1\}$ for all $i$, according to $\mathcal{G}$.\\

In section (ref), I will restrict to discrete instruments and represent the coefficients $\alpha_k$ as a vector in $\mathbbm{R}^{|\mathcal{Z}|}$. But for generality, this section continues to allow the instruments to have arbitrary support and we can think of the coefficients $\alpha_k$ as a function $\alpha(z_k)=\alpha_k$ from $\mathcal{Z}_K\rightarrow\mathbbm{R}$ where $\mathcal{Z}_K=\{z_k\}_{k=1}^K$, or equivalently a function $\alpha$ on all of $\mathcal{Z}$ whose support (the set of points $z$ where $\alpha(z)$ differs from zero) is contained within a finite set $\mathcal{Z}_K \subseteq \mathcal{Z}$.

Binary combinations can thus be indexed by the pair $(t, \alpha)$. Note that given a binary combination, the value of $\sum_{k}\alpha_{k}\cdot D^{[t]}_i(z_k)$ depends only on the response type $G_i$ of individual $i$, so given a binary combination we can write the event that $\sum_{k}\alpha_{k}\cdot D^{[t]}_i(z_k)=1$ as $c^{[t,\alpha]}(G_i)=1$, where $c^{[t,\alpha]}: \mathcal{G} \rightarrow \{0,1\}$ is a function whose definition depends on the treatment $t$ and coefficients $\alpha$ of that binary combination.

With this terminology we can formalize the result of Equation ((ref)) as follows:

theoremGiven independence Eq. ((ref)) and a binary combination $(t, \alpha)$, $P(c^{[t,\alpha]}(G_i)=1)$ is identified by the LHS of (ref). If additionally $P(c^{[t,\alpha]}(G_i)=1)>0$, the conditional counterfactual mean $\mathbbm{E}[Y_i(t)|c^{[t,\alpha]}(G_i)=1]$ is identified by the RHS of ((ref)). Since the only restrictions placed on $\mathcal{P}$ in deriving (ref) are that Eq. (ref) holds, that $\mathcal{Z}_K \subseteq \mathcal{Z}$, and that $P(c^{[t,\alpha]}(G_i)=1)>0$, identification of $\mathbbm{E}[Y_i(t)|c^{[t,\alpha]}(G_i)=1]$ is outcome nonrestrictive.

The key to the observation that $\mathbbm{E}[Y_i(t)|c^{[t,\alpha]}(G_i)=1]$ is outcome-nonrestrictive identified is that whether the requirement $P(c^{[t,\alpha]}(G_i)=1)>0$ holds depends only on the marginal distribution of $G_i$, and whether $\mathcal{Z}_K \subseteq \mathcal{Z}$ depends only on the marginal distribution of $Z_i$. This implies nothing about the distribution of potential outcomes or their relation to $G_i$.

Identifying treatment effects through “binary collections”

While Theorem (ref) yields identification of conditional counterfactual means, we can furthermore identify treatment effects when two treatment values $t$ and $t'$ admit of binary combinations that yield the same conditioning events.

Consider a collection of binary combinations that apply to at least two distinct values $t \in \mathcal{T}$. Let us denote set of coefficients $\alpha$ in each binary combination by $\alpha^{[t]}$, indexed by the treatment value $t$ it will be applied to. In this notation, $\alpha_k^{[t]}$ is the coefficient on $z_k$ in the binary combination corresponding to treatment $t$.\\

Definition. A binary collection is a set of binary combinations $\{(t,\alpha^{[t]})\}_{t \in \psi}$ for treatment values in set $\psi \subseteq \mathcal{T}$ where $|\psi| \ge 2$, with the property that given the selection model $\mathcal{G}$, the functions $c^{[t,\alpha^{[t]}]}$ and $c^{[t',\alpha^{[t']}]}$ are identical, for any $t,t' \in \psi$.\\

For a given binary collection, let us for brevity denote the common function $c^{[t,\alpha^{[t]}]}$ for all $t \in \psi$ as $c$. It follows immediately from Theorem (ref) that treatment effects $\mathbbm{E}[Y_i(t')-Y_i(t)|c(G_i)=1] = \mathbbm{E}[Y_i(t')|c(G_i)=1] - \mathbbm{E}[Y_i(t)|c(G_i)=1]$ are identified for any pair $t, t' \in \psi$.\\

Note: By replacing $Y_i(t)$ by $\mathbbm{1}(Y_i(t) \le y)$ we can also identify the conditional distributions $F_{Y(t)|c(G)=1}$ for all $t \in \psi$ and compute e.g. quantile treatment effects or establish bounds on the distribution of treatment effects among the $c(G_i)=1$ group fanpark.\\

When treatment is itself binary, we can generate binary collections from any binary combination where the coefficients sum to zero:

propositionLet $\mathcal{T} = \{0,1\}$, and suppose $(t,\alpha)$ is a binary combination such that $\sum_k \alpha_k = 0$. Then there exists a binary collection with $\psi = \mathcal{T}$. In particular, the coefficients for $t=0$ are simply $-1$ times the corresponding coefficients for $t=1$.
proofSee alternative statement of this result in Section (ref).

The restriction that $\sum_k \alpha_k = 0$ is a natural one, in the following sense:

propositionLet $\Delta_c^{t,t'}=\mathbbm{E}[Y_i(t')-Y_i(t)|c(G_i)=1]$ be outcome-nonrestrictive identified from a binary collection with $t' \ne t$. Then if $\mathcal{G}$ contains a group $g_0$ that always takes treatment $t$, it must be the the case that $\sum_k \alpha^{[t]}_k = 0$.
proofSince $P(T_i=t'|G_i=g_0) = 0$, the data provide no information on $Y(t')|G_i=g_0$, so we must have $c(g_0)=0$ (see proof of Proposition (ref)). Thus $c(g_0)=\sum_{k} \alpha^{[t]}_k \cdot \mathbbm{1}(T_{g_0}(z_k)=t) = \sum_{k} \alpha^{[t]}_k = 0$.

For example, in the LATE model of Imbens2018, allowing for “always-takers” (who always take treatment $t=1$, regardless of $Z_i$) implies that $\sum_z \alpha^{[1]}_z = 0$, while allowing for “never-takers” (who always take treatment $t=0$) implies that $\sum_z \alpha^{[0]}_z = 0$. Consistent with this, identification of the compliers LATE follows from the binary collection in which $\alpha^{[1]}_1 = 1$, $\alpha^{[1]}_0 = -1$, $\alpha^{[0]}_1 = -1$, and $\alpha^{[0]}_0 = 1$.

Using binary combinations and collections for testing the model

The existence of binary combinations with $K>1$ generally yields overidentification restrictions that can used to test the IV model (including exclusion, independence, and the choice of selection model $\mathcal{G}$). In particular, suppose that $|\mathcal{G}| < \infty$ and note that for any Borel set $\mathcal{B}$ of $\mathbbm{R}$ and binary combination $(t, \alpha)$, we have that:

equation[equation omitted — 141 chars of source]

using Eq. (ref) and that $P(Y_i(t) \in \mathcal{B}, T_i(z_k)=t) = \sum_{g \in \mathcal{G}} P(G_i=g)\cdot P(Y_i(t) \in \mathcal{B}|G_i=g) \cdot A^{[t]}_{z_k,g}$. Since the RHS of Eq. (ref) represents a probability, the LHS must be weakly positive. Provided that not all of the $\alpha_k$ are positive, the implication that $\sum_{k=1}^K \alpha_k \cdot P(Y_i \in \mathcal{B}, T_i=t|Z_i=z_k) \ge 0$ is not guaranteed and therefore can be used to test the model assumptions.

Furthermore, finding binary collections may yield further overidentification restrictions that make use of the “first stage” data alone. Depending on the selection model, the equality $\sum_{k=1}^K\alpha^{[t]}_{k}\cdot \mathbbm{E}\left[D^{[t]}_i|Z_i=z_k\right] =\sum_{k=1}^K\alpha^{[t']}_{k}\cdot \mathbbm{E}\left[D^{[t']}_i|Z_i=z_k\right]$ may not be trivially satisfied, even in the case of a binary treatment. See Section (ref) for an example of such equality restrictions in the context of an empirical application, and Appendix (ref) for further linear inequality constraints that are based upon first stage empirical moments. Still further testable restrictions hold if one has a binary collection and Eq. (ref) holds conditional on observed covariates $X_i$. See Appendix (ref) for details.

Outcome-nonrestrictive identification with discrete instruments

The remainder of this paper now specializes to settings in which the instruments $Z_i$ are discrete and take only a finite number of values $\mathcal{Z}$. This simplification allows us to establish that binary combinations are also necessary for outcome-nonrestrictive identification to occur. Combining with Theorem (ref) and Proposition (ref), this provides a full characterization of outcome-nonrestrictive identification which yields a simple geometric interpretation.\\

Notation: selection models with finite instrument values. When $\mathcal{Z}$ is discrete and finite, any binary combination $\alpha$ can be associated with a vector in $\mathbbm{R}^{|\mathcal{Z}|}$(having non-zero components only for values $z \in \mathcal{Z}_K$). Across a finite set of treatments, note that $\mathcal{G}$ can then also only take finitely many values, i.e. $|\mathcal{G}| \le |\mathcal{T}|^{|\mathcal{Z}|}$. A function $c: \mathcal{G} \rightarrow \{0,1\}$ defining a causal parameter like $\mu_c^t=\mathbbm{E}[Y_i(t)|c(G_i)=1]$ can now be associated with a $|\mathcal{G}|$-component vector $c$ with components $c_g = c(g)$ for each $g \in \mathcal{G}$.

In this setting, we can also express the content of the selection model $\mathcal{G}$ through a $|\mathcal{Z}| \times |\mathcal{G}|$ matrix $A$, where component $A_{zg}$ gives the common treatment $T_i(z)$ that all units $i$ with $G_i=g$ take, when the instruments are equal to $z$. The restrictions imposed by selection model $\mathcal{G}$ correspond to deleting columns from a $|\mathcal{Z}| \times |\mathcal{T}|^{|\mathcal{Z}|}$ matrix that would include all $|\mathcal{T}|^{|\mathcal{Z}|}$ imaginable response types given $\mathcal{T}$ and $\mathcal{Z}$.

Define for any treatment $t$ the binary matrix $A^{[t]}$ having components $[A^{[t]}]_{zg} = \mathbbm{1}(A_{zg}=t)$, which records the common value of $D^{[t]}_i(z)$ for any individual $i$ having $G_i=g$. The matrix $A^{[t]}$ simply tells us whether units take treatment $t$ versus any other treatment.\footnote{A matrix analogous to $A^{[t]}$ is used heavily in Heckman2018.}

A necessary condition for outcome-nonrestrictive identification

Consider a parameter of the form $E[Y_i(t)|c(G_i)=1]$. Recall from above the representation of $c(\cdot)$ as a binary vector $c \in \mathbbm{R}^{|\mathcal{G}|}$ with $c_g \in \{0,1\}$ for each $g \in \mathcal{G}$. For any matrix $B$ let $rowspace(B)$ or $rs(B)$ denote its rowspace, and $B'$ its transpose.

theoremSuppose that $\mathcal{Z}$ and $\mathcal{T}$ are finite. Then if $\mu_c^t:=E[Y_i(t)|c(G_i)=1]$ is outcome-nonrestrictive identified, then $c' = \alpha'A^{[t]}$, for some $\alpha \in \mathbbm{R}^{|\mathcal{Z}|}$, i.e. $c \in rowspace(A^{[t]})$.

Note that given finite $|\mathcal{Z}|$ and hence a finite space of response types, $c \in rs(A^{[t]})$ occurs exactly when there exists a binary combination $(t,\alpha)$ with conditioning function $c=c^{[\alpha,t]}$. Theorem (ref) thus establishes that if the instruments are finite, then Theorem (ref) covers all instances in which a counterfactual mean that conditions on response types can be identified in an outcome-nonrestrictive way. In other words, binary combinations are both necessary and sufficient for outcome-nonrestrictive identification of counterfactual means. Together with Proposition (ref), it follows that binary collections are similarly both necessary and sufficient for outcome-nonrestrictive identification of local average treatment effects.\\

Remark: Theorems (ref) and (ref) both extend to the more general family of target parameters that can be defined by functions $c$ that depend on $Z_i$ in addition to response types $G_i$. This is useful to nest parameters like the average treatment effect on the treated, or certain parameters that can arise in settings with multiple instruments. See Appendix (ref) for details.

Examples to which Theorem (ref) does not apply

Although Theorem (ref) synthesizes a wide variety of existing IV identification results (detailed in Appendix (ref)), the presumption that outcome-nonrestrictive identification holds---rather than point identification in general---is important.

The structure of Theorem (ref) can be summarized as follows. With $M:=\{\mathcal{P}_{latent} \times \mathcal{P}_{Z}: \mathcal{P}_{latent} \in \mathscr{P}_{latent,c}(\mathcal{G}), \mathcal{P}_{Z} \in \mathscr{P}_{Z}\}$, outcome-nonrestrictive identification says that $\{\theta(\mathcal{P}): \mathcal{P} \in M \textrm{ and } \phi(\mathcal{P}) = \mathcal{P}_{obs}\}$ is a singleton for all $\mathcal{P}_{obs} \in \mathscr{P}_{obs,c}(\mathcal{G})$. This requires that there be no $\mathcal{P},\mathcal{P}' \in M$ such that $\phi(\mathcal{P})=\phi(\mathcal{P}')$ but $\theta(\mathcal{P}) \ne \theta(\mathcal{P}')$. The proof of Theorem (ref) uses that there always exist $\mathcal{P} \in M$ that satisfy a certain regularity condition, which enables us to construct from $\mathcal{P}$ such a $\mathcal{P}'$, provided that $c \notin rs(A^{[t]})$. But if one restricts the model space $M$ of permissible DGPs by maintaining further assumptions about the distribution of potential outcomes (or how they are correlated with response types $G_i$), it can be that the constructed distribution $\mathcal{P}'$ violates those assumptions and therefore do not belong to $M$.

For example, it is known for example that $\mathbbm{E}[Y_i(t')-Y_i(t)]$---the unconditional average treatment effect (ATE) between $t$ and $t'$---is identified under an assumption of “no selection on gains” (NSOG): that is that $Y_i(t)-Y_i(t')$ is mean independent of $T_i$ and $Z_i$ for all $t,t' \in \mathcal{T}$ Kolesar2013, aroragoffhjort. Homogeneous treatment effects are a special case of NSOG.\footnote{Another stronger restriction is when potential outcomes $Y_i(t)$ alone---and not just treatment effects---are mean independent of $T_i$ and $Z_i$ for all $t$. This essentially rules out endogeneity: $\mathbbm{E}[Y_i(t)]=\mathbbm{E}[Y_i|T_i=t]$, so identification is unsurprising under this stronger restriction.} In Appendix (ref), I show that the result that the ATE between $t$ and $t'$ is identified under NSOG can be extended to see that unconditional counterfactual means $\mathbbm{E}[Y_i(t)]$ are also identified under NSOG, requiring no assumptions on selection (beyond an order condition that can be verified in the data).

Without a selection model $\mathcal{G}$ that imposes substantive restrictions, the vector $c = (1,1 \dots 1)'$ will not be in the row space of $A^{[t]}$. In particular, as long as there is a “never-takers” group $g_0(t)$ for treatment $t$ such that $T_i(z) \ne t$ for all $z \in \mathcal{Z}$ when $G_i = g_0(t)$, then $(1,1 \dots 1)' \in rs(A^{[t]})$ cannot hold. However there is no contradiction with Theorem (ref), since identification based on NSOG does not hold for all joint distributions between $\tilde{Y}=\{Y_i(t)\}_{t \in \mathcal{T}}$ and $G_i$. Rather, the assumption of NSOG eliminates some such distributions that are compatible with the data and the basic model of Eq. (ref), shrinking $M$.\footnote{In fact, the NSOG assumption is strong enough to let the researcher impute the value of $\mathbbm{E}[Y_i(t)|G_i = g_0(t)]$ if such a never-taker group exists, thus eliminating the dependence of the estimand on the distribution of $Y_i$ among individuals such that $G_i = g_0(t)$ and $T_i=t$ (which would not be identified from observable data).} Indeed I show explicitly in Section (ref) that absent a selection model $\mathcal{G}$ such that $(1,1 \dots 1)' \in rs(A^{[t]})$ for all $t \in \mathcal{T}$, the construction $\mathcal{P}'$ in the proof of Theorem (ref) will not satisfy the additional restriction of NSOG.

Another example of an IV identification result that is not covered by Theorem (ref) is the “compliers--defiers” result of toleratingdefiance that the local average treatment effect among a subset of compliers is identified in a setting with a binary treatment and instrument, if there are more compliers than defiers and a subset of the compliers have the same average treatment effect as the defiers. Again, this additional assumption places restrictions on the joint distribution of response types $G_i$ and potential outcomes $\tilde{Y}_i$. Further, the identified parameter conditions on an event (a particular subgroup of the compliers) that is less course than the groups $G_i$ that are defined simply by counterfactual selection behavior, so does not fit the form $\Delta_{c}^{t,t'}=\mu_c^{t'}-\mu_c^{t}$ that Theorem (ref) and Proposition (ref) speak to. Similar considerations apply to recent results of comey2023supercompliers that show identification of the local average treatment effect among “supercompliers” in a setting in which $\mathcal{Y}=\mathcal{T}=\mathcal{Z}=\{0,1\}$, where the supercompliers are defined as the subset of compliers that have a strictly positive treatment effect. This model imposes monotonicity in the outcome equation, and the conditioning event for the supercomplier LATE conditions both on selection behavior and a property of outcomes, namely that $Y_i(1) > Y_i(0)$.

Another type of identification result that is not covered by Theorem (ref) above---although it is outcome-nonrestrictive---is identification of a treatment effect parameter that does not maintain two fixed treatment values $t$ and $t'$ across all units included in the parameter. An example of this kind arises in klinewalters, in which the identified causal parameter compares the effect of Head Start to one of two next-best alternatives (either traditional pre-school or no pre-school). This estimand combines two response types for which this next-best alternative is generally different. See Section (ref) for details.

When $c \notin rs(A^{[t]})$, Theorem (ref) establishes that the parameter $\mu_c^t$ is not point identified in an outcome-nonrestrictive manner. However, the data may still provide identifying information about the value of $\mu_c^t$ if auxiliary conditions are maintained, for example that the support of $Y_i$ is bounded with known bounds. Appendix (ref) considers partial identification of $\mu_c^t$ in such settings, and also relates the results of this paper to recent results by bai2024identifyingpowermonotonicityaverage, who focus on bounding the ATE and unconditional means in particular.

Summary: combining the necessary and sufficient conditions for identification

Combining Theorems (ref) and (ref), we have in the case of finite discrete instruments that given a selection model $\mathcal{G}$, a conditional counterfactual mean of the form $\mathbbm{E}[Y_i(t)|c(G_i)=1]$ is outcome-nonrestrictive identified if and only if the vector representation $c \in \{0,1\}^{|\mathcal{G}|}$ of $c(\cdot)$ lies in the rowspace of $A^{[t]}$, i.e. $c \in rs(A^{[t]})$. Analogously, a treatment effect parameter $\mathbbm{E}[Y_i(t')-Y_i(t)|c(G_i)=1]$ is outcome-nonrestrictive identified if and only if $c$ lies in the rowspaces of both $A^{[t]}$ and $A^{[t']}$, i.e. $c \in (rs(A^{[t']}) \cap rs(A^{[t]}))$.

Appendix (ref) discusses how Theorems (ref) and (ref) relate to recent necessary and sufficient conditions for identification in IV models by navjeevan2023identification. While Theorem (ref) can be seen a special case of their results, Theorem (ref) cannot. Further, navjeevan2023identification do not derive the condition $c \in (rs(A^{[t']}) \cap rs(A^{[t]}))$ for the identification of treatment effect parameters (though sufficiency of this condition would represent a corollary to their results). I turn now to the analysis of this key condition, which enables an exhaustive search for outcome-nonrestrictive identification results for treatment effects when the instruments are discrete and finite.

Making use of the equivalence result for identification

Continuing our focus on settings in which the instruments have finite points of support, this Section shows how the condition $c \in rs(A^{[t]})$ characterizing identification can be useful in understanding existing identification results, generating new ones, and ruling out further opportunities for identification in a given selection model.

A geometric characterization of identification with discrete instruments

Let $\mathcal{C}(t)$ be the set of $c$ in the rowspace of $A^{[t]}$ that have entries of only zero or one: i.e.

equation[equation omitted — 109 chars of source]

It is always the case that $C(t) \ne \emptyset$ provided that $P(T_i=t)>0$.\footnote{To see this, note that $\mathbbm{E}[Y_i(t)|T_i(z)=t] = \frac{\mathbbm{E}[Y_i\cdot D^{[t]}_i|Z_i=z]}{\mathbbm{E}[D^{[t]}_i|Z_i=z]}$, which considers all units that take treatment $t$ when $Z_i=z$. This corresponds to a binary combination with $\alpha_{z'} = \mathbbm{1}(z'=z)$ and $c_g = \mathbbm{1}(T_g(z)=t)$.} A binary collection in turn occurs when $\mathcal{C}(t) \cap \mathcal{C}(t') \ne \emptyset$.\footnote{Binary combinations can occur in choice models where no binary collections exist. The choice model described in Proposition 8 of lee2023treatment for example has this property.} When treatment is binary, this observation yields a simple proof of Proposition (ref). In this discrete setting, we can rewrite Proposition (ref) as

proposition*[(alternative statement of Proposition (ref))] Let $\mathbbm{1}_{n}$ denote a vector of ones in $\mathbbm{R}^{n}$. If $\mathcal{T} = \{0,1\}$ and $\alpha'\mathbbm{1}_{|\mathcal{Z}|}=0$ and ${A^{[1]}}'\alpha \in \mathcal{C}(1)$, then ${A^{[1]}}'\alpha=-{A^{[0]}}'\alpha$ so that ${A^{[0]}}'(-\alpha) \in \mathcal{C}(0)$ and hence $\mathbb{E}[Y_i(1)-Y_i(0)|c_{G_i}]$ is outcome-nonrestrictive identified.
proofSince $A^{[0]} = \mathbbm{1}_{|\mathcal{Z}|}\mathbbm{1}_{|\mathcal{G}|}'-A^{[1]}$, so $\alpha'A^{[0]}=\cancel{\alpha'\mathbbm{1}_{|\mathcal{Z}|}}\mathbbm{1}_{|\mathcal{G}|}'-\alpha'A^{[1]}=(-\alpha')A^{[1]}$.

Example selection model: the classic LATE model with a binary instrument

Consider the model of Imbens2018 with a binary instrument, where $\mathcal{Z} = \mathcal{T} = \{0,1\}$ and we rule out “defiers”, i.e. those who would have $T_i(0)=1, T_i(1)=0$. Then: $$A^{[1]} =

bmatrix[bmatrix omitted — 38 chars of source]

\quad \quad \quad and \quad \quad \quad A^{[0]} =

bmatrix[bmatrix omitted — 38 chars of source]

$$ where the first row of each matrix represents $z=0$ and the second $z=1$, while the columns correspond to never-takers, always-takers, and compliers, respectively.

The rowspaces of $A^{[1]}$ and $A^{[0]}$ can be found by row-reducing each matrix, yielding: $$rs(A^{[1]}) = span\left\{

pmatrix[pmatrix omitted — 28 chars of source]

,

pmatrix[pmatrix omitted — 28 chars of source]

\right\} \quad \quad \quad and \quad \quad \quad rs(A^{[0]}) = span\left\{

pmatrix[pmatrix omitted — 28 chars of source]

,

pmatrix[pmatrix omitted — 28 chars of source]

\right\}$$

figure[figure omitted — 614 chars of source]

By Theorem (ref), we can thus identify the mean of $Y_i(1)$ among always-takers or among compliers (or among both), and we can identify the mean of $Y_i(0)$ among never-takers or among compliers (or both). As depicted in Figure (ref), these correspond to the non-zero vertices of the unit cube in $\mathbbm{R}^3$ that take a value of zero in the never-takers “direction”, or a value of zero in the always-takers “direction”, respectively.

Note that $(0,0,1)'$ the unique non-zero vertex of the unit cube in $\mathbbm{R}^3$ that belongs to both $rs(A^{[1]})$ and to $rs(A^{[0]})$. Theorem (ref) demonstrates that the LATE among compliers is then in fact the only treatment effect parameter $\Delta_c$ that is outcome-nonrestrictive identified in the LATE model. The local average treatment effect $\Delta_c$ is outcome-nonrestrictive identified for the compliers $c=(0,0,1)'$, because this this $c$ belongs to both $rs(A^{[1]})$ and to $rs(A^{[0]})$.\\

Remark: melowinter study the cardinality of the intersection between the unit cube in $\mathbbm{R}^n$ and any linear subspace of $\mathbbm{R}^n$. For a matrix $A$ with $rs(A)$ of dimension $k$, their result implies that the $rs(A) \cap \{0,1\}^{n}$ has a cardinality of at most $2^k$. In the binary-binary LATE model, $k=|\mathcal{Z}|=2$ for either of $A^{[0]}$ or $A^{[1]}$, and in either case $|rs(A^{[t]}) \cap \{0,1\}^{n}|=2$, which does not meet this upper bound of $2^k=4$. However the result does imply that there can be no more than $2^{|\mathcal{Z}|}$ binary combinations, even though typically $2^{|\mathcal{Z}|} < 2^{|\mathcal{G}|}$ and there are $2^{|\mathcal{G}|}$ potential values of $c$ to consider ex-ante.

Example target parameter: the unconditional average treatment effect

With a binary treatment, a well-studied parameter of interest is the overall population average treatment effect (ATE): $\Delta^{0,1}=\mathbbm{E}[Y_i(1)-Y_i(0)]$. This can be seen as a parameter $\mu_{c}^{t',t}=\mathbbm{E}[Y_i(t')-Y_i(t)|c(G_i)=1]$ in which the function $c(g)=1$ for all $g \in \mathcal{G}$, or in vector form $c=(1,1, \dots 1)'$.

As a Corollary of Theorems (ref) and (ref) we thus have the following:

corollarySuppose that $\mathcal{Z}$ and $\mathcal{T}$ are finite. Then the unconditional counterfactual mean $\mathbbm{E}[Y_i(t)]=\mu^t_{(1,1, \dots 1)'}$ is outcome-nonrestrictive identified if and only if $(1,1, \dots 1)' \in rs(A^{[t]})$, and the unconditional average treatment effect $\mathbbm{E}[Y_i(t')-Y_i(t)]=\Delta^{t,t'}_{(1,1, \dots 1)'}$ is outcome-nonrestrictive identified if and only if $(1,1, \dots 1)' \in (rs(A^{[t']}) \cap rs(A^{[t]}))$.

Since the presence of never-takers with respect to treatment $t$ implies that $(1,1, \dots 1)' \notin rs(A^{[t]}), $\footnote{If such never-takers are alowed in $\mathcal{G}$, this introduces a column of all zeroes in the matrix $A^{[t]}$.} Corollary (ref) implies that ATEs and unconditional counterfactual means are never point-identified in an outcome-nonrestrictive manner absent restrictions on selection.

Corollary (ref) also relates my results to recent work by bai2024identifyingpowermonotonicityaverage on the partial identification power of monotonicity for these parameters, as described in Appendix (ref). bai2024identifyingpowermonotonicityaverage show that selection models can have limited additional identifying power for ATEs provided that they include a restriction that the authors call generalized monotonicity, and the outcome is discrete and bounded. These results underscore the upside to focusing on target parameters beyond the ATE (i.e. $c \ne (1,1,\dots 1)'$) when one is willing to impose restrictions on selection, or finding restrictions on selection that do not imply generalized monotonicity but still aid in identification. Section (ref) shows some examples of this kind.

Applying the characterization to search for identified treatment effect parameters

What then can we say about the set of possible identified treatment effect parameters for a given $t' \ne t$, that is: $\mathbbm{E}[Y_i(t')-Y_i(t)|c_{G_i}=1]$ where $c \in rs(A^{[t]}) \cap rs(A^{[t']}) \cap \{0,1\}^{|\mathcal{G}|}$?

For ease of notation, let us for the moment label the treatment values of interest $t'=1$ and $t=0$, without loss of generality. Let us similarly denote $\alpha^{[t']}$ by $\alpha_1$ and $\alpha^{[t]}$ by $\alpha_0$ (each of these is a $|\mathcal{Z}|$-component vector). Then for some $c \in \{0,1\}^{|\mathcal{G}|}$ and $\alpha_0, \alpha_1 \in \mathbbm{R}^{|\mathcal{G}|}$, we have a binary collection when: $$c' = \alpha_1'A^{[1]} = \alpha_0'A^{[0]}$$ which occurs if and only if

equation[equation omitted — 160 chars of source]

where we let $A^{[1,0]}$ denote a $2\cdot|\mathcal{Z}| \times |\mathcal{G}|$ matrix composed of the rows of $A^{[1]}$ followed by the rows of $A^{[0]}$, and $\alpha = (\alpha_1',-\alpha_0')'$ is a $2\cdot|\mathcal{Z}| \times 1$ vector. For any $\alpha$ in the left null-space $ns({A^{[1,0]}})$ of ${A^{[1,0]}}$, let $c(\alpha)$ denote the value $c = {A^{[1]}}'\alpha_1 = {A^{[0]}}'\alpha_0$ where $\alpha_1$ is a vector of the the first $|\mathcal{Z}|$ components of $\alpha$ and $\alpha_0$ is a vector of minus one times each of the last $|\mathcal{Z}|$ components of $\alpha$. In general then $\mathcal{C}(t) \cap \mathcal{C}(t') = \{c(\alpha): \alpha \in ns(A^{[t',t]})\} \cap \{0,1\}^{|\mathcal{G}|}$, where $A^{[t',t]}$ is composed from $A^{[t']}$ and $A^{[t]}$ as above. This characterization proves useful in the search for new IV identification results to follow.

The following result further aids in implementing a practical search for binary combinations (and hence binary collections):

propositionIf $c \in rs(A^{[t]})$ for some $t$ and $c \in \{0,1\}^{|\mathcal{G}|}$, then the equation $c'=\alpha'A^{[t]}$ can be satisfied by a vector $\alpha$ having elements that are rational and belong to the set $$\mathcal{C}_{n} := \left\{\frac{a}{b}: a,b, \in \mathcal{D}_{|\mathcal{Z}|}\right\}$$ where $\mathcal{D}_{n}:=\{det(B): B \in \{0,1\}^{n \times n}\}$ is the set of possible determinant values for an $n \times n$ matrix $B$ having entries in $\{0,1\}$.

Proposition (ref) implies that when searching for binary combinations, we can always restrict the components of $\alpha$ to belong to the finite set $\mathcal{D}_{|\mathcal{Z}|}$. For $n \le 7$, the set $\mathcal{D}_{n}$ is known to consist of consecutive integers symmetric about zero craigen:\footnote{For $n \ge 8$, $\mathcal{D}_n$ remains a bounded set of integers for any given $n$, but $\mathcal{D}_n$ generally skips some consecutive integers. For example, it is not possible for a $7 \times 7$ binary matrix to have a determinant of $28$ but one can achieve a determinant of $32$ craigen.}

align*[align* omitted — 264 chars of source]

As an implication it for example follows that for $|\mathcal{Z}|\le 2$, all $\alpha_z$ must be in the set $\mathcal{C}_{1}=\mathcal{C}_{2}=\{-1,0,1\}$, in the set $\mathcal{C}_{3}=\{-2,-1,-1/2,0,1/2,1,2\}$ for $|\mathcal{Z}|= 3$, and for $|\mathcal{Z}|= 4$: $$\mathcal{C}_{4}=\{-3,-2,-3/2,-1,-2/3,-1/2,-1/3,0,1/3,1/2,2/3,1,3/2,2,3\}$$

Remark: Since each $\alpha_z$ is rational, we can without loss of generality rewrite Eq. (ref) with integer coefficients $\alpha_k$ (by multiplying the numerator and denominator of (ref) by the least common multiple of the denominators of all $\alpha_z$). However, these integer coefficients do not necessary need to add up to zero, as they do in the binary treatment LATE model and various extensions of it.

Algorithms for generating binary collections

In this section I implement a brute-force algorithm that uses the results thus far to perform an exhaustive search for binary collections in settings with $|\mathcal{Z}|, |\mathcal{T}| \le 3$, uncovering several novel identification results for treatment effects in IV models.\footnote{I am grateful to Simon Lee for suggesting this idea to me.}

I compare two versions of the algorithm, which are laid out explicitly in Appendix (ref). The first is a “naive” approach that iterates over all possible selection models $\mathcal{G}$ given $\mathcal{Z}$ and $\mathcal{T}$ and then finds binary collections within that selection model. For a given selection model $\mathcal{G}$, there are $2^{|\mathcal{G}|}$ possible values of the vector $c$, and a certificate of whether $c$ corresponds to a binary collection for a given $t',t$ can be verified by testing whether $c=c(\alpha)$ for some $\alpha$ in the left nullspace of matrix $A^{[t',t]}$ defined in Eq. (ref). A second algorithm makes use of Proposition (ref) to instead iterate over the possible $2 \cdot |\mathcal{Z}|$-component vectors $\alpha$, rather than over selection models $\mathcal{G}$. This comes at great computational benefit, as computations for a single $\alpha$ are useful for studying many selection models at once. Given the results of Section (ref), we can without loss of generality restrict the search over $\alpha$ to those having components in the discrete and finite set $\mathcal{C}_{|\mathcal{Z}|}$. Compared with Algorithm 1 above, which quickly becomes infeasible for $|\mathcal{Z}|\ge 3$, this second approach runs on $|\mathcal{Z}|=3$ within minutes. The reason is that the number of possible selection models $2^{|\mathcal{T}|^{|\mathcal{Z}|}}$ scales much more quickly with $|\mathcal{Z}|$ than the number $(\mathcal{C}_{|\mathcal{Z}|})^{2|\mathcal{Z}|}$ of possible $\alpha$ vectors, as shown in Appendix Table (ref).

Overview of computational results

Table (ref) presents an overview of results of the two algorithms for settings with $|\mathcal{Z}| ,|\mathcal{T}| \le 3$.\footnotemark While the next section highlights examples from each combination $(|\mathcal{Z}|,|\mathcal{T}|)$ in detail, a full catalog of the identification results is provided in Appendix (ref). While the settings reported in Table (ref) are “small”, they turn out to contain a rich structure of identification results, which varies considerably by $\mathcal{T}$ and $\mathcal{Z}$.

table[table omitted — 841 chars of source]

The third column in Table (ref) counts the number of distinct selection models for a given support of the instruments and treatments, that are maximal for some binary collection. The detailed description of Algorithm 2 in Appendix (ref) describes how given a binary collection indexed by $\alpha \in \mathbbm{R}^{2 \cdot |\mathcal{Z}|}$ (and a choice of $t,t'$, implicit), we can define a maximal selection model $\mathcal{G}(\alpha)$ with the property that $\alpha$ continues to deliver a binary collection for $t',t$ within any smaller selection model $\mathcal{G} \subseteq \mathcal{G}(\alpha)$ that is more restrictive that $\mathcal{G}(\alpha)$.\footnotetext{Run times are with R version 4.3.2 with a 3600MHz processor (AMD Ryzen Threadripper PRO 5975WX), 128GB RAM. While Algorithm 1 is parallelized across 31 cores, Algorithm 2 computation uses a single core. Algorithm 2 is not trivial to parallelize across processors given the need to check for redundancies, but does enable Algorithm 2 to be feasibly extended to $|\mathcal{Z}|=4$ on this computer setup.}\footnote{For example, let $\mathcal{G}$ be the choice model described in Section (ref) from case ii of Proposition 2 of kirkeboenleuvenmogstad. After removing two response types from $\mathcal{G}$, a second treatment effect parameter becomes identified, which is listed under a different selection model $\mathcal{G'} \subset \mathcal{G}$ counted in Table (ref).}

The fourth column in Table (ref) counts the number of distinct binary collection vectors $\alpha$ for a given support of the instruments and treatments. Although a given $\alpha$ generates a valid binary collection under any $\mathcal{G} \subseteq \mathcal{G}(\alpha)$, this column only counts a given $\alpha$ one time, to avoid double counting of the same identification result. Note that for some $|\mathcal{T}|,|\mathcal{Z}|$ there are the same number of distinct identification results for conditional average treatment effects as there are distinct selection models admitting such identification results, this does not mean that exactly one treatment effect parameter is identified in any given selection model. The reason is that the selection models may be nested as described above. The preamble to Appendix (ref) provides a detailed example.

Detailed examples and new identification results

Appendix (ref) makes several illustrative observations from identification results that are summarized in Table (ref), and reported in full in the catalog of Appendix (ref).

Application: interaction effects in cross-randomized designs

This section applies Theorems (ref) and (ref) to study the identification of complementarities between two binary treatment variables. This represents a setting in which $|\mathcal{T}|=|\mathcal{Z}| = 4$. Appendix (ref) considers a second application for this case, which studies treatment effects when there can be spillovers between pairs of observational units.

Background and empirical practice

In many experimental settings, researchers cross randomize two treatments $A$ and $B$, and investigate whether there are interaction effects between the treatments, i.e. whether the effect of receiving both $A$ and $B$ differs from the sum of the effects of each of $A$ and $B$ alone. In some such settings instrumental variables methods are not needed, because compliance is perfect or the intent-to-treat effect is the policy-relevant effect of direct interest (see e.g. dufloetal,mbitietal). Given randomization, intent-to-treat (ITT) effects can be straightforwardly estimated by the regression:

equation[equation omitted — 168 chars of source]

where $Z_i=C$ indicates the treatment arm for both treatments $A$ and $B$. Such cross-randomized experiments are often referred to as “factorial designs”.\footnote{See crosscuts for a review of empirical practice in factorial designs, especially regarding the $\gamma_3$ term.}

However in many factorial designs the treatment arms $Z_i \in \{A,B,C\}$ represent offers for treatments $A$ or $B$ or both, respectively, and researchers obtain data on whether the treatments were actually received. For example, depression study the effects of pharmacotherapy (medication) and livelihood assistance (personalized training and support around income generation), among adults with depression in Karnataka, India. Across the three treatment arms of the cross-randomized experiment, roughly 65% of participants actually undertake pharmacotherapy (defined as attending at least one psychiatric consultation), receive livelihoods assistance (attending at least one livelihoods workshop), or both. Further many adults assigned to receive both pharmacotherapy and assistance undertake only one of the two treatments, although they are offered both. The population studied does not typically have access to pharmacotherapy or livelihoods assistance except through the field experiment, so the non-compliance is one-sided.

To move beyond analysis that is limited to intent-to-treat effects, let us denote the possible treatments as $\mathcal{T} = \{0,A,B,C\}$, with associated potential outcomes $Y_i(t)$ for $t \in \mathcal{T}$. For example, $Y_i(0)$ is the outcome $i$ would experience with neither of the two treatments $A$ and $B$. Meanwhile the set of instrument values is $$\mathcal{Z} = \{\textrm{offered neither},\textrm{offered just A},\textrm{offered just B},\textrm{offered both}\}$$ If subjects are offered both $A$ and $B$, they may choose to take treatment $A$ only, treatment $B$ only, or both treatments $C$. Assume that treatments $A$ and $B$ are otherwise not available to participants, so non-compliance is one-sided.

For a single individual, we can say that $A$ and $B$ exhibit complementarity if $|Y_i(C)-Y_i(0)| > |Y_i(A)-Y_i(0)| + |Y_i(B)-Y_i(0)|$. Of course, testing for complementarity at the individual is infeasible due to the fundamental problem of causal inference. Let

equation[equation omitted — 83 chars of source]

instead be the two-sided hypothesis of no interaction on average, where the interaction effect is $\{Y_i(C)-Y_i(0)\}-\{Y_i(A)-Y_i(0)\} + \{Y_i(B)-Y_i(0)\}=Y_i(C)-Y_i(A)-Y_i(B)+Y_i(0)$. Under perfect compliance, $H_0$ is equivalent to the hypothesis $\gamma_3-\gamma_1-\gamma_2 > 0$ from the ITT regression (ref). This test is employed for example by depression, using only data on assignment and ignoring information about compliance. However, the interpretation of this test may be misleading if compliance is not perfect:

propositionIf there is imperfect compliance, the parameter $\gamma_3-\gamma_1-\gamma_2$ in Eq. (ref) may be zero even when $H_0$ does not hold, and may be non-zero even when $H_0$ holds.

The intuition behind Proposition (ref) is that regression (ref) tells us nothing about complementarity effects among individuals who do not align their actual treatments $T_i$ with their treatment assignment $Z_i$. The result suggests that the common empirical practice of using ITT regressions (rather than focusing on treatment effects per-se) is problematic given that compliance is often known to be far from perfect. However, there are limited identification results for researchers to make use of to estimate interaction effects with imperfect compliance and effect heterogeneity blackwell2017,interactingtreatments.

One solution is to restrict outcomes, assuming sufficient treatment effect homogeneity to get around Proposition (ref). For example, if we assume that no selection on gains (NSOG) holds, the four unconditional counterfactual means $\mathbbm{E}[Y_i(C)]$, $\mathbbm{E}[Y_i(B)]$, $\mathbbm{E}[Y_i(A)]$, and $\mathbbm{E}[Y_i(0)]$ are identified under general conditions given in Appendix (ref). Identification is constructive and corresponds to the estimand of a two-stage least squares (2SLS) regression of $Y_i$ on indicators for each of the four treatments (and no constant), instrumented by indicators for each of the four treatment assignment arms.\footnote{In particular, since there are four instrument values and four treatment values, we can use a result derived in Appendix (ref) under NSOG, that $\mathbbm{E}[Y_i(t)] = \sum_z \Sigma^{-1}_{tz} \cdot \mathbbm{E}[Y_i\cdot \mathbbm{1}(Z_i=z)]$, provided that the matrix with entries $\Sigma_{zt} = P(Z_i=z,T_i=t)$ is invertible. Some algebra shows that this coincides with the two-stage least squares estimand mentioned above.} We can then test $H_0$ by testing $\beta_3=0$ in the equation $Y_i = \beta_0+\beta_1 \cdot \mathbbm{1}(T_i \in \{A,C\})+\beta_2 \cdot \mathbbm{1}(T_i \in \{B,C\}) +\beta_3 \cdot \mathbbm{1}(T_i=C) + \epsilon_i$, estimated using the instruments $\mathbbm{1}(Z_i=A)$, $\mathbbm{1}(Z_i=B)$, $\mathbbm{1}(Z_i=\textrm{both})$ and a constant.

Nevertheless, NSOG is a very restrictive assumption. It suggests for example that individuals do not have some knowledge of their specific gains from the various treatments that informs their selection behavior. When NSOG does not hold, interactingtreatments detail how the 2SLS estimand $\beta_3$ generally mixes interaction effects with terms that simply reflect treatment effect heterogeneity. It is thus desirable to pursue an alternative approach that leads to an interpretable causal estimand without restricting outcomes.

Identifying the local average interaction effect among compliers

We now use Theorems (ref) and (ref) to examine to what extent NSOG can be meaningfully relaxed. Ex-ante, there are $2\times 2\times 4 = 16$ response types that respect one-sided non-compliance.\footnote{Response types correspond to the choices individuals would make across three decisions: whether to take treatment A if A only is offered, B if B only is offered, and which of the four treatment combinations to take if both are offered.} However, assuming that the weak-axiom of revealed preference (WARP) holds, we obtain the additional restrictions that $\{T_{i}(\textrm{offered both})=A \implies T_{i}(\textrm{offered A})=A\}$, $\{T_{i}(\textrm{offered both})=B \implies T_{i}(\textrm{offered B})=B\}$, $\{T_{i}(\textrm{offered both})=0 \implies T_{i}(\textrm{offered A})=T_{i}(\textrm{offered B})=0\}$, $\{T_{i}(\textrm{offered A})=A \implies T_{i}(\textrm{offered both}) \ne 0\}$, and $\{T_{i}(\textrm{offered B})=B \implies T_{i}(\textrm{offered both}) \ne 0\}$. These restrictions eliminate seven response types in total. The nine that remain are enumerated in Table (ref). I refer to the nine remaining response types as $\mathcal{G}^{WARP}$. $\mathcal{G}^{WARP}$ represents the weakest selection model consistent with rational choice and one-sided non-compliance in a factorial design.

table[table omitted — 1,056 chars of source]

Given a selection model $\mathcal{G}$ and a function $c: \mathcal{G} \rightarrow \{0,1\}$, let us refer to $LAIE(c):=\mathbbm{E}[Y_i(C)-Y_i(A)-Y_i(B)+Y_i(0)|c(G_i)=1]$ as the local average interaction effect among the subgroup of $g \in \mathcal{G}$ such that $c(g)=1$. LAIEs are causal quantities like the local treatment effect parameters introduced in Section (ref), except that they involve the potential outcomes for all four treatments rather than just two. The following Proposition uses Theorems (ref) and (ref) to establish when $LAIE(c)$ is identified in a manner that does not restrict outcomes:

propositionGiven one-sided noncompliance and WARP, a local average interaction effect parameter $LAIE(c)$ is outcome-nonrestrictive identified if and only if $c(g)=\mathbbm{1}(g=\textrm{complier})$ and $\mathcal{G} \subseteq \{\textrm{n.t.}, \textrm{complier}, \textrm{A only}, \textrm{B only}\}$.

Proposition (ref) follows from a brute-force enumeration over all of the 511 selection models $\mathcal{G} \subseteq \mathcal{G}^{WARP}$, and the $c \in \{0,1\}^{|\mathcal{G}|}$ within each of them. Theorems (ref) and (ref) along with an extension of Proposition (ref) to parameters that involve more than two treatment states (proved in Appendix (ref)) shows that $LAIE(c)$ is outcome non-restrictive identified iff $c \in rs(A^{[0]}) \cap rs(A^{[A]}) \cap rs(A^{[B]}) \cap rs(A^{[C]})$. Thus it is sufficient to enumerate all binary collections with $\psi = \{0,A,B,C\}$.\footnote{The search also shows that average interaction effects among compliers can also be identified in a selection model in which the “A only” group is traded for a group that takes treatment $A$ only when offered both, and takes neither treatment under all other treatment assignments. I omit an extended discussion of this novel result here for brevity since this final group violates WARP and the model is asymmetric with respect to the treatments (details are available upon request).}

Proposition (ref) establishes that identifying complementarities in a cross-randomized design without outcome restrictions requires substantive restrictions on selection: many of the response types in $\mathcal{G}^{WARP}$ not included in $\{\textrm{n.t.}, \textrm{complier}, \textrm{A only}, \textrm{B only}\}$ are ex-ante plausible. For example, let $U_i(t)$ denote the interpret $U_i(t)$ as the net utility of treatment $t \in \mathcal{T}$ relative to no treatment for individual $i$ (thus normalizing $U_i(0)=0$). Without loss of generality, consider a random coefficients form for the utility function: $U_i(t) = \pi_{Ai} \cdot \mathbbm{1}(t=A)+\pi_{Bi} \cdot \mathbbm{1}(t=B)+\pi_{Ci} \cdot \mathbbm{1}(t=C)$. If the vector $\pi_i = (\pi_{Ai},\pi_{Bi},\pi_{Ci})'$ has support in an open neighborhood of the origin in $\mathbbm{R}^3$, all nine groups from Table (ref) will be present in the population.

Appendix (ref) shows that we can rationalize the restriction made in Proposition (ref) by supposing that individuals choose separately whether to receive treatment $A$ or $B$, rather than as a single joint decision. That is, individuals choose as if they evaluate the costs and benefits of each treatment $A$ or $B$ separately, and choose all treatments offered to them for which benefits outweigh costs. For this reason, let us denote the largest selection model in which the local average interaction effect among compliers is identified as $\mathcal{G}^{sep}:=\{\textrm{n.t.}, \textrm{complier}, \textrm{A only}, \textrm{B only}\}$. Given one-sided noncompliance, the selection model $\mathcal{G}^{sep}$ is also equivalent to what blackwell2017 calls a “treatment exclusion” restriction that the instrument for treatment $A$ does not affect uptake of treatment $B$ (and vice versa). blackwell2017 shows that in this case the interaction coefficient $\beta_3$ from a 2SLS regression identifies the local average interaction effect among compliers. Proposition (ref) shows that treatment exclusion is furthermore necessary to identify this parameter without restricting outcomes. Given WARP and one-sided noncompliance, treatment exclusion cannot be relaxed without restricting outcomes.\% Theorem (ref) then implies that given $\mathcal{G}^{sep}$, the the effects of treatments $A$ and $B$ alone can also be identified without restricting outcomes.\\

Estimating treatment effects and interaction effects among compliers: In the language of Section (ref), the positive side of Proposition (ref) can be understood as the existence of a binary collection $\{(t,\alpha^{[t]})\}_{t \in \psi}$ in which $\psi = \mathcal{T} = \{0,A,B,C\}$. The function $c$ describing this binary collection is $c(g)=\mathbbm{1}(g=\textrm{complier})$, or in vector notation $c' = (0,1,0,0)'$. Let us denote the functions $\alpha^{[t]}(z)$ in vector form as $\alpha_t$, in which case

align[align omitted — 157 chars of source]

One can verify directly that for each $t \in \mathcal{R}$, $\alpha_t'A^{[t]} = (0,1,0,0)$ where the matrix $A$ is defined from the first four columns of Table (ref). For brevity, let us denote $LAIE((0,1,0,0)') = \mathbbm{E}[Y_i(C)-Y_i(A)-Y_i(B)+Y_i(0)|g = \textrm{complier}]$ as simply LAIE (with this $c$ implicit).

The binary collection (ref) implies a cumbersome expression for $LAIE$, but some simplification shows that $LAIE=\theta^{ITT}/p$, where $p=P(G_i=\textrm{complier})$ and $\theta^{ITT}:=\gamma_3-\gamma_1-\gamma_2$ is the measure of average complimentary from the intent-to-treat regression Eq. (ref). This delivers the following important consequence of Proposition (ref):

corollaryGiven $\mathcal{G} \subseteq \mathcal{G}^{sep}$, the sign of the local average interaction effect among compliers $LAIE$ is the same as $\gamma_3-\gamma_1-\gamma_2$ from the intent-to-treat regression (ref).

The algebra that leads to $LAIE=\theta^{ITT}/p$ is given in Appendix (ref), where it is also extended to the case in which covariates are included in Eq. (ref).

While the ITT condition $\gamma_3-\gamma_1-\gamma_2$ cannot be used to test the hypothesis of overall unconditional complementarity (without outcome restrictions), it can by Propisition (ref) be used to test for the sign of local average interaction effect among compliers (the only group for whom interaction effects can be point identified in any way without restricting outcomes). This latter interpretation requires no outcome restrictions, but instead the non-trivial selection model $\mathcal{G} \subseteq \mathcal{G}^{sep}$. Corollary (ref) formally justifies the test for complementarity used by depression within this selection model.

Note finally that the binary collection (ref) implied by the selection model $\mathcal{G} \subseteq \mathcal{G}^{sep}$ yields three overidentification restrictions for the share of compliers:

align[align omitted — 309 chars of source]

for some value $p \in [0,1]$ which identifies $P(G_i = \textrm{complier})$. This testable implication is new to the literature and can be used to assess the substantive assumption $\mathcal{G} \subseteq \mathcal{G}^{sep}$.\footnote{These are stronger than testable implications mentioned by blackwell2017, which give $P(A\textrm{ or }C|\textrm{both})=P(A|\textrm{just A})$ and $P(B\textrm{ or }C|\textrm{both})=P(B|\textrm{just B})$ in the case of one-sided noncompliance. Those do not imply the last line of Eq. (ref).}

Empirical application

I use the replication data from depression to implement the above findings empirically. First, we assess the testable implications of $\mathcal{G} \subseteq \mathcal{G}^{sep}$. Unable to reject the over-identifying restrictions, I then estimate $LAIE$ following Proposition (ref).

In Appendix (ref) I consider a second empirical setting, from angristlangoreopoulos, in which students were cross-randomized into academic support and financial incentives for good grades. In that setting, I find that the testable implications of $\mathcal{G} \subseteq \mathcal{G}^{sep}$ are rejected, and therefore local average interaction effect parameters cannot be identified absent restrictions on outcomes. This illustrates that the general over-identifying restrictions highlighted in Section (ref) have power in an empirically relevant way.

Since the experiment reported in depression stratifies randomization into nine strata (defined by district and terciles of a village poverty index), the implementation below requires some extensions to the basic results of this paper that allow randomization to hold based on observed covariates $X_i$. Implementation is described in Appendix (ref). For simplicity, conditional expectations that need to be estimated are assumed to be additively separable between instruments $Z_i$ and indicators for the strata $X_i$.\\

Testing the overidentification restrictions: Each of the four expressions the proportion $p$ of compliers in Eq. (ref) can be estimated using regressions of the various $D_i^{[t]}$ on instrument indicators as well as indicators for strata $X_i$. The extension of (ref) to the case with strata fixed effects is given in Appendix (ref). Following depression, I use cluster robust inference by village (the level of treatment assignment).

The point estimates for $p:=P(G_i=\textrm{complier})$ are $36.7\%$, $40.2\%$, $39.7\%$, and $43.7\%$, respectively. A chi-squared test for equality of all four estimates of $p$ can be implemented using standard seemingly unrelated regression routines, and yields a p-value of 65%. This indicates that we cannot reject these overidentification restrictions at all conventional levels. This provides some initial evidence in favor of the choice model $\mathcal{G} \subseteq \mathcal{G}^{sep}$.

However, the equality restrictions (ref) are not the only observable implications of $\mathcal{G} \subset \mathcal{G}^{sep}$. In Appendix (ref), I describe how all of the observable first-stage information can be aggregated into a system of linear equations $\mathcal{A}x = \beta$, where $\mathcal{A}$ is a known matrix defined from the $A^{[t]}$, $\beta$ is a vector of observed treatment choice probabilities, and $x$ is a vector of the (non-negative) unobserved occupancies $x_g = P(G_i=g)$ of each response type. Maintaining the weaker assumption of $\mathcal{G} \subseteq \mathcal{G}^{WARP}$, we can test whether $\mathcal{G}$ is furthermore a subset of $\mathcal{G}^{set}$ by computing a lower bound on the sum of the components of $x_g$ for $g \in \mathcal{G}^{WARP}-\mathcal{G}^{sep}$, subject to the constraints that $\mathcal{A}x = \beta$, each $x_g \ge 0$ and the $x_g$ sum to unity. In principle, this exercise could be implemented by strata to test $\mathcal{G} \subseteq \mathcal{G}^{sep}$ among the individuals within each. To increase statistical power given the small sample, I pool the data across all strata for this exercise. This is valid under the assumption that the response-type distribution is common across strata.\footnote{Note that this resrtiction does not require potential outcomes to be uncorrelated with stratum.}

Ignoring sampling uncertainty in the observed treatment choice probabilities, the data suggest that $P(G_i \in \mathcal{G}^{WARP}-\mathcal{G}^{sep})$ is at least $6.3\%$ (and is no more than $80.8\%$). Since $6.3\% > 0$, this provides some evidence against the restriction $\mathcal{G} \subseteq \mathcal{G}^{sep}$.\footnote{Point estimates further yield $p_{oboth} \in [0,3\%]$, $p_{A+} \in [0,3\%]$, $p_{B+} \in [0,34\%]$, $p_{favor A} \in [3\%,6\%]$, $p_{favor B} \in [3\%,38\%]$.} However, accounting for uncertainty in $\beta$ suggests that this lower bound for $P(G_i \in \mathcal{G}^{WARP}-\mathcal{G}^{sep})$ is not statistically significant. fsst provide a method for testing whether there exists componentwise non-negative solutions $x$ to systems of the form $\mathcal{A}x=\beta$ like the above, when $\beta$ is estimated from the data. This method yields a 95% confidence interval of $[0, 0.83]$ for the share of offending response types. This confidence interval includes zero (up to machine-precision), indicating that we cannot reject that $\mathcal{G} \subseteq \mathcal{G}^{sep}$ within the weaker assumption $\mathcal{G} \subseteq \mathcal{G}^{WARP}$, even using the full observable information on treatment uptake. Appendix (ref) provides details on this procedure.\footnote{I implement the FSST method using the $\mathtt{R}$ package $\mathtt{lpinfer}$. This method does not involve any clustering and is designed for $i.i.d.$ data, so the confidence interval reported above may undercover the parameter $P(G_i \in \mathcal{G}^{WARP}-\mathcal{G}^{sep})$ if one considers uncertainty as arising from treatment assignment as well. Since the proportion of each cluster (in this case village) that is sampled is small (on average about two individuals), results for OLS suggest that the influence of clustering in treatment assignment may be minimal, even considering both uncertainty arising from clustered treatment assignment as well as sampling abadieetalclustering. Similar results are provided using alternative methods for inference on linear systems introduced by romanoshaikh and chorussell.}\\

Estimates of local average interaction among compliers: The data from depression follow 1,000 respondents over five survey waves. I use their main outcome variable, which is a standardized version of the PHQ-9 score for depression, with higher values indicating more severe depression. I focus on longer-run outcomes in the fourth and fifth waves, which occured between one and two years after treatment. In these longer waves, the authors estimate $\gamma_3-\gamma_1-\gamma_2$ from the ITT regression to be marginally significant at the 10% level with a p-value of $.10$. Meanwhile, they find that the combination (treatment “C”) of pharmacotherapy (treatment “A”) and livelihoods assistance (treatment “B”) reduces depression symptoms even after the intervention that is significant at the 95% level, while the effects of treatments A or B alone are insignificant (cf. their Table 2, panel B). However, these estimates come from an ITT regression that ignores actual treatment uptake, and may be attenuated or otherwise distorted when interpreted as effects of the treatments themselves rather than as effects of assignment.

table[table omitted — 786 chars of source]

Column (1) of Table (ref) implements this ITT regression of the outcome on instrument indicators (and strata fixed effects). Departing slightly from depression, I focus on a minimal specification and do not control for baseline values of the outcome as they do. However, the findings are qualitatively the same and quantitatively similar. In line with depression only the effect of treatment $C$ (pharmacotherapy and livelihoods assistance) is statistically significant. Column (2) uses data on treatment uptake and implements a 2SLS regression as described in Section (ref). Consistent with intuition given imperfect compliance, these treatment effects estimates are larger in magnitude and have the same pattern of significance. However, the main treatment effect estimates (besides the LAIE) invoke the strong assumption of no-selection-on-gains (NSOG) to be interpreted causally and as averaging over the same response types.\footnote{Thm. 2 of blackwell2017 shows how the various 2SLS coefficients average effects over different groups of response types even given $\mathcal{G} \subseteq \mathcal{G}^{sep}$, making them not comparable to one another without outcome restrictions like NSOG.}

By contrast, none of the estimates reported in Columns (3) and (4) require restrictions on outcomes to be causally interpreted and compared. Column (3) uses simple sample estimators of the expectations from Eq. (ref) (extended for strata fixed effects) along with the binary combinations Eq. (ref) that isolate compliers. See Appendix (ref) for details. Column (4) re-estimates Column (3) while further imposing the overidentification restrictions (ref) for the share of compliers, using a generalized method of moments (GMM) estimator (see Appendix (ref) for details). This estimator penalizes the differing numerical estimates of $p=P(i \textit{ is complier})$ in each of the four terms that make up $LAIE$. While the GMM estimator does not end up reducing the standard error of $LAIE$ in this setting, it does restore statistical significance to $\mathbbm{E}[Y_i(C)-Y_i(0)|i \textit{ is complier}]$ at the 95% level.

The three estimates of $LAIE$ in columns (3)-(5) are valid under the same assumption that $\mathcal{G} \subseteq \mathcal{G}^{sep}$, and suggest that pharmacotherapy and livelihoods assistance are complementary: they have an interaction effect of about half of a standard deviation of PHQ-9 among compliers. 2SLS provides the most precise estimate of this parameter, which is signficant at the 10% level. The relatively large positive magnitude of the effect of livelihoods assistance (treatment B) in columns (3) and (4) raises the question of whether this intervention may in fact exacerbate depression symptoms among compliers, when it is not accompanied by pharmacotherapy (treatment A). This economically meaningful magnitude is not evident in the ITT estimates from depression that do not adjust for non-compliance. However, the estimate is not quite significant at the 10% level even with the GMM estimator (t-statistic $2.47/1.67=1.48$), and thus should be interpreted with caution. Otherwise, the estimates reported in Table (ref) confirm the qualitative findings of depression, while offering quantitative treatment effects that account for the partial compliance.

Conclusion

This paper has formalized the notion of “outcome-nonrestrictive” identification in IV models, and shown that it is equivalent (with discrete instruments) to the existence of linear combinations of counterfactual treatment indicators that add up to zero or one for all response types in the assumed selection model. A selection model only allows for treatment effects to be identified in an outcome-nonrestrictive way when a particular matrix that summarizes the selection behavior allowed by the model with respect by two treatment values has a non-trivial null-space that intersects the unit cube in the space of types allowed by the model. This insight yields a systematic approach to enumerating all selection models that afford identification of treatment effects in a manner that does not restrict outcomes. The search delivers a multiplicity of new identification results, despite its computational complexity scaling rapidly with the size of support of the instruments and treatment. Future work could leverage the algorithms proposed in this paper to construct a searchable database of identification results for still more complex settings.

\printbibliography \nocite{depressiondata}