Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,406 characters · 17 sections · 22 citation commands
A Vector Monotonicity Assumption for Multiple Instruments
The local average treatment effects (LATE) framework of Imbens2018 allows instrumental variables to be used for causal inference even when there is arbitrary heterogeneity in treatment effects. However, the model makes an important assumption about homogeneity in individuals' selection behavior, referred to as monotonicity. When the researcher has a single instrumental variable at their disposal, this LATE monotonicity assumption is typically quite a natural one to make. But when multiple instruments are combined, LATE monotonicity can become hard to justify---a point recently emphasized by \citet*{Mogstad2020a} (henceforth MTW).
This paper considers a natural alternative assumption, which is that monotonicity holds on an instrument-by-instrument basis: what I call vector monotonicity (VM).\footnote{A leading special case of VM is discussed by MTW under the name actual monotonicity (see Section (ref)).} Vector monotonicity assumes that each instrument has an impact on treatment uptake in a direction that is common across units, regardless of the values of the other instruments. This direction need not be known ex-ante by the researcher, but is often implied by economic theory. For example, two instruments for college enrollment might be: i) proximity to a college; and ii) affordability of nearby colleges. VM assumes that increasing either instrument induces some individuals towards going to college, while discouraging none, i.e. proximity to a college weakly encourages college attendance regardless of price, and lower tuition weakly encourages college attendance regardless of distance. This contrasts with the traditional monotonicity assumption of the LATE model, which requires that either proximity or affordability dominates in the selection behavior of all individuals: in particular, it implies that all individuals who would go to college if it were far but cheap would also go if it were close but expensive, or that the reverse is true.
I provide a simple approach to estimating causal effects under VM. In a setting with any number of binary instruments satisfying VM, I show that average treatment effects can be point identified for subgroups of the population if and only if that subgroup satisfies a certain condition.\footnote{In Appendix (ref), I show how discrete instruments more generally can be accommodated by re-expressing them as a larger number of binary instruments, while preserving vector monotonicity and without loss of information.} The condition is met by, for example, the group of all units (e.g. individuals) that move into treatment when any fixed subset of the instruments are switched “on”. As special cases, this includes for example the set of units that would respond to changing a single particular instrument, or those units for whom treatment status would vary in any way given changes to the available instruments. I propose a two-step estimator for this family of identified causal parameters.\footnote{This estimator is implemented in the companion Stata package ivcombine, available from \href{https://github.com/leonardgoff/ivcombine}{https://github.com/leonardgoff/ivcombine}.} Notably, the estimator has the same computational cost as the popular two-stage least squares (2SLS) estimator, despite the rapid proliferation of potential selection patterns compatible with VM as one increases the number of instruments.
Vector monotonicity represents a special case of what MTW refer to as partial monotonicity (PM). VM and PM are very similar, but PM is ex-ante weaker: it allows the “direction” in which treatment uptake increases for each instrument to depend on the values of the other instruments. However given PM and the standard instrumental variables (IV) independence assumption, the additional restriction made by VM is testable. In particular, VM implies that the propensity score function is component-wise monotonic in the instruments. VM and PM thus coincide when this testable restriction is satisfied, and VM can be thought of as an application of PM within a class of settings that can be distinguished empirically. Further, VM is also often implied by natural choice-theoretic considerations, making monotonicity of the propensity score reasonable to expect provided that the instruments are valid.
In their paper, MTW focus on the causal interpretation of the 2SLS estimand under PM, and show that 2SLS is not guaranteed to recover a convex combination of heterogeneous treatment effects under PM or VM.\footnote{For example, Proposition 5 of MTW demonstrates this in the case of two binary instruments satisfying VM.} This motivates the question of what identifying power remains for instrumental variables satisfying PM or VM to uncover causal effects. In a second paper (\citealt*{Mogstad2020b}, henceforth MTW2), these same authors discuss identification more generally under partial monotonicity. MTW2 adapt the marginal treatment effects (MTE) framework of Heckman2005 for use under PM, and construct identified sets for a broad class of causal parameters that are typically only partially identified by IV methods (absent parametric assumptions and/or continuous instruments).
By contrast, my results maintain VM and characterize the class of causal parameters that are point identified without any auxiliary assumptions and even with discrete instruments. I show in Appendix (ref) that when VM does hold, the class of treatment effect parameters identified by my approach coincides with those of the same form that would be point identified under the approach of MTW2, if the method of MTW2 is applied using all identifying moments provided by the data but without additional maintained assumptions (e.g. parametric forms for MTEs).\footnote{My results thus also confirm a conjecture of MTW2---that their approach leads to identified sets that are sharp---in the setting I consider and when point identification holds.} In view of this, a desirable feature of my approach is that it is able to guarantee upfront to the researcher that their chosen target parameter is identified, and give a menu of such parameters that one could estimate. By leveraging constructive estimands for the target parameter, my results also lead to an easy-to-implement estimator and associated confidence intervals.
The estimator I propose in this paper can thus be seen as an alternative to the method of MTW2, but also to 2SLS, which is the most popular method to make use of multiple instruments in applied work. MTW derive additional testable conditions which are sufficient for the 2SLS estimand to deliver positive weights under PM, but the number of conditions to be verified grows combinatorially with the number of instruments. Targeting a particular treatment effect parameter that is identified under VM avoids the need for such tests. My estimator couples this advantage with the computational ease of a simple “2SLS-like” estimator.
In Section (ref), I review the basic IV setup with a binary treatment, and compare VM to the traditional LATE monotonicity assumption and PM. In Section (ref) I set the stage for the identification analysis by showing how with any number of binary instruments VM partitions the population into well-defined “response groups”, nesting results from MTW for the two-instrument case. I then use this taxonomy of response groups in Section (ref) to characterize the family of identified parameters under VM with binary instruments, which leads to the estimator proposed in Section (ref). Section (ref) applies my method to study the labor market returns to college.
Suppose the researcher has a scalar outcome variable $Y$, a binary treatment variable $D$, and a vector $Z = (Z_1,Z_2, \dots ,Z_J)'$ of $J$ instrumental variables that can take values in $\mathcal{Z} \subseteq (\mathcal{Z}_1 \times \mathcal{Z}_2 \times \dots \times \mathcal{Z}_J)$, where $\mathcal{Z}_j$ denotes the set of values that instrument $Z_j$ can take. A typical value in $\mathcal{Z}$ will be denoted with the boldface notation $\mathbf{z}$, with $z_j$ denoting the component corresponding to the $j^{th}$ instrument. I employ the standard definitions of potential outcomes and potential treatments, letting $D_i(\mathbf{z})$ denote the counterfactual treatment status of observational unit $i$ (e.g. an individual) when the vector of instruments takes value $\mathbf{z}$, and $Y_i(d,\mathbf{z})$ the outcome that would occur with treatment $d \in \{0,1\}$ and value $\mathbf{z}$ for the instruments. Let $Z_i=(Z_{1i},\dots, Z_{Ji})'$ denote unit $i$'s realized value of all $J$ instruments.
The following assumption states that the $J$ available instrumental variables are valid:
The first part of Assumption 1 states that the instruments satisfy the exclusion restriction that potential outcomes do not depend on instrument values once treatment status is fixed. The second part of Assumption 1 states that the instruments are statistically independent of potential outcomes and potential treatments.\footnote{It's worth noting that whether or not to use multiple instruments may not be “optional”, in the sense that if a collection of instruments are valid, this does not imply that a subset of the instruments are as well.} In practice, it is common to maintain a version of this independence assumption that holds conditional on a set of observed covariates. I implicitly condition on any such covariates and discuss incorporating them in estimation in Appendix (ref).
It is well-known that when treatment effects are heterogeneous, Assumption 1 alone is not sufficient for instrument variation to identify treatment effects. The seminal LATE model of Imbens2018 introduces the additional assumption of monotonicity:
I have referred to IAM as “LATE monotonicity” in the introduction, but for the remainder of the paper I follow MTW and call it IAM for short (for “Imbens and Angrist monotonicity”).
To appreciate the sense in which IAM can be strong when $\mathbf{z}$ is a vector, let us code the two instruments for college from the introduction as binary variables (“far”/“close” and “cheap”/“expensive”). As emphasized by MTW, IAM says that a given counterfactual change to the proximity and/or tuition instruments can either move some students into college attendance, or some students out, but not both. In particular, this requires that all units who would go to college when it is far but cheap would also go to college if it was close and expensive, or the other way around. This implication will generally fail to hold if individuals differ in how much each of the instruments “matters” to them: for example, if some students are primarily sensitive to distance while others are primarily sensitive to tuition.\footnote{MTW also show that with continuous instruments, IAM implies the very strong restriction that marginal rates of substitution are identical among individuals indifferent between treatment and non-treatment.}
Vector monotonicity instead captures monotonicity as the notion that increasing the value of any one instrument weakly encourages (or discourages) all units to take treatment, regardless of the values of the other instruments.
When each $\ge_j$ is the standard ordering on real numbers, MTE call VM “actual monotonicity”, or AM.\footnote{Mountjoy2018 imposes a version of VM in a setting with continuous instruments and a ternary treatment.} I instead use the term “VM” to emphasize that $\ge_j$ need not be this order for identification results to hold, but I will typically restrict to AM (which represents a simple relabeling of the instrument values) for ease of exposition.
Assumption IAM implies the existence of a (total) order on $\mathcal{Z}$, where if $\mathbf{z} \ge \mathbf{z}'$ with respect to that order then $D_i(\mathbf{z}) \ge D_i(\mathbf{z'})$ for all $i$.\footnote{This follows since if $D_i(\mathbf{z}) \ge D_i(\mathbf{z}')$ and $D_i(\mathbf{z}') \ge D_i(\mathbf{z}'')$, then $D_i(\mathbf{z}) \ge D_i(\mathbf{z}'')$. Any two points in $\mathcal{Z}$ can be ranked in this way, yielding a weak total ordering on $\mathcal{Z}$.} In the returns-to-schooling example, this order might be the following, where an arrow from $\mathbf{z}'$ to $\mathbf{z}$ indicates that $D_i(\mathbf{z}) \ge D_i(\mathbf{z}')$ for all $i$: \tikzstyle{line} = [draw]
An alternative ordering to the one depicted above would be that instead $D_i(expensive, far) \le D_i(expensive, close) \le D_i(cheap, far) \le D_i(cheap, close)$. While either of these two orders may seem equally plausible ex-ante, Assumption IAM requires that only one or the other holds, common to all $i$ in the population.
By contrast, VM ascribes a partial order on $\mathcal{Z}$---only some pairs $(\mathbf{z},\mathbf{z}')$ are ranked. In the returns to schooling example, the obvious partial order under VM is:
The absence of vertical arrows between $(cheap, far)$ and $(expensive, close)$ above means that under VM, it could be the case that $D_i(cheap, far) > D_i(expensive, close)$ for some $i$, while $D_i(cheap, far) < D_i(expensive, close)$ for some other $i$.
The partial monotonicity assumption (PM) introduced by MTW is weaker than both IAM and VM. Like VM, it implies a partial order on $\mathcal{Z}$. Let $(z_j, \mathbf{z}_{-j})$ denote a given value in $\mathcal{Z}$ as the combination of a value $z_j \in \mathcal{Z}_j$ for the $j^{th}$ instrument and value $\mathbf{z}_{-j} \in \mathcal{Z}_{-j}$ for the other instruments, where $\mathcal{Z}_{-j}$ denotes the set of possible values for all instruments aside from $Z_j$.
Under PM, there exists for any instrument $j$ an ordering on the points $z \in \mathcal{Z}_j$ such that $D_i(z,\mathbf{z}_{-j})$ is weakly increasing along the order, for that fixed choice of $\mathbf{z}_{-j}$. The key (and only) additional restriction made by VM beyond PM is that under VM, this ordering must be the same across all values of $\mathbf{z}_{-j}$ for a given $j$. For example, close proximity to a college encourages going to college, whether or not nearby colleges are cheap. By contrast, PM could capture a situation in which college proximity encourages attendance when nearby colleges are cheap, but discourages attendance when they are expensive.
While VM is ex-ante stronger than PM, the additional restriction made by VM over PM is empirically testable, by inspecting the propensity score function $\mathcal{P}(\mathbf{z}):=\mathbbm{E}[D_i|Z_i=\mathbf{z}]$. Call $\mathcal{Z}$ non-disjoint when for any two $\mathbf{z},\mathbf{z'} \in \mathcal{Z}$ there exists a sequence of vectors $\mathbf{z}_1, \dots, \mathbf{z}_M \in \mathcal{Z}$ where each $\mathbf{z}_m$ and $\mathbf{z}_{m-1}$ differ on only one component, and $\mathbf{z}_1=\mathbf{z}$, $\mathbf{z}_M=\mathbf{z}'$.\footnote{This property rules out atypical cases such as $\mathcal{Z}$ consisting only of the points $(0,0)$ and $(1,1)$ e.g. if $J=2$.}
Unlike VM, PM (like IAM) is compatible with any propensity score function $\mathcal{P}(\mathbf{z})$. Since IAM implies PM, it also follows from Proposition (ref) that if IAM and Assumption 1 hold and $\mathcal{P}(\mathbf{z})$ is component-wise monotonic in $\mathbf{z}$, then VM holds. Thus if a researcher has verified that the propensity score function is monotonic,\footnote{With two binary instruments for example, one could test the four inequalities $\mathcal{P}(1,1) \ge \mathcal{P}(1,0)$, $\mathcal{P}(1,1) \ge \mathcal{P}(0,1)$, $\mathcal{P}(1,0) \ge \mathcal{P}(0,0)$, and $\mathcal{P}(0,1) \ge \mathcal{P}(0,0)$. This can be accomplished through a regression $D_i = \beta_0 + \beta_1 Z_{1i}+\beta_2 Z_{2i}+\beta_3 Z_{1i}Z_{2i}+\epsilon_i$ and testing that $\beta_1, \beta_2, \beta_3+\beta_1$ and $\beta_3 + \beta_2$ are all positive.} VM becomes a strictly weaker assumption than IAM. The overall relationship between Assumptions IAM, VM and PM is depicted in Figure (ref).\\
Remark 1: Note that if Assumption 1 holds conditional on covariates $X_i$, Proposition (ref) also need only hold with respect to the conditional propensity score $\mathbbm{E}[D_i|Z_i=\mathbf{z},X_i=x]$ (see Section (ref)). If VM is maintained, this property could in principle be used to test Assumption 1 conditional on a given set of covariates $X$.\\
Remark 2: Another sufficient condition for VM given PM is the existence of individuals for each instrument that are responsive only to the value of that instrument. For example, suppose Alice only cares about proximity (going to college if and only if it is close), and Bob only cares about tuition (going to college if and only if it is cheap). If Alice and Bob are both present in the population, PM then requires that all other units in the population exhibit (weakly) the same directions of response to both instruments that Alice and Bob do, implying VM. The existence of both Alice and Bob in the population would also imply that IAM does not hold.
To set the stage for analysis of identification under VM, I in this section show that VM partitions the population of interest into a set of groups that generalize the familiar taxonomy of “always-takers”, “never-takers”, and “compliers” from Imbens2018, and also nests a taxonomy of six groups introduced by MTW for the case of two binary instruments.
To simplify notation, let $G_i$ represent an individual's entire vector of counterfactual treatments $\{D_i(\mathbf{z})\}_{\mathbf{z} \in \mathcal{Z}}$. For example, with a single binary instrument $G_i=\textit{always-taker}$ indicates that $D_i(0)=D_i(1)=1$. I refer to $G_i$ as unit $i$'s “response group”.\footnote{This language follows Lee2020. Heckman2018 use response-types or strata.} Response groups partition individuals in the population based on their selection behavior over all counterfactual values of the instruments. VM can be thought of as a restriction on the support $\mathcal{G}$ of $G_i$, limiting the types of response groups that can coexist in the population.
While this section describes the structure of $\mathcal{G}$ under VM when the instruments are each binary, Appendix (ref) shows how one can re-code a set of discrete but non-binary instruments as a larger number of binary instruments, while preserving VM. Also without loss of generality, let us let the value labeled “1” for each binary instrument be the direction in which potential treatments are increasing. These “up” values might be predicted ex-ante, but by Proposition (ref) they are also empirically identified from the propensity score function.
With one binary instrument, VM and IAM coincide. $\mathcal{G}$ then contains the three groups (see e.g. Angrist2008): “compliers” (for whom $D_i(1)>D_i(0)$), “always-takers” (for whom $D_i(1)=D_i(0)=1$) and “never-takers” (for whom $D_i(1)=D_i(0)=0$).
In the case of two binary instruments satisfying VM, MTW show that $\mathcal{G}$ contains six distinct response groups, enumerated in Table (ref) below. In the returns to college setting, a “$Z_1$ complier”, for example, would go to college if and only if college is cheap, regardless of whether it is close (like Bob). A $Z_2$ complier, by contrast, would go to college if and only if college is close, regardless of whether it is cheap (like Alice). A reluctant complier requires college to both be cheap and close to attend, while an eager complier goes to college so long as it is either cheap or close. Never and always takers are defined in the same way as they are under IAM: by $\max_{\mathbf{z} \in \mathcal{Z}} D_i(\mathbf{z})=0$ or $\min_{\mathbf{z} \in \mathcal{Z}} D_i(\mathbf{z})=1$.
Now consider any number $J$ of binary instruments. For simplicity, let $\mathcal{Z} = \{0,1\}^J$, where $\{0,1\}^J = \{(z_1, z_2, \dots z_J): z_j \in \{0,1\}\}$ denotes the $J-$times Cartesian product of $\{0,1\}$.\footnote{If $\mathcal{Z}$ is a strict subset of $\{0,1\}^J$, the response groups can be defined from the restrictions to $\mathcal{Z}$ of the groups defined here. Appendix (ref) generalizes identification results to such cases, which arise with discrete instruments.} There are ex-ante $2^{|\mathcal{Z}|}=2^{2^{J}}$ distinct possible mappings between vectors of instrument values and treatment status, and we wish to characterize the subset of these that satisfy VM. The number of such response groups $G_i$ is the number of isotone boolean functions on $J$ variables, which is known to follow the so-called Dedekind sequence:\footnote{An analytical expression for the $\textrm{Ded}_J$ is given by A.Kisielewicz1988, but only the first eight have been calculated numerically. While the Dedekind sequence explodes quite rapidly, it does so much more slowly than $2^{2^{J}}$ does. For example while $3/4=75\%$ of conceivable response groups for $J=1$ satisfy VM, only $20/256 \approx 7.8\%$ do for $J=3$, and just $ 7581/4294967296 \approx 1.7 *10^{-4}$ do for $J=5$. Thus the “bite” of VM is increasing with $J$, ruling out an increasing fraction of conceivable selection patterns.} $3, 6, 20, 168, 7581, 7828354, \dots $ A.Kisielewicz1988. Letting $\mathrm{Ded}_J$ denote the $J^{th}$ value in this sequence, there are e.g. $\mathrm{Ded}_3=20$ response groups when there are three instruments, and $\mathrm{Ded}_4=168$ groups when there are four.
For an arbitrary $J$, we can enumerate these $\mathrm{Ded}_J$ groups as follows. One group that always satisfies VM is composed of “never-takers”: those units for whom $D_i(\mathbf{z})=0$ for all values $\mathbf{z} \in \mathcal{Z}$. Each of the remaining response groups can be associated with a collection of sets of instruments. These sets represent minimal sets of the instruments that are sufficient for that unit to take treatment, if all instuments in the set take a value of one. For example, in a setting with three instruments, one response group would be the units that take treatment if either $Z_1=1$, or if $Z_2=Z_3=1$. We associate this response group with the collection of sets $\{1\}, \{2,3\}$. Note that by VM, any unit in this group must also take treatment if $Z_1=Z_2=Z_3=1$. Another response group might take treatment only if $Z_1=Z_2=Z_3=1$, and is associated with the single set $\{1,2,3\}$. This response group is more “reluctant” than the former. The group of always-takers are the least “reluctant”: they require no instruments to equal one in order for them to take treatment.
Formally, we can associate each response group aside from never-takers with a collection $F$ (which I refer to as a “family”) of subsets $S \subseteq \{1\dots J\}$ of the instruments. A unit in the response group associated with family $F$ takes treatment when all instruments in any of the $S \in F$ are equal to one. However, we need only consider families for which no $S \in F$ is a subset of some other $S' \in F$. Families of sets having this property are referred to as Sperner families (see e.g. Kleitman1973). Families that are not Sperner would be redundant under VM: for example, if $F$ consists of the set $\{2,3\}$ and $\{1,2,3\}$, then given VM the set $\{1,2,3\}$ could be dropped from $F$ without affecting the implied selection function $D_i(\mathbf{z})$.
The response groups satisfying VM with $J$ binary instruments are thus: i) the never-takers group; and ii) $Ded_J-1$ further groups $g(F)$ corresponding to each distinct Sperner family (one such family is the null-set, which corresponds to always-takers).\footnote{For an explicit proof these exhaust all distinct $D:\{0,1\}^J \rightarrow \{0,1\}$, see e.g. Anderson1987 (Sec. 3.4.1).}
In the simplest example when $J=1$, the only Sperner families are the null set and the singleton $\{1\}$: corresponding to always-takers and compliers, respectively. Together with never-takers, we have the familiar three groups from LATE analysis with a single binary instrument. For $J=2$, the five groups (apart from never takers) map to Sperner families shown in the rightmost column of Table (ref). With $J=3$ there are 19 Sperner families,\footnote{These are (listed each within bold brackets for legibility): $ \pmb{\{}\emptyset \pmb{\}}, \pmb{\{}\{1\}\pmb{\}}, \pmb{\{}\{2\}\pmb{\}}, \pmb{\{}\{3\}\pmb{\}}, \pmb{\{}\{1,2\}\pmb{\}}, \pmb{\{}\{1,3\}\pmb{\}},\pmb{\{}\{2,3\}\pmb{\}},$ $\pmb{\{}\{1,2,3\}\pmb{\}},\pmb{\{}\{1\},\{2\}\pmb{\}}, \pmb{\{}\{2\},\{3\}\pmb{\}}, \pmb{\{}\{1\},\{3\}\pmb{\}}, \pmb{\{}\{1\},\{2\},\{3\}\pmb{\}}, \pmb{\{}\{1,2\},\{3\}\pmb{\}}, \pmb{\{}\{1,3\},\{2\}\pmb{\}}, \pmb{\{}\{2,3\},\{1\}\pmb{\}},$\\$\pmb{\\{1,2\},\{1,3\}\pmb{\}}, \pmb{\\{1,2\},\{2,3\}\pmb{\}}, \pmb{\\{1,3\},\{2,3\}\pmb{\}}, \pmb{\\{1,2\},\{1,3\},\{2,3\}\pmb{\}}$.} An individual with $G_i = g\left(\{1,2\},\{1,3\},\{2,3\}\right)$, for instance, would take treatment so long as any two of the instruments take a value of one.
In a slight abuse of notation, let $D_g(\mathbf{z})$ the potential treatments function $D_i(\mathbf{z})$ that is common to all units sharing a value $g$ of $G_i$. A key difference between VM and IAM for identification is that under VM, the functions $D_g(\mathbf{z})$ for various response groups $g$ are not all linearly independent of one another. Indeed, as functions of $J$ binary variables, only $2^J$ such $D_g(\mathbf{z})$ could be independent, while $\mathrm{Ded}_J$ is strictly larger than $2^J$ for $J>1$. Let $\mathcal{G}^{c}:=\mathcal{G}/\{a.t., n.t.\}$ denote the set of $\mathrm{Ded}_J-2$ response groups aside from the never-takers and always takers that are compatible with VM. All of the groups in $\mathcal{G}^c$ can be thought of as generalized “compliers” of some kind: units that would vary treatment uptake in some way in response to counterfactual changes to the values of the instruments.
We can construct a natural linear basis for the set of selection functions $\{D_g(\mathbf{z})\}_{g \in \mathcal{G}^c}$ by considering response groups $g(F)$ corresponding to Sperner families that consist of a single set $S$. I refer to such response groups, denoted $g(S)$, as simple.\footnote{Note that a similar construction plays a central role in Lee2018a.} For $J=2$, for example, we have: $$D_{Z_1}(\mathbf{z})=z_1 \hspace{1cm} D_{Z_2}(\mathbf{z})=z_2\hspace{1cm} D_{reluctant}(\mathbf{z})=z_1\cdot z_2$$ The selection function for the remaining group, eager compliers, can then be obtained as: $$D_{eager}(\mathbf{z}) = z_1+z_2-z_1\cdot z_2= D_{Z_1}(\mathbf{z})+D_{Z_2}(\mathbf{z})-D_{reluctant}(\mathbf{z})$$ We can express this linear dependency across all groups by the matrix $M$ in the system:
Let $\mathcal{G}^s$ be the set of all simple response groups in $\mathcal{G}^c$. The set $\mathcal{G}^s$ is isomorphic to the collection of all subsets of $\{1 \dots J\}$ aside from the empty set (which corresponds to always-takers). For arbitrary $J$, we can define a $|\mathcal{G}^c| \times |\mathcal{G}^s|$ matrix $M$ that generalizes the linear system ((ref)): $$D_{g}(\mathbf{z}) = \sum_{g' \in \mathcal{G}^s} M_{gg'} \cdot D_{g'}(\mathbf{z}) \quad \textrm{ for all } g \in \mathcal{G}^c \textrm{ and } \mathbf{z} \in \mathcal{Z}$$ Let $F(g)$ denote the Sperner family corresponding to a given $g \in \mathcal{G}^c$ (i.e. the inverse of the function $g(G)$ in Definition 1). For any $g \in \mathcal{G}^s$, let $S(g)$ denote the lone set $S$ in $F(g)$. Given this notation, the entries of $M$ are given explicitly by the following expression:
Fore completeness, note that for $g \in \mathcal{G}^s$, we have $D_{g}(\textbf{z}) = \left(\prod_{j \in S(g)}z_{j}\right) = \mathbbm{1}\left(z_j=1 \textrm{ for all }j \in S(g)\right)$.
In this section I define and characterize the class of causal parameters that are point identified under vector monotonicity, assuming that the instruments are binary and have full support.
Appendix (ref) shows that the assumption of binary instruments is without loss of generality in the sense that if one begins with finite discrete instruments satisfying vector monotonicity, these discrete instruments can be re-expressed as a larger number of binary instruments in a way that preserves VM. Appendix (ref) also relaxes the assumption of full-support, which is necessary to make use of this mapping.
To build up parameters of interest, I consider conditional averages of either potential outcome $Y_i(0)$ or $Y_i(1)$, after possible transformation by a function $f$. For $d\in \{0,1\}$, let
where $C_i = c(G_i,Z_i)$ is any function $c: \mathcal{G} \times \mathcal{Z} \rightarrow \{0,1\}$ of individual $i$'s response group and their realization of the instruments. Intuitively, the event $C_i=1$ will indicate that unit $i$ belongs to a particular subgroup of generalized “compliers”. Allowing $c$ to depend on $Z_i$ in addition to $G_i$ lets the practitioner focus attention on compliers that are responsive to some rather than all of the instruments, as described in Section (ref). Functions of the form $c(g,z)$ are the most general type of conditioning event that depends on the primitives of the IV model given in Section (ref), without depending directly on potential outcomes.\footnote{However, no restrictions are imposed on the joint distribution of $(Y_i(1),Y_i(0),G_i)$, so the model is compatible with $G_i$ being arbitrarily correlated with potential outcomes or with treatment effects, as in Roy-type models.}
Most of the discussion will center on the class of average treatment effect parameters: $$\Delta_c := \mathbbm{E}[Y_i(1)-Y_i(0)|C_i=1] = \mu^{1}_c - \mu^{0}_c$$ with $f(y)=y$ the identity function (for this reason I leave $f$ implicit in the notation $\mu^{d}_c$). The form $\Delta_c$ nests many treatment effect parameters familiar both from the LATE Imbens2018 and marginal treatment effects Heckman2005 literatures. For instance, with a single binary instrument the LATE sets $c(g,\mathbf{z})=\mathbbm{1}(g=complier)$.
I now characterize the family of $c(g,\mathbf{z})$ under which identification of $\mu^{d}_c$ is possible. In particular, a necessary and sufficient condition will be what I call “Property M”:
I'll also say that a parameter $\mu^{d}_c$ or $\Delta_c$ “satisfies Property M” if its underlying function $c(g,\mathbf{z})$ does. Recall that the matrix $M$ is defined in Proposition (ref).
While Property M is somewhat abstract, the discussion in Section (ref) will give intuition for its role in identification. Additionally, the following result connects Property M to the basic logic of of Imbens2018 underlying identification under IAM:
Proposition (ref) shows that average treatment effects that satisfy Property M can be written in the form $\Delta_c = \mathbbm{E}\left[Y_i(1)-Y_i(0)\left|i \in \bigcup_{k=1}^K \left\{i': D_{i'}(u_k(Z_i))>D_{i'}(l_k(Z_i))\right\}\right.\right]$. Specific examples are discussed in Section (ref). The restriction $\mathbf{u}_k(\mathbf{z}) \ge \mathbf{l}_k(\mathbf{z})$ implies that the expansion of $c(g,z)$ is into terms that each take a value of zero or one given VM, and $\mathbf{l}_k(\mathbf{z}) \ge \mathbf{u}_{k+1}(\mathbf{z})$ implies that only one of these can be equal to one for a given $(g,\mathbf{z})$. The proof of Proposition (ref) shows that we can also take $K \le J/2$ without loss of generality.
As an example of a function $c$ that does not satisfy Property M, consider $c(g',\mathbf{z}) = \mathbbm{1}(g'=g)$, a function that picks out a single response group $g$. This cannot be written in the form of Proposition (ref) under VM when $J>1$, and we cannot in general identify the average treatment effect $\mathbbm{E}[Y_i(1)-Y_i(0)|G_i=g]$ within single response groups $g$.\footnote{We can see this in a simple example with $J=2$ and $g$ being a $Z_1$ complier. In this case Property M would require that $c(\textrm{eager},\mathbf{z}) = c(Z_1,\mathbf{z})+c(Z_2 ,\mathbf{z})-c(\textrm{reluctant},\mathbf{z})$, i.e. that $0=1+0-0$, by Eq. ((ref)).} Another example is the ATE, as $c(g,\mathbf{z})=1$ for all $g \in \mathcal{G}$ including always- and never-takers violates item i) of Property M.
Causal parameters that satisfy Property M are identified under VM with binary instruments, provided the various instruments provide sufficient independent variation in treatment uptake. A simple sufficient condition for this is that the instruments have full (rectangular) support. This assumption is stronger than necessary (Appendix (ref) gives a generalization), but simplifies presentation. Let $\mathbb{S}_Z:=\{\mathbf{z} \in \mathcal{Z}: P(Z_i=\mathbf{z})>0\}$ be the support of the random variable $Z_i$.\footnote{I distinguish between $\mathbb{S}_Z$ and the set of conceivable values $\mathcal{Z}$ because some results (e.g. Proposition (ref)) can leverage the assumption that VM holds for values $\mathbf{z} \in \mathcal{Z}$ even if they have zero probability in the population.}
An alternative expression of Assumption 3 is useful for stating the constructive identification result below. For an arbitrary ordering $g_1 \dots g_k$ of the $k:=2^J-1$ simple response groups in $\mathcal{G}^s$, define a $k$-component random vector $\Gamma_i=(D_{g_1}(Z_i), \dots, D_{g_k}(Z_i))'$ where each component gives the treatment status for a particular response group given realization $Z_i$ of the instruments.\footnote{Equivalently, $\Gamma_i= (Z_{S_1i} \dots, Z_{S_ki})'$ for some arbitrary ordering of the $k=2^J-1$ non-empty subsets $S \subseteq \{1\dots J\}$, where $Z_{Si} := \prod_{j \in S} Z_{ji}$ and $g_\ell = g(S_\ell)$ for $\ell=1 \dots k$.} Let $\Sigma$ be the $k \times k$ variance-covariance matrix of $\Gamma_i$.
Lemma (ref) demonstrates that full support of the instruments is equivalent to there being linearly independent variation in treatment take-up among all of the simple response groups.
Theorem (ref) provides an explicit estimand for $\mu^{d}_c$ when the function $c$ satisfies Property M:
It follows immediately from Theorem (ref) that conditional average treatment effects $\Delta_c=\mu^{1}_c - \mu^{0}_c$ satisfying Property M are identified, and the expression simplifies to: $\Delta_c = \mathbbm{E}[h(Z_i)Y_i]/\mathbbm{E}[h(Z_i)D_i]$ (using that $ \mathbbm{1}(D_i=0)+\mathbbm{1}(D_i=1)=1$). Note that as the numerator of $\Delta_c$ depends on $Z_i$ and $Y_i$ only and the denominator depends on $Z_i$ and $D_i$ only, identification of $\Delta_c$ would hold in a “split-sample” setting where $Y_i$ and $D_i$ are not necessarily known for the same individual.
Now I show that Theorem (ref) has a converse: any identified $\Delta_c$ must satisfy Property M. In this sense, Property M is both necessary and sufficient for identification. To state this result, let us consider so-called “IV-like estimands” introduced by Mogstad2018, which are any cross moment $\mathbbm{E}[s(D_i,Z_i)Y_i]$ between $Y_i$ and a function of treatment and instruments for some function $s$. Let $\mathcal{P}_{DZ}$ denote the joint distribution of $D_i$ and $Z_i$, which is identified. Then:
Since the identification approach of MTW2 relies on IV-like estimands for identification, Theorem (ref) implies that any parameter of the form $\Delta_c$ that is identified by MTW2's approach is also identified by Theorem (ref) (provided no additional restrictions are leveraged with MTW2's approach).\footnote{In saying that a parameter $\theta$ is identified by a particular set of empirical estimands, I mean that the set of values of $\theta$ that are compatible with the empirical estimands and maintained assumptions is a singleton, for any joint distribution of the model primitives---in this case $(G_i,Y_i(1), Y_i(0),Z_i)$---that is compatible with those assumptions idzoo.} In Appendix (ref) I show that the reverse is also true: when MTW2's approach to identification is leveraged with a “complete” set of IV-like estimands, it also point identifies all parameters $\Delta_c$ that satisfy Property M.
Before turning to examples, this section provides an algebraic intuition for Theorem (ref). For simplicity, I focus on average treatment effect parameters $\Delta_c$.
By Assumption 1 and the law of iterated expectations, we can write any $\Delta_c$ as a weighted average over response-group specific average treatment effects $\Delta_g:=\mathbbm{E}[Y_i(1)-Y_i(0)|G_i=g]$:
where notice that the weight on $\Delta_g$ is proportional to the quantity $\mathbbm{E}[c(g,Z_i)]$, as well as $P(G_i=g)$. Now consider a general type of IV estimand in which a single scalar $h(Z_i)$ is constructed from the vector of instruments $Z_i$ according to a function $h$, and then used as a single “instrument” in linear IV regression.\footnote{Special cases of this form include 2SLS: $h(\mathbf{z}) = \mathcal{P}(\mathbf{z})$, and Wald-type estimands: $h(\mathbf{z}) = \frac{\mathbbm{1}(Z_i=\mathbf{z})}{P(Z_i=\mathbf{z})}-\frac{\mathbbm{1}(Z_i=\mathbf{z'})}{P(Z_i=\mathbf{z'})}$.} Some algebra shows that under Assumption 1:
These estimands therefore also uncover a weighted average of the $\Delta_g$, similar to ((ref)). In ((ref)) however, the weight placed on each response group $g$ is governed by the covariance between $D_g(Z_i)$ and $h(Z_i)$. Thus a simple IV estimand using $h(Z_i)$ can identify $\Delta_c$ if the function $h$ is chosen in such a way that $Cov(D_{g}(Z_i), h(Z_i))=\mathbbm{E}[c(g,Z_i)]$ for each of the response groups $g$. Since the covariance operator is linear, the linear dependencies among the functions $D_g(\cdot)$ captured by the matrix $M$ in Section (ref) translate into a linear restrictions that must hold among the $\mathbbm{E}[c(g,Z_i)]$. Property M guarantees that the $\mathbbm{E}[c(g,Z_i)]$ satisfy these restrictions, regardless of the distribution of $Z_i$. What remains is then to simply “tune” the covariances $Cov(D_{g}(Z_i), h(Z_i))$ for each simple response group $g \in \mathcal{G}^s$ by careful choice of $h(\cdot)$. This is possible when the instruments have full support via the function $h(\cdot)$ defined in Theorem (ref). A direct proof of Theorem (ref) along these lines is provided in the Online Appendix. The main proof provided in Appendix (ref) is more involved, and is structured around building a foundation for the comparison to the identification approach of MTW2 in Appendix (ref).
The need for Property M in Theorem (ref) thus arises from there being under VM more response groups in $\mathcal{G}^c$ than there are independent pairs of points in the support of the instruments. This contrasts with IAM, under which both are generally equal (with binary instruments) to $2^J-1$.\footnote{Under IAM, there is an order on the $2^J$ points in $\mathcal{Z}$ such that between any two adjacent instrument values $\mathbf{z}, \mathbf{z}'$ along that order, there is a type of complier $g$ that first takes treatment at $\mathbf{z}$, and $\Delta_g = \frac{\mathbbm{E}[Y_i|Z_i=\mathbf{z}']-\mathbbm{E}[Y_i|Z_i=\mathbf{z}]}{\mathbbm{E}[D_i|Z_i=\mathbf{z}']-\mathbbm{E}[D_i|Z_i=\mathbf{z}]}$.} As a result, it is possible under IAM to identify $\Delta_{g}$ for any single such response group $g \in \mathcal{G}^c$. However, under VM the corresponding choice $c(g',\mathbf{z}) = \mathbbm{1}(g'=g)$ fails to satisfy Property M, as described in Section (ref).
This section highlights some of parameters $\Delta_c$ that are identified under VM according to Theorem (ref), and discusses their interpretation in the returns to schooling setting mentioned in the introduction. Let $\Delta_i = Y_i(1)-Y_i(0)$ be the treatment effect for unit $i$, and $\mathcal{J} \subseteq \{1, \dots J\}$ be any subset of the instruments. Proposition (ref) shows that each of the parameters introduced in Table (ref) below satisfy Property M when $\mathcal{Z} = \{0,1\}^J$, and are hence identified by Theorem (ref).\footnote{ Some further examples of identified parameters from those mentioned in Table (ref) can be constructed using a closure property of the set of $c$ satisfying Property M. Let $\mathcal{C}$ denote the set of $c: \mathcal{G} \times \mathcal{Z} \rightarrow \{0,1\}$ that satisfy Property M, and let $c_a(g,\mathbf{z})$ and $c_b(g,\mathbf{z})$ be two functions in $\mathcal{C}$. Then $c_a(g,\mathbf{z})-c_b(g,\mathbf{z}) \in \mathcal{C}$ iff $c_b(g,\mathbf{z}) \le c_a(g,\mathbf{z})$ for all $\mathbf{z} \in \mathcal{Z},g \in \mathcal{G}^c$. We can use this observation to generate identified parameters that condition on the complement of the complier group for ${c_b}$ within the larger complier group for ${c_a}$. For example with $J=2$, consider the average treatment effect among individuals who are counted in the ACLATE but not in $SLATE_{\{1\}}$: $\mathbbm{E}[\Delta_i|G_i \in \mathcal{G}^c \textrm{ but } \left\{D_i(1,Z_{2i}) = D_i(0,Z_{2i})\right\}]$. This represents the average effect among individuals that would not respond to reduction in college tuition alone, but would respond if both tuition and proximity were shifted in concert.}
I call the first item in Table (ref) the “all-compliers LATE” (ACLATE). The ACLATE is the average treatment effect among all units who would change their treatment uptake in any way in response to the instruments, and is the largest subgroup of the population for which treatment effects can be generally point identified from instrument variation alone. In the returns to schooling example, the ACLATE can be described as the average treatment effect among individuals who would go to college were it close and cheap, but would not were it far and expensive.
A set local average treatment effect, or $SLATE_\mathcal{J}$, captures the average treatment effect among units that move into treatment when all instruments in some fixed set $\mathcal{J}$ are changed from 0 to 1, with the other instruments not in $\mathcal{J}$ remaining at that unit's realized values. The ACLATE is a special case of SLATE when $\mathcal{J} = \{1, \dots J\}$. In the other extreme where $\mathcal{J}$ contains just one instrument index, SLATE recovers treatment effects among those who would “comply” with variation in that instrument alone. For example, $SLATE_{\{2\}}$ is the average treatment effect among individuals who don't go to college if it is far, but do if it is close (given their realized value of the tuition instrument).\footnote{Note that a single-instrument SLATE like $SLATE_{\{2\}}$ does not generally correspond to using $Z_2$ alone as an instrument, e.g. $Cov(Y,Z_2)/Cov(D,Z_2)$, unless $Z_1$ and $Z_2$ are independent.} This parameter may for example be of interest to policymakers considering whether to expand a community college to a new campus, and is related to the marginal treatment effect curve for instrument $j$ (see Appendix (ref)).
The treatment effect parameters $SLATT_{\mathcal{J}}$ and $SLATU_{\mathcal{J}}$ are similar to $SLATE_{\mathcal{J}}$ but additionally condition on units' realized treatment status. For example $SLATT_{\{1,2\}}$ with our two instruments averages over individuals who do go to college, but wouldn't have gone were it far and expensive.\footnote{Note that with a single binary instrument, $SLATT_{\{1\}}$ coincides with $ACLATE=SLATE_{\{1\}}$, as $\mathbbm{E}[\Delta_i|D_i=1, G_i = complier]=\mathbbm{E}[\Delta_i|Z_i=1, complier] = \mathbbm{E}[\Delta_i|complier]$, using Assumption 1. However, when the group $\mathcal{G}^c$ consists of more than one group, the “all-compliers” version of $SLATT$ generally differs from $ACLATE$.} The final row of Table (ref) gives the most disaggregated type of identified parameter that can be identified under VM, what might be called a partial treatment effect $PTE_j(\mathbf{z}_{-j})$. This is the average treatment effect among individuals that move into treatment when a single instrument $j$ is shifted from zero to one, while the other instrument values are held fixed at some explicit vector of values $\mathbf{z}_{-j}$. An example is the average treatment effect among individuals who go to college if it is close and cheap, but do not if it is far and cheap.
Section (ref) discusses estimation of the parameters listed in Table (ref). The next section first outlines some further remarks on identification under VM.
1) Linear dependency among the instruments. Assumption 3 is stronger than is strictly necessary for identification, since linear dependencies between products of the instruments pose no problem if the corresponding “weights” in $\Delta_c$ do not need be tuned independently from one another. In Appendix (ref), I give a version of Assumption 3 and generalization of Theorem (ref) that can accommodate instrument support restrictions and instruments that are not binary.
2) Conditional distributions of the potential outcomes. By choosing $f(Y) = \mathbbm{1}(Y \le y)$ for a value $y$ in the support of $Y_i$, we can through Theorem (ref) identify the CDF of each potential outcome at $y$ conditional on $C_i=1$ (provided that $(Y_i,Z_i,D_i)$ are all observed in the same sample). This allows for the identification of quantile treatment effects or bounds on the distribution of treatment effects Theory2019, in either case conditional on $C_i=1$.
3) Identified sets for ATE, ATT, and ATU. When $Y_i$ has bounded support, we can generate sharp bounds in the spirit of Manski2019 for parameters like the average treatment effect (ATE), using the identified parameters in this paper. For example, the ATE can be written as $ATE := \mathbbm{E}[Y_i(1)-Y_i(0)] = p_a \Delta_{a} + p_n \Delta_{n} + (1-p_n-p_a)ACLATE$, where $\Delta_a=\mathbbm{E}[Y_i(1)-Y_i(0)|G_i=a.t.]$, and $p_a=P(G_i=a.t.)$ (and analogously for $\Delta_n$ and $p_n$). Both $p_a$ and $p_n$ are point identified, while bounds on $\Delta_n$ and $\Delta_a$ can be obtained from the support of $Y_i$. The SLATT and SLATU can similarly be used to construct bounds on the average treatment effect on the treated (ATT) or untreated (ATU).
This section proposes a simple two-step estimator for the family of identified causal parameters introduced in Section (ref), focusing on the conditional average treatment effects $\Delta_c$. The estimator is asymptotically normal and converges at the parametric rate.
Theorem (ref) establishes that a $\Delta_c$ satisfying Property M is identified by a ratio of two population expectations $\mathbbm{E}[h(Z_i)Y_i]/\mathbbm{E}[h(Z_i)D_i]$, so a natural estimator simply replaces these expectations with their sample counterparts, plugging in a first-step estimate of $h(Z_i)$. Recall that the function $h(\cdot)$ is defined from the vector $\mathbf{\lambda} =(\mathbbm{E}[c(g_1,Z_i)], \dots \mathbbm{E}[c(g_k,Z_i)])'$ where $k=2^J-1$. Given an $i.i.d.$ sample of size $n$ and a consistent estimator $\hat{\mathbf{\lambda}}$ of $\mathbf{\lambda}$, we can estimate $\Delta_c$ by $\hat{\rho}(\hat{\mathbf{\lambda}})$, where
and we introduce a $n \times 2^J$ matrix $\Gamma$ comprised of rows $(0,\Gamma_i')$ for each observation $i$, as well as $n \times 1$ vectors $D$ and $Y$ comprised of observations of $D_i$ and $Y_i$.\footnote{To obtain ((ref)), recall that $h(Z_i) = (\Gamma_i - \mathbbm{E}[\Gamma_i])'\mathbbm{E}[(\Gamma_i - \mathbbm{E}[\Gamma_i])(\Gamma_i - \mathbbm{E}[\Gamma_i])']^{-1} \mathbf{\lambda}$. Accordingly, let $\hat{H} = n \tilde{\Gamma}(\tilde{\Gamma}'\tilde{\Gamma})^{-1}\hat{\mathbf{\lambda}}$ be a vector $\hat{H}$ of estimates for $h(Z_i)$, where $\tilde{\Gamma}$ is a $n \times k$ matrix with entries $\tilde{\Gamma}_{il} = D_{g_\ell}(Z_i) - \frac{1}{n}\sum_{j=1}^n D_{g_\ell}(Z_j)$, and $g_\ell$ is the $\ell^{th}$ response group for some arbitrary ordering of the $k:=2^J-1$ groups $g_\ell \in \mathcal{G}^s$. Now consider $\hat{\Delta}_c = (\hat{H}'D)^{-1}(\hat{H}'Y)$, where $Y$ and $D$ are $n \times 1$ vectors of observations of $Y_i$ and $D_i$, respectively. By the Frisch-Waugh-Lovell theorem, $(\tilde{\Gamma}'\tilde{\Gamma})^{-1}\tilde{\Gamma}'D$ is the same as the final $k$ components of the vector $(\Gamma'\Gamma)^{-1}\Gamma'D$. Thus $\mathbf{\lambda}'(\tilde{\Gamma}'\tilde{\Gamma})^{-1}\tilde{\Gamma}'D = (0,\mathbf{\lambda}')(\Gamma'\Gamma)^{-1}\Gamma'D$, and similarly for $Y$.} Note that the population analog of $(\Gamma'\Gamma)^{-1}$ exists by Assumption 3. However the RHS of ((ref)) is still consistent for $\Delta_c$ when Assumption 3 is relaxed as in Appendix (ref) (with $\lambda$ and $\Gamma$ modified as described therein).
Table (ref) below gives examples of $\hat{\mathbf{\lambda}}$ for leading treatment effect parameters. With $\hat{\lambda} \stackrel{p}{\rightarrow} \lambda$ in each of these examples, we have that $\hat{\rho}(\hat{\mathbf{\lambda}}) \stackrel{p}{\rightarrow} \sum_{g \in \mathcal{G}^c}\frac{P(G_i=g)[M\mathbf{\lambda}]_g}{\sum_{g' \in \mathcal{G}^c}P(G_i=g')[M\mathbf{\lambda}]_{g'}}\cdot \Delta_g$ under standard regularity conditions. Matching this estimand to particular parameters $\Delta_c$ that satisfy Property M is achieved by choosing $\hat{\mathbf{\lambda}}$ appropriately for that $\Delta_c$. Asymptotic normality $\hat{\rho}(\hat{\mathbf{\lambda}})$ follows as a special case of Theorem 3 in Imbens2018, which provides an expression of the estimator's asymptotic variance. The estimator $\hat{\Delta}_c=\hat{\rho}(\hat{\mathbf{\lambda}})$ and accompanying confidence intervals for $\Delta_c$ are implemented in the Stata package ivcombine.\\
Comparison with 2SLS: The estimator $\hat{\Delta}_c$ has a similar form to a “fully-saturated” 2SLS estimator that includes an indicator for each value of $Z_i$ in the first stage. Indeed, this version of 2SLS can be written in the form $\hat{\rho}(\mathbf{\lambda})$ where the components of $\mathbf{\lambda}$ are sample covariances between $D_i$ and a given component of $\Gamma_i$, corresponding to the estimand: $\rho_{2sls}= \sum_{g \in \mathcal{G}^c} \frac{ P(G_i=g)\cdot Cov(D_i,D_g(Z_i))}{\sum_{g'} P(G_i=g')\cdot Cov(D_i,D_{g'}(Z_i))}\cdot \Delta_g$. The weights that 2SLS uses to aggregate over linear projection coefficients $(\Gamma'\Gamma)^{-1}\Gamma'D$ and $(\Gamma'\Gamma)^{-1}\Gamma'Y$ are thus determined asymptotically by the joint distribution of $D_i$ and $Z_i$, which the researcher has no control over. MTW show that the implied weight on some $\Delta_g$ under VM may be negative, depending on the DGP. By contrast, $\hat{\Delta}_c$ uses $\hat{\lambda}$ chosen to match the desired parameter of interest, guaranteeing that the estimator recovers a well-defined causal parameter under VM. Even with a large number of instruments, $\hat{\Delta}_c$ is no more “expensive” than 2SLS: both involve computing two linear projections each with $2^J$ of terms (despite the fact that number of underlying selection groups is much larger under VM compared with IAM).\\
Estimation of the ACLATE from a single Wald ratio: The population estimand corresponding to the all-compliers LATE takes on a particularly simple form, a single “Wald ratio”:
where $\bar{Z} = (1\dots1)'$ and $\underbar{Z} = (0\dots0)'$, provided that $P(Z_i=\bar{Z})>0$ and $P(Z_i=\underbar{Z})>0$, and the denominator is non-zero. This can be shown via a Corollary to Theorem (ref) presented in Appendix (ref). By ((ref)), a very simple consistent estimator of the ACLATE is thus: $\widehat{ACLATE} := \frac{\hat{\mathbbm{E}}[Y_i|Z_i=\bar{Z}]-\hat{\mathbbm{E}}[Y_i|Z_i=\underline{Z}]}{\hat{\mathbbm{E}}[D_i|Z_i=\bar{Z}]-\hat{\mathbbm{E}}[D_i|Z_i=\underline{Z}]}$. It turns out that $\widehat{ACLATE}$ is in fact numerically equivalent in finite sample to $\hat{\Delta}_c=\hat{\rho}((1\dots 1)')$ obtained via Eq. ((ref)).\footnote{To see this, note that the vector $H$ of $H_i$ solves the system of equations $\Gamma' H_i = (1 \dots 1)'$. Among vectors that are in the column space of $\Gamma$, $H$ is the unique such solution, given that the design matrix $\Gamma$ has full column rank. One can readily verify that $\Gamma' H = (1\dots 1)$ with the choice $H_i = \frac{\mathbbm{1}(Z_i=(1\dots1))}{\hat{P}(Z_i=(0\dots0))} -\frac{\mathbbm{1}(Z_i=(0\dots0))}{\hat{P}(Z_i=(0\dots0))}$, and that this $H = \Gamma \eta$ with $\eta=(1/\hat{P}(Z_i=(1\dots1)), 0, \dots 0, -1/\hat{P}(Z_i=(0\dots0)))$.} In situations where $Z_i$ has non-zero but small probability for the points $\bar{Z}$ and $\underbar{Z}$, we may thus expect that $\hat{\Delta}_c$ may perform poorly as an estimator of ACLATE in small samples, since it effectively ignores all of the data for which $Z_i \notin \{\underline{Z},\bar{Z}\}$. This issue also arises in the context of IAM Frolich2007, in which case $\hat{\rho}_{\bar{Z}, \underline{Z}}$ is also consistent for the ACLATE with finite $\mathcal{Z}$.\footnote{An analogous result to Eq. ((ref)) holds under IAM with finite instruments, where in that case we take any $\bar{Z} \in \textrm{argmax}_{z}\mathbbm{E}[D_i|Z_i=\mathbf{z}]$ and $\underbar{Z} \in \textrm{argmin}_{z}\mathbbm{E}[D_i|Z_i=\mathbf{z}]$, and define $\mathcal{G}^c:= \{g \in \mathcal{G}: \mathbbm{E}[D_g(Z_i)] \in (0,1)\}$.} Regularization of the estimator to make use of other points in the support of $Z_i$---at the expense of some finite-sample bias---may be useful in improving performance in such cases.\\
Covariates: In Appendix (ref), I describe how covariates can be accommodated in estimation when instrument independence holds only after conditioning on observed variables $X$. The main result is that while conditional average treatment effects $\Delta_c(x):=\mathbbm{E}[Y_i(1)-Y_i(0)|C_i=c,X_i=x]$ can be identified for each $x$ in the support of $X_i$, the unconditional $\Delta_c$ can be easier to estimate. A particularly simple case occurs when the conditional expectation functions $\mathbbm{E}[Y_i|Z_i=\mathbf{z},X_i=x]$ and $\mathbbm{E}[D_i|Z_i=\mathbf{z},X_i=x]$ are each additively separable between $\mathbf{z}$ and $x$, and linear in $x$. A simple consistent estimator of $\Delta_c$ is then: $ \left((0,\hat{\mathbf{\lambda}}')(\Gamma'\mathcal{M}_X\Gamma)^{-1}\Gamma'\mathcal{M}_XD\right)^{-1}(0,\hat{\mathbf{\lambda}}')(\Gamma'\mathcal{M}_X\Gamma)^{-1}\Gamma'\mathcal{M}_XY$, where $\mathcal{M}_X$ is an orthogonal projection matrix for observations of $X_i$. In this case, the only modification to $\hat{\Delta}_c$ required is to add $X_i$ as additional regressors to the linear projections of $Y_i$ and $D_i$ onto the instruments $\Gamma_i$. I implement this estimator in the empirical application below.
In this section I apply the results of this paper to a well-known setting in which multiple instruments have been used: the labor market returns to college. While most existing literature has in this setting bases IV methods on the traditional IAM notion of monotonicity (or on an assumption of homogeneous treatment effects), I instead base estimates on the identification results of this paper that hold under VM. This approach reveals new evidence of heterogeneity in treatment effects across groups that differ in their counterfactual selection behavior, under more plausible assumptions.
I use the dataset from \citet*{Carneiro2011} (henceforth CHV) constructed from the 1979 National Longitudinal Survey of Youth. This setting is also considered by MTW (under assumption of PM). The sample consists of 1,747 white males in the U.S., first interviewed in 1979 at ages that ranged from 14 to 22, and then again annually. The outcome of interest $Y_i$ is the log of individual $i$'s wage in 1991, and treatment $D_i=1$ indicates $i$ attended at least some college. As in CHV, treatment effects are expressed in approximate per-year equivalents by dividing them by four.
CHV consider four separate instruments for schooling. In a baseline setup, I use the two binary instruments discussed throughout this paper: tuition and proximity. In particular, I let $Z_{1i}$ indicate that average tuition rates local to $i$'s residence around age 17 falls below the sample median, which corresponds to about \$2,170 in 1993 dollars. I let $Z_{2i}$ indicate the presence of a public college in $i$'s county of residence at age 14. I later add two additional instruments used by CHV, which capture local labor market conditions when a student is in high school.
While VM is a natural assumption for the tuition and proximity instruments, a conditional version of instrument validity is more plausible than Assumption 1. I follow CHV and include a set of control variables $X_i$,\footnote{In particular, a student's corrected Armed Forces Qualification Test score, mother's years of education, number of siblings, “permanent” local earnings in county of residence at 17, “permanent” unemployment in county of residence at 17, earnings in county of residence in 1991, and unemployment in state of residence in 1991, along with an indicator for urban residence at 17 and cohort dummies (see CHV for variable definitions and construction). Also following CHV, I include as components of $X_i$ the squares of continuous control variables. All together, these represent the union of variables that CHV use in their first stage and outcome equation, with one exception: As MTW do, I drop years of experience in 1991 since it may itself be affected by schooling. In the two instrument setup, I also add to $X_i$ the two “unused” instruments from CHV and their squares: long-run local earnings in county of residence at 17 and long run unemployment in state of residence at 17.} implemented as described in Section (ref) and Appendix (ref). Standard errors are computed by applying the delta method to the system of estimated regression equations (allowing for heteroscedasticity and cross-correlation between the equations).
The left panel of Table (ref) reports a cross tabulation of the two instruments, which have a weak positive correlation, though the observations are fairly evenly distributed across the four cells.
The right panel of Table (ref) reports predictions from the estimated conditional propensity score function $\mathbbm{E}[D_i|Z_i=\mathbf{z}, X_i=x]$ estimated via a linear regression of $D_i$ on the instruments (and their interaction) as well as $X_i$, then evaluated at the mean $\bar{x}$ of $X_i$.\footnote{I note that when all controls are omitted from this regression, the estimated propensity score function is no longer monotonic in $Z_1$ and $Z_2$. This underscores the potential of VM to be used to evaluate the validity of instruments given a set of conditioning variables (in contrast to PM and IAM, who lack this testable implication).}
The top-left value of $\hat{\mathcal{P}}(expensive, far,\bar{x}) = 45.1\%$ provides an estimate of the overall proportion of always-takers in the population, while the share of never-takers is estimated to be $1-0.53=47.0\%$. The remaining roughly $8\%$ of the population are generalized “compliers” consisting of the tuition ($Z_1$), proximity ($Z_2$), eager and reluctant compliers. From the table we can also see that $P(D_i(expensive,close,x)>D_i(expensive,far,x))$ is estimated to be $5.7\%$, and $P(D_i(cheap,far,x)>D_i(expensive,far,x))$ is estimated to be $3.6\%$ (these quantities are the same for all values of $x$, under the maintained assumption that $\mathbbm{E}[D_i|Z_i,X_i]$ is additively separable between $Z_i$ and $X_i$). Combining these figures and the response group definitions from Section (ref), we see that between 1.5% and 3.6% of the population are estimated to be eager compliers, while no more than 2.1% are reluctant compliers. Similarly, no more than 3.6% are tuition compliers, and between 2.1% and 5.7% are proximity compliers.
Figure (ref) reports estimates of several of the parameters introduced in Section (ref), alongside fully-saturated 2SLS for comparison. Consider first the ACLATE: the point estimate of $0.14$ indicates that having attended a year of college increases 1991 wages of all compliers by roughly 14% on average. This estimate is within the range of roughly $-0.1$ to $0.3$ of the marginal treatment effect (MTE) function estimated by CHV under the assumption of IAM, and is similar to their point estimate of the average treatment on the treated under a parametric normal selection model. The 2SLS estimate from Figure (ref) yields a similar value at $0.12$. Note that given the limited sample size none of the estimates are quite significant at the 90% level. I focus discussion on the point estimates for the sake of illustration with this caveat.
The point estimates from the remaining rows in Figure (ref) suggest that the ACLATE aggregates over substantial heterogeneity in the population. For example, the proximity SLATE indicates that a year of college has no average effect on the wages of individuals who move into treatment if and only if a college is nearby, given local affordability. This group includes proximity compliers, eager compliers for whom college is expensive, and reluctant compliers for whom it is cheap. On the other hand, the SLATE for tuition is about three times as large as the ACLATE. These results suggest that the average treatment effect among tuition compliers is larger than it is among proximity compliers, however the sign of the difference is not identified.\footnote{Note however that in the $J=2$ case, if $\Delta_g$ and the corresponding group size $p_g$ is known ex-ante for one group $g \in \mathcal{G}^c$, then the remaining three group specific treatment effects and group sizes can be point identified.} Note finally that the point estimates for $SLATU$ and $SLATT$ suggest that among the compliers averaged over by the ACLATE, those who in fact go to college have greater treatment effects on average than those who do not, which is consistent with students selecting on the basis of their heterogeneous future gains (as in a Roy-type model).
I now add the additional two instruments from CHV, to increase comparability and emphasize the scalability of my method to several instruments. Let $Z_{3i}$ indicate that local earnings in $i$'s county of residence at 17 is below the sample median, and $Z_{4i}$ that unemployment in $i$'s state of residence at 17 is above the sample 25% percentile (this threshold is chosen as it yields a stronger predictor of college as compared with using the median). The two labor market variables and their squares are removed from the controls $X_i$.
With all four instruments, over 17% of the population are now some type of “complier” and counted in $\mathcal{G}^c$, which now contains 167 underlying response groups (compared with just 7.8% of the population for the four groups in $\mathcal{G}^c$ with the two instruments used before). Nevertheless, computing the treatment effect estimates involves regressions with at most 16 terms in addition to the controls, keeping implementation manageable.
Table (ref) shows that the $ACLATE$ is not much changed from the case with only two instruments, and we again have that the tuition SLATE is much larger and that the proximity SLATE is close to zero. The SLATE for low local wages occupies an intermediate value, while the SLATE for high unemployment is estimated to be negative (but with a much larger standard error). The unemployment SLATE is so imprecisely estimated in part because its corresponding complier group is the smallest of the estimands considered: with just $2\%$ of the population.
In this paper, I have characterized the causal parameters that can be point identified using multiple instruments under a monotonicity assumption that is often motivated by economic theory: vector monotonicity (VM). I accomplish this by focusing on binary instruments, but results are applicable to discrete instruments more generally, as shown in Appendix (ref).
The estimator I propose targets well-defined causal-parameters at no additional computational cost relative to the popular 2SLS estimator, which is not guaranteed to recover an interpretable causal parameter under VM. In an application to the labor market returns to college education, I find leveraging VM that underlying groups in the population which exhibit different selection behavior also have very different average returns to college.
\printbibliography