Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,069 characters · 15 sections · 0 citation commands
Interacting Treatments with Endogenous Takeup
\thispagestyle{empty}
The experimental approach to establishing the causal effect of a treatment is based on allocating units randomly to treatment and control, thereby precluding any systematic difference between the two groups other than the treatment itself. While unit-level treatment effects may be heterogeneous, comparing the average outcome in the treated and control groups gives a consistent estimate of the average treatment effect. This deceptively simple description of the experimental ideal, which originates in the work of Fisher (1925), embodies several further assumptions, formalized later by Rubin (1974, 1978) and others. A lot of subsequent work on causal inference has sought to extend the analysis of experimental data to more complicated situations with some of the following features.
First, in real world experiments perfect compliance with the intended treatment assignment is not always possible or ethical to enforce. If non-compliance is endogenous (i.e., it depends on unobserved confounders), then the average difference between the treated and non-treated values does not represent the treatment effect alone but selection effects as well. Second, randomization itself does not ensure that the treatment status of an individual does not interfere with the potential outcomes of another, violating what is called the stable unit treatment value assumption (SUTVA) in the Rubin causal model. Third, there are experimental setups in which units are potentially subject to multiple, but not mutually exclusive, treatments that may interact with each other.
In this paper we derive the causal interpretation of various instrumental variable (IV) estimands in an experimental or quasi-experimental setup that extends the basic model in all three directions mentioned above. More concretely:
We subsequently present two examples which illustrate these ideas.
Motivated by Example (ref), it will be convenient in the rest of the paper to represent population units as pairs $(A,B)$, where member $A$ is targeted by $D_A/Z_A$ and member $B$ is targeted by $D_B/Z_B$. The outcome of interest may be be associated with member $A$, member $B$ or the pair itself. Thus, we can equate treatment interaction with interference across pair members (a violation of SUTVA). The representation of units in terms of pairs comes without loss of generality, as a single individual targeted by two treatments (such as in Example (ref)) can always be thought of as a “pair” with identical members.
The estimands we consider in this framework derive from two simple IV regressions and a saturated IV regression. Specifically, we study the causal interpretation of the standard Wald estimand associated with treatment/instrument $A$ conditional on $Z_B=0$ and $Z_B=1$, respectively, as well as the IV (2SLS) regression of the outcome on $D_A$, $D_B$ and $D_AD_B$, instrumented by $Z_A$, $Z_B$ and $Z_AZ_B$. There is now a substantial econometrics literature on causal inference in similar settings. Our results contribute to this growing body of work in the following ways.\footnote{In Section (ref) below we provide a brief literature review and position our paper more carefully relative to the most relevant subset of papers.}
First, we employ weaker restrictions on instrument spillovers than other studies in the literature. Specifically, we allow for a compliance type for whom the presence of, say, $Z_B$ represents a strong incentive against taking treatment $A$; we call this type the cross-defiers. At the same time, our framework also accommodates joint compliers---a type that reacts positively to the presence of the partner instrument and ultimately takes the given treatment if (and only if) both instruments are present. While identification results with joint compliers are available in other studies (e.g., Vazquez-Bare 2022), cross-defiers are typically ruled out, despite our application in Section (ref) showing that this is also an empirically relevant type.
Second, our general identification results make explicit the difficulties that instrument spillovers cause in identifying “standalone” average treatment effects and, even more starkly, in separating treatment interaction from treatment effect heterogeneity. In particular, we provide the causal interpretation of the interaction term in the saturated IV regression, and show that a generally interesting local average interaction effect is confounded by terms that depend on how different the average effects of the two treatments are across various subgroups. These confounding terms however vanish if there are no interactive types or there is no treatment effect heterogeneity.
Third, given the general lack of interpretability of the interaction term in the IV regression, we also provide partial identification results---that is, bounds---for a parameter measuring the average interaction effect of the two treatments in a specific subgroup. We bound the interaction effect directly, based on its definition, as well as indirectly by bounding the confounding heterogeneity terms that show up in the interpretation of the interaction term. We further consider a formal Manski-type bounding strategy, which only uses the data and weak auxiliary assumptions, and a less formal strategy that relies on heuristic restrictions on treatment effect heterogeneity. The application illustrates all bounding approaches.
Fourth, while the results discussed in the main text impose one-sided noncompliance on the individual instruments (similarly to many results in the literature), we explore the consequences of relaxing this powerful but restrictive assumption in an appendix to the paper.\footnote{One-sided noncompliance means that the treatment cannot be accessed without receiving the instrument.} Specifically, we extend the analysis of the two Wald estimands to the case in which only one of the two instruments satisfies one-sided noncompliance, and even provide causal interpretations that do not require any monotonicity conditions.
The rest of the paper is organized as follows. In Section (ref) we position our paper in the literature. Section (ref) presents a formal potential outcome framework for pairs with endogenous (non-)compliance. We state and discuss our identification results in Section (ref). Section (ref) applies our theory in the context of Example (ref). Section (ref) concludes. There is a substantial amount of supplementary material relegated to appendices. Appendix (ref) provides supporting identification results. Appendix (ref) explores the relaxation of one-sided non-compliance. Appendix (ref) contains the proofs of all results in the main text, and, finally, Appendix (ref) supplements the empirical application.
The seminal paper by Imbens and Angrist (1994) investigates endogenous non-compliance with a binary instrument, representing intended assignment to a binary treatment, when individual treatment effects are heterogeneous. The study establishes the now well-known result that a simple IV regression identifies the local average treatment effect (LATE) among compliers, the subpopulation whose treatment status complies with the instrument, under conditions that include the weak monotonicity of the treatment in the instrument.
Several studies extend this framework to multiple treatments or instruments, which typically entails refinements of the monotonicity assumption. For example, Mogstad, Torgovitsky, and Walters (2020) impose partial monotonicity of the treatment in one instrument conditional on the other instrument(s), while Goff (2022) invokes vector monotonicity, which implies that each instrument affects treatment uptake in a direction that is common across subjects. Closer to our framework, Behaghel, Cr\'{e}pon and Gurgand (2013) consider two mutually exclusive binary treatments along with two binary instruments. Ruling out instrument spillovers, they identify the LATEs among the two groups of compliers. Kirkeboen, Leuven, and Mogstad (2016) exploit information on next-best treatment alternatives to ease the assumption of no instrument spillovers. Heckman and Pinto (2018) allow for instrument spillovers but impose an unordered monotonicity assumption. The latter requires that if some subjects move into (out of) a treatment when the instrument values are switched, then no subjects can at the same time move out of (into) that respective treatment. Lee and Salani\'{e} (2018) discuss LATE identification without an unordered monotonicity assumption in the presence of sufficiently many continuous (rather than binary) instruments when treatment choice is governed by threshold-crossing models.
Another strand of related literature is concerned with relaxations of SUTVA, allowing for specific forms of interference among treatments, in various (quasi-) experimental settings; see for instance Sobel (2006), Hong and Raudenbush (2006), Hudgens and Halloran (2008), Ferracci, Jolivet and van den Berg (2014) or Huber and Steinmayr (2021). Particularly relevant in our context are the studies by Kang and Imbens (2016), Imai, Jiang and Malani (2021), and Vazquez-Bare (2022), who combine relaxations of SUTVA with treatment non-compliance. In addition, Blackwell (2017) studies the interaction between two randomized treatments with non-compliance with an application in political science. We now discuss the relationship between this last set of papers and ours in more detail.
Kang and Imbens (2016) and Imai, Jiang and Malani (2021) consider a partial interference framework, where interference between a given unit's outcome and a peer's treatment occurs within well-defined clusters, such as geographic regions. (In our setting the clusters are the pairs, i.e., there is only one peer.) In addition to conventional IV assumptions, Kang and Imbens (2016) impose a treatment exclusion restriction. This assumption posits that each unit's treatment uptake depends solely on their own instrument, but not on the peers' instruments. Ruling out coordinated compliance, this restriction permits identification of the direct LATE, i.e., the average effect of the own treatment among compliers. Under an additional one-sided non-compliance assumption, one can also identify the average interference (spillover) effect of the peers' treatments, in the absence of the own treatment, within the whole population. Imai, Jiang, and Malani (2021) also study the identification of the direct LATE and present a condition that holds under the treatment exclusion restriction, but is even satisfied under the weaker condition that a unit's treatment status does not depend on the instrument values of those peers who are non-compliers with respect to their own instruments. Our paper allows for even more general forms of instrument spillovers and makes explicit the resulting difficulties in identifying easily interpretable causal effects.
The framework of our paper is most closely related to Blackwell (2017; henceforth, BW) and Vazquez-Bare (2022; henceforth, VB), but it extends each in certain directions.\footnote{The basic features of our framework were originally developed in the MA thesis of Kormos (2018).} BW does not explicitly consider (partial) interference, but a scenario where a unit's outcome may be affected by two interacting binary treatments, \( D_A \) and \( D_B \), each with a distinct instrument, \( Z_A \) and \( Z_B \), respectively. As discussed above, this setup can be viewed as a special case of our framework where pair members \( A \) and \( B \) are each targeted by a corresponding treatment and instrument, with one unit's treatment potentially having an interference effect on the other unit's outcome. Just as Kang and Imbens (2016), BW imposes a treatment exclusion restriction to identify the direct LATE of a given treatment (say, \( D_A \)) conditional on the other treatment (\( D_B = 0,1\)), among complier units who follow the respective instruments in both their takeup decisions. Again, our framework and VB's are more general in that instrument spillovers are allowed; on the other hand, BW's results do not impose one-sided noncompliance.
VB explicitly considers a paired design and imposes a weaker first stage condition than the treatment exclusion restriction. Specifically, VB assumes that a pair member's treatment status (say, \( D_A \)) is weakly higher when both the unit's and the peer's instruments are activate (\( Z_A = 1 \), \( Z_B = 1 \)) compared to only the unit's instrument being activate (\( Z_A = 1 \), \( Z_B = 0 \)). Furthermore, compared to the latter instrument configuration, the value of $D_A$ is weakly smaller when only the peer's instrument is activated (\( Z_A = 0 \), \( Z_B = 1 \)), and weakly smaller still when neither instrument is active (\( Z_A = 0 \), \( Z_B = 0 \)). In deriving his main results, VB also imposes one-sided non-compliance, which implies the second part of the monotonicity condition above, and facilitates identification of LATE among compliers with nontreated peers as well as the interference effect on untreated units induced by complying peers.
Importantly, our general results are derived under even weaker monotonicity restrictions on the instruments than in VB. While we also maintain one-sided non-compliance with respect to a treatment's “own” instrument, we do not impose any monotonicity restrictions on how the partner instrument affects treatment takeup. As noted in the introduction, we accommodate a non-standard compliance type called cross-defiers, necessary for the application in Example (ref), but not present in any other compliance framework we are aware of. In a supplement to this paper, we also investigate relaxations of one-sided noncompliance.
Given our general assumptions, we show that the Wald estimand for treatment $A$, conditional on $Z_B=0$, reflects the average treatment effect in the union of two groups: compliers with $Z_A$ and cross-defiers for $Z_B$. The second Wald estimand, conditional on $Z_B=1$, has a complicated interpretation in general. However, under auxiliary conditions this estimand also lends itself to an insightful causal interpretation, given by the weighted average of three local average treatment effects. Similarly to our paper, both BW and VB consider the saturated IV regression but under different sets of conditions---BW of course rules out instrument spillovers and assumes statistical independence between the instruments while VB does not allow for cross-defiers and also uses one-sided noncompliance. In BW's framework the coefficient on $D_AD_B$ does identify a meaningful local average interaction effect between the two treatments, while VB does not attempt to draw out any interesting causal parameters buried in this estimand.\footnote{He states a formula in the Supplemental Appendix and only notes that it does not lend itself to a direct causal interpretation.} It is true that in our general framework this coefficient also does not lend itself to a “clean” interpretation. Instead, the interaction effect, often of central interest in applications, is bound up with terms that result from instrument spillovers and treatment effect heterogeneity across various compliance types. We state partial identification results that provide bounds for the interaction effect.
A final paper we must acknowledge here is a recent contribution by Bhuller and Sigstad (2024), which also considers a multiple treatment framework. However, the study differs from ours, as well as from BW and VB, in that it does not focus on identifying a specific LATE or interaction effect within a well-defined group of compliers. Instead, it aims to provide conditions under which a 2SLS regression with multiple treatments consistently estimates the weighted average effect of a given treatment across multiple compliance types, ensuring proper (in particular, non-negative) weights under arbitrary treatment effect heterogeneity. The authors demonstrate that it is necessary and sufficient to have an “average conditional monotonicity” and “no cross effects” condition hold, encompassing the treatment exclusion restriction (along with treatment monotonicity in the own instrument) as a special case. In Section (ref) below we will use more specific insights from this paper in explaining the structure of our own results.
The population consists of ordered pairs of individuals (e.g., married couples); we will refer to the first member of a pair as member $A$ and the second as member $B$. There are two potentially different binary treatments: $D_A$ is targeted at member $A$ and $D_B$ is targeted at member $B$. By representing individual units as pairs with identical members, the setup also accommodates the analysis of two (interacting) treatments received by a single unit.
We are interested in the effect of $D_A$ and/or $D_B$ on some dependent variable $Y$. This outcome may be associated with member $A$ alone, member $B$ alone, or the pair itself. The observed value of $Y$ is given by one of four potential outcomes: $Y(d_A,d_B)$ for $d_A,d_B\in\{0,1\}$. For example, $Y(1,0)$ is the potential outcome if one imposes $D_A=1$ and $D_B=0$, i.e., member $A$ is exposed to treatment $A$, but member $B$ is not exposed to treatment $B$. To make the notation less cluttered, we will omit the comma and simply write $Y(10)$ whenever actual figures ('1' and/or '0') are used in the argument. Using the potential outcomes and the treatment status indicators, we can formally express the observed outcome as
Treatment effect identification is facilitated by a pair of binary instruments, $Z_A$ and $Z_B$, assigned to pair members $A$ and $B$, respectively. We think of these instruments as indicators of (randomly assigned) treatment eligibility or the presence of an exogenous incentive to take the corresponding treatment. The leading example is a randomized control trial, where $Z_A$ and $Z_B$ are the experimenter's intended treatment assignments for pair member $A$ and $B$, respectively. Compliance with these assignments is, however, endogenous and possibly coordinated across pair members. We refer to $Z_A$ as member/treatment $A$'s own instrument and $Z_B$ as the partner instrument. The labels are of coursed reversed for treatment $B$.
Thus, there are four potential treatment status indicators associated with each pair member; they are denoted as $D_A(z_A, z_B)$ for member $A$ and $D_B(z_A, z_B)$ for member $B$, $z_A, z_B\in\{0,1\}$. For example, $D_A(01)$ indicates whether member $A$ of a pair takes up treatment $A$ when they are not assigned ($Z_A=0$) but their partner is assigned to treatment $B$ ($Z_B=1$). The actual treatment status of member $A$ can be written as
There is of course a corresponding formula for $D_B$.
We now formally impose standard IV assumptions on $Z_A$ and $Z_B$.
The exclusion restriction stated in part (i) of Assumption (ref) is one of the defining properties of an instrument, and it justifies (ex-post) the potential outcomes being indexed by $(d_A,d_B)$ only. Part (ii), known as “random assignment,” states that the instrument values $(Z_A,Z_B)$ are exogenously determined. This assumption holds, by design, in an experimental setting where intended treatment assignments are explicitly randomized. Part (iii) states that the intended treatment assignments follow a $2\times 2$ factorial design, i.e., there is a positive fraction of pairs assigned to each of the following four categories: treatment $A$ alone, treatment $B$ alone, both treatments, or neither treatment.
We will impose further assumptions on the potential treatment status indicators in Section (ref).
Let $\mathcal{P}$ be a subset of the population of pairs. We define the following treatment effect parameters and notation:
The parameters $ATE_{A|\bar B}(\mathcal{P})$ and $ATE_{A|B}(\mathcal{P})$ are called local average conditional effects, or LACEs, by BW, while $ATE_{AB}(\mathcal{P})$ is called the local average joint effect (LAJE). For a given group $\mathcal{P}$, the difference between the two conditional effects measures the interaction between the two treatments within $\mathcal{P}$, and is hence termed the local average interaction effect (LAIE) by ibid. That is, \[ LAIE(\mathcal{P})=ATE_{A|B}(\mathcal{P})-ATE_{A|\bar B}(\mathcal{P}). \] If the LAIE is positive, the two treatments reinforce each other, while if it is negative, then they work against each other. One can define analogous LACE parameters for treatment $B$ by interchanging the roles of $A$ and $B$ in the definitions above. The associated joint and interaction effects stay unchanged.
In case the pairs have distinct members, the interpretation of these parameters also depends on the definition of the outcome $Y$. In particular, if $Y$ is associated with pair member $A$ alone, then $ATE_{A|\bar B}(\mathcal{P})$ and $ATE_{A|B}(\mathcal{P})$ measure what is called the direct effect of treatment $A$ by Hudgens and Halloran (2008). On the other hand, if $Y$ is associated with member $B$ alone, then $ATE_{A|\bar B}(\mathcal{P})$ and $ATE_{A|B}(\mathcal{P})$ measure the indirect or spillover effect of treatment $A$ on pair member $B$. For example, if the treatment is vaccination, and the outcome is the incidence of a disease, then the vaccination of member $A$ confers protection on member $A$, but also indirectly protects his or her partner.
The setup presented in Section (ref) assigns four potential treatment indicators to each pair member, corresponding to the four possible incentive schemes represented by $(Z_A, Z_B)$. Without any further restrictions on treatment takeup, the possible configurations of these 8 potential treatment variables partition the population of pairs into $2^8=256$ different compliance profiles. At this level of generality a couple of regression-based estimands can hardly be a meaningful summary of the various average treatment effects across types. Therefore, similarly to VB, we impose one-sided noncompliance with respect to the treatment's own instrument, which dramatically reduces the number of possible compliance profiles.
Assumption (ref) states that neither member of the pair has access to their own treatment unless they have been “randomized in,” i.e., the value of their own instrument is 1. In other words, one-sided noncompliance presumes that the experimenter is able to exclude individuals from all sources of the treatment. Whether or not this assumption is reasonable depends on the institutional setting and details of the underlying experiment, but it often fails in practice. Therefore, we consider relaxations of Assumption (ref) in Appendix (ref).
Under Assumption (ref), each pair member may belong to one of only four compliance types, summarized by the following definition.
\paragraph{Remarks}
Given the four individual compliance types, every pair belongs to one of the 16 compliance profiles $\{s,j,d,n\}\times \{s,j,d,n\}$. For example, $(s,j)$ is the set of pairs where $A$ is a self-complier and $B$ is a joint complier, etc. Furthermore, we will use the notation $(c,\cdot)$ to denote the set of pairs where member $A$ is a complier, etc., and $P(s,n)$ to denote the probability that for a randomly drawn pair $A$ is a self-complier and $B$ is a never-taker, etc.
The following assumption ensures that some of the compliance categories are not vacuous (e.g., there are at least some individuals who respond to their own instrument).
Part (i) of Assumption (ref) means that $P(s\cup d,\cdot)>0$ and $P(\cdot, s\cup d)>0$ while part (ii) means $P(c,c)>0$. These conditions ensure that the IV estimands considered in Section (ref) are well defined.
A testable implication of the existence of joint compliers (with respect to treatment $A$) is that if one runs a simple OLS regression of $D_A$ on $Z_B$ in the $Z_A=1$ subsample, then the coefficient of $Z_B$ should be positive. Conversely, if the coefficient is negative, then cross-defiers must be present in the population.\footnote{A zero coefficient means that takeup is consistent with the treatment exclusion restriction, i.e., only the own instrument matters.} Clearly, joint compliers treat the presence of $Z_B$ as a positive incentive to take treatment $A$, while cross-defiers treat it as a (strong) disincentive. Given the situation, it may be possible to argue that the two behaviors do not exist simultaneously, i.e., one could rule out joint compliers or cross-defiers for treatment $A$, depending on which type is not needed to explain the sign of the regression coefficient. Our empirical application in Section (ref) illustrates how to exploit such simplifications in practice to enhance the interpretation of IV estimands.
Our framework can also accommodate simplifying assumptions on pair formation. For example, one might postulate that there are no $(n,j)$ or $(j,n)$ pairs. This assumption is plausible if member $A$'s utility of taking treatment $A$ is affected by $Z_B$ only through member $B$'s actual treatment status $D_B$, which more formally means that $D_A(z_A,z_B)$ is of the form $f(z_A, D_B(z_A,z_B))$. If $B$ is a never taker then $D_B=0$, and the value of $Z_B$ is not relevant for $A$'s decision. Hence $A$ cannot be a joint complier. For example, $Z_B$ could be a randomized monetary reward payable only on actual takeup of treatment $B$. If $B$ is never treated, the reward is not paid out, and should be irrelevant to $A$. On the other hand, $(n,j)$ or $(j,n)$ pairs may well exist if $A$ has direct access to the incentive represented by $Z_B$. This is the case, for example, if the pair stands for a single unit targeted by two treatments.\footnote{Suppose that individuals participate in a study on the health benefits of physical exercise. Specifically, there are two treatments: running $(D_A)$ and swimming $(D_B)$. The instrument $Z_A$ is a seminar on the health benefits of running and $Z_B$ a seminar on the benefits of swimming. A person who cannot swim will be a never taker with respect to $D_B$. Nevertheless, it is conceivable that for the same person $D_A(10)=0$ but $D_A(11)=1$. This means that a single lecture is not sufficient to convince this person to take up running but after hearing more about the health benefits of exercise, he eventually decides to do so.}
The exact type of a given pair is generally unobserved as it depends on the pair's behavior in counterfactual scenarios. Nevertheless, the observed conditional probabilities
can be used to identify the relative frequency of a number of compliance profiles in the population. Nevertheless, not all probabilities under ((ref)) carry independent information. This is for two reasons: first, for any given $(z_A, z_B)$, the corresponding probabilities add up to 1, and, second, $Z_A=0$ automatically implies $D_A=0$, and $Z_B=0$ automatically implies $D_B=0$ by Assumption (ref). It follows that there are only five independently informative moments, which of course makes it impossible to identify the relative frequencies of all 16 compliance profiles separately.
Lemma (ref) presents the interpretation of five independent conditional probabilities of the form ((ref)). The subsequent corollary provides further results.
\paragraph{Remarks}
Each pair in the target population is associated with an observed 5-vector $(Y, D_{A}, D_{B}, Z_{A}, Z_{B})$. Given a sample of observations on this vector, one may run several different IV regressions using the full sample or a suitable subsample.
The following three theorems state the causal interpretation of these estimands. The proofs are provided in Appendix (ref).
\paragraph{Remarks}
In Appendix (ref) we consider an extension of Theorem (ref) to the case in which $Z_B$ continues to satisfy one-sided noncompliance but $Z_A$ only obeys a general monotonicity condition. The result is similar to Theorem (ref) except that an additional compliance profile must be added to the set of pairs $(s\cup d,\cdot)$; namely, pairs where member $A$ is a “cross-complier” in the sense that they take the treatment in response to any one of the instruments being turned on (possibly the partner instrument $Z_B$ alone). The extension is in fact derived from a very general representation theorem that gives causal interpretations to $\delta_{A0}$ and $\delta_{A1}$ without imposing any monotonicity conditions on the instruments. This latter result is rather too general to be useful in practice; its main value lies in the fact that it can readily be “customized” via auxiliary restrictions that fit the application at hand. For example, the general result also shows that the extension of Theorem (ref) to the case in which $Z_B$ does not satisfy one-sided noncompliance is much more complicated, even when $Z_A$ does obey this restriction.
The next result states the causal interpretation of $\delta_{A1}$.
\paragraph{Remarks}
To understand the causal effects appearing in the special case ((ref)), consider changing $Z_A$ from 0 to 1 conditional on $Z_B=1$. In this case $(c,j)$ pairs will switch from no treatment at all to both treatments, contributing the first term in ((ref)). For $(c,s)$ pairs, member $A$ switches from no treatment to treatment $A$, while member $B$ continues to take treatment $B$ throughout. This contributes the second term. Finally, among $(c,n)$ pairs member $A$ switches from no treatment to treatment $A$, while member $B$ continues to abstain from treatment. This option contributes the last term. As in the absence of cross-defiers $P(c,\cdot)=P(c,j)+P(c,s)+P(c,n)$, the probability weights in ((ref)) sum to one, and are identified from the observed data using Lemma (ref) and Corollary (ref). The general expression ((ref)) follows the same logic --- it reflects the reaction of various types of pairs to changing $Z_A$ from 0 to 1 while maintaining $Z_B=1$. However, there are now more possibilities, including member $B$ cross-defiers dropping $D_B$ when $Z_A$ is turned on. The end result is a linear combination of average treatment effects where some of the weights are negative and do not sum to one.
Finally, Theorem (ref) states the causal interpretation of the elements in the coefficient vector $\beta=(\beta_0,\beta_A,\beta_B,\beta_{AB})'$.
\paragraph{Remarks}
The coefficient $\beta_{AB}$ in Theorem (ref) does not have a clean interpretation because it conflates the interaction between the two treatments with the heterogeneity of the treatment effects across various compliance types. We now show that it is still possible to learn about $LAIE(c,c)$ through bounds constructed under some auxiliary conditions. There are two different approaches. First, it is possible to bound $LAIE(c,c)$ directly, based on the moments in its definition. Second, one can take the causal interpretation of $\beta_{AB}$ in Theorem (ref) as a starting point, and bound the influence of the heterogeneity terms ((ref)) through ((ref)). In doing so, one obtains an indirect bound on $LAIE(c,c)$ as well. We present the direct bounds here in the main text; the indirect ones are stated in Appendix (ref). The application in Section (ref) illustrates both approaches.
There are four conditional means involved in the definition of $LAIE(c,c)$:
The first quantity under ((ref)) is identified directly from the data by the conditional expectation of $Y$ given $D_A=D_B=1$ and $Z_A=Z_B=1$; see Lemma (ref) in Appendix A. We bound the remaining moments in the spirit of Manski (1989, 1990), using the following assumption.
Part $(i)$ states that the potential outcomes are bounded; the fact that the lower bound is set to zero is a normalization. Part $(ii)$ postulates that when treatment $A$ is applied in isolation, it has a positive effect on any individual unit, and the same is assumed about treatment $B$. While researchers often hold prior expectations about the sign of an average treatment effect, the requirement that the sign applies uniformly in the population is a non-trivial homogeneity restriction known as monotone treatment response (Manski (1997)). Importantly, however, part $(ii)$ does not restrict the sign of the interaction effect.
As a first step toward bounding the interaction effect, we provide bounds for the joint effect of the two treatments.
\paragraph{Remarks}
The relationship between the average joint and interaction effect can be written as
While the average effects of treatments $A$ and $B$, applied in isolation, are not identified in the $(c,c)$ subgroup, they are identified in subgroups $(s\cup d,\cdot)$ and $(\cdot,s\cup d)$, respectively (see Theorem (ref)). It is natural to take the identified subgroup effects as a reference point and speculate about other groups on the basis of these. In particular, we may write $ATE_{A|\bar B}(c,c)=\lambda_A\cdot ATE_{A|\bar B}(s\cup d,\cdot)$ for some (unknown) multiplier $\lambda_A\ge 0$, and define $\lambda_B$ similarly for treatment $B$. Combining these expressions with ((ref)) and ((ref)) bounds the interaction effect in terms of $\lambda_A$ and $\lambda_B$. One can then consider various hypotheses about these parameters; for example, it may be reasonable to postulate in a given application that $ATE_{A|\bar B}(c,c)$ is at most three times as large as $ATE_{A|\bar B}(s\cup d,\cdot)$, implying $\lambda_A\in [0,3]$. Or, one may plot those $(\lambda_A,\lambda_B)$ pairs for which the upper bound of LAIE is zero, etc. Boundaries of this type are common in the econometrics literature on sensitivity analysis (e.g., Masten and Poirier 2020, Martinez-Iriarte 2021). We demonstrate the construction and use of such heuristic bounds in the context of our application in Sections (ref) and (ref).
One can bound the interaction effect in a more formal way by also bounding the last two conditional means under ((ref)). For these bounds to be potentially tighter than the interval $[0,K]$, we impose further compliance type restrictions. Motivated by the discussion following Assumption (ref) in Section (ref), we assume that treatment $A$ does not admit cross defiers while treatment $B$ does not admit joint compliers. (We will argue that such a restrictions are reasonable in our application, at least as a polar case.) This leads to the following result.
In this section, we present an empirical illustration of our theory based on data from the Student Achievement and Retention Project first analysed by Angrist, Lang and Oreopoulos (2009). This program, implemented on a campus in Canada in Fall 2005, randomly assigned two treatments, namely academic services (in the form of tutoring) and financial incentives among first year college students whose high school grade point average was lower than the upper quartile. Tutoring included both access to more experienced students trained to provide academic support, as well as sessions aiming at improving study habits. The financial incentives consisted of conditional cash payments, ranging from 1,000 to 5,000 Canadian Dollars, which were paid out if a student reached a specific average grade target in college, as a function of the grade point average previously attained in high school. In our empirical application, $D_A$ is a binary variable indicating the takeup of any form of tutoring, while $D_B$ is an indicator for signing up to receive financial incentives. We are interested in the impact of these treatments on the average grade at the end of the fall semester, which is our outcome variable $Y$. The latter is measured as a credit-weighted average on a 0-100 grading scale for students taking at least one one-semester course.
The random offer of the treatments in the project was partly overlapping in the sense that some students were invited to either one of the treatments, to both, or neither. In this context, the instruments $Z_A$ and $Z_B$ correspond to binary indicators for being invited (and thus, being eligible) for tutoring and financial incentives. Thus, we are in the previously mentioned special case (covered by our framework) where the very same individual is targeted by up to two distinct treatments, rather than having a pair of individuals that might be targeted by the same or separate treatments.
Treatment takeup $D_A$ and $D_B$ may endogenously differ from the random assignment $Z_A$ and $Z_B$, respectively, because unobserved background characteristics such as personality traits likely drive both the treatment decision and academic performance. For instance, among those students offered tutoring and/or financial incentives, less motivated individuals satisfied with lower exam grades might not be willing to take the treatment(s), regardless of having received an offer or not. Due, in part, to such never takers, not all subjects comply with the random assignment, and the groups taking and not taking the treatment(s) generally differ in terms of outcome-relevant characteristics. On the other hand, among those students not offered the respective treatment, nobody managed to not comply with the assignment and take that treatment anyway. For this reason, non-compliance in our data is one-sided, as postulated in Assumption (ref).
Applying traditional IV approaches (ruling out relaxations of the treatment exclusion restriction), the findings of Angrist, Lang and Oreopoulos (2009) point to positive effects of financial incentives or combined treatments among females, but not among males. For this reason, our empirical illustration here only focuses on female students, leaving all in all 948 observations. However, for 150 females the outcome is missing, implying that these students did not take any exams in the fall semester. As the missing outcomes indicator is not statistically significantly associated with $Z_A$ or $Z_B$ (with p-values exceeding 20%), we drop those observations from the sample, leaving us with 798 females for which the outcomes (as well as the treatments and the instruments) are observed.
Table (ref) provides descriptive statistics for our evaluation sample, namely the treatment and outcome means in the total sample and in the subsamples defined by the instrument values $Z_A$ and $Z_B$. We see that the treatment frequencies observed in the data are consistent with one-sided noncompliance, since the probability of treatment $D_A$ as well as treatment $D_B$ is zero whenever $Z_A$ and $Z_B$, respectively, is zero. As a further observation, the average grade ($Y$) is highest among female students receiving both instruments and lowest among those receiving neither. The difference in average outcomes between the two groups is statistically significant at the 1% level, pointing to a non-zero reduced form effect of the joint instruments $Z_A$ and $Z_B$ on $Y$. Moreover, the average outcome is somewhat higher among students exclusively eligible for financial incentives than among those exclusively eligible for tutoring (but this difference is not statistically significant at the 10% level).
To present the effect of the instruments on the treatments (the “first stage”), Tables (ref), (ref) and (ref) report conditional probabilities of the first, second, and joint treatments, respectively, and relate them to specific compliance types. In analyzing treatment takeup patterns, we impose the restriction that there are no cross-defier types with respect to the tutor treatment arm. We consider a similar restriction---no joint compliers---with respect to the financial incentive treatment, but we are more agnostic about this condition and impose it selectively in our bounding exercise later on. As discussed in Section (ref) (and demonstrated by Corollary (ref)), such additional restrictions can greatly simplify the interpretation of IV estimands.
Table (ref) shows the takeup statistics for $D_A$. The absence of cross-defiers means that when being eligible for it, nobody is discouraged from actually taking up tutoring services by additionally being offered financial incentives. We see from Table (ref) that the nonexistence of cross-defiers is consistent with the data since our estimates suggest that $P(D_A=1|\, Z_A=1,Z_B=1)-P(D_A=1|\, Z_A=1,Z_B=0)>0$ (statistically significant at the 1% level). Yet, we emphasize that this is only a necessary, but not a sufficient condition for the absence of cross-defiers, which requires that everyone who receives tutoring when eligible for tutoring alone would also receive tutoring when additionally being eligible for financial incentives. This appears plausible if one agrees that, if anything, financial incentives for good grades should encourage (rather than discourage) the takeup of tutoring given that the latter is expected to increase academic performance. In the absence of cross-defiers, the estimated shares of self-compliers (taking tutoring if and only if eligible for it), joint compliers (taking tutoring if and only if eligible for both treatments) and never takers amount to 28%, 21%, and 51%, respectively.
Similarly, Table (ref) shows the takeup statistics for $D_B$. Not surprisingly, 93% sign up for the conditional payment when eligible for it (and nothing else). However, the estimates also suggest that $P(D_B=1|\, Z_A=1,Z_B=1) < P(D_B=1|\, Z_A=0,Z_B=1)$, the 12pp difference being statistically significant at the 5% level. For this inequality to hold, cross-defiers must be present in the population, and the prevalence of joint compliers must be limited. Cross-defiers with respect to treatment $D_B$ (i.e., the instrument $Z_A$) behave oddly in that they accept the financial incentive if this is the only treatment they are eligible for, but they refuse it if they are additionally eligible for tutoring. While it is not clear why the availability of tutoring should be a disincentive for taking the conditional payment, these types are clearly present in the data.
Some of the subsequent analysis is simpler and more informative under the restriction that there are no joint compliers with respect to the financial incentive. Table (ref) also shows that the combined share of never-takers and joint compliers is estimated to be only 7%. Still, ruling out joint compliers altogether is a rather strong assumption, since this implies that access to tutoring cannot positively affect sign-up decisions for the financial incentive given eligibility for the latter. This would be violated if some individuals judged their chances of obtaining good grades and realizing the financial rewards to be highly dependent on tutoring, so much so that they would not even sign up for the financial incentive without access to tutoring. This behavior actually seems more reasonable than that of the cross-defiers. On the other hand, $P(\cdot,d)-P(\cdot,j)=0.12$, so the higher the share of joint compliers, the higher the share of cross-defiers must be to explain the data. Given the odd behavior of the latter type, one could argue that their assumed prevalence should be as low as possible, which happens when there are no joint compliers.
There are additional moments of the data which we present in Appendix (ref). Specifically, Table (ref) shows the joint distribution of $D_A$ and $D_B$, conditional on eligibility for both treatments. These probabilities identify the shares of specific joint compliance profiles in the population, in accordance with Lemma (ref) and Corollary (ref). For example, 49% of the female students are estimated to have a $(c,c)$ profile, meaning that their takeup behavior is either in line with the randomized eligibility for tutoring and financial incentives, respectively, or they take both treatments when, and only when, they are eligible for both. In addition, Table (ref) shows the mean outcome (GPA) conditional on all configurations of the treatment dummies and their instruments. As shown by Lemma (ref) in Appendix (ref), these moments identify the mean outcome in subgroups with various compliance profiles.
Finally, in Table (ref) we provide the results of a two stage least squares regression of $Y$ (GPA) on treatments $D_A$ and $D_B$, as well as the interaction $D_AD_B$, when using $Z_A$, $Z_B$ and $Z_AZ_B$ as instruments. Theorem 3 shows that the constant term provides an estimate for the mean potential outcome $E[Y(00)]$ when receiving no treatment, suggesting that the average grade amounts to 62.83 points when female students neither take up tutoring, nor sign up for financial incentives. In the absence of cross-defiers with respect to $D_A$, the estimate of $\beta_A$ corresponds to the average effect of among self-compliers when $D_B$ is switched off. Thus, among those complying with eligibility for tutoring, receiving tutoring alone increases the average grade by 2.58 points. However, this impact is far from being statistically significant at any conventional level, as the p-value (based on heteroscedasticity-robust standard errors) is equal to 55%.
The estimate of $\beta_B$ suggests that among those who either (i) comply with their eligibility for financial incentives (self-compliers) or (ii) refuse the financial incentive when both treatments are available (cross-defiers), signing up for the financial incentive has a positive effect of 3.15 points (without tutoring). This effect is statistically significant at the 1% level. In contrast, the estimate of the interaction term $\beta_{AB}$ is small and statistically insignificant. Nevertheless, as Theorem (ref) shows, this term is not straightforward to interpret; the fact that it is close to zero does not, by itself, imply that their is no interference across the two treatments.
We now illustrate how to employ the results in Section (ref) to learn about the local average interaction effect in the context of our application. In these calculations we will ignore standard errors and treat all point estimates as if they were probability limits. We maintain Assumption (ref) throughout, but impose type restrictions only as indicated.
We start by applying Theorem (ref). Using the estimates from Tables (ref) and (ref), we obtain
Given that $E[Y(11)|(c,c)]=66.94$, this yields $ATE_{AB}(c,c)\in[-33.93,7.09]$. Clearly, $U_{00}$ could be replaced by 100, but it may also be replaced by 66.94 under the additional assumption that $Y(11)\ge Y(00)$, i.e., that taking the two treatments jointly cannot hurt anybody's GPA. This improves the bounds for the joint effect to the reasonably tight interval $[0,7.09]$ without requiring any type restrictions.
\paragraph{Heuristic analysis} As suggested in Section (ref), we can parameterize the standalone effects of $D_A$ and $D_B$ in the $(c,c)$ subgroup as $ATE_{A|\bar B}(c,c)=\lambda_A ATE_{A|\bar B}(s\cup d,\cdot)=2.58\lambda_A$ and $ATE_{B|\bar A}(c,c)=\lambda_B ATE_{B|\bar A}(\cdot, s\cup d)=3.15\lambda_B$ for some multipliers $\lambda_A, \lambda_B\ge 0$. Equation ((ref)) and the tightened bound on the joint effect then gives
Figure (ref) depicts the level sets of the upper bound in ((ref)) for $\lambda_A,\lambda_B \in [0,3]$. Any combination $(\lambda_A,\lambda_B)$ that lies above (i.e., northeast of) the zero line implies a negative LAIE, while for combinations below this line the sign of the interaction effect is unidentified. For example, if the average effect of $D_A$ and $D_B$ is about 25% percent larger among $(c,c)$ types than among $(s\cup d,\cdot)$ and $(\cdot, s\cup d)$ types, respectively, then LAIE$(c,c)$ is negative.
\paragraph{Formal analysis} Adopting the restriction that there are no $(d,\cdot)$ and $(\cdot,j)$ pairs and computing the bounds in Theorem (ref) yields $E[Y(10)|(c,c)]\in [37.28,80.14]$ and $E[Y(01)|(c,c)]\in [59.38,83.87]$. Even with the improved upper bound for $E[Y(00)|\, (c,c)]$, the implied LAIE lies in the interval $[-37.21,37.22]$, which is very wide. If we impose the additional assumption that $Y(11)\ge \max\{Y(10),Y(01)\}$, then $E[Y(11)|\, (c,c)]=66.94$ may be used as a tightened upper bound for $E[Y(10)|(c,c)]$ and $E[Y(01)|(c,c)]$ as well.\footnote{The condition $Y(11)\ge \max\{Y(10),Y(01)\}$ implies that, for any individual, the joint effect is (weakly) larger than the standalone effect of each treatment. This assumption is rather strong but it still allows for the interaction effect to be potentially negative; see equation ((ref)).} This shrinks the bound on the local interaction effect to $[-7.09,37.22]$, but the sign remains unidentified.
We impose the auxiliary condition $P(d,\cdot)=0$ (see Section (ref)), but for the time being allow for joint compliers in the financial incentive treatment arm ($P(\cdot,j)>0$). The expression for the interaction coefficient $\beta_{AB}$ stated in Theorem (ref) simplifies, and the local average interaction effect for $(c,c)$ pairs can be expressed as
where both average treatment effects under ((ref)) are identified along with their the probability weights. By contrast, the average treatment effects under ((ref)) are not identified and $P(\cdot,j)$ and $P(\cdot,d)$ are also not identified separately. For the identified quantities, we can substitute the point estimates from Tables (ref) through (ref) into ((ref)) and ((ref)) to obtain
Just as in case of the direct bounds, we can proceed in two ways.
\paragraph{Heuristic analysis} As suggested in Section (ref), we use the identified local average treatment effects $ATE_{A|\bar B}(s,\cdot)$ and $ATE_{B|\bar A}(\cdot, s\cup d)$ as reference points, and write
where the $\lambda_i$, $i=1,2,3$ are scalar multipliers. Using the identified values of $ATE_{A|\bar B}(s,\cdot)$ and $ATE_{B|\bar A}(\cdot, s\cup d)$, and substituting into ((ref)) gives
If it is hypothesized that all the $\lambda_i$ fall into, say, the interval $[0,3]$, then ((ref)) gives the following bounds on the interaction effect: \[ -2.30-19.29P(\cdot,j)\le ATE_{A|B}(c,c)-ATE_{A|\bar B}(c,c)\le 3.34+19.29P(\cdot,j). \] Clearly, the bounds are the tightest when $P(\cdot,j)=0$, i.e., there are no joint compliers with respect to the financial incentive treatment. The sign of the interaction effect is not identified even in this case, which is not too surprising given that $\beta_{AB}$ is so close to zero.
If one is willing to work with the assumption that $P(\cdot,j)=0$, expression ((ref)) reduces to a function of $\lambda_1$ and $\lambda_2$ only, and it becomes more straightforward to evaluate $LAIE(c,c)$ in various hypothetical scenarios. In particular, Figure (ref) shows the isoquants (level sets) of $LAIE(c,c)$ as a function of $\lambda_1$ and $\lambda_2$. From this graph one can read off the $(\lambda_1,\lambda_2)$ pairs that are consistent with, say, a negative interaction between the treatments.
\paragraph{Formal analysis} Alternatively, one can bound the unknown treatment effects in ((ref)) by computing the Manski-type bounds stated in Theorem (ref), while also imposing $P(\cdot,j)=0$. This yields, after replacing any uninformative bounds with trivial ones, $ATE_{A|\bar B}(j,\cdot)\in [0,43.88]$ and $ATE_{A|\bar B}(\cdot,d)\in [0,100]$. Combining these bounds with equation ((ref)) gives LAIE$(c,c)\in[-17.85,25.51]$, which is rather too loose to be useful. One can intersect this interval with the formal direct bounds for LAIE, but the sign remains unidentified. Nevertheless, these Manski-type bounds can still be informative in other applications.
We study randomized experiments (or quasi-experiments) in which the experimental units are potentially exposed to one of two different treatments, both, or none. Compliance with the intended treatment assignments, described by two binary instruments, is allowed to be endogenous. Our setup allows for the presence of compliance types that, to our knowledge, have not been considered in the literature, but are needed to accommodate some applications. In particular, there can be individuals in the population for whom the presence of $Z_B$ (or, resp.\ $Z_A$) represents a negative incentive to take treatment $D_A$ (or, resp.\ $D_B$); we call this type cross-defiers. At the same time, we allow for joint compliers as well --- a type that reacts positively to the partner instrument and ultimately takes a given treatment whenever both instruments are present.
We develop the causal interpretation of three IV estimands in our framework. The price of generality is that some of the identification results are weak in the sense that interesting causal parameters are inextricably tied up with terms arising from treatment effect heterogeneity, and auxiliary conditions are needed to obtain more useful interpretations. Alternatively, we provide partial identification results with the goal of bounding the interaction effect between the two treatments, which is frequently of interest in applications.
A clear advantage of the general approach is that one does not need to pre-commit to a theoretical framework that does not quite fit the data, and any further auxiliary conditions can be tailored to the application at hand. (For example, Blackwell (2017) needs to drop a small set of data points because they violate the treatment exclusion restriction.) Our empirical application, which analyzes a program randomly offering tutoring services (treatment $A$) and financial incentives (treatment $B$) to female college students, illustrates the advantages of starting from a general interpretative framework as well as the use of our partial identification results.