Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
84,294 characters · 14 sections · 105 citation commands
Sharp Testable Implications of Encouragement Designs
KEYWORDS: Multi-valued treatment, instrumental variable, encouragement design, random utility model, moment inequalities
JEL classification codes: C14, C31, C35, C36
\thispagestyle{empty} \setcounter{page}{1}
The analysis of potential outcome models with a discrete multi-valued treatment and a discrete multi-valued instrument has gained considerable interest recently in economics. In the setting where both the treatment and the instrument are binary, the well-known monotonicity assumption of imbens1994identification serves a key role in the identification of causal effects. Beyond this setting, the most natural extension to the monotonicity assumption for modeling choice behavior is arguably to assume that each value of the instrument (e.g., subsidy or voucher towards a program) increases the appeal of at most one unique treatment choice (e.g., the corresponding program). In fact, such an assumption on choice behavior serves as the key identifying assumption in several prominent recent empirical papers kirkeboen2016field,kline2016evaluating. Borrowing the terminology used in randomized experiments, we call such a setting an encouragement design powers1984effects,holland1988causal,duflo2007chapter. We derive sharp, closed-form testable implications, in the form of inequalities on the conditional distributions of choices and the outcome given the instrument, that characterize when the distribution of the observed data is consistent with an encouragement design. These inequalities are sharp in the sense that they exhaust all the information in the model. Because the testable implications are in closed form, if in fact the implications are violated, then we can pinpoint which substitution patterns lead to the violation. In an empirical application to behaghel2013robustness,behaghel2014private, we apply a test based on our sharp testable implications and demonstrate that the data is not compatible with assuming that the instrument does not affect the appeal of the control group, which is often thought of as a harmless normalization. Moreover, our method identifies which substitution pattern leads to the violation.
To motivate the assumptions on potential treatments that we consider, suppose there are three preschool programs (the treatment or choice), and each person receives a voucher (the instrument) towards one of them. Because we assume that the voucher towards a program only increases the appeal of the corresponding program, receiving such a voucher should not change the comparison among the remaining programs. Therefore, for each person, there exists a “default” choice (which possibly differs across people), which is the choice they would have made if the instrument didn't exist; when the instrument equals $j$, then the person chooses either treatment $j$ or the default choice. These restrictions on potential treatments immediately lead to a set of inequalities on conditional distributions given the instrument which are easy to interpret. To the best of our knowledge, these inequalities are new to the literature beyond the setting where both the treatment and the instrument are binary. We then show through a novel constructive argument that they are sharp; that is, for each distribution of the observed data that satisfies these inequalities, we construct a distribution of the potential outcomes and potential treatments that generates the observed distribution while satisfying the restrictions discussed above. We note that these restrictions immediately imply the following substitution patterns: when the instrument changes from $j$ to $k$, $k$ becomes more appealing but $j$ becomes less appealing, and hence one may stay at their original choice if it is the default choice, switch to $k$, or fall back to the default choice, which may be neither $j$ nor $k$. The sharp testable inequalities, however, involve more complex restrictions on substitution patterns across multiple values of the instrument beyond simple pairwise comparison.
To accommodate a larger class of empirical examples, we further allow researchers to impose that the values of certain choices are not affected by the instrument at all, and that there exists a “base state” of the instrument which does not affect the values of any choice. Examples of such settings include kline2016evaluating and kirkeboen2016field.\footnote{Both examples are also studied in lee2024treatment, who focus instead on the identification of treatment effect parameters conditional on “response groups” defined as sets of possible values of potential treatments.} Further examples include, for instance, feller2016compared,wu2024generalized,dahl2023high,altmejd2024inheritance,heinesen2024instrumental,humlum2025what. In kline2016evaluating, the treatment takes three values: no preschool, other preschools, Head Start; and the instrument is a binary indicator of whether the household receives an offer for Head Start. In this case, not receiving the Head Start offer is a “base state” which does not change the appeal of any choice, and the values of no preschool and other preschools are never shifted by the instrument. As a result, we recover the key identifying restriction in kline2016evaluating, that receiving an offer to Head Start may shift the household to participate in Head Start, but not lead them to switch between no preschool and other preschools. Although the restriction that the values of some choices are not affected by the instrument is reasonable in kline2016evaluating, it is otherwise often motivated as a harmless “normalization” that the value of the control status is not affected by the instrument. As we demonstrate in this paper, however, this “normalization” is not innocuous, and potentially refutable in the data. The key insight is that when there is a base state of the instrument, the default choice further coincides with the choice made under this base state. Furthermore, the substitution patterns become more restrictive: when the instrument changes from the base state to $j$, one might only switch to $j$ if their choice changes. Therefore, the testable inequalities simplify, although the construction to show sharpness requires further modification. Surprisingly, only the pairwise substitution patterns between the base state and other values of the instrument appear in the sharp testable implications.
For the setting with a binary treatment, a binary instrument, and a binary outcome, balke1997bounds,balke1997probabilistic provide the first set of inequalities that sharply characterize when the distribution of the observed data is consistent with instrument exogeneity and monotonicity. Their results are obtained through a linear programming formulation. kitagawa2015test generalizes these inequalities when the outcome is allowed to be continuous and importantly, shows they are sharp constructively. He further proposes a corresponding test. mourifie2017testing leverage the intersection bounds framework of chernozhukov2013intersection to construct an alternative test based on the same testable implications. When the treatment and the instrument are binary, the model in this paper is equivalent to the model studied in imbens1994identification. In that case, we recover the inequalities in balke1997bounds,balke1997probabilistic and kitagawa2015test. However, both the inequalities and our construction to establish their sharpness beyond this special case are, to our knowledge, novel to the literature. As shown by vytlacil2002independence, the model considered in the papers above is equivalent to the nonparametric selection model in heckman2005structural, who also discuss testable implications in the binary setting with a possibly continuous instrument. kedagni2020generalized derive a set of inequalities for instrument exogeneity with a binary outcome, a possibly multi-valued treatment and a multi-valued instrument, and further show they are sharp when both the treatment and the instrument are binary. kitagawa2021identification derives a set of sharp inequalities for instrument exogeneity with a continuous outcome, a binary treatment, and a binary instrument. sun2023instrument derives a set of inequalities with a possibly multi-valued treatment and instrument, under instrument exogeneity and the “unordered monotonicity” assumption of heckman2018unordered, but does not establish that these are necessarily sharp. As we explain in Remark (ref) below, however, the “unordered monotonicity” assumption is not implied by, nor does it imply, our assumption. kwon2024testingmechanisms characterize testable implications in the setting with a multi-valued treatment and a binary instrument\footnote{However, we note that they frame their contribution in the context of developing tests for mediation analysis.}.
Our paper is also related to a vast literature that studies identification and inference for treatment effects using instrumental variables. See, for example, bhattacharya2008treatment,bhattacharya2012treatment, machado2019instrumental, sloczynski2020should, mogstad2021causal, goff2024vector, as well as the comprehensive review article on instrumental variables by mogstad2024instrumental. Of particular relevance to our setting are the papers which study multi-valued treatments and instruments: see, for instance, lee2018identifying, kamat2023identification, bai2024inference, bai2024identifying, and bhuller20242sls.
The remainder of the paper is organized as follows. In Section (ref), we describe our setup and notation. In Section (ref), we characterize the set of inequalities implied by our model and show that these are sharp. For simplicity, we first state versions of our results in a setting with only a treatment and an instrument. In Section (ref), we extend the results to a setting with an additional, possibly continuous, outcome variable. We propose tests of the model based on these sharp implications in Section (ref) and examine the performance of these tests through simulations in Section (ref). We apply tests based on our sharp testable implications to behaghel2013robustness,behaghel2014private in Section (ref), where we demonstrate that the data is not compatible with assuming that the instrument does not affect the appeal of the control group. Moreover, our method identifies which choice patterns lead to the violation.
Let $J \geq 2$ be an integer. Let $D \in \{0, \dots, J - 1\}$ denote a multi-valued treatment choice and $Z \in \mathcal Z \subseteq \{0, \dots, J - 1 \}$ denote a multi-valued instrument. (In Section (ref), we additionally consider an outcome variable $Y \in \mathcal{Y}$, but we ignore it for the time being.) In the models that we consider, each value of the instrument encourages towards at most one unique choice. As we explain further below, to accommodate a larger class of empirical examples, we do not require that the support of $Z$ be the same as that of $D$. Specifically, we consider two forms of $\mathcal Z$: (1) $\mathcal Z = \{0, \dots, J - 1\}$ and (2) $\mathcal Z = \{0, J_0, \dots, J - 1\}$, for some $1 \leq J_0 \leq J - 1$. The first form corresponds to a setting where every choice has a corresponding instrument value which encourages towards it. The second form corresponds to a setting where the first $J_0$ choices are not affected by the instrument, and in this case we interpret $Z = 0$ as the “base state” of the instrument. To avoid notational ambiguity, we will always explicitly state the value of $J_0$\footnote{This is particularly important when $\mathcal{Z} = \{0, 1, \dots, J-1\}$, which is possible when either $J_0 = 0$ or $J_0 = 1$.}. Let $D_z$ for $z \in \mathcal Z$ denote the potential treatment choice when assigned to the instrument value $z$. As usual, the observed choice is related to potential treatment choices and the instrument through
Let $Q$ denote the distribution of $((D_z: z \in \mathcal Z), Z)$ and $P$ denote the distribution of $(D, Z)$. Note that given the mapping $T$ such that $D = T((D_z: z \in \mathcal Z), Z)$ as implied by (ref), we obtain by construction that $P = Q T^{-1}$. Throughout, we will impose the assumption that the instrument is exogenous; formally:
We will further rule out degenerate situations by requiring that the instrument takes on each value in its support with strictly positive probability:
In what follows, we will frequently use the following facts about the relationship between $P$ and $Q$. Suppose $P = Q T^{-1}$ for some $Q$ that satisfies Assumptions (ref)--(ref). Then, for $z \in \mathcal Z$, \[ P \{Z = z\} = Q \{Z = z\} > 0~.\] Therefore, the conditional choice probabilities can be defined and they satisfy
where the first equality follows because $D = D_z$ when $Z = z$ by (ref), and the second equality follows from Assumption (ref).
Our goal is to study the necessary and sufficient conditions for $P$ to be consistent with a class of restrictions on potential treatments that are commonly imposed when analyzing encouragement designs. Loosely speaking, these restrictions dictate that each value of the instrument encourages towards at most one unique choice. As we explain below, these restrictions, or stronger versions of them, are used as the key identifying restrictions for the causal interpretation of regression estimands. To state the class of restrictions, fix $0 \leq J_0 \leq J - 1$, where $J_0$ is the number of choices that are not affected by the instrument. Recall from the beginning of this section that the support of the instrument is $\mathcal{Z} = \{0, J_0, \dots, J-1\}$.
To interpret Assumption (ref), first consider the case $J_0 = 0$. Because $Z = j$ is an encouragement towards $D = j$, it should not affect the comparison among all other choices $\{0, \dots, J - 1\} \backslash \{j\}$. As a result, if we think of $j^\ast$ as a “default” choice (which is a random variable so could differ across people) that the person would have made if the instrument didn't exist, then $Z = j$ either pushes them to choose $j$ or stay at $j^\ast$; in particular, the person cannot choose any choice that is not $j$ or $j^\ast$. Furthermore, the substitution patterns must be as follows: with a change from $Z = j$ to $Z = k$, the person may stay with the original choice if it is the default choice, switch to $k$, or fall back to the default choice $j^\ast$, which may be neither $j$ nor $k$.
Note that Assumption (ref) rules out a large number of vectors of potential treatments. Consider for instance the setting where $J = 3$ and $J_0 = 0$. The restriction in (ref) implies \[ Q \{D_0 = 1, D_1 = 2\} = 0 ~.\] To see why, suppose $D_0(\omega) = 1$ and $D_1(\omega) = 2$ for some individual $\omega \in \Omega$, where $\Omega$ denotes the underlying probability space. Then, $D_0(\omega) \in \{0, j^\ast(\omega)\}$ implies $j^\ast(\omega) = 1$; at the same time, $D_1(\omega) \in \{1, j^\ast(\omega)\}$ implies $j^\ast(\omega) = 2$, a contradiction. Following similar arguments, we can conclude that under $Q$, $(D_0, D_1, D_2)$ can at most take 10 values with positive probabilities, listed in Table (ref), instead of $3^3 = 27$ values.
In some settings, one may further want to restrict the model so that the appeal of the first $J_0 > 0$ choices are not affected by the instrument. In this case, we interpret $Z = 0$ as the “base state” as if the instrument didn't exist. As a result, $j^\ast = D_0$. As illustrated through Examples (ref)--(ref) below, this additional restriction may be reasonable in some examples but not others, and is in general not a harmless normalization. When $J = 2$, $J_0 = 0$ and $J_0 = 1$ are equivalent, and in both cases, Assumption (ref) simply states $Q \{(D_0, D_1) = (1, 0)\} = 0$, i.e., defiers are ruled out. When $J > 2$, however, setting $J_0 = 1$ is no longer without loss of generality, because $J_0 = 1$ implies more restrictive substitution patterns than $J_0 = 0$. Indeed, consider $J = 3$ as an example. If $J_0 = 1$, then $Q \{(D_0, D_1, D_2) = (0, 1, 1)\} = 0$, because (ref) requires that $D_2 \in \{D_0, 2\}$ with probability one. On the other hand, if $J_0 = 0$, then (ref) allows for $Q \{(D_0, D_1, D_2) = (0, 1, 1)\} > 0$. The reason is that when $J_0 = 0$, changing $Z = 0$ to $Z = 2$ not only increases the appeal of $D = 2$, but decreases the appeal of $D = 0$ as well (because $D = 0$ is encouraged by $Z = 0$ but not $Z = 2$), making it possible for the person to fall back to the default choice, which in this case is $D = 1$. When $J_0 = 1$, however, changing $Z = 0$ to $Z = 2$ only increases the appeal of $D = 2$ but does not decrease the appeal of $D = 0$, so one can only switch to $D = 2$ instead of $D = 1$. As we show below, the testable implications in Sections (ref) and (ref) are different when $J_0 = 0$ and $J_0 > 0$, enabling us to test directly for whether the instrument does not affect the appeal of some choices.
Let $\mathbf Q_1$ denote the set of all distributions of $((D_z: z \in \mathcal Z), Z)$ that satisfy Assumptions (ref)--(ref). We now present a series of empirical examples in which Assumption (ref) or some strengthened version is used as the key identifying assumption for the causal interpretation of regression estimands. The first example will be revisited in the empirical application of Section (ref).
In this section, we present our main results on sharp testable implications of Assumptions (ref)--(ref). In order to do so, in Section (ref), we first derive inequalities in terms of the conditional choice probabilities. Then, in Section (ref), for each $P$ that satisfies these inequalities, we explicitly construct a distribution $Q \in \mathbf Q_1$ such that $P = Q T^{-1}$, thus showing the inequalities are sharp.
The following theorem characterizes a set of necessary conditions in order for $P$ to be consistent with $\mathbf Q_1$. Recall $\mathcal Z = \{0, J_0, \ldots, J-1\}$. We further define \[ \mathcal Z(j) =
\] Here, when $J_0 = 0$, it is understood that $\mathcal Z(j) = \mathcal Z \backslash \{j\}$ for $0 \leq j \leq J - 1$.
The inequalities in (ref) are direct consequences of the restriction in (ref). To see why, first suppose $J_0 = 0$, so that $z(j) \neq j$ for $0 \leq j \leq J - 1$, and consider the events \[ \{D_{z(0)} = 0\}, \dots, \{D_{z(J - 1)} = J - 1\}~. \] Fix $\omega \in \Omega$. Because $z(j) \neq j$ for all $j$, $D_{z(j)}(\omega) = j$ and (ref) imply that the default choice $j^\ast(\omega) = j$. As a result, the events listed above are disjoint across $0 \leq j \leq J - 1$, so their probabilities sum up to less than one, and (ref) follows. When $J_0 > 0$, $D_0(\omega) = j$ implies that $j^\ast(\omega) = j$, and hence $\{D_0 = j\}$ is disjoint from all other events as well. Therefore, (ref) holds in addition when $z(j) = 0$ for $0 \leq j \leq J_0 - 1$. \rule{2mm}{2mm}
The inequalities in (ref) restrict the substitution patterns jointly at different values of the instrument $z(0), \dots, z(J - 1)$. In particular, it does not suffice to consider pairwise substitution patterns when the instrument changes from one value to another. Note that $z(0), \dots, z(J - 1)$ do not have to be distinct. Therefore, for $j \neq k$, by setting $z(j) = k$ and $z(\ell) = j$ for all $\ell \neq j$, we have \[ P \{D = j | Z = k\} + \sum_{\ell \neq j} P \{D = \ell | Z = j\} \leq 1~, \] which implies
The inequality in (ref) states that the conditional probability of choosing $j$ is maximized at $Z = j$, which aligns with the intuition that $Z = j$ “encourages” towards $D = j$.
When $J_0 > 0$, we obtain the following simplification of the inequalities:
To see why the inequalities in Corollary (ref) follow from the ones in Theorem (ref), first note that for $0 \leq j \leq J - 1$ and $k \in \mathcal Z$ such that $j \neq k$, by setting $z(j) = k \in \mathcal Z(j)$ and $z(\ell) = 0 \in \mathcal Z(\ell)$ for $\ell \neq j$ in (ref), we get \[ \sum_{\ell \neq j} P \{D = \ell | Z = 0\} + P \{D = j | Z = k\} \leq 1~, \] which implies (ref) immediately. On the other hand, suppose (ref) holds for $0 \leq j \leq J - 1$ and $k \in \mathcal Z$ such that $j \neq k$. For $z(j) \in \mathcal Z(j)$ for $0 \leq j \leq J - 1$, we have $z(j) \neq j$ for $J_0 \leq j \leq J - 1$ and $z(j) \in \{0, J_0, \dots, J - 1\}$ for $0 \leq j \leq J_0 - 1$, and hence (ref) implies \[ \sum_{0 \leq j \leq J - 1} P \{D = j | Z = z(j)\} \leq \sum_{0 \leq j \leq J - 1} P \{D = j | Z = 0\} \leq 1~, \] so (ref) follows. \rule{2mm}{2mm}
The inequality in (ref) follows immediately from the restrictions in (ref). To see that, suppose $D_k = j$ for $k \neq j$. Because (ref) implies $D_k \in \{k, D_0\}$ and $k \neq j$, it has to be the case that $D_0 = j$. Therefore, \[ \{D_k = j\} \implies \{D_0 = j\}~, \] which immediately implies (ref). An interesting feature of Corollary (ref) is that only the substitution patterns between $Z = 0$ and $Z = k$ appear in the testable implications, but the substitution patterns between $Z = k$ and $Z = \ell$ for $k, \ell \neq 0$ and $k \neq \ell$ do not. Surprisingly, as we show in the next section, these inequalities exhaust all the information in the restrictions imposed by the model $\mathbf Q_1$. That is, as long as $P$ satisfies (ref), $P = Q T^{-1}$ for some $Q \in \mathbf Q_1$. As a result, when $J_0 > 0$, all information in the data about its consistency with the model is contained in the pairwise comparison between the choices when $Z = 0$ versus $Z = k$. Before proceeding, we revisit the Examples (ref)--(ref) and apply Corollary (ref).
Our next theorem is the converse of Theorem (ref)---namely, for each $P$ that satisfies (ref), there exists a distribution $Q \in \mathbf Q_1$ such that $P = Q T^{-1}$. In other words, the inequalities in Theorem (ref) are sharp in the sense that they exhaust all the information in the behavioral restrictions imposed by the model $\mathbf Q_1$. In the special case of $J = 2$, a constructive proof was provided by kitagawa2015test, but it does not extend to the case when $J > 2$. In particular, the construction in kitagawa2015test relies crucially on the observation that if $D = 1$ when $Z = 0$, then $(D_0, D_1) = (1, 1)$; this subgroup of individuals are referred to as the “always takers,” and the group that takes $D = 0$ when $Z = 1$ are called the “never takers.” The remaining probability mass is then assigned to the “compliers,” for whom $(D_0, D_1) = (0, 1)$. When $J > 2$, however, it is in general impossible to pin down the joint distribution of potential treatments using this argument, and hence the proof requires an entirely new strategy. Importantly, the proof that we present is still constructive, in that we construct $Q$ explicitly from the given distribution $P$.
To illustrate why (ref) is sufficient for determining whether $P$ is consistent with the model $\mathbf Q_1$, we start by sketching the construction when $J = 3$ and $J_0 = 0$. Let $Q^\ast$ denote a candidate distribution for which we wish to show $P = Q^\ast T^{-1}$ and $Q^\ast \in \mathbf Q_1$. We separately consider four classes of potential treatment vectors according to the value of the default choice $j^\ast$ and assign $Q^\ast$ separately for each class, such that $P = Q^\ast T^{-1}$.
$Q^\ast$ is clearly a probability measure. We now show $P = Q^\ast T^{-1}$. It suffices to verify $Q^\ast \{D_z = j\} = P \{D = j | Z = z\}$ for $j \in \{0, 1, 2\}$ and $z \in \{0, 1, 2\}$. We start by verifying that $Q^\ast \{D_z = 0\} = P \{D = 0 | Z = z\}$ for all $z$. Note $D_1 = 0$ and $D_2 = 0$ is only allowed in the events in (a) above, so that
Following similar arguments, we can show that \[ Q^\ast \{D_z = j\} = P \{D = j | Z = z\} \] for $0 \leq z \leq 2$, $0 \leq j \leq 2$, and $z \neq j$. Given these equalities, for each $k \in \{0, 1, 2\}$,
We have therefore successfully shown that $Q^\ast \{D_z = j\}$ for $0 \leq z \leq 2$ and $0 \leq j \leq 2$, and hence $P = Q^\ast T^{-1}$. \rule{2mm}{2mm}
The proof for general $J$ when $J_0 = 0$ follows similar arguments as in the previous sketch. Although we won't present the full proof in the main text, here we present some intuition on why the proof works in general. Note that although individuals cannot be classified into the three subgroups beyond the binary setting, each potential treatment vector permitted by (ref) can still be characterized by the default choice $j^\ast$ together with the set \[ \{z: D_z = z\}~. \] In other words, any potential treatment vector permitted by (ref) can be completely characterized by the default choice as well as the set of choices towards which the individual complies with the encouragement. For example, a person with $(D_0, D_1, D_2, D_3, D_4) = (0, 0, 2, 0, 4)$ can be thought of as a “0-default, $\{2, 4\}$-complier.” Similarly, we can call someone with $(D_0, D_1, D_2, D_3, D_4) = (0, 0, 0, 0, 0)$ a “0-always taker.” For each default value $0 \leq j \leq J - 1$, we order $0 \leq z \leq J - 1$ so that \[ P \{D = j | Z = z_1(j)\} \leq \dots \leq P \{D = j | Z = z_J(j)\}~. \] For this specific $j$, in step 1, we first pin down the probability of “$j$-always takers” as \[ P \{D = j | Z = z_1(j)\} = \min_{0 \leq z \leq J - 1} P \{D = j | Z = z\}~. \] Then, in step $\ell$ for $2 \leq \ell \leq J - 1$, for $\mathcal J_\ell = \{z_1(j), \dots, z_{\ell - 1}(j)\}$, we define the probability of “$j$-default, $\mathcal J_\ell$-compliers” as \[ P \{D = j | Z = z_\ell(j)\} - P \{D = j | Z = z_{\ell - 1}(j)\}~. \] Because we conclude at step $\ell = J - 1$, the total probability assigned for this specific $j$ is then \[ P \{D = j | Z = z_{J - 1}(j)\}~. \] The key reason why the construction guarantees $P = Q^\ast T^{-1}$ is as follows. If $z \neq j$, then (ref) implies $z = z_\ell(j)$ for some $1 \leq \ell \leq J - 1$. As summarized in Table (ref), $D_z = j$ happens only for “$j$-always takers,” “$j$-default, $\mathcal J_2$-compliers,” through “$j$-default, $\mathcal J_\ell$-compliers,” whose probabilities sum up to \[ P \{D = j | Z = z_\ell(j)\} = P \{D = j | Z = z\}~, \] as can be seen from Table (ref).
After carrying out this construction for each $j$, the remaining mass of $Q^\ast$ is assigned to the “diagonal” event that $(D_0, \dots, D_{J - 1}) = (0, \dots, J - 1)$, and the rest of the proof follows similarly as that for $J = 3$. The proof when $J_0 > 0$ builds on the proof when $J_0 = 0$ and requires further modifications.
In this section, we present the general results with an outcome variable. The discussion runs mostly in parallel with Section (ref). Let $Y \in \mathbf R$ denote an observed outcome and $Y_d$ for $0 \leq d \leq J - 1$ denote the potential outcome under treatment choice $d$. We allow $Y$ to be continuous or discrete and denote its support\footnote{Following pp.73--74 of lifshits1995gaussian, we define the (topological) support of $Y$ as the smallest closed set with probability one under $P$, i.e., $\mathcal Y := \bigcap \big \{ F \subseteq \mathbf R: F \text{ closed }, P \{Y \in F\} = 1 \big \}$.} by $\mathcal Y$. In addition to (ref), the observed outcome and the potential outcomes are related through \[ Y = \sum_{0 \leq d \leq J - 1} Y_d I \{D = d\}~. \] With some abuse of notation, we continue letting $T(\cdot)$ denote the mapping defined by the equation above together with (ref). Let $P$ denote the distribution of $(Y, D, Z)$ and $Q$ denote the distribution of $(Y_0, \dots, Y_{J - 1}, (D_z: z \in \mathcal Z), Z)$. We modify Assumption (ref) to include the potential outcomes:
Let $\mathbf Q_1^Y$ denote the set of all distributions $Q$ for which Assumptions (ref) and (ref) as well as (ref) hold. We first present the counterpart to Theorems (ref) and (ref).
For $z(j) \in \mathcal Z(j)$, $0 \leq j \leq J - 1$, note that by taking $B_{z(j)}(j) = \mathcal Y$ and $B_z(j) = \emptyset$ for $0 \leq j \leq J - 1$ and $z \neq z(j)$ in (ref), we recover (ref).
As in Section (ref), when $J_0 > 0$, we obtain the following simplification of the inequalities:
To see why Corollary (ref) holds, first note for $0 \leq j \leq J - 1$ and $J_0 \leq k \leq J - 1$, by taking $B_k(j) = B$, $B_0(j) = \mathcal Y \backslash B$, and $B_0(\ell) = \mathcal Y$ and $B_z(\ell) = \emptyset$ for $\ell \neq j$ and $z \neq 0$, we get \[ P \{Y \in B, D = j | Z = k\} + P \{Y \notin B, D = j | Z = 0\} + \sum_{\ell \neq j} P \{Y \in \mathcal Y, D = \ell | Z = 0\} \leq 1~, \] from which we immediately obtain (ref). On the other hand, suppose (ref) holds for all $0 \leq j \leq J - 1$, $J_0 \leq k \leq J - 1$, and $j \neq k$, and for $0 \leq j \leq J - 1$, Borel sets $\{B_z(j): z \in \mathcal Z(j)\}$ form a partition of $\mathcal Y$. Then, (ref) holds because \[ \sum_{0 \leq j \leq J - 1} \sum_{z \in \mathcal Z(j)} P \{Y \in B_z(j), D = j | Z = z\} \leq \sum_{0 \leq j \leq J - 1} \sum_{z \in \mathcal Z(j)} P \{Y \in B_z(j), D = j | Z = 0\} = 1~. \]
In this section, we extend our results to settings with one-sided noncompliance. Formally, let $\mathcal Z = \{0, \dots, J - 1\}$ and $\mathbf Q_{1, 0}$ denote the collection of all distributions of $(D_0, \dots, D_{J - 1}, Z)$ such that Assumptions (ref)--(ref) hold and
The condition in (ref) requires that when assigned $Z = j$, the subject either takes up $D = j$ or the control status $D = 0$. Therefore, noncompliance can only be one-sided ($j$ to $0$) instead of the other way around. The only difference between (ref) and (ref) is that we additionally require $D_0 = 0$. Such a setting is prevalent in economics, especially if the instrument is the “gate-keeper” or eligibility for each program, so that one either takes up the program they are eligible for or falls back to the control status. As discussed in Example (ref), assuming the “next-best” condition in kirkeboen2016field together with (ref)--(ref) and (ref)--(ref) is equivalent to (ref). See angrist2009incentives for another example. The following theorem presents the sharp testable implications of (ref) without and with an outcome. As in Section (ref), let $\mathbf Q_{1, 0}^Y$ denote set of all distributions $Q$ for which Assumptions (ref) and (ref) as well as (ref) hold.
We note that the results in Theorem (ref) follow from Corollary (ref) once we impose that $P \{D = j | Z = z\} = 0$ for $j \notin \{z, 0\}$. Indeed, with this additional restriction, (ref) is vacuous because the left-hand side is always 0. At the same time, both sides of (ref) are 0 for $j \neq 0$, so (ref) is only meaningful when $j = 0$, becoming (ref).
In this section, using the characterizations in Theorem (ref) or Corollary (ref), we construct tests to assess if the model is consistent with the distribution of the data. Formally, we test
in a way that is uniform in level across a large class of distributions. We now separately discuss tests for (ref) according to whether $\mathcal Y$ is discrete or continuous.
If $\mathcal Y$ is discrete (or is discretized ex-ante), then Theorem (ref) and Corollary (ref) generate a finite number of inequalities. In particular, (ref) becomes
for $z(j, y) \in \mathcal Z(j)$, $0 \leq j \leq J - 1$. The inequalities in (ref) become that for each $y \in \mathcal Y$,
In addition, the inequalities in (ref) become that for each $y \in \mathcal Y$,
The inequalities in (ref)--(ref) can be tested using any off-the-shelf inference method for a finite number of moment inequalities. See, for instance, pakes2017practical for an overview. Here we sketch how we can convert the problem into testing the feasibility of a linear program, so that we can directly apply recent results in, for instance, fang2023inference. Suppose $\mathcal{Y}$ is discrete and let $p$ denote the vector of $(P \{Y = y, D = j | Z = z\}: y \in \mathcal Y, 0 \leq j \leq J - 1, z \in \mathcal Z)$. We can represent the inequalities in Theorem (ref) and Corollary (ref) as
where $\Gamma$ and $\gamma$ have known entries which lie in $\{-1, 0, 1\}$. With a vector of slack variables $x$, (ref) is equivalent to
where $A$ is the identity matrix and $\beta(P) = \gamma - \Gamma p$. This formulation maps into the notation of fang2023inference, and their tests apply immediately.
If $\mathcal Y$ is continuous and we do not wish to discretize the outcome, then we focus on testing the inequalities in Corollary (ref), which apply when $J_0 > 0$; it seems difficult to test the general inequalities described in (ref) without first discretizing the outcome. Here we discuss how to test (ref) with a continuous outcome. We could develop a test based on the K-S statistic in kitagawa2015test, but following mourifie2017testing, we discuss a method that transforms the infinite number of inequalities in (ref) into a conditional moment inequality where the conditioning variable is $Y$ instead of $Z$. Indeed, note (ref) holds if and only if \[ E[I \{Y \in B\} I \{D = j, Z = k\}] P \{Z = 0\} \leq E [I\{Y \in B\} I \{D = j, Z = 0\}] P \{Z = k\}~, \] which holds for all Borel sets $B \subseteq \mathcal Y$ if and only if
with probability one for $Y$. Similarly, (ref) holds for all Borel sets $B \subseteq \mathcal Y$ if and only if
with probability one for $Y$. The conditional moment inequalities in (ref)--(ref) could then be tested using any off-the-shelf inference method for conditional moment inequalities. See, for instance, andrews2013inference, chernozhukov2013intersection, armstrong2016multiscale, and chetverikov2018adaptive.
In this section, we study the properties of the inference procedures described in Section (ref). Our goal is to illustrate the size control and power properties of our tests for (ref), and to compare them with other procedures that are based on an implicit characterization of the inequalities, as discussed in Remark (ref). We present the results separately for $J_0 > 0$ and $J_0 = 0$.
In this subsection, we study tests of the null hypothesis in (ref) for $J = 4$ and $J_0 = 1$, so that $\mathcal D = \mathcal Z = \{0, 1, 2, 3\}$. Throughout this subsection, $Z$ is uniformly distributed on $\mathcal Z$. In the simulation design, the potential treatments are generated by an additive random utility model, so that $D_z(\omega) \in \operatorname*{argmax}_{0 \leq j \leq J - 1} (\beta_j I\{j = z\} + \epsilon_j)$, where $(\epsilon_0, \dots, \epsilon_{J - 1})$ and $Z$ are independent. Let $(\beta_1, \beta_2, \beta_3) = (1.5, 1, 0.5)$ and $\beta_0$ be specified below. Further let $\epsilon = (\epsilon_0, \epsilon_1, \epsilon_2, \epsilon_3)' \sim N(\mu, I_4)$, where $\mu = (0.5, 1, 1.5, 2)'$ and $I_4$ is the $4 \times 4$ identity matrix. Note that when $\beta_0 = 0$, the distribution $P$ satisfies the null in (ref) with $J_0 = 1$; furthermore, for each $z \in \mathcal Z$, $E[\max_{j \in \mathcal D} (U_j(z) + \epsilon_j)] = E[U_z(z) + \epsilon_z] = 2.5$, so that the mean utility of the choice encouraged by $z$, $d=z$, stays constant across different values of the instrument $z$, but the mean utility of the alternative choice $d\neq z$ varies across $z$. The potential outcomes are determined as $Y_d = I \{d \geq 1\} + \xi$, where $\xi \perp \!\!\! \perp (\epsilon, Z)$ and $\xi \sim N(0, 1)$. Next, to assess the power of the tests, we further consider $\beta_0 \in \{-0.5, -1, -1.5\}$, so that it can be verified through direct calculation that (ref) is violated and hence the null in (ref) is violated.
We implement several tests of (ref) with $J_0 = 1$ at the 5% level. First, similarly as in mourifie2017testing, we test (ref)--(ref) using chernozhukov2013intersection and the accompanying clrtest package in stata chernozhukov2015implementing. Table (ref) presents the rejection probabilities in percentages. We implement both the parametric and local options for the estimation of the conditional moments and follow the choices of tuning parameters in mourifie2017testing. Second, we binarize $Y$ to $I \{Y \geq 1\}$ and test (ref)--(ref) using fang2023inference and the accompanying lpinfer package in R conroylau_2021_5506545. The results are displayed in Table (ref). Finally, with the discretized $Y$ and using fang2023inference, we directly test the implicit linear programming formulation in Remark (ref), i.e., whether there exists a probability measure $Q$ that satisfies (ref)\footnote{With an outcome $Y$, (ref) strengthens to $P\{Y=y, D=j | Z=z\} = Q\{Y_j=y, D_z=j\}$, which is the restriction we consider whenever there is an outcome.} and (ref). The results are displayed in Table (ref). We perform each test at sample sizes of 500, 1000, and 2000. For each value of $\beta_0$, and each sample size, we calculate the rejection probabilities across 5000 replications.
We begin by noting in Table (ref) that the results for testing (ref)--(ref) using chernozhukov2013intersection are very sensitive to the methods for estimating the conditional moments. With the local approach, the test fails to control size even at $n = 2000$. The seemingly higher power is likely a result of its poor size control. The parametric approach controls size well and has nontrivial power against $\beta_0 \in \{-1, -1.5\}$. The drastic difference in the size control of the two approaches may be because the conditional moments in (ref)--(ref) can be reasonably approximated by linear functions in the current data generating process, as can be seen from plotting the conditional moments. The over-rejection is also observed in the appendix to mourifie2017testing when $J = 2$, in which case our model is equivalent to the one in imbens1994identification. Consequently, we test the same inequalities presented in kitagawa2015test as mourifie2017testing do, and we further replicate this behavior in Appendix (ref) for a range of designs with $J = 2$. In any case, chernozhukov2013intersection require choosing a number of tuning parameters. Next, in Tables (ref) and (ref), after discretizing $Y$, both testing the moment inequalities in (ref)--(ref) and testing the linear program defined by (ref) and (ref) using fang2023inference control size well and have nontrivial power for $\beta = -1.5$ at all sample sizes, for $\beta = -1$ when $n = 1000$ and $n = 2000$, and also for $\beta_0 = -0.5$ when $n = 2000$. Between the two methods, testing the closed-form inequalities in (ref)--(ref) is more powerful than testing the linear program defined by (ref) and (ref). Finally, comparing Table (ref)(a) and Table (ref), although (ref)--(ref) based on discretizing $Y$ do not sharply characterize the model compared to the conditional moment inequalities in (ref)--(ref), the test based on the former discretization is still more powerful.
Next, we study tests of the hypothesis in (ref) for $J_0 = 0$. The data is generated in the same way as in Section (ref), except now that $\beta_0 = 2$ under the null hypothesis. Throughout this subsection, we also binarize $Y$ as in Section (ref). In this case, testing (ref)--(ref) using fang2023inference is computationally prohibitive when $J > 4$, and testing the linear program defined by (ref) and (ref) using fang2023inference is also computationally prohibitive when $J > 5$. As a result, we compare the two methods when $J = 3$, and the design therefore uses only the first three entries of $\beta$ and $\epsilon$. Here, (ref)--(ref) generate 76 inequalities with 18 variables, while (ref) and (ref) define a linear system with 18 equalities with 80 latent variables. The results are presented in Tables (ref) and (ref). As in Section (ref), both tests control size well, and the test based on the closed-form characterization in (ref)--(ref) is more powerful.
Because the datasets in kline2016evaluating and kirkeboen2016field are confidential, we apply our methods to the dataset from behaghel2013robustness,behaghel2014private. They study a randomized controlled trial with three job search counseling programs in France. In their setting, $Z = 0$ denotes assignment (of eligibility) to the usual public program without intensive counseling, which is thought of as the control group; $Z = 1$ denotes assignment to the public program $Z$ with intensive counseling; and $Z = 2$ denotes assignment to the private program with intensive counseling. $D = 0, 1, 2$ denotes participation in the corresponding programs. Assumption (ref) holds because $Z$ is randomly assigned. Job seekers did not necessarily comply with their assignment, i.e., they may not enter the program they were assigned to, and the noncompliance rate was as high as 60% on average behaghel2014private. As can be seen in Table (ref) below, some job seekers assigned to the control group participated in the two treatment programs, and other job seekers assigned to either of the two treatment programs entered the control or the other treatment program. As a result, the assignment $Z = j$ becomes an encouragement towards $D = j$ which one may or may not take up.
We consider three outcomes in their dataset: “EMPLOI 6MOIS,” which indicates exit from PES registers to employment; “EMPLOI AR110 6MOIS,” which indicates any employment; and “SUCCES OPP 6MOIS,” which indicates employment eligible for payment. All three outcomes are binary and were measured six months after the beginning of the program following their analysis. To recover a causal interpretation of the IV estimand, behaghel2013robustness impose an assumption called “extended monotonicity.” As shown in Appendix (ref), this assumption is strictly stronger than setting $J_0 = 1$ in our model. As a result, we focus on testing (ref) for $J_0 = 0$ and $J_0 = 1$. The test for $J_0 = 0$ indicates whether the assumption that each value of the instrument encourages towards an unique choice is consistent with the data. The test for $J_0 = 1$ further indicates whether it is truly without loss of generality to assume that assignment to the control group does not in fact encourage people to take up the control program. This may be a tempting assumption if the assignment to the control program is thought of as the “base state” of the instrument. Because all outcomes are discrete, we test (ref)--(ref) for $J_0 = 0$ and (ref)--(ref) for $J_0 = 1$ using fang2023inference. To illustrate the gains from the refinement by including an outcome, we also include the result for testing (ref) for $J_0 = 0$ and (ref) for $J_0 = 1$ which do not make use of the outcome. Table (ref) presents the $p$-values of the test in percentages for each null hypothesis and outcome combination.
Without the outcome, the tests for $J_0 = 0$ and $J_0 = 1$ both fail to reject at the 10% level. With any of the three outcomes, the test for $J_0 = 0$ fails to reject at the 10% level, but the test for $J_0 = 1$ rejects at the 10% level. For the outcome “EMPLOI AR110 6MOIS,” the test for $J_0 = 1$ further rejects at both the 5% and 1% levels. Therefore, including the outcome in the tests helps us reject the null hypothesis that $J_0 = 1$. To further investigate which moment inequalities among (ref)--(ref) are violated, for each outcome, we report the point estimates for the violated moments, i.e., moments that are estimated to violate (ref)--(ref). The results are presented in Table (ref). Among all moments, the one that is consistently violated across outcomes is $P \{Y = y, D = 1 | Z = 2\} \leq P \{Y = y, D = 1 | Z = 0\}$. This moment is also violated without an outcome, but the violation is more pronounced with an outcome. Together with the results in Table (ref), they indeed illustrate the gains from the refinement by including an outcome in the test. To understand the violated moments, note that in Assumption (ref), $D_2 = 1$ and $D_2 \in \{D_0, 2\}$ implies $D_0 = 1$. The violation of this condition implies that one cannot think of the assignment to the control program as the “base state.” In other words, at least for some people, $Z = 0$ strictly increases the appeal of the public program $D = 0$. Our test clearly pinpoints the violation is through the substitution pattern that $D_2 = 1$ and $D_0 \neq 1$, i.e., some subjects choose program 1 when encouraged towards program 2 but do not choose program 1 when there is no encouragement. As discussed in Section (ref), when $J_0 = 0$, changing $Z = 0$ to $Z = 2$ not only increases the appeal of $D = 2$, but decreases the appeal of $D = 0$ as well, making it possible for the person to switch to $D = 1$. When $J_0 = 1$, however, changing $Z = 0$ to $Z = 2$ only increases the appeal of $D = 2$ but does not decrease the appeal of $D = 0$, so one can only switch to $D = 2$ instead of $D = 1$.
In this paper, we proposed sharp, closed-form testable implications of a potential outcome model that assumes each value of the instrument only encourages towards one choice. Because the testable implications are in closed form, we can immediately check if they are violated for a given dataset and more importantly, pinpoint where the violation occurs. In an empirical application to the dataset for behaghel2013robustness,behaghel2014private, we find that the data is not compatible with the assumption that the control program is not encouraged by its corresponding assignment, which may be a tempting assumption if the control program is thought of as the “base state.” The identity of the violated inequalities further indicates which choice patterns are incompatible with the data.
Our assumption that each value of the instrument only encourages towards one choice is satisfied in many examples, sometimes with the additional restriction that some choices do not have an encouragement. That said, it would be interesting to study settings in which each value of the instrument possibly encourages towards multiple choices. We leave this question open for future work.