EconBase
← Back to paper

Sharp Testable Implications of Encouragement Designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

84,294 characters · 14 sections · 105 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sharp Testable Implications of Encouragement Designs

abstractThis paper studies a potential outcome model with a continuous or discrete outcome, a discrete multi-valued treatment, and a discrete multi-valued instrument. We derive sharp, closed-form testable implications for a class of restrictions on potential treatments where each value of the instrument encourages towards at most one unique treatment choice; such restrictions serve as the key identifying assumption in several prominent recent empirical papers. Borrowing the terminology used in randomized experiments, we call such a setting an encouragement design. The testable implications are inequalities in terms of the conditional distributions of choices and the outcome given the instrument. Through a novel constructive argument, we show these inequalities are sharp in the sense that any distribution of the observed data that satisfies these inequalities is compatible with this class of restrictions on potential treatments. Based on these inequalities, we propose tests of the restrictions. In an empirical application, we show some of these restrictions are violated and pinpoint the substitution pattern that leads to the violation.

KEYWORDS: Multi-valued treatment, instrumental variable, encouragement design, random utility model, moment inequalities

JEL classification codes: C14, C31, C35, C36

\thispagestyle{empty} \setcounter{page}{1}

Introduction

The analysis of potential outcome models with a discrete multi-valued treatment and a discrete multi-valued instrument has gained considerable interest recently in economics. In the setting where both the treatment and the instrument are binary, the well-known monotonicity assumption of imbens1994identification serves a key role in the identification of causal effects. Beyond this setting, the most natural extension to the monotonicity assumption for modeling choice behavior is arguably to assume that each value of the instrument (e.g., subsidy or voucher towards a program) increases the appeal of at most one unique treatment choice (e.g., the corresponding program). In fact, such an assumption on choice behavior serves as the key identifying assumption in several prominent recent empirical papers kirkeboen2016field,kline2016evaluating. Borrowing the terminology used in randomized experiments, we call such a setting an encouragement design powers1984effects,holland1988causal,duflo2007chapter. We derive sharp, closed-form testable implications, in the form of inequalities on the conditional distributions of choices and the outcome given the instrument, that characterize when the distribution of the observed data is consistent with an encouragement design. These inequalities are sharp in the sense that they exhaust all the information in the model. Because the testable implications are in closed form, if in fact the implications are violated, then we can pinpoint which substitution patterns lead to the violation. In an empirical application to behaghel2013robustness,behaghel2014private, we apply a test based on our sharp testable implications and demonstrate that the data is not compatible with assuming that the instrument does not affect the appeal of the control group, which is often thought of as a harmless normalization. Moreover, our method identifies which substitution pattern leads to the violation.

To motivate the assumptions on potential treatments that we consider, suppose there are three preschool programs (the treatment or choice), and each person receives a voucher (the instrument) towards one of them. Because we assume that the voucher towards a program only increases the appeal of the corresponding program, receiving such a voucher should not change the comparison among the remaining programs. Therefore, for each person, there exists a “default” choice (which possibly differs across people), which is the choice they would have made if the instrument didn't exist; when the instrument equals $j$, then the person chooses either treatment $j$ or the default choice. These restrictions on potential treatments immediately lead to a set of inequalities on conditional distributions given the instrument which are easy to interpret. To the best of our knowledge, these inequalities are new to the literature beyond the setting where both the treatment and the instrument are binary. We then show through a novel constructive argument that they are sharp; that is, for each distribution of the observed data that satisfies these inequalities, we construct a distribution of the potential outcomes and potential treatments that generates the observed distribution while satisfying the restrictions discussed above. We note that these restrictions immediately imply the following substitution patterns: when the instrument changes from $j$ to $k$, $k$ becomes more appealing but $j$ becomes less appealing, and hence one may stay at their original choice if it is the default choice, switch to $k$, or fall back to the default choice, which may be neither $j$ nor $k$. The sharp testable inequalities, however, involve more complex restrictions on substitution patterns across multiple values of the instrument beyond simple pairwise comparison.

To accommodate a larger class of empirical examples, we further allow researchers to impose that the values of certain choices are not affected by the instrument at all, and that there exists a “base state” of the instrument which does not affect the values of any choice. Examples of such settings include kline2016evaluating and kirkeboen2016field.\footnote{Both examples are also studied in lee2024treatment, who focus instead on the identification of treatment effect parameters conditional on “response groups” defined as sets of possible values of potential treatments.} Further examples include, for instance, feller2016compared,wu2024generalized,dahl2023high,altmejd2024inheritance,heinesen2024instrumental,humlum2025what. In kline2016evaluating, the treatment takes three values: no preschool, other preschools, Head Start; and the instrument is a binary indicator of whether the household receives an offer for Head Start. In this case, not receiving the Head Start offer is a “base state” which does not change the appeal of any choice, and the values of no preschool and other preschools are never shifted by the instrument. As a result, we recover the key identifying restriction in kline2016evaluating, that receiving an offer to Head Start may shift the household to participate in Head Start, but not lead them to switch between no preschool and other preschools. Although the restriction that the values of some choices are not affected by the instrument is reasonable in kline2016evaluating, it is otherwise often motivated as a harmless “normalization” that the value of the control status is not affected by the instrument. As we demonstrate in this paper, however, this “normalization” is not innocuous, and potentially refutable in the data. The key insight is that when there is a base state of the instrument, the default choice further coincides with the choice made under this base state. Furthermore, the substitution patterns become more restrictive: when the instrument changes from the base state to $j$, one might only switch to $j$ if their choice changes. Therefore, the testable inequalities simplify, although the construction to show sharpness requires further modification. Surprisingly, only the pairwise substitution patterns between the base state and other values of the instrument appear in the sharp testable implications.

For the setting with a binary treatment, a binary instrument, and a binary outcome, balke1997bounds,balke1997probabilistic provide the first set of inequalities that sharply characterize when the distribution of the observed data is consistent with instrument exogeneity and monotonicity. Their results are obtained through a linear programming formulation. kitagawa2015test generalizes these inequalities when the outcome is allowed to be continuous and importantly, shows they are sharp constructively. He further proposes a corresponding test. mourifie2017testing leverage the intersection bounds framework of chernozhukov2013intersection to construct an alternative test based on the same testable implications. When the treatment and the instrument are binary, the model in this paper is equivalent to the model studied in imbens1994identification. In that case, we recover the inequalities in balke1997bounds,balke1997probabilistic and kitagawa2015test. However, both the inequalities and our construction to establish their sharpness beyond this special case are, to our knowledge, novel to the literature. As shown by vytlacil2002independence, the model considered in the papers above is equivalent to the nonparametric selection model in heckman2005structural, who also discuss testable implications in the binary setting with a possibly continuous instrument. kedagni2020generalized derive a set of inequalities for instrument exogeneity with a binary outcome, a possibly multi-valued treatment and a multi-valued instrument, and further show they are sharp when both the treatment and the instrument are binary. kitagawa2021identification derives a set of sharp inequalities for instrument exogeneity with a continuous outcome, a binary treatment, and a binary instrument. sun2023instrument derives a set of inequalities with a possibly multi-valued treatment and instrument, under instrument exogeneity and the “unordered monotonicity” assumption of heckman2018unordered, but does not establish that these are necessarily sharp. As we explain in Remark (ref) below, however, the “unordered monotonicity” assumption is not implied by, nor does it imply, our assumption. kwon2024testingmechanisms characterize testable implications in the setting with a multi-valued treatment and a binary instrument\footnote{However, we note that they frame their contribution in the context of developing tests for mediation analysis.}.

Our paper is also related to a vast literature that studies identification and inference for treatment effects using instrumental variables. See, for example, bhattacharya2008treatment,bhattacharya2012treatment, machado2019instrumental, sloczynski2020should, mogstad2021causal, goff2024vector, as well as the comprehensive review article on instrumental variables by mogstad2024instrumental. Of particular relevance to our setting are the papers which study multi-valued treatments and instruments: see, for instance, lee2018identifying, kamat2023identification, bai2024inference, bai2024identifying, and bhuller20242sls.

The remainder of the paper is organized as follows. In Section (ref), we describe our setup and notation. In Section (ref), we characterize the set of inequalities implied by our model and show that these are sharp. For simplicity, we first state versions of our results in a setting with only a treatment and an instrument. In Section (ref), we extend the results to a setting with an additional, possibly continuous, outcome variable. We propose tests of the model based on these sharp implications in Section (ref) and examine the performance of these tests through simulations in Section (ref). We apply tests based on our sharp testable implications to behaghel2013robustness,behaghel2014private in Section (ref), where we demonstrate that the data is not compatible with assuming that the instrument does not affect the appeal of the control group. Moreover, our method identifies which choice patterns lead to the violation.

Setup and Notation

Let $J \geq 2$ be an integer. Let $D \in \{0, \dots, J - 1\}$ denote a multi-valued treatment choice and $Z \in \mathcal Z \subseteq \{0, \dots, J - 1 \}$ denote a multi-valued instrument. (In Section (ref), we additionally consider an outcome variable $Y \in \mathcal{Y}$, but we ignore it for the time being.) In the models that we consider, each value of the instrument encourages towards at most one unique choice. As we explain further below, to accommodate a larger class of empirical examples, we do not require that the support of $Z$ be the same as that of $D$. Specifically, we consider two forms of $\mathcal Z$: (1) $\mathcal Z = \{0, \dots, J - 1\}$ and (2) $\mathcal Z = \{0, J_0, \dots, J - 1\}$, for some $1 \leq J_0 \leq J - 1$. The first form corresponds to a setting where every choice has a corresponding instrument value which encourages towards it. The second form corresponds to a setting where the first $J_0$ choices are not affected by the instrument, and in this case we interpret $Z = 0$ as the “base state” of the instrument. To avoid notational ambiguity, we will always explicitly state the value of $J_0$\footnote{This is particularly important when $\mathcal{Z} = \{0, 1, \dots, J-1\}$, which is possible when either $J_0 = 0$ or $J_0 = 1$.}. Let $D_z$ for $z \in \mathcal Z$ denote the potential treatment choice when assigned to the instrument value $z$. As usual, the observed choice is related to potential treatment choices and the instrument through

equation[equation omitted — 76 chars of source]

Let $Q$ denote the distribution of $((D_z: z \in \mathcal Z), Z)$ and $P$ denote the distribution of $(D, Z)$. Note that given the mapping $T$ such that $D = T((D_z: z \in \mathcal Z), Z)$ as implied by (ref), we obtain by construction that $P = Q T^{-1}$. Throughout, we will impose the assumption that the instrument is exogenous; formally:

assumption$(D_z: z \in \mathcal Z) \perp \!\!\! \perp Z$ under $Q$.

We will further rule out degenerate situations by requiring that the instrument takes on each value in its support with strictly positive probability:

assumption$Q \{Z = z\} > 0$ for $z \in \mathcal Z$.

In what follows, we will frequently use the following facts about the relationship between $P$ and $Q$. Suppose $P = Q T^{-1}$ for some $Q$ that satisfies Assumptions (ref)--(ref). Then, for $z \in \mathcal Z$, \[ P \{Z = z\} = Q \{Z = z\} > 0~.\] Therefore, the conditional choice probabilities can be defined and they satisfy

equation[equation omitted — 92 chars of source]

where the first equality follows because $D = D_z$ when $Z = z$ by (ref), and the second equality follows from Assumption (ref).

Our goal is to study the necessary and sufficient conditions for $P$ to be consistent with a class of restrictions on potential treatments that are commonly imposed when analyzing encouragement designs. Loosely speaking, these restrictions dictate that each value of the instrument encourages towards at most one unique choice. As we explain below, these restrictions, or stronger versions of them, are used as the key identifying restrictions for the causal interpretation of regression estimands. To state the class of restrictions, fix $0 \leq J_0 \leq J - 1$, where $J_0$ is the number of choices that are not affected by the instrument. Recall from the beginning of this section that the support of the instrument is $\mathcal{Z} = \{0, J_0, \dots, J-1\}$.

assumptionThere exists a random variable $j^\ast$ such that $0 \leq j^\ast \leq J-1$ and \begin{equation} Q \{D_j \in \{j, j^\ast\}\} = 1 for 0 \leq j \leq J - 1 . \end{equation} Furthermore, when $J_0 > 0$, $Q\{j^\ast = D_0\} = 1$, so that (ref) becomes \begin{equation} Q \{D_j \in \{j, D_0\} \} = 1 for J_0 \leq j \leq J-1 . \end{equation}

To interpret Assumption (ref), first consider the case $J_0 = 0$. Because $Z = j$ is an encouragement towards $D = j$, it should not affect the comparison among all other choices $\{0, \dots, J - 1\} \backslash \{j\}$. As a result, if we think of $j^\ast$ as a “default” choice (which is a random variable so could differ across people) that the person would have made if the instrument didn't exist, then $Z = j$ either pushes them to choose $j$ or stay at $j^\ast$; in particular, the person cannot choose any choice that is not $j$ or $j^\ast$. Furthermore, the substitution patterns must be as follows: with a change from $Z = j$ to $Z = k$, the person may stay with the original choice if it is the default choice, switch to $k$, or fall back to the default choice $j^\ast$, which may be neither $j$ nor $k$.

Note that Assumption (ref) rules out a large number of vectors of potential treatments. Consider for instance the setting where $J = 3$ and $J_0 = 0$. The restriction in (ref) implies \[ Q \{D_0 = 1, D_1 = 2\} = 0 ~.\] To see why, suppose $D_0(\omega) = 1$ and $D_1(\omega) = 2$ for some individual $\omega \in \Omega$, where $\Omega$ denotes the underlying probability space. Then, $D_0(\omega) \in \{0, j^\ast(\omega)\}$ implies $j^\ast(\omega) = 1$; at the same time, $D_1(\omega) \in \{1, j^\ast(\omega)\}$ implies $j^\ast(\omega) = 2$, a contradiction. Following similar arguments, we can conclude that under $Q$, $(D_0, D_1, D_2)$ can at most take 10 values with positive probabilities, listed in Table (ref), instead of $3^3 = 27$ values.

table[table omitted — 492 chars of source]

In some settings, one may further want to restrict the model so that the appeal of the first $J_0 > 0$ choices are not affected by the instrument. In this case, we interpret $Z = 0$ as the “base state” as if the instrument didn't exist. As a result, $j^\ast = D_0$. As illustrated through Examples (ref)--(ref) below, this additional restriction may be reasonable in some examples but not others, and is in general not a harmless normalization. When $J = 2$, $J_0 = 0$ and $J_0 = 1$ are equivalent, and in both cases, Assumption (ref) simply states $Q \{(D_0, D_1) = (1, 0)\} = 0$, i.e., defiers are ruled out. When $J > 2$, however, setting $J_0 = 1$ is no longer without loss of generality, because $J_0 = 1$ implies more restrictive substitution patterns than $J_0 = 0$. Indeed, consider $J = 3$ as an example. If $J_0 = 1$, then $Q \{(D_0, D_1, D_2) = (0, 1, 1)\} = 0$, because (ref) requires that $D_2 \in \{D_0, 2\}$ with probability one. On the other hand, if $J_0 = 0$, then (ref) allows for $Q \{(D_0, D_1, D_2) = (0, 1, 1)\} > 0$. The reason is that when $J_0 = 0$, changing $Z = 0$ to $Z = 2$ not only increases the appeal of $D = 2$, but decreases the appeal of $D = 0$ as well (because $D = 0$ is encouraged by $Z = 0$ but not $Z = 2$), making it possible for the person to fall back to the default choice, which in this case is $D = 1$. When $J_0 = 1$, however, changing $Z = 0$ to $Z = 2$ only increases the appeal of $D = 2$ but does not decrease the appeal of $D = 0$, so one can only switch to $D = 2$ instead of $D = 1$. As we show below, the testable implications in Sections (ref) and (ref) are different when $J_0 = 0$ and $J_0 > 0$, enabling us to test directly for whether the instrument does not affect the appeal of some choices.

Let $\mathbf Q_1$ denote the set of all distributions of $((D_z: z \in \mathcal Z), Z)$ that satisfy Assumptions (ref)--(ref). We now present a series of empirical examples in which Assumption (ref) or some strengthened version is used as the key identifying assumption for the causal interpretation of regression estimands. The first example will be revisited in the empirical application of Section (ref).

examplebehaghel2013robustness,behaghel2014private study a randomized controlled trial involving three job search counseling programs in France. In their setting, $Z = 0$ denotes encouragement towards the usual public program without intensive counseling, $Z = 1$ denotes encouragement towards the public program with intensive counseling, and $Z = 2$ denotes encouragement towards the private program. $D = 0, 1, 2$ denotes participation in the corresponding programs. Assumption (ref) holds because $Z$ is randomly assigned and Assumption (ref) holds as long as a nontrivial portion of people are assigned to each arm. To recover a causal interpretation of the IV estimand as a local average treatment effect (LATE), however, behaghel2013robustness impose an assumption called “extended monotonicity.” As shown in Appendix (ref), this assumption is strictly stronger than setting $J_0 = 1$ in our model. However, because each value of $Z$ encourages towards one value of $D$, it is reasonable to expect we are in a situation where $J_0 = 0$ instead of $J_0 = 1$; in particular, there is no compelling reason to handle $Z = 0$ asymmetrically with $Z \in \{1, 2\}$ purely because it is the encouragement towards a control group. The discussion following Assumption (ref) demonstrated that $\{D_0 \neq 1, D_2 = 1\}$ is not allowed when $J_0 = 1$, but is allowed when $J_0 = 0$. In the empirical application in Section (ref), we reject $J_0 = 1$, and thus also their extended monotonicity assumption. Moreover, we show that the substitution pattern $\{D_0 \neq 1, D_2 = 1\}$ is exactly the reason for rejection. At the same time, we cannot reject $J_0 = 0$. The findings therefore suggest that at least for some people, $Z = 0$ strictly increases the appeal of the public program $D = 0$.
examplekline2016evaluating consider an RCT with a “close substitute” to study the effects of preschooling on educational outcomes and impose (ref) with $J = 3$ and $J_0 = 2$. In their setting, $D \in \{0,1,2\}$, where $D=0$ denotes home care (no preschool), $D = 2$ denotes participation in a preschool program called Head Start, and $D = 1$ denotes participation in preschools other than Head Start, namely the close substitute. Here, $Z \in \{0, 2\}$, where $Z = 2$ denotes that the household receives an offer to attend Head Start, and $Z = 0$ denoted otherwise. Assumption (ref) holds because $Z$ is randomly assigned and Assumption (ref) holds as long as a nontrivial portion of households are assigned to each arm. In their equation (1), kline2016evaluating impose the restriction that \begin{equation} Q \{ D_2 = 2 | D_0 \neq D_2\}=1 . \end{equation} The restriction in (ref) states that if a household switches their choice upon receiving a Head Start offer, then they must be switching to Head Start. In other words, receiving an offer to Head Start does not change the comparison between no preschool and preschools other than Head Start. The restriction in (ref) is equivalent to $Q \{D_2 \in \{D_0, 2\}\} = 1$, which is exactly (ref) with $J = 3$ and $J_0 = 2$. The support of $(D_0, D_2)$ is summarized in Table (ref). The restriction in (ref) is plausible, although it may be violated because of, for instance, a salience effect, where receiving the offer to Head Start makes the household more aware of the merit of preschools in general and causes them to send the kid to other preschools. In that case, a change of $Z = 0$ to $Z = 2$ may also decrease the appeal of home care or increase the appeal of other preschools. \begin{table}[ht!] \begin{tabular}{cc} \toprule $D_0$ & $D_2$ \\ \cmidrule(lr){1-2} 0 & 0 \\ 1 & 1 \\ 2 & 2 \\ 0 & 2 \\ 1 & 2 \\ \bottomrule \end{tabular} \caption{The support of $(D_0, D_2)$ in kline2016evaluating.} \end{table} The key insight behind the empirical results in kline2016evaluating is that under the assumption in (ref), the Wald estimand from the IV regression identifies a weighted combination of what they call “sub-LATEs.” kline2016evaluating further establish conditions under which optimal policy depends upon these “sub-LATEs.” Our results in Sections (ref) and (ref) will characterize when the distribution of the data is consistent with (ref).
examplekirkeboen2016field study the effects of fields of study on earnings and impose (ref) with $J = 3$ and $J_0 = 1$. In their setting, $D \in \{0, 1, 2\}$ represent three fields of study, ordered by their (soft) admission cutoffs from the lowest to the highest. The instrument $Z \in \{0, 1, 2\}$ is called “expected offer” in their paper. Roughly speaking, $Z = 1$ when the student crosses the (soft) admission cutoff for field 1, $Z = 2$ when the student crosses the (soft) admission cutoff for field 2, and $Z = 0$ otherwise. The authors assume that $Z$ is exogenous in the sense that $Q$ satisfies Assumption (ref), and also assume Assumption (ref) holds. They further impose the following monotonicity conditions: \begin{align} Q \{D_1 = 1 | D_0 = 1\} & = 1 , \\ Q \{D_2 = 2 | D_0 = 2\} & = 1 . \end{align} The conditions in (ref)--(ref) require that crossing the cutoff for field 1 or 2 weakly encourages them towards that field. They further impose the following “irrelevance” conditions: \begin{align} Q \{I \{D_1 = 2\} & = I \{D_0 = 2\} | D_0 \neq 1, D_1 \neq 1\} = 1 , \\ Q \{I \{D_2 = 1\} & = I \{D_0 = 1\} | D_0 \neq 2, D_2 \neq 2\} = 1 . \end{align} The condition in (ref) states that if crossing the cutoff for field 1 does not cause the student to switch to field 1, then it does not cause them to switch to or away from field 2. A similar interpretation applies to (ref). The restrictions in (ref)--(ref) are equivalent to (ref) with $J = 3$ and $J_0 = 1$. We now show they imply (ref) and the other direction is straightforward. Suppose by contradiction that $D_1 \notin \{D_0, 1\}$. Under this assumption, if $D_1 = 2$, then (ref) implies that $D_0 \neq 1$; (ref) therefore implies that $D_0 = 2$, so $D_1 = D_0$, a contradiction to $D_1 \notin \{D_0, 1\}$. If instead $D_1 = 0$, then (ref) again implies that $D_0 \neq 1$; because $D_1 \notin \{D_0, 1\}$, we know $D_0 = 2$, which by (ref) implies $D_1 = 2$, a contradiction to $D_1 = 0$. Therefore, (ref) is satisfied for $j = 1$. Similar arguments show it is satisfied for $j = 2$ as well. The support of $(D_0, D_1, D_2)$ is summarized in Table (ref). The restrictions in (ref)--(ref) may be violated if, for instance, the change from $Z = 0$ to $Z = 1$ not only increases the appeal of field 1 but also decreases the appeal of field 0. \begin{table}[ht!] \begin{tabular}{ccc} \toprule $D_0$ & $D_1$ & $D_2$ \\ \cmidrule(lr){1-3} 0 & 0 & 0 \\ 0 & 1 & 0 \\ 0 & 0 & 2 \\ 1 & 1 & 1 \\ 1 & 1 & 2 \\ 2 & 2 & 2 \\ 2 & 1 & 2 \\ 0 & 1 & 2 \\ \bottomrule \end{tabular} \caption{The support of $(D_0, D_1, D_2)$ in kirkeboen2016field.} \end{table} kirkeboen2016field derive causal interpretations of the IV estimand under (ref)--(ref) plus the restriction that $D_0 = 0$, which they call the “next-best” condition. The full set of assumptions are then equivalent to one-sided noncompliance, meaning that $D_j \in \{j, 0\}$ for each $j$. Under all these assumptions together with Assumptions (ref)--(ref), they show that an IV regression identifies the average treatment effects for the “compliers" of each instrument value relative to $Z = 0$. Our results in Sections (ref) and (ref) will characterize when the distribution of the data is consistent with (ref)--(ref), and the results in Section (ref) apply to the case when the “next best” condition is additionally imposed.
remarkAlthough we feel that Assumption (ref) is intuitive, it can also further be motivated through the lens of a fairly general (non-separable) random utility model where each instrument increases the utility of at most one unique choice. In particular, in Appendix (ref) we introduce such a model and establish its equivalence to Assumptions (ref)--(ref) in our setting. We note that the same model was shown by kline2016evaluating to imply the restrictions in (ref). We will in turn show that they are in fact equivalent.
remarkAssumption (ref) is distinct from the restriction considered in bai2024identifying, which in the current context states \begin{equation} Q \{D_j = j | D_k = j for some k \neq j\} = 1 . \end{equation} The restriction (ref) can be shown to be equivalent to (ref) when $J = 3$, but is weaker than (ref) when $J \geq 4$. Indeed, $(D_0, D_1, D_2, D_3) = (1, 1, 2, 2)$ is not ruled out by (ref), but is ruled out by (ref), because $D_0(\omega) = 1 \neq 0$ implies $j^\ast(\omega) = 1$, whereas $D_3(\omega) = 2 \neq 3$ implies $j^\ast(\omega) = 2$, a contradiction.
remarkIn a setting with a multi-valued treatment and a multi-valued instrument, heckman2018unordered propose a condition called “unordered monotonicity,” which requires that for each $0 \leq j \leq J - 1$ and $z, z' \in \mathcal Z$, either $Q\{I \{D_z = j\} \leq I \{D_{z'} = j\}\} = 1$ or $Q\{I \{D_{z'} = j\} \leq I \{D_z = j\}\} = 1$. sun2023instrument derives a set of inequalities that are implied by this assumption. We note that this assumption is not implied by, nor does it imply, Assumption (ref).

Main Results

In this section, we present our main results on sharp testable implications of Assumptions (ref)--(ref). In order to do so, in Section (ref), we first derive inequalities in terms of the conditional choice probabilities. Then, in Section (ref), for each $P$ that satisfies these inequalities, we explicitly construct a distribution $Q \in \mathbf Q_1$ such that $P = Q T^{-1}$, thus showing the inequalities are sharp.

Testable implications

The following theorem characterizes a set of necessary conditions in order for $P$ to be consistent with $\mathbf Q_1$. Recall $\mathcal Z = \{0, J_0, \ldots, J-1\}$. We further define \[ \mathcal Z(j) =

cases\mathcal Z, & for 0 \leq j \leq J_0 - 1 \\ \mathcal Z \backslash \{j\}, & for J_0 \leq j \leq J - 1 .

\] Here, when $J_0 = 0$, it is understood that $\mathcal Z(j) = \mathcal Z \backslash \{j\}$ for $0 \leq j \leq J - 1$.

theoremSuppose $P = Q T^{-1}$ for $Q \in \mathbf Q_1$. Then, for $z(j) \in \mathcal{Z}(j)$, $0 \leq j \leq J-1$, \begin{equation} \sum_{0 \leq j \leq J - 1} P \{D = j| Z = z(j)\} \leq 1 . \end{equation}

The inequalities in (ref) are direct consequences of the restriction in (ref). To see why, first suppose $J_0 = 0$, so that $z(j) \neq j$ for $0 \leq j \leq J - 1$, and consider the events \[ \{D_{z(0)} = 0\}, \dots, \{D_{z(J - 1)} = J - 1\}~. \] Fix $\omega \in \Omega$. Because $z(j) \neq j$ for all $j$, $D_{z(j)}(\omega) = j$ and (ref) imply that the default choice $j^\ast(\omega) = j$. As a result, the events listed above are disjoint across $0 \leq j \leq J - 1$, so their probabilities sum up to less than one, and (ref) follows. When $J_0 > 0$, $D_0(\omega) = j$ implies that $j^\ast(\omega) = j$, and hence $\{D_0 = j\}$ is disjoint from all other events as well. Therefore, (ref) holds in addition when $z(j) = 0$ for $0 \leq j \leq J_0 - 1$. \rule{2mm}{2mm}

The inequalities in (ref) restrict the substitution patterns jointly at different values of the instrument $z(0), \dots, z(J - 1)$. In particular, it does not suffice to consider pairwise substitution patterns when the instrument changes from one value to another. Note that $z(0), \dots, z(J - 1)$ do not have to be distinct. Therefore, for $j \neq k$, by setting $z(j) = k$ and $z(\ell) = j$ for all $\ell \neq j$, we have \[ P \{D = j | Z = k\} + \sum_{\ell \neq j} P \{D = \ell | Z = j\} \leq 1~, \] which implies

equation[equation omitted — 87 chars of source]

The inequality in (ref) states that the conditional probability of choosing $j$ is maximized at $Z = j$, which aligns with the intuition that $Z = j$ “encourages” towards $D = j$.

exampleWhen $J = 2$ and $J_0 = 0$, Theorem (ref) implies only one inequality: \[ P \{D = 1 | Z = 0\} \leq P \{D = 1 | Z = 1\}~. \] When $J = 3$ and $J_0 = 0$, (ref) leads to six inequalities: \begin{align*} P \{D = 0 | Z = 1\} & \leq P \{D = 0 | Z = 0\} \\ P \{D = 0 | Z = 2\} & \leq P \{D = 0 | Z = 0\} \\ P \{D = 1 | Z = 2\} & \leq P \{D = 1 | Z = 1\} \\ P \{D = 1 | Z = 0\} & \leq P \{D = 1 | Z = 1\} \\ P \{D = 2 | Z = 0\} & \leq P \{D = 2 | Z = 2\} \\ P \{D = 2 | Z = 1\} & \leq P \{D = 2 | Z = 2\} . \end{align*} In addition, by considering the cases where $z(0), z(1), z(2)$ are all distinct in (ref), we end up with two additional inequalities: \begin{align*} P \{D = 1 | Z = 0\} + P \{D = 2 | Z = 1\} + P \{D = 0 | Z = 2\} & \leq 1 \\ P \{D = 2 | Z = 0\} + P \{D = 0 | Z = 1\} + P \{D = 1 | Z = 2\} & \leq 1 . \end{align*} These two inequalities involve the substitution patterns across triplets of values of the instrument instead of simple pairwise comparison. In total, we obtain eight inequalities. For a general $J$, when $J_0 = 0$, we obtain $(J-1)^J$ inequalities.

When $J_0 > 0$, we obtain the following simplification of the inequalities:

corollarySuppose $P = Q T^{-1}$ for $Q \in \mathbf Q_1$ and $J_0 > 0$. Then, the inequalities described by (ref) are equivalent to the statement that, for $0 \leq j \leq J - 1$ and $k \in \mathcal Z$ such that $j \neq k$, \begin{equation} P \{D = j | Z = k\} \leq P \{D = j | Z = 0\} . \end{equation}

To see why the inequalities in Corollary (ref) follow from the ones in Theorem (ref), first note that for $0 \leq j \leq J - 1$ and $k \in \mathcal Z$ such that $j \neq k$, by setting $z(j) = k \in \mathcal Z(j)$ and $z(\ell) = 0 \in \mathcal Z(\ell)$ for $\ell \neq j$ in (ref), we get \[ \sum_{\ell \neq j} P \{D = \ell | Z = 0\} + P \{D = j | Z = k\} \leq 1~, \] which implies (ref) immediately. On the other hand, suppose (ref) holds for $0 \leq j \leq J - 1$ and $k \in \mathcal Z$ such that $j \neq k$. For $z(j) \in \mathcal Z(j)$ for $0 \leq j \leq J - 1$, we have $z(j) \neq j$ for $J_0 \leq j \leq J - 1$ and $z(j) \in \{0, J_0, \dots, J - 1\}$ for $0 \leq j \leq J_0 - 1$, and hence (ref) implies \[ \sum_{0 \leq j \leq J - 1} P \{D = j | Z = z(j)\} \leq \sum_{0 \leq j \leq J - 1} P \{D = j | Z = 0\} \leq 1~, \] so (ref) follows. \rule{2mm}{2mm}

The inequality in (ref) follows immediately from the restrictions in (ref). To see that, suppose $D_k = j$ for $k \neq j$. Because (ref) implies $D_k \in \{k, D_0\}$ and $k \neq j$, it has to be the case that $D_0 = j$. Therefore, \[ \{D_k = j\} \implies \{D_0 = j\}~, \] which immediately implies (ref). An interesting feature of Corollary (ref) is that only the substitution patterns between $Z = 0$ and $Z = k$ appear in the testable implications, but the substitution patterns between $Z = k$ and $Z = \ell$ for $k, \ell \neq 0$ and $k \neq \ell$ do not. Surprisingly, as we show in the next section, these inequalities exhaust all the information in the restrictions imposed by the model $\mathbf Q_1$. That is, as long as $P$ satisfies (ref), $P = Q T^{-1}$ for some $Q \in \mathbf Q_1$. As a result, when $J_0 > 0$, all information in the data about its consistency with the model is contained in the pairwise comparison between the choices when $Z = 0$ versus $Z = k$. Before proceeding, we revisit the Examples (ref)--(ref) and apply Corollary (ref).

exampleRecall in Example (ref) that $\mathcal Z = \{0, 2\}$ and $J_0 = 2$. In this case, two inequalities follow from (ref): \begin{align*} P \{D = 0 | Z = 2\} & \leq P \{D = 0 | Z = 0\} \\ P \{D = 1 | Z = 2\} & \leq P \{D = 1 | Z = 0\} . \end{align*} There is no additional inequality for $j = 2$.
exampleRecall in Example (ref) that $\mathcal Z = \{0, 1, 2\}$ and $J_0 = 1$. In this case, the following four inequalities follow from (ref): \begin{align*} P \{D = 0 | Z = 1\} & \leq P \{D = 0 | Z = 0\} \\ P \{D = 0 | Z = 2\} & \leq P \{D = 0 | Z = 0\} \\ P \{D = 1 | Z = 2\} & \leq P \{D = 1 | Z = 0\} \\ P \{D = 2 | Z = 1\} & \leq P \{D = 2 | Z = 0\} . \end{align*} These inequalities and their counterparts with an outcome will be used in Section (ref) to test $J_0 = 1$ in an empirical application to the dataset for behaghel2013robustness,behaghel2014private.

Sharpness of the Implications

Our next theorem is the converse of Theorem (ref)---namely, for each $P$ that satisfies (ref), there exists a distribution $Q \in \mathbf Q_1$ such that $P = Q T^{-1}$. In other words, the inequalities in Theorem (ref) are sharp in the sense that they exhaust all the information in the behavioral restrictions imposed by the model $\mathbf Q_1$. In the special case of $J = 2$, a constructive proof was provided by kitagawa2015test, but it does not extend to the case when $J > 2$. In particular, the construction in kitagawa2015test relies crucially on the observation that if $D = 1$ when $Z = 0$, then $(D_0, D_1) = (1, 1)$; this subgroup of individuals are referred to as the “always takers,” and the group that takes $D = 0$ when $Z = 1$ are called the “never takers.” The remaining probability mass is then assigned to the “compliers,” for whom $(D_0, D_1) = (0, 1)$. When $J > 2$, however, it is in general impossible to pin down the joint distribution of potential treatments using this argument, and hence the proof requires an entirely new strategy. Importantly, the proof that we present is still constructive, in that we construct $Q$ explicitly from the given distribution $P$.

theoremLet $P$ be a probability distribution on $\{0, \dots, J - 1\} \times \mathcal Z$ such that $P \{Z = z\} > 0$ for every $z \in \mathcal Z$. Further suppose (ref) holds for all $z(j)\in \mathcal{Z}(j)$, $0 \leq j \leq J - 1$. Then, there exists a $Q \in \mathbf Q_1$ such that $P = Q T^{-1}$.

To illustrate why (ref) is sufficient for determining whether $P$ is consistent with the model $\mathbf Q_1$, we start by sketching the construction when $J = 3$ and $J_0 = 0$. Let $Q^\ast$ denote a candidate distribution for which we wish to show $P = Q^\ast T^{-1}$ and $Q^\ast \in \mathbf Q_1$. We separately consider four classes of potential treatment vectors according to the value of the default choice $j^\ast$ and assign $Q^\ast$ separately for each class, such that $P = Q^\ast T^{-1}$.

enumerate[(a)] • $j^\ast = 0$. Consider $P \{D = 0 | Z = z\}$. First note that if $P = Q^\ast T^{-1}$, then for $z \in \{0, 1, 2\}$, \[ P \{D = 0 | Z = z\} = Q^\ast \{D_z = 0\} \geq Q^\ast \{(D_0, D_1, D_2) = (0, 0, 0)\}~. \] Respecting this constraint, we set \[ Q^\ast \{(D_0, D_1, D_2) = (0, 0, 0)\} = \min_{z \in \mathcal Z} P \{D = 0 | Z = z\}~. \] The inequalities in (ref) imply the minimum on the right-hand side is attained at $z \in \{1, 2\}$, and without loss of generality suppose it is attained at $z = 1$. If it is attained at $z = 2$, then a symmetric construction applies. Next, note any candidate $Q^\ast$ has to satisfy \begin{align*} P \{D = 0 | Z = 2\} = Q^\ast \{D_2 = 0\} & = Q^\ast \{(D_0, D_1, D_2) = (0, 0, 0)\} + Q^\ast \{(D_0, D_1, D_2) = (0, 1, 0)\} \\ & = P \{D = 0 | Z = 1\} + Q^\ast \{(D_0, D_1, D_2) = (0, 1, 0)\} , \end{align*} and hence we have to define \[ Q^\ast \{(D_0, D_1, D_2) = (0, 1, 0)\} = P \{D = 0 | Z = 2\} - P \{D = 0 | Z = 1\}~. \] This step stops here. Note we have assigned no mass to $(0, 0, 2)$, so that \[ Q^\ast \{(D_0, D_1, D_2) = (0, 0, 2)\} = 0~. \] The total mass we have assigned in this step is \begin{align*} Q^\ast \{(D_0, D_1, D_2) = (0, 0, 0)\} + Q^\ast \{(D_0, D_1, D_2) = (0, 1, 0)\} = \max_{z \neq 0} P \{D = 0 | Z = z\} . \end{align*} In all the events we have considered so far, $j^\ast = 0$, so that 0 is the default choice. It is chosen for at least two values of the instrument. • $j^\ast = 1$. Similarly as in (a), carry out the construction for events corresponding to $P \{D = 1 | Z = z\}$. The total mass assigned in this step is \[ \max_{z \neq 1} P \{D = 1 | Z = z\}~. \]$j^\ast = 2$. Similarly as in (a), carry out the construction for events corresponding to $P \{D = 2 | Z = z\}$. The total mass assigned in this step is \[ \max_{z \neq 2} P \{D = 2 | Z = z\}~. \] All of the events we have considered so far are disjoint. To see it, note the default choice is different in each class (a), (b) and (c), and it is chosen for at least two values of the instrument. For all other values of the instrument, the choice has to coincide with the instrument. Therefore, these events cannot intersect across classes. They are furthermore all disjoint from the final event: • Diagonal: Note the sum of the masses that we have assigned when considering $j = 0, 1, 2$ is \[ \sum_{0 \leq j \leq 2} \max_{z(j) \neq j} P \{D = j | Z = z(j)\} \leq 1~, \] because of (ref). We then assign all of the remaining mass to \[ Q^\ast \{(D_0, D_1, D_2) = (0, 1, 2)\} = 1 - \sum_{0 \leq j \leq 2} \max_{z(j) \neq j} P \{D = j | Z = z(j)\} \geq 0~. \]

$Q^\ast$ is clearly a probability measure. We now show $P = Q^\ast T^{-1}$. It suffices to verify $Q^\ast \{D_z = j\} = P \{D = j | Z = z\}$ for $j \in \{0, 1, 2\}$ and $z \in \{0, 1, 2\}$. We start by verifying that $Q^\ast \{D_z = 0\} = P \{D = 0 | Z = z\}$ for all $z$. Note $D_1 = 0$ and $D_2 = 0$ is only allowed in the events in (a) above, so that

align*[align* omitted — 319 chars of source]

Following similar arguments, we can show that \[ Q^\ast \{D_z = j\} = P \{D = j | Z = z\} \] for $0 \leq z \leq 2$, $0 \leq j \leq 2$, and $z \neq j$. Given these equalities, for each $k \in \{0, 1, 2\}$,

align*[align* omitted — 185 chars of source]

We have therefore successfully shown that $Q^\ast \{D_z = j\}$ for $0 \leq z \leq 2$ and $0 \leq j \leq 2$, and hence $P = Q^\ast T^{-1}$. \rule{2mm}{2mm}

The proof for general $J$ when $J_0 = 0$ follows similar arguments as in the previous sketch. Although we won't present the full proof in the main text, here we present some intuition on why the proof works in general. Note that although individuals cannot be classified into the three subgroups beyond the binary setting, each potential treatment vector permitted by (ref) can still be characterized by the default choice $j^\ast$ together with the set \[ \{z: D_z = z\}~. \] In other words, any potential treatment vector permitted by (ref) can be completely characterized by the default choice as well as the set of choices towards which the individual complies with the encouragement. For example, a person with $(D_0, D_1, D_2, D_3, D_4) = (0, 0, 2, 0, 4)$ can be thought of as a “0-default, $\{2, 4\}$-complier.” Similarly, we can call someone with $(D_0, D_1, D_2, D_3, D_4) = (0, 0, 0, 0, 0)$ a “0-always taker.” For each default value $0 \leq j \leq J - 1$, we order $0 \leq z \leq J - 1$ so that \[ P \{D = j | Z = z_1(j)\} \leq \dots \leq P \{D = j | Z = z_J(j)\}~. \] For this specific $j$, in step 1, we first pin down the probability of “$j$-always takers” as \[ P \{D = j | Z = z_1(j)\} = \min_{0 \leq z \leq J - 1} P \{D = j | Z = z\}~. \] Then, in step $\ell$ for $2 \leq \ell \leq J - 1$, for $\mathcal J_\ell = \{z_1(j), \dots, z_{\ell - 1}(j)\}$, we define the probability of “$j$-default, $\mathcal J_\ell$-compliers” as \[ P \{D = j | Z = z_\ell(j)\} - P \{D = j | Z = z_{\ell - 1}(j)\}~. \] Because we conclude at step $\ell = J - 1$, the total probability assigned for this specific $j$ is then \[ P \{D = j | Z = z_{J - 1}(j)\}~. \] The key reason why the construction guarantees $P = Q^\ast T^{-1}$ is as follows. If $z \neq j$, then (ref) implies $z = z_\ell(j)$ for some $1 \leq \ell \leq J - 1$. As summarized in Table (ref), $D_z = j$ happens only for “$j$-always takers,” “$j$-default, $\mathcal J_2$-compliers,” through “$j$-default, $\mathcal J_\ell$-compliers,” whose probabilities sum up to \[ P \{D = j | Z = z_\ell(j)\} = P \{D = j | Z = z\}~, \] as can be seen from Table (ref).

table[table omitted — 1,040 chars of source]

After carrying out this construction for each $j$, the remaining mass of $Q^\ast$ is assigned to the “diagonal” event that $(D_0, \dots, D_{J - 1}) = (0, \dots, J - 1)$, and the rest of the proof follows similarly as that for $J = 3$. The proof when $J_0 > 0$ builds on the proof when $J_0 = 0$ and requires further modifications.

remarkIn the special cases of $(J, J_0) = (3, 1)$ and $(J, J_0) = (3, 2)$, lee2024treatment present the testable implications in (ref) (along with some redundant inequalities), but do not discuss their sharpness. The testable implications they present also do not include an outcome. In Section (ref) below, we derive the sharp testable implications when an outcome is additionally considered.
remarkOne may also consider obtaining the sharp inequalities using a random set approach beresteanu2012partial based on Artstein's inequalities, or through a linear programming approach. These approaches provide implicit characterizations of the problem, by stating that $P$ is consistent with $\mathbf Q_1$ as long as certain linear systems have nonnegative solutions. They need to be implemented case-by-case for each value of $J$ and preclude the consideration of a continuous outcome. Consider the case $J_0 = 0$ as an example. To introduce the random set approach, following luo2024selecting, let $\mathcal S$ denote the support of $(D_0, \dots, D_{J - 1})$ allowed by (ref). For $0 \leq z \leq J - 1$, further define \[ B_z(D) = \{(D_0, \dots, D_{J - 1}) \in \{0, \dots, J - 1\}^J: D_z = D\}~. \] Define \[ G(D, Z) = \sum_{z \in \mathcal Z} I \{Z = z\} B_z(D) \cap \mathcal S~. \] The model predicts that $(D_0, \dots, D_{J - 1}) \in G(D, Z)$, which by Artstein's inequalities is equivalent to requiring for all $A \subseteq \{0, \dots, J - 1\}^J$, that \[ Q \{(D_0, \dots, D_{J - 1}) \in A | Z = z\} \geq Q \{G(D, Z) \subseteq A | Z = z\} \] for $0 \leq z \leq J - 1$. One would then need to find the core-determining class of the sets $A$, and the problem then becomes characterizing when a linear system has a nonnegative solution. Another approach is to determine whether there exists a probability measure $Q$ that satisfies (ref) and (ref) (or (ref) $J_0 > 0$) for $0 \leq j \leq J - 1$ and $0 \leq z \leq J - 1$. See, for example, bai2024inference for a detailed description. Such an approach is again equivalent to determining whether a linear system has a nonnegative solution. Through solving what is called a facet enumeration problem, one can further convert these implicit characterizations into a closed-form characterization like ours, but such a step is computationally prohibitive unless $J$ is very small.

Extensions

Results with an Outcome

In this section, we present the general results with an outcome variable. The discussion runs mostly in parallel with Section (ref). Let $Y \in \mathbf R$ denote an observed outcome and $Y_d$ for $0 \leq d \leq J - 1$ denote the potential outcome under treatment choice $d$. We allow $Y$ to be continuous or discrete and denote its support\footnote{Following pp.73--74 of lifshits1995gaussian, we define the (topological) support of $Y$ as the smallest closed set with probability one under $P$, i.e., $\mathcal Y := \bigcap \big \{ F \subseteq \mathbf R: F \text{ closed }, P \{Y \in F\} = 1 \big \}$.} by $\mathcal Y$. In addition to (ref), the observed outcome and the potential outcomes are related through \[ Y = \sum_{0 \leq d \leq J - 1} Y_d I \{D = d\}~. \] With some abuse of notation, we continue letting $T(\cdot)$ denote the mapping defined by the equation above together with (ref). Let $P$ denote the distribution of $(Y, D, Z)$ and $Q$ denote the distribution of $(Y_0, \dots, Y_{J - 1}, (D_z: z \in \mathcal Z), Z)$. We modify Assumption (ref) to include the potential outcomes:

assumption$(Y_0, \dots, Y_{J - 1}, (D_z: z \in \mathcal Z), Z) \perp \!\!\! \perp Z$ under $Q$.

Let $\mathbf Q_1^Y$ denote the set of all distributions $Q$ for which Assumptions (ref) and (ref) as well as (ref) hold. We first present the counterpart to Theorems (ref) and (ref).

theoremLet $P$ be a probability distribution on $\mathcal Y \times \{0, \dots, J - 1\} \times \mathcal Z$ such that $P \{Z = z \} > 0$ for every $z \in \mathcal Z$. Then, $P = Q T^{-1}$ for $Q \in \mathbf Q_1^Y$ if and only if both of the following sets of conditions hold: \begin{enumerate}[\rm (a)] • If for each $0 \leq j \leq J - 1$, the Borel sets $\{B_z(j): z \in \mathcal Z(j)\}$ form a partition of $\mathcal Y$, then \begin{equation} \sum_{0 \leq j \leq J - 1} \sum_{z \in \mathcal Z(j)} P \{Y \in B_z(j), D = j | Z = z\} \leq 1 . \end{equation} • For each Borel set $B \subseteq \mathcal Y$, $J_0 \leq j \leq J - 1$, and $k \neq j$, \begin{equation} P \{Y \in B, D = j | Z = k\} \leq P \{Y \in B, D = j | Z = j\} . \end{equation} \end{enumerate}

For $z(j) \in \mathcal Z(j)$, $0 \leq j \leq J - 1$, note that by taking $B_{z(j)}(j) = \mathcal Y$ and $B_z(j) = \emptyset$ for $0 \leq j \leq J - 1$ and $z \neq z(j)$ in (ref), we recover (ref).

As in Section (ref), when $J_0 > 0$, we obtain the following simplification of the inequalities:

corollarySuppose $J_0 > 0$. Then, $P \in \mathbf Q_1^Y T^{-1}$ if and only if Theorem (ref)(b) holds and for all Borel sets $B \subseteq \mathcal Y$, $0 \leq j \leq J - 1$ and $j \neq k$, \begin{equation} P \{Y \in B, D = j | Z = k\} \leq P \{Y \in B, D = j | Z = 0\} . \end{equation}

To see why Corollary (ref) holds, first note for $0 \leq j \leq J - 1$ and $J_0 \leq k \leq J - 1$, by taking $B_k(j) = B$, $B_0(j) = \mathcal Y \backslash B$, and $B_0(\ell) = \mathcal Y$ and $B_z(\ell) = \emptyset$ for $\ell \neq j$ and $z \neq 0$, we get \[ P \{Y \in B, D = j | Z = k\} + P \{Y \notin B, D = j | Z = 0\} + \sum_{\ell \neq j} P \{Y \in \mathcal Y, D = \ell | Z = 0\} \leq 1~, \] from which we immediately obtain (ref). On the other hand, suppose (ref) holds for all $0 \leq j \leq J - 1$, $J_0 \leq k \leq J - 1$, and $j \neq k$, and for $0 \leq j \leq J - 1$, Borel sets $\{B_z(j): z \in \mathcal Z(j)\}$ form a partition of $\mathcal Y$. Then, (ref) holds because \[ \sum_{0 \leq j \leq J - 1} \sum_{z \in \mathcal Z(j)} P \{Y \in B_z(j), D = j | Z = z\} \leq \sum_{0 \leq j \leq J - 1} \sum_{z \in \mathcal Z(j)} P \{Y \in B_z(j), D = j | Z = 0\} = 1~. \]

remarkWhen $J = 2$, the only inequalities implied by Theorem (ref) are \begin{align*} P \{Y \in B, D = 1 | Z = 0\} & \leq P \{Y \in B, D = 1 | Z = 1\} \\ P \{Y \in B, D = 0 | Z = 1\} & \leq P \{Y \in B, D = 0 | Z = 0\} . \end{align*} Note that these inequalities coincide with the simplification described in Corollary (ref) when $J_0 = 1$. Furthermore, these inequalities are exactly those derived by balke1997bounds,balke1997probabilistic and kitagawa2015test. We emphasize, however, that both the inequalities and the arguments to establish sharpness are novel beyond this setting.

One-Sided Noncompliance

In this section, we extend our results to settings with one-sided noncompliance. Formally, let $\mathcal Z = \{0, \dots, J - 1\}$ and $\mathbf Q_{1, 0}$ denote the collection of all distributions of $(D_0, \dots, D_{J - 1}, Z)$ such that Assumptions (ref)--(ref) hold and

equation[equation omitted — 83 chars of source]

The condition in (ref) requires that when assigned $Z = j$, the subject either takes up $D = j$ or the control status $D = 0$. Therefore, noncompliance can only be one-sided ($j$ to $0$) instead of the other way around. The only difference between (ref) and (ref) is that we additionally require $D_0 = 0$. Such a setting is prevalent in economics, especially if the instrument is the “gate-keeper” or eligibility for each program, so that one either takes up the program they are eligible for or falls back to the control status. As discussed in Example (ref), assuming the “next-best” condition in kirkeboen2016field together with (ref)--(ref) and (ref)--(ref) is equivalent to (ref). See angrist2009incentives for another example. The following theorem presents the sharp testable implications of (ref) without and with an outcome. As in Section (ref), let $\mathbf Q_{1, 0}^Y$ denote set of all distributions $Q$ for which Assumptions (ref) and (ref) as well as (ref) hold.

theorem\begin{enumerate}[\rm (a)] • $P = Q T^{-1}$ for some $Q \in \mathbf Q_{1, 0}$ if and only if $P \{D = j | Z = z\} = 0$ for $j \notin \{z, 0\}$. • $P = Q T^{-1}$ for some $Q \in \mathbf Q_{1, 0}^Y$ if and only if $P \{D = j | Z = z\} = 0$ for $j \notin \{z, 0\}$, and for each Borel set $B \subseteq \mathcal Y$ and each $0 \leq j \leq J - 1$, \begin{equation} P \{Y \in B, D = 0 | Z = j\} \leq P \{Y \in B, D = 0 | Z = 0\} . \end{equation} \end{enumerate}

We note that the results in Theorem (ref) follow from Corollary (ref) once we impose that $P \{D = j | Z = z\} = 0$ for $j \notin \{z, 0\}$. Indeed, with this additional restriction, (ref) is vacuous because the left-hand side is always 0. At the same time, both sides of (ref) are 0 for $j \neq 0$, so (ref) is only meaningful when $j = 0$, becoming (ref).

Inference

In this section, using the characterizations in Theorem (ref) or Corollary (ref), we construct tests to assess if the model is consistent with the distribution of the data. Formally, we test

equation[equation omitted — 105 chars of source]

in a way that is uniform in level across a large class of distributions. We now separately discuss tests for (ref) according to whether $\mathcal Y$ is discrete or continuous.

If $\mathcal Y$ is discrete (or is discretized ex-ante), then Theorem (ref) and Corollary (ref) generate a finite number of inequalities. In particular, (ref) becomes

equation[equation omitted — 133 chars of source]

for $z(j, y) \in \mathcal Z(j)$, $0 \leq j \leq J - 1$. The inequalities in (ref) become that for each $y \in \mathcal Y$,

equation[equation omitted — 108 chars of source]

In addition, the inequalities in (ref) become that for each $y \in \mathcal Y$,

equation[equation omitted — 106 chars of source]

The inequalities in (ref)--(ref) can be tested using any off-the-shelf inference method for a finite number of moment inequalities. See, for instance, pakes2017practical for an overview. Here we sketch how we can convert the problem into testing the feasibility of a linear program, so that we can directly apply recent results in, for instance, fang2023inference. Suppose $\mathcal{Y}$ is discrete and let $p$ denote the vector of $(P \{Y = y, D = j | Z = z\}: y \in \mathcal Y, 0 \leq j \leq J - 1, z \in \mathcal Z)$. We can represent the inequalities in Theorem (ref) and Corollary (ref) as

equation[equation omitted — 72 chars of source]

where $\Gamma$ and $\gamma$ have known entries which lie in $\{-1, 0, 1\}$. With a vector of slack variables $x$, (ref) is equivalent to

align*[align* omitted — 55 chars of source]

where $A$ is the identity matrix and $\beta(P) = \gamma - \Gamma p$. This formulation maps into the notation of fang2023inference, and their tests apply immediately.

If $\mathcal Y$ is continuous and we do not wish to discretize the outcome, then we focus on testing the inequalities in Corollary (ref), which apply when $J_0 > 0$; it seems difficult to test the general inequalities described in (ref) without first discretizing the outcome. Here we discuss how to test (ref) with a continuous outcome. We could develop a test based on the K-S statistic in kitagawa2015test, but following mourifie2017testing, we discuss a method that transforms the infinite number of inequalities in (ref) into a conditional moment inequality where the conditioning variable is $Y$ instead of $Z$. Indeed, note (ref) holds if and only if \[ E[I \{Y \in B\} I \{D = j, Z = k\}] P \{Z = 0\} \leq E [I\{Y \in B\} I \{D = j, Z = 0\}] P \{Z = k\}~, \] which holds for all Borel sets $B \subseteq \mathcal Y$ if and only if

equation[equation omitted — 120 chars of source]

with probability one for $Y$. Similarly, (ref) holds for all Borel sets $B \subseteq \mathcal Y$ if and only if

equation[equation omitted — 122 chars of source]

with probability one for $Y$. The conditional moment inequalities in (ref)--(ref) could then be tested using any off-the-shelf inference method for conditional moment inequalities. See, for instance, andrews2013inference, chernozhukov2013intersection, armstrong2016multiscale, and chetverikov2018adaptive.

Simulations

In this section, we study the properties of the inference procedures described in Section (ref). Our goal is to illustrate the size control and power properties of our tests for (ref), and to compare them with other procedures that are based on an implicit characterization of the inequalities, as discussed in Remark (ref). We present the results separately for $J_0 > 0$ and $J_0 = 0$.

Simulations for $J_0 > 0$

In this subsection, we study tests of the null hypothesis in (ref) for $J = 4$ and $J_0 = 1$, so that $\mathcal D = \mathcal Z = \{0, 1, 2, 3\}$. Throughout this subsection, $Z$ is uniformly distributed on $\mathcal Z$. In the simulation design, the potential treatments are generated by an additive random utility model, so that $D_z(\omega) \in \operatorname*{argmax}_{0 \leq j \leq J - 1} (\beta_j I\{j = z\} + \epsilon_j)$, where $(\epsilon_0, \dots, \epsilon_{J - 1})$ and $Z$ are independent. Let $(\beta_1, \beta_2, \beta_3) = (1.5, 1, 0.5)$ and $\beta_0$ be specified below. Further let $\epsilon = (\epsilon_0, \epsilon_1, \epsilon_2, \epsilon_3)' \sim N(\mu, I_4)$, where $\mu = (0.5, 1, 1.5, 2)'$ and $I_4$ is the $4 \times 4$ identity matrix. Note that when $\beta_0 = 0$, the distribution $P$ satisfies the null in (ref) with $J_0 = 1$; furthermore, for each $z \in \mathcal Z$, $E[\max_{j \in \mathcal D} (U_j(z) + \epsilon_j)] = E[U_z(z) + \epsilon_z] = 2.5$, so that the mean utility of the choice encouraged by $z$, $d=z$, stays constant across different values of the instrument $z$, but the mean utility of the alternative choice $d\neq z$ varies across $z$. The potential outcomes are determined as $Y_d = I \{d \geq 1\} + \xi$, where $\xi \perp \!\!\! \perp (\epsilon, Z)$ and $\xi \sim N(0, 1)$. Next, to assess the power of the tests, we further consider $\beta_0 \in \{-0.5, -1, -1.5\}$, so that it can be verified through direct calculation that (ref) is violated and hence the null in (ref) is violated.

We implement several tests of (ref) with $J_0 = 1$ at the 5% level. First, similarly as in mourifie2017testing, we test (ref)--(ref) using chernozhukov2013intersection and the accompanying clrtest package in stata chernozhukov2015implementing. Table (ref) presents the rejection probabilities in percentages. We implement both the parametric and local options for the estimation of the conditional moments and follow the choices of tuning parameters in mourifie2017testing. Second, we binarize $Y$ to $I \{Y \geq 1\}$ and test (ref)--(ref) using fang2023inference and the accompanying lpinfer package in R conroylau_2021_5506545. The results are displayed in Table (ref). Finally, with the discretized $Y$ and using fang2023inference, we directly test the implicit linear programming formulation in Remark (ref), i.e., whether there exists a probability measure $Q$ that satisfies (ref)\footnote{With an outcome $Y$, (ref) strengthens to $P\{Y=y, D=j | Z=z\} = Q\{Y_j=y, D_z=j\}$, which is the restriction we consider whenever there is an outcome.} and (ref). The results are displayed in Table (ref). We perform each test at sample sizes of 500, 1000, and 2000. For each value of $\beta_0$, and each sample size, we calculate the rejection probabilities across 5000 replications.

table[table omitted — 1,184 chars of source]
table[table omitted — 679 chars of source]
table[table omitted — 652 chars of source]

We begin by noting in Table (ref) that the results for testing (ref)--(ref) using chernozhukov2013intersection are very sensitive to the methods for estimating the conditional moments. With the local approach, the test fails to control size even at $n = 2000$. The seemingly higher power is likely a result of its poor size control. The parametric approach controls size well and has nontrivial power against $\beta_0 \in \{-1, -1.5\}$. The drastic difference in the size control of the two approaches may be because the conditional moments in (ref)--(ref) can be reasonably approximated by linear functions in the current data generating process, as can be seen from plotting the conditional moments. The over-rejection is also observed in the appendix to mourifie2017testing when $J = 2$, in which case our model is equivalent to the one in imbens1994identification. Consequently, we test the same inequalities presented in kitagawa2015test as mourifie2017testing do, and we further replicate this behavior in Appendix (ref) for a range of designs with $J = 2$. In any case, chernozhukov2013intersection require choosing a number of tuning parameters. Next, in Tables (ref) and (ref), after discretizing $Y$, both testing the moment inequalities in (ref)--(ref) and testing the linear program defined by (ref) and (ref) using fang2023inference control size well and have nontrivial power for $\beta = -1.5$ at all sample sizes, for $\beta = -1$ when $n = 1000$ and $n = 2000$, and also for $\beta_0 = -0.5$ when $n = 2000$. Between the two methods, testing the closed-form inequalities in (ref)--(ref) is more powerful than testing the linear program defined by (ref) and (ref). Finally, comparing Table (ref)(a) and Table (ref), although (ref)--(ref) based on discretizing $Y$ do not sharply characterize the model compared to the conditional moment inequalities in (ref)--(ref), the test based on the former discretization is still more powerful.

Simulations for $J_0 = 0$

Next, we study tests of the hypothesis in (ref) for $J_0 = 0$. The data is generated in the same way as in Section (ref), except now that $\beta_0 = 2$ under the null hypothesis. Throughout this subsection, we also binarize $Y$ as in Section (ref). In this case, testing (ref)--(ref) using fang2023inference is computationally prohibitive when $J > 4$, and testing the linear program defined by (ref) and (ref) using fang2023inference is also computationally prohibitive when $J > 5$. As a result, we compare the two methods when $J = 3$, and the design therefore uses only the first three entries of $\beta$ and $\epsilon$. Here, (ref)--(ref) generate 76 inequalities with 18 variables, while (ref) and (ref) define a linear system with 18 equalities with 80 latent variables. The results are presented in Tables (ref) and (ref). As in Section (ref), both tests control size well, and the test based on the closed-form characterization in (ref)--(ref) is more powerful.

table[table omitted — 676 chars of source]
table[table omitted — 653 chars of source]

Empirical Application

Because the datasets in kline2016evaluating and kirkeboen2016field are confidential, we apply our methods to the dataset from behaghel2013robustness,behaghel2014private. They study a randomized controlled trial with three job search counseling programs in France. In their setting, $Z = 0$ denotes assignment (of eligibility) to the usual public program without intensive counseling, which is thought of as the control group; $Z = 1$ denotes assignment to the public program $Z$ with intensive counseling; and $Z = 2$ denotes assignment to the private program with intensive counseling. $D = 0, 1, 2$ denotes participation in the corresponding programs. Assumption (ref) holds because $Z$ is randomly assigned. Job seekers did not necessarily comply with their assignment, i.e., they may not enter the program they were assigned to, and the noncompliance rate was as high as 60% on average behaghel2014private. As can be seen in Table (ref) below, some job seekers assigned to the control group participated in the two treatment programs, and other job seekers assigned to either of the two treatment programs entered the control or the other treatment program. As a result, the assignment $Z = j$ becomes an encouragement towards $D = j$ which one may or may not take up.

table[table omitted — 499 chars of source]

We consider three outcomes in their dataset: “EMPLOI 6MOIS,” which indicates exit from PES registers to employment; “EMPLOI AR110 6MOIS,” which indicates any employment; and “SUCCES OPP 6MOIS,” which indicates employment eligible for payment. All three outcomes are binary and were measured six months after the beginning of the program following their analysis. To recover a causal interpretation of the IV estimand, behaghel2013robustness impose an assumption called “extended monotonicity.” As shown in Appendix (ref), this assumption is strictly stronger than setting $J_0 = 1$ in our model. As a result, we focus on testing (ref) for $J_0 = 0$ and $J_0 = 1$. The test for $J_0 = 0$ indicates whether the assumption that each value of the instrument encourages towards an unique choice is consistent with the data. The test for $J_0 = 1$ further indicates whether it is truly without loss of generality to assume that assignment to the control group does not in fact encourage people to take up the control program. This may be a tempting assumption if the assignment to the control program is thought of as the “base state” of the instrument. Because all outcomes are discrete, we test (ref)--(ref) for $J_0 = 0$ and (ref)--(ref) for $J_0 = 1$ using fang2023inference. To illustrate the gains from the refinement by including an outcome, we also include the result for testing (ref) for $J_0 = 0$ and (ref) for $J_0 = 1$ which do not make use of the outcome. Table (ref) presents the $p$-values of the test in percentages for each null hypothesis and outcome combination.

table[table omitted — 661 chars of source]

Without the outcome, the tests for $J_0 = 0$ and $J_0 = 1$ both fail to reject at the 10% level. With any of the three outcomes, the test for $J_0 = 0$ fails to reject at the 10% level, but the test for $J_0 = 1$ rejects at the 10% level. For the outcome “EMPLOI AR110 6MOIS,” the test for $J_0 = 1$ further rejects at both the 5% and 1% levels. Therefore, including the outcome in the tests helps us reject the null hypothesis that $J_0 = 1$. To further investigate which moment inequalities among (ref)--(ref) are violated, for each outcome, we report the point estimates for the violated moments, i.e., moments that are estimated to violate (ref)--(ref). The results are presented in Table (ref). Among all moments, the one that is consistently violated across outcomes is $P \{Y = y, D = 1 | Z = 2\} \leq P \{Y = y, D = 1 | Z = 0\}$. This moment is also violated without an outcome, but the violation is more pronounced with an outcome. Together with the results in Table (ref), they indeed illustrate the gains from the refinement by including an outcome in the test. To understand the violated moments, note that in Assumption (ref), $D_2 = 1$ and $D_2 \in \{D_0, 2\}$ implies $D_0 = 1$. The violation of this condition implies that one cannot think of the assignment to the control program as the “base state.” In other words, at least for some people, $Z = 0$ strictly increases the appeal of the public program $D = 0$. Our test clearly pinpoints the violation is through the substitution pattern that $D_2 = 1$ and $D_0 \neq 1$, i.e., some subjects choose program 1 when encouraged towards program 2 but do not choose program 1 when there is no encouragement. As discussed in Section (ref), when $J_0 = 0$, changing $Z = 0$ to $Z = 2$ not only increases the appeal of $D = 2$, but decreases the appeal of $D = 0$ as well, making it possible for the person to switch to $D = 1$. When $J_0 = 1$, however, changing $Z = 0$ to $Z = 2$ only increases the appeal of $D = 2$ but does not decrease the appeal of $D = 0$, so one can only switch to $D = 2$ instead of $D = 1$.

table[table omitted — 1,148 chars of source]

Conclusion

In this paper, we proposed sharp, closed-form testable implications of a potential outcome model that assumes each value of the instrument only encourages towards one choice. Because the testable implications are in closed form, we can immediately check if they are violated for a given dataset and more importantly, pinpoint where the violation occurs. In an empirical application to the dataset for behaghel2013robustness,behaghel2014private, we find that the data is not compatible with the assumption that the control program is not encouraged by its corresponding assignment, which may be a tempting assumption if the control program is thought of as the “base state.” The identity of the violated inequalities further indicates which choice patterns are incompatible with the data.

Our assumption that each value of the instrument only encourages towards one choice is satisfied in many examples, sometimes with the additional restriction that some choices do not have an encouragement. That said, it would be interesting to study settings in which each value of the instrument possibly encourages towards multiple choices. We leave this question open for future work.