EconBase
← Back to paper

Inference for Treatment Effects Conditional on Generalized Principal Strata using Instrumental Variables

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

77,784 characters · 8 sections · 84 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference for Treatment Effects Conditional on Generalized Principal Strata using Instrumental Variables

spacing{1.25}
spacing{1.2} \begin{abstract} We propose a general approach to inference for a broad class of models that arise in the analysis of treatment effects with discrete-valued treatments and instruments and a general-valued outcome. In addition to instrument exogeneity, the main substantive assumption in our class of models rules out certain response types by assuming that they occur with probability zero. Here, the response type refers to the vector of potential outcomes and potential treatments, and we refer to a set of possible values for the response type as a generalized principal stratum. Through a series of examples, we show that this framework encompasses a wide variety of assumptions that have been considered in the previous literature. Our framework allows inference on any treatment effect parameter that can be expressed as the expectation of a function of the response type conditional on a generalized principal stratum. We develop methods for inference on such parameters under these assumptions, as well as methods for testing the validity of the assumptions themselves. A key result of our analysis is a characterization of the identified set for such parameters under these assumptions and the testable restrictions for the assumptions themselves in terms of existence of a non-negative solution to linear systems of equations with a special structure. We propose methods for inference exploiting this special structure and recent results in fang2023inference. \end{abstract} KEYWORDS: Multi-valued Treatments, Model Specification, Model Validity, Randomized Controlled Trial, Principal Strata, Instrumental Variables, Partial Identification.

\thispagestyle{empty} \setcounter{page}{1}

Introduction

We propose a general approach to inference for a broad class of models that arise in the analysis of treatment effects with discrete-valued treatments and instruments and a general-valued outcome. In addition to instrument exogeneity, the main substantive restriction imposed in our analysis is that certain values for the response types occur with probability zero. Here, the response type refers to the vector of potential outcomes and potential treatments, and we refer to a set of possible values for the response type as a generalized principal stratum. Through a series of examples, we show that our framework encompasses a wide variety of assumptions that have been considered in the previous literature, including the following: (i) restrictions in the analysis of randomized controlled trials (RCTs) under noncompliance considered, e.g., in imbens1994identification and cheng2006bounds; (ii) generalizations of these restrictions considered in bai2024identifying; (iii) revealed preference-type restrictions on response types considered in kirkeboen2016field, kline2016evaluating, and heckman2018unordered; and (iv) restrictions on the ordering of potential treatments or on the ordering of potential outcomes considered in manski1997monotone, manski1998monotone, and machado2019instrumental.\footnote{In the context of mediation analysis, kwon2024testing show that “full mediation” is equivalent to certain assumptions like those we consider.}

Our framework allows inference on any treatment effect parameter that can be expressed as the expectation of a function of the response type conditional on a generalized principal stratum. In this way, our framework accommodates many parameters that have been considered previously in the literature, including both average and distributional treatment effect parameters, such as the probability of being strictly helped and the probability of not being hurt by the treatment. It further accommodates versions of these parameters conditional on sets of possible values of potential treatments, such as the local average treatment effect in imbens1994identification, average effects conditional on principal strata in frangakis2002principal, and average effects conditional on sets of possible values of potential outcomes, such as the parameters considered in heckman1997making and heckman1998evaluating. By contrast, each of the papers cited in the preceding paragraph tailors their method to specific examples of such treatment effect parameters. Furthermore, because the types of restrictions considered in these papers are typically insufficient for identification of many parameters of interest, some of these papers focus specifically on those treatment effect parameters that are identified from the two-stage least squares estimand. A characterization of the parameters that are identified by such restrictions is given by navjeevan2023identification and goff2025doesividentificationrestrict.

A key result in our analysis is a characterization of the identified set for such parameters under these assumptions and the testable restrictions for the assumptions themselves in terms of existence of a non-negative solution to linear systems of equations with a special structure. Using this result, we build on results in fang2023inference to develop a test for the null hypothesis that a pre-specified value for the parameter of interest lies in the identified set as well as the null hypothesis that the testable restrictions are satisfied. Importantly, the resulting tests remain well behaved even in “high-dimensional” settings, meaning it is uniformly consistent in level over a large class of distributions satisfying weak assumptions and permitting, among other things, the support of the treatment and instrument or the support of the response types to be “large” relative to the sample size. Through test inversion, we may also construct confidence regions for such parameters that are uniformly consistent in level. We highlight that while our test builds on fang2023inference, it leverages the structure of our problem to obtain novel strong approximations that significantly improve on the coupling rates of fang2023inference. These coupling results for high-dimensional vectors of estimates of conditional probabilities and their bootstrap counterparts may be of independent interest.

Other methods for inference have been proposed in the prior literature for some special cases of our framework. For inference on certain treatment effect parameters, see, e.g., cheng2006bounds, bhattacharya2008treatment,bhattacharya2012treatment, and machado2019instrumental; for inference on the validity of the assumptions, see, e.g., kitagawa2015test, heckman2018unordered, sun2023instrument, and bai2025sharp. These approaches rely upon closed-form expressions for the identified set for the parameter of interest or testable inequalities, which must be proved or computed on a case-by-case basis. When such expressions need to be computed, the methods are feasible only when the support of response types is small. In contrast, our approach does not rely on such closed-form expressions and remains computationally feasible even in high-dimensional settings. A key strength of our approach is its broad applicability, including in settings for which no existing methods apply.

The remainder of the paper is organized as follows. Section (ref) introduces our formal setup and notation. In Section (ref), we provide several examples of parameters and restrictions previously considered in the literature that are nested by our framework. We present our inference method when the support of the outcome variable is restricted to be finite in Section (ref), and without such a restriction in Section (ref). We illustrate our method through a simulation study in Section (ref). Proofs of all results can be found in Appendix (ref).

Setup and Notation

Denote by $Y \in \mathcal Y$ an outcome, by $D \in \mathcal D$ a discrete valued endogenous regressor (i.e., the treatment), and by $Z \in \mathcal Z$ a discrete valued instrumental variable. To rule out degenerate cases, we assume throughout that $2 \leq |\mathcal Y|$, $2 \leq |\mathcal D| < \infty$, and $2 \leq |\mathcal Z| < \infty$. We emphasize that neither $\mathcal D$ and $\mathcal Z$ needs to be a subset of the real line; in this way, our framework accommodates vector-valued treatments and instruments provided they only take a finite number of values. In what follows, we will consider both settings in which $|\mathcal Y|$ is finite and those in which it is not. Further denote by $Y(d) \in \mathcal Y$ the potential outcome if $D = d \in \mathcal D$ and by $D(z) \in \mathcal D$ the potential treatment if $Z = z \in \mathcal Z$. As usual, we assume that

equation[equation omitted — 142 chars of source]

In our discussion below, it will be convenient to define $R_o \equiv (Y(d) : d \in \mathcal D)$ and $R_t \equiv (D(z) : z \in \mathcal Z)$, and set $R \equiv (R_o,R_t)$. Following heckman2018unordered, we will refer to $R_t$ as the “treatment response type.” By analogy with this terminology, we will also refer to $R_o$ as the “outcome response type” and to $R$ as simply the “response type.” We let $Q$ denote the distribution of $(R_o,R_t,Z)$ and note that by (ref) we have

equation[equation omitted — 52 chars of source]

for a transformation $T$ implicitly defined through (ref). Letting $P$ denote the distribution of $(Y, D, Z)$, we therefore have that $P = QT^{-1}$.

Below we will require that $Q \in \mathbf Q$, where $\mathbf Q$ is a class of distributions satisfying assumptions that we will specify. Different specifications of $\mathbf Q$ represent different assumptions that we impose on the distribution of potential outcomes and potential treatments. In this sense, $\mathbf Q$ may be viewed as a model for potential outcomes and potential treatments.

Given a distribution $P$ of $(Y,D,Z)$ and a model $\mathbf Q$, we define the set of $Q \in \mathbf Q$ that can rationalize $P$ as \[ \mathbf Q_0(P, \mathbf Q) = \{Q \in \mathbf Q: P = Q T^{-1}\}~. \] We say $\mathbf Q$ is consistent with $P$ if and only if $\mathbf Q_0(P, \mathbf Q ) \ne \emptyset$. For every model $\mathbf Q$ considered in this paper, every $Q \in \mathbf Q$ is assumed to satisfy the restriction:

assumption[Instrument Exogeneity] $R \perp \!\!\! \perp Z$ under $Q$.

Our remaining restrictions on $\mathbf Q$ will be formulated in terms of restrictions on possible values of the response type $R$. These restrictions will be expressed by specifying a set $\mathcal R \subseteq \mathcal Y^{|\mathcal D|} \times \mathcal D^{|\mathcal Z|}$ characterizing the possible values of $R$. In Section (ref), we provide several different choices of $\mathcal R$ that have been previously considered in the literature. For a given choice of $\mathcal R$, we will impose the following on every $Q \in \mathbf Q$:

assumption[Response Type Restrictions] $Q\{R \in \mathcal R \} = 1$.

Following frangakis2002principal, sets of the form $\{R_t = r_t\}$ for $r_t$ a possible value of $R_t$ are referred to as “principal strata.” By analogy with this terminology, we will refer to sets of the form $\{R \in \mathcal R'\}$ for $\mathcal R' \subseteq \mathcal R$ as “generalized principal strata.”

Finally, we define our parameters of interest. The parameters we consider can be written as

equation[equation omitted — 86 chars of source]

for different choices of function $g : \mathcal R \rightarrow \mathbf R$ and generalized principal strata $\mathcal R' \subseteq \mathcal R$. Note that for $\theta(Q)$ in (ref) to be well defined we require $Q$ to satisfy $Q\{R \in \mathcal R'\} > 0$. A wide variety of parameters can be accommodated in this way. In Section (ref), we will show specific parameters that have been considered previously in the literature have the structure in (ref). We note now, however, that natural choices of $g$ correspond to the effect of one treatment versus another (i.e., $g(R) = Y(d) - Y(d')$ for $d \in \mathcal D$ and $d' \in \mathcal D$) and the probability that one treatment leads to a larger outcome than another treatment (i.e., $g(R) = I\{Y(d) > Y(d')\}$ for $d \in \mathcal D$ and $d' \in \mathcal D$). We further note that when $\mathcal R' = \mathcal R$, $\theta(Q)$ defined in (ref) simplifies to $E_Q[g(R)]$.

For a given distribution $P$ and model $\mathbf Q$, note that the identified set for $\theta(Q)$ under $P$ relative to $\mathbf{Q}$ is

equation[equation omitted — 156 chars of source]

The set $\Theta_0(P, \mathbf Q) $ is nonempty whenever there exists at least one $Q \in \mathbf Q_0(P, \mathbf Q)$ with $Q\{R \in \mathcal R'\} > 0$. By construction, this set is “sharp” in the sense that for any value in the set there exists a distribution $Q$ that is consistent with $P$, satisfies the restrictions of the model, and for which $\theta(Q)$ equals the prescribed value.

Examples

In this section, we show how to accommodate several examples from the previous literature in our framework. Our discussion focuses in particular on Assumption (ref) and the specification of $\mathcal R$ in Assumption (ref), but, where the cited literature has emphasized specific parameters of interest, we additionally describe how those parameters of interest can be expressed as (ref) for suitable choices of $g$ and $\mathcal R'$.

example[RCT with one-sided noncompliance] Consider a multi-arm randomized controlled trial (RCT) with noncompliance, where $\mathcal D = \mathcal Z = \{0, \dots, K\}$, $Z=d$ denotes random assignment to treatment $d$, and $D(d)=d$ denotes that the subject would comply with assignment if assigned to treatment $d$. In this example, $Q$ satisfies Assumption (ref) because $Z$ is randomly assigned. Suppose noncompliance to the assignment is one-sided in the sense that one can always take the control $d=0$, but, for any other treatment $d \ne 0$, one can only take that treatment if assigned to it. This restriction can be expressed in terms of Assumption (ref) with $\mathcal R = \{( y(0), \dots, y(K), d(0), \dots ,d(K )) : d(j) \in \{0, j\} \text{ for all } j \in \mathcal D\}$. With the notable exception of cheng2006bounds, discussed below, analyses of causal parameters that condition on generalized principal strata in this context follow imbens1994identification in using the Wald estimand to identify local average treatment effect (LATE) parameters of the form $E_Q[Y(j)-Y(0) \mid D(j)=1]$ for $ j \in \{1, \dots, K\}$. Note that identification of such parameters does not allow comparison of $Y(j)$ to $Y(k)$ for $j, k \ne 0$ for any subgroup of subjects. Our framework nests these identified parameters as well as partially identified parameters including the relative treatment effectiveness for any subgroup defined by a generalized principal stratum.
examplecheng2006bounds study the special case of Example (ref) in which $\mathcal{D} = \mathcal{Z} = \{0,1,2\}$. Their “Monotonicity I” assumption corresponds to Assumption (ref) with $\mathcal R$ defined as in Example (ref). They further consider imposing the restriction that $Q\{ D(1)=1 \mid D(2)=2\}=1$, i.e., that subjects who would comply with assignment to treatment $2$ would also comply with assignment to treatment $1$. They argue that this assumption is plausible in contexts where the “cost” of compliance with treatment $1$ is lower than that with treatment $2$. The combination of these two restrictions can be formulated in terms of Assumption (ref) with $\mathcal R = \{(y(0), y(1),y(2), d(0), d(1), d(2)) : (d(0), d(1), d(2)) \in \{(0, 0, 0), (0, 1, 0), (0, 1, 2)\}\}$. In their application, cheng2006bounds focus on the following parameters: (i) $E_Q[Y(j) - Y(0) \mid (D(0),D(1),D(2)) = (0,1,2)]$ for $j \in \{1,2\}$; and (ii) $Q\{(D(0),D(1),D(2)) = r_t\}$ for different values of $r_t$. Each of these parameters can be expressed in the form of (ref) for appropriate choices of $g$ and $\mathcal R'$. For example, the parameter in (i) equals $E_Q[g(R) \mid R \in \mathcal R']$ for $g(R) = Y(j) - Y(0)$ and $\mathcal R' = \{(y(0), y(1),y(2), d(0), d(1), d(2)) \in \mathcal R: (d(0), d(1), d(2)) = (0, 1, 2)\}$. Our framework nests their parameters within their context, while also allowing for inference on the corresponding parameters for trials with more than three treatment arms, for which their approach becomes computationally infeasible.
example[Encouragement design] Consider a multi-arm RCT with possibly two-sided noncompliance, where $\mathcal D = \mathcal Z = \{0, \dots, K\}$, $Z=d$ denotes random assignment to treatment $d$, and $D(d)=d$ denotes that the subject would comply with assignment if assigned to treatment $d$. More generally, not necessarily in the context of an RCT, one can interpret $Z=d$ as random encouragement to treatment $d$ and interpret $D(d)=d$ as the subject would take treatment $d$ if encouraged to do so. In this example, $Q$ satisfies Assumption (ref) because $Z$ is randomly assigned. bai2024identifying generalize the “no-defier” restriction of imbens1994identification to \begin{equation} Q\{D(d) \ne d , D(d')=d for some d^{\prime} \ne d \}=0 , \end{equation} i.e., a subject that would not take treatment $d$ if assigned to (encouraged to take) $d$ but would also not take $d$ if assigned (encouraged) to some other treatment $d^{\prime} \ne d$. This restriction can be formulated in terms of Assumption (ref) with $\mathcal R = \{(y(0), \dots, y(K), d(0), \dots ,d(K) ) : d(j) \ne j ~ \Rightarrow d(k) \ne j ~ \forall ~ j, k \in \mathcal{D} \}$. bai2024identifying derive the identified set on unconditional average treatment effects in this context under (ref). While the two-stage least squares (TSLS) estimand in this context generally does not correspond to any well-defined causal parameter when $K \ge 2$, BHULLER2024105785 show that under strong, additional assumptions, it identifies a particular weighted average of strata-specific causal effects that may not be of a priori interest. In contrast, our framework allows inference on a broad class of parameters of substantive interest including those that condition on generalized principal strata and that need not coincide with the TSLS estimand.
example[RCT with close substitute] kline2016evaluating consider an RCT with a “close substitute” to study the effects of preschooling on educational outcomes. In their setting, $D \in \mathcal{D} = \{c,h,n\}$, where $D= c$ denotes home care (no preschool), $D = h$ denotes a preschool program called Head Start, and $D = n$ denotes preschools other than Head Start, i.e., the close substitute. Let $Z \in \mathcal{Z}=\{0,1\}$ denote an indicator variable for a randomized offer to attend Head Start. Assumption (ref) holds because $Z$ is randomly assigned. In evaluating the cost-effectiveness of Head Start, kline2016evaluating impose the restriction \begin{equation} Q \{ D(1) = h \mid D(0) \ne D(1)\}=1 . \end{equation} The condition in (ref) states that if a family's schooling choice changes upon receiving a Head Start offer, then they must choose Head Start when receiving the offer. In other words, it cannot be the case that upon receiving a Head Start offer, a family switches from no preschool to preschools other than Head Start, or the other way around. The restriction can be formulated in terms of Assumption (ref) with \[ \mathcal R = \{(y(c), y(h),y(n), d(0), d(1)) : (d(0), d(1)) \in \{(c, c), (h, h), (n, n), (c, h), (n, h)\}\}~. \] kline2016evaluating show that the Wald estimand identifies a weighted combination of “sub-LATEs”: \[ \frac{E[Y \mid Z=1] - E[Y \mid Z=0]}{P\{D = h \mid Z=1\} - P\{D=h \mid Z=0\}} = S_c \mathrm{SubLATE}_{ch} + (1 - S_c) \mathrm{SubLATE}_{nh}~, \] where \begin{align*} \mathrm{SubLATE}_{ch} & = E_Q[Y(h)-Y(c) \mid D(1)=h, D(0)=c] \\ \mathrm{SubLATE}_{nh} & = E_Q[Y(h)-Y(n) \mid D(1)=h, D(0)=n] \end{align*} and $S_c = Q\{ D(1) = h, D(0)=c \mid D(1) = h, D(0) \ne h\}$. Here, $S_c$ and the sub-LATEs can all be written as (ref) for appropriate choices of $g$ and $\mathcal R'$. kline2016evaluating show that these separate sub-LATE parameters are of substantive interest, and identify them under additional assumptions including imposing a multinomial normal selection model. Our framework allows inference on these sub-LATE parameters without imposing such additional assumptions sufficient to identify them.
examplekirkeboen2016field study the effects of college field of study on earnings. In their setting, $\mathcal D = \{0, 1, 2\}$ represents three fields of study, ordered by their (soft) admission cutoffs from lowest to highest. The instrument is $Z \in \{0,1,2\}$, where $Z = 1$ when the student crosses the (soft) admission cutoff for field 1 but not for field 2, $Z = 2$ when the student crosses the (soft) admission cutoff for field 2, and $Z = 0$ otherwise. The authors assume that $Z$ is exogenous in the sense that $Q$ satisfies Assumption (ref), and they impose: \begin{align} Q \{D(1) = 1 \mid D(0) = 1\} & = 1 , \\ Q \{D(2) = 2 \mid D(0) = 2\} & = 1 . \end{align} The monotonicity conditions in (ref)--(ref) require that crossing the cutoff for field 1 or 2 weakly encourages students toward that field. They further impose the following “irrelevance” conditions: \begin{align} Q \{I \{D(1) = 2\} & = I \{D(0) = 2\} \mid D(0) \neq 1, D(1) \neq 1\} = 1 , \\ Q \{I \{D(2) = 1\} & = I \{D(0) = 1\} \mid D(0) \neq 2, D(2) \neq 2\} = 1 . \end{align} The condition in (ref) states that if crossing the cutoff for field 1 does not cause a student to switch to field 1, then it also does not cause them to switch to or away from field 2. A similar interpretation applies to (ref). The restrictions in (ref)--(ref) can be formulated in terms of Assumption (ref) with $\mathcal R = \{(y(0), \allowbreak y(1),y(2), d(0), d(1), d(2)) : (d(0), d(1) , d(2)) \in \{(0, 0, 0), (0, 0, 2), (0,1,0), (0,1,2), (1, 1, 1), (1, 1, 2), (2, 1, 2), \allowbreak (2,2,2)\} \}$. kirkeboen2016field regard the irrelevance conditions (ref)–(ref) as strong but, under them, show that the TSLS estimands identify $E_Q[Y(2)-Y(0) \mid D(1)=1, D(0)=0]$ and $E_Q[Y(2)-Y(0) \mid D(2)=2, D(0)=0]$. Under their assumptions (ref)–(ref), our analysis allows one to analyze additional parameters that do not correspond to the TSLS estimands, including the average relative effects of field 1 versus field 2 for any given generalized principal strata. Our framework further allows one to investigate what can be learned about these parameters without imposing (ref)–(ref). For further discussion of this example and Example (ref), see lee2023treatment and bai2025sharp.
example[Restrictions from WARP] heckman2018unordered consider a setting in which there is a voucher $Z$ that subsidizes in different ways that we specify below the purchase of three different cars, that we denote by $A$, $B$ and $C$. They further assume that the voucher is randomly assigned, so that Assumption (ref) holds. The treatment $D$ corresponds to the purchase of the different cars; let $D = A$ correspond to the purchase of car $A$, $D = B$ correspond to the purchase of car $B$, and $D = C$ correspond to the purchase of car $C$. In this setting, heckman2018unordered consider a series of examples in which they use the Weak Axiom of Revealed Preference (WARP) to restrict treatment response types, each of which can be formulated in terms of Assumption (ref) for appropriate choice of $\mathcal{R}.$ \begin{enumerate}[\rm (i)] • In their leading example, $Z = 0$ corresponds to no voucher, $Z = 1$ corresponds to a voucher that subsidizes the purchase of car $A$, and $Z = 2$ corresponds to a voucher that subsidizes the purchase of either $B$ or $C$. WARP generates the restriction in Table III of heckman2018unordered. These restrictions can be formulated in terms of Assumption (ref) with $\mathcal R = \{(y(A), y(B),y(C), d(0), d(1), d(2)) : (d(0), d(1) , d(2)) \in \{(A, A, A), (A, A, B), (A, A, C), (B, A, B), (B, B, B), (C, A, C), (C, C, C)\}\}$. • In a second example, $Z = 0$ corresponds to no voucher, $Z = 1$ corresponds to a voucher that subsidizes the purchase of $B$, and $Z = 2$ corresponds to a voucher that subsidizes the purchase of $B$ or $C$. WARP generates the restriction in Table V of heckman2018unordered, which can be formulated in terms of Assumption (ref) with $\mathcal R = \{(y(A), y(B),y(C), d(0), d(1), d(2)) : (d(0), d(1) , d(2)) \in \{(A, A, A), (A, A, C), (A, B, B), \\(A, B, C), (B, B, B), (C, C, C), (C, B, C)\} \}$. • Finally, in a third example, $Z = 0$ corresponds to a voucher that subsidizes the purchase of $C$, $Z = 1$ corresponds to a voucher that subsidizes the purchase of $B$, and $Z = 2$ corresponds to a voucher that subsidizes the purchase of $B$ or $C$. WARP generates the restriction in Table VI of heckman2018unordered, which can be formulated in terms of Assumption (ref) with $\mathcal R = \{(y(A), y(B),y(C), d(0), d(1), d(2)) : (d(0), d(1) , d(2)) \in\{(A, A, A), (A, B, B), (B, B, B), (C, A, C), (C, B, B), (C, B, C), (C, C, C)\}\}$. \end{enumerate} heckman2018unordered provide several additional examples that also fit within our framework. They establish conditions for identification of average treatment effect parameters that condition on generalized principal strata in these examples; in comparison, our analysis allows inference for a more general class of parameters that may be only partially identified.
example[Ordered monotonicity with known direction] Consider a setting in which both the treatment and the instrument admit natural orderings, and write $\mathcal D = \{0, \ldots, K_d\}$ and $\mathcal Z = \{0, \ldots, K_z\}$. Suppose further that $Z$ is randomly assigned, so Assumption (ref) holds. In many applications, it is reasonable to assume the monotonicity restriction $Q\{D(j) \geq D(k)\} = 1$ for any pair of instrument values $j,k \in \mathcal Z$ with $j \geq k$. For example, dupas2014short reports an experiment in which $Z$ represents different randomly assigned subsidy levels for insecticide-treated bed nets, and $D$ denotes the total number of bednets purchased over a two-year period. The monotonicity restriction in this context is that one would purchase at least as many insecticide-treated bed nets with a greater subsidy as with a lower subsidy. This restriction can be expressed in terms of Assumption (ref) as $\mathcal R = \{(y(0), \dots, y(K_d), d(0), \dots ,d(K_z) ) : d(K_z) \geq \ldots \geq d(0)\}$. The results of angrist1995two-stage imply that the TSLS estimand in this context identifies a specific weighted average of strata-specific causal effects, whereas our analysis allows inference on a broader class of parameters—including partially identified parameters—that need not coincide with the TSLS estimand.
example[Monotone treatment response] Consider a setting in which the treatment admits a natural ordering and write $\mathcal D = \{0, \ldots, K_d\}$. Suppose further that $Z$ is randomly assigned, so Assumption (ref) holds. Following manski1997monotone, it is often reasonable to assume that $Q\{Y(j) \geq Y(k)\} = 1$ for any pair of treatments $j,k \in \mathcal D$ with $j \geq k$. For example, in manski1998monotone, the treatment is years of schooling and the outcome is log(wage), and the restriction is that additional schooling leads to (weakly) higher wages. This restriction can be formulated in terms of Assumption (ref) with $\mathcal R = \{(y(0), \dots, y(|\mathcal D| - 1), d(0), \dots ,d(|\mathcal Z|-1 ) ) : y(K_d) \geq \ldots \geq y(0)\}$. While manski1997monotone develops partial-identification results for unconditional treatment parameters in this context without restrictions on treatment-response types, our analysis additionally allows such restrictions to be incorporated and enables inference on parameters that condition on generalized principal strata.
example[Harmless treatment] Consider a setting in which $Z\in \mathcal Z = \{0,\ldots, K_z\}$ is randomly assigned, so Assumption (ref) holds, and there is a baseline treatment, i.e., the control, corresponding to $0 \in \mathcal D = \{0,\ldots, K_d\}$. In many applications, it is reasonable to assume that the remaining treatments are harmless relative to the control. For example, in angrist2009incentives, the outcome is academic performance, the control is no treatment, and the noncontrol treatments are (i) providing students with academic peer-advising service; (ii) financial incentives for good academic performance; and (iii) both (i) and (ii). It is therefore natural to assume that $Q\{Y(d) \geq Y(0)\} =1$ for all $d \in \mathcal D$. This restriction can be formulated in terms of Assumption (ref) with $\mathcal R = \{(y(0), \dots, y(K_d), d(0), \dots ,d(K_z) ) : y(d) \geq y(0) \text{ for all } d \in \mathcal D \}$. This restriction is an example of what manski1997monotone termed a semi-monotone ordering of outcomes. As discussed above in Example (ref), our analysis differs from that of manski1997monotone by additionally allowing restrictions on treatment-response types to be incorporated and enabling inference on parameters that condition on generalized principal strata.

In many of the examples discussed above, $Y$ may be an ordinal outcome. In such instances, average treatment effects may not be interpretable. Nonetheless, researchers may consider other parameters that fit our framework, such as: (i) $Q\{Y(j) > Y(k) \mid R \in \mathcal R' \}$, the conditional probability of benefit of treatment $j$ versus treatment $k$; (ii) $Q\{Y(j) \ge Y(k) \mid R \in \mathcal R' \}$, the conditional probability of no harm of treatment $j$ versus treatment $k$; or (iii) $Q\{Y(j) > Y(k) \mid R \in \mathcal R'\} - Q\{Y(k) > Y(j) \mid R \in \mathcal R'\}$, the conditional relative treatment effect of treatment $j$ versus $k$. Each of these parameters can be written as (ref) for appropriate choices of $g$. These parameters have been previously studied for the special case in which $\mathcal{R}'=\mathcal{R}$ and $|\mathcal{D}| = 2$ or $3$ by lu2018treatment, huang2019constructing, and gabriel2024sharp.

It is also instructive to discuss examples of restrictions that do not fit our framework. angrist1995two-stage, for instance, consider a generalization of the the monotonicity assumption in Example (ref) in which the direction of the monotonicity is not known a priori, i.e., they consider the restriction that, for any pair of values for the instrument $j,k \in \mathcal Z$, either $Q\{D(j) \geq D(k)\} = 1$ or $Q\{D(j) \leq D(k)\} = 1$. Such a restriction is also analyzed in vytlacil2006ordered and in heckman2018unordered, where the latter paper terms it “ordered monotonicity.” This restriction is equivalent to imposing $Q \{R \in \mathcal R_\pi\} = 1$ for some $\pi \in \Pi(|\mathcal Z|)$, where $\Pi(|\mathcal Z|)$ is the set of all permutations of $\{0, \dots, |\mathcal Z| - 1\}$ and $\mathcal R_\pi = \{(y(0), \dots, y(|\mathcal D| - 1), d(0), \dots ,d(|\mathcal D|-1 ) ) : d(\pi(|\mathcal Z| - 1)) \geq \ldots \geq d(\pi(0))\}$. A similar generalization of the assumption in Example (ref) can also be considered. In particular, machado2019instrumental study such a generalization with $|\mathcal D| = |\mathcal Z| = 2$. These models, in which the permutation $\pi$ is not known a priori, fall outside the class of models we consider.

Inference with Discrete Outcomes

In this section, we propose a test for conducting inference on the parameter of interest $\theta(Q)$. To this end, recall that $P$ denotes the distribution of $(Y,D,Z)$ and set $\mathcal M \equiv \mathcal Y \times \mathcal D \times \mathcal Z$. In this section only, we suppose $|\mathcal Y| < \infty$. Doing so allows us to introduce our method succinctly. This assumption will be removed in Section (ref), where we allow for a more general outcome through discretization. For $(y, d, z) \in \mathcal M$, define $P_{ydz} = P \{Y = y, D = d, Z = z\}$ and $P_z = P \{Z = z\}$. In what follows, we assume that $P\in \mathbf P$, where $\mathbf P$ is a “large” class of distributions that we will specify below. The class $\mathbf P$ can depend on the sample size $n$, but we suppress the dependence from our notation. We consider the problem of testing

equation[equation omitted — 122 chars of source]

at level $\alpha \in (0,1)$, where, for a pre-specified value of $\theta_0$, the null hypothesis we consider is given by

equation[equation omitted — 106 chars of source]

As we show below, a test of this null hypothesis can be used to construct confidence regions for $\theta(Q)$ through test inversion over the pre-specified value $\theta_0$.

The key insight underlying the construction of our test is the following theorem, which provides a convenient reformulation of $\mathbf P_0$ in terms of existence of a non-negative solution to a (possibly under-determined) system of linear equations in which the “coefficients” are known. The statement of the theorem involves the parameter

equation[equation omitted — 95 chars of source]

where $P_{yd|z} \equiv P\{Y = y, D = d | Z = z\}$. Using this notation, we obtain the following result:

theoremSuppose $|\mathcal Y| < \infty$. Let $\mathbf Q$ be the set of all distributions of $(R,Z)$ satisfying Assumptions (ref) and (ref), and $\mathbf P$ be the set of all distributions of $(Y,D,Z)$ satisfying $P\{Z=z\} > 0$. Then, for $\beta(P)$ defined in (ref) and a matrix $A$ defined in the beginning of Appendix (ref) that depends only on $\mathcal R$ in Assumption (ref), $\theta_0$ in (ref), and the quantities $g$ and $\mathcal R'$ in the definition of $\theta(Q)$ in (ref), it follows \begin{equation} \mathbf P_0 \subseteq \{P \in \mathbf P : Ax = \beta(P) for some x \geq 0 \} . \end{equation} Furthermore, if $\mathbf P_0\neq \emptyset$, then $ \{P\in \mathbf P : Ax = \beta(P) \text{ for some } x \geq 0\}\subseteq {\rm cl}_{\mathbf P}(\mathbf P_0) $, where $\mathrm{cl}_{\mathbf P}({\mathbf P_0})$ denotes the closure of $\mathbf P_0$ in $\mathbf P$ (understood as a subset of $\mathbf R^{|\mathcal M|}$).

The key implication of Theorem (ref) is that we may test the null hypothesis of interest by examining whether there exists a positive vector $x$ satisfying $Ax=\beta(P)$. Because $\theta(Q)$ need not always be well defined when $\mathcal R' \neq \mathcal R$, (ref) involves an inclusion rather than an equality. From a testing perspective, however, $\mathbf P_0$ and its closure are indistinguishable, in the sense that any level $\alpha$ test of whether $P$ belongs to $\mathbf P_0$ necessarily has power no larger than $\alpha$ against any $P$ in the closure $\text{cl}(\mathbf P_0)$ romano2004non. Therefore, the second part of Theorem (ref) may be interpreted as showing that testing whether $P\in \mathbf P_0$ is in fact equivalent to testing whether $\beta(P) = Ax$ for some $x\geq 0$. This conclusion holds under the condition that $\mathbf P_0$ is not empty. However, by definition, $\mathbf P_0$ is empty whenever there is no possible distribution of the data $P\in \mathbf P$ that is compatible with the null hypothesis. In other words, we may focus on testing whether $\beta(P) = Ax$ for some $x\geq 0$ unless the null hypothesis specifies restrictions that are incompatible with any $P$ and, in this sense, may be interpreted as being logically inconsistent with each other.

Theorem (ref) also has implications for deriving an analytical expression for the identified set of $\theta(Q)$. We discuss these implications in the next two remarks, and elaborate on them in the appendix.

remarkWhen $Q\{R \in \mathcal R'\}$ is known or identified, in the sense that $\{Q\{R \in \mathcal R'\} : Q \in \mathbf Q_0(P,\mathbf Q)\}$ is a singleton for all $P \in \mathbf P$ for which $\mathbf Q_0(P,\mathbf Q) \neq \emptyset$, it is possible to use the characterization of $\mathbf P_0$ in Theorem (ref) to derive closed-form expressions $L(P)$ and $U(P)$ such that $\Theta_0(P,\mathbf Q) = [L(P), U(P)]$ for all $P \in \mathbf P$ for which $\mathbf Q_0(P,\mathbf Q) \neq \emptyset$ and $Q\{R\in \mathcal R'\} > 0$. As in balke1997bounds, the key idea is to express $L(P)$ and $U(P)$ as the values of linear programs and to use the duals of these programs to obtain expressions for $L(P)$ and $U(P)$ in terms of $P_{yd|z}$ through vertex enumeration. Computing $L(P)$ and $U(P)$ in this way rapidly becomes computationally prohibitive as the support of $(Y,D,Z)$ becomes large. We describe this procedure in Appendix (ref) and further develop a related procedure when $Q\{R \in \mathcal R'\}$ is not identified. Following the approach in machado2019instrumental, it is possible to use such expressions for inference, but a virtue of the approach we pursue here is that it does not rely on knowledge of these expressions.
remarkThe identified set for the parameter of interest in general depends on the distribution of the full vector $(Y,D,Z)$. This is the case even for parameters that only depend on the distribution of treatment response types $R_t$, $Q\{R_t=r_t\}$---parameters that, when point-identified, typically depend only on the distribution of $(D,Z)$ heckman2018unordered, navjeevan2023identification. In Appendix (ref), we illustrate this phenomenon by deriving the closed-form expression for the identified set in an example using the method described in Remark (ref), and comparing it with the bounds using only the distribution of $(D,Z)$.

Theorem (ref) immediately permits us to test whether the model is correctly specified in the sense that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$. To see this, consider $g \equiv 1$, $\theta_0 = 1$, and $\mathcal R' = \mathcal R$. For such values of these quantities,

eqnarray*[eqnarray* omitted — 164 chars of source]

Therefore, $\mathbf P_0 = \{P \in \mathbf P : 1 \in \Theta_0(P,\mathbf Q)\} = \{P \in \mathbf P : \mathbf Q_0(P,\mathbf Q) \neq \emptyset \}$. In this case, the equality in the last row of $Ax = \beta(P)$ in (ref) is always true, so it can be omitted. In order to state the result formally, denote by $\tilde A$ and $\tilde \beta(P)$ all but the last rows of $A$ and $\beta(P)$, respectively, in Theorem (ref).

corollarySuppose $|\mathcal Y| < \infty$. Let $\mathbf Q$ be the set of all distributions of $(R,Z)$ satisfying Assumptions (ref) and (ref), and $\mathbf P$ be the set of all distributions of $(Y,D,Z)$ satisfying $P\{Z=z\} > 0$. For $g \equiv 1$, $\theta_0 = 1$, and $\mathcal R' = \mathcal R$, \begin{equation} \mathbf P_0 = \{P \in \mathbf P : \mathbf Q_0(P, \mathbf Q) \neq \emptyset \} = \{P \in \mathbf P : \tilde A x = \tilde \beta(P) for some x \geq 0 \} . \end{equation}

Because $\theta(Q)$ is always well defined when $\mathcal R' = \mathcal R$, (ref) in Corollary (ref) differs from its counterpart (ref) in Theorem (ref) in that it involves an equality instead of an inclusion.

remarkBased on the characterization in Corollary (ref), we show in Appendix (ref) how a strategy like that described in Remark (ref) can be used to derive a set of analytical inequalities that hold if and only if $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$. We emphasize, however, that our inference methods below do not rely on these analytical inequalities.

By Theorem (ref), we may test the null hypothesis of interest by testing whether $P$ is such that

equation[equation omitted — 72 chars of source]

In recent work, fang2023inference derived a test for whether $P$ satisfies the restriction in (ref) for a general class of problems; see also bai2022testing for other approaches. In what follows, we build on fang2023inference by employing the structure of our problem to sharpen rate conditions and establish the validity of their test under weaker requirements. Following Corollary (ref), it is straightforward to modify the test described below to test the null hypothesis that the model is correctly specified in the sense that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$.

In order to describe the test statistic, we require some additional notation. We will assume the availability of an i.i.d.\ sample $\{V_i\}_{i=1}^n$ with $V_i=(Y_i,D_i,Z_i)$, $V_i \sim P$, and let $\hat P_n$ denote the empirical distribution corresponding to the sample $\{V_i\}_{i=1}^n$. We further set $\hat \beta_n \equiv \beta(\hat P_n)$ to denote the plug-in estimator for the parameter $\beta(P)$ and let $\hat \Omega \equiv \Omega(\hat P_n)$, where $\Omega(P)$ is a diagonal weighting matrix whose $(i,i)$-entry equals the asymptotic standard deviation of the $i^{th}$ row of $\sqrt n(\hat \beta_n - \beta(P))$. Specifically, the first $|\mathcal M|$ diagonal entries of $\Omega(P)$ equal $$\left(\frac{P_{yd|z}(1-P_{yd|z})}{P_z}\right)^{1/2}$$ for some $(y,d,z)\in \mathcal M$, and the final two diagonal entries of $\Omega(P)$ are equal to zero. Finally, we let $A^\dagger$ denote the Moore-Penrose pseudoinverse of $A$.

Given the introduced notation, we define the set $\hat{\mathcal V}_n \equiv \{s : A^\dagger s \leq 0, \|\hat \Omega_n (AA')^\dagger s\|_1 \leq 1\}$ and set

equation[equation omitted — 139 chars of source]

as our test-statistic -- here, $\|a\|_1 = \sum_{i=1}^d |a_i|$ for any vector $a = (a_1,\ldots, a_d) \in \mathbf R^d$ and $\langle a,b\rangle = a^\prime b$. From results in fang2019inference, $T_n$ represents a sample analogue to an equivalent formulation of the restriction in (ref) obtained from Farkas' lemma. In particular, the population analogue to $T_n$ equals zero whenever (ref) indeed holds and, as a result, under the null hypothesis $T_n$ should not be “too large.” We highlight that the test statistic can be computed through linear programming. As a result, $T_n$ can be reliably and quickly computed even in high-dimensional applications.

Unfortunately, the asymptotic distribution of $T_n$ is not pivotal. In particular, its asymptotic distribution depends on the “directions” $s \in \hat {\mathcal V}_n$ at which the population analogue $\langle A^\dagger s, A^\dagger \beta(P)\rangle$ is “close” to zero. In order to obtain a valid critical value, we therefore rely on the nonparametric bootstrap and a construction that asymptotically excludes directions $s$ that do not play a role in the distribution of $T_n$; i.e., directions $s$ for which $\langle A^\dagger s, A^\dagger \beta(P)\rangle $ is “far” from zero. Specifically, for the latter purpose we introduce

align*[align* omitted — 283 chars of source]

which represents an estimator for $\beta(P)$ that is restricted to satisfy the null hypothesis. Letting $\{V_i^*\}_{i=1}^n$ denote a bootstrap sample, (i.e., a sample drawn i.i.d.\ from $\hat P_n$), and $\hat \beta_n^*$ the bootstrap analogue to $\hat \beta_n$, we set

equation[equation omitted — 241 chars of source]

as our bootstrap statistic. Here, $\lambda_n \in [0,1]$ represents a penalty satisfying $\lambda_n =o(1)$, and whose role is to ensure that directions $s$ that do not impact the finite-sample distribution of our test statistic do not impact the bootstrap statistic either. We again highlight that, just like $T_n$, the bootstrap statistic $T_n^*$ can also be computed through linear programming.

The critical value for a level $\alpha$ test is then simply the $1-\alpha$ quantile of the distribution of $T_n^*$ conditionally on the sample $\{V_i\}_{i=1}^n$. Formally, the critical value for our test is given by

equation[equation omitted — 147 chars of source]

As usual, the critical value $\hat c_n(1-\alpha)$ can be computed through simulation by drawing multiple samples $\{V_i^*\}_{i=1}^n$ i.i.d.\ from $\hat P_n$ and computing the $1-\alpha$ quantile of the corresponding sample of bootstrap statistics. Given the critical value, we next formally introduce our test, which rejects whenever $T_n$ exceeds $\hat c_n(1-\alpha)$:

equation[equation omitted — 90 chars of source]

We establish the asymptotic validity of our test under the following main assumption.

assumption(i) $V_i = (Y_i,D_i,Z_i), i = 1,\ldots, n$ is an i.i.d.\ sample with $V_i \sim P\in \mathbf P$; (ii) $(\hat \beta_n - \beta(P)) \in \text{range}(A)$ with probability tending to one uniformly in $P\in \mathbf P$; (iii) For some constants $0 < c_1 < c_2$, $$c_1 \leq \inf_{P\in \mathbf P} \min_{(y,d,z)\in \mathcal M} |\mathcal M| P_{ydz} \leq \sup_{P\in \mathbf P} \max_{(y,d,z)\in \mathcal M} |\mathcal M| P_{ydz} \leq c_2.$$

Assumption (ref)(i) simply formalizes our requirement that the sample be independent and identically distributed. In Assumption (ref)(ii) we impose the additional condition that $(\hat \beta_n - \beta(P))$ belong to the range of the matrix $A$ with probability tending to one. This requirement is satisfied in our examples, where $(\hat \beta_n - \beta(P))$ belongs to the range of $A$ whenever $\hat P_z > 0$ for all $z\in \mathcal Z$ -- i.e., whenever $\hat \beta_n$ is well defined. We note that Assumption (ref)(ii) is not imposed in fang2023inference, but we deploy it here due to its wide applicability and its importance in obtaining simple primitive conditions for establishing the validity of our test. In applications in which Assumption (ref)(ii) is violated, it is still possible to conduct inference by using the results in fang2023inference. Finally, Assumption (ref)(iii) represents our main regularity condition on the set $\mathbf P$. Assumption (ref)(iii) requires that the probability of each support point $(y,d,z)$ be proportional to each other. We note that this restriction implies that no conditional probability is dominant, in the sense that $P_{yd|z}$ is bounded away from one uniformly in the $(y,d,z)$ and $P\in \mathbf P$.

We next establish the validity of our test under Assumption (ref) and an additional regularity condition that we discuss after the theorem.

theoremSuppose Assumption (ref) holds, $\alpha \in (0, 0.5)$, $\log^3(n)|\mathcal M|/n = o(1)$, and $0 \leq \lambda_n \leq 1$ satisfies $\lambda_n\sqrt{\log(|\mathcal M|)} = o(1)$. Further suppose Assumption (ref) in the appendix holds. Then, the test in (ref) satisfies $$\limsup_{n \rightarrow \infty} \sup_{P \in \mathbf P_0} E_P[\phi_n] \leq \alpha.$$

Theorem (ref) establishes the validity of our test under weak conditions on the number of support points for $(Y,D,Z)$. The main rate requirement imposed by Theorem (ref) is that $|\mathcal M|/n$ tend to zero (up to logs). We view this rate condition as close to minimal, in the sense that it is necessary for the empirical distribution $\hat P_n$ to be consistent for $P$. Our rate restriction is in addition a significant improvement over fang2023inference, whose best attainable rate is that $|\mathcal M|^2/n $ tend to zero (up to logs). We improve on their results by employing the specific structure of our problem to derive better coupling rates for our test and bootstrap statistics. In particular, by relying on results in massart1989strong we are able to couple our test statistic to a Gaussian counterpart at a rate $r_n$ given by $$r_n \equiv \frac{\log^{3/2}(n)\sqrt{|\mathcal M|}}{\sqrt n}.$$ The bootstrap statistic is in turn shown to be coupled to a Gaussian distribution at a rate of $r_n^{1/2}$; see Lemma (ref) in the Appendix for additional details. We conjecture that the bootstrap statistic can in fact be coupled at a rate $r_n$ as well. However, we do not pursue this refinement because it does not impact the conditions on how $|\mathcal M|$ relates to $n$, which is our primary concern.

Showing weak convergence of a test statistic and its bootstrap counterpart to a common distribution is in general not sufficient for establishing consistency of the corresponding critical values. In high-dimensional settings, consistency of the critical values requires a condition termed anti-concentration by chernozhukov2015comparison. Intuitively, anti-concentration ensures that the c.d.f.\ of the bootstrap statistic is suitably well behaved at the quantile of interest. Assumption (ref), stated in the Appendix for ease of exposition, is a sufficient condition for anti-concentration to hold in our setting. We note that Assumption (ref) is automatically satisfied in an asymptotic framework in which $|\mathcal M|$ is fixed with the sample size. Alternatively, Assumption (ref) can be dispensed with by modifying our critical value to include a “correction factor” introduced by andrews2013inference. We opt to introduce Assumption (ref) instead, however, to avoid introducing an additional tuning parameter.

Finally, we emphasize that while we have focused our discussion on hypothesis testing, our results readily deliver confidence regions as well. In particular, through test inversion it is straightforward to show that

equation[equation omitted — 86 chars of source]

is a valid confidence region, in the sense that it covers the parameter of interest with asymptotic probability of $1-\alpha$ uniformly in $P\in \mathbf P$.

remarkAssumption (ref) imposes that response types take on certain values with probability zero. Our framework can easily be modified, however, to accommodate the restriction that response types take on certain values with at most (or at least) some pre-specified probability. This modification is useful for conducting sensitivity analysis, as in masten2020inference and kline2013sensitivity. For instance, in Example (ref), one may relax restriction (ref) so that it instead specifies that the left-hand side is at most some pre-specified amount $\epsilon > 0$, and explore how inferences on $\theta(Q)$ change as one varies $\epsilon$. In this way, we can, for example, generalize the analysis of noack2021sensitivity to multi-valued instruments and treatments.
remarkIn applications with closed-form expressions of the identified set, $\Theta_0(P,\mathbf Q) = [L(P), U(P)]$, the plug-in estimator $[L(\hat P_n),U(\hat P_n)]$ provides a natural approach for estimating the identified set. However, the plug-in bootstrap in general does not provide correct coverage. Formally, denote by $\ell(\alpha/2, P)$ the $\alpha/2$ quantile of the distribution of $L(\hat P_n)$ under $P$ and by $u(1 - \alpha/2,P)$ the $1 - \alpha/2$ quantile of the distribution of $U(\hat P_n)$ under $P$. Consider the confidence region given by the bootstrap analogue $[\ell(\alpha/2, \hat P_n), u(1 - \alpha/2,\hat P_n)]$. Because the functionals $L(P)$ and $U(P)$ generally involve the maximum and minimum over functionals of $P$, they are only directionally differentiable and not fully differentiable when the maximum and minimum are attained at multiple functionals. Hence, it follows from results in fang2019inference that the bootstrap quantiles $\ell(\alpha/2, \hat P_n)$ and $u(1 - \alpha/2, \hat P_n)$ are not consistent for the population quantiles $\ell(\alpha/2,P)$ and $u(1-\alpha/2,P)$ that justify the validity of the proposed confidence region.

Inference with General-valued Outcomes

Section (ref) proposed an inference framework with the support $\mathcal Y$ of $Y$ restricted to be finite. We now extend that framework by relaxing the restriction that $\mathcal Y$ be finite, in particular allowing for continuous outcomes. The treatment $D$ and instrument $Z$ remain discrete with supports $\mathcal D$ and $\mathcal Z$ satisfying $2\leq |\mathcal D|, |\mathcal Z| < \infty$.

Our approach essentially partitions the outcome space $\mathcal Y$ into a finite number of subsets $\{B_{\ell,L}\}_{\ell=1}^L$ and reduces the problem into a discrete one. The analysis mimics the discrete case studied in Theorem (ref), but replaces the parameter $\beta(P)$ by a vector $\beta_L(P)$ that is given by

equation[equation omitted — 162 chars of source]

where $\mathcal L \equiv \{1, \ldots, L\}$. In what follows, it is helpful to note that because $R_o \in \mathcal Y^{|\mathcal D|}$, any partition $\{B_{\ell,L}\}_{\ell = 1}^L$ induces a partition for the support of $R_o$ as well. Specifically, writing any $c_o\in \mathcal L^{|\mathcal D|}$ as $c_o \equiv (c_o(0),\ldots, c_o(|\mathcal D|-1))$, and setting $B_{c_o} \equiv B_{c_o(0),L}\times \ldots \times B_{c_o(|\mathcal D|-1)}$ we have that $\{B_{c_o}\}_{c_o\in \mathcal L^{|\mathcal D|}}$ forms a partition for $\mathcal Y^{|\mathcal D|}$. Our next assumption uses this notation in introducing the main regularity conditions we require:

assumption(i) $\mathcal Y$ is compact; (ii) $g$ is continuous; (iii) $\{B_{\ell,L}\}_{\ell=1}^L$ is a partition of $\mathcal Y$ satisfying $\max_{1\leq \ell\leq L} \text{diam}\{B_{\ell,L}\} = o(1)$ as $L\to \infty$; (iv) $\mathcal R_o^\prime(r_t) \equiv \{r_o\in \mathcal Y^{|\mathcal D|} : (r_o,r_t)\in \mathcal R^\prime\}$ is closed and $B_{c_o}\cap \mathcal R_o^\prime(r_t)$ is connected for all $c_o\in \mathcal L^{|\mathcal D|}$ and $r_t\in \mathcal D^{|\mathcal Z|}$.

Assumption (ref)(i) imposes that $\mathcal Y$ be closed and bounded, which ensures that the partition $\{B_{\ell,L}\}_{\ell=1}^L$ can be chosen so that each set $B_{\ell,L}$ is bounded. In Assumption (ref)(ii), we further restrict attention to the case in which $g$ is continuous. As we discuss in Remark (ref) below, some forms of discontinuity can be accommodated, though they can require specific choices of partitions $\{B_{\ell,L}\}_{\ell=1}^L$ that satisfy restrictions beyond those imposed by Assumption (ref). Finally, Assumptions (ref)(iii)-(iv) impose the main requirements on the partition $\{B_{\ell,L}\}_{\ell=1}^L$, which include that it become finer as $L$ becomes large. In applications in which $\mathcal R$ and $\mathcal R^\prime$ leave the outcome response type $R_o$ unrestricted, $\mathcal R^\prime_o(r_t) = \mathcal Y^{|\mathcal D|}$ and Assumption (ref)(iv) is satisfied whenever the sets $\{B_{\ell,L}\}_{\ell=1}^L$ are chosen to be convex.

The next result may be seen as a generalization of Theorem (ref).

theoremLet $\mathbf Q$ be the set of all distributions of $(R,Z)$ satisfying Assumptions (ref) and (ref) and $\mathbf P$ be the set of all distributions of $(Y,D,Z)$ satisfying $P \{Z = z\} > 0$ for every $z \in \mathcal Z$. Further suppose $g$ in the definition of $\theta(Q)$ in (ref) is bounded. Then, for $\beta_L(P)$ defined in (ref) and a matrix $A_L$ defined in the beginning of Appendix (ref) that depends only on $\mathcal R$ in Assumption (ref), $\theta_0$ in (ref), and $\mathcal R'$ in the definition of $\theta(Q)$ in (ref), it follows that \begin{equation} \mathbf P_0 \subseteq \{P \in \mathbf P : A_Lx = \beta_L(P) for some x \geq 0 \}\equiv \mathbf P_{0,L} . \end{equation} If in addition Assumption (ref) holds and $\mathbf P_0$ is not empty, then $$\limsup_{L \to \infty} \mathbf P_{0,L} \equiv \bigcap_{M=1}^\infty \bigcup_{L\geq M} \mathbf P_{0,L} \subseteq {\rm cl}_{\mathbf P}(\mathbf P_0)~,$$ where ${\rm cl}_{\mathbf P}(\mathbf P_0)$ denotes the closure of $\mathbf P_0$ in $\mathbf P$ under weak convergence.

The first part of Theorem (ref) shows that $\mathbf P_0 \subseteq \mathbf P_{0,L}$ for $\mathbf P_{0,L}$ as defined in (ref). Therefore, any test of the null hypothesis that $P\in \mathbf P_{0,L}$ is also a valid test for the null hypothesis that $P\in \mathbf P_0$. Importantly, this conclusion does not depend on any of the regularity conditions imposed in Assumption (ref). The second part of Theorem (ref) shows, under Assumption (ref), that $\limsup_{L \to \infty} \mathbf P_{0,L}$ is included in the closure of $\mathbf P_0$. The second part of Theorem (ref) therefore shows that, as our discretization becomes finer, we are able to detect whether $P$ belongs to $\mathbf P_0$ or not (up to closure).

The main implication of Theorem (ref) is that, with a continuous outcome, we still can test the null hypothesis of interest by testing whether $P$ is such that $A_Lx = \beta_L(P)$ for some $x \geq 0$. The latter null hypothesis can be tested by using an analogous approach to that employed in (ref) for the finite support case. The asymptotic validity of the resulting test holds under similar assumptions to those required in Theorem (ref), but applied with ${\mathcal M}_L \equiv \mathcal L \times \mathcal D \times \mathcal Z$ instead of $\mathcal M$. This requires, among other assumptions, that $| {\mathcal M}_L|/n$ tend to zero (up to logs). In particular, if $\mathcal D$ and $\mathcal Z$ are fixed, then the condition that $|\mathcal M_L|/n$ tend to zero is equivalent to requiring that $L/n$ tending to zero. The partition $\{B_{\ell,L}\}_{\ell=1}^L$ is therefore allowed to quickly become finer with the sample size which, in light of Theorem (ref), can be desirable for power considerations.

remarkAn important example of a parameter $g$ that corresponds to a discontinuous $g$, and hence violates Assumption (ref)(i), is the probability of an event conditional on generalized principal strata. However, in this case, $g$ is piecewise constant and the conclusion of the second part of Theorem (ref) continues to hold if we redefine $A_L$ suitably. To see why, note by an inspection of the proof of Theorem (ref) that a solution to the linear system in (ref) is a vector of probabilities the cells determined by the intersections of $B_{c_o} \times \{r_t\}$ and $\mathcal R'$. If $\mathcal Y^{|\mathcal D|} = \bigcup_{1 \leq j \leq K} \mathcal S_j$ and $g$ is constant on $\mathcal S_j$ for each $j$, then we redefine $A_L$ so that the solution to the linear system is instead a vector of probabilities on each $B_{c_o} \times \{r_t\}$ and $\mathcal R'$ with $\mathcal S_j$ for $1 \leq j \leq K$.

By arguing as in Section (ref), Theorem (ref) immediately permits us to test whether the model is correctly specified in the sense that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$. Recall that by setting $g \equiv 1$, $\theta_0 = 1$ and $\mathcal R' = \mathcal R$, we have that $\mathbf P_0 = \{P \in \mathbf P : \mathbf Q_0(P,\mathbf Q) \neq \emptyset \}$. In this case, the equalities in the last two rows of $A_Lx = \beta_L(P)$ in (ref) are always true, so they can be omitted. In order to state the result, denote by $\tilde A_L$ and $\tilde \beta_L(P)$ all but the last two rows of $A_L$ and $\beta_L(P)$, respectively, in Theorem (ref).

corollaryLet $\mathbf Q$ be the set of all distributions of $(R,Z)$ satisfying Assumptions (ref) and (ref) and $\mathbf P$ be the set of all distributions of $(Y,D,Z)$ satisfying $P \{Z = z\} > 0$ for every $z \in \mathcal Z$. For $g \equiv 1$, $\theta_0 = 1$ and $\mathcal R' = \mathcal R$, $$\mathbf P_0 = \{P \in \mathbf P : \mathbf Q_0(P,\mathbf Q) \neq \emptyset \} \subseteq \{P \in \mathbf P : \tilde A_L x = \tilde \beta_L(P)\} \equiv \tilde {\mathbf P}_{0,L}~.$$ If in addition Assumption (ref) holds and $\mathbf P_0$ is not empty, then $\limsup_{L \to \infty} \tilde {\mathbf P}_{0,L} \equiv \\\bigcap_{M=1}^\infty \bigcup_{L\geq M} \tilde {\mathbf P}_{0,L}$ satisfies $\limsup_{L\to \infty} \tilde {\mathbf P}_{0,L} \subseteq {\rm cl}_{\mathbf P}(\mathbf P_0)$ where ${\rm cl}_{\mathbf P}(\mathbf P_0)$ denotes the closure of $\mathbf P_0$ in $\mathbf P$ under weak convergence.
remarkAs in Remark (ref), when $Q \{R \in \mathcal R'\}$ is known or identified for each $L \geq 1$, it is possible to use the characterization of $\mathbf P_0$ in Theorem (ref) to derive closed-form expressions $L_L(P)$ and $U_L(P)$ such that $\Theta_0(P,\mathbf Q) \subseteq [L_L(P), U_L(P)]$ for all $P \in \mathbf P$ for which $\mathbf Q_0(P,\mathbf Q) \neq \emptyset$. The inclusion $\Theta_0(P,\mathbf Q) \subseteq [L_L(P), U_L(P)]$ is generally strict for each fixed $L$ because of the loss of information from discretization. Therefore, as opposed to settings in which the support of $Y$ is discrete, for a fixed $L$ we typically obtain closed-form expressions for an outer set of the identified set instead of the identified set itself. The second part of Theorem (ref) implies, however, that under appropriate assumptions $\limsup_{L \to \infty} [L_L(P), U_L(P)] = \Theta_0(P,\mathbf Q)$. Similarly, as in Remark (ref), for each $L \geq 1$, it is possible to use Corollary (ref) to obtain a set of analytical inequalities $a_L(P) \leq 0$, such that $P \in \mathbf P$ satisfies $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$ if and only if $a_L(P) \leq 0$ for every $L \geq 1$.

Simulations

In this section we present simulation results that illustrate the finite-sample performance of our proposed inference methods. Throughout this section, $\mathcal D = \mathcal Z = \{0, 1, 2\}$. We set $\mathcal Y$ to be discrete in Section (ref) and an interval in Section (ref). The parameters of interest are \[ \theta_{\mathrm{ATE}}(Q) = E_Q[Y(2) - Y(1) \mid R \in \mathcal R'] \quad \mathrm{and} \quad \theta_{\mathrm{Prob}}(Q) = Q\{Y(2) \geq Y(1) \mid R \in \mathcal R'\}~, \] where $\mathcal R' = \{(y(0),y(1),y(2),d(0),d(1),d(2)): (d(0),d(1),d(2)) = (0,1,2)\}$. We consider the following two specifications of $\mathcal R$, corresponding to two different models:

enumerate[] • {\it One-sided noncompliance}: $\mathcal R$ is as specified in Example (ref) and denoted by $\mathcal R_{\mathrm{1s}}$, with the corresponding model $\mathbf Q_{\mathrm{1s}} = \{Q: Q \text{ satisfies Assumptions \ref{as:exog} and \ref{ass:generalizedstrata} with } \mathcal R = \mathcal R_{\mathrm{1s}}\}$; • {\it Encouragement design}: $\mathcal R$ is as specified in Example (ref) and denoted by $\mathcal R_{\mathrm{enc}}$, with the corresponding model $\mathbf Q_{\mathrm{enc}} = \{Q: Q \text{ satisfies Assumptions \ref{as:exog} and \ref{ass:generalizedstrata} with } \mathcal R = \mathcal R_{\mathrm{enc}}\}$.

Discrete Outcomes

We first present simulation results for a discrete outcome with $\mathcal Y = \{-1, 0, 1\}$. From each of $\mathbf Q_{\mathrm{1s}}$ and $\mathbf Q_{\mathrm{enc}}$, we pick a $Q$ distribution and fix them. The two selected $Q$'s are described in Appendices (ref) and (ref), respectively. For each $Q$, we set $Q\{Z=z\} = 1/3$ for each $z \in \mathcal Z$ and define $P = Q T^{-1}$. The simulation goes as follows. For each $Q$, we repeatedly draw an i.i.d.\ sample $\{V_i\}_{i = 1}^n$ of size $n = 3000$ or $12000$ from $P$, where $V_i = (Y_i, D_i, Z_i)$. For each sample, we obtain the confidence interval for $\theta_{\mathrm{ATE}}$ and $\theta_{\mathrm{Prob}}$ by inverting the test of the feasibility of the linear system in (ref) at the 5% level.

The results are reported in Tables (ref) and (ref) for $\mathbf Q_{\mathrm{1s}}$ and $\mathbf Q_{\mathrm{enc}}$ respectively. For each parameter and each model, we report the length and the coverage rate of the confidence interval in the last two columns, averaged over 3000 replications. We also report the length of the identified set $|\Theta_0(P, \mathbf{Q}_{\rm 1s})|$ or $|\Theta_0(P, \mathbf{Q}_{\rm enc})|$. We note that the confidence interval covers the parameter with probability one in all cases. This phenomenon is to be expected because in each case the true value of the parameter lies strictly in the interior of the identified set. For example, $\Theta_0(P,\mathbf Q_{1s}) = [0.1482, 1.0295]$ and $\theta_{\rm ATE}(Q) = 0.4919$. In this example, for $n = 3000$, the coverage probabilities of the lower and upper endpoints of the identified set are 0.947 and 0.956, respectively.

table[table omitted — 822 chars of source]
table[table omitted — 827 chars of source]

General-valued Outcomes

We now present simulation results for a continuous outcome. In this subsection, $\mathcal D = \mathcal Z = \{0,1,2\}$ and $\mathcal Y = [-5,5]$. We focus on $\mathcal R = \mathcal R_{\mathrm{1s}}$ and $\theta_{\mathrm{ATE}}(Q) = E_Q[Y(2) - Y(1) \mid R \in \mathcal R']$. As in Section (ref), $Q\{Z = z\} = 1/3$ for each $z \in \mathcal Z$. A distribution $Q \in \mathbf Q_{\mathrm{1s}}$ is generated as follows. With a slight abuse of notation, let $\tilde Q$ on $\{-1,0,1\}^3 \times \mathcal D^3$ be the discrete distribution in Appendix (ref). We then define $Q$ as a mixture of truncated normal distributions where the mixture weights are given by $\tilde Q$. In particular, let $\Phi_{\mathcal Y}(\mu(d))$ denote the truncation of normal distribution $N(\mu(d), 1)$ on $\mathcal Y$. Then,

equation[equation omitted — 315 chars of source]

In words, for each mixture component, the value of $D(z)$ is fixed at $d(z)$ for $z \in \mathcal Z$ and independently across $d \in \mathcal D$, $Y(d)$ has a truncated normal distribution. Finally let $P = QT^{-1}$. As in Section (ref), we repeatedly draw an i.i.d.\ sample $\{V_i\}_{i = 1}^n$ of size $n = 3000$ or $12000$ from $P$, where $V_i = (Y_i, D_i, Z_i)$, and for each sample obtain confidence intervals by inverting the test of the feasibility of the linear system in (ref) at the 5% level. We consider the following three partitions of $\mathcal Y$:

enumerate[] • {\tt bin1}: $\mathcal Y = [-5, 0) \cup [0, 5]$; • {\tt bin2}: $\mathcal Y = [-5,-3) \cup [-3, -1) \cup \cdots \cup [3,5]$; • {\tt bin3}: $\mathcal Y = [-5, -4) \cup [-4,-3) \cup \cdots \cup [4,5]$.
table[table omitted — 1,019 chars of source]

The results are reported in Table (ref). For each level of discretization, we report the length and the coverage rate of the confidence intervals in the last two columns, averaged over 3000 replications. To illustrate the computational costs, we report in the second column of Table (ref) the number of rows and columns of the $A$ matrix in (ref) for each level of discretization, which correspond to the number of constraints and the number of decision variables in each of the linear systems being tested. Next, recall Theorem (ref) shows that for each level of discretization $L$, $\mathbf P_{0, L}$ contains $\mathbf P_0$. Let $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$ be the set of $\theta_0 \in \mathbf R$ such that $\mathbf P_{0, L}$ is nonempty. Theorem (ref) implies that $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$ is a superset of the identified set $\Theta_0(P,\mathbf Q_{\mathrm{1s}})$. In the third column of Table (ref), we display the length of $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$ for each $L$. A comparison of the length of our confidence intervals to the length of $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$ reveals how much of the length of the confidence interval is attributed to sampling uncertainty.

We note the following findings from Table (ref). First, as in Tables (ref) and (ref), the confidence intervals cover the parameter with probability one for each level of discretization we consider. This phenomenon is again to be expected because in each case the true value of the parameter lies strictly in the interior of $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$; coverage probabilities of the lower and upper endpoints of $\Theta_{0,L}(P, \mathbf Q_{1s})$ are approximately equal to the nominal level. Next, as the number of intervals in the partition increases, the dimensionality of the problem grows exponentially. Specifically, the number of constraints and decision variables scale at the rates $O(|\mathcal L||\mathcal D||\mathcal Z|)$ and $O(|\mathcal L|^{|\mathcal D|} |\mathcal R_t|)$, respectively. For a fixed sample size, a finer partition yields a shorter confidence interval, although we note it comes at a higher computational cost.