Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,784 characters · 8 sections · 84 citation commands
Inference for Treatment Effects Conditional on Generalized Principal Strata using Instrumental Variables
\thispagestyle{empty} \setcounter{page}{1}
We propose a general approach to inference for a broad class of models that arise in the analysis of treatment effects with discrete-valued treatments and instruments and a general-valued outcome. In addition to instrument exogeneity, the main substantive restriction imposed in our analysis is that certain values for the response types occur with probability zero. Here, the response type refers to the vector of potential outcomes and potential treatments, and we refer to a set of possible values for the response type as a generalized principal stratum. Through a series of examples, we show that our framework encompasses a wide variety of assumptions that have been considered in the previous literature, including the following: (i) restrictions in the analysis of randomized controlled trials (RCTs) under noncompliance considered, e.g., in imbens1994identification and cheng2006bounds; (ii) generalizations of these restrictions considered in bai2024identifying; (iii) revealed preference-type restrictions on response types considered in kirkeboen2016field, kline2016evaluating, and heckman2018unordered; and (iv) restrictions on the ordering of potential treatments or on the ordering of potential outcomes considered in manski1997monotone, manski1998monotone, and machado2019instrumental.\footnote{In the context of mediation analysis, kwon2024testing show that “full mediation” is equivalent to certain assumptions like those we consider.}
Our framework allows inference on any treatment effect parameter that can be expressed as the expectation of a function of the response type conditional on a generalized principal stratum. In this way, our framework accommodates many parameters that have been considered previously in the literature, including both average and distributional treatment effect parameters, such as the probability of being strictly helped and the probability of not being hurt by the treatment. It further accommodates versions of these parameters conditional on sets of possible values of potential treatments, such as the local average treatment effect in imbens1994identification, average effects conditional on principal strata in frangakis2002principal, and average effects conditional on sets of possible values of potential outcomes, such as the parameters considered in heckman1997making and heckman1998evaluating. By contrast, each of the papers cited in the preceding paragraph tailors their method to specific examples of such treatment effect parameters. Furthermore, because the types of restrictions considered in these papers are typically insufficient for identification of many parameters of interest, some of these papers focus specifically on those treatment effect parameters that are identified from the two-stage least squares estimand. A characterization of the parameters that are identified by such restrictions is given by navjeevan2023identification and goff2025doesividentificationrestrict.
A key result in our analysis is a characterization of the identified set for such parameters under these assumptions and the testable restrictions for the assumptions themselves in terms of existence of a non-negative solution to linear systems of equations with a special structure. Using this result, we build on results in fang2023inference to develop a test for the null hypothesis that a pre-specified value for the parameter of interest lies in the identified set as well as the null hypothesis that the testable restrictions are satisfied. Importantly, the resulting tests remain well behaved even in “high-dimensional” settings, meaning it is uniformly consistent in level over a large class of distributions satisfying weak assumptions and permitting, among other things, the support of the treatment and instrument or the support of the response types to be “large” relative to the sample size. Through test inversion, we may also construct confidence regions for such parameters that are uniformly consistent in level. We highlight that while our test builds on fang2023inference, it leverages the structure of our problem to obtain novel strong approximations that significantly improve on the coupling rates of fang2023inference. These coupling results for high-dimensional vectors of estimates of conditional probabilities and their bootstrap counterparts may be of independent interest.
Other methods for inference have been proposed in the prior literature for some special cases of our framework. For inference on certain treatment effect parameters, see, e.g., cheng2006bounds, bhattacharya2008treatment,bhattacharya2012treatment, and machado2019instrumental; for inference on the validity of the assumptions, see, e.g., kitagawa2015test, heckman2018unordered, sun2023instrument, and bai2025sharp. These approaches rely upon closed-form expressions for the identified set for the parameter of interest or testable inequalities, which must be proved or computed on a case-by-case basis. When such expressions need to be computed, the methods are feasible only when the support of response types is small. In contrast, our approach does not rely on such closed-form expressions and remains computationally feasible even in high-dimensional settings. A key strength of our approach is its broad applicability, including in settings for which no existing methods apply.
The remainder of the paper is organized as follows. Section (ref) introduces our formal setup and notation. In Section (ref), we provide several examples of parameters and restrictions previously considered in the literature that are nested by our framework. We present our inference method when the support of the outcome variable is restricted to be finite in Section (ref), and without such a restriction in Section (ref). We illustrate our method through a simulation study in Section (ref). Proofs of all results can be found in Appendix (ref).
Denote by $Y \in \mathcal Y$ an outcome, by $D \in \mathcal D$ a discrete valued endogenous regressor (i.e., the treatment), and by $Z \in \mathcal Z$ a discrete valued instrumental variable. To rule out degenerate cases, we assume throughout that $2 \leq |\mathcal Y|$, $2 \leq |\mathcal D| < \infty$, and $2 \leq |\mathcal Z| < \infty$. We emphasize that neither $\mathcal D$ and $\mathcal Z$ needs to be a subset of the real line; in this way, our framework accommodates vector-valued treatments and instruments provided they only take a finite number of values. In what follows, we will consider both settings in which $|\mathcal Y|$ is finite and those in which it is not. Further denote by $Y(d) \in \mathcal Y$ the potential outcome if $D = d \in \mathcal D$ and by $D(z) \in \mathcal D$ the potential treatment if $Z = z \in \mathcal Z$. As usual, we assume that
In our discussion below, it will be convenient to define $R_o \equiv (Y(d) : d \in \mathcal D)$ and $R_t \equiv (D(z) : z \in \mathcal Z)$, and set $R \equiv (R_o,R_t)$. Following heckman2018unordered, we will refer to $R_t$ as the “treatment response type.” By analogy with this terminology, we will also refer to $R_o$ as the “outcome response type” and to $R$ as simply the “response type.” We let $Q$ denote the distribution of $(R_o,R_t,Z)$ and note that by (ref) we have
for a transformation $T$ implicitly defined through (ref). Letting $P$ denote the distribution of $(Y, D, Z)$, we therefore have that $P = QT^{-1}$.
Below we will require that $Q \in \mathbf Q$, where $\mathbf Q$ is a class of distributions satisfying assumptions that we will specify. Different specifications of $\mathbf Q$ represent different assumptions that we impose on the distribution of potential outcomes and potential treatments. In this sense, $\mathbf Q$ may be viewed as a model for potential outcomes and potential treatments.
Given a distribution $P$ of $(Y,D,Z)$ and a model $\mathbf Q$, we define the set of $Q \in \mathbf Q$ that can rationalize $P$ as \[ \mathbf Q_0(P, \mathbf Q) = \{Q \in \mathbf Q: P = Q T^{-1}\}~. \] We say $\mathbf Q$ is consistent with $P$ if and only if $\mathbf Q_0(P, \mathbf Q ) \ne \emptyset$. For every model $\mathbf Q$ considered in this paper, every $Q \in \mathbf Q$ is assumed to satisfy the restriction:
Our remaining restrictions on $\mathbf Q$ will be formulated in terms of restrictions on possible values of the response type $R$. These restrictions will be expressed by specifying a set $\mathcal R \subseteq \mathcal Y^{|\mathcal D|} \times \mathcal D^{|\mathcal Z|}$ characterizing the possible values of $R$. In Section (ref), we provide several different choices of $\mathcal R$ that have been previously considered in the literature. For a given choice of $\mathcal R$, we will impose the following on every $Q \in \mathbf Q$:
Following frangakis2002principal, sets of the form $\{R_t = r_t\}$ for $r_t$ a possible value of $R_t$ are referred to as “principal strata.” By analogy with this terminology, we will refer to sets of the form $\{R \in \mathcal R'\}$ for $\mathcal R' \subseteq \mathcal R$ as “generalized principal strata.”
Finally, we define our parameters of interest. The parameters we consider can be written as
for different choices of function $g : \mathcal R \rightarrow \mathbf R$ and generalized principal strata $\mathcal R' \subseteq \mathcal R$. Note that for $\theta(Q)$ in (ref) to be well defined we require $Q$ to satisfy $Q\{R \in \mathcal R'\} > 0$. A wide variety of parameters can be accommodated in this way. In Section (ref), we will show specific parameters that have been considered previously in the literature have the structure in (ref). We note now, however, that natural choices of $g$ correspond to the effect of one treatment versus another (i.e., $g(R) = Y(d) - Y(d')$ for $d \in \mathcal D$ and $d' \in \mathcal D$) and the probability that one treatment leads to a larger outcome than another treatment (i.e., $g(R) = I\{Y(d) > Y(d')\}$ for $d \in \mathcal D$ and $d' \in \mathcal D$). We further note that when $\mathcal R' = \mathcal R$, $\theta(Q)$ defined in (ref) simplifies to $E_Q[g(R)]$.
For a given distribution $P$ and model $\mathbf Q$, note that the identified set for $\theta(Q)$ under $P$ relative to $\mathbf{Q}$ is
The set $\Theta_0(P, \mathbf Q) $ is nonempty whenever there exists at least one $Q \in \mathbf Q_0(P, \mathbf Q)$ with $Q\{R \in \mathcal R'\} > 0$. By construction, this set is “sharp” in the sense that for any value in the set there exists a distribution $Q$ that is consistent with $P$, satisfies the restrictions of the model, and for which $\theta(Q)$ equals the prescribed value.
In this section, we show how to accommodate several examples from the previous literature in our framework. Our discussion focuses in particular on Assumption (ref) and the specification of $\mathcal R$ in Assumption (ref), but, where the cited literature has emphasized specific parameters of interest, we additionally describe how those parameters of interest can be expressed as (ref) for suitable choices of $g$ and $\mathcal R'$.
In many of the examples discussed above, $Y$ may be an ordinal outcome. In such instances, average treatment effects may not be interpretable. Nonetheless, researchers may consider other parameters that fit our framework, such as: (i) $Q\{Y(j) > Y(k) \mid R \in \mathcal R' \}$, the conditional probability of benefit of treatment $j$ versus treatment $k$; (ii) $Q\{Y(j) \ge Y(k) \mid R \in \mathcal R' \}$, the conditional probability of no harm of treatment $j$ versus treatment $k$; or (iii) $Q\{Y(j) > Y(k) \mid R \in \mathcal R'\} - Q\{Y(k) > Y(j) \mid R \in \mathcal R'\}$, the conditional relative treatment effect of treatment $j$ versus $k$. Each of these parameters can be written as (ref) for appropriate choices of $g$. These parameters have been previously studied for the special case in which $\mathcal{R}'=\mathcal{R}$ and $|\mathcal{D}| = 2$ or $3$ by lu2018treatment, huang2019constructing, and gabriel2024sharp.
It is also instructive to discuss examples of restrictions that do not fit our framework. angrist1995two-stage, for instance, consider a generalization of the the monotonicity assumption in Example (ref) in which the direction of the monotonicity is not known a priori, i.e., they consider the restriction that, for any pair of values for the instrument $j,k \in \mathcal Z$, either $Q\{D(j) \geq D(k)\} = 1$ or $Q\{D(j) \leq D(k)\} = 1$. Such a restriction is also analyzed in vytlacil2006ordered and in heckman2018unordered, where the latter paper terms it “ordered monotonicity.” This restriction is equivalent to imposing $Q \{R \in \mathcal R_\pi\} = 1$ for some $\pi \in \Pi(|\mathcal Z|)$, where $\Pi(|\mathcal Z|)$ is the set of all permutations of $\{0, \dots, |\mathcal Z| - 1\}$ and $\mathcal R_\pi = \{(y(0), \dots, y(|\mathcal D| - 1), d(0), \dots ,d(|\mathcal D|-1 ) ) : d(\pi(|\mathcal Z| - 1)) \geq \ldots \geq d(\pi(0))\}$. A similar generalization of the assumption in Example (ref) can also be considered. In particular, machado2019instrumental study such a generalization with $|\mathcal D| = |\mathcal Z| = 2$. These models, in which the permutation $\pi$ is not known a priori, fall outside the class of models we consider.
In this section, we propose a test for conducting inference on the parameter of interest $\theta(Q)$. To this end, recall that $P$ denotes the distribution of $(Y,D,Z)$ and set $\mathcal M \equiv \mathcal Y \times \mathcal D \times \mathcal Z$. In this section only, we suppose $|\mathcal Y| < \infty$. Doing so allows us to introduce our method succinctly. This assumption will be removed in Section (ref), where we allow for a more general outcome through discretization. For $(y, d, z) \in \mathcal M$, define $P_{ydz} = P \{Y = y, D = d, Z = z\}$ and $P_z = P \{Z = z\}$. In what follows, we assume that $P\in \mathbf P$, where $\mathbf P$ is a “large” class of distributions that we will specify below. The class $\mathbf P$ can depend on the sample size $n$, but we suppress the dependence from our notation. We consider the problem of testing
at level $\alpha \in (0,1)$, where, for a pre-specified value of $\theta_0$, the null hypothesis we consider is given by
As we show below, a test of this null hypothesis can be used to construct confidence regions for $\theta(Q)$ through test inversion over the pre-specified value $\theta_0$.
The key insight underlying the construction of our test is the following theorem, which provides a convenient reformulation of $\mathbf P_0$ in terms of existence of a non-negative solution to a (possibly under-determined) system of linear equations in which the “coefficients” are known. The statement of the theorem involves the parameter
where $P_{yd|z} \equiv P\{Y = y, D = d | Z = z\}$. Using this notation, we obtain the following result:
The key implication of Theorem (ref) is that we may test the null hypothesis of interest by examining whether there exists a positive vector $x$ satisfying $Ax=\beta(P)$. Because $\theta(Q)$ need not always be well defined when $\mathcal R' \neq \mathcal R$, (ref) involves an inclusion rather than an equality. From a testing perspective, however, $\mathbf P_0$ and its closure are indistinguishable, in the sense that any level $\alpha$ test of whether $P$ belongs to $\mathbf P_0$ necessarily has power no larger than $\alpha$ against any $P$ in the closure $\text{cl}(\mathbf P_0)$ romano2004non. Therefore, the second part of Theorem (ref) may be interpreted as showing that testing whether $P\in \mathbf P_0$ is in fact equivalent to testing whether $\beta(P) = Ax$ for some $x\geq 0$. This conclusion holds under the condition that $\mathbf P_0$ is not empty. However, by definition, $\mathbf P_0$ is empty whenever there is no possible distribution of the data $P\in \mathbf P$ that is compatible with the null hypothesis. In other words, we may focus on testing whether $\beta(P) = Ax$ for some $x\geq 0$ unless the null hypothesis specifies restrictions that are incompatible with any $P$ and, in this sense, may be interpreted as being logically inconsistent with each other.
Theorem (ref) also has implications for deriving an analytical expression for the identified set of $\theta(Q)$. We discuss these implications in the next two remarks, and elaborate on them in the appendix.
Theorem (ref) immediately permits us to test whether the model is correctly specified in the sense that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$. To see this, consider $g \equiv 1$, $\theta_0 = 1$, and $\mathcal R' = \mathcal R$. For such values of these quantities,
Therefore, $\mathbf P_0 = \{P \in \mathbf P : 1 \in \Theta_0(P,\mathbf Q)\} = \{P \in \mathbf P : \mathbf Q_0(P,\mathbf Q) \neq \emptyset \}$. In this case, the equality in the last row of $Ax = \beta(P)$ in (ref) is always true, so it can be omitted. In order to state the result formally, denote by $\tilde A$ and $\tilde \beta(P)$ all but the last rows of $A$ and $\beta(P)$, respectively, in Theorem (ref).
Because $\theta(Q)$ is always well defined when $\mathcal R' = \mathcal R$, (ref) in Corollary (ref) differs from its counterpart (ref) in Theorem (ref) in that it involves an equality instead of an inclusion.
By Theorem (ref), we may test the null hypothesis of interest by testing whether $P$ is such that
In recent work, fang2023inference derived a test for whether $P$ satisfies the restriction in (ref) for a general class of problems; see also bai2022testing for other approaches. In what follows, we build on fang2023inference by employing the structure of our problem to sharpen rate conditions and establish the validity of their test under weaker requirements. Following Corollary (ref), it is straightforward to modify the test described below to test the null hypothesis that the model is correctly specified in the sense that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$.
In order to describe the test statistic, we require some additional notation. We will assume the availability of an i.i.d.\ sample $\{V_i\}_{i=1}^n$ with $V_i=(Y_i,D_i,Z_i)$, $V_i \sim P$, and let $\hat P_n$ denote the empirical distribution corresponding to the sample $\{V_i\}_{i=1}^n$. We further set $\hat \beta_n \equiv \beta(\hat P_n)$ to denote the plug-in estimator for the parameter $\beta(P)$ and let $\hat \Omega \equiv \Omega(\hat P_n)$, where $\Omega(P)$ is a diagonal weighting matrix whose $(i,i)$-entry equals the asymptotic standard deviation of the $i^{th}$ row of $\sqrt n(\hat \beta_n - \beta(P))$. Specifically, the first $|\mathcal M|$ diagonal entries of $\Omega(P)$ equal $$\left(\frac{P_{yd|z}(1-P_{yd|z})}{P_z}\right)^{1/2}$$ for some $(y,d,z)\in \mathcal M$, and the final two diagonal entries of $\Omega(P)$ are equal to zero. Finally, we let $A^\dagger$ denote the Moore-Penrose pseudoinverse of $A$.
Given the introduced notation, we define the set $\hat{\mathcal V}_n \equiv \{s : A^\dagger s \leq 0, \|\hat \Omega_n (AA')^\dagger s\|_1 \leq 1\}$ and set
as our test-statistic -- here, $\|a\|_1 = \sum_{i=1}^d |a_i|$ for any vector $a = (a_1,\ldots, a_d) \in \mathbf R^d$ and $\langle a,b\rangle = a^\prime b$. From results in fang2019inference, $T_n$ represents a sample analogue to an equivalent formulation of the restriction in (ref) obtained from Farkas' lemma. In particular, the population analogue to $T_n$ equals zero whenever (ref) indeed holds and, as a result, under the null hypothesis $T_n$ should not be “too large.” We highlight that the test statistic can be computed through linear programming. As a result, $T_n$ can be reliably and quickly computed even in high-dimensional applications.
Unfortunately, the asymptotic distribution of $T_n$ is not pivotal. In particular, its asymptotic distribution depends on the “directions” $s \in \hat {\mathcal V}_n$ at which the population analogue $\langle A^\dagger s, A^\dagger \beta(P)\rangle$ is “close” to zero. In order to obtain a valid critical value, we therefore rely on the nonparametric bootstrap and a construction that asymptotically excludes directions $s$ that do not play a role in the distribution of $T_n$; i.e., directions $s$ for which $\langle A^\dagger s, A^\dagger \beta(P)\rangle $ is “far” from zero. Specifically, for the latter purpose we introduce
which represents an estimator for $\beta(P)$ that is restricted to satisfy the null hypothesis. Letting $\{V_i^*\}_{i=1}^n$ denote a bootstrap sample, (i.e., a sample drawn i.i.d.\ from $\hat P_n$), and $\hat \beta_n^*$ the bootstrap analogue to $\hat \beta_n$, we set
as our bootstrap statistic. Here, $\lambda_n \in [0,1]$ represents a penalty satisfying $\lambda_n =o(1)$, and whose role is to ensure that directions $s$ that do not impact the finite-sample distribution of our test statistic do not impact the bootstrap statistic either. We again highlight that, just like $T_n$, the bootstrap statistic $T_n^*$ can also be computed through linear programming.
The critical value for a level $\alpha$ test is then simply the $1-\alpha$ quantile of the distribution of $T_n^*$ conditionally on the sample $\{V_i\}_{i=1}^n$. Formally, the critical value for our test is given by
As usual, the critical value $\hat c_n(1-\alpha)$ can be computed through simulation by drawing multiple samples $\{V_i^*\}_{i=1}^n$ i.i.d.\ from $\hat P_n$ and computing the $1-\alpha$ quantile of the corresponding sample of bootstrap statistics. Given the critical value, we next formally introduce our test, which rejects whenever $T_n$ exceeds $\hat c_n(1-\alpha)$:
We establish the asymptotic validity of our test under the following main assumption.
Assumption (ref)(i) simply formalizes our requirement that the sample be independent and identically distributed. In Assumption (ref)(ii) we impose the additional condition that $(\hat \beta_n - \beta(P))$ belong to the range of the matrix $A$ with probability tending to one. This requirement is satisfied in our examples, where $(\hat \beta_n - \beta(P))$ belongs to the range of $A$ whenever $\hat P_z > 0$ for all $z\in \mathcal Z$ -- i.e., whenever $\hat \beta_n$ is well defined. We note that Assumption (ref)(ii) is not imposed in fang2023inference, but we deploy it here due to its wide applicability and its importance in obtaining simple primitive conditions for establishing the validity of our test. In applications in which Assumption (ref)(ii) is violated, it is still possible to conduct inference by using the results in fang2023inference. Finally, Assumption (ref)(iii) represents our main regularity condition on the set $\mathbf P$. Assumption (ref)(iii) requires that the probability of each support point $(y,d,z)$ be proportional to each other. We note that this restriction implies that no conditional probability is dominant, in the sense that $P_{yd|z}$ is bounded away from one uniformly in the $(y,d,z)$ and $P\in \mathbf P$.
We next establish the validity of our test under Assumption (ref) and an additional regularity condition that we discuss after the theorem.
Theorem (ref) establishes the validity of our test under weak conditions on the number of support points for $(Y,D,Z)$. The main rate requirement imposed by Theorem (ref) is that $|\mathcal M|/n$ tend to zero (up to logs). We view this rate condition as close to minimal, in the sense that it is necessary for the empirical distribution $\hat P_n$ to be consistent for $P$. Our rate restriction is in addition a significant improvement over fang2023inference, whose best attainable rate is that $|\mathcal M|^2/n $ tend to zero (up to logs). We improve on their results by employing the specific structure of our problem to derive better coupling rates for our test and bootstrap statistics. In particular, by relying on results in massart1989strong we are able to couple our test statistic to a Gaussian counterpart at a rate $r_n$ given by $$r_n \equiv \frac{\log^{3/2}(n)\sqrt{|\mathcal M|}}{\sqrt n}.$$ The bootstrap statistic is in turn shown to be coupled to a Gaussian distribution at a rate of $r_n^{1/2}$; see Lemma (ref) in the Appendix for additional details. We conjecture that the bootstrap statistic can in fact be coupled at a rate $r_n$ as well. However, we do not pursue this refinement because it does not impact the conditions on how $|\mathcal M|$ relates to $n$, which is our primary concern.
Showing weak convergence of a test statistic and its bootstrap counterpart to a common distribution is in general not sufficient for establishing consistency of the corresponding critical values. In high-dimensional settings, consistency of the critical values requires a condition termed anti-concentration by chernozhukov2015comparison. Intuitively, anti-concentration ensures that the c.d.f.\ of the bootstrap statistic is suitably well behaved at the quantile of interest. Assumption (ref), stated in the Appendix for ease of exposition, is a sufficient condition for anti-concentration to hold in our setting. We note that Assumption (ref) is automatically satisfied in an asymptotic framework in which $|\mathcal M|$ is fixed with the sample size. Alternatively, Assumption (ref) can be dispensed with by modifying our critical value to include a “correction factor” introduced by andrews2013inference. We opt to introduce Assumption (ref) instead, however, to avoid introducing an additional tuning parameter.
Finally, we emphasize that while we have focused our discussion on hypothesis testing, our results readily deliver confidence regions as well. In particular, through test inversion it is straightforward to show that
is a valid confidence region, in the sense that it covers the parameter of interest with asymptotic probability of $1-\alpha$ uniformly in $P\in \mathbf P$.
Section (ref) proposed an inference framework with the support $\mathcal Y$ of $Y$ restricted to be finite. We now extend that framework by relaxing the restriction that $\mathcal Y$ be finite, in particular allowing for continuous outcomes. The treatment $D$ and instrument $Z$ remain discrete with supports $\mathcal D$ and $\mathcal Z$ satisfying $2\leq |\mathcal D|, |\mathcal Z| < \infty$.
Our approach essentially partitions the outcome space $\mathcal Y$ into a finite number of subsets $\{B_{\ell,L}\}_{\ell=1}^L$ and reduces the problem into a discrete one. The analysis mimics the discrete case studied in Theorem (ref), but replaces the parameter $\beta(P)$ by a vector $\beta_L(P)$ that is given by
where $\mathcal L \equiv \{1, \ldots, L\}$. In what follows, it is helpful to note that because $R_o \in \mathcal Y^{|\mathcal D|}$, any partition $\{B_{\ell,L}\}_{\ell = 1}^L$ induces a partition for the support of $R_o$ as well. Specifically, writing any $c_o\in \mathcal L^{|\mathcal D|}$ as $c_o \equiv (c_o(0),\ldots, c_o(|\mathcal D|-1))$, and setting $B_{c_o} \equiv B_{c_o(0),L}\times \ldots \times B_{c_o(|\mathcal D|-1)}$ we have that $\{B_{c_o}\}_{c_o\in \mathcal L^{|\mathcal D|}}$ forms a partition for $\mathcal Y^{|\mathcal D|}$. Our next assumption uses this notation in introducing the main regularity conditions we require:
Assumption (ref)(i) imposes that $\mathcal Y$ be closed and bounded, which ensures that the partition $\{B_{\ell,L}\}_{\ell=1}^L$ can be chosen so that each set $B_{\ell,L}$ is bounded. In Assumption (ref)(ii), we further restrict attention to the case in which $g$ is continuous. As we discuss in Remark (ref) below, some forms of discontinuity can be accommodated, though they can require specific choices of partitions $\{B_{\ell,L}\}_{\ell=1}^L$ that satisfy restrictions beyond those imposed by Assumption (ref). Finally, Assumptions (ref)(iii)-(iv) impose the main requirements on the partition $\{B_{\ell,L}\}_{\ell=1}^L$, which include that it become finer as $L$ becomes large. In applications in which $\mathcal R$ and $\mathcal R^\prime$ leave the outcome response type $R_o$ unrestricted, $\mathcal R^\prime_o(r_t) = \mathcal Y^{|\mathcal D|}$ and Assumption (ref)(iv) is satisfied whenever the sets $\{B_{\ell,L}\}_{\ell=1}^L$ are chosen to be convex.
The next result may be seen as a generalization of Theorem (ref).
The first part of Theorem (ref) shows that $\mathbf P_0 \subseteq \mathbf P_{0,L}$ for $\mathbf P_{0,L}$ as defined in (ref). Therefore, any test of the null hypothesis that $P\in \mathbf P_{0,L}$ is also a valid test for the null hypothesis that $P\in \mathbf P_0$. Importantly, this conclusion does not depend on any of the regularity conditions imposed in Assumption (ref). The second part of Theorem (ref) shows, under Assumption (ref), that $\limsup_{L \to \infty} \mathbf P_{0,L}$ is included in the closure of $\mathbf P_0$. The second part of Theorem (ref) therefore shows that, as our discretization becomes finer, we are able to detect whether $P$ belongs to $\mathbf P_0$ or not (up to closure).
The main implication of Theorem (ref) is that, with a continuous outcome, we still can test the null hypothesis of interest by testing whether $P$ is such that $A_Lx = \beta_L(P)$ for some $x \geq 0$. The latter null hypothesis can be tested by using an analogous approach to that employed in (ref) for the finite support case. The asymptotic validity of the resulting test holds under similar assumptions to those required in Theorem (ref), but applied with ${\mathcal M}_L \equiv \mathcal L \times \mathcal D \times \mathcal Z$ instead of $\mathcal M$. This requires, among other assumptions, that $| {\mathcal M}_L|/n$ tend to zero (up to logs). In particular, if $\mathcal D$ and $\mathcal Z$ are fixed, then the condition that $|\mathcal M_L|/n$ tend to zero is equivalent to requiring that $L/n$ tending to zero. The partition $\{B_{\ell,L}\}_{\ell=1}^L$ is therefore allowed to quickly become finer with the sample size which, in light of Theorem (ref), can be desirable for power considerations.
By arguing as in Section (ref), Theorem (ref) immediately permits us to test whether the model is correctly specified in the sense that $\mathbf Q_0(P, \mathbf Q) \neq \emptyset$. Recall that by setting $g \equiv 1$, $\theta_0 = 1$ and $\mathcal R' = \mathcal R$, we have that $\mathbf P_0 = \{P \in \mathbf P : \mathbf Q_0(P,\mathbf Q) \neq \emptyset \}$. In this case, the equalities in the last two rows of $A_Lx = \beta_L(P)$ in (ref) are always true, so they can be omitted. In order to state the result, denote by $\tilde A_L$ and $\tilde \beta_L(P)$ all but the last two rows of $A_L$ and $\beta_L(P)$, respectively, in Theorem (ref).
In this section we present simulation results that illustrate the finite-sample performance of our proposed inference methods. Throughout this section, $\mathcal D = \mathcal Z = \{0, 1, 2\}$. We set $\mathcal Y$ to be discrete in Section (ref) and an interval in Section (ref). The parameters of interest are \[ \theta_{\mathrm{ATE}}(Q) = E_Q[Y(2) - Y(1) \mid R \in \mathcal R'] \quad \mathrm{and} \quad \theta_{\mathrm{Prob}}(Q) = Q\{Y(2) \geq Y(1) \mid R \in \mathcal R'\}~, \] where $\mathcal R' = \{(y(0),y(1),y(2),d(0),d(1),d(2)): (d(0),d(1),d(2)) = (0,1,2)\}$. We consider the following two specifications of $\mathcal R$, corresponding to two different models:
We first present simulation results for a discrete outcome with $\mathcal Y = \{-1, 0, 1\}$. From each of $\mathbf Q_{\mathrm{1s}}$ and $\mathbf Q_{\mathrm{enc}}$, we pick a $Q$ distribution and fix them. The two selected $Q$'s are described in Appendices (ref) and (ref), respectively. For each $Q$, we set $Q\{Z=z\} = 1/3$ for each $z \in \mathcal Z$ and define $P = Q T^{-1}$. The simulation goes as follows. For each $Q$, we repeatedly draw an i.i.d.\ sample $\{V_i\}_{i = 1}^n$ of size $n = 3000$ or $12000$ from $P$, where $V_i = (Y_i, D_i, Z_i)$. For each sample, we obtain the confidence interval for $\theta_{\mathrm{ATE}}$ and $\theta_{\mathrm{Prob}}$ by inverting the test of the feasibility of the linear system in (ref) at the 5% level.
The results are reported in Tables (ref) and (ref) for $\mathbf Q_{\mathrm{1s}}$ and $\mathbf Q_{\mathrm{enc}}$ respectively. For each parameter and each model, we report the length and the coverage rate of the confidence interval in the last two columns, averaged over 3000 replications. We also report the length of the identified set $|\Theta_0(P, \mathbf{Q}_{\rm 1s})|$ or $|\Theta_0(P, \mathbf{Q}_{\rm enc})|$. We note that the confidence interval covers the parameter with probability one in all cases. This phenomenon is to be expected because in each case the true value of the parameter lies strictly in the interior of the identified set. For example, $\Theta_0(P,\mathbf Q_{1s}) = [0.1482, 1.0295]$ and $\theta_{\rm ATE}(Q) = 0.4919$. In this example, for $n = 3000$, the coverage probabilities of the lower and upper endpoints of the identified set are 0.947 and 0.956, respectively.
We now present simulation results for a continuous outcome. In this subsection, $\mathcal D = \mathcal Z = \{0,1,2\}$ and $\mathcal Y = [-5,5]$. We focus on $\mathcal R = \mathcal R_{\mathrm{1s}}$ and $\theta_{\mathrm{ATE}}(Q) = E_Q[Y(2) - Y(1) \mid R \in \mathcal R']$. As in Section (ref), $Q\{Z = z\} = 1/3$ for each $z \in \mathcal Z$. A distribution $Q \in \mathbf Q_{\mathrm{1s}}$ is generated as follows. With a slight abuse of notation, let $\tilde Q$ on $\{-1,0,1\}^3 \times \mathcal D^3$ be the discrete distribution in Appendix (ref). We then define $Q$ as a mixture of truncated normal distributions where the mixture weights are given by $\tilde Q$. In particular, let $\Phi_{\mathcal Y}(\mu(d))$ denote the truncation of normal distribution $N(\mu(d), 1)$ on $\mathcal Y$. Then,
In words, for each mixture component, the value of $D(z)$ is fixed at $d(z)$ for $z \in \mathcal Z$ and independently across $d \in \mathcal D$, $Y(d)$ has a truncated normal distribution. Finally let $P = QT^{-1}$. As in Section (ref), we repeatedly draw an i.i.d.\ sample $\{V_i\}_{i = 1}^n$ of size $n = 3000$ or $12000$ from $P$, where $V_i = (Y_i, D_i, Z_i)$, and for each sample obtain confidence intervals by inverting the test of the feasibility of the linear system in (ref) at the 5% level. We consider the following three partitions of $\mathcal Y$:
The results are reported in Table (ref). For each level of discretization, we report the length and the coverage rate of the confidence intervals in the last two columns, averaged over 3000 replications. To illustrate the computational costs, we report in the second column of Table (ref) the number of rows and columns of the $A$ matrix in (ref) for each level of discretization, which correspond to the number of constraints and the number of decision variables in each of the linear systems being tested. Next, recall Theorem (ref) shows that for each level of discretization $L$, $\mathbf P_{0, L}$ contains $\mathbf P_0$. Let $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$ be the set of $\theta_0 \in \mathbf R$ such that $\mathbf P_{0, L}$ is nonempty. Theorem (ref) implies that $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$ is a superset of the identified set $\Theta_0(P,\mathbf Q_{\mathrm{1s}})$. In the third column of Table (ref), we display the length of $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$ for each $L$. A comparison of the length of our confidence intervals to the length of $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$ reveals how much of the length of the confidence interval is attributed to sampling uncertainty.
We note the following findings from Table (ref). First, as in Tables (ref) and (ref), the confidence intervals cover the parameter with probability one for each level of discretization we consider. This phenomenon is again to be expected because in each case the true value of the parameter lies strictly in the interior of $\Theta_{0,L}(P,\mathbf Q_{\mathrm{1s}})$; coverage probabilities of the lower and upper endpoints of $\Theta_{0,L}(P, \mathbf Q_{1s})$ are approximately equal to the nominal level. Next, as the number of intervals in the partition increases, the dimensionality of the problem grows exponentially. Specifically, the number of constraints and decision variables scale at the rates $O(|\mathcal L||\mathcal D||\mathcal Z|)$ and $O(|\mathcal L|^{|\mathcal D|} |\mathcal R_t|)$, respectively. For a fixed sample size, a finer partition yields a shorter confidence interval, although we note it comes at a higher computational cost.