Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
109,217 characters · 21 sections · 76 citation commands
Inference for Interval-Identified Parameters Selected from an Estimated Set
There is now a large and growing literature on partial identification of optimal treatments and policies under practically-relevant assumptions for observational data and experimental data with imperfect compliance (e.g., Sto12,KZ21, PZ21, DAd21,Yat21,Han24 to name just a few). Interval identification of average treatment effects (ATEs), average partial effects and welfare is particularly common in these settings due to the endogeneity of individuals' treatment uptake. In order to use the partial identification results for treatment or policy choice in practice, a researcher must typically estimate a set of best-performing treatments or policies from data. Consequently, the researcher is typically interested in a treatment or policy that is either selected from the estimated set of best-performers or arises from a data-dependent selection rule. It is now well-known that selecting an object of interest from data invalidates standard inference tools (e.g., AKM24). The failure of standard inference tools after data-dependent selection is only compounded by the presence of partially-identified parameters.
In this paper, we develop new inference tools for interval-identified parameters corresponding to selection from either an estimated set or arising from a data-dependent selection rule. Estimating identified sets for the best-performing treatments/policies or forming data-dependent selection rules in these settings is important for choosing which treatments/policies to implement in practice. Therefore, the ability to infer how well these treatments or policies should be expected to perform when selected, for instance to gauge whether their implementation is worthwhile, is of primary practical importance.
The current literature has not yet developed valid post-selection inference tools in partially-identified contexts, an important deficiency that the methods proposed in this paper aim to correct. The methods we propose here build upon the ideas of conditional and hybrid inference employed in various point-identified contexts by, e.g., LSST16, FST17, TRTW18, AKM24 and McC24 to produce confidence intervals (CIs) for interval-identified parameters such as welfare or ATEs chosen from an estimated set or via a data-dependent selection rule. Although ARP23 also propose conditional and hybrid inference methods in the partial identification context of moment inequality models, they do not allow for data-dependent selection of objects of interest, one of the main focuses of the present paper. Finally, this paper directly relies upon results in the literature on interval identification of welfare, treatment effects and partial outcomes such as Man90, BP97,BP13, MP00, MST18, HY24 and Han24. We apply our inference methods to a general class of problems nesting these examples.
After sketching the ideas behind our inference methods in a simple example, we introduce the general inference framework to which our methods can be applied. We show that our general inference framework incorporates several problems of interest for data-dependent selection and treatment rules for parameters belonging to an identified set such as Manski bounds for average potential outcomes or ATEs, bounds on parameters derived from linear programming (e.g., BP97,BP13; MST18; Han24; HY24), bounds on welfare for treatment allocation rules that are partially identified by observational data (e.g., Sto12,KZ21, PZ21, DAd21,Yat21) and bounds for dynamic treatment effects (e.g., Han24). We also show how to incorporate inference on parameters chosen via asymptotically optimal treatment choice rules (e.g., CMS23) into our general inference framework. Our framework can also be applied to settings where welfare is partially identified for reasons other than treatment endogeneity (e.g., IK21,AC22,BGIJ22,CH24).
Within the general inference framework, we develop three types of CIs for data-dependent and interval-identified parameters. As the name suggests, the conditional CIs are asymptotically valid conditional on the parameter of interest corresponding to a treatment or policy chosen from an estimated set. The construction of this CI does not require a specific rule for choosing the parameter of interest from the estimated set. In addition, the sampling framework underlying its conditional validity is most appropriate in contexts for which a researcher will only be interested in the parameter because it is chosen from the estimated set. Importantly, we show that these CIs are asymptotically valid uniformly across a large class of data-generating processes (DGPs). Uniform asymptotic validity is especially important for approximately correct finite-sample coverage in post-selection contexts like those in this paper (see, e.g., AG09).
The second and third types of CIs we develop in this paper are designed for inference on parameters chosen by data-dependent selection rules for which the object of interest is uniquely determined by the selection rule. The projection CIs do not require knowledge of the selection rule to be asymptotically valid, whereas the hybrid CI construction utilizes the particular form of a selection rule to improve upon the length properties of the projection CI. The conditional CIs are short for selections that occur with high probability but can become exceptionally long when this probability is small (see, e.g., KL21). Conversely, projection CIs are overly conservative when selection probabilities are high. Although conditional CIs can be used in this setting, hybrid CIs interpolate the length properties of the conditional and projection CIs in order to attain good length properties regardless of the value selection probabilities take. In analogy with the conditional CIs, we formally show that both projection and hybrid CIs are asymptotically valid in a uniform sense.
We analyze the coverage and length properties of our proposed CIs in finite samples. Since, to our knowledge, these are the first uniformly valid CIs for data-dependent selections of partially-identified parameters, there are no existing CIs to which we can directly compare. Nevertheless, since our CIs can also be used for inference on a priori chosen interval-identified parameters, we conduct a power comparison with one of the leading methods for inference on a partially-identified parameter. In particular, we compare the power of the test implied by our hybrid CIs to the power of the hybrid test of ARP23, a test that applies to a general class of moment-inequality models that is also based upon a (different) hybrid between conditional and projection-based inference. Encouragingly, the power of the test implied by our hybrid CI is quite competitive even in this environment for which it was not designed. We also find that the finite-sample coverage of all of our CIs is approximately correct in a simple Manski bound example. Finally, we analyze the length tradeoffs between the three different CIs across different DGPs, finding the hybrid CI to perform best overall.
The remainder of this paper is structured as follows. Section (ref) sketches the ideas behind our general CI constructions in the context of a simple Manski bound example. Section (ref) lays down the general high-level inference framework we are interested in, while Section (ref) details how the general framework applies in several different examples. Section (ref) then details the various CI constructions in the general setting. Sections (ref) and (ref) are devoted to finite-sample comparisons of the properties of the different CIs in the context of a simple Manski bound example. The final section, Section (ref), contains an empirical application where we apply our procedures to dynamic policies of schooling and post-school training. Appendix (ref) contains additional examples that fit our general framework that are not covered in Section (ref), while Appendix (ref) contains additional simulation results corresponding to the dynamic treatment regime example detailed in Section (ref). Mathematical proofs are relegated to a Technical Appendix at the end of the paper.
We first provide a simple example to illustrate our proposed methods. Consider a binary outcome of interest $Y$, a binary treatment indicator $D$ and a binary treatment assignment $Z$. Furthermore, let $Y(1)$ and $Y(0)$ denote potential outcomes under treatment ($D=1$) and no treatment ($D=0$). Assuming $\mathbb{E}[Y(d)|Z]=\mathbb{E}[Y(d)]$, Man90 shows that we can bound the average potential outcomes $W(d)=\mathbb{E}[Y(d)]$ in the absence and presence of treatment as follows:
for $d=0,1$ and $p^{ydz} \equiv Pr(Y=y,D=d|Z=z)$. Given (ref) it is natural to define the set of best-performing options $\mathcal{D}^{*}$, as a subset of the two options of treatment and no treatment, to be those that are undominated options. From an observed dataset of outcomes, treatments and treatment assignments $\{(Y_i,D_i,Z_i)\}_{i=1}^n$, such a set can be estimated as: \[ \widehat{\mathcal{D}}=\left\{d\in\{0,1\}:\widehat U(d)\geq \widehat L(d^ \prime)\text{ }\forall d^{ \prime}\in\{0,1\} \text{ s.th. }d^{ \prime}\neq d\right\}, \] where $\widehat L(d)\equiv \max\{\hat p^{1d0},\hat p^{1d1}\}$ and $\widehat U(d)\equiv\min\{1-\hat p^{0d0},1-\hat p^{0d1}\}$ with $\hat p^{ydz}$ being an empirical estimate of the fitted probability $p^{ydz} \equiv \mathbb{P}(Y=y,D=d|Z=z)$.
We are interested in inference on the identified interval $[L(d),U(d)]$ for the average potential outcome $W(d)$ of option $d\in\{0,1\}$ after the researcher selects this option from $\widehat{\mathcal{D}}$. In other words, we would like to provide statistically precise statements about the true average potential outcome of an option selected from the data to give the researcher an idea of how well this selected option should be expected to perform in the population. We first provide some intuition for why standard inference techniques based upon asymptotic normality fail and then sketch our proposals for valid inference in this context.
To fix ideas, let us focus for now on inference for the lower bound $L(d)$ of a selected option rather than the entire identified set $[L(d),U(d)]$. More specifically, since $L(d)$ is a lower bound for the average potential outcome $W(d)$, we would like to obtain a probabilistic lower bound for $L(d)$. Under standard conditions, a central limit theorem implies $\hat p=(\hat p^{100},\hat p^{010},\hat p^{110},\hat p^{101},\hat p^{011},\hat p^{111})^\prime$ is normally distributed in large samples. So why not form a CI using $\widehat L(d)$ and quantiles from a normal distribution as the basis for inference? There are two reasons such an approach is (asymptotically) invalid:
Reason 1. is easy to see since $\widehat L(d)$ is the maximum of two normally distributed random variables in large samples when $d$ is chosen a priori. To better understand reason 2., note that the distribution of $\widehat L(d)$ given $d\in \widehat{\mathcal{D}}$ is the conditional distribution of the maximum of two normally distributed random variables given that the minimum of two other normally distributed random variables, $\widehat U(d)\equiv\min\{1-\hat p^{0d0},1-\hat p^{0d1}\}$, exceeds the maximum of yet another set of two normally distributed random variables, $\widehat L(d')=\max\{\hat p^{1d'0},\hat p^{1d'1}\}$ for $d'\neq d$. Unconditionally, $\widehat L(\hat d)$ for any data-dependent choice of $\hat d$ is distributed as a mixture of the distributions of $\widehat L(0)$ and $\widehat L(1)$, neither of which are themselves normally distributed.
Suppose that a researcher's interest in inference on $L(d)$ only arises when $d$ is estimated to be in the set of best-performing options, viz., $d\in \widehat{\mathcal{D}}$. In such a case, we are interested in a probabilistic lower bound for $L(d)$ that is approximately valid across repeated samples for which $d\in \widehat{\mathcal{D}}$, i.e., we would like to form a conditionally valid lower bound $\widehat L(d)_\alpha^C$ such that\footnote{See AKM24 for an extensive discussion of when conditional vs unconditional validity is desirable for inference after selection.}
for some $\alpha\in(0,1)$ in large samples. To do so we characterize the conditional distribution of $\widehat L(d)$. Specifically, let $\hat j_L(d)\equiv \operatorname*{argmax\;}_{j\in\{0,1\}}\hat p^{1dj}$ be the value of $Z$ at which the maximum between the two estimated probabilities is achieved. Then $\widehat L(d)=\hat p^{1d\hat j_L(d)}$. Also, since the conditioning event $\{d\in \widehat{\mathcal{D}},\hat j_L(d)=j_L^*\}$ can be written as a polyhedron in $\hat p$ and $\widehat L(d)$ is equal to an element of $\hat p$, Lemma 5.1 of LSST16 implies
for some known functions $\widehat{\mathcal{L}}_L(\cdot)$ and $\widehat{\mathcal{U}}_L(\cdot)$, where $\hat p^{1dj_L^*}\sim\mathcal{N}(p^{1dj_L^*},\operatorname*{Var}(\hat p^{1dj_L^*}))$ in large samples and $\mathcal{Z}_L$ is a sufficient statistic for the nuisance parameter $p=( p^{100}, p^{010}, p^{110}, p^{101}, p^{011}, p^{111})^\prime$ that is asymptotically independent of $\hat p^{1dj_L^*}$. Using results in Pfa94, the characterization in (ref) permits the straightforward computation of a conditionally quantile-unbiased estimator for $p^{1d\hat j_L(d)}=p^{1dj_L^*}$, since the latter is equal to the mean of the underlying normally distributed random variable $\hat p^{1dj_L^*}$ that is subject to truncation. Denoting this quantile-unbiased estimator as $\widehat L(d)_{\alpha}^C$, we have
in large samples. However, noting that $L(d)\geq p^{1d\hat j_L(d)}$ with probability one, we can see that (ref) holds for this choice of $\widehat L(d)_{\alpha}^C$.
Although (ref) does not hold with exact equality, we note that the left-hand side cannot be much larger than the right-hand side. In other words, although $\widehat L(d)_{\alpha}^C$ is a conservative probabilistic lower bound for $L(d)$, it is not very conservative. This can be seen heuristically by working through the two possible values that $L(d)$ can take:
Finally, a construction analogous to that described above for producing a probabilistic lower bound for $L(d)$ produces a conditionally valid probabilistic upper bound $\widehat U(d)_{1-\alpha}^C$ for $U(d)$ that satisfies
for some $\alpha\in(0,1)$ in large samples. The probabilistic lower and upper bounds can then be combined to form a CI, $[\widehat L(d)_{\alpha/2}^C,\widehat U(d)_{1-\alpha/2}^C]$, that is conditionally valid for $[L(d),U(d)]$ in large samples since
Suppose now that the researcher uses a data-dependent rule to select a unique option of inferential interest. For example, suppose the researcher is interested in choosing the option with the highest potential outcome in the worst case across its identified set so that she chooses $\hat d=\operatorname*{argmax\;}_{d\in\{0,1\}} \widehat L(d)$. In such a case, it is natural to form a probabilistic lower bound $\widehat L(\hat d)_{\alpha}^U$ for $L(\hat d)$ that is unconditionally valid across repeated samples such that
for some $\alpha\in(0,1)$ in large samples. Given its conditional validity (ref), the conditional lower bound $\widehat L(d)_{\alpha}^C$ also satisfies (ref) upon changing the definition of $\widehat{\mathcal{D}}$ to $\widehat{\mathcal{D}}=\{\hat d\}$ in its construction. However, it is well known in the literature on selective inference that conditionally-valid probabilistic bounds can be very uninformative (i.e., far below the true value) when the probability of the conditioning event is small (see e.g., KL21, AKM24 and McC24). Here, we propose two additional forms of probabilistic bounds that are only unconditionally valid but do not suffer from this drawback.
First, we can form a probabilistic lower bound for $L(\hat d)$ by projecting a one-sided rectangular simultaneous confidence lower bound for all possible values $L(\hat d)$ can take: $\widehat L(\hat d)_\alpha^P \equiv\widehat L(\hat d)-\hat c_{1-\alpha,L}\sqrt{\widehat\Sigma_{L,2\hat d+1+\hat j_L(\hat d)}},$ where $\hat c_{1-\alpha,L}$ is the $1-\alpha$ quantile of $\max_i\hat\zeta_i/\sqrt{\widehat\Sigma_{L,i}}$ for $\hat\zeta\sim\mathcal{N}(0,\widehat\Sigma_L)$, $\widehat\Sigma_L$ is a consistent estimator of $\Sigma_L\equiv\operatorname*{Var}(\hat p^{100},\hat p^{101},\hat p^{110},\hat p^{111})$ and $\Upsilon_i$ denotes the $i^{th}$ element of the main diagonal of any square matrix $\Upsilon$. Here, the maximum is taken to guarantee simultaneous coverage of all possible values of $L(\hat d)$. Since $p^{1\hat d\hat j_L(\hat d)}\in\{\hat p^{100},\hat p^{101},\hat p^{110},\hat p^{111}\}$ with probability one,
in large samples and (ref) holds for $\widehat L(\hat d)_\alpha^U=\widehat L(\hat d)_\alpha^P$ because $L(\hat d)\geq p^{1\hat d\hat j_L(\hat d)}$. However, $\widehat L(\hat d)_\alpha^P$ suffers from a converse drawback to that of $\widehat L(\hat d)_\alpha^C$: it is unnecessarily conservative when $\hat d=d$ is chosen with high probability (see e.g., AKM24 and McC24).
We propose a second probabilistic lower bound for $L(\hat d)$ that combines the complementary strengths of $\widehat L(d)_\alpha^C$ and $\widehat L(\hat d)_\alpha^P$. Construction of this hybrid lower bound $\widehat L(\hat d)_\alpha^H$ proceeds analogously to the construction of $\widehat L( d)_\alpha^C$ after adding the additional condition $\{p^{1\hat d\hat j_L(\hat d)}\geq \widehat L(\hat d)_\beta^P\}$ for $\beta<\alpha$ to the conditioning event and instead computing a conditionally quantile-unbiased estimator for $p^{1\hat d\hat j_L(\hat d)}$, denoted as $\widehat L(\hat d)_{\alpha}^H$, satisfying
in large samples, where $d^*$ is any realized value of the random variable $\hat d$. Imposing this additional condition in the formation of the hybrid bound ensures that $\widehat L(\hat d)_{\alpha}^H$ is always greater than $\widehat L(\hat d)_\beta^P$, limiting its worst-case performance relative to $L(\hat d)_\beta^P$ when $\mathbb{P}(\hat d=d^*)$ is small. On the other hand, when $\mathbb{P}(\hat d=d^*)$ is large, the additional condition $\{p^{1\hat d\hat j_L(\hat d)}\geq \widehat L(\hat d)_\beta^P\}$ is far from binding with high probability so that $\widehat L(\hat d)_\alpha^H$ becomes very close to $\widehat L(d)_{(\alpha-\beta)/(1-\beta)}^C$. In this case, $\widehat L(d)_{(\alpha-\beta)/(1-\beta)}^C$ is close to the naive lower bound based upon the normal distribution $\widehat L(\hat d)-z_{(1-\alpha)/(1-\beta)}\sqrt{\operatorname*{Var}(\hat p^{1dj_L^*})}$ because the truncation bounds in (ref) are very wide (Proposition 3 in AKM24).
To see how (ref) holds for $\widehat L(\hat d)_\alpha^U=\widehat L(\hat d)_\alpha^H$, note first that
for all $d^*\in\{0,1\}$. Then, note that
by the law of total probability.
By similar reasoning to that used for the conditional CIs in Section (ref) above, $\widehat L(\hat d)_{\alpha}^H$ is not very conservative as a probabilistic lower bound for $L(\hat d)$. The researcher's choice of $\beta\in(0,\alpha)$ trades off the performance of $\widehat L(\hat d)_{\alpha}^H$ across scenarios for which $\mathbb{P}(\hat d=d^*)$ is large and small with a small $\beta$ corresponding to better performance when $\mathbb{P}(\hat d=d^*)$ is large. See McC24 for an in-depth discussion of these tradeoffs. We recommend $\beta=\alpha/10$.
Finally, analogous constructions to those above produce unconditional projection and hybrid probabilistic upper bounds $\widehat U(\hat d)_{1-\alpha}^P$ and $\widehat U(\hat d)_{1-\alpha}^H$ that can then be combined with the lower bounds to form CIs $[\widehat L(d)_{\alpha/2}^P,\widehat U(d)_{1-\alpha/2}^P]$ and $[\widehat L(d)_{\alpha/2}^H,\widehat U(d)_{1-\alpha/2}^H]$ for $[L(d),U(d)]$ that are unconditionally valid in large samples by the same arguments as those used in (ref) above.
We now introduce the general inference framework that we propose, nesting the Manski bound example of the previous section as a special case. After introducing the general framework, we describe several additional example applications that fall within this framework.
We are interested in performing inference on a parameter $W(d)$ that is indexed by a finite set $d\in \mathcal D \equiv \{d^0,\ldots,d^K\}$ for some $K>0$. The index $d$ may correspond to a particular treatment, treatment allocation rule or policy, depending upon the application. We assume that $W(d)$ belongs to an identified set taking a particular interval form that is common to many applications of interest.
The lower and upper endpoints of identified sets for the welfare, average potential outcome or ATE typically take the form of $L(d)$ and $U(d)$, especially when (sequences of) outcomes, treatments and instruments are discrete.
In the setting of this paper, a researcher's interest in $W(d)$ arises when $d$ belongs to a set $\widehat{\mathcal{D}}\subset \mathcal D$ that is estimated from a sample of $n$ observations. It is often the case that $\widehat{\mathcal{D}}$ is an estimate of the identified set of best performers $\mathcal{D}^{*}$. This set could correspond to an estimated set of optimal treatments or policies or other data-dependent index sets of interest. The estimated set is determined by an estimator $\hat p$ of the finite-dimensional parameter $p$ that determines the bounds on $W(d)$ according to Assumption (ref). Let
which are the indices at which the estimated lower and upper bounds are realized. Then, the estimated lower and upper bounds for $W(d)$ are equal to $\tilde\ell_{d,\hat j_L(d)}+\ell_{d,\hat j_L(d)}\hat p$ and $\tilde u_{d,\hat j_U(d)}+u_{d,\hat j_U(d)}\hat p$. We work under the high-level assumption that the following event can be written as a polyhedron in $\hat p$: (i) an option index $d$ is in the set of interest $\widehat{\mathcal{D}}$, (ii) the estimated bounds on $W(d)$ are realized at a given value and (iii) (optionally) an additional random vector is realized at any given value.
Depending upon the application, $\hat\gamma_L(d)$ and $\hat\gamma_U(d)$ (and thus $\gamma_L^*$ and $\gamma_U^*$) in this assumption may not be necessary to condition on, in which case they can be vacuously set to constants. Although not immediately obvious, this assumption holds in a variety of settings; see the examples below. In many cases, this assumption can be simplified because, consistent with Assumption 3.1, $d\in\widehat{\mathcal{D}}$ if and only if $A_{\mathcal{D}}\hat{p}\leq c_{\mathcal{D}}$ for some fixed and known matrix $A_{\mathcal{D}}$ and vector $c_{\mathcal{D}}$. For these cases, $\hat\gamma_L(d)$ and $\hat\gamma_U(d)$ are not needed and can be vacuously set to fixed constants and
and
A leading example of this special case is \[ \widehat{\mathcal{D}}=\left\{d\in\{d^0,\ldots,d^K\}:\widehat U(d)\geq \max_{d\in\{d^0,\ldots,d^K\}}\widehat L(d)\right\}, \] where $\widehat L(d)\equiv \max_{j\in \{1,\ldots,J_L\}}\{\tilde\ell_{d,j}+\ell_{d,j}\hat p\}$ and $\widehat U(d)\equiv \min_{j\in \{1,\ldots,J_U\}}\{\tilde u_{d,j}+u_{d,j}\hat p\}$, since $d\in\widehat{\mathcal{D}}$ if and only if \[(\ell_{d',j'}-u_{d,j})\hat p \leq \tilde u_{d,j}-\tilde\ell_{d',j'}\] for all $d'\in\{d^0,\ldots,d^K\}$, $j\in \{1,\ldots,J_U\}$ and $j'\in\{1,\ldots,J_L\}$.
We also note that Assumption (ref) is compatible with the absence of data-dependent selection for which the researcher is interested in forming a CI for an identified interval $[L(d^*),U(d^*)]$ chosen by the researcher a priori. In these cases, $\widehat{\mathcal{D}}=\{d^*\}$, $\hat\gamma_M(d^*)$ can be vacuously set to a fixed constant, $A^M(d^*,j,\gamma)=A_M(d^*,j)$ and $c^M(d^*,j,\gamma)=c_M(d^*,j)$ for $M=L,U$. Indeed, we examine an example of this special case when conducting a finite-sample power comparison in Section (ref) below.
In general, less conditioning is more desirable in terms of the lengths of the CIs we propose. Although conditioning on the events $d\in\widehat{\mathcal{D}}$, $\hat j_L(d)=j_L^*$ and $\hat j_U(d)=j_U^*$ is necessary to construct our CIs (see Section (ref) below), the researcher should therefore minimize the number of elements in $\hat \gamma_L(d)$ and $\hat \gamma_U(d)$ subject to satisfying Assumption (ref) when constructing our CIs. In some cases it is necessary to condition on these additional random vectors in order to satisfy Assumption (ref). But in many cases, such as the example given immediately above, additional conditioning random vectors are unnecessary and can be vacuously set to fixed constants.
We impose the following assumption for our unconditional hybrid CIs in order for the object of inferential interest to be well-defined unconditionally.
In conjunction, Assumptions (ref) and (ref) hold naturally when the object of interest $\hat d$ is selected by uniquely maximizing a linear combination of the estimates of the bounds characterizing the identified intervals and the additional conditioning vectors $\hat\gamma_L(d)$ and $\hat\gamma_U(d)$ are defined appropriately. Leading examples of this form of selection include when $\hat d$ corresponds to the largest estimated lower bound, upper bound or weighted average of lower and upper bounds.
Expressions for $A^M(d,j_M^*,\gamma_M^*)$ and $c^M(d,j_M^*,\gamma_M^*)$ for $M=L,U$ in the settings of Proposition (ref) are available for reference in its proof in Appendix (ref). As this proposition makes clear, the additional conditioning vectors needed for Assumption (ref) to hold depend upon the particular form of selection rule used by the researcher. For example, when $\hat d$ is chosen to maximize the estimated lower bound of the identified set $\widehat L( d)$, one must condition not only on the realized value of $\hat j_U(\hat d)$ when forming a probabilistic upper bound for $U(\hat d)$ but also $\hat j_L(\hat d)$. On the other hand, the formation of either a probabilistic lower bound for $L(\hat d)$ or upper bound for $U(\hat d)$ when $\hat d$ is chosen to maximize the estimated upper bound of the identified set $\widehat U( d)$ requires conditioning on the entire vector $(\hat j_U(0),\ldots,\hat j_U(T))'$.
Although intuitively appealing, the treatment choice rules of the form described in Proposition (ref) can be sub-optimal from a statistical decision-theoretic point of view (see, e.g., Man21,Man23 and CMS23). In Section (ref), we show how proper definition of $\hat\gamma_L(d)$ and $\hat\gamma_U(d)$ satisfies Assumptions (ref) and (ref) in the context of the optimal selection rules of CMS23.
We suppose that the sample of data is drawn from some unknown distribution $\mathbb{P}\in\mathcal{P}_n$. As an estimator for $p$, we assume that $\hat p$ is uniformly asymptotically normal under $\mathbb{P}\in\mathcal{P}_n$.
The notation of this assumption makes explicit that the parameter $p$ and the asymptotic variance $\Sigma$ depend upon the unknown distribution of the data $\mathbb{P}$. It holds naturally for standard estimators $\hat p$ under random sampling or weak dependence in the presence of bounds on the moments and dependence of the underlying data.
Next, we assume that the asymptotic variance of $\hat p$ can be uniformly consistently estimated by an estimator $\widehat\Sigma$.
This assumption is again naturally satisfied when using a standard sample analog estimator of $\Sigma$ under random sampling or weak dependence in the presence of moment and dependence bounds.
In addition, we restrict the asymptotic variance of $\hat p$ to be positive definite.
This assumption is naturally satisfied, for example, when $\hat p$ is a standard sample analog estimator of reduced-form probabilities composing $p$ that are non-redundant and bounded away from zero and one.
In this section, we show that the proposed inference method is applicable to various examples for which parameters are interval-identified. In particular, we show that Assumptions (ref), (ref) and (ref) are satisfied in these examples. See Appendix (ref) for additional examples.
In more complex settings, calculating analytical bounds on $W(d)$ or $W(d)-W(\tilde{d})$ may be cumbersome. This is especially true when the researcher wants to incorporate additional identifying assumptions. In this situation, the computational approach using linear programming can be useful (MST18,HY24).
To incorporate many complicated settings, suppose that $W(d)=A_{d}q$ and $p=Bq$ for some known row vector $A_{d}$ and matrix $B$, an unknown vector $q$ in a simplex $\mathcal{Q}$, and a vector $p$ that is estimable from data. Typically $q$ is a vector of probabilities of a latent variable that governs the DGP; see BP97,BP13, Han24 and HY24. The linearity in this assumption is usually implied by the nature of a particular problem (e.g., discreteness). Then we have
and ATE bounds for a change from treatment $d$ to treatment $\tilde d$
Note that $L(\tilde d,{d})\neq L(\tilde d)-U({d})$ in general because the $q$ that solves (ref) for $L(\tilde d)$ and $U({d})$ may be different (and similarly for $U(\tilde d,{d})$). As before, the identified set of optimal treatments here is characterized as $\mathcal{D}^{*}\equiv \{d:L(\tilde{d},d)\le0,\forall\tilde{d}\neq d\}$.
An example of this setting can be found in HY24. Let $(Y,D,Z)$ be a vector of a binary outcome, treatment and instrument and let $p$ be a vector with entries $p(y,d|z)\equiv \mathbb{P}(Y=y,D=d|Z=z)$ across $(y,d,z)\in\{0,1\}^{3}$.\footnote{See HY24 for the use of linear programming with continuous $Y$.} Suppose $W(d)=\mathbb{E}[Y(d)]$ for $d\in\{0,1\}$. Then, we can define the response type $\varepsilon\equiv(Y(1),Y(0),D(1),D(0))$ with a realized value $e\equiv(y(1),y(0),d(1),d(0))$, where $Y(d)$ denotes the potential outcome under treatment $d$ and $D(z)$ denotes the potential treatment under instrument value $z$. Let $q(e)\equiv \mathbb{P}(\varepsilon=e)$ be the latent distribution. Then
where $q$ is the vector of $q(e)$'s and $A_{d}$ is an appropriate selector (a row vector).
Assume that $(Y(d),D(z))$ is independent of $Z$ for $d,z\in\{0,1\}$. The data distribution $p$ is related to the latent distribution by
where the first equality follows by the independence assumption, $q$ is a vector of $q(e)$'s and $B_{d,z}$ is an appropriate selector (a row vector). Now define
so that all of the constraints relating the data distribution to the latent distribution can be expressed as $Bq=p$.
To verify Assumption (ref), it is helpful to invoke strong duality for the primal problems (ref) (under regularity conditions) and write the following dual problems:
where $\tilde{B}\equiv\left[
\right]$ is a $(d_{p}+1)\times d_{q}$ matrix with $\boldsymbol{1}$ being a $d_{q}\times1$ vector of ones, and $\tilde{p}\equiv\left[
\right]$ is a $(d_{p}+1)\times1$ vector. By using a vertex enumeration algorithm (e.g., \citet{AF91}), one can find all (or a relevant subset) of vertices of the polyhedra $\{\lambda:\tilde{B}'\lambda\ge-A_{d}'\}$ and $\{\lambda:\tilde{B}'\lambda\ge A_{d}'\}$. Let $\Lambda_{L,d}\equiv\{\lambda_{1},...,\lambda_{J_{L,d}}\}$ and $\Lambda_{U,d}\equiv\{\lambda_{1},...,\lambda_{J_{U,d}}\}$ be the sets that collect such vertices, respectively. Then, it is easy to see that $L(d)=\max_{\lambda\in\Lambda_{L,d}}-\tilde{p}'\lambda$ and $U(d)=\min_{\lambda\in\Lambda_{U,d}}\tilde{p}'\lambda$, and thus Assumption (ref) holds.
To verify Assumption (ref), we use the dual problems to (ref):
where $\Delta_{\tilde d,{d}}\equiv A_{\tilde d}-A_{{d}}$. Analogous to the vertex enumeration argument above, let $\Lambda_{L,\tilde d,{d}}\equiv\{\lambda_{1},...,\lambda_{J_{L,\tilde d,{d}}}\}$ and $\Lambda_{U,\tilde d,{d}}\equiv\{\lambda_{1},...,\lambda_{J_{U,\tilde d,{d}}}\}$ be the sets that collect all (or a relevant subset) of vertices of the polyhedra $\{\lambda:\tilde{B}'\lambda\ge-\Delta_{\tilde d,{d}}'\}$ and $\{\lambda:\tilde{B}'\lambda\ge\Delta_{\tilde d,{d}}'\}$, respectively. Then, $L(\tilde d,{d})=\max_{\lambda\in\Lambda_{L,\tilde d,{d}}}-\tilde{p}'\lambda$ and $U(\tilde d,{d})=\min_{\lambda\in\Lambda_{U,\tilde d,{d}}}\tilde{p}'\lambda$. Let $\widehat{\mathcal{D}}=\{d:\widehat{L}(\tilde d,{d})\le0,\forall\tilde{d}\neq d\}$, where $\widehat{L}(\tilde d,{d})$ is the sample counterpart of $L(\tilde d,{d})$ with $\widehat{\tilde{p}}\equiv\left[
\right]$ replacing $\tilde{p}\equiv\left[
\right]$. Partition $\lambda$ as $\lambda=(\lambda^{1\prime},\lambda^{0})'$ where $\lambda^{0}$ is the last element of $\lambda$. Note that $d\in\widehat{\mathcal{D}}$ if and only if
where $\tilde{\Lambda}_{L,d}=\bigcup_{\tilde{d}\neq d}\Lambda_{L,\tilde d,{d}}$. Also let $\hat{\lambda}$ be such that $-\widehat{\tilde{p}}'\hat{\lambda}=\max_{\lambda\in\tilde{\Lambda}_{L,d}}-\widehat{\tilde{p}}'\lambda$. Then, $\hat{\lambda}=\lambda_{j_{L}^{*}}$ if and only if
so that Assumption (ref) holds.
Finally, $\hat p$ is again equal to a vector of sample means so that Assumption (ref) is satisfied if $\hat{p}$ is calculated using the random sample $\{Y_{i},D_{i},Z_{i}\}$$_{i=1}^{n}$.
Consider allocating a binary treatment based on observed covariates $X\in\mathcal{X}$. A treatment allocation rule can be defined as a function $\delta:\mathcal{X}\rightarrow\{0,1\}$ in a class of rules $\mathcal{D}$. Consider the utilitarian welfare of deploying $\delta$ relative to treating no one. The optimal allocation $\delta^{*}$ satisfies
Note that $\mathbb{E}[Y(\delta(X))-Y(0)] =E\left[\delta(X)\Delta(X)\right]$, where $\Delta(X)\equiv \mathbb{E}[Y(1)-Y(0)|X]$. This problem is considered in KT18 and AW21, among others. When only observational data for $(Y,D,X)$ are available with $D$ being endogenous, $W(\delta)$ is only partially identified unless strong treatment effect homogeneity is assumed. This problem has been studied in KZ21, PZ21, DAd21, Bya22, among others. Using instrumental variables, one can consider bounds on the conditional ATE based on conditional versions of the bounds considered in Sections (ref) and (ref) (i.e., Manski's bounds and bounds produced by linear programming).
In particular, assume that $(Y(d),D(z))$ is independent of $Z$ given $X$. Let $L(X)$ and $U(X)$ be conditional Manski bounds on $\Delta(X)$. Then, bounds on $W(\delta)$ can be characterized as
Similarly, bounds on $W(\tilde\delta)-W({\delta})=\mathbb{E}[(\tilde\delta(X)-{\delta}(X))\Delta(X)]$ can be characterized as
Note that $L(\tilde\delta,{\delta})\neq L(\tilde\delta)-U({\delta})$ in general (and similarly for $U(\tilde\delta,{\delta})$).
Suppose $\mathcal{X}$ is finite and $\mathcal{X}=\{x_{1},...,x_{K}\}$ where $x_{k}$ can be a vector and $K$ can potentially be large. For simplicity of exposition, suppose $\mathcal{X}=\{0,1,2\}$. Then $\mathcal{D}=\{\delta_{1},...,\delta_{8}\}$ where each $\delta_{j}$ corresponds to a mapping type from $\{0,1,2\}$ to $\{0,1\}$. To verify Assumptions (ref) and (ref), we proceed as follows. For given $x\in\mathcal{X}$, by arguments analogous to those in Section (ref) (and Section (ref)), bounds $L_x$ and $U_x$ on $\Delta(x)$ satisfy, for some scalars $\tilde{\ell}_{j}$ and $\tilde{u}_{j}$ and row vectors $\ell_{j}$ and $u_{j}$,
where $p(x)$ is the vector of $p(y,d|z,x)$'s across $(y,d,z)$ fixing $x$. Then, by Jensen's inequality, for each $\delta\in\mathcal{D}$,
Note that $\tilde{L}(\delta)$ and $\tilde{U}(\delta)$ are non-sharp bounds; for calculation of sharp bounds, see Section (ref). We can verify Assumption (ref) with $\tilde{L}(\delta)$ and $\tilde{U}(\delta)$ by defining
and, for $\delta=\delta_{1}$ as an example, by using $\ell_{\delta_{1},j} =(
)$. Similarly, we can verify Assumptions \ref{ass: estimated ID'd set} and \ref{ass: joint normality} by estimating $p(X)$ and $\mathbb{E}[\delta(X)]$ with sample means and ${E}[\delta(X){p}(X)]$ with $\frac{1}{n}\sum_{i}^{n}\delta(X_{i})\hat{p}(X_{i}).$ If the data $\{Y_{i},D_{i},Z_{i},X_{i}\}$$_{i=1}^{n}$ form a random sample, and $\widehat{\mathcal{D}}=\{\delta\in\mathcal{D}:\widehat{L}(\tilde{\delta},\delta)\le0\,\forall\tilde{\delta}\neq \delta\}$ for $\widehat{L}(\tilde{\delta},\delta)$ defined the same as ${L}(\tilde{\delta},\delta)$ in (ref) after substituting $\hat p$ for $p$, Assumptions (ref) and (ref) hold.
This framework can be generalized to settings where $W(\delta)$ is partially identified, not necessarily due to treatment endogeneity but because $W(\delta)$ is a non-utilitarian welfare defined as a functional of the joint distribution of potential outcomes (e.g., CH24): $W(\delta)=f(F_{Y(1),Y(0)|X})$ where $f$ is some functional and $F_{Y(1),Y(0)|X}$ is the joint distribution of $(Y(1),Y(0))$ conditional on $X$.
Consider binary $Y_{t}$ and $D_{t}$ for $t=1,...,T$. Let $Y\equiv(Y_{1},...,Y_{T})$ and $D\equiv(D_{1},...,D_{T})$. Suppose that we are equipped with a sequence of binary instruments $Z\equiv(Z_{t_{1}},...,Z_{t_{K}})$, which is a subvector of $(Z_{1},...,Z_{T})$. For $t=1,...,T$, let $Y_{t}(d_{1},...,d_{t})$ be the potential outcome at $t$ and $Y(d)\equiv(Y_{1}(d_{1}),...,Y_{T}(d_{1},...,d_{T}))$. We assume that the instruments $Z$ are independent of the potential outcomes $Y(d)$.
Let $T=2$. Then $Y\equiv(Y_{1},Y_{2})$ and $D\equiv(D_{1},D_{2})$. For given welfare $W(d)$ with $d\equiv (d_1,d_2)$, we are interested in the optimal policy $d^{*}$ that satisfies
where $\mathcal{D}\equiv\{(1,1),(1,0),(0,1),(0,0)\}$. The sign of the welfare difference, $W(d)-W(\tilde{d})$ for $d,\tilde{d}\in\mathcal{D}$, is useful for establishing the ordering of $W(d)$ with respect to $d$ and thus to identifying $d^{*}$. However, without additional identifying assumptions, we can only establish a partial ordering of $W(d)$ based on the bounds on the welfare difference Han24. This will produce the identified set $\mathcal{D}^{*}$ for $d^{*}$.
An example of the welfare is $W(d)\equiv \mathbb{E}[Y_{2}(d)]$, namely, the average potential terminal outcome. The bounds on welfare $W(d)$ are
where
which have forms analogous to those in the static case. Define the dynamic ATE in the terminal period for a change in treatment from $d$ to $\tilde d$ as
Then the bounds on the dynamic ATE are as follows:
Another example of welfare is the joint distribution $W(d)\equiv \mathbb{P}(Y(d)=(1,1))$ where $Y(d)\equiv(Y_{1}(d_{1}),Y_{2}(d))$. The bounds on $W(d)$ in this case are $L(d)\equiv\max_{z}L(d;z)$ and $U(d)\equiv\min_{z}U(d;z)$ where
with $d'_1\neq d_1$ and $d'_2\neq d_2$. Consider the effect of treatment on the joint distribution, for example,
Then, with $\tilde d=(1,1)$ and ${d}=(1,0)$, the bounds on this parameter are
In these examples, the identified set $\mathcal{D}^{*}$ can be characterized as a set of maximal elements:
These examples are special cases of the model in Han24.\footnote{See also HL24 for other examples of dynamic causal parameters that can be used to define optimal treatments.}
In both cases, it is easy to see that $L(d)$ and $U(d)$ satisfy Assumption (ref) with $p$ being the vector of probabilities $p(y,d|z)\equiv \mathbb{P}(Y=y,D=d|Z=z)$. To verify Assumption (ref), let $\widehat{L}(\tilde{d},d)$ be the estimator of $L(\tilde{d},d)$ where the sample frequency replaces the population probability. Then $\widehat{\mathcal{D}}=\{d:\widehat{L}(\tilde{d},d)\le0,\forall\tilde{d}\neq d\}$.
Continuing with the first example, let $Z=Z_{1}\in\{0,1\}$, that is, the researcher is only equipped with a binary instrument in the first period and no instrument in the second period. We focus on inference for $L(d)$ for $d\in\widehat{\mathcal{D}}$. Let $\widehat{z}(d)\in\arg\max_{z\in\{0,1\}}\widehat{L}(d;z)$ so that $\widehat{L}(d)=\widehat{L}(d;\widehat{z}(d))$. Then we can write the data-dependent event that $d$ is an element of $\widehat{\mathcal{D}}$, $\widehat{L}(d)=\widehat{L}(d;z^{*})$ as a polyhedron
for some matrix $A^L$, where $\hat{p}$ is the vector of probabilities $\widehat{p}(y,d|z)$, so that Assumption (ref) holds. This is due to the forms of $L(d;z)$ and $U(d;z)$ above and $\widehat{L}(\tilde{d},d)\le0$ $\forall\tilde{d}\neq d$ if and only if
and, for example, $\widehat{z}(d)=1$ if and only if $\widehat{L}(d;0)-\widehat{L}(d;1) \le0$. A similar formulation follows for the second example. In fact, this approach applies to a general parameter $W(d)$ with bounds that are minimum and maximums of linear combinations of $p(y,d|z)$'s, such as parameters that have the following form:
for some linear functional $f$, where $q_{d}(y)\equiv \mathbb{P}(Y(d)=y)$.
Finally, Assumption (ref) is satisfied in both examples when the data $\{Y_i,D_i,Z_i\}_{i=1}^n$ form a random sample since the entries of $\hat p$ are sample means.
This framework can be further generalized to incorporate treatment choices adaptive to covariates or past outcomes as in Han24, analogous to Section (ref). This generalization is considered in our empirical application in Section (ref). Sometimes this generalization prevents the researcher from deriving analytical bounds, in which case the linear programming approach can be used.
In recent work, CMS23 note that “plug-in” rules for determining treatment choice can be sub-optimal when ATEs are not point-identified since the bounds on the ATE are not smooth functions of the reduced-form parameter $p$. Using an optimality criterion that minimizes maximum regret over the identified set for the ATE, conditional on $p$, they advocate bootstrap and quasi-Bayesian methods for optimal treatment choice. More specifically, they consider settings for which the ATE of a treatment is identified via intersection bounds: \[b_L(p)\equiv \max_{j\in\{1,\ldots,J_L\}}\{\tilde{\ell}_j+\ell_jp\}\leq ATE\leq \min_{j\in\{1,\ldots,J_U\}}\{\tilde{u}_k+u_kp\}\equiv b_U(p)\] for some fixed and known $J_L$, $J_U$, $\tilde{\ell}_1,\ldots,\tilde{\ell}_{J_L}$, $\tilde{u}_1,\ldots,\tilde{u}_{J_U}$, $\ell_1,\ldots,\ell_{J_L}$ and $u_1,\ldots,u_{J_U}$.\footnote{Although CMS23 do not write the form of their bounds as they are written here, the representation here is equivalent to the one in that paper upon proper definition of $p$ since the elements in the intersection bounds are smooth functions of a reduced-form parameter.} Therefore, Assumption (ref) trivially holds for the ATE.
CMS23 advocate a quasi-Bayesian implementation of their optimal treatment choice rule taking the form
for some large $m$, where $\varepsilon_1,\ldots,\varepsilon_m\overset{i.i.d.}\sim\mathcal{N}(0,\widehat \Sigma)$ are independent of $\hat p$. As the following proposition shows, this form of $\hat d$ satisfies Assumption (ref) when $\hat\gamma_L(d)=\hat\gamma_U(d)$ are specified properly.
We now generalize the CI construction described in Section (ref) to apply in the general framework of Section (ref), covering all example applications discussed above and in Appendix (ref). We start with conditional CIs and then move to unconditional CIs.
We first generalize the conditional CI construction described in Section (ref). As in Section (ref), we are interested in forming probabilistic lower and upper bounds $\widehat L(\hat d)_{\alpha}^C$ and $\widehat U(\hat d)_{1-\alpha}^C$ that satisfy (ref) and (ref) for all $d\in\{d^0,\ldots,d^K\}$ as endpoints in the formation of a conditionally valid CI. This is because the researcher's interest in inference on option $d$ only arises when it is a member of the estimated set $\widehat{\mathcal{D}}$.
To begin, we characterize the conditional distributions of $\widehat{L}(d)$ given the event $\{d\in\widehat{\mathcal{D}}$, $\hat j_L(d)=j_L^*$ and $\hat\gamma_L(d)=\gamma_L^*\}$ characterized by Assumption (ref). These conditional distributions depend upon the nuisance parameter $p$. As a first step, we form sufficient statistics for $p$ that are asymptotically independent of $\widehat{L}(d)$ given $\hat j_L(d)=j_L^*$ and $\widehat{U}(d)$ given $\hat j_U(d)=j_U^*$. Since $\widehat{L}(d)=\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p$ given $\hat j_L(d)=j_L^*$ by Assumption (ref) and (ref) (and similarly for $\widehat{U}(d)$), such sufficient statistics can be constructed as \[\widehat{\mathcal{Z}}_{L}(d,j_L^*)\equiv \sqrt{n}\hat p-\hat b_{L}(d,j_L^*)\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p)\] and \[\widehat{\mathcal{Z}}_{U}(d,j_U^*)\equiv \sqrt{n}\hat p-\hat b_{U}(d,j_U^*)\sqrt{n}(\tilde u_{d,j_U^*}+u_{d,j_U^*}\hat p)\] with \[\hat b_{L}(d,j_L^*)\equiv \widehat\Sigma\ell_{d,j_L^*}^{\prime}\left(\ell_{d,j_L^*}\widehat\Sigma\ell_{d,j_L^*}^{\prime}\right)^{-1} \quad \text{and} \quad \hat b_{U}(d,j_U^*)\equiv \widehat\Sigma u_{d,j_U^*}^{\prime}\left(u_{d,j_U^*}\widehat\Sigma u_{d,j_U^*}^{\prime}\right)^{-1},\] by the asymptotic normality of $\hat p$ and asymptotic absence of correlation between $\sqrt{n}\hat p$ and $\widehat{\mathcal{Z}}_{L}(d,j_L^*)$ and $\widehat{\mathcal{Z}}_{U}(d,j_U^*)$. Next, Assumption (ref) characterizes the conditioning events $\{d\in\widehat{\mathcal{D}},\hat j_L(d)=j_L^*,\hat \gamma_L(d)=\gamma_L^*\}$ and $\{d\in\widehat{\mathcal{D}},\hat j_U(d)=j_U^*,\hat \gamma_U(d)=\gamma_U^*\}$ as polyhedra in $\hat p$, which can in turn be expressed as intervals for $\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p$ and $\tilde u_{d,j_U^*}+u_{d,j_U^*}\hat p$; see, e.g., (ref) and (ref). Then, again since $\widehat{L}(d)=\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p$ given $\hat j_L(d)=j_L^*$ by Assumption (ref), Lemma 1 of McC24 implies
and
with
for $M=L,U$.
Now, under Assumptions (ref) and (ref), the distribution of $\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p)$ can be approximated by a $\mathcal{N}(\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*} p),\ell_{d,j_L^*}\widehat\Sigma \ell_{d,j_L^*}')$-distributed random variable that is asymptotically independent of $\widehat{\mathcal{Z}}_{L}(d,j_L^*)$. Using the distributional characterization in (ref), we can therefore use the corresponding truncated normal cumulative distribution function to produce quantile-unbiased estimators of the underlying mean $\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*} p)$. Let $F_{TN}(\cdot;\mu,\sigma^2|\mathcal V^-,\mathcal V^+)$ denote the truncated normal cumulative distribution function for an underlying normally-distributed random variable with mean $\mu$ and variance $\sigma^2$ that is truncated to lie between $\mathcal V^-$ and $\mathcal V^+$. For $\alpha\in(0,1)$, define $\widehat L(d)_{\alpha}^C$ to solve
in $\mu$. Similarly, define $\widehat U(d)_{\alpha}^C$ to solve
in $\mu$. Then, results in Pfa94 imply that $\widehat L(d)_{\alpha}^C$ and $\widehat U(d)_{\alpha}^C$ are optimal $\alpha$ quantile-unbiased estimators of $\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*} p)$ and $\sqrt{n}(\tilde u_{d,j_L^*}+u_{d,j_L^*} p)$ asymptotically.
Finally, combine these quantile-unbiased estimators to form a conditional CI for the identified interval $[L(d),U(d)]$:
We establish the conditional uniform asymptotic validity of this CI.
In parallel with the previous subsection, we now generalize the unconditional CI constructions described in Section (ref) to the general framework of Section (ref). Note that conditional inference on $W(d)$ is well-defined for any given $d\in\widehat{\mathcal{D}}$, when conditioning on $d\in\widehat{\mathcal{D}}$. In contrast, unconditional inference on a data-dependent $W(d)$ requires it to be uniquely defined, as $W(\hat d)$ in our notation. This is implied by Assumption (ref). Here, we would like to construct CIs that unconditionally cover the identified interval corresponding to a unique data-dependent object of inferential interest. As mentioned in Section (ref), if only unconditional coverage of $[L(\hat d),U(\hat d)]$ is desired the conditional CI (ref) with $d=\hat d$ can be unnecessarily wide. We describe two different methods---projection and hybrid methods---to form the unconditional probabilistic bounds that constitute the endpoints of these unconditional CIs in this general framework.
The general formation of the probabilistic bounds based upon projecting simultaneous confidence bounds for all possible values of $\sqrt{n}L(\hat d)$ and $\sqrt{n}U(\hat d)$ proceeds by computing $\hat{c}_{1-\alpha,M}$, the $1-\alpha$ quantile of $\max_{i\in \{1,\ldots,(T+1)J_M\}}\hat \zeta_{M,i}/\sqrt{\widehat{\Sigma}_{M,i}}$, where $\hat \zeta_M\sim\mathcal{N}(0,\widehat{\Sigma}_M)$ for $M=L,U$ with $\widehat{\Sigma}_L=\ell^{mat}\widehat{\Sigma}\ell^{mat\prime}$, $\widehat{\Sigma}_U=u^{mat}\widehat{\Sigma}u^{mat\prime}$, $\ell^{mat}=(\ell_{0,1}',\hdots,\ell_{0,J_L}',\hdots,\ell_{T,1}',\hdots,\ell_{T,J_L}')'$ and $u^{mat}=(u_{0,1}',\hdots,u_{0,J_U}',\hdots,u_{T,1}',\hdots,u_{T,J_U}')'$, recalling that $\Upsilon_i$ denotes the $i^{th}$ element of the main diagonal of any square matrix $\Upsilon$. Here, the maximum is taken to guarantee simultaneous coverage. The lower level $1-\alpha$ projection confidence bound for $\sqrt{n}L(\hat d)$ is $\widehat L(\hat d)_\alpha^P=\sqrt{n}\widehat L(\hat d)-\hat{c}_{1-\alpha,L}\sqrt{\widehat{\Sigma}_{L,\hat dJ_L+\hat{j}_L(\hat d)}}$ and the upper level $1-\alpha$ projection confidence bound for $\sqrt{n}U(\hat d)$ is $\widehat{U}_{1-\alpha}^P(\hat d)=\sqrt{n}\widehat U(\hat d)+\hat{c}_{1-\alpha,U}\sqrt{\widehat{\Sigma}_{U,\hat dJ_U+\hat{j}_U(\hat d)}}$, because e.g., $\sqrt{n}L(\hat d)$ can take value equal to any entry of the vector $\sqrt{n}(\tilde\ell_{0,1},\hdots,\tilde\ell_{0,J_L},\hdots,\tilde\ell_{T,1},\hdots,\tilde\ell_{T,J_L})'+\sqrt{n}\ell^{mat} p$, $\sqrt{n}(\ell^{mat}\hat p-\ell^{mat}p)$ is asymptotically distributed $\mathcal{N}(0,\Sigma_L)$ for $\Sigma_L=\ell^{mat}\Sigma\ell^{mat\prime}$ by Assumption (ref) and $\widehat\Sigma_L$ is consistent for $\Sigma_L$ by Assumption (ref).
Combining these two confidence bounds at appropriate levels yields an unconditional CI for the identified interval $[L(\hat d),U(\hat d)]$ of the selected $\hat d$,
with uniformly correct asymptotic coverage, regardless of how $\hat d$ is selected from the data.
Note that the projection CI (ref) has the benefit of correct coverage regardless of how $\hat d$ is chosen from the data. In this sense, it is more robust than the other CIs we propose in this paper. On the other hand, by using the common selection structure of Assumption (ref), we are able to produce a hybrid CI that combines the strengths of the conditional CI (ref) and the projection CI (ref) which, as described in Section (ref), are shorter under complementary scenarios.
In analogy with the construction of the conditional CIs, to construct the hybrid CIs we begin by characterizing the conditional distributions of $\widehat L(\hat d)$ and $\widehat U(\hat d)$ but now adding an additional component to the conditioning events. More specifically, under Assumptions (ref) and (ref), by intersecting the events
and
for some $0<\beta<\alpha<1$, we have
where \[\widehat{\mathcal{V}}_{L}^{+,H}(z,d,j,\gamma,\mu)\equiv\min\left\{\widehat{\mathcal{V}}_{L}^{+}(z,d,j,\gamma),\mu+\hat{c}_{1-\beta,L}\sqrt{\widehat{\Sigma}_{L,dJ_L+j}}\right\}.\] Similarly,
where \[\widehat{\mathcal{V}}_{U}^{-,H}(z,d,j,\gamma,\mu)\equiv\max\left\{\widehat{\mathcal{V}}_{U}^{-}(z,d,j,\gamma),\mu-\hat{c}_{1-\beta,U}\sqrt{\widehat{\Sigma}_{U,dJ_U+j}}\right\}.\] Since the distribution of $\sqrt{n}(\tilde \ell_{d^*,j_L^*}+\ell_{d^*,j_L^*}\hat p)$ can be approximated by $\mathcal{N}(\sqrt{n}(\tilde \ell_{d^*,j_L^*}+\ell_{d^*,j_L^*} p),\ell_{d^*,j_L^*}\Sigma \ell_{d^*,j_L^*}')$ asymptotically and $\widehat{\mathcal{Z}}_{L}(d^*,j_L^*)$ is asymptotically independent, we again work with the truncated normal distribution to compute a hybrid probabilistic lower bound for $\sqrt{n}L(\hat d)$: for $0<\beta<\alpha<1$, define $\widehat L(d)_{\alpha}^H$ to solve
in $\mu$. Similarly, define $\widehat U(d)_{\alpha}^H$ to solve
in $\mu$.
Here, $\widehat L(\hat d)_{\alpha}^H$ is an unconditionally valid probabilistic lower bound for $L(\hat d)$ and $\widehat U(\hat d)_{1-\alpha}^H$is an unconditionally valid probabilistic upper bound for $U(\hat d)$. Combining these two confidence bounds at appropriate levels yields an unconditional CI for the identified interval $[L(\hat d),U(\hat d)]$ of the selected $\hat d$,
with uniformly correct asymptotic coverage when $\hat d$ is selected from the data according to Assumptions (ref) and (ref).
To our knowledge, the CIs proposed in this paper are the first with proven asymptotic validity for interval-identified parameters selected from the data. Therefore, we have no existing inference method to compare the performance of our CIs to when the interval-identified parameter is data-dependent. However, as discussed in Section (ref) above, our inference framework covers cases for which the interval-identified parameter is chosen a priori. For these special cases, there is a large literature on inference on partially-identified parameters or their identified sets that can be applied to form CIs. Although these special cases are not of primary interest for this paper, in this section we compare the performance of our proposed inference methods to one of the leading inference methods in the partial identification literature as a “reality check” on whether our proposed methods are reasonably informative.
In particular, we compare the power of the test implied by our hybrid CI (i.e., a test that rejects when the value of the parameter under the null hypothesis lies outside of the hybrid CI) to the power of the hybrid test of ARP23, which applies to a general class of moment-inequality models. When $d$ is chosen a priori and the parameter $p$ is equal to a vector of moments of underlying data, Assumption (ref) implies that $L(d)\leq W(d)\leq U(d)$ can be written as a set of (unconditional) moment inequalities, a special case of the general framework of that paper. Of the many papers on inference for moment inequalities, we choose the test of ARP23 for comparison for two reasons: (i) it has been shown to be quite competitive in terms of power and (ii) it is also based upon a (different) inference method that is a hybrid between conditional and projection-based inference.
We compare the power of tests on the ATE in the same setting of the Manski bounds example of Section (ref), strengthening the mean independence assumption $\mathbb{E}[Y(d)|Z]=\mathbb{E}[Y(d)]$ to full statistical independence $(Y(1),Y(0))\perp Z$ and using the sharp bounds on the ATE $W(1)-W(0)=\mathbb{E}[Y(1)]-\mathbb{E}[Y(0)]$ derived by BP97,BP13:
and
For a sample size of $n=100$, we generated $\hat p$ from a $\mathcal{N}(p,\Sigma)$ distribution.\footnote{Note that in this problem, the value of $\Sigma$ is implied by the value of $p$.} Figure (ref) plots the power curves of the hybrid ARP23 test and the test implied by our hybrid CI for three different DGPs within the framework of this example, as well as the true identified interval for the ATE. The DGP corresponding to $p=(.08,.001,.001,.073,.139,.473)'$ is calibrated to the probabilities estimated by BP13 in the context of a treatment for high-cholesterol (specifically, by the drug cholestyramine).\footnote{BP13 estimate $p$ to equal $(.081,0,0,.073,.139,.473)'$. If the true DGP is set exactly equal to this, Assumption (ref) would be violated. We therefore slightly alter the calibrated probabilities.} The DGP corresponding to $p=(.25,.25,.25,.25,.25,.25)'$ generates completely uninformative bounds for the ATE. And the DGP corresponding to $p=(.01,.44,.01,.01,.01,.54)'$ generates quite informative bounds.
We can see that overall, the power of the test implied by our hybrid CI is quite competitive with that of ARP23. Interestingly, it appears that our test tends to be more powerful than that of ARP23 when the true ATE is larger than the hypothesized one, which can be seen from the hypothesized ATE lying to the left of the lower bound of the identified interval, and vice versa. Although the main innovation of our inference procedures is really their validity in the presence of data-dependent selection, the exercise of this section is reassuring for the informativeness of the procedures we propose as they are quite competitive in the absence of selection.
Moving now to a context in which the object of interest is selected from the data, we compare the finite sample performance of our conditional, projection and hybrid CIs again in the setting of the Manski bounds example of Section (ref). In this case, we are interested in inference on the average potential outcome $W(\hat d)$ for $W(d)=\mathbb{E}[Y(d)]$, where interest arises either in the average potential outcome for treatment ($\hat d=1$) or control ($\hat d=0$) depending upon which has the largest estimated lower bound: $\hat d=\operatorname*{argmax\;}_{d\in\{0,1\}}\widehat L(d)$. This form of $\hat d$ corresponds to case 1. of Proposition (ref) and we use the corresponding result of the proposition to specify $\hat\gamma_L(\hat d)$ and $\hat\gamma_U(\hat d)$ in the construction of the conditional and hybrid CIs. We report analogous simulation results for the dynamic treatment regime example in Appendix (ref) with $\hat p$ generated from a multinomial, rather than normal, distribution.
For the same DGPs as in Section (ref), we compute the unconditional coverage frequencies of the conditional, projection and hybrid CIs as well as that of the conventional CI based upon the normal distribution. These coverage frequencies are reported in Table (ref). Consistent with the asymptotic results of Theorems (ref), (ref) and (ref), the conditional, projection and hybrid CIs all have correct coverage for all DGPs and the modest sample size of $n=100$. Also consistent with Theorem (ref), we note that the projection CI tends to be conservative with true coverage above the nominal level of 95%. Finally, we note that the conventional CI can substantially under-cover, consistent with the discussion in Section (ref).
Next, we compare the length quantiles of the CIs with correct coverage for these same DGPs. Figure (ref) plots the ratios of the $5^{th}$, $25^{th}$, $50^{th}$, $75^{th}$ and $95^{th}$ quantiles of the length of the conditional, projection and hybrid CIs relative to those same length quantiles of the projection CI. As can be seen from the figure, the conditional CI has the tendency to become very long, especially at high quantile levels for certain DGPs, whereas the hybrid CI tends to perform the best overall by limiting the worst-case length performance of the conditional CI relative to the projection CI. Relative to projection, the hybrid CI enjoys length reductions of 20-30% for favorable DGPs while only showing length increases of 5-10% for unfavorable DGPs.
We revisit Han24's (Han24) application. Han24 considers schooling and post-school training as a sequence of treatments and estimates the partial ordering of treatment regimes and the identified set of the optimal regime. Han24 considers the Job Training Partnership Act (JTPA) program for post-school training and a high school diploma (HS) (or its equivalents) for schooling. He considers high school diplomas rather than college degrees because the former is more relevant for the disadvantaged population of Title II of the JTPA program. In this paper, we are interested in conducting inference on welfare---earnings---evaluated at the regime chosen in a data-driven manner. The dataset is constructed by combining the JTPA data with the US Census and the National Center for Education Statistics (NCES). The following is the set of variables: $Y_{2}$ is an indicator for whether the individual is above or below the median of 30-month earnings, $D_{2}$ is an indicator for whether the individual participated in the job training program, $Z_{2}$ is an indicator for whether the individual was (randomly) assigned to the program, $Y_{1}$ is an indicator for whether the individual is above or below the 80th percentile of pre-program earnings, $D_{1}$ is an indicator for whether the individual received a HS diploma or GED, and $Z_{1}$ is an indicator for the density of high schools Nea97.\footnote{Specifically, $Z_{1}=1$ if the number of high schools per square mile in each training site (e.g. a city) is above 35.} The number of individuals in the sample is 9,223. We assume $Z\bot(Y(d),D(z))$.
Following Han24, consider the dynamic treatment regime $\boldsymbol{\delta}(\cdot)\equiv(\delta_{1},\delta_{2}(\cdot))\in \mathcal{D}^*$, where $\delta_{1}$ indicates receipt of a HS diploma and $\delta_{2}(y_{1})$ indicates receipt of the job training program given pre-program earning type $y_{1}$. By having $\delta_{2}$ as a function of $y_{1}$, the allocation decision adaptively incorporates information about unobserved characteristics of the individuals reflected in the response $Y_{1}$ to allocation $\delta_{1}$. The counterfactual earning type in the terminal stage given $\boldsymbol{\delta}(\cdot)$ is defined as $Y_{2}(\boldsymbol{\delta}(\cdot))\equiv Y_{2}(\delta_{1},\delta_{2}(Y_{1}(\delta_{1})))$, where $Y_{1}(\delta_{1})$ is the counterfactual earning type in the first stage given $\delta_{1}$. All possible regimes in $\mathcal{D}^*$ are listed in Table (ref). Suppose $\boldsymbol{\delta}^{*}$ is the optimal regime that maximizes the average terminal earning $W(\boldsymbol{\delta})\equiv \mathbb{E}[Y_{2}(\boldsymbol{\delta}(\cdot))]$ as welfare. We are interested in constructing CIs for $W(\boldsymbol{\delta})$ evaluated at $\boldsymbol{\delta}=\hat{\boldsymbol{\delta}}$, where $\hat{\boldsymbol{\delta}}$ is calculated from the estimated identified set of $\boldsymbol{\delta}^{*}$.
We first derive analytical bounds on the welfare under no additional assumptions. The distribution of data is expressed as the vector $p$ of the form
The welfare $W(\boldsymbol{\delta})\equiv \mathbb{E}[Y_{2}(\boldsymbol{\delta}(\cdot))]$ can be expressed as
by the law of iterated expectation. To derive bounds on $W(\boldsymbol{\delta})$, first consider bounds on $W_{y}(d)\equiv \mathbb{P}[Y(d)=y]$ for $d\equiv(d_{1},d_{2})$, which are $L_{y}(d)\equiv\max_{z}L_{y}(d;z)$ and $U_{y}(d)\equiv\min_{z}U_{y}(d;z)$ where
Using these bounds, we can calculate bounds on
which are
For example, for the fourth regime in Table (ref),
is bounded by
Since $\max_{z}L(z)+\max_{z}\tilde{L}(z)=\max_{z,\tilde{z}}\left\{ L(z)+\tilde{L}(\tilde{z})\right\} $ for any functions $L$ and $\tilde{L}$, we can express (ref) as
and, analogously, (ref) as $U(\boldsymbol{\delta})\equiv\max_{z,\tilde{z}}U(\boldsymbol{\delta};z,\tilde{z})$. Therefore the bounds satisfy Assumption (ref). Now, consider choosing a single optimal regime that maximizes $L(\boldsymbol{\delta})$. Then, for example, we can show that the following event can be characterized as a polyhedron in the space of $p$
where $(z^{*},\tilde{z}^{*})=\arg\max_{z,\tilde{z}}L(\boldsymbol{\delta}_{(4)};z,\tilde{z})$. Note that this event is equivalent to $\{L(\boldsymbol{\delta}_{(4)};z^{*},\tilde{z}^{*}) \ge L(\boldsymbol{\delta};z,\tilde{z})\text{ for all }\boldsymbol{\delta}\text{ and }z,\tilde{z}\}$ or equivalently,
Note that each $L(\boldsymbol{\delta};z,\tilde{z})$ is a linear combination of the elements in $p$ as shown in (ref) and (ref). More formally, we can show that (ref) is equivalent to $A_{L}p \le0$ for some $A_{L}$. Note that
for some row vectors $A^{1}$ and $A^{0}$ because, for $\boldsymbol{\delta}=\boldsymbol{\delta}_{(4)}$,
An analogous argument can be applied to $U(\boldsymbol{\delta};z,\tilde{z})$. This characterization implies that Assumption (ref) holds in this setting. Also, this characterization facilitates the CI calculations in Sections (ref)--(ref).
Table (ref) reports 95% CIs for the welfare selected by maximizing the estimated lower bound on the welfare. Recall that the welfare is the probability that the 30-month earnings is above the median. It is notable that the hybrid CI is shorter than the conventional CI even though the latter does not have (uniformly) correct coverage. We can also see that although the hybrid CI is not quite contained in the projection CI, it is somewhat shorter. Finally, the conditional CI has infinite length in this example, demonstrating an extreme example of how the conditional CIs can be uninformatively long (see KL21).
Appendix (ref) contains Monte Carlo simulated coverage frequencies of the CIs with various DGPs in the setting of this application. Overall, the findings are consistent with the ones in Section (ref).
This appendix contains examples in addition to Section (ref), showing how they fall under the general framework of this paper.
We revisit the example with Manski bounds in Section (ref), now allowing for a continuous outcome with bounded support. Let $Y\in[y^{l},y^{u}]$ be continuously distributed and assume $\mathbb{E}[Y(d)|Z]=\mathbb{E}[Y(d)]$ for $d\in\{0,1\}$. Note that the sharp bounds on $W(d)\equiv \mathbb{E}[Y(d)]$ are
and the sharp bounds on the ATE are $L(1,0)$ and $U(1,0)$ for $L(\tilde{d},d) \equiv L(\tilde{d})-U(d)$ and $U(\tilde{d},d) \equiv U(\tilde{d})-L(d)$. We can define the identified set $\mathcal D^* \subseteq \mathcal D$ of optimal treatments as
Then, Assumption (ref) holds with
and $\tilde{\ell}_{0,j} =y^{l}$ and $\tilde{\ell}_{1,j} =0$ for $j\in\{0,1\}$ and
and symmetrically for $\tilde{u}_{d,j}$ and $u_{d,j}$. For $\widehat{L}(d)$ and $\widehat{U}(d)$ being the sample counterparts of $L(d)$ and $U(d)$ upon replacing $p$ with $\hat p$, Assumption (ref) holds by the argument in the paragraph after Assumption (ref) since
For Assumption (ref), let $[y^{l},y^{u}]=[0,1]$ for simplicity. Note that
where the first equality uses integration by parts. Therefore, we can estimate the elements of $p$ by sample means, forming $\hat{p}$.
Continuing from Section (ref), we show how sharp bounds on $W(\delta)$ can be computed using linear programming. This can be done by extending the example in Section (ref) with binary $Y$. Again, let $\mathcal{X}=\{x_{1},...,x_{K}\}$. Let $q(e|x)\equiv \mathbb{P}(\varepsilon=e|X=x)$. In analogy,
where $q(x)$ is a vector with entries $q(e|x)$ across $e$, and
where the first equality holds by the independence assumption in the previous section. Then the constraint for each $x\in\mathcal{X}$ becomes
where $p(x)$ is the vector of $p(y,d|z,x)$'s across $(y,d,z)$ fixing $x$. Now we can construct a linear program for welfare:
Therefore $W(\delta)$ satisfies the structure of Section (ref) and by analogous arguments, Assumptions (ref), (ref) and (ref) hold.
In analogy with Section (ref), we compute the unconditional coverage frequencies of the conditional, projection and hybrid CIs for DGPs in the dynamic treatment regime setting of the empirical application (Section (ref)). In particular, we consider two DGPs: DGP 1 generates $\hat p$ from a multinomial distribution based on $p_{d_1,y_1,d_2,y_2|z_1,z_2}=0.25$ for $(d_1,y_1,d_2,y_2)\in\{(1,0,0,1),(1,1,1,1)\}$ and all $(z_1,z_2)$ and $p_{d_1,y1,d_2,y_2|z_1,z_2}=0.0357$ for all other $(d_1,y_1,d_2,y_2)$ and all $(z_1,z_2)$; DGP 2 generates $\hat p$ based on $p_{d_1,y_1,d_2,y_2|z_1,z_2}=0.375$ for $(d_1,y_1,d_2,y_2)\in\{(1,0,0,1),(1,1,1,1)\}$ and all $(z_1,z_2)$ and $p_{d_1,y1,d_2,y_2|z_1,z_2}=0.0179$ for all other $(d_1,y_1,d_2,y_2)$ and all $(z_1,z_2)$. The coverage frequencies of the CIs are reported in Table (ref). Again, consistent with the asymptotic results of Theorems (ref)--(ref), the conditional, projection and hybrid CIs all have correct coverage for all DGPs and the sample size of $n=1000$. Also, the projection CI tends to be conservative with true coverage above the nominal level of 95%, and the conventional CI can substantially under-cover.
Figure (ref) plots the ratios of the $5^{th}$, $25^{th}$, $50^{th}$, $75^{th}$ and $95^{th}$ quantiles of the length of the conditional, projection and hybrid CIs relative to those same length quantiles of the projection CI. The figure shows that the conditional CI has the tendency to become very long, especially at high quantile levels for certain DGPs, whereas the hybrid CI tends to perform the best overall by limiting the worst-case length performance of the conditional CI relative to the projection CI. Relative to projection, the hybrid CI enjoys length reductions of 10-20% for favorable DGPs while showing length increases of 5-10% for unfavorable DGPs.