EconBase
← Back to paper

Inference for Interval-Identified Parameters Selected from an Estimated Set

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

109,217 characters · 21 sections · 76 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference for Interval-Identified Parameters Selected from an Estimated Set

abstractInterval identification of parameters such as average treatment effects, average partial effects and welfare is particularly common when using observational data and experimental data with imperfect compliance due to the endogeneity of individuals' treatment uptake. In this setting, the researcher is typically interested in a treatment or policy that is either selected from the estimated set of best-performers or arises from a data-dependent selection rule. In this paper, we develop new inference tools for interval-identified parameters chosen via these forms of selection. We develop three types of confidence intervals for data-dependent and interval-identified parameters, discuss how they apply to several examples of interest and prove their uniform asymptotic validity under weak assumptions. Keywords: Partial Identification, Post-Selection Inference, Selective Inference, Conditional Inference, Uniform Validity, Treatment Choice.

Introduction

There is now a large and growing literature on partial identification of optimal treatments and policies under practically-relevant assumptions for observational data and experimental data with imperfect compliance (e.g., Sto12,KZ21, PZ21, DAd21,Yat21,Han24 to name just a few). Interval identification of average treatment effects (ATEs), average partial effects and welfare is particularly common in these settings due to the endogeneity of individuals' treatment uptake. In order to use the partial identification results for treatment or policy choice in practice, a researcher must typically estimate a set of best-performing treatments or policies from data. Consequently, the researcher is typically interested in a treatment or policy that is either selected from the estimated set of best-performers or arises from a data-dependent selection rule. It is now well-known that selecting an object of interest from data invalidates standard inference tools (e.g., AKM24). The failure of standard inference tools after data-dependent selection is only compounded by the presence of partially-identified parameters.

In this paper, we develop new inference tools for interval-identified parameters corresponding to selection from either an estimated set or arising from a data-dependent selection rule. Estimating identified sets for the best-performing treatments/policies or forming data-dependent selection rules in these settings is important for choosing which treatments/policies to implement in practice. Therefore, the ability to infer how well these treatments or policies should be expected to perform when selected, for instance to gauge whether their implementation is worthwhile, is of primary practical importance.

The current literature has not yet developed valid post-selection inference tools in partially-identified contexts, an important deficiency that the methods proposed in this paper aim to correct. The methods we propose here build upon the ideas of conditional and hybrid inference employed in various point-identified contexts by, e.g., LSST16, FST17, TRTW18, AKM24 and McC24 to produce confidence intervals (CIs) for interval-identified parameters such as welfare or ATEs chosen from an estimated set or via a data-dependent selection rule. Although ARP23 also propose conditional and hybrid inference methods in the partial identification context of moment inequality models, they do not allow for data-dependent selection of objects of interest, one of the main focuses of the present paper. Finally, this paper directly relies upon results in the literature on interval identification of welfare, treatment effects and partial outcomes such as Man90, BP97,BP13, MP00, MST18, HY24 and Han24. We apply our inference methods to a general class of problems nesting these examples.

After sketching the ideas behind our inference methods in a simple example, we introduce the general inference framework to which our methods can be applied. We show that our general inference framework incorporates several problems of interest for data-dependent selection and treatment rules for parameters belonging to an identified set such as Manski bounds for average potential outcomes or ATEs, bounds on parameters derived from linear programming (e.g., BP97,BP13; MST18; Han24; HY24), bounds on welfare for treatment allocation rules that are partially identified by observational data (e.g., Sto12,KZ21, PZ21, DAd21,Yat21) and bounds for dynamic treatment effects (e.g., Han24). We also show how to incorporate inference on parameters chosen via asymptotically optimal treatment choice rules (e.g., CMS23) into our general inference framework. Our framework can also be applied to settings where welfare is partially identified for reasons other than treatment endogeneity (e.g., IK21,AC22,BGIJ22,CH24).

Within the general inference framework, we develop three types of CIs for data-dependent and interval-identified parameters. As the name suggests, the conditional CIs are asymptotically valid conditional on the parameter of interest corresponding to a treatment or policy chosen from an estimated set. The construction of this CI does not require a specific rule for choosing the parameter of interest from the estimated set. In addition, the sampling framework underlying its conditional validity is most appropriate in contexts for which a researcher will only be interested in the parameter because it is chosen from the estimated set. Importantly, we show that these CIs are asymptotically valid uniformly across a large class of data-generating processes (DGPs). Uniform asymptotic validity is especially important for approximately correct finite-sample coverage in post-selection contexts like those in this paper (see, e.g., AG09).

The second and third types of CIs we develop in this paper are designed for inference on parameters chosen by data-dependent selection rules for which the object of interest is uniquely determined by the selection rule. The projection CIs do not require knowledge of the selection rule to be asymptotically valid, whereas the hybrid CI construction utilizes the particular form of a selection rule to improve upon the length properties of the projection CI. The conditional CIs are short for selections that occur with high probability but can become exceptionally long when this probability is small (see, e.g., KL21). Conversely, projection CIs are overly conservative when selection probabilities are high. Although conditional CIs can be used in this setting, hybrid CIs interpolate the length properties of the conditional and projection CIs in order to attain good length properties regardless of the value selection probabilities take. In analogy with the conditional CIs, we formally show that both projection and hybrid CIs are asymptotically valid in a uniform sense.

We analyze the coverage and length properties of our proposed CIs in finite samples. Since, to our knowledge, these are the first uniformly valid CIs for data-dependent selections of partially-identified parameters, there are no existing CIs to which we can directly compare. Nevertheless, since our CIs can also be used for inference on a priori chosen interval-identified parameters, we conduct a power comparison with one of the leading methods for inference on a partially-identified parameter. In particular, we compare the power of the test implied by our hybrid CIs to the power of the hybrid test of ARP23, a test that applies to a general class of moment-inequality models that is also based upon a (different) hybrid between conditional and projection-based inference. Encouragingly, the power of the test implied by our hybrid CI is quite competitive even in this environment for which it was not designed. We also find that the finite-sample coverage of all of our CIs is approximately correct in a simple Manski bound example. Finally, we analyze the length tradeoffs between the three different CIs across different DGPs, finding the hybrid CI to perform best overall.

The remainder of this paper is structured as follows. Section (ref) sketches the ideas behind our general CI constructions in the context of a simple Manski bound example. Section (ref) lays down the general high-level inference framework we are interested in, while Section (ref) details how the general framework applies in several different examples. Section (ref) then details the various CI constructions in the general setting. Sections (ref) and (ref) are devoted to finite-sample comparisons of the properties of the different CIs in the context of a simple Manski bound example. The final section, Section (ref), contains an empirical application where we apply our procedures to dynamic policies of schooling and post-school training. Appendix (ref) contains additional examples that fit our general framework that are not covered in Section (ref), while Appendix (ref) contains additional simulation results corresponding to the dynamic treatment regime example detailed in Section (ref). Mathematical proofs are relegated to a Technical Appendix at the end of the paper.

Basic Ideas: Inference with Manski Bounds Example

We first provide a simple example to illustrate our proposed methods. Consider a binary outcome of interest $Y$, a binary treatment indicator $D$ and a binary treatment assignment $Z$. Furthermore, let $Y(1)$ and $Y(0)$ denote potential outcomes under treatment ($D=1$) and no treatment ($D=0$). Assuming $\mathbb{E}[Y(d)|Z]=\mathbb{E}[Y(d)]$, Man90 shows that we can bound the average potential outcomes $W(d)=\mathbb{E}[Y(d)]$ in the absence and presence of treatment as follows:

equation[equation omitted — 150 chars of source]

for $d=0,1$ and $p^{ydz} \equiv Pr(Y=y,D=d|Z=z)$. Given (ref) it is natural to define the set of best-performing options $\mathcal{D}^{*}$, as a subset of the two options of treatment and no treatment, to be those that are undominated options. From an observed dataset of outcomes, treatments and treatment assignments $\{(Y_i,D_i,Z_i)\}_{i=1}^n$, such a set can be estimated as: \[ \widehat{\mathcal{D}}=\left\{d\in\{0,1\}:\widehat U(d)\geq \widehat L(d^ \prime)\text{ }\forall d^{ \prime}\in\{0,1\} \text{ s.th. }d^{ \prime}\neq d\right\}, \] where $\widehat L(d)\equiv \max\{\hat p^{1d0},\hat p^{1d1}\}$ and $\widehat U(d)\equiv\min\{1-\hat p^{0d0},1-\hat p^{0d1}\}$ with $\hat p^{ydz}$ being an empirical estimate of the fitted probability $p^{ydz} \equiv \mathbb{P}(Y=y,D=d|Z=z)$.

We are interested in inference on the identified interval $[L(d),U(d)]$ for the average potential outcome $W(d)$ of option $d\in\{0,1\}$ after the researcher selects this option from $\widehat{\mathcal{D}}$. In other words, we would like to provide statistically precise statements about the true average potential outcome of an option selected from the data to give the researcher an idea of how well this selected option should be expected to perform in the population. We first provide some intuition for why standard inference techniques based upon asymptotic normality fail and then sketch our proposals for valid inference in this context.

Why Does Standard Inference Fail?

To fix ideas, let us focus for now on inference for the lower bound $L(d)$ of a selected option rather than the entire identified set $[L(d),U(d)]$. More specifically, since $L(d)$ is a lower bound for the average potential outcome $W(d)$, we would like to obtain a probabilistic lower bound for $L(d)$. Under standard conditions, a central limit theorem implies $\hat p=(\hat p^{100},\hat p^{010},\hat p^{110},\hat p^{101},\hat p^{011},\hat p^{111})^\prime$ is normally distributed in large samples. So why not form a CI using $\widehat L(d)$ and quantiles from a normal distribution as the basis for inference? There are two reasons such an approach is (asymptotically) invalid:

enumerate• Even in the absence of selection, $\widehat L(d)= \max\{\hat p^{1d0},\hat p^{1d1}\}$ is not normally distributed in large samples. • Data-dependent selection of $d$ further complicates the distribution of $\widehat L(d)$.

Reason 1. is easy to see since $\widehat L(d)$ is the maximum of two normally distributed random variables in large samples when $d$ is chosen a priori. To better understand reason 2., note that the distribution of $\widehat L(d)$ given $d\in \widehat{\mathcal{D}}$ is the conditional distribution of the maximum of two normally distributed random variables given that the minimum of two other normally distributed random variables, $\widehat U(d)\equiv\min\{1-\hat p^{0d0},1-\hat p^{0d1}\}$, exceeds the maximum of yet another set of two normally distributed random variables, $\widehat L(d')=\max\{\hat p^{1d'0},\hat p^{1d'1}\}$ for $d'\neq d$. Unconditionally, $\widehat L(\hat d)$ for any data-dependent choice of $\hat d$ is distributed as a mixture of the distributions of $\widehat L(0)$ and $\widehat L(1)$, neither of which are themselves normally distributed.

Conditional Confidence Intervals

Suppose that a researcher's interest in inference on $L(d)$ only arises when $d$ is estimated to be in the set of best-performing options, viz., $d\in \widehat{\mathcal{D}}$. In such a case, we are interested in a probabilistic lower bound for $L(d)$ that is approximately valid across repeated samples for which $d\in \widehat{\mathcal{D}}$, i.e., we would like to form a conditionally valid lower bound $\widehat L(d)_\alpha^C$ such that\footnote{See AKM24 for an extensive discussion of when conditional vs unconditional validity is desirable for inference after selection.}

equation[equation omitted — 157 chars of source]

for some $\alpha\in(0,1)$ in large samples. To do so we characterize the conditional distribution of $\widehat L(d)$. Specifically, let $\hat j_L(d)\equiv \operatorname*{argmax\;}_{j\in\{0,1\}}\hat p^{1dj}$ be the value of $Z$ at which the maximum between the two estimated probabilities is achieved. Then $\widehat L(d)=\hat p^{1d\hat j_L(d)}$. Also, since the conditioning event $\{d\in \widehat{\mathcal{D}},\hat j_L(d)=j_L^*\}$ can be written as a polyhedron in $\hat p$ and $\widehat L(d)$ is equal to an element of $\hat p$, Lemma 5.1 of LSST16 implies

equation[equation omitted — 289 chars of source]

for some known functions $\widehat{\mathcal{L}}_L(\cdot)$ and $\widehat{\mathcal{U}}_L(\cdot)$, where $\hat p^{1dj_L^*}\sim\mathcal{N}(p^{1dj_L^*},\operatorname*{Var}(\hat p^{1dj_L^*}))$ in large samples and $\mathcal{Z}_L$ is a sufficient statistic for the nuisance parameter $p=( p^{100}, p^{010}, p^{110}, p^{101}, p^{011}, p^{111})^\prime$ that is asymptotically independent of $\hat p^{1dj_L^*}$. Using results in Pfa94, the characterization in (ref) permits the straightforward computation of a conditionally quantile-unbiased estimator for $p^{1d\hat j_L(d)}=p^{1dj_L^*}$, since the latter is equal to the mean of the underlying normally distributed random variable $\hat p^{1dj_L^*}$ that is subject to truncation. Denoting this quantile-unbiased estimator as $\widehat L(d)_{\alpha}^C$, we have

equation[equation omitted — 178 chars of source]

in large samples. However, noting that $L(d)\geq p^{1d\hat j_L(d)}$ with probability one, we can see that (ref) holds for this choice of $\widehat L(d)_{\alpha}^C$.

Although (ref) does not hold with exact equality, we note that the left-hand side cannot be much larger than the right-hand side. In other words, although $\widehat L(d)_{\alpha}^C$ is a conservative probabilistic lower bound for $L(d)$, it is not very conservative. This can be seen heuristically by working through the two possible values that $L(d)$ can take:

enumerate• If $L(d)=p^{1d\hat j_L(d)}$, then (ref) holds with equality in large samples by (ref). • If $L(d)\neq p^{1d\hat j_L(d)}$, then $L(d)\approx p^{1d\hat j_L(d)}$ since $\widehat L(d)=\hat p^{1d\hat j_L(d)}$ so that the left-hand side of (ref) cannot be much larger than the right-hand side.

Finally, a construction analogous to that described above for producing a probabilistic lower bound for $L(d)$ produces a conditionally valid probabilistic upper bound $\widehat U(d)_{1-\alpha}^C$ for $U(d)$ that satisfies

equation[equation omitted — 162 chars of source]

for some $\alpha\in(0,1)$ in large samples. The probabilistic lower and upper bounds can then be combined to form a CI, $[\widehat L(d)_{\alpha/2}^C,\widehat U(d)_{1-\alpha/2}^C]$, that is conditionally valid for $[L(d),U(d)]$ in large samples since

gather[gather omitted — 414 chars of source]

Unconditional Confidence Intervals

Suppose now that the researcher uses a data-dependent rule to select a unique option of inferential interest. For example, suppose the researcher is interested in choosing the option with the highest potential outcome in the worst case across its identified set so that she chooses $\hat d=\operatorname*{argmax\;}_{d\in\{0,1\}} \widehat L(d)$. In such a case, it is natural to form a probabilistic lower bound $\widehat L(\hat d)_{\alpha}^U$ for $L(\hat d)$ that is unconditionally valid across repeated samples such that

equation[equation omitted — 130 chars of source]

for some $\alpha\in(0,1)$ in large samples. Given its conditional validity (ref), the conditional lower bound $\widehat L(d)_{\alpha}^C$ also satisfies (ref) upon changing the definition of $\widehat{\mathcal{D}}$ to $\widehat{\mathcal{D}}=\{\hat d\}$ in its construction. However, it is well known in the literature on selective inference that conditionally-valid probabilistic bounds can be very uninformative (i.e., far below the true value) when the probability of the conditioning event is small (see e.g., KL21, AKM24 and McC24). Here, we propose two additional forms of probabilistic bounds that are only unconditionally valid but do not suffer from this drawback.

First, we can form a probabilistic lower bound for $L(\hat d)$ by projecting a one-sided rectangular simultaneous confidence lower bound for all possible values $L(\hat d)$ can take: $\widehat L(\hat d)_\alpha^P \equiv\widehat L(\hat d)-\hat c_{1-\alpha,L}\sqrt{\widehat\Sigma_{L,2\hat d+1+\hat j_L(\hat d)}},$ where $\hat c_{1-\alpha,L}$ is the $1-\alpha$ quantile of $\max_i\hat\zeta_i/\sqrt{\widehat\Sigma_{L,i}}$ for $\hat\zeta\sim\mathcal{N}(0,\widehat\Sigma_L)$, $\widehat\Sigma_L$ is a consistent estimator of $\Sigma_L\equiv\operatorname*{Var}(\hat p^{100},\hat p^{101},\hat p^{110},\hat p^{111})$ and $\Upsilon_i$ denotes the $i^{th}$ element of the main diagonal of any square matrix $\Upsilon$. Here, the maximum is taken to guarantee simultaneous coverage of all possible values of $L(\hat d)$. Since $p^{1\hat d\hat j_L(\hat d)}\in\{\hat p^{100},\hat p^{101},\hat p^{110},\hat p^{111}\}$ with probability one,

gather*[gather* omitted — 459 chars of source]

in large samples and (ref) holds for $\widehat L(\hat d)_\alpha^U=\widehat L(\hat d)_\alpha^P$ because $L(\hat d)\geq p^{1\hat d\hat j_L(\hat d)}$. However, $\widehat L(\hat d)_\alpha^P$ suffers from a converse drawback to that of $\widehat L(\hat d)_\alpha^C$: it is unnecessarily conservative when $\hat d=d$ is chosen with high probability (see e.g., AKM24 and McC24).

We propose a second probabilistic lower bound for $L(\hat d)$ that combines the complementary strengths of $\widehat L(d)_\alpha^C$ and $\widehat L(\hat d)_\alpha^P$. Construction of this hybrid lower bound $\widehat L(\hat d)_\alpha^H$ proceeds analogously to the construction of $\widehat L( d)_\alpha^C$ after adding the additional condition $\{p^{1\hat d\hat j_L(\hat d)}\geq \widehat L(\hat d)_\beta^P\}$ for $\beta<\alpha$ to the conditioning event and instead computing a conditionally quantile-unbiased estimator for $p^{1\hat d\hat j_L(\hat d)}$, denoted as $\widehat L(\hat d)_{\alpha}^H$, satisfying

equation*[equation* omitted — 209 chars of source]

in large samples, where $d^*$ is any realized value of the random variable $\hat d$. Imposing this additional condition in the formation of the hybrid bound ensures that $\widehat L(\hat d)_{\alpha}^H$ is always greater than $\widehat L(\hat d)_\beta^P$, limiting its worst-case performance relative to $L(\hat d)_\beta^P$ when $\mathbb{P}(\hat d=d^*)$ is small. On the other hand, when $\mathbb{P}(\hat d=d^*)$ is large, the additional condition $\{p^{1\hat d\hat j_L(\hat d)}\geq \widehat L(\hat d)_\beta^P\}$ is far from binding with high probability so that $\widehat L(\hat d)_\alpha^H$ becomes very close to $\widehat L(d)_{(\alpha-\beta)/(1-\beta)}^C$. In this case, $\widehat L(d)_{(\alpha-\beta)/(1-\beta)}^C$ is close to the naive lower bound based upon the normal distribution $\widehat L(\hat d)-z_{(1-\alpha)/(1-\beta)}\sqrt{\operatorname*{Var}(\hat p^{1dj_L^*})}$ because the truncation bounds in (ref) are very wide (Proposition 3 in AKM24).

To see how (ref) holds for $\widehat L(\hat d)_\alpha^U=\widehat L(\hat d)_\alpha^H$, note first that

equation*[equation* omitted — 361 chars of source]

for all $d^*\in\{0,1\}$. Then, note that

gather*[gather* omitted — 360 chars of source]

by the law of total probability.

By similar reasoning to that used for the conditional CIs in Section (ref) above, $\widehat L(\hat d)_{\alpha}^H$ is not very conservative as a probabilistic lower bound for $L(\hat d)$. The researcher's choice of $\beta\in(0,\alpha)$ trades off the performance of $\widehat L(\hat d)_{\alpha}^H$ across scenarios for which $\mathbb{P}(\hat d=d^*)$ is large and small with a small $\beta$ corresponding to better performance when $\mathbb{P}(\hat d=d^*)$ is large. See McC24 for an in-depth discussion of these tradeoffs. We recommend $\beta=\alpha/10$.

Finally, analogous constructions to those above produce unconditional projection and hybrid probabilistic upper bounds $\widehat U(\hat d)_{1-\alpha}^P$ and $\widehat U(\hat d)_{1-\alpha}^H$ that can then be combined with the lower bounds to form CIs $[\widehat L(d)_{\alpha/2}^P,\widehat U(d)_{1-\alpha/2}^P]$ and $[\widehat L(d)_{\alpha/2}^H,\widehat U(d)_{1-\alpha/2}^H]$ for $[L(d),U(d)]$ that are unconditionally valid in large samples by the same arguments as those used in (ref) above.

General Inference Framework

We now introduce the general inference framework that we propose, nesting the Manski bound example of the previous section as a special case. After introducing the general framework, we describe several additional example applications that fall within this framework.

We are interested in performing inference on a parameter $W(d)$ that is indexed by a finite set $d\in \mathcal D \equiv \{d^0,\ldots,d^K\}$ for some $K>0$. The index $d$ may correspond to a particular treatment, treatment allocation rule or policy, depending upon the application. We assume that $W(d)$ belongs to an identified set taking a particular interval form that is common to many applications of interest.

assumptionFor all $d\in\{d^0,\ldots,d^K\}$ and an unknown finite-dimensional parameter $p$, \begin{enumerate} • $L(d)\equiv \max_{j\in \{1,\ldots,J_L\}}\{\tilde\ell_{d,j}+\ell_{d,j}p\}\leq W(d)$ for some fixed and known $J_L$, $\tilde\ell_{d,1},\ldots,\tilde\ell_{d,J_L}$ and nonzero row vectors $\ell_{d,1},\ldots,\ell_{d,J_L}$ such that $\ell_{d,j}\neq\ell_{d,j'}$ for $j\neq j'$. • $U(d)\equiv \min_{j\in \{1,\ldots,J_U\}}\{\tilde u_{d,j}+u_{d,j}p\}\geq W(d)$ for some fixed and known $J_U$, $\tilde u_{d,1},\ldots,\tilde u_{d,J_U}$ and nonzero row vectors $u_{d,1},\ldots,u_{d,J_U}$ such that $u_{d,j}\neq u_{d,j'}$ for $j\neq j'$. \end{enumerate}

The lower and upper endpoints of identified sets for the welfare, average potential outcome or ATE typically take the form of $L(d)$ and $U(d)$, especially when (sequences of) outcomes, treatments and instruments are discrete.

In the setting of this paper, a researcher's interest in $W(d)$ arises when $d$ belongs to a set $\widehat{\mathcal{D}}\subset \mathcal D$ that is estimated from a sample of $n$ observations. It is often the case that $\widehat{\mathcal{D}}$ is an estimate of the identified set of best performers $\mathcal{D}^{*}$. This set could correspond to an estimated set of optimal treatments or policies or other data-dependent index sets of interest. The estimated set is determined by an estimator $\hat p$ of the finite-dimensional parameter $p$ that determines the bounds on $W(d)$ according to Assumption (ref). Let

equation[equation omitted — 238 chars of source]

which are the indices at which the estimated lower and upper bounds are realized. Then, the estimated lower and upper bounds for $W(d)$ are equal to $\tilde\ell_{d,\hat j_L(d)}+\ell_{d,\hat j_L(d)}\hat p$ and $\tilde u_{d,\hat j_U(d)}+u_{d,\hat j_U(d)}\hat p$. We work under the high-level assumption that the following event can be written as a polyhedron in $\hat p$: (i) an option index $d$ is in the set of interest $\widehat{\mathcal{D}}$, (ii) the estimated bounds on $W(d)$ are realized at a given value and (iii) (optionally) an additional random vector is realized at any given value.

assumption\begin{enumerate} • For some fixed and known matrix $A^L(d,j_L^*,\gamma_L^*)$, some fixed and known vector $c^L(d,j_L^*,\gamma_L^*)$ and some finite-valued random vector $\hat\gamma_L(d)$, the event $\{d\in\widehat{\mathcal{D}}$, $\hat j_L(d)=j_L^*$ and $\hat\gamma_L(d)=\gamma_L^*\}$ is equivalent to $\{A^L(d,j_L^*,\gamma_L^*)\hat p\leq c^L(d,j_L^*,\gamma_L^*)\}$, where $j_L^*\in \{1,\ldots,J_L\}$ and $\gamma_L^*$ is in the support of $\hat\gamma_L(d)$. • For some fixed and known matrix $A^U(d,j_U^*,\gamma_U^*)$, some fixed and known vector $c^U(d,j_U^*,\gamma_U^*)$ and some finite-valued random vector $\hat\gamma_U(d)$, the event $\{d\in\widehat{\mathcal{D}}$, $\hat j_U(d)=j_U^*$ and $\hat\gamma_U(d)=\gamma_U^*\}$ is equivalent to $\{A^U(d,j_U^*,\gamma_U^*)\hat p\leq c^U(d,j_U^*,\gamma_U^*)\}$, where $j_U^*\in \{1,\ldots,J_U\}$ and $\gamma_U^*$ is in the support of $\hat\gamma_U(d)$. \end{enumerate}

Depending upon the application, $\hat\gamma_L(d)$ and $\hat\gamma_U(d)$ (and thus $\gamma_L^*$ and $\gamma_U^*$) in this assumption may not be necessary to condition on, in which case they can be vacuously set to constants. Although not immediately obvious, this assumption holds in a variety of settings; see the examples below. In many cases, this assumption can be simplified because, consistent with Assumption 3.1, $d\in\widehat{\mathcal{D}}$ if and only if $A_{\mathcal{D}}\hat{p}\leq c_{\mathcal{D}}$ for some fixed and known matrix $A_{\mathcal{D}}$ and vector $c_{\mathcal{D}}$. For these cases, $\hat\gamma_L(d)$ and $\hat\gamma_U(d)$ are not needed and can be vacuously set to fixed constants and

align[align omitted — 292 chars of source]

and

align[align omitted — 349 chars of source]

A leading example of this special case is \[ \widehat{\mathcal{D}}=\left\{d\in\{d^0,\ldots,d^K\}:\widehat U(d)\geq \max_{d\in\{d^0,\ldots,d^K\}}\widehat L(d)\right\}, \] where $\widehat L(d)\equiv \max_{j\in \{1,\ldots,J_L\}}\{\tilde\ell_{d,j}+\ell_{d,j}\hat p\}$ and $\widehat U(d)\equiv \min_{j\in \{1,\ldots,J_U\}}\{\tilde u_{d,j}+u_{d,j}\hat p\}$, since $d\in\widehat{\mathcal{D}}$ if and only if \[(\ell_{d',j'}-u_{d,j})\hat p \leq \tilde u_{d,j}-\tilde\ell_{d',j'}\] for all $d'\in\{d^0,\ldots,d^K\}$, $j\in \{1,\ldots,J_U\}$ and $j'\in\{1,\ldots,J_L\}$.

We also note that Assumption (ref) is compatible with the absence of data-dependent selection for which the researcher is interested in forming a CI for an identified interval $[L(d^*),U(d^*)]$ chosen by the researcher a priori. In these cases, $\widehat{\mathcal{D}}=\{d^*\}$, $\hat\gamma_M(d^*)$ can be vacuously set to a fixed constant, $A^M(d^*,j,\gamma)=A_M(d^*,j)$ and $c^M(d^*,j,\gamma)=c_M(d^*,j)$ for $M=L,U$. Indeed, we examine an example of this special case when conducting a finite-sample power comparison in Section (ref) below.

In general, less conditioning is more desirable in terms of the lengths of the CIs we propose. Although conditioning on the events $d\in\widehat{\mathcal{D}}$, $\hat j_L(d)=j_L^*$ and $\hat j_U(d)=j_U^*$ is necessary to construct our CIs (see Section (ref) below), the researcher should therefore minimize the number of elements in $\hat \gamma_L(d)$ and $\hat \gamma_U(d)$ subject to satisfying Assumption (ref) when constructing our CIs. In some cases it is necessary to condition on these additional random vectors in order to satisfy Assumption (ref). But in many cases, such as the example given immediately above, additional conditioning random vectors are unnecessary and can be vacuously set to fixed constants.

We impose the following assumption for our unconditional hybrid CIs in order for the object of inferential interest to be well-defined unconditionally.

assumption$\widehat{\mathcal{D}}=\{\hat d\}$ almost surely for a random variable $\hat d$ with support $\{d^0,\ldots,d^K\}$.

In conjunction, Assumptions (ref) and (ref) hold naturally when the object of interest $\hat d$ is selected by uniquely maximizing a linear combination of the estimates of the bounds characterizing the identified intervals and the additional conditioning vectors $\hat\gamma_L(d)$ and $\hat\gamma_U(d)$ are defined appropriately. Leading examples of this form of selection include when $\hat d$ corresponds to the largest estimated lower bound, upper bound or weighted average of lower and upper bounds.

propositionSuppose $\widehat{\mathcal{D}}=\{\hat d\}$, where $\hat d=\operatorname*{argmax\;}_{d\in\{d^0,\ldots,d^K\}}\{w_L\widehat L(d)+w_U\widehat U(d)\}$ is unique almost surely for some fixed known weights $w_L,w_U\geq 0$. Then Assumptions (ref) and (ref) are satisfied for \begin{enumerate} • $\hat\gamma_L(d)$ equal to any fixed constant and $\hat\gamma_U(d)=\hat j_L(d)$ when $w_U=0$, • $\hat\gamma_L(d)=\hat\gamma_U(d)=(\hat j_U(0),\ldots,\hat j_U(T))'$ when $w_L=0$, • $\hat\gamma_L(d)=\hat\gamma_U(d)=(\hat j_L(0),\ldots,\hat j_L(T),\hat j_U(0),\ldots,\hat j_U(T))'$ when $w_L,w_U\neq 0$. \end{enumerate}

Expressions for $A^M(d,j_M^*,\gamma_M^*)$ and $c^M(d,j_M^*,\gamma_M^*)$ for $M=L,U$ in the settings of Proposition (ref) are available for reference in its proof in Appendix (ref). As this proposition makes clear, the additional conditioning vectors needed for Assumption (ref) to hold depend upon the particular form of selection rule used by the researcher. For example, when $\hat d$ is chosen to maximize the estimated lower bound of the identified set $\widehat L( d)$, one must condition not only on the realized value of $\hat j_U(\hat d)$ when forming a probabilistic upper bound for $U(\hat d)$ but also $\hat j_L(\hat d)$. On the other hand, the formation of either a probabilistic lower bound for $L(\hat d)$ or upper bound for $U(\hat d)$ when $\hat d$ is chosen to maximize the estimated upper bound of the identified set $\widehat U( d)$ requires conditioning on the entire vector $(\hat j_U(0),\ldots,\hat j_U(T))'$.

Although intuitively appealing, the treatment choice rules of the form described in Proposition (ref) can be sub-optimal from a statistical decision-theoretic point of view (see, e.g., Man21,Man23 and CMS23). In Section (ref), we show how proper definition of $\hat\gamma_L(d)$ and $\hat\gamma_U(d)$ satisfies Assumptions (ref) and (ref) in the context of the optimal selection rules of CMS23.

We suppose that the sample of data is drawn from some unknown distribution $\mathbb{P}\in\mathcal{P}_n$. As an estimator for $p$, we assume that $\hat p$ is uniformly asymptotically normal under $\mathbb{P}\in\mathcal{P}_n$.

assumptionFor the class of Lipschitz functions that are bounded in absolute value by one and have Lipschitz constant bounded by one, $BL_1$, there exist functions $p(\mathbb{P})$ and $\Sigma(\mathbb{P})$ such that for $\xi_{\mathbb{P}}\sim\mathcal{N}(0,\Sigma(\mathbb{P}))$ with \[\lim_{n\rightarrow\infty}\sup_{\mathbb{P}\in \mathcal{P}_n}\sup_{f\in BL_1}\left|{E}_{\mathbb{P}}\left[f\left(\sqrt{n}(\hat p-p(\mathbb{P}))\right)\right]-{E}_{\mathbb{P}}\left[f\left( \xi_{\mathbb P}\right)\right]\right|=0.\]

The notation of this assumption makes explicit that the parameter $p$ and the asymptotic variance $\Sigma$ depend upon the unknown distribution of the data $\mathbb{P}$. It holds naturally for standard estimators $\hat p$ under random sampling or weak dependence in the presence of bounds on the moments and dependence of the underlying data.

Next, we assume that the asymptotic variance of $\hat p$ can be uniformly consistently estimated by an estimator $\widehat\Sigma$.

assumptionFor all $\varepsilon>0$, the estimator $\widehat\Sigma$ satisfies \[\lim_{n\rightarrow\infty}\sup_{\mathbb{P}\in\mathcal{P}_n} \mathbb{P}\left(\left\|\widehat\Sigma-\Sigma(\mathbb{P})\right\|>\varepsilon\right)=0.\]

This assumption is again naturally satisfied when using a standard sample analog estimator of $\Sigma$ under random sampling or weak dependence in the presence of moment and dependence bounds.

In addition, we restrict the asymptotic variance of $\hat p$ to be positive definite.

assumptionFor some finite $\bar\lambda>0$, $1/\bar \lambda\leq\lambda_{\min}(\Sigma(\mathbb P))\leq\lambda_{\max}(\Sigma(\mathbb P))\leq \bar\lambda$ for all $\mathbb{P}\in \mathcal{P}_n$.

This assumption is naturally satisfied, for example, when $\hat p$ is a standard sample analog estimator of reduced-form probabilities composing $p$ that are non-redundant and bounded away from zero and one.

Examples

In this section, we show that the proposed inference method is applicable to various examples for which parameters are interval-identified. In particular, we show that Assumptions (ref), (ref) and (ref) are satisfied in these examples. See Appendix (ref) for additional examples.

Bounds Derived from Linear Programming

In more complex settings, calculating analytical bounds on $W(d)$ or $W(d)-W(\tilde{d})$ may be cumbersome. This is especially true when the researcher wants to incorporate additional identifying assumptions. In this situation, the computational approach using linear programming can be useful (MST18,HY24).

To incorporate many complicated settings, suppose that $W(d)=A_{d}q$ and $p=Bq$ for some known row vector $A_{d}$ and matrix $B$, an unknown vector $q$ in a simplex $\mathcal{Q}$, and a vector $p$ that is estimable from data. Typically $q$ is a vector of probabilities of a latent variable that governs the DGP; see BP97,BP13, Han24 and HY24. The linearity in this assumption is usually implied by the nature of a particular problem (e.g., discreteness). Then we have

align[align omitted — 150 chars of source]

and ATE bounds for a change from treatment $d$ to treatment $\tilde d$

align[align omitted — 207 chars of source]

Note that $L(\tilde d,{d})\neq L(\tilde d)-U({d})$ in general because the $q$ that solves (ref) for $L(\tilde d)$ and $U({d})$ may be different (and similarly for $U(\tilde d,{d})$). As before, the identified set of optimal treatments here is characterized as $\mathcal{D}^{*}\equiv \{d:L(\tilde{d},d)\le0,\forall\tilde{d}\neq d\}$.

An example of this setting can be found in HY24. Let $(Y,D,Z)$ be a vector of a binary outcome, treatment and instrument and let $p$ be a vector with entries $p(y,d|z)\equiv \mathbb{P}(Y=y,D=d|Z=z)$ across $(y,d,z)\in\{0,1\}^{3}$.\footnote{See HY24 for the use of linear programming with continuous $Y$.} Suppose $W(d)=\mathbb{E}[Y(d)]$ for $d\in\{0,1\}$. Then, we can define the response type $\varepsilon\equiv(Y(1),Y(0),D(1),D(0))$ with a realized value $e\equiv(y(1),y(0),d(1),d(0))$, where $Y(d)$ denotes the potential outcome under treatment $d$ and $D(z)$ denotes the potential treatment under instrument value $z$. Let $q(e)\equiv \mathbb{P}(\varepsilon=e)$ be the latent distribution. Then

align*[align* omitted — 74 chars of source]

where $q$ is the vector of $q(e)$'s and $A_{d}$ is an appropriate selector (a row vector).

Assume that $(Y(d),D(z))$ is independent of $Z$ for $d,z\in\{0,1\}$. The data distribution $p$ is related to the latent distribution by

align*[align* omitted — 109 chars of source]

where the first equality follows by the independence assumption, $q$ is a vector of $q(e)$'s and $B_{d,z}$ is an appropriate selector (a row vector). Now define

align*[align* omitted — 209 chars of source]

so that all of the constraints relating the data distribution to the latent distribution can be expressed as $Bq=p$.

To verify Assumption (ref), it is helpful to invoke strong duality for the primal problems (ref) (under regularity conditions) and write the following dual problems:

align*[align* omitted — 182 chars of source]

where $\tilde{B}\equiv\left[

array[array omitted — 35 chars of source]

\right]$ is a $(d_{p}+1)\times d_{q}$ matrix with $\boldsymbol{1}$ being a $d_{q}\times1$ vector of ones, and $\tilde{p}\equiv\left[

array[array omitted — 21 chars of source]

\right]$ is a $(d_{p}+1)\times1$ vector. By using a vertex enumeration algorithm (e.g., \citet{AF91}), one can find all (or a relevant subset) of vertices of the polyhedra $\{\lambda:\tilde{B}'\lambda\ge-A_{d}'\}$ and $\{\lambda:\tilde{B}'\lambda\ge A_{d}'\}$. Let $\Lambda_{L,d}\equiv\{\lambda_{1},...,\lambda_{J_{L,d}}\}$ and $\Lambda_{U,d}\equiv\{\lambda_{1},...,\lambda_{J_{U,d}}\}$ be the sets that collect such vertices, respectively. Then, it is easy to see that $L(d)=\max_{\lambda\in\Lambda_{L,d}}-\tilde{p}'\lambda$ and $U(d)=\min_{\lambda\in\Lambda_{U,d}}\tilde{p}'\lambda$, and thus Assumption (ref) holds.

To verify Assumption (ref), we use the dual problems to (ref):

align*[align* omitted — 237 chars of source]

where $\Delta_{\tilde d,{d}}\equiv A_{\tilde d}-A_{{d}}$. Analogous to the vertex enumeration argument above, let $\Lambda_{L,\tilde d,{d}}\equiv\{\lambda_{1},...,\lambda_{J_{L,\tilde d,{d}}}\}$ and $\Lambda_{U,\tilde d,{d}}\equiv\{\lambda_{1},...,\lambda_{J_{U,\tilde d,{d}}}\}$ be the sets that collect all (or a relevant subset) of vertices of the polyhedra $\{\lambda:\tilde{B}'\lambda\ge-\Delta_{\tilde d,{d}}'\}$ and $\{\lambda:\tilde{B}'\lambda\ge\Delta_{\tilde d,{d}}'\}$, respectively. Then, $L(\tilde d,{d})=\max_{\lambda\in\Lambda_{L,\tilde d,{d}}}-\tilde{p}'\lambda$ and $U(\tilde d,{d})=\min_{\lambda\in\Lambda_{U,\tilde d,{d}}}\tilde{p}'\lambda$. Let $\widehat{\mathcal{D}}=\{d:\widehat{L}(\tilde d,{d})\le0,\forall\tilde{d}\neq d\}$, where $\widehat{L}(\tilde d,{d})$ is the sample counterpart of $L(\tilde d,{d})$ with $\widehat{\tilde{p}}\equiv\left[

array[array omitted — 27 chars of source]

\right]$ replacing $\tilde{p}\equiv\left[

array[array omitted — 21 chars of source]

\right]$. Partition $\lambda$ as $\lambda=(\lambda^{1\prime},\lambda^{0})'$ where $\lambda^{0}$ is the last element of $\lambda$. Note that $d\in\widehat{\mathcal{D}}$ if and only if

align*[align* omitted — 94 chars of source]

where $\tilde{\Lambda}_{L,d}=\bigcup_{\tilde{d}\neq d}\Lambda_{L,\tilde d,{d}}$. Also let $\hat{\lambda}$ be such that $-\widehat{\tilde{p}}'\hat{\lambda}=\max_{\lambda\in\tilde{\Lambda}_{L,d}}-\widehat{\tilde{p}}'\lambda$. Then, $\hat{\lambda}=\lambda_{j_{L}^{*}}$ if and only if

align*[align* omitted — 186 chars of source]

so that Assumption (ref) holds.

Finally, $\hat p$ is again equal to a vector of sample means so that Assumption (ref) is satisfied if $\hat{p}$ is calculated using the random sample $\{Y_{i},D_{i},Z_{i}\}$$_{i=1}^{n}$.

Empirical Welfare Maximization with Observational Data

Consider allocating a binary treatment based on observed covariates $X\in\mathcal{X}$. A treatment allocation rule can be defined as a function $\delta:\mathcal{X}\rightarrow\{0,1\}$ in a class of rules $\mathcal{D}$. Consider the utilitarian welfare of deploying $\delta$ relative to treating no one. The optimal allocation $\delta^{*}$ satisfies

align*[align* omitted — 71 chars of source]

Note that $\mathbb{E}[Y(\delta(X))-Y(0)] =E\left[\delta(X)\Delta(X)\right]$, where $\Delta(X)\equiv \mathbb{E}[Y(1)-Y(0)|X]$. This problem is considered in KT18 and AW21, among others. When only observational data for $(Y,D,X)$ are available with $D$ being endogenous, $W(\delta)$ is only partially identified unless strong treatment effect homogeneity is assumed. This problem has been studied in KZ21, PZ21, DAd21, Bya22, among others. Using instrumental variables, one can consider bounds on the conditional ATE based on conditional versions of the bounds considered in Sections (ref) and (ref) (i.e., Manski's bounds and bounds produced by linear programming).

In particular, assume that $(Y(d),D(z))$ is independent of $Z$ given $X$. Let $L(X)$ and $U(X)$ be conditional Manski bounds on $\Delta(X)$. Then, bounds on $W(\delta)$ can be characterized as

align*[align* omitted — 109 chars of source]

Similarly, bounds on $W(\tilde\delta)-W({\delta})=\mathbb{E}[(\tilde\delta(X)-{\delta}(X))\Delta(X)]$ can be characterized as

align[align omitted — 192 chars of source]

Note that $L(\tilde\delta,{\delta})\neq L(\tilde\delta)-U({\delta})$ in general (and similarly for $U(\tilde\delta,{\delta})$).

Suppose $\mathcal{X}$ is finite and $\mathcal{X}=\{x_{1},...,x_{K}\}$ where $x_{k}$ can be a vector and $K$ can potentially be large. For simplicity of exposition, suppose $\mathcal{X}=\{0,1,2\}$. Then $\mathcal{D}=\{\delta_{1},...,\delta_{8}\}$ where each $\delta_{j}$ corresponds to a mapping type from $\{0,1,2\}$ to $\{0,1\}$. To verify Assumptions (ref) and (ref), we proceed as follows. For given $x\in\mathcal{X}$, by arguments analogous to those in Section (ref) (and Section (ref)), bounds $L_x$ and $U_x$ on $\Delta(x)$ satisfy, for some scalars $\tilde{\ell}_{j}$ and $\tilde{u}_{j}$ and row vectors $\ell_{j}$ and $u_{j}$,

align*[align* omitted — 145 chars of source]

where $p(x)$ is the vector of $p(y,d|z,x)$'s across $(y,d,z)$ fixing $x$. Then, by Jensen's inequality, for each $\delta\in\mathcal{D}$,

align*[align* omitted — 291 chars of source]

Note that $\tilde{L}(\delta)$ and $\tilde{U}(\delta)$ are non-sharp bounds; for calculation of sharp bounds, see Section (ref). We can verify Assumption (ref) with $\tilde{L}(\delta)$ and $\tilde{U}(\delta)$ by defining

align*[align* omitted — 187 chars of source]

and, for $\delta=\delta_{1}$ as an example, by using $\ell_{\delta_{1},j} =(

array[array omitted — 63 chars of source]

)$. Similarly, we can verify Assumptions \ref{ass: estimated ID'd set} and \ref{ass: joint normality} by estimating $p(X)$ and $\mathbb{E}[\delta(X)]$ with sample means and ${E}[\delta(X){p}(X)]$ with $\frac{1}{n}\sum_{i}^{n}\delta(X_{i})\hat{p}(X_{i}).$ If the data $\{Y_{i},D_{i},Z_{i},X_{i}\}$$_{i=1}^{n}$ form a random sample, and $\widehat{\mathcal{D}}=\{\delta\in\mathcal{D}:\widehat{L}(\tilde{\delta},\delta)\le0\,\forall\tilde{\delta}\neq \delta\}$ for $\widehat{L}(\tilde{\delta},\delta)$ defined the same as ${L}(\tilde{\delta},\delta)$ in (ref) after substituting $\hat p$ for $p$, Assumptions (ref) and (ref) hold.

This framework can be generalized to settings where $W(\delta)$ is partially identified, not necessarily due to treatment endogeneity but because $W(\delta)$ is a non-utilitarian welfare defined as a functional of the joint distribution of potential outcomes (e.g., CH24): $W(\delta)=f(F_{Y(1),Y(0)|X})$ where $f$ is some functional and $F_{Y(1),Y(0)|X}$ is the joint distribution of $(Y(1),Y(0))$ conditional on $X$.

Bounds for Dynamic Treatment Effects

Consider binary $Y_{t}$ and $D_{t}$ for $t=1,...,T$. Let $Y\equiv(Y_{1},...,Y_{T})$ and $D\equiv(D_{1},...,D_{T})$. Suppose that we are equipped with a sequence of binary instruments $Z\equiv(Z_{t_{1}},...,Z_{t_{K}})$, which is a subvector of $(Z_{1},...,Z_{T})$. For $t=1,...,T$, let $Y_{t}(d_{1},...,d_{t})$ be the potential outcome at $t$ and $Y(d)\equiv(Y_{1}(d_{1}),...,Y_{T}(d_{1},...,d_{T}))$. We assume that the instruments $Z$ are independent of the potential outcomes $Y(d)$.

Let $T=2$. Then $Y\equiv(Y_{1},Y_{2})$ and $D\equiv(D_{1},D_{2})$. For given welfare $W(d)$ with $d\equiv (d_1,d_2)$, we are interested in the optimal policy $d^{*}$ that satisfies

align*[align* omitted — 56 chars of source]

where $\mathcal{D}\equiv\{(1,1),(1,0),(0,1),(0,0)\}$. The sign of the welfare difference, $W(d)-W(\tilde{d})$ for $d,\tilde{d}\in\mathcal{D}$, is useful for establishing the ordering of $W(d)$ with respect to $d$ and thus to identifying $d^{*}$. However, without additional identifying assumptions, we can only establish a partial ordering of $W(d)$ based on the bounds on the welfare difference Han24. This will produce the identified set $\mathcal{D}^{*}$ for $d^{*}$.

An example of the welfare is $W(d)\equiv \mathbb{E}[Y_{2}(d)]$, namely, the average potential terminal outcome. The bounds on welfare $W(d)$ are

align[align omitted — 92 chars of source]

where

align*[align* omitted — 141 chars of source]

which have forms analogous to those in the static case. Define the dynamic ATE in the terminal period for a change in treatment from $d$ to $\tilde d$ as

align*[align* omitted — 130 chars of source]

Then the bounds on the dynamic ATE are as follows:

align[align omitted — 124 chars of source]

Another example of welfare is the joint distribution $W(d)\equiv \mathbb{P}(Y(d)=(1,1))$ where $Y(d)\equiv(Y_{1}(d_{1}),Y_{2}(d))$. The bounds on $W(d)$ in this case are $L(d)\equiv\max_{z}L(d;z)$ and $U(d)\equiv\min_{z}U(d;z)$ where

align*[align* omitted — 248 chars of source]

with $d'_1\neq d_1$ and $d'_2\neq d_2$. Consider the effect of treatment on the joint distribution, for example,

align*[align* omitted — 81 chars of source]

Then, with $\tilde d=(1,1)$ and ${d}=(1,0)$, the bounds on this parameter are

align*[align* omitted — 485 chars of source]

In these examples, the identified set $\mathcal{D}^{*}$ can be characterized as a set of maximal elements:

align*[align* omitted — 146 chars of source]

These examples are special cases of the model in Han24.\footnote{See also HL24 for other examples of dynamic causal parameters that can be used to define optimal treatments.}

In both cases, it is easy to see that $L(d)$ and $U(d)$ satisfy Assumption (ref) with $p$ being the vector of probabilities $p(y,d|z)\equiv \mathbb{P}(Y=y,D=d|Z=z)$. To verify Assumption (ref), let $\widehat{L}(\tilde{d},d)$ be the estimator of $L(\tilde{d},d)$ where the sample frequency replaces the population probability. Then $\widehat{\mathcal{D}}=\{d:\widehat{L}(\tilde{d},d)\le0,\forall\tilde{d}\neq d\}$.

Continuing with the first example, let $Z=Z_{1}\in\{0,1\}$, that is, the researcher is only equipped with a binary instrument in the first period and no instrument in the second period. We focus on inference for $L(d)$ for $d\in\widehat{\mathcal{D}}$. Let $\widehat{z}(d)\in\arg\max_{z\in\{0,1\}}\widehat{L}(d;z)$ so that $\widehat{L}(d)=\widehat{L}(d;\widehat{z}(d))$. Then we can write the data-dependent event that $d$ is an element of $\widehat{\mathcal{D}}$, $\widehat{L}(d)=\widehat{L}(d;z^{*})$ as a polyhedron

align*[align* omitted — 165 chars of source]

for some matrix $A^L$, where $\hat{p}$ is the vector of probabilities $\widehat{p}(y,d|z)$, so that Assumption (ref) holds. This is due to the forms of $L(d;z)$ and $U(d;z)$ above and $\widehat{L}(\tilde{d},d)\le0$ $\forall\tilde{d}\neq d$ if and only if

align*[align* omitted — 101 chars of source]

and, for example, $\widehat{z}(d)=1$ if and only if $\widehat{L}(d;0)-\widehat{L}(d;1) \le0$. A similar formulation follows for the second example. In fact, this approach applies to a general parameter $W(d)$ with bounds that are minimum and maximums of linear combinations of $p(y,d|z)$'s, such as parameters that have the following form:

align*[align* omitted — 36 chars of source]

for some linear functional $f$, where $q_{d}(y)\equiv \mathbb{P}(Y(d)=y)$.

Finally, Assumption (ref) is satisfied in both examples when the data $\{Y_i,D_i,Z_i\}_{i=1}^n$ form a random sample since the entries of $\hat p$ are sample means.

This framework can be further generalized to incorporate treatment choices adaptive to covariates or past outcomes as in Han24, analogous to Section (ref). This generalization is considered in our empirical application in Section (ref). Sometimes this generalization prevents the researcher from deriving analytical bounds, in which case the linear programming approach can be used.

Optimal Treatment Assignment with Interval-Identified ATE

In recent work, CMS23 note that “plug-in” rules for determining treatment choice can be sub-optimal when ATEs are not point-identified since the bounds on the ATE are not smooth functions of the reduced-form parameter $p$. Using an optimality criterion that minimizes maximum regret over the identified set for the ATE, conditional on $p$, they advocate bootstrap and quasi-Bayesian methods for optimal treatment choice. More specifically, they consider settings for which the ATE of a treatment is identified via intersection bounds: \[b_L(p)\equiv \max_{j\in\{1,\ldots,J_L\}}\{\tilde{\ell}_j+\ell_jp\}\leq ATE\leq \min_{j\in\{1,\ldots,J_U\}}\{\tilde{u}_k+u_kp\}\equiv b_U(p)\] for some fixed and known $J_L$, $J_U$, $\tilde{\ell}_1,\ldots,\tilde{\ell}_{J_L}$, $\tilde{u}_1,\ldots,\tilde{u}_{J_U}$, $\ell_1,\ldots,\ell_{J_L}$ and $u_1,\ldots,u_{J_U}$.\footnote{Although CMS23 do not write the form of their bounds as they are written here, the representation here is equivalent to the one in that paper upon proper definition of $p$ since the elements in the intersection bounds are smooth functions of a reduced-form parameter.} Therefore, Assumption (ref) trivially holds for the ATE.

CMS23 advocate a quasi-Bayesian implementation of their optimal treatment choice rule taking the form

equation[equation omitted — 191 chars of source]

for some large $m$, where $\varepsilon_1,\ldots,\varepsilon_m\overset{i.i.d.}\sim\mathcal{N}(0,\widehat \Sigma)$ are independent of $\hat p$. As the following proposition shows, this form of $\hat d$ satisfies Assumption (ref) when $\hat\gamma_L(d)=\hat\gamma_U(d)$ are specified properly.

propositionSuppose $\widehat{\mathcal{D}}=\{\hat d\}$, where $\hat d$ is defined by (ref). Then Assumptions (ref) and (ref) are satisfied for \[\hat\gamma_L(d)=\hat\gamma_U(d)=(\varepsilon_1^\prime,\ldots,\varepsilon_m^\prime,\underbar k_1,\ldots,\underbar k_m,\bar k_1,\ldots,\bar k_m,s_1^{\ell},\ldots,s_m^{\ell},s_1^{u},\ldots,s_m^{u})',\] where $\underbar k_i\equiv\operatorname{argmin}_{k\in\{1,\ldots,J_U\}}\{\tilde{u}_k+u_k(\hat p+\varepsilon_i)\}$, $\bar k_i\equiv\operatorname*{argmax\;}_{k\in\{1,\ldots,J_L\}}\{\tilde{\ell}_k+\ell_k(\hat p+\varepsilon_i)\}$, $s_i^{\ell}\equiv\operatorname*{sign}(\tilde{\ell}_{\bar k_i}+\ell_{\bar k_i}(\hat p+\varepsilon_i))$ and $s_i^{u}\equiv\operatorname*{sign}(\tilde{u}_{\underbar k_i}+u_{\underbar k_i}(\hat p+\varepsilon_i))$ for $i=1,\ldots,m$.

Confidence Interval Construction

We now generalize the CI construction described in Section (ref) to apply in the general framework of Section (ref), covering all example applications discussed above and in Appendix (ref). We start with conditional CIs and then move to unconditional CIs.

Conditional Confidence Intervals

We first generalize the conditional CI construction described in Section (ref). As in Section (ref), we are interested in forming probabilistic lower and upper bounds $\widehat L(\hat d)_{\alpha}^C$ and $\widehat U(\hat d)_{1-\alpha}^C$ that satisfy (ref) and (ref) for all $d\in\{d^0,\ldots,d^K\}$ as endpoints in the formation of a conditionally valid CI. This is because the researcher's interest in inference on option $d$ only arises when it is a member of the estimated set $\widehat{\mathcal{D}}$.

To begin, we characterize the conditional distributions of $\widehat{L}(d)$ given the event $\{d\in\widehat{\mathcal{D}}$, $\hat j_L(d)=j_L^*$ and $\hat\gamma_L(d)=\gamma_L^*\}$ characterized by Assumption (ref). These conditional distributions depend upon the nuisance parameter $p$. As a first step, we form sufficient statistics for $p$ that are asymptotically independent of $\widehat{L}(d)$ given $\hat j_L(d)=j_L^*$ and $\widehat{U}(d)$ given $\hat j_U(d)=j_U^*$. Since $\widehat{L}(d)=\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p$ given $\hat j_L(d)=j_L^*$ by Assumption (ref) and (ref) (and similarly for $\widehat{U}(d)$), such sufficient statistics can be constructed as \[\widehat{\mathcal{Z}}_{L}(d,j_L^*)\equiv \sqrt{n}\hat p-\hat b_{L}(d,j_L^*)\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p)\] and \[\widehat{\mathcal{Z}}_{U}(d,j_U^*)\equiv \sqrt{n}\hat p-\hat b_{U}(d,j_U^*)\sqrt{n}(\tilde u_{d,j_U^*}+u_{d,j_U^*}\hat p)\] with \[\hat b_{L}(d,j_L^*)\equiv \widehat\Sigma\ell_{d,j_L^*}^{\prime}\left(\ell_{d,j_L^*}\widehat\Sigma\ell_{d,j_L^*}^{\prime}\right)^{-1} \quad \text{and} \quad \hat b_{U}(d,j_U^*)\equiv \widehat\Sigma u_{d,j_U^*}^{\prime}\left(u_{d,j_U^*}\widehat\Sigma u_{d,j_U^*}^{\prime}\right)^{-1},\] by the asymptotic normality of $\hat p$ and asymptotic absence of correlation between $\sqrt{n}\hat p$ and $\widehat{\mathcal{Z}}_{L}(d,j_L^*)$ and $\widehat{\mathcal{Z}}_{U}(d,j_U^*)$. Next, Assumption (ref) characterizes the conditioning events $\{d\in\widehat{\mathcal{D}},\hat j_L(d)=j_L^*,\hat \gamma_L(d)=\gamma_L^*\}$ and $\{d\in\widehat{\mathcal{D}},\hat j_U(d)=j_U^*,\hat \gamma_U(d)=\gamma_U^*\}$ as polyhedra in $\hat p$, which can in turn be expressed as intervals for $\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p$ and $\tilde u_{d,j_U^*}+u_{d,j_U^*}\hat p$; see, e.g., (ref) and (ref). Then, again since $\widehat{L}(d)=\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p$ given $\hat j_L(d)=j_L^*$ by Assumption (ref), Lemma 1 of McC24 implies

gather[gather omitted — 628 chars of source]

and

gather*[gather* omitted — 578 chars of source]

with

gather*[gather* omitted — 531 chars of source]

for $M=L,U$.

Now, under Assumptions (ref) and (ref), the distribution of $\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*}\hat p)$ can be approximated by a $\mathcal{N}(\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*} p),\ell_{d,j_L^*}\widehat\Sigma \ell_{d,j_L^*}')$-distributed random variable that is asymptotically independent of $\widehat{\mathcal{Z}}_{L}(d,j_L^*)$. Using the distributional characterization in (ref), we can therefore use the corresponding truncated normal cumulative distribution function to produce quantile-unbiased estimators of the underlying mean $\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*} p)$. Let $F_{TN}(\cdot;\mu,\sigma^2|\mathcal V^-,\mathcal V^+)$ denote the truncated normal cumulative distribution function for an underlying normally-distributed random variable with mean $\mu$ and variance $\sigma^2$ that is truncated to lie between $\mathcal V^-$ and $\mathcal V^+$. For $\alpha\in(0,1)$, define $\widehat L(d)_{\alpha}^C$ to solve

gather*[gather* omitted — 415 chars of source]

in $\mu$. Similarly, define $\widehat U(d)_{\alpha}^C$ to solve

gather*[gather* omitted — 404 chars of source]

in $\mu$. Then, results in Pfa94 imply that $\widehat L(d)_{\alpha}^C$ and $\widehat U(d)_{\alpha}^C$ are optimal $\alpha$ quantile-unbiased estimators of $\sqrt{n}(\tilde \ell_{d,j_L^*}+\ell_{d,j_L^*} p)$ and $\sqrt{n}(\tilde u_{d,j_L^*}+u_{d,j_L^*} p)$ asymptotically.

Finally, combine these quantile-unbiased estimators to form a conditional CI for the identified interval $[L(d),U(d)]$:

equation[equation omitted — 110 chars of source]

We establish the conditional uniform asymptotic validity of this CI.

theoremSuppose Assumptions (ref), (ref) and (ref)--(ref) hold. Then, for any $d\in\{d^0,\ldots,d^K\}$ and $0<\alpha_1,\alpha_2<1/2$, \begin{gather*} \liminf_{n\rightarrow\infty}\inf_{\mathbb{P}\in\mathcal{P}_n}\left\{\left[\mathbb{P}\left(\left.[L(d),U(d)]\subseteq \left(n^{-1/2}\widehat L(d)_{\alpha_1}^C,n^{-1/2}\widehat U(d)_{1-\alpha_2}^C\right)\right|d\in\widehat{\mathcal{D}}\right)-(1-\alpha_1-\alpha_2)\right]\cdot\mathbb{P}(d\in\widehat{\mathcal{D}})\right\}\geq 0 \end{gather*} for all $d\in\{d^0,\ldots,d^K\}$.

Unconditional Confidence Intervals

In parallel with the previous subsection, we now generalize the unconditional CI constructions described in Section (ref) to the general framework of Section (ref). Note that conditional inference on $W(d)$ is well-defined for any given $d\in\widehat{\mathcal{D}}$, when conditioning on $d\in\widehat{\mathcal{D}}$. In contrast, unconditional inference on a data-dependent $W(d)$ requires it to be uniquely defined, as $W(\hat d)$ in our notation. This is implied by Assumption (ref). Here, we would like to construct CIs that unconditionally cover the identified interval corresponding to a unique data-dependent object of inferential interest. As mentioned in Section (ref), if only unconditional coverage of $[L(\hat d),U(\hat d)]$ is desired the conditional CI (ref) with $d=\hat d$ can be unnecessarily wide. We describe two different methods---projection and hybrid methods---to form the unconditional probabilistic bounds that constitute the endpoints of these unconditional CIs in this general framework.

The general formation of the probabilistic bounds based upon projecting simultaneous confidence bounds for all possible values of $\sqrt{n}L(\hat d)$ and $\sqrt{n}U(\hat d)$ proceeds by computing $\hat{c}_{1-\alpha,M}$, the $1-\alpha$ quantile of $\max_{i\in \{1,\ldots,(T+1)J_M\}}\hat \zeta_{M,i}/\sqrt{\widehat{\Sigma}_{M,i}}$, where $\hat \zeta_M\sim\mathcal{N}(0,\widehat{\Sigma}_M)$ for $M=L,U$ with $\widehat{\Sigma}_L=\ell^{mat}\widehat{\Sigma}\ell^{mat\prime}$, $\widehat{\Sigma}_U=u^{mat}\widehat{\Sigma}u^{mat\prime}$, $\ell^{mat}=(\ell_{0,1}',\hdots,\ell_{0,J_L}',\hdots,\ell_{T,1}',\hdots,\ell_{T,J_L}')'$ and $u^{mat}=(u_{0,1}',\hdots,u_{0,J_U}',\hdots,u_{T,1}',\hdots,u_{T,J_U}')'$, recalling that $\Upsilon_i$ denotes the $i^{th}$ element of the main diagonal of any square matrix $\Upsilon$. Here, the maximum is taken to guarantee simultaneous coverage. The lower level $1-\alpha$ projection confidence bound for $\sqrt{n}L(\hat d)$ is $\widehat L(\hat d)_\alpha^P=\sqrt{n}\widehat L(\hat d)-\hat{c}_{1-\alpha,L}\sqrt{\widehat{\Sigma}_{L,\hat dJ_L+\hat{j}_L(\hat d)}}$ and the upper level $1-\alpha$ projection confidence bound for $\sqrt{n}U(\hat d)$ is $\widehat{U}_{1-\alpha}^P(\hat d)=\sqrt{n}\widehat U(\hat d)+\hat{c}_{1-\alpha,U}\sqrt{\widehat{\Sigma}_{U,\hat dJ_U+\hat{j}_U(\hat d)}}$, because e.g., $\sqrt{n}L(\hat d)$ can take value equal to any entry of the vector $\sqrt{n}(\tilde\ell_{0,1},\hdots,\tilde\ell_{0,J_L},\hdots,\tilde\ell_{T,1},\hdots,\tilde\ell_{T,J_L})'+\sqrt{n}\ell^{mat} p$, $\sqrt{n}(\ell^{mat}\hat p-\ell^{mat}p)$ is asymptotically distributed $\mathcal{N}(0,\Sigma_L)$ for $\Sigma_L=\ell^{mat}\Sigma\ell^{mat\prime}$ by Assumption (ref) and $\widehat\Sigma_L$ is consistent for $\Sigma_L$ by Assumption (ref).

Combining these two confidence bounds at appropriate levels yields an unconditional CI for the identified interval $[L(\hat d),U(\hat d)]$ of the selected $\hat d$,

equation[equation omitted — 120 chars of source]

with uniformly correct asymptotic coverage, regardless of how $\hat d$ is selected from the data.

theoremSuppose Assumptions (ref) and (ref)--(ref) hold. Then, for any (random) $\hat{d}\in\{d^0,\ldots,d^K\}$ and $0<\alpha_1,\alpha_2<1/2$, \begin{equation*} \liminf_{n\rightarrow\infty}\inf_{\mathbb{P}\in\mathcal{P}_n}\mathbb{P}\left([L(\hat d),U(\hat d)]\subseteq \left(n^{-1/2}\widehat L(\hat d)_{\alpha_1}^P,n^{-1/2}\widehat U(\hat d)_{1-\alpha_2}^P\right)\right)\geq 1-\alpha_1-\alpha_2. \end{equation*}

Note that the projection CI (ref) has the benefit of correct coverage regardless of how $\hat d$ is chosen from the data. In this sense, it is more robust than the other CIs we propose in this paper. On the other hand, by using the common selection structure of Assumption (ref), we are able to produce a hybrid CI that combines the strengths of the conditional CI (ref) and the projection CI (ref) which, as described in Section (ref), are shorter under complementary scenarios.

In analogy with the construction of the conditional CIs, to construct the hybrid CIs we begin by characterizing the conditional distributions of $\widehat L(\hat d)$ and $\widehat U(\hat d)$ but now adding an additional component to the conditioning events. More specifically, under Assumptions (ref) and (ref), by intersecting the events

gather*[gather* omitted — 590 chars of source]

and

gather*[gather* omitted — 407 chars of source]

for some $0<\beta<\alpha<1$, we have

gather*[gather* omitted — 781 chars of source]

where \[\widehat{\mathcal{V}}_{L}^{+,H}(z,d,j,\gamma,\mu)\equiv\min\left\{\widehat{\mathcal{V}}_{L}^{+}(z,d,j,\gamma),\mu+\hat{c}_{1-\beta,L}\sqrt{\widehat{\Sigma}_{L,dJ_L+j}}\right\}.\] Similarly,

gather*[gather* omitted — 762 chars of source]

where \[\widehat{\mathcal{V}}_{U}^{-,H}(z,d,j,\gamma,\mu)\equiv\max\left\{\widehat{\mathcal{V}}_{U}^{-}(z,d,j,\gamma),\mu-\hat{c}_{1-\beta,U}\sqrt{\widehat{\Sigma}_{U,dJ_U+j}}\right\}.\] Since the distribution of $\sqrt{n}(\tilde \ell_{d^*,j_L^*}+\ell_{d^*,j_L^*}\hat p)$ can be approximated by $\mathcal{N}(\sqrt{n}(\tilde \ell_{d^*,j_L^*}+\ell_{d^*,j_L^*} p),\ell_{d^*,j_L^*}\Sigma \ell_{d^*,j_L^*}')$ asymptotically and $\widehat{\mathcal{Z}}_{L}(d^*,j_L^*)$ is asymptotically independent, we again work with the truncated normal distribution to compute a hybrid probabilistic lower bound for $\sqrt{n}L(\hat d)$: for $0<\beta<\alpha<1$, define $\widehat L(d)_{\alpha}^H$ to solve

gather*[gather* omitted — 439 chars of source]

in $\mu$. Similarly, define $\widehat U(d)_{\alpha}^H$ to solve

gather*[gather* omitted — 428 chars of source]

in $\mu$.

Here, $\widehat L(\hat d)_{\alpha}^H$ is an unconditionally valid probabilistic lower bound for $L(\hat d)$ and $\widehat U(\hat d)_{1-\alpha}^H$is an unconditionally valid probabilistic upper bound for $U(\hat d)$. Combining these two confidence bounds at appropriate levels yields an unconditional CI for the identified interval $[L(\hat d),U(\hat d)]$ of the selected $\hat d$,

equation[equation omitted — 119 chars of source]

with uniformly correct asymptotic coverage when $\hat d$ is selected from the data according to Assumptions (ref) and (ref).

theoremSuppose Assumptions (ref)--(ref) hold. Then, for any $0<\alpha_1,\alpha_2<1/2$, \begin{equation*} \liminf_{n\rightarrow\infty}\inf_{\mathbb{P}\in\mathcal{P}_n}\mathbb{P}\left([L(\hat d),U(\hat d)]\subseteq \left(n^{-1/2}\widehat L(\hat d)_{\alpha_1}^H,n^{-1/2}\widehat U(\hat d)_{1-\alpha_2}^H\right)\right)\geq 1-\alpha_1-\alpha_2. \end{equation*}

Reality Check Power Comparison

To our knowledge, the CIs proposed in this paper are the first with proven asymptotic validity for interval-identified parameters selected from the data. Therefore, we have no existing inference method to compare the performance of our CIs to when the interval-identified parameter is data-dependent. However, as discussed in Section (ref) above, our inference framework covers cases for which the interval-identified parameter is chosen a priori. For these special cases, there is a large literature on inference on partially-identified parameters or their identified sets that can be applied to form CIs. Although these special cases are not of primary interest for this paper, in this section we compare the performance of our proposed inference methods to one of the leading inference methods in the partial identification literature as a “reality check” on whether our proposed methods are reasonably informative.

In particular, we compare the power of the test implied by our hybrid CI (i.e., a test that rejects when the value of the parameter under the null hypothesis lies outside of the hybrid CI) to the power of the hybrid test of ARP23, which applies to a general class of moment-inequality models. When $d$ is chosen a priori and the parameter $p$ is equal to a vector of moments of underlying data, Assumption (ref) implies that $L(d)\leq W(d)\leq U(d)$ can be written as a set of (unconditional) moment inequalities, a special case of the general framework of that paper. Of the many papers on inference for moment inequalities, we choose the test of ARP23 for comparison for two reasons: (i) it has been shown to be quite competitive in terms of power and (ii) it is also based upon a (different) inference method that is a hybrid between conditional and projection-based inference.

We compare the power of tests on the ATE in the same setting of the Manski bounds example of Section (ref), strengthening the mean independence assumption $\mathbb{E}[Y(d)|Z]=\mathbb{E}[Y(d)]$ to full statistical independence $(Y(1),Y(0))\perp Z$ and using the sharp bounds on the ATE $W(1)-W(0)=\mathbb{E}[Y(1)]-\mathbb{E}[Y(0)]$ derived by BP97,BP13:

equation*[equation* omitted — 317 chars of source]

and

equation*[equation* omitted — 319 chars of source]

For a sample size of $n=100$, we generated $\hat p$ from a $\mathcal{N}(p,\Sigma)$ distribution.\footnote{Note that in this problem, the value of $\Sigma$ is implied by the value of $p$.} Figure (ref) plots the power curves of the hybrid ARP23 test and the test implied by our hybrid CI for three different DGPs within the framework of this example, as well as the true identified interval for the ATE. The DGP corresponding to $p=(.08,.001,.001,.073,.139,.473)'$ is calibrated to the probabilities estimated by BP13 in the context of a treatment for high-cholesterol (specifically, by the drug cholestyramine).\footnote{BP13 estimate $p$ to equal $(.081,0,0,.073,.139,.473)'$. If the true DGP is set exactly equal to this, Assumption (ref) would be violated. We therefore slightly alter the calibrated probabilities.} The DGP corresponding to $p=(.25,.25,.25,.25,.25,.25)'$ generates completely uninformative bounds for the ATE. And the DGP corresponding to $p=(.01,.44,.01,.01,.01,.54)'$ generates quite informative bounds.

figure[figure omitted — 645 chars of source]

We can see that overall, the power of the test implied by our hybrid CI is quite competitive with that of ARP23. Interestingly, it appears that our test tends to be more powerful than that of ARP23 when the true ATE is larger than the hypothesized one, which can be seen from the hypothesized ATE lying to the left of the lower bound of the identified interval, and vice versa. Although the main innovation of our inference procedures is really their validity in the presence of data-dependent selection, the exercise of this section is reassuring for the informativeness of the procedures we propose as they are quite competitive in the absence of selection.

Finite Sample Performance of Confidence Intervals

Moving now to a context in which the object of interest is selected from the data, we compare the finite sample performance of our conditional, projection and hybrid CIs again in the setting of the Manski bounds example of Section (ref). In this case, we are interested in inference on the average potential outcome $W(\hat d)$ for $W(d)=\mathbb{E}[Y(d)]$, where interest arises either in the average potential outcome for treatment ($\hat d=1$) or control ($\hat d=0$) depending upon which has the largest estimated lower bound: $\hat d=\operatorname*{argmax\;}_{d\in\{0,1\}}\widehat L(d)$. This form of $\hat d$ corresponds to case 1. of Proposition (ref) and we use the corresponding result of the proposition to specify $\hat\gamma_L(\hat d)$ and $\hat\gamma_U(\hat d)$ in the construction of the conditional and hybrid CIs. We report analogous simulation results for the dynamic treatment regime example in Appendix (ref) with $\hat p$ generated from a multinomial, rather than normal, distribution.

For the same DGPs as in Section (ref), we compute the unconditional coverage frequencies of the conditional, projection and hybrid CIs as well as that of the conventional CI based upon the normal distribution. These coverage frequencies are reported in Table (ref). Consistent with the asymptotic results of Theorems (ref), (ref) and (ref), the conditional, projection and hybrid CIs all have correct coverage for all DGPs and the modest sample size of $n=100$. Also consistent with Theorem (ref), we note that the projection CI tends to be conservative with true coverage above the nominal level of 95%. Finally, we note that the conventional CI can substantially under-cover, consistent with the discussion in Section (ref).

table[table omitted — 964 chars of source]

Next, we compare the length quantiles of the CIs with correct coverage for these same DGPs. Figure (ref) plots the ratios of the $5^{th}$, $25^{th}$, $50^{th}$, $75^{th}$ and $95^{th}$ quantiles of the length of the conditional, projection and hybrid CIs relative to those same length quantiles of the projection CI. As can be seen from the figure, the conditional CI has the tendency to become very long, especially at high quantile levels for certain DGPs, whereas the hybrid CI tends to perform the best overall by limiting the worst-case length performance of the conditional CI relative to the projection CI. Relative to projection, the hybrid CI enjoys length reductions of 20-30% for favorable DGPs while only showing length increases of 5-10% for unfavorable DGPs.

figure[figure omitted — 650 chars of source]

Application to Dynamic Treatment Choice

We revisit Han24's (Han24) application. Han24 considers schooling and post-school training as a sequence of treatments and estimates the partial ordering of treatment regimes and the identified set of the optimal regime. Han24 considers the Job Training Partnership Act (JTPA) program for post-school training and a high school diploma (HS) (or its equivalents) for schooling. He considers high school diplomas rather than college degrees because the former is more relevant for the disadvantaged population of Title II of the JTPA program. In this paper, we are interested in conducting inference on welfare---earnings---evaluated at the regime chosen in a data-driven manner. The dataset is constructed by combining the JTPA data with the US Census and the National Center for Education Statistics (NCES). The following is the set of variables: $Y_{2}$ is an indicator for whether the individual is above or below the median of 30-month earnings, $D_{2}$ is an indicator for whether the individual participated in the job training program, $Z_{2}$ is an indicator for whether the individual was (randomly) assigned to the program, $Y_{1}$ is an indicator for whether the individual is above or below the 80th percentile of pre-program earnings, $D_{1}$ is an indicator for whether the individual received a HS diploma or GED, and $Z_{1}$ is an indicator for the density of high schools Nea97.\footnote{Specifically, $Z_{1}=1$ if the number of high schools per square mile in each training site (e.g. a city) is above 35.} The number of individuals in the sample is 9,223. We assume $Z\bot(Y(d),D(z))$.

Following Han24, consider the dynamic treatment regime $\boldsymbol{\delta}(\cdot)\equiv(\delta_{1},\delta_{2}(\cdot))\in \mathcal{D}^*$, where $\delta_{1}$ indicates receipt of a HS diploma and $\delta_{2}(y_{1})$ indicates receipt of the job training program given pre-program earning type $y_{1}$. By having $\delta_{2}$ as a function of $y_{1}$, the allocation decision adaptively incorporates information about unobserved characteristics of the individuals reflected in the response $Y_{1}$ to allocation $\delta_{1}$. The counterfactual earning type in the terminal stage given $\boldsymbol{\delta}(\cdot)$ is defined as $Y_{2}(\boldsymbol{\delta}(\cdot))\equiv Y_{2}(\delta_{1},\delta_{2}(Y_{1}(\delta_{1})))$, where $Y_{1}(\delta_{1})$ is the counterfactual earning type in the first stage given $\delta_{1}$. All possible regimes in $\mathcal{D}^*$ are listed in Table (ref). Suppose $\boldsymbol{\delta}^{*}$ is the optimal regime that maximizes the average terminal earning $W(\boldsymbol{\delta})\equiv \mathbb{E}[Y_{2}(\boldsymbol{\delta}(\cdot))]$ as welfare. We are interested in constructing CIs for $W(\boldsymbol{\delta})$ evaluated at $\boldsymbol{\delta}=\hat{\boldsymbol{\delta}}$, where $\hat{\boldsymbol{\delta}}$ is calculated from the estimated identified set of $\boldsymbol{\delta}^{*}$.

table[table omitted — 615 chars of source]

We first derive analytical bounds on the welfare under no additional assumptions. The distribution of data is expressed as the vector $p$ of the form

align*[align* omitted — 153 chars of source]

The welfare $W(\boldsymbol{\delta})\equiv \mathbb{E}[Y_{2}(\boldsymbol{\delta}(\cdot))]$ can be expressed as

align*[align* omitted — 331 chars of source]

by the law of iterated expectation. To derive bounds on $W(\boldsymbol{\delta})$, first consider bounds on $W_{y}(d)\equiv \mathbb{P}[Y(d)=y]$ for $d\equiv(d_{1},d_{2})$, which are $L_{y}(d)\equiv\max_{z}L_{y}(d;z)$ and $U_{y}(d)\equiv\min_{z}U_{y}(d;z)$ where

align[align omitted — 438 chars of source]

Using these bounds, we can calculate bounds on

align*[align* omitted — 361 chars of source]

which are

align[align omitted — 342 chars of source]

For example, for the fourth regime in Table (ref),

align*[align* omitted — 361 chars of source]

is bounded by

align*[align* omitted — 191 chars of source]

Since $\max_{z}L(z)+\max_{z}\tilde{L}(z)=\max_{z,\tilde{z}}\left\{ L(z)+\tilde{L}(\tilde{z})\right\} $ for any functions $L$ and $\tilde{L}$, we can express (ref) as

align*[align* omitted — 245 chars of source]

and, analogously, (ref) as $U(\boldsymbol{\delta})\equiv\max_{z,\tilde{z}}U(\boldsymbol{\delta};z,\tilde{z})$. Therefore the bounds satisfy Assumption (ref). Now, consider choosing a single optimal regime that maximizes $L(\boldsymbol{\delta})$. Then, for example, we can show that the following event can be characterized as a polyhedron in the space of $p$

align*[align* omitted — 231 chars of source]

where $(z^{*},\tilde{z}^{*})=\arg\max_{z,\tilde{z}}L(\boldsymbol{\delta}_{(4)};z,\tilde{z})$. Note that this event is equivalent to $\{L(\boldsymbol{\delta}_{(4)};z^{*},\tilde{z}^{*}) \ge L(\boldsymbol{\delta};z,\tilde{z})\text{ for all }\boldsymbol{\delta}\text{ and }z,\tilde{z}\}$ or equivalently,

align[align omitted — 195 chars of source]

Note that each $L(\boldsymbol{\delta};z,\tilde{z})$ is a linear combination of the elements in $p$ as shown in (ref) and (ref). More formally, we can show that (ref) is equivalent to $A_{L}p \le0$ for some $A_{L}$. Note that

align*[align* omitted — 231 chars of source]

for some row vectors $A^{1}$ and $A^{0}$ because, for $\boldsymbol{\delta}=\boldsymbol{\delta}_{(4)}$,

align*[align* omitted — 351 chars of source]

An analogous argument can be applied to $U(\boldsymbol{\delta};z,\tilde{z})$. This characterization implies that Assumption (ref) holds in this setting. Also, this characterization facilitates the CI calculations in Sections (ref)--(ref).

Table (ref) reports 95% CIs for the welfare selected by maximizing the estimated lower bound on the welfare. Recall that the welfare is the probability that the 30-month earnings is above the median. It is notable that the hybrid CI is shorter than the conventional CI even though the latter does not have (uniformly) correct coverage. We can also see that although the hybrid CI is not quite contained in the projection CI, it is somewhat shorter. Finally, the conditional CI has infinite length in this example, demonstrating an extreme example of how the conditional CIs can be uninformatively long (see KL21).

table[table omitted — 499 chars of source]

Appendix (ref) contains Monte Carlo simulated coverage frequencies of the CIs with various DGPs in the setting of this application. Overall, the findings are consistent with the ones in Section (ref).

appendix

Additional Examples

This appendix contains examples in addition to Section (ref), showing how they fall under the general framework of this paper.

Revisiting Manski Bounds with a Continuous Outcome

We revisit the example with Manski bounds in Section (ref), now allowing for a continuous outcome with bounded support. Let $Y\in[y^{l},y^{u}]$ be continuously distributed and assume $\mathbb{E}[Y(d)|Z]=\mathbb{E}[Y(d)]$ for $d\in\{0,1\}$. Note that the sharp bounds on $W(d)\equiv \mathbb{E}[Y(d)]$ are

align*[align* omitted — 249 chars of source]

and the sharp bounds on the ATE are $L(1,0)$ and $U(1,0)$ for $L(\tilde{d},d) \equiv L(\tilde{d})-U(d)$ and $U(\tilde{d},d) \equiv U(\tilde{d})-L(d)$. We can define the identified set $\mathcal D^* \subseteq \mathcal D$ of optimal treatments as

align*[align* omitted — 185 chars of source]

Then, Assumption (ref) holds with

align*[align* omitted — 274 chars of source]

and $\tilde{\ell}_{0,j} =y^{l}$ and $\tilde{\ell}_{1,j} =0$ for $j\in\{0,1\}$ and

align*[align* omitted — 329 chars of source]

and symmetrically for $\tilde{u}_{d,j}$ and $u_{d,j}$. For $\widehat{L}(d)$ and $\widehat{U}(d)$ being the sample counterparts of $L(d)$ and $U(d)$ upon replacing $p$ with $\hat p$, Assumption (ref) holds by the argument in the paragraph after Assumption (ref) since

align*[align* omitted — 131 chars of source]

For Assumption (ref), let $[y^{l},y^{u}]=[0,1]$ for simplicity. Note that

align*[align* omitted — 203 chars of source]

where the first equality uses integration by parts. Therefore, we can estimate the elements of $p$ by sample means, forming $\hat{p}$.

Empirical Welfare Maximization via Linear Programming

Continuing from Section (ref), we show how sharp bounds on $W(\delta)$ can be computed using linear programming. This can be done by extending the example in Section (ref) with binary $Y$. Again, let $\mathcal{X}=\{x_{1},...,x_{K}\}$. Let $q(e|x)\equiv \mathbb{P}(\varepsilon=e|X=x)$. In analogy,

align*[align* omitted — 99 chars of source]

where $q(x)$ is a vector with entries $q(e|x)$ across $e$, and

gather*[gather* omitted — 148 chars of source]

where the first equality holds by the independence assumption in the previous section. Then the constraint for each $x\in\mathcal{X}$ becomes

align*[align* omitted — 28 chars of source]

where $p(x)$ is the vector of $p(y,d|z,x)$'s across $(y,d,z)$ fixing $x$. Now we can construct a linear program for welfare:

align*[align* omitted — 194 chars of source]

Therefore $W(\delta)$ satisfies the structure of Section (ref) and by analogous arguments, Assumptions (ref), (ref) and (ref) hold.

Additional Monte Carlo Simulations

In analogy with Section (ref), we compute the unconditional coverage frequencies of the conditional, projection and hybrid CIs for DGPs in the dynamic treatment regime setting of the empirical application (Section (ref)). In particular, we consider two DGPs: DGP 1 generates $\hat p$ from a multinomial distribution based on $p_{d_1,y_1,d_2,y_2|z_1,z_2}=0.25$ for $(d_1,y_1,d_2,y_2)\in\{(1,0,0,1),(1,1,1,1)\}$ and all $(z_1,z_2)$ and $p_{d_1,y1,d_2,y_2|z_1,z_2}=0.0357$ for all other $(d_1,y_1,d_2,y_2)$ and all $(z_1,z_2)$; DGP 2 generates $\hat p$ based on $p_{d_1,y_1,d_2,y_2|z_1,z_2}=0.375$ for $(d_1,y_1,d_2,y_2)\in\{(1,0,0,1),(1,1,1,1)\}$ and all $(z_1,z_2)$ and $p_{d_1,y1,d_2,y_2|z_1,z_2}=0.0179$ for all other $(d_1,y_1,d_2,y_2)$ and all $(z_1,z_2)$. The coverage frequencies of the CIs are reported in Table (ref). Again, consistent with the asymptotic results of Theorems (ref)--(ref), the conditional, projection and hybrid CIs all have correct coverage for all DGPs and the sample size of $n=1000$. Also, the projection CI tends to be conservative with true coverage above the nominal level of 95%, and the conventional CI can substantially under-cover.

table[table omitted — 836 chars of source]

Figure (ref) plots the ratios of the $5^{th}$, $25^{th}$, $50^{th}$, $75^{th}$ and $95^{th}$ quantiles of the length of the conditional, projection and hybrid CIs relative to those same length quantiles of the projection CI. The figure shows that the conditional CI has the tendency to become very long, especially at high quantile levels for certain DGPs, whereas the hybrid CI tends to perform the best overall by limiting the worst-case length performance of the conditional CI relative to the projection CI. Relative to projection, the hybrid CI enjoys length reductions of 10-20% for favorable DGPs while showing length increases of 5-10% for unfavorable DGPs.

figure[figure omitted — 493 chars of source]