Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
41,560 characters · 8 sections · 26 citation commands
On the power properties of inference for parameters with interval identified sets
\relax \hypersetup{pageanchor=false} \hypersetup{pageanchor=true}
\thispagestyle{empty}
KEYWORDS: bounds, interval identified set, partial identification, confidence intervals, hypothesis testing, local power analysis.
JEL classification codes: C01, C12.
\thispagestyle{empty}
This paper contributes to the literature of inference in partially-identified econometric models. Our setup is as in imbens/manski:2004 and stoye:2009 where the econometric model indicates that for a data distribution $P$ the real-valued parameter of interest $\theta_{0}(P)$ belongs to an interval identified set $[ \theta_{l}(P), \theta_{u}(P) ]$. We focus on the case in which the endpoints of the interval do not cross, i.e., $\theta_{l}(P) \leq \theta_{u}(P)$, so the identified set is non-empty.
We assume the researcher can implement asymptotically exact inference for the partially identified parameters based on asymptotically normal estimators of the identified set's endpoints. That is, we assume that the availability of a pair of estimators $( \hat{\theta}_{l},\hat{\theta}_{u}) $ such that
uniformly in a suitable set of distributions, along with a uniformly consistent estimator of $\Sigma ( P) $. Furthermore, we assume that the bounds estimators are “ordered”, in the sense that $\hat{\theta}_{l}\leq \hat{\theta}_{u}$ with probability one. We refer to these conditions as the “ordered bounds setup” or OBS.
Under our conditions, the researcher has several options to construct a confidence interval (CI, henceforth) for $\theta _{0}( P) $ with a confidence level of $(1-\alpha)$. The first option is the CI proposed by imbens/manski:2004, which we denote by $CI_{\alpha }^{1}$. The other two options were introduced by stoye:2009, and we denote them as $CI_{\alpha }^{2}$ and $CI_{\alpha }^{3}$. See Section (ref) for a detailed description of these CIs. The results in stoye:2009 shows that $CI_{\alpha }^{1}$, $CI_{\alpha }^{2}$, and $CI_{\alpha }^{3}$ are all uniformly asymptotically valid and exact. This implies that the three CIs are equivalent in terms of uniform coverage of parameters in the identified set, which are all the valid candidates for $\theta_{0}(P)$.
The results mentioned above do not speak about the ability of these CIs to rule out parameters outside of the identified set, which are not valid candidates for $\theta_{0}(P)$. That is, they are silent about the statistical power of inference associated with the three CIs. In fact, to our knowledge, the literature has not compared these CIs in terms of power. Our first contribution is to conduct this comparison. To this end, we study the limiting coverage probability of the three CIs for all possible sequences of parameters that do not belong to the identified set. A higher limiting coverage rate for these parameters translates into a lower power. We formally show that $CI_{\alpha }^{1}$ and $CI_{\alpha }^{2}$ are equally powerful, and both dominate $CI_{\alpha }^{3}$.
For our second result, we consider the favorable situation in which the researcher has not one but two pairs of estimators that satisfy (ref), and one of them is known to be more efficient than the other, in the sense of having smaller diagonal elements in their asymptotic variance.\footnote{This is weaker than the standard definition of efficiency, where the difference of the asymptotic covariance matrices of the efficient and inefficient estimators, respectively, is negative semi-definite.} While one could implement asymptotically exact CIs for $\theta _{0}( P) $ using either one of these pairs of estimators, it is reasonable to expect that inference based on the more efficient pair is preferable. In particular, one would expect that the more efficient estimator would result in more powerful inference. We formally demonstrate that this result generally holds for $CI_{\alpha }^{1}$ and $CI_{\alpha }^{2}$, but not for $CI_{\alpha }^{3}$. Specifically, for $CI_{\alpha }^{1}$ and $CI_{\alpha }^{2}$, inference based on the more efficient bounds estimators is always at least as powerful as inference based on the less efficient estimators. In contrast, for $CI_{\alpha }^{3}$, the reverse can occur. That is, inference based on the more efficient estimators can be strictly less powerful than inference based on the less efficient estimators. We explain this counterintuitive phenomenon and provide the conditions under which it arises.
The motivation behind our contributions arises from a concrete empirical application. In bugni/gao/obradovic/velez:2024a, we investigate inference for treatment effect parameters such as the average treatment effect (ATE) in the context of a randomized controlled trial (RCT) with imperfect compliance. In this setting, the ATE is partially identified, with its identified set being an interval. We can also propose consistent and asymptotically normal estimators of these bounds that satisfy the OBS. Accordingly, we can implement asymptotically valid and exact inference using either $CI_{\alpha}^{1}$, $CI_{\alpha}^{2}$, or $CI_{\alpha}^{3}$. Our first contribution demonstrates that $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$ are equally powerful, and both of these are more powerful than $CI_{\alpha}^{3}$. Based on this result, we choose to conduct inference using either $CI_{\alpha}^{1}$ or $CI_{\alpha}^{2}$, instead of $CI_{\alpha}^{3}$.
Within the same empirical application, we have two possible implementations of our bounds estimators for the ATE: we can estimate treatment probabilities using sample frequencies or exact probabilities (known in an RCT). By analogous arguments to those in hahn:1998 and hirano/imbens/ridder:2003, the bounds estimators that use sample analogs are shown to be more efficient than those using exact probabilities. Intuitively, one would expect that the more efficient implementation of the bounds estimator produces a more powerful inference of the partially-identified parameter value. Our second contribution shows that this intuitive result holds for $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$, but may fail to hold for $CI_{\alpha}^{3}$. That is, it is possible for the more efficient bounds estimator to result in less powerful inference when $CI_{\alpha}^{3}$ is used. Based on the two contributions, we suggest using $CI_{\alpha}^{1}$ or $CI_{\alpha}^{2}$ under OBS.
The rest of the paper is organized as follows. Section (ref) describes the econometric model. Section (ref) presents our setup (i.e., OBS), and Section (ref) details the three CIs. We present our main results in Section (ref), which is divided into two subsections. Section (ref) compares the power of inference across the CIs. Section (ref) compares the power of inference for each CI when two bounds estimators are available, with one being more efficient than the other. Section (ref) concludes. The paper's appendix collects all of the proofs, along with intermediate results.
For a data distribution $P$, the real-valued parameter of interest is denoted by $\theta _{0}(P)$. The econometric model indicates that $\theta _{0}(P)$ belongs to an interval identified set $\Theta _{I}(P) = [\theta _{l}(P), \theta _{u}(P)]$. Moreover, we assume that the identified set is non-empty, i.e., $\theta _{l}(P) \leq \theta _{u}(P)$.
We consider a CI $CI_{\alpha }$ that covers $\theta _{0}(P)$ with a minimum prespecified coverage probability of $(1-\alpha)$ as the sample size $N$ increases. Furthermore, we require this coverage condition to be satisfied uniformly for all parameters in the identified set $\Theta _{I}(P)$ and for all probability distributions $P$ in a suitable space $\mathcal{P}$. Specifically, we require our CI to be {\it uniformly asymptotically valid} and, if possible, {\it uniformly asymptotically exact}, which we define next.
$CI_{\alpha }$ for $\theta _{0}(P) \in \Theta _{I}(P)$ is {\it uniformly asymptotically valid} if it satisfies the following property:
Moreover, $CI_{\alpha }$ is {\it uniformly asymptotically exact} if (ref) holds with equality, i.e.,
By definition, a CI that is uniformly asymptotically exact is also uniformly asymptotically valid.
The remainder of this section is organized as follows. Section (ref) specifies our main assumptions, referred to as OBS. Section (ref) describes the three CIs considered in this paper. Under OBS, these three CIs are uniformly asymptotically exact.
Throughout our paper, we assume that the researcher can construct estimators of the bounds of the identified set and its limiting distribution that satisfy the following condition.
The set of distributions $\mathcal{P}$ in Definition (ref) encodes high-level assumptions that facilitate our asymptotic analysis. In applications, one would typically impose low-level conditions (e.g., i.i.d.\ sampling and bounded moments) to establish conditions (b)-(c) in Definition (ref) for a suitable estimator $(\hat{\theta}_{l},\hat{\theta}_{u},\hat{\sigma}_{l},\hat{\sigma}_{u},\hat{\rho})$ and set of distributions $\mathcal{P}$. In this paper, we choose to leave $\mathcal{P}$ unspecified to maintain generality.
The OBS implies all of the assumptions required by stoye:2009. In particular, the ordering condition $P(\hat{\theta}_{l} \leq \hat{\theta}_{u}) = 1$ implies his superefficiency condition; see stoye:2009. This implies that one can use an estimator $(\hat{\theta}_{l}, \hat{\theta}_{u}, \hat{\sigma}_{l}, \hat{\sigma}_{u}, \hat{\rho})$ that satisfies OBS to implement any of the CIs proposed in that paper to achieve uniformly asymptotically exact inference. The next section reviews these CIs.
This paper considers three CIs. Our first CI is $CI_{\alpha }^{1}$, originally proposed by imbens/manski:2004, and revisited by stoye:2009. Given an estimator $(\hat{\theta}_{l},\hat{\theta}_{u},\hat{\sigma}_{l},\hat{\sigma}_{u},\hat{\rho})$ and a confidence level $(1-\alpha)$, $CI_{\alpha }^{1}$ is defined as follows:
where $c^{1}$ solves
Provided that $\max \{\hat{\sigma}_{l},\hat{\sigma}_{u}\}>0$, it follows that $c^{1}$ is uniquely determined by (ref).\footnote{Under our OBS assumptions, $\max \{\hat{\sigma}_{l},\hat{\sigma}_{u}\}>0$ occurs with probability approaching one.} Under OBS, imbens/manski:2004 implies that $CI_{\alpha }^{1}$ is uniformly asymptotically valid, while stoye:2009 shows that $CI_{\alpha }^{1}$ is uniformly asymptotically exact.
Our second CI is $CI_{\alpha }^{2}$, developed by stoye:2009. Given a generic estimator $(\hat{\theta}_{l},\hat{\theta}_{u},\hat{\sigma}_{l},\hat{\sigma}_{u},\hat{\rho})$ and a confidence level $(1-\alpha)$, $CI_{\alpha }^{2}$ is defined as follows:
where $(c_{l}^{2},c_{u}^{2})$ solves {
} where $(z_{1},z_{2})\sim N(\mathbf{0}_{2\times 1},\mathbf{I}_{2\times 2})$. It is unclear to us whether (ref) is guaranteed to have a unique solution.\footnote{stoye:2009 states that typically $(c_{l}^{2},c_{u}^{2})$ is uniquely determined by the fact that both constraints in (ref) hold with equality.} Be that as it may, our formal arguments will only require the researcher to choose $ (c_{l}^{2},c_{u}^{2})$ arbitrarily whenever (ref) has multiple solutions.
As explained in stoye:2009, $CI_{\alpha }^{2}$ calibrates the critical values $(c_{l}^{2},c_{u}^{2})$ taking into account that the underlying problem is bivariate. In this sense, $CI_{\alpha }^{2}$ is considered an improvement upon $CI_{\alpha }^{1}$. In fact, stoye:2009 argues that $CI_{\alpha }^{2}$ is the shortest CI with correct nominal size. stoye:2009 shows that $CI_{\alpha }^{2}$ is uniformly asymptotically exact under OBS.
Our last CI is $CI_{\alpha }^{3}$, also developed by stoye:2009. Unlike $ CI_{\alpha }^{1}$ and $CI_{\alpha }^{2}$, $CI_{\alpha }^{3}$ was introduced as a CI that does not require the so-called superefficiency condition (Assumption 3 in stoye:2009) for its validity. The superefficiency condition is guaranteed under OBS, but may fail in other contexts. To implement $CI_{\alpha }^{3}$, the researcher must define a preassigned sequence of constants $\left\{ b_{N}\right\} _{N\geq 1}$ that satisfies $ b_{N} \to 0$ and $b_{N}\sqrt{N} \to \infty $. These types of sequences are common in the partial identification literature; e.g., see andrews/soares:2010,bugni:2010,bugni:2015. For example, these papers suggest sequences such as $b_{N}=\ln N/\sqrt{N}$, $b_{N}=\sqrt{ \ln \ln N}/\sqrt{N}$, or $b_{N}=N^{-c}$ for any $c\in (0,1/2)$.
Given an $(\hat{\theta}_{l},\hat{\theta}_{u},\hat{\sigma}_{l}, \hat{\sigma}_{u},\hat{\rho})$, a confidence level $(1-\alpha)$, and a sequence $\left\{ b_{N}\right\} _{N\geq 1}$, $CI_{\alpha }^{3}$ is defined as follows:
where $(c_{l}^{3},c_{u}^{3})$ solves {
} where $(z_{1},z_{2})\sim N(\mathbf{0}_{2\times 1},\mathbf{I}_{2\times 2})$. Once again, we are unclear whether (ref) is guaranteed to have a unique solution. In any case, our arguments only require the researcher to choose $(c_{l}^{3},c_{u}^{3})$ arbitrarily whenever (ref) has multiple solutions. stoye:2009 shows that $CI_{\alpha }^{3}$ is uniformly asymptotically exact under OBS.
Our goal in this paper is to compare the power of inference based on CIs for the partially identified parameter $\theta_{0}(P) \in \Theta_{I}(P)$. By the duality between CIs and hypothesis testing, we can investigate the power of an inference method based on a CI by deriving its limiting coverage probability for a parameter value $\theta$ outside $\Theta_{I}(P)$. With an interval identified set $\Theta_{I}(P) = [\theta_l(P), \theta_u(P)]$, this means that either $\theta < \theta_l(P)$ or $\theta > \theta_u(P)$. Since $\theta$ does not belong to the identified set $\Theta_{I}(P)$, it cannot be the true parameter value $\theta_{0}(P)$. Thus, a lower limiting coverage probability for $\theta$ implies higher statistical power against the (incorrect) null hypothesis $H: \theta_{0}(P) = \theta$.
Based on the previous discussion, our paper compares the limiting coverage probability of parameter values outside $\Theta_{I}(P)$ for various CIs. Importantly, our analysis allows both the parameter value and the data distribution to drift arbitrarily with the sample size. Specifically, we consider all possible sequences $\{(P_{N}, \theta_{N}) \in \mathcal{P} \times \Theta_{I}(P_{N})^{c}\}_{N\in \mathbb{N}}$. Thus, our results include power analysis for fixed alternatives, i.e., $\theta_{N} = \bar{\theta} \not\in \Theta_{I}(P_{N})$, as well as local alternatives, i.e., $\theta_{N} \uparrow \theta_{l}(P_{N})$ or $\theta_{N} \downarrow \theta_{u}(P_{N})$.
Our sole result in this section is Theorem (ref). This result compares the limiting coverage probability of the three CIs for sequences of parameters outside $\Theta_{I}(P)$.
The proof of this theorem, along with all other proofs, appears in the Appendix. Theorem (ref) compares the limiting behavior of the coverage probabilities for any sequence $\{(P_{N}, \theta_{N}) \in \mathcal{P} \times \Theta_{I}(P_{N})^{c}\}_{N \in \mathbb{N}}$ across the three CIs. As mentioned earlier, the result allows both parameters $\theta_{N}$ and distributions $P_{N}$ to drift arbitrarily with the sample size, so it encompasses the power comparison for fixed or local alternatives.
The result has several parts. Part (a) states that $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$ are equivalent in terms of power. For all sequences of parameters such that $\theta_{N} \not\in \Theta_{I}(P_{N})$, the difference in the coverage rates of these CIs converges to zero. To explain this result, it is useful to split the analysis into two mutually exclusive cases: “short” identified sets and “long” identified sets. We say that the identified set is “short” if the distance between the lower and upper bounds is of order $O(1/\sqrt{N})$, and we say that the identified set is “long” otherwise. For long identified sets, $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$ agree on considering the situation as a one-sided testing problem. This implies that both CIs agree on using the $(1-\alpha)$-quantiles of the normal distribution, leading to identical limiting coverage rates. In the case of short identified sets, the ordered nature of the bounds implies that the estimators have a degenerate asymptotic distribution, or would otherwise cross with positive probability.\footnote{See Lemma (ref) for a precise statement of this result.} This degeneracy in the asymptotic distribution implies that $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$ also agree on the critical values, leading to identical limiting coverage rates. Combining both cases, we deduce that the two CIs have identical limiting coverage rates.
Parts (b)-(d) compare the power of $CI_{\alpha}^{3}$ with the other CIs for all sequences of parameters $\theta_{N} \not\in \Theta_{I}(P_{N})$. Notably, our conclusions hold regardless of the choice of $\{b_{N}\}_{N \in \mathbb{N}}$ used in the implementation of $CI_{\alpha}^{3}$, provided, of course, that $b_{N} \to 0$ and $b_{N}\sqrt{N} \to \infty$. Part (b) considers models with short identified sets. In this case, $CI_{\alpha}^{3}$ is equivalent to the other two CIs in terms of power. The intuition is similar to that given earlier: the ordered bounds lead to a degenerate asymptotic distribution, resulting in all three CIs agreeing on their critical values. Parts (c)-(d) pertain models with long identified sets. Part (c) shows that, in general, the $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$ are equally or more powerful than $CI_{\alpha}^{3}$. The intuition behind this result stems from the fact that $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$ both use critical values equal to the $(1-\alpha)$-quantiles of the normal distribution. While the critical values used by $CI_{\alpha}^{3}$ are complex functions of the limits of local parameters, they are always larger than or equal to the $(1-\alpha)$-quantiles of the normal distribution to guarantee uniform asymptotic validity. Finally, part (d) confirms that there are sequences of models in which $CI_{\alpha}^{3}$ is strictly less powerful than the other two CIs. Our argument is based on considering data-generating processes for which the critical values of $CI_{\alpha}^{3}$ are strictly larger than the $(1-\alpha)$-quantiles of the normal distribution, resulting in a power loss relative to the other two CIs. Importantly, these data-generating processes exist for any choice of the sequence $\{b_{N}\}_{N \in \mathbb{N}}$.
Theorem (ref) provides a comprehensive comparison of the relative power properties of the CIs. It states that, under OBS, $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$ are equally powerful, and both dominate $CI_{\alpha}^{3}$.
The proof of Theorem (ref) is based on three auxiliary results that characterize the limiting coverage rates for suitable sequences $\{(P_N, \theta_{N})\}_{N \in \mathbb{N}}$ with $P_N \in \mathcal{P}$ and $\theta_{N} \not\in \Theta_{I}(P_{N})$ for all $N \in \mathbb{N}$. These results are presented in Lemmas (ref), (ref), and (ref), corresponding to $CI_{\alpha}^{1}$, $CI_{\alpha}^{2}$, and $CI_{\alpha}^{3}$, respectively. We believe these auxiliary results may be of independent interest beyond this study.
This section considers a situation where the researcher has two bounds estimators to construct the CIs for $\theta_{0}(P) \in \Theta_{I}(P)$, with one of these bounds estimators being more efficient than the other. The goal of this section is to compare the power of the inference resulting from these two bounds estimators. The next assumption formalizes the setup.
Assumption (ref) describes the two bounds estimators that satisfy OBS. We use the superscript $E$ for the efficient estimator and $I$ for the inefficient estimator. Assumptions (ref)(a)-(b) imply that both $(\hat{\theta}_{l}^{E}, \hat{\theta}_{u}^{E})$ and $(\hat{\theta}_{l}^{I}, \hat{\theta}_{u}^{I})$ are asymptotically normal estimators of the bounds of the identified set (i.e., $(\theta_{l}(P), \theta_{u}(P))$), and that we consistently estimate their limiting variance. Moreover, these results hold uniformly for all distributions $P \in \mathcal{P}$. This means that either estimator can be used to construct CIs that are uniformly asymptotically exact. Assumption (ref)(c) specifies the sense in which $(\hat{\theta}_{l}^{E}, \hat{\theta}_{u}^{E})$ is more efficient than $(\hat{\theta}_{l}^{I}, \hat{\theta}_{u}^{I})$: for both the lower and upper bounds, the asymptotic variance of each efficient estimator is less than or equal to that of the inefficient estimator. We note that this condition is weaker than the usual relative asymptotic efficiency comparison, which states that, for all $P\in \mathcal{P}$, {
} is a negative semi-definite matrix.
We seek to compare the power of inference based on CIs constructed from the inefficient and efficient bounds estimators. As discussed in Section (ref), we compare their power by contrasting the limiting coverage rates of the corresponding CIs for sequences of parameter values outside the identified set. Since these parameters lie outside the identified set, a lower limiting coverage probability implies higher statistical power against (incorrect) null hypotheses.
As in Section (ref), our analysis considers all possible sequences $\{(P_{N}, \theta_{N}) \in \mathcal{P} \times \Theta_{I}(P_{N})^{c}\}_{N \in \mathbb{N}}$. Thus, our results include power analysis for fixed alternatives, i.e., $\theta_{N} = \bar{\theta} \not\in \Theta_{I}(P_{N})$, as well as local alternatives, i.e., $\theta_{N} \uparrow \theta_{l}(P_{N})$ or $\theta_{N} \downarrow \theta_{u}(P_{N})$.
Our first result in this section compares the limiting coverage rates of inference based on $CI_{\alpha}^{1}$ when implemented with the efficient and inefficient estimators.
In simple terms, Theorem (ref) shows that $CI_{\alpha}^{1}$ based on the more efficient bounds estimator (i.e., $CI_{1}^{E}$) is more powerful than when it is based on the less efficient bounds estimator (i.e., $CI_{1}^{I}$). While both CIs satisfy the coverage goal in (ref) with equality, Theorem (ref) demonstrates that the former dominates the latter in terms of power.
Our second result compares the limiting coverage rates of inference based on $CI_{\alpha}^{2}$ when implemented with the efficient and inefficient estimators.
Our takeaway from Theorem (ref) is analogous to that of Theorem (ref): $CI_{\alpha}^{2}$ based on the more efficient bounds estimator (i.e., $CI_{2}^{E}$) is more powerful than when it is based on the less efficient bounds estimator (i.e., $CI_{2}^{I}$). Both CIs satisfy the coverage goal in (ref) with equality, but the former is more powerful than the latter.
Results of Theorems (ref) and (ref), while novel, may not seem surprising. However, as the following result shows, conventional intuition need not necessarily hold in the OBS setting for certain CIs, such as $CI_{\alpha}^{3}$. The result compares limiting coverage rates of inference based on $CI_{\alpha}^{3}$ when implemented with the efficient and inefficient estimators.
The result has two parts. Part (a) considers models with “short” identified sets, in which the lower and upper bounds of the identified set are within $O(1/\sqrt{N})$ of each other. For these models, our results are similar to those obtained for $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$. That is, $CI_{\alpha}^{3}$ based on the more efficient bounds estimator (i.e., $CI_{3}^{E}$) is more powerful than when it is based on the less efficient bounds estimator (i.e., $CI_{3}^{I}$).
In turn, part (b) considers models with “long” identified sets, where the lower and upper bounds of the identified set are not within $O(1/\sqrt{N})$ of each other. In this case, it is possible to find data-generating processes in which implementing $CI_{\alpha}^{3}$ with the more efficient bounds estimator (i.e., $CI_{3}^{E}$) is less powerful than when it is based on the less efficient bounds estimator (i.e., $CI_{3}^{I}$). We achieve this result by considering sequences of models with long identified sets where these lengths, given by $\sqrt{N}(\theta_{u}(P_{N}) - \theta_{l}(P_{N}))$, are “comparable” with $\sqrt{N} b_N$. Under this condition, it is possible to generate a situation where the critical values used by $CI_{\alpha}^{3}$ switch between two quantities. Furthermore, it is possible for $CI_{3}^{I}$ to use the smaller critical values more frequently than $CI_{3}^{E}$, generating a power advantage in favor of the inefficient bounds. This phenomenon is counterintuitive and undesirable, as one would hope that using more efficient bounds estimators would always result in more powerful inference.
This paper studies the power properties of CIs for a partially-identified parameter of interest with an interval identified set. We assume that the researcher has bounds estimators to construct the CIs proposed by imbens/manski:2004 and stoye:2009, known as $CI_{\alpha }^{1}$, $CI_{\alpha }^{2}$, and $CI_{\alpha }^{3}$. We also assume these estimators are “ordered” in the sense that the estimator of the lower bound is less than or equal to the estimator of the upper bound.
Under our conditions, the literature has established that $CI_{\alpha}^{1}$, $CI_{\alpha}^{2}$, and $CI_{\alpha}^{3}$ are all uniformly asymptotically valid and exact. However, the literature does not address the ability of these CIs to rule out parameters outside the identified set, which are not valid candidates for $\theta_{0}(P)$. Put differently, the results do not speak to the power of inference associated with these CIs.
In this context, this paper makes two contributions. Our first contribution is to compare the coverage probability of the three CIs for all possible sequences of parameters that do not belong to the identified set. A higher coverage rate for these parameters translates into lower power. We formally show that $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$ are equally powerful, and both dominate $CI_{\alpha}^{3}$.
For our second contribution, we consider a favorable situation in which the researcher has two pairs of estimators to implement these CIs, and one of these pairs is known to be more efficient than the other. In this context, it is reasonable to expect that inference based on the more efficient pair leads to more powerful results. We formally demonstrate that this conclusion holds generally for $CI_{\alpha}^{1}$ and $CI_{\alpha}^{2}$, but not for $CI_{\alpha}^{3}$.