Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
104,379 characters · 22 sections · 113 citation commands
Numerical Analysis of Test Optimality
\thispagestyle{empty} \setcounter{page}{0}
In nonstandard testing environments involving composite null and/or alternative hypotheses, a uniformly most powerful test often does not exist. It is therefore common for researchers to devise implementable ad hoc tests that are designed to have provably correct (asymptotic) size. Although researchers may go to great lengths to demonstrate that their devised ad hoc tests have “good” power properties via Monte Carlo simulation, the theoretical power optimality properties of these tests are frequently unknown and elusive. It is therefore difficult for both those who design hypothesis tests and those who are meant to use them to assess whether the power of a given test can be improved upon by a different test, a central question for its usefulness.
Instances for which researchers have sought to show that an ad hoc test is theoretically optimal in some sense have been limited. Such approaches can be roughly divided into two types of power comparisons. The first compares the power function of the ad hoc test to a point-wise power envelope. This point-wise power envelope is constructed as the power of a collection of point-optimal tests computed at each point in the corresponding collection of point alternative hypotheses for points lying in the composite alternative space (e.g., AMS06; AMY19 and GKM19). Although it is possible for the power function of an ad hoc test to coincide with a power envelope constructed in this fashion, this is rarely the case because this form of power envelope may be unattainable by the power of any single test. To avoid or at least mitigate such unfair comparisons, researchers typically impose additional constraints on the point-optimal tests that are used to produce the point-wise power envelope such as unbiasedness (e.g., the two-sided $t$-test and MM13,MM19), rotational invariance (e.g., AMS06 and MSR21) and similarity (e.g., AMS06 and MO20). However, the imposition of constraints on the underlying collection of (constrained) point-optimal tests still does not generally produce a power envelope that is necessarily attainable by a single test and leads to an optimality analysis that is limited to tests satisfying the constraints. So when the power function of an ad hoc test falls short of a power envelope constructed in this way, what is a researcher to conclude? Is the ad hoc test suboptimal? Or is the power envelope simply an unfair and unattainable comparison because it does not correspond to the power function of any single test of the composite hypothesis?
The second approach to test optimality assessment in the literature compares either (i) the power function of the ad hoc test to the power function of a test that is known to be weighted average power (WAP) maximizing for a particular weight function over the parameter values in the composite alternative or (ii) the WAPs of these two tests, or both (e.g., AP94; AMS08; EM14; EMW15 and Cox24). This approach is also incomplete, leading to its own set of questions. How should the researcher choose the weight function when constructing the WAP-maximizing test under comparison? If the power function or WAP of the ad hoc test is not dominated by that of the WAP-maximizing test, should we conclude that it is optimal or could it be dominated by a WAP-maximizing test for a different set of weights?
In this paper, we propose to numerically determine whether there exists a weight function over the composite alternative that justifies the ad hoc test. More precisely, we propose a numerical procedure for automatically finding the weight function that minimizes the WAP difference between the WAP-maximizing test corresponding to that weight function and the ad hoc test over a discretization of the composite alternative space. This eliminates the ambiguity of the second approach to test optimality implicit in the questions above. When this WAP-difference is (not) approximately equal to zero, this approach immediately allows us to conclude that the ad hoc test under study is (not) effectively optimal in the sense of (not) being an approximately WAP-maximizing test itself. A side benefit to our approach is that in the case for which the ad hoc test is determined to be effectively optimal, we know which weights over the alternative space make it approximately WAP-maximizing, allowing us to interpret how the test “directs power”.
In addition to determining whether an ad hoc test is approximately WAP-maximizing for some weight function, we show the surprising result that our procedure immediately produces an attainable power envelope for the ad hoc test. Over a discretization of the composite alternative space, we provide a set of sufficient conditions under which the power function of the WAP-maximizing test using the weights that minimize the WAP difference with the ad hoc test is a power envelope for the ad hoc test. Since this power envelope is constructed from a single test, it is indeed attainable. In addition, this power envelope is “most favorable” to the ad hoc test in the sense that it is constructed from a single WAP-maximizing test with WAP as close as possible to that of the ad hoc test.
Our numerical procedure is composed of an inner and an outer loop algorithm. The inner loop algorithm computes an approximate WAP-maximizing test for a given set of weights over the discretized alternative space, in analogy with MM13, EMW15 and AFBMOQST25. Our outer loop algorithm, which has no counterparts that we are aware of in the literature, approximates the weight function that minimizes the WAP difference between the corresponding WAP-maximizing test and the ad hoc test over the discretized alternative space.\footnote{KM25 made the first attempt we are aware of at numerically finding weights that minimize the WAP difference between an ad hoc test and a WAP-maximizing test. However, their approach was confined to the particular problem of testing with inequality restrictions on nuisance parameters, carried no formal convergence guarantees and did not contain any results on producing power envelopes. Nevertheless, it helped motivate us to write the current paper.} Using the results mentioned above, the power function of this corresponding WAP-maximizing test can then be used as a power envelope for the ad hoc test. We provide new convergence results for both the inner and outer loop algorithms, formally guaranteeing that they approximate the WAP-maximizing test and weights mentioned above. Our algorithms and their convergence guarantees are general and not case-specific.
One may wonder why it would be more desirable to construct an ad hoc test with correct size and subsequently analyze its optimality properties rather than simply constructing a WAP-maximzing test directly using, e.g., MM13 or EMW15. There are two main reasons: one conceptual and one computational. The first reason for constructing an ad hoc test is that it is often unclear which weight function one should use when constructing a WAP-maximizing test. Our procedure instead simply tells the researcher whether there exists a weight function that rationalizes the ad hoc test they have devised (as well as providing an approximation to it). The second reason is that even with a clear idea of a desirable weight function, computing the corresponding WAP-maximizing test can become computationally prohibitive when the dimension of the alternative parameter space is too high. Since our inner loop algorithm also computes WAP-maximizing tests iteratively, it is also of course subject to this criticism. However, our numerical procedure can be used to provide partial evidence on the optimality of a test if the researcher applies it to lower-dimensional special cases and is able to show that the ad hoc test is optimal in those special cases. On the other hand, a WAP-maximizing test for a lower-dimensional special case does not readily generalize up to higher dimensions.
Finally, we note that discovering an ad hoc test to be effectively optimal for a particular testing problem can be used to motivate the development of tests for related problems. For example, the theoretical optimality results established by AMS06,AMS08 in the specific setting of a homoskedastic linear IV model with randomly sampled data for Mor03's (Mor03) conditional likelihood ratio (CLR) test, apparently helped motivate AM16 and AG19 to produce generalizations of this test to settings that allow for dependent data, nonlinear models and singular variance matrices.
We apply our results and algorithms to two different testing problems, the first of which has been previously analyzed using the first approach to optimality assessment described above and the second of which has been analyzed using the second. For the first problem, we substantively contribute to the ongoing debate over the optimality properties of the CLR test. Specifically, we revisit classic optimality results from AMS06,AMS08 and more recent critiques by AMY19 and VdSW23 in the context of the Gaussian homoskedastic linear IV model. Using the first existing approach to test optimality assessment described above, AMS06,AMS08 show that the CLR test effectively attains the point-wise power envelope for invariant (similar) tests under a design that holds the reduced-form error variance matrix constant. On the other hand, AMY19 argue that the CLR test falls short of the point-wise power envelope obtained by fixing the true value of the parameter of interest and instead varying its hypothesized value, which VdSW23 show is equivalent to producing the more standard point-wise power envelope that fixes the hypothesized value and varies the true value under a design that instead holds the variance matrix of the structural errors constant. Using our new algorithms, we show that the CLR test's power function is virtually indistinguishable from that of a single WAP-maximzing test for both designs, leading us to conclude that the CLR test is in fact effectively optimal in the class of invariant asymptotically efficient tests for this problem.
For our second application, we analyze the optimality properties of the new inequality-imposed confidence interval (IICI) of Cox24, designed to improve inference when a scalar nuisance parameter may lie on or near the boundary of the parameter space. Using the second existing approach to test optimality assessment described above, Cox24 compares the WAP of the test implied by the IICI to that of the WAP-maximizing test proposed for this problem by EMW15, finding them to be very close. We plot the power functions of both tests and find that they intersect, making the optimality properties of the IICI-implied test unclear. Instead, we show that while the power function of the IICI-implied test is very close to the power function of a single most favorable WAP-maximizing test, it does fall slightly short with a maximum gap of about 0.3 percentage points. Nevertheless, its WAP is extremely close to that of the WAP-maximizing test, implying that it is highly-competitive but slightly short of optimal for problems with a scalar nuisance parameter that may lie on or near its boundary.
The scope for other applications of our results and numerical procedure is very wide and we just briefly mention some additional examples here. In addition to analyzing the CLR test, one may seek to analyze the (unknown) optimality properties of other weak IV-robust tests in the literature such as the Lagrange multiplier test of Kle02, which could be particularly interesting due to its non-monotonic power function. CY06's (CY06) test of stock return predictability was not found to be dominated by the WAP-maximizing test with weights chosen by EMW15. Using our approach, it would be interesting to learn if there is a WAP-maximizing test that indeed dominates it. We could use our approach to determine whether the apparent power deficiencies of the test of GKM19 when compared to the point-wise power envelope disappear when using our approach that makes the attainable comparison to a most favorable single WAP-maximizing test. More broadly, our approach could be used to analyze tests in the large literatures on inference for moment inequality models, inference robust to identification failure, inference in structural change models and inference with highly-persistent (e.g., local-to-unit root) processes. We intend to analyze some of these examples in follow-up work. And we of course hope that our results and numerical procedures will be useful for the optimality analysis of new ad hoc tests that have yet to be developed.
The remainder of the paper proceeds as follows. In Section (ref), we provide the intuition underlying our numerical approach by illustrating how to find the weight function that justifies the two-sided $t$-test as a WAP-maximizing test and contrasting it with the point-wise power envelope that is not attained by the two-sided $t$-test. Section (ref) formalizes the general hypothesis testing framework we study, defines the numerical problem of finding weights that make an ad hoc test approximately WAP-maximizing and shows that the resulting WAP-maximizing test yields a “most favorable” approximate power envelope under a set of sufficient conditions. Section (ref) describes our numerical procedure, detailing the inner loop algorithm that computes approximate WAP-maximizing tests and the outer loop algorithm that adjusts weights, along with practical considerations such as switching tests, thresholding, adding support points, and Monte Carlo smoothing. We provide the theoretical convergence guarantees for these algorithms in Section (ref), establishing that our numerical method yields valid approximate WAP-maximizing tests and power envelopes. In Section (ref), we apply our results and numerical procedures to the two testing applications described above. Mathematical proofs and implementation details for the applications are contained in the Appendix.
We begin by analyzing a simple canonical hypothesis testing example to impart intuition and motivate our approach to assessing the optimality of a hypothesis test. Consider a two-sided test of the mean of a Gaussian random variable with known unit variance: $H_0:\beta=0$ vs.\ $H_1:\beta\neq 0$ for $Y\sim\mathcal{N}(\beta,1)$.\footnote{The analysis of this section applies without loss of generality to all two-sided $t$-tests of the mean of a Gaussian random variable with known variance via a simple scale transformation.} The standard level-$\alpha$ two-sided $t$-test that rejects when $|Y|$ exceeds $z_{1-\alpha/2}$ is well-known to be the uniformly most powerful test amongst all unbiased tests. However, suppose that we would not like to confine our analysis to unbiased tests. Since there is no uniformly most powerful test, we take the common approach of comparing the $t$-test to the point-wise power envelope for this testing problem. Specifically, the point-wise power envelope is equal to the rejection probabilities of the collection of the most powerful point-wise tests of $H_0$ vs.\ $H_{\beta'}:\beta=\beta'$, as a function of a collection of $\beta'$ values. By the Neyman-Pearson lemma, we know that each of these point-wise tests is the likelihood ratio test of $H_0$ vs.\ $H_{\beta'}$.
Figure (ref) plots the power function of the two-sided $t$-test and the point-wise power envelope. Notably, the power function of the two-sided $t$-test lies substantially below the power envelope. Using the point-wise power envelope could thus lead us to believe that the two-sided $t$-test is suboptimal. However, note that the point-wise power envelope does not produce a power comparison that is necessarily attainable because it corresponds to the power of a collection of point-optimal tests that are each optimal against a point alternative $H_{\beta'}$ with $\beta'\in\mathbb{R}\setminus\{0\}$, none of which correspond to the composite alternative hypothesis of interest $H_1$. In other words, it is still possible for the two-sided $t$-test of the composite alternative $H_1$ to be optimal even though its power function lies below the power function of a collection of point-wise optimal tests of $H_{\beta'}$\textemdash we may simply be comparing its power function to an unattainable upper bound.
In the context of this testing problem, our proposed approach is to instead determine whether there exists a set of weights over the composite alternative space $\Theta_1=\mathbb{R}\setminus\{0\}$ such that the WAP of the two-sided $t$-test is well-approximated by the WAP of a WAP-maximizing test with this set of weights. If so, we may conclude that the two-sided $t$-test is “nearly optimal” for this set of weights. To address this task, we discretize the composite alternative space $\Theta_1$ into support points $\bar\Theta_1\subset \Theta_1$ and introduce an algorithm that successively adjusts the weights on each support point to produce a WAP-maximizing test with WAP as close as possible to that of the two-sided $t$-test. Surprisingly, we show below that the rejection probabilities of this WAP-maximizing test produce a power envelope for the two-sided $t$-test. An important feature of this power envelope is that it corresponds to the rejection probabilities of a single test and is therefore necessarily attainable, in contrast to the point-wise power envelope. This power envelope is not only attainable but it is a “most favorable” power envelope in the sense that it is based upon a WAP-maximizing test with WAP as close as possible to that of the two-sided $t$-test itself.
For the sake of illustration, let us work with the coarse discretization $\bar\Theta_1=\{-1,1\}$. The WAP-maximizing test of $H_0$ vs.\ the discretized composite alternative $\bar H_1:\beta\in \bar{\Theta}_1$ that weights $\beta=1$ by $\omega_1$ (and $\beta=-1$ by $1-\omega_1$) rejects when \[ LR(Y) = \frac{(1-\omega_1) f(Y;-1) + \omega_1 f(Y;1)}{f(Y;0)} > \text{cv} \] for $\text{cv}$ satisfying $\mathbb{P}_{H_0}(LR(Y)>\text{cv}) = \alpha$, where $f(\cdot;\mu)$ denotes the density function of a $\mathcal{N}(\mu,1)$ random variable and $\mathbb{P}_{H_0}$ denotes the probability under $H_0$. Our approach begins with some initial weight $\omega_1$, computes the power function of the WAP-maximizing test using this weight and adjusts the value of $\omega_1$ according to the difference between the power function of the $t$-test and the WAP-function at $\beta=1$: if the power of the $t$-test lies below that of the WAP-maximizing test at $\beta=1$, we adjust $\omega_1$ downwards, putting more weight on $\beta=-1$, and vice versa. We then repeat this procedure iteratively until the power function of the $t$-test lies weakly below that of the WAP-maximizing test at all points $\beta\in \bar\Theta_1$. If the power function of the $t$-test lies weakly below that of the WAP-maximizing test at all points $\beta\in \Theta_1$, the power function of this final WAP-maximizing test produces a power envelope that is based upon a single test. If the power function of the $t$-test lies within a small value of the power function of this final WAP-maximizing test, we deem the $t$-test to be “effectively optimal” since its rejection probabilities nearly coincide with those of a test with known optimality. The weights given in the final iteration tell us how the nearly optimal $t$-test weights points in the alternative space since they are the weights for which the corresponding WAP-maximizing test has WAP nearly equal to that of the $t$-test. Otherwise, we deem the $t$-test to be suboptimal since its power function is dominated by that of another test.
We provide the specifics of how the weight $\omega_1$ is adjusted at each iteration in the outer loop algorithm in Section (ref), along with a theoretical justification for the algorithm in Section (ref), but the intuition is as follows. Since any WAP-maximizing test is optimal for its set of weights by definition, we know (i) its power function cannot be dominated by that of the $t$-test on $\bar\Theta_1$ and (ii) if the $t$-test is itself WAP-optimal, there exists a set of weights for which its power function must match that of a WAP-maximizing test. These two facts justify our iterative weighting adjustment: at each iteration, we want to increase the weight at support points for which the WAP-maximizing test has lower power than the $t$-test and therefore decrease the weights at other support points. Naturally, these decreased weights will be at points for which the WAP-maximizing test has greater power than the $t$-test. The five panels of Figure (ref) illustrate this re-weighting principle in five iterative steps, starting at $\omega_1=0.1$, for which $\omega_1$ is successively increased until the power function of the WAP-maximizing test coincides with that of the $t$-test. The upward-pointing (downward-pointing) arrows indicate that the weight on the support point for the WAP-maximizing test needs to be adjusted downward (upward) to bring the power function of the WAP-maximizing test closer to lying (weakly) above that of the $t$-test. Since the power function of the WAP-maximizing test with weight $\omega_1=0.5$ in the final panel coincides with that of the two-sided $t$-test, we can say that the two-sided $t$-test is optimal in the class of all tests of $H_0$ vs.\ $H_1$ while no longer needing to constrain ourselves to the class of unbiased tests.
Having illustrated the intuition behind our approach in a simple problem, we now move to the general hypothesis testing framework of interest.
Suppose that we observe a random element $Y$ taking values in the metric space $\mathcal{Y}$ and that $Y$ has a probability density function $f_\theta(\cdot)$ relative to some sigma-finite measure $\nu$, where the parameter governing its distribution $\theta\in\Theta\subset\mathbb{R}^k$ is finite-dimensional. We are interested in testing
where $\Theta_0, \Theta_1\subset\Theta$, $\Theta_0 \cap \Theta_1 = \emptyset$ and $\Theta_0$ is not a singleton. This testing problem with a single observation $Y$ typically arises as the limiting problem under an asymptotic approximation to a finite-sample problem with many observations via a local asymptotic embedding corresponding to a limit experiment or by using the asymptotic equivalence approach of Muller:11. See EMW15 for a more detailed discussion.
A generic test of (ref) is a measurable function $\varphi:\mathcal{Y}\mapsto [0,1]$ for which $\varphi(y)$ is the probability of rejecting $H_0$ upon observing the realization $Y = y$. For the parameter value $\theta\in\Theta$, $\int\varphi f_\theta d\nu$ is thus equal to the rejection probability of the test when the true value of the parameter is $\theta$. The starting point of our analysis is to suppose that we have an ad hoc test with correct size, a measurable function $\varphi_{ah}:\mathcal{Y}\mapsto [0,1]$ that is known to satisfy the (uniform) size constraint $\sup_{\theta\in\Theta_0}\int\varphi_{ah}f_\theta d\nu\leq \alpha$. We would like to assess whether $\varphi_{ah}$ is “nearly optimal” among tests that control size. For this problem to be nontrivial, we focus on tests of the hypothesis (ref) for which no uniformly most powerful test is known to exist.
As illustrated above in the simple context of a two-sided $t$-test for the mean of a Gaussian random variable with known variance, one could attempt to assess the optimality of $\varphi_{ah}$ by comparing it to the point-wise power envelope obtained by computing the rejection probabilities of a collection of level-$\alpha$ Neyman-Pearson tests indexed by $\theta'\in\Theta_1$ of $H_{0,\Lambda_{\theta'}}:$ the density of $Y$ is $\int f_\theta d\Lambda_{\theta'}(\theta)$ vs.\ $H_{\theta'}:\theta=\theta'$ given by \[ \varphi_{\Lambda_{\theta'},\theta'}(y) =
, \] for some cv $\geq 0$ and $0 \leq \varkappa \leq 1$ satisfying $\int\varphi_{\Lambda_{\theta'},\theta'}(\int f_\theta d\Lambda_{\theta'}(\theta)) d\nu= \alpha$, where $\Lambda_{\theta'}$ is the least-favorable probability distribution over $\Theta_0$ corresponding to the point alternative $H_{\theta'}$.\footnote{See LR05 for details on least-favorable distributions and AMS08, EMW15 and AFBMOQST25 for details on computing approximations to them.} Indeed, this approach has been applied in the literature, oftentimes after imposing additional side constraints such as similarity or invariance (e.g., AMS06,AMS08; MM13,MM19; GKM19; AMY19; MO20). However, as in the simple example of Section (ref), comparison to this collection of Neyman-Pearson tests may produce an unattainable power envelope.
Instead, we propose to determine whether the WAP of $\varphi_{ah}$ is nearly equal to that of a single WAP-maximizing test corresponding to some set of weights over $\Theta_1$ and is thus “nearly optimal”. More formally, we seek to determine whether there exists a weight function $\Omega$, a probability distribution with support on $\Theta_1$, such that $\varphi_{ah}$ nearly maximizes the WAP criterion (Wal43)
within the set of level-$\alpha$ tests \[\Phi_\alpha\equiv\left\{\varphi:\mathcal{Y}\mapsto [0,1]:\varphi \text{ is measurable, }\sup_{\theta\in\Theta_0}\int \varphi f_\theta d\nu\leq \alpha\right\}.\] If a least-favorable probability distribution $\Lambda_\Omega$ over $\Theta_0$ corresponding to the simple alternative $H_{1,\Omega}:$ the density of $Y$ is $\int f_\theta d\Omega(\theta)$ exists, this WAP-maximizing test takes the familiar Neyman-Pearson form of a test of $H_{0,\Lambda_{\Omega}}:$ the density of $Y$ is $\int f_\theta d\Lambda_{\Omega}(\theta)$ vs.\ $H_{1,\Omega}$ given by
Here, $\text{cv}_{\Omega}\geq 0$ and $0 \leq \varkappa_\Omega \leq 1$ are defined to satisfy $\int\varphi_{\Lambda_{\Omega},\Omega}(\int f_\theta d\Lambda_{\Omega}(\theta)) d\nu= \alpha$. As noted by EMW15, even if this least-favorable distribution does not exist, we can find an “approximate least-favorable distribution” to approximate the test the maximizes (ref).
It is not typically possible to determine the existence of a weight function $\Omega$ for which an ad hoc test $\varphi_{ah}$ maximizes the WAP criterion WAP$(\varphi)$ analytically. We thus propose a numerical approach that computes the weights that make the WAP of a WAP-maximizing test as close as possible to the WAP of the ad hoc test for those weights over a discretization of the alternative space. Specifically, let $\bar\Theta_1=(\theta_1,\ldots,\theta_{M_1})\subset\Theta_1$ denote a finite set of support points for the alternative space with $M_1$ elements and $\bar\Omega=(\omega_1,\ldots,\omega_{M_1})\in \Delta_{M_1}$ denote a probability distribution over those support points, where $\Delta_{M_1}$ is the $M_1$-dimensional unit simplex. Formally, the numerical problem we wish to solve is given by
A test $\varphi^*\in\Phi_\alpha$ that solves (ref) is a WAP-maximizing test with corresponding WAP as close as possible to that of $\varphi_{ah}$ since
A natural byproduct of our approach is that a weight function $\bar\Omega$ solving (ref) provides valuable information about the power properties of the ad hoc test: higher weight over a set of points in $\bar\Theta_1$ indicates that the test prioritizes rejection in that set relative to other points in $\bar\Theta_1$.
In addition to determining whether $\varphi_{ah}$ nearly maximizes a WAP criterion for some set of weights via (ref), our numerical approach has the added benefit of being able to produce an approximate power envelope (APE) for $\varphi_{ah}$, whether or not it is WAP-maximizing. This power envelope is again obtained from a single test and is therefore necessarily attainable, in contrast to power envelopes derived from a collection of tests. To see this, first note that the rejection probabilities over $\Theta_1$ of the test that solves \[\sup_{\varphi\in\Phi_\alpha}\inf_{\theta\in\Theta_1}\int(\varphi-\varphi_{ah})f_\theta d\nu\] constitute a power envelope for $\varphi_{ah}\in\Phi_\alpha$ since \[\sup_{\varphi\in\Phi_\alpha}\inf_{\theta\in\Theta_1}\int(\varphi-\varphi_{ah})f_\theta d\nu\geq\inf_{\theta\in\Theta_1}\int(\varphi_{ah}-\varphi_{ah})f_\theta d\nu=0. \] For an appropriately chosen $\bar\Theta_1=(\theta_1,\ldots,\theta_{M_1})\subset\Theta_1$\textemdash see Section (ref) for details, we can then view the rejection probabilities of the test that solves
as an APE for $\varphi_{ah}$ since it constitutes a power envelope for $\varphi_{ah}$ over $\bar \Theta_1$ by definition. In Theorem (ref) below, we show that the value of the maximin problem in (ref) is equal to the value of the minimax problem in (ref) and any test that solves (ref) also solves (ref). Since we show how to approximate the solution to (ref) in Sections (ref) and (ref) below, this latter result provides a practical strategy for obtaining an APE for $\varphi_{ah}$: solve (ref) and numerically check whether the solution yields an APE for $\varphi_{ah}$. Absent further structure on the problem, this strategy is not guaranteed to produce an APE but it can still provide a useful guide for obtaining an APE in practice. In Theorem (ref) below, we impose additional structure on the problem that is sufficient for guaranteeing that this strategy indeed produces an APE.
Let \[ \Phi^\ast = \operatorname*{arg\,max}_{\varphi\in \Phi_\alpha} \min_{j=1,\ldots,M_1}\int(\varphi-\varphi_{ah})f_{\theta_j} d\nu \] and \[ \Delta_{M_1}^\ast = \operatorname*{arg\,min}_{\bar\Omega\in \Delta_{M_1}} \sup_{\varphi\in \Phi_\alpha} \sum_{j=1}^{M_1} \omega_j \int(\varphi - \varphi_{ah}) f_{\theta_j} d\nu. \] We can now state our first result.
In words, equation (ref) states that any $\varphi^* \in \Phi^*$, i.e., any test that solves the maximin problem in (ref), is WAP-maximizing with respect to any $\bar \Omega^* \in \Delta_{M_1}^*$. Therefore, if one follows the strategy mentioned above and finds that a test that solves (ref) yields an APE for $\varphi_{ah}$, this strategy comes with the added benefit that the APE is “most-favorable” for $\varphi_{ah}$ in the sense that it corresponds to the rejection probabilities of a WAP-maximizing test with WAP as close as possible to that of $\varphi_{ah}$ over $\bar\Theta_1$.
It can be shown that Theorem (ref) further implies that if the least-favorable distribution $\Lambda_{\bar \Omega^*}$ exists for all $\bar \Omega^* \in \Delta_{M_1}^\ast$, then any test $\varphi^* \in \Phi^*$\textemdash which serves as a power envelope for $\varphi_{ah}$ on $\bar\Theta_1$\textemdash is equal to a Neyman-Pearson test $\varphi_{\Lambda_{\bar \Omega^*},\bar \Omega^*}$ for some $\bar \Omega^* \in \Delta_{M_1}^\ast$, except possibly on the event:\footnote{This follows from Lemma (ref) under the assumption that $\Lambda^\ast(\Theta_0) > 0$ for all $\Lambda^\ast \in \mathcal{M}_0^\ast$.}
We impose the commonly-satisfied sufficient condition that this event has $\nu$-measure zero for all $\bar \Omega^*$ in Theorem (ref) below and formally show that any $\varphi^* \in \Phi^*$ is equal to a Neyman-Pearson test $\varphi_{\Lambda_{\bar \Omega^*},\bar \Omega^*}$ $\nu\text{-almost everywhere}$ (a.e.), and this Neyman-Pearson test is unique $\nu\text{-a.e.}$ This uniqueness result trivially implies that the test that solves (ref) also solves (ref) so that the test obtained by solving the minimax problem according to our algorithms in Section (ref) readily provides a most-favorable APE for $\varphi_{ah}$.
The condition that (ref) has $\nu$-measure zero typically holds when $Y$ is an absolutely continuous random vector since (ref) is typically a lower-dimensional submanifold of $\mathcal Y$. The existence of a least favorable distribution has been established under weak conditions for a Euclidean sample space $\mathcal Y$ when $f_\theta$ is continuous in $\theta$ and $\Theta_0$ is a closed Borel set in a finite-dimensional Euclidean space. See LR05 and references therein.
To numerically approximate the weight function that justifies an ad hoc test $\varphi_{ah}$ in terms of WAP or produce a power function that dominates it, we aim to solve (ref). Our numerical algorithm for solving (ref) is composed of an inner loop and an outer loop. The inner loop computes an approximation to a WAP-maximizing test $\phi_{\Lambda_\Omega,\Omega}$ for a given weight function $\Omega$ in the spirit of MM13, EMW15 and AFBMOQST25.\footnote{GH24 also briefly note that their numerical algorithm for approximating minimax regret treatment rules could also potentially be modified to numerically approximate a WAP-maximizing test for a given weight function.} The outer loop approximates a weight function $\bar\Omega$ that solves (ref) using the inner loop as input at each step since it involves searching over WAP-maximizing tests via the relation (ref). In Section (ref), we provide theoretical convergence results justifying the use of our algorithm for solving (ref).
Discretize the null parameter space $\Theta_0$ into a finite set of support points $\bar\Theta_0=\{\tilde\theta_1,\ldots,\tilde\theta_{M_0}\}\subset\Theta_0$ with $M_0$ elements. In light of (ref) (and Fubini's Theorem), we aim to solve the following discretized optimization problem:
where $\Phi = \{ \varphi: \mathcal{Y} \to [0,1]: \varphi \text{ is measurable}\}$ and $g(y)=\int f_\theta(y)d\Omega(\theta)$ for a given weight function $\Omega$. Its dual problem is given by
where $\widetilde\Lambda=(\tilde\lambda_1,\ldots,\tilde\lambda_{M_0})$. It is not hard to see that Slater's condition is satisfied in this setting and therefore strong duality holds between these two problems.\footnote{See Theorem 1 in Chapter 8.3 in luenberger1997optimization.} This implies that for any solution $\widetilde\Lambda^\ast$ of the dual problem, any solution $\tilde\varphi_{\widetilde\Lambda^\ast}$ of
is also a solution of the primal problem (ref). We can operationalize this observation via the following simple result.
Since $\tilde\varphi_{\widetilde\Lambda^\ast}$ is given in closed form as soon as we know $\widetilde\Lambda^\ast$, it is sufficient to solve the dual problem in order to find a tractable solution to the primal problem (ref). Also, note that given the above, we have
This observation motivates the following algorithm for approximating $\tilde\varphi_{\widetilde\Lambda^\ast}$.
The intuition for this algorithm is essentially the same as that of EMW15. However, in contrast to the algorithm in EMW15, our algorithm carries proven convergence guarantees\textemdash see Section (ref).\footnote{AFBMOQST25 show that a slight modification to the algorithm of EMW15 produces an algorithm with proven convergence guarantees.}
For $k$ large enough, a $\tilde\varphi_{\widetilde\Lambda^{(k)}}$ approximately solves the original optimization problem (ref) in a sense we make precise in Section (ref). By Lemma 1 of EMW15, if $\sup_{\theta\in\Theta_0}\int\tilde\varphi_{\widetilde\Lambda^{(k)}}f_\theta d\nu\leq \alpha$, then $\tilde\varphi_{\widetilde\Lambda^{(k)}}$ attains an upper bound on the WAP for the weight function $\Omega$ for any test $\varphi\in\Phi_\alpha$, no matter how large $k$ is. This size constraint needs to be checked numerically but approximately holds by design if $\bar\Theta_0\subset\Theta_0$ is rich enough. Finally, to map any test $\tilde\varphi_{\widetilde\Lambda}$ given in Lagrange multiplier form back to the Neyman-Pearson form (ref), simply set $\text{cv}_{\Omega}=\sum_{i=1}^{M_0}\tilde\lambda_i$, $\varkappa_\Omega=1$ and $\Lambda=(\lambda_1,\ldots,\lambda_{M_0})\in\Delta_{M_0}$ with $\lambda_i=\tilde\lambda_i/\text{cv}_{\Omega}$ for $i=1,\ldots,M_0$.
Let
and
The function $\phi$ may fail to be fully differentiable but is convex, therefore enabling us to work with the following projected subgradient descent algorithm to solve (ref).
The intuition underlying the algorithm is simple. Since our goal is to obtain a power envelope for $\varphi_{ah}$ from a WAP-maximizing test, we aim to find weights $\bar\Omega$ such that the power of $\varphi_{\bar\Omega}^*$ is (weakly) greater than that of $\varphi_{ah}$ at all support points in $\bar\Theta_1$. So at each iteration $k$, we aim to increase the weight given to support points for which $\varphi_{\bar\Omega^{(k)}}^*$ currently has lower power than $\varphi_{ah}$. Since the weights must add to one, we therefore must decrease the weight given to other support points. Naturally, we do so at points for which $\varphi_{\bar\Omega^{(k)}}^*$ currently has higher power than $\varphi_{ah}$. Step 2(a) computes the support points for which we aim to increase (decrease) the corresponding weights while Step 2(b) computes the weight adjustment.
For $k$ large enough, a $\bar\Omega^{(k)}$ approximately solves the original optimization problem (ref) in a sense we make precise in Section (ref). Step 2(b) of the algorithm relies on a Euclidean projection onto the unit simplex $\Delta_{M_1}$. This is a quadratic programming problem and there are efficient computational algorithms in the literature designed for solving this problem (e.g., wang2013simplexefficient).
In this section, we provide precise recommendations on how to implement our numerical procedure. We follow these recommendations ourselves when analyzing the applications in Section (ref).
We follow EMW15 and implement switching tests when the parameter space for the nuisance parameter is unbounded, albeit with an additional related objective. Tests that do not (approximately) reduce to a standard test in the “standard” part of the parameter space, typically characterized by large values of nuisance parameters, tend to significantly sacrifice power in the “standard” part of the parameter space, a conclusively undesirable property (see Section 4 of EMW15). Therefore, we do not wish to limit the scope of our analysis on test optimality to WAP-maximizing tests that are not able to “switch” to standard tests in the “standard” parts of the parameter space. Indeed, ad hoc tests are typically purposefully designed to reduce to standard tests that are known to be optimal in some sense in the “standard” part of the parameter space for this very reason. For example, in the linear instrumental variables (IV) model, the CLR test reduces to the two-sided $t$-test when the concentration parameter is large. Since our optimality-assessment procedure relies upon approximating WAP-maximizing tests on a finite set of support points $\bar\Theta_1\subset\Theta_1$, comparing to WAP-maximzing tests that do no allow for “switching” would inherently disadvantage an ad hoc test that reduces to a standard test in the “standard” part of the parameter space because the WAP-maximizing test would place zero weight over any region in $\Theta_1$ outside of $\bar\Theta_1$. Due to the typical noncompactness of the parameter spaces $\Theta_0$ and $\Theta_1$, there are also numerical benefits to focusing on switching tests, for which we also refer the interested reader to Section 4 of EMW15.
Since an ad hoc test can immediately be seen as suboptimal when it does not reduce to a known “standard best test” in the “standard” part of the parameter space, we focus here on ad hoc tests that do. Section 4.1 of EMW15 formalizes how the “standard” part of the parameter space can be characterized in terms of large values of a parameter $\delta$. Let $\delta_\text{S}$ be the point where the “standard” part of the parameter space begins in the sense that for $\delta > \delta_\text{S}$ the ad hoc test and the standard test have essentially the same rejection profile. Let $D(Y)$ and $\delta_\text{SP}$ denote a statistic and a “switching point” such that with probability very close to zero, $D(Y)>\delta_\text{SP}$ whenever $\delta\leq \delta_\text{S}$. The motivation for this choice is that we only want the test to which we compare $\varphi_{ah}$ to switch to a standard test $\varphi_S$ that has the best rejection profile in the standard part of the parameter space when we know that the ad hoc test has essentially the same rejection profile, enabling us to analyze its optimality properties outside of this region. In practice, all this amounts to changing $\tilde\varphi_{\widetilde\Lambda}(y)$ in the inner loop algorithm from the expression in Proposition (ref) to the switching form \[ \tilde\varphi_{\widetilde\Lambda}(y)=\mathbbm{1}(D(y)>\delta_{SP})\varphi_S(y)+\mathbbm{1}(D(y)\leq\delta_{SP})\mathbbm{1}\left(g(y) \geq \sum_{i=1}^{M_0} \tilde\lambda_i f_{\tilde\theta_i}(y)\right). \] See EMW15 for a formal definition of the standard best test $\varphi_S$ and further details on switching tests.
Acknowledging that we cannot perfectly compute $\bar\Omega^*$ (due to the numerical approximation of the outer loop) or $\varphi_{\Omega}^*$ for any $\Omega\in\Delta_{M_1}$ (due to the numerical approximation of the inner loop), let $\hat\varphi_{\widehat\Omega}^*$ denote the test obtained from the outer loop algorithm. In addition to potential numerical error arising from the use of our algorithms, we must rely on approximations of rejection probabilities within step 2.(a) of the inner loop algorithm and step 2.(a) of the outer loop algorithm. Nevertheless, the theoretical results in the following section and standard laws of large numbers imply that if we apply our algorithms with enough iterations and simulate rejection probabilities from enough Monte Carlo replications, these approximation errors will be small so that (i) $\int\hat\varphi_{\widehat\Omega}^*f_{\tilde\theta}d\nu$ cannot lie substantially above $\alpha$ for any $\tilde\theta\in\bar\Theta_0$ and (ii) $\int(\hat\varphi_{\widehat\Omega}^*-\varphi_{ah})f_{\theta}d\nu$ cannot lie substantially below zero for any $\theta\in\bar\Theta_1$. We therefore use the following numerical approximation threshold rules for some small threshold $\epsilon>0$:
Using the numerical approximation threshold rules above, we implement the algorithm by potentially adding support points according to the following steps:
The motivation for step 2. is to ensure $\hat\varphi_{\widehat\Omega}^*$ has approximately correct size, in analogy with step 8. of EMW15's (EMW15) algorithm. Similarly, step 3. is used to ensure that the power function of $\hat\varphi_{\widehat\Omega}^*$ provides a good approximation to that of a WAP-maximizing test that yields a power envelope for $\varphi_{ah}$.
To approximate rejection probabilities, we rely on Monte Carlo simulation but incorporate several modifications designed to improve the numerical stability and convergence of our algorithms. These adjustments impose qualitative features that the true rejection probabilities are known to satisfy, thereby reducing simulation noise that would otherwise distort the optimization.
First, in all examples we consider, the power function of the ad hoc test is continuous in $\theta$. Independent simulation draws across different $\theta$ values produce artificially jagged power functions. To avoid this, we employ common random numbers by generating a single set of baseline draws and obtaining simulated values of $Y$ for each parameter value by transforming these baseline draws (i.e., by translating them by $\theta$). This induces smoothness in the simulated rejection probabilities.
Second, we ensure that key moments of the distribution of $Y$ are matched exactly in the simulations. For example, when $Y$ is standard bivariate normal, we standardize the baseline draws to have mean zero and identity covariance. In some instances we also impose symmetry by symmetrizing the baseline draws. For the bivariate normal case, this guarantees that the marginal distributions are symmetric around zero. In the boundary-robust testing application of Section (ref), this ensures that the estimated power function of the two-sided $t$-test is symmetric, as implied by theory.
Third, we choose the random seed so that the estimated null rejection probabilities are close to the nominal level whenever the true null rejection probabilities are known to equal the nominal level. For example, in the linear IV model the CLR test is similar by construction. If the simulation draws happen to make the CLR test appear to overreject at some points in $\bar\Theta_0$, then the algorithm---which searches for a WAP-maximizing test that weakly dominates the CLR test while satisfying the size constraint---may struggle to converge. Ensuring that simulated null rejection rates are close to their theoretical values avoids these problems and yields more reliable numerical results.
In this section, we present the theoretical justification for both the inner and outer loop algorithms for computing approximate WAP-maximizing tests that can produce a power envelope for a given ad hoc test.
We now present the theoretical result guaranteeing the convergence of the inner loop algorithm for minimizing $\tilde\phi$ in (ref), which as discussed in Section (ref), also solves the primal problem (ref). This algorithm is a dual (projected) subgradient descent algorithm and therefore its convergence properties readily follow from known results in the literature. Let
The convergence rate is controlled by $\widetilde\Gamma_k$, which depends on the starting value $\widetilde\Lambda^{(0)}$ through $\lVert \widetilde\Lambda^{(0)} - \widetilde\Lambda^\ast\rVert_2^2$ and on the tuning parameter sequence $\{\tilde h_k\}$. It holds that $\widetilde\Gamma_k \to 0$ as $k\to \infty$ if $\sum \tilde h_k = \infty$ and $\sum \tilde h_k^2 < \infty$. The particular proposed choice of the tuning parameters $\{\tilde h_k\}$ is taken from Section 3.2.3 of nesterov2018lecture. This choice is convenient as it only depends on known quantities and guarantees that the algorithm finds an $\varepsilon$-solution if the number of iterations is sufficiently large.
EMW15 also propose an algorithm for approximating $\tilde\varphi_{\widetilde\Lambda^\ast}$ but they do not provide any convergence guarantees analogous to our Theorem (ref) for our inner loop algorithm. However, while we were writing this paper, AFBMOQST25 proposed an algorithm that slightly modifies that of EMW15 and produces a test that provably approximates $\tilde\varphi_{\widetilde\Lambda^\ast}$ with convergence guarantees analogous to our Theorem (ref) as well as bounds the Monte Carlo error. Their results also enable them to bound the size-distortions of their approximately optimal test and provide specific theoretically-grounded tuning parameter recommendations for an updating parameter, number of iterations and initial weights. Either the original algorithm of EMW15, the new modified algorithm of AFBMOQST25 or the linear-programming algorithm of MM13 can be used in place of our inner loop algorithm if desired.
Moving to the theoretical result guaranteeing the convergence of the outer loop algorithm for approximating the weight function that solves (ref), we again note that since this algorithm is a (projected) subgradient descent algorithm, its convergence properties readily follow from known results in the literature. Let
The convergence rate is controlled by $\Gamma_k$, which depends on the starting value $\Omega^{(0)}$ through $\lVert \Omega^{(0)} - \bar\Omega^\ast\rVert_2^2$ and on the tuning parameter sequence $\{h_k\}$. It holds $\Gamma_k \to 0$ as $k\to \infty$ if $\sum h_k = \infty$ and $\sum h_k^2 < \infty$, which as with the inner loop is a convenient choice taken from Section 3.2.3 of nesterov2018lecture.
In this section, we apply our results and algorithms to shed new light on the optimality properties of a test whose optimality properties have already been thoroughly analyzed in the literature as well as a brand-new test that has not yet. Specifically, we analyze the optimality of the CLR test of Moreira (2003) and the test implied by the inequality-imposed confidence interval of Cox24. Throughout this section, the nominal level $\alpha$ is taken equal to 5%.
As alluded to in the introduction, AMS06,AMS08 (AMS06 and AMS08, henceforth) investigate the optimality of the CLR test in the homoskedastic linear IV model, holding the variance matrix for the reduced-form errors fixed in their analysis (coined as the “fixed-$\Omega$ design” by VdSW23). In this setting, AMS06 construct an asymptotically efficient two-sided point-wise power envelope for invariant similar tests as the rejection probabilities arising from a collection of point-optimal invariant similar tests, finding the striking result that the CLR test numerically attains this power envelope. AMS08 further strengthen this result by showing that one essentially obtains the same power envelope without imposing similarity. However, AMY19 (AMY, henceforth) provide a counterpoint to these optimality results by finding that the CLR test falls short of a point-wise power envelope that is constructed by varying the value of the parameter of interest and keeping its hypothesized value fixed, rather than the more standard analysis that keeps the value of the parameter of interest fixed and varies its hypothesized value. In recent work, VdSW23 show that this latter analysis is essentially the same as the standard analysis that varies the parameter of interest while keeping its hypothesized value fixed but instead of holding the variance matrix for the reduced-form errors fixed, fixes the variance matrix of the structural and first-stage errors (coined as the “fixed-$\Sigma$ design” by VdSW23). VdSW23 further argue that the fixed-$\Omega$ design implicitly favors the power function of the CLR test and that the fixed-$\Sigma$ design is better suited for analyzing power in cases of low to moderate endogeneity as well as differing signs of the parameter of interest and the correlation between the structural errors in the IV model.
These recent results of AMY and VdSW23 motivate us to revisit the optimality analysis of the CLR test, looking at both the fixed-$\Omega$ and fixed-$\Sigma$ designs. Analyzing the CLR test under the fixed-$\Sigma$ design using our new numerical optimality assessment enables us to assess whether the previously examined point-wise power envelopes of AMY may simply be setting an unattainable power bound. Before proceeding, we formally introduce the linear IV model and the CLR test, and briefly describe the power envelope of AMS06. We closely follow the notation of AMS06.
The linear IV model is given by
where $y_1,y_2, u,v_2 \in \mathbbm{R}^n$, $Z \in \mathbbm{R}^{n\times k}$, $\beta \in \mathbbm{R}$, and $\pi \in \mathbbm{R}^k$. Here, $y_1$, $y_2$ and $Z$ are observed, where $y_1$ denotes the outcome of interest, $y_2$ the (potentially) endogenous regressor of interest, and $Z$ the instruments.\footnote{We omit additional exogenous regressors without loss of generality, as the above model can always be obtained by partialling them out.} The random vectors $u$ and $v_2$ are unobserved structural error terms. We assume that $Z$ is fixed. Plugging the reduced-form equation for $y_2$ into the structural equation for $y_1$, we obtain the reduced-form equation for $y_1$, i.e., \[ y_1 = Z \pi \beta + v_1, \] where $v_1 = u + v_2\beta$. The resulting set of reduced-form equations can be written as\footnote{Here, we use $Y$ to be consistent with the notation in AMS06. This $Y$ should, however, not be confused with the random element $Y$ that enters the general testing problem in Section (ref).} \[ Y = Z\pi a' + V, \] where \[ Y = [y_1, y_2], \ V = [v_1, v_2] \text{ and } a = (\beta,1)'. \] Let $V_i$ denote the $i^\text{th}$ row of $V$ and assume that $V_i$ is iid across $i$ with \[ V_i \sim \mathcal{N}(0,\Omega). \] Having defined the linear IV model, we can now formally state the testing problem of interest, which is given by \[ H_0: \ \beta = \beta_0, \ \pi \in \mathbbm{R}^k \text{ vs.\ } H_1: \ \beta \in \mathbbm{R}\backslash \{ \beta_0 \}, \ \pi \in \mathbbm{R}^k. \] We follow AMS06 and further simplify this testing problem. To that end, we define the following transformations of $Y$: \[ S = (Z'Z)^{-1/2}Z'Yb_0\cdot(b_0'\Omega b_0)^{-1/2} \text{ and } T = (Z'Z)^{-1/2}Z'Y \Omega^{-1}a_0 \cdot (a_0'\Omega a_0)^{-1/2}, \] where $b_0 = (1,-\beta_0)'$ and $a = (\beta_0,1)'$. Lemma 2 of AMS06 shows that $S$ and $T$ are jointly normally distributed and independent. AMS06 argue that the coordinate system used to specify $S$ and $T$ should not affect inference and, therefore, only consider statistics that are invariant to rotation of the coordinate system. This is achieved by considering test statistics that only depend on $S$ and $T$ through \[ Q = \left[
\right] = \left[
\right]; \] see Theorem 1 in AMS06 and the surrounding discussion. The distribution of $Q$ is noncentral Wishart and, importantly, depends on $\pi$ only through \[ \lambda = \pi'Z'Z\pi, \] which implies that if we restrict our attention to test statistics based on $Q$, the testing problem of interest simplifies to
The LR statistic for testing (ref) can be written as \[ \text{LR} = \frac{1}{2} \left( Q_S - Q_T + \sqrt{(Q_S - Q_T)^2 + 4 Q_{ST}^2} \right). \] Moreira (2003) observes that $Q_T$ is a sufficient statistic for $\lambda$ under $H_0$. This, in turn, allows the construction of a size $\alpha$ test using conditional critical values. In particular, the CLR test rejects when $\text{LR} > \text{cv}_{1-\alpha}(Q_T)$, where $\text{cv}_{1-\alpha}(Q_T)$ is such that $P_{\beta_0}(\text{LR} > \text{cv}_{1-\alpha}(Q_T)|Q_T) = \alpha$.
The power envelope proposed by AMS06 is constructed from a collection of point-optimal invariant similar two-sided (POIS2) tests. Invariance is imposed by restricting attention to tests that only depend on the data through $Q$. In the context at hand, imposing similarity is tantamount to relying on conditional critical values, conditional on the observed value of $Q_T$. This avoids the need for approximating the least-favorable distribution, as done in AMS08. The two-sidedness that AMS06 consider is such that the resulting test is asymptotically efficient, i.e., the test has the same power as the two-sided $t$-test when instruments are strong ($\lambda = \infty$). The corresponding point-optimal test, POIS2, puts equal weight on $(\beta^*,\lambda^*)$ and $(\beta^*_2,\lambda^*_2)$, where the latter is a function of the former, ensuring asymptotic efficiency. The corresponding power envelope is then mapped out by the POIS2 test as $(\beta^*,\lambda^*)$ varies.
AMS06 consider a fixed-$\Omega$ design in their simulations, setting $\Omega_{11} = \Omega_{22} = 1$ and $\beta_0 = 0$. They numerically compare the power of the CLR test to their point-wise power envelope as functions of $\beta$ and $\lambda$ for various values of $k$ and $\Omega_{12}$ ($\rho$ in their paper). For given values of $k$ and $\Omega_{12}$, they consider several $\lambda$-slices of the CLR power function and the power envelope, for which $\beta$ is varied over a fine grid for a given value of $\lambda$. An example of such a $\lambda$-slice is given in Figure (ref).
Panel (a) of Figure (ref) is a screenshot of Figure 1(c) in AMS06, which shows the power of the CLR test (and the Anderson-Rubin and Lagrange multiplier tests of AR49 and Kle02) together with their proposed power envelope as $\beta$ varies and $\lambda = 5$ for $k = 5$ and $\Omega_{12} = 0.5$. The power of the CLR test and the power envelope are virtually indistinguishable (or “very close” in this and many other figures in AMS06 where $k$, $\Omega_{12}$, and $\lambda$ are varied), which leads AMS06 to conclude that the CLR test attains the power envelope “in a numerical sense”.
We reconsider the above testing problem for $k = 5$ and $\Omega_{12} = 0.5$. In the case at hand, $Q$ takes the role of $Y$ and $\lambda$ replaces $\delta$.\footnote{In principle, we could use $(S,T)$ as $Y$, replace $\delta$ by $\pi$ and investigate the optimality of the CLR test in the class of asymptotically efficient tests. However, the computational cost of approximating of $\Lambda$ increases rapidly in the dimension of the nuisance parameter.} We set $\lambda_\text{S} = 75$ and $\lambda_\text{SP} = 160$, where $Q_T$ takes the role of $D(Y)$, since the probability that $Q_T>\lambda_{SP}$ is less than $0.01$ whenever $\lambda\leq\lambda_S$. The standard test is implemented as the Lagrange multiplier test, which rejects $H_0$ when $Q_{ST}^2/Q_T > \chi^2_{1-\alpha}(1)$.\footnote{For $\lambda > \lambda_\text{S} = 75$, the CLR test, the Lagrange multiplier test and the two-sided $t$-test all approximately coincide.} Note that switching to the standard test (for large $\lambda$) is equivalent to imposing asymptotic efficiency as defined by AMS06. We obtain our APE using $\bar \Theta_0 = \{ (\beta,\lambda) : \beta = 0,\ \lambda \in \{1,5,10,\dots,30,40,\dots,170\} \}$ and $\bar \Theta_1 = \{ (b/\sqrt{\lambda},\lambda) : b\in \{-4,-3,-2,2,3,4\},\ \lambda \in \{ 1,5,10,\dots,30,40,\dots,170 \} \}$. Implementation details, including the number of simulation draws and the choices of $\{h_k\}$, $\Theta_0^f$ and $\Theta_1^f$ are given in Appendix (ref).
Panel (b) of Figure (ref) shows the power function of the CLR test (for the same parameter constellation as in panel (a)) together with our APE. As in panel (a), the power of the CLR test is indistinguishable from the power envelope for the $\lambda$-slice under consideration. Instead of considering multiple $\lambda$-slices, we produce a heatmap showing the difference between our APE and the power of the CLR test over a grid of values for $\beta$ and $\lambda$.\footnote{The grid is given by $\{ (b/\sqrt{\lambda},\lambda) : b\in \{-3.5,-3,\dots,3.5\},\ \lambda \in \{ 0.1,10,20,\dots,170 \} \}$. Note that this grid does not coincide with $\bar \Theta_0 \cup \bar{\Theta}_1$ and that the underlying rejection probabilities are evaluated using draws that are independent from those used to obtain the APE. This protects from a potential winner's curse, cf. AKM24.} This heatmap is shown in Figure (ref).
The differences between the APE and the power of the CLR test are very close to zero over the entire grid; note that the scale on the right ranges only from -0.5 percentage points (pp) to 0.5pp. In fact, the largest difference (in absolute value) is below 0.1pp. We therefore conclude that the CLR test is effectively optimal (for $k = 5$ and $\Omega_{12} = 0.5$). Our finding is complementary to that of AMS06: we find that the CLR test is (effectively) optimal in the class of invariant asymptotically efficient tests, while AMS06 find the stronger result of point-optimality in the smaller class of tests that are similar and two-sided (in the sense that they impose).\footnote{Our finding is also complementary to that of AMS08: while we find a weaker result (optimality as opposed to point-optimality), we do not impose two-sidedness (in the sense that they impose).}
As shown by VdSW23, AMY consider a fixed-$\Sigma$ design in their simulations, where the variance matrix of the {structural} errors $u$ and $v_2$, $\Sigma$, is held constant. AMY set $\Sigma_{11} = \Sigma_{22} = 1$ and consider several values for $\Sigma_{12}$ ($\rho_{uv}$ in their paper). Their Figure 1 shows the power functions of the CLR test (and the Anderson-Rubin test) and the power envelope of AMS06 based on the POIS2 test for different values of $\Sigma_{12}$, keeping $k=10$ and $\lambda = 15$ fixed. In contrast to AMS06, however, AMY set $\beta = 0$ and vary $\beta_0$.
Panel (a) of Figure (ref) is a screenshot of Figure 1(b) in AMY, where $\Sigma_{12} = 0.5$. The power of the CLR test is on the power envelope for values of $\beta_0 \sqrt{\lambda}$ between 0 and 5, but drops below for values further away from 0. The maximal gap of the CLR test's power function with the POIS2 power envelope is around 3--4pp. Based on this finding, AMY conclude that the finding of AMS06 “that the CLR test is essentially on the [...] power envelope does not hold...”. However, as discussed extensively above, this does not mean that the CLR test is not optimal since it is unclear if the point-wise power envelope used by AMY is attainable. We therefore reconsider the testing problem considered in AMY, where $k = 10$ and $\Sigma_{12} = 0.5$, and construct our APE. We set $\lambda_\text{S} = 160$ and $\lambda_\text{SP} = 320$. The underlying $\bar \Theta_0$ and $\bar \Theta_1$ as well as additional implementation details are provided in Appendix (ref).
Panel (b) of Figure (ref) displays the power function of the CLR test (for the same parameter constellation as in panel (a)) together with our APE. We find that the CLR test attains our APE, at least for the $\lambda$-slice under consideration. As before, we produce a heatmap showing the difference between our APE and the power of the CLR test over a grid of values for $\beta$ and $\lambda$, while setting $\beta_0 = 0$. While this setting seemingly differs from the setting of AMY (where $\beta = 0$ while $\beta_0$ varies), it follows from Corollary 1 of VdSW23 and the subsequent discussion that the resulting power curves are simply mirror images of one another.\footnote{That is, the power of the CLR test for testing $H_0: \beta = \beta_0$ when the true value of $\beta$, say $\beta^*$, is equal to 0, is equal to the power of the CLR test for testing $H_0: \beta = 0$ when $\beta^* =- \beta_0$.} The heatmap is shown in Figure (ref).
The heatmap in Figure (ref) uses the same scale as the heatmap in Figure (ref) and leads us to the same conclusion: the difference between our APE and the power function of the CLR test is very close to zero across the entire grid. We conclude that the CLR test is, in fact, effectively optimal under both the fixed-$\Omega$ and fixed-$\Sigma$ designs.
A line of recent work has examined testing problems involving uniformly-valid inference when nuisance parameters may lie on or near the boundary of the parameter space.\footnote{See, for example, AG10b, Ket18, KM25 and Cavaliere2025.} Focusing on the special case of a scalar nuisance parameter to analyze the properties of a new confidence interval proposed by Cox24, we work with the Gaussian experiment \[ Y = \left(
\right) \sim N\left( \left(
\right),\left(
\right)\right), \] where $\beta$ is the scalar parameter of interest and $\delta \ge 0$ is a scalar nuisance parameter. As discussed in Section (ref), this formulation corresponds to the asymptotic behavior of a broad class of finite-sample models (see, e.g., EMW15 and KM25). Formally, the testing problem of interest is given by\footnote{This limiting Gaussian formulation also applies to problems for which nuisance parameters are restricted to be greater/less than or equal to any known value via a simple affine transformation.}
Cox24 proposes a new confidence interval for $\beta$, called the inequality-imposed confidence interval (IICI). The IICI has several desirable properties: (i) it is easy to compute, (ii) it does not require simulation or tuning parameters, (iii) it is adaptive, and (iv) it has weakly shorter length than the standard two-sided confidence interval that ignores the information contained in $Y_2$. The IICI is constructed as the union of the standard two-sided CI when $Y_2$ is large and positive, the two-sided confidence interval that imposes $\delta = 0$ when $Y_2$ is large and negative, and the intersection of the two when $Y_2$ is close to zero; see Panel A of Figure 1 in Cox24 for a visual illustration. Formally, the lower and upper bounds of the $(1-\alpha)$-nominal IICI are given by \[
\] and \[
\] respectively, where $c = \frac{1 - \sqrt{1-\rho^2}}{\rho} z_{1-\alpha/2}$ and $z_{1-\alpha/2}$ denotes the $1-\alpha/2$ quantile of the $\mathcal{N}(0,1)$ distribution. The test implied by the IICI rejects $H_0$ in (ref) when $\beta_0$ lies outside of it.
To gauge the optimality of the IICI, Cox24 compares the WAP of the test implied by the IICI with the power bound on WAP derived in EMW15 (EMW, henceforth), using equal weight on $\beta=-2$ and $\beta=2$ and uniform weights on $\delta \in [0,9]$. For $\rho = 0.7$, the WAP of the IICI is 53.1% and the WAP of the EMW power bound is 53.5%. Using $\epsilon = 0.005$ (following EMW), Cox24 concludes that the test implied by the IICI is nearly optimal in the sense of EMW, as its WAP lies within $\epsilon$ of the power bound. However, some ambiguity with respect to the optimality of the test remains:\footnote{This is related to the fact that $\epsilon$ is given as an absolute (rather than a relative) value, making it difficult to interpret what it means for the difference in WAPs to be “small”.} as shown in Figure (ref), the power functions of the test implied by the IICI and the nearly optimal test of EMW (under the above weights) cross, which prevents us from concluding whether the test implied by the IICI is optimal or dominated.
To address this ambiguity, we compute our APE for the testing problem with $\rho = 0.7$. We follow EMW and choose $\delta_{\text{SP}} = 6$. The standard test is the two-sided $t$-test. We use $\bar{\Theta}_1 = \{ (b,d) : b \in \{-3,-2,-1,1,2,3\},\ d \in \{0,0.5,\dots,8\} \}$ and follow EMW in discretizing $\Theta_0$ in terms of “base” distributions. Details on this and other implementation choices are provided in Appendix (ref). As in Section (ref), we produce a heatmap showing the difference between our APE and the power of the test implied by the IICI over a grid of values for $\beta$ and $\delta$, fixing $\beta_0 = 0$; see Figure (ref).
The scale of the heatmap is the same as for the heatmaps in Section (ref). The difference between the APE and the power of the test implied by the IICI is close to zero over a large portion of the grid. However, there are some values for which the test implied by the IICI falls short of the APE—the maximal difference is about 0.3pp and occurs at $(\beta,\delta) = (2,1)$. We, therefore, conclude that the test implied by the IICI is effectively dominated, albeit by a very small margin. In fact, the WAPs of the test underlying the APE and the test implied by the IICI are 52.532% and 52.529%, respectively, with the corresponding weights given in Appendix (ref). Abstracting from numerical approximations, the WAP of the test implied by the IICI is thus within 0.00003 of the lowest possible upper bound (given $\bar \Theta_1$), improving on the finding in Cox24 and providing a strong argument in favor of using the IICI in practice. Although the test is effectively dominated, the loss in WAP is negligible considering the simplicity of the procedure.