EconBase
← Back to paper

Power Bounds and Efficiency Loss for Asymptotically Optimal Tests in IV Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

64,149 characters · 11 sections · 46 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Power Bounds and Efficiency Loss for Asymptotically Optimal Tests in IV Regression

abstractWe characterize the maximal attainable power–size gap in overidentified instrumental variables models with heteroskedastic or autocorrelated (HAC) errors. Using total variation distance and Kraft’s theorem, we define the decision theoretic frontier of the testing problem. We show that Lagrange multiplier and conditional quasi likelihood ratio tests can have power arbitrarily close to size even when the null and alternative are well separated, because they do not fully exploit the reduced-form likelihood. In contrast, the conditional likelihood ratio (CLR) test uses the full reduced-form likelihood. We prove that the power--size gap of CLR converges to one if and only if the testing problem becomes trivial in total variation distance, so that CLR attains the decision theoretic frontier whenever any test can. An empirical illustration based on Yogo (2004) shows that these failures arise in empirically relevant configurations. \noindentKeywords: Endogenous regressor, Instrumental variable, Conditional likelihood ratio test, Lagrange Multiplier test, HAC errors, Total variation distance JEL classification: C14, C36

Introduction

This paper studies the power--size tradeoff in instrumental variable models under weak identification and heteroskedastic or autocorrelated (HAC) errors. We characterize the intrinsic difficulty of the testing problem using total variation distance and derive the maximal attainable power--size gap, which defines the decision theoretic frontier for the problem. We then study whether commonly used weak-identification-robust procedures attain this frontier. In particular, we show that the conditional likelihood ratio (CLR) test attains this frontier whenever any test can do so, whereas Lagrange multiplier (LM) and conditional quasi likelihood ratio (CQLR) tests may fail to attain it in HAC environments.

Inference in instrumental variable models is complicated by weak identification. The test of AndersonRubin49 was the first procedure shown to remain valid when instruments are arbitrarily weak because its null distribution does not depend on instrument strength. However, the Anderson Rubin (AR) test is not efficient under the usual strong instrument asymptotics.

This inefficiency motivated the development of procedures that retain robustness to weak identification while achieving asymptotic efficiency when instruments are strong. Standard Wald and likelihood ratio tests are efficient under strong instruments but can exhibit severe size distortions when instruments are weak; see NelsonStartz90, Dufour97, StaigerStock97, and WangZivot98. These distortions arise because their null distributions depend on first-stage coefficients when instruments are weak. A conditioning argument resolves this problem by replacing the usual chi--square critical value with a null quantile conditional on a statistic sufficient for the first-stage coefficients, as proposed by Moreira03. The resulting conditional Wald (CW) and conditional likelihood ratio (CLR) tests, together with the LM test, are asymptotically efficient under strong identification and valid under weak identification.

In homoskedastic models these procedures have well-understood properties. AndrewsMoreiraStock07 show that CW can exhibit severe power distortions under weak identification. In contrast, LM and CLR are unbiased and efficient under strong identification. In this setting, the efficiency gain relative to AR does not compromise asymptotic power properties, as AndrewsMoreiraStock04, AndrewsMoreiraStock06 establish that both LM and CLR are consistent. However, these results do not determine whether existing procedures fully exploit the information available in more general environments.

The situation changes in overidentified models with heteroskedastic or autocorrelated errors. In these models the reduced-form covariance structure contains information that is not fully exploited by tests constructed from AR and LM statistics. Tests designed to improve upon AR can therefore lose power for reasons that do not arise under homoskedasticity. We show that this loss of information can be structural rather than local.

In instrumental variable models with heteroskedastic or autocorrelated errors the reduced-form covariance matrix introduces many nuisance parameters. When the number of instruments increases, the number of covariance parameters grows roughly with the square of the number of instruments, making the testing problem inherently high dimensional. Classical power comparisons, such as those of AndrewsMoreiraStock06, depend on a small number of parameters even when the number of instruments increases. In contrast, the HAC IV framework allows a much richer set of covariance configurations. Showing that power failures arise in such a rich environment is therefore not straightforward.

Our first contribution is to characterize the maximal attainable power--size gap using the total variation distance between the null and alternative distributions. BertanhaMoreira20 show that total variation distance is the relevant metric for determining whether any test can deliver nontrivial power. By Kraft55, the minimum total variation distance between the convex hulls of the null and alternative models determines the largest power--size gap achievable by any measurable test. We derive an explicit bound for this frontier in the linear IV model and show that it provides a necessary and sufficient condition for asymptotic distinguishability.

We then show that the power--size gap of the conditional likelihood ratio (CLR) test converges to one if and only if the decision theoretic frontier converges to one. Thus, CLR attains the maximal attainable power--size gap whenever any test can do so. This result reflects the fact that CLR exploits the full reduced-form information available in the HAC model.

It is useful to contrast this property with existing procedures. AR also attains the frontier, although it is not efficient under strong identification. The LM test and the CQLR test were developed to improve upon AR while preserving robustness to weak identification. The CQLR test can be viewed as an HAC extension of CLR following Kleibergen05. However, the two procedures are not equivalent in HAC environments. CQLR is constructed from statistics based on AR and LM components, whereas CLR is defined directly from the Gaussian reduced-form likelihood; see AndrewsMikusheva16 and MoreiraMoreira19. Consequently, CQLR may discard information contained in the reduced-form covariance structure that CLR fully exploits.

We identify a class of data generating processes, which we call impossibility designs, under which the LM noncentrality parameter remains bounded even when the null and alternative distributions are arbitrarily well separated. In such designs LM and CQLR can be inconsistent, with power arbitrarily close to size. This phenomenon has no analog in the homoskedastic case.

These designs are not pathological. The set of data generating processes that generate bounded noncentrality parameters has positive measure, and the associated power losses extend to neighborhoods of these designs. Using data of Yogo04, we show that empirically plausible parameter configurations can lie near this region. In such cases AR--LM based conditional procedures can suffer substantial power losses, whereas CLR maintains high power.

Taken together, our results establish a sharp separation between CLR and both LM and CQLR tests in instrumental variable models with HAC errors. CLR exploits the full reduced-form likelihood and achieves the decision theoretic frontier, whereas LM and CQLR operate in a restricted information space and may fail to convert statistical separation into power.

Dufour97, building on GleserHwang87, shows that conventional Wald tests can have size arbitrarily close to one when instruments are arbitrarily weak. This finding supported the use of the AR test and the development of other procedures robust to weak instruments. Our results reveal a parallel issue for power in HAC IV models. When the covariance structure is near an impossibility design, LM and CQLR tests can have power arbitrarily close to size even though the null and alternative distributions are well separated. Just as BoundJaegerBaker95, Dufour97, and StaigerStock97 shifted empirical practice away from Wald tests toward procedures robust to weak instruments, our results suggest that procedures based on AR and LM statistics may discard relevant information in HAC environments. This motivates the use of procedures such as the CLR test that exploit the full reduced-form likelihood and also encourages the development of new tests that fully use the information available in the reduced-form.

The remainder of the paper is organized as follows. Section (ref) introduces the model and test statistics. Section (ref) studies the CLR test using total variation distance. Section (ref) characterizes the noncentrality parameter of the LM statistic and defines impossibility designs for both LM and CQLR tests. Section (ref) provides diagnostics and an empirical application. Section (ref) concludes the paper. An online supplement contains the proofs of all theoretical results, additional lemmas and a proposition, further power comparisons, and supplementary results for the empirical application.

Model and test statistics

We consider the instrumental variable regression model

align[align omitted — 74 chars of source]

where $y_{1}$ and $y_{2}$ are $n \times1$ vectors, $Z$ is an $n \times k$ matrix of nonrandom instruments with full column rank, and $\pi$ is a $k \times1$ vector of first stage coefficients. For simplicity we omit additional covariates. They can be partialled out using orthogonal projections according to the Frisch–Waugh–Lovell (FWL) theorem; see AndrewsMoreiraStock06 for the IV model. The disturbances $u$ and $v_{2}$ have mean zero and may exhibit heteroskedasticity and autocorrelation.

Our objective is to test \[ H_{0}: \beta^{\ast} = \beta_{0} \quad\text{against}\quad H_{1}: \beta^{\ast} \neq\beta_{0}, \] treating $\pi$ as a nuisance parameter.

Let $Y=[y_{1}\ y_{2}]$. The reduced form can be written as

equation[equation omitted — 44 chars of source]

where $a^{\ast}=(\beta^{\ast},1)^{\prime}$ and $V=[v_{1}\ v_{2}]$ with $v_{1}=u+\beta^{\ast}v_{2}$.

Define

equation[equation omitted — 89 chars of source]

where $\mu=(Z^{\prime}Z)^{1/2}\pi$ and $\widetilde{V}=(Z^{\prime} Z)^{-1/2}Z^{\prime}V$.

We impose the following normalized average conditions.

\noindentAssumption NA.

enumerate$n^{-1}Z^{\prime}Z \to D$, where $D$ is positive definite. • $n^{-1}V^{\prime}V \overset{p}{\to} \Omega$, where $\Omega$ is positive definite. • $(Z^{\prime}Z)^{-1/2}Z^{\prime}V \overset{d}{\to} N(0,\Sigma)$, where $\Sigma$ is positive definite.

Under Assumption NA, the finite dimensional distribution of $R$ converges to a Gaussian limit experiment with covariance matrix $\Sigma$. This limit experiment coincides with the weak instrument asymptotics of StaigerStock97 and provides a sharp approximation to finite-sample behavior when first-stage coefficients are local to zero. It also forms the basis for our fixed-alternative comparisons in later sections.

Accordingly, we conduct our analysis under the Gaussian reduced-form model

equation[equation omitted — 69 chars of source]

treating $\Sigma$ as known. This formulation isolates the information content of the reduced form and allows us to analyze the geometry of the testing problem directly.

Fix $\beta_{0}$ and define \[ a_{0}=(\beta_{0},1)^{\prime}, \qquad b_{0}=(1,-\beta_{0})^{\prime}. \] Define

align[align omitted — 298 chars of source]

Under $H_{0}$, $T$ is sufficient for $\mu$, while $S$ is pivotal and independent of $T$.

Let \[ B_{0}=

pmatrix[pmatrix omitted — 37 chars of source]

, \qquad R_{0}=RB_{0}. \] Then \[ \mathrm{vec}(R_{0}) \sim N\!\left( \mathrm{vec}(\mu a_{\Delta}^{\prime}), \Sigma_{0} \right) , \] where $a_{\Delta}=(\Delta,1)^{\prime}$ with $\Delta=\beta^{\ast}-\beta_{0}$ and \[ \Sigma_{0} = (B_{0}^{\prime}\otimes I_{k}) \Sigma(B_{0}\otimes I_{k}) =

pmatrix[pmatrix omitted — 68 chars of source]

. \]

The Anderson Rubin statistic is

equation[equation omitted — 41 chars of source]

The one-sided and two-sided LM statistics are

equation[equation omitted — 328 chars of source]

where

equation[equation omitted — 108 chars of source]

Another commonly used statistic is the quasi likelihood ratio (QLR):

equation[equation omitted — 87 chars of source]

where natural choices for the rank statistic $r(T)$ are

align*[align* omitted — 143 chars of source]

as outlined by AndrewsGuggenberger17. We refer to these variants as $QLR1$ and $QLR2$.

The null distribution of $QLR$ depends on $\mu$. Since $T$ is sufficient for $\mu$ under $H_{0}$, we follow the conditioning argument of Moreira03 and define conditional versions that reject the null hypothesis when the statistic exceeds its null quantile conditional on $T$. The conditional critical value function based on a statistic $\psi$ is

equation[equation omitted — 156 chars of source]

We refer to this as the conditional QLR (CQLR) test. Andrews16 shows that CQLR is a special case of conditional linear combination (CLC) tests of the form

equation[equation omitted — 64 chars of source]

where $0\le w(T)\le1$ and proposes the plug-in conditional linear combination test (PI-CLC).

Finally, following Moreira03 for the homoskedastic case, AndrewsMikusheva16 and MoreiraMoreira19 define the likelihood ratio statistic for the HAC case:

equation[equation omitted — 88 chars of source]

where

equation[equation omitted — 183 chars of source]

with $b=(1,-\beta)^{\prime}$.

The conditional likelihood ratio (CLR) test rejects when $LR$ exceeds its null quantile conditional on $T$.

CQLR extends the algebraic structure of CLR from the homoskedastic setting to the GMM framework with heteroskedastic or autocorrelated errors, whereas CLR is defined directly from the Gaussian reduced-form likelihood. The two procedures coincide under homoskedasticity but are generally distinct in HAC environments.

When $k>2$ in HAC settings, the likelihood ratio statistic does not admit a closed-form solution to the minimization problem in (ref). MoreiraNeweySharifvaghefi24 establish this result and develop computational algebra methods that allow the CLR statistic to be computed reliably in such environments.

CLR test

In this section, we show that the power--size gap of the CLR test, which rejects the null hypothesis when the $LR$ statistic defined in (ref) exceeds its conditional critical value given by (ref), converges to one as the minimum total variation distance between the convex hulls of the distributions under the null and alternative hypotheses for $R_{0}$ converges to one. By Kraft55's theorem, discussed below, this result implies that the power--size gap of any test can converge to one only if the same holds for the CLR test.

To place Kraft55's theorem in the context of linear IV models, let $\phi_{\Delta,\mu}(\cdot)$ denote the normal probability density function of $\mathrm{vec}(R_{0})$ with mean $\mathrm{vec}(\mu a_{\Delta}^{\prime})$ and variance matrix $\Sigma_{0}$. Denote the convex hull of the set of normal densities under the null hypothesis by

equation[equation omitted — 181 chars of source]

Similarly, let

equation[equation omitted — 250 chars of source]

denote the convex hull of the set of normal densities under the alternative hypothesis satisfying

equation[equation omitted — 63 chars of source]

where $d>0$ is an arbitrary constant.

The separation condition in (ref) reflects the intrinsic statistical distance between the null and alternative distributions in the Gaussian reduced-form model. The scalar \[ \Delta^{2}\mu^{\prime}\Sigma_{11}^{-1}\mu \] is central for three reasons. First, it equals the noncentrality parameter of the Anderson--Rubin statistic under the alternative, measuring the structural deviation $\Delta$ scaled by instrument strength as summarized by $\mu^{\prime}\Sigma_{11}^{-1}\mu$. Thus alternatives satisfying $\Delta^{2}\mu^{\prime}\Sigma_{11}^{-1}\mu \ge d$ are separated from the null by at least a fixed AR signal. Second, as shown below, a useful bound for the total variation distance over nuisance mean vectors under the null depends only on this scalar, so it directly indexes intrinsic distinguishability between the null and alternative models. Third, it has a natural econometric interpretation: $\Delta^{2} = (\beta^{\ast}-\beta_{0})^{2}$ is the squared structural distance from the null, while $\mu^{\prime}\Sigma_{11}^{-1}\mu$ summarizes instrument strength. Their product therefore measures structural separation scaled by instrument relevance. For these reasons, defining the alternative convex hull through $\Delta^{2}\mu^{\prime}\Sigma_{11}^{-1}\mu \ge d$ provides a model-based and decision theoretic notion of separation that allows a sharp characterization of the maximal attainable power--size gap.

The total variation (TV) distance between $f(\cdot)\in\mathcal{C}_{0}$ and $g(\cdot)\in\mathcal{C}_{1}$ is defined as

equation[equation omitted — 136 chars of source]

Kraft55 shows that, in order for there to exist a test $\psi(\cdot):\mathbb{R}^{2k}\to[0,1]$ such that

equation[equation omitted — 157 chars of source]

it is necessary and sufficient that

equation[equation omitted — 99 chars of source]

That is, the power--size gap of a test can be at least $q$ if and only if the minimum total variation distance between $\mathcal{C}_{0}$ and $\mathcal{C}_{1}$ equals $q$\footnote{The proof of this result is nonconstructive: it relies on an infinite-dimensional separating hyperplane argument that establishes the existence of a test achieving the bound but does not provide an explicit form for such a test.}. Consequently, a power--size gap equal to $1$ is possible if and only if the minimum total variation distance equals $1$.

Total variation distance plays a central role in our analysis because it is the decision theoretic metric that determines what any hypothesis test can accomplish. Kraft55's theorem is sharp: the largest power--size gap attainable by any measurable test, including randomized tests, coincides with the minimal TV distance between the convex hulls of the null and alternative models. Other notions of distance, such as the L\'evy--Prokhorov distance, can be useful for specific arguments but, as BertanhaMoreira20 show, they characterize only more restricted classes of tests and therefore do not capture the full difficulty of the testing problem. For this reason, TV distance provides the natural benchmark for the intrinsic difficulty of distinguishing the null from the alternative. Using this metric, we derive convenient bounds showing that the power of the CLR test approaches the decision theoretic frontier when the null and alternative distributions become well separated in TV, a property that AR-LM-based procedures cannot generally replicate in HAC environments.

The minimum total variation distance between $\mathcal{C}_{0}$ and $\mathcal{C}_{1}$ is smaller than the minimum total variation distance between the extreme points of these sets. Therefore,

equation[equation omitted — 282 chars of source]

It can be shown that

equation[equation omitted — 257 chars of source]

where $\Phi(\cdot)$ and $F_{\chi^{2}_{1}}(\cdot)$ denote the cumulative distribution functions of a standard normal random variable and a chi--square random variable with one degree of freedom, respectively. Moreover,

equation[equation omitted — 211 chars of source]

where $\Sigma^{22}$ is given by (ref), \[ \Sigma^{11} = \left( \Sigma_{11}-\Sigma_{12}\Sigma_{22}^{-1}\Sigma _{21}\right) ^{-1}, \qquad \text{and} \qquad\Sigma^{21} = -\Sigma ^{22}\Sigma_{21}\Sigma_{11}^{-1}. \] Therefore,

equation[equation omitted — 290 chars of source]

Since $F_{\chi^{2}_{1}}(\cdot)$ is increasing, we can write

equation[equation omitted — 162 chars of source]

where

equation[equation omitted — 182 chars of source]

The solution to this constrained optimization problem is $\delta_{\min}^{2} =d$. Consequently,

equation[equation omitted — 146 chars of source]

Lemma (ref) formalizes this finding.

lemmaLet $\mathcal{C}_{0}$, as defined in (ref), be the convex hull of distributions under the null hypothesis, and let $\mathcal{C}_{1}$, as defined in (ref), be the convex hull of distributions under the alternative hypothesis. Consider the total variation distance between $f(\cdot)\in\mathcal{C}_{0}$ and $g(\cdot)\in\mathcal{C}_{1}$ as defined in (ref). Under Assumption NA, \[ \min_{\substack{f\in\mathcal{C}_{0}\\g\in\mathcal{C}_{1}}} D_{\mathrm{TV} }(f,g) \leq F_{\chi^{2}_{1}}\left( \frac{d}{4}\right) . \]

$F_{\chi^{2}_{1}}\left(d/4\right) $ is always less than one and converges to one if and only if $d \to \infty$. Therefore, as the total variation distance between $\mathcal{C}_{0}$ and $\mathcal{C}_{1}$ converges to one, $d\to\infty$. In the next step, we show that as $d\to\infty$, the power--size gap of the CLR test converges to one. Consequently, Kraft55's theorem implies that the power--size gap of any test can converge to one only if the power--size gap of the CLR test does.

Finding the exact power of the CLR test is complicated because the $LR$ statistic depends on an optimization problem with no closed-form solution. Moreover, the critical value function of the CLR test is random. To address these issues, we consider a lower bound for CLR power.

Since $\inf_{\beta\in\mathbb{R}}Q(\beta)\ge0$, the conditional critical value function defined in (ref) satisfies

equation[equation omitted — 146 chars of source]

Given that $Q(\beta_{0})=S^{\prime}S$ is independent of $T$, we can further write

equation[equation omitted — 131 chars of source]

Under the null hypothesis $\Delta=0$, $Q(\beta_{0})$ follows a chi--square distribution with $k$ degrees of freedom. Therefore,

equation[equation omitted — 49 chars of source]

where $q_{1-\alpha}(k)$ denotes the $(1-\alpha)$ quantile of a chi--square distribution with $k$ degrees of freedom. Lemma (ref) states this upper bound.

lemmaUnder Assumption NA and for a given nominal size $\alpha $, the conditional critical value function $c_{\alpha}(T)$ of the CLR test, defined in (ref), is bounded above by $q_{1-\alpha}(k)$.

It follows by Lemma (ref) that

equation[equation omitted — 222 chars of source]

Moreover, since $\inf_{\beta\in\mathbb{R}}Q(\beta)\le Q(\beta^{\ast})$,

equation[equation omitted — 171 chars of source]

Thus, the CLR power is bounded from below by \[ \Pr\!\left[ Q(\beta_{0})-Q(\beta^{\ast})>q_{1-\alpha}(k) \right] . \]

We have

equation[equation omitted — 204 chars of source]

Moreover, $Q(\beta_{0})=S^{\prime}S$ follows a noncentral chi--square distribution with $k$ degrees of freedom and noncentrality parameter $d=\Delta^{2}\mu^{\prime}\Sigma_{11}^{-1}\mu$. Therefore, as shown in Proposition (ref) below, the lower bound on CLR power converges to one at an exponential rate in $d-q_{1-\alpha}(k)$.

propositionConsider the CLR test that rejects the null hypothesis when the $LR$ statistic given by (ref) exceeds its conditional critical value function given by (ref). Suppose Assumption NA holds. Then, \begin{align} \Pr\!\left[ LR>c_{\alpha}(T)\right] & \geq1-2k\exp\!\left( -\frac{1} {4k}\max\{0,d-q_{1-\alpha}(k)\}\right) \\ & \quad-2k\exp\!\left( -\frac{1}{8k}\max\{0,d-q_{1-\alpha}(k)\}\right) -2\exp\!\left( -\frac{(\max\{0,d-q_{1-\alpha}(k)\})^{2}}{128\,d} \right) ,\nonumber \end{align} where $d=\Delta^{2}\mu^{\prime}\Sigma_{11}^{-1}\mu$.

To ensure the power--size gap of the CLR test converges to one as $d\to\infty $, it suffices that $d-q_{1-\alpha}(k)\to\infty$ (for example, by choosing critical values so that $q_{1-\alpha}(k)$ grows slower than $d$). Under this choice, as $d\to\infty$, both the CLR power--size gap and the minimum total variation distance between $\mathcal{C}_{0}$ and $\mathcal{C}_{1}$ converge to one at an exponential rate in $d$. By Kraft55, the power--size gap of any test can converge to one only if the power--size gap of the CLR test does. Theorem (ref) formalizes this conclusion.

theoremConsider the CLR test with test statistic $LR$ given by (ref) and conditional critical value function $c_{\alpha}(T)$ chosen such that (ref) holds. Suppose Assumption NA holds. Then the power--size gap of the CLR test converges to one if and only if the minimum total variation distance between the convex hulls of probability densities under the null and alternative hypotheses converges to one. By Kraft55, this implies that the power--size gap of any test can converge to one only if the power--size gap of the CLR test converges to one.

LM-based tests and failure under impossibility designs

The LM test is asymptotically efficient under the conventional strong-instrument local asymptotic framework. However, this approximation can be misleading when instruments are weak or when we consider fixed alternatives. In such cases, the LM statistic can exhibit low power even in situations where distinguishing the null from the alternative should be straightforward in a decision theoretic sense. This section characterizes the behavior of the LM statistic under fixed alternatives and shows how the resulting limitations propagate to CLC and CQLR procedures.

We work with the one-sided LM statistic $LM_{1}$ in (ref). Under Assumption NA, the joint normality of $(S,T)$ implies the representation

equation[equation omitted — 133 chars of source]

where $U_{S}$ and $U_{T}$ are independent $N(0,I_{k})$ random vectors. Hence

equation[equation omitted — 344 chars of source]

Benchmark: strong instruments and local alternatives

\noindentAssumption SIV LA. (a) $\Delta_{n}=h_{\Delta}/n^{1/2}$ for some constant $h_{\Delta}$. (b) $\pi$ is a fixed nonzero $k$ vector.

propositionUnder Assumptions SIV LA and NA, \[ LM_{1}\rightarrow_{d} N\!\left( h_{\Delta} \left( \pi^{\prime}D^{1/2} \Sigma_{11}^{-1}D^{1/2}\pi\right) ^{1/2}, 1 \right) . \]

This is the standard efficiency result. Under strong identification and local alternatives, the LM noncentrality increases with both the local distance from the null and instrument strength. In this regime, the LM test is asymptotically optimal.

Fixed alternatives

We now keep $\Delta$ fixed. In this regime, terms that vanish under local asymptotics remain first order. The resulting approximation reveals that the LM drift need not increase with $|\Delta|$.

\noindentAssumption SIV FA. (a) $\Delta$ is fixed. (b) $\pi$ is a fixed nonzero $k$ vector.

theoremDefine \begin{equation} c(\Delta ,\mu )=\frac{\Delta \mu ^{\prime }\Sigma _{11}^{-1}\mu -\Delta ^{2}\mu ^{\prime }\Sigma _{11}^{-1}\Sigma _{21}\Sigma _{11}^{-1}\mu }{\left( \mu ^{\prime }(I_{k}-\Delta \Sigma _{11}^{-1}\Sigma _{12})\Sigma _{11}^{-1}(I_{k}-\Delta \Sigma _{21}\Sigma _{11}^{-1})\mu \right) ^{1/2}}. \end{equation} Under Assumptions SIV FA and NA, \begin{equation*} LM_{1}-c(\Delta ,\mu ) \rightarrow_{d} N \left( 0,\, 1+\|\gamma\|^{-2}\Delta^{2} \pi' D^{1/2}\Sigma_{11}^{-1/2} M_{\gamma}\Sigma_{11}^{-1/2} (\Sigma^{22})^{-1} \Sigma_{11}^{-1/2} M_{\gamma}\Sigma_{11}^{-1/2} D^{1/2}\pi \right), \end{equation*} where \begin{equation*} \gamma =\Sigma _{11}^{-1/2}(I_{k}-\Delta \Sigma _{21}\Sigma _{11}^{-1})D^{1/2}\pi, \qquad and \qquad M_{\gamma } = I_{k} - \frac{\gamma \gamma ^{\prime }}{\Vert \gamma \Vert^{2}}. \end{equation*}

The behavior of the statistic is driven by two forces: a stochastic component with controlled variance and a deterministic drifting term $c(\Delta, \mu )$.

The variance does not explode when $\Delta \to \infty$. Although $\Delta$ is fixed in the present asymptotic regime, this observation highlights that the stochastic component of the statistic remains well behaved even for large values of $\Delta$.

Furthermore, the variance does not depend on $\pi$ through its norm, but only through its direction. In principle, one could characterize the smallest and largest values of this variance as $\pi$ varies, since the expression is a ratio of quadratic forms in $\pi$, but we choose not to pursue this here. The main message is that this variance is well controlled, while an important component of the statistic is the drifting term $c(\Delta,\mu)$, which need not diverge as $\Delta$ grows.

We can further decompose this asymptotic distribution into two components that correspond to the two sources of randomness in the LM statistic. The first component, \[ N \!\left( 0,\, 1+\|\gamma\|^{-2}\Delta^{2}\pi' D^{1/2}\Sigma_{11}^{-1}(\Sigma^{22})^{-1}\Sigma_{11}^{-1} D^{1/2}\pi \right), \] originates from the random variation in the numerator of the $LM_{1}$ statistic. The second component, \[ N \left( 0,\, \|\gamma\|^{-2}\Delta^{2}\pi' D^{1/2}\Sigma_{11}^{-1/2} \Big( M_{\gamma}\Sigma_{11}^{-1/2}(\Sigma^{22})^{-1}\Sigma_{11}^{-1/2}M_{\gamma} - \Sigma_{11}^{-1/2}(\Sigma^{22})^{-1}\Sigma_{11}^{-1/2} \Big) \Sigma_{11}^{-1/2}D^{1/2}\pi \right), \] arises from the randomness in the denominator of the $LM_{1}$ statistic.

Under both the SIV FA and SIV LA approximations, we obtain a normal approximation to $\mathrm{LM}_{1}$, rather than the more complicated mixed normal distribution that arises under weak IV approximations. Unlike the SIV LA approximation, the SIV FA approximation retains all terms required for an accurate analysis of the deterministic drifting component $c(\Delta,\mu)$ of $\mathrm{LM}_{1}$.

The drifting term $c(\Delta, \mu )$ depends on \[ \zeta= \Delta \mu^{\prime}\Sigma_{11}^{-1} \mu- \Delta^{2} \mu^{\prime}\Sigma_{11}^{-1} \Sigma_{21}\Sigma_{11}^{-1}\mu, \] which need not have the same sign as $\Delta$. Thus, even for large $|\Delta|$, the one-sided LM test need not become powerful. This phenomenon can arise even under homoskedasticity; see AndrewsMoreiraStock06. Although the two-sided LM statistic remains consistent under homoskedastic errors, we show below that a deeper issue arises in HAC models: the LM noncentrality parameter need not diverge even when the null and alternative distributions become well separated. Consequently, the LM statistic may fail to translate statistical separation into power.

The decomposition above highlights two distinct forces that determine the behavior of the LM statistic. The stochastic component has a well-behaved variance that remains bounded even when $\Delta$ becomes large. In contrast, the behavior of the statistic may be dominated by the deterministic drifting term $c(\Delta,\mu)$. Whether this drift becomes large depends on the structure of the quadratic form that defines $c(\Delta,\mu)$ and cannot be determined from the approximation alone. In particular, the drift need not increase with $|\Delta|$. We next study the behavior of this drifting component in detail. We show that, under HAC covariance structures, the drifting term may fail to diverge even when the null and alternative distributions become well separated.

Impossibility designs

The drift term in Theorem (ref) is a ratio of a quadratic polynomial in $\Delta$ divided by the square root of another quadratic polynomial in $\Delta$. Under generic covariance structures, the leading quadratic term in the numerator dominates and the drift grows linearly in $|\Delta|$. In that case, the LM statistic separates the null and alternative as $|\Delta|\to\infty$.

However, this need not occur. If the coefficient on the leading quadratic term vanishes, then the numerator grows only linearly in $\Delta$ while the denominator grows proportionally to $|\Delta|$. Consequently, the drift remains bounded as $|\Delta|\to\infty$.

This motivates the following definition.

\noindentAssumption ID (Impossibility Design). \[ \mu^{\prime}\Sigma_{11}^{-1}\Sigma_{21}\Sigma_{11}^{-1}\mu= 0. \]

Under Assumption ID, the quadratic term in $\Delta^{2}$ in the numerator of the LM drift vanishes. As a result, the noncentrality parameter of the LM statistic does not diverge even when the null and alternative are arbitrarily well separated in total variation distance.

We now characterize when such designs can occur. A data generating process is an impossibility design if there exists $\mu\neq0$ satisfying Assumption ID. Since $\mu$ represents the standardized first-stage coefficients, this is a purely algebraic restriction on the reduced-form covariance structure. Define \[ A = \Sigma_{11}^{-1}\Sigma_{21}\Sigma_{11}^{-1}. \] Because $A$ need not be symmetric, we consider its Hermitian part.

propositionLet $A$ be a $k\times k$ matrix and define its Hermitian part: \[ H=\frac{A+A^{\prime}}{2}. \] Then there exists $\mu\neq0$ such that $\mu^{\prime}A \mu=0$ if and only if the convex hull of the spectrum of $H$ contains zero.

Proposition (ref) gives a geometric characterization of impossibility designs. The condition \[ \mu^{\prime}A\mu= 0 \] admits a nontrivial solution if and only if the Hermitian part $H$ is not definite. In particular, impossibility designs arise whenever $H$ has eigenvalues of opposite sign or zero lies in the convex hull of its spectrum.

Thus impossibility designs are not knife-edge choices of $\mu$. They reflect the covariance geometry encoded in $\Sigma$ and $\mu$. Whenever the reduced-form covariance structure is sufficiently nonorthogonal, the LM drift can be bounded.

corollaryUnder Assumptions SIV FA, NA, and ID, \begin{align*} &LM_{1} - \frac{ \Delta \mu^{\prime} \Sigma_{11}^{-1} \mu} { \left(\mu^{\prime} \Sigma_{11}^{-1} \mu+ \Delta^{2} \mu^{\prime} \Sigma_{11}^{-1} \Sigma_{12} \Sigma_{11}^{-1} \Sigma_{21} \Sigma_{11}^{-1} \mu \right) ^{1/2} } \\ \rightarrow_{d} & N \left( 0,\, 1+\|\gamma\|^{-2}\Delta^{2} \pi' D^{1/2}\Sigma_{11}^{-1/2} M_{\gamma}\Sigma_{11}^{-1/2} (\Sigma^{22})^{-1} \Sigma_{11}^{-1/2} M_{\gamma}\Sigma_{11}^{-1/2} D^{1/2}\pi \right). \end{align*}

Under Assumption ID, as $|\Delta| \to \infty$, the LM drift converges to the finite limit

equation[equation omitted — 207 chars of source]

This bound can be made arbitrarily small by replacing $\mu$ by $\eta\cdot\mu$ with $\eta$ small while maintaining the impossibility design restriction. In contrast, the AR statistic has noncentrality parameter \[ \Delta^{2}\mu^{\prime}\Sigma_{11}^{-1}\mu, \] which diverges as $|\Delta|\to\infty$ whenever $\|\Sigma_{11}^{-1/2}\mu\|$ is bounded away from zero. Hence there exist designs in which AR strongly separates the null and alternative while LM remains weak. This separation is structural and reflects the covariance geometry determined jointly by the instrument design and the first-stage coefficients, as encoded in $H$.

Implications for CQLR

Conditional linear combination tests weight AR and LM as functions of $T$. When LM becomes nearly uninformative, any procedure that mixes LM information may lose power relative to procedures that rely primarily on AR. Proposition S.1 in the online supplement provides an illustrative example in which the LM statistic is asymptotically ancillary while the noncentrality of AR is bounded away from zero. In this example, CLC procedures cannot outperform a test that behaves essentially like AR. Since AR-type procedures are not efficient under strong identification, this creates a sharp separation between CLR and AR--LM based tests in these designs.

The same logic applies to CQLR because Andrews16 shows that CQLR is a special case of a CLC test. Moreover, we can find a sharper bound for CQLR directly.

propositionFor the QLR statistic in (ref): (i) if $r(T)\to\infty$, then $QLR\to LM$; (ii) if $r(T)>\text{AR}$, then \[ QLR \le\frac{\text{LM}\cdot r(T)}{r(T)-\text{AR}} = \frac{\text{LM} }{1-\text{AR}/r(T)}. \]

When $r(T)$ diverges, the CQLR statistic collapses to LM. In such designs, CQLR inherits the same power limitations as LM.

figure[figure omitted — 635 chars of source]

Figure (ref) illustrates the mechanism behind the impossibility designs and shows that LM and CQLR can fail even when the null and alternative distributions are nearly perfectly distinguishable. We assume normally distributed errors and set $\Sigma_{11}$ and $\Sigma_{22}$ proportional to the identity matrix and $\Sigma_{12}$ proportional to the anti diagonal identity matrix (that is, its $(i,j)$ element is a positive constant when $i+j=k+1$ and zero otherwise). The proportionality constants are chosen so that $\Sigma_{0}$ is positive definite. This simple covariance structure is chosen for transparency and satisfies the conditions for an impossibility design described above. This design is used only to illustrate the mechanism. The theory in the previous section shows that such failures arise for a nontrivial set of covariance structures. We take $k=10$ and set $\mu=\lambda^{1/2}e_{1}$, where $e_{1}$ denotes the first unit vector. The parameter $\lambda$ is a population $F$ statistic robust to HAC errors and measures instrument strength. We set $\lambda=100$.

We choose a very small nominal size $\alpha=0.001$ to illustrate a case in which the null and alternative are nearly perfectly distinguishable in total variation distance. The testing problem is therefore essentially trivial and the CLR power-size gap is close to one. In contrast, the LM and CQLR tests can fail to exploit this separation and may have power close to their nominal size despite the fact that the null and alternative are nearly perfectly distinguishable. The AR test separates the null and alternative but is dominated by CLR.

Section S-3 in the supplement considers several additional cases and parameter configurations for power comparisons, including impossibility design (ID) setups and near-ID setups, different values of $\alpha$, different numbers of instruments $k$, and different levels of instrument strength.

Empirical Application

Section (ref) shows that under impossibility designs the LM noncentrality parameter can remain bounded even when the null and alternative are well separated in total variation distance. A natural question is whether such designs are empirically relevant or merely theoretical curiosities.

In practice, we do not observe $\mu$ directly. Instead, we estimate $\pi$ and hence $\mu=(Z^{\prime} Z)^{1/2} \pi$. Under weak instruments, we may write $\pi=h_{\pi}/\sqrt{n}$. Because $h_{\pi}$ is not consistently estimable, even large samples do not allow us to determine with certainty whether the data generating process lies on or near an impossibility design.

The empirically relevant question is therefore geometric. Does a standard confidence region for $\mu$ intersect the set of impossibility designs? If so, then parameter values consistent with the data can lie in regions where LM and CQLR lose power, even though the underlying testing problem is not intrinsically difficult.

We study the intersection between a $1-\alpha$ confidence set for $\mu$ and the impossibility design restriction. In particular, we consider

equation[equation omitted — 223 chars of source]

The set described in (ref) is defined by a polynomial inequality and a polynomial equality, and hence is semi algebraic. General results such as the Positivstellensatz give conditions under which a semi algebraic set is nonempty; see Krivine64 and Stengle74. However, verifying those conditions is not convenient in applications. We instead exploit the specific quadratic structure of (ref) and reduce the question of feasibility to a simple optimization problem:

equation[equation omitted — 145 chars of source]

where the Hermitian matrix is

equation[equation omitted — 118 chars of source]

We call the minimum value of the objective function in (ref) the confidence bound, because it is the smallest confidence set cutoff such that the confidence region intersects the impossibility design. If the confidence bound exceeds $q_{1-\alpha}(k)$, then the intersection of the $1-\alpha$ confidence set and the impossibility design is empty.

If the convex hull of the spectrum of $H$ does not contain zero, then trivially the only solution to (ref) is $\widetilde{\mu}=0$. Otherwise, a solution $\widetilde{\mu}$ satisfies the first order condition

equation[equation omitted — 115 chars of source]

where $\kappa$ is a Lagrange multiplier.

The matrix $\Sigma_{22}^{-1}+\kappa H$ may not be invertible. If it is invertible, then

equation[equation omitted — 132 chars of source]

In this case, the constraint in (ref) yields the scalar constraint equation

equation[equation omitted — 205 chars of source]

where $\overline{H}=\Sigma_{22}^{1/2}H\Sigma_{22}^{1/2}$. In ridge regression the left-hand side of the corresponding constraint is decreasing in $\kappa$, so that the solution for $\kappa$ is unique.\footnote{We refer the reader to see DraperNostrand79 for further details.} Here, because $H$ can have eigenvalues of opposite signs, the left-hand side of (ref) need not be monotonic. In practice, we solve for all values of $\kappa$ numerically and then select the $\widehat{\kappa}$ that minimizes the objective function.\footnote{Section S-4 in the online supplement provides more details on our search algorithm.} For any such $\kappa$, the objective value can be written as

equation[equation omitted — 258 chars of source]
table[table omitted — 855 chars of source]

The confidence bound admits a simple geometric interpretation. It is the squared Mahalanobis distance, measured using $\Sigma_{22}^{-1}$, from the point estimate $\widehat{\mu}$ to the nearest point in the impossibility design set. If this distance is smaller than the chi--square cutoff $q_{1-\alpha}(k)$, then the data are statistically consistent with parameter values that lie on an impossibility design. If it exceeds the cutoff, the data rule out such designs at level $\alpha$.

For practitioners, this provides a direct diagnostic. The confidence bound measures how far the estimated first stage lies from the region in which LM and CQLR may lose power. It translates the algebraic condition $\mu^{\prime}\Sigma_{11}^{-1}\Sigma_{21}\Sigma_{11}^{-1}\mu = 0$ into a computable distance comparison based on standard first-stage output. In practice, this issue can also be avoided by simply using the CLR test.

As an example, we consider the estimation of the intertemporal elasticity of substitution (IES) of \Citet{Yogo04}. He considers four instruments and three models. Like \Citet{MoreiraMoreira19}, we focus on the specification in which the dependent variable is real consumption growth and the endogenous regressor is the real stock return, and we use the same estimator of the $\Sigma$ matrix as Andrews16 does. Of the eleven countries considered by \Citet{Yogo04}, nine have eigenvalues with opposite signs, i.e., all countries except the United States and Australia. For these countries, the necessary condition for the LM noncentrality parameter to be bounded is satisfied.

Table (ref) shows, for all countries that satisfy the necessary condition for an impossibility design, the AR noncentrality parameter when $\Delta=1$ (second column), the minimum value of the objective function (third column, that is, the confidence bound), and the value of the objective function at $\mu=0$ (fourth column).

The second column shows that the AR noncentrality parameter when $\Delta=1$ is at least $46.737$. Thus, separating the null from the alternative is not intrinsically difficult. Nevertheless, the third column shows that for all such countries the $95\%$ confidence region intersects the impossibility design set.

In this data $k=4$, so the cutoff for the $95\%$ confidence region is $9.49$. Since the confidence bound is below $9.49$ for each country listed, the data do not rule out configurations in which LM and CQLR lose power.

However, $\widetilde{\mu}$ is not the only point that belongs to this intersection. The vector $\eta\cdot\widetilde{\mu}$ for scalar $\eta$ satisfies the impossibility design restriction and lies in the confidence region if

equation[equation omitted — 165 chars of source]

As is clear from (ref) in Section (ref), choosing $\eta$ sufficiently small can make the upper bound for the noncentrality parameter of the LM statistic arbitrarily small. To find the smallest $\eta$ for which $\eta.\widetilde{\mu}$ lies within the confidence set, we minimize $\eta^{2}$ subject to (ref). The fourth column of Table (ref) shows that, except for Japan, $\eta=0$ is a solution. This implies that the LM noncentrality parameter can be arbitrarily close to zero within the confidence set. For Japan the restriction (ref) is binding, so that

equation[equation omitted — 418 chars of source]

For Japan the solution is $\eta=0.1309$.\footnote{Instead we could search over all $\mu$ in the confidence region satisfying the impossibility design restriction to minimize the LM bound. The semi algebraic nature of this setup implies that the minimization problem in (ref) is a polynomial problem for which we can find a global minimum; see Lasserre15. We do not pursue this approach here because the power of the LM test is already very low in this application.}

figure[figure omitted — 1,159 chars of source]

Figure (ref) shows power curves for Japan. We select this country because it is farthest removed from the impossibility design of all countries reported in Table (ref). For $\mu$ we choose the coefficients corresponding to the confidence bound $\widetilde{\mu}$ and the scaled value $0.1309\,\widetilde{\mu}$ on the boundary of the $95\%$ confidence set.

The power curves are consistent with the simulation results. As in the simulations, the CLR test is not much affected by the power loss at or in the neighborhood of an impossibility design. That conclusion holds for both choices of $\mu$. For $k>1$ the AR test is not optimal. The LM test suffers a substantial loss of power. Section S-5 in the online supplement shows that $r_{2}(T)>AR$, so that Proposition (ref) , part (ii), applies. Consequently, the CQLR2 test, for which the test statistic behaves like the LM statistic, also suffers a substantial loss of power. The CQLR1 dominates the CQLR2 test because the weight on the AR component is much larger in CQLR1 than in CQLR2. The PI CLC test has power similar to the AR test. Moreover, when the LM noncentrality shrinks toward zero, the power of the PI CLC test is bounded by the power of the $J$ overidentification test, with statistic $J = AR - LM$, as predicted by Proposition S.1 in Supplement S-1.

Up to this point, we have not characterized the full set of $\mu$ that are consistent with an impossibility design and that appear within the confidence set. The Tarski Seidenberg theorem guarantees that projections of semi algebraic sets are also semi algebraic sets. The Cylindrical Algebraic Decomposition (CAD) algorithm finds these projections. Figure (ref) presents all six projections in $\mathbb{R}^{2}$ of the set in $\mathbb{R} ^{4}$ for Japan at the $95\%$ confidence level.\footnote{Throughout this paper we take $\Sigma$ as known. If instead $\Sigma$ were estimated and allowed to vary, the set of potentially problematic first-stage estimates would be larger.} The graphs are close to symmetric near zero. This is not surprising because if $\mu$ satisfies the impossibility design restriction, then so does $-\mu$.

figure[figure omitted — 1,079 chars of source]

Conclusion

This paper characterizes the maximal attainable power--size gap in linear IV models using total variation distance and studies which weak identification robust procedures attain this decision theoretic frontier in the presence of heteroskedastic or autocorrelated (HAC) errors.

Our first result establishes that the minimum total variation distance between the convex hulls of the null and alternative models provides a sharp benchmark for the intrinsic difficulty of the testing problem. By Kraft55, this distance determines the largest power--size gap achievable by any measurable test. We show that the conditional likelihood ratio (CLR) test attains this frontier: its power--size gap converges to one if and only if the testing problem becomes trivial in total variation distance. Thus, whenever statistical separation is sufficient to permit near-perfect discrimination between the null and the alternative, the CLR test achieves it.

The contrast with AR-LM-based conditional procedures is sharp in HAC IV models. Under fixed alternatives, retaining all terms that contribute to the finite sample noncentrality parameter of the LM statistic reveals a class of covariance structures, which we call impossibility designs, under which the LM noncentrality parameter remains bounded even as the null and alternative distributions become arbitrarily well separated in total variation distance. In such designs, LM and CQLR can have power arbitrarily close to size despite the fact that the underlying testing problem is not intrinsically difficult. The empirical illustration based on Yogo04 shows that confidence sets for first-stage parameters can intersect this region, so these designs arise in empirically plausible configurations.

Taken together, these results establish a structural separation in HAC IV models. The CLR test exploits the full reduced-form information and achieves the decision theoretic frontier. Procedures constructed from AR and LM statistics operate in a restricted information space and may fail to convert statistical separation into power. This loss is not a local artifact of asymptotic approximations. It can persist even when the null and alternative are strongly separated.

Although our analysis is developed for HAC IV regression and extends naturally to GMM, the broader lesson concerns the evaluation of robust procedures more generally. Efficiency under standard asymptotics and robustness to weak identification do not by themselves guarantee good performance in richer environments. When covariance structures become more complex, procedures that discard information may suffer substantial power losses while continuing to control size. Decision theoretic benchmarks based on total variation distance provide a disciplined way to assess whether a test fully exploits the information available in the model.

FGV EPGE, Praia de Botafogo, 190, 11th floor, Rio de Janeiro, RJ 22250-040, Brasil; [email removed]

Department of Economics, University of Southern California, Kaprielian Hall, Los Angeles, USA, CA 90089. Electronic; [email removed]

Department of Economics, University of Pittsburgh, 4927 Wesley W. Posvar Hall, 230 S Bouquet St., Pittsburgh, PA 15260; [email removed]

Acknowledgments

Preliminary results of this paper were presented at seminars organized by BU, Brown, Caltech, Harvard-MIT, PUC-Rio, University of California (Berkeley, Davis, Irvine, Los Angeles, Santa Barbara, and Santa Cruz campuses), UCL, USC, and Yale, at the FGV Data Science workshop, and at conferences organized by CIREq (in honor of Jean-Marie Dufour), Harvard University (in honor of Gary Chamberlain), Oxford University (New Approaches to the Identification of Macroeconomic Models), and the Tinbergen Institute (Inference Issues in Econometrics). We thank Marinho Bertanha, Leandro Gorno, Michael Jansson, Pierre Perron, and Jack Porter for helpful comments; and Pedro Melgar\'{e} for excellent research assistance. This study was financed in part by the Coordena\c{c}\ {a}o de Aperfei\c{c}oamento de Pessoal de N\' ivel Superior - Brasil (CAPES) - Finance Code 001. It was also supported in part by the University of Pittsburgh Center for Research Computing and Data, RRID:SCR_022735, through the resources provided. Specifically, this work used the H2P cluster, which is supported by NSF award number OAC-2117681.