EconBase
← Back to paper

Best Feasible Conditional Critical Values for a More Powerful Subvector Anderson-Rubin Test

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

38,233 characters · 9 sections · 16 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Best Feasible Conditional Critical Values for a More Powerful Subvector Anderson-Rubin Test

abstract\baselineskip=15pt For subvector inference in the linear instrumental variables model under homoskedasticity but allowing for weak instruments, \citet*{GuggenbergerKleiMavrQE2019} (GKM) propose a conditional subvector AndersonRubin1949 ($AR$) test that uses data-dependent critical values that adapt to the strength of the parameters not under test. This test has correct size and strictly higher power than the test that uses standard asymptotic chi-square critical values. The subvector $AR$ test is the minimum eigenvalue of a data dependent matrix. The GKM critical value function conditions on the largest eigenvalue of this matrix. We consider instead the data dependent critical value function conditioning on the second-smallest eigenvalue, as this eigenvalue is the appropriate indicator for weak identification. We find that the data dependent critical value function of GKM also applies to this conditioning and show that this test has correct size and power strictly higher than the GKM test when the number of parameters not under test is larger than one. Our proposed procedure further applies to the subvector $AR$ test statistic that is robust to an approximate kronecker product structure of conditional heteroskedasticity as proposed by \citet*{Guggenberger_Kleibergen_Mavroeidis_2024}, carrying over its power advantage to this setting as well.

{\smallJEL Classification:}{ C12, C26.}

{\smallKeywords:}{ Instrumental Variables, Weak-Instrument Robust Inference, Subvector $AR$ Test, LIML, Overidentification Test, Power.}

\thispagestyle{empty}

\baselineskip=20pt

\setcounter{page}{1}

Introduction

For the linear instrumental variables model under homoskedasticity, \citet*{GuggenbergerKleiMavrQE2019} (GKM) introduced a subvector $AR$ testing procedure that uses data-dependent critical values that adapt to the identification strength of the parameters not under test. This test has correct size and strictly higher power than the test that uses standard asymptotics chi-square critical values. The subvector $AR$ test is the minimum eigenvalue of a data dependent matrix and the GKM critical-value function conditions on the largest eigenvalue of this non-central Wishart distributed matrix. If this largest eigenvalue is small then the nuisance parameters not under test suffer from a weak-instruments problem. The $AR$ test based on the chi-square critical values is then undersized and the conditional approach of GKM decreases the critical value. It does this maintaining size control, and thus improves power compared to the standard approach.

Whilst a small largest eigenvalue does indicate identification problems, a large largest eigenvalue does not necessarily imply strong identification. For example, the standard weak-instruments test for the two-stage least squares estimator is the CraggDon1993 test. This is a rank test on the first-stage parameters, which is the minimum eigenvalue of the first-stage concentration matrix, see also StockYogo2005. Under the maintained assumption of instrument validity, the equivalent in the LIML based $AR$ subvector test is the second-smallest eigenvalue. A small second-smallest eigenvalue then indicates weak-instrument problems for the estimation of the parameters not under test, affecting the size and power of the test.

We therefore investigate the use of data-dependent critical values that condition not on the largest but on the second-smallest eigenvalue. Our main finding is that we can use the same critical-value function as derived by GKM. We obtain this result by first showing that, for $p>2$, if the $p-2$ largest eigenvalues of the $p\times p$ non-central Wishart distributed matrix are all large and increasing to $\infty$, then the joint distribution of the two smallest eigenvalues is the same as that of a $2\times2$ Wishart matrix, but with a change in degrees of freedom, which is the difference between the number of instruments and the number of parameters not under test. The two eigenvalues of the associated non-centrality matrix are the same as the two smallest ones for the $p\times p$ matrix. Therefore, in this large $p-2$ largest-eigenvalues setting, the results and tabulated critical values of GKM for different degrees of freedom apply directly, simply substituting the largest for the second-smallest eigenvalue as the conditioning variable. This implies that size is controlled in this case. Our second contribution is then to show that when the $p-2$ largest eigenvalues decrease in value to be less than $\infty$, the rejection probability of the test gets smaller, hence the test will have size smaller or equal to nominal size under the null in all circumstances. This second part is shown by direct calculation of the joint distribution of the two smallest eigenvalues, based on the results of james1964distributions and muirhead2009aspects, from which the conditional distribution can be obtained. Simulations, drawing from the non-central Wishart distribution, give the same results. As the direct calculations are very time consuming, we confirm the second finding further with a range of simulations.

For $p>2$, as the critical value when conditioning on the second smallest eigenvalue is smaller than that conditioning on the largest eigenvalue, this $AR$ subvector testing procedure has strictly higher power than the GKM test.

The distribution of the subset $AR$ test statistic is of course a function of all $p-1$ largest eigenvalues of the non-central Wishart distributed matrix, and other approaches that consider a larger subset of these eigenvalues could increase power, but likely at the expense of much higher computational cost (GKM, pp 496-497). Our method retains one of the main advantages of the GKM test, namely its computational simplicity. The second-smallest eigenvalue captures the information regarding weak identification better than the largest eigenvalue. We therefore consider our approach to be the best feasible subset $AR$ test in this linear IV setting -- it maintains computational simplicity while properly incorporating weak identification of the parameters not under test.

In the next section, Section (ref), we first present the finite sample model setup and analysis of GKM and present their conditioning on the largest eigenvalue $AR$ testing procedure in Section (ref). In Section (ref) we present our main results for the $AR$ testing procedure conditioning on the second-smallest eigenvalue. Section (ref) presents some simulation results, confirming the correct size and superior power properties of our proposed test. Guggenberger_Kleibergen_Mavroeidis_2024 show for a model with heteroskedasticity of a kronecker form that their conditional approach applies to a robust version of the $AR$ test statistic. In Appendix (ref) we show that these results also extend to our conditional approach. Section (ref) concludes.

Finite Sample Analysis

We follow the finite sample analysis of GKM and consider the linear model with endogenous variables $X$ and $W$, given by the equations

align[align omitted — 128 chars of source]

where $y\in\mathbb{R}{}^{n}$, $X\in\mathbb{R}{}^{n\times m_{X}}$, $W\in\mathbb{R}{}^{n\times m_{W}}$, the instruments $Z\in\mathbb{R}{}^{n\times k}$ and $k-m_{W}\geq1$. As in GKM, we assume that the instruments $Z$ are fixed, $Z^{\prime}Z$ is positive definite and that $\left(\varepsilon_{i},V_{Xi}^{\prime},V_{Wi}^{\prime}\right)^{\prime}\sim i.i.d.~\mathcal{N}\left(0,\Sigma\right)$, $i=1,\ldots,n$, for some $\Sigma\in\mathbb{R}{}^{(m+1)\times(m+1)}$, with $m=m_{X}+m_{W}$, and \[ \Sigma=\left(

array[array omitted — 280 chars of source]

\right). \]

For testing the hypothesis \[ H_{0}:\beta=\beta_{0}\,\,\,\text{against}\,\,\,H_{1}:\beta\neq\beta_{0}, \] let \[ y_{0}:=y-X\beta_{0}, \] and consider the restricted model given by the equations

align[align omitted — 115 chars of source]

with $\varepsilon\left(\beta_{0}\right)=\varepsilon+X\left(\beta-\beta_{0}\right)$.

The reduced form for $y_{0}$ is given by \[ y_{0}=Z\pi_{y_{0}}+v_{y_{0}}, \] with $\pi_{y_{0}}=\Pi_{X}\left(\beta-\beta_{0}\right)+\Pi_{W}\gamma$ and $v_{y_{0}}=\varepsilon+V_{X}\left(\beta-\beta_{0}\right)+V_{W}\gamma$. It follows that $\left(v_{y_{0}i},V_{Wi}^{\prime}\right)\sim i.i.d.\,\mathcal{N}\left(0,\Omega\left(\beta_{0}\right)\right)$, with

align*[align* omitted — 244 chars of source]

where we assume, as in GKM, that $\Omega\left(\beta_{0}\right)$ is known and positive definite.

The subvector AndersonRubin1949 test statistic is then defined as

align*[align* omitted — 441 chars of source]

where $P_Z=Z(Z'Z)^{-1}Z'$ and $\min\,\text{eval}\left(A\right)$ denotes the minimum eigenvalue of $A$. The resulting estimator for $\tilde{\gamma}$ is the so-called LIMLK estimator of $\gamma$ in the restricted model ((ref)), an infeasible estimator which requires the reduced form error covariance matrix to be known, see Anderson1977 and the discussion in StaigerStockEcta1997.\footnote{The feasible statistic as evaluated later in Section (ref) is given by $AR_{n,f}\left(\beta_{0}\right)=\frac{\left(y_{0}-W\widehat{\tilde{\gamma}}_{L}\right)^{\prime}P_{Z}\left(y_{0}-W\widehat{\tilde{\gamma}}_{L}\right)}{\left(y_{0}-W\widehat{\tilde{\gamma}}_{L}\right)^{\prime}M_{Z}\left(y_{0}-W\widehat{\tilde{\gamma}}_{L}\right)/\left(n-k\right)}$, where $\widehat{\tilde{\gamma}}_{L}$ is the LIML estimator of $\gamma$ in the restricted model ((ref)) and $M_Z=I_p-P_Z$}

Denote the $p:=m_{W}+1$ eigenvalues of the matrix $\left(y_{0},W\right)^{\prime}P_{Z}\left(y_{0},W\right)\left(\Omega\left(\beta_{0}\right)\right)^{-1}$ by $\widehat{\kappa}_{1}\geq\widehat{\kappa}_{2}\geq\ldots\geq\widehat{\kappa}_{p}$, so that $AR_{n}\left(\beta_{0}\right)=\widehat{\kappa}_{p}$. GKM show that the eigenvalues $\widehat{\kappa}_{j}$ for $j=1,\ldots,p$ are equal to the eigenvalues of a $p\times p$ non-central Wishart matrix $\Xi^{'}\Xi$, with the $k$ rows of the $k\times p$ matrix $\Xi$ are independently normally distributed with common covariance matrix $I_{p}$ and $\mathbb{E}\left(\Xi\right)=\mathcal{M}$. Therefore, $\Xi^{'}\Xi\sim\mathcal{W}_{p}\left(k,I_{p},\mathcal{M}^{\prime}\mathcal{M}\right)$, a non-central Wishart distribution with $k$ degrees of freedom, covariance matrix $I_{p}$ and noncentrality matrix $\mathcal{M}^{\prime}\mathcal{M}$.

As GKM show, under the null $H_{0}:\beta=\beta_{0}$, $\mathcal{M}=\left(0,\Theta_{W}\right)$, where the $k\times m_{W}$ matrix $\Theta_{W}$ is given by \[ \Theta_{W}:=\left(Z^{\prime}Z\right)^{1/2}\Pi_{W}\Sigma_{V_{W}V_{W}.\varepsilon}^{-1/2}, \] where \[ \Sigma_{V_{W}V_{W}.\varepsilon}=\Sigma_{V_{W}V_{W}}-\Sigma_{\varepsilon V_{W}}^{\prime}\Sigma_{\varepsilon V_{W}}\sigma_{\varepsilon\varepsilon}^{-1}. \] It follows that then \[ \mathcal{M}^{\prime}\mathcal{M}=\left(

array[array omitted — 58 chars of source]

\right), \] with \[ \Theta_{W}^{\prime}\Theta_{W}=\Sigma_{V_{W}V_{W}.\varepsilon}^{-1/2}\Pi_{W}^{\prime}Z^{\prime}Z\Pi_{W}\Sigma_{V_{W}V_{W}.\varepsilon}^{-1/2}. \] The joint distribution of $\left(\widehat{\kappa}_{1},\ldots,\widehat{\kappa}_{p}\right)$ under the null then only depends on the eigenvalues of $\Theta_{W}^{\prime}\Theta_{W}$, which GKM denote by \[ \kappa_{j}:=\kappa_{j}\left(\Theta_{W}^{\prime}\Theta_{W}\right),\,\,\,\,j=1,\ldots,m_{W}. \]

Concentration Parameter

It is at this point interesting to compare the $\Theta_{W}^{\prime}\Theta_{W}$ concentration matrix to the standard concentration matrix which is defined using the first-stage model for $W$ and given by $\Sigma_{V_{W}V_{W}}^{-1/2}\Pi_{W}^{\prime}Z^{\prime}Z\Pi_{W}\Sigma_{V_{W}V_{W}}^{-1/2}$, see \citet*{StockWrightYogoJBES2002} and StockYogo2005. For the $m_{W}=1$ case, and writing the variance matrix under the null in the restricted model ((ref)) as $\Sigma_{r}=\left(

array[array omitted — 120 chars of source]

\right)$ and the correlation coefficient $\rho_{\varepsilon v_{w}}=\frac{\sigma_{\varepsilon v_{w}}}{\sigma_{\varepsilon}\sigma_{v_{w}}}$, the standard concentration parameter is given by \[ \mu_{w}^{2}=\frac{\pi_{w}^{\prime}Z^{\prime}Z\pi_{w}}{\sigma_{v_{w}}^{2}}, \] whereas \[ \theta_{w}^{\prime}\theta_{w}=\frac{\pi_{w}^{\prime}Z^{\prime}Z\pi_{w}}{\sigma_{v_{w}}^{2}-\frac{\sigma_{\epsilon v_{w}}^{2}}{\sigma_{\varepsilon}^{2}}}=\frac{\pi_{w}^{\prime}Z^{\prime}Z\pi_{w}}{\sigma_{v_{w}}^{2}\left(1-\rho_{\varepsilon v_{w}}^{2}\right)}. \] $\mu_{w}^{2}$ is the noncentrality parameter for the Wald test testing $H_{0}:\pi_{w}=0$ based on the OLS estimator $\widehat{\pi}_{w}=\left(Z^{\prime}Z\right)^{-1}Z^{\prime}w$, whereas $\theta_{w}^{\prime}\theta_{w}$ is the noncentrality parameter for the Wald test based on, here, the infeasible LIMLK estimator of $\pi_{w}$ in the restricted model ((ref)). As this is essentially part of a FIML type system estimator, the structural errors enter this expression.

Under the weak instrument, local-to-zero setting $\pi_{w}=\pi_{w,n}=c/\sqrt{n}$, with $c$ a vector of constants, it follows that $\theta_{w}^{\prime}\theta_{w}=$$\frac{c^{\prime}\left(Z^{\prime}Z/n\right)c}{\sigma_{v_{w}}^{2}\left(1-\rho_{\varepsilon v_{w}}^{2}\right)}$, and so, assuming that $c^{\prime}\left(Z^{\prime}Z/n\right)c<a$, $\forall n$, for a finite $a\in\mathbb{R}^{+}$, under weak instruments this noncentrality/concentration parameter $\theta_{w}^{\prime}\theta_{w}=\kappa_{1}\rightarrow\infty$ when $\rho_{\varepsilon v_{w}}^{2}\rightarrow1$, and hence in that case the weak-instrument distribution of the $AR_{n}\left(\beta_{0}\right)$ statistic under the null is the standard strong-instrument $\chi_{k-1}^{2}$ distribution, following the results of Theorem 1 in GKM, see also VandeSijpe2023.

Conditioning on Largest Eigenvalue

$m_{W}=1$

For the $m_{W}=1$ case, GKM show that an approximate conditional density of the smallest eigenvalue, $\widehat{\kappa}_{2}$, given the largest, $\widehat{\kappa}_{1}$, is given by

equation[equation omitted — 295 chars of source]

where $f_{\chi_{k-1}^{2}}\left(.\right)$ is the density of a $\chi_{k-1}^{2}$ and $g\left(\widehat{\kappa}_{1}\right)$ is a function that does not depend on any unknown parameters. Using this conditional density function GKM derive and tabulate data dependent, conditional on $\widehat{\kappa}_{1}$, critical values $c_{1-\alpha}\left(\widehat{\kappa}_{1},k-1\right)$ for the subvector $AR_{n}\left(\beta_{0}\right)$ statistic. As here $AR_{n}\left(\beta_{0}\right)=\widehat{\kappa}_{2}$, the test function for rejecting $H_{0}$ at nominal size $\alpha$ is written as

equation[equation omitted — 171 chars of source]

where $\mathbf{1}\left[.\right]$ is the indicator function and, for general $s,r$, $\widehat{\kappa}_{s,r}:=\left\{ \widehat{\kappa}_{s},\widehat{\kappa}_{r}\right\} $.

GKM (Theorem 2, p 496) show by numerical integration and Monte Carlo simulation that the conditional critical values $c_{1-\alpha}\left(\widehat{\kappa}_{1},k-1\right)$ guarantee size control for the subvector $AR$ test defined in ((ref)). This result holds for the critical value functions that are tabulated in GKM, for $\alpha\in\{0.01,0.05,0.10\}$ and $k-m_W\in\{1,\ldots,20\}$ for a grid of values for $\widehat\kappa_1$, with the conditional critical value for any particular observed value of $\widehat\kappa_{1}$ obtained by linear interpolation.

As $c_{1-\alpha}\left(\widehat{\kappa}_{1},k-1\right)<c_{1-\alpha}\left(\infty,k-1\right)$, $\phi_{c}$ has strictly higher power than the unconditional approach, the subvector $AR$ test $\phi_{\chi^{2}}$ that uses the $1-\alpha$ critical values of the $\chi_{k-1}^{2}$ distribution, the latter discussed in \citet*{GugKleiMavrChenEcta2012}.

$m_{W}>1$

The GKM approach for $m_{W}>1$ is to use the test function \[ \phi_{c_{1}}\left(\widehat{\kappa}_{1,p}\right):=\mathbf{1}\left[\widehat{\kappa}_{p}>c_{1-\alpha}\left(\widehat{\kappa}_{1},k-m_{W}\right)\right], \] where the subscript in $c_{1}$ makes it clear that conditioning is on the largest eigenvalue. GKM (Corollary 4, p 500) show that $\phi_{c_{1}}$ has correct size for the tabulated critical value functions as mentioned in Section (ref), with $k-m_W\in\{1,\ldots,20\}$, and has power strictly higher than $\phi_{\chi^{2}}$ that uses the critical values of the $\chi_{k-m_{W}}^{2}$ distribution.

Conditioning on Second-Smallest Eigenvalue

Instead of the conditioning on the largest eigenvalue, we propose conditioning on the second-smallest eigenvalue, $\widehat{\kappa}_{m_{W}}$, with the test function given by \[ \phi_{c_{p-1}}\left(\widehat{\kappa}_{p-1,p}\right):=\mathbf{1}\left[\widehat{\kappa}_{p}>c_{1-\alpha}\left(\widehat{\kappa}_{p-1},k-m_{W}\right)\right]. \] The reason for considering conditioning on the second-smallest eigenvalue is that the smallest eigenvalue of the concentration matrix $\Theta_{W}^{\prime}\Theta_{W}$ is a better indicator of the strength of the identification of $\gamma$ than the largest eigenvalue. This is also why the standard CraggDon1993 rank test for underidentification is based on the minimum eigenvalue of the concentration matrix $\widehat{\Sigma}_{V_{W}V_{W}}^{-1/2}\widehat{\Pi}_{W}^{\prime}Z^{\prime}Z\widehat{\Pi}_{W}\widehat{\Sigma}_{V_{W}V_{W}}^{-1/2}$, with $\widehat{\Pi}_{W}=\left(Z^{\prime}Z\right)^{-1}Z^{\prime}W$ and $\widehat{\Sigma}_{V_{W}V_{W}}=W^{\prime}M_{Z}W/\left(n-k\right)$, see also StockYogo2005.

For example, for $m_{W}=2$, we can have the situation that $\kappa_{1}=\infty$, but $\kappa_{2}$ is small, in which case $\gamma$ is weakly identified. Conditioning on $\widehat{\kappa}_{1}$ would then result in conditional critical values close to those of the $\chi_{k-m_{W}}^{2}$, whereas conditioning on $\widehat{\kappa}_{2}$ would properly address the weak identification problem.

\sloppy As $\kappa_{1}\geq\kappa_{2}\geq...\geq\kappa_{p}$, we have that if $\kappa_{p-2}=\infty$, it follows that also $\kappa_{j}=\infty$, $j=1,\ldots,p-3$. For $\Xi^{\prime}\Xi\sim\mathcal{W}_{p}\left(k,I_{p},\mathcal{M}^{\prime}\mathcal{M}\right)$, our main finding is that the joint distribution of $\left(\widehat{\kappa}_{p-1},\widehat{\kappa}_{p}\right)$ when the eigenvalues of $\mathcal{M}^{\prime}\mathcal{M}$ are given by $\left\{ \kappa_{1}=\ldots=\kappa_{p-2}=\infty,\kappa_{p-1},\kappa_{p}=0\right\} $ is the same as that of $\left(\widehat{\kappa}_{1}^{*},\widehat{\kappa}_{2}^{*}\right)$, which are the eigenvalues of $\Xi^{*\prime}\Xi^{*}\sim\mathcal{W}_{2}\left(k-m_{W}+1,I_{2},\mathcal{M^{*}}^{\prime}\mathcal{M}^{*}\right)$, with the two eigenvalues of $\mathcal{M^{*}}^{\prime}\mathcal{M}^{*}$ equal to $\kappa_{1}^{*}=\kappa_{p-1}$ and $\kappa_{2}^{*}=0$. It therefore follows that \[ f_{\widehat{\kappa}_{p}|\widehat{k}_{p-1};\kappa_{p-2}=\infty}^{*}\left(x_{p}|\widehat{\kappa}_{p-1};\kappa_{p-2}=\infty\right)=f_{\chi_{k-m_{W}}^{2}}\left(x_{p}\right)\left(\widehat{\kappa}_{p}-x_{p}\right)^{1/2}g\left(\widehat{\kappa}_{p-1}\right),\,\,\,x_{p}\in\left[0,\widehat{\kappa}_{p-1}\right], \] and so the conditional critical values of GKM apply directly to this conditioning, only adjusting the degrees of freedom commensurate with $m_{W}$. We state this result formally in the following Proposition.

propLet $\Xi^{\prime}\Xi\sim\mathcal{W}_{p}\left(k,I_{p},\mathcal{M}^{\prime}\mathcal{M}\right)$, with $p=m_{W}+1>2$, and let $\widehat{\kappa}_{1}\geq\widehat{\kappa}_{2}\geq\ldots\geq\widehat{\kappa}_{p}$ denote the ordered eigenvalues of $\Xi^{\prime}\Xi$. The joint distribution of $\left(\widehat{\kappa}_{p-1},\widehat{\kappa}_{p}\right)$ when the eigenvalues of $\mathcal{M}^{\prime}\mathcal{M}$ are given by $\left\{ \kappa_{1}=\ldots=\kappa_{p-2}=\infty,\kappa_{p-1},\kappa_{p}=0\right\} $ is the same as that of $\left(\widehat{\kappa}_{1}^{*},\widehat{\kappa}_{2}^{*}\right)$, which are the eigenvalues of $\Xi^{*\prime}\Xi^{*}\sim\mathcal{W}_{2}\left(k-m_{W}+1,I_{2},\mathcal{M^{*}}^{\prime}\mathcal{M}^{*}\right)$, with the two eigenvalues of $\mathcal{M^{*}}^{\prime}\mathcal{M}^{*}$ equal to $\kappa_{1}^{*}=\kappa_{p-1}$ and $\kappa_{2}^{*}=0$.
proofSee Appendix (ref).

It follows directly from Proposition (ref) that the critical value functions tabulated by GKM for different values of $k-m_{W}$ apply directly to this setting, simply replacing in their tables the values of $\widehat{\kappa}_{1}$ by $\widehat{\kappa}_{p-1}$.

In Appendix (ref), we further show how to calculate the joint probability density function (pdf) of the ordered eigenvalues $\widehat{\kappa}_{1}\geq\widehat{\kappa}_{2}\geq\ldots\geq\widehat{\kappa}_{p}$, from which we can calculate the conditional pdf $f_{\widehat{\kappa}_{p}|\widehat{\kappa}_{p-1}}\left(x_{p}|\widehat{\kappa}_{p-1}\right)$. We find that the test procedure $\phi_{c_{p-1}}\left(\widehat{\kappa}_{p-1,p}\right)$ controls size also for the cases where all or some of the $\kappa_{j}$, $j=1,\ldots,p-2$ are less then $\infty$, with the main result as follows.

resultConsider, under the null $H_{0}:\beta=\beta_{0}$ and for given values of $\kappa_{p-1}$ and $\kappa_{p}=0$, the rejection probability of the subset $AR$ test using the critical values based on conditioning on the second-smallest eigenvalue. The rejection probability when $\kappa_{1:(p-2)}<\infty$ is bounded from above by the rejection probability when $\kappa_{1:(p-2)}=\infty$, \[ \mathbb{P}\left(AR_{n}\left(\beta_{0}\right)_{\kappa_{1:(p-2)}<\infty,\kappa_{p-1},\kappa_{p}=0}>c_{1-\alpha}\left(\widehat{\kappa}_{p-1},k-m_{W}\right)\right) \] \begin{equation} \leq\mathbb{P}\left(AR_{n}\left(\beta_{0}\right)_{\kappa_{1:(p-2)}=\infty,\kappa_{p-1},\kappa_{p}=0}>c_{1-\alpha}\left(\widehat{\kappa}_{p-1},k-m_{W}\right)\right)\leq\alpha. \end{equation}

Result (ref) again holds for the critical value functions as tabulated in GKM, as described in Section (ref). We can see the combination of Proposition (ref) and Result (ref) as an extension of Theorem 1 in GKM, which states for the $p=2$ case that the distribution of the minimum eigenvalue $\widehat{\kappa}_{2}$ is monotonically decreasing in $\kappa_{1}$, and converging to $\chi_{k-1}^{2}$ as $\kappa_{1}\rightarrow\infty$.

To confirm Result (ref), we show in Table (ref) some calculated points on the conditional cdfs for $p=3$, \[ \mathbb{P}\left(\widehat{\kappa}_{3}\leq a|\widehat{\kappa}_{2}=10;\kappa_{1},\kappa_{2}=2,\kappa_{3}=0\right), \] for the values $a=\left(4,5,6\right)$, and different values of $\kappa_{1}=\left(2,5,10,\infty\right)$. It is clear that the conditional cdfs are monotonically decreasing in $\kappa_{1}$, from which Result (ref) follows.

table[table omitted — 615 chars of source]

\sloppy The numerical calculations of the conditional cdfs are further confirmed by simulations. We create large samples drawn from the Wishart distribution, $\Xi^{\prime}\Xi\sim\mathcal{W}_{p}\left(k,I_{p},\mathcal{M}^{\prime}\mathcal{M}\right)$, with \[ \mathcal{M}^{\prime}\mathcal{M}=\left(

array[array omitted — 132 chars of source]

\right), \] and obtain from these samples the conditional cdfs as the observed frequencies $P\left(\widehat{\kappa}_{p}\leq a|\widehat{\kappa}_{p-1}=b\pm h\right)$, for a small bandwidth $h=0.1$.\footnote{The code for these simulations can be found at \url{https://github.com/jesse-hoekstra/subvector_AR_HW}.} Figure (ref) shows the results for the same setup as in Table 1. The monotonicity is again confirmed across the range $a\in[0,10]$. The right panel further shows that our calculated conditional cdf points from Table (ref) align exactly with the simulated ones. In Appendix (ref), we show simulated conditional cdfs for a range of settings, all showing the same kind of monotonicity and confirming Result (ref).

figure[figure omitted — 281 chars of source]

Using the same simulation design, we next present the null rejection probabilities (NRP) for the $AR$ test procedures $\phi_{\chi^{2}}$, $\phi_{c_{1}}$ and $\phi_{c_{p-1}}$at 5% level, as a function of $\kappa_{p-1}$. For each $m_{W}>1$, we first set $\kappa_{1}=\ldots=\kappa_{p-2}=100,000$ and these results are presented in Figure (ref). It is clear that the behaviour of $\phi_{c_{p-1}}$ is the same for all values of $m_{W}$, whereas here the behaviours of $\phi_{\chi^{2}}$ and $\phi_{c_{1}}$ are the same for $m_{W}>1$, as $\kappa_{1}$ is set very large.

figure[figure omitted — 764 chars of source]

In Figure (ref) we next show the NRPs of the tests for $m_{W}=2$, in the left panel as a function of $\kappa_{2}$ with $\kappa_{1}=\kappa_{2}$ and in the right panel as a function of $\kappa_{1}$ for a fixed value of $\kappa_{2}=10$. We see that the rejection frequencies of the tests are monotonically increasing in $\kappa_{1}$ when keeping $\kappa_{2}$ fixed and the NRPs as a function of $\kappa_{2}$ with $\kappa{}_{1}=\kappa_{2}$ are hence all below those when $\kappa_{1}=100,000$ as displayed in Figure (ref).

figure[figure omitted — 482 chars of source]

Figure (ref) shows similar results for the $m_{W}=3$ case, confirming that the $\phi_{c_{p-1}}$ test has correct size.

As, for $p>2$, $c_{1-\alpha}\left(\widehat{\kappa}_{p-1},k-m_W\right)<c_{1-\alpha}\left(\widehat\kappa_1,k-m_W\right)$, $\phi_{c_{p-1}}$ has strictly higher power than the GKM test $\phi_{c_1}$ when $m_W>1$.

figure[figure omitted — 700 chars of source]

Feasible $AR$ Test Asymptotic Power Comparisons

GKM specify in their Section 3 the mild conditions under which their finite sample results extend to asymptotic results in a setting where the instruments are random, the reduced-form and first-stage errors not necessarily normally distributed and $\Omega$ is unknown. Instead they assume that the random vectors $\left(\varepsilon_i,Z_i',V_{Xi}',V_{Wi}'\right)$ for $i=1,\ldots,n$ are i.i.d.\ with distribution $F$.

Let $U_i=\left(\varepsilon_i+V_{Wi}'\gamma,V_{Wi}'\right)'$. Then GKM consider the parameter space $\mathcal{F}$ for $\left(\gamma,\Pi_W,\Pi_X,F\right)$ under the null hypothesis $H_0:\beta=\beta_0$,

align[align omitted — 607 chars of source]

for some $\delta_1>0$, $\delta_2>0$ and $B<\infty$.

The feasible subset $AR$ test is given by

align*[align* omitted — 415 chars of source]

where \[ \widehat{\Omega}_0=\left(y_{0},W\right)^{\prime}M_{Z}\left(y_{0},W\right)/(n-k-1), \] and $\widehat{\tilde{\gamma}}_{L}$ is the LIML estimator of $\gamma$ in restricted model ((ref)). The additional degree of freedom loss is because we include a constant in the model by taking all variables in deviations from sample means. Note that the feasible test is identical to the LIML-based Basmann1960 test for overidentifying restrictions in the restricted model.

As in GKM, denote the $p$ eigenvalues of the matrix $\left(y_{0},W\right)^{\prime}P_{Z}\left(y_{0},W\right)\widehat\Omega_0^{-1}$ by $\widehat{\kappa}_{1n}\geq\widehat{\kappa}_{2n}\geq\ldots\geq\widehat{\kappa}_{pn}$, so that $AR_{n,f}\left(\beta_{0}\right)=\widehat{\kappa}_{pn}$. GKM (Theorem 5, p 502) then show the result that for the parameter space $\mathcal{F}$ defined in ((ref)) the feasible conditional subvector $AR$ test rejects $H_0$ at nominal size $\alpha$ asymptotically if $$ AR_{n,f}\left(\beta_0\right)>c_{1-\alpha}\left(\widehat\kappa_{1n},k-m_W\right), $$ where $c_{1-\alpha}(.,.)$ is the same conditional critical value function as for the finite sample results in Section (ref). Again, these critical value functions are tabulated in GKM for $\alpha\in\{0.01,0.05,0.10\}$ and $k-m_W\in\{1,\ldots,20\}$ for a grid of values for $\widehat\kappa_1$, with the conditional critical values for any value of $\widehat\kappa_{1n}$ obtained by linear interpolation.

GKM show in the proof of their Theorem 5 that the limiting null rejection probabilities equal the finite sample ones as derived in Section (ref), see GKM (Comment 1, p 502). Because of this convergence, it follows from the results in Section (ref) that the feasible conditional subvector $AR$ test also rejects $H_0$ at nominal size $\alpha$ asymptotically if we condition on the second-smallest eigenvalue, that is if \[ AR_{n,f}\left(\beta_0\right)>c_{1-\alpha}\left(\widehat\kappa_{(p-1)n},k-m_W\right). \]

For $m_W>1$, as the critical value when conditioning on the second smallest eigenvalue is smaller than that conditioning on the largest eigenvalue, this $AR$ subvector testing procedure has strictly higher power than the GKM test. This is confirmed in the following simulations.

For the full model specification ((ref)), we set $m_{X}=1$, $m_{W}=3$, $k=7$, $\beta=0$, $\gamma=\left(-1,1,1\right)^{\prime}$, $n=250$, $Z_{i}\sim i.i.d.~\mathcal{N}\left(0,I_{k}\right)$ and $\left(\varepsilon_{i},V_{Xi},V_{Wi}^{\prime}\right)^{\prime}\sim i.i.d.~\mathcal{N}\left(0,\Sigma\right)$. The specific parameter values of $\Sigma$, $\pi_{x}$ and $\Pi_{W}$ are given in Appendix (ref). The resulting expectation of the concentration matrix is given by \[ \mathbb{E}_{Z}\left(\Theta_{W}^{\prime}\Theta_{W}\right)=n\Sigma_{V_{W}V_{W}.\varepsilon}^{-1/2}\Pi_{W}^{\prime}\Pi_{W}\Sigma_{V_{W}V_{W}.\varepsilon}^{-1/2}=\left(

array[array omitted — 78 chars of source]

\right). \] Figure (ref) presents power curves for $\phi_{\chi^{2}}$, $\phi_{1}$ and $\phi_{p-1}$, for testing $H_{0}:\beta=0$ at level $\alpha=0.05$, for 100,000 replications at each value of $\beta=-2,-1.95,\ldots,2$. We consider the weakly identified situations with $\kappa=\left(\kappa_{1},\kappa_{2},\kappa_{3}\right)=\left(35,25,15\right)$ and $\kappa=\left(100,30,15\right)$ and a strongly identified case with $\kappa=\left(100,95,90\right)$. The $\phi_{p-1}$ test controls size and clearly dominates power in the weak identification settings, with the tests behaving similarly in the strong identified setting, as expected.

figure[figure omitted — 477 chars of source]

Concluding Remarks

We have shown that the GuggenbergerKleiMavrQE2019 conditional critical value function for the subvector $AR$ test in the linear model under homoskedasticy applies when conditioning on the second-smallest eigenvalue instead of the largest eigenvalue of a non-central Wishart distributed matrix. This test procedure that conditions on the second-smallest eigenvalue controls size and has power strictly higher than the GKM test.

Guggenberger_Kleibergen_Mavroeidis_2024 show further that the same conditional critical value function applies to a situation with conditional heteroskedasticity, where the covariance matrix has an approximate Kronecker product structure. The subvector $AR$ statistic is redefined to take account of this Kronecker product structure and the eigenvalues are from the associated concentration matrix. We present further details in Appendix (ref). As for the homoskedastic case, it follows that also for this particular structure of conditional heteroskedasticy, conditioning on the second-smallest eigenvalue controls size and has power strictly higher than when conditioning on the largest eigenvalue.