EconBase
← Back to paper

A Powerful Subvector Anderson Rubin Test in Linear Instrumental Variables Regression with Conditional Heteroskedasticity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,345 characters · 13 sections · 74 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Powerful Subvector Anderson-Rubin Test in Linear Instrumental Variables Regression with Conditional Heteroskedasticity

array[array omitted — 118 chars of source]

$ \and $

array[array omitted — 122 chars of source]

$ \and \vspace{0.25in}$

array[array omitted — 110 chars of source]

$}

abstractWe introduce a new test for a two-sided hypothesis involving a subset of the structural parameter vector in the linear instrumental variables (IVs) model. GuggenbergerKleibergenMavroeidis2019, GKM19 from now on, introduce a subvector Anderson-Rubin (AR) test with data-dependent critical values that has asymptotic size equal to nominal size for a parameter space that allows for arbitrary strength or weakness of the IVs and has uniformly nonsmaller power than the projected AR test studied in GuggenbergerKleibergenMavroeidisChen2012. However, GKM19 imposes the restrictive assumption of conditional homoskedasticity. The main contribution here is to robustify the procedure in GKM19 to arbitrary forms of conditional heteroskedasticity. We first adapt the method in GKM19 to a setup where a certain covariance matrix has an approximate Kronecker product (AKP) structure which nests conditional homoskedasticity. The new test equals this adaption when the data is consistent with AKP structure as decided by a model selection procedure. Otherwise the test equals the AR/AR test in Andrews2017 that is fully robust to conditional heteroskedasticity but less powerful than the adapted method. We show theoretically that the new test has asymptotic size bounded by the nominal size and document improved power relative to the AR/AR test in a wide array of Monte Carlo simulations when the covariance matrix is not too far from AKP. Keywords: Asymptotic size, conditional heteroskedasticity, Kronecker product, linear IV regression, subvector inference, weak instruments JEL codes: C12, C26

Introduction

Robust and powerful subvector inference constitutes an important problem in Econometrics. For instance, it is standard practice to report confidence intervals on each of the coefficients in a linear regression model. By robust we mean a testing procedure for a hypothesis of (or a confidence region for) a subset of the structural parameter vector such that the asymptotic size is bounded by the nominal size for a parameter space that allows for weak or partial identification. Recent contributions to robust subvector inference have been made in the context of the linear instrumental variables (IVs from now on) model (see, for example, DufourTaamouti2005, GuggenbergerKleibergenMavroeidisChen2012 (GKMC from now on), GuggenbergerKleibergenMavroeidis2019 (GKM19 from now on), and Kleibergen2021), GMM models (see, for example, ChaudhuriZivot2011, AndrewsCheng2014, AndrewsMikusheva2016, Andrews2017, and HanMcCloskey2019), and also models defined by moment (in)equalities (see, for example, BugniCanayShi2017, Gafarov2017, and KaidoMolinariStoye2019). GKM19 introduce a new subvector test that compares the AR subvector statistic to conditional critical values that adapt to the strength or weakness of identification and verify that the resulting test has correct asymptotic size for a parameter space that imposes conditional homoskedasticity (CHOM from now on) and uniformly improves on the power of the projected AR test studied in DufourTaamouti2005.

The contribution of the current paper is to provide a robust subvector test that improves the power of another robust subvector test by combining it with a more powerful test that is robust for only a smaller parameter space. More specifically, in the context of the linear IV model, we first provide a modification of the subvector AR test of GKM19, called the AR$_{AKP,\alpha}$ test, where $\alpha$ denotes the nominal size. We verify that it has correct asymptotic size for a parameter space that nests the setup with CHOM and also allows for particular cases of conditional heteroskedasticity (CHET from now on), namely setups where a particular covariance matrix has a Kronecker product (KP from now on) structure. For example, the data generating process (DGP from now on) has a KP structure if the vector of structural and reduced-form errors equals a random vector independent of the IVs times a scalar function of the IVs. In particular then, the variances of all the errors depend on the IVs by the same multiplicative constant given as a scalar function of the IVs. In the companion paper GuggenbergerKleibergenMavroeidis2022 (GKM22 from now on) we find that KP structure is not rejected at the 5% nominal size in more than 63% of empirical data sets we studied of several recently published empirical papers (namely, 38 of 60 specifications are not rejected; and, including cases with clustering, 56 out of 118 are not rejected). For comparison, CHOM is rejected for 57 of the 60 specifications considered at the 5% nominal size, using the test in Kelejian (1982). Of course, these findings do not prove that empirical data sets do have KP\ structure as the low number of rejections of KP structure may be due to type II errors of the test. However, coupled with the quite favorable finite sample power results of the test of KP structure reported in GKM22 (for sample sizes of $n=200$) we believe that KP\ structure might be compatible with a sizable subset of empirical data sets.

Second, depending on a model selection mechanism that determines whether the data are compatible with KP, the recommended test then equals the AR$_{AKP,\alpha}$ test or the AR/AR test in Andrews2017 that is robust to arbitrary forms of CHET. We show that the recommended test has correct asymptotic size. An important ingredient in establishing that is showing that the AR$_{AKP,\alpha}$ test does not reject less often under the null hypothesis than the AR/AR test when the data are close to KP structure.

We propose two different model selection methods. One is based on the KPST test statistic introduced in GKM22 for testing the null hypothesis that a covariance matrix has KP structure. The other one is based on the standardized norm of the distance between the covariance matrix estimator and its closest KP approximation. As in the model selection method proposed in AndrewsSoares2010, we compare the test statistic to a user chosen threshold that, in the asymptotics, is let go to infinity. The thresholds can be chosen differently depending on the number of IVs $k$ and parameters not under test. Based on comprehensive finite sample simulations we provide choices for the thresholds for several values of $k$ that lead to good control of the finite sample size.

As the main contribution of the paper, we verify that the resulting test, called $\varphi_{MS-AKP,\alpha}$ test, has asymptotic size bounded by the nominal size $\alpha$ under certain conditions on the selection mechanism and implementation of the AR/AR test at nominal size $\alpha-\delta$ for some arbitrarily small $\delta>0$.

In a Monte Carlo study, we compare the suggested new test $\varphi _{MS-AKP,\alpha}$ with several alternatives given in Andrews2017, in particular, the AR/AR and the AR/QLR1 tests. Andrews2017 fills a very important gap in the literature on subvector inference by providing two-step Bonferroni-like methods\footnote{See McCloskey2017 for a general reference on Bonferroni methods in nonstandard testing setups.} for a rich class of models that nests GMM, that i) control the asymptotic size under relatively mild high-level conditions that allow for CHET, ii) are asymptotically non-conservative (in contrast to standard Bonferroni methods) and iii) for the case of AR/QLR1 is asymptotically efficient under strong identification (while the AR/AR test is not asymptotically efficient under strong identification in overidentified situations). In contrast, the test considered here, $\varphi_{MS-AKP,\alpha}$, can only be used in the linear IV model and is not asymptotically efficient under strong identification. The Monte Carlo study finds that $\varphi_{MS-AKP,\alpha}$ has uniformly higher rejection probabilities than the AR/AR test for all the DGPs considered. That includes the null rejection probabilities (NRPs from now on) with the $\varphi _{MS-AKP,\alpha}$ test having finite sample size of 6% versus the 5.4% of the AR/AR test at nominal size 5%. Based on the Monte Carlo study we conclude that relative to the AR/QLR1 test, $\varphi_{MS-AKP,\alpha}$ can be a useful alternative in terms of power in situations of weak or mixed identification strengths when the degree of overidentification is small and the covariance matrix of the data is not too far from KP structure. Whenever the data are compatible with KP structure, it also offers an important computational advantage because the AR$_{AKP,\alpha}$ test is given in closed form. In contrast, implementation of the two-step Bonferroni-like methods require minimization of a statistic over a set that has dimension equal to the number of parameters not under test. The computation time should grow exponentially in the dimension of that set which constitutes a computational challenge especially when an applied researcher uses the proposed methods for the construction of a confidence region by test inversion. This being said, an applied researcher who uses the $\varphi_{MS-AKP,\alpha}$ test has to be ready to implement the AR/AR test in case it is determined that KP structure is not compatible with the data. Given the construction of the AR$_{AKP,\alpha}$ test it is not surprising to find the relative best performance of the $\varphi_{MS-AKP,\alpha}$ test to occur under weak identification. Namely, the critical values of the former test adapt to the strength of identification and can be substantially lower than the corresponding chi-square critical values when identification is deemed to be weak.

The rest of the paper is organized as follows. In Section (ref) we introduce a version of a subvector AndersonRubin1949 test that has correct asymptotic size for a parameter space that imposes an approximate Kronecker product (AKP) structure for the covariance matrix. In Section (ref) we introduce the new test that has correct asymptotic size for a parameter space that does not impose any structure on the covariance matrix and therefore, in particular, allows for arbitrary forms of conditional heteroskedasticity. Finally, in Section (ref) we study the finite sample properties of the test. Proofs are given in the Appendix at the end.

Notation: Throughout the paper, we denote by \textquotedblleft$\otimes $\textquotedblright\ the KP of two matrices, by $vec(\cdot)$ the column vectorization of a matrix, and by $||\cdot||$ the Frobenius norm.\footnote{Recall the Frobenius norm for a matrix $A=(a_{ij})\in \Re^{m\times n}$ is defined as $||A||^{2}:= {\textstyle\sum\nolimits_{i=1}^{m}} {\textstyle\sum\nolimits_{j=1}^{n}} a_{ij}^{2}$. When $A$ is a vector the Frobenius and the Euclidean norm are numerically equivalent.} We use the notation $M_{A}:=I_{n}-P_{A}$ and $P_{A}:=A(A^{\prime}A)^{-1}A^{\prime}$ for any full rank matrix $A\in \Re^{n\times k}.$

Subvector AR Test under Approximate Kronecker Product Structure

Assume the linear IV model is given by the equations

align[align omitted — 245 chars of source]

where $y\in\Re^{n},$ $Y\in\Re^{n\times m_{Y}},$ $W\in\Re^{n\times m_{W}},$ and $\overline{Z}\in\Re^{n\times k}.$ Here, $W$ contains endogenous regressors, while the regressors $Y$ may be endogenous or exogenous. We assume that $k-m_{W}\geq1$ and $m_{W}\geq1.$ The reduced form can be written as

equation[equation omitted — 403 chars of source]

where $v_{y}:=V_{Y}\beta+$ $V_{W}\gamma+\varepsilon$ (which depends on the true $\beta$ and $\gamma),$ $V_{W}^{\prime}=(V_{W,1},\ldots,V_{W,n}),$ $V_{Y}^{\prime}=(V_{Y,1},\ldots,V_{Y,n}),$ $\overline{Z}^{\prime} =(\overline{Z}_{1},\ldots,\overline{Z}_{n}).$ By $V_{i},$ for $i=1,...,n,$ we denote the $i$-th row of $V$ written as a column vector and similarly for other matrices.

The objective is to test the subvector hypothesis

equation[equation omitted — 102 chars of source]

using tests whose size, i.e. the highest NRP over a large class of distributions for $(\varepsilon_{i},\overline{Z}_{i}^{\prime},V_{Y,i}^{\prime },V_{W,i}^{\prime})$ and the unrestricted nuisance parameters $\Pi_{Y},$ $\Pi_{W},$ and $\gamma$, equals the nominal size $\alpha$, at least asymptotically. In particular, weak identification and non-identification of $\beta$ and $\gamma$ are allowed for. The setup allows testing the coefficients of exogenous or endogenous regressors $Y$ in the presence of endogenous regressors $W$. We impose the following assumption as in GKM19 (from where the name of the assumption is inherited).

\noindentAssumption B: The random vectors $(\varepsilon_{i} ,\overline{Z}_{i}^{\prime},V_{Y,i}^{\prime},V_{W,i}^{\prime})$ for $i=1,...,n$ in ((ref)) are i.i.d. with distribution $F.$

For a given sequence $a_{n}=o(1)$ in $\Re_{\geq0},$ we define a sequence of parameter spaces $\mathcal{F}_{AKP,a_{n}}$ for $(\gamma,\Pi_{W},\Pi_{Y},F)$ under the null hypothesis $H_{0}:\beta=\beta_{0}$ that is larger than the corresponding ones in GKMC and GKM19 in that general forms of AKP structures for the variance matrix

equation[equation omitted — 159 chars of source]

are allowed for.\footnote{Regarding the notation $(\gamma,\Pi_{W},\Pi_{Y},F)$ and elsewhere, note that we allow as components of a vector column vectors, matrices (of different dimensions), and distributions.} Namely, for $U_{i}:=(\varepsilon_{i}+V_{W,i}^{\prime}\gamma,V_{W,i}^{\prime})^{\prime}$ (which equals $(v_{yi}-V_{Y,i}^{\prime}\beta,V_{W,i}^{\prime})^{\prime}$), $p:=1+m_{W},$ and $m:=m_{Y}+m_{W}$ let

align[align omitted — 713 chars of source]

for symmetric matrices $\Upsilon_{n}\in\Re^{kp\times kp}$ such that

equation[equation omitted — 63 chars of source]

positive definite (pd from now on) symmetric matrices $G_{F}\in\Re^{p\times p}$ (whose upper left element is normalized to 1) and $\overline{H}_{F}\in \Re^{k\times k},$ $\delta_{1},\delta_{2}>0,$ $B<\infty$. Note that the factors in the KP $G_{F}\otimes\overline{H}_{F}$ are not uniquely defined due to the summand $\Upsilon_{n}$. Note that no restriction is imposed on the variance matrix of $vec(\overline{Z}_{i}V_{Y,i}^{\prime})$ and, in particular, $E_{F}(vec(\overline{Z}_{i}V_{Y,i}^{\prime})(vec(\overline{Z}_{i} V_{Y,i}^{\prime}))^{\prime})$ does not need to factor into a KP.

The factorization of the covariance matrix into an AKP in line three of ((ref)) is a weaker assumption than CHOM. Under CHOM, we have $G_{F}=E_{F}\left( U_{i}U_{i}^{\prime}\right) $ and $\overline{H}_{F} =E_{F}(\overline{Z}_{i}^{\prime}\overline{Z}_{i})$ (prior to the normalization of the upper left element of $G_{F}$) and $\Upsilon_{n}=0^{kp\times kp}.$ The AKP structure allowed for here (but not in GKMC and GKM19) also covers some important cases of CHET involving $vec(\overline{Z}_{i}U_{i}^{\prime})$.

Examples. i)\ Consider the case in ((ref)) where $(\widetilde{\varepsilon}_{i},\widetilde{V}_{W,i}^{\prime})^{\prime}\in\Re ^{p}$ are i.i.d. zero mean with a pd variance matrix, independent of $\overline{Z}_{i},$ and $(\varepsilon_{i},V_{W,i}^{\prime})^{\prime }:=f(\overline{Z}_{i})(\widetilde{\varepsilon}_{i},\widetilde{V}_{W,i} ^{\prime})^{\prime}$ for some scalar valued function $f$ of $\overline{Z} _{i}.$\footnote{For example, Andrews2017 considers $f(Z_{i} )=||Z_{i}||/k^{1/2}.$} In that case, the covariance matrix $\overline{R}_{F}$ can be written

align[align omitted — 878 chars of source]

and thus has KP structure even though, obviously, CHOM is not satisfied because

equation[equation omitted — 310 chars of source]

depends on $\overline{Z}_{i}.$

We can construct illustrative examples where the proportionality $(\varepsilon_{i},V_{W,i}^{\prime})^{\prime }:=f(\overline{Z}_{i})(\widetilde{\varepsilon}_{i},\widetilde{V}_{W,i} ^{\prime})^{\prime}$ (that would imply KP\ structure) holds. Consider e.g. the model

align*[align* omitted — 90 chars of source]

where the covariates $Y_{i}$ and $X_{i}$ are exogenous, the variables $ y_{i},W_{i}$ are endogenous, and $Y_{i}$ has heterogeneous causal effects on $y_{i},W_{i},$ denoted by $\beta _{i},\phi _{i},$ respectively. Let $\beta :=E\left( \beta _{i}\right) ,$ $\phi :=E\left( \phi _{i}\right) ,$ define $ \widetilde{\varepsilon }_{i}:=\beta _{i}-\beta $ and $\widetilde{V} _{Wi}:=\phi _{i}-\phi ,$ and assume that $\widetilde{\varepsilon }_{i}, \widetilde{V}_{Wi}$ are orthogonal to $Z_{i}:=\left( Y_{i},X_{i}\right) .$ Then, this fits exactly into the setup above with $f\left( Z_{i}\right) =Y_{i}.$ In other words, KP structure can result as a special case of heterogeneous causal effects.

ii) In a wage regression to assess the effect of "years of education", the assumption of CHOM would require that e.g. the variance of "wage" does not depend on the included regressor "race". This assumption is incompatible with recent US data where the wage dispersion is largest for Asians. Instead, the construction $(\varepsilon_{i},V_{W,i}^{\prime})^{\prime }:=f(\overline{Z}_{i})(\widetilde{\varepsilon}_{i},\widetilde{V}_{W,i} ^{\prime})^{\prime}$ in i) allows for dependence of the variances of the regressand and all endogenous regressors on a scalar function of $\overline {Z}_{i}.$ The maintained restriction is that all these variances are affected approximately by the same scalar function of $\overline{Z}_{i}.$ In the related paper, GKM22, we test the null hypothesis of KP structure for 118 specifications in about a dozen highly cited papers and find that at the 5% nominal size in 47.5% of the cases the null is not rejected.

In this section we will introduce a new conditional subvector AR$_{AKP}$ test and show it has asymptotic size with respect to the parameter space $\mathcal{F}_{AKP,a_{n}}$ equal to the nominal size. We next define the new test statistic and the critical value for the case considered here of AKP\ structure.

Estimation of the two factors in the AKP structure: Define

equation[equation omitted — 122 chars of source]

and $Z\in\Re^{n\times k}$ with rows given by $Z_{i}^{\prime}$ for $i=1,...,n.\footnote{For simplicity, we do not use the more precise notation $Z_{in}$ for $Z_{i}.$ It is explained in detail in Comment 3 below Theorem \ref{correct asy size} why we introduce $Z_{i},$ namely to obtain invariance of the testing procedure with respect to nonsingular transformations of the IVs.}$ Define an estimator of the matrix

equation[equation omitted — 213 chars of source]

by

align[align omitted — 295 chars of source]

Note that $\widehat{R}_n$ is automatically a centered estimator because, as straightforward calculations show, $n^{-1} {\textstyle\sum\nolimits_i} f_i=0.$ From $\overline{R}_F=G_F\otimes\overline{H}_F+\Upsilon_n,$ it follows that $R_F=G_F\otimes H_F+o(1)$ for

equation[equation omitted — 164 chars of source]

Let

equation[equation omitted — 105 chars of source]

where the minimum is taken over $(G,H)$ for $G\in\Re^p\times p,$ $H\in \Re^k\times k$ being pd, symmetric matrices, and normalized such that the upper left element of $G$ equals 1.

It can be shown that $(\widehat{G}_{n},\widehat{H}_{n})$ are given in closed form by the following construction.\footnote{This follows from a combination of Lemma (ref) below and Theorem 5.8 in VanLoanPitsianis1993.} First, for a pd matrix $A\in \Re^{kp\times kp}$ define the rearrangement of $A$ as

align[align omitted — 328 chars of source]

where $A_{lj}\in\Re^{k\times k}$ denotes the $(l,j)$ submatrix of dimensions $k\times k,$ where $l,j=1,...,p.$ Second, denote by

equation[equation omitted — 125 chars of source]

a singular value decomposition of $\mathcal{R}(A)$,\footnote{In VanLoanPitsianis1993, the orthogonal matrices $\widehat{L}\in\Re^{pp\times pp}$ and $\widehat{N}\in\Re^{kk\times kk}$ are called $U$ and $V,$ respectively, notation that we have already used for other objects.} where the singular values $\widehat{\sigma}_{l}$ for $l=1,...,p^{2}$ are ordered non-increasingly. Finally, denote by $\widehat{L}(:,1)$ and $\widehat{N}(:,1)$ singular vectors corresponding to the largest singular value $\widehat{\sigma}_{1}$ and let $\widehat{L}(1,1)$ denote the first component of $\widehat{L}(:,1).$ Then, letting the role of $A$ be played by $\widehat{R}_{n}$ in ((ref)), minimizers $(\widehat{G}_{n} ,\widehat{H}_{n})$ to ((ref)) are defined by

equation[equation omitted — 182 chars of source]

where $\widehat{L}(1,1)>0$ whenever $\widehat{R}_{n}$ is pd. By Lemma (ref) below, the definition given in ((ref)) is unique for all large enough $n$ wp1\footnote{Note that it would not be unique if the eigenspace associated with the largest singular value had dimension larger than 1.} and

equation[equation omitted — 159 chars of source]

under certain sequences $F_{n}$ as defined in $\mathcal{F}_{AKP,a_{n}}$ for which $R_{F_{n}}=G_{F_{n}}\otimes H_{F_{n}}+o(1)$ (where $R_{F_{n}}$ is defined in ((ref)) with $F$ replaced by $F_{n}$), $H_{F_{n}} :=(E_{F_{n}}\overline{Z}_{i}\overline{Z}_{i}^{\prime})^{-1/2}\overline {H}_{F_{n}}(E_{F_{n}}\overline{Z}_{i}\overline{Z}_{i}^{\prime})^{-1/2}$ (as defined in ((ref))), and the upper left element of $G_{F_{n}}$ is normalized to 1.

Definition of the conditional subvector test:\ We denote the subvector AR statistic when the variance matrix has AKP structure by $AR_{AKP,n}(\beta_{0})$ and define it as the smallest root $\hat{\kappa}_{pn}$ of the roots $\hat{\kappa}_{in},$ $i=1,...,p$ (ordered nonincreasingly) of the characteristic polynomial

equation[equation omitted — 250 chars of source]

The conditional subvector test AR$_{AKP,\alpha}$ rejects $H_{0}$ at nominal size $\alpha$ if

equation[equation omitted — 113 chars of source]

where $c_{1-\alpha}\left( \cdot,\cdot\right) $ is defined as follows. Muirhead1978, in the case where $m_{W}=1$ and assuming normality, provides an approximate, nuisance parameter free, conditional density of the smaller eigenvalue $\hat{\kappa}_{2n}$ given the larger one $\hat{\kappa} _{1n}$ for any degree of overidentification $k-m_{W},$ see (2.12) in GKM19 for the conditional pdf$.$ For given $\hat{\kappa}_{1n}$ and arbitrary $m_{W},$ $c_{1-\alpha}(\hat{\kappa}_{1n},k-m_{W})$ denotes the $1-\alpha$-quantile of that approximation. GKM19 (Table 1 and Supplement C) provide $c_{1-\alpha }(\hat{\kappa}_{1n},k-m_{W})$ for $\alpha=1,5,10\%,$ $k-m_{W}=1,...,20$ and a fine grid of values for $\hat{\kappa}_{1n},$ say $\hat{\kappa}_{1,1} \leq...\leq\hat{\kappa}_{1,j}\leq...\leq\hat{\kappa}_{1,J}$ for some large $J.$ We reproduce Table 1 (that covers the case $\alpha=5\%$ and $k-m_{W}=4)$ from GKM19 below. Conditional critical values for values of $\hat{\kappa} _{1n}$ not reported in the tables are obtained by linear interpolation. Specifically, let $q_{1-\alpha,j}(k-1)$ denote the $1-\alpha$ quantile of the distribution whose density is given by (2.12) in GKM19 with $\hat{\kappa} _{1n}$ replaced by $\hat{\kappa}_{1,j}$. The end point of the grid $\hat{\kappa}_{1,J}$ should be chosen high enough so that $q_{1-\alpha ,J}(k-m_{W})\approx\chi_{k-m_{W},1-\alpha}^{2}$. For any realization of $\hat{\kappa}_{1n}\leq\hat{\kappa}_{1,J}$, find $j$ such that $\hat{\kappa }_{1n}\in\left[ \hat{\kappa}_{1,j-1},\hat{\kappa}_{1,j}\right] $ with $\hat{\kappa}_{1,0}=0$ and $q_{1-\alpha,0}\left( k-m_{W}\right) =0$, and let

equation[equation omitted — 351 chars of source]

Table 1: $cv=c_{1-\alpha}(\hat{\kappa}_{1},k-m_{W})$ for $\alpha=5\%,$ $k-m_{W}=4$ for various values of $\hat{\kappa}_{1}$

{

tabular[tabular omitted — 894 chars of source]

}

Denote by $P_{(\gamma,\Pi_{W},\Pi_{Y},F)}(\cdot)$ the probability of an event under the null hypothesis when the true values of the structural and reduced form parameters and the distribution of the random variables are given by $(\gamma,\Pi_{W},\Pi_{Y},F).$ Recall the definition of the parameter space $\mathcal{F}_{AKP,a_{n}}$ in ((ref)). We can now formulate the main result of this section.

theoremUnder Assumption B, the conditional subvector test AR$_{AKP,\alpha}$ defined in $($(ref)$)$ implemented at nominal size $\alpha$ has asymptotic size, i.e. \[ \lim\sup_{n\rightarrow\infty}\sup_{(\gamma,\Pi_{W},\Pi_{Y},F)\in \mathcal{F}_{AKP,a_{n}}}P_{(\gamma,\Pi_{W},\Pi_{Y},F)}(AR_{AKP,n}(\beta _{0})>c_{1-\alpha}(\hat{\kappa}_{1n},k-m_{W})) \] equal to $\alpha$ for $\alpha\in\left\{ 1\%,5\%,10\%\right\} $ and $k-m_{W}\in\left\{ 1,...,20\right\} .$

Comment. 1. The conditional subvector test AR$_{AKP,\alpha}$ adapts the test in GKM19 from a setup under CHOM to AKP structure. The modification involves replacing the matrices $\left( \overline{Y}_{0},W\right) ^{\prime }M_{\overline{Z}}\allowbreak \left( \overline{Y}_{0},W\right) /(n-k)$ and $n^{-1} \overline{Z}^{\prime}\overline{Z}$ in GKM19 by the matrices $\widehat{G}_{n}$ and $\widehat{H}_{n},$ respectively, in ((ref)) to account for the more general structure of the covariance matrix. Some portions of the proof follow similar steps as the proof of Theorem 5 in GKM19. In particular, one portion of the proof relies on a one-dimensional simulation exercise to prove that the NRPs are bounded by the nominal size. This exercise could be extended to choices of $\alpha$ and $k-m_{W}$ beyond those in the theorem and likely the theorem would extend to many more choices.

2. Trivially, under the same assumptions as in Theorem (ref), we obtain that \[ \lim\sup_{n\rightarrow\infty}\sup_{(\gamma,\Pi_{W},\Pi_{Y},F)\in \mathcal{F}_{AKP,a_{n}}}P_{(\gamma,\Pi_{W},\Pi_{Y},F)}(AR_{AKP,n}(\beta _{0})>\chi_{k-m_{W},1-\alpha}^{2})=\alpha. \] That is, the generalization of the subvector test in GKMC to AKP structure has correct asymptotic size. This result is obtained fully analytically; its proof does not require any simulations.

3. Invariance with respect to nonsingular transformations of the IV matrix. The identifying power of the model comes from the moment condition $E_{F}\varepsilon_{i}\overline{Z}_{i}=E_{F}(y_{i}-Y_{i}^{\prime}\beta -W_{i}^{\prime}\gamma)\overline{Z}_{i}=0.$ This moment condition obviously still holds when the instrument vector is premultiplied by a nonrandom nonsingular matrix $A\in\Re^{k\times k},$ i.e. $E_{F}\varepsilon_{i} A\overline{Z}_{i}=0.$ It then seems reasonable to look for testing procedures whose outcome is invariant to such nonsingular transformations. In the weak IV literature, e.g. AndrewsMoreiraStock06 and AndrewsMarmerYu2019 and references therein, the class of (similar) invariant tests to orthogonal transformations $A$, that is, changes of the coordinate system, has been studied. The transformation of the IVs in ((ref)) is performed in order for the test to be invariant to nonsingular transformations of the IVs.

If the conditional subvector AR$_{AKP}$ test defined in ((ref)) (and $\widehat{R}_{n}$ in ((ref))) was defined with $\overline{Z}_{i}$ in place of $Z_{i}$ it would be invariant to orthogonal transformations but not necessarily to nonsingular ones. To see the former, denote by $\widehat{R}_{nA}$ the matrix $\widehat{R}_{n}$ when the instrument vector has been transformed to $A\overline{Z}_{i}$ (and consequently $\overline{Z}$ is changed to $\overline{Z}A^{\prime}).$ Then the claim follows from $\mathcal{R}(\widehat{R}_{nA})=\mathcal{R}(\widehat{R} _{n})(A^{\prime}\otimes A^{\prime})$ (which holds for any nonsingular matrix $A$ by straightforward calculations using $vec(ABC)=(C^{\prime}\otimes A)vec(B)$ for any conformable matrices $A,B,$ and $C$ and $M_{\overline{Z} }=M_{\overline{Z}A^{\prime}})$ which implies $\widehat{G}_{nA}=\widehat{G} _{n}$ and $\widehat{H}_{nA}=A\widehat{H}_{n}A^{\prime}$ when $A$ is orthogonal, where again $\widehat{G}_{nA}$ and $\widehat{H}_{nA}$ denote the matrices $\widehat{G}_{n}$ and $\widehat{H}_{n}$ when the instrument vector $\overline{Z}_{i}$ has been transformed to $A\overline{Z}_{i}$. It then follows that the matrix $n^{-1}\widehat{G}_{n}^{-1/2}\left( \overline{Y} _{0},W\right) ^{\prime}\overline{Z}\widehat{H}_{n}^{-1}\overline{Z}^{\prime }\left( \overline{Y}_{0},W\right) \widehat{G}_{n}^{-1/2}$ in ((ref)) (and thus its eigenvalues) remain invariant under orthogonal transformations $\overline{Z}_{i}\rightarrow A\overline{Z}_{i}$ of the instrument matrix. This test however is not invariant in general to arbitrary nonsingular transformations.

But with the replacement of $\overline{Z}_{i}$ by $Z_{i}$ as done in ((ref)) and, correspondingly, $\overline{Z}$ by $\overline{Z} (n^{-1}\overline{Z}^{\prime}\overline{Z})^{-1/2}$ in ((ref)), the test is invariant against nonsingular transformations $A$. The invariance of this test to arbitrary nonsingular transformations $\overline{Z}_{i}\rightarrow A\overline{Z}_{i}$ of the instrument matrix (which leads to a transformation of $Z_{i}$ to $(A\overline{Z}^{\prime}\overline{Z}A^{\prime})^{-1/2}A\overline{Z}_{i}$) follows from straightforward calculations and the fact that the matrix

equation[equation omitted — 157 chars of source]

is orthogonal. In particular, one can easily show that the matrices $\mathcal{R}(\widehat{R}_{n}),$ $\widehat{G}_{n},$ and $\widehat{H}_{n}$ that appear as ingredients in the conditional subvector test AR$_{AKP,\alpha}$ with $A=I_{k}$ are related to the corresponding matrices $\mathcal{R} (\widehat{R}_{nA}),$ $\widehat{G}_{nA},$ and $\widehat{H}_{nA}$, when $A$ is an arbitrary nonsingular matrix, via

equation[equation omitted — 245 chars of source]

which immediately implies the desired invariance result.

4. The conditional subvector test can be generalized to a stationary time series setting, see the Appendix, Section (ref), for details. In the context of a time series setting we offer another example of AKP structure. Namely, consider a structural vector autoregression $AX_{t}=BX_{t-1}+\eta_{t},$ where $\dim X_{t}=\dim\eta_{t}=n,$ $E\left( \eta_{t}|X_{t-1}\right) =0$ and suppose that $var\left( \eta_{t} |X_{t-1}\right) =\allowbreak var\left( \eta_{t}\right) =\allowbreak \Sigma_{t}=\allowbreak diag\left( \sigma_{1t}^{2},...,\sigma_{nt}^{2}\right) $. If $\sigma_{it}^{2}=a_{t}\sigma_{i}^{2}$ for some scalar function of time $a_{t},$ i.e., the volatilities of all the shocks change over time in a proportional manner, then the variance of $X_{t-1}\eta_{t}$ has KP structure. In this model, identification can be achieved by exclusion restrictions Sims80 that render some of $X_{t-1}$ valid instruments. It can also be achieved with external instruments if available StockWatson2018. Time-variation in volatilities has been reported in many contexts. For instance, the `great moderation' is a well-documented phenomenon of a fall in macroeconomic volatility in the US in the early 1980s (cf. Bernanke2004, ch. 4). AKP would result if the fall in the volatilities were similar across variables.

5. Note that under the null hypothesis the test does not depend on the value of the reduced form matrix $\Pi_{Y}$ because the test statistic and the critical value are affected by $Y$ only through $\overline{Y}_{0}=y-Y\beta _{0}.$

6. GKM19 establish that the conditional subvector AR test introduced there enjoys near optimality properties in the linear IV\ model with conditional homoskedasticity in a certain class of tests that depend on the data only through the roots $\hat{\kappa}_{in},$ $i=1,...,p$ when $k-m_{W}=1.$ On the other hand, when $k-m_{W}$ gets bigger the test may be quite conservative. The power gains over the projected AR subvector test discussed in DufourTaamouti2005 arise in weakly identified scenarios while under strong identification these two tests become identical. Similarly, we expect the power properties of the new conditional subvector test AR$_{AKP,\alpha}$ to be most competitive for small $k-m_{W},$ in particular, when $k-m_{W}=1,$ in weakly identified situations.

Intuition behind the result derived in GKM19 that conditioning on the largest eigenvalue when $m_{W}>1$ leads to a test with correct size is based on i) the corresponding result for $m_{W}=1$ and ii) the so-called "inclusion principle" that provides a ranking of the corresponding eigenvalues of a Hermitian matrix and its principal submatrices (see GKM19 bottom p.499-500, in particular eq (2.23)).

Subvector Testing under Arbitrary Forms of Conditional Heteroskedasticity

We now allow for arbitrary forms of CHET, that is, the parameter space does not impose an AKP structure for $\overline{R}_{F}$. We describe a testing procedure under high level assumptions that we then verify in the next subsections for particular implementations of the test. In particular, Lemma (ref) below verifies Assumptions RT and RP below for a particular implementation of the AR/AR test.

In what follows, $\mathcal{F}_{Het}$ is a generic parameter space for $(\gamma,\Pi_{W},\Pi_{Y},F)$ that does not impose an AKP structure, but if the restriction $\overline{R}_{F}=G_{F}\otimes\overline{H}_{F}+\Upsilon_{n}$ as in $\mathcal{F}_{AKP,a_{n}}$ in ((ref)) was added to the conditions in $\mathcal{F}_{Het}$ then $\mathcal{F}_{Het}\subset\mathcal{F}_{AKP,a_{n}}.$ For example, the null parameter space $\mathcal{F}_{Het}$ may impose stronger moment conditions than $\mathcal{F}_{AKP,a_{n}}$ so that certain Lyapunov CLTs apply. See the definitions of $\mathcal{F}_{Het}$ in the next subsections. We summarize the restrictions on the parameter space (PS) in the following assumption.

Assumption PS: $\mathcal{F}_{Het}\subset\widetilde{\mathcal{F} }_{AKP,a_{n}},$ where $\widetilde{\mathcal{F}}_{AKP,a_{n}}$ is equal to $\mathcal{F}_{AKP,a_{n}}$ without the condition $\overline{R}_{F}=G_{F} \otimes\overline{H}_{F}+\Upsilon_{n}$ (AKP structure) and without the assumptions $\kappa_{\min}(A)\geq\delta_{2}$ for $A\in\{G_{F},\overline{H} _{F}\}.$

We assume there exists a robust test (RT) $\varphi_{Rob,\alpha}$ that has asymptotic size for the parameter space $\mathcal{F}_{Het}$ bounded by the nominal size $\alpha$. For example, in the next subsection we consider a particular implementation of the AR/AR test in Andrews2017. In general, we think of $\varphi_{Rob,\alpha}$ as a test whose power can be substantially improved on by the test $\varphi_{AKP,\alpha}$ when $\overline{R}_{F}$ has AKP structure.

\noindentAssumption RT: The test $\varphi_{Rob,\alpha}$ of ((ref)) has asymptotic size bounded by the nominal size $\alpha$ for the parameter space $\mathcal{F}_{Het}$.

We now define a new test that, roughly speaking, coincides with $\varphi _{AKP,\alpha}$ or $\varphi_{Rob,\alpha}$ depending on whether the data seems consistent or not with AKP structures. We now provide the details.

Consider a given sequence of constants $c_{n}$ such that

equation[equation omitted — 95 chars of source]

e.g. $c_{n}=cn^{1/2}/\ln(n)$ or $c_{n}=cn^{1/2}/\ln\ln(n)$ for some constant $c>0$ and define

equation[equation omitted — 116 chars of source]

where the minimum (here and in analogous expressions below) is taken over $(G,H)$ for $G\in\Re^{p\times p},$ $H\in\Re^{k\times k}$ being pd, symmetric matrices, normalized such that the upper left element of $G$ equals 1.\footnote{The expression $G\otimes H-R_{F_{n}}$ is pre- and postmultiplied by $R_{F_{n}}^{-1/2}$ for invariance reasons.} The quantity $\lambda_{9n}$ measures how far from KP structure the covariance matrix $R_{F_{n}}$ in ((ref)) when $F=F_{n}$ is. To show that the new test $\varphi _{MS-AKP,\alpha}$ defined below has asymptotic significance level $\alpha$, it is sufficient (as proven in the Appendix) to consider two types of drifting sequences of DGPs in $\mathcal{F}_{Het}$ and to establish that the test has limiting NRP bounded by the nominal size $\alpha$ in each case. The first type of sequences are those for which

equation[equation omitted — 80 chars of source]

that is sequences where the covariance matrix $R_{F_{n}}$ is "far away" from KP structure. We assume that there is a model selection (MS) method $\varphi_{MS,c_{n}}\in\{0,1\}$ such that when $R_{F_{n}}$ is "far from" KP structure\ it will choose the robust test wpa1. The next assumption makes that statement more precise. To properly formulate the assumption we require terminology that is provided in the Appendix because it requires a lot of space. In particular, we need to consider particular sequences of drifting parameters $\lambda_{w_{n},h}$ (defined in ((ref)) in the Appendix) where $w_{n}$ denotes a subsequence of $n$.

\noindentAssumption MS: The model selection method $\varphi _{MS,c_{n}}\in\{0,1\}$ satisfies $\varphi_{MS,c_{n}}=1$ wpa1 under parameter sequences $\lambda_{w_{n},h}$ (with underlying parameter space $\mathcal{F} _{Het})$ with $h_{9}=\infty$.

By definition, along $\lambda_{w_{n},h},$ $w_{n}^{1/2}\lambda_{9w_{n} }\rightarrow h_{9}$ and thus when $h_{9}=\infty$ the sequence is not local to KP structure.

Definition of the fully robust test: Let $\delta\geq0.$ The new suggested test $\varphi_{MS-AKP,\delta,c_{n},\alpha}$ of nominal size $\alpha$ of the null hypothesis ((ref)) is defined as

equation[equation omitted — 123 chars of source]

We typically write $\varphi_{MS-AKP,\alpha}$ rather than $\varphi _{MS-AKP,\delta,c_{n},\alpha}$ to simplify notation. Ideally, $\delta=0$ can be chosen in this construction. To verify Assumption RP below using the AR/AR test as $\varphi_{Rob,\alpha-\delta}$ we need to have $\delta>0.$ (Potentially, Assumption RP may hold with $\delta=0$ but our current proof technique does not allow verifying it).

By Assumption MS, $\varphi_{MS-AKP,\alpha}=\varphi_{Rob,\alpha-\delta}$ wpa1 in case ((ref)). Thus, by Assumption RT, the new test $\varphi_{MS-AKP,\alpha}$ has limiting NRP bounded by $\alpha-\delta$ of the test in that case.

For the model selection methods introduced below, the sequence of constants $c_{n}$ reflects a trade-off between size and power. Large values of $c_{n}$ will imply frequent use of $\varphi_{AKP,\alpha}$ which should translate into good power properties. On the other hand, use of $\varphi_{AKP,\alpha}$ could distort the NRPs in finite samples if the test is used in a scenario where the covariance matrix does not have AKP structure. Below we make a recommendation regarding the choice of $c_{n}$ based on comprehensive Monte Carlo studies. Note that $c_{n}$ can also depend on observed nonrandom quantities such as e.g. $k$ and $m_{W}$ but for the sake of notational simplicity we do not make that explicit.

To guarantee correct asymptotic significance level $\alpha$ of the test $\varphi_{MS-AKP,\alpha}$ and to rule out any potential pretesting issue, we have to implement the test $\varphi_{Rob,\alpha}$ at a nominal size infinitesimally smaller than $\alpha.$ For example, we can pick $\delta =10^{-6},$ which should not make any practical difference in terms of power relative to using the test with $\delta=0.$

In addition, we have to impose one additional assumption regarding the relative NRPs (Assumption RP below) of the robust test $\varphi_{Rob,\alpha -\delta}$ and $\varphi_{AKP,\alpha}$ under sequences with AKP structure in order to make sure that $\varphi_{MS-AKP,\alpha}$ has limiting NRP bounded by $\alpha.$ More precisely, consider a sequence of DGPs in $\mathcal{F}_{Het}$ such that

equation[equation omitted — 88 chars of source]

Using $n^{1/2}/c_{n}\rightarrow\infty,$ one can then show that $\min||G\otimes H-R_{F_{n}}||\rightarrow0$ and the sequences are of AKP structure. Therefore, under such sequences the test $\varphi_{AKP,\alpha}$ has limiting NRP bounded by $\alpha$. The notation $P_{\lambda_{w_{n},h}}(A)$ denotes probability of an event $A$ when the true DGP is characterized by $\lambda_{w_{n},h}$. By definition, along $\lambda_{w_{n},h},$ $w_{n}^{1/2}\lambda_{9w_{n}}\rightarrow h_{9}$ and thus when $h_{9}<\infty$ the sequence is local to KP structure.

\noindentAssumption RP: Under sequences of DGPs $(\gamma_{w_{n}} ,\Pi_{Ww_{n}},\Pi_{Yw_{n}},F_{w_{n}})$ in $\mathcal{F}_{Het}$ for subsequences $w_{n}$ for which $\lambda_{w_{n},h}$ satisfies $h_{9}\in\lbrack0,\infty),$ $P_{\lambda_{w_{n},h}}(\varphi_{Rob,\alpha-\delta}\leq\varphi_{AKP,\alpha })\rightarrow1.$

Assumption RP says that under null sequences local to KP\ structure the robust test $\varphi_{Rob,\alpha-\delta}$ has critical region that is contained in the critical region of $\varphi_{AKP,\alpha}$ with probability going to one. Even when $\delta=0$ this does not need to imply that the two tests are asymptotically identical because the robust test may have limiting NRP strictly smaller than $\alpha$ and may be more conservative than $\varphi_{AKP,\alpha}.$ Under Assumption RP one can show that in case ((ref)) (i.e. under drifting sequences of DGPs $\lambda_{w_{n},h}$ with finite $h_{9})$ $\varphi_{MS-AKP,\alpha}$ has limiting NRP bounded by the nominal size of the test (because from the proof of Theorem (ref) the test $\varphi_{AKP,\alpha}$ has limiting NRP bounded by $\alpha$; and the limiting NRP of the new test $\varphi _{MS-AKP,\alpha}$ is then bounded by $\alpha$ by the assumption that $\varphi_{Rob,\alpha-\delta}$ has asymptotic size bounded by $\alpha-\delta$.)

From the above, it then follows quite straightforwardly, that the asymptotic size of $\varphi_{MS-AKP,\alpha}$ is bounded by the nominal size for the parameter space $\mathcal{F}_{Het}$. Also, the new test is at most as nonsimilar asymptotically as $\varphi_{Rob,\alpha-\delta}$ which translates into favorable power properties of the new test.

theoremSuppose Assumptions PS, RT, MS, and RP\ hold. Then the test $\varphi_{MS-AKP,\delta,c_{n},\alpha}$ defined in $($(ref)$)$ with $\delta>0$ and $c_{n}$ satisfying the conditions in $($(ref)$)$ has asymptotic size bounded by the nominal size $\alpha$ for the parameter space $\mathcal{F}_{Het}$ for $\alpha\in\left\{ 1\%,5\%,10\%\right\} $ and $k-m_{W}\in\left\{ 1,...,20\right\} .$

Comments. 1. If $\lim\inf_{n\rightarrow\infty}\inf_{(\gamma,\Pi _{W},\Pi_{Y},F)\in\mathcal{F}_{Het}}E_{(\gamma,\Pi_{W},\Pi_{Y},F)} \varphi_{MS-AKP,\delta,c_{n},\alpha}$ is continuous in $\delta$ at $\delta=0$ then as $\delta\rightarrow0$ the new test $\varphi_{MS-AKP,\delta,c_{n} ,\alpha}$ is asymptotically not more nonsimilar (i.e. less conservative) than $\varphi_{Rob,\alpha}$, i.e.

align[align omitted — 367 chars of source]

See the proof of Theorem (ref) for a proof. Property ((ref)) should translate into improved power of $\varphi _{MS-AKP,\delta,c_{n},\alpha}$ relative to $\varphi_{Rob,\alpha}.$

2. The restriction to $\alpha\in\left\{ 1\%,5\%,10\%\right\} $ and $k-m_{W}\in\left\{ 1,...,20\right\} $ in the formulation of Theorem (ref) is an artifact of Theorem (ref) where the conditional subvector test $\varphi _{AKP,\alpha}$ was shown to have correct asymptotic size for these cases only. The same is true for other theorems formulated below.

In the next subsection we specifically use the AR/AR subvector procedure due to Andrews2017 as $\varphi_{Rob,\alpha-\delta}.$

Model selection methods $\varphi_{MS,c_{n}}\label{2 model sel}$

In this subsection we discuss two methods that can be used for $\varphi _{MS,c_{n}}$ as model selection procedures. The first one is akin to the moment selection method in AndrewsSoares2010 to check which moment inequalities are binding in a model defined by moment inequalities. The second one is based on the test for KP structure introduced in GKM22.

Method 1: Define

equation[equation omitted — 181 chars of source]

with $\widehat{G}_{n}$ and $\widehat{H}_{n}$ defined in ((ref)), to evaluate how far the true model is away from KP structure. Define the first choice for model selection as

equation[equation omitted — 79 chars of source]

Recall the definition of $\widetilde{\mathcal{F}}_{AKP,a_{n}}$ given in Assumption PS. Here we take

align[align omitted — 270 chars of source]

It is easy to show using the formulae in ((ref)) and the analogous one $\widehat{R}_{nA}=(I_{p}\otimes T_{A}^{\prime} )\widehat{R}_{n}(I_{p}\otimes T_{A})$ for $\widehat{R}_{n}$, orthogonality of $T_{A}$, and using the fact that the Frobenius norm is invariant to orthogonal transformations, that $\widehat{K}_{n}$ is invariant to nonsingular transformations of the instrument vector. Crucial for this result is again that $f_{i}$ in ((ref)) in the definition of $\widehat{R}_{n}$ (and as a result in the definition of $\widehat{G}_{n}$ and $\widehat{H}_{n}$ in ((ref))) is implemented with the transformed instrument vector $Z_{i}$ (rather than with $\overline{Z}_{i}$).

Method 2: Define

equation[equation omitted — 73 chars of source]

where $KPST$ is the test statistic introduced in GKM22 to test the null of a KP structure of $R_{F}.\footnote{The test statistic is defined in (19) and (22) in GKM20 and not reproduced here for brevity. In their notation our $f_{i}$ is $\widehat{f}_{i},$ compare the formula below (7) in GKM20 to our (\ref{fi}).}$ To employ this method, we need to strengthen the moment restrictions in $\mathcal{F}_{Het}$ to $E_F(||T_i||^2+\delta_1)\leq B,$ for $T_i\in\{||\overline{Z}_i||^4||U_i||^4,||\overline{Z}_i||^4\},$ see Theorem 3 in GKM22.

We verify Assumption MS in the Appendix, Section (ref), for these two choices of $\varphi_{MS,c_{n}}$ and for the parameter space defined in ((ref)).

Choice for $\varphi_{Rob,\alpha}:$ The AR/AR test in Andrews2017

In this subsection we define one particular version of the various weak IVs and heteroskedasticity robust subvector tests suggested in Andrews2017, namely the so-called AR/AR test and verify that it satisfies Assumptions RT and RP from the previous subsection. We define it for nominal size $\alpha.$

To do so, we use the following quantities. For $\theta=\left( \beta ,\gamma\right) $ let\footnote{To simplify notation we write $(\beta,\gamma)$ here and in other situations, rather than the more correct $(\beta^{\prime },\gamma^{\prime})^{\prime}.$}

equation[equation omitted — 247 chars of source]

Define

equation[equation omitted — 297 chars of source]

The heteroskedasticity-robust AR statistic for testing hypotheses involving the full parameter vector $\theta$, evaluated at $\left( \beta_{0} ,\gamma\right) ,$ is defined as

equation[equation omitted — 235 chars of source]

For $s=1,...,m_{W}$ denote by $W^{s}\in\Re^{n}$ the $s$-th column of $W.$ Next, as in Andrews2017 let

align[align omitted — 1,048 chars of source]

where $HAR_{\beta,n}\left( \beta_{0},\gamma\right) $ is a $C\left( \alpha\right) $-AR statistic, obtained as a quadratic form in the moment conditions projected onto the space orthogonal to the orthogonalized Jacobian with respect to $\gamma$. The random perturbation $an^{-1/2}\zeta_{1}$ (with $\zeta_{1}\in\Re^{k\times m_{W}}$ a random matrix of independent standard normal random variables that are independent of all other statistics considered) in the last line of ((ref)) is introduced in Andrews2017, to guarantee that the space projected on has rank $m_{W}$ a.s. Here $a\in\Re$ is a tiny positive constant.

Let $\alpha\in(0,1)$. The AR/AR test at nominal size $\alpha$ is defined as follows.

enumerate• Fix an $\alpha_{1}\in\left( 0,\alpha\right) .$ As in Andrews2017 define \begin{equation} CS_{1n}^{+}:=\{\widetilde{\gamma}\in\Re^{m_{W}}:HAR_{n}\left( \beta _{0},\widetilde{\gamma}\right) <\chi_{k,1-\alpha_{1}}^{2}\}\cup \widetilde{\Gamma}_{1n}, \end{equation} where for $\widehat{Q}_{n}\left( \theta\right) :=\widehat{g}_{n}\left( \theta\right) ^{\prime}(n^{-1} {\textstyle\sum\nolimits_{i=1}^{n}} \overline{Z}_{i}\overline{Z}_{i}^{\prime})^{-1}\widehat{g}_{n}\left( \theta\right) ,$ \begin{align} \widetilde{\Gamma}_{1n} & :=\left\{ \gamma\in\Re^{m_{W}}:W^{\prime }\overline{Z}( {\textstyle\sum\nolimits_{i=1}^{n}} \overline{Z}_{i}\overline{Z}_{i}^{\prime})^{-1}\widehat{g}_{n}\left( \beta_{0},\gamma\right) =0^{m_{W}} &\right. \\ & \left. \widehat{Q}_{n}\left( \beta_{0},\gamma\right) \leq \min_{\widetilde{\gamma}\in\Re^{m_{W}}}\widehat{Q}_{n}\left( \beta _{0},\widetilde{\gamma}\right) +\frac{\ln n}{n}\right\} \nonumber \end{align} is the so-called \textquotedblleft estimator set\textquotedblright, see Andrews2017. If $W^{\prime}P_{\overline{Z}}W$ is invertible (which would happen wpa1 under the assumption (not imposed here) that $E_{F}\overline{Z}_{i}W_{i}^{\prime}$ is full column rank) then the first condition in $\widetilde{\Gamma}_{1n}$ has the unique solution $\overline {\gamma}_{n}:=(W^{\prime}P_{\overline{Z}}W)^{-1}W^{\prime}P_{\overline{Z} }(y-Y\beta_{0})$ and therefore $\widetilde{\Gamma}_{1n}=\{\overline{\gamma }_{n}\}.$ (Note that along certain sequences for which $||\gamma ||\rightarrow\infty$ it follows that $||\widehat{g}_{n}\left( \beta _{0},\gamma\right) ||\rightarrow\infty$ and therefore if the function $\widehat{Q}_{n}\left( \beta_{0},\gamma\right) \geq0$ only has one local extremum it must be a global minimum.) • For $\alpha_{2,n}(\theta)$ defined below (and depending on $\alpha$ and $\alpha_{1})$, $H_{0}$ in ((ref)) is rejected if \[\inf _{\widetilde{\gamma}\in CS_{1n}^{+}}\left(HAR_{\beta,n}\left( \beta_{0} ,\widetilde{\gamma}\right) -\chi_{k-m_{W},1-\alpha_{2,n}(\beta_{0} ,\widetilde{\gamma})}^{2}\right)>0.\] That is \begin{equation} \varphi_{AR/AR,\alpha,\alpha_{1}}=1_{\left\{ \inf_{\widetilde{\gamma}\in CS_{1n}^{+}}\left(HAR_{\beta,n}\left( \beta_{0},\widetilde{\gamma}\right) -\chi_{k-m_{W},1-\alpha_{2,n}(\beta_{0},\widetilde{\gamma})}^{2}\right)>0\right\} }, \end{equation} see Andrews2017. We typically write $\varphi_{AR/AR,\alpha}$ instead of $\varphi_{AR/AR,\alpha,\alpha_{1}}.$

The second step size $\alpha_{2,n}(\theta)$ is chosen as

equation[equation omitted — 209 chars of source]

for some positive number $K_{L}$, e.g., $K_{L}=0.05$ and $\alpha_{1}=.005,$ see Andrews2017\footnote{Andrews2017 allows for more involved definitions of $\alpha_{2,n}(\theta).$ We choose the version that takes $K_{U}=K_{L}$ in the notation of Andrews2017 that is also used in the Monte Carlos in Andrews2017. Regarding the definition of $\widehat{\Phi}_{n}(\theta),$ note that it constitutes a slight modification compared with the definitions in Andrews2017. In particular, the modification in the definition of $\hat{\sigma}_{sn}^{2}$ is necessary to make the procedure invariant to nonsingular transformations of the instrument vector. We thank Donald Andrews for suggesting this updated version of his test statistic.}, where

align[align omitted — 791 chars of source]

see Andrews2017, where $W_{i}^{s}\in\Re$ denotes the $s$-th component of $W_{i}$.

Coming back to the statistic $AR_{AKP,n}(\beta_{0})$ given in ((ref)) note that

align[align omitted — 531 chars of source]

using the fact that the minimal eigenvalue of any symmetric square matrix $A\in\mathfrak{R}^{p\times p}$ is obtained as $\min_{x\in\mathfrak{R} ^{p},||x||=1}x^{\prime}Ax.$ Furthermore,

align[align omitted — 970 chars of source]

and $(\widehat{G}_{n},\widehat{H}_{n})$ defined in ((ref)).

Let $\gamma_{n}^{+}$ be an element in $\arg\min_{\widetilde{\gamma} \in\mathfrak{R}^{m_{W}}}\widetilde{AR}_{AKP,n}(\beta_{0},\widetilde{\gamma}).$ We impose a mild technical condition below, namely that

equation[equation omitted — 95 chars of source]

and $\gamma_{n}^{+}=O_{p}(1)$ under sequences in $\mathcal{F}_{Het}$ (defined in ((ref)) below)\ that are of AKP structure, i.e. under sequences $\lambda_{n,h}$ for which $h_{9}\in\lbrack0,\infty)$.

Condition ((ref)) has been established for several closely related estimators. E.g. $\gamma_{n}^{+}-\gamma_{n}=O_{p}(1)$ holds under weak IV sequences $\Pi_{Wn}=C/n^{1/2}$ (for some fixed matrix $C)$ and homoskedasticity when $\gamma_{n}^{+}$ is the LIML estimator, see StaigerStock1997. Results in HahnKuersteiner2002 imply ((ref)) for the 2SLS\ estimator under a setup where $\Pi_{Wn}=C/n^{\delta}$ for $\delta>0.$ StockWright2000 and GuggenbergerSmith2005 implies ((ref)) for the CU estimator under mixed weak/strong IV asymptotics $\Pi _{Wn}=(C/n^{1/2},D)$ for a fixed full rank matrix $D\in\Re^{k\times m_{W}^{\prime}}$ with $m_{W}^{\prime}\leq m_{W}$ (using high level assumptions, such as Assumptions B and D in StockWright2000) and possible CHET.

StockWright2000 can also be applied in the current situation to show ((ref)) under sequences $\lambda_{n,h}$ for which $h_{9}\in\lbrack0,\infty).$ Given $\varepsilon>0$ we need to show that for some compact set $K_{\varepsilon},$ $\Pi_{Wn}n^{1/2}(\gamma_{n} ^{+}-\gamma_{n})\in K_{\varepsilon}$ with probability at least $1-\varepsilon$ for all large enough sample sizes. Assuming $\gamma_{n}^{+}=O_{p}(1),$ then for all $\varepsilon>0,$ $\gamma_{n}^{+}$ is contained in a compact set $K_{\varepsilon}$ with probability at least $1-\varepsilon$ for all large enough sample sizes. Consider the estimator $\gamma_{n}^{K_{\varepsilon}}$ that is defined as a minimizer of $\widetilde{AR}_{AKP,n}(\beta_{0} ,\widetilde{\gamma})$ in $\widetilde{\gamma}$ over $K_{\varepsilon}.$ Thus $\gamma_{n}^{K_{\varepsilon}}$ and $\gamma_{n}^{+}$ are numerically identical for all sample sizes large enough with probability at least $1-\varepsilon.$ Note that $\widetilde{AR}_{AKP,n}(\beta_{0},\widetilde{\gamma})$ has the same structure as the criterion function $S_{T}(\theta,\theta)$ in (2.2) in StockWright2000 with $\widetilde{\Sigma}_{n}\left( \beta _{0},\widetilde{\gamma}\right) ^{-1}$ playing the role of the weighting matrix $W_{T}(\theta)$ and $n^{1/2}\widehat{g}_{n}\left( \beta_{0} ,\widetilde{\gamma}\right) $ playing the role of $n^{-1/2} {\textstyle\sum\nolimits_{s=1}^{T}} \phi_{s}\left( \theta\right) $. Therefore, under drifting sequences of mixed weak/strong IVs, namely $\Pi_{Wn}=(C/n^{1/2},D),$ the limiting distribution of $\gamma_{n}^{K_{\varepsilon}}$ is given in StockWright2000 if Assumptions B and D in StockWright2000 hold for parameter space $K_{\varepsilon}$ for $\widetilde{\gamma}$ and $\widetilde{AR}_{AKP,n}(\beta_{0},\widetilde{\gamma })$ has a unique minimum. StockWright2000 states that those components of $\gamma_{n}^{K_{\varepsilon}}-\gamma_{n}$ that correspond to the columns of $C/n^{1/2}$ in $\Pi_{Wn}$ are $O_{p}(1)$ and those that correspond to the columns of $D$ in $\Pi_{Wn}$ are $O_{p}(n^{-1/2})$ which establishes $\Pi_{Wn}n^{1/2}(\gamma_{n}^{K_{\varepsilon}}-\gamma_{n} )=O_{p}(1)$.

Assumption B in StockWright2000 holds for $\gamma_{n}^{K_{\varepsilon }}$ (in fact, Assumption B' in StockWright2000, which is sufficient for Assumption B, holds by linearity of $g_{i}\left( \theta\right) ,$ the moment conditions in $\mathcal{F}_{Het},$ and compactness of $K_{\varepsilon }).$ To establish Assumption D note that under sequences $\lambda_{n,h},$ $n^{-1}\overline{Z}^{\prime}\overline{Z}\rightarrow_{p}\lim E_{F_{n}} \overline{Z}_{i}\overline{Z}_{i}^{\prime},$ $\widehat{G}_{n}\rightarrow _{p}\lim G_{F_{n}},$ and $\widehat{H}_{n}\rightarrow_{p}\lim H_{F_{n}}$ (note that the right hand side limits exist by definition of $\lambda_{n,h}).$ Therefore, under sequences $\lambda_{n,h}$

equation[equation omitted — 399 chars of source]

uniformly over $\widetilde{\gamma}$ (noting that $||\left( 1,-\widetilde{\gamma}^{\prime}\right) ||\geq1,$ $\lim G_{F_{n}}>0$, and $\lim H_{F_{n}}>0)$ with the limit matrix being nonrandom, continuous, symmetric, and pd for all $\widetilde{\gamma}.$ Thus, by StockWright2000, if $\widetilde{AR}_{AKP,n}(\beta _{0},\widetilde{\gamma})$ has a unique minimum in $K_{\varepsilon}$ and $\gamma_{n}^{+}=O_{p}(1),$ it follows that $\Pi_{Wn}n^{1/2}(\gamma _{n}^{K_{\varepsilon}}-\gamma_{n})=O_{p}(1).$ Thus there exists a compact set $K$ such that $\Pi_{Wn}n^{1/2}(\gamma_{n}^{K_{\varepsilon}}-\gamma_{n})\in K$ at least with probability $1-\varepsilon$ for all $n$ large enough. Because $\gamma_{n}^{+}$ and $\gamma_{n}^{K_{\varepsilon}}$ coincide at least with probability $1-\varepsilon$ for all large enough sample sizes, it then follows that $\Pi_{Wn}n^{1/2}(\gamma_{n}^{+}-\gamma_{n})\in K$ at least with probability $1-2\varepsilon$ for all $n$ large enough.

Deriving ((ref)) under all possible drifting sequences $\Pi_{Wn}$ is technically tedious and involves e.g. also consideration of so-called sequences of non-standard weak identification, see {Andrews and Guggenberger (2019, AG from now on) \nocite{AndrewsGuggenberger2019} }for more discussion. If ((ref)) is not already implied by the restrictions in the parameter space $\mathcal{F}_{Het}$ below then the asymptotic size results should simply be interpreted for sequences of parameter spaces $\mathcal{F}_{Het,n}$ that impose additional restrictions on $\mathcal{F}_{Het}$ such that ((ref)) holds.

The null parameter space is restricted by the conditions in $\mathcal{F} _{AR/AR}$ of Andrews2017 and some weak additional ones, namely,

align[align omitted — 879 chars of source]

for constants $B<\infty,$ and $\delta_{1},\delta_{2}>0$ and a bounded set $\Theta_{\gamma\ast}$ such that for some $\epsilon>0$ we have $B(\Theta _{\gamma\ast},\epsilon)\subset\Theta_{\gamma},$ where $\Theta_{\gamma}$ denotes the null nuisance parameter space for $\gamma$ and $B(\Theta _{\gamma\ast},\epsilon)$ denotes the union of closed balls in $\Re^{m_{W}}$ with radius $\epsilon$ centered at points in $\Theta_{\gamma\ast}.$

lemmaAssume that under any sequence of DGPs $(\gamma_{w_{n}} ,\Pi_{Ww_{n}},\Pi_{Yw_{n}},F_{w_{n}})$ in $\mathcal{F}_{Het}$ defined in $($(ref)$)$ for subsequences $w_{n}$ for which $\lambda_{w_{n},h}$ satisfies $h_{9}\in\lbrack0,\infty)$ we have $\gamma_{w_{n}}^{+}=O_{p}(1)$ and $\Pi_{Ww_{n}}^{1/2}w_{n}(\gamma_{w_{n}}^{+}-\gamma_{w_{n}})=O_{p}(1).$ Then, for any $\delta>0,$ the AR/AR test $\varphi_{AR/AR,\alpha-\delta,\alpha_{1}}$ in $($(ref)$)$ satisfies Assumptions RT and RP for the parameter space $\mathcal{F}_{Het}$.

Main result

We obtain the following corollary of Lemma (ref), Theorem (ref), and the verification of Assumption MS in subsection (ref) for the two model selection methods $\varphi_{MS,c_{n}}$ suggested there.

Define the parameter space $\mathcal{F}_{Het}$ as the intersection of the parameter spaces defined in $($(ref)$)$ and $($(ref)$)$ when the method in ((ref)) is used as $\varphi_{MS,c_{n}}$ (and a slightly more restricted parameter space when ((ref)) is used, as explained below ((ref)).)

corollaryAssume the same condition as in Lemma (ref). Then the test $\varphi_{MS-AKP,\alpha}$ defined in $($(ref)$)$ with $\delta>0$ and $c_{n}$ satisfying the conditions in $($(ref)$)$ implemented with the AR/AR test $\varphi_{AR/AR,\alpha-\delta,\alpha_{1}}$ of Andrews2017 playing the role of $\varphi_{Rob,\alpha-\delta}$ and either of the two model selection methods described above used for $\varphi_{MS,c_{n}},$ has asymptotic size bounded by the nominal size $\alpha$ for the parameter space $\mathcal{F}_{Het}$ defined in the paragraph above for $\alpha\in\left\{ 1\%,5\%,10\%\right\} $ and $k-m_{W}\in\left\{ 1,...,20\right\} .$

Comment. Note that under the null hypothesis the test does not depend on the value of the reduced form matrix $\Pi_{Y}.$

Monte Carlo study

In this section we investigate the finite sample performance in model ((ref)) of the suggested new test $\varphi_{MS-AKP,\alpha}$ defined in ((ref)) and juxtapose it to the performance of alternative methods suggested in the extant literature, namely the two-step tests AR/AR, AR/LM, and AR/QLR1 in Andrews2017. For the implementation of $\varphi _{MS-AKP,\alpha}$ we use both methods considered in Section (ref) and call the resulting tests MS-AKP1 and MS-AKP2 for the remainder of this section. We also simulate the performance of the test AR$_{AKP,\alpha}$ (which is of course size distorted in the setups with CHET that are outside of KP structure).

All results below are for nominal size $\alpha=5\%.$ Unless otherwise stated, we take $m_{W}=1.$ We consider the case $\beta\in\Re$ and $\gamma\in\Re$ and pick $\gamma=0$ and test the null hypothesis in ((ref)) with $\beta_{0}=0.$ $\smallskip$

Choice of tuning parameters

The implementation of the various tests depends on a large number of user chosen constants. In particular, to implement the AR/AR, AR/LM,\ and the AR/QLR1 tests we pick $\alpha_{1}=.005,$ $K_{L}=K_{U}=0.05$ as already mentioned above after ((ref)). To calculate the estimator set $\widetilde{\Gamma}_{1n}$ we employ the closed form solution provided below ((ref)). We choose $a=.001$ and pick the elements of the random matrix $\zeta_{1}\in\Re^{k\times m_{W}}$ as i.i.d. $N(0,1)$ independent of all other variables considered, see the last line of ((ref)).\footnote{Note that by choosing $a\neq0$ the tests are no longer invariant to nonsingular transformations of the IV vector. However, for small $a$ the differences after a transformation are usually very small.} The confidence interval (or region) for $\gamma$ that appears in ((ref)) is obtained by grid search over an interval (or rectangle) of length 20 centered at the true value of $\gamma$ with 100 equally spaced gridpoints.\footnote{When the dimension of $\gamma$ grows then the implementation of that step by grid search will cause an exponential increase in computation time for each of the two-step methods.} To implement the AR/QLR1 test, as in Andrews2017 we pick $K_{L}^{\ast }=K_{U}^{\ast}=0.005$ and $K_{rk}=1$. We refer to Table II in Andrews2017 that provides the results of a comprehensive sensitivity analysis on most of the user chosen constants above. To calculate the data-dependent critical values for the AR/QLR1 test we use 10,000 i.i.d chi-square random variables. There was no noticeable difference between $\delta=0$ and $\delta=10^{-6}$ for $\delta$ given in ((ref)); therefore, for the sake of computational simplicity, we pick the former in the simulations. Finally, $c_{n}$ has to be chosen, which we do in the next subsection.

Recommended choices for $c_{n}$

First, we perform a large number of simulations in order to determine recommendations for the sequence of constants $c_{n}$ satisfying ((ref)). We make recommendations for $c_{n,k,m_{W}}=c_{n}$ as a function of the number $k$ of IVs and the subvector dimension $m_{W}$ and consider choices from $k\in\{2,3,4\}$ and $m_{W}\in\{1,2\}$.

For each $k$, sample size $n\in\{250,500\},$ and $(\Pi_{Y},\Pi_{W})\in\Re^{k\times2}$ with

equation[equation omitted — 72 chars of source]

with $\pi_{W}\in\{2,4,40\},$ corresponding to \textquotedblleft very weak\textquotedblright, \textquotedblleft weak\textquotedblright,\ and \textquotedblleft strong\textquotedblright\ identification of $\gamma$ $($and, relevant for the power results below, $\Pi_{Y}=\widetilde{1}^{k}\pi _{Y}/(nk)^{1/2}$ with $\pi_{Y}\in\{2,4,40\}$ and $\widetilde{1}^{k}$ equal to $(1^{k/2\prime},-1^{k/2\prime})^{\prime}$ when $k$ is even and equal to $(1,-1^{2\prime})^{\prime}$ when $k=3)$ we randomly generate $1,000$ different DGPs (that is a choice for the covariance matrix) as described below and simulate the NRPs (using $5,000$ i.i.d samples of each given DGPs) of MS-AKP1 and MS-AKP2 for choices of $c_{n}$ given as

equation[equation omitted — 74 chars of source]

with $c(k,1)$ taken from the set $C:=\{.05,.1,...,3\}.$

In finite sample simulations for the DGPs considered here, the AR/AR test sometimes slightly overrejects. For example, under CHOM, $n=250,$ $k=3,$ strong IVs, and covariance matrix $\Sigma$ being chosen as below ((ref)), where $(u_{i},v_{Y,i},v_{W,i})^{\prime }\sim$ i.i.d. $N(0^{3},\Sigma),$ the AR/AR test has NRP equal to 5.4%. From our theory we also know that the test AR$_{AKP,\alpha}$ (at least under AKP structures) has nonsmaller NRP than the AR/AR test. Define as the "simulated size of a test when there are $k$ IVs" the highest empirical NRP of the test over all choices of $n$, $\Pi$, and ($1,000$) random DGPs considered. For each of the two methods MS-AKP1 and MS-AKP2 and for each $k\in\{2,3,4\},$ our recommendation for $c_{n,k,1}$ then is to take the largest $c(k,1)$ in $C$ such that the simulated size does not exceed 6% (that is, we allow for a distortion of 1% in the "simulated size"). It turns out that in our simulations this criterion for $c_{n,k,1}$ always leads to well defined choice of $c(k,1)$ (when a priori it could be that even for the smallest/largest choice of $c(k,1)$ in $C$ the simulated size exceeds/is still below 6%)$.$

To generate random DGPs we consider the following mechanism. Given all tests considered above, including AR$_{AKP,\alpha},$ have correct asymptotic size under AKP structure we focus attention on designs with conditional heteroskedasticity that are not of AKP\ structure. In particular, we choose

align[align omitted — 262 chars of source]

with $(u_{i},v_{Y,i},v_{W,i})^{\prime}\sim$ i.i.d. $N(0^{3},\Sigma)$ and independent of $\overline{Z}_{i}\sim$ i.i.d. $N(0^{k},I_{k})$ for $i=1,...,n$. Each of the 1,000 random DGPs is determined by choosing $\alpha_{\varepsilon },\alpha_{V}\in\Re,$ $Q_{\varepsilon},Q_{V}\in\Re^{k\times k},$ and $\Sigma \in\Re^{3\times3},$ where $\Sigma$ has diagonal elements equal to 1. The scalars $\alpha_{\varepsilon},\alpha_{V}$ and the components of $Q_{\varepsilon},Q_{V}\in\Re^{k\times k}$ are obtained by i.i.d. draws from a $U[0,10],$ and the off-diagonal ones of $\Sigma\in\Re^{3\times3}$ are obtained by i.i.d. draws from a $U[0,1]$ (subject to the restriction that the resulting matrix $\Sigma$ is pd). Note that the setup in ((ref)) nests KP structure when e.g. $\alpha_{\varepsilon}=\alpha_{V}=0,$ $Q_{\varepsilon }=Q_{V}=I_{k}$ and CHOM when e.g. $\alpha_{\varepsilon}=\alpha_{V}=1,$ $Q_{\varepsilon}=Q_{V}=0^{k\times k}.$

For each $k=2,3,4$ the binding constraint on $c(k,1)$ always came from the combination $n=250$ and \textquotedblleft strong\textquotedblright \ identification, while for the \textquotedblleft very weakly\textquotedblright\ identified scenario even the largest choice of $c(k,1)\in C$ typically did not yield overrejection for any of the sample sizes considered. Based on the above setup we recommend the following choices for $c_{n,k,1}.$ For Method 1 in Section (ref), MS-AKP1, that is for $\varphi_{MS-AKP,\alpha}$ based on the distance in Frobenius norm statistic$,$ we suggest

equation[equation omitted — 94 chars of source]

while for Method 2, MS-AKP2, that is for $\varphi_{MS-AKP,\alpha}$ based on the KPST statistic in GKM22, we suggest

equation[equation omitted — 94 chars of source]

Recall that with these choices of $c(k,1)$ and $c_{n}$ chosen as in ((ref)) the tests and MS-AKP1 and MS-AKP2 have correct asymptotic size for a parameter space with arbitrary forms of conditional heteroskedasticity.

Next we consider $m_{W}=2.$ We take $\gamma=(0,0)^{\prime}$. As pointed out above already, the computational effort in the above exercise increases exponentially in the dimension of $m_{W}$ if we use the same number of gridpoints in each dimension in the calculation of the confidence interval for $\gamma$ that appears in ((ref)). Therefore, we use a grid of a product of two intervals of length 20 centered at the true value of $\gamma$ with only 50 equally spaced gridpoints in each dimension (rather than 100 in the case $m_{W}=1$.) Everything else is the same, mutatis mutandis (e.g. $\Sigma$ is now a 4x4 matrix), as described in the case $m_{W}=1$ except that $\Pi_{W}\in\Re^{k\times2}$ is taken as $(\pi_{W1}e_{1},\pi_{W2}e_{2})/(nk)^{1/2}$ with $\pi_{W1},\pi_{W2} \in\{2,40\}$ and that the search set for $c(k,2)$ is increased to $C:=\{.05,.1,...,7.5\}$.

For Method 1 in Section (ref), MS-AKP1 we suggest

equation[equation omitted — 90 chars of source]

while for Method 2, MS-AKP2 we suggest

equation[equation omitted — 88 chars of source]

Just like in the case $m_{W}=1$ the highest null rejection probabilities occur in the strongly identified case.

Size results

All results below are for the case where $m_{W}=1.$ Under a setup with CHET outside of KP, the tests MS-AKP1 and MS-AKP2 equal the AR/AR test wpa1. We therefore first consider the KP setup in Andrews2017 in Section 9.1 which is obtained from ((ref)) with $\alpha_{\varepsilon}=\alpha_{V}=0$ and $Q_{\varepsilon}=Q_{V}=I_{k}$. We also consider the setup with CHOM obtained from ((ref)) with $\alpha_{\varepsilon}=\alpha_{V}=1$ and $Q_{\varepsilon}=Q_{V}=0^{k\times k}.$ Then, finally, below we also examine how power is affected as the DGP transitions from CHOM to CHET outside of KP.

In both cases of CHET and CHOM, we take the matrix

equation[equation omitted — 77 chars of source]

to have diagonal elements equal to one, and the (1,2) and (1,3) elements equal to .8 and the (2,3)\ element equal to .3, as in Andrews2017. We consider $\pi_{W}=\pi_{Y}\in\{2,4,40\}$ in ((ref)), again, representing "very weak", "weak", and "strong" IVs, also see Andrews2017. Finally, we take $k\in\{2,3,4\}$ and sample sizes $n\in\{250,500\}.$ Altogether, that makes for 36 different specifications. In addition, we also obtain results for certain cases of mixed identification strength, e.g. when $\pi_{W}\neq\pi_{Y}\in\{2,40\}$ and also some results for larger sample sizes.

As reported in Andrews2017, we also find that in an overall sense the AR/AR and AR/LM tests are dominated by the AR/QLR1 test. For instance, regarding the AR/LM test, its power function (even in the strong IV context under CHOM) is not always U-shaped and suffers from power dips against certain alternatives. For example, for the KP setup for $n=250,$ $k=4,$ with weak IVs, the power of the AR/LM and AR/QLR1 tests when $\beta=-2$ are 8.6% and 75.6%, respectively, while in the setup with CHOM when $\beta=-1.43$ the power of the AR/LM test is 34.9% while all the other tests have power equal to 100%. On the other hand, the AR/AR test fares worse than the AR/QLR1 test in strongly identified overidentified situations. In what follows, we do not therefore discuss the AR/LM test in much detail.

We consider rejection probabilities under the null $\beta_{0}=0$ and (for power) under a grid of seven $\beta$ values on each side of 0 with distances from the hypothesized value 0 chosen depending on the strength of identification. For example, in the very weakly, weakly, and strongly identified cases we take $\beta$ in the interval $[-2,2],$ $[-2,2]$, and $[-.2,.2]$, respectively, around the true value of 0. Results are obtained from $10,000$ i.i.d samples from each DGP.

First, we discuss the NRPs. Over the 18 DGPs of the KP setups, the NRPs of MS-AKP1, MS-AKP2$,$ AR/AR, AR/LM, and AR/QLR1 lie in the intervals (all numbers in %): [3.5,5.9], [3.3,6.0], [1.9,5.1], [.6,5.2], and [1.5,4.9]. As set up above, the tests MS-AKP1 and MS-AKP2 slightly overreject the null for small sample sizes (especially in the strongly identified case), but the size distortion disappears as $n$ grows. For example, the NRPs of MS-AKP2 in the KP setup with $k=3$ and strong identification is 6.0, 5.5, 5.2, and 5.1%, respectively, when $n=250,$ $500,$ $1,000,$ and $1,500.$ On the other hand, the tests AR/AR, AR/LM, and AR/QLR1, while controlling the NRP very well, underreject the null in weakly identified scenarios. This leads to relatively poor power properties relative to the tests MS-AKP1 and MS-AKP2 in weakly identified situations.

Regarding the 18 DGPs with CHOM, the one important difference relative to the KP\ setup is that the three tests AR/AR, AR/LM, and AR/QLR1 are less conservative with NRPs over the 18 DGPs in the intervals [4.1,5.4], [3.5,5.4], and [3.7,5.1], respectively. As a consequence, these tests have relatively better power properties than in the KP setup.

Power results

Next we discuss the power results. We focus again on the case where $m_{W}=1.$ Power for MS-AKP1, MS-AKP2$,$ AR/AR, and AR/QLR1 increases as the IVs become stronger. On the other hand, by the local-to-zero design considered here (see ((ref)) and below), as $n$ increases, power for these three tests changes only slightly. We therefore only provide details for the case where $n=250$. Power of all the tests is much higher in the setting with CHOM compared to the KP setting and especially so for the AR/QLR1 test (because it underrejects the null hypothesis less under CHOM than under KP). As one example, consider the case $n=250,$ $k=2,$ with weak identification. In that case, when $\beta=-.571$ the tests MS-AKP2, AR/AR, and AR/QLR1 have power 48.7, 46.3, and 45.4% under KP, but power equal to 95.9, 95.6, and 95.4% under CHOM!

figure[figure omitted — 340 chars of source]

A representative selection of power curves in four different cases is plotted in Figure (ref). Note that in the figures corresponding to the different cases, both the scale of the horizontal and the vertical axes vary by a lot depending on the strength of identification.

The key takeaways from the power study are as follows:

i) Based on the DGPs considered here we cannot make a clear recommendation as to which one of the two tests MS-AKP1 and MS-AKP2 is preferable. In most cases, they have virtually identical power. In few cases, one dominates the other, but only by a small difference. In the Figures below we only report results for MS-AKP2.

ii) Regarding the comparison between the tests MS-AKP1, MS-AKP2 and AR/AR we find that the former two virtually uniformly dominate the latter in all the designs considered. This is not surprising given the construction of the new tests and given they satisfy Assumption RP above. The relative power advantage of the tests MS-AKP1, MS-AKP2 over AR/AR partly stems from the underrejection of the latter test under the null. See e.g. Figure (ref)a that contains power curves for $n=250,$ $k=2,$ very weak identification, and KP structure for MS-AKP2, AR/AR, and AR/QLR1. (The NRPs of the three tests reported here are 4.2, 2.0, and 1.6%, respectively. Note that the x-axis in Figures I-IV plots the true value $\beta.$)

iii) Regarding the comparison between the tests MS-AKP1, MS-AKP2 and AR/QLR1 in the case of equal identification strength $\pi_{W}=\pi_{Y}$ we find that the former two are generally more powerful under weak identification and small $k$ while the reverse is true under strong identification and larger $k,$ see Figures (ref)a and b for the cases \textquotedblleft$k=2$ and very weak identification\textquotedblright\ and \textquotedblleft$k=4$ and strong identification,\textquotedblright\ respectively, both for $n=250$ and KP. (In Figure (ref)b, the NRPs of the tests MS-AKP2, AR/AR, and AR/QLR1 are 5.9, 5.1, and 4.6%, respectively.) These two figures show the best relative performances for the MS-AKP1, MS-AKP2 and AR/QLR1 tests in the \textquotedblleft equal identification\textquotedblright\ settings where $\pi_{W}=\pi_{Y}$. In Figure (ref)a the power advantage of MS-AKP2 over AR/QLR1 is as high as 5.2%, while in Figure (ref)b the power of AR/QLR1 can be up to 13.1% more powerful than MS-AKP2.

In the \textquotedblleft intermediate\textquotedblright\ case between these extremes, namely \textquotedblleft$k=3$ and weak identification\textquotedblright\ (again with $n=250$ and KP; not reported in Figure (ref)), the MS-AKP1 and MS-AKP2 tests have slightly higher power than AR/QLR1 when $\beta$ is positive while the reverse is true for negative values of $\beta$. In all cases, the relative performance of the AR/QLR1 test improves under CHOM; under CHOM, for the \textquotedblleft intermediate\textquotedblright\ case \textquotedblleft$k=3$ and weak identification\textquotedblright\ (again with $n=250$) the AR/QLR1 test has uniformly higher power than the MS-AKP1 and MS-AKP2 tests, see Figure (ref)c. (In Figure (ref)c, the NRPs of the tests MS-AKP2, AR/AR, and AR/QLR1 are 5.5, 4.7, and 5.1%, respectively.)

In cases of mixed identification strength, $\pi_{W}\neq\pi_{Y}\in\{2,40\},$ we find that when $\pi_{W}=2$ and $\pi_{Y}=40$ the tests MS-AKP1 and MS-AKP2 have uniformly higher power than AR/QLR1 for all $k$ considered whereas in the case $\pi_{W}=40$ and $\pi_{Y}=2$ all tests have comparable power. See Figure (ref)d that contains the case $\pi_{W}=2$ and $\pi_{Y}=40,$ $n=250,$ $k=4,$ with KP structure where the power gap between the new tests and AR/QLR1 is as high as 13.4%. (In Figure (ref)d, the NRPs of the tests MS-AKP2, AR/AR, and AR/QLR1 are 3.3, 1.9, and 0.9%, respectively.) It seems that in these cases of mixed identification strength the new tests enjoy their most competitive relative performance.

Results for non-KP DGPs

Finally, we examine how rejection probabilities are affected as the DGP transitions from KP to CHET outside of KP. To do so, we report rejection probabilities under the null and certain alternatives and probabilities with which MS-AKP1 and MS-AKP2 equals the AR/AR test in the second stage for a class of DGPs that under KP coincide with the ones considered in Figure (ref)b ($\pi_{W}=\pi_{Y}=40)$ and Figure (ref)d ($(\pi_{W},\pi_{Y})=(2,40))$. In particular, we choose $k=4,$ $n=250,$ $\gamma=0,$ and the matrix $\Sigma$ equals the one in ((ref)). In ((ref) ) we take

equation[equation omitted — 172 chars of source]

for $\varrho\in\{0,.01,...,.1\},$ $\alpha_{\varepsilon}=\alpha_{V}=0,$ and $Q_{V}=I_{4}.$ Note that for $\varrho=0,$ the design leads to KP while for $\varrho>0,$ it leads to CHET outside of KP. In particular, $\arg\min _{G,H>0}||G\otimes H-\overline{R}_{F}||$ (with $\overline{R}_{F}$ defined in ((ref)) with $U_{i}=(\varepsilon_{i},V_{W,i}^{\prime})^{\prime})$ equals $0,$ $.14,$ $.29,$ $.45,$ $.60,$ $.74,$ $.88,$ $1.01,$ $1.13,$ $1.24$, and $1.34$ when $\varrho\in\{0,.01,.02,....,.1\},$ respectively. The latter numbers are found by simulations based on $10^{7}$ simulation repetitions using Theorem 1 in GKM22.

As before, we report results for $10,000$ simulation repetitions at nominal size $5\%$.

Null rejection probabilities. Here we report results when $\beta=0,$ that is, we report NRPs$.$

First, in the setup of Figure (ref)b the probability with which MS-AKP1 and MS-AKP2 coincide with AR/AR is strictly increasing in $\varrho$ and e.g. equals 66.2% and 64.9%, 82.9% and 85.1%, and 98.8% and 99.5%, respectively for $\varrho=0,.03,$ and $.1$, respectively. (We also simulated these probabilities when $\varrho=0$ for $n=500$ and they equal 20.8% and 20.2%, respectively). The highest NRP of both the MS-AKP1 and MS-AKP2 tests is 5.9% which occurs when $\varrho=0$ and is caused by a 7.4% NRP of the conditional subvector test AR$_{AKP,\alpha}$ (even though by Theorem (ref) this test has correct asymptotic NRP for this DGP; interestingly, this test has NRP equal to 5.8% when $\varrho=.1,$ a case that is not covered by Theorem (ref)). Given the MS-AKP1 and MS-AKP2 tests equal the AR/AR test with increasing probability as $\varrho$ increases, their NRPs get closer (but not monotonically so) to 5% as $\varrho$ increases. As $\varrho=.1$ both tests have NRP equal to 5.2%.

Second, in the setup of Figure (ref)d the probability with which MS-AKP1 and MS-AKP2 coincide with AR/AR are identical as just reported for the setup in Figure (ref)b. The highest NRPs of the MS-AKP1 and MS-AKP2 tests are 3.5% and 3.2% respectively, which occur when $\varrho=0$. The AR/AR and AR/QLR1 tests have NRPs in the intervals [1.2%,1.9%] and [.3%,.8%], respectively, and therefore, quite substantially underreject the null hypothesis. As the MS-AKP1 and MS-AKP2 tests equal the AR/AR test with increasing probability as $\varrho$ increases, their NRPs approach 1.2% as $\varrho$ gets closer to $.1.$

Power results. Here we examine how power is affected as the DGP transitions from KP to CHET outside of KP.

First, in the setup of Figure (ref)b) we consider the alternative $\beta=.1.$ The probabilities with which MS-AKP1 and MS-AKP2 coincide with AR/AR are increasing in $\varrho$ and are very similar to the corresponding values when $\beta=0;$ e.g. the probabilities equal 68.6% and 68.9% when $\varrho=.01$, respectively, and equal 98.2% and 99.1% when $\varrho=.1.$ Power for all tests monotonically decreases as $\varrho$ increases, e.g. for AR/QLR1, AR/AR, and MS-AKP1 from 83.5% to 44.4%, from 71.5% to 32.4% and from 72.5% to 32.4%, respectively, when $\varrho$ goes from 0 to .1.

Second, in the setup of Figure (ref)d) we consider the alternative $\beta=-1.$ When $\varrho=.01,$ MS-AKP1 and MS-AKP2 coincide with AR/AR with probability 78.3% and 75.5%, respectively and for $\varrho\geq.04$ both MS-AKP1 and MS-AKP2 coincide with AR/AR at least 99.6% of the cases. The power of the AR/AR and the AR/QLR1 for all values of $\varrho \in\{0,.01,.02,....,.1\}$ are in the intervals [27.9%,29.4%] and [21.2%,23.5%], respectively, with neither test's power being monotonic in $\varrho.$ While the power of MS-AKP1 and MS-AKP2 slightly exceeds the power of the AR/AR test for $\varrho<.04$ their power is identical to the one of the AR/AR test for larger values of $\varrho.\smallskip$

In sum, as one would expect given the construction of MS-AKP1 and MS-AKP2 tests, when moving from KP to CHET outside of KP, their rejection probabilities get closer and closer to those of the AR/AR test.

Conclusion

We propose the construction of a robust test that improves the power of another robust test by combining it with a powerful test that is only robust for a subset of the parameter space. We implement this construction in the context of the linear IV model applied to the AR$_{AKP,\alpha}$ test that has correct asymptotic size for a parameter space that imposes AKP structure and the AR/AR test that is robust even when allowing for arbitrary forms of CHET. We believe that the particular construction and implementation suggested here, namely combining a powerful but non fully robust test with a less powerful fully robust test in order to obtain a fully robust more powerful test, might be successfully applied in other scenarios and also in the current scenario based on different choices of testing procedures. For instance, it might be feasible to combine the LR type subvector test of Kleibergen2021 with the AR/QLR1 of Andrews2017 but it would be technically substantially more challenging to verify the assumptions given above that are sufficient for control of the asymptotic size of the resulting test. Other extensions include improving the power of the AR$_{AKP,\alpha}$ test by making the conditional critical value depend on more than just the largest eigenvalue.