EconBase
← Back to paper

A Test for Kronecker Product Structure Covariance Matrix

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

63,005 characters

A Test for Kronecker Product Structure Covariance Matrix



\title{\textbf{A Test for Kronecker Product Structure Covariance Matrix}
\thanks{Guggenberger gratefully acknowledges the hospitality of the EUI in
Florence while parts of the paper were drafted. Mavroeidis gratefully
acknowledges the research support of the European Research Council via
Consolidator grant number 647152. We would like to thank the Editor Serena Ng, Associate Editor and two anonymous referees for helpful comments and Lewis McLean for
research assistance.}}
\author{$
\begin{array}
[c]{c}
\text{Patrik Guggenberger}\\
\text{Department of Economics}\\
\text{Pennsylvania State University}
\end{array}
$
\and $
\begin{array}
[c]{c}
\text{Frank Kleibergen}\\
\text{Amsterdam School of Economics}\\
\text{University of Amsterdam}
\end{array}
$
\and \vspace{0.25in}$
\begin{array}
[c]{c}
\text{Sophocles Mavroeidis}\\
\text{Department of Economics}\\
\text{University of Oxford}
\end{array}
$}
\date{First Version: October, 2019\\
Revised: \today\vspace{0.25in}}
\maketitle

\begin{abstract}
We propose a test for a covariance matrix to have Kronecker Product Structure
(KPS). KPS implies a reduced rank restriction on a certain transformation of
the covariance matrix and the new procedure is an adaptation of the
\cite{kleibergen2006grr} reduced rank test. To derive the limiting
distribution of the Wald type test statistic proves challenging partly because
of the singularity of the covariance matrix estimator that appears in the
weighting matrix. We show that the test statistic has a $\chi^{2}$ limiting
null distribution with degrees of freedom equal to the number of restrictions
tested. Local asymptotic power results are derived. Monte Carlo simulations
reveal good size and power properties of the test. Re-examining fifteen highly
cited papers conducting instrumental variable regressions, we find that KPS is
not rejected in 56 out of 118 specifications at the 5\% nominal size. \medskip

Keywords: covariance matrix, heteroskedasticity, invariance, Kronecker product
structure, linear instrumental variables regression model, reduced rank, weak identification

JEL codes: C12, C26

\end{abstract}

\section{Introduction}

The robustness properties of nonparametric covariance matrix estimators, like
those proposed by \cite{White80} against heteroskedasticity and by, for
example, \cite{Newe87} and \cite{Andr91} against heteroskedasticity and
autocorrelation, have led to the current default of conducting
semi-parametric inference in econometrics. It is well understood that compared
to parametrically specified covariance matrix estimators, these robustness
properties come at the cost of a large number of additional estimated
components. The latter affects the precision of semi-parametric estimators of
the structural parameters compared to parametric ones.

For some structural models estimated by the generalized method of moments
(GMM), see \cite{Hans82l}, use of nonparametric covariance matrix estimators
may also lead to computational challenges for estimation of the structural
parameters when using the continuous updating estimator (CUE) of
\cite{Hans96l}. Prominent examples of such models are the linear instrumental
variables (IV) regression model and the linear factor model in asset pricing.
When using a nonparametric covariance matrix estimator, the CUE objective
function is often ill behaved, that is, it is flat and/or has many local
extrema, making the CUE difficult to compute. Part of the appeal of the CUE
stems from a number of weak-identification-robust tests based on statistics
centered around the CUE for hypotheses involving the structural parameters,
see e.g. \cite{Kleibergen2005}.

When one uses a Kronecker Product Structure (KPS) covariance matrix estimator
instead of a nonparametric one in the CUE\ objective function in linear IV and
factor asset pricing models, the CUE (which is then typically referred to as
the limited information maximum likelihood (LIML) estimator) is
straightforward to compute. Furthermore, weak-identification-robust tests
specified on a subvector of the structural parameter vector with uniformly
better power than projected robust full-vector tests are available, see e.g.
\cite{gkm19}, \cite{GKM21}, and \cite{kf21}. The KPS structure of the
covariance matrix also allows for an analytical computation of the confidence
sets of the structural parameters using the algorithm from
\cite{dufour2005projection}.\footnote{\cite{dufour2005projection} actually
assume homoskedasticity, which is a special case of KPS covariance, but their
algorithm can be modified to cover the more general case of KPS covariance.}

The above illustrates the trade-off between, on the one hand, the robustness
provided by a nonparametric covariance matrix estimator and, on the other
hand, the computational ease and accurate statistical inference provided by a
KPS covariance matrix estimator. To help empirical researchers decide when the
use of a KPS covariance matrix estimator is justified, we develop a test for
the null hypothesis that the covariance matrix $R=E\left(  \frac{1}{n}
\sum_{i=1}^{n}f_{i}f_{i}^{\prime}\right)  $ has KPS, where $f_{i}=V_{i}\otimes
Z_{i}$ and $V_{i}\in\mathbb{R}^{p}$ and $Z_{i}\in\mathbb{R}^{k}$ are
uncorrelated random vectors. Here $V_{i}$ are unobserved error variables (for
which consistent estimators are available) and $Z_{i}$ are observed
regressors. This setup encompasses, for example, the linear IV and factor asset
pricing models.

The test is based on the insight that KPS implies that a certain invertible
transformation $\mathcal{R}\left( R\right)  $ of $R$ has rank one, see
\cite{vanloanpit93} and (\ref{eq: red rank cov}) below, and our procedure
adapts the \cite{kleibergen2006grr} reduced rank statistic to test for
KPS.\footnote{Another adaptation of the \cite{kleibergen2006grr} reduced-rank
statistic is by \cite{dfp07}, who develop a test for singularity of a
symmetric matrix.} More precisely, the new test statistic is given as a
quadratic form in $vec(\hat{\Lambda})$ with weighting matrix that depends on
$vec(\mathcal{R}(\hat{R})),$ where $\hat{R}$ is a sample analogue of $R$ and
$\hat{\Lambda}$ is an estimator for a certain matrix that is known to be rank
restricted under the null hypothesis (see (\ref{eq: lambda}
)-(\ref{eq: hyp rank}) below). The adaptation of \cite{kleibergen2006grr} is
nontrivial partly because the covariance matrix of $vec(\mathcal{R}(\hat
{R})),$ that appears in the modified test statistic, is singular. As a
consequence, it is a priori not obvious whether the use of the Moore-Penrose
generalized inverse in the expression of the \cite{kleibergen2006grr} reduced
rank statistic still leads to a $\chi^{2}$ limiting distribution. To answer
that question, we first derive the limiting distribution of $\hat{\Lambda}$
and show the limit to be degenerate Normal. We next establish that the
probability limit of the Moore-Penrose inverse of the covariance matrix
involved in the \cite{kleibergen2006grr} rank statistic is such that it
offsets this degeneracy. As the final result, we conclude that the new KPS
test statistic has a $\chi^{2}$ limiting null distribution with degrees of
freedom equal to the number of tested restrictions. We also consider an
asymptotic setup where $p,$ $k,$ and $n$ jointly go to infinity and show that
the asymptotic null rejection probability of the test is controlled as long as
$(pk)^{16}=o(n^{3}).$ For power considerations, we establish that under
sequences of covariance matrices local to KPS the test statistic has a
limiting noncentral chi square distribution.

As an important property we show that the proposed test is invariant to
orthonormal transformations of the data. In contrast, we show that this is not
true for certain alternative tests for KPS that are based on an application of
the \cite{kleibergen2006grr} test statistic to a different transformation of
the sample covariance matrix estimator that does not lead to a singular
covariance matrix (and may therefore a priori seem the more natural choice).

We provide comprehensive Monte Carlo simulations that document good size and
power properties of the suggested test. Finally, we apply the new KPS test to
various specifications of linear IV models employed in fifteen highly cited
empirical studies recently published in top ranked economic journals. We find
that for the specifications with independent data and moderate numbers of
observations, KPS is not rejected in 24 out of 30 cases at the 5\%
significance level, while for smaller numbers of observations, it is rejected in
14 out of 28 cases. In specifications with clustered data, KPS is not rejected
in 7 out of 17 cases with moderate sample sizes, and 11 out of 35 cases with
smaller samples. Overall, KPS is not rejected in 56 out of the 118
specifications that we tested. The relatively high number of non-rejections
illustrates the potential importance of the KPS test for applied work.

In a companion paper, \cite{GKM21}, we show how the new KPS\ test can be used
as a key ingredient in a testing procedure with correct asymptotic size for a
null hypothesis that restricts the values of a subvector of the structural
parameter vector in the linear IV model with a general covariance matrix. The
first step of the algorithm uses the KPS test to test the null of a KPS of the
covariance matrix of the unrestricted reduced-form sample moment vector. In
the second step of the algorithm, the null hypothesis involving the structural
parameter is tested using the improved subvector Anderson-Rubin test from
\cite{gkm19} when the test in the first step does not reject and using the
size correct AR
$\backslash$
AR test procedure from \cite{and17} otherwise. The AR
$\backslash$
AR procedure from \cite{and17} is an asymptotically size correct inference
procedure for testing hypotheses on a subvector of the structural parameters
for general covariance matrices but is less powerful than the improved
subvector Anderson-Rubin test from \cite{gkm19} in the linear IV regression
model. However, the latter test is asymptotically size correct only when the
covariance matrix has KPS. \cite{GKM21} establish that the resulting two-step
procedure has correct asymptotic size and conduct Monte-Carlo experiments
which show that it leads to more powerful subvector inference than the AR
$\backslash$
AR test in \cite{and17}.

As in the linear IV\ regression model, a KPS structure of the covariance
matrix of the sample moment vector of the linear regression model encompassing
linear asset pricing models also leads to improvements in terms of the power
of identification robust tests on individual elements of the vector of risk
premia and computational ease of obtaining the estimator of the risk premia.
There is increasing awareness that risk premia of many risk factors are only
weakly identified, see e.g. \cite{kz99}, \cite{kf09}, and \cite{kz2020}.
Therefore, it is important to analyze them using inference methods that are
robust to weak identification. The current state of the art for conducting
weak-factor-robust inference on risk premia is to assume homoskedasticity.
Extending homoskedasticity to KPS or even further by extending the switching
test procedure from \cite{GKM21} would extend the scope of the
weak-factor-robust inference methods for analyzing the individual risk premia
in linear asset pricing models. The KPS test would be an integral part of such extensions.

KPS or separability, which is how other fields sometimes refer to KPS, of the
covariance matrix is also studied in the statistics and signal processing
literature. The distance to a covariance matrix with KPS\ is considered in
\cite{Genton07} and \cite{VH17}, while \cite{LuZim05} and \cite{mgg06} analyze
the likelihood ratio test of KPS of the covariance matrix of Normally
distributed data. They estimate the elements of the KPS covariance matrix
using a switching algorithm. Exploiting the reduced rank restriction imposed
on the reordered covariance matrix by KPS is also done in \cite{wjs08}. Their
results are, however, based on a complex Gaussian distribution for the data,
which leads to a degrees of freedom parameter of the $\chi^{2}$ limiting
distribution of their test that is different from the one derived here.

KPS is an example of dimension reduction of a covariance matrix. Other
examples of dimension reduction result from shrinking the covariance matrix to
a matrix with (much) fewer unrestricted elements to estimate, for example, a
scalar multiple of the identity matrix, see e.g. \cite{lw12}, or by shrinking
the population eigenvalues, see e.g. \cite{lw15} and \cite{lw18}.

The paper is organized as follows. In the second section, we introduce the new
test for a KPS covariance matrix and derive the asymptotic null distribution
of the test statistic, which we denote as KPST. The third section contains the
limiting distribution of the KPST statistic under local alternatives while the
fourth section conducts a simulation study to analyze the size and power of
the new KPS test. The fifth section summarizes the extensive analysis of
testing for a KPS\ reduced-form covariance matrix in a considerable number of
prominent articles. The final sixth section concludes. Proofs and detailed
empirical results are given in the Appendix.

We use the vec operator of the matrix $A$, $vec(A):=(a_{1}^{\prime}\ldots
a_{k}^{\prime})^{\prime}\in\mathbb{R}^{mk}$ for an $m\times k$ dimensional
matrix $A=(a_{1},\ldots,a_{k}).$ For a symmetric $m\times m$ dimensional
matrix $A,$ we also use the $m^{2}\times\frac{1}{2}m(m+1)$ dimensional,
so-called, duplication matrix $D_{m}$ which selects the $\frac{1}{2}m(m+1)$
unique elements of $A$ in the $\frac{1}{2}m(m+1)$ dimensional vector $vech(A)$
that vectorizes only the lower triangular part of $A$:
\[
vech(A)=(D_{m}^{\prime}D_{m})^{-1}D_{m}^{\prime}vec(A)\text{ \ \ \ and
\ \ }vec(A)=D_{m}vech(A).
\]


\section{A Test for Kronecker Product Structure Covariance Matrix}

We propose a test for a covariance matrix $R\in\mathbb{R}^{kp\times kp}$ to
have KPS, where
\begin{equation}
\begin{array}
[c]{c}
R:=E\left(  \frac{1}{n}\sum_{i=1}^{n}f_{i}f_{i}^{\prime}\right)  ,
\end{array}
\label{eq: cov}
\end{equation}
for mean zero, independently distributed random vectors $f_{i}\in
\mathbb{R}^{kp},$ $i=1,\ldots,n,$ which satisfy
\begin{equation}
f_{i}:=(V_{i}\otimes Z_{i}) \label{eq: fi}
\end{equation}
with $V_{i}\in\mathbb{R}^{p}$ and $Z_{i}\in\mathbb{R}^{k}$ uncorrelated random
vectors.\footnote{The matrix $R$ can depend on the sample size $n$ but for
simplicity of notation we do not index $R$ by $n.$} The specification of
$f_{i}$ fits, for example, a setting where $V_{i}$ contains the errors of a
number of regression equations and $Z_{i}$ contains the regressors, so that
$R$ is then the covariance matrix of the sample covariance between these
errors and the regressors.

It follows that the covariance matrix has a block structure
\begin{equation}
R:=\left(
\begin{array}
[c]{ccc}
R_{11} & \cdots & R_{1p}\\
\vdots & \ddots & \vdots\\
R_{p1} & \cdots & R_{pp}
\end{array}
\right)  , \label{eq: block}
\end{equation}
where $R_{jl}\in\mathbb{R}^{k\times k},$ $j,l=1,\ldots,p$. Because
$R_{jl}=E\left(  \frac{1}{n}\sum_{i=1}^{n}V_{ij}V_{il}Z_{i}Z_{i}^{\prime
}\right)  \allowbreak=R_{lj}^{\prime},$ for $V_{i}=(V_{i1}\ldots
V_{ip})^{\prime},$ it follows that $R_{jl}$ is symmetric. We are interested in
testing if the covariance matrix $R$ has KPS:
\begin{equation}
\text{H}_{0}:R=G_{1}\otimes G_{2} \label{eq: KPS hyp}
\end{equation}
with $G_{1}\in\mathbb{R}^{p\times p}$ and $G_{2}\in\mathbb{R}^{k\times k}$
symmetric positive definite matrices, against the alternative hypothesis of
not having KPS. For normalization purposes, we set one diagonal element equal
to one (say the upper left element of $G_{1}$
).\footnote{\label{foot: normalization}Normalizing $G_{1,11}$ to one is an
obvious normalization because $G_{1}$ is a positive definite matrix (because
$(G_{1}\otimes G_{2})$ is positive definite) so its diagonal elements are all
strictly larger than zero. The normalization does therefore not imply a
restriction.} When $p=1$ or $k=1$ the null is always true and from now on we
assume that $\min\{p,k\}\geq2.$ To measure the distance of the sample
covariance matrix estimator from a KPS\ covariance matrix, we use a convenient
(invertible) transformation proposed by \cite{vanloanpit93}.

\noindent For a matrix $A\in\mathbb{R}^{kp\times kp}$ with block structure as
in (\ref{eq: block}) define
\begin{equation}
\mathcal{R}\left(  A\right)  :=
\begin{pmatrix}
A_{1}\\
\vdots\\
A_{p}
\end{pmatrix}
\in\mathbb{R}^{p^{2}\times k^{2}}\quad\text{with }A_{j}:=
\begin{pmatrix}
vec\left(  A_{1j}\right)  ^{\prime}\\
\vdots\\
vec\left(  A_{pj}\right)  ^{\prime}
\end{pmatrix}
\in\mathbb{R}^{p\times k^{2}}, \label{eq: trans cov}
\end{equation}
for $j=1,...,p.$ One can easily show that
\begin{equation}
\mathcal{R}\left(  G_{1}\otimes G_{2}\right)  =vec(G_{1})vec(G_{2})^{\prime}
\label{eq: red rank cov}
\end{equation}
and by Theorem 2.1 in \cite{vanloanpit93}, we have
\[
\left\Vert R-G_{1}\otimes G_{2}\right\Vert _{F}=\allowbreak\left\Vert
\mathcal{R}(R)-\allowbreak vec(G_{1})vec(G_{2})^{\prime}\right\Vert _{F},
\]
with $\left\Vert .\right\Vert _{F}$ the Frobenius or trace norm of a matrix,
$\left\Vert A\right\Vert _{F}^{2}:=tr(A^{\prime}A)=vec(A)^{\prime}vec(A),$ for
any rectangular matrix $A$. Because $\mathcal{R}\left(  G_{1}\otimes
G_{2}\right)  $ is a matrix of rank one, when testing for a KPS, it is more
convenient to test for the rank of $\mathcal{R}\left(  R\right)  $ to be one
instead of directly testing for KPS\ of $R$.

Consider the covariance matrix estimator
\begin{equation}
\begin{array}
[c]{c}
\hat{R}:=\frac{1}{n}
{\textstyle\sum\limits_{i=1}^{n}}
\hat{f}_{i}\hat{f}_{i}^{\prime}\in\mathbb{R}^{kp\times kp}
\end{array}
\label{eq: sample cov}
\end{equation}
which uses sample values $\hat{f}_{i}:=\hat{V}_{i}\otimes Z_{i}$ of the random
vectors $f_{i}$ for some estimated residuals $\hat{V}_{i}.$ We assume that
$\hat{f}_{i}=f_{i}+o_{p}(1),$ uniformly over $i=1,\ldots,n,$ as $n\rightarrow
\infty.$ Define the distance from a KPS covariance matrix by the Frobenius
norm
\begin{equation}
DS:=\min_{G_{1}>0,G_{2}>0, G_{1,11}=1}\left\Vert \mathcal{R}(\hat{R})-vec(G_{1}
)vec(G_{2})^{\prime}\right\Vert _{F}, \label{eq: dis KPS}
\end{equation}
where $G_{1},$ $G_{2}>0$ indicates that $G_{1}\in\mathbb{R}^{p\times p}$ and
$G_{2}\in\mathbb{R}^{k\times k}$ are positive definite symmetric matrices, and $G_{1,11}=1$ states that the upper left element of $G_1$ is normalized to 1.

We test for $\mathcal{R}(\hat{R})$ being a rank one matrix using the
\cite{kleibergen2006grr} rank statistic. To describe the
\cite{kleibergen2006grr} rank statistic consider first a singular value
decomposition (SVD) of $\mathcal{R}(\hat{R})$:
\begin{equation}
\mathcal{R(}\hat{R})=\hat{L}\hat{\Sigma}\hat{N}^{\prime},
\label{eq: SVD general}
\end{equation}
where $\hat{\Sigma}:=diag(\hat{\sigma}_{1},\ldots,\hat{\sigma}_{\min
(p^{2},k^{2})})$ denotes a $p^{2}\times k^{2}$ dimensional diagonal matrix
with the singular values $\hat{\sigma}_{j}$ ($j=1,...,\min(p^{2},k^{2})$) on
the main diagonal ordered non-increasingly, and with $\hat{L}\in
\mathbb{R}^{p^{2}\times p^{2}}$ and $\hat{N}\in\mathbb{R}^{k^{2}\times k^{2}}$
orthonormal matrices. Decompose
\begin{equation}
\hat{L}:=
\begin{pmatrix}
\hat{L}_{11} & \hat{L}_{12}\\
\hat{L}_{21} & \hat{L}_{22}
\end{pmatrix}
=\left(  \hat{L}_{1}\text{ }\vdots\text{ }\hat{L}_{2}\right)  ,\text{ }
\hat{\Sigma}:=
\begin{pmatrix}
\hat{\sigma}_{1} & 0\\
0 & \hat{\Sigma}_{2}
\end{pmatrix}
,\text{ }\hat{N}:=
\begin{pmatrix}
\hat{N}_{11} & \hat{N}_{12}\\
\hat{N}_{21} & \hat{N}_{22}
\end{pmatrix}
=(\hat{N}_{1}\text{ }\vdots\text{ }\hat{N}_{2}), \label{eq: svdpar}
\end{equation}
with $\hat{L}_{11}:1\times1,$ $\hat{L}_{12}:1\times(p^{2}-1),$ $\hat{L}
_{21}:(p^{2}-1)\times1,$ $\hat{L}_{22}:(p^{2}-1)\times(p^{2}-1),$ $\hat
{\sigma}_{1}:1\times1,$ $\hat{\Sigma}_{2}:(p^{2}-1)\times(k^{2}-1),$ $\hat
{N}_{11}:1\times1,$ $\hat{N}_{12}:1\times(k^{2}-1),$ $\hat{N}_{21}
:(k^{2}-1)\times1,$ $\hat{N}_{22}:(k^{2}-1)\times(k^{2}-1)$ dimensional matrices.

\begin{theorem}
\label{th: frob norm}Suppose $\hat{R}$ is positive definite. The distance
measure DS in (\ref{eq: dis KPS}) equals the square root of the sum of squares
of all but the largest singular value of $\mathcal{R}(\hat{R})\in
\mathbb{R}^{p^{2}\times k^{2}},$ i.e. $DS^{2}=\sum_{i=2}^{\min(p^{2},k^{2}
)}\hat{\sigma}_{i}^{2},$ where $\hat{\sigma}_{1}\geq...\geq\hat{\sigma}
_{\min(p^{2},k^{2})}$ are the ordered singular values of $\mathcal{R}(\hat
{R}).$ Furthermore, if $\hat{\sigma}_{1}>\hat{\sigma}_{2},$ then the positive definite symmetric
minimizers $\hat{G}_{1},\hat{G}_{2}$ of (\ref{eq: dis KPS}) have the following unique expression:
\begin{equation}
\begin{array}
[c]{ll}
vec(\hat{G}_{1}):=\hat{L}_{1}/\hat{L}_{11} & \in\mathbb{R}^{p^{2}\times1},\\
vec(\hat{G}_{2})^{\prime}:=\hat{L}_{11}\hat{\sigma}_{1}\hat{N}_{1}^{\prime
}\quad & \in\mathbb{R}^{1\times k^{2}}.
\end{array}
\label{eq: g1g2 est}
\end{equation}

\end{theorem}

\begin{proof}
See the Appendix.\smallskip
\end{proof}

If $R$ is positive definite, and $\hat{R}\overset{p}{\rightarrow}R$, then
$\hat{R}$ will be positive definite with probability approaching one
(w.p.a.1). The choice of normalization in (\ref{eq: g1g2 est}) conforms with
the normalization of $G_{1},G_{2}$ in (\ref{eq: KPS hyp}) discussed in
Footnote \ref{foot: normalization} above.

\paragraph{The KPST statistic}

We use the distance between $\mathcal{R}(\hat{R})$ and a matrix of rank one to
test for a KPS of $R$. The test is based on the limiting distribution of the
unique elements of $\hat{R}$ or equivalently $\mathcal{R}(\hat{R}).$ These
elements result from using the $k^{2}\times\frac{1}{2}k(k+1)$ and $p^{2}
\times\frac{1}{2}p(p+1)$ dimensional duplication matrices $D_{k}$ and
$D_{p}:$
\begin{equation}
\begin{array}
[c]{rl}
\mathcal{R}(\hat{R})= & \mathcal{R}\left(  \frac{1}{n}\sum_{i=1}^{n}(\hat
{V}_{i}\hat{V}_{i}^{\prime}\otimes Z_{i}Z_{i}^{\prime})\right) \\
= & \frac{1}{n}\sum_{i=1}^{n}vec(\hat{V}_{i}\hat{V}_{i}^{\prime}
)vec(Z_{i}Z_{i}^{\prime})^{\prime}\\
= & D_{p}\hat{R}^{\ast}D_{k}^{\prime},
\end{array}
\label{eq: R-spec}
\end{equation}
with
\begin{equation}
\begin{array}
[c]{c}
\hat{R}^{\ast}:=\frac{1}{n}\sum_{i=1}^{n}vech(\hat{V}_{i}\hat{V}_{i}^{\prime
})vech(Z_{i}Z_{i}^{\prime})^{\prime}.
\end{array}
\label{eq: R star}
\end{equation}
The $\frac{1}{2}p(p+1)\times\frac{1}{2}k(k+1)$ dimensional matrix $\hat
{R}^{\ast}$ contains the unique elements of $\hat{R}$ and $\mathcal{R}(\hat
{R}).$ We assume $vec(\hat{R}^{\ast})$ satisfies a central limit theorem:
\begin{equation}
\begin{array}
[c]{cc}
\sqrt{n}(vec(\hat{R}^{\ast})-vec(R^{\ast})) & \underset{d}{\rightarrow}
\psi=vec(\Psi),
\end{array}
\label{eq: CLT R_n}
\end{equation}
with $\psi\sim N(0,V_{R^{\ast}}),$ $\Psi$ a $\frac{1}{2}p(p+1)\times\frac
{1}{2}k(k+1)$ dimensional normally distributed random matrix and
\begin{equation}
\begin{array}
[c]{rl}
R^{\ast}:= & E\left(  \frac{1}{n}\sum_{i=1}^{n}vech(V_{i}V_{i}^{\prime
})vech(Z_{i}Z_{i}^{\prime})^{\prime}\right)  ,\\
V_{R^{\ast}}:= & \lim_{n\rightarrow\infty}\left[  E\left(  \frac{1}{n}
\sum_{i=1}^{n}\left(  vech(Z_{i}Z_{i}^{\prime})vech(Z_{i}Z_{i}^{\prime
})^{\prime}\otimes vech(V_{i}V_{i}^{\prime})vech(V_{i}V_{i}^{\prime})^{\prime
}\right)  \right)  \right. \\
& \left.  -E\left(  vec(\frac{1}{n}\sum_{i=1}^{n}vech(V_{i}V_{i}^{\prime
})vech(Z_{i}Z_{i}^{\prime})^{\prime})\right)  E\left(  vec(\frac{1}{n}
\sum_{i=1}^{n}vech(V_{i}V_{i}^{\prime})vech(Z_{i}Z_{i}^{\prime})^{\prime
})\right)  ^{\prime}\right]  .
\end{array}
\label{eq: lim_mean_var1n}
\end{equation}
In fact, we assume a slightly stronger result, namely, that $\hat{R}^{\ast
}=R^{\ast}+\frac{1}{\sqrt{n}}\Psi+o_{p}(n^{-\frac{1}{2}}),$ holds. A central
limit theorem (\ref{eq: CLT R_n}) for (possibly) non-identical distributed
independent random variables holds under mild conditions, e.g. under the
Liapounov's or Lindeberg's condition, see \cite{wh84}.


Define
\begin{equation}
\begin{array}
[c]{rll}
\hat{\Lambda}:= & \left(  \hat{L}_{22}\hat{L}_{22}^{\prime}\right)
^{-1/2}\hat{L}_{22}\hat{\Sigma}_{2}\hat{N}_{22}^{\prime}\left(  \hat{N}
_{22}\hat{N}_{22}^{\prime}\right)  ^{-1/2} & :(p^{2}-1)\times(k^{2}-1).
\end{array}
\label{eq: lambda}
\end{equation}
It can be shown that $\hat{\Lambda}=vec(\hat{G}_{1})_{\perp}^{\prime
}\mathcal{R}(\hat{R})vec(\hat{G}_{2})_{\perp},$ where
\[
\begin{array}
[c]{ll}
vec(\hat{G}_{1})_{\perp}:=\hat{L}_{2}\hat{L}_{22}^{-1}(\hat{L}_{22}\hat{L}
_{22}^{\prime})^{1/2} & :p^{2}\times(p^{2}-1),\\
vec(\hat{G}_{2})_{\perp}^{\prime}:=\left(  \hat{N}_{22}\hat{N}_{22}^{\prime
}\right)  ^{1/2}\hat{N}_{22}^{\prime-1}\hat{N}_{2}^{\prime} & :(k^{2}-1)\times
k^{2},
\end{array}
\]
see
\cite[page 102]{kleibergen2006grr}. We then have
\begin{equation}
\mathcal{R}(\hat{R})=vec(\hat{G}_{1})vec(\hat{G}_{2})^{\prime}+vec(\hat{G}
_{1})_{\perp}\hat{\Lambda}vec(\hat{G}_{2})_{\perp}^{\prime}.
\label{eq: rank speci}
\end{equation}


Using $\mathcal{R(}R),$ our hypothesis of interest H$_{0}$ (\ref{eq: KPS hyp})
is transformed into
\begin{equation}
\text{H}_{0}:\mathcal{R(}R)=vec(G_{1})vec(G_{2})^{\prime}\text{ or H}
_{0}:vec(G_{1})_{\perp}^{\prime}\mathcal{R(}R)vec(G_{2})_{\perp}=0,
\label{eq: hyp rank}
\end{equation}
where $vec(G_{1})_{\perp}$ and $vec(G_{2})_{\perp}$ are $p^{2}\times(p^{2}-1)$
and $k^{2}\times(k^{2}-1)$ dimensional matrices that contain the orthogonal
complements of $vec(G_{1})$ and $vec(G_{2}),$ $vec(G_{1})_{\perp}^{\prime
}vec(G_{1})\equiv0,$ $vec(G_{1})_{\perp}^{\prime}vec(G_{1})_{\perp}\equiv
I_{p^{2}-1},$ $vec(G_{2})_{\perp}^{\prime}vec(G_{2})\equiv0,$ $vec(G_{2}
)_{\perp}^{\prime}vec(G_{2})_{\perp}\equiv I_{k^{2}-1}.$ The KPST test uses
the sample analog of the last component in (\ref{eq: hyp rank}) to test
H$_{0}.$ It further results from identifying $vec(G_{1})$ and $vec(G_{2})$
using the eigenvectors associated with the first singular value of
$\mathcal{R(}R)$.

The \cite{kleibergen2006grr} rank test statistic is a quadratic form of the
vectorization of $\hat{\Lambda}$ in (\ref{eq: lambda}). Its specification
directly extends to the new KPS test but because the covariance matrix of
$vec\left(  \mathcal{R}(\hat{R})\right)  $ is singular, the (degenerate)
asymptotic normal distribution of $vec(\hat{\Lambda})$ and the resulting
degrees of freedom parameter of the $\chi^{2}$ limiting distribution of the
\cite{kleibergen2006grr} rank test statistic are not obvious.

We define the statistic KPST\ for testing H$_{0}$ in (\ref{eq: KPS hyp}) as
\begin{equation}
KPST:=n\times vec\left(  \hat{\Lambda}\right)  ^{\prime}\left(  \hat
{J}^{\prime}\hat{V}\hat{J}\right)  ^{-}vec\left(  \hat{\Lambda}\right)  ,
\label{eq: kpst}
\end{equation}
where
\begin{equation}
\hat{J}:=\left(  vec(\hat{G}_{2})_{\perp}\otimes vec(\hat{G}_{1})_{\perp
}\right)  ,\quad\hat{V}:=\widehat{\text{cov}}\left(  vec\left(  \mathcal{R}
(\hat{R})\right)  \right)  \in\mathbb{R}^{p^{2}k^{2}\times p^{2}k^{2}},
\label{eq: Jhat Vhat}
\end{equation}
and
\begin{equation}
\begin{array}
[c]{rl}
\widehat{\text{cov}}\left(  vec\left(  \mathcal{R}(\hat{R})\right)  \right)
= & \frac{1}{n}\sum_{i=1}^{n}\left(  vec(Z_{i}Z_{i}^{\prime})vec(Z_{i}
Z_{i}^{\prime})^{\prime}\otimes vec(\hat{V}_{i}\hat{V}_{i}^{\prime}
)vec(\hat{V}_{i}\hat{V}_{i}^{\prime})^{\prime}\right)  \\
& -vec\left(  \mathcal{R}(\hat{R})\right)  vec\left(  \mathcal{R}(\hat
{R})\right)  ^{\prime}\\
= & (D_{k}\otimes D_{p})\widehat{\text{cov}}\left(  vec\left(  \hat{R}^{\ast
}\right)  \right)  (D_{k}\otimes D_{p})^{\prime},\\
\widehat{\text{cov}}\left(  vec\left(  \hat{R}^{\ast}\right)  \right)  = &
\frac{1}{n}\sum_{i=1}^{n}\left(  vech(Z_{i}Z_{i}^{\prime})vech(Z_{i}
Z_{i}^{\prime})^{\prime}\otimes vech(\hat{V}_{i}\hat{V}_{i}^{\prime}
)vech(\hat{V}_{i}\hat{V}_{i}^{\prime})^{\prime}\right) \\ & -vec\left(  \hat
{R}^{\ast}\right)  vec\left(  \hat{R}^{\ast}\right)  ^{\prime}.
\end{array}
\label{eq: covhat}
\end{equation}

In the Appendix it is shown that the KPST statistic in (\ref{eq: kpst}) can
be simplified as follows:
\begin{equation}
\begin{array}
[c]{rl}
KPST= & n\times\left(  vec\left(  \hat{\Sigma}_{2}\right)  \right)  ^{\prime
}\left[  (\hat{N}_{2}\otimes\hat{L}_{2})^{\prime}\hat{V}\left(  \hat{N}
_{2}\otimes\hat{L}_{2}\right)  \right]  ^{-}\left(  vec\left(  \hat{\Sigma
}_{2}\right)  \right)  .
\end{array}
\label{eq: simple KPST}
\end{equation}
This provides an expression for KPST which is easier to compute. On the other
hand, it cannot be directly used to obtain the $\chi^{2}$ limiting
distribution because $\hat{\Sigma}_{2}$ does not have an asymptotic normal
distribution while $vec(\hat{\Lambda})$ does.

\paragraph{\textbf{The }KPST$^{\ast}$\textbf{ statistic}}

For comparison, we now introduce an alternative test statistic KPST$^{\ast}$
that fits more naturally into the \cite{kleibergen2006grr} framework. However,
unlike KPST, KPST$^{\ast}$ turns out not to be invariant to orthonormal
transformations of the data. Because $vec(G_{1})=D_{p}vech(G_{1}),$
$vec(G_{2})=D_{k}vech(G_{2}),$ the hypothesis of interest (\ref{eq: hyp rank})
can also be specified as:
\begin{equation}
\text{H}_{0}:R^{\ast}=vech(G_{1})vech(G_{2})^{\prime}\text{ or H}
_{0}:vech(G_{1})_{\perp}^{\prime}R^{\ast}vech(G_{2})_{\perp}=0,
\label{eq: hyp rankvech}
\end{equation}
where $vech(G_{1})_{\perp}$ and $vech(G_{2})_{\perp}$ are $\frac{1}
{2}p(p+1)\times(\frac{1}{2}p(p+1)-1)$ and $\frac{1}{2}k(k+1)\times(\frac{1}
{2}k(k+1)-1)$ dimensional matrices that contain the orthogonal complements of
$vech(G_{1})$ and $vech(G_{2}),$ $vech(G_{1})_{\perp}^{\prime}vech(G_{1}
)=0,$ $vech(G_{1})_{\perp}^{\prime}vech(G_{1})_{\perp}= I_{\frac
{1}{2}p(p+1)-1},$ $vech(G_{2})_{\perp}^{\prime}vech(G_{2})=0,$
$vec(G_{2})_{\perp}^{\prime}\allowbreak vec(G_{2})_{\perp}= I_{\frac
{1}{2}k(k+1)-1}.$ This specification of the hypothesis fits directly in the
setup of the \cite{kleibergen2006grr} rank test because the covariance matrix
of $\hat{R}^{\ast}$ is non-singular. Therefore, the corresponding
specification of $vec(\hat{\Lambda})$ converges to a Normally distributed
random vector. The specification of null hypothesis in (\ref{eq: hyp rankvech}
) allows us to easily infer the number of restrictions tested, which equals
$\left(  \frac{1}{2}k(k+1)-1\right)  \left(  \frac{1}{2}p(p+1)-1\right)  $,
but the resulting rank statistic does not equal KPST in (\ref{eq: kpst}).
Specifically, define the SVD of $\hat{R}^{\ast}
=\hat{L}^{\ast}\hat{\Sigma}^{\ast}\hat{N}^{\ast\prime}$, where
\[
\hat{L}^{\ast}:=
\begin{pmatrix}
\hat{L}_{11}^{\ast} & \hat{L}_{12}^{\ast}\\
\hat{L}_{21}^{\ast} & \hat{L}_{22}^{\ast}
\end{pmatrix}
,\text{ }\hat{\Sigma}^{\ast}:=
\begin{pmatrix}
\hat{\sigma}_{1}^{\ast} & 0\\
0 & \hat{\Sigma}_{2}^{\ast}
\end{pmatrix}
,\text{ }\hat{N}^{\ast}:=
\begin{pmatrix}
\hat{N}_{11}^{\ast} & \hat{N}_{12}^{\ast}\\
\hat{N}_{21}^{\ast} & \hat{N}_{22}^{\ast}
\end{pmatrix}
,
\]
with $\hat{L}_{11}^{\ast}:1\times1,$ $\hat{L}_{12}^{\ast}:1\times\left(
\frac{1}{2}p(p+1)-1\right)  ,$ $\hat{L}_{21}^{\ast}:(\frac{1}{2}
p(p+1)-1)\times1,$ $\hat{L}_{22}^{\ast}:(\frac{1}{2}p(p+1)-1)\times(\frac
{1}{2}p(p+1)-1),$ $\hat{\sigma}_{1}^{\ast}:1\times1,$ $\hat{\Sigma}_{2}^{\ast
}:(\frac{1}{2}p(p+1)-1)\times(\frac{1}{2}k(k+1)-1),$ $\hat{N}_{11}^{\ast
}:1\times1,$ $\hat{N}_{12}^{\ast}:1\times(\frac{1}{2}k(k+1)-1),$ $\hat{N}
_{21}^{\ast}:(\frac{1}{2}k(k+1)^{2}-1)\times1,$ $\hat{N}_{22}^{\ast}:(\frac
{1}{2}k(k+1)-1)\times(\frac{1}{2}k(k+1)-1)$ dimensional matrices and $\hat
{L}_{2}^{\ast}=(\hat{L}_{12}^{\ast\prime}$ $\vdots$ $\hat{L}_{22}^{\ast\prime
})^{\prime},$ $\hat{N}_{2}^{\ast}=(\hat{N}_{12}^{\ast\prime}$ $\vdots$
$\hat{N}_{22}^{\ast\prime})^{\prime}$. The \cite{kleibergen2006grr} statistic
for testing (\ref{eq: hyp rankvech}) using $\hat{R}^{\ast}$ is
\begin{equation}
\begin{array}
[c]{l}
KPST^{\ast}:=n\times vec\left(  \hat{\Lambda}^{\ast}\right)  ^{\prime}\left(
\hat{J}^{\ast\prime}\hat{V}^{\ast}\hat{J}^{\ast}\right)  ^{-}vec\left(
\hat{\Lambda}^{\ast}\right)  ,\text{ where}\\
\hat{\Lambda}^{\ast}:=\left(  \hat{L}_{22}^{\ast}\hat{L}_{22}^{\ast\prime
}\right)  ^{-1/2}\hat{L}_{22}^{\ast}\hat{\Sigma}_{2}^{\ast}\hat{N}_{22}
^{\ast\prime}\left(  \hat{N}_{22}^{\ast}\hat{N}_{22}^{\ast\prime}\right)
^{-1/2},\\
\hat{J}^{\ast}:=\left(  \left(  N_{22}^{\ast}N_{22}^{\ast\prime}\right)
^{1/2}N_{22}^{\ast\prime-1}\left[  N_{12}^{\ast\prime}\text{ }\vdots\text{
}N_{22}^{\ast\prime}\right]  \otimes\left(  L_{22}^{\ast}L_{22}^{\ast\prime
}\right)  ^{1/2}L_{22}^{\ast\prime-1}\left[  L_{12}^{\ast\prime}\text{ }
\vdots\text{ }L_{22}^{\ast\prime}\right]  \right)  ,\\
\hat{V}^{\ast}:=\widehat{\text{cov}}\left(  vec\left(  \hat{R}^{\ast}\right)
\right)  \in\mathbb{R}^{\left(  \frac{1}{4}p(p+1)k(k+1)\right)  \times\left(
\frac{1}{4}p(p+1)k(k+1)\right)  },
\end{array}
\label{eq: kpst*}
\end{equation}
see Corollary 1 in \cite{kleibergen2006grr}.

\paragraph{Asymptotic theory and invariance to orthonormal transformations}

The statistics KPST in (\ref{eq: kpst}) and KPST$^{\ast}$ in
(\ref{eq: kpst*}) are not identical, and, unlike the proposed KPST
statistic, tests of the KPS hypothesis based on the KPST$^{\ast}$ statistic
are not invariant to orthonormal transformations of the data, as stated in the
following Theorem.\smallskip

\begin{theorem}
\label{th: kps test}Assume $E\left(  \left\Vert f_{i}\right\Vert ^{8}\right)
<\kappa$ for some $\kappa<\infty,$ $\hat{f}_{i}=f_{i}+o_{p}(1),$ uniformly for
$i=1,\ldots,n,$ as $n\rightarrow\infty,$ and the central limit theorem in
(\ref{eq: CLT R_n}) holds in the slightly stronger version $\hat{R}^{\ast
}=R^{\ast}+\frac{1}{\sqrt{n}}\Psi+o_{p}(n^{-\frac{1}{2}})$. Then, under
H$_{0},$ for KPST and KPST$^{\ast}$ defined in (\ref{eq: kpst}) and
(\ref{eq: kpst*}), respectively, the following hold:\newline\textbf{a.}\[
KPST\underset{d}{\rightarrow}\chi_{df}^{2}
\]  as
$n\rightarrow\infty$ (for fixed $p$ and $k$) with degrees of freedom
\begin{equation}
\begin{array}
[c]{c}
df:=\left(  \frac{1}{2}k(k+1)-1\right)  \left(  \frac{1}{2}p(p+1)-1\right)  .
\end{array}
\label{eq: a_spec}
\end{equation}
\newline\textbf{b. }
\[KPST^{\ast}\underset{d}{\rightarrow}\chi_{df}^{2}\]  as
$n\rightarrow\infty$ (for fixed $p$ and $k$) with $df$ as given in
(\ref{eq: a_spec}). \newline\textbf{c. }The statistics KPST and KPST$^{\ast}$ are in general not
numerically identical. While KPST is invariant to orthonormal
transformations of the data in $\hat{V}_{i}$ and $Z_{i},$ KPST$^{\ast}$ is not
invariant to such transformations.\newline\textbf{d. }For sequences $p,$ $k,$
$n$ that satisfy
\begin{equation}
\begin{array}
[c]{c}
\frac{(pk)^{16}}{n^{3}}\rightarrow0,
\end{array}
\label{eq: joint conv}
\end{equation}
we have
\begin{equation}
\begin{array}
[c]{c}
\lim_{n,p,k\rightarrow\infty}\Pr\left[  KPST<\chi_{df,1-\alpha}^{2}\right]
\leq\alpha,
\end{array}
\label{eq: lim many pkn}
\end{equation}
where $\chi_{df,1-\alpha}^{2}$ denotes the $1-\alpha$ quantile of a $\chi
_{df}^{2}$ distribution.
\end{theorem}

\begin{proof}
see the Appendix.\footnote{We thank an anonymous associate editor for pointing
at the vech operator and duplication matrix to simplify the proof and
exposition.}\smallskip
\end{proof}

We define the new KPST test as follows: it rejects H$_{0}$ in
(\ref{eq: KPS hyp}) at nominal size $\alpha$ if
\begin{equation}
KPST>\chi_{df,1-\alpha}^{2}. \label{eq: KPST test}
\end{equation}
Based on Theorem \ref{th: kps test}a and d, the resulting test has limiting
null rejection probability bounded by $\alpha.$

Theorem \ref{th: kps test}c shows that the rank-one tests KPST and
KPST$^{\ast}$ that are based on $\mathcal{R(}\hat{R})$ and $\hat{R}^{\ast}$
respectively are not identical even though they are testing the same
underlying hypothesis that the covariance $R$ matrix has KPS. This difference
occurs because these are Wald tests, and Wald statistics are in general not
invariant to non-linear transformations.

Theorem \ref{th: kps test}d provides a sufficient condition for uniform
convergence of $\hat{\Lambda}$ and its covariance matrix estimator for
settings where $p,$ $k,$ and $n$ jointly go to infinity so the main results
for the limiting distribution of KPST remain unaltered. It is needed to assess
the validity of the asymptotic approximation for settings where $p$ and $k$
are relatively large compared to the number of observations $n.$

The conditions in Theorem \ref{th: kps test}d are weaker than those in
\cite{newey2009gmm}. \cite{newey2009gmm} prove the validity of the asymptotic
approximation of test statistics where the number of observations grows faster
than the cube of the number of moment restrictions. The number of moment
restrictions here is proportional to $(pk)^{2}$ so their rate would be
$(pk)^{6}/n\rightarrow0$ which is more restrictive than the rate in
(\ref{eq: joint conv}).

\paragraph{Invariance to nonsingular transformations}

Theorem \ref{th: kps test}c shows the KPST is invariant to orthonormal
transformations, but it is still not invariant to general nonsingular
transformations of the data. To ensure invariance to nonsingular
transformations, we need to normalize the data as in \cite{kleibergen2006grr}.
Specifically, the KPST statistic is computed using the moment vector
\begin{equation}
\hat{f}_{i}=C_{1}^{\prime}\widehat{V}_{i}\otimes C_{2}^{\prime}Z_{i},
\label{eq: f par boot}
\end{equation}
where $C_{1}$ and $C_{2}$ are the Choleski factors of the inverse of the
second moments of $\hat{V}_{i}$ and $Z_{i},$ i.e., $C_{1}C_{1}^{\prime
}=\left(  \frac{1}{n}\sum_{i=1}^{n}\widehat{V}_{i}\widehat{V}_{i}^{\prime
}\right)  ^{-1}$ and $C_{2}C_{2}^{\prime}=\left(  \frac{1}{n}\sum_{i=1}
^{n}Z_{i}Z_{i}^{\prime}\right)  ^{-1}$. To see why this normalization yields
invariance, let $A$ be a nonsingular $k\times k$ matrix, and define the
transformed instruments $Z_{A,i}:=AZ_{i}.$ Let $C_{2A}$ denote the Choleski
factor of $\left(  \frac{1}{n}\sum_{i=1}^{n}AZ_{i}Z_{i}^{\prime}A^{\prime
}\right)  ^{-1}.$ The KPST statistic with the original instruments $Z_{i}$ is
computed using $C_{2}^{\prime}Z_{i}$ in the moment vector
(\ref{eq: f par boot}), while the KPST with the transformed instruments
$AZ_{i}$ uses $C_{2A}^{\prime}AZ_{i}$ in the same formula
(\ref{eq: f par boot}). Therefore, the transformation from $C_{2}^{\prime
}Z_{i}$ to $C_{2A}^{\prime}Z_{A,i}$ is given by $T_{A}:=\allowbreak
C_{2A}^{\prime}AC_{2}^{-1\prime},$ i.e., $C_{2A}^{\prime}Z_{A,i}
\allowbreak=\allowbreak C_{2A}^{\prime}AZ_{i}\allowbreak=\allowbreak
T_{A}\left(  C_{2}^{\prime}Z_{i}\right)  .$ Now, observe that $T_{A}$ is an
orthonormal matrix, because $T_{A}^{\prime}T_{A}\allowbreak=\allowbreak
C_{2}^{-1}A^{\prime}C_{2A}\allowbreak C_{2A}^{\prime}AC_{2}^{-1\prime
}\allowbreak=\allowbreak C_{2}^{-1}A^{\prime}\allowbreak\left(  \frac{1}
{n}\sum_{i=1}^{n}AZ_{i}Z_{i}^{\prime}A^{\prime}\right)  ^{-1}\allowbreak
AC_{2}^{-1\prime}\allowbreak=\allowbreak C_{2}^{-1}\allowbreak\left(  \frac
{1}{n}\sum_{i=1}^{n}Z_{i}Z_{i}^{\prime}\right)  ^{-1}\allowbreak
C_{2}^{-1\prime}\allowbreak=\allowbreak I_{k}.$ Hence, invariance follows from
Theorem \ref{th: kps test}c. The exact same argument can be made about
rotations of the reduced form errors $\hat{V}.$

\paragraph{Clustered data}

In case of clustered data, we assume there are $n$ clusters of $N_{i}$
observations each, so the total number of data points is $\sum_{i=1}^{n}
N_{i}:$
\begin{equation}
\begin{array}
[c]{c}
f_{i}=\sum_{j=1}^{N_{i}}f_{ij},
\end{array}
\label{eq: f_i}
\end{equation}
for mean zero $kp$ dimensional random vectors $f_{ij},$ $j=1,\ldots,N_{i},$
$i=1,\ldots,n.$ Observations $f_{ij}$ within cluster $i$ can be arbitrarily
dependent, i.e., $E\left(  f_{ij}f_{is}\right)  $ is unrestricted for all $j,$
$s=1,...,N_{i},$ while observations across clusters are independent. The
$kp\times kp$ dimensional (positive semi-definite) covariance matrix of the
sample moments then results as:
\begin{equation}
\begin{array}
[c]{c}
R=\frac{1}{n}\sum_{i=1}^{n}E(f_{i}f_{i}^{\prime}).
\end{array}
\label{eq: cluster covariance}
\end{equation}


\section{Limiting distribution of KPST under local alternatives}

To analyze the power of KPST under local alternatives, we construct the
limiting distribution of the KPST statistic under alternatives where the
covariance matrix of the moments $R\in\mathbb{R}^{kp\times kp}$ is local to
KPS:
\begin{equation}
\begin{array}
[c]{c}
\text{H}_{1}:R=(G_{1}\otimes G_{2})+\frac{1}{\sqrt{n}}A_{0},
\end{array}
\label{eq: R_n local}
\end{equation}
where $G_{1}\in\mathbb{R}^{p\times p}$ and $G_{2}\in\mathbb{R}^{k\times k}$
are symmetric positive definite matrices, and $A_{0} \in\mathbb{R}^{kp\times kp}$
is a fixed symmetric matrix. The
best-fitting KPS approximation of $R$ under H$_{1}$ w.r.t. Frobenius norm,
defined as $\bar{G}_{1,n}\otimes\bar{G}_{2,n}$, where $\bar{G}_{1,n},\bar
{G}_{2,n}$ solve $\min_{\bar{G}_{1}>0,\bar{G}_{2}>0}\left\Vert (G_{1}\otimes
G_{2})\allowbreak+\frac{1}{\sqrt{n}}A_{0}\allowbreak-\bar{G}_{1}\otimes\bar
{G}_{2}\right\Vert _{F}$,$\,$will differ from $G_{1}\otimes G_{2}$. That is,
$\bar{G}_{1,n}\neq G_{1}$ and $\bar{G}_{2,n}\neq G_{2},$ unless $A_{0}$ lies
in the span of the orthogonal complement of $G_{1}\otimes G_{2}$. However,
under the local alternatives (\ref{eq: R_n local}), $\bar{G}_{1,n}\rightarrow
G_{1}$ and $\bar{G}_{2,n}\rightarrow G_{2}$. This needs to be taken into
account when we characterize the asymptotic distribution of the KPST statistic
under the local alternatives in (\ref{eq: R_n local}).

The re-arranged matrix $\mathcal{R}(R)$ under H$_{1}$ is:
\begin{equation}
\begin{array}
[c]{rl}
\mathcal{R}(R)= & vec(G_{1})vec(G_{2})^{\prime}+\frac{1}{\sqrt{n}}
\mathcal{R}(A_{0})\\
= & vec(\bar{G}_{1,n})vec(\bar{G}_{2,n})^{\prime}+
vec(\bar{G}_{1,n})_{\perp}\Lambda_{n}vec(\bar{G}_{2,n})_{\perp}^{\prime},
\end{array}
\label{eq: rearranged A_n}
\end{equation}
with
\begin{equation}
\Lambda _{n}=vec(\bar{G}_{1,n})_{\perp }^{\prime }\mathcal{R}(R)vec(\bar{G}
_{2,n})_{\perp }.  \label{eq: Lambda_n}
\end{equation}
The decomposition in the last line of
(\ref{eq: rearranged A_n}) is identical to the one in (\ref{eq: rank speci}).

\begin{theorem}
\label{th: kps test power}Under local to KPS sequences of covariance
matrices as in (\ref{eq: R_n local}) and for mean zero, independently
distributed random vectors $f_{i}\in \mathbb{R}^{kp}$ with finite eighth
moments,
\begin{equation*}
KPST\underset{d}{\rightarrow }\chi _{df}^{2}(\delta )
\end{equation*}
as $n\rightarrow \infty $ (with $k,p$ fixed), where
\begin{equation}
\begin{array}{rl}
\delta := & vec(a_{0})^{\prime }\left[ \left( vec(G_{2})_{\perp }^{\prime
}\otimes vec(G_{1})_{\perp }^{\prime }\right) (D_{k}\otimes D_{p})V_{R^{\ast
}}\right. \\
& \left. (D_{k}\otimes D_{p})^{\prime }\left( vec(G_{2})_{\perp }\otimes
vec(G_{1})_{\perp }\right) \right] ^{-}vec(a_{0}),
\end{array}
\label{eq: explicit delta}
\end{equation}
$V_{R^{\ast }}$ has been defined in (\ref{eq: lim_mean_var1n}), and
\begin{equation*}
a_{0}:=vec(G_{1})_{\perp }^{\prime }\mathcal{R}(A_{0})vec(G_{2})_{\perp }\in
\mathbb{R}^{\left( \frac{1}{2}k(k+1)-1\right) \times \left( \frac{1}{2}
p(p+1)-1\right) }.
\end{equation*}
\end{theorem}

\begin{proof}
See the Appendix. \smallskip
\end{proof}

\section{Simulation study on size and power}

\paragraph{Size}

We evaluate the accuracy of the limiting distribution in Theorem
\ref{th: kps test} to approximate the finite sample distribution of the
KPST\ statistic. We do so in a small simulation experiment using the linear
regression model:
\begin{equation}
Y_{i}=Z_{i}^{\prime}\Pi+V_{i},\qquad i=1,\ldots,n, \label{eq: regression}
\end{equation}
where $Y_{i}$ is a $p$ dimensional vector of dependent variables, $Z_{i}$ is a
$k$ dimensional vector of explanatory (exogenous) variables and $V_{i}$ is a
$p$ dimensional vector of errors. We further set $\Pi$ to zero (which is
without loss of generality because KPST uses the residual vectors) and
generate the $Z_{i}$'s independently from $N(0,I_{k})$ distributions and
$V_{i}$ given $Z_{i}$ independently from a $N\left(  0,h\left(  Z_{i}\right)
I_{p}\right)  $ distribution. We consider two different specifications of
$h(Z_{i}).$ The first leads to homoskedasticity and has $h(Z_{i})=1$ while the
second leads to (scalar) heteroskedasticity and has $h\left(  Z_{i}\right)
=\left\Vert Z_{i}\right\Vert ^{2}/k.$ For each case, we compute null rejection
probabilities (NRPs) using the three conventional nominal significance levels
of 10\%, 5\% and 1\%. The NRPs are computed using 40,000 Monte Carlo
replications for the KPST\ test that uses chi-square critical values based on
the results from Theorem \ref{th: kps test}. Table \ref{Tab: NRP} reports the
NRPs when the sample size depends on the dimensions $p$ and $k,$ specifically
$n=\left(  kp\right)  ^{16/3}$, in accordance with Theorem \ref{th: kps test}.
We notice only a slight underrejection in some cases, but in the remaining
cases the NRPs are not significantly different from the test's nominal levels.
Table \ref{Tab: NRP 2} reports NRPs with a smaller sample size $n=\left(
pk\right)  ^{4}$. In this case, we find some modest deviations from the
nominal size but these are generally quite small.

To investigate NRPs in smaller samples, Figures \ref{fig: p_2} to
\ref{fig: p_4_5} show the NRPs as a function of the sample size $n$ for
smaller sample sizes than in Tables \ref{Tab: NRP}, \ref{Tab: NRP 2} for
different settings of $p$ and $k.$ Depending on the value of the latter, the
NRPs are close to the nominal level for values of $n$ much smaller than
$\left(  pk\right)  ^{4}.$ For larger values of $pk,$ we therefore do not
(like for the smaller values of $pk)$ show the rejection frequencies all the
way up to $n=(pk)^{\frac{16}{3}},$ i.e. the value indicated by Theorem
\ref{th: kps test}d, but just to $(pk)^{4},$ which is for $p=2,$ $k=7$ at the
bottom right hand side of Figure \ref{fig: p_2}, equal to approximately
40,000, and for $p=5,$ $k=4$ at the bottom right hand side of Figure
\ref{fig: p_3} equal to 160,000 (note that the horizontal axis is in
log-scale). In many cases, the NRPs are still much closer to their nominal
significance levels than indicated by this rate. For example, when $p=k=2$ and
testing at the 5\% significance level, the NRP is close to the nominal level
for sample size of around 100. More striking is that when $p=2$ and $k=5$ the
KPST test at 5\% nominal size has NRPs close to the nominal size for values of
$n$ around 200. Figures \ref{fig: p_2}-\ref{fig: p_4_5} also show that the
KPST test generally over-rejects for small $n$. Moreover, the over-rejection is increasing in the dimensions $k$ and $p$ and can be very substantial for very small $n$, see Figure \ref{fig: p_4_5}, as is the case for any Wald test when the number of restrictions is large relative to the sample size. Therefore, it is of interest to investigate the possibility of small-sample corrections, e.g., following the bootstrap approach of \cite{Chen2019}. From a practical perspective, this over-rejection means that rejection of
KPS with small sample sizes, which happens only a few times in the
applications reported in Section \ref{s: empirical}, could be due to a
significantly higher type 1 error probability than the nominal size of the test.
\footnote{When KPST is used as a pre-test in a two-step procedure, such as the subvector Anderson-Rubin test of \cite{GKM21}, that involves choosing a second-step test that is robust to violation of KPS when the KPST rejects in the first step, over-rejection will only affect the power but not the overall size of the two-step procedure.}

\begin{table}[tbp] \centering
\begin{tabular}
[c]{rrrrr|rrr|rrr}\hline\hline
\multicolumn{5}{r}{\textit{Data Generating Process}:} &
\multicolumn{3}{|c|}{homoskedastic} & \multicolumn{3}{c}{scalar hetero}
\\\hline
p & k & n & a & m & 10\% & 5\% & 1\% & 10\% & 5\% & 1\%\\\hline
2 & 2 & 1626 & 4 & 9 & 10.0 & 5.1 & 1.0 & 9.7 & 4.4 & 0.7\\
2 & 3 & 14130 & 10 & 18 & 10.0 & 5.0 & 0.8 & 9.3 & 4.2 & 0.7\\
2 & 4 & 65536 & 18 & 30 & 9.4 & 5.0 & 0.9 & 9.7 & 4.9 & 0.9\\
2 & 5 & 215444 & 28 & 45 & 9.8 & 4.7 & 0.9 & 9.8 & 5.1 & 1.0\\
3 & 2 & 14130 & 10 & 18 & 10.2 & 5.0 & 0.9 & 10.0 & 4.7 & 0.9\\\hline
3 & 3 & 122827 & 25 & 36 & 9.7 & 4.9 & 1.0 & 9.8 & 5.0 & 0.9\\\hline\hline
\end{tabular}
\caption{Rejection frequencies (in percentages) of KPST test at various
significance levels. $\chi^2_{df}$ critical values.
$[n=(pk)^{16/3}]$, $df$: number of restrictions given in eq.~(\ref{eq:
a_spec}), $m$: number of estimated parameters. Computed using 40,000 MC replications.}\label{Tab: NRP}
\end{table}

\begin{table}[tbp] \centering
\begin{tabular}
[c]{rrrrr|rrr|rrr}\hline\hline
\multicolumn{5}{r}{\textit{Data Generating Process}:} &
\multicolumn{3}{|c|}{homoskedastic} & \multicolumn{3}{c}{scalar hetero}
\\\hline
p & k & n & a & m & 10\% & 5\% & 1\% & 10\% & 5\% & 1\%\\\hline
2 & 2 & 256 & 4 & 9 & 11.2 & 5.3 & 0.9 & 11.4 & 4.8 & 0.5\\
2 & 3 & 1296 & 10 & 18 & 10.2 & 4.9 & 0.9 & 9.3 & 4.0 & 0.5\\
2 & 4 & 4096 & 18 & 30 & 9.9 & 5.1 & 1.0 & 9.1 & 4.2 & 0.8\\
2 & 5 & 10000 & 28 & 45 & 9.7 & 4.6 & 0.8 & 8.8 & 4.0 & 0.6\\
2 & 6 & 20736 & 40 & 63 & 10.0 & 5.1 & 1.0 & 9.5 & 4.5 & 0.7\\
2 & 7 & 38416 & 54 & 84 & 9.8 & 4.8 & 0.9 & 9.5 & 4.5 & 0.8\\
3 & 2 & 1296 & 10 & 18 & 9.9 & 4.8 & 0.7 & 9.0 & 3.7 & 0.5\\
3 & 3 & 6561 & 25 & 36 & 9.8 & 5.0 & 0.9 & 9.6 & 4.4 & 0.7\\
3 & 4 & 20736 & 45 & 60 & 10.7 & 5.6 & 1.2 & 10.2 & 5.1 & 0.9\\
3 & 5 & 50625 & 70 & 90 & 10.4 & 5.2 & 1.0 & 10.2 & 5.0 & 0.7\\\hline
3 & 6 & 104976 & 100 & 126 & 10.2 & 5.0 & 1.1 & 10.1 & 5.0 & 1.0\\\hline
3 & 7 & 194481 & 135 & 168 & 10.2 & 5.0 & 1.0 & 10.0 & 5.0 & 1.0\\\hline\hline
\end{tabular}
\caption{Rejection frequencies (in percentages) of KPST test at various
significance levels. $\chi^2_{df}$ critical values.
$n=(pk)^4$, $df$: number of restrictions given in eq.~(\ref{eq:
a_spec}), $m$: number of estimated parameters. Computed using 40,000 MC replications.}\label{Tab: NRP 2}
\end{table}


\begin{figure}[ptb]
\centering
\includegraphics[
height=3.6288in,
width=5.4293in
]
{Figure1.eps}
\caption{Null rejection probabilities of KPST test as a function of sample size $n$ at different significance levels: 10\% (red), 5\% (blue) and 1\% (green); and different data generating processes: homoskedastic (solid) and scalar heteroskedastic (dashed). Computed using 40,000 MC replications.}
\label{fig: p_2}
\end{figure}

\begin{figure}[ptb]
\centering
\includegraphics[
height=3.6288in,
width=5.4293in
]
{Figure2.eps}
\caption{Null rejection probabilities of KPST test as a function of sample size $n$ at different significance levels: 10\% (red), 5\% (blue) and 1\% (green); and different data generating processes: homoskedastic (solid) and scalar heteroskedastic (dashed). Computed using 40,000 MC replications.}
\label{fig: p_3}
\end{figure}

\begin{figure}[ptb]
\centering
\includegraphics[
height=3.224in,
width=4.83in
]
{Figure3.eps}
\caption{Null rejection probabilities of KPST test as a function of sample size $n$ at different significance levels: 10\% (red), 5\% (blue) and 1\% (green); and different data generating processes: homoskedastic (solid) and scalar heteroskedastic (dashed). Computed using 40,000 MC replications.}
\label{fig: p_4_5}
\end{figure}

\paragraph{Power}

We simulate the power of the KPST test using the asymptotic $\chi^{2}$
critical values stated in Theorem \ref{th: kps test}. The Data Generating
Process (DGP) is generated by a model with $p=k=2,$ where $Y_{i}=Z_{i}
\Pi+V_{i}$ and $\Pi=0,$ see (\ref{eq: regression}). The two dimensional
vectors containing the regressors $Z_{i}$ and errors $V_{i}$ are simulated
according to:
\begin{equation}
V_{i}\sim iid\left\{
\begin{array}
[c]{l}
N\left(  0,\Omega_{1}\right)  ,\\
N\left(  0,\Omega_{2}\right)  ,
\end{array}
\right.  \text{ }Z_{i}\sim iid\left\{
\begin{array}
[c]{l}
N\left(  0,Q_{zz,1}\right)  ,\quad i=1,...,\left[  n/2\right] \\
N\left(  0,Q_{zz,2}\right)  ,\quad i=[n/2]+1,...,n,
\end{array}
\right.  \label{eq: inst_errors}
\end{equation}
with $\Omega_{1}=diag\left(  b,1\right)  ,$ $\Omega_{2}=diag\left(
1,b\right)  ,$ $Q_{zz,1}=diag\left(  1,c\right)  ,$ $Q_{zz,2}=diag\left(
c,1\right)  ,$ and
\begin{equation}
b:=\frac{1}{2}\frac{\sigma}{\sqrt{n}}-\frac{1}{2}\sqrt{\frac{\sigma}{\sqrt{n}
}\left(  \frac{\sigma}{\sqrt{n}}+8\right)  }+1,\quad c:=\frac{1}{2}
\frac{\sigma}{\sqrt{n}}+\frac{1}{2}\sqrt{\frac{\sigma}{\sqrt{n}}\left(
\frac{\sigma}{\sqrt{n}}+8\right)  }+1,
\end{equation}
for $\sigma\in\lbrack0,\sqrt{n}).$ The covariance matrix $R$ is then such
that:
\begin{equation}
\begin{array}
[c]{rl}
R= & \frac{1}{n}var\left(  \sum_{i=1}^{n}(V_{i}\otimes Z_{i})\right)
=\frac{1}{2}diag\left(  b+c,1+bc,1+bc,b+c\right) \\
= & \underbrace{I_{4}}_{G_{1}\otimes G_{2}}+\frac{\sigma}{\sqrt{n}}{\times
}diag\left(  1,-1,-1,1\right)  ,
\end{array}
\label{eq: R_N}
\end{equation}
and $G_{1}=G_{2}=I_{2}$. Because
\begin{equation}
\mathcal{R(}diag\left(  1,-1,-1,1\right)  )=\left(
\begin{array}
[c]{cccc}
1 & 0 & 0 & -1\\
0 & 0 & 0 & 0\\
0 & 0 & 0 & 0\\
-1 & 0 & 0 & 1
\end{array}
\right)  ,
\end{equation}
$vec(G_{1})^{\prime}\mathcal{R(}diag\left(  1,-1,-1,1\right)  )vec(G_{2})=0,$
the re-arranged specification of $R$ in (\ref{eq: rearranged A_n}) equals$:$
\begin{equation}
\begin{array}
[c]{rl}
\mathcal{R}(R)= & vec(G_{1})vec(G_{2})^{\prime}+\frac{\sigma}{\sqrt{n}}\left(
\begin{array}
[c]{cccc}
1 & 0 & 0 & -1\\
0 & 0 & 0 & 0\\
0 & 0 & 0 & 0\\
-1 & 0 & 0 & 1
\end{array}
\right) \\
= & vec(G_{1})vec(G_{2})^{\prime}+\frac{1}{\sqrt{n}}vec(G_{1})_{\perp}
a_{0}vec(G_{2})_{\perp}^{\prime},
\end{array}
\end{equation}
where
\begin{equation}
\begin{array}
[c]{c}
vec(G_{1})_{\perp}=vec(G_{2})_{\perp}=\frac{1}{\sqrt{2}}\left(
\begin{array}
[c]{ccc}
1 & 0 & 0\\
0 & \sqrt{2} & 0\\
0 & 0 & \sqrt{2}\\
-1 & 0 & 0
\end{array}
\right)  ,\text{ }a_{0}=\sigma\left(
\begin{array}
[c]{ccc}
1 & 0 & 0\\
0 & 0 & 0\\
0 & 0 & 0
\end{array}
\right)  =\sigma e_{1}e_{1}^{\prime},\text{ }e_{1}=(1,0,0)^{\prime},
\end{array}
\end{equation}
is such that the local deviation from KPS lies in the orthogonal complement of
$vec(G_{1})$ and $vec(G_{2}).$ The non-centrality parameter of the non-central
$\chi^{2}$ limiting distribution follows from (\ref{eq: explicit delta}). Note
that
\begin{equation}
\begin{array}
[c]{l}
(e_{1}\otimes e_{1})^{\prime}\left[  (\left[  vec(G_{2})\right]  _{\perp
}^{\prime}\otimes\left[  vec(G_{1})\right]  _{\perp}^{\prime})\right. \\
\left.  \text{cov}\left(  vec\left(  \mathcal{R}(\hat{R})\right)  \right)
(\left[  vec(G_{2})\right]  _{\perp}\otimes\left[  vec(G_{1})\right]  _{\perp
})\right]  ^{-}(e_{1}\otimes e_{1})=\frac{1}{4},
\end{array}
\label{Eq: non-central1}
\end{equation}
where $G_{i}=I_{2}$ for $i=1,2$. Note also that $vec\left(  a_{0}\right)
=2\sigma(e_{1}\otimes e_{1})$. Thus, the non-centrality parameter is
\begin{equation}
\begin{array}
[c]{c}
\delta=\frac{1}{4}\sigma^{2}.
\end{array}
\label{Eq: non-central2}
\end{equation}
For $\sigma=0,$ $R$ has KPS, so the null hypothesis in (\ref{eq: KPS hyp})
holds. For the limiting case of $\sigma=\sqrt{n}:$ $b=0,$ so $\Omega_{1}$ and
$\Omega_{2}$ are singular.

We compute the power function of the KPST test at three significance levels
10\%, 5\% and 1\% using 10,000 Monte Carlo replications. For comparison, we
also compute the power of the non-invariant KPST$^{\ast}$ test that rejects
H$_{0}$ if the statistic KPST$^{\ast}$ in (\ref{eq: kpst*}) exceeds the
corresponding $1-\alpha$ quantile of $\chi_{df}^{2}$ with degrees of freedom
$df$ given in Theorem \ref{th: kps test}b, which are the same critical values
as for the KPST test in (\ref{eq: KPST test}). The results are reported
graphically in Figure \ref{fig: power}. The left-hand-side graphs in Figure
\ref{fig: power} show that for a moderate sample of size $n=200$ both tests
have good and essentially identical power. Moreover, as the sample size
increases, the power function of both tests approaches the noncentral
$\chi^{2}$ asymptotic approximation in Theorem \ref{th: kps test power}
indicated in blue on the right-hand-side graphs of Figure \ref{fig: power} for
$n=100,000.$ Results for other sample sizes are qualitatively similar and are
omitted in the interest of brevity. In particular, KPST has nontrivial power
even for small samples.

\begin{figure}[ptb]
\centering
\includegraphics[
height=4.4222in,
width=6.5303in
]
{Figure4.eps}
\caption{Power of KPST (solid red) and KPST$^{\ast}$ (dashed green) tests with
sample size $n=200$ (left) and $n=10^{5}$ (right). Asymptotic approximation
from Theorem \ref{th: kps test power} (dotted blue) superimposed on the right.
$\sigma$ measures deviation from KPS in Frobenius norm. Computed using 10,000
Monte Carlo replications.}
\label{fig: power}
\end{figure}


\section{Empirical applications\label{s: empirical}}

We investigate whether KPS covariance matrices are potentially relevant for
applied work. To do so, we apply the KPST\ test to the covariance matrices of
estimators in published empirical studies. We consider fifteen highly cited
papers conducting linear IV regressions from top journals in economics and
test for KPS of the joint covariance matrix of the (unrestricted reduced form)
least squares estimators which result from regressing all endogenous variables
on the instruments.\footnote{Both the endogenous variables and the instruments
are first regressed on the control, or included exogenous, variables and only
the residuals from these regressions are used.} Tables \ref{tab:KPS app} and
\ref{tab:clust} in the Supplementary Appendix report the results of the KPST
test for the 118 different specifications we analyzed. Table \ref{tab:KPS app}
does so for the studies using independent data (sixty specifications) while
Table \ref{tab:clust} lists the results for studies with clustered data (fifty
eight specifications). Because these tables are rather extensive, Tables
\ref{TableKey nocluster} and \ref{TableKey cluster} report a summary of our
findings on the KPST\ tests.

Table \ref{TableKey nocluster}, summarizing our results on KPS tests for the
papers using independent data, shows considerable support for KPS covariance
matrices especially when the number of observations is not too large. For the
60 different specifications using independent data reported in Table
\ref{TableKey nocluster}, KPS is rejected at the 5\% nominal size for only
about one third of them, namely for 22.

Table \ref{TableKey cluster}, summarizing the test results for papers using
clustered data, shows that for the 58 different specifications with clustered
data, KPS\ is rejected at the 5\% nominal size for 46 specifications when
using the unrestricted covariance matrix estimator (\ref{eq: sample cov}) and
for 40 when using the clust qered covariance matrix estimator
(\ref{eq: cluster covariance}). The number of observations in the involved
papers using clustered data is typically much larger than for the papers using
independent observations which largely explains our different findings for
independent compared to clustered observations.

Summarising, our analysis of the KPS of covariance matrices of moment
condition vectors in a considerable number of prominent empirical studies
shows that KPS is often not rejected especially for moderate sample sizes.

\nocite{ACJR2011}\nocite{AD2013}\nocite{ADG2013}\nocite{AGN2013}
\nocite{AJ2005}\nocite{AJRY2008}

\nocite{DL2012}\nocite{DT2011}\nocite{HG2010}\nocite{JPS2006}\nocite{MSS2004}
\nocite{Nunn2008}

\nocite{PSJM2013}\nocite{TCN2010}\nocite{Voors2012}\nocite{Yogo2004}\bigskip

\begin{table}[tbp] \centering
\begin{tabular}
[c]{|l|c|c|c|}\hline
\textbf{Paper} & \#\textbf{ specifications} & \textbf{KPS rejection} &
\textbf{\# observations}\\\hline
\cite{TCN2010} & 2 & none & moderate\\
\cite{Nunn2008} & 4 & 4 & small\\
\cite{AJ2005} & 24 & 10 & small\\
\cite{HG2010} & 2 & 2 & huge\\
\cite{AGN2013} & 6 & 1 & moderate\\
\cite{Yogo2004} & 22 & 5 & moderate\\\hline
\end{tabular}
\caption{Summary of results of 5\% significance level KPST tests for specifications in papers using independent observations}
\label{TableKey nocluster}
\end{table}

\begin{table}[tbp] \centering
\begin{tabular}
[c]{|l|c|c|c|c|c|}\hline
&  &  &  & \textbf{clustered} & \\
\textbf{Paper} & \#\textbf{ specific.} & \textbf{KPS rej.} & \textbf{\# obs.}
& \textbf{KPS rej.} & \textbf{\# clusters}\\\hline
\cite{DT2011} & 8 & 6 & large & 5 & moderate\\
\cite{AJRY2008} & 9 & 7 & large & 5 & moderate\\
\cite{JPS2006} & 4 & 4 & huge & 4 & huge\\
\cite{PSJM2013} & 2 & 2 & huge & 2 & huge\\
\cite{ADG2013} & 18 & 18 & large & 13 & small\\
\cite{AD2013} & 7 & 7 & huge & 7 & small\\
\cite{ACJR2011} & 1 & 1 & small & 1 & very small\\
\cite{MSS2004} & 3 & 0 & large & 3 & small\\
\cite{Voors2012} & 6 & 1 & moderate & 0 & small\\\hline
\end{tabular}
\caption{Summary of results of 5\% signficance level KPST tests for specifications in papers using clustered observations}
\label{TableKey cluster}
\end{table}


\section{Conclusion}

We propose a test for the null of a covariance matrix of a vector of moment
equations to have a KPS. The test is an extension of the
\cite{kleibergen2006grr}\ rank test and is easy to use. We apply it to data
used in a considerable number of prominent applied studies conducting IV
regressions and find that KPS of the covariance matrix of the least squares
estimator of the unrestricted reduced form is often not rejected for moderate
sample sizes. In linear IV regression, a KPS covariance matrix brings
considerable advantages for both computation and inference in weakly
identified settings. Given the common occurrence of weak identification in
applications, our empirical findings underscore the contribution that the use
of KPS covariance matrices can make in applied work.

In a companion paper, \cite{GKM21}, we develop a two-step test procedure that
in the first step uses the new KPS covariance matrix test and, depending on
its outcome, in the second step conducts a weak-identification-robust test on
a subset of the structural parameters based either on an improved powerful
subvector AR test or based on the AR/AR test that is robust to arbitrary forms
of conditional heteroskedasticity. The two-step procedure is constructed such
that its asymptotic size is bounded by the nominal size. A promising area for
application of testing for KPS is in linear factor models for establishing
risk premia. The default setting in this area is to assume homoskedasticity
and weak identification is often present.

To further improve the approximation of the finite sample distribution of the
KPST statistic, it would also be of interest to investigate whether the
bootstrap can deliver refinements as \cite{Chen2019} show for rank tests on
general matrices. We leave this important extension for future work.

\newpage