Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,005 characters · 6 sections · 74 citation commands
A Test for Kronecker Product Structure Covariance Matrix
$ \and $
$ \and \vspace{0.25in}$
$}
The robustness properties of nonparametric covariance matrix estimators, like those proposed by White80 against heteroskedasticity and by, for example, Newe87 and Andr91 against heteroskedasticity and autocorrelation, have led to the current default of conducting semi-parametric inference in econometrics. It is well understood that compared to parametrically specified covariance matrix estimators, these robustness properties come at the cost of a large number of additional estimated components. The latter affects the precision of semi-parametric estimators of the structural parameters compared to parametric ones.
For some structural models estimated by the generalized method of moments (GMM), see Hans82l, use of nonparametric covariance matrix estimators may also lead to computational challenges for estimation of the structural parameters when using the continuous updating estimator (CUE) of Hans96l. Prominent examples of such models are the linear instrumental variables (IV) regression model and the linear factor model in asset pricing. When using a nonparametric covariance matrix estimator, the CUE objective function is often ill behaved, that is, it is flat and/or has many local extrema, making the CUE difficult to compute. Part of the appeal of the CUE stems from a number of weak-identification-robust tests based on statistics centered around the CUE for hypotheses involving the structural parameters, see e.g. Kleibergen2005.
When one uses a Kronecker Product Structure (KPS) covariance matrix estimator instead of a nonparametric one in the CUE\ objective function in linear IV and factor asset pricing models, the CUE (which is then typically referred to as the limited information maximum likelihood (LIML) estimator) is straightforward to compute. Furthermore, weak-identification-robust tests specified on a subvector of the structural parameter vector with uniformly better power than projected robust full-vector tests are available, see e.g. gkm19, GKM21, and kf21. The KPS structure of the covariance matrix also allows for an analytical computation of the confidence sets of the structural parameters using the algorithm from dufour2005projection.\footnote{dufour2005projection actually assume homoskedasticity, which is a special case of KPS covariance, but their algorithm can be modified to cover the more general case of KPS covariance.}
The above illustrates the trade-off between, on the one hand, the robustness provided by a nonparametric covariance matrix estimator and, on the other hand, the computational ease and accurate statistical inference provided by a KPS covariance matrix estimator. To help empirical researchers decide when the use of a KPS covariance matrix estimator is justified, we develop a test for the null hypothesis that the covariance matrix $R=E\left( \frac{1}{n} \sum_{i=1}^{n}f_{i}f_{i}^{\prime}\right) $ has KPS, where $f_{i}=V_{i}\otimes Z_{i}$ and $V_{i}\in\mathbb{R}^{p}$ and $Z_{i}\in\mathbb{R}^{k}$ are uncorrelated random vectors. Here $V_{i}$ are unobserved error variables (for which consistent estimators are available) and $Z_{i}$ are observed regressors. This setup encompasses, for example, the linear IV and factor asset pricing models.
The test is based on the insight that KPS implies that a certain invertible transformation $\mathcal{R}\left( R\right) $ of $R$ has rank one, see vanloanpit93 and ((ref)) below, and our procedure adapts the kleibergen2006grr reduced rank statistic to test for KPS.\footnote{Another adaptation of the kleibergen2006grr reduced-rank statistic is by dfp07, who develop a test for singularity of a symmetric matrix.} More precisely, the new test statistic is given as a quadratic form in $vec(\hat{\Lambda})$ with weighting matrix that depends on $vec(\mathcal{R}(\hat{R})),$ where $\hat{R}$ is a sample analogue of $R$ and $\hat{\Lambda}$ is an estimator for a certain matrix that is known to be rank restricted under the null hypothesis (see ((ref) )-((ref)) below). The adaptation of kleibergen2006grr is nontrivial partly because the covariance matrix of $vec(\mathcal{R}(\hat {R})),$ that appears in the modified test statistic, is singular. As a consequence, it is a priori not obvious whether the use of the Moore-Penrose generalized inverse in the expression of the kleibergen2006grr reduced rank statistic still leads to a $\chi^{2}$ limiting distribution. To answer that question, we first derive the limiting distribution of $\hat{\Lambda}$ and show the limit to be degenerate Normal. We next establish that the probability limit of the Moore-Penrose inverse of the covariance matrix involved in the kleibergen2006grr rank statistic is such that it offsets this degeneracy. As the final result, we conclude that the new KPS test statistic has a $\chi^{2}$ limiting null distribution with degrees of freedom equal to the number of tested restrictions. We also consider an asymptotic setup where $p,$ $k,$ and $n$ jointly go to infinity and show that the asymptotic null rejection probability of the test is controlled as long as $(pk)^{16}=o(n^{3}).$ For power considerations, we establish that under sequences of covariance matrices local to KPS the test statistic has a limiting noncentral chi square distribution.
As an important property we show that the proposed test is invariant to orthonormal transformations of the data. In contrast, we show that this is not true for certain alternative tests for KPS that are based on an application of the kleibergen2006grr test statistic to a different transformation of the sample covariance matrix estimator that does not lead to a singular covariance matrix (and may therefore a priori seem the more natural choice).
We provide comprehensive Monte Carlo simulations that document good size and power properties of the suggested test. Finally, we apply the new KPS test to various specifications of linear IV models employed in fifteen highly cited empirical studies recently published in top ranked economic journals. We find that for the specifications with independent data and moderate numbers of observations, KPS is not rejected in 24 out of 30 cases at the 5% significance level, while for smaller numbers of observations, it is rejected in 14 out of 28 cases. In specifications with clustered data, KPS is not rejected in 7 out of 17 cases with moderate sample sizes, and 11 out of 35 cases with smaller samples. Overall, KPS is not rejected in 56 out of the 118 specifications that we tested. The relatively high number of non-rejections illustrates the potential importance of the KPS test for applied work.
In a companion paper, GKM21, we show how the new KPS\ test can be used as a key ingredient in a testing procedure with correct asymptotic size for a null hypothesis that restricts the values of a subvector of the structural parameter vector in the linear IV model with a general covariance matrix. The first step of the algorithm uses the KPS test to test the null of a KPS of the covariance matrix of the unrestricted reduced-form sample moment vector. In the second step of the algorithm, the null hypothesis involving the structural parameter is tested using the improved subvector Anderson-Rubin test from gkm19 when the test in the first step does not reject and using the size correct AR $\backslash$ AR test procedure from and17 otherwise. The AR $\backslash$ AR procedure from and17 is an asymptotically size correct inference procedure for testing hypotheses on a subvector of the structural parameters for general covariance matrices but is less powerful than the improved subvector Anderson-Rubin test from gkm19 in the linear IV regression model. However, the latter test is asymptotically size correct only when the covariance matrix has KPS. GKM21 establish that the resulting two-step procedure has correct asymptotic size and conduct Monte-Carlo experiments which show that it leads to more powerful subvector inference than the AR $\backslash$ AR test in and17.
As in the linear IV\ regression model, a KPS structure of the covariance matrix of the sample moment vector of the linear regression model encompassing linear asset pricing models also leads to improvements in terms of the power of identification robust tests on individual elements of the vector of risk premia and computational ease of obtaining the estimator of the risk premia. There is increasing awareness that risk premia of many risk factors are only weakly identified, see e.g. kz99, kf09, and kz2020. Therefore, it is important to analyze them using inference methods that are robust to weak identification. The current state of the art for conducting weak-factor-robust inference on risk premia is to assume homoskedasticity. Extending homoskedasticity to KPS or even further by extending the switching test procedure from GKM21 would extend the scope of the weak-factor-robust inference methods for analyzing the individual risk premia in linear asset pricing models. The KPS test would be an integral part of such extensions.
KPS or separability, which is how other fields sometimes refer to KPS, of the covariance matrix is also studied in the statistics and signal processing literature. The distance to a covariance matrix with KPS\ is considered in Genton07 and VH17, while LuZim05 and mgg06 analyze the likelihood ratio test of KPS of the covariance matrix of Normally distributed data. They estimate the elements of the KPS covariance matrix using a switching algorithm. Exploiting the reduced rank restriction imposed on the reordered covariance matrix by KPS is also done in wjs08. Their results are, however, based on a complex Gaussian distribution for the data, which leads to a degrees of freedom parameter of the $\chi^{2}$ limiting distribution of their test that is different from the one derived here.
KPS is an example of dimension reduction of a covariance matrix. Other examples of dimension reduction result from shrinking the covariance matrix to a matrix with (much) fewer unrestricted elements to estimate, for example, a scalar multiple of the identity matrix, see e.g. lw12, or by shrinking the population eigenvalues, see e.g. lw15 and lw18.
The paper is organized as follows. In the second section, we introduce the new test for a KPS covariance matrix and derive the asymptotic null distribution of the test statistic, which we denote as KPST. The third section contains the limiting distribution of the KPST statistic under local alternatives while the fourth section conducts a simulation study to analyze the size and power of the new KPS test. The fifth section summarizes the extensive analysis of testing for a KPS\ reduced-form covariance matrix in a considerable number of prominent articles. The final sixth section concludes. Proofs and detailed empirical results are given in the Appendix.
We use the vec operator of the matrix $A$, $vec(A):=(a_{1}^{\prime}\ldots a_{k}^{\prime})^{\prime}\in\mathbb{R}^{mk}$ for an $m\times k$ dimensional matrix $A=(a_{1},\ldots,a_{k}).$ For a symmetric $m\times m$ dimensional matrix $A,$ we also use the $m^{2}\times\frac{1}{2}m(m+1)$ dimensional, so-called, duplication matrix $D_{m}$ which selects the $\frac{1}{2}m(m+1)$ unique elements of $A$ in the $\frac{1}{2}m(m+1)$ dimensional vector $vech(A)$ that vectorizes only the lower triangular part of $A$: \[ vech(A)=(D_{m}^{\prime}D_{m})^{-1}D_{m}^{\prime}vec(A)\text{ \ \ \ and \ \ }vec(A)=D_{m}vech(A). \]
We propose a test for a covariance matrix $R\in\mathbb{R}^{kp\times kp}$ to have KPS, where
for mean zero, independently distributed random vectors $f_{i}\in \mathbb{R}^{kp},$ $i=1,\ldots,n,$ which satisfy
with $V_{i}\in\mathbb{R}^{p}$ and $Z_{i}\in\mathbb{R}^{k}$ uncorrelated random vectors.\footnote{The matrix $R$ can depend on the sample size $n$ but for simplicity of notation we do not index $R$ by $n.$} The specification of $f_{i}$ fits, for example, a setting where $V_{i}$ contains the errors of a number of regression equations and $Z_{i}$ contains the regressors, so that $R$ is then the covariance matrix of the sample covariance between these errors and the regressors.
It follows that the covariance matrix has a block structure
where $R_{jl}\in\mathbb{R}^{k\times k},$ $j,l=1,\ldots,p$. Because $R_{jl}=E\left( \frac{1}{n}\sum_{i=1}^{n}V_{ij}V_{il}Z_{i}Z_{i}^{\prime }\right) \allowbreak=R_{lj}^{\prime},$ for $V_{i}=(V_{i1}\ldots V_{ip})^{\prime},$ it follows that $R_{jl}$ is symmetric. We are interested in testing if the covariance matrix $R$ has KPS:
with $G_{1}\in\mathbb{R}^{p\times p}$ and $G_{2}\in\mathbb{R}^{k\times k}$ symmetric positive definite matrices, against the alternative hypothesis of not having KPS. For normalization purposes, we set one diagonal element equal to one (say the upper left element of $G_{1}$ ).\footnote{Normalizing $G_{1,11}$ to one is an obvious normalization because $G_{1}$ is a positive definite matrix (because $(G_{1}\otimes G_{2})$ is positive definite) so its diagonal elements are all strictly larger than zero. The normalization does therefore not imply a restriction.} When $p=1$ or $k=1$ the null is always true and from now on we assume that $\min\{p,k\}\geq2.$ To measure the distance of the sample covariance matrix estimator from a KPS\ covariance matrix, we use a convenient (invertible) transformation proposed by vanloanpit93.
For a matrix $A\in\mathbb{R}^{kp\times kp}$ with block structure as in ((ref)) define
for $j=1,...,p.$ One can easily show that
and by Theorem 2.1 in vanloanpit93, we have \[ \left\Vert R-G_{1}\otimes G_{2}\right\Vert _{F}=\allowbreak\left\Vert \mathcal{R}(R)-\allowbreak vec(G_{1})vec(G_{2})^{\prime}\right\Vert _{F}, \] with $\left\Vert .\right\Vert _{F}$ the Frobenius or trace norm of a matrix, $\left\Vert A\right\Vert _{F}^{2}:=tr(A^{\prime}A)=vec(A)^{\prime}vec(A),$ for any rectangular matrix $A$. Because $\mathcal{R}\left( G_{1}\otimes G_{2}\right) $ is a matrix of rank one, when testing for a KPS, it is more convenient to test for the rank of $\mathcal{R}\left( R\right) $ to be one instead of directly testing for KPS\ of $R$.
Consider the covariance matrix estimator
which uses sample values $\hat{f}_{i}:=\hat{V}_{i}\otimes Z_{i}$ of the random vectors $f_{i}$ for some estimated residuals $\hat{V}_{i}.$ We assume that $\hat{f}_{i}=f_{i}+o_{p}(1),$ uniformly over $i=1,\ldots,n,$ as $n\rightarrow \infty.$ Define the distance from a KPS covariance matrix by the Frobenius norm
where $G_{1},$ $G_{2}>0$ indicates that $G_{1}\in\mathbb{R}^{p\times p}$ and $G_{2}\in\mathbb{R}^{k\times k}$ are positive definite symmetric matrices, and $G_{1,11}=1$ states that the upper left element of $G_1$ is normalized to 1.
We test for $\mathcal{R}(\hat{R})$ being a rank one matrix using the kleibergen2006grr rank statistic. To describe the kleibergen2006grr rank statistic consider first a singular value decomposition (SVD) of $\mathcal{R}(\hat{R})$:
where $\hat{\Sigma}:=diag(\hat{\sigma}_{1},\ldots,\hat{\sigma}_{\min (p^{2},k^{2})})$ denotes a $p^{2}\times k^{2}$ dimensional diagonal matrix with the singular values $\hat{\sigma}_{j}$ ($j=1,...,\min(p^{2},k^{2})$) on the main diagonal ordered non-increasingly, and with $\hat{L}\in \mathbb{R}^{p^{2}\times p^{2}}$ and $\hat{N}\in\mathbb{R}^{k^{2}\times k^{2}}$ orthonormal matrices. Decompose
with $\hat{L}_{11}:1\times1,$ $\hat{L}_{12}:1\times(p^{2}-1),$ $\hat{L} _{21}:(p^{2}-1)\times1,$ $\hat{L}_{22}:(p^{2}-1)\times(p^{2}-1),$ $\hat {\sigma}_{1}:1\times1,$ $\hat{\Sigma}_{2}:(p^{2}-1)\times(k^{2}-1),$ $\hat {N}_{11}:1\times1,$ $\hat{N}_{12}:1\times(k^{2}-1),$ $\hat{N}_{21} :(k^{2}-1)\times1,$ $\hat{N}_{22}:(k^{2}-1)\times(k^{2}-1)$ dimensional matrices.
If $R$ is positive definite, and $\hat{R}\overset{p}{\rightarrow}R$, then $\hat{R}$ will be positive definite with probability approaching one (w.p.a.1). The choice of normalization in ((ref)) conforms with the normalization of $G_{1},G_{2}$ in ((ref)) discussed in Footnote (ref) above.
\paragraph{The KPST statistic}
We use the distance between $\mathcal{R}(\hat{R})$ and a matrix of rank one to test for a KPS of $R$. The test is based on the limiting distribution of the unique elements of $\hat{R}$ or equivalently $\mathcal{R}(\hat{R}).$ These elements result from using the $k^{2}\times\frac{1}{2}k(k+1)$ and $p^{2} \times\frac{1}{2}p(p+1)$ dimensional duplication matrices $D_{k}$ and $D_{p}:$
with
The $\frac{1}{2}p(p+1)\times\frac{1}{2}k(k+1)$ dimensional matrix $\hat {R}^{\ast}$ contains the unique elements of $\hat{R}$ and $\mathcal{R}(\hat {R}).$ We assume $vec(\hat{R}^{\ast})$ satisfies a central limit theorem:
with $\psi\sim N(0,V_{R^{\ast}}),$ $\Psi$ a $\frac{1}{2}p(p+1)\times\frac {1}{2}k(k+1)$ dimensional normally distributed random matrix and
In fact, we assume a slightly stronger result, namely, that $\hat{R}^{\ast }=R^{\ast}+\frac{1}{\sqrt{n}}\Psi+o_{p}(n^{-\frac{1}{2}}),$ holds. A central limit theorem ((ref)) for (possibly) non-identical distributed independent random variables holds under mild conditions, e.g. under the Liapounov's or Lindeberg's condition, see wh84.
Define
It can be shown that $\hat{\Lambda}=vec(\hat{G}_{1})_{\perp}^{\prime }\mathcal{R}(\hat{R})vec(\hat{G}_{2})_{\perp},$ where \[
\] see kleibergen2006grr. We then have
Using $\mathcal{R(}R),$ our hypothesis of interest H$_{0}$ ((ref)) is transformed into
where $vec(G_{1})_{\perp}$ and $vec(G_{2})_{\perp}$ are $p^{2}\times(p^{2}-1)$ and $k^{2}\times(k^{2}-1)$ dimensional matrices that contain the orthogonal complements of $vec(G_{1})$ and $vec(G_{2}),$ $vec(G_{1})_{\perp}^{\prime }vec(G_{1})\equiv0,$ $vec(G_{1})_{\perp}^{\prime}vec(G_{1})_{\perp}\equiv I_{p^{2}-1},$ $vec(G_{2})_{\perp}^{\prime}vec(G_{2})\equiv0,$ $vec(G_{2} )_{\perp}^{\prime}vec(G_{2})_{\perp}\equiv I_{k^{2}-1}.$ The KPST test uses the sample analog of the last component in ((ref)) to test H$_{0}.$ It further results from identifying $vec(G_{1})$ and $vec(G_{2})$ using the eigenvectors associated with the first singular value of $\mathcal{R(}R)$.
The kleibergen2006grr rank test statistic is a quadratic form of the vectorization of $\hat{\Lambda}$ in ((ref)). Its specification directly extends to the new KPS test but because the covariance matrix of $vec\left( \mathcal{R}(\hat{R})\right) $ is singular, the (degenerate) asymptotic normal distribution of $vec(\hat{\Lambda})$ and the resulting degrees of freedom parameter of the $\chi^{2}$ limiting distribution of the kleibergen2006grr rank test statistic are not obvious.
We define the statistic KPST\ for testing H$_{0}$ in ((ref)) as
where
and
In the Appendix it is shown that the KPST statistic in ((ref)) can be simplified as follows:
This provides an expression for KPST which is easier to compute. On the other hand, it cannot be directly used to obtain the $\chi^{2}$ limiting distribution because $\hat{\Sigma}_{2}$ does not have an asymptotic normal distribution while $vec(\hat{\Lambda})$ does.
\paragraph{The KPST$^{\ast}$ statistic}
For comparison, we now introduce an alternative test statistic KPST$^{\ast}$ that fits more naturally into the kleibergen2006grr framework. However, unlike KPST, KPST$^{\ast}$ turns out not to be invariant to orthonormal transformations of the data. Because $vec(G_{1})=D_{p}vech(G_{1}),$ $vec(G_{2})=D_{k}vech(G_{2}),$ the hypothesis of interest ((ref)) can also be specified as:
where $vech(G_{1})_{\perp}$ and $vech(G_{2})_{\perp}$ are $\frac{1} {2}p(p+1)\times(\frac{1}{2}p(p+1)-1)$ and $\frac{1}{2}k(k+1)\times(\frac{1} {2}k(k+1)-1)$ dimensional matrices that contain the orthogonal complements of $vech(G_{1})$ and $vech(G_{2}),$ $vech(G_{1})_{\perp}^{\prime}vech(G_{1} )=0,$ $vech(G_{1})_{\perp}^{\prime}vech(G_{1})_{\perp}= I_{\frac {1}{2}p(p+1)-1},$ $vech(G_{2})_{\perp}^{\prime}vech(G_{2})=0,$ $vec(G_{2})_{\perp}^{\prime}\allowbreak vec(G_{2})_{\perp}= I_{\frac {1}{2}k(k+1)-1}.$ This specification of the hypothesis fits directly in the setup of the kleibergen2006grr rank test because the covariance matrix of $\hat{R}^{\ast}$ is non-singular. Therefore, the corresponding specification of $vec(\hat{\Lambda})$ converges to a Normally distributed random vector. The specification of null hypothesis in ((ref) ) allows us to easily infer the number of restrictions tested, which equals $\left( \frac{1}{2}k(k+1)-1\right) \left( \frac{1}{2}p(p+1)-1\right) $, but the resulting rank statistic does not equal KPST in ((ref)). Specifically, define the SVD of $\hat{R}^{\ast} =\hat{L}^{\ast}\hat{\Sigma}^{\ast}\hat{N}^{\ast\prime}$, where \[ \hat{L}^{\ast}:=
, \hat{\Sigma}^{\ast}:=
, \hat{N}^{\ast}:=
, \] with $\hat{L}_{11}^{\ast}:1\times1,$ $\hat{L}_{12}^{\ast}:1\times\left( \frac{1}{2}p(p+1)-1\right) ,$ $\hat{L}_{21}^{\ast}:(\frac{1}{2} p(p+1)-1)\times1,$ $\hat{L}_{22}^{\ast}:(\frac{1}{2}p(p+1)-1)\times(\frac {1}{2}p(p+1)-1),$ $\hat{\sigma}_{1}^{\ast}:1\times1,$ $\hat{\Sigma}_{2}^{\ast }:(\frac{1}{2}p(p+1)-1)\times(\frac{1}{2}k(k+1)-1),$ $\hat{N}_{11}^{\ast }:1\times1,$ $\hat{N}_{12}^{\ast}:1\times(\frac{1}{2}k(k+1)-1),$ $\hat{N} _{21}^{\ast}:(\frac{1}{2}k(k+1)^{2}-1)\times1,$ $\hat{N}_{22}^{\ast}:(\frac {1}{2}k(k+1)-1)\times(\frac{1}{2}k(k+1)-1)$ dimensional matrices and $\hat {L}_{2}^{\ast}=(\hat{L}_{12}^{\ast\prime}$ $\vdots$ $\hat{L}_{22}^{\ast\prime })^{\prime},$ $\hat{N}_{2}^{\ast}=(\hat{N}_{12}^{\ast\prime}$ $\vdots$ $\hat{N}_{22}^{\ast\prime})^{\prime}$. The kleibergen2006grr statistic for testing ((ref)) using $\hat{R}^{\ast}$ is
see Corollary 1 in kleibergen2006grr.
\paragraph{Asymptotic theory and invariance to orthonormal transformations}
The statistics KPST in ((ref)) and KPST$^{\ast}$ in ((ref)) are not identical, and, unlike the proposed KPST statistic, tests of the KPS hypothesis based on the KPST$^{\ast}$ statistic are not invariant to orthonormal transformations of the data, as stated in the following Theorem.
We define the new KPST test as follows: it rejects H$_{0}$ in ((ref)) at nominal size $\alpha$ if
Based on Theorem (ref)a and d, the resulting test has limiting null rejection probability bounded by $\alpha.$
Theorem (ref)c shows that the rank-one tests KPST and KPST$^{\ast}$ that are based on $\mathcal{R(}\hat{R})$ and $\hat{R}^{\ast}$ respectively are not identical even though they are testing the same underlying hypothesis that the covariance $R$ matrix has KPS. This difference occurs because these are Wald tests, and Wald statistics are in general not invariant to non-linear transformations.
Theorem (ref)d provides a sufficient condition for uniform convergence of $\hat{\Lambda}$ and its covariance matrix estimator for settings where $p,$ $k,$ and $n$ jointly go to infinity so the main results for the limiting distribution of KPST remain unaltered. It is needed to assess the validity of the asymptotic approximation for settings where $p$ and $k$ are relatively large compared to the number of observations $n.$
The conditions in Theorem (ref)d are weaker than those in newey2009gmm. newey2009gmm prove the validity of the asymptotic approximation of test statistics where the number of observations grows faster than the cube of the number of moment restrictions. The number of moment restrictions here is proportional to $(pk)^{2}$ so their rate would be $(pk)^{6}/n\rightarrow0$ which is more restrictive than the rate in ((ref)).
\paragraph{Invariance to nonsingular transformations}
Theorem (ref)c shows the KPST is invariant to orthonormal transformations, but it is still not invariant to general nonsingular transformations of the data. To ensure invariance to nonsingular transformations, we need to normalize the data as in kleibergen2006grr. Specifically, the KPST statistic is computed using the moment vector
where $C_{1}$ and $C_{2}$ are the Choleski factors of the inverse of the second moments of $\hat{V}_{i}$ and $Z_{i},$ i.e., $C_{1}C_{1}^{\prime }=\left( \frac{1}{n}\sum_{i=1}^{n}\widehat{V}_{i}\widehat{V}_{i}^{\prime }\right) ^{-1}$ and $C_{2}C_{2}^{\prime}=\left( \frac{1}{n}\sum_{i=1} ^{n}Z_{i}Z_{i}^{\prime}\right) ^{-1}$. To see why this normalization yields invariance, let $A$ be a nonsingular $k\times k$ matrix, and define the transformed instruments $Z_{A,i}:=AZ_{i}.$ Let $C_{2A}$ denote the Choleski factor of $\left( \frac{1}{n}\sum_{i=1}^{n}AZ_{i}Z_{i}^{\prime}A^{\prime }\right) ^{-1}.$ The KPST statistic with the original instruments $Z_{i}$ is computed using $C_{2}^{\prime}Z_{i}$ in the moment vector ((ref)), while the KPST with the transformed instruments $AZ_{i}$ uses $C_{2A}^{\prime}AZ_{i}$ in the same formula ((ref)). Therefore, the transformation from $C_{2}^{\prime }Z_{i}$ to $C_{2A}^{\prime}Z_{A,i}$ is given by $T_{A}:=\allowbreak C_{2A}^{\prime}AC_{2}^{-1\prime},$ i.e., $C_{2A}^{\prime}Z_{A,i} \allowbreak=\allowbreak C_{2A}^{\prime}AZ_{i}\allowbreak=\allowbreak T_{A}\left( C_{2}^{\prime}Z_{i}\right) .$ Now, observe that $T_{A}$ is an orthonormal matrix, because $T_{A}^{\prime}T_{A}\allowbreak=\allowbreak C_{2}^{-1}A^{\prime}C_{2A}\allowbreak C_{2A}^{\prime}AC_{2}^{-1\prime }\allowbreak=\allowbreak C_{2}^{-1}A^{\prime}\allowbreak\left( \frac{1} {n}\sum_{i=1}^{n}AZ_{i}Z_{i}^{\prime}A^{\prime}\right) ^{-1}\allowbreak AC_{2}^{-1\prime}\allowbreak=\allowbreak C_{2}^{-1}\allowbreak\left( \frac {1}{n}\sum_{i=1}^{n}Z_{i}Z_{i}^{\prime}\right) ^{-1}\allowbreak C_{2}^{-1\prime}\allowbreak=\allowbreak I_{k}.$ Hence, invariance follows from Theorem (ref)c. The exact same argument can be made about rotations of the reduced form errors $\hat{V}.$
\paragraph{Clustered data}
In case of clustered data, we assume there are $n$ clusters of $N_{i}$ observations each, so the total number of data points is $\sum_{i=1}^{n} N_{i}:$
for mean zero $kp$ dimensional random vectors $f_{ij},$ $j=1,\ldots,N_{i},$ $i=1,\ldots,n.$ Observations $f_{ij}$ within cluster $i$ can be arbitrarily dependent, i.e., $E\left( f_{ij}f_{is}\right) $ is unrestricted for all $j,$ $s=1,...,N_{i},$ while observations across clusters are independent. The $kp\times kp$ dimensional (positive semi-definite) covariance matrix of the sample moments then results as:
To analyze the power of KPST under local alternatives, we construct the limiting distribution of the KPST statistic under alternatives where the covariance matrix of the moments $R\in\mathbb{R}^{kp\times kp}$ is local to KPS:
where $G_{1}\in\mathbb{R}^{p\times p}$ and $G_{2}\in\mathbb{R}^{k\times k}$ are symmetric positive definite matrices, and $A_{0} \in\mathbb{R}^{kp\times kp}$ is a fixed symmetric matrix. The best-fitting KPS approximation of $R$ under H$_{1}$ w.r.t. Frobenius norm, defined as $\bar{G}_{1,n}\otimes\bar{G}_{2,n}$, where $\bar{G}_{1,n},\bar {G}_{2,n}$ solve $\min_{\bar{G}_{1}>0,\bar{G}_{2}>0}\left\Vert (G_{1}\otimes G_{2})\allowbreak+\frac{1}{\sqrt{n}}A_{0}\allowbreak-\bar{G}_{1}\otimes\bar {G}_{2}\right\Vert _{F}$,$\,$will differ from $G_{1}\otimes G_{2}$. That is, $\bar{G}_{1,n}\neq G_{1}$ and $\bar{G}_{2,n}\neq G_{2},$ unless $A_{0}$ lies in the span of the orthogonal complement of $G_{1}\otimes G_{2}$. However, under the local alternatives ((ref)), $\bar{G}_{1,n}\rightarrow G_{1}$ and $\bar{G}_{2,n}\rightarrow G_{2}$. This needs to be taken into account when we characterize the asymptotic distribution of the KPST statistic under the local alternatives in ((ref)).
The re-arranged matrix $\mathcal{R}(R)$ under H$_{1}$ is:
with
The decomposition in the last line of ((ref)) is identical to the one in ((ref)).
\paragraph{Size}
We evaluate the accuracy of the limiting distribution in Theorem (ref) to approximate the finite sample distribution of the KPST\ statistic. We do so in a small simulation experiment using the linear regression model:
where $Y_{i}$ is a $p$ dimensional vector of dependent variables, $Z_{i}$ is a $k$ dimensional vector of explanatory (exogenous) variables and $V_{i}$ is a $p$ dimensional vector of errors. We further set $\Pi$ to zero (which is without loss of generality because KPST uses the residual vectors) and generate the $Z_{i}$'s independently from $N(0,I_{k})$ distributions and $V_{i}$ given $Z_{i}$ independently from a $N\left( 0,h\left( Z_{i}\right) I_{p}\right) $ distribution. We consider two different specifications of $h(Z_{i}).$ The first leads to homoskedasticity and has $h(Z_{i})=1$ while the second leads to (scalar) heteroskedasticity and has $h\left( Z_{i}\right) =\left\Vert Z_{i}\right\Vert ^{2}/k.$ For each case, we compute null rejection probabilities (NRPs) using the three conventional nominal significance levels of 10%, 5% and 1%. The NRPs are computed using 40,000 Monte Carlo replications for the KPST\ test that uses chi-square critical values based on the results from Theorem (ref). Table (ref) reports the NRPs when the sample size depends on the dimensions $p$ and $k,$ specifically $n=\left( kp\right) ^{16/3}$, in accordance with Theorem (ref). We notice only a slight underrejection in some cases, but in the remaining cases the NRPs are not significantly different from the test's nominal levels. Table (ref) reports NRPs with a smaller sample size $n=\left( pk\right) ^{4}$. In this case, we find some modest deviations from the nominal size but these are generally quite small.
To investigate NRPs in smaller samples, Figures (ref) to (ref) show the NRPs as a function of the sample size $n$ for smaller sample sizes than in Tables (ref), (ref) for different settings of $p$ and $k.$ Depending on the value of the latter, the NRPs are close to the nominal level for values of $n$ much smaller than $\left( pk\right) ^{4}.$ For larger values of $pk,$ we therefore do not (like for the smaller values of $pk)$ show the rejection frequencies all the way up to $n=(pk)^{\frac{16}{3}},$ i.e. the value indicated by Theorem (ref)d, but just to $(pk)^{4},$ which is for $p=2,$ $k=7$ at the bottom right hand side of Figure (ref), equal to approximately 40,000, and for $p=5,$ $k=4$ at the bottom right hand side of Figure (ref) equal to 160,000 (note that the horizontal axis is in log-scale). In many cases, the NRPs are still much closer to their nominal significance levels than indicated by this rate. For example, when $p=k=2$ and testing at the 5% significance level, the NRP is close to the nominal level for sample size of around 100. More striking is that when $p=2$ and $k=5$ the KPST test at 5% nominal size has NRPs close to the nominal size for values of $n$ around 200. Figures (ref)-(ref) also show that the KPST test generally over-rejects for small $n$. Moreover, the over-rejection is increasing in the dimensions $k$ and $p$ and can be very substantial for very small $n$, see Figure (ref), as is the case for any Wald test when the number of restrictions is large relative to the sample size. Therefore, it is of interest to investigate the possibility of small-sample corrections, e.g., following the bootstrap approach of Chen2019. From a practical perspective, this over-rejection means that rejection of KPS with small sample sizes, which happens only a few times in the applications reported in Section (ref), could be due to a significantly higher type 1 error probability than the nominal size of the test. \footnote{When KPST is used as a pre-test in a two-step procedure, such as the subvector Anderson-Rubin test of GKM21, that involves choosing a second-step test that is robust to violation of KPS when the KPST rejects in the first step, over-rejection will only affect the power but not the overall size of the two-step procedure.}
\paragraph{Power}
We simulate the power of the KPST test using the asymptotic $\chi^{2}$ critical values stated in Theorem (ref). The Data Generating Process (DGP) is generated by a model with $p=k=2,$ where $Y_{i}=Z_{i} \Pi+V_{i}$ and $\Pi=0,$ see ((ref)). The two dimensional vectors containing the regressors $Z_{i}$ and errors $V_{i}$ are simulated according to:
with $\Omega_{1}=diag\left( b,1\right) ,$ $\Omega_{2}=diag\left( 1,b\right) ,$ $Q_{zz,1}=diag\left( 1,c\right) ,$ $Q_{zz,2}=diag\left( c,1\right) ,$ and
for $\sigma\in\lbrack0,\sqrt{n}).$ The covariance matrix $R$ is then such that:
and $G_{1}=G_{2}=I_{2}$. Because
$vec(G_{1})^{\prime}\mathcal{R(}diag\left( 1,-1,-1,1\right) )vec(G_{2})=0,$ the re-arranged specification of $R$ in ((ref)) equals$:$
where
is such that the local deviation from KPS lies in the orthogonal complement of $vec(G_{1})$ and $vec(G_{2}).$ The non-centrality parameter of the non-central $\chi^{2}$ limiting distribution follows from ((ref)). Note that
where $G_{i}=I_{2}$ for $i=1,2$. Note also that $vec\left( a_{0}\right) =2\sigma(e_{1}\otimes e_{1})$. Thus, the non-centrality parameter is
For $\sigma=0,$ $R$ has KPS, so the null hypothesis in ((ref)) holds. For the limiting case of $\sigma=\sqrt{n}:$ $b=0,$ so $\Omega_{1}$ and $\Omega_{2}$ are singular.
We compute the power function of the KPST test at three significance levels 10%, 5% and 1% using 10,000 Monte Carlo replications. For comparison, we also compute the power of the non-invariant KPST$^{\ast}$ test that rejects H$_{0}$ if the statistic KPST$^{\ast}$ in ((ref)) exceeds the corresponding $1-\alpha$ quantile of $\chi_{df}^{2}$ with degrees of freedom $df$ given in Theorem (ref)b, which are the same critical values as for the KPST test in ((ref)). The results are reported graphically in Figure (ref). The left-hand-side graphs in Figure (ref) show that for a moderate sample of size $n=200$ both tests have good and essentially identical power. Moreover, as the sample size increases, the power function of both tests approaches the noncentral $\chi^{2}$ asymptotic approximation in Theorem (ref) indicated in blue on the right-hand-side graphs of Figure (ref) for $n=100,000.$ Results for other sample sizes are qualitatively similar and are omitted in the interest of brevity. In particular, KPST has nontrivial power even for small samples.
We investigate whether KPS covariance matrices are potentially relevant for applied work. To do so, we apply the KPST\ test to the covariance matrices of estimators in published empirical studies. We consider fifteen highly cited papers conducting linear IV regressions from top journals in economics and test for KPS of the joint covariance matrix of the (unrestricted reduced form) least squares estimators which result from regressing all endogenous variables on the instruments.\footnote{Both the endogenous variables and the instruments are first regressed on the control, or included exogenous, variables and only the residuals from these regressions are used.} Tables (ref) and (ref) in the Supplementary Appendix report the results of the KPST test for the 118 different specifications we analyzed. Table (ref) does so for the studies using independent data (sixty specifications) while Table (ref) lists the results for studies with clustered data (fifty eight specifications). Because these tables are rather extensive, Tables (ref) and (ref) report a summary of our findings on the KPST\ tests.
Table (ref), summarizing our results on KPS tests for the papers using independent data, shows considerable support for KPS covariance matrices especially when the number of observations is not too large. For the 60 different specifications using independent data reported in Table (ref), KPS is rejected at the 5% nominal size for only about one third of them, namely for 22.
Table (ref), summarizing the test results for papers using clustered data, shows that for the 58 different specifications with clustered data, KPS\ is rejected at the 5% nominal size for 46 specifications when using the unrestricted covariance matrix estimator ((ref)) and for 40 when using the clust qered covariance matrix estimator ((ref)). The number of observations in the involved papers using clustered data is typically much larger than for the papers using independent observations which largely explains our different findings for independent compared to clustered observations.
Summarising, our analysis of the KPS of covariance matrices of moment condition vectors in a considerable number of prominent empirical studies shows that KPS is often not rejected especially for moderate sample sizes.
\nocite{ACJR2011}\nocite{AD2013}\nocite{ADG2013}\nocite{AGN2013} \nocite{AJ2005}\nocite{AJRY2008}
\nocite{DL2012}\nocite{DT2011}\nocite{HG2010}\nocite{JPS2006}\nocite{MSS2004} \nocite{Nunn2008}
\nocite{PSJM2013}\nocite{TCN2010}\nocite{Voors2012}\nocite{Yogo2004}
We propose a test for the null of a covariance matrix of a vector of moment equations to have a KPS. The test is an extension of the kleibergen2006grr\ rank test and is easy to use. We apply it to data used in a considerable number of prominent applied studies conducting IV regressions and find that KPS of the covariance matrix of the least squares estimator of the unrestricted reduced form is often not rejected for moderate sample sizes. In linear IV regression, a KPS covariance matrix brings considerable advantages for both computation and inference in weakly identified settings. Given the common occurrence of weak identification in applications, our empirical findings underscore the contribution that the use of KPS covariance matrices can make in applied work.
In a companion paper, GKM21, we develop a two-step test procedure that in the first step uses the new KPS covariance matrix test and, depending on its outcome, in the second step conducts a weak-identification-robust test on a subset of the structural parameters based either on an improved powerful subvector AR test or based on the AR/AR test that is robust to arbitrary forms of conditional heteroskedasticity. The two-step procedure is constructed such that its asymptotic size is bounded by the nominal size. A promising area for application of testing for KPS is in linear factor models for establishing risk premia. The default setting in this area is to assume homoskedasticity and weak identification is often present.
To further improve the approximation of the finite sample distribution of the KPST statistic, it would also be of interest to investigate whether the bootstrap can deliver refinements as Chen2019 show for rank tests on general matrices. We leave this important extension for future work.