Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
55,690 characters · 9 sections · 40 citation commands
A Canonical Representation of Block Matrices with Applications to Covariance and Correlation Matrices
JEL Classification:{ C10; C22; C58 }{}
We derive a canonical representation for a broad class of block matrices, which includes block covariance and block correlation matrices as special cases. The representation is a semi-spectral decomposition of a block matrix, which is diagonalized with the exception of a single diagonal block, whose dimension is given by the number of blocks. The canonical representation facilitates simple computations of several matrix functions, such as the matrix inverse, the matrix exponential, and the matrix logarithm. Interestingly, we show that these transformations preserve the block structure of the original matrix. Consequently, the decomposition greatly simplifies the evaluation of Gaussian log-likelihood functions when the covariance matrix, or the correlation matrix, has a block structure. The canonical representation can also be used in regressions with many regressors, instrumental variables, and dependent variables when a block structure is appropriate.
We contribute to the literature on block correlation models by providing simple expressions for the inverse of any (invertible) block correlation matrix, as well as a simple expression for its determinant. The results apply to block correlation matrices with an arbitrary number of blocks. For block correlation matrices with two blocks, an expression for its inverse was obtained in EngleKelly2012, and related results can be found in VianaOlkin1997.
We apply the block structure to estimate annual covariance matrices of a large panel of assets using daily returns for 26 calendar years. The block structure makes it possible to estimate and manipulate large covariance matrices. The latter can be used to compute partial correlations. A preview of our empirical results is presented in Figure (ref), where we use color codes to present estimated correlation matrices for 3,340 US stocks for 2019 (left) and 2020 (right). These are estimated with a block structure that assumes that the correlation between two assets is defined by the sub-industries they belong to. The sub-industries are sorted according to their Global Industry Classification Standard (GICS) code as of 2020.\footnote{https://en.wikipedia.org/wiki/Global_Industry_Classification_Standard}
The solid lines indicate the boundaries of the 11 GICS sectors: Energy (10), Materials (15), Industrials (20), Consumer Discretionary (25), Consumer Staples (30), Health Care (35), Financials (40), Information Technology (45), Communication Services (50), Utilities (55), and Real Estate (60). Each of these $3340\times3340$ correlation matrices are estimate with just 252 vectors of daily returns in 2019 and 253 daily returns in 2020. Correlations were generally larger in 2020 than in 2019. This is not surprising given the impact that the COVID-19 pandemic had on financial markets in 2020. The estimated correlations reveal interesting differences within the sectors and sub-industries (related to gold and other precious metals, biotechnology, and pharmaceuticals) that are largely uncorrelated with other sub-industries. More details are presented in Section (ref).
As a preview of our theoretical results, consider the $n\times n$ equicorrelation matrix, \[ C=\left[
\right], \] which has the eigenvalues, $1+\rho(n-1)$ and $1-\rho$, where the latter has multiplicity $n-1$. This follows directly from the spectral decomposition,
where $Q$ is an orthonormal matrix, i.e., $Q^{\prime}Q=I_{n}$, see OlkinPratt:1958. Here $I_{n}$ denotes the $n\times n$ identity matrix. The matrix $Q$ is given by $Q=(v_{n},v_{n\bot})$, where $v_{n}$ is the $n$-dimensional vector, $v_{n}=(\tfrac{1}{\sqrt{n}},\ldots,\tfrac{1}{\sqrt{n}})^{\prime}$, and $v_{n\bot}$ is an $n\times(n-1)$ matrix that is orthogonal to $v_{n}$, i.e., $v_{n\bot}^{\prime}v_{n}=0$, and orthonormal, i.e., $v_{n\bot}^{\prime}v_{n\bot}=I_{n-1}$.\footnote{When $n=1$, $v_{1\bot}$ is an $1\times0$ “matrix” and we use the conventions: $v_{1\bot}^{\prime}v_{1\bot}=\emptyset$ (dimension $0\times0$) and $v_{1\bot}v_{1\bot}^{\prime}=0$ (dimension $1\times1$). This ensures that our expressions also hold in the special case, where one or more blocks has size one.} It can now be verified that $QQ^{\prime}=I_{n}$ and $C=QQ^{\prime}CQQ^{\prime}=QDQ^{\prime}$. In this example, $D$ is the canonical form of $C$, which is obtained via a rotation of $C$, where the rotation matrix, $Q$, does not depend on $\rho$.
In this paper, we derive a similar decomposition for a broad class of block matrices, which includes block covariance matrices and block correlation matrices as special cases. In the general case with multiple blocks, $K\geq2$, the canonical representation does not fully disentangle all eigenvalues. The canonical representation decomposes any block matrix into a $K\times K$ matrix and $n-K$ real-valued eigenvalues, where $K$ is the number of blocks. We can illustrate the general results with a $2\times2$ block correlation matrix, $$C= \left[
\right] $$where $C_{\rho_{11}}$ and $C_{\rho_{22}}$ are equicorrelation matrices with correlations $\rho_{11}$ and $\rho_{22}$, respectively, and dimensions $n_{1}\times n_{1}$ and $n_{2}\times n_{2}$, respectively, and $\mathbf{1}_{n_{1}\times n_{2}}$ is the $n_{1}\times n_{2}$ matrix whose elements are all equal to one. Now define \[ Q=\left[
\right]. \] For this correlation matrix with $2\times2$ blocks, we have the following representation,
We denote the upper-left $2\times2$ matrix by $A$. In general, $A$ will be a $K\times K$ matrix, whose eigenvalues are also eigenvalues of $C$. The general result for block matrices with $K$ blocks will be presented in Theorem (ref), with a structure similar to that in ((ref)). Importantly, the matrix $Q$ does not depend on the elements in block matrix, but is solely determined by the block partition, $(n_{1},\ldots,n_{K})$, where $n=n_{1}+\cdots+n_{K}$.
The canonical representation is obtained for general block matrices, which need not be symmetric, nor positive semidefinite. In fact, our results are applicable to non-square matrices. Block covariance matrices and block correlation matrices are important special cases. For block correlation matrices, the $A$-matrix, which emerges in ((ref)), was previously established in HuangYang:2010 and CadimaCalheirosPreto2010, as we will discuss in Section (ref). We derive additional results for block correlation matrices that simplify various matrix transformations and the evaluation of the Gaussian log-likelihood function.
The rest of this paper is organized as follows. We present the main result in Section 2, where the canonical representation is established for a broad class of block matrices, along with related results for some matrix functions. We also cover aspects of block structures in regressions with many variables. In Section 3, we consider the special case with block covariance matrices and block correlation matrices. Many of these results are useful for maximum likelihood estimation with a Gaussian log-likelihood function, as we show in Section 4. In Section (ref), we present our empirical analysis of block covariance matrices for a large panel of daily stock returns. We conclude in Section 6 and all proofs are presented in the Appendix. A separate Web Appendix with additional empirical results, expressions for computing partial correlations from block correlation matrices, and Matlab code used in our empirical analysis.
Let $B$ be a square $n\times n$ matrix. The extension to rectangular matrices is trivial and will be addressed towards the end of this section. The matrix, $B$, is called a block matrix with block partition, $n_{1},\ldots,n_{K}$, if it can be expressed as: \[ B=\left[
\right], \] where $B_{[k,l]}$ is an $n_{k}\times n_{l}$ matrix with the following structure
for some constants, $d_{k}$ and $b_{kl}$, $k,l=1,\ldots,K$. So the diagonal elements of the diagonal blocks, $B_{[k,k]}$, can take a different value than the off-diagonal elements, whereas all elements in an off-diagonal block, $B_{[k,l]}$, $k\neq l$, are identical.
We introduce the following notation, which relates to orthogonal projections. Let $P_{[k,l]}=v_{n_{k}}v_{n_{l}}^{\prime}$ be the $n_{k}\times n_{l}$ matrix whose elements are all equal to $\frac{1}{\sqrt{n_{k}n_{l}}}$. It is simple to verify that $P_{[k,m]}P_{[m,l]}=P_{[k,l]}$, and with $k=l=m$ it follows that $P_{[k,k]}P_{[k,k]}=P_{[k,k]}$, such that $P_{[k,k]}$ is a projection matrix. It then follows that $P_{[k,k]}^{\bot}=I_{n_{k}}-P_{[k,k]}$ is a projection matrix, and it can be verified that $P_{[k,k]}^{\bot}=v_{n_{k}\bot}v_{n_{k}\bot}^{\prime}$, where the matrix, $v_{n\bot}$, was characterized in the introduction.
Finally, we define the $n\times n$ matrix \[ Q=\left[
\right], \] and observe that $Q$ is an orthonormal matrix, characterized by the identity $Q^{\prime}Q=I$. The first $K$ columns of $Q$ can be used to form averages within each of the $K$ blocks, whereas the remaining columns of $Q$ capture “differences” within each block. The two sets of columns span orthogonal subspaces, which correspond to distinct components of the block decomposition. Note that $Q$ is solely defined by the block partition, $n_{1},\ldots,n_{K}$, and it is therefore invariant to the actual values taken by the elements in the block matrix.
The matrix $Q$ rotates $B$ into its canonical form, $D$. The first $K$ columns of $Q$ span the eigenspace of $B$ that is associated with the eigenvalues that $A$ and $B$ have in common. The last $n-K$ columns of $Q$ are the remaining eigenvectors of $B$.
For an arbitrary matrix, $B$, we would not expect $Q^{\prime}BQ$ to have zeroes in particular entries. The block structure imposes a type of sparsity. Rather than imposing the sparsity on $B$, it is imposed on $Q^{\prime}BQ$.\footnote{The repeated diagonal elements of $D$ (for $k\geq3$) implies additional structure.}
Theorem (ref) can be used to characterize properties of $B$ and simplifies the computation of some matrix transformations. These include the matrix exponential, $\exp(B)=\sum_{j=0}^{-\infty}\tfrac{1}{j!}B^{j}$, and the matrix logarithm, denoted $\log B$, which is the inverse of $\exp(B)$. The matrix exponential and matrix logarithm are used in spatial models, see LeSagePace:2007, in multivariate volatility models, see e.g., Kawakatsu:2006, maheu-mccurdy:2011, AsaiSo:2015, and ArchakovHansenLundeMRG, and play a central role in Markov processes. The matrix logarithm is not always well-defined, but for a positive definite symmetric matrix, $B=V\Lambda V^{\prime},$ where $\Lambda$ is the diagonal matrix with $B$'s eigenvalues, $\xi_{1},\ldots,\xi_{n}$, we simply have $\log B=V\mathrm{diag}(\log\xi_{1},\ldots,\log\xi_{n})V^{\prime}$.
It follows that $B^{q}$ is well-defined for all positive integers of $q$, and the matrix inverse, $B^{-1}$, exists whenever $A$ is invertible and $\lambda_{k}\neq0$, for all $k=1,\ldots,K$, in which case $B^{q}$ is also well-defined for other negative integers of $q$. The logarithms, $\log A$ and $\log(d_{k}-b_{kk})$, exist provided that $A$ is invertible and $d_{k}-b_{kk}\neq0$. This may result in a complex-valued solution to the matrix logarithm. If a real-valued solution is required, then the conditions are that $A$ is positive definite and that $d_{k}-b_{kk}>0$ for all $k=1,\ldots,K$.
Many of the expressions can be simplified further in the special case, where all block sizes are identical, i.e., $n_{1}=n_{2}=\ldots=n_{K}=n$, with $n=N/K$. In this situation, we have $B=A\otimes P+\Lambda\otimes P_{\bot}$, where $P$ is the $n\times n$ matrix with $1/n$ in all entries, $P_{\bot}=I_{n}-P$, and $\Lambda=\mathrm{diag}(\lambda_{1},\ldots,\lambda_{K})$. In this case, it follows that $h(B)=h(A)P+h(\Lambda)P_{\bot}$, where $h(\cdot)$ represents the matrix inverse, the matrix exponential, or the matrix logarithm, provided these are well-defined.
Suppose that $B$ has blocks, $B_{[k,l]}\in\mathbb{R}^{n_{k}\times n_{l}}$, as specified in ((ref)), where $k=1,\ldots,K_{1}$ and $l=1,\ldots,K_{2}$, and $K_{1}\neq K_{2}$, such that $B$ is a non-square matrix. Set $K=\max(K_{1},K_{2})$ and suppose that $K_{1}>K_{2}$. We conjoin zero-matrix to $B$, such that $\tilde{B}=(B,0)$ is a square matrix with the block partition, $n_{1},\ldots,n_{K}$. Our results apply to $\tilde{B}$, such that $\tilde{B}=QDQ^{\prime}$ has the canonical form and $B=QD\tilde{Q}^{\prime}$, where $\tilde{Q}^{\prime}$ is made up of the first $n_{1}+\cdots+n_{K_{2}}$ columns of $Q^{\prime}$. If $K_{2}>K_{1}$, we can instead define $\tilde{B}=(B^{\prime},0)^{\prime}$, and the results follow similarly. We will make use of rectangular block matrices in the next subsection.
Block matrices can be used to impose structure in regression models with many regressors, many instrumental variables, and many dependent variables. Consider the standard regression model with stationary variables, $Y_{t}=\beta^{\prime}X_{t}+\varepsilon_{t}$, $t=1,\ldots,n$ and $X_{t}\in\mathbb{R}^{q}$. If $q$ is large relative to $n$, it may be desirable to estimate $\Sigma_{xx}=\mathbb{E}[X_{t}X_{t}^{\prime}]$ with a suitable block structure. If $Y_{t}\in\mathbb{R}^{p}$ is also high-dimensional, one might also want to impose a block structure on $\mathbb{E}[X_{t}Y_{t}^{\prime}]$. A similar problem arises in regressions with many instrument, $Z_{t}$, where we can impose a block structure on $\mathbb{E}[X_{t}Z_{t}^{\prime}]$.
Theorem (ref) is applicable to regression-type estimators, such as the two-stage least squares (TSLS) estimator, $\hat{\beta}_{\mathrm{TSLS}}=[\hat{\Sigma}_{xz}\hat{\Sigma}_{zz}^{-1}\hat{\Sigma}_{zx}]^{-1}[\hat{\Sigma}_{xz}\hat{\Sigma}_{zz}^{-1}\hat{\Sigma}_{zy}]$, where the same, or different, block structures can be imposed on the matrices, $\Sigma_{zx}$, $\Sigma_{zz}$, and $\Sigma_{zy}$.
Imposing a block structure entails a bias-variance trade-off, because the block structure will induce a bias, if $\Sigma_{xz}$ does not have the block structure. Meanwhile, a large reduction in the number of parameters will reduce the variance of the estimator. In this context, the unrestricted estimate, $\frac{1}{T}\sum_{t=1}^{T}Z_{t}X_{t}^{\prime}$, is also consistent, but has a larger estimation error and many other problems, such as those arising with many instrumental variables. It would be interesting to develop formal tests for block structures in order to avoid block structures that are at odds with data. The bias-variance trade-off can motivate shrinkage methods that are based on block structures. For instance, the use of block structures could be combined with regularization methods, such as those in LedoitWolf:2004, as we elaborate on in our concluding remarks.
A block correlation matrix is characterized by correlation coefficients that form a block structure, where the correlation between two variables is solely determined by the blocks to which the two variables belong.
Block correlation matrices offer a way to parameterize large covariance matrices in a parsimonious manner. This structure is used in some multivariate GARCH models, see EngleKelly2012 and ArchakovHansenLundeMRG.
An $n\times n$ block correlation matrix, $C$, with $K$ blocks, is a symmetric block matrix with,
where $\rho_{kk}$, $k=1,\ldots,K$ are within-block correlations, and $\rho_{kl}=\rho_{lk}$, $k\neq l$, are between-block correlations, for $k,l=1,\ldots,K$. For $C$ to be a correlation matrix, we obviously need $\rho_{kl}\in[-1,1]$ for all $k,l=1,\ldots,K$, but this is insufficient to guarantee a valid correlation matrix. Negative eigenvalues can arise with some combinations of correlation coefficients, even if these are all strictly smaller than one in absolute value.
Block equicorrelation matrices corresponds to the case where the diagonal elements of all diagonal blocks, $B_{[k,k]}$ equal $d_{k}=1$, for $k=1,\ldots,K$. So, Theorem (ref) fully characterizes the set of correlation coefficients that yields a positive (semi-) definite correlation matrix. We formulate this result as a separate Corollary. Note that the canonical form, ((ref)), for $C$ in ((ref)), yields a symmetric $A$ with elements $a_{kl}=\rho_{kl}\sqrt{n_{k}n_{l}}$, for $k\neq l$, $a_{kk}=1+\rho_{kk}(n_{k}-1)$, and $\lambda_{k}=1-\rho_{kk}$.
For $C$ in ((ref)) to be a correlation matrix (possibly singular), we need that $A$ is positive semidefinite and that $|\rho_{kk}|\leq1$, and Corollary (ref) characterizes the set of positive definite block equicorrelation matrices. The additional requirements are that $A$ is positive definite and $|\rho_{kk}|<1$, $k=1,\ldots,K$.
In this context with block correlation matrices, the expression for $A$ was previously obtained in HuangYang:2010 and in CadimaCalheirosPreto2010. HuangYang:2010 focused on computational issues, which might explain that their paper is overlooked in much of the literature.\footnote{We were, until recently, also unaware of the results in HuangYang:2010 and CadimaCalheirosPreto2010. An anonymous referee (on a different paper than the present one) directed us to RoustantDeville2017 and we subsequently discovered the more detailed results in HuangYang:2010 and CadimaCalheirosPreto2010. Some of their results, e.g., HuangYang:2010, were rediscovered in RoustantDeville2017, who do not cite HuangYang:2010 or CadimaCalheirosPreto2010. In fact, none of the papers, CadimaCalheirosPreto2010, HuangYang:2010, EngleKelly2012, and RoustantDeville2017 cite any of the other papers listed here. The results in RoustantDeville2017 appear to have been absorbed and extended in RoustantEtAl2020.} Their results add valuable insight about the block-DECO model by EngleKelly2012. For instance, their results provide a simple way to evaluate if a block correlation matrix is positive definite (or semidefinite). The expression for the determinant of a correlation matrix in Corollary (ref) is a simple implication of the eigenvalues derived in HuangYang:2010 and CadimaCalheirosPreto2010, whereas the expressions for the inverse and logarithmically transformed correlation matrices are new, and so is our results on the preservation of block structures for certain matrix functions.
In this section, we focus on covariance and correlation matrices for normally distributed random variables. We derive simplified expressions for the corresponding log-likelihood functions, which greatly reduce the computational burden when $n$ is large relative to $K$. We derive the maximum likelihood estimator and provide a simple expression for the first derivatives of the log-likelihood function with respect to the unknown parameters (the scores).
We will follow the conventional notation for covariances and variances, we write $\sigma_{kl}$ in place of $b_{kl}$, $k,l=1,\ldots,K$, and $\sigma_{k}^{2}$ in place of $d_{k}$, $k=1,\ldots,K$. Similarly, for correlation matrices we write $\rho_{kl}$ in place of $b_{kl}$, and have $d_{k}=1$.
The density function for the multivariate Gaussian distribution with mean zero and an $n\times n$ covariance matrix, $\Sigma$, is $f(x)=(2\pi)^{-\frac{n}{2}}(\det\Sigma)^{-\frac{1}{2}}\exp(-\tfrac{1}{2}x^{\prime}\Sigma^{-1}x)$. Suppose that $\Sigma$ has the block structure given by $(n_{1},\ldots,n_{K})$, and let $\Sigma=QDQ^{\prime}$ be its canonical representation. The corresponding log-likelihood function (multiplied by $-2$) can be expressed as \[ -2\ell=n\log2\pi+\log\det D+X^{\prime}QD^{-1}Q^{\prime}X, \] where $D=\mathrm{diag}(A,\lambda_{1}I_{n_{1}-1},\ldots,\lambda_{K}I_{n_{K}-1})$, with $\lambda_{k}=\sigma_{k}^{2}-\sigma_{k,k}$, and \[ a_{kl}=
\] So, if we define $Y=(y_{0}^{\prime},y_{1}^{\prime},\ldots,y_{K}^{\prime})^{\prime}=Q^{\prime}X$, where $y_{0}$ is $K$-dimensional and $y_{k}$ is $n_{k}-1$ dimensional, $k=1,\ldots,K$, then it follows that
This expression shows that the block structure yields a considerable simplification for log-likelihood evaluation. Instead of inverting the $n\times n$ matrix $\Sigma$ and computing $\det\Sigma$, it suffices to invert the smaller $K\times K$ matrix, $A$, and evaluate $\det A$. Moreover, the maximum likelihood estimator based on a random sample, $X_{1},\ldots,X_{N}$, is easily expressed in terms of the transformed variables, $Y_{1}=Q^{\prime}X_{1},\ldots,Y_{N}=Q^{\prime}X_{N}$, as formulated in the following theorem.
The maximum likelihood estimates of the individual parameters can be obtained directly from $\hat{A}$ and $\hat{\lambda}_{k}$, $k=1,\ldots,K$. From the definition of $A$ it follows that $\hat{\sigma}_{k,l}=\hat{a}_{kl}/\sqrt{n_{k}n_{l}}$ for $k\neq l$, and for $k=l$, we have $\hat{\sigma}_{k,k}=(\hat{a}_{kk}-\hat{\lambda}_{k})/n_{k}$ and $\hat{\sigma}_{k}^{2}=\hat{\lambda}_{k}+\hat{\sigma}_{k,k}=\tfrac{1}{n_{k}}\hat{a}_{kk}+\tfrac{n_{k}-1}{n_{k}}\hat{\lambda}_{k}$.\footnote{This follows from the definitions, $\sigma_{k}^{2}=\lambda_{k}+\sigma_{kk}$ and $a_{kk}=\sigma_{k}^{2}+(n_{k}-1)\sigma_{kk}$, such that $a_{kk}=\lambda_{k}+\sigma_{kk}+(n_{k}-1)\sigma_{kk}=\lambda_{k}+n_{k}\sigma_{kk}$, and the invariance of the maximum likelihood estimator.}
In the special case where a block has size one, we have $\Sigma_{[k,k]}=\sigma_{k}^{2}$ and $\sigma_{k,k}$ is undefined. In this situation, the corresponding, $y_{k,t}$, is also undefined, and hence, so is $\hat{\lambda}_{k}$. Yet the expressions for the maximum likelihood estimators continue to be valid, including the expression for $\hat{\Sigma}$ in Theorem (ref). If $n_{k}=1$, then $\hat{\sigma}_{k}^{2}=\hat{a}_{kk}$, while the expression for $\hat{\sigma}_{k,k}$ is redundant and can be ignored.
Estimation when the correlation matrix is assumed to have a block structure, as opposed to the covariance matrix having a block structure, is more convoluted. A block correlation matrix is entirely given by the $A$-matrix, because the eigenvalues $\lambda_{1},\ldots,\lambda_{K}$ are given from $A$. Below we use the notation, $\mathcal{I}_{k}\equiv\{i_{k-1}+1,\ldots,i_{k}\}$, with $i_{k}=\sum_{j=1}^{k}n_{j}$, which contains the $n_{k}$ indices associated with the $k$-th block.
Thus, the estimate of $D$ can be obtained solely from the $K\times K$ matrix $\tilde{A}$. For the individual correlations we have $\tilde{\rho}_{k,l}=\tilde{a}_{kl}/\sqrt{n_{k}n_{l}}$, for $k\neq l$, and $\tilde{\rho}_{k,k}=(\tilde{a}_{kk}-1)/(n_{k}-1)$, for $k=l$. Corollary (ref) does not fully specify the maximum likelihood estimates of $(\sigma_{1}^{2},\ldots,\sigma_{n}^{2})$, but these are given in the proof in implicit form, and some additional details are stated immediately after the proof.
A simple and consistent estimator, as $T\rightarrow\infty$, is to set $\hat{\sigma}_{i}^{2}=s_{i}^{2}$, $i=1,\ldots,n$, which satisfy ((ref)), and compute the corresponding $\hat{C}$, with the expression for $A$ and $\lambda_{k}$, $k=1,\ldots,K$, which maximizes the log-likelihood function subject to $\sigma_{i}^{2}=s_{i}^{2}$, $i=1,\ldots,n$. We refer to this estimator as the two-stage estimator.
The score of the log-likelihood function is often of separate interest. For instance, the score is used for the computation of robust standard errors, in Lagrange multiplier tests, in tests for structural breaks, see e.g., Nyblom89, and in dynamic models with time-varying parameters (the so-called score-driven models), see CrealKoopmanLucasGAS_JAE. So we provide the expressions for the score in this context with a block covariance matrix.
Suppose that $\Sigma$ is a block covariance matrix and let $\Sigma=QDQ^{\prime}$ be its canonical representation. Because $Q$ is entirely given by the block partition $(n_{1},\ldots,n_{K})$, and does not depend on the unknown parameters in $\Sigma$, the expressions for the partial derivatives are relatively simple.
The hessian could be derived similarly. In some applications, it might be preferable to parametrize the block covariance matrix with $A$ and ($\lambda_{1},\ldots,\lambda_{K})$. In this case, one can use $\partial(-2\ell)/\partial A=M$, and $\partial(-2\ell)/\partial\lambda_{k}=\tfrac{n_{k}-1}{\lambda_{k}}-\tfrac{y_{k}^{\prime}y_{k}}{\lambda_{k}^{2}}$, for $k=1,\ldots,K$.
We proceed to illustrate how high-dimensional covariance matrices with a block structure are straightforward to estimate in practice. We estimate block correlation matrices for a large panel of daily asset returns for each of the years from 1995 to 2020. We include all stocks from the Center for Research in Security Prices (CRSP) database that could be matched with a unique permanent number (PERMNO from the Compustat data). Stocks with missing observations in a calendar year were excluded from the estimation in that calendar year. Across years, we have between $n=3,340$ and $n=6,637$ stocks, with an average of 4,446 stocks per year. Each calendar year has $T\simeq250$ daily returns, which we use to estimate the $n\times n$ correlation matrix for that year.
The objective of this empirical application is to demonstrate that high-dimensional covariance matrices can be estimated with relatively few observations once a block structure is imposed, and that the canonical representation makes it simple to obtain consistent estimates and evaluate the Gaussian log-likelihood function. Because variances and covariances vary over time, our estimates reflect correlations implied by the average covariance matrices over each calendar year, rather than an accurate description of the data generating process.
In our analysis, we inspect five nested block structures for the correlation matrix, where the equicorrelation structure ($K=1$) is the simplest and most restrictive model. The other four correlation models use block structures defined by GICS Sectors, Groups, Industries, and Sub-Industries. The numbers of blocks are increasing from Sectors to Sub-Industries, but may change from year to year within a given category. These correspond to $K=11$ for Sectors, and across years $K$ ranges between 24 and 26 for Groups, between 69 and 76 for Industries, and between 152 and 182 for Sub-Industries.
We estimate the canonical correlation matrix for each calendar year using the two-stage estimator. Hence, the individual variances are estimated with the sample variances, $\hat{\sigma}_{i}^{2}=T^{-1}\sum_{t=1}^{T}r_{i,t}^{2}$, where the daily returns, $r_{i,t}$, are adjusted for dividends and stock splits. Then the block correlation matrix is estimated from standardized returns, $\hat{z}_{i,t}=r_{i,t}/\hat{\sigma}_{i}$, using Corollary (ref), and we obtain $\hat{\rho}_{k,l}=\hat{a}_{kl}/\sqrt{n_{k}n_{l}}$, for $k\neq l$, and $\hat{\rho}_{k,k}=(\hat{a}_{kk}-1)/n_{k}$, where $\hat{A}=\frac{1}{T}\sum_{t=1}^{T}\hat{y}_{0,t}\hat{y}_{0,t}^{\prime}$ and the $k$-th element of $\hat{y}_{0,t}$ is given by $\frac{1}{\sqrt{n_{k}}}\sum_{i\in\mathcal{I}_{k}}\hat{z}_{i,t}$. So, the entire $n\times n$ covariance matrix with the block correlation structure is estimated by $n$ (univariate) variances and the $K\times K$ matrix, $\hat{A}$. The unrestricted estimate would obviously be singular because the dimension, $n$, is an order of magnitude larger than $T$ in all calendar years. The block assumption imposes enough structure for $\hat{C}$ to be invertible, which requires an inverse of the $K\times K$ matrix, $\hat{A}$.
To conserve space, we only provide the most detailed estimation results for the last two calendar years, 2019 and 2020, and partial correlations will be presented for six calendar years. Detailed results for all 26 calendar years (1995-2018) are presented in the Web Appendix, ArchakovHansen:BlockAppendix. Comparing 2019 and 2020 is interesting because some effects of the COVID-19 pandemic can be observed in 2020. Both years have the same number of assets, $n=3,340$ assets, and the same number of blocks for all five block structures.
In Table (ref), we report the range of estimated correlations for each of the five types of block structure and for both calendar years, 2019 and 2020. The range of estimated correlations was larger in 2020 than in 2019, and it increases as the number of blocks increases. The latter is expected because the number of distinct correlation coefficients in $C$ increases as $K$ increases. Each estimated correlation represents an average correlation, subject to both time averaging (over a calendar year) and cross-sectional averaging within the corresponding sector/group/industry/sub-industry. We also report the log-likelihood function (scaled by $-2/(nT)$) evaluated at the parameter estimates, and the corresponding value of the Bayesian Information Criterion (BIC).\footnote{These are approximate BIC statistics, because they are based on the two-estimator.} The minimum BIC is obtained with a block structures based on Groups in both 2019 and 2020.\footnote{The BIC adds the penalty $p\log(nT)$ to $-2\ell$, where $p$ is the number of free parameters. For comparison, the AIC, which uses the penalty $2p$, selects the most general specification in both years.} The last column reports the number $K(K+1)/2$ of unique correlations within a block structure with $K$ blocks, and while this number increases rapidly with $K$, the gains in the log-likelihood are relatively modest. Consequently, the BIC increases substantially when blocks are defined by sub-industries.
The estimated block correlation matrices based on Sectors, Groups, Industries and Sub-industries are shown in Figure (ref). Along the diagonal are the estimated correlation coefficients for assets in the same block (within block correlations). Other estimates are for pairs of assets from different blocks (between block correlations). Due to the larger number of estimated correlation coefficients, we present most estimates using color coding. A darker shade of red denotes a stronger correlation.
From Figure (ref) we can see that correlations were generally higher in 2020 than in 2019. This can be attributed to the COVID-19 Pandemic. When COVID-19 cases began to spread worldwide, beyond isolated cases, the market experienced a large decline that continued as lockdowns were imposed in most countries. The S&P 500 index declined by more than 33% from February 19, 2020 to March 23, 2020. This was followed by a strong rally where the market, from March 23, 2020 to the end of the year, increased by more than 63%. The block structure is somewhat more visible in 2020, which could be due to the differentiated effect the pandemic had on different sectors of the economy. For instance, in Figure (ref) we can see that the correlation between Utilities (55) and other sectors increased substantially. Finance (40) tended to have the highest average correlation with other sectors, but Utilities (55) had the highest average correlation with other sectors in 2020. The low correlation between Health Care (35) and other sectors is also very pronounced in 2020. The blocks are listed in order of their GICS codes and we include solid black lines to separate different sectors (the number of blocks is too large to include labels for Industries and Sub-Industries individually). The block partition based on Sub-Industries in panel (d)\footnote{These are identical to the results presented in the Introduction in Figure (ref).} reveals additional details about the correlation structure. Within the Health Care sector (35) there is a distinct stripe that signifies near zero correlations with all other blocks. Interestingly, this stripe corresponds to two subindustries, Biotechnology (35201010) and Pharmaceuticals (35202010). In the Figure (ref)(d) there is also a band with low correlations that is associated with sub-industries in Materials (15), and these are Gold (15104030), Precious Metals and Minerals (15104040), and Silver (15104045).
To further illustrate the usefulness of the block correlation structure, we compute partial correlations for pairs of stocks. Partial correlations require inversion of a high-dimensional matrix, a computation that is greatly simplified with the block structure. In Figure (ref) we report the partial correlation for a pair of stocks, where we have conditioned on all other stocks in other sectors. These partial correlations are based on the estimated correlation matrices using the sector-block structure.\footnote{The block structure greatly simplifies the computation of this type of partial correlation. The formulae are derived and presented in the Web Appendix.} Figure (ref) includes results for six calendar years, the results for all other calendar years are presented in the Web Appendix.
One feature that stands out from the partial correlation matrices is the similarity across calendar years, of which six years are shown in Figure (ref). Had the annual estimates of the correlation matrices been very noisy, we would not expect to see very similar structures in estimates based on different data sets (daily returns from different calendar years). The estimate from one calendar year, tend to be similar to that of the neighboring years, with some exceptions associated with the Global Financial Crisis and the COVID-19 pandemic. This is precisely what we would expect if the correlation structure is time-varying but typically evolves in a relatively smooth manner. It is interesting that two sectors, Energy and Utilities, have large degrees of residual correlation that are left unexplained after having conditioned on all stocks in other sectors. This indicates that these sectors need a sector specific factor to explain their correlation structure. A potentially interesting application of the partial correlation analysis, would be to extend the set of assets with a set of “factors”, such as the three Fama-French factors, and other candidate factors. Computing partial correlations, where the conditioning is on the factors, could be used to identify correlation structures that are left unexplained by the factors. We leave this for future research.
Figure (ref) presents selected results for all 26 calendar years. In the upper panel (a), we present the estimated correlations (left) and partial correlations (right) between assets in the Energy and Utilities based on the sector-block structure. The shaded areas represent the average correlation and average partial correlation based on the equicorrelation structure for all assets. Figure (ref) shows that the correlations have been trending upwards and there has also been a great deal of variation in the correlation between these two sectors. The partial correlations are interesting, because they indicate that these two sectors, Energy and Utilities, have large idiosyncratic components, because a large fraction of the correlations between stocks within either of these two sectors is unexplained by the thousands of stocks in other sectors.
In the lower panel (b) of Figure (ref), we present the BIC for each of the calendar years and each block structure. Before 2000, the BIC always selected the block structures based on sectors and after 2000 it systematically favors the block structure based on Groups. The most heavily parametrized specification, which is based on sub-industries, has the worst BIC in all calendar years.
We have derived a canonical representation of block matrices. The representation provides valuable simplifications for models with block matrices, such as stochastic block models for large networks, and models with block covariance and block correlation matrices. We derived a number of expressions that greatly simplify the computation of the Gaussian log-likelihood function with block covariance/correlation matrices. We illustrated this in an empirical application, where we estimate large covariance matrices for a vector with thousands of assets, with daily returns over a single calendar year. Inverting the covariance matrix, computing partial correlations, and evaluating the Gaussian log-likelihood is straightforward once a block structure is imposed.
The canonical representation and the related results are potentially useful for regularizing large covariance matrices. For instance, one could shrink the sample correlation matrix towards a block correlation matrix, analogous to the way LedoitWolf:2004 proposed to shrink towards the equicorrelation matrix with the MacGyver method, see also Engle:2009. This could possibly be extended to shrinkage involving a convex combination of several block correlation matrices.
The canonical representation also paves new ways to testing block structures in covariance and correlation matrices. This predominantly amounts to testing a large number of zero-restrictions in the canonical representation. We identified a number of transformations that preserves the block structures, so testing of block structures could be based on any of the transformations, rather than the original matrix. For instance, block structures in a correlation matrix $C$ would be tested on the canonical representation for $\log C$. This is potentially interesting, because the connection between logarithmically transformed correlation matrix and the Fisher transformation, see ArchakovHansen:Correlation. Finally, the group assignments, and hence $K$, will be unknown in many empirical applications. The literature has therefore proposed various classification methods to determine an appropriate block structure. It is possible that the canonical representation will be useful for this type of classification problems.