Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
21,758 characters · 7 sections · 20 citation commands
Characterizing Correlation Matrices that Admit a Clustered Factor Representation
Keywords:{ Block correlation matrix, copula, clustering, factor model.}
JEL Classification:{ C38 }
In empirical models involving high-dimensional covariance matrices, it is often beneficial to introduce a parsimonious structure to mitigate issues stemming from overfitting. A commonly used approach is the Clustered Factor (CF) model, which is characterized by a linear factor structure and group clustering, where factor loadings are shared within each group.
The clustered factor model emerged from high-dimensional copula models with a multi-factor structure, see e.g KrupskiiJoe2013 and CrealTsay2015. Clustering became a popular additional structural component, because it is easy to interpret and makes the model more parsimonious, see KrupskiiJoe2015, OhPatton2017JBES, MannerStarkWied2019, and OpschoorLucasBarraVanDick:2021.\nocite{OhPatton2017}\nocite{OhPatton2023}
It is well known that the CF model induces a block structure on the correlation matrix (whenever variables are ordered by group assignments). However, the CF model is not consistent with all block correlation matrices. It cannot generate negative within-group correlations. However, this is unlikely to be problematic in empirical applications, because variables in the same cluster are expected to share traits and be positively correlated. It is more problematic that the CF model imposes other restrictions that rules out a class of block correlation matrices, which seem perfectly reasonable and ordinary from an empirical viewpoint. Below we fully characterize the class of block correlation matrices that are coherent with the CF model, and express the superfluous restrictions the CF model imposes as a testable hypothesis. We then proceed to highlight an alternative parametrization of block correlation matrices by ArchakovHansen:CanonicalBlockMatrix, which is based on the matrix logarithm of the correlation matrix of ArchakovHansen:Correlation. This parametrization uniquely parametrizes any non-singular block correlation matrix, without imposing additional structure. A simple linear factor structure is also applicable to this parametrization, if a more parsimonious model is required.
The rest of this paper is organized as follows. In Section (ref), we introduce the required notation and the CF model followed by the theoretical results that characterizes the block correlation matrices that admit CF model. We express the CF structure as a testable hypothesis, and detail a testing procedure for selecting the number of factors. In Section 3, we presents the alternative parametrization by ArchakovHansen:CanonicalBlockMatrix and illustrate it with a simple example. We conclude in Section 4 with a brief summary.
Consider an $n$-dimensional random variable, $X\in\mathbb{R}^{n}$, with non-singular covariance matrix $\Sigma=\mathrm{var}(X)$. Then the corresponding correlation matrix is given by \[ C=\Lambda_{\sigma}^{-1}\Sigma\Lambda_{\sigma}^{-1}, \] where $\Lambda_{\sigma}=\mathrm{diag}(\sqrt{\Sigma_{11}},\ldots,\sqrt{\Sigma_{nn}})$, and the vector of standardized variables, \[ Z=\Lambda_{\sigma}(X-\mu),\qquad\mu=\mathbb{E}X, \] is such that $C=\mathbb{E}(ZZ^{\prime})$.
In the copula literature, it is common to parametrize $C$ with the clustered factor (CF) model, which is characterized by a linear factor structure for $Z$ and a partitioning of the variables into clusters/groups. The factor structure is defined by,
where $f\in\mathbb{R}^{r}$ is a common vector with $r$ orthogonal factors, $\mathbb{E}f=0$ and $\mathrm{var}(f)=I_{r}$, and the idiosyncratic variables $(\varepsilon_{1},\ldots,\varepsilon_{n})$ are mutually uncorrelated and uncorrelated with the factors, $\mathrm{cov}(f,\varepsilon_{i})=0$ and $\mathrm{cov}(\varepsilon_{i},\varepsilon_{j})=0$ for all $i\neq j$. The vectors of factor loadings are given by $b_{i}={\rm corr}(f,Z_{i})\in\mathbb{R}^{r}$, for $i=1,\ldots,n$, and it follows that $\mathrm{var}(\varepsilon_{i})=1-b_{i}^{\prime}b_{i}$.
A partitioning of the $n$ elements into $K<n$ groups adds additional structure. It is assumed that the factor loadings are common within each group, such that
where $({\rm G}_{1},\ldots,{\rm G}_{K})$ is a partition of $\{1,\ldots,n\}$, with ${\rm G}_{k}$ containing $n_{k}$ elements, $k=1,\ldots,K$, such that $n=\sum_{k}n_{k}$.
The CF model implies that the correlation between two variables is solely determined by their group classification. For $i\in{\rm G}_{k}$ and $j\in{\rm G}_{l}$ with $i\neq j$ we have \[ C_{ij}={\rm corr}(z_{i},z_{j})=\rho_{kl}\equiv\beta_{k}^{\prime}\beta_{l}. \] By rearranging $X_{1},\ldots,X_{n}$ according to their group assignments, we obtain a block correlation matrix that can be expresses as
where $C_{[k,l]}$ is an $n_{k}\times n_{l}$ matrix given by \[ C_{[k,k]}=\left[
\right]\ensuremath{\quad}{\rm and}\ensuremath{\quad C_{[k,l]}=\left[
\right]\quad}{\rm if}\ \ensuremath{k\neq l}. \]
The advantage of a block correlation matrix, is that its number of unique correlations is (at most) $d=K\left(K+1\right)/2$, whereas the number of distinct correlations in an unrestricted $n\times n$ correlation matrix is $n\left(n-1\right)/2$. If $K$ is fixed, then $d$ does not increase with $n$, which makes it possible to model high-dimensional correlation matrices with relatively few parameters.
An interesting question is whether a give block correlation matrix can be expressed as a CF model. Any block correlation matrix is clearly coherent with a group partitioning ((ref)), but additional structure is needed for it to have the representation in ((ref)). In some applications, it will be relevant to know if the factor structure rules out empirically relevant correlation matrices. It is therefore interesting to characterize the set of block correlation matrices that are compatible with the CF model.
For later use, we define the matrix $B=\left(\beta_{1},\beta_{2},\ldots,\beta_{K}\right)^{\prime}\in\mathbb{R}^{K\times r}$, where $\beta_{k}\in\mathbb{R}^{r\times1}$ is the vector of group-specific factor loadings, as defined in ((ref)) and ((ref)).
From a block correlation matrix, $C$, we define the following $K\times K$ matrix
In this paper we focus, without loss of generality, on the case where $C$ is nonsingular, which implies that $\rho_{kk}<1$ for all within-group correlations, $k=1,\ldots,K$.
Although Lemma (ref) shows that a psd $A^{*}$ with $\rho_{kk}<1$ implies that $A$ is positive definite, the converse is not true. We have that $A^{*}=\Lambda_{n}^{-1/2}\left(A-\Psi_{1-\rho}\right)\Lambda_{n}^{-1/2}$, which is psd if and only if $A-\Psi_{1-\rho}$ is psd. The latter is not guaranteed. The simplest example is if a within-group correlation is negative, since the diagonal elements of $A-\Psi_{1-\rho}$ equal $n_{k}\rho_{kk}$, $k=1,\ldots,K$. Another, less obvious, example is the following block correlation matrix:
which is a positive definite, it's smallest eigenvalue is $\lambda_{{\rm min}}(C)=0.26$. However, \[ A^{*}=\left[
\right], \] has a negative eigenvalue, $\lambda_{\min}(A^{\ast})=-0.02$, and $C$ can therefore not be expresses as a CF model. It seems unjustified that the Clustered Factor (CF) model precludes the correlation matrix in ((ref)) (and similar matrices) in advance, because there does not appear to be anything bizarre or unusual about this particular block correlation matrix.
Another question is how the number of factors, $r$, are needed to to generate a given correlation matrix. We address this in the following Theorem.
Theorem 2 shows that increasing the number of factors beyond $K$ does not broaden the range of attainable correlation matrices. The (maximum) number of free parameters in the CF model is $K(K+1)/2$, because we can always rotate the latent factors, $f$, such that $B$ is upper (or lower) triangular.
Theorem (ref) shows that the block correlation matrices that admits a CF representation, are characterized the smallest eigenvalue of $A^{\ast}$ being nonnegative, i.e. $\lambda_{\min}(A^{\ast})\geq0$. Here $A^{\ast}$ is defined by ((ref)) where we set $\rho_{kk}=1$ whenever $n_{k}=1$.
In practice, it is therefore relatively straight forward to test the CF representation using the null hypothesis: \[ H_{F}:\lambda_{\min}(A^{\ast})\geq0. \]
Given a block correlation matrix with a CF structure, we can proceed to estimate the minimum number of factors, $r$, as defined by the rank of $A^{\ast}$. This could be done with the sequential procedure propose in Pantula:1989, which is commonly used to determine the cointegration rank in VAR models, see Johansen88.\footnote{The same testing principle is used for lag-length selection in time series model, see NgPerron01, and in multiple comparisons to determine the model confidence set, see HansenLundeNasonMCS.}\nocite{Johansen91}
Let $\lambda_{1}\geq\lambda_{2}\geq\cdots\geq\lambda_{K}$ denote the eigenvalues of $A^{\ast}$, and consider the hypothesis $H_{F,q}:\lambda_{r+1}=\cdots=\lambda_{K}=0$, implying that $A^{\ast}$ is psd with $\mathrm{rank}(A^{\ast})\leq r$. The sequential procedure begins by testing $H_{F,0}$ against $H_{F,K}$. In the event of a rejection, we proceed to test $H_{F,1}$ against $H_{F,K}$, and so forth until a null hypothesis is not rejected. We can set $\hat{r}=r^{\ast}$, where $r^{\ast}$ is the first instance (the smallest $r$) where $H_{F,r}$ is not rejected. The degrees of freedom in $B$ for a model with $r\leq K$ factors is $r[K+(K-r)+1]/2$.
A new parametrization of correlation matrices was proposed in ArchakovHansen:Correlation, and is defined by
where ${\rm vecl}(\cdot)$ extracts and vectorizes the elements below the diagonal and $\log C$ is the matrix logarithm of the correlation matrix.\footnote{For a nonsingular correlation matrix, we have $\log C=Q\log\Lambda Q^{\prime}$, where $C=Q\Lambda Q^{\prime}$ is the spectral decomposition of $C$, so that $\Lambda$ is a diagonal matrix with the eigenvalues of $C$.} The identity ((ref)) defines a one-to-one mapping between $\mathbb{R}^{d}$ and the set of non-singular $n\times n$ correlation matrices.
Because the matrix logarithm preserves the block structure in $C$, see ArchakovHansen:CanonicalBlockMatrix, $\gamma$ will contain many “duplicates”. Thus, we can parametrize block correlation matrices with $\eta\in\mathbb{R}^{q}$, where $\eta$ is a subvector of $\gamma$ and $q\leq K(K+1)/2$. The follow example will serve as an illustration,{ \[ \ensuremath{\underbrace{\left[
\right]}_{=C}\quad\underbrace{\left[
\right]}_{=\log C}}. \] }Here we can use $\eta=(
)^{\prime}$ as the condensed vector parametrization of the block correlation matrix. For a given block partitioning, $(n_{1},\ldots,n_{K})$, any non-singular block correlation matrix will map to a unique vector $\eta$, and any vector $\eta\in\mathbb{R}^{q}$ will map to a unique non-singular block correlation matrix, see ArchakovHansen:CanonicalBlockMatrix.
It is straightforward to add additional structure onto the $\eta$-parametrization, for instance by restricting $\eta$ to be in a subspace of lower dimension than $q$. For a multivariate GARCH model, these ideas are explored in ArchakovHansenLundeMRG.
CrealKim:BayesianBlockCorr recently adopted this parametrization for Bayesian modeling of block correlation matrices. An attractive feature of this parametrization, it that it facilitates priors with full support on the entire set of non-singular block correlation matrices.
We have characterized the class of block correlation matrices that can be expressed as a clustered factor model. While the clustered factor model serves as a valuable tool for generating clustered correlation structures, it does introduce unnecessary constraints on the correlation matrix, which may limit its practical relevance for some empirical problems. The alternative parametrization, which is based on the matrix logarithm of the correlation matrix, seems better suited for the modeling of block correlation matrices. It avoids the imposition of superfluous constraints on the correlation matrix and, if needed, it provides a flexible framework for further reduction of the degrees of freedom.