EconBase
← Back to paper

Cluster GARCH

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

120,504 characters · 31 sections · 68 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Cluster GARCH

abstractWe introduce a novel multivariate GARCH model with flexible convolution-$t$ distributions that is applicable in high-dimensional systems. The model is called Cluster GARCH because it can accommodate cluster structures in the conditional correlation matrix and in the tail dependencies. The expressions for the log-likelihood function and its derivatives are tractable, and the latter facilitate a score-drive model for the dynamic correlation structure. We apply the Cluster GARCH model to daily returns for 100 assets and find it outperforms existing models, both in-sample and out-of-sample. Moreover, the convolution-$t$ distribution provides a better empirical performance than the conventional multivariate $t$-distribution.

Keywords:{ Multivariate GARCH, Score-Driven Model, Cluster Structure, Block Correlation Matrix, Heavy Tailed Distributions.}

JEL Classification:{ G11, G17, C32, C58 }

Introduction

Univariate GARCH models have enjoyed considerable empirical success since they were introduced in Engle:1982 and refined in bollerslev:86. In contrast, the success of multivariate GARCH models has been more moderate due to a number of challenges, see e.g. BauwensLaurentRombouts:2006. A common approach to modeling covariance matrices is to model variances and correlations separately, as is the case in the Constant Conditional Correlation (CCC) model by Bollerslev1990 and the Dynamic Conditional Correlation (DCC) model by Engle2002. See also EngleSheppard:2001, TseTsui:2002, Aielli:2013, EngleLedoitWolf:2019, and PakelShephardSheppardEngle:2021. While univariate conditional variances can be effectively modeled using standard GARCH models, the modeling of dynamic conditional correlation matrices necessitates less intuitive choices to be made. One challenge is that the number of correlations increases with the square of the number of variables, a second challenge is that the conditional correlation matrix must be positive semidefinite, and a third challenge is to determine how correlations should be updated in response to sample information.

In this paper, we develop a novel dynamic model of the conditional correlation matrix, the Cluster GARCH model, which has three main features. First, use convolution-$t$ distributions, which is a flexible class of multivariate heavy-tailed distributions with tractable likelihood expressions. The multivariate $t$-distributions are nested in this framework, but a convolution-$t$ distribution can have heterogeneous marginal distributions and cluster-based dependencies. For instance, convolution-$t$ distributions can generate the type of sector-specific price jumps reported in AndersenDingTodorov:2024. Second, the dynamic model is based on the score-driven framework by CrealKoopmanLucas:2013, which leads to closed-form expressions for all key quantities. Third, the model can be combined with a block correlation structure that makes the model applicable to high-dimensional systems. This partitioning, defining the block structure, can also be interpreted as a second type of cluster structure.

Heavy-tailed distributions are common in financial returns, and many empirical studies adopt the multivariate $t$-distribution to model vectors of financial series, e.g., Kotz2004, Harvey2013, and IbragimovIbragimovWalden:2015. An implication of the multivariate $t$-distribution is that all standardized returns have identical and time-invariant marginal distributions. This is a restrictive assumption, especially in high dimensions. The convolution-$t$ distributions by HansenTong:2024 relax these assumptions, and one of the main contributions of this paper is to incorporate this class of distributions into a tractable multivariate GARCH model. A convolution-$t$ distribution is a convolution of multivariate $t$-distributions. In the Cluster GARCH model, standardized returns are time-varying linear combinations of independent $t$-distributions, which can have different degrees of freedom. This leads to dynamic and heterogeneous marginal distributions for standardized returns, albeit the conventional multivariate $t$-distribution is nested in this framework as a special case. We focus on three particular types of convolution-$t$ distributions, labelled Canonical-Block-$t$, Cluster-$t$, and Hetero-$t$. These all have relatively simple log-likelihood functions, such that we can obtain closed-form expressions for the first two derivatives, score and information matrix, of the conditional log-likelihood functions. These are used in our score-driven model for the time-varying correlation structure, which is a key component of the Cluster GARCH model.

High-dimensional correlation matrices can be modeled using a parsimonious block structure for the conditional correlation matrix. The DCC model is parsimonious but greatly limits the way the conditional covariance matrix can be updated. Without additional structure, the number of latent variables increases with $n^{2}$, where $n$ is the number of assets. This number becomes unmanageable once $n$ is more than a single digit, and maintaining a positive definite correlation matrix can be challenging too. The correlation structure in the Block DECO model by EngleKelly:2012 is an effective way to reduce the dimension of the estimated parameters. However, the estimation strategy in EngleKelly:2012 was based on an ad-hoc averaging of within-block correlations for an auxiliary DCC model, and they did not fully utilize the simplifications offered by the block structure.\footnote{They derived likelihood expressions for the case with $K=2$ blocks. For more two blocks, $K>2$, they resort to a composite likelihood evaluation.} The model proposed in this paper draws on recent advances in correlation matrix analysis by ArchakovHansen:Correlation\nocite{ArchakovHansen:CanonicalBlockMatrix}. We will, in some specifications, adopt the block parameterization of the conditional correlation matrix, used in ArchakovHansenLundeMRG, which has (at most) $K\left(K+1\right)/2$ free parameters where $K$ is the number of blocks. This approach guarantees a positive definite correlation matrix and the likelihood evaluation is greatly simplified. Overall, the Cluster GARCH offers a good balance between flexibility and computational feasibility in high dimensions.

We adopt the convenient parametrization of the conditional correlation matrix, $\gamma(C)$, which is defined by taking the matrix logarithm of the correlation matrix, $C$, and stacking the off-diagonal elements of $\log C$ into the vector, $\gamma\in\mathbb{R}^{d}$, where $d=n(n-1)/2$. This parametrization was introduced in ArchakovHansen:Correlation and the mapping $C\mapsto\gamma(C)$ is one-to-one between the set of non-singular correlation matrices $\mathcal{C}_{n\times n}$ and $\mathbb{R}^{d}$. So, the inverse mapping, $C(\gamma)$, will always yield a positive definite correlation matrix and any non-singular correlation matrix can be generated in this way. The parametrization can be viewed as a generalization of Fisher\textquoteright s Z-transformation to the multivariate case. It has attractive finite sample properties, which makes it suitable for an autoregressive model structure, see ArchakovHansen:Correlation.

commentThe two coincide when $n=2$ and $\hat{\gamma}=\gamma(\hat{C})$ has attractive finite sample properties, as is the case for the Fisher transformed correlation, see ArchakovHansen:Correlation, and this makes it very suitable for dynamic modeling with an autoregressive structure.

A block correlation structure arises when variables can be partitioned into clusters, $K$ say, and the correlation between two variables is determined by their cluster assignments. When $C$ has a block structure, then $\log C$ also has a block structure. This leads to a new parametrization of block correlation matrices, which defines a one-to-one mapping $C\mapsto\eta(C)$ between the set of non-singular block correlation matrices $\mathcal{C}_{n\times n}$ and $\mathbb{R}^{d}$ with $d=K\left(K+1\right)/2$. We adopt the canonical representation by ArchakovHansen:CanonicalBlockMatrix, which is a quasi-spectral decomposition of block matrices that diagonalizes the matrix with the exception of a small $K\times K$ submatrix. This decomposition makes the model parsimonious and greatly simplifies the evaluations of the log-likelihood function. This parameterization of block correlation matrices is more general than the factor-based approach to parametrizing block correlation matrices.\footnote{The factor-induced block structure, see CrealTsay2015, OpschoorLucasBarraVanDick:2021, and OhPatton2023, entails superfluous restrictions on $C$, see TongHansen:2023. Both approach simplifies the computation of $\det C$ and $C^{-1}$, but only the parametrization based on the canonical representation simplifies the evaluation of the likelihood function for the convolution-$t$ distributions.}

Our paper contributes to the literature on score-driven model for dynamics of covariance matrices. Using the multivariate $t$-distribution, CrealKoopmanLucasJBES:2012 and HafnerWang:2023 proposed score-driven model for time-varying covariance and correlation matrix, respectively.\footnote{The model by HafnerWang:2023 update parameters using the unscaled score, i.e., they did not use the information matrix.} OhPatton2023 proposed a score-driven dynamic factor copula models with skew-$t$ copula function, however, the analytical information matrices in these copula models are not available. Using realized measures of the covariance matrix, GorgiHansenJanusKoopman:2019 proposed the Realized Wishart-GARCH, which relies on a Wishart distribution for realized covariance matrices and on a Gaussian distribution for returns. OpschoorJanusLucasVanDick:2017 constructed a multivariate HEAVY model based on Heavy-tailed distributions for both returns and the realized covariances. An aspect, which sets the Cluster GARCH apart from the existing literature, is that the model is based on the convolution-$t$ distributions, which includes the Gaussian distribution and the multivariate $t$-distributions as special cases. The block structures we impose on the correlation matrix in some specifications, was previously used in ArchakovHansenLundeMRG. Their model used the Realized GARCH framework with a Gaussian specification, whereas we adopt the score-driven framework for convolution-$t$ distributions, and do require realized volatility measures in the modeling.

We conduct an extensive empirical investigation on the performance of our dynamic model for correlation matrices. The sample period spans the period from January 3, 2005 to December 31, 2021. The new model is applicable to high dimensions, and we consider a “small universe” with $n=9$ assets and a “large universe” with $n=100$ assets. The small universe allows us compare the new models with a range of existing models, as most of these are not applicable to the large universe. We also undertake an more detailed specification analysis with the small universe. The nine stocks are from three sectors, three from each sector, which motivates certain block and cluster structures. First, we find that the convolution-$t$ distribution offers a better fit than the conventional $t$-distribution. Overall, the Cluster-$t$ distribution has the largest log-likelihood value. Second, we find that score-driven models successfully captures the dynamic variation in the conditional correlation matrix. The new score-driven models outperform traditional DCC models when based on the same distributional assumptions, and the proposed score-driven model with a sector motivated block correlation matrix has the smallest BIC.

The large universe with $n=100$ stocks poses no obstacles for the Cluster GARCH model. We used the sector classification of the stocks to define the block structure in the correlation matrix. We also used the sector classification to explore possible cluster structures in the tail-dependencies, which are related to parameters in the convolution-$t$ distribution. With $K=10$ sectors this reduces the number of free parameters in the correlation matrix from 4950 to 55, and the model estimation is very fast and stable, in part because the required computations only involve $K\times K$ matrices (instead of $n\times n$ matrices). For the large universe, the empirical results favor the Hetero-$t$ specification, which entails a convolutions of a large number of univariate $t$-distributions. We also find that correlation targeting, which is analogous to variance targeting in GARCH models, is beneficial.

The rest of this paper is organized as follows: In Section (ref) we introduce a new parametrization of block correlation matrices, based on ArchakovHansen:Correlation and ArchakovHansen:CanonicalBlockMatrix. In Section 3, we introduce the convolution-$t$ distributions. We derive the score-driven models in Section 4, and we obtain analytical expressions for the score and information matrix for the convolution-$t$ distributions, including the special case where $C$ has a block structure. Some details about practical implementation are given in Section 5. The empirical analysis is presented in Section 6 and includes in-sample and out-of-sample evaluations and comparisons. All proofs are given in the Appendix.

The Theoretical Model

Consider an $n$-dimensional time-series, $R_{t}$, $t=1,2,\ldots,T$, and let $\{\mathcal{F}_{t}\}$ be a filtration to which $R_{t}$ is adapted, i.e. $R_{t}\in\mathcal{F}_{t}$. We denote the conditional mean by $\mu_{t}=\mathbb{E}(R_{t}|\mathcal{F}_{t-1})$ and the conditional covariance matrix by $\Sigma_{t}=\mathrm{var}(R_{t}|\mathcal{F}_{t-1})$. With $\Lambda_{\sigma_{t}}\equiv\mathrm{diag}(\sigma_{1t},\ldots,\sigma_{nt})$, where $\sigma_{it}^{2}=\mathrm{var}(R_{it}|\mathcal{F}_{t-1})$, $i=1,\ldots,n$, it follows that the conditional correlation matrix is given by \[ C_{t}=\Lambda_{\sigma_{t}}^{-1}\Sigma_{t}\Lambda_{\sigma_{t}}^{-1}. \] Initially, we take $\mu_{t}$ and $\Lambda_{\sigma_{t}}$ as given and focus on the dynamic modeling of $C_{t}$. We are particularly interested in the case where $n$ is large. We define the following standardized variables with a dynamic correlation matrix $C_{t}$, \[ Z_{t}=\Lambda_{\sigma_{t}}^{-1}(X_{t}-\mu_{t}). \]

To simplify the notation, we omit subscript-$t$ in most of Sections (ref) and (ref) and reintroduce it again in Section (ref) where the dynamic model is presented.

Block Correlation Matrix

If $n$ is relatively small, we can model the dynamic correlation matrix using $d=n(n-1)/2$ latent variables. Additional structure on $C$ is required when $n$ is larger, because the number of latent variables becomes unmanageable. Additional structure can be imposed using a block structures on $C$, as in EngleKelly:2012.

A block correlation matrix is characterized by a partitioning of the variables into clusters, such that the correlation between two variables is solely determined by their cluster assignments. Let $K$ be the number of clusters, and let $n_{k}$ be the number of variables in the $k$-th cluster, $k=1,\ldots,K$, such that $n=\sum_{k=1}^{K}n_{k}$. We let $\bm{n}=\left(n_{1},n_{2},\ldots,n_{K}\right)^{\prime}$ be the vector with cluster sizes and sort the variables such that the first $n_{1}$ variables are those in the first cluster , the next $n_{2}$ variables are those in the second cluster, and so forth. Then $C=\mathrm{corr}(Z)$ will have the following block structure

equation[equation omitted — 211 chars of source]

where $C_{[k,l]}$ is an $n_{k}\times n_{l}$ matrix given by \[ C_{[k,l]}=\left[

array[array omitted — 108 chars of source]

\right], for k\neq l\qquadand C_{[k,k]}=\left[

array[array omitted — 127 chars of source]

\right]. \] Each block, $C_{[k,l]}$, has just one correlation coefficient, such that the block structure reduces the number of unique correlations from $n\left(n-1\right)/2$ to at most $K\left(K+1\right)/2$.\footnote{This is based on the general case that the number of assets in each group is at least two. When there are $\tilde{K}\leq K$ groups with only one asset, this number become $K\left(K+1\right)/2-\tilde{K}$. The reason for the distinction between these two cases is that an $1\times1$ diagonal block has no correlation coefficients.} This number does not increase with $n$, and this makes it possible to scale the model to accommodate high-dimensional correlation matrices.

Below we derive score-driven models for unrestricted correlation matrices and for the case where $C$ has a block structure. time{]}.\footnote{It is unproblematic to extend the model to allow for some missing observations and occasional changes in the cluster assignments. }

Parametrizing the Correlation Matrix

We parameterize the correlation matrix with the vector

equation[equation omitted — 116 chars of source]

where ${\rm vecl}(\cdot)$ extracts and vectorizes the elements below the diagonal and $\log C$ is the matrix logarithm of the correlation matrix.\footnote{For a nonsingular correlation matrix, we have $\log C=Q\log\Lambda Q^{\prime}$, where $C=Q\Lambda Q^{\prime}$ is the spectral decomposition of $C$, so that $\Lambda$ is a diagonal matrix with the eigenvalues of $C$.} The following example illustrates this parametrization for an $3\times3$ correlation matrix:{ \[ {\rm vecl}\left[\log\left(

array[array omitted — 81 chars of source]

\right)\right]={\rm vecl}\left[\left(

array[array omitted — 156 chars of source]

\right)\right]=\left(

array[array omitted — 34 chars of source]

\right)=:\gamma. \] }This parametrization is convenient because it guarantees a unique positive definiteness correlation matrix, $C(\gamma)$ for any vector $\gamma,$ without imposing superfluous restrictions on the correlation matrix, see ArchakovHansen:Correlation.

For a block correlation matrix the logarithmic transformation preserves the block structure as illustrated in the following example:{ \[ \ensuremath{\underbrace{\left[

array[array omitted — 893 chars of source]

\right]}_{=C}\quad\underbrace{\left[

array[array omitted — 942 chars of source]

\right]}_{=\log C}}. \] }The parameter vector, $\gamma$ will only have as many unique elements as there are different blocks in $C$. This number is $(K+1)K/2$, and we can therefore condense $\gamma$ into a subvector, $\eta$, such that

equation[equation omitted — 51 chars of source]

where $B$ is a known bit-matrix with a single one in each row and $\eta\in\mathbb{R}^{K(K+1)/2}$. This factor structure for $\gamma$ was first proposed in ArchakovHansenLundeMRG.

For later use, we define the condensed log-correlation matrix, $\tilde{C}\in\mathbb{R}^{K\times K}$, whose $(k,l$)-th element is the off-diagonal element from the $(k,l)$-th block of $\log C$, $k,l=1,\ldots,K$, and we can set $\eta=\mathrm{vech}(\tilde{C})\in\mathbb{R}^{K(K+1)/2}$. In the example above, we have \[ \tilde{C}=\left[

array[array omitted — 178 chars of source]

\right], \] such that $\eta=\left[1.02,0.251,0.115,0.626,0.036,0.259\right]^{\prime}$ has dimension six whereas $\gamma$ has dimension 21. Since the block correlation matrix, $C$, is only a function of $\eta$ we can model the time-variation in $C$ using a dynamic model for the unrestricted vector $\eta$. This will be our approach below.

Canonical Form for the Block Correlation Matrix

Block matrices has a canonical representation that resembles the eigendecomposition of matrices, see ArchakovHansen:CanonicalBlockMatrix. For a block correlation matrix with block-sizes, $(n_{1},\ldots,n_{K})$, we have

equation[equation omitted — 302 chars of source]

where the upper left block, $A$, is an $K\times K$ matrix with elements $A_{kl}=\rho_{kl}\sqrt{n_{k}n_{l}}$, for $k\neq l,$and $A_{kk}=1+\left(n_{k}-1\right)\rho_{kk}$. The matrix $Q$ is a group-specific orthonormal matrix, i.e., $Q^{\prime}Q=QQ^{\prime}=I_{n}$. Importantly, $Q$ is solely determined by the block sizes, $\ensuremath{(n_{1},\ldots,n_{K}})$, and does not depend on the elements in $C$. This matrix is given by \[ \ensuremath{Q=\left[

array[array omitted — 246 chars of source]

\right]}, \] where $v_{n_{k}}=(1/\sqrt{n_{k}},\ldots,1/\sqrt{n_{k}})^{\prime}\in\mathbb{R}^{n_{k}}$ and $v_{n_{k}}^{\perp}$ is an $n_{k}\times(n_{k}-1)$ matrix, which is orthogonal to $v_{n_{k}}$, i.e., $v_{n_{k}}^{\prime}v_{n_{k}}^{\perp}=0$, and orthonormal, such that $v_{n_{k}}^{\perp\prime}v_{n_{k}}^{\perp}=I_{n_{k}-1}.$\footnote{The Gram-Schmidt process can be used to obtain $v_{n\perp}$ from $v_{n}$.} The canonical representation enables us to rotate $Z$ with $Q$ and define

equation[equation omitted — 134 chars of source]

where $Y_{0}$ is $K$-dimensional with ${\rm var}(Y_{0})=A$, and $Y_{k}$ is $n_{k}-1$ dimensional with ${\rm var}(Y_{k})=\lambda_{k}I_{n_{k}-1}$ for $k=1,\ldots,K$. The block-diagonal structure of $D$ implies that $Y_{0},Y_{1},\ldots$, and $Y_{K}$ are uncorrelated, which simplifies several expressions. For instance, we have the following identities:

equation[equation omitted — 186 chars of source]

such that the computation of the determinant and any power of $C$ is greatly simplified. The square-root of the $n\times n$ correlation matrix, $C^{1/2}$, is straight forward to compute. From the eigendecomposition of $A$, $A=P\Lambda_{a}P^{\prime}$, we define the block diagonal matrix: $D^{1/2}=\mathrm{diag}(P\Lambda_{a}^{1/2}P^{\prime},\lambda_{1}^{1/2}I_{n_{1}-1},\ldots,\lambda_{K}^{1/2}I_{n_{K}-1})$, and set $C^{1/2}\equiv QD^{1/2}Q^{\prime}$. It is easy to verify that $C=C^{1/2}C^{1/2}$ and that $C^{1/2}$ is symmetric. Computing $C^{1/2}$ therefore only requires an eigendecomposition of the symmetric and positive definite $K\times K$ matrix, $A$, rather than the eigendecomposition of $C$, which is $n\times n$. Computing other power of $C$ can be done similarly.

We can use ArchakovHansen:CanonicalBlockMatrix to recover the elements of the condensed log-correlation matrix, \[ \tilde{C}=\ensuremath{\ensuremath{\Lambda_{n}^{-1}W\Lambda_{n}^{-1}}},\quad W=\log A-\log\Lambda_{\lambda}, \] where \[ \Lambda_{\lambda}=\left[

array[array omitted — 96 chars of source]

\right],\qquadand\qquad\ensuremath{\Lambda_{n}=\left[

array[array omitted — 71 chars of source]

\right]}. \] The unique values in $\tilde{C}$, which are the elements in $\eta$, can be expressed as \[ \eta={\rm vech}(\tilde{C})=L_{K}\left(\Lambda_{n}^{-1}\otimes\Lambda_{n}^{-1}\right){\rm vec}(W), \] where $L_{K}$ is the elimination matrix, that solves $\mathrm{vech}(A)=L_{k}\mathrm{vec}(A)$. This parametrization of block correlation matrices does not impose additional superfluous restrictions, and the canonical representation facilitates simple computation of the determinant, the matrix inverse, and any other power, as well as the matrix logarithm and the matrix exponential. This is very useful for the evaluation of the likelihood function, especially for the more complicated models with heterogeneous heavy tails and complex dependencies, which we pursue in the next section.

Distributions

The next step is to specify a distribution for the $n$-dimensional random vector $Z$, from which the log-likelihood function, $\ell$, is defined. We consider several specifications, ranging from the multivariate normal distribution to convolutions of multivariate $t$-distributions. The convolution-$t$ distributions by HansenTong:2024 have simple log-likelihood functions and the canonical representation of a block correlation matrix motivates some particular specifications of the convolution-$t$ distribution.

We define \[ U=C^{-1/2}Z, \] such that $\mathrm{var}(U)=I_{n}$,\footnote{An advantage of having defined $C^{1/2}$ from the eigendecomposition, is that the normalized variables in $U$ are invariant to reordering of the elements in $Z$, which would not be the case if a Cholesky form was used to define $C^{1/2}$.} and a convenient property of any log-likelihood function, $\ell$, is that

equation[equation omitted — 91 chars of source]

This shows that the log-likelihood function will be in closed-form if we adopt a distribution for $U$ with a closed-form expression for $\ell(U)$, and this is important for obtaining tractable score-driven models. It is well known that the multivariate $t$-distribution and the Gaussian distribution have simple expression for $\ell(U)$. Fortunately, so does the multivariate convolution-$t$ distributions, which has different and interesting statistical properties for $Z$.

Multivariate $t$-Distributions

We begin with the simplest heavy-tailed distribution, a scaled multivariate $t$-distribution, which nests the Gaussian distribution as a limited case. The multivariate $t$-distribution is widely used to model vectors of returns with heavy tailed distributions, see e.g. CrealKoopmanLucasJBES:2012, OpschoorJanusLucasVanDick:2017, and HafnerWang:2023.

The $n$-dimensional multivariate $t$-distribution with $\nu$ degrees of freedom, location $\mu\in\mathbb{R}^{n}$, and scale matrix $\Sigma\in\mathbb{R}^{n\times n}$, typically written $X\sim t_{\nu}(\mu,\Sigma)$, has density \[ f_{X}(x)=\tfrac{\Gamma(\tfrac{\nu+n}{2})}{\Gamma(\tfrac{\nu}{2})}[\nu\pi]^{-\frac{n}{2}}|\Sigma|^{-\frac{1}{2}}\left[1+\tfrac{1}{\nu}(x-\mu)^{\prime}\Sigma^{-1}(x-\mu)\right]^{-\frac{\nu+n}{2}}. \] The variance is well-defined when $\nu>2$, in which case $\mathrm{var}(X)=\tfrac{\nu}{\nu-2}\Sigma$. The parameter $\nu$ governs the heaviness of the tail and the multivariate $t$-distribution converges to the multivariate normal distribution, $N(\mu,\Sigma)$, as $\nu\rightarrow\infty$.

To simplify the notation, we will use a scaled multivariate $t$-distribution, denoted $t_{\nu}^{\mathrm{std}}(0,\Sigma)$, which is defined for $\nu>2$. Its density is given by,

equation[equation omitted — 242 chars of source]

The relation between the two distributions is as follows: If $X\sim t_{\nu}(0,\Sigma)$ with $\nu>2$, then $Y=\sqrt{\tfrac{\nu-2}{\nu}}X\sim t_{\nu}^{\mathrm{std}}(0,\Sigma)$. The main advantage of the scaled $t$-distribution is that $\mathrm{var}(Y)=\Sigma$. Thus, if $U\sim t_{\nu}^{\mathrm{std}}(0,I_{n})$ then $Z=C^{1/2}U\sim t_{\nu}^{\mathrm{std}}(0,C)$, and the corresponding log-likelihood function is given by

equation[equation omitted — 154 chars of source]

where $c(\nu,n)=\log\left(\Gamma(\tfrac{\nu+n}{2})/\Gamma(\tfrac{\nu}{2})\right)-\frac{n}{2}\log\left[\left(\nu-2\right)\pi\right]$ is a normalizing constant that does not depend on the correlation matrix, $C$. If $C$ has a block structure we can use the identities in ((ref)), and obtain the following simplified expression,

equation[equation omitted — 331 chars of source]

The multivariate $t$-distribution has two implications for all elements of the vector $Z$. First, all elements of a multivariate $t$-distribution are dependent, because they share a common random mixing variable. Second, all elements of $U$ are identically distributed, because they are $t$-distributed with the same degrees of freedom. Both implications may be too restrictive in many applications, especially if the dimension, $n$, is large. Below we consider the convolution-$t$ distribution proposed in HansenTong:2024, which allows for heterogeneity and cluster structures in the tail properties and the tail dependencies.

Multivariate Convolution-$t$ Distributions

The multivariate convolution-$t$ distribution is a suitable rotations of a random vector that is made up of independent multivariate $t$-distributions. More specific, let $V_{1},\ldots,V_{G}$ be mutually independent standardized multivariate $t$-distributed variables, $V_{g}\sim t_{\nu_{g}}^{\mathrm{std}}(0,I_{m_{g}})$, with $\nu_{g}>2$ for all $g=1,\ldots,G$ and $n=\sum_{g=1}^{G}m_{g}$.

Then $V=(V_{1}^{\prime},\ldots,V_{G}^{\prime})^{\prime}\in\mathbb{R}^{n}$ has the standardized convolution-$t$ distribution (with zero location vector and identity scale-rotation matrix) that is denoted by \[ V\sim\mathrm{CT}_{\boldsymbol{m},\boldsymbol{\nu}}^{{\rm std}}(0,I_{n}), \] where $\boldsymbol{\nu}=(\nu_{1},\ldots,\nu_{G})^{\prime}$ is the vector with degrees of freedom and $\boldsymbol{m}=(m_{1},\ldots,m_{G})^{\prime}$ is the vector with the dimensions for the $G$ multivariate $t$-distributions. We can think of the partitioning of elements in $V$ as a second cluster structure, as we discuss below.

We will model the distribution of $U$ using $U=PV$, where $P\in\mathbb{R}^{n\times n}$ is an orthonormal matrix, i.e. $P^{\prime}P=I_{n}$, and we use the notation $U\sim\mathrm{CT}_{\boldsymbol{m},\boldsymbol{\nu}}^{{\rm std}}(0,P)$. While $=\mathrm{var}(U)=\mathrm{var}(V)=I_{n}$, they will not have the same distribution, unless $P$ has a particular structure, such as $P=I_{n}$. Similarly, we use the following notation for the distribution of \[ Z=C^{1/2}PV\sim\mathrm{CT}_{\boldsymbol{m},\boldsymbol{\nu}}^{{\rm std}}(0,C^{1/2}P), \] which is a convolution-$t$ distribution with location zero and scale-rotation matrix $C^{1/2}P$. Note that we have $\mathrm{var}(Z)=C$, for any orthonormal matrix, $P$, but different choices for $P$ lead to different distributions with distinct non-linear dependencies that arise from the cluster structure in $V$.

Conveniently, we have the expression, $V=P^{\prime}C^{-1/2}Z=P^{\prime}U$, and if we partition the columns in $P$, using the same cluster structure as in $V$, i.e. $P=(P_{1},\ldots,P_{G})$ with $P_{g}\in\mathbb{R}^{n\times m_{g}}$, then it follows that $V_{g}=P_{g}^{\prime}U\in\mathbb{R}^{m_{g}}$, for $g=1,\ldots,G$. Next, $U$ and $V$ have the exact same log-likelihoods, $\ell(U)=\ell(P^{\prime}U)=\ell(V)$, and we can use ((ref)) to express the log-likelihood function for $Z$ as

align[align omitted — 174 chars of source]

where $c_{g}=c(\nu_{g},m_{g})$. When $C$ has a block structure, then we also have \[ V=P^{\prime}Q\left[

array[array omitted — 90 chars of source]

\right], \] where $Y=Q^{\prime}Z$, and some interesting special cases emerge from this structure.

We have previously used a partitioning of the variables to form a block correlation structure, which arises from a cluster structure for the variables. The convolution-$t$ distribution involves a second partitioning that defines the $G$ independent multivariate $t$-distributions. This is a cluster structure in the underlying random innovations in the model. The two cluster structures can be identical, or can be different, as we illustrate with examples and in our empirical application. Next, we highlight six distributional properties that are the product of this model design.

enumerate• Each element of $V_{g}\in\mathbb{R}^{m_{g}}$, has the same marginal $t$-distribution with $\nu_{g}$ degrees of freedom. This does not carry over to the same elements of $Z$ (even if $P=I$). In general, the marginal distribution of an element of $Z$, will be a unique convolution of (as many as) $G$ independent $t$-distributions with different degrees of freedom. • While the (multivariate) $t$-distributions are independent across groups, this does not carry over to the corresponding sub-vectors of $Z$. • The convolution for each element of $Z$ is, in part, defined by the correlation matrix, $C$. So, time-variation in $C$ will induce time-varying marginal distributions for the elements of $Z$. • The partitioning of $V=P^{\prime}U$ into $G$ clusters ($G$-clusters) induces heterogeneity in tail dependencies and the heavyness of the tails. The $G$-clusters can be entirely different from the $K$-clusters (partitioning of $Z$ variables) that define the block structure in the correlation matrix, and the two numbers of clusters can be different. • Increasing the number of $G$-clusters, does not necessarily improve the empirical fit. While increasing $G$ will increase the number parameters (degrees of freedom) in the model, it also entails dividing $V$ into additional subvectors, which eliminates the innate dependence between elements of $V$, which apply to elements from the same multivariate $t$-distribution. • Sixth, this model framework nests the conventional multivariate $t$-distribution as the special case, $G=1$, which facilitates simple comparisons with a natural benchmark model.

Density and CDF of Convolution-$t$ Distribution

The marginal distributions of the elements of $Z$ are convolutions of independent $t$-distributed variables, and neither their densities nor their cumulative distribution function have simple expressions.\footnote{Even for the simplest case -- a convolution of two univariate $t$-distributions -- the resulting density does not have a simple closed-form expression.} However, using HansenTong:2024 we obtain the following semi-analytical expressions, where ${\rm Re}\left[x\right]$ and ${\rm Im}\left[x\right]$ denote the real and imaginary part of $x\in\mathbb{C}$, respectively, and $e_{j,n}$ is the $j$-th column of identity matrix $I_{n}$.

propSuppose $Z\sim\mathrm{CT}_{\boldsymbol{m},\boldsymbol{\nu}}^{{\rm std}}(0,C^{1/2}P)$. Then the marginal density and cumulative distribution function for $Z_{j}$, $j=1,\ldots,n$, are given by \begin{align*} f_{Z_{j}}(z) & =\frac{1}{\pi}\int_{0}^{\infty}{\rm Re}\left[e^{-isz}\varphi_{Z_{j}}(s)\right]\mathrm{d}s,\quad\quad F_{Z_{j}}(z)=\ensuremath{\frac{1}{2}-\frac{1}{\pi}\int_{0}^{\infty}\frac{{\rm Im}\left[e^{-isz}\varphi_{Z_{j}}(s)\right]}{s}\mathrm{d}s}, \end{align*} respectively, where $\varphi_{Z_{j}}(s)=\prod_{g=1}^{G}\varphi_{\nu_{g}}^{\mathrm{std}}\left(is\|P_{g}^{\prime}C^{\frac{1}{2}}e_{j,n}\|\right)$ is the characteristic function for $Z_{j}$, and \[ \varphi_{\nu}^{\mathrm{std}}(s)=\frac{K_{\frac{\nu}{2}}(\sqrt{\nu-2}|s|)(\sqrt{\nu-2}|s|)^{\frac{1}{2}\nu}}{\Gamma\left(\frac{\nu}{2}\right)2^{\frac{\nu}{2}-1}}, \] is the characteristic function of the univariate $t_{\nu}^{\mathrm{std}}$-distribution.

To gain some insight about convolution-$t$ distributions and the expressions in Proposition (ref) we present features of two densities in Figure (ref). We specifically consider convolutions, $\tfrac{1}{\sqrt{G}}\sum_{g=1}^{G}V_{g}$, for $G=2$ and $G=10$, where $V_{1},\ldots,V_{g}$ are independent and standardized $t$-distributed with six degrees of freedom.

figure[figure omitted — 594 chars of source]

The upper panel of Figure (ref), Panel (a), shows the log-densities of the (left) tail of the distribution, and how they compare to those of a standardized $t_{(6)}^{{\rm std}}$-distribution and a standard normal distribution. As $G\rightarrow\infty$ the convolution-$t$ distribution will approach the normal distribution. So, it is not surprisingly that the log-densities for the convolutions are between that of a $t_{(6)}^{{\rm std}}$ and that of a standard normal. Unsurprisingly, the convolution of $G=10$ standardized $t$-distributions is closer to the normal distribution than the convolution of $G=2$ distributions. However, the convolution-$t$ distribution is not a $t$-distribution for $G>1$. In terms of Kullback-Leibler discrepancy, the best approximating $t$-distribution to the convolution-$t$ distribution is a $t_{(8.75)}^{{\rm std}}$-distribution when $G=2$ and a $t_{(26.15)}^{{\rm std}}$-distribution when $G=10$, see the Q-Q plots in Panels (b) and (c) in Figure (ref).

The expression for the marginal density of convolution-$t$ distributions is particularly useful in our empirical analysis, because it gives us a factorization of the joint density into marginal densities and the copula density by Sklar's theorem. This leads to the decomposition of the log-likelihood, $\ell(Z)=\sum_{j=1}^{n}\ell(Z_{j})+\log(c(Z))$, where $c(Z)$ denotes the copula density, and we can see if gains in the log-likelihood are primarily driven by gains in the marginal distributions or by gains in the copula density.

Three Special Types of Convolution-$t$ Distributions

The convolution-$t$ distributions define a broad class of distributions, with many possible partitions of $V$ and choices for $P$. Below we elaborate on som particular details of three special types of convolution-$t$ distributions. For latter use, we use $e_{k}\in\mathbb{R}^{K\times1}$ to denote the $k$-th column of identity matrix $I_{K}$.

Special Type 1: Cluster-$t$ Distribution

The first special type of convolution-$t$ distribution has $P=I$, such that $U=V$, and a single cluster structure. The cluster structure, $\boldsymbol{m}$, is imposed on $V$, whereas $C$ can be unrestricted, or have block structure based on the the same clustering, in which case $\bm{n}=\bm{m}$ and $G=K$.

Without a block correlation structure on $C$, we have $V=C^{-1/2}Z$ and the log-likelihood function is simply computed using ((ref)). If the block structure is imposed on $C$, then we can express the multivariate $t$-distributed variables as linear combinations on the canonical variables, $Y_{0},\ldots,Y_{K}$,

equation[equation omitted — 180 chars of source]

We therefore have the expression for the quadratic terms, \[ U_{k}^{\prime}U_{k}=Y_{0}^{\prime}A^{-\tfrac{1}{2}}e_{k}e_{k}^{\prime}A^{-\tfrac{1}{2}}Y_{0}+\lambda_{k}^{-1}Y_{k}^{\prime}Y_{k},\quad k=1,\ldots,K, \] and the log-likelihood function simplifies to

align[align omitted — 217 chars of source]

where $c_{k}=c(\nu_{k},n_{k})$. The block structure simplifies implementation of the score-driven model for this specification, and makes it possible to implement the model with a large number of stocks.

Special Type 2: Hetero-$t$ Distribution

A second special type of convolution-$t$ distributions has $P=I$ and $G=n$. So, the elements of $U$ are made up of $n$ independent univariate $t$-distributions with degrees of freedom, $\nu_{i}$, $i=1,\cdots,n$. This distribution can accommodate a high degree of heterogeneity in the tail properties of $Z_{i}$, $i=1,\ldots,n$, which are different convolutions of the $n$ independent $t$-distributions. For this reason, we refer to these distributions as the Hetero-$t$ distributions. The number of degrees of freedom increases from $G$ to $n$, but the additional parameters do not guarantee a better in-sample log-likelihood, because all dependence between elements of $V$ is eliminated. The Cluster-$t$ distribution has dependence between $V$-variables within the same cluster. This has implications the linear combinations of $U$, including those that define $Z$.

For the case with a general correlation matrix, the Hetero-$t$ distribution simplifies the log-likelihood function in ((ref)) to \[ \ell(Z)=-\tfrac{1}{2}\log|C|+\sum_{i=1}^{n}c_{i}-\tfrac{\nu_{i}+1}{2}\log\left(1+\tfrac{1}{\nu_{i}-2}U_{i}^{2}\right), \] where $c_{i}=c(\nu_{i},1)$.\footnote{Note that we can obtain preliminary estimates (starting values) of the $n$ degrees of freedom parameters, by estimating $\nu_{i}$ from $e_{i}^{\prime}\tilde{U}_{t}$, where $\tilde{U}_{t}=\tilde{C}^{-\frac{1}{2}}Z_{t}$ and $\tilde{C}$ is an estimate of the unconditional correlation matrix, for $i=1,\ldots,n$. }

We can combine the heterogenous $t$-distributions with a block correlation matrix, in which case the log-likelihood function simplifies to

equation[equation omitted — 250 chars of source]

where $c=\sum_{i=1}^{n}c(\nu_{i},1)$ and $U_{k,j}$ is the $j$-th element of the vector $U_{k}$ expressed by ((ref)).

Special Type 3: Canonical-Block-$t$ Distribution

A third special type of convolution-$t$ distributions is based on the canonical canonical variables, as defined by the canonical representation of the block correlation matrix. The Canonical-Block-$t$ distribution has $P=Q$ and $\boldsymbol{m}=(K,n_{1}-1,\ldots,n_{K}-1)^{\prime}$, such that $V=Q^{\prime}U$ is composed of $G=K+1$ independent multivariate $t$-distributions. So, \[ Q^{\prime}U=\left(V_{0}^{\prime},V_{1}^{\prime},\cdots,V_{K}^{\prime}\right)^{\prime},\quad{\rm where}\ V_{0}\sim t_{\nu_{0}}(0,I_{K}),\quad{\rm and}\text{ }V_{k}\sim t_{\nu_{k}}(0,I_{n_{k}-1}). \] This construction is motivated by the $K+1$ canonical variables, $Y_{0},\ldots,Y_{K}$, that arises from the canonical representation of block correlation matrices. Interestingly, this type of convolution-$t$ distribution can be used, regardless of $C$ having a block structure or not. For a general correlation matrix, $C$, the log-likelihood function is given by

align*[align* omitted — 232 chars of source]

where $V=Q^{\prime}U=Q^{\prime}C^{-1/2}Z$.

From a practical viewpoint, a more interesting situation is when $C$ has a block structure, such that $C=QDQ^{\prime}$. With this structure, the log-likelihood function simplifies to

align[align omitted — 358 chars of source]

which is computationally advantageous, because it does note require an inverse (nor a determinant) of an $n\times n$ matrix.

The expression for the log-likelihood function shows that this distribution is equivalent to assuming that $Y_{0},Y_{1},\ldots,Y_{K}$ are independent and distributed as $\ensuremath{Y_{0}\sim t_{\nu_{0}}^{\mathrm{std}}(0,A)}$ and $\ensuremath{Y_{k}\sim t_{\nu_{k}}^{\mathrm{std}}(0,\lambda_{k}I_{n_{k}-1})}$, for $k=1,\cdots,K$. This yields insight about the standardized returns within each block. Let $Z_{k}$ be the $n_{k}$-dimensional subvector of $Z=(Z_{1}^{\prime},\ldots,Z_{K}^{\prime})^{\prime}$. From $Z=QQ^{\prime}Z=QY$ it follows that \[ Z_{k}=v_{n_{k}}Y_{0,k}+v_{n_{k}}^{\perp}Y_{k}, \] such that a standardized return in the $k$-th block has the same loading on the common variable $Y_{0,k}$, and orthogonal loadings on the vector $Y_{k}$.

Additional convolution-$t$ distributions could be bases on this structure. For instance, we could combine $P=Q$ with heterogeneous univariate $t$-distributions, for some or all of the canonical variables. For instance, the canonical variable, $V_{0}$, could be made up of $K$heterogeneous $t$-distributions, while other canonical variables, $V_{1},\ldots,V_{K}$ have multivariate $t_{\nu_{k}}^{\mathrm{std}}$-distributions.

Score-Driven Models

We turn to the dynamic modeling of the conditional correlation matrix in this section. To this end we adopt the score-drive framework by CrealKoopmanLucas:2013, to model the dynamic properties of $\gamma_{t}={\rm vecl}(\log C_{t})\in\mathbb{R}^{d}$, with $d=n\left(n-1\right)/2$. Specifically, we adopt the vector autoregressive model of order one, VAR(1):

equation[equation omitted — 99 chars of source]

where $\beta$ and $\alpha$ are $d\times d$ matrices of coefficients, $\mu=\mathbb{E}(\gamma_{t})$, and $\varepsilon_{t}$ will be defined by the first-order conditions of the log-likelihood at times $t$.\footnote{It is straightforward to include additional lagged values of $\eta_{t}$, such that ((ref)) has a higher-order VAR(p) structure, and adding $q$ lagged values of $\varepsilon_{t}$, would generalize ((ref)) to a VARMA(p,q) model, we do not pursue these extensions in this paper.} The key aspect of a score-driven model is that the score of the predictive likelihood function is used to define the innovation $\varepsilon_{t}$, specifically

equation[equation omitted — 167 chars of source]

and $\mathcal{S}_{t}$ is a scaling matrix. The score $\nabla_{t}$ is the first-order derivative of log-likelihood with respect to $\gamma_{t}$, and $\nabla_{t}$ is a martingale difference process if the model is correctly specified. The Fisher information matrix, $\mathcal{I}_{t}=\mathbb{E}_{t-1}\left(\nabla_{t}\nabla_{t}^{\prime}\right)$, is often used as the scaling matrix, in which case the time-varying parameter vector is updated in a manner that resembles a Newton-Raphson algorithm, see CrealKoopmanLucas:2013.\footnote{One exception is HafnerWang:2023, who used an unscaled score, i.e. $\mathcal{S}_{t}=I$, which does not take any curvature of the log-likelihood into account when parameter values are revised.}

A potential drawback of using $\mathcal{S}_{t}^{-1}$ as the scaling matrix in ((ref)) is that the precision of the inverse deteriorates as the dimension increases. We will therefore approximate $\mathcal{S}_{t}^{-1}$ by imposing a diagonal structure, and simply inverting the diagonal elements of $\mathcal{I}_{t}$. This is equivalent to using the scaling matrix, \[ \mathcal{S}_{t}={\rm diag}\left(\mathcal{I}_{t,11},\ldots,\mathcal{I}_{t,dd}\right). \] In this manner, each element of the parameter vector is updated with a scaled version of the corresponding element of the score. Computing the inverse, $\mathcal{S}_{t}^{-1}$, is now straightforward and simple to implement.

The score is computed using the following decomposition,

equation[equation omitted — 215 chars of source]

The expression for the last term was derived in ArchakovHansen:Correlation using results from LintonMcCrorie:1995. The drawback of this approach is that it requires an eigendecomposition of $n^{2}\times n^{2}$ matrix and this is impractical and unstable when $n$ is large. Moreover, the computational burden for the corresponding information matrix is even worse. Fortunately, when $C$ has a block structure, we have the following simplified expression, \[ \frac{\partial\ell}{\partial\eta^{\prime}}=\frac{\partial\ell}{\partial{\rm vec}\left(A\right)^{\prime}}\frac{\partial{\rm vec}\left(A\right)}{\partial{\rm vec}\left(W\right)^{\prime}}\left(\Lambda_{n}\otimes\Lambda_{n}\right)D_{K}. \] The first term can be computed very fast for all the variants of the convolution-$t$ distributions we consider. The second term only requires an eigendecomposition of $A$ (the upper-left $K\times K$ submatrix of $D$), and this greatly reduces the computational burden for evaluating both the score and the information matrix.

For block correlation matrices, we use the vector autoregression of order one for the subvector,

equation[equation omitted — 108 chars of source]

where $\mu=\mathbb{E}(\text{\ensuremath{\eta}}_{t})\in\mathbb{R}^{d}$, and $\alpha$ and $\beta$ are $d\times d$ matrices with $d=K\left(K+1\right)/2$.

To implement the score-driven model we need to derive the appropriate score and scaling matrix for each of the log-likelihoods. For this purpose, we will adopt the following notation involving matrices and matrix operators, with some notation adopted from CrealKoopmanLucasJBES:2012. Let $A$ and $B$ be two matrices with suitable dimensions. The Kronecker product is denoted by $A\otimes B$ and we use $A_{\otimes}\equiv A\otimes A$ and $A\oplus B\equiv A\otimes B+B\otimes A$. We let $K_{k}$ denote the commutation matrix, $D_{k}$ the duplication matrix, and $L_{k}$, $E_{l}$, $E_{u}$, are $E_{d}$ elimination matrices. These are defined by the following identities: \[

array[array omitted — 299 chars of source]

\] for any symmetric matrix, $A\in\mathbb{R}^{k\times k}$, and any matrix, $B\in\mathbb{R}^{k\times k}$.

Scores and Information Matrices for a General Correlation Matrix

We first derive expressions for $\nabla$ and $\mathcal{I}$ with a general correlation matrix. Recall that the log-likelihood function, based on a convolution-$t$ distribution, is given by ((ref)), and in the special case with a multivariate $t$-distribution, the log-likelihood simplifies to the expression in ((ref)).

Score-Driven Model with Multivariate $t$-Distribution

thmSuppose that $Z\sim t_{n,\nu}^{\mathrm{std}}(0,C)$. Then the score vector and information matrix with respect to $\gamma={\rm vecl}\left(\log C\right)$, are given by: \begin{align} \nabla & =\tfrac{1}{2}M^{\prime}C_{\otimes}^{-1}\left[W{\rm vec}\left(ZZ^{\prime}\right)-{\rm vec}\left(C\right)\right],\\ \mathcal{I} & =\tfrac{1}{4}M^{\prime}\left[\phi C_{\otimes}^{-1}H_{n}+(\phi-1){\rm vec}(C^{-1}){\rm vec}(C^{-1})^{\prime}\right]M, \end{align} respectively, with $H_{n}=I_{n^{2}}+K_{n}$, \[ W=\frac{\nu+n}{\nu-2+Z^{\prime}C^{-1}Z},\quad\phi=\frac{\nu+n}{\nu+n+2}, \] and \[ M=\partial{\rm vec}\left(C\right)/\partial\gamma^{\prime}=\left(E_{l}+E_{u}\right)^{\prime}E_{l}\left(I_{n^{2}}-\Gamma E_{d}^{\prime}\left(E_{d}\Gamma E_{d}^{\prime}\right)^{-1}E_{d}\right)\Gamma\left(E_{l}+E_{u}\right)^{\prime}, \] where the expression for $\Gamma=\ensuremath{\partial{\rm vec}(C)/\partial{\rm vec}\left(\log C\right)^{\prime}}$ is presented in the appendix, see ((ref)).

The expression of $W$ shows that the impact of extreme values (outliers) is dampened by the degrees of freedom, however this mitigation subsides as $\nu\rightarrow\infty$. The result for the Gaussian distribution is obtained by setting $W=\phi=1$, which are their limits as $\nu\rightarrow\infty$.

Score-Driven Model with Convolution-$t$ Distributions

thmSuppose that $Z\sim\mathrm{CT}_{\boldsymbol{m},\boldsymbol{\nu}}^{{\rm std}}(0,C^{1/2}P)$. Then the score vector and information matrix with respect to $\gamma={\rm vecl}\left(\log C\right)$, are given by: \begin{align*} \nabla & =M^{\prime}\Omega\left[\sum_{g=1}^{G}W_{g}{\rm vec}\left(P_{g}V_{g}U^{\prime}\right)-{\rm vec}\left(I_{n}\right)\right],\\ \mathcal{I} & =M^{\prime}\Omega\left(K_{n}+\Upsilon_{G}\right)\Omega M, \end{align*} respectively, where $M$ is defined in Theorem (ref), $\Omega=\ensuremath{(I_{n}\otimes C^{-\frac{1}{2}})(C^{\frac{1}{2}}\oplus I_{n})^{-1}}$, and $\Upsilon_{G}=\sum_{g=1}^{G}\Psi_{g}$ with \begin{align*} \Psi_{g} & =\psi_{g}\left(I_{n}\otimes J_{g}\right)+\ensuremath{\left(\phi_{g}-\psi_{g}\right)J_{g\otimes}+\left(\phi_{g}-1\right)\left[J_{g\otimes}K_{n}+{\rm vec}\left(J_{g}\right){\rm vec}\left(J_{g}\right)^{\prime}\right],} \end{align*} where $J_{g}=P_{g}P_{g}^{\prime}$, \[ W_{g}=\frac{\nu_{g}+m_{g}}{\nu_{g}-2+V_{g}^{\prime}V_{g}},\quad\phi_{g}=\frac{\nu_{g}+m_{g}}{\nu_{g}+m_{g}+2},\quad\psi_{g}=\phi_{g}\frac{\nu_{g}}{\nu_{g}-2}, \] for $g=1,\ldots,G$.

The inverse of $C^{\frac{1}{2}}\oplus I_{n}$ (an $n^{2}\times n^{2}$ matrix) is available in closed form (see Appendix A) and is computationally inexpensive because it relies on an eigendecomposition of $C$, which is already needed for computing $\Gamma$ in the expression of $M$.

Some insight can be gained from considering the case $P=I$. A key component of $\nabla$ is $\sum_{g=1}^{G}\left(W_{g}P_{g}V_{g}\right)=\left(W_{1}V_{1}^{\prime},W_{2}V_{2}^{\prime},\ldots,W_{G}V_{G}^{\prime}\right)^{\prime}$, which shows that the impact that $g$-th cluster, $V_{g}$, has one the score is controlled by the coefficient $W_{g}$.

Scores and Information Matrices for a Block Correlation Matrix

Next, we derive the corresponding expression for the case where $C$ has a block structure. For the score we have the following expression \[ \nabla^{\prime}=\frac{\partial\ell}{\partial\eta^{\prime}}=\nabla_{A}^{\prime}\Pi_{A},\quad{\rm where}\quad\nabla_{A}=\frac{\partial\ell}{\partial{\rm vec}(A)}, \] and the expression for $\Pi_{A}$ is given in the following Lemma.

lemLet $\Pi_{A}=\partial{\rm vec}(A)/\partial\eta^{\prime}$, then \begin{equation} \Pi_{A}=\left[\Gamma_{A}-\Gamma_{A}E_{d}^{\prime}\left(\Phi+E_{d}\Gamma_{A}E_{d}^{\prime}\right)^{-1}E_{d}\Gamma_{A}\right]\Lambda_{n\otimes}D_{k}, \end{equation} where $\Phi$ is a $K\times K$ diagonal matrix with $\Phi_{kk}=\lambda_{k}\left(n_{k}-1\right)$, $k=1,\ldots,K$, and $\Gamma_{A}=\ensuremath{\partial{\rm vec}(A)/\partial{\rm vec}\left(\log A\right)^{\prime}}$ has the expression given in ((ref)).

Conveniently, the computation of $\Pi_{A}$ only requires the inverse of a $K\times K$ matrix. From the results for $\nabla_{A}$ we have $\nabla=\Pi_{A}^{\prime}\nabla_{A}$ and similarly, \[ \mathcal{I}=\Pi_{A}^{\prime}\mathcal{I}_{A}\Pi_{A},\quad{\rm where}\quad\mathcal{I}_{A}=\mathbb{E}\left(\nabla_{A}\nabla_{A}^{\prime}\right). \]

Score-Driven Model with Block Correlation and Multivariate $t$-Distribution

With a block correlation structure, we define the standardized canonical variables \[ X=\ensuremath{\left(X_{0}^{\prime},X_{1}^{\prime},\ldots,X_{K}^{\prime}\right)^{\prime}=Q^{\prime}U=D^{-\frac{1}{2}}Y}, \] such that $X_{0}=A^{-\frac{1}{2}}Y_{0}$ with ${\rm var}(X_{0})=I_{K}$ and $X_{k}=\lambda_{k}^{-\frac{1}{2}}Y_{k}$ with ${\rm var}(X_{k})=I_{n_{k}-1}$ for $k=1,\ldots,K$.

thmSuppose that $Z\sim t_{\nu,n}^{{\rm std}}(0,C)$. Then the score vector and information matrix with respect to the dynamic parameters, ${\rm vec}(A)$, are given by: \begin{align*} \nabla_{A} & =\tfrac{1}{2}\ensuremath{A_{\otimes}^{-\frac{1}{2}}}\left[W{\rm vec}\left(X_{0}X_{0}^{\prime}\right)-{\rm vec}\left(I_{K}\right)\right]+\tfrac{1}{2}E_{d}^{\prime}S,\\ \mathcal{I}_{A} & =\tfrac{1}{4}\left[\ensuremath{\ensuremath{\phi A_{\otimes}^{-1}H_{K}+(\phi-1){\rm vec}(A^{-1}){\rm vec}(A^{-1})^{\prime}}}\right]+\tfrac{\phi}{2}E_{d}^{\prime}\Xi E_{d}\\ & \quad+\tfrac{1-\phi}{4}\left[{\rm vec}(A^{-1})\xi^{\prime}E_{d}+E_{d}^{\prime}\xi{\rm vec}(A^{-1})^{\prime}-E_{d}^{\prime}\xi\xi^{\prime}E_{d}\right], \end{align*} respectively, where \[ \phi=\frac{\nu+n}{\nu+n+2},\quad W=\frac{\nu+n}{\nu-2+X_{0}^{\prime}X_{0}+\sum_{k=1}^{K}X_{k}^{\prime}X_{k}}, \] and $S\in\mathbb{R}^{K}$, $\xi\in\mathbb{R}^{K}$, and the diagonal matrix, $\Xi$, are defined by \[ S_{k}=\frac{1}{\lambda_{k}}-\frac{WX_{k}^{\prime}X_{k}}{\lambda_{k}\left(n_{k}-1\right)},\quad\xi_{k}=\lambda_{k}^{-1},\quad\ensuremath{\Xi_{kk}=\lambda_{k}^{-2}\left(n_{k}-1\right)^{-1},} \] for $k=1,\ldots,K$. In the special case where $Z$ has a multivariate Gaussian distribution ($\nu=\infty$, $\phi=1$), the expression for the information matrix simplifies to $\mathcal{I}_{A}=\tfrac{1}{4}\ensuremath{A_{\otimes}^{-1}}H_{K}+\tfrac{1}{2}E_{d}^{\prime}\Xi E_{d}$.

Score-Driven Model with Block Correlation and Cluster-$t$ Distribution

thm[Cluster-$t$ with Block-$C$] Suppose that $Z\sim\mathrm{CT}_{\boldsymbol{n},\boldsymbol{\nu}}^{{\rm std}}(0,C^{1/2})$ where $C$ has the block structure defined by $\boldsymbol{n}$. Then the score vector and information matrix with respect to dynamic parameters, ${\rm vec}(A)$, are given by: \begin{align*} \nabla_{A} & =\Omega_{A}\left[\sum_{k=1}^{K}W_{k}X_{0,k}{\rm vec}\left(e_{k}X_{0}^{\prime}\right)-{\rm vec}\left(I_{K}\right)\right]+\tfrac{1}{2}E_{d}^{\prime}S,\\ \mathcal{I}_{A} & =\Omega_{A}\left(K_{K}+\Upsilon_{K}\right)\Omega_{A}+\ensuremath{\tfrac{1}{4}E_{d}^{\prime}\Xi E_{d}}+\ensuremath{\tfrac{1}{2}E_{d}^{\prime}\Theta\Omega_{A}}+\ensuremath{\tfrac{1}{2}\Omega_{A}\Theta^{\prime}E_{d}}, \end{align*} respectively, where $\Omega_{A}=\ensuremath{(I_{K}\otimes A^{-\frac{1}{2}})(A^{\frac{1}{2}}\oplus I_{K})^{-1}}$, and vector $e_{k}$ is the $k$-th column of the identity matrix $I_{K}$. The vector $S\in\mathbb{R}^{K}$, the diagonal matrix, $\Xi$, and $\Theta$ are defined as \begin{align*} S_{k} & =\frac{1}{\lambda_{k}}-\frac{W_{k}X_{k}^{\prime}X_{k}}{\lambda_{k}\left(n_{k}-1\right)},\quad\quad\ \ensuremath{W_{k}=\frac{\nu_{k}+n_{k}}{\nu_{k}-2+\ensuremath{X_{0,k}^{2}+X_{k}^{\prime}X_{k}}},}\\ \Xi_{kk} & =\ensuremath{\frac{\phi_{k}-1}{\lambda_{k}^{2}}+\frac{2\phi_{k}}{\lambda_{k}^{2}\left(n_{k}-1\right)}},\quad\Theta=\sum_{k=1}^{K}\lambda_{k}^{-1}\left(1-\phi_{k}\right)e_{k}{\rm vec}\left(J_{k}^{e}\right)^{\prime}, \end{align*} for $k=1,\ldots,K$. The matrix $\Upsilon_{K}$ is defined analogously to $\Upsilon_{G}$ in Theorem (ref).

Score-Driven Model with Block Correlation and Hetero-$t$ Distribution

thm[Heterogeneous-Block Convolution-$t$] Suppose that $Z\sim\mathrm{CT}_{\boldsymbol{n},\boldsymbol{\nu}}^{{\rm std}}(0,C^{1/2})$, where $C$ has the block structure defined by $\boldsymbol{n}$. Then the score vector and information matrix with respect to the dynamic parameters, ${\rm vec}(A)$, are given by: \begin{align*} \nabla_{A} & =\Omega_{A}\left[\sum_{k=1}^{K}\sum_{i=1}^{n_{k}}W_{k,i}U_{k,i}{\rm vec}\left(e_{k}X_{0}^{\prime}\right)n_{k}^{-\frac{1}{2}}-{\rm vec}(I_{K})\right]+\tfrac{1}{2}E_{d}^{\prime}S,\\ \mathcal{I}_{A} & =\ensuremath{\Omega_{A}\left(K_{K}+\Upsilon_{K}^{e}\right)\Omega_{A}+\ensuremath{\tfrac{1}{4}E_{d}^{\prime}\Xi E_{d}}+\ensuremath{\tfrac{1}{2}E_{d}^{\prime}\Theta\Omega_{A}}+\ensuremath{\tfrac{1}{2}\Omega_{A}\Theta^{\prime}E_{d}}}, \end{align*} respectively, where \[ \ensuremath{S_{k}=\frac{1}{\lambda_{k}}-\frac{\sum_{i=1}^{n_{k}}W_{k,i}U_{k,i}F_{k,i}U_{k}}{\left(n_{k}-1\right)\lambda_{k}}},\qquad k=1,\ldots,K, \] with \[ W_{k,i}=\frac{\nu_{k,i}+1}{\nu_{k,i}-2+U_{k,i}^{2}},\qquad F_{k,i}=\ensuremath{\tilde{e}_{i}^{\prime}\left(I_{n_{k}}-v_{n_{k}}v_{n_{k}}^{\prime}\right)}, \] and $\tilde{e}_{i}$ is the $i$-th column of identity matrix $I_{n_{k}}$. The matrix $\Upsilon_{K}^{e}=\sum_{k=1}^{K}\Psi_{k}^{e}$ is given by: \[ \Psi_{k}^{e}=\ensuremath{n_{k}^{-1}\left(3\bar{\phi}_{k}-2-\bar{\psi}_{k}\right)J_{k\otimes}^{e}+\bar{\psi}_{k}\left(I_{K}\otimes J_{k}^{e}\right)}, \] where $J_{k}^{e}=e_{k}e_{k}^{\prime}$, and \[ \ensuremath{\bar{\phi}_{k}=\frac{1}{n_{k}}\sum_{i=1}^{n_{k}}\phi_{k,i}},\quad\bar{\psi}_{k}=\frac{1}{n_{k}}\sum_{i=1}^{n_{k}}\psi_{k,i}. \] The diagonal matrix $\Xi$ and $\Theta$ are given by: \begin{align*} \Xi_{kk} & =\ensuremath{\lambda_{k}^{-2}n_{k}^{-1}\left[3\bar{\phi}_{k}-1+\left(\bar{\psi}_{k}+1\right)\left(n_{k}-1\right)^{-1}\right]},\\ \Theta & =\sum_{k=1}^{K}\ensuremath{\lambda_{k}^{-1}n_{k}^{-1}\left(\bar{\psi}_{k}+2-3\bar{\phi}_{k}\right)}e_{k}{\rm vec}\left(J_{k}^{e}\right)^{\prime}. \end{align*}

Score-Driven Model with Block Correlation and Canonical-Block-$t$ Distribution

thm[Canonical-Block Convolution-$t$] Suppose that $Z\sim\mathrm{CT}_{\boldsymbol{m},\boldsymbol{\nu}}^{{\rm std}}(0,C^{1/2}Q)$, where $C$ has the block structure defined by $\boldsymbol{n}$ and $\boldsymbol{m}=(K,n_{1}-1,\ldots,n_{K}-1)^{\prime}$. Then the score vector and information matrix with respect to the dynamic parameters, ${\rm vec}(A)$, are given by: \begin{align*} \nabla_{A} & =\tfrac{1}{2}A_{\otimes}^{-\frac{1}{2}}\left[W_{0}{\rm vec}(X_{0}X_{0}^{\prime})-{\rm vec}(I_{K})\right]+\tfrac{1}{2}E_{d}^{\prime}S,\\ \mathcal{I}_{A} & =\tfrac{1}{4}\left[\ensuremath{\ensuremath{\phi_{0}A_{\otimes}^{-1}H_{K}+(\phi_{0}-1){\rm vec}(A^{-1}){\rm vec}(A^{-1})^{\prime}}}+E_{d}^{\prime}\Xi E_{d}\right], \end{align*} where the expressions for $S$ and diagonal matrix, $\Xi$, are those given in Theorem (ref) with \begin{align*} W_{0} & =\frac{\nu_{0}+K}{\nu_{0}-2+X_{0}^{\prime}X_{0}},\quad W_{k}=\frac{\nu_{k}+n_{k}-1}{\nu_{k}-2+X_{k}^{\prime}X_{k}},\\ \phi_{0} & =\frac{\nu_{0}+K}{\nu_{0}+K+2},\quad\quad\phi_{k}=\frac{\nu_{k}+n_{k}-1}{\nu_{k}+n_{k}+1}, \end{align*} for $k=1,\ldots,K$.

Some details about practical implementation

Obtaining the $A$-matrix from the vector $\eta$

The $K\times K$ matrix, $A=\mathrm{var}(Y_{0})$, plays a central role in the score models with block-correlation matrices. Below we show how $A_{t}$ can be computed from $\eta_{t}$.

In order to obtain $A$ from $\eta$, we adopt the algorithm developed in ArchakovHansenLuo-RandomCorr:2024 to generate random block correlation matrices. The algorithm has three steps.

enumerate• Compute the elements of the $K\times K$ matrix, $\tilde{A}$, using \[ \tilde{A}_{k,l}=\begin{cases} \tilde{c}_{kk}\left(n_{k}-1\right) & \text{ for }k=l,\\ \tilde{c}_{kl}\sqrt{n_{k}n_{l}} & \text{ for }k\neq l, \end{cases} \] where $\tilde{c}_{kl}$ are elements of $\eta$, as defined by the identity, $\eta={\rm vech}(\tilde{C})$. • From an arbitrary starting value, $y^{(0)}\in\mathbb{R}^{K}$, e.g. a vector of zeroes, evaluate the recursion, \[ y_{k}^{(N+1)}=y_{k}^{(N)}+\log n_{k}-\log\left(\left[\exp\left\{ \tilde{A}+{\rm diag}\left(y^{(N)}\right)\right\} \right]_{kk}+\left(n_{k}-1\right)e^{y_{k}^{(N)}-\tilde{c}_{kk}}\right), \] repeatedly, until convergence. Let $y$ denote the final value. (The convergences tends to be quick because $y$ is a fixed point to a contraction). • Compute $A={\rm exp}\left(\tilde{A}+{\rm diag}(y)\right)$.

Correlation/Moment Targeting of Dynamic Parameters

The dimension of $\eta$ in the score-driven model with $K$ groups is $d=K\left(K+1\right)/2$. For this model we adopt the following dynamic model \[ \eta_{t+1}=\left(I_{d}-\beta\right)\mu+\beta\eta_{t}+\alpha s_{t}, \] where $\beta$ and $\alpha$ are diagonal matrices. This makes the total number of parameters to be estimated $K\left(K+1\right)/2\times3$ when we use the Gaussian specification. Specifications with $t$-distributions will have additional degrees of freedom parameters.

So-called variance targeting is often used when estimating multivariate GARCH models, where the expected value of the conditional covariance matrix is estimated in an initial step.\footnote{Targeting is often found to be beneficial for prediction but can have drawbacks, e.g. for inference, see Pedersen:2016.} This idea can also be applied to the transformed correlations with an estimate of $\mu=\mathbb{E}(\eta_{t})$ as the target. In the present context, it would be more appropriate to call it correlation targeting, or moment targeting that encompasses many variations of this method. For the initial estimation of the target, $\mathbb{E}(\eta_{t})$, we follow ArchakovHansen:CanonicalBlockMatrix and estimate the unconditional sample block-correlation matrix with \[ \hat{C}=Q\hat{D}Q^{\prime},\quad\hat{D}=\left[

array[array omitted — 175 chars of source]

\right], \] where \[ \ensuremath{Y_{t}=Q^{\prime}X_{t}=}\ensuremath{\left(Y_{0,t}^{\prime},Y_{1,t}^{\prime},\ldots,Y_{K,t}^{\prime}\right)^{\prime}}\ensuremath{,\quad\hat{A}=\sum_{t=1}^{T}\ensuremath{Y_{0,t}Y_{0,t}^{\prime}},\quad\hat{\lambda}_{k}=\frac{n_{k}-\hat{A}_{kk}}{n_{k}-1}}. \] We then proceed to compute $\hat{\mu}=\gamma(\hat{C})$. Because $\gamma(C)$ is non-linear, $\hat{\mu}$ is only a first-order approximation of $\mu$, but our empirical results suggest that it is a good approximation.

Benchmark Correlation Model: The DCC Model

The original DCC model was proposed by Engle2002, see also EngleSheppard:2001. The original form of variance targeting could result in inconsistencies, see Aielli:2013, who proposed a modification that resolves this issue. This model is known as cDCC model and is given by:

align*[align* omitted — 73 chars of source]

where $Q_{t}$ is a symmetric positive definite matrix (whose dynamic properties are defined below) and $\Lambda_{Q_{t}}$ is the diagonal matrix with the same diagonal elements as $Q_{t}$. This structure ensures that $C_{t}$ is a valid correlation matrix. The dynamic properties of $C_{t}$ are defined from those of $Q_{t}$, which are defined by

align[align omitted — 218 chars of source]

where $\iota$ is the vector of ones, $Z_{t}$ is a $n\times1$ vector with standardized return shocks, $\odot$ is the Hadamard product (element by element multiplication), and $\bar{C}$, $\beta$ and $\alpha$ are unknown $n\times n$ matrices. Here $\bar{C}$ is the unconditional correlation matrix, which can be parametrized as $\mu={\rm vecl}(\log\bar{C})$. Note that this model has $n(n+1)/2$ time-varying parameters, as defined by the unique elements of ${\rm vech}(Q_{t})$. However, $C_{t}$ only has $n(n-1)/2$ distinct correlations, so there are $n$ redundant variable in $Q_{t}$.

Empirical Analysis

We estimate and evaluate the models using nine stocks (small universe) as well as 100 stocks (large universe). We will use industry sectors, as defined by the Global Industry Classification Standard (GICS), to form block structures in the correlation matrix and/or the heavy tail index. The ticker symbols for all 100 stocks are listed in Table (ref), organized by industry sectors. The nine stocks in the small universe are highlighted with bold font.

table[table omitted — 2,815 chars of source]

The sample period spans the period from January 3, 2005 to December 31, 2021, with a total of $T=4,280$ trading days. We obtained daily close-to-close returns from the CRSP daily stock files in the WRDS database.

The focus of this paper concerns the dynamic modeling of correlations, but in practice we also need to estimate the conditional variances. In our empirical analysis, we estimated each of the univariate time series of conditional variances using the EGARCH models by Nelson91, where the conditional mean has an AR(1) structure, as is common in this literature. Thus, the model for the $i$-th asset return on day $t$, $r_{i,t}$, is given by:

align[align omitted — 224 chars of source]

The parameter, $\tau_{i}$, is related to the well-known leverage effect, whereas $\theta_{i}$ is tied to the degree of volatility clustering. By modeling the logarithm of conditional volatility, the estimated volatility paths are guaranteed to be positive, which in conjunction with the parametrization of the correlation matrix, $C(\gamma)$, guarantees a positive definite conditional covariance matrix. At this stage of the estimation, we do not want to select a particular type of heavy tail distributions for $z_{i,t}$. So, we simply estimate the EGARCH models by quasi maximum likelihood estimation using a Gaussian specification. From the estimated time series for $h_{i,t}$, we obtain the vector of standardized returns, $Z_{t}=\left[z_{1,t},z_{2,t},\cdots,z_{n,t}\right]$, which are common to all the multivariate models we consider below.

table[table omitted — 3,400 chars of source]

Small Universe: Dynamic Correlations for Nine Stocks

We begin by analyzing nine stocks and we refer to this data set as the small universe. The nine stocks are: Marathon Oil (MRO), Occidental Petroleum (OXY), and Devon Energy (DVN) from the energy sector, Bank of America (BAC), Citigroup (C), and JPMorgan Chase & Co (JPM) from the Financial sector, and Microsoft (MSFT), Intel (INTC), and Cisco (CSCO) from the Information Technology sector. Table (ref) reports the full-sample unconditional correlation matrix (lower triangle) and its logarithm (upper-triangle) with the sector-based block structure illustrated with the shaded regions. Note that the estimated unconditional correlations within each of the blocks have similar averages. The assets within the Energy sector and Financial sector are highly correlated, with an average correlation of about 0.80. Within-sector correlations for Information Technology stock returns tend to be smaller, with an average of about 0.58. The between-sector correlations tend to be smaller and range from 0.36 to 0.51. A similar pattern is observed for the corresponding elements of the logarithm of the unconditional correlation matrix, as the logarithm transformation preserves the block structure.

sidewaystable\caption{Small Universe Estimation Results: $C_{t}$ Unrestricted} \begin{centering} \begin{footnotesize} \begin{tabularx}{\textwidth}{p{0.5cm}p{0.3cm}p{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Y} \toprule \midrule & & & \multicolumn{9}{c}{Score-Driven Model} & & \multicolumn{9}{c}{DCC Model} \\ \\[-0.2cm] & & & Gaussian & & Multiv.-$t$ & & Canon-$t$ & & Cluster-$t$ & & Hetero-$t$ & & Gaussian & & Multiv.-$t$ & & Canon-$t$ & & Cluster-$t$ & & Hetero-$t$ \\ \cmidrule{4-12}\cmidrule{14-22} \\[-0.2cm] \multirow{6}[0]{*}{$\mu$} & Mean & & 0.248 & & 0.265 & & 0.268 & & 0.269 & & 0.268 & & 0.239 & & 0.255 & & 0.250 & & 0.266 & & 0.266 \\ & Min & & 0.022 & & 0.027 & & 0.091 & & 0.076 & & 0.037 & & 0.077 & & 0.097 & & 0.093 & & 0.088 & & 0.095 \\ & $Q_{25}$ & & 0.127 & & 0.124 & & 0.124 & & 0.123 & & 0.129 & & 0.106 & & 0.116 & & 0.116 & & 0.117 & & 0.129 \\ & $Q_{50}$ & & 0.164 & & 0.169 & & 0.166 & & 0.165 & & 0.169 & & 0.136 & & 0.144 & & 0.149 & & 0.150 & & 0.155 \\ & $Q_{75}$ & & 0.227 & & 0.327 & & 0.307 & & 0.325 & & 0.306 & & 0.240 & & 0.331 & & 0.261 & & 0.317 & & 0.305 \\ & Max & & 0.816 & & 0.777 & & 0.805 & & 0.807 & & 0.815 & & 0.874 & & 0.783 & & 0.829 & & 0.848 & & 0.856 \\ \\[-0.2cm] \multirow{6}[0]{*}{$\beta$} & Mean & & 0.917 & & 0.962 & & 0.952 & & 0.970 & & 0.970 & & 0.967 & & 0.973 & & 0.974 & & 0.971 & & 0.972 \\ & Min & & 0.503 & & 0.601 & & 0.502 & & 0.803 & & 0.817 & & 0.937 & & 0.945 & & 0.938 & & 0.938 & & 0.939 \\ & $Q_{25}$ & & 0.881 & & 0.976 & & 0.951 & & 0.966 & & 0.964 & & 0.964 & & 0.972 & & 0.973 & & 0.967 & & 0.968 \\ & $Q_{50}$ & & 0.980 & & 0.990 & & 0.988 & & 0.988 & & 0.985 & & 0.968 & & 0.975 & & 0.977 & & 0.975 & & 0.976 \\ & $Q_{75}$ & & 0.995 & & 0.996 & & 0.994 & & 0.994 & & 0.994 & & 0.972 & & 0.979 & & 0.980 & & 0.978 & & 0.978 \\ & Max & & 0.999 & & 0.999 & & 0.999 & & 0.999 & & 0.999 & & 0.979 & & 0.984 & & 0.991 & & 0.986 & & 0.981 \\ \\[-0.2cm] \multirow{6}[0]{*}{$\alpha$} & Mean & & 0.019 & & 0.014 & & 0.016 & & 0.015 & & 0.016 & & 0.016 & & 0.013 & & 0.011 & & 0.012 & & 0.012 \\ & Min & & 0.001 & & 0.001 & & 0.001 & & 0.001 & & 0.001 & & 0.011 & & 0.008 & & 0.005 & & 0.007 & & 0.008 \\ & $Q_{25}$ & & 0.007 & & 0.005 & & 0.004 & & 0.005 & & 0.005 & & 0.012 & & 0.010 & & 0.008 & & 0.009 & & 0.011 \\ & $Q_{50}$ & & 0.014 & & 0.011 & & 0.009 & & 0.009 & & 0.011 & & 0.016 & & 0.013 & & 0.012 & & 0.012 & & 0.012 \\ & $Q_{75}$ & & 0.025 & & 0.019 & & 0.023 & & 0.023 & & 0.024 & & 0.019 & & 0.015 & & 0.014 & & 0.014 & & 0.014 \\ & Max & & 0.071 & & 0.069 & & 0.071 & & 0.057 & & 0.063 & & 0.028 & & 0.021 & & 0.016 & & 0.022 & & 0.019 \\ \\[-0.2cm] $\nu_{0}$ & & & & & 6.232 & & 6.391 & & & & & & & & 6.029 & & 6.501 & & & & \\ \\[0.0cm] & & & & & & & & & & & 6.193 & & & & & & & & & & 5.907 \\ $\nu_{1}$ & & & & & & & 5.517 & & 6.022 & & 5.320 & & & & & & 5.097 & & 5.839 & & 5.162 \\ & & & & & & & & & & & 4.597 & & & & & & & & & & 4.533 \\ & & & & & & & & & & & 4.552 & & & & & & & & & & 4.330 \\ $\nu_{2}$ & & & & & & & 4.919 & & 4.871 & & 4.614 & & & & & & 4.397 & & 4.671 & & 4.340 \\ & & & & & & & & & & & 5.165 & & & & & & & & & & 4.985 \\ & & & & & & & & & & & 3.289 & & & & & & & & & & 3.237 \\ $\nu_{3}$ & & & & & & & 3.714 & & 4.175 & & 3.807 & & & & & & 3.315 & & 4.046 & & 3.740 \\ & & & & & & & & & & & 4.023 & & & & & & & & & & 3.910 \\ \\[0.0cm] $p$ & & & 108 & & 109 & & 112 & & 111 & & 117 & & 126 & & 127 & & 130 & & 129 & & 135 \\ \\[-0.2cm] $\ell$ & & & -42603 & & -39971 & & {-39762} & & -39282 & & -39348 & & -42675 & & -40054 & & -39827 & & -39370 & & -39465 \\ $\ell_m$ & & & -54653 & & -52866 & & -52893 & & -52774 & & -52770 & & -54653 & & -52857 & & -52858 & & -52767 & & -52746 \\ $\ell_c$ & & & 12050 & & 12895 & & 13131 & & 13492 & & 13422 & & 11978 & & 12803 & & 13031 & & 13397 & & 13281 \\ \\[-0.2cm] AIC & & & 85422 & & 80160 & & {79748} & & \textbf{78786} & & 78930 & & 85602 & & 80362 & & 79914 & & \textbf{78998} & & 79200 \\ BIC & & & 86109 & & 80853 & & {80461} & & \textbf{79492} & & 79674 & & 86404 & & 81170 & & 80741 & & \textbf{79819} & & 80059 \\ \\[-0.2cm] \\[-0.5cm] \midrule \bottomrule \end{tabularx} \end{footnotesize} \end{centering} {Note: Parameter estimates for the full sample period, January 2005 to December 2021. The Score-Driven model and DCC model are both estimated with five distributional specifications, without imposing a block structure on $C_{t}$. We report summary statistics for the estimates of $\mu$, $\alpha$, and $\beta$, and report all estimates of the degrees of freedom. $p$ is the number of parameters and we report the maximized log-likelihoods, $\ell=\ell_{m}+\ell_{c}$, and its two components: the log-likelihoods for the nine marginal distributions, $\ell_{m}$, and the corresponding log-copula density, $\ell_{c}$. We also report the AIC $=-2\ell+2p$ and BIC $=-2\ell+p\ln T$. Bold font is used to identify the “best performing” specification in each row among Score-Driven models and among DCC models. }

We estimate three types of dynamic correlation models using five different distributions. The first type of model is the DCC model, see ((ref)). The second model is the new score-driven model for $C_{t}$, which we introduced in Section (ref). The third model is the score-driven model for a block correlation matrix, see Section (ref). We consider five distributional specifications for $U$, for each of these models. The distributions are: Gaussian, multivariate $t$, Canonical-Block-$t$, Cluster-$t$, and Hetero-$t$ distributions. We impose a diagonal structure on the matrices, $\alpha$ and $\beta$. In Tables (ref) and (ref) we report means and quantiles for the estimated parameters, $\mu$, ${\rm diag}\left(\beta\right)$, ${\rm diag}\left(\alpha\right)$ for score-driven model, and $\mu$, ${\rm vech}\left(\beta\right)$, ${\rm vech}\left(\alpha\right)$ for DCC model, i.e. the DCC models have more parameters. We denote $p$ as the number of parameters, $\ell$ is the full log-likelihood function, $\ell_{m}$ and $\ell_{c}$ are the log-likelihood for marginal densities and copula functions. We also report the Akaike and Bayesian information criteria (AIC and BIC) to compare the performance of models with different number of parameters.

Table (ref) reports the estimation results for the DCC model and score-driven models for general correlation matrix (Score-Full model). There are several interesting findings: First, the score-driven model provides superior performance relative to the simple DCC model for all five specifications of distributions. Second, the models with heavy-tailed distributions perform better than the corresponding model with a Gaussian distribution. For the score-driven models we see that persistence parameter, $\beta$, is larger for with heavy tailed specifications, as the existence of $W$ would mitigate the effect from extreme value in updating interested parameters. Third, introducing the structured heavy tails greatly improve the model performances, as indicated by higher likelihood values $\ell$. That this improves the empirical fit is supported by the estimated degree of freedoms, which are different for different asset groups. The Information Tech sector is estimated to have the heaviest tails, follow by the Financial and Energy sectors. Fourth, the degree of freedoms estimated from Cluster-$t$ distribution is larger than the averages of each group from Hetero-$t$ distribution, as we have explained earlier. Fifth, from the decomposition of $\ell$, we could observe that the improvements of Canonical-Block-$t$ relative to the multivariate $t$-distribution are all driven by the copula part. This is also the case for comparing Hetero-$t$ and Cluster-$t$ distributions. Although the former provides more flexibility in fitting marginal distribution of individual asset, it doesn't necessarily lead to a better dependence structure. In this dataset, the Cluster-$t$ provides the largest copula functions, as it allows for a common $\chi^{2}$ shock among assets within the same group.

table[table omitted — 4,855 chars of source]

Table (ref) presents the estimation results for the score-driven models for block correlation matrix (Score-Block model). We report all the estimated coefficients with subscripts referring to the parameters for within/between groups with \textquotedblleft Energy=1, Financial=2, Information Tech=3\textquotedblright . Results are similar to the Table (ref). When compare with the (ref), we could find although the DCC models the general correlation matrix, the restricted Score-Block models provide superior performances with the last three convolution-$t$ specifications. And compared with the Score-Full models, the Score-Block models delivery smaller BIC for all specifications, and smaller AIC for the last three cases. We plot the time series of correlations in Figure (ref) filtered by Cluster-$t$ distributions. Several heterogenous patterns are observed: First, expect for the within correlations for financial sector, other correlations have a sharp decline in late 2010 and increase in early 2011. Second, the inter-group correlations that involves Energy sector have a evident decline in late 2008 and the recovered.

sidewaysfigure[ph] \caption{{Within-sector and between-sector conditional correlations implied by the estimated Score-Driven model with a sector-based cluster structure in correlations and the Convolution-$t$ distribution (Cluster-$t$ with block correlation matrix).}}

Large Universe: Dynamic Correlation Matrix for $100$ Assets

Next, we estimate the model with the large universe, where $C_{t}$ has dimension $100\times100$. We use the sector classification, see Table (ref), to define the block structure on $C_{t}$. Ten (of the eleven) sectors represented in the Large Universe, such that $K=10$, and the number of unique correlations in $C_{t}$ is reduced from 4,950 to 55. We estimate the score-driven model with and without correlation targeting, see Section (ref). With correlation targeting, the intercept, $\mu$, is estimated first, and the remaining parameters are estimated in a second stage.

sidewaystable\caption{Large Universe Estimation Results: $C_{t}$ with Block Structure} \begin{centering} \begin{footnotesize} \begin{tabularx}{\textwidth}{p{0.5cm}p{0.3cm}p{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Yp{-0.2cm}Y} \toprule \midrule & & & \multicolumn{9}{c}{Score-Block Model with Full Parametrization} & & \multicolumn{9}{c}{Score-Block Model with Correlation Targeting} \\ \\[-0.2cm] & & & Gaussian & & Multiv.-$t$ & & Canon-$t$ & & Cluster-$t$ & & Hetero-$t$ & & Gaussian & & Multiv.-$t$ & & Canon-$t$ & & Cluster-$t$ & & Hetero-$t$ \\ \cmidrule{4-12}\cmidrule{14-22} \\[-0.2cm] \multirow{6}[0]{*}{$\mu$} & Mean & & 0.053 & & 0.051 & & 0.056 & & 0.056 & & 0.055 & & 0.053 & & 0.053 & & 0.053 & & 0.053 & & 0.053 \\ & Min & & 0.004 & & 0.001 & & 0.008 & & 0.008 & & 0.007 & & 0.002 & & 0.002 & & 0.002 & & 0.002 & & 0.002 \\ & $Q_{25}$ & & 0.026 & & 0.025 & & 0.026 & & 0.027 & & 0.026 & & 0.025 & & 0.025 & & 0.025 & & 0.025 & & 0.025 \\ & $Q_{50}$ & & 0.038 & & 0.037 & & 0.038 & & 0.039 & & 0.040 & & 0.036 & & 0.036 & & 0.036 & & 0.036 & & 0.036 \\ & $Q_{75}$ & & 0.048 & & 0.049 & & 0.048 & & 0.048 & & 0.050 & & 0.049 & & 0.049 & & 0.049 & & 0.049 & & 0.049 \\ & Max & & 0.333 & & 0.327 & & 0.343 & & 0.348 & & 0.342 & & 0.349 & & 0.349 & & 0.349 & & 0.349 & & 0.349 \\ & & & & & & & & & & & & & & & & & & & & & \\ \multirow{6}[0]{*}{$\beta$} & Mean & & 0.842 & & 0.886 & & 0.888 & & 0.903 & & 0.887 & & 0.850 & & 0.891 & & 0.902 & & 0.915 & & 0.905 \\ & Min & & 0.443 & & 0.432 & & 0.612 & & 0.654 & & 0.619 & & 0.401 & & 0.415 & & 0.614 & & 0.651 & & 0.621 \\ & $Q_{25}$ & & 0.794 & & 0.859 & & 0.817 & & 0.853 & & 0.808 & & 0.799 & & 0.854 & & 0.827 & & 0.861 & & 0.846 \\ & $Q_{50}$ & & 0.897 & & 0.983 & & 0.935 & & 0.940 & & 0.944 & & 0.898 & & 0.983 & & 0.941 & & 0.951 & & 0.954 \\ & $Q_{75}$ & & 0.975 & & 0.996 & & 0.989 & & 0.985 & & 0.989 & & 0.976 & & 0.996 & & 0.991 & & 0.989 & & 0.990 \\ & Max & & 0.999 & & 0.999 & & 0.999 & & 0.999 & & 0.999 & & 0.999 & & 1.000 & & 0.999 & & 0.999 & & 0.999 \\ & & & & & & & & & & & & & & & & & & & & & \\ \multirow{6}[0]{*}{$\alpha$} & Mean & & 0.030 & & 0.023 & & 0.042 & & 0.041 & & 0.041 & & 0.030 & & 0.023 & & 0.041 & & 0.039 & & 0.041 \\ & Min & & 0.001 & & 0.003 & & 0.005 & & 0.006 & & 0.004 & & 0.001 & & 0.004 & & 0.008 & & 0.005 & & 0.003 \\ & $Q_{25}$ & & 0.013 & & 0.007 & & 0.015 & & 0.014 & & 0.016 & & 0.013 & & 0.007 & & 0.013 & & 0.015 & & 0.015 \\ & $Q_{50}$ & & 0.028 & & 0.015 & & 0.035 & & 0.038 & & 0.036 & & 0.028 & & 0.015 & & 0.032 & & 0.029 & & 0.036 \\ & $Q_{75}$ & & 0.045 & & 0.029 & & 0.057 & & 0.058 & & 0.055 & & 0.045 & & 0.028 & & 0.058 & & 0.055 & & 0.055 \\ & Max & & 0.102 & & 0.108 & & 0.137 & & 0.127 & & 0.140 & & 0.103 & & 0.106 & & 0.138 & & 0.128 & & 0.139 \\ & & & & & & & & & & & & & & & & & & & & & \\ $\nu_0$ & & & & & 10.06 & & 12.25 & & & & & & & & 10.11 & & 12.20 & & & & \\ $\nu_1$ & & & & & & & 7.111 & & 7.320 & & 4.852$^\dagger$ & & & & & & 7.158 & & 7.341 & & 4.882$^\dagger$ \\ $\nu_2$ & & & & & & & 5.778 & & 6.012 & & 4.723$^\dagger$ & & & & & & 5.489 & & 5.812 & & 4.635$^\dagger$ \\ $\nu_3$ & & & & & & & 6.640 & & 6.903 & & 4.451$^\dagger$ & & & & & & 6.397 & & 6.644 & & 4.368$^\dagger$ \\ $\nu_4$ & & & & & & & 5.373 & & 5.579 & & 3.884$^\dagger$ & & & & & & 5.306 & & 5.546 & & 3.849$^\dagger$ \\ $\nu_5$ & & & & & & & 5.657 & & 5.896 & & 4.049$^\dagger$ & & & & & & 5.627 & & 5.847 & & 4.018$^\dagger$ \\ $\nu_6$ & & & & & & & 6.024 & & 6.263 & & 4.070$^\dagger$ & & & & & & 5.983 & & 6.183 & & 4.021$^\dagger$ \\ $\nu_7$ & & & & & & & 6.018 & & 6.159 & & 4.280$^\dagger$ & & & & & & 6.008 & & 6.127 & & 4.289$^\dagger$ \\ $\nu_8$ & & & & & & & 5.581 & & 5.784 & & 3.829$^\dagger$ & & & & & & 5.516 & & 5.706 & & 3.779$^\dagger$ \\ $\nu_9$ & & & & & & & 7.022 & & 7.286 & & 5.347$^\dagger$ & & & & & & 7.043 & & 7.283 & & 5.337$^\dagger$ \\ $\nu_{10}$ & & & & & & & 5.693 & & 6.007 & & 4.366$^\dagger$ & & & & & & 5.658 & & 5.945 & & 4.334$^\dagger$ \\ & & & & & & & & & & & & & & & & & & & & & \\ $p$ & & & 165 & & 166 & & 176 & & 175 & & 265 & & 110 & & 111 & & 121 & & 120 & & 210 \\ \\[-0.2cm] $\ell$ & & & -481966 & & -464247 & & -448572 & & -446633 & & -436613 & & -482019 & & -464307 & & -448640 & & -446711 & & -436704 \\ $\ell_m$ & & & -607256 & & -588971 & & -589777 & & -587713 & & -586615 & & -607256 & & -589006 & & -589625 & & -587506 & & -586446 \\ $\ell_c$ & & & 125292 & & 124724 & & 141204 & & 141079 & & 150002 & & 125236 & & 124699 & & 140986 & & 140795 & & 149742 \\ \\[-0.2cm] AIC & & & 964262 & & 928826 & & 897496 & & 893616 & & 873756 & & 964258 & & 928836 & & 897522 & & 893662 & & 873828 \\ BIC & & & 965312 & & 929882 & & 898616 & & 894729 & & 875442 & & 964958 & & 929542 & & 898292 & & 894425 & & 875164 \\ \\[-0.2cm] \\[-0.5cm] \midrule \bottomrule \end{tabularx} \end{footnotesize} \end{centering} {Note: Parameter estimates for the full sample period, January 2005 to December 2021. Score-Driven models with a block correlation structure and five distributional specifications are estimated without correlation targeting (left panel) and with correlation targeting (right panel). We report summary statistics for the estimates of $\mu$, $\alpha$, and $\beta$, and all estimates of the degrees of freedom, except for the Heterogeneous Convolution-$t$ specifications where we report the average estimate within each cluster., as identified with the $\dagger$-superscript. $p$ is the number of parameters and we report the maximized log-likelihoods, $\ell=\ell_{m}+\ell_{c}$, and its two components: the log-likelihoods for the nine marginal distributions, $\ell_{m}$, and the corresponding log-copula density, $\ell_{c}$. We also report the AIC $=-2\ell+2p$ and BIC $=-2\ell+p\ln T$. Bold font is used to identify the “best performing” specification in each row for models with and without correlation targeting.}

Table (ref) reports the estimation results for the the score-driven models with block correlation matrices. The left panel has estimation results for models without correlation targeting, and the right panel has the estimation results based on correlation targeting. The estimates identified with a $\dagger$-superscript, are the average degrees of freedom within each cluster. These are used for specifications with heterogeneous Convolution-$t$ specifications (Hetero-$t$), which estimates 100 degrees of freedom parameters. Compared with the results for the Small Universe, we note some interesting difference. First, different from the results on small universe, the model with hetero-$t$ distribution now provides the best fitting performance, and compared with Cluster-$t$ distribution, its improvement concentrates on the copula part. This may due to the high level of heterogeneity across the large dataset, and the simple classification based GICS is poor.\footnote{One could estimates the group structure by using the method in OhPatton2023, here we only focus such simple classification to assess our score-driven model in modeling high-dimensional assets.} Second, the models estimated with targeting perform well and have the smallest BIC across all distributional specifications.

Out-of-sample Results

We next compare the out-of-sample (OOS) performance of the different models/specifications. We estimate all models (once) using data from 2005-2014 and evaluate the estimated models with (out-of-sample) data that spans the years: 2015-2021.

The OOS results for the Small Universe are shown in Panel A of Table (ref). We decompose the predicted log-likelihood, $\ell$, into the marginal, $\ell_{m}$, and copula, $\ell_{c}$, components. For each of the five distributional specifications, we have highlighted the largest predicted log-likelihood, which is the Score-Driven model without a block structure on $C_{t}$, for all five distributions. This is consistent with our in-sample results, where this model also had the largest (in-sample) log-likelihood for each of the five distributional specifications, see Tables (ref) and (ref). Overall, the Convolution-$t$ distribution with a sector-based cluster structure, Cluster-$t$, has the largest predictive log-likelihood. We also note that the DCC model is has the worst performance across all distributional specifications. In sample, the DCC model was slightly better than the Score-Driven model with a block correlation matrix, for two of the five distributions (Gaussian and multivariate $t$). This suggests that the DCC suffer from an overfitting problem.

table[table omitted — 3,515 chars of source]

We report the OOS results for the Large Universe in Panel B of Table (ref), where all model-specifications employ a block structure on $C_{t}$. The empirical results favor correlation targeting, because the Score-Driven model with correlation targeting has the largest predicted log-likelihood for each of the five distributions. Across the five distributions, the Convolution-$t$ distribution based on $100$, independent $t$-distributions, Hetero-$t$, has the largest predictive log-likelihood.

Summary

We have introduced the Cluster GARCH model, which is a novel multivariate GARCH model, with two types of cluster structures. One that relates to the correlation structure and one that define non-linear dependencies. The Cluster GARCH framework combines several useful components from the existing literature. For instance, we incorporate the block correlation structure by EngleKelly:2012, the correlation parametrization by ArchakovHansen:Correlation, and the convolution-$t$ distributions by HansenTong:2024. A convolution-$t$ distribution is a multivariate heavy-tailed distribution with cluster structures, flexible nonlinear dependencies, and heterogeneous marginal distributions. We also adopted the score-driven framework by CrealKoopmanLucas:2013 to model the dynamic variation in the correlation structure. The convolution-$t$ distributions are well-suited for score-driven models, because their density functions are sufficiently tractable, allowing us to derive closed-form expressions for the key ingredients in score-driven models: the score and the Hessian. We derived detailed results for three special types of convolution-$t$ distributions. These are labelled Canonical-Block-$t$, Cluster-$t$, and Hetero-$t$, and their score functions and Fisher informations are all available in closed-form.

Applying the model to high-dimensional systems is possible when the block correlation structure is imposed. This was pointed out in ArchakovHansenLundeMRG, but the present paper is first to demonstrate this empirically with $n=100$. This was achieved with $K=10$ sector-based clusters that was used to define the block structure on the correlation matrix. The block structure is advantages for several reason. First, it reduces the number of distinct correlations in $C_{t}$ from 4,950 to 55 ($n(n-1)/2$ to $K(K+1)/2$). Second, many likelihood computations are greatly simplified due to the canonical representation of block correlation matrix, see ArchakovHansen:CanonicalBlockMatrix. An important implication for the dynamic model is that computations only involve inverses, determinants, square-roots of $K\times K$ matrices rather than $n\times n$ matrices.

We conduct an extensive empirical investigation on the performance of our dynamic model for correlation matrices. And we consider a “small universe” with $n=9$ assets and a “large universe” with $n=100$ assets. The empirical results find strong support for convolution-$t$ distributions that outperforms conventional distributions, in-sample as well as out-of-sample. Moreover, the score-driven framework out-performs the standard DCC model in all cases (dimensions and choice of distribution). The score-driven model with a sector-based block correlation matrix has the smallest BIC.