Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
113,503 characters · 17 sections · 52 citation commands
Multivariate GARCH and portfolio variance prediction: a forecast reconciliation perspective
\doublespacing
Multivariate conditional heteroskedastic models belonging to the Generalized Auto Regressive Conditional Heteroskedasticity (MGARCH) class are standard tools in financial econometrics. They represent classic instruments for a range of applications, from risk management to hedging and asset allocation; see surveys by bauwensetal2005, SilvennoinenTerasvirta2009, and FrancqZakoian2019. Among the several specifications proposed by the literature (see the surveys previously cited for details), only a few are frequently adopted, and these include the Dynamic Conditional Correlation (DCC) model of Engle2002, the Orthogonal GARCH (OGARCH) model of alexander2002, and the Scalar BEKK of DingEngle2001, which are feasible even in large dimensional cases. A relevant challenge for MGARCH models is, in fact, related to the so-called curse of dimensionality, that is, the large number of parameters of the most flexible models, which is sensibly limiting their use. For instance, the BEKK model of englekroner1995 is commonly considered only under strong parametric restrictions, such as in the just mentioned Scalar-BEKK case, while the VECH specification (see again englekroner1995) has received limited attention in empirical analyzes. However, restrictions imply a reduction in model flexibility, which generally corresponds to a limit in the interdependence across assets shocks and/or covariances (or correlations). This is in contrast with the empirical evidence suggesting that interdependence is present and relevant, as demonstrated by the various works re-introducing limited flexibility in MGARCH specifications; see, among many others, CaporinParuolo2015 and billioetal2023. From a different perspective, the existence of interdependence is also at the base of the growing literature on systemic risk and financial connectedness, originating from the seminal contribution of DieboldYilmaz2009; see also DieboldYilmaz2023 and references therein cited.
An aspect not yet fully explored by the literature is the possibility of indirectly taking into account the interdependence in a MGARCH framework and of exploiting its potential in forecasting, thus paving the way for possible applications either in risk management and portfolio allocation. Our intuition takes the point of view of a risk manager that has to evaluate the risk (i.e., to predict the variance) of an existing portfolio. By construction, the weights of the assets are known, and if sufficient historical data are available for the returns or all the assets included in the portfolio, backward simulated, synthetic, portfolio returns can be derived. The portfolio returns are then accounting for the interdependence across the assets risk (either in terms of shocks or variance spillovers), and this will be captured implicitly by any univariate GARCH-type model we might fit on the portfolio returns. Of course, a univariate model might represent a too restrictive specification, not taking properly into account the heterogeneity in the dynamic of assets variances, covariances and/or correlations, something that could be provided only by a MGARCH model. Fortunately, the forecasting literature includes tools for combining a fit of a univariate model in a synthetic portfolio returns series, with a MGARCH model adapted to the collection of portfolio assets returns; we refer to forecast reconciliation approaches Hyndman2011, Wickramasuriya2019, Panagiotelis2021, DifonzoGirolimetto2022, Girolimetto2024, Girolimettoetal2023, Athanasopoulos2024.
In this work, we show that the prediction of the variance of a univariate GARCH model on portfolio returns can be efficiently combined with the prediction of the covariance of an MGARCH specification to improve the prediction of a given portfolio risk. We thus contribute both to the literature on forecast reconciliation, extending the application of its tools in finance (see Caporin2024 and Matteraetal2024), and to the MGARCH literature, showing how assets interdependence can be accounted for. For what concerns the forecast reconciliation contribution, we design a reconciliation approach tailored to the portfolio-variance setting, adapting shrinkage-based methods to the aggregation constraints implied by portfolio weights. Our procedure ensures coherence and preserves the structure of the covariance matrix, making the reconciled forecasts more accurate and economically interpretable. Using simulations, we show how the combination of univariate and multivariate forecasts improves prediction accuracy, in line with the findings of Wickramasuriya2019 and Panagiotelis2021. Specifically, when focusing on the financial econometrics side, in addition to the introduction of interdependence by means of a univariate fit, we provide further contributions. First, we show how the predictive performance of MGARCH models might deteriorate under misspecification. Second, and more relevant, we demonstrate that the use of noisy proxies masks the possible presence of interdependence, making MGARCH models neglecting interdependence equivalent to the correctly specified ones.
The paper proceeds as follows. Section 2 is devoted to the methods that show how forecast reconciliation can be used for portfolio variances, describing both the MGARCH models and the forecast comparison tools that we will consider. Section 3 reports the simulation design and the results, while Section 4 includes an empirical example with real data. Section 5 concludes the paper. An extensive Online Appendix includes detailed results for the simulations and the empirical study.
As we noted in the introduction, the portfolio variance can be forecast either through a univariate approach applied directly to portfolio returns or via a bottom-up strategy leveraging a multivariate model of asset returns. This dual structure provides an ideal setting for the application of forecast reconciliation techniques.
We first define the framework we refer to. Assume that the interest is in the prediction of the portfolio variance for the next period, a common need in financial risk management. If this is the case, the portfolio composition is assumed to be known. Therefore, in a setting where $N$ assets are available, we denote by $\bm{\omega}$ the (known) $N-$dimensional vector of portfolio weights. We also assume that the returns\footnote{For simplicity, we assume that the returns are de-meaned.} of the $N$ assets are available over a common sample, that is, the $N$-dimensional vectors $\boldsymbol r_t \in \mathbb{R}^N$ are observed for $t=1,2,\ldots, T$.
The goal is to produce a forecast of the portfolio variance for $T+1$. Two approaches are commonly pursued. In the first, one computes the portfolio returns $r_{p,t} = \bm{\omega}^{\prime}\boldsymbol r_t$ and fits a univariate GARCH-type model to obtain the one-step-ahead forecast $\sigma_{p,T+1}^2$ of the portfolio variance.\footnote{For simplicity, we do not denote predicted quantities with hats or conditioning on the available information set, unless it is needed for clarity of the discussion.} In the second approach, a multivariate GARCH (MGARCH) model is fitted to the asset return vector $\boldsymbol r_t$, providing a forecast $\bm{\Sigma}_{T+1}$ of the conditional covariance matrix. Then, we recover the predicted portfolio variance by aggregating as $\gamma_{p,T+1}^2=\bm{\omega}^{\prime} \bm{\Sigma}_{T+1}\bm{\omega}$.
The presence of a reference forecast recovered from a univariate approach and of a second forecast obtained by aggregation of a collection of elements (the predictions of asset variances and covariances) represents a standard situation in forecast reconciliation. Similar settings appear in non-financial applications such as hierarchical forecasting in demography Shang2017, Li2019, macroeconomic Petropoulos2014, Mircetic2022 energy Silva2018, Wang2021, Abolghasemi2025. In the financial literature, notable examples are Caporin2024 and Matteraetal2024. To simplify notation, all forecast quantities are implicitly assumed to refer to the next period, and the time subscript $T+1$ is omitted throughout the discussion of this subsection: $\sigma_{p,T+1}^2 \equiv \widehat{\sigma}_{p}^2$ and $\gamma_{p,T+1}^2 \equiv \gamma_{p}^2$, $\bm{\Sigma}_{T+1} \equiv\widehat{\bm{\Sigma}}$.
Moving back to our setting, the prediction of portfolio variance using a univariate approach and of the covariance matrix, $\widehat{\sigma}_{p}^2$ and $\widehat{\bm{\Sigma}}$, represent the so-called base forecasts. The construction of the portfolio variance prediction by aggregation gives the bottom-up alternative $\gamma_{p}^2$, exploiting the informative content of the risk predictions coming from the single components of the portfolio and accounting for their interdependence (measured by both the correlations and the links between variances and covariances). In general, the base forecast and the bottom-up forecast do not coincide, that is, $\widehat{\sigma}_{p}^2 \neq \gamma_{p}^2$.
To connect our framework with the results of the forecast reconciliation literature, we express the bottom-up forecast as a linear combination of the univariate underlying elements (variances and covariances) of $\widehat{\bm{\Sigma}}$:
where $\mathit{vech}(\cdot)$ denotes the half-vectorization operator, $\bm{D}_N$ is the duplication matrix,\footnote{The duplication matrix satisfies $\mathit{vec}\left(\widehat{\bm{\Sigma}}\right)=\bm{D}_N \mathit{vech}\left(\widehat{\bm{\Sigma}}\right)$ with $\mathit{vec}$ being the matrix vectorization operator and has size $\left(N^2 \times \frac{N(N+1)}{2}\right)$.} and $\widehat{\bm{\sigma}}$ is the vector of dimension $m = \frac{N(N+1)}{2}$ collecting the different elements of $\widehat{\bm{\Sigma}}$. The $(m \times 1)$ vector $\bm{a} = \left[(\bm{\omega}^{\prime} \otimes \bm{\omega}^{\prime}) \bm{D}_N\right]^{\prime} = \bm{D}^{\prime}_N (\bm{\omega} \otimes \bm{\omega}) $ is also called the aggregation Girolimetto2024 vector that is equivalent to the aggregation matrix in the reconciliation literature Hyndman2011. Then, the portfolio variance forecasts obtained from the univariate GARCH-type approach and from the MGARCH model can be reconciled by solving the following generalized least squares problem:
with the solution given by Stone1942, Byron1978, Byron1979: $$ \widetilde{\bm{y}} = \left[\bm{I}_{m+1} - \bm{\Omega} \bm{c} (\bm{c}^{\prime} \bm{\Omega} \bm{c})^{-1} \bm{c}^{\prime} \right] \widehat{\bm{y}}, $$ where $$ \bm{y} =
, \qquad \widehat{\bm{y}} =
, \qquad \widetilde{\bm{y}} =
, \qquad and \qquad \bm{c} =
, $$ are, respectively, the $[(m+1) \times 1]$ stacked variance vector, the base forecast vector collecting the incoherent variance forecasts from both the univariate and multivariate models, the reconciled forecast vector, and the constraints vector. In addition, $\bm{\Omega}$ is a $[(m+1) \times (m+1)]$ positive definite matrix representing the covariance matrix of the (time series of) forecast errors of both the bottom-level components (i.e., the elements of vector $\widehat{\bm{\sigma}}$) and the portfolio series (i.e., $\widehat{\sigma}_{p}^2$).
Several approaches have been proposed in the literature to estimate $\bm{\Omega}$, see Athanasopoulos2024. However, in this work, we consider the state-of-the-art shrinkage estimators proposed by Wickramasuriya2019, based on the in-sample residuals of the individual forecasting models Hyndman2016, Panagiotelis2021, Caporin2024.
After reconstructing the reconciled covariance matrix as $\widetilde{\bm{\Sigma}} = \mathit{vech}^{-1}(\widetilde{\bm{\sigma}})$, the corresponding correlation matrix $\widetilde{\bm{R}}$ may not satisfy the required properties of a proper correlation matrix. In particular, it may contain off-diagonal elements with absolute values that exceed one, that is, $|\widetilde{\rho}_{i,j}| > 1$, thus violating the mathematical definition of a correlation coefficient. Therefore, we propose two novel reconciliation strategies designed to specifically address this issue:
Taking into account the previous elements, we propose a method for reconciling portfolio variance forecasts to enhance the accuracy of risk estimation by combining univariate and multivariate approaches. The steps of this reconciliation procedure are described in Algorithm (ref).
We stress that the main objective of our research is to determine if forecast reconciliation tools lead to potential improvements in the prediction of portfolio variance. Moreover, we are also interested in determining if, by exploiting the information contained in the direct portfolio variance prediction (based on univariate methods), we will improve the prediction based on the commonly adopted MGARCH models, which might be miss-specified due to the lack of variance interdependence. In both cases, the evaluation will compare the baseline portfolio variance forecast with the bottom-up one and forecast reconciliation, for a given pair of fitted univariate GARCH and MGARCH specifications.
As mentioned in the previous section, the construction of reconciled forecasts for the portfolio variance, starting from either portfolio or asset returns, requires the specification of a univariate model for the former and of a multivariate model for the latter.
For simplicity, in the case of the portfolio returns, we specify a simple GARCH(1,1) model:
We are aware that a more standard approach is now given by a model with asymmetry in the variances, but we prefer to maintain now the coherence in the features captured by the model providing the base forecast and the multivariate model behind the bottom-up forecast. Further generalizations are left for future work.
Moving to the Multivariate GARCH models used to produce forecasts, we restrict our attention to two specific cases: the BEKK model of englekroner1995 and the Dynamic Conditional Correlation (DCC) model of Engle2002. For the BEKK model, we consider the simplest specification:
where $\mathcal{C}$ is lower triangular, while $\mathcal{A}$ and $\mathcal{B}$ should be full matrices. However, in empirical studies, to deal with the so-called curse of dimensionality, both $\mathcal{A}$ and $\mathcal{B}$ are restricted to be diagonal or even driven by a single parameter, see, among many others, englekroner1995,DingEngle2001,CaporinMcAleer2008.
The second model builds on the DCC dynamic for correlations, accompanied by a peculiar dynamic over marginals to allow, potentially, for variance interdependence. The model we consider might be seen as a special case of the Vector ARMA-GARCH of LingMcAleer2003 with DCC dynamic, and has been proposed, in the case of constant conditional correlation, by HeTerasvirta2004, and by Caporale2014 for the DCC model. Our target is to maintain model feasibility (that is, allowing for parameter estimation in a two-step procedure, first univariate GARCH on the marginals and then the correlation dynamic) and to allow for interdependence, limiting attention to positive spillover effects; the generalization allowing for negative parameters as in ConradKaranasos2010 is left to future research. In our simulations, the univariate models for the single assets are thus specified as follows:
We define the standardized innovations $\eta_{i,t}=\sigma_{i,t}^{-1}r_{i,t-1}$ for $i=1,2,\ldots n$ and then specify a DCC model
In the following, we refer to this model as Extended DCC (DCC with variance interaction, EDCC).
In the simulations, we will also consider the restricted specifications of both the BEKK and the DCC models. In detail, we will consider the BEKK model with scalar parameters and the baseline DCC model without interactions (i.e., the marginals are simple GARCH(1,1)).
Following the standard practice, the models are estimated by Quasi Maximum Likelihood, and in the case of the DCC specification using a multi-step procedure, starting from the marginals, then the $\boldsymbol\Gamma$ matrix (with a sample correlation estimator on $\boldsymbol\eta_t$), and finally the parameters governing the correlation dynamic.
In this section, we describe the methodologies used to compare the forecast performance of different portfolio variance forecast approaches. These methods include the accuracy of the point forecast using both absolute and relative indicators, and employ hypothesis testing based on two different well-known statistical procedures: the Diebold1995 test and the Model Confidence Set Hansen2011.
Five different approaches are compared based on their accuracy in forecasting the portfolio variance. These approaches include both univariate and multivariate models, as well as different reconciliation techniques, aligning with the steps described in Algorithm 1 for forecast reconciliation. The approaches considered are as follows:
To evaluate the point forecast accuracy of the different models, we employ three widely used metrics: the Mean Squared Error Davydenko2013, the Mean Absolute Error Davydenko2013, and the Quasi-Likelihood score Patton2011a, Caporin2024. Each of these metrics provides a distinct perspective on the accuracy of the predicted variance. The expressions for these measures are given by the following:
where $\sigma_{p,i}^2$ denotes the true portfolio variance at time $i$, $h_{i,j,q}^2$ denotes the corresponding portfolio variance forecast from approach $j$ at replication $q$, $M$ denotes the size of the test set,\footnote{This is set to 250 for simulations in Section (ref), and 4037 for applications in Section (ref).} and $Q$ denotes the number of replications.\footnote{We use 500 for simulations in Section (ref) and 1 for the application in Section (ref).} The index $j$ identifies the forecasting approach, with $j \in \{base, bu, shr, shr_A, shr_B\}$. We note that while MSE and MAE are symmetric loss functions, QLIKE is asymmetric, with a larger loss associated to under-prediction. To obtain the overall accuracy measures, we compute the average of each metric across all replications: $$ \text{IND}_j = \frac{1}{Q}\sum_{q = 1}^{Q} \text{IND}_{j,q}, $$ and for relative accuracy Davydenko2013, we calculate: $$ \text{AvgRelIND}_{j,x} = \left(\prod_{q = 1}^{Q} \frac{\text{IND}_{j,q}}{ \text{IND}_{x,q}}\right)^{\frac{1}{Q}}, $$ where $x \in \{base, bu\}$ and $\text{IND}$ represents one of the evaluation metrics (MSE, MAE, or QLIKE). Specifically, $\text{AvgRelIND}_{bu}$ and $\text{AvgRelIND}_{base}$ are two relative indicators that show an improvement in forecast compared to the bu and base models, respectively.
In addition to these overall and relative measures, we conduct hypothesis testing to compare forecast accuracy across different models. We employ the Diebold and Mariano Diebold1995 (DM) test to evaluate the null hypothesis of equal predictive accuracy (EPA) between competing models, with an overall significance level set at $\alpha=0.05$. To account for multiple pairwise comparisons, the DM test is implemented using the Bonferroni correction Dunn1961-xm. In addition, we apply the Model Confidence Set (MCS) procedure proposed by Hansen2011 to identify the subset of models that cannot be statistically distinguished in terms of forecast accuracy, reporting results at confidence levels of 70%, 75%, 80%, 85%, 90% and 95%. These tests are crucial in determining whether the observed differences in forecast performance are statistically significant or are merely due to randomness in the data.
In the following sections, we will consider the forecast comparison for both simulated and observed data. In the first case, the true portfolio variance $\sigma^2$ is known and will be used to evaluate the competing approaches. However, with real data, the true variance is not observed and a proxy must be used. Two solutions are commonly considered, the first being squared observed de-meaned returns. As shown by PattonSheppard2009 for the univariate case and Laurentetal2013 for multivariate models, this noisy proxy might lead to non-robust model selection for some loss functions, in particular for the MAE (while MSE and QLIKE are loss functions robust the presence of noise in the proxy). An alternative proxy might be recovered using high frequency data, and the daily variance (covariance) is estimated using the Realized Variance (Realized Covariance). Given the complexity of simulating high frequency data under a Multivariate GARCH model, we chose to first consider the noisy proxy in the simulation experiments, and then to mimic the existence of a better proxy by contaminating with additive noise the true covariance. In contrast, realized measures will be used, together with the squared returns, as proxies for the unknown variance when considering real data.
We consider several scenarios with varying sample sizes and number of variables. In particular, simulated return series are generated for portfolios comprising $9$ and $24$ assets using four distinct data generating processes: a fully parametrized BEKK, a scalar BEKK, the standard DCC-GARCH, and the EDCC-GARCH (incorporating interactions). We begin by illustrating the DGP used in the simulations for the case of $9$ assets.
The covariance matrix in the full BEKK specification is generated according to equation (ref) and setting $\boldsymbol r_t = \boldsymbol \Sigma_t^{\frac{1}{2}}\boldsymbol z_t$ with $\boldsymbol z_t$ sampled from a Multivariate Normal density with zero mean and covariance equal to the identity matrix. In terms of parameters, this formulation represents the most general approach, capturing a broad range of interdependencies. In contrast, the scalar BEKK restricts the dynamics by imposing that the matrices $\mathcal{A}$ and $\mathcal{B}$ assume the forms $\sqrt{\alpha} I$ and $\sqrt{\beta} I$, respectively, where $\alpha$ and $\beta$ are scalars and $I$ denotes identity.
In the scalar BEKK model, covariance stationarity is guaranteed when $\alpha + \beta < 1$. Accordingly, in each simulation we draw $\alpha \sim \mathcal{U}(0.05, 0.20)$ and $\beta \sim \mathcal{U}(0.70, 0.95)$. We then verify the stationarity condition and, if $\alpha + \beta \ge 1$, we redraw the parameters until $\alpha + \beta < 1$ holds.
In the full BEKK, the covariance stationarity condition depends on the eigenvalues:
$$ P_k\,(A \otimes A)D_k + P_k\,(B \otimes B)D_k $$
which are required to be strictly below unity in modulus; in the previous equation $D_k$ represents the duplication matrix of size $k$ and $P_k$ its generalized inverse.\footnote{The duplication matrix $D_k$ satisfies $\mathit{vec}\left(M\right)=D_k \mathit{vech}\left(M\right)$ for a $k-$dimensional square symmetric matrix $M$.}
We considered specific designs for the $\mathcal{B}$ matrix. In particular, we partition $\mathcal{B}$ into $3\times 3$ blocks and, for each diagonal block, we set the entries as follows: the main-diagonal elements are drawn from a uniform distribution on $[0.70,0.95]$, whereas the off-diagonal elements are drawn from a uniform distribution on $[0,0.10]$. In contrast, $\mathcal{A}$ is generated without imposing any block structure, and all of its entries are drawn from a uniform distribution on $[0,0.10]$. After generating $\mathcal{A}$ and $\mathcal{B}$, we check the stationarity conditions; if they are not satisfied, the parameter draw is discarded and the matrices are resampled until the conditions hold. This procedure is repeated independently for each simulation.
For both scalar and full BEKK models, the matrix $\mathcal{C}$ is generated as a lower triangular matrix with random entries, constructed to ensure that the product $\mathcal{C}\mathcal{C}'$ maintains full rank and has a positive diagonal.
The remaining generators are based on the DCC-GARCH framework, as given in (ref) and (ref). Similarly to the BEKK case, we randomly draw parameters. In each simulation, the parameters governing the correlation dynamics are drawn from uniform distributions, with $\theta_1 \sim \mathcal{U}(0.05,0.30)$ and $\theta_2 \sim \mathcal{U}(0.70,0.85)$. To enforce stationarity, we discard the draw if the condition $\theta_1 + \theta_2 < 1$ is not satisfied and resample the parameters until it holds. Moreover, the elements of $\boldsymbol\Gamma$ are generated as follows: first, we sample a matrix $A$ of random Normal numbers with mean $-0.15$ and standard deviation $0.6$; then we compute $Q=A^{\prime}A$, and normalize it so that it has unit elements on the main diagonal; finally, we set $\boldsymbol\Gamma=0.5\left(Q+Q^{\prime}\right)$ and retain the simulated matrix only if the smallest eigenvalue is larger than $1e-10$. The initial generation of random numbers ensures that the average correlation level is slightly higher than $0.4$.
The difference between the standard DCC-GARCH and the variant with interactions lies in the specification of individual dynamics. In the simpler DCC model, the evolution of variances follows (ref), with the DCC-GARCH coefficients $\alpha_i$ and $\beta_i$ sampled from the distributions $\alpha_i \sim U(0.05, 0.15)$ and $\beta_i \sim U(0.7, 0.85)$, with $\beta_j=0$ for $j\neq i$. In the model incorporating interactions, we represent the collection of variances in a matrix form as $$ \bm{\sigma}^2_t = \boldsymbol{\nu} + A\,r^2_{t-1} + B\,\bm{\sigma}^2_t, $$ where $\boldsymbol{\nu}$ is the vector of intercepts, $\bm{\sigma}^2_t$ is the vector of variances, $B$ is diagonal and $A$ is unrestricted, permitting nonzero off-diagonal elements. The diagonal entries for $A$ and $B$ are selected from uniform distributions in $[0, 0.2]$ and $[0.7, 0.85]$, respectively, while the off-diagonal elements of $A$ are drawn from $U(0, 0.02)$. The coefficients are generated repeatedly until the eigenvalues of $A+B$ are all within the unit circle, thus guaranteeing covariance stationarity as discussed in LingMcAleer2003.
For the simulations with $24$ assets, we adopt fixed coefficient designs. For models such as the full BEKK, drawing fully random parameter matrices that (i) satisfy covariance stationarity and (ii) yield reasonable parameter values (e.g., larger and positive diagonal entries) would typically require repeated draws for each simulation, substantially increasing the computational burden. We therefore fix the parameters of the data-generating processes.
Specifically, in the scalar BEKK we set $\alpha=0.15$ and $\beta=0.80$. In the full BEKK, the coefficient matrices $\mathcal{A}$ and $\mathcal{B}$ are set to \[ \mathcal{B}= I_{8}\otimes
,\qquad \mathcal{A}=
. \]
For the DCC--GARCH model, we set $\theta_1=0.15$ and $\theta_2=0.80$, and we fix the univariate GARCH parameters to $\alpha_i=0.15$ and $\beta_i=0.80$ for all $i$. In the model incorporating interactions, we choose $\theta_1$ and $\theta_2$ analogously, and we set $B$ accordingly. Conversely, the matrix $A$ is specified with $0.08$ on the main diagonal and $0.05$ elsewhere.
Given the parameters, we simulate returns sequences from the various data generating processes, also storing the true conditional covariance and correlation matrices. Subsequently, on the returns simulated from each DGP, we estimate three multivariate models, the Scalar BEKK, a standard DCC-GARCH, and the extended DCC (EDCC); the last model is not estimated in all cases. In addition, we aggregate the multivariate returns into a portfolio, either using equal weights, $1/n$, where $n$ denotes the total number of assets ($8$ or $24$), or with randomly generated weights (ensuring that their sum equals $1$). On the portfolio returns, we estimate a univariate GARCH model.
For each scenario, $100+T+250$ observations are simulated, the first $100$ are then discarded to avoid dependence on starting values, $T$ are employed for model estimation and the final $250$ reserved for out-of-sample evaluation. All forecasts are one-step ahead, meaning that $\Sigma_{t+1}$ is estimated using the observed return $r_t$. For the portfolios with 24 assets we set $T=1000$, while for the 9-asset scenario we also consider the alternative sample sizes of $T=500$ and $T=2000$. For all simulation designs, we performed $500$ experiments.
In this subsection, we analyze the simulation results. We first focus on univariate versus multivariate models, then highlight the benefits of forecast reconciliation, and later discuss the impact of a noisy proxy. We stress that we are not contrasting fitted DCC to fitted BEKK or EDCC but rather contrasting the portfolio variance forecast from a univariate GARCH to those of a MGARCH model and the forecast reconciliation based on the given MGARCH.
We start by contrasting the portfolio variance forecasts made following a bottom-up approach (bu), that is, estimating a MGARCH model, with those obtained by directly working on the portfolio returns, that is, with a univariate GARCH model (base). At first we consider a correctly specified MGARCH model and look at the relative accuracy indexes setting the univariate model as the reference (with a relative index thus equal to $1$): see Table (ref), columns labeled base and bu; full results including the value of the accuracy indexes are reported in the Online Appendix.
Notably, if we estimate a correctly specified MGARCH model, the prediction of portfolio risk is closer to the true value than the one obtained from a univariate GARCH model (miss-specified) fit on portfolio returns, that is, we observe a value lower than $1$ for the bu case. Note that, as reported in the Appendix, the accuracy measures decrease with the sample size in all cases, as expected. The pattern observed on relative accuracy, lower bu values, does not depend on the portfolio weights (equal or randomly generated), on the sample size, and on the accuracy measure. We observe that, for the Scalar-BEKK DGP, the preference for the bottom-up approach is attenuated as the sample size increases (the bu values increase) and is, overall, less pronounced than in the DCC-GARCH case. We interpret this as a consequence of aggregating returns generated under a Scalar-BEKK model with Gaussian innovations. In fact, if returns follow $\boldsymbol r_t \vert \mathcal{I}_{t-1}\sim \mathcal{N}\left(\boldsymbol 0, \bm{\Sigma}_t\right)$, with $\mathcal{I}_{t-1}$ being the time $t-1$ information set, and we set the portfolio weights to be time-invariant, $\bm{\omega}$ (and summing up to one), we have
with $\overline{\boldsymbol r}_{t-1}$ being the returns of the portfolio (a weighted average). The preference of the MGARCH specification for shorter samples might be linked to the larger amount of information used for parameters estimation ($N$ series against $1$) an effect that tends to disappear asymptotically, as suggested by the closer performances for $T=2000$ with random weights. A similar aggregation result does not hold for the DCC-GARCH model due to the heterogeneity in the volatility dynamic and the standardization required to obtain the dynamic correlations. In line with this pattern, for increasing sample size, the performance of the correctly specified MGARCH improves, with bu values decreasing with $T$. Overall, these first outcomes are somewhat expected.
Results start to be more interesting if we consider a misspecified model, still focusing only on DGPs without any form of interdependence; see Table (ref), again columns labeled base and bu. In terms of levels of accuracy indexes, they decrease with increasing sample size (as expected); see the Appendix. Moving to relative indicators, if we estimate a Scalar BEKK on series generated from a DCC-GARCH, and focus on small sample sizes, the wrongly specified model is providing superior forecasts compared to a univariate GARCH fitted on portfolio returns. However, as the sample size increases, the base and bu approaches tend to converge, and then bu starts deteriorating with respect to base. This behaviour does not depend on portfolio weights. The outcome is different if the DGP is a Scalar BEKK and the fitted model is a DCC-GARCH: the univariate model on portfolio returns is inferior to the misspecified MGARCH irrespective of the sample size (and again this does not depend on the portfolio weights). We link such evidence, again, to the aggregation of Scalar BEKK into a univariate GARCH on portfolio returns. In this case, the DCC-GARCH is clearly miss-specified but its heterogeneity helps in capturing the differences in unconditional variances (as driven by the differences across elements in the intercept).
Extending the cross-sectional dimension from $N=9$ to $N=24$ produces some notable shifts in the simulations. Under correctly specified estimation, when the DGP is DCC-GARCH, the univariate baseline now outperforms the correctly specified multivariate model—reversing the $N=9$ ranking. By contrast, when the DGP is Scalar BEKK, the correctly specified multivariate model continues to dominate the univariate benchmark, in line with the $N=9$ evidence.
Under misspecification, dimensionality also reshuffles the cross-cases. If the DGP is Scalar BEKK but the estimation is performed with DCC-GARCH, the multivariate specification performs better at $N=24$ as in the $N=9$ case. Conversely, if the DGP is DCC-GARCH and the multivariate estimator is a Scalar BEKK, the univariate baseline is now preferred. Overall, moving from nine to twenty-four assets strengthens the bias–variance trade-off and aggregation effects, eroding the advantage of correctly specified high-dimensional DCC in one case while preserving the edge of Scalar BEKK in the other. Results referring to the $N=24$ case are available in the Online Appendix.
Consider now the two DGPs that include interdependence, namely, the EDCC-GARCH and the BEKK models. Even in this case, we estimate the two specifications without interdependence, the DCC-GARCH and the Scalar BEKK, and we contrast the portfolio variance forecast obtained from a univariate GARCH with that of the misspecified MGARCH models; see Tables (ref) and (ref), still focusing only on columns base and bu. For completeness, we also estimate the correctly specified model under the EDCC-GARCH DGP (the fully parameterized BEKK estimation remains computationally unfeasible). For the BEKK DGP the univariate model is providing superior forecast abilities compared to either the Scalar BEKK and the DCC-GARCH. Moreover, the performance of MGARCH models worsens, in relative terms, with increasing sample size. This is valid irrespective of the loss function considered. On the other hand, for the EDCC-GARCH DGP, the use of a misspecified DCC-GARCH provides better forecasts than a univariate model on portfolio returns. This is in striking contrast with the result of the fitted Scalar BEKK (on the EDCC-GARCH DGP) whose forecasts are clearly inferior to those of the univariate model for sample size above 1000. We link such a finding to the relatively limited size of the coefficients of other asset shocks in the EDCC-GARCH DGP. The correctly specified model clearly favors the bottom-up approach compared to the univariate baseline. All results are equivalent for equally weighted portfolios and randomly generated portfolios. See the Online Appendix for detailed results for both the $N=9$ and the $N=24$ cases.
Re-running the simulations with $N=24$ assets yields patterns broadly consistent with the $N=9$ case when the DGP is a BEKK: the baseline univariate forecast (base) remains preferable to the bottom-up aggregation (bu), with the gap especially pronounced when the estimated multivariate model is a Scalar BEKK. By contrast, under the EDCC–GARCH DGP, moving to $N=24$ makes the univariate estimate outperform the bottom-up approach both when the multivariate estimator is Scalar BEKK and when it is a standard DCC–GARCH. Overall, higher dimensionality amplifies estimation risk and misspecification costs in parsimonious multivariate structures, keeping the univariate baseline a particularly competitive benchmark at $N=24$.
Tables from (ref) to (ref) contain the results of the three forecast reconciliation approaches that we consider. From the estimation of the correctly specified MGARCH model without interdependence (Table (ref)), we have a first interesting result: in terms of relative forecast accuracy indexes, forecast reconciliation always improves with respect to the bottom-up case, that is, the forecast from the correctly specified model. In addition, differences across forecast reconciliation methods are minimal, and not affected by the portfolio weights (fixed vs. random). Furthermore, in the DCC-GARCH case, we observe that the distance between the forecast reconciliation cases and the bu tends to reduce with increasing sample size, an effect not present for the Scalar BEKK case; again, we link this to the aggregation results available for the latter model. The evidence in favor of forecast reconciliation methods is consistent with the literature Wickramasuriya2019, showing that reconciliation enforces aggregation constraints and generally improves forecast accuracy, with differences in out-of-sample performance relative to the base forecasts reflecting mainly estimation error effects.
If we estimate a misspecified model (without interdependence) when the DGP does not include interdependence, the preference for forecast reconciliation methods is cleare and present under both the Scalar BEKK and DCC-GARCH data generating processes. In both cases, the preference for reconciliation methods slightly worsens with the increase in the sample size and, interesting, remains better than the base case when a misspecified Scalar BEKK is estimated on a DCC-GARCH DGP. Similarly to the estimation of a correctly specified model, the three forecast reconciliation approaches are very close to each other.
{Moving to the case of DGPs with interdependence, the results differ according to the data generating model. If we simulate data from an EDCC-GARCH, forecast reconciliation generally improves the bottom-up approach, irrespective of the fitted model and of the portfolio weights; only in the case of a fitted Scalar BEKK and large sample size, reconciliation becomes closer to the univariate case (bu). Moreover, as in the previous cases, reconciliation methods are similar, with no clear preference for one of the approaches. Differently, if we simulate data from a BEKK, forecast reconciliation improves under MSE and MAE losses, with a gain that decreases slightly with sample size. However, under the QLIKE loss when estimating a Scalar BEKK or an EDCC, the forecast reconciliation seems to not improve over the use of a univariate model, even though, we must admit, the values of the loss functions are really close (bu vs. forecast reconciliation).
We consider now the case where we estimate a correctly specified model in the presence of interdependence, namely EDCC-GARCH.\footnote{We highlight that the BEKK model cannot be considered due to its large number of parameters, more than $200$ with $N=9$.} The bu approach dominates on the univariate one, as in other correctly specified models (with preference becoming more clear with increasing $T$). Forecast reconciliation improves prediction accuracy, with a decreasing benefit for increasing sample size, and the three methods are extremely close to each other.
For $N=24$, the findings are largely consistent with the $N=9$ case. When the DGP does not present interdependence, even if the estimated multivariate model is correctly specified, forecast reconciliation still improves performance, and differences between reconciliation schemes remain small. With $N=24$ this preference is still strong, although slightly less pronounced than for $N=9$. Under interdependence, if the DGP is EDCC-GARCH and the multivariate estimator is DCC-GARCH, reconciliation yields clear gains; however, when the estimator is Scalar BEKK, the picture changes with respect to $N=9$: under the QLIKE loss, the univariate model is marginally preferred, even relative to reconciled forecasts. Under the Full BEKK DGP, the results for $N=9$ are confirmed: forecast reconciliation improves performance, with a slight advantage for the $shr_A$ method, while under the QLIKE loss, there is no significant gain when the multivariate estimation relies on a Scalar BEKK specification (the base method performs comparable to reconciled forecasts). Detailed results are available in the Online Appendix.
In the previous sections we contrasted the prediction of the portfolio variance to the true value of the portfolio variance. However, the latter is not observed and is usually replaced by a proxy in empirical analyzes. The early literature on variance forecasting has focused on the use of squared observed returns to proxy for the latent conditional variance. Similarly, when multiple assets are analyzed, the returns cross-product proxies for the conditional covariance. Since early 2000's with the availability of high frequency data, realized variances and covariances, replaced the squared and cross-product of returns, being more precise proxies of the returns conditional variances and covariances. To assess the effect of using a proxy in evaluating the benefits of forecast reconciliation, we rely on the same simulation settings considered above. While the returns cross-product is readily available from our simulation, recovering high frequency data coherent with the existence of a given GARCH-type dynamic on the daily returns is complex and challenging. We chose a simplified approach, mimicking the reduction in the noise provided by the realized covariance estimators using a contaminated proxy. In detail, we generate a noisy proxy of the true covariance by contaminating the true covariance with the returns cross-product. Therefore, we define the proxy as follows:
with $\delta \in \left(0,1\right]$, and $\bm{\Sigma}_t$ the true covariance (obtained from the simulations), $\boldsymbol r_t$ the simulated daily returns. When $\delta=1$ we are setting the proxy to the cross-product of returns, the most noisy proxy. On the other hand, with $\delta <1$ we are assuming the noise in the proxy is smaller, thus mimicking the improvement associated with the use of high frequency data. In the following, we set $\delta \in \left\lbrace 0.25, 0.5, 0.75, 1\right\rbrace$.
Given that we have four different $\delta$ values for each tuple of DGP, fitted model, forecast method and sample size, we graphically represent the average relative MSE; see Figures (ref) to (ref). In all plots, the baseline case is the estimation of the portfolio variance using a univariate GARCH model; notably, this approach, for a given data generating process is not changing across fitted Multivariate GARCH models. Therefore, if it represents the reference forecasting approach, the relative indicators can be contrasted both across forecasting approaches and over models.
In Figures (ref) and (ref) the DGPs are without interdependence, and the patterns share some similarity: if the noise in the covariance proxy is increasing, the difference between adopted modeling choices and forecast methods (base, bu and reconciliation) tends to disappear for increasing sample sizes. In summary, if one believes that the data are not characterized by variance spillovers, then whatever the model and the forecast approach is used, the results will be very close to those of the correctly specified model and close to a univariate GARCH (all relative indicators are close to 1). Moreover, when the noise is smaller, forecast reconciliation benefits are limited, smaller than those observed under the true covariance, and might also disappear, such as in the case of Scalar BEKK. This shows that the improvement of forecast reconciliation might also be hidden by even a limited noise in the covariance proxy.
When moving to the DGPs with interdependence, the results differ. In the case of the EDCC, Figure (ref), when fitting the Scalar BEKK and when the noise in the proxy and the sample size both increases, the results are very close to a baseline univariate GARCH and in some cases even worse, with no effect of reconciliation. Unlikely, if a DCC-type model is adopted for portfolio variance forecasting, we do observe a (limited) reduction in the relative MSE compared to the baseline. Moreover, forecast reconciliation becomes comparable to the fit of a misspecified model. This suggests that the noise in the proxy also masks the possible presence of assets interdependence, making it difficult to identify the correct model (we remind that we can perform a comparison across models). In terms of forecasts precision, such a finding is clearly dependent on the strength of the relation across conditional variances. The values we considered in the simulations are in line with values recovered from the observed data. If the DGP is a Full BEKK, Figure (ref), the results show one relevant element, the multivariate models, the bu forecasts, are the worst even if the noise in the proxy is not large, and the reconciliation improves the forecast quality. However, if the noise in the proxy increases, the fitted models and the forecast approaches tend to converge. Again, this suggests that if we use a noisy covariance proxy the presence of interdependence in the data might not be detected and thus efficiently exploited in forecast construction; the final outcome seems a suggestion for the use of a simpler model (for instance a DCC) when in reality the data dynamic is much more intricate. Moreover, we confirm previous findings in the literature on the crucial role of using less noisy covariance proxies for model comparison in a forecasting framework.
In the case $N=24$ ($T=1000$ only), Figure (ref), results observed under the Scalar BEKK and BEKK data generating processes are confirmed: in the SBEKK DGP case, correctly specified and misspecified models are really close one to the other and reconciliation is not improving over either a univariate fit or a bottom-up approach; for the full BEKK DGP, the noise is making the bottom-up the worst approach (due to misspecification) and forecast reconciliation is not adding much to univariate models. For the DCC DGP, forecast reconciliation is improving, with a more effective outcome in the correctly specified model case. Differently, for the EDCC, miss-specification matters, and both DCC and SBEKK are very close to the univariate case and slightly better than the bottom up, in particular for increasing noise in the proxy. This further confirms the importance of using a less noisy proxy for the covariance matrix when contrasting forecasting models and approaches, the potential advantage of forecast reconciliation, and the risk of not recognizing the existence and relevance of interdependence in multivariate GARCH models.
We close with a note regarding similar plots to those here reported and included in the Appendix for QLIKE and MAE. Interestingly, even if MAE is an inconsistent loss function for model ranking (see Laurentetal2013), the results it provides are in line with those of both MSE. QLIKE provides results aligned with the other two indicators only in the case $N=9$. When the cross-sectional dimension increases, QLIKE shows larger preference for the univariate approach.
We further evaluate differences in predictive accuracy using pairwise Diebold--Mariano tests. The Online Appendix reports summaries in Figures A2--A4, while the corresponding comparisons obtained using noisy covariance proxies are shown in Figures A5--A7.
The figures highlights meaningful differences across data-generating processes. When the DGP follows the Full BEKK specification (BKF), the bottom-up approach (bu) tends to outperform the univariate portfolio-based forecast (base) in most pairwise comparisons. In contrast, under the DCC-GARCH generator (DCG) the opposite pattern typically emerges, with the base forecast more frequently preferred to bu. This advantage of the univariate approach becomes weaker when the multivariate model used for estimation differs from the true DGP, such as when Scalar BEKK (SCB) or DIG specifications are employed, although the base forecast generally remains competitive. Under the Scalar BEKK DGP, the relative ranking between base and \textit{bu} is less systematic and depends more strongly on the multivariate GARCH used. Across all DGPs, however, forecast reconciliation methods consistently match or outperform the best among the \textit{base} and \textit{bu} forecasts in pairwise comparisons. This result indicates that reconciliation effectively combines the information contained in the univariate and multivariate predictions.
When noisy proxies for the covariance matrix are used in the evaluation stage (Figures A5--A7), the differences between base and bu become less pronounced. Nevertheless, reconciliation methods remain competitive and often provide improvements relative to the individual forecasts, particularly when less stringent significance thresholds are considered.
To complement the pairwise comparisons, we apply the Model Confidence Set (MCS) procedure of Hansen2005 with detailed results reported in the Online Appendix.
When the true covariance matrix is used for evaluation (Tables A3--A5), forecast reconciliation methods are included in the MCS with very high frequency across most simulation designs. In many cases the inclusion frequency of the reconciliation approaches exceeds 90%, and often remains above 80%. This pattern holds across sample sizes and portfolio weighting schemes. Among the reconciliation strategies, the differences are generally small: the shrinkage-based reconciled forecasts ($shr$ and $shr_A$) typically display very similar inclusion frequencies, while $shr_B$ tends to enter the confidence set slightly less often, although it remains comparable in magnitude.
In contrast, the bottom-up forecast (bu) rarely belongs to the MCS in several simulation settings. This is particularly evident in designs based on the Full BEKK generator combined with DCC-type estimation, where the inclusion frequency of the bottom-up approach is close to zero in most cases. The univariate portfolio-based forecast (base) performs better than the bottom-up alternative and is often included in the MCS, although it typically remains below the reconciliation methods.
Differences across data-generating processes nevertheless emerge. Under the BEKK-based generator, reconciliation clearly dominates both base and bu. For DCC-based generators the gap between reconciliation and the univariate approach becomes smaller, and the base forecast often enters the confidence set with moderate frequency. Under the Scalar BEKK generator the relative ranking between base and bu becomes less systematic and depends more strongly on the multivariate model used for estimation.
The same qualitative conclusions hold when the number of assets is increased (Tables A11--A13). Forecast reconciliation remains the class of methods most frequently included in the MCS, confirming that combining univariate and multivariate information provides robust improvements in predictive accuracy.
When noisy covariance proxies are used for evaluation (Tables A7--A9), the evidence becomes less clear-cut. In these settings the inclusion frequencies of the different approaches become more similar, and both the univariate and multivariate forecasts appear more often in the confidence set. Nevertheless, reconciliation methods remain competitive and continue to belong to the MCS in a large fraction of the simulation designs.
We perform an empirical analysis using daily returns from 28 constituents of the Dow Jones Industrial Average index, which are continuously available from January 2, 2003, to December 31, 2024. Excluding weekends and bank holidays, the sample includes a total of 5,537 observations. For the same period, we also collect the daily Realized Covariance, based on 1-minute returns. All data is sourced from the Capire database and is freely available at \url{capire.stat.unipd.it}.
Following the approach used in the simulations, we estimate both multivariate models on the set of the 28 assets as well as a univariate model on an equally weighted portfolio. The multivariate models we consider are: the DCC-GARCH, with univariate GARCH(1,1) as marginals; the Extended DCC-GARCH; the Scalar BEKK. In the univariate approach also considers a GARCH(1,1) model.
Estimates are obtained using a rolling window of the most recent 1,500 observations. All the models are re-estimated on the first day of each month, and the estimated coefficients are then used to generate variance forecasts with a horizon up to one month (22 observations), starting from each day of the month. Forecasts are produced on a one-step-ahead basis (i.e., fixed parameters within the month, but moving the information set on a daily basis).
Tables (ref) to (ref) report forecast evaluation results for a single model, contrasting competing forecast construction approaches, namely, the univariate GARCH on equally weighted portfolio returns (baseline), the bottom-up of the multivariate GARCH models, and the three forecast reconciliation cases. The tables provide comparisons across different forecasting horizons, one day, one week (5 days) and one month (22 days).
When forecasts are obtained by fitting a DCC model, the results differ across loss functions. In fact, for QLIKE the baseline univariate approach has the lowest loss function value and is always included in the confidence set; moreover, the base approach is always statistically superior to the bottom-up case (see the Diebold-Mariano test outcomes). This evidence is consistent across forecast horizons. Differently, for MAE and MSE the use of forecast reconciliation provides better results, even though the model confidence set always includes the univariate-based forecasts. The bottom-up case is part of the confidence set only for the daily horizon. For the Scalar-BEKK model, Table (ref), the results are more favorable to the univariate model (except for the MSE and one day horizon, where $shr_B$ is preferred). Notably, the MCS never includes the bottom-up case (i.e., the direct use of the multivariate GARCH). Finally, when fitting the EDCC, the only model partially accounting for interdependence, while the loss function values suggest a preference for forecast reconciliation ($shr$) in most cases (QLIKE is tilted toward the univariate approach for weekly and monthly horizons, while MAE suggest bottom-up at the weekly horizon), the MCS includes most forecast approaches, and the Diebold-Mariano tests suggest, with larger frequencies, that the bottom-up approach is inferior to the others. These results show some heterogeneity, and are focuses on a single model, across forecasting approaches. Taking a more general viewpoint, and applying Model Confidence Set across models and forecasts, using a 20% threshold for exclusion from the set, a more clear picture emerges, see Table (ref). The only model which is almost always included in the confidence set is the EDCC (it is excluded under the QLIKE function for weekly and monthly horizons), while the most excluded model is the Scalar BEKK. Forecast reconciliation approaches when they are excluded from the confidence set, they have larger p-values compared to the \textit{bottom-up} cases. Under QLIKE the univariate approach seems the best at the monthly horizon, while for other horizons and for the other loss functions, the confidence set always include the EDCC under forecast reconciliation ($shr$ case). In summary, the use of forecast reconciliation in general improves the forecast ability of MGARCH models, in particular, in situations where we cannot exclude the existence of interdependence and the proxy of the covariance matrix is not too noisy.
Through simulations, we show that forecast reconciliation approaches might improve forecast accuracy of portfolio risk when the underlying model is a Multivariate GARCH, possibly including interdependence. Our results also shed light on the role exerted by noisy proxies in forecast evaluation, highlighting the fundamental impact of the noise that could mask the possible presence of interdependence across conditional variances and covariances. An empirical example illustrates the use of reconciliation with real data and a proxy based on realized covariances.
Our conclusions provide a first look at the role of forecast reconciliation in risk management settings, where the focus moves toward quantiles and conditional expectations. Although current research is moving in that direction, we believe that this topic represents an interesting area for future developments.