EconBase
← Back to paper

Forecasting realized covariances using HAR-type models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

59,038 characters · 10 sections · 75 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Forecasting realized covariances using HAR-type models

abstractWe investigate methods for forecasting multivariate realized covariances matrices applied to a set of 30 assets that were included in the DJ30 index at some point, including two novel methods that use existing (univariate) log of realized variance models that account for attenuation bias and time-varying parameters. We consider the implications of some modeling choices within the class of heterogeneous autoregressive models. The following are our key findings. First, modeling the logs of the marginal volatilities is strongly preferred over direct modeling of marginal volatility. Thus, our proposed model that accounts for attenuation bias (for the log-response) provides superior one-step-ahead forecasts over existing multivariate realized covariance approaches. Second, accounting for measurement errors in marginal realized variances generally improves multivariate forecasting performance, but to a lesser degree than previously found in the literature. Third, time-varying parameter models based on state-space models perform almost equally well. Fourth, statistical and economic criteria for comparing the forecasting performance lead to some differences in the models' rankings, which can partially be explained by the turbulent post-pandemic data in our out-of-sample validation dataset using sub-sample analyses. \\ \\ Keywords: State space model, Heterogeneous autoregressive, Realized measures, Volatility forecasting.

Introduction

Accurate modeling and forecasting of volatility are pivotal in risk management, pricing derivatives, and guiding investment decisions in financial markets. The stock market's inherent association with the national economy underscores the importance of volatility, evident in major financial events like the stock market crash of 1987, the 2008 Lehman Brothers bankruptcy, and subsequent global financial crises. Volatility modeling, particularly with high-frequency data, has evolved over the years, moving from early models like generalized autoregressive conditional heteroskedasticity (GARCH) and stochastic volatility to more sophisticated approaches based on realized variance measures. Intraday volatility, exemplified by events such as the crash in 2010 and trading errors by Knight Capital in 2012, underscores regulators' need to thoroughly investigate the relationship between high-frequency trading and intraday fluctuations to mitigate potential risks. The ability to accurately forecast volatility becomes increasingly crucial in today's financial markets, particularly during instability as triggered by, e.g., Brexit and the COVID-19 pandemic. Financial market volatility quantifies the variability in asset returns and constitutes a fundamental component in pivotal aspects of financial decision-making, including trading of volatility, asset allocation, risk management, and derivatives and asset pricing; see, e.g., caporin2017chasing, harvey2023score and kikuchi2018minimum. Early attempts at modeling financial market volatility, dating back to the 1960s, employed temporal models such as those proposed by fama1965behavior. These models, often based on autoregressive conditional heteroskedasticity (GARCH) frameworks engle1982autoregressive, bollerslev1986, bollerslev1994arch, demonstrated robust in-sample parameter estimation but exhibited limitations in predicting daily squared returns (poon2003forecasting).

For a price process lacking jumps, volatility quantifies the magnitude of price fluctuations and is defined as the square root of integrated volatility (IV). Historically, researchers relied on low-frequency data for IV estimation, such as daily squared returns, despite the resulting high variance of such estimates. With the advent of high-frequency (intraday) data, andersen1998answering and andersen1998deutsche introduced realized variance (RV) as an estimator of latent variance. This model-free metric is calculated by aggregating the squared price changes observed within each trading day from high-frequency data, offering a significant improvement in volatility estimation and prediction accuracy compared to GARCH and stochastic volatility models (andersen2003modeling). Furthermore, RV provides an ex-post estimate of asset return variance and is a consistent estimator of IV under the assumption of a diffusion process for the logarithmic asset price kinnebrock2008note.

RV is advantageous for high-frequency data but faces challenges due to microstructure noise, as highlighted by studies like figlewski1997forecasting and andersen2001distribution. This noise, stemming from market data imperfections, introduces biases in RV estimates. To address this, researchers have developed microstructure noise-robust IV estimators to enhance accuracy in high-frequency volatility modeling. However, eliminating noise may be impractical. Alternatively, using lower-frequency data reduces the impact of noise. Still, it raises the variance of the RV estimator, posing a key challenge in achieving the optimal balance between noise reduction and variance control in high-frequency data-based volatility estimation. bandi2008microstructure, patton2015good, and liu2015does shed light on a trade-off between the effect of the estimator and microstructure noise. They found that using 5-minute data for estimating RV is the recommendable frequency. Furthermore, another commonly used method is the realized kernel suggested by barndorff2008designing, which is popular for reducing the impact of microstructure noise when dealing with high-frequency data. The realized kernel approach is particularly suitable for estimating the realized covariance matrix, ensuring positive definiteness; see barndorff2011multivariate. This is the method we rely on for realized covariance estimation in this paper, and we use it in combination with 1-minute returns due to its ability to mitigate microstructure noise.

The use of high-frequency measures results in more accurate variance forecasts compared to those derived from GARCH or stochastic volatility (SV) models tailored for daily returns, as demonstrated by francq2019garch and koopman2005forecasting. Moreover, enhancing GARCH and SV models with RV measures derived from high-frequency data leads to enhanced model fit and forecasting accuracy. This enhancement is evidenced in studies such as engle2006multiple and hansen2012realized for observation-driven models, as well as dobrev2010information, and takahashi2009estimating for stochastic volatility models. However, an approach for forecasting RV directly has become extremely popular, namely the heterogeneous autoregressive (HAR) model by corsi2009simple. The HAR model is widely adopted due to its simplicity and excellent predictive performance, making it a prominent choice in the literature (e.g., busch2011role, souvcek2013realized,cubadda2017vector,liu2023trading). HAR models can effectively replicate volatility persistence by capturing aggregated volatility across various interval sizes. Since the HAR model can be represented as a constrained autoregressive model of the order 20 model, parameter estimation using ordinary least squares (OLS) is straightforward. In the context of multivariate volatility, a realized covariance (RCov) matrix is formed from the returns of financial assets. chiriac2011modelling expanded the univariate HAR to the multivariate domain, namely multivariate HAR, to model and predict the RCov. Like its univariate counterpart, the multivariate HAR model boasts a straightforward and easily implementable structure. An alternative approach for modeling realized covariance matrices is the Conditional Autoregressive Wishart by Golosnoy2012.

huang2019volatility observe that volatility risk fluctuates over time and influences asset pricing. However, despite this, models applied for realized volatility have primarily remained static and have not explored the dynamics of variance, skewness, or the potential evolution of information flow pace. This limitation became evident during the 2008 financial crisis when volatility experienced a sudden and significant increase. Models with constant parameters proved inadequate in capturing this abrupt rise in volatility, potentially harming their forecasting accuracy. The widespread use of constant-parameter models is attributed to their simplicity in estimation and standard inference under accurate specification. However, this convenience comes at the cost of neglecting crucial structural changes in the volatility process, particularly over extended sample periods. Early models seeking to move away from the assumption of constant parameters include the nonparametric and parametric model in fan2003nonlinear and cai2000efficient, change-point ARCH models in davis2006structural, regime-switching models in bollerslev1994arch and ang2002international, and the mixture GARCH model developed by haas2009asymmetric. More recently, engle2013stock have proposed multiplicative component structures, while amado2013modelling explore both additive and multiplicative components. In a multivariate context, bauwens2013multivariate and bauwens2016forecasting introduce component structures for capturing lower-frequency movements.

For modeling and forecasting RV using HAR models, in addition to the arguments above for time-variation in model parameters, bollerslev2016exploiting point out that, after all, RV is only an estimator for the underlying volatility and therefore is characterized by measurement errors. Consequently, the OLS estimator suffers from an attenuation bias. This is solved by conditioning the coefficients on realized quarticity, an estimator for RV's variance, which results in the so-called HARQ model having superior model fit and forecasting performance. The multivariate extension for predicting realized covariance matrices using a variety of approaches is studied in bollerslev2018modeling. Adapting the attenuation bias approach to the natural logarithm of univariate RV is suggested in WLWH20. We propose to use this to model the log of the marginal RV in the multivariate realized covariance approach in bollerslev2018modeling and show that this improves the resulting forecast. An alternative approach to incorporate time-varying parameters in the univariate setting is the state-space approach in bekierman2018forecasting, modelling the persistence in the HAR models as a latent autoregressive process of order one.

To summarize our contributions, we compare and extend forecasting models for realized covariance. We build on the recommended approach from bollerslev2018modeling that breaks the problem of forecasting realized covariances into separate models for variances and correlations. We extend this by comparing models for RV vs. $log RV$, models with and without controlling for attenuation bias, and state-space HAR models as in bekierman2018forecasting. The models are compared for predicting daily covariances of 30 assets over approximately 12 years using statistical and economic measures for forecast comparison.

This paper is set up as follows. In Section (ref), we introduce the models and present the theoretical framework. Section (ref) evaluates the benchmark models based on their fit and forecasting capabilities, while Section (ref) provides the conclusion.

Methodology

Models for univariate realized volatility

In this work, we consider an asset with a price process, $P_{t}$, governed by a stochastic differential equation that captures both long-term trends and random fluctuations. We focus on the daily integrated variance, $I V_{t}$ calculated through the integral of squared instantaneous volatility $\sigma^{2}_{s}$ as follows

equation[equation omitted — 58 chars of source]

Assume that there are $M$ intraday returns in a trading day $t$ and denote the $j$th intra-day return by $r_{t,j}$. Then the Realized Volatility for day $t$ is defined as

equation[equation omitted — 54 chars of source]

Realized volatility $RV$ is inherently heterogeneous, encompassing multiple volatility components. It has been observed that volatility over extended time intervals tends to exert a greater influence on short-term volatility than the reverse. This can be understood intuitively: intraday speculators are influenced by long-term volatility since it shapes the expected trend, whereas long-term traders are less impacted by the fluctuations caused by intraday activities. Recognizing these characteristics of realized volatility, corsi2009simple introduced the heterogeneous autoregressive (HAR) model. This model, which is autoregressive in nature, captures the volatility dynamics by incorporating three distinct components. The daily realized volatility $RV_{t-1}$, the weekly mean realized volatility $RV_{t-1:t-5}$ and the monthly mean realized volatility $RV_{t-1:t-20}$\footnote{More precisely, $RV_{t-1:t-j}=\frac{1}{j}\sum_{i=1}^{j} RV_{t-i}$}. The HAR model has the following form:

equation[equation omitted — 112 chars of source]

with $\epsilon_t$ a mean zero error term. The same model may be specified for $\log RV_t$, which we denote as the HARL model.

bollerslev2016exploiting note that the HAR model suffers from an attenuation bias. Their key insight is that realized volatility is a noisy measurement of the integrated volatility, and that the variance of the measurement is not homogeneous in time as assumed in early work (see, e.g., koopman2012analysis) in the standard HAR model in corsi2009simple, but depends on the integrated quarticity $IQ_t$. bollerslev2016exploiting account for the heterogeneity by allowing time-varying coefficients by including a covariate $RQ_t$, which is the realized quarticity at time $t$. The authors demonstrate that accounting for time-varying coefficients is of less importance for the weekly and monthly lags, and thus propose the model

align[align omitted — 152 chars of source]

which is known as the HARQ model. Realized quarticity is defined as $RQ_t=\frac{M}{3}\sum_{i=1}^M r_{t,i}^4$.\footnote{It is important to note, however, that not only is $I V_{t-1}$ measured with uncertainty, but $R Q_{t-1}$ also acts as a noisy estimator for $I Q_{t-1}$.} The underlying concept of this model is that when the variance of the measurement error is substantial, resulting in a large $R Q_{t-1}^{1 / 2}$, the model exhibits reduced persistence for $\gamma<0$. WLWH20 extend this idea to a model for the natural logarithm of $RV$, motivated by the common approach to model $\log RV_t$ instead of $RV_t$, as

align[align omitted — 200 chars of source]

This model, termed HARQL, has been shown to perform very well empirically.

bekierman2018forecasting suggested a different approach to introduce a time-varying component to the autoregressive parameter by relying on the state space model

align[align omitted — 125 chars of source]

where the error term $\epsilon_t$ is assumed to be iid normally distributed. The state variable is driven by a Markovian process with Gaussian noise

equation[equation omitted — 120 chars of source]

The idea behind their model is that the time-varying coefficient $\lambda_t$ may capture time-variation due to the measurement error and other variations. The model can directly be applied to $\log RV_t$ by replacing (ref) with

align[align omitted — 148 chars of source]

Note that this model for $\log \left(\mathrm{RV}_{t}\right)$ instead of $\mathrm{RV}_{t}$ has the advantage that the assumption $\epsilon_{t} \sim N\left(0, \sigma_{\varepsilon}^{2}\right)$ is more likely to hold. Forecasts for $RV_t$ use properties of the log-normal distribution; see bekierman2018forecasting. bekierman2018forecasting report much better forecasting performance for models based on $\log RV_t$. We term these state space models as HARS and HARSL, respectively. When predicting the realized variance (on the ordinary scale) for the log models, we use the standard approach that bias corrects $\exp(\widehat{\log RV_t})$ assuming the estimate is normally distributed.

The models above are univariate in nature, but they form the basis of the multivariate forecasting models realized covariance matrices that we introduce next.

Models for realized covariance

Now consider an $N$-dimensional price process $\bm P_t$ with spot covariance matrix $\Sigma(u)$ and integrated covariance matrix for day $t$ \[ \Sigma_t=\int_{t-1}^t \Sigma(u)du, \] which can, e.g., be estimated using the multivariate kernel estimator by barndorff2011multivariate. We denote the corresponding estimates by $S_t$ and let $s_t$ be the (half-vectorisation of the) realized covariance $S_t \in \mathbb{R}_{+}^{N^\star}$ of $N$ assets at time $t$, where $N^\star = N(N+1)/2$ (the number of unique elements in the covariance matrix). To extend the HAR model to a multivariate setting, chiriac2011modelling propose a parsimonious extension by modeling the vectorized covariance matrix, $s_t$, by

align[align omitted — 133 chars of source]

where $\varepsilon_t$ is a mean-zero error vector, $\bm \alpha_0$ is of dimension $N^\star$ and $\alpha_1, \alpha_2, \alpha_3$ are scalar parameters. bollerslev2018modeling note that this suffers from the attenuation bias in a similar way as in the univariate case and propose a vech HARQ model that directly extends the HARQ model above to the multivariate case by allowing $\alpha_1$ to be time-varying. The time variation is driven by the diagonal elements of the measurement error covariance matrix $\Pi_t$. Specifically, the multivariate HARQ model is given by

align[align omitted — 180 chars of source]

with the $N^\star$ dimensional vector $\pi_t=\sqrt{\mathrm{diag}(\Pi_t)}$ ($\circ$ denotes element-wise multiplication), which can be estimated straightforwardly from the data, $\boldsymbol{\iota}$ an $N^\star$ vector of ones and $\alpha_{1Q}$ a scalar parameter.

Another approach is to model the marginal variances and the cross-correlations separately, the approach we heavily rely on in this paper, by using the following decomposition of the realized covariance matrix

align[align omitted — 56 chars of source]

where $D_t$ is a diagonal matrix with standard deviations and $R_t$ is the correlation matrix. oh2016high propose the HAR-DRD model, where the individual variances are modeled via a univariate HAR model, and the correlation matrix is modeled analogously to (ref). We follow this approach and model the correlations parsimoniously through the following scalar HAR model

align*[align* omitted — 169 chars of source]

with $\bar{\bm r}= \frac{1}{T}\sum_{i=1}^{T}\bm r_t$ and the components of $\bm \epsilon_t$ are iid mean zero errors. oh2016high show that the predictions using this model formulation are valid correlation matrices if the estimates $\widehat{\gamma}_1, \widehat{\gamma}_2, \widehat{\gamma}_3 > 0$ and $\widehat{\gamma}_1 + \widehat{\gamma}_2 + \widehat{\gamma}_3 < 1$. The forecast is then produced by separately forecasting each part and combined to a realized covariance forecast using (ref). bollerslev2018modeling propose the HARQ-DRD model, which uses the univariate HARQ specification (ref) for the individual variances and the above model for realized correlations. bollerslev2018modeling point out that it is possible to model the correlation matrix using (ref), thus allowing for time-varying coefficients, but note that the heteroskedasticity in the measurement errors of the correlations tend to be somewhat limited. Therefore, constant parameters are recommended for modeling and forecasting the realized correlations.

We propose to build on the DRD decomposition of the covariance matrix in (ref) and replace the univariate HARQ model with models that have been shown to provide superior empirical performance in the univariate case. To be specific, $RV_t$ is modeled using the HARQL model (ref) as well as the state-space versions HARS in (ref) and HARSL in (ref).

Estimation

The estimation of the HAR and HARQ models, both based on $RV_t$ and $\log RV_t$, and for the correlation model is straightforwardly done by OLS. The state space model, on the other hand, requires maximum likelihood estimation using the Kalman filter. Consider, e.g., the HARSL model in (ref), which we can rewrite as

align[align omitted — 255 chars of source]

Setting $y_t=\log RV_t$, $\alpha = (\alpha_0,\;\ \alpha_1,\;\ \alpha_2)^\top$, $$x_t = (1, \;\ \log RV_{t-1}, \;\ \log RV_{t-5:t-1}, \;\ \log RV_{t-20:t-1} )^\top$$ and $f_t=\log RV_{t-1} $, we can simplify (ref)

align[align omitted — 163 chars of source]

The Kalman filtering is implemented using the dlm package in R (see, e.g., durbin2012time, petris2010r), which estimates state space models of the form

align[align omitted — 148 chars of source]

Note that $F_t = f_t$ is time-varying and that $G_t = \phi$ is static.

Forecast evaluation

The models described above deliver one-step-ahead predictions $\widehat{S}_t$ for the realized covariance matrix $S_t$. To evaluate the predictions, we use the Frobenius norm and the quasi-likelihood (Q-Like) loss function; see LRV13 for their motivation. The Frobenius norm is commonly used to measure the distance between two matrices,

align*[align* omitted — 100 chars of source]

The Q-Like measure is based on the negative of the log-likelihood of a multivariate normal distribution and is defined as,

align*[align* omitted — 78 chars of source]

We take the average losses over the entire out-of-sample period, and lower values are preferable.

Next, we consider the economic evaluation of the covariance predictions using some criteria suggested in the literature and nicely motivated and summarized in bollerslev2018modeling. Their economic evaluations are based on the use in the construction of global minimum variance (GMV) portfolios and portfolios designed to track the aggregate market. Based on the covariance forecast $\widehat{S}_t$ of the returns on the assets, to minimize the conditional volatility, the optimal portfolio allocation vector is

align*[align* omitted — 120 chars of source]

where $\boldsymbol{\iota}$ is a $n \times 1$ vector of ones. The variance of this portfolio should be smaller for better forecasts.

Let $r_t^{(j)}$ represent the return on asset $j$ in day $t$. The turnover from day $t$ to day $t+1$ is given by

align*[align* omitted — 98 chars of source]

With proportional transaction costs $cTO_t$ (with $c$ being, e.g., 0, 1% or 2%), the portfolio excess return net of transaction costs is

align*[align* omitted — 44 chars of source]

To assess how extreme the portfolio allocations are, we use the portfolio concentrations

align*[align* omitted — 57 chars of source]

and the total portfolio short positions,

align*[align* omitted — 79 chars of source]

Using a quadratic utility function, the economic value of the different models is determined by solving for $\Delta_\gamma$ in

align*[align* omitted — 99 chars of source]

where the utility of the investor with risk aversion $\gamma$ is assumed to be

align*[align* omitted — 92 chars of source]

$\Delta_\gamma$ is the return that an investor with risk aversion $\gamma$ would be willing to pay to switch from model $k$ to $l$.

Application

The application is based on high-frequency returns for a set of assets contained in the Dow Jones 30 index at some point in our sample period.\footnote{The data were obtained from EOD historical data via www.eodhd.com. The ticker symbols of the included assets can be found in Table (ref).} We use a multivariate realized kernel to estimate the daily return covariance matrix from intraday 1-minute returns from March 2008 until June 2024, resulting in 4051 daily observations of realized covariance matrices. We also computed the realized quarticities for the same period. Both quantities are computed using the highfrequency package in R highfrequency2022. Returns are in percentage form (that is, multiplied by 100) in all calculations unless otherwise stated.

The data have some outliers, which we treat as follows. First, we consider a (daily) realized covariance matrix an outlier if any of its elements is more than 20 standard deviations away from its mean. Only 36 of 4051 observations were classified as outliers. We also found 44 outliers detected for the realized quarticities. The union between these two sets of outliers is 53 out of 4051 for the 30 assets. For these 53 days, we replace the realized covariance matrix with the covariance matrix of the previous day. Note that we cannot simply replace the outlier element with its mean because this does not guarantee that the covariance matrix remains positive definite. We treat the realized quarticities in the same way.

The realized covariances are then modeled using the following 7 models, HARSL-DRD, HARS-DRD, multivariate HAR (M-HAR), univariate HARL, univariate HAR, HARQL, and HARQ, (only HARQL and HARQ use realized quarticities) outlined in Section (ref).

The models use up to 20 lags as regressors, and hence we lose the first 20 observations, leaving a total of 4051 observations. Table (ref) shows summary statistics of variances, quarticities and correlations for all $n=30$ assets.

{\tiny

table[table omitted — 6,278 chars of source]

}

Table (ref) reports the average parameter estimates over all 30 assets using all the available data. The table also shows the Frobenius norm and quasi-likelihood measures, which are losses indicating the quality of the fit. However, it is important to note that these are computed in-sample. The dramatically lower values for HARS indicate overfitting, which we examine later.

landscape{\tiny \begin{table}\caption{In-sample results} \begin{tabular}{lcccccccccc} \toprule {Models} &{$\overline{\alpha}_0$} & {$\overline{\alpha}_1$} & {$\overline{\alpha}_2$} & {$\overline{\alpha}_3$} &{$\overline{\phi}$} &{$\overline{\sigma}_\epsilon$} &{$\overline{\sigma}_\eta$} & {$\overline{\alpha}_{1,q}$} & {$\overline{L}^{\mathrm{Frobenius}}$} & {$\overline{L}^{\mathrm{Q-Like}}$} \\ \midrule M-HAR & 0.1593 & 0.1983 & 0.3884 & 0.2847 & - & - & - & - & 48.1023 & 42.4565 \\ & {-} & {-} & {-} & {-} & {-} & {-} & {-} & {-} & & \\ HAR & 0.4184 & 0.2147 & 0.3311 & 0.3125 & - & 4.9211 & - & - & 46.654 & 40.3348 \\ & {(0.2436)} & {(0.1098)} & {(0.1287)} & {(0.1202)} & {-} & {(2.4239)} & {-} & {-} & & \\ HARL & -0.1138 & 0.2762 & 0.2182 & 0.3701 & - & 0.7609 & - & - & 46.1667 & 40.2713 \\ & {(0.0443)} & {(0.0293)} & {(0.0388)} & {(0.0484)} & {-} & {(0.0303)} & {-} & {-} & & \\ HARQ & 0.2955 & 0.4245 & 0.2744 & 0.2446 & - & 4.8483 & - & -0.0006 & 46.3588 & 40.6368 \\ & {(0.251)} & {(0.1191)} & {(0.1294)} & {(0.1034)} & {-} & {(2.3991)} & {-} & {(0.0004)} & &\\ HARQL & -0.0300 & 0.5663 & 0.1939 & 0.3010 & - & 0.7421 & - & -0.0522 & 45.2730 & 40.2152 \\ & {(0.0576)} & {(0.0774)} & {(0.0440)} & {(0.0420)} & {-} & {(0.0280)} & {-} & {(0.0078)} & & \\ HARS & 0.4541 & 0.5463 & 0.1008 & 0.2588 & -0.0885 & 2.5564 & 1.0768 & - & 36.8527 & 33.7369 \\ & {(0.1996)} & {(0.1666)} & {(0.0616)} & {(0.0990)} & {(0.1549)} & {(1.7944)} & {(0.6525)} & {-} & & \\ HARSL & -0.0888 & 0.2656 & 0.1640 & 0.3239 & 0.9564 & 0.7399 & 0.0425 & - & 44.2440 & 39.4222 \\ & {(0.0740)} & {(0.0275)} & {(0.0332)} & {(0.0556)} & {(0.0450)} & {(0.0315)} & {(0.0250)} & {-} & & \\ \bottomrule \end{tabular}\\ \begin{flushleft} {Note: Parameter estimates and in-sample loss measures for all models using all the available data for the 30 assets. For the multivariate HAR (M-HAR), $\alpha_0 \in \mathbb{R}^{n(n+1)/2}$ ($=465$ for $n=30$ assets), and $\overline{\alpha}_0$ denotes the average of the the $465$ estimates. The sample standard deviation of the $465$ estimates is shown in parentheses. Moreover, for M-HAR, $\alpha_1, \alpha_2, \alpha_3\in \mathbb{R}$, i.e.\ the bar notation is not needed for M-HAR (and the standard deviation of the single estimate is omitted). For the other models, all parameters are scalar-valued, with the parameters being different for each asset: the bar notation indicates averaging over the $30$ assets. The sample standard deviations of the $30$ estimates are shown in parentheses. Empty cells in the table indicate that the parameter is not available for the corresponding model. The Frobenius and Q-like measures are computed using the mean over the in-sample period.}\\ \end{flushleft} \end{table} }

Forecasting performance

Model selection based on in-sample loss measures may suffer from overfitting. We saw in the previous section that the HARS model achieved a dramatically lower loss compared to the other models. It is thus preferable to fit the model with a training set (in-sample) of the data and use a test set (out-of-sample) for evaluation, which we now do in a forecasting setting.

Recall that the statistical quality of the forecast is evaluated using the Frobenius norm of the predicted covariance matrix minus the true covariance matrix and the Q-Like measure based on the negative log density of a multivariate normal. For both metrics, lower values are preferable.

To evaluate the forecasting abilities of the different models, we therefore consider an out-of-sample evaluation. For each period, we consider the accuracy of the one-step-ahead forecast via a cross-validation approach with a rolling window of fixed size of 1000 observations (equal to the number of in-sample observations) that rolls forward one step at a time. The parameter estimates remain fairly stable as the rolling window moves forward, and thus we re-estimate the model only every 30th observation for computational convenience. This results in predictions for the period April 2012 until June 2024, a total of 3031 predictions for which we compute the loss functions. We additionally divide the losses of losses into two sub-samples: those that correspond to time periods with low-quarticity (defined as the smallest 50% quarticity values) and high-quarticity (defined as the largest 50% quarticity values), respectively. The results are shown in Table (ref). We computed the 90% model confidence sets (MCS) for each loss and sub-sample, respectively, and marked the included models with an asterisk.

The results show that the HARQL model generally has the best performance with the lowest loss in 3 out of 6 cases and 5 inclusions in the MCS. The second best model appears to be the HARSL with the lowest loss twice and five inclusions in the MCS. The other models perform clearly worse and are only included in the MCS for the cases that include most models. The worst performance is found for the HARS and M-HAR models with clearly higher losses and each only one inclusion in the MCS. The poor out-of-sample performance of the HARS model confirms the suspected overfitting observed in the in-sample results. Moreover, the poor performance of the HARQ model is in contrast with the findings in bollerslev2016exploiting,bollerslev2018modeling, whereas the strong performance of the HARQL model highlights the importance of addressing attenuation bias.

{\tiny

table[table omitted — 1,669 chars of source]

}

Next, we performed pairwise Diebold-Mariano tests in order to evaluate specific model extensions against their respective baseline models presented in Table (ref). In particular, we compare models along three dimensions: $log RV$ models vs models for $RV$ in levels, Q-models against the non-attenuated counterparts of the models, and state-space models against fixed parameter versions. The results are not clear-cut, and we conclude the following: The $log$ versions outperform the models for $RV$ in the low quarticity periods. The HARQ model does not outperform its non-attenuated counterpart, and similarly, the evidence that the HARQL is better than the HARL is quite weak. Finally, time-variation based on the state-space approach is better only for the $log RV$ models, and even here, the evidence is not perfectly clear.

Our sample period includes the highly volatile Covid-19 period, which may drive the results. Therefore, in the appendix, we report the results after splitting the out-of-sample periods into the pre-Covid time 2012-2019 and the (post) Covid period 2020-2024; see Tables (ref)-(ref). For the pre-Covid subsample, the ranking of the losses is very similar to the full sample ranking, and the HARQL emerges as the best model. However, the results are more mixed for the 2020-2024 subsample, with HARQL and HARSL performing comparably well, albeit with a slight advantage for HARSL. The pairwise Diebold-Mariano tests for the pre-Covid period show more rejections than for the full-sample period and confirm the superiority of the models for $log RV$. For the 2020-2024 period the results are more mixed, but give evidence in favor of the $log RV$ models for the low-quarticity periods and the HARSL over the HARL.

In general, we conclude that the HARQL model provides the best overall forecast performance, suggesting that modeling $log RV$ while addressing attenuation bias is the preferred approach for modeling and forecasting. Nevertheless, the similar performance of the HARSL model makes it a viable alternative as well.

{\tiny

table[table omitted — 1,263 chars of source]

}

Economic evaluation

For evaluating the economic relevance of the competing forecasting models, we follow the methodology in bollerslev2018modeling summarized in Section (ref) for the same evaluation periods as above. Based on the different forecasts, we construct the global minimum variance portfolio (GMVP), with and without short-selling restrictions, and evaluate its performance using some measures. These measures are turnover (TO), portfolio concentration (CO), short positions (SP), mean and standard deviations of the realized portfolio returns, the Sharpe ratio, and $\Delta_{\gamma}$, which is the amount an investor is willing to pay to switch from a given model to the HARQL model. We selected this model as the basis for its convincing performance in terms of statistical losses. The Sharpe ratio and $\Delta_{\gamma}$ are computed for different transaction costs $c=0,1,2$ percent of the turnover and $\gamma=1,10$. Note that a positive $\Delta$ implies that HARQL is the superior portfolio, whereas negative values show that the corresponding other portfolio is preferred. The results are in Tables (ref) and (ref).

{\tiny

table[table omitted — 2,009 chars of source]

}

{\tiny

table[table omitted — 2,097 chars of source]

}

The results for portfolios without a short-selling restriction, reported in Table (ref), show that different models achieve a comparably low portfolio standard deviation. The HARQ is the only model with a notably higher standard deviation. The model confidence set for the standard deviations (not reported) includes all models, which is also true for (almost) all scenarios considered below. The turnover is lowest for the M-HAR, followed by HAR, and highest for HARQL. Portfolio concentration and short positions are very similar across all models. The mean returns vary significantly over the models, which is attributed to chance, given that means are treated as unpredictable and left unmodeled. Their large variation contrasts with the results in bollerslev2018modeling and drives the results, sometimes dominating the performance of the volatility predictions. Portfolio performance, considering Sharpe ratios and $\Delta_{\gamma}$, is best for HARS and HARQL for low transaction costs, but HAR and HARL are preferable for 2% transaction costs.

The results when imposing a short-selling restriction can be found in Table (ref). Regarding TO, M-HAR and HARS models give the more stable portfolio weights, whereas portfolio concentration is again comparable across models. The HAR, HARL, and HARQL models achieve the lowest portfolio standard deviation. As in the previous case, the variation in mean returns is noticeable and strongly drives the economic value of the competing portfolios. Again, the portfolio based on the HARS model has the highest mean returns and has the largest Sharpe ratios for all levels of transaction costs. For low transaction costs, the second best model is the HARQL, but for larger transaction costs, the basic HAR model is second-best in terms of Sharpe-ration and $\Delta_{\gamma}$.

In the appendix, the corresponding results for the 2012-2019 and 2020-2024 subsamples can be found; see Tables (ref)-(ref). For the first subsample, the HARQL performs well but is characterized by high turnover that results in large transaction costs for higher values of $c$. For the 2020-2024 period, the results are mixed, but the HARS model has the best economic performance due to significantly higher mean returns. In terms of portfolio standard deviations, it stands out that the HARQ model performs significantly worse than all other models for the more turbulent 2020-2024 period. The basic HAR model performs fairly well for both subsamples for large transaction costs due to its low turnover and low portfolio variance.

Overall, these results are not as clear-cut as the statistical evaluation. We can conclude that the simpler M-HAR and HAR models lead to more stable portfolio weights and, hence, less turnover than the more sophisticated models. The mean returns vary unexpectedly much across models, which strongly drives the economic performance, but all models, except the HARQ, result in similar portfolio standard deviations. As global minimum variance portfolios are constructed to minimize portfolio volatility, this indicates that multiple models can result in fairly reliable predictions for the covariance matrix. The HARQL model, our preferred specification for statistical losses, performs fairly well and may be recommended if transaction costs are low. However, the basis HAR model does a good job for high transaction costs.

Conclusion

We studied the problem of forecasting realized covariance matrices by extending the approach proposed by bollerslev2018modeling, which relies on forecasting realized correlations and variances separately. Researchers face the decision of how to precisely model $RV$, and it is not clear which approach is best in the multivariate asset case, so we tried to shed light on this issue by focusing mainly on combining different approaches for univariate $RV$ forecasting while fixing the model for correlations. The evaluation of the models' forecasting performance was based on commonly used statistical and economic criteria using a dataset of 30 stocks that were part of the Dow Jones 30 index at some point during the sample period. The predictions were evaluated for the period 2012-2024 using a rolling window scheme, and as a robustness check, additionally split the evaluation period into two sub-periods (pre-Covid and Covid/post-Covid).

For the statistical losses the results are clear cut and suggest that our proposed multivariate HARQL model for predicting realized covariances based on a univariate attenuation bias approach WLWH20 and the log-transformation of $RV$ is superior. However, the state-space model for $log RV$ also performs well, confirming the finding for the univariate case in bekierman2018forecasting. Moreover, the results suggest that, in general, forecasting models based on log $RV$ are preferable over models for $RV$ itself and that accounting for the attenuation bias is only advantageous for log $RV$. In contrast, the performance of the $HARQ$ model in levels is disappointing compared to the results in bollerslev2016exploiting, bollerslev2018modeling.

The economic evaluation showed mixed and partially counterintuitive results, contrasting the clear-cut findings in bollerslev2018modeling and in our statistical evaluation. The HARQL model performs well but is associated with high portfolio turnover, leading to large transaction costs. On the other hand, the basic HAR has fairly low turnover while still leading to low portfolio variance. The results are strongly driven by large differences in mean portfolio returns, which one would expect to be similar across models. The mixed results can partially be explained by the volatile post-pandemic sample period, arguably making the problem significantly more difficult, and the fact that we consider a set of 30 assets, in contrast to only 10 assets studied in bollerslev2018modeling. Combining the findings of the statistical and economic evaluations, the HARQL model can be regarded as a recommendable specification.

Future research may investigate under which conditions statistical and economic evaluation of covariance forecasts disagree and whether realized covariance forecasting models are helpful for portfolio construction in larger-dimensional settings and during adverse market conditions. Additionally, research might look further into the predictions of realized correlations and whether time-varying parameter models can improve forecast performance.