EconBase
← Back to paper

A Vine-copula extension for the HAR model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

88,424 characters · 27 sections · 93 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Vine-copula extension for the HAR model

abstractThe heterogeneous autoregressive (HAR) model is revised by modeling the joint distribution of the four partial-volatility terms therein involved. Namely, today's, yesterday's, last week's and last month's volatility components. The joint distribution relies on a (C-) Vine copula construction, allowing to conveniently extract volatility forecasts based on the conditional expectation of today's volatility given its past terms. The proposed empirical application involves more than seven years of high-frequency transaction prices for ten stocks and evaluates the in-sample, out-of-sample and one-step-ahead forecast performance of our model for daily realized-kernel measures. The model proposed in this paper is shown to outperform the HAR counterpart under different models for marginal distributions, copula construction methods, and forecasting settings.

Introduction

Volatility estimation and forecasting have been a major research area in financial econometrics. In the last decades, the availability of high-frequency financial data led to prolific research on volatility estimation in high-frequency settings, in particular to the development of the so-called realized measures mcaleer2008realized. From merton1980estimating who first suggested that the volatility can be estimated arbitrary well employing finely sampled high-frequency returns, andersen1998answering,andersen2001distribution,meddahi2002theoretical developed the theory of the widespread realized variance (RV) estimator for the integrated variance. The early studies of andersen2003modeling,andersen2004analytical show that indeed simple models of realized variance outperform the popular GARCH and other SV models in out-of-sample forecasting.

The slowly decreasing autocorrelation and long persistence in squared returns, along with their slow convergence to the normal distribution associated with fat tails and leptokurtic return distributions constitute well known stylized facts. These challenges for the econometric modeling empirically outline the importance of long-memory dependencies in markets' volatility cont2005long.

Several ARCH and SV models have been specifically formulated to deal with this phenomenon, usually by incorporating long-memory patterns through fractional differencing. Fractionally integrated long-memory formulations (e.g. ARFIMA or FIGARCH models) are generally complex, lacking intuitive economic interpretation and mixing long and short memory features of difficult disentangling comte1998long. The estimation is not straightforward and requires long estimation windows bollerslev1996modeling. In forecasting high-frequency volatility measures, the heteroskedastic auto-regressive (HAR) model of corsi2009simple stands out as the main tool, attractive for its simplicity in construction, interpretation, and estimation. Importantly, the HAR model effectively approximates the long-range dependence observed in volatility series and is able to reproduce several stylized facts. Indeed the aggregation as a sum of different processes, like the one the HAR and the earlier HARCH model of muller1997volatilities consider, conveys long-memory features granger1980long,lebaron2001stochastic.

For its linear structure, immediate OLS estimation and remarkable performance in out-of-sample analyses, the HAR model constitutes a well-established benchmark for volatility forecasting with realized measures. Several extensions to the HAR model have been proposed. andersen2007roughing includes a jump component in the regressors, showing that short-lived bursts in volatility are associated with jumps in the process. By following the results of barndorff2008measuring decomposing the RV in semi-variances due to positive and negative returns, patton2015good includes asymmetries based on the return sign, noting that negative returns are of greater impact on RV and have longer persistence than positive ones. corsi2012discrete accounts for both the continuous and jump components of RV and leverage effects too. In a similar specification liu2007there finds strong empirical evidence of structural breaks in realized variance. Further models and applications allowing for structural breaks include mcaleer2008multiple,hillebrand2010benefits,mcaleer2011forecasting,wen2016forecasting,gong2018structural. Structural breaks and leverage effects are tackled under a MEM model perspective in gallo2015forecasting. Further extensions allow for time-varying parameters. Among them, bollerslev2016exploiting implicitly reaches time-variation by accounting for the measurement error between the RV and the integrated variance. In chen2018nonparametric the time variation is free of a functional form but locally approximated with a kernel function.

Recently, buccheri2017hark introduced an autoregressive model with time-varying parameters driven by an autoregressive component and scores of the conditional density. Their time-varying specification can be employed as an alternative representation for general non-linear autoregressive specifications blasques2014optimal. Indeed, the structurally non-linear smooth transition model of mcaleer2008multiple with multiple volatility regimes is shown to be outperformed in out-of-sample forecasting. Considering non-linearities is an important aspect: long-memory features can be misinterpreted as non-linearities and the other way around mcaleer2011forecasting, inflating coefficient's estimates. Hence, the relevance of non-linear models embedding long-memory features for disentangling these two aspects. hillebrand2010benefits proposes a neural network extension where a mixture of logistic functions approximates the unknown function linking log-RV and state variables, extended to dummies for weekdays, macroeconomic announcements and cumulative returns. Lagged variables are selected via bagging, a predictor selection strategy shown to improve, to different extents, the forecasting accuracy for all the models therein investigated. mcaleer2011forecasting further expands hillebrand2010benefits by randomly selecting the number of sums in the bagging algorithm. Similarly, but for a fixed set of predictors arneric2018neural discusses different neural networks alternatives for different HAR specifications, having evidence of statistically significant in-sample and out-of-sample outperformance over their respective linear counterparts.

Besides the particular vector of regressors $\bm{X}_t$ for a day $t$, inclusive or not of possible leverage or jump-variation terms, the problem of determining a suitable functional form $f$ linking $RV_t$ and $\bm{X}_t$ is challenging. This can be either retrieved by a functional approximating strategies based on machine-learning inspired methods like hillebrand2010benefits,lebaron2018forecasting, or with time-varying parameters leading to specifications of equivalent nonlinear representations for some unknown function $f$ buccheri2017hark. A general regression problem takes form $RV_t = f\left(\bm{X}_t,\bm{\beta } \right)$, where $\operatorname*{\mathbb{E}}\left[ RV_t|\bm{X}_t\right] = f\left(\bm{X}_t,\bm{\beta} \right)$ and $\bm{\beta} $ is a vector of parameters. Conversely, linear HAR-like specifications can generically be reduced to a form such as $f\left(\bm{X}_t,\bm{\beta} \right) = \bm{X}_t^{'}\bm{\beta}$. In this work $\operatorname*{\mathbb{E}}\left[ RV_t|\bm{X}_t\right]$ is directly achieved from the conditional distribution $F_{RV_t|\bm{X}_t}$, retrieved from the joint distributions $F_{RV_t,\bm{X}_t}$. This formulation does not restrict the regression over a particular $f$, but allows for general functionals, implicitly determined by the joint distribution $F$, driven by the underlying copula and the complexity of the dependencies between the regressors. Copulas are invariant under transforms of the margins nelsen2007introduction. For alternative log- and square-root HAR specifications, the joint is readily obtained in virtue of Sklar's theorem by simply updating the margins. Importantly, conditional expectations for strictly positive multivariate distributions are by construction nonnegative: forecasts are always positive. The rich information on the joint distribution $F_{RV_t,\bm{X}_t}$ is readily accessible when estimating any of the HAR model formulations. Yet, this information has not been so far exploited, and there are no copula-based approaches in the HAR-related literature. Equivalence in-sample and out-of-sample forecasts wrt. the HAR model would implicitly unveil whether HAR's linear assumption is perhaps misspecified or not. Results favoring our specification would indicate that the information conveyed in the join distribution is by itself highly informative for tomorrow's volatility, directly exploitable without further underlying assumptions.

This paper revisits the HAR model retaining its original formulation involving three interacting volatility components at different time scales, but models their joint distribution with copulas and extracting one-step-ahead forecasts accordingly. The closest work is the bivariate framework of sokolinskiy2011forecasting. We extend it in a setting fully resembling the spirit of the HAR model and adopt a more flexible distribution modeling approach. We adopt the recent advances in multivariate modeling provided by the so-called Vine copulas joe1994multivariate,joe2011dependence. Setting apart from sokolinskiy2011forecasting conditional distributions and expectations are retrieved via numerical integration over multivariate distributions, without relying on simulation. Some of our results are further discussed wrt. a simple neural network benchmark model.

The empirical application considers more than seven years of high-frequency transaction data for 10 stocks and implements different estimation and forecasting schemes, exploiting several different measures to asses the performance of our model with respect to the standard HAR. The model we develop seems to outperform the HAR model both in-sample and in one-day-ahead forecasting.

Section (ref) shortly introduces the realized measure used in our applications and recalls the concept of realized variance. Section (ref) introduces the HAR model and motivates the non-linear copula-based model presented in Section (ref). Vine copula construction is presented in Section (ref). Section (ref) merges all the earlier discussion, and specifies the setting of the empirical application. Section (ref) discusses the results, while Section (ref) concludes. Figures (and tables on test statistics) are conveniently organized in the Appendix.

Volatility estimation with high frequency data

Realized variance

A major problem in high-frequency econometrics consists of the non-parametric estimation of the volatility of a price process. The advantage of rich tick-by-tick data allows in the high-frequency setting to accurately estimate the so-called integrated variance (IV). A widely employed specification models the logarithmic price according to the continuous time diffusion:

equation[equation omitted — 102 chars of source]

Note that this specification involves a time interval $[0,t]$, say e.g. a day. Despite the specific problem considered, a number of further operational assumption can be taken. Equation (ref) is commonly required to be such that the variation of the drift term is neglectful with respect to the stochastic integral, $\sigma_s$ is assumed to be positive, have a continuous sample path, and be independent of the Brownian motion $W_s$. Often $\mu_s \equiv 0$ constant, or predictable and of finite-variation.

The object of interest in the realized variance theory is the so called integrated variance. Let $r\left(0,t\right)$ be the compound return over the period $\left[0,t \right]$, and $\mathcal{F}_t = \left\lbrace \mu_{s}, \sigma_{s} \right\rbrace_{s \in \left[0,t \right]}$ the $\sigma$-algebra generated by the sample paths of drift and diffusion processes. The integrated variance is defined as: $$ IV_{t} = \int_0^t\sigma^2_{s}ds = \text{Var}\left(r \left(0,t \right) \mid \mathcal{F}_{t} \right) $$ This is a key-ingredient commonly taken as and adequate volatility measure over the period $[0,t]$, representing a synthesis of the volatility path through the time interval under consideration.\\ Suppose the log-price process is observed over $[0,t]$. For convenience the price is sampled at regular sub-intervals of size $\delta$, be $\lbrace p_0,...,p_{i\delta},...,p_{n\delta} \rbrace$ the observed prices, ${i = 1,...,n }$ and $n\delta=t$. The realized variance (RV) is defined as:

equation[equation omitted — 174 chars of source]

The realized variance is in fact the second sample moment of the return process over the interval, scaled by the number of observations to provide a measure calibrated to the length of the measurement interval andersen2010parametric.

A well-known implication is that the realized variance is a consistent estimator for the increments in quadratic variation of a process andersen2010parametric, barndorff2002estimating. Under eq.(ref), this implies that the RV consistently estimates the corresponding IV, i.e. $\operatorname*{plim} RV_{t,n} =IV_t$. At high sampling frequencies, non-negligible market microstructure noise (MMS) effects turns the estimator biased and inconsistent. Whether one would like to sample at the highest possible frequency in virtue of the above consistency result, microstructure noise constitutes a limit. 5-minute sampling is a common threshold andersen1998answering, however the realized variance is inevitably affected by discretization error barndorff2003realized,barndorff2006limit.

Market microstructure noise (MMS) is a key concept in high-frequency econometrics. MMS is an error source contaminating the ideal price process of eq.(ref), dominant at high frequencies. Empirical evidence shows that the ideal model in eq.(ref) is inappropriate when prices are sampled at high frequencies: random MMS noise affecting the price should be taken into account to consistently estimate the integrated variance with no bias hansen2006realized.

Realized kernel

Among the feasible IV estimators developed for MMS regimes, we adopt the realized kernel (RK). Exhaustive references on the RK are barndorff2006limit,barndorff2008designing,barndorff2009realized,barndorff2011multivariate, while the general idea of kernel-based estimator for the integrated variance in MMS setting goes back to zhou1996high,barndorff2004regular,hansen2006realized.

We discuss the univariate kernel and of barndorff2009realized, whose theoretical foundations comes from barndorff2008designing. With $r_j$ represents the $j$-th high-frequency return calculated over the interval $\left[ t_{j-1},t_j\right]$ (tick-by-tick return). The realized kernel estimator takes form:

equation[equation omitted — 156 chars of source]

where $k$ is a kernel weighting function. In applications the Parzen kernel is the preferred choice barndorff2008designing,barndorff2009realized. Within the RK analyzed in barndorff2008designing the kernel estimator in eq.(ref) is the so-called non-flat-top realized kernel, which is guaranteed to produce non-negative estimates. Indeed the assumption on the efficient price process is more relaxed wrt. eq.(ref) (e.g. allowing for jumps), while the estimator is consistent also under serially-dependent noise. barndorff2008designing also develops the theory for the selection of an optimal bandwidth $H^*$ in terms of best trade-off between asymptotic bias and variance. The estimation of the quantities involved in determining $H^*$ and the overall implementation of the RK in eq.(ref), are in detail discussed in barndorff2009realized. In our implementation, we adopt the Parzen kernel function and use one observation at each of the sample endpoints for jittering.

HAR model

The HAR model of corsi2009simple stands as a generalization of earlier HARCH models muller1997volatilities, heuristically motivated by the heterogeneous market hypothesis, which assumes the existence of different type of agents, heterogeneous over the different investment horizons they trade. corsi2009simple shows that a simple linear model obtained by mixing three volatility components is able to reproduce the typical slow decay in volatility autocorrelation, stylized facts about returns' and volatility distributions and, although its simplicity, has been shown to be difficult to beat in terms volatility forecasting.

The HAR model assumes a three-factor stochastic volatility model for the latent volatility, identified by the daily integrated variance, captured with an appropriate measure\footnote{The original model of corsi2009simple adopts the RV, but applies to general intraday realized measures, e.g. the RK hillebrand2010benefits,gallo2015forecasting.}. The model assumes that the daily volatility process is a function of the past daily realized volatility and of longer-term partial volatility components \textendash daily component ($d$), weekly component ($w$), and monthly component ($m$). corsi2009simple suggests the following simple time-series representation:

equation[equation omitted — 226 chars of source]

In particular, $RK^{\autobracket*{w}}_t=\frac{1}{5}\sum_{i=0}^4 RK^{\autobracket*{d}}_{t-i}$ and $RK^{\autobracket*{m}}_t=\frac{1}{22}\sum_{i=0}^{21} RK^{\autobracket*{d}}_{t-i}$ are respectively interpretable as weekly and monthly partial volatility terms\footnote{In eq.(ref) $t+1d$ reads as \enquote{(end of) day $t$ plus one day}.}. Volatility innovations are serially independent and zero-mean, with a truncated left tail to guarantee positivity. Eq.(ref) corresponds to an autoregressive model with autoregressive weights taking a step-function form, restricted in a parsimonious way such that the three emerging components are economically meaningful and interpretable.

Although the HAR model stands out as the preferred choice for daily volatility modeling with intraday data, several authors have proposed modifications or improvements, e.g. by different or additional regressors andersen2007roughing,patton2015good,bollerslev2016exploiting, or considering non-linear specifications hillebrand2010benefits,arneric2018neural. In the following, we provide some arguments that motivate the Vine-copula research direction developed in this paper.

The crucial feature of the HAR model is its linearity. Although linearity is attractive in terms of interpretability and eventually in model estimation, this stems as an assumption on the functional linkage between the components. Forecasts from the HAR model are conditional expectations day-t volatility given the observed past terms. Such conditional expectation is assumed as being a linear combination of lagged RK terms. Moving aside from this specification and deal with potential non-linearities, instead of specifying alternative functions to link the RK terms, or applying machine learning -like algorithms to flexibly approximate an unknown functional, it appears interesting to directly look at the joint distribution between the three RK components, by focusing e.g. on their multivariate joint distribution $F_{RK_t^{\autobracket*{d}},RK_t^{\autobracket*{w}},RK_t^{\autobracket*{m}}}$, and consequently by considering the expectation of the conditional distribution $F_{RK_{t+1d}^{\autobracket*{d}}|RK_t^{\autobracket*{d}}, RK_t^{\autobracket*{w}}}$ for forecasting $RK_t^{\autobracket*{m}}$ sokolinskiy2011forecasting, without further assumptions, perhaps on the functional forms and errors' distribution of any potential model. Implicitly this framework accounts for possible non-linearities in the conditional expectation, since not constrained to a specific functional form, but directly recovered from $F$ and driven by the dependence relationships therein involved between its variables.

On the other hand, an OLS-estimated linear model like eq.(ref) is not guaranteed to produce positive estimates of tomorrow's volatility. As a turnaround, a log-specification of the HAR model can be adopted, however, the forecasts are not of direct use (Jensen inequality) and need to be e.g. bootstrapped. Also in our data, we have spurious evidence of negative estimates and confidence intervals, which are economically of no sense. In a non-log framework RK, the positiveness of RK implies a non-normal error term in (ref). This is not affecting the OLS estimator in terms best linear unbiased estimator (BLUE) of the regression coefficients, but poses issues in inference and e.g. in predicting confidence intervals. Non-linearity and positivity issues in the original HAR model, favor the discussion over a structural-free model that directly exploits the joint distribution of the four volatility components. Such an alternative is discussed in the next Section.

CV-HAR model

Motivated by the discussion in Section (ref), we look at the joint distribution of the RK-based day-$\autobracket*{t+1d}$, day-$t$, last week's and last month's volatility measures, $F_{RK_{t+1d}^{\autobracket*{d}},\, RK_{t}^{\autobracket*{d}}, \, RK_{t}^{\autobracket*{w}}, \, RK_{t}^{\autobracket*{m}}}$ to extract the conditional distribution $F_{RK_{t+1d}^{\autobracket*{d}} | RK_{t}^{\autobracket*{d}}, \, RK_{t}^{\autobracket*{w}}, \, RK_{t}^{\autobracket*{m}}}$ and compare the HAR model against the alternative:

equation[equation omitted — 244 chars of source]

The error term accounts for both measurement errors and the variability in $RK_{t+1d}^{\autobracket*{d}}$. Whereas the day-ahead volatility forecasts of the HAR model are obtained by evaluating the right side of eq.(ref) by plugging the observed RK values, here the forecast is the expectation of day-$\autobracket*{t+1d}$ volatility given the other volatility components begin equal to the empirically observed RK counterparts. To obtain the forecasts, the model requires nothing but evaluating the integral involved in the conditional expectation. At time $t$ as a forecast for the volatility at $t+1d$, with the HAR and the above model we respectively have:

align[align omitted — 518 chars of source]

where $x$ are the sampled (observed) RK values at day $t$ for the different volatility components.

Indeed is eq.(ref) we should look at (rather than eq.(ref)) to get the intuition behind our model: \enquote{forecast tomorrow's volatility with the conditional expectation of tomorrow's volatility given today's, last week's and last month's}.

The model in eq.(ref) can be seen a generalization of the framework in eq.(ref): with $\bm{X}$ being a vector of regressors, a general regression problem takes the form $Y = f\autobracket*{\bm{X},\bm{\beta}}$, with $f\autobracket*{\bm{X},\bm{\beta}} = \operatorname*{\mathbb{E}}\left[ Y \mid \bm{X} \right]$. Whether in the HAR model $f$ is a linear function of the parameters, in the CV-HAR setting, $f$ already embraces a conditional expectation and does not imply any strict structure (e.g. linear relationship between the regressors): $f\autobracket*{\bm{X},\bm{\beta}} = \operatorname*{\mathbb{E}} \left[RK_{t+1d}^{\autobracket*{d}} | RK_{t}^{\autobracket*{d}}, \, RK_{t}^{\autobracket*{w}}, \, RK_{t}^{\autobracket*{m}} \right]$. In this light, the HAR model can be seen as a specific parametrization of a more general model which directly exploits conditional expectation $\operatorname*{\mathbb{E}}\left[RK_{t+1d}^{\autobracket*{d}} | RK_{t}^{\autobracket*{d}}, \, RK_{t}^{\autobracket*{w}}, \, RK_{t}^{\autobracket*{m}} \right]$, as our specification aims to. Modelling the volatility at $t+1d$ as the expectation of a conditional distribution $F_{RK_{t+1d}^{\autobracket*{d}}|RK_t^{\autobracket*{d}}, RK_t^{\autobracket*{w}}, RK_t^{\autobracket*{m}}}$, by construction constraints the forecasts to the positive domain of $RK_{t+1d}^{\autobracket*{d}}$. Logarithmic specifications of the HAR model are no longer attractive, since this alternative approach circumvents positivity and normality issues, being naively suitable to cope with non-transformed realized measures. Furthermore, by modelling the joint distribution with copulas, transformations of the variables are not affecting the underlying copula nelsen2007introduction, so that also for modelling purposes transformations are irrelevant. Also, the regression in eq.(ref) as such, produces symmetric confidence intervals based on t-distribution quantiles for the forecast values, while skewness and heavy-taildness are often observed in volatility. On the contrary, the above approach leads to confidence intervals that are immediately identified by the actual quantiles of the same conditional distribution, potentially non-symmetric and showing kurtosis. Since the joint distribution modeling based on C-Vine copulas presented in the following Section, we call such model CV-HAR\footnote{Where \enquote{CV} stands for \enquote{C-Vine} and \enquote{HAR} is {\it nothing more than a reminder} of the motivation related to the HAR model. It does not stand for \enquote{heterogeneous auto-regressive}.}.

Modelling joint distributions with Vine copulas

Although the wide range of flexible bivariate parametric copulas, the number of copula families available for multivariate (three or more variables) modeling is rather limited in contrast to the bivariate case. In the last two decades, a number of methods have been developed to construct high dimensional copulas with desirable proprieties joe2014dependence. Among them we mention: hierarchical (or nested) Archimedean copulas mai2012h, mixtures of max-min infinitely divisible distributions joe1996multivariate, Factor copulas hull2004valuation, mcneil2005quantitative,oh2017modeling and Vine copulas joe1994multivariate, joe1996families, cooke1997markov. The pair construction method we adopt in this paper leads to the so-called Vine copulas (or Vines), see e.g. czado2010pair,joe2011dependence for an introduction to Vine copulas.

Vine copulas

This Section aims at providing a short introduction to Vine copulas and in particular to the recursive pair-copulas construction method. Further details can be found e.g. in the monograph of joe2011dependence. The starting point for constructing multivariate distributions is the well known recursive decomposition of a multivariate density into products of conditional densities. Let $(X_1, ...,X_d)$ be a set of random variables with joint distribution $F$ and density $f$, let $F(\cdot|\cdot)$ and $f(\cdot|\cdot)$ denote conditional CDFs and densities respectively, then:

align[align omitted — 217 chars of source]

As second ingredient we need Sklar's theorem (in its density form) to conveniently factorize a bivariate density $f (x_1,x_2)$ into a product of (unconditional) marginals $f_1$, $f_2$ and a bivariate copula density $c_{12}(\cdot,\cdot)$:

equation[equation omitted — 88 chars of source]

Using (ref) we can express the conditional density of $X_1$ given $X_2$ as

equation[equation omitted — 79 chars of source]

For distinct indices $i,j$, $i_1,...,i_k$ with $i<j$ and $i_1< ...<i_k$ we use the abbreviation: $$ c_{ij|i_1,...,i_k} = c_{ij|i_1,...,i_k} \left( F(x_i|x_{i_1},...,x_{i_k} ),F(x_j |x_{i_1} ,...,x_{i_k} )\right) $$ From equation (ref) we can express $f\left( x_i|x_1,...,x_{i-1}\right)$ recursively, this yields to the expression:

equation[equation omitted — 130 chars of source]

Taking equation (ref) in (ref), it follows that:

align[align omitted — 420 chars of source]

This is called C-(canonical) Vine distribution.

Note that C-Vine decomposition of the joint density in eq.(ref) consists of pair-copula densities $c_{ij|i_1,...,i_k}$ specified for the variables indices $i,j$, conditioned to variables $i_1,...,i_k$, evaluated at the conditional CDFs $F\left( x_i|x_{i_1},...,x_{i_k}\right)$, $F\left( x_j|x_{i_1},...,x_{i_k}\right)$ and marginal densities. This is why such a decomposition is called pair-copula decomposition and the above construction leading to eq.(ref) is called pair-copula construction. Importantly note that the decomposition is not unique, there are indeed $\frac{d\left(d-1\right)}{d}$ different sets of copulas to chose from and thus structures that build up to the joint distribution of variables $1,..,d$. Therefore, based on the specific problem under consideration, a specific tree needs to be identified, see Subsection (ref).

Estimation

The standard framework for Vine estimation is likelihood maximization. From eq.(ref) the log-likelihood $l$ is immediately recovered for a C-vine copula with parameter $\bm{\theta}_{CV}$, for a sample $\bm{u} = \autobracket*{u_{k,j}}$, with $k=1,...,N$ and $ j=1,...,d$:

equation[equation omitted — 300 chars of source]

where $F_{j|i_1:i_m} = F\autobracket*{u_{k,j}|u_{k,i_1},...,u_{k,i_m}}$ and the marginal distributions are uniform, i.e., $f_k\autobracket*{u_k} = \bm{1}_{\left[ 0,1 \right]}\autobracket*{u_k}$. $\bm{\theta}_{i,i+j|1:\autobracket*{i-1}}$ is the parameter set corresponding to the copula $c_{i,i+j|1:\autobracket*{i-1}}$. Note that according to eq.(ref) $F_{j|i_1:i_m}$ depends on the parameters of pair-copula terms in tree 1 up to tree $i_m$ brechmann2013cdvine. For extensions to general R-Vine structures see e.g. joe2011dependence.

The likelihood maximization is not limited to the best parameter selection but defines the copula parametric specifications for the copulas in all the trees. While the unconditional copulas on tree 1 one can be easily specified by the researcher simply by using the input data, the copulas on the conditional CDF in higher trees are not directly available, since the conditional sample is not observed. Therefore the most common way for the parametric copula specification at trees $m>1$ is by recursively trying different copulas families for each conditional copula and retain the one leading to the highest likelihood in its fitted parameters. I.e. a given parametric copula $C_{i,i+j|1:\autobracket*{i-1}}$ is selected by recursively evaluating the likelihood of the data with different copulas specifications (Archimedeans, Gaussian and t- copulas). The Vine copula specification leading to the highest overall likelihood is then retained (actually the criterion we use is based on This copula selection proceeds tree by tree, since the conditional pairs in trees $2,...,m-1$ depend on the specification of the previous trees. Fig.(ref) (lower panel) provides an illustration of an estimated Vine.

It is a good practice to inspect the unconditional copulas specification at the first tree with alternative methods, since misspecifications would propagate through the whole tree and affect all the other parameters' estimates. These include graphical methods such as quantile dependence plots oh2017modeling, contour plots for the fitted copula, $\chi$ and $k$ plots and likelihood ratio tests such as Voung and Clarke vuong1989likelihood,clarke2007simple. See brechmann2013cdvine and references therein for further details.

Pair copula construction therefore allows to sequentially combine specific pair copulas to build up a multivariate copula (and thus distribution) by identifying suitable bivariate pair-copulas:

itemize• Use the sample to model the marginal distributions $F_i$, $i=1,...,d$ involved in the first tree. Use the margins to reduce the sample to the $\left[ 0,1 \right]$ interval. • Specify the tree structure, i.e. type of vine and variable order, Subsection (ref)) • Determine conditional and unconditional copulas $c_{ij|i_1,...,i_k} $ families and parameters on the basis of the above discussion. • Apply eq.(ref) to build the multivariate distribution based on vine copula construction.

Importantly, the above procedures can be exploited to test for independence by using the independent copula $C\autobracket*{u,v} = uv$. along with test statistics for the Spearman's rho and Kendall's tau dependence measures genest2007everything.

Tree-structure representation

As earlier mentioned, a density $f\left(x_1,...,x_d\right)$ can be represented by a product of pair-copula densities and marginal densities. The decomposition is however not unique. For instance, with $d = 3$, a possible decomposition for $f\left(x_1,x_2,x_3\right)$ is:

align*[align* omitted — 1,111 chars of source]

However by applying a different conditioning,

align*[align* omitted — 178 chars of source]

which is clearly a different decomposition of $f\left(x_1,x_2,x_3\right)$.

In general $d$-dimensional joint distributions allows for $d\left(d-1\right)/2$ different pair-copulas. bedford2001probability introduced a tool called regular vine (R-Vine) structure to help to organized them. The formal definition of a regular vine they provide allows for a convenient graphical representation of the structure in terms of trees and nodes. Importantly, among the regular vines, an important sub-class is that of the so-called canonical vines (C-Vine). Fig.(ref) provides a graphical representation of a C-Vine. Canonical vines have a typical \enquote{star} structure, where each tree has a unique node connected to all the other nodes. By defining the degree of a node as the number of nodes attaching to it, for a $d$-dimensional problem, a C-Vine is a structure such that each node in tree $T_j$, $j=1,\dots d-1$ is of maximal degree, i.e. each tree $T_j$ has a unique node of degree $j-1$.

Fig.(ref) clarifies the above statement. The C-vine density therein depicted, corresponds to the factorization; $$ f_{1234} = f_1 \cdot f_2 \cdot f_3 \cdot f_4 \cdot c_{12} \cdot c_{13} \cdot c_{14} \cdot c_{23|1} \cdot c_{24|1} \cdot c_{34|12} $$ where $X_1$ is set as a node in tree $T_1$, and the dependence with any other variable is considered with respect to it. I.e. the involved pair of variables are 12, 13, 14 with their respective pair-copulas $c_{12}$, $c_{13}$, $c_{14}$. In three $T_2$, the connection involves $F_{12}$, $F_{13}$, $F_{14}$, with $F_{12}$ as node. Conditional on the common variable (in the Fig.(ref) always 1, corresponding to $X_1$) the dependence between $F_{2|1}$ and $F_{3|1}$ is captured by the copula $c_{23|1}$ (which would naturally arise given a proper factorization of $f\autobracket*{x_1,.x_1,x_2,x_4}$ by applying the law of total probability, similarly as eq.(ref)). Analogously $c_{24|1}$ connects $F_{2|1}$ and $F_{4|1}$. In the last tree $T_3$ a path connecting 23\textbar1 and 24\textbar1 ($F\autobracket*{X_2|X_1,X_3|X_1}$ and $F\autobracket*{X_2|X_1,X_4|X_1}$) is given by the copula $c_{23|12}$. Coherently with the above definition, C-Vines always show a path connection (rather than a star) in the last tree $T_{d-1}$.

C-Vines log-likelihood has the convenient form of eq.(ref), that is of immediate evaluation given the result discussed in Section (ref) about the computation of conditional CDFs. About the structure selection, some guidelines are provided e.g. in dissmann2013selecting,czado2013selection. The intuition of selecting the structure leading to maximum likelihood is unfeasible for-large dimensional problems. Some hypothesis on the structure, i.e. on the relationship between variables must be accounted for in order to simplify the selection problem. Thus the nature of the problem and its interpretability are important drivers in structure selection. For the volatility modeling problem here analyze we chose a C-Vine structure, although other structures might be applicable and of feasible interpretation too. See section (ref) for further details in this regard.

Conditional distributions and expectations from Vine copulas

In our CV-HAR specification the quantity of interest is an expectation having form $\operatorname*{\mathbb{E}} \left[ X_1|X_2,X_3,X_4 \right ]$. Here we discuss how such an expectation is obtained from the conditional distribution extracted from the overall Vine joint.

Conditional distribution

By now, consider the simplest case of $X_1$ and $X_2$ being uniformly distributed, $C$ is their copula and $Y$ their joint CDF. Be $0 \leq \epsilon \leq 1- x_2$, with $x_1, x_2 \in \mathbb{R}$.

align*[align* omitted — 383 chars of source]

Since $X_2$ is uniform, $\operatorname*{P}\autobracket*{X_2 \in \autobracket*{x_2, x_2 + \epsilon}}\epsilon}} = \epsilon$, by Bayes the theorem we compute the conditional probability:

equation*[equation* omitted — 195 chars of source]

By letting $\epsilon \rightarrow 0$ (provided that the limit exists) the partial derivative arises and the conditional CDF $X_1|X_2 = x_2$ solves to nelsen2007introduction:

equation*[equation* omitted — 130 chars of source]

The conditional CDF $F\autobracket*{x_1|x_2}$ turns to have an immediate expression in terms of the copula $C$ between $X_1$ and $X_2$: just take its partial derivative with respect to the conditioning variable and evaluate it in $\autobracket*{x_1,x_2}$.

The case with $X_1$ and $X_2$ distributed according to $F_1$ and $F_2$ (in general different and not uniform), by Sklar's theorem, is immediately solved by updating the arguments in which the copula is evaluated:

equation[equation omitted — 257 chars of source]

Similar results can be proved for the conditional PDF, instead of CDFs. By the same reasoning, expanding for $d = 3$, e.g. $F_{3|12}$ can be evaluated as follows:

align*[align* omitted — 529 chars of source]

where $C_{12}$ is the copula associated with the pair $\autobracket*{X_1,X_2}$, $C_{13}$ the copula for $\autobracket*{X_1,X_3}$ and $C_{32|1}$ the copula for $\autobracket*{X_3|X_1,X_2|X_1}$. It clearly emerges that the earlier pair copulas $C_{12}$ and $C_{13}$ are sequentially used to estimate the conditional CDFs. And that these constitute the arguments of the higher-order conditional copula $C_{32|1}$, from which the higher-order conditional CDF $F_{3|12}$ is constructed.

In a general setting joe1996families obtains the following recursive relationship, for $F\autobracket*{x|\bm{v}}$:

equation[equation omitted — 253 chars of source]

where $\bm{v}$ is a $m$-dimensional vector, $v_j$ any arbitrary component of $\bm{v}$ and $\bm{v}_{-j}$ the $(m-1)$-dimensional vector obtained by excluding $v_j$ from $\bm{v}$. Importantly, note that $C_{xv_j|\bm{v}_{-j}}$ is always a bivariate copula function. \\ This is a crucial result for the conditional CDF construction, since it shows that $F\autobracket*{x|\bm{v}}$ can be obtained by {\it sequentially mixing} conditional CDFs with copulas, where the conditional copula $C_{xv_j|\bm{v}_{-j}}$ depends on the copulas $C_{xv_i|\bm{v}_{-ij}}$ and $C_{v_j v_i|\bm{v}_{-ij}}$, conditional on the smaller set $\bm{v}_{-ij}$, and so backwards up to the unconditional copulas on the first tree.

It is clear how Vine copulas are particularly attractive in terms of the recursive relation in eq.(ref). Once the structure is specified, all the conditional copulas are settled in terms of their parametric characterization. All the conditional distributions $F\autobracket*{x|\bm{v}}$, recursively determined by the partial derivatives of the copulas identified in the previous tree, are determined as well. Once the vine structure is estimated, all the pair-copula parameters are known and the Vine is completely determined. By simple substitution of the conditioning values the variables' CDFs, $F\autobracket*{x|\bm{v}}$ is readily computed.

Conditional expectation

The conditional CDF can be recursively evaluated given the convenient Vine decomposition. This stands as a starting point to evaluate conditional expectations of the type $\operatorname*{\mathbb{E}} \left[ X|\bm{v} \right ]$. Solving $\operatorname*{\mathbb{E}} \left[ x|\bm{v} \right ]$ is in general an integration problem, whose complexity depends on the parametric copulas involved in eq.(ref). In this research, we follow a non simulation-based approach for computing the conditional expectation sokolinskiy2011forecasting. Although the most used definition for computing the expectation is in terms of integral with respect to the conditional PDF ($f\autobracket*{x|\bm{v}}$), eq.(ref) involves CDFs. Not to differentiate twice the CDF to extract the corresponding PDF, but to use eq.(ref) directly we adopt the following alternative in terms of CDF:

equation[equation omitted — 172 chars of source]

The joint implementation of the nested structure and integration of eq.(ref) and eq.(ref) is complex. We verify the above implementation by comparing conditional expectation computed via simulation for a number of different Vines.

Empirical application

Data

This research uses trade data for 10 of the 30 stocks constituting the Dow Jones industrial average index. The stocks under consideration in this analysis are, AAPL, AXP, BA, CAT, CSCO, CVX, DIS, GS, HD, IBM. In order to have a data-sample large enough for in-sample and out-of-sample analyses, the data covers a long span of 1634 days\footnote{The length of the data used in the application reduces to 1626 days, by excluding 22 days for constructing the first $RK_{t}^{\autobracket*{w}}$.}, from January 1\textsuperscript{st} 2012 to June 30\textsuperscript{th} 2018. The data is extracted from the TAQ database, consisting of raw trade prices, their respective timestamps, quantities and other fields identifying e.g. the exchange. For each stock raw prices have been preprocessed and cleaned according to the guidelines presented in barndorff2009realized. To avoid biases induced by non-regular trading hours entries outside 9:30am-4pm have been removed, as well as entries with transaction price equal to zero. For the 10 stocks the exchanges on which the trading activity took place have been recorded, and their absolute frequencies analyzed. We retain the data from the exchange \enquote{D} - Financial Industry Regulatory Authority, Inc. (FINRA ADF)\footnote{TAQ manual: Daily TAQ client specification. Version 2.2a.}- it collects 25.13% of the total number of transactions that occurred in the period analyzed (and traded all the above 10 stocks, for every day). Entries with abnormal sale condition (not corrected or canceled by the participant), were also removed. Multiple transactions sharing the same timestamp have been replaced by their median price. The timestamp resolution varied in the period under investigation, however, the accuracy is up a millisecond. Entries for which the transaction price mean absolute deviation deviated by more than 10 times the average MAD computed over a centered window of 2 minutes length., have been removed as well\footnote{This partially resembles the rule \enquote{Q4} of barndorff2009realized: besides applying the previous cleaning steps, occasionally there are extreme outliers, visually not coherent with the average daily behavior of the price series.}. By averages over the last 5 and 22 RK measures, we compute $RK_{t}^{\autobracket*{w}}$ and $RK_{t}^{\autobracket*{m}}$.

Margins

Modelling the CDFs corresponding to the four\footnote{Note that unconditionally $RK_{t+1d}^{\autobracket*{d}}$ and $RK_{t}^{\autobracket*{d}}$ share the same distribution. In practice, we deal with tree marginal distributions.} volatility terms is a sensible step. On first instance this leads to the transformed sample in $\left[0,1 \right]$ interval upon which the copula model is built, secondly it drives the implementation of eq.(ref) and eq.(ref). Not to obtain results that are specific and valid under a given procedure for computing the CDFs, for each of the four volatility components $RK_{t+1d}^{\autobracket*{d}},\, RK_{t}^{\autobracket*{d}}, \, RK_{t}^{\autobracket*{w}}, \, RK_{t}^{\autobracket*{m}}$ we perform the analysis by the use of (i) parametric CDFs, (ii) Kernel-based CDFs (estimated over positive domains) and (iii) ECDFs. In the parametric case we fit an Inverse-Gaussian (IG) distribution, motivated by the arguments of barndorff2002econometric,forsberg2002bridging.

Vine-Copula construction

The C-Vine copula construction follows the methodology described in Section (ref). The copula models considered in the analyses are the Archimedean copulas (Gumbel, Frank, Joe, Clayton), the Gaussian copula, and the t-copula joe2014dependence. The implementation of eq.(ref) does not pose problems for the Archimedean copulas (all the copulas are smooth functions), while the differentiation of the Gaussian and particularly the t-copula is achieved numerically with the approximation $\partial C\autobracket*{u,v}/\partial v = \autobracket*{C\autobracket*{u,v+h}-C\autobracket*{u,v}}C\autobracket*{u,v}}/h$ with $h = 0.001$. The use of Gaussian and t-copula constitutes a computational complication, which is however not granted to lead to a considerable gain in fitting performance and, in the latter, forecast improvements with respect to the Archimedean set. Therefore the analyses are conducted separately, for Vine constructions estimated using Archimedeans copulas only (\enquote{A} set) and for Vines allowing for of all the six alternatives (\enquote{AGT} set).

The copula selection in the different trees is automated based on the AIC criterion. However, misspecifications in the first tree would propagate through the whole Vine structure. Therefore, for a selected number of stocks and a range of days, we double-check by use of the alternative selection methods described Section (ref).

As mentioned in Section (ref), the choice of the structure is not unique. We use a C-Vine copula representation. The C-Vine structure is built in such a way that $RK_{t}^{\autobracket*{m}}$ constitutes a node on the first tree. This is a feasible choice, assuming a cascade effect of past volatilities to the current one, i.e. that the past volatility and its history drive today's - modeling in the first tree last moth's volatility given today's is not logical as modeling today's given last month's. Note that the intuition of using $RK_{t+1d}^{\autobracket*{d}}$ as a node on the first tree is not feasible since our target is the estimation of $\operatorname*{\mathbb{E}}\left[RK_{t+1d}^{\autobracket*{d}} | RK_{t}^{\autobracket*{d}}, \, RK_{t}^{\autobracket*{w}}, \, RK_{t}^{\autobracket*{m}} \right]$ by eq.(ref): within this framework $RK_t$ can only by conditioned and not conditioning. In this way, the first tree considers all the dependencies between the volatility variables and $RK_{t}^{\autobracket*{m}}$ . On the last tree, in the conditional copula $c_{RK_{t+1d}^{\autobracket*{d}} RK_{t}^{\autobracket*{d}} | RK_{t}^{\autobracket*{w}} RK_{t}^{\autobracket*{m}}}$ today's and yesterday's volatility are conditioned to last week's and month's, which is coherent and intuitive from a logical point of view (see Fig.(ref), upper panel). The estimation follows Section (ref). The conditional CDF $F_{RK_{t+1d}^{\autobracket*{d}} | RK_{t}^{\autobracket*{d}}, \, RK_{t}^{\autobracket*{w}}, \, RK_{t}^{\autobracket*{m}}}$ is retrieved by applying eq.(ref), whereas the conditional expectation $\operatorname*{\mathbb{E}} \left[ RK_{t+1d}^{\autobracket*{d}} \vert RK_{t}^{\autobracket*{d}} = x_t^{\autobracket*{d}}, \, RK_{t}^{\autobracket*{w}} = x_t^{\autobracket*{w}}, \, RK_{t}^{\autobracket*{m}} = x_t^{\autobracket*{m}} \right]$ is evaluated with eq.(ref).

A simple non-linear benchmark model

As a benchmark for comparing the CV-HAR model, we implement a simple neural network (NN) model. We follow the methodological approach of arneric2018neural, by implementing a feed-forward neural network with a single hidden layer embracing two neurons. arneric2018neural argues that such a network design is optimal in terms of in-sample MSE, prevents over-fitting and is parsimonious. A logistic activation function is applied between input and hidden layers, and a linear function between the hidden and output layers medeiros2006building. The very same regressors as for the HAR and CV-HAR models are used for the NN estimation, namely $RK_{t}^{\autobracket*{d}}$, $RK_{t}^{\autobracket*{w}}$, $RK_{t}^{\autobracket*{m}}$. Following medeiros2006building,hillebrand2010benefits, the network is estimated via Bayesian regularization mackay1992practical, along with the Levenberg-Marquardt optimization algorithm hagan1994training. 70% of the data is used as a training sample, the remaining 30% for validation. Splits are randomly initialized, as well as the initial parameters. Hence, each of the NN forecasts is computed by averaging over 500 bootstraps.

Results

Estimation and forecasting schemes

The estimation of both the HAR and CV-HAR models, and thus the corresponding measures of forecast accuracy, are developed under three different schemes.

itemize• Fixed window (FW). Estimation of the models at days 250, 500, 750 and 1250. Thus, we estimate the models by using the first $W=\left\lbrace 250,500,750,1250\right\rbrace$ observations respectively. $W$ splits the sample in two. The part involving observations from day one to $W$, upon which the model is estimated, is used for in-sample analysis. The remaining part, from day $W+1$ to the last, is used for out-of-sample analysis. Coherently, in-sample days from day one to $W$, constitute the training-validation set of the NN implementation (with a 70%-30% split). The remaining out-of-sample days, constitute the training set. • Increasing window (IW). We first estimate the HAR and CV-HAR models on days 1 to $W$, then sequentially for each day $d=W+i$, $ i=1 \dots \autobracket*{1626-1}-W$ we re-estimate the model by using the whole dataset up to day $d$. With this procedure, for any day between $W+1$ and 1625, one-step-ahead forecasts are constructed, and the effect of sequentially increasing the sample size analyzed. Results are based on the following sizes of the first window: $W=\left\lbrace 250,500,750 \right\rbrace$. • Rolling window (RW). Similarly, as in (ii), we re-calibrate the model and construct one-step-ahead forecasts by using rolling windows of size $W=\left\lbrace 250, 500, 750\right\rbrace$ days. I.e. forecasts at day $d+1$ are based on models estimates with observations from $d-W+1$ to $d$, for $W+1\leq d \leq 1625$.

These three approaches provide different copula estimates. The IW and RW schemes allow for time-variation of the copula and therefore of the dependence between the pair-variables in the underlying copulas. For the pair-copulas involved in the C-Vine tree, in Fig.(ref) we show the estimated copulas with the increasing window approach. Since the parameters' space for different copulas is different, we plot the corresponding Kendall's-$\tau$ implied from the estimated copula. Fig.(ref) uncovers a time-varying nature of the dependence between the variables' pairs that the FW approach is unable to capture. Therefore, although IW and RW are much more demanding than FW from a computational point of view, including IW and RW provides a deeper insight into the complex dependence dynamics between the variables.

In the results' tables, we adopt the following measures to compare the performance of the CV-HAR against the HAR model. (i) Mean squared error (MSE), (ii) mean absolute error (MAE), (iii) median absolute deviation (MAD), (iv) mean absolute scaled error (MASE) hyndman2006another, (v) mean absolute percentage error (MAPE), (vi) mean directional accuracy (MDA) and (vii) Qlik patton2009evaluating which has been extensively used in similar applications patton2015good,bollerslev2016exploiting\footnote{ With $\lbrace y_t \rbrace$ being sample values and $ \lbrace \hat{y}_t \rbrace $ their respective forecasts, with $i = 1,...,T$: $MASE = \frac{1}{T} \sum_{t=1}^T \frac{\mid \hat{y}_t - y_t \mid }{\frac{1}{T-1}\sum_{t=2}^T \mid y_t - y_{t-1} \mid}$, $MAPE = \frac{1}{T} \sum_{t=1}^T \mid \frac{y_t-\hat{y}_t}{y_t} \mid$, $ MDA = \frac{1}{T} \sum_{t=1}^T \bm{1}_{sign\autobracket*{y_t-y_{t-1}} == sign \autobracket*{\hat{y}_t - y_{t-1}}}$, $Qlik = \frac{y_t}{\hat{y}_t}-\log\autobracket*{\frac{y_t}{\hat{y}_t}}-1$.}. Results report the average measures over the 10 stocks. Whereas MSE, MAE, MAD, MAPE and Qlik are actual overall measures, MDA and MASE are averages of the 10 individual measures computed for each stock\footnote{MDA and MASE involve lagged values: it is not possible to compute them on an overall basis from an general time-series constructed by stacking the individual ones.}.

One-step-ahead forecasts of the two models have been tested to be statistically different with the Diebold-Mariano (DM) test diebold2002comparing and the conditional predictive ability (CPA) test giacomini2006tests. The well-known DM test applies to non-nested models only. For the IW scheme, the CPA test constitutes a suitable alternative. Under RW we apply both the DM and CPA tests. Tests are applied to squared, absolute and Qlik loss functions. This corresponds to test for differences in one-step-ahead MSE, MAE, and Qlik for the two models. For the FW case, analyses are separate for the sample up to $W$, and from $W$ to 1625. These correspond to in-sample and out-of-sample analyses. For the IW and RW cases, measures refer to one-step-ahead forecasts. Accordingly, note that the measures are interpreted differently, e.g. MSE under FW is the in-sample MSE, while under RW is the one-step-ahead out-of-sample MSE.

Common NNs, and more generally machine-learning implementations are difficult to scale over RW and IW schemes. The implementation of the NN for the RW and IW schemes is computationally challenging and very demanding, if not unfeasible. E.g. just for the IW case with $W=250$ this would require $6.88\cdot 10^5$ estimations over an increasing data-sample (1363 forecasting days and 500 bootstraps) to average out the effects of the initial random sample split and weights. Such a complex and time-consuming estimation is out-of-scope wrt. the objectives and motivation of the present research: the NN is discussed under the FW scheme only. Broader NN applications and analyses are left for future research.

In-sample analysis

We present the results relative to the estimation of the CV-HAR model against the HAR model. Epochs at which the models are estimated under a fixed-window approach, corresponding to the width $W$ of the estimation period, are $W=\lbrace 250, 500,750,1250\rbrace$. This analysis focuses on the in-sample accuracy of the two models under investigation. Since this does not involve any (one-step-)ahead forecasting, no formal testing of forecasting accuracy is here developed. Results in tab.(ref) show that on a general level the window size has an important impact. Indeed, individual measures are generally improving as the size of the window widens. In this regard, under $W=250$ we do not observe a uniform improvement of the CV-HAR model over the HAR, which is remarkable for wider windows. However, at the shortest window $W=250$, improvements over the HAR model are observed corresponding to copula constructions allowing for Gaussian and t-copulas besides Archimedeans. This indicates that in small samples, flexibility on pair-copulas leads to clear improvements over a constrained framework allowing for Archimedeans only. This pattern applies in general too, i.e. measures corresponding to constructions relying on Archimedeans only are outperformed by those extending the set to Gaussian copula and t-copula too. This is however not crucial under wider windows, where the model is always satisfactory, suggesting that the role of the extended set at short window lengths is that of correcting for small-size distortions in the joint distribution of the volatility terms captured by the C-Vine copula. On the other hand, models for margins seems not to have a role in this analysis, since all the constructions lead to similar ratios. Furthermore, note that the mean directional accuracy (MDA) is generally close to unity. This indicates that the HAR and CV-HAR models are equally capable of forecasting the direction of tomorrows' volatility movement wrt. to today's (i.e. increase or decrease).

Overall, on an in-sample basis, these results seem to favor the CV-HAR flexibility allowing for conditional means of generic non-linear nature. The linear combination between the volatility components of the HAR model, besides being outperformed for most of the measures, is not guaranteed to produce positive volatility forecasts. As it appears from Fig.(ref), HAR estimates are in general around an average volatility level, not accurately tracking neither days of low nor high volatility. In fact, a simple linear model in the three volatility terms, without any further dummies, approximates today's volatility overall average behavior. On the other hand, the CV-HAR seems capable of reacting, although less promptly than the HAR model, to high volatility periods by generating appropriate estimates and non-linearly adapting to low the different regimes observed in the sample.

That the CV-HAR successfully identifies a non-linear pattern is confirmed by the vicinity of the performance measures wrt. to those from the NN. On an in-sample basis, the NN seems to outperform the CV-HAR model wrt. MSE measure. This is not surprising considering that the NN estimation explicitly seeks for the set of parameters minimizing the MSE. On the contrary, the CV-HAR model is based on the apparently unrelated copula fitting via likelihood, from which forecasts are indirectly extracted. Because of the completely different modeling and underlying assumptions, construction, and estimation approaches, and because of the very-different algorithm complexities, is remarkable, the CV-HAR setting leads to comparable results wrt. the NN implementation. The out-of-sample panel of Tab.(ref), on the contrary, is not favoring the NN. On the test-set over which the NN has not been optimized, its actual forecasting performance emerges. The non-linear regression functional that the CV-HAR model approximates seems to lead to better performance measures wrt. the NN, whose approximation appears to be constrained on the training set, not generalizing on new data. This is as stronger as the out-of-sample window size $W$ grows. The growing information conveyed in the joint distribution that the CV-HAR model exploits seem to be valuable in predicting future's RK dynamics.

table[table omitted — 6,432 chars of source]

Out-of-sample analysis

Fixed window

With respect to the fixed window approach, out-of-sample results from Tab.(ref) confirm a general improvement of the ratios favoring the CV-HAR model by the widening of the window size. Also, the set of pair-copulas, allowing for Archimedeans, Gaussian and t-copula, improves the forecasts over the set allowing for of Archimedean copulas only. Similarly, as for in-sample analyses, all the measures at longer windows are supporting the CV-HAR model. There are however exceptions wrt. MSE and R2 measures. In this regard, the best performing window is that of length 750 days. Indeed, as Fig.(ref) illustrates, volatility at the very beginning of 2017 (the sixth year, i.e. sample days after 1250) is particularly quiet, as never in the preceding sample. Under the longest window, the model closely adapts to the dependence structure earlier observed, becoming gradually rigid in capturing deviations from the actual time series's behavior on which it was estimated. Furthermore, margins may provide an unsatisfactory fit, since such low volatility values have been rarely encountered and the left tail of the modeled distribution can potentially provide a poor fit for the actual data.

The grater rations we uniformly observe for the MSE wrt. to MAE and MAD measures are interpretable by the presence of outliers heavily penalizing the squared loss in the MSE measure wrt. to whose the HAR model appears to be less sensitive. Moreover, as for low quantiles, for rare and very high volatilities, both kernel and parametric distributions could provide an inadequate fitting, since they are extrapolated from a sample that is poorly representative of the actual distribution in the very upper quantiles. However, this discrepancy is mild, based on the detected ratios close to unity.

Increasing window

Tab.(ref) reports one-step-ahead forecasts measures for increasing window estimations of the HAR and CV-HAR models. One-step-ahead forecasts are tested to be statistically different by means of the CPA tests. Test statistics and P-values are reported in Tab.(ref). On a general level, results still favor the CV-HAR model. For MAE, MAD, MASE, MAPE and Qlik performance measures, CV-HAR forecasts seem to greatly improve over the HAR ones. In general, AGT copulas are preferred, but easier constructions based on Archimedean copulas only are outperforming the HAR model too. It is not the complexity of the copulas involved in the Vine model that drives the performance, but the flexible CV-HAR model itself seems to provide a more attractive alternative for one-step-ahead forecasting wrt. to the linear specification of the HAR model.

Although MSE ratios are around the unity, by looking at the significance levels of the CPA test (Tab. (ref)), differences in MSE between the two models are largely not statistically significant, however unbalanced in favor of the HAR model. As observed for the fixed window analyses, it appears that the squared loss penalizes the CV-HAR specification more than the HAR one. This is not surprising by observing that on low volatility periods the CV-HAR model provides a very satisfactory fit (Fig.(ref)), while it deteriorates at high volatility regimes. The squared loss shrinks very small residuals and amplifies the large one observed on days of high volatility, leading to an overall squared loss favoring the HAR model. However, the close tracking of the CV-HAR model to the actual volatility observed at low-to-medium volatility days (the vast majority) is much more satisfactory for the CV-HAR model, confirmed both by MAD, MAE, Qlik measures and tests statistics. These are jointly interpretable as an overall ability of the CV-HAR model in forecasting volatility better than HAR in absolute terms, while under a squared loss, high residuals at seldom high-volatility days are covering CV-HAR's overall very satisfactory performance. Indeed the log-term in the Qlik loss highlights the good performance over the majority of days, as opposed to wide residuals on high-volatility days. Hence, the importance of adopting a wide number of performance measures to unveil such behaviors. As for the fixed window analyses, the mean directional accuracy (MDA) is similar between the two models: both the models capture the direction of the realized measure with an accuracy of about 62%. This comment holds for all the sizes of the minimum window.

Rolling window

Lastly, Tab.(ref) reports the results for the rolling window one-step-ahead forecasts, while Tab.(ref) reports the DM and CPA test statistics for a significant difference between HAR and CV-HAR forecast errors. Tab.(ref) shows a clear cut-off between $W=250$ and wider windows, confirmed by the respective DM statistics. Under this window size, a rolling window approach seems to be not feasible for outperforming the HAR forecasts. This is due to the short data span over which the joint distribution of the volatility terms is modeled via Vine copulas, which, in small-samples, might likely provide an unsatisfactory representation, especially in conveying a proper description of the margins at their tails. Results look very different for $W=500$ and $W=750$. MAE and Qlik forecasts are statistically different and favoring the HAR model, whereas differences in MSE appear to be non-significant. Other measures such as MAD, MASE, and MAPE clearly support the CV-HAR model as a preferable choice for one-step-ahead volatility forecasting, as for the increasing window approach. The discussion about close-to-unity MSE ratios opposed to below-the-unity MAE and MAD losses for the increasing window estimation applies here as well. On a general level, we observe coherence between the patterns of strong statistical significance for the DM and CPA tests, further validating our analyses. With AGT pair-copulas there are no major boosts in forecasting performance, suggesting that the simpler modeling estimation involving Archimedean copulas only is a feasible option. Also, models for marginal distributions appear not to play a central role, since holding on a general level: beyond the specific implementation, the CV-HAR alternative seems to capture some non-linear dynamics.

Under RW and IW, the effects of the non-linear flexibility the CV-HAR model appear well-visible. This suggests, along with the in-sample analyses, that the linear specification of the HAR model, through which the conditional expectation of today's volatility given the past terms is implicitly provided by a linear function of the regressors, might be too rigid.

table[table omitted — 4,248 chars of source]

Conclusion

This research shows how the use of a C-Vine copula construction to model the joint distribution of volatility components over different scales can be exploited to model and forecast realized measures. Motivated by the earlier literature investigating the importance of structural breaks and non-linearities within the HAR setting, this paper proposes a non-linear approach that readily exploits the conditional structure arising from the joint distribution of the volatility components involved in the HAR model of corsi2009simple. Following a general regression framework, we do not impose any structure on the conditional expectation (i.e. linearity between the variables) but we recover it from the joint (Vine) distribution. Importantly, this approach guarantees positivity of realized measures' forecasts, standing out as one of the very few approaches (if perhaps not the only one) naturally suitable for non-logarithmic realized measures modeling and forecasting, overcoming, by construction, positivity issues.

This research is inspired by the work of sokolinskiy2011forecasting, but it goes beyond their bivariate framework by fully recalling the HAR modeling spirit. We set apart from sokolinskiy2011forecasting wrt. marginal distribution modeling, multivariate copula construction, analytic computation of the forecasts (opposed to simulation) and forecasting framework.

We apply the CV-HAR method on real high-frequency financial data, by extracting realized kernel intraday volatility measures for 10 stocks, over almost seven years. As a general result, the CV-HAR improves all the performance measures considered, in all the three estimation-forecasting schemes examined. A partial exception, applying in particular to small estimation windows, are MSE point-estimates. Opposed to different performance measures (such as MAD, MAE, Qlike) the quadratic loss seems to be overlooking the very satisfactory performance (small squared residuals) of our approach for the broad majority of low-to-moderate volatility days, while excessively penalizing large losses on days of high volatility.

The CDF modeling plays a secondary role in the results, as well as the copulas involved in the C-Vine construction. Indeed, the CV-HAR specification seems to outperform the HAR alternative, not because of its flexible marginal modeling and copula construction, but rather because the joint distribution modeling approach itself seems to improve over the HAR specification. Our results pinpoint that by relaxing original the linear form of the HAR model by allowing for a more general functional linking the volatility components, considerable in-sample and out-of-sample improvements can be achieved. This suggests that the linear form of the HAR model is restrictive, whereas the conditional expectation of today's volatility - given its past terms directly retrieved from their joint distribution, with no functional assumptions - captures a relationship of more complex nature. Under a fixed window approach we report that a simple neural network (NN) seems not to outperform the CV-HAR specification. This suggests that the information conveyed by the joint distribution plays a central role: the non-linear regressor retrieved from the joint distribution of the lagged volatility terms seems preferable over the complex non-linear regressor retrieved with a NN architecture.

Future research may extend the analyses to different realized measures, include D- and R- Vine specifications, and consider a wider set of copula families. The CV-HAR model can considered in value-at-risk applications since naturally leading to asymmetric confidence intervals and quantiles around the expected conditional mean. Importantly, it would be interesting to widen the current analysis against NN alternatives, and extend it to different models where the non-linear functional is retrieved with machine-learning approaches lebaron2018forecasting. This could further shed light on the role of the rich information on the joint distribution of the lagged volatility terms plays in forecasting.

Acknowledgments

The research leading to this manuscript received funding from the European Union's Horizon 2020 research and innovation program under Marie Sk\lodowska-Curie grant agreement No. 675044. The author is grateful to Professor Juho Kanniainen, Assoc. Professors Marcelo C. Medeiros and Bezirgen Veliyev for their helpful comments and insightful discussions.