Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
83,287 characters · 35 sections · 87 citation commands
Autoencoder Enhanced Realised GARCH on Volatility Forecasting
\listoffigures
\listoftables
\pagenumbering{arabic}
Financial volatility modelling and forecasting are crucial tools for understanding and managing risk in financial markets. The 2008 Global Financial Crisis (GFC) and the COVID-19 pandemic are stark reminders of the devastating consequences of inadequate volatility assessment. During the GFC, flawed volatility models failed to capture the mounting risk within the subprime mortgage market, leading to underestimation of potential losses, institutional failures and widespread economic fallout. Similarly, during COVID-19, the unprecedented volatility drove sharp and extreme price fluctuations in the stock market, resulting in significant losses for investors. These events clearly demonstrate the important link between financial market volatility and the real economy. Effective volatility analysis provides valuable insights for financial institutions and policymakers, acting as a barometer for the vulnerability of financial markets and aiding in risk management and timely interventions. More importantly, advanced volatility forecasting can detect early warning signals, allowing stakeholders to anticipate potential shocks and deal with similar financial crises more effectively in the future. Given these roles in maintaining financial stability, accurate volatility forecasting is a meaningful topic in financial research.
One may ask why we do not focus on predicting returns directly to control outcomes. The challenge lies in the nature of financial returns, specifically of stocks, which are affected by many factors, such as the efficient market hypothesis and the random walk nature of price movement, both of which make predicting returns notoriously challenging McAleer_2008_Realized. Financial volatility, on the other hand, exhibits more predictable behaviours, such as mean reversion, making it mathematically more feasible to forecast than returns Poon_2003_forecasting.
Another complexity is that volatility is a latent variable, it cannot be explicitly observed from a single piece of data Andersen_2009_Realized. Although we can observe closing prices and changes, these are the results of volatility rather than direct measures of it. Volatility is not directly observable as it represents the underlying variability in a market's returns over a period. To estimate this hidden volatility, we infer it from the known values, such as point-to-point returns. Daily squared returns are widely used in financial research as a proxy for volatility, approximating the latent volatility by reflecting the magnitude of price changes. However, the daily squared return is typically calculated at the end of a day using the closing price. They can reflect the volatility observed in the entire previous day, but cannot provide real-time information.
Extensive effort has been dedicated to finding good real-time estimates of current volatility. With the rapidly increasing availability of high-frequency transaction data across financial assets, researchers have developed a more accurate ex-post volatility estimator, realised volatility. The realised volatility calculated from high-frequency intraday returns captures price changes at very fine time scales, such as minute-by-minute or even seconds, providing a near real-time estimate of “true" volatility Andersen_2009_Realized. Now, this measure is widely used as it offers a dynamic view of volatility in near real-time, offering a more immediate estimate of volatility than daily returns.
With these measures in place, researchers and practitioners also recognise that volatility has both persistence and heteroskedasticity features Poon_2003_forecasting, Engle_2004_Risk. To capture these characteristics, volatility measures are often integrated into advanced econometric models, such as Generalised Autoregressive Conditional Heteroskedasticity (GARCH) bollerslevGeneralizedAutoregressiveConditional1986 and Realised GARCH models hansenRealizedGARCHJoint2012, which employs squared returns and realised volatility, respectively, to account for the inherent latent and dynamic nature of volatility Andersen_2009_Realized.
With realised volatility becoming increasingly prevalent, numerous measures for realised volatility have been developed, such as Realised Variance, Realised Kernel, and Bipower Variation, et al. This variety leads to an initial research question: which realised volatility measure should be selected for volatility forecasting analysis? Naimoli_2022_Improving handle this by generating a synthetic realised measure from a range of candidate measures using linear combination techniques, including Principal Component Analysis and Independent Component Analysis, and incorporate them into the Realised GARCH model to improve the accuracy of volatility forecasting. Inspired by the idea of combining various measures into a single synthetic measure, and recognising the limitations of linear transformations, we propose to explore the potential benefits of a nonlinear combination for generating this synthetic measure. This leads to the final research question: does employing a nonlinear dimension reduction approach enhance the effectiveness of the synthetic realised volatility measure in Realised GARCH compared to linear methods?
Autoencoder, as a deep learning algorithm based on neural network architecture, has an outstanding ability for nonlinear dimension reduction and numerous applications in financial time series analysis. In this context, our research objective is designed to: develop a framework that uses nonlinear dimension reduction through an autoencoder to generate a synthetic realised volatility measure with essential information and incorporate it into Realised GARCH for volatility forecasting.
The primary contribution of this thesis lies in the novel use of autoencoder-based nonlinear dimension reduction to create a flexible synthetic measure. This enhanced measure enables Realised GARCH to achieve more adaptable parameter estimations at each forecasting step, ultimately improving volatility forecasting accuracy compared to linear methods.
The rest of this thesis is presented as follows: Chapter 2 reviews the relevant literature on volatility forecasting models, realised volatility measures, and dimension reduction techniques. Chapter 3 presents the essential knowledge of background models, along with the proposed framework and its parameter estimation method. Chapter 4 provides an empirical study of the proposed model, comparing in-sample estimates and out-of-sample forecasting performance against competing models. Chapter 5 summarises the conclusions, discusses limitations, and suggests future research directions.
Since the groundbreaking paper of engleAutoregressiveConditionalHeteroscedasticity1982, Autoregressive Conditional Heteroskedasticity (ARCH) offers a new perspective for understanding volatility. bollerslevGeneralizedAutoregressiveConditional1986 further expands it to include the lagged conditional variance to predict future volatility, developing the Generalised ARCH (GARCH). Since their introduction, ARCH and GARCH have been recognised as foundational approaches to financial volatility modelling and are commonly used to describe and predict volatility. Specifically, GARCH uses squared daily returns as a proxy for the ex-post volatility. Let innovation in returns be written as $r_t=\sigma_t z_t$, where $z_t$ represents an independent stochastic process with mean zero and unit variance, while the latent volatility $\sigma_t$ evolves according to the specific model formula. This approach provides an unbiased estimate of volatility as $E(r^2_{t}) = E(\sigma^2_{t} z^2_{t}) = \sigma^2_{t} $. However, it is prone to noise from idiosyncratic errors $z^2_{t}$, especially during periods of significant volatility changes andersenAnsweringSkepticsYes1998. This noise limits the model's effectiveness in adjusting to new levels of volatility, highlighting the need for more efficient volatility measures based on high-frequency data hansenRealizedGARCHJoint2012.
As high-frequency financial data become increasingly available and standard in volatility forecasting. The latent volatility can be measured by analysing the variability of intraday squared returns based on high-frequency data, which provides more information than daily squared returns andersenModelingForecastingRealized2003. Therefore, Realised Variance (RV) was introduced by andersenAnsweringSkepticsYes1998, as a realised volatility measure, dramatically outperforming daily squared returns. In addition to RV, several more informative measures for volatility have been proposed, such as the Realised Kernel (RK) Barndorff-Nielsen_2002_Econometric, and Bipower Variation (BV) barndorff-nielsenPowerBipowerVariation2004.
RV is calculated by summing the intraday squared returns, a method that has been extensively used for volatility measurement. However, high-frequency data based RV is contaminated by the market microstructure effect, such as discreteness of prices, bid-ask bounce and irregular trading, which introduce noise into the data. This noise arises because the microstructure effect is not reflective of actual price changes but from the mechanics of trading. As a result, using high-frequency data may lead to inaccurate volatility estimates as the observed fluctuations are affected by factors unrelated to “true" volatility. On the other hand, using low-frequency data may miss important volatility information from price changes. Therefore, RV faces the problem of deriving the appropriate frequency to balance the noise impact and the loss of information christensenRealizedRangebasedEstimation2007.
Barndorff-Nielsen_2002_Econometric focus on kernel-based estimators and design the class of RK. RK estimates the quadratic variation from high-frequency noisy data, similar to the sum-of-squares variance estimator. However, it combines the estimation of intraday volatility and kernel smoothing, resulting in robustness and efficiency in the presence of market frictions, therefore being more robust to microstructure noise Barndorff-Nielsen_2008_Designing.
BV is introduced by barndorff-nielsenPowerBipowerVariation2004 as a realised volatility measure that effectively distinguishes between the continuous and jump components of price movements. Unlike RV, which captures both continuous price movements and rare price jumps, BV is designed to specifically capture the continuous component of the price process, calculated by summing the products of absolute returns over adjacent time intervals. The difference between RV and BV, known as the jump variation, provides an estimate of the jump component of volatility, capturing the effect of large and infrequent price changes which is a key element in volatility modelling barndorff-nielsenPowerBipowerVariation2004.
In addition, various other realised measures exist and serve as candidate measures in this thesis. Realised Semivariance (RSV), which divides RV into downside and upside, emphasises the magnitude of downside risk Barndorff-Nielsen_2008_Measuring. Median Realised Variance (MedRV) squares the median absolute return among three consecutive intraday returns, asymptotically avoiding the impact of a jump in the measure while reducing the sensitivity to the extreme return values caused by the microstructure effect within the trading day andersenJumprobustVolatilityEstimation2012. Two-scale Realised Kernel (TwoScale RK) proposed by ikedaTwoScaleRealizedKernels2015 is a bias-corrected realised kernel estimator, given by a particular convex combination of two RK estimators with different bandwidths.
Although some of the measures introduced above can partially mitigate the impact of microstructure, their reliance on high-frequency data still exposes them to microstructure noise christensenInferenceHighfrequencyData2017. Consequently, the techniques for smoothing out noise can also be used in realised volatility measures.
The subsampling approach introduced by Zhang_2005_Tale involves taking samples every 5 minutes, starting with the first observation, then the second, and so forth. Then averaging the results across those subgrids significantly decreases bias. The Parzen Weighted RK introduced by barndorff-nielsenRealisedKernelsPractice2008 assigns varying weights to observations based on their distance from the centre of the sample, focusing on central and reducing noisy observations, thus dampening the influence of high-frequency noise.
Based on the informative characteristics of the realised volatility measures, they provide significant improvements in modelling volatility compared to squared returns. These improvements have led researchers to incorporate realised measures directly into volatility models. engleNewFrontiersArch2002 includes a realised measure like RV in the GARCH framework (GARCH-X), effectively illustrating the feasibility of using RV in forecasting volatility. In the GARCH-X model, the high-frequency data-based realised measure is used as an exogenous variable (X).
Building on GARCH-X, hansenRealizedGARCHJoint2012 introduce the Realised GARCH model (RealGARCH), which addresses the gap by incorporating an explicit volatility explanation within the realised measure. RealGARCH adds a measurement equation, relating the observed realised volatility as a dependent variable on contemporaneous latent volatility. watanabeQuantileForecastsFinancial2012 verifies the predictive likelihood improvement of RealGARCH over GARCH and GARCH-X models.
To extend RealGARCH further, hansenExponentialGARCHModeling2016 develop the Realised Exponential GARCH (RealEGARCH), allowing the incorporation of multiple realised volatility measures through multiple measurement equations. Compared with RealGARCH, an additional feature of the RealEGARCH model is the introduction of two leverage terms within both GARCH and measurement equations.
Given the numerous volatility measures proposed in the existing literature, one of the new questions is which realised measure should be selected in the model. Although RealEGARCH can incorporate multiple realised measures for volatility, this added additional flexibility is achieved at the price of a higher-dimensional problem Naimoli_2022_Improving.
To overcome the high-dimensional problem when considering multiple realised measures, dimension reduction can be employed as a solution. There are many famous linear dimensionality-reduced techniques, such as Principal Component Analysis (PCA) and Independent Component Analysis (ICA). PCA is a nonparametric method that is applied to reduce the dimensionality of the data while maintaining as much variability as possible. It was invented by Pearson_1901_LIII and further developed by Hotelling_1933_Analysis. Given the calculation of eigenvectors of the covariance matrix of the input data, PCA linearly transforms high-dimension inputs into low-dimension ones whose components are not correlated Cao_2003_comparison. In addition, these components are also ordered according to the variance, which means that the first principal component can explain the largest variance Tahmasebi_2020_Machine.
ICA is considered as an extension of PCA. Instead of focusing on optimising the covariance matrix of the data, ICA optimises higher-order statistics such as kurtosis Tharwat_2021_Independent. ICA targets independent components and can extract statistically independent sources when the higher-order correlations are insignificant Hyvarinen_2000_Independent. Hyvarinen_1999_Fast propose an algorithm of ICA called FastICA, which extracts the independent component by maximising non-Gaussianity. Later, ICA was generalised to a well-known method like PCA for feature extraction Jang_1999_Feature.
With the promotion of PCA and ICA, there are many applications of these techniques applied in Multivariate Factor GARCH. For PCA, such as Orthogonal-GARCH Alexander_2001_Market, Generalised Orthogonal-GARCH VanDerWeide_2002_GO-GARCH, and conditionally uncorrelated component model Fan_2008_Modelling. For ICA, it has been widely applied in financial econometrics, such as papers of Chen_2007_Portfolio and Garcia-Ferrer_2012_conditionally, among others.
In order to overcome the problem of selecting the optimal set of realised volatility measures and integrating it in a univariate RealGARCH framework, Naimoli_2022_Improving propose an extension of RealGARCH with the use of PCA and ICA, respectively, denoted as PC-RealGARCH and IC-RealGARCH. By incorporating dimension reduction techniques, a synthetic realised volatility measure is generated from a wide range of candidate realised measures, and it is used to substitute a single realised measure in RealGARCH. The performance of PC-RealGARCH and IC-RealGARCH with a mixed realised measure leads to significant accuracy gains in tail risk forecasting Naimoli_2022_Improving.
While PCA is widely used for dimension reduction, its linear structure may reduce its effectiveness in capturing the complexity of nonlinear data. In this case, nonlinear methods such as locally linear embedding (LLE) Roweis_2000_Nonlinear, stochastic neighbour embedding (SNE) Hinton_2002_Stochastic, and autoencoder Hinton_2006_Reducing were developed.
In this thesis, we employ the autoencoder to reduce the dimension of realised measures nonlinearly. Because autoencoder is a deep learning approach based on a neural network architecture that can use local adjustments in connection weights to reduce the multidimensional data of raw input to a cohesive and lower dimension, and then reconstruct this transformed data back to its original form as accurately as possible Baldi_2012_Autoencoders. By minimising the distortion, the layers of autoencoder work together to capture the essential information, forming a self-organisation process.
The autoencoder is also widely used to handle dimension reduction problems in the financial time series field. Cortes_2024_Autoencoder-Enhanced develop an innovative clustering framework leveraging autoencoder to compress financial time series data. Bao_2017_deep propose stacked autoencoders for financial time series, which consisting multiple single-layer autoencoders where the output of each layer is successively transformed to the following layer. Zhang_2020_Stacked combine the autoencoder with the ensemble learning algorithm to predict stock returns, reducing the dimension, eliminating redundant noise, and resulting in more robust prediction performance.
Building on the robust volatility modelling capabilities of RealGARCH hansenRealizedGARCHJoint2012, we choose it as our foundation of the volatility model. Starting with our initial question of selecting the optimal realised volatility measure from the wide range of available candidate measures to incorporate into RealGARCH, we propose moving beyond an individual measure by generating a synthetic measure. Motivated by PC-RealGARCH and IC-RealGARCH Naimoli_2022_Improving, which illustrate combining measures through linear dimension reduction can enhance accuracy. Our objective is to integrate a nonlinear approach to aggregate information from multiple realised measures more flexibly. In particular, we employ an autoencoder which is well-suited for dimension reduction and financial time series, to synthesise and embed a realised measure within RealGARCH, enhancing the volatility modelling and forecasting performance.
Building on the theoretical foundation of GARCH-type models and drawing inspiration from the successful dimension reduction using PCA and ICA on RealGARCH Naimoli_2022_Improving, we propose an autoencoder enhanced RealGARCH. This approach aims to use the synthetic realised measure generated by the autoencoder to further improve volatility forecasting performance, compared to existing methods that rely on linear combinations.
The standard GARCH model can be written as:
where $r_t=\left[\log C_t-\log C_{t-1}\right] \times 100$, representing the scaled percentage log return, calculated from the closing price $C_t$ on day $t$ and $C_{t-1}$ on day $t-1$. $\sigma_{t}$ denotes the conditional volatility on day $t$. The random error term $z_t$ is distributed independently and identically, as a normal distribution with zero mean and one variance $z_t \stackrel{\text { i.i.d. }}{\sim} N(0,1)$. $\omega > 0$, $\alpha \geq 0$, $\beta\geq0$, as positivity condition for conditional variance. and the stationary condition requires $\alpha + \beta < 1$.
Inspired by the more informative advantage of realised volatility measures, replacing the squared returns with a realised measure in Equation ((ref)), leads to the construction of the GARCH-X model:
where $x_{t-1}$ is a realised volatility measure such as RV. Both $x_{t-1}$ and $r_{t-1}^2$ are expressed on the variance scale, as they quantify the magnitude of squared fluctuations.
The GARCH model traditionally uses squared returns as a proxy for hidden volatility. GARCH-X improves it by replacing returns with the realised volatility measure. The key contribution of RealGARCH is introducing a measurement equation that links realised volatility with latent volatility, further improving the volatility modelling framework. With its log specification, RealGARCH ensures a strictly positive conditional variance and stabilises the variance process. Here, we use RealGARCH with log specification to integrate the synthetic realised measure and forecast volatility. RealGARCH models the volatility through three key components: the return equation, the GARCH equation, and the measurement equation. Three equations, presented in the order below, form a cohesive volatility framework:
where $\sigma_t^2$ is always positive because of the log transformation, the sign of $r_t$ is the same with $z_t$, ensuring they share the same positive or negative direction. Here, $z_t \stackrel{\text { i.i.d. }}{\sim} N(0,1), \varepsilon_t \stackrel{\text { i.i.d. }}{\sim} N(0,1)$, $\sigma_\varepsilon$ is the standard deviation of error term $\varepsilon_t$. $x_t$ is a realised volatility measure, and the term $\tau_1 z_t+\tau_2\left(z_t^2-1\right)$ is regarded as the leverage function, capturing the asymmetric influence of positive and negative returns $r_t$ on volatility. More specifically, the asymmetric impact arises from the term $\tau_1 z_t$, depending on whether $z_t$ is negative or positive, while the term $\tau_2\left(z_t^2-1\right)$ remains unchanged to different signs of $z_t$ as $z_t^2$ yields the same value regardless.
In the measurement equation, observed realised volatility measure $x_t$ is related to contemporary hidden variance $\sigma_t^2$ using a linear regression with leverage effect term $\tau_1 z_t+\tau_2\left(z_t^2-1\right)$ and a random error term $\sigma_{\varepsilon} \varepsilon_t$.
In addition to the model structure, there are some conditions that constrain the RealGARCH model. Since RealGARCH with log specification has already ensured a positive conditional variance $\sigma_t^2$, the stationary condition remains justified in this model.
Combining the measurement equation into the GARCH equation in Equation ((ref)) results in an autoregressive process of order 1 (AR(1)), incorporating a lag 1 log conditional variance:
where $\mu=\omega+\gamma \xi$, $\pi=\beta+\gamma \varphi$, and $w_{t}=\gamma \left[ \tau_1 z_t+\tau_2\left(z_t^2-1\right)+\sigma_{\varepsilon} \varepsilon_t \right]$. Since $E (w_{t-1}) =0$, the stationary condition for an AR(1) model requires that $ \left| \pi \right| < 1$, so that the required condition for the RealGARCH is:
The above discussed how RealGARCH models volatility using a realised measure. However, recent research has made it increasingly available to access a growing range of volatility estimators such as RV, RK, BV, et al., as introduced in Section (ref). Therefore, Naimoli_2022_Improving propose to replace the single realised measure in RealGARCH with a combined measure obtained by aggregating the principal information from a wide set of different realised volatility measures using PCA and ICA, respectively.
Given a times series $\{\mathbf{x}_t \in \mathbb{R}^D \}_{t = 1}^T$, where $\mathbf{x}_t^{\intercal}$ represents a $D$-dimensional row vector of realised volatility measures at time $t$. The entire series over period $T$ can be written as a matrix $\mathbf{X} \in \mathbb{R}^{T \times D}$ (usually $T>D$):
PCA first conducts the eigen decomposition on the covariance matrix of $\mathbf{X}$, generating an orthogonal matrix $\mathbf{U}$. Then, linearly transforms the input measures matrix $\mathbf{X}$ into a principal component matrix $\mathbf{P}$ by
where $\mathbf{U}\in \mathbb{R}^{D\times D}$, is the orthogonal matrix whose $i$th column $\mathbf{u}_i\in \mathbb{R}^D$ is the $i$th eigenvector. Eigenvectors provide the weights for each principal component and show the direction in the feature space along which the data varies the most. These weights (eigenvectors) determine how much each realised volatility measure contributes to the principal component.
Since the first principal component can explain the largest variance as illustrated in Tahmasebi_2020_Machine, Naimoli_2022_Improving use the first principal component generated from PCA as the synthetic measure in univariate RealGARCH. The first principal component vector $\mathbf{x}_{PC}$ can be written as:
where only the first column in $\mathbf{U}$ (first eigenvector), $\mathbf{u}_1$, is used as the weight to project the realised measures. This results in the vector $\mathbf{x}_{PC}\in \mathbb{R}^T$, which represents a series of linearly combined measures derived from the realised measure matrix and the first weights. Finally, $\mathbf{x}_{PC}$ is rescaled to match the range defined by the maximum and minimum values of $\mathbf{X}$:
where each element $x_{PC,t}$ in $\mathbf{x}_{PC}$ represents the projected first principal component at time $t$.
The PC-RealGARCH replaces the single realised measure $x_t$ with the first principal component $x_{PC,t}$, which can be defined as:
Different from PCA, which attempts to find an orthogonal linear transformation of realised measures, ICA operates based on the inverse of the central limit theorem (CLT), decomposing independent components by maximising their non-Gaussianity. Consequently, ICA eliminates the correlations and higher-order dependencies, whereas PCA only removes correlations Naimoli_2022_Improving.
One necessary condition for ICA to work is that it requires non-Gaussianity in the underlying sources Tharwat_2021_Independent. Since the realised volatility measures are inherently computed on the variance scale, strictly non-negative, and exhibit clear non-Gaussian behaviour, this asymmetric non-Gaussianity allows ICA to explore the independent directions for each component. FastICA, as introduced in Section (ref), is therefore used here as an alternative to PCA and a suitable technique when applied in dimension reduction for realised volatility measures.
FastICA estimates the un-mixing matrix $\mathbf{V}$ by maximising the non-Gaussianity of the components Hyvarinen_1999_Fast. The independent component matrix $\mathbf{S}$ is obtained by:
where $\mathbf{V}\in \mathbb{R}^{D\times D}$, is the un-mixing matrix whose $i$th column is $\mathbf{v}_i\in \mathbb{R}^D$, analogous to $\mathbf{u}_i$ in PCA, playing the role of weights. Similarly, the first independent component vector $\mathbf{x}_{IC}\in \mathbb{R}^T$ is obtained by the product of the $\mathbf{X}$ and $\mathbf{v}_1$,:
Then, $\mathbf{x}_{IC}$ is also rescaled according to Equation ((ref)) to the range of maximum and minimum values of realised measures $\mathbf{X}$, similar to the re-scaling process of $\mathbf{x}_{PC}$ from PCA. Each element $x_{IC,t}$ in $\mathbf{x}_{IC}$ represents the first independent component at time $t$. Finally, replace the $x_t$ with $x_{IC,t}$ in RealGARCH to construct IC-RealGARCH:
In the spirit of PC-RealGARCH and IC-RealGARCH, a simpler combined realised measure generated from directly taking an average of a range of different realised measures can be used to replace the single measure in RealGARCH. This combined realised measure $\mathbf{\Bar{x}}=\{\Bar{x}_t\}_{t=1}^T$ is calculated as the average of $D$ realised measures:
where $x_{t d}$ is the element in $\mathbf{X}$. Incorporating the averaged realised measure $\Bar{x}_t$ in the RealGARCH results AVG-RealGARCH:
The autoencoder is a special neural network model designed to generate the output that approximates the input data. This autoencoder is composed of two main components: an encoder and a decoder Gu_2021_Autoencoder. The inputs are transformed through a relatively smaller number of neurons in the hidden layer(s) than the dimension of inputs, resulting in a low-dimensional encoded representation of the inputs. This process is called encoding and happens in the “encoder". The encoded data are then transformed through a “decoder" network with an opposite structure of “encoder", to recover the encoded data, resulting in a reconstruction of inputs. The whole system is called an “autoencoder". Since the autoencoder generates output to approximate its input, no labelled data is needed. The autoencoder is trained as an unsupervised process. However, the training process is still based on the optimisation of minimising the discrepancy between the input data and its reconstruction in the output Hinton_2006_Reducing.
Autoencoder with L hidden layers can be written as the recursive formula. Denote $K^{l}$ as the number of neurons in layer $l$, where $l=1,...,L$. The input and output of a single neuron $k$ in layer $l$ can be written as:
where $s^{l}_k$ is the input sent to neuron $k$, $o_k^{l}$ is the output. $w_{i k}^{l}$ is the weight that connects neuron $i$ in layer $l-1$ to neuron $k$ in layer $l$, $b_{k}^{l}$ is the bias for neuron $k$. The input is transformed through an activation function $g^{l}(\cdot)$, becoming the output and being passed to the next layer. Define the vector of all inputs to layer $l$ as $\mathbf{s}^{l}\in\mathbb{R}^{K^l}$, and outputs of this layer as $\mathbf{o}^{l}\in\mathbb{R}^{K^l}$. The output in layer $l$ can then be written in a matrix notation as
where $\mathbf{W}^l\in \mathbb{R} ^{K^{l} \times K^{l-1}}$, is a matrix of weight parameters, and $b^{l}$ is a $K^{l} \times 1 $ bias vector. The final output at layer $L$ of the autoencoder is:
where the final output has the same dimension as the input.
Ideally, an autoencoder could employ multiple hidden layers and neurons, allowing for deep learning and more nuanced feature extraction. However, increasing network complexity can make training significantly harder. Without proper pretraining, a deep autoencoder often fails to capture patterns effectively, instead outputting a result that approximates the average of the input data. As Hinton_2006_Reducing note, this issue can persist despite extensive fine-tuning, whereas a shallower autoencoder with a single hidden layer can learn without the need for pretraining. Given the complexity and time-consuming nature of the deep autoencoder, this thesis employs a single hidden-layer autoencoder as a starting point to explore the effectiveness of nonlinear dimension reduction, with plans to extend to a deeper architecture in future work.
In this design, the layer between the input and the hidden layer works as an encoder, transforming the input to the hidden representation. The layer between the hidden layer and the output works as a decoder, reconstructing the original input. To initialise, the input is the realised volatility measures vector $\mathbf{x}_t$, $t=1,...,T$, as defined in Equation ((ref)). Thus, $\mathbf{o}^{0}=\mathbf{x}_t\in \mathbb{R}^D$ with $K^0=D$. There is one neuron in the hidden layer $K^1 = 1$, as its output is used to incorporate into univariate RealGARCH. We use the sigmoid function as nonlinear activation functions $g^l(\cdot)$ in two layers, $g^l(\mathbf{s}^l):=sig(\mathbf{s}^l)=1/(1+e^{-\mathbf{s}^l})$. Let $\mathbf{x}_t$ denote the initial input $\mathbf{o}^0$, $x_{AE,t}$ denote the output of the encoder layer $\mathbf{o}^1$, and $\hat{\mathbf{x}}_t$ denote the output of the decoder layer $\mathbf{o}^2$. The autoencoder in this thesis can be written as:
where $x_{AE,t}$ is a encoded value (scalar) at time $t$, $\mathbf{w}^1$ is a $1 \times D$ vector of weight parameters for the encoder, and $b^1$ is a scalar of bias parameter. $\mathbf{w}^2$ is a $D \times 1$ vector of weight parameters for the decoder, and $\mathbf{b}^2$ is a $D \times 1$ vector of bias parameters. $\mathbf{x}_t\in \mathbb{R}^D$ and $\hat{\mathbf{x}}_t\in \mathbb{R}^D$. Over period $T$, the autoencoder generates a synthetic realised measure vector $\mathbf{x}_{AE}\in \mathbb{R}^T$, composed of scalar outputs $x_{AE,t}$ for each time $t$. Like $\mathbf{x}_{PC}$, $\mathbf{x}_{IC}$ and $\mathbf{\Bar{x}}$, $\mathbf{x}_{AE}$ serves as the synthetic realised measure series. But $\mathbf{x}_{AE}$ is obtained through nonlinear dimension reduction of realised measures, different to the linear dimension reduction in PCA, ICA and average. Figure (ref) illustrates the architecture of the autoencoder employed in this thesis, where the input and output are displayed in a horizontal layout for visualisation purpose.
Empirically, the weights and biases of the autoencoder can be estimated by minimising the loss function $L$ between input $\mathbf{x}_t$ and its reconstruction $\mathbf{\hat{x}}_t$, we adopt Mean Squared Error (MSE) here:
where $x_{t d}$ represents the $d$th feature of the input realised measures input $\mathbf{x}_t$ at time $t$, while $\hat{x}_{t d}$ represents the corresponding reconstructed feature at time $t$.
The Autoencoder shares a similar essence with neural networks. The high capacity of neural networks gives the autoencoder more flexibility in extracting significant features from the data. However, this highly flexible characteristic is directly related to the complexity of the model. The standard autoencoder is trained based on minimising the loss function in Equation ((ref)), which typically leads to the underlying overfitting problem. Motivated by Gu_2021_Autoencoder, we considered the use of regularisation to prevent overfitting in two aspects: weights and sparsity.
The most common type of regularisation is to add a complexity penalty to the optimisation objective. We define the optimisation objective as:
where $\lambda_1$ is a hyperparameter and $C(\cdot)$ is the regulariser that measures the complexity of the model as a function of weight parameters.
There are many choices for penalty functions, such as Ridge and Lasso. Given that the input data consists of realised volatility measures, which are different approaches to measuring realised volatility, these features are closely relevant to each other. Compared to Lasso, which assigns substantial coefficients to a small number of features, Ridge is more suitable in this context, as it assigns similar and small coefficients to all features, better capturing the correlated relationships among measures. Therefore, we use Ridge regularisation which takes the form:
where $w_{i 1}^{1}$ is the weight connecting input $i$ and the neuron in the first layer, $w_{1 k}^{2}$ is the weight connecting the neuron and output $k$ in the second layer. The Ridge regularisation term is the sum of squared elements of the weight vectors for each layer.
In addition to weight regularisation, sparsity regularisation also encourages the autoencoder to find a more general and robust representation of the input data by constraining the activation value. This helps prevent the autoencoder from fitting to noise and learning undesired variations, ensuring that the encoded data captures consistent patterns in realised measures instead of noise. This regularisation works on the average output activation value of the neuron in the hidden layer, the average of this neuron from the training samples with a size of $T$ can be expressed as:
The sparsity regularisation aims to restrict the value of $\hat{\rho}$ to be low, preventing the encoded data from fluctuating significantly when the original data do not show extreme variations.
The sparsity regularisation is designed to give a large penalty when the average activation value $\hat{\rho}$ is not close to the settled value $\rho$. We use the Kullback-Leibler divergence as the function, which has the form:
where KL provides a value of 0 only when $\rho$ and $\hat{\rho}$ are equal to each other, but provides a large value when they diverge from each other. Therefore, minimising this sparsity regularisation term enforces $\hat{\rho}$ approaching $\rho$ as close as possible.
Finally, the parameters in the autoencoder $\mathbf{w}^1$, $b^1$, $\mathbf{w}^2$ and $\mathbf{b}^2$ for a dataset with size $T$ can be estimated by minimising the overall loss function $L$, which is designed as:
where $\lambda_1$ and $\lambda_2$ control the magnitude of the penalty in the loss function, and $\rho$ controls the desired value of the average activation value.
The autoencoder-produced realised measure series $\mathbf{x}_{AE}$, consisting of $\{x_{AE,t}\}_{t=1}^T$, can be obtained by first estimating parameters through minimising the loss function in Equation ((ref)), and then applying these estimated parameters to the autoencoder formula in Equation ((ref)). Similarly to realised measure series $\mathbf{x}_{PC}$, $\mathbf{x}_{IC}$ and $\mathbf{\Bar{x}}$, $\mathbf{x}_{AE}$ includes a linear weighted combination process but also has an additional nonlinear transformation process using the sigmoid activation function. Additionally, in terms of decomposition purposes, PCA and ICA are designed to produce uncorrelated and independent components, respectively. However, the autoencoder encoded the realised measures into a synthetic measure which can be decoded back to the original realised measures with minimum accuracy loss. Here, we denote $\mathbf{x}_{PC}$, $\mathbf{x}_{IC}$, $\mathbf{\Bar{x}}$ and $\mathbf{x}_{AE}$ in row vector formats as shorthands for their transposed forms for later reference.
The nonlinear dimensionality-reduced realised measure $\mathbf{x}_{AE}$ is then rescaled by Equation ((ref)), to fall within the typical range of realised measures. Replace the single realised measure in RealGARCH, we propose an autoencoder enhanced RealGARCH (AE-RealGARCH):
In the empirical study section, we will compare the volatility forecasting performance of AE-RealGARCH, PC-RealGARCH, IC-RealGARCH and AVG-RealGARCH to evaluate the improvements offered by the autoencoder-produced synthetic measure.
The AE-RealGARCH, PC-RealGARCH, IC-RealGARCH and AVG-RealGARCH models employ different processes on realised measure $x_t$, but their underlying volatility modelling principles are all rooted in the RealGARCH framework. The estimation of these models adheres to the procedures established for RealGARCH, as outlined by hansenRealizedGARCHJoint2012, where parameters as estimated by maximising the likelihood function. Accordingly, the following section presents the likelihood function for RealGARCH as a representative example.
Given the return $r_t$ and realised volatility measure $\mathbf{x}_t$ can be observed at day $t$, the Maximum Likelihood (ML) can be used here to find the parameters in RealGARCH that make these two observed data most probable under our specified model. Following the approach of hansenRealizedGARCHJoint2012, the joint likelihood in RealGARCH is equal to the sum of two likelihood functions.
Let $\varepsilon_t$ denote $\sigma_{\varepsilon} \varepsilon_t$ here for consistent presentation, we assume Gaussian distributions $z_t \stackrel{\text { i.i.d. }}{\sim} N(0,1)$, and $\varepsilon_t \stackrel{\text { i.i.d. }}{\sim} N(0,\sigma_{\varepsilon}^2)$. Therefore the zero-mean returns are given by $r_t=\sigma_t z_t$ and the “de-meaned" realised measure is modelled as: $\log(x_t)-\xi-\varphi \log(\sigma_t^2)-\tau_1 z_t-\tau_2\left(z_t^2-1\right)=\varepsilon_t$. The distributions can be expressed as: $r_t \sim N\left(0, \sigma_t^2\right)$, $\varepsilon_t \sim N\left(0, \sigma_{\varepsilon}^2\right)$. In this case, the probability density function (PDF) of $r_{t}$ and $\varepsilon_t$ are:
The joint log-likelihood function for RealGARCH, omitting constant terms, can then be written as the sum of two log-likelihood functions for returns $\mathbf{r}$ and realised measure $\mathbf{x}$:
where $\mathbf{r}=\left\{r_1,...,r_T \right\}$, $\mathbf{x}=\left\{x_1,...,x_T \right\}$. In AE-RealGARCH, replace $\mathbf{x}$ with $\mathbf{x}_{AE}$; in PC-RealGARCH, replace $\mathbf{x}$ with $\mathbf{x}_{PC}$; and similarly, use $\mathbf{x}_{IC}$ and $\mathbf{\Bar{x}}$ for IC-RealGARCH and AVG-RealGARCH, respectively. $T$ is the training sample size. $\boldsymbol{\theta}$ summarise the parameters $\boldsymbol{\theta} = [\omega, \beta, \gamma, \xi, \varphi, \tau_1, \tau_2, \sigma_\varepsilon]^{\intercal}$. $l(\mathbf{r};\boldsymbol{\theta})$ and $l(\mathbf{x}\mid\mathbf{r};\boldsymbol{\theta})$ represent the log-likelihood functions for returns component and realised measure component, respectively. Consequently, the parameters in the RealGARCH model can be estimated by maximising the log-likelihood function $l(\mathbf{r,x}; \boldsymbol{\theta})$, while considering the constraint in Equation ((ref)).
It is worth noting that, GARCH-type models such as GARCH and GARCH-X do not have the log-likelihood function for realised measure $\mathbf{x}$, their parameters $\boldsymbol{\theta}_{garch}=[\omega, \alpha, \beta]^{\intercal}$ are estimated based on the maximisation of $l(\mathbf{r};\boldsymbol{\theta})$ only. In this case, differences in the log-likelihood function specifications between RealGARCH and GARCH-type models make the direct comparison of the overall maximised log-likelihood challenging. In this case, the common log-likelihood term for returns $l(\mathbf{r};\boldsymbol{\theta})$, allows for a meaningful comparison for assessment purpose Naimoli_2022_Improving.
Since the optimisation objective includes an inequality constraint as in Equation ((ref)), Sequential Least Squares Programming (SLSQP) is used to solve the constrained optimisation problem. Developed by kraft1988software, SLSQP is implemented as an optimisation method in Python's Scipy library. Compared to standard gradient descent, SLSQP does not require specifying the learning rate and batch size at the beginning. Known for reaching global convergence and superlinear convergence speed Xie_2020_Energy, SLSQP ensures the resulting parameters satisfy the constraint condition in Equation ((ref)).
All GARCH-type models are estimated using the Python code developed for this thesis. The synthetic realised measures from PCA and ICA, $\mathbf{x}_{PC}$, $\mathbf{x}_{IC}$, are computed using Scikit-learn library in Python, while $\mathbf{x}_{AE}$ is obtained with Matlab's Deep Learning Toolbox.
Four daily international stock market indices are analysed in this study: the S&P 500 (US); FTSE 100 (UK); AORD (Australia); Hang Seng (Hong Kong). Data for the daily closing price index from January 2000 to June 2022 are collected from The Oxford-Man Institute's Realised Library. These closing prices are used to calculate daily returns. The daily percentage log returns are calculated as $r_t = [\log(C_t)-\log(C_{t-1})]\times100$, where $C_t$ is the closing price on day $t$. The returns are then de-meaned. Since the return $r_t$ requires the closing price $C_{t-1}$ on the previous day, the first row of each market data is removed to match the series of returns. The market-specific non-trading days are removed from each market data. After these adjustments, the whole dataset sizes for four markets are approximately 5600 days, with the largest size being 5668 days in AORD and the smallest size being 5504 days in Hang Seng.
In addition, 12 realised volatility measures introduced in Section (ref), are included in the dataset: Realised Variance (RV) andersenAnsweringSkepticsYes1998, Realised Semivariance (RSV) Barndorff-Nielsen_2008_Measuring, Median Realised Variance (MedRV) andersenJumprobustVolatilityEstimation2012, Realised Kernel (RK) Barndorff-Nielsen_2002_Econometric, Two-scale Realised Kernel (TwoScale RK) ikedaTwoScaleRealizedKernels2015, Parzen Weighted Realised Kernel (Parzen RK) barndorff-nielsenRealisedKernelsPractice2008, Bipower Variation (BV) barndorff-nielsenPowerBipowerVariation2004, computed at 5 and 10-minute frequencies and combined with subsampling technique Zhang_2005_Tale, resulting 12 different realised measures. These 12 realised measures are calculated based on the high-frequency data within the trading time, and they are used as the input to generate the synthetic realised measure using PCA, ICA, average, and autoencoder, respectively. Hence, the input dimension $D$ in Equation ((ref)) is 12 in this thesis.
Table (ref) reports the summary statistics of daily percentage log returns and 12 log-transformed daily realised volatility measures in the S&P 500. Similar patterns are observed in the other three markets, where the daily percentage log returns following a de-mean process have a mean value of 0, exhibiting positive excess kurtosis, persistent heteroskedasticity, and negative skewness. The original realised volatility measures are all positive and obvious right-skewed distributions, the log-transformed realised measures display the properties of mitigated positive skewness and positive excess kurtosis. Furthermore, the statistics for the realised measures closely align with those of the sub-sampled realised measures.
Figure (ref) displays the plot of absolute log percentage returns and the square root of 5-min RV for the full dataset of S&P 500. The consistent pattern between these two series indicates the effectiveness of the realised volatility measure as a proxy for the underlying volatility influencing returns. Notably, the high volatility periods are associated with significant fluctuation events. The burst of the tech bubble caused market volatility from 2000 to 2002, while the GFC of 2007-2008, triggered by the collapse of the US housing market, resulted in significant volatility during that period. Additionally, the crisis that spread from the banking system to sovereign debt issues in Europe was evident in late 2011. Finally, the immediate global impact of COVID-19 led to an unprecedented surge in market volatility during 2019-2020.
The full dataset is divided into two disjoint time periods, comprising approximately 80% and 20% of the data while maintaining the temporal order. The first subsample, in-sample, from January 2000 to December 2017, including the period of the GFC, is used to estimate the models. The second subsample, out-of-sample, from January 2018 to June 2022, including the period of COVID-19, is used for comparing the predictive performance of models. There are some small differences in the size of in-sample and out-of-sample across markets due to the specific non-trading days in each market. The resulting in-sample sizes for all four markets are approximately 4500 days, with the largest size being 4537 days in AORD and the smallest being 4412 days in Hang Seng. The out-of-sample sizes for four markets are around 1100 days, with the largest size being 1133 days in AORD and the smallest size being 1092 days in Hang Seng.
This section illustrates the synthetic realised measures derived from PCA, ICA, average and autoencoder approaches for in-sample data of all markets. As the data split described in Section (ref), the in-sample size for S&P 500 is 4517 days, and estimated model parameters for the S&P 500 are displayed here for discussion.
Table (ref) shows the statistical summary of synthetic measures after log transformation. Since these synthetic measures undergo log transformation when incorporated into the RealGARCH model, the table provides their log-transformed statistics, facilitating comparison with Table (ref). Specifically, $\log(\mathbf{x}_{AE})$, $\log(\mathbf{x}_{PC})$ and $\log(\mathbf{x}_{IC})$ share the same minimum and maximum values, because they all have been rescaled according to Equation ((ref)) to fit within the minimum and maximum values of realised measures, shown in Table (ref). Moreover, the highly close statistics of $\log(\mathbf{x}_{PC})$ and $\log(\mathbf{x}_{IC})$ indicate a strong similarity between the two series, suggesting that both PCA and ICA capture similar patterns in decomposition.
Figure (ref) shows that the synthetic realised measure generated from the autoencoder $\mathbf{x}_{AE}$ has a similar pattern to that of the realised volatility measure, such as 5-min RV, effectively capturing the changing pattern in absolute returns. This means that the autoencoder can successfully generate a synthetic realised measure through nonlinear dimension reduction from multiple realised measures, to serve as a proxy for the underlying volatility.
Figure (ref) displays that the nonlinear combined realised measure from autoencoder shares a similar pattern to those of linear-combined measures, such as PCA and average. Given $\log(\mathbf{x}_{PC})$ and $\log(\mathbf{x}_{IC})$ are highly close as shown in Table (ref), only the realised measure from PCA, $\mathbf{x}_{PC}$, is displayed here to maintain visual clarity. This similarity suggests that the autoencoder captures core dynamic volatility, which is comparable to traditional linear combined methods. This novel representation of volatility offers an opportunity to improve the predictive performance and robustness of volatility modelling.
The estimated parameters of two GARCH-type models and five RealGARCH-type models are presented in Table (ref). Overall, the GARCH equation parameters are comparable among these models, and the consistency of measurement equation parameters can be observed.
In terms of the GARCH equation, $\alpha$ in standard GARCH and GARCH-X is the coefficient of lagged squared returns and realised measure, respectively. Since GARCH-X replaces the squared returns with the more informative realised measure, specifically 5-min RV here, the value of $\alpha$ is much higher than in standard GARCH. This higher $\alpha$ indicates that GARCH-X assigns more weight to the realised measure than squared returns, enhancing its contribution to volatility modelling. Therefore, GARCH-X has a lower training negative log-likelihood value than standard GARCH, indicating the improvement when using the more efficient realised measure during the training process. In RealGARCH models, $\gamma$ is the coefficient for the lagged realised measure term, comparable to the $\alpha$ in GARCH and GARCH-X models. The values of $\gamma$ in PC-RealGARCH, IC-RealGARCH, AVG-RealGARCH, and AE-RealGARCH are around 0.36 to 0.38, illustrating the comparable effectiveness of synthetic realised measures.
In terms of the measurement equation, there are 5 parameters in RealGARCH models specifically. The values for these 5 parameters are consistent with the estimates in Table $\mathrm{II}$: “Results for the log-linear specification" of hansenRealizedGARCHJoint2012. The estimates of $\varphi$ in RealGARCH models are all close to unity, suggesting that the synthetic realised measure is roughly proportional to the conditional variance of daily returns. The always negative estimates of $\xi$ reflect that a negative bias correction is needed in the regression of realised measure and the conditional variance of returns. One potential explanation for this is that the returns calculated in this thesis are based on close-to-close prices, including the overnight period (e.g. 24 hours) for prices to fluctuate. In contrast, the synthetic realised measure is derived from various realised measures, which are captured during the opening hours of the market (e.g., 6.5 hours). This shorter calculation period may lead to an underestimation of the underlying volatility of returns on average. Therefore, $\xi < 0$ should be expected, and a downside bias correction is required to align the conditional variance of returns with realised measures, resulting in the log-transformed realised measure being a biased estimator of the true log-transformed volatility.
Regarding the leverage effect term $\tau_1 z_t+\tau_2 (z_t^2-1)$, the estimates of $\tau_1$ are all below 0, which is consistent with the famous phenomenon in stock markets of negative correlation between today's return and tomorrow's volatility Nelson_1991_Conditional. To be specific, given $\tau_1<0$, $z_t$ is negative when the return $r_t<0$, then $\tau_1 z_t$ is a positive value. Combining this positive value with $\tau_2 (z_t^2-1)$ produces a larger term, resulting in a larger $\log(x_t)$ in the measurement equation than that of $z_t\ge 0$. Then this larger $\log(x_t)$ will lead to a larger volatility when predicting $\log(\sigma_{t+1}^2)$ through the GARCH equation as the coefficient of $\lambda$ is positive. Consequently, these estimates of parameters successfully reflect the existing leverage effect in markets.
For each market, the out-of-sample period spans approximately 1090 to 1130 days, starting on 2 January 2018 and ending on 28 June 2022. The parameters mentioned above, generated from the first estimation window (in-sample, 4 January 2000 to 31 December 2017), are used to produce the first one-step-ahead forecast for 2 January 2018 in out-of-sample. Then, the estimation window is slid forward by one day, maintaining a fixed window size. The synthetic realised measure series $\mathbf{x}_{AE}$, $\mathbf{x}_{PC}$, $\mathbf{x}_{IC}$ and $\mathbf{\Bar{x}}$ are re-generated based on the updated estimation window to re-estimate each model and produce the forecast for the next day. This rolling process is repeated daily until forecasts are generated for all days in the out-of-sample period across all models. Eventually, we obtained a total of approximately 1090 to 1130 one-step-ahead forecasts for each market.
Specifically, the hyperparameter values for $\lambda_1$, $\lambda_2$ and $\rho$ in Equation ((ref)) are set to 0.001, 0.001, and 0.05, respectively, which correspond to the default values of hyperparameters in “trainAutoencoder" function in Matlab. To pursue a more accurate encoding effect, the hyperparameters should ideally be selected through the validation process within each rolling window. However, due to the computational cost, we use the default values in this case, which work well in the empirical study with comprehensive tests. The encoded series $\mathbf{x}_{AE}$ has a desired pattern with realised measures in each rolling window, as shown in Figure (ref). It is noted that in some rolling window forecasting steps, the autoencoder may generate encoded series with unintended patterns by estimating negative weights $\mathbf{w}^1$ in the encoder. This is because the negative weights can cause the input values to approach negative infinity more readily than positive weights would, thereby resulting in the average activation value that is more easily closer to the desired level $\rho$ through the sigmoid activation function $sig(\cdot)$. In this case, we re-run those steps, which then yield series $\mathbf{x}_{AE}$ with desired patterns.
As introduced in Section (ref), the common log-likelihood function part $l(\mathbf{r};\boldsymbol{\theta})$ is used to compare between models. To assess the volatility forecasting performance, $l(\mathbf{r};\boldsymbol{\theta})$ can be used to calculate the predictive log-likelihood values. According to Gerlach_2016_Forecasting, the predictive log-likelihood function can be written as:
where $t$ starts from the first day in out-of-sample. For the first predictive log-likelihood value, $\hat{\sigma}_{t}^2$ is the forecasted variance based on estimated parameters $\hat{\boldsymbol{\theta}}$ ($\hat{\boldsymbol{\theta}}_{garch}$) and all available information up to and including time $T_{in}$. $r_{t}^2$ is the squared return observed at time $T_{in}+1$. Generally, the forecast volatility $\hat{\sigma}_{t}$ is generated at time $T_{in}$; when time advances to $T_{in}+1$, the observed return $r_{t}$ allows $l$ to be used as a measure of predictive accuracy. These calculations are performed for each day in the out-of-sample period and eventually summed to estimate the overall predictive log-likelihood for each model.
Table (ref) reports the predictive log-likelihood values for GARCH, GARCH-X and RealGARCH using 5-min RV, PC-RealGARCH, IC-RealGARCH, AVG-RealGARCH, and AE-RealGARCH in four markets.
First, in all markets, GARCH-X replacing squared returns with realised volatility measure, 5-min RV, consistently shows a lower negative predictive log-likelihood compared to standard GARCH using squared returns. This aligns with the findings of engleNewFrontiersArch2002. The in-sample parameters analysis in Section (ref) may partially explain this improvement. Additionally, RealGARCH extends GARCH-X by adding a measurement equation, while still using the same 5-min RV, further improving the predictive accuracy in three out of four markets.
Second, PC-RealGARCH and IC-RealGARCH, which use the synthetic realised measure $\mathbf{x}_{PC}$ and $\mathbf{x}_{PC}$ derived from PCA and ICA, respectively, exhibit the same predictive log-likelihood values across all markets. This similarity could be explained by the similarity of log-transformed series, $\log(\mathbf{x}_{PC})$ and $\log(\mathbf{x}_{IC})$, in Section (ref), as well as the estimated parameters incorporated into PC-RealGARCH and IC-RealGARCH, as described in Section (ref). The statistical summaries of $\log(\mathbf{x}_{PC})$ and $\log(\mathbf{x}_{IC})$, along with the parameter estimates and training negative log-likelihood values, are highly similar in PC-RealGARCH and IC-RealGARCH. Therefore, the high similarity between the two models suggests that, although PCA and ICA are based on different principles, targeting uncorrelation and independence, respectively, their decompositions of realised measures produce highly coincident results. This illustrates that PCA and ICA may not provide significantly different benefits when dealing with realised volatility measures.
Finally, the proposed model AE-RealGARCH ranks first in volatility forecasting for S&P 500 and FTSE, and second for AORD, demonstrating its solid performance among the competing models. The other fairly competitive model is the AVG-RealGARCH, which consistently ranks first or second for all markets, outperforming linear dimension reduction techniques like PCA and ICA. This observation motivates the exploration of nonlinear dimension reduction techniques for realised measures. In this regard, $\mathbf{x}_{AE}$ generated through a nonlinear transformation in autoencoder, gaining additional flexibility and capturing more information about volatility, provides a more effective and informative realised volatility measure than linear methods, resulting in a reliable AE-RealGARCH.
Figure (ref) shows the forecasted volatility series of AE-RealGARCH, AVG-RealGARCH, PC-RealGARCH and IC-RealGARCH, respectively. All models exhibit forecasted volatility series that consistently align with the absolute returns in the out-of-sample period, indicating that RealGARCH models using synthetic realised measures effectively capture and forecast dynamic volatility. Specifically, the forecasted volatility series in PC-RealGARCH and IC-RealGARCH are nearly identical, meaning that RealGARCH using whether $\mathbf{x}_{PC}$ or $\mathbf{x}_{IC}$ produces highly close forecasted volatility results, which is in line with the predictive log-likelihood analysis discussed above.
To further illustrate the impact of linear and nonlinear dimension reduction techniques on the synthetic realised measure within the RealGARCH model, Figure (ref) presents parameter estimates of $\omega$, $\beta$, $\gamma$ in the GARCH equation, as well as $\sigma_\varepsilon$ in measurement equation, across each forecasting step for S&P 500 out-of-sample period. Results are shown for the PC-RealGARCH, IC-RealGARCH, AE-RealGARCH and AVG-RealGARCH models, respectively.
First, regarding the parameters in the GARCH equation, the estimated value of $\gamma$ continuously increases over time, while the estimate of $\beta$ decreases correspondingly. This indicates that as RealGARCH forecasts volatility, it gradually relies on the realised measure during the out-of-sample period. The realised measure contributes progressively more to the model. Additionally, the estimates of parameters $\omega$, $\beta$ and $\gamma$ in the AE-RealGARCH exhibit significantly larger fluctuations than those of other models. One potential reason for this is that AE-RealGARCH incorporates the realised measure produced by autoencoder $\mathbf{x}_{AE}$, keeping a high sensitivity to each day's volatility. As a result, AE-RealGARCH has more flexibility to adjust the parameters in each forecasting step, effectively capturing the time-varying nature of volatility. The parameter estimates in PC-RealGARCH and IC-RealGARCH are almost identical in each rolling forecast window, which can further confirm the similar effect of PCA and ICA on realised measures discussed in Section (ref).
Second, regarding the measurement equation, the frequently fluctuated volatility $\sigma_\varepsilon$ in AE-RealGARCH indicates the synthetic realised measure $\mathbf{x}_{AE}$ has a more fluctuating volatility in each rolling window. Compared with $\mathbf{x}_{PC}$, $\mathbf{x}_{IC}$ and $\mathbf{\Bar{x}}$ which are obtained by linearly combining series with weights, $\mathbf{x}_{AE}$ is derived from a linear combination with weights first, followed by a nonlinear activation function $g(\cdot)$, specifically the sigmoid function $sig(\cdot)$ in this thesis. This nonlinear activation function provides $\mathbf{x}_{AE}$ additional flexibility, enabling it to effectively capture the underlying volatility across numerous realised measures, thus resulting in greater adaptability in each rolling window.
In this thesis, we propose an autoencoder enhanced RealGARCH that relies on a synthetic realised volatility measure obtained through the nonlinear dimension reduction of the autoencoder. This choice led to significant improvements in the out-of-sample predictive performance, compared to RealGARCH employing a single realised measure; linearly dimension-reduced synthetic realised measure from PCA, ICA and average, respectively, as well as traditional GARCH models. Furthermore, this study provides a new perspective for volatility modelling when researchers face the challenge of selecting a single volatility measure from multiple candidates. Instead of subjective selection, AE-RealGARCH synthesises the information from different realised measures into a robust component for model fitting and forecasting. The intended flexibility of the nonlinear transformation is evident through the more flexible parameter estimates in each rolling window, highlighting the ability of the autoencoder to capture complex patterns beyond the limitations of linear methods. This adaptability confirms the advantage of using nonlinear synthesis for volatility forecasting.
This study could be extended in the following ways. First, the autoencoder designed in this thesis has a single hidden layer for dimension reduction, and the input realised measures are directly transformed into a one-dimensional measure for RealGARCH. It is worthwhile to extend the single hidden layer autoencoder into a multilayer format, as in Hinton_2006_Reducing. Additionally, incorporating uncertainty measures based on financial news, such as economic policy uncertainty Baker_2016_Measuring, alongside realised volatility measures could enrich the model's input. This approach could allow the dimension-reduced output to be more than one dimension, thereby enhancing its integration with the RealEGARCH model. Second, the hyperparameter values for the autoencoder are set to the default values in Matlab, and the same hyperparameters are used for the subsequent rolling window steps to minimise computational cost. Further improvement involves validation setting for more advanced hyperparameter selection in each rolling window step, to tune based on different volatility conditions on each day. Lastly, this study employs Gaussian distribution for errors in returns to estimate models. Future research could consider more appropriate distributions, such as Student-$t$ errors, to further enhance forecasting accuracy.
\addcontentsline{toc}{section}{References}