EconBase
← All papers

A machine learning approach to volatility forecasting

Kim Christensen, Mathias Siggaard, Bezirgen Veliyev

arXiv 19 Jan 2026 · Econometrics

arXiv:2601.13014 · PDF · Extracted main text

Abstract

We inspect how accurate machine learning (ML) is at forecasting realized variance of the Dow Jones Industrial Average index constituents. We compare several ML algorithms, including regularization, regression trees, and neural networks, to multiple Heterogeneous AutoRegressive (HAR) models. ML is implemented with minimal hyperparameter tuning. In spite of this, ML is competitive and beats the HAR lineage, even when the only predictors are the daily, weekly, and monthly lags of realized variance. The forecast gains are more pronounced at longer horizons. We attribute this to higher persistence in the ML models, which helps to approximate the long-memory of realized variance. ML also excels at locating incremental information about future volatility from additional predictors. Lastly, we propose a ML measure of variable importance based on accumulated local effects. This shows that while there is agreement about the most important predictors, there is disagreement on their ranking, helping to reconcile our results.

Citation extraction

70
references
88
in-text mentions
70
distinct cited
3
self-citations
17,637
main-text words

appendix boundary found by appendix_command · 84% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1D. P. Kingma and J. Ba (2014) Adam: A method for stochastic optimization0.7374275%
2T. Bollerslev and A. J. Patton and R. Quaedvlieg (2016) Exploiting the errors: A simple approach for improved volatility forecasting0.73732100%
3F. Corsi (2009) A simple approximate long-memory model of realized volatility0.73732100%
4T. G. Andersen and T. Bollerslev and F. X. Diebold (2007) Roughing it up: Including jump components in the measurement, modeling and forecasting of return volatility0.64422100%
5F. Corsi and R. Renò (2012) Discrete-time volatility forecasting with persistent leverage effect and the link with continuous-time volatility modeling0.64422100%
6F. X. Diebold and R. S. Mariano (1995) Comparing predictive accuracy0.64422100%
7J. H. Friedman (2001) Greedy function approximation: A gradient boosting machine0.64422100%
8S. Gu and B. Kelly and D. Xiu (2020) Empirical asset pricing via machine learning0.64422100%
9P. R. Hansen and A. Lunde and J. M. Nason (2011) The model confidence set0.64422100%
10A. J. Patton and K. Sheppard (2015) Good volatility, bad volatility: Signed jumps and the persistence of volatility0.64422100%

Showing the top 10 of 70 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1HARd to Beat: The Overlooked Impact of Rolling Windows in the Era of Machine Learning0.981185
2Forecasting Realized Volatility with Time Series Foundation Models: A Comparison with Econometric Benchmarks0.84333
3HARNet: A Convolutional Neural Network for Realized Volatility Forecasting0.40511
4Generalized Autoregressive Score Trees and Forests0.40511
5Sparse Tree-Based Aggregation for Time Series Regressions0.40511
6Forecasting of volatility and risk premia in electricity markets0.40511