Kim Christensen, Mathias Siggaard, Bezirgen Veliyev
arXiv 19 Jan 2026 · Econometrics
arXiv:2601.13014 · PDF · Extracted main text
We inspect how accurate machine learning (ML) is at forecasting realized variance of the Dow Jones Industrial Average index constituents. We compare several ML algorithms, including regularization, regression trees, and neural networks, to multiple Heterogeneous AutoRegressive (HAR) models. ML is implemented with minimal hyperparameter tuning. In spite of this, ML is competitive and beats the HAR lineage, even when the only predictors are the daily, weekly, and monthly lags of realized variance. The forecast gains are more pronounced at longer horizons. We attribute this to higher persistence in the ML models, which helps to approximate the long-memory of realized variance. ML also excels at locating incremental information about future volatility from additional predictors. Lastly, we propose a ML measure of variable importance based on accumulated local effects. This shows that while there is agreement about the most important predictors, there is disagreement on their ranking, helping to reconcile our results.
appendix boundary found by appendix_command · 84% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | D. P. Kingma and J. Ba (2014) Adam: A method for stochastic optimization | 0.737 | 4 | 2 | 75% |
| 2 | T. Bollerslev and A. J. Patton and R. Quaedvlieg (2016) Exploiting the errors: A simple approach for improved volatility forecasting | 0.737 | 3 | 2 | 100% |
| 3 | F. Corsi (2009) A simple approximate long-memory model of realized volatility | 0.737 | 3 | 2 | 100% |
| 4 | T. G. Andersen and T. Bollerslev and F. X. Diebold (2007) Roughing it up: Including jump components in the measurement, modeling and forecasting of return volatility | 0.644 | 2 | 2 | 100% |
| 5 | F. Corsi and R. Renò (2012) Discrete-time volatility forecasting with persistent leverage effect and the link with continuous-time volatility modeling | 0.644 | 2 | 2 | 100% |
| 6 | F. X. Diebold and R. S. Mariano (1995) Comparing predictive accuracy | 0.644 | 2 | 2 | 100% |
| 7 | J. H. Friedman (2001) Greedy function approximation: A gradient boosting machine | 0.644 | 2 | 2 | 100% |
| 8 | S. Gu and B. Kelly and D. Xiu (2020) Empirical asset pricing via machine learning | 0.644 | 2 | 2 | 100% |
| 9 | P. R. Hansen and A. Lunde and J. M. Nason (2011) The model confidence set | 0.644 | 2 | 2 | 100% |
| 10 | A. J. Patton and K. Sheppard (2015) Good volatility, bad volatility: Signed jumps and the persistence of volatility | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 70 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.