arXiv 6 Jul 2026 · Finance — Statistical Finance
arXiv:2607.05291 · PDF · DOI · OpenAlex · Extracted main text
We ask whether pretrained time series foundation models (TSFMs) improve on established econometric benchmarks for forecasting realized volatility. Using the VOLARE dataset, we conduct the first systematic comparison of nine zero-shot TSFMs against eight econometric specifications, including the Heterogeneous Autoregressive (HAR) family, across 50 assets in equities, foreign exchange, and futures, and three forecast horizons, with formal pairwise and multi-model forecast-comparison tests. Foundation models do not deliver a uniform gain. Pooled losses favor them, but the advantage is concentrated in a few outlier assets; averaging each asset's loss ratio to a well-specified Log-HAR benchmark, so that no single asset dominates, only one small model, Tiny Time Mixers (TTM), beats the benchmark at every horizon, and by a narrow margin. The other foundation models do not improve on Log-HAR, and the econometric benchmarks remain competitive throughout. A Mincer--Zarnowitz recalibration, which removes level and scale bias from every forecast, shows that much of the short-horizon advantage reflects better-scaled forecasts rather than better prediction of volatility dynamics, and only at the monthly horizon does a genuine informational gain remain. Because this edge is thin and even TTM is not best on every asset, a simple equal-weight average of TTM and Log-HAR matches the best single model and enters the Model Confidence Set for 98 to 100% of assets, more often than either component alone, so a forecaster need not identify the best model for each asset in advance. Our most durable finding is that performance varies so much across foundation-model architectures that choosing the right architecture matters more than the broader choice between foundation and econometric models.
appendix boundary found by appendix_command · 88% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bollerslev, Tim and Patton, Andrew J. and Quaedvlieg, Rogier (2016) Exploiting the Errors: A Simple Approach for Improved Volatility Forecasting | 1.000 | 5 | 3 | 100% |
| 2 | Corsi, Fulvio (2009) A Simple Approximate Long-Memory Model of Realized Volatility | 0.928 | 4 | 3 | 100% |
| 3 | Giacomini, Raffaella and Rossi, Barbara (2010) Forecast Comparisons in Unstable Environments | 0.928 | 4 | 3 | 100% |
| 4 | Ansari, Abdul Fatir and Stella, Lorenzo and Turkmen, Caner and Zhang… (2024) Chronos: Learning the Language of Time Series | 0.874 | 6 | 4 | 67% |
| 5 | Carriero, Andrea and Pettenuzzo, Davide and Shekhar, Shubhranshu (2024) Macroeconomic Forecasting with Large Language Models | 0.843 | 4 | 4 | 75% |
| 6 | Liu, Chenghao and Aksu, Taha and Liu, Juncheng and Liu, Xu and Yan,… (2025) Moirai 2.0: When Less Is More for Time Series Forecasting | 0.843 | 4 | 4 | 75% |
| 7 | Rasul, Kashif and Ashok, Arjun and Williams, Andrew Robert and Ghoni… (2024) Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting | 0.843 | 4 | 4 | 75% |
| 8 | Ekambaram, Vijay and Jati, Arindam and Dayama, Pankaj and Mukherjee,… (2024) Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series | 0.843 | 4 | 3 | 75% |
| 9 | Andersen, Torben G. and Bollerslev, Tim and Diebold, Francis X (2007) Roughing It Up: Including Jump Components in the Measurement, Modeling, and Forecasting of Return Volatility | 0.843 | 3 | 3 | 100% |
| 10 | Christensen, Kim and Siggaard, Mathias and Veliyev, Bezirgen (2023) A Machine Learning Approach to Volatility Forecasting | 0.843 | 3 | 3 | 100% |
Showing the top 10 of 80 scored citations.