Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
106,077 characters · 22 sections · 128 citation commands
Forecasting Realized Volatility with Time Series Foundation Models: A Comparison with Econometric Benchmarks
\thispagestyle{empty}
\noindentKeywords: Long memory time series; Econometric models; Foundation models; Model selection; Evaluating forecasts; Realized volatility.
Forecasting realized volatility is central to risk management, derivative pricing, and portfolio allocation. Since the work of andersen1998 and barndorffnielsen2002, realized volatility constructed from high-frequency returns has become the standard model-free measure of ex-post price variation. The Heterogeneous Autoregressive (HAR) model of corsi2009 is the dominant forecasting benchmark: its three-component structure, aggregating past realized volatility at daily, weekly, and monthly frequencies, provides a parsimonious approximation to the long-memory dynamics that characterize volatility. Extensions incorporating jump variation andersen2007roughing, semivariance asymmetries patton2015, and measurement error corrections bollerslev2016harq refine the model. Machine learning methods, including neural networks bucci2020, zhang2024intraday, random forests luong2018, graph-based approaches zhang2025graph, brini2025spotv2net, and convolutional architectures morenopino2024deepvol, have also been applied, and a broader assessment by christensen2023ml finds that machine learning beats the HAR family, with the gains most pronounced at longer horizons, which they attribute to the higher persistence of the machine-learning forecasts approximating the long memory of realized variance.
The models discussed so far, econometric and machine-learning alike, are estimated on the target volatility series itself. A different class of models dispenses with this step: time series foundation models (TSFMs). These are large pretrained transformer neural networks vaswani2017, which learn dependencies between positions in a sequence through an attention mechanism without imposing a fixed lag structure. They are trained on large and diverse corpora of time series from multiple domains and can produce forecasts for previously unseen series without any retraining, a capability known as zero-shot forecasting, which gruver2023 also demonstrated for general-purpose large language models. Leading examples of TSFMs include Chronos ansari2024chronos, Moirai woo2024moirai, and Lag-Llama rasul2024lagllama, several of which have released second-generation versions with improved architectures ansari2025chronos2, liu2025moirai2. Surveys document the growth of this area liang2024tsfmsurvey, ye2024tsfmsurvey, miller2024tsfmsurvey. Outside finance, carriero2024macro applied zero-shot TSFMs to macroeconomic forecasting and found that these models are not yet a clear replacement for macroeconometric baselines, raising the question of whether similar conclusions hold for other financial time series, such as the realized volatility we study.
Applications of TSFMs to financial time series remain limited. goel2025rv tested TimesFM 2.0 on realized volatility for 21 global equity indices and found that fine-tuning was necessary for the model to compete with HAR, with zero-shot performance not consistently better than the benchmark. rahimikia2025revisiting evaluated several TSFMs on daily excess returns and reported uniformly negative zero-shot results, with fine-tuning yielding only limited gains that did not close the gap with benchmark ensembles. Realized volatility differs from returns in ways that may favor foundation models, being strictly positive, mean-reverting, and long-memory andersen2001exchange, andersen2003. Whether these properties make a general-purpose pretrained model competitive with a benchmark designed for volatility is an open question that a single-model study cannot settle.
Since no study has yet evaluated multiple TSFMs on realized volatility with formal statistical testing, in this paper we conduct the first systematic comparison of zero-shot TSFMs against established econometric benchmarks. We evaluate nine TSFMs, spanning eight distinct architectures, against eight econometric specifications across 50 assets in three asset classes (equities, foreign exchange (FX), futures) and three forecast horizons ($h = 1, 5, 22$ days), using the VOLARE dataset volare2026. The forecast target is the point-in-time realized volatility. The zero-shot setting, with no domain-specific training, isolates the value of general time series pretraining. We apply Diebold--Mariano (DM) tests diebold1995 and Model Confidence Set (MCS) analysis hansen2011mcs to control for multiple comparisons, and supplement these with Mincer--Zarnowitz (MZ) regressions mincer1969 for forecast efficiency, Giacomini--Rossi (GR) fluctuation tests giacomini2010 for time-varying relative performance, sub-sample analysis across pre- and post-COVID regimes, and context window sensitivity checks for the foundation models.
We find that pretrained foundation models do not deliver a uniform gain over the econometric benchmarks. Pooled-mean quasi-likelihood (QLIKE) losses appear to favor several TSFMs, but they are sensitive to a few outlier assets and overstate the typical advantage. Under average QLIKE loss ratios relative to Log-HAR, which weight each asset equally so that no single asset dominates, only Tiny Time Mixers (TTM), the smallest model in the evaluation ($<$1M parameters), beats Log-HAR at every horizon on the raw zero-shot forecasts, and only by a small margin of roughly 1.3 to 1.8%. The other eight TSFMs do not beat a well-specified Log-HAR on average. Log-HAR, HAR, the Autoregressive Fractionally Integrated Moving Average (ARFIMA) model, the Autoregressive Moving Average (ARMA) model, and the multiplicative error model (MEM) all cluster within a few percent of Log-HAR and remain competitive throughout, a ranking that the Model Confidence Set corroborates, while the jump- and quarticity-augmented HAR variants fall well behind. A uniform MZ recalibration then decomposes this edge into a calibration component and an information component: at the daily horizon Log-HAR is in fact the more efficient forecast and TTM's edge is largely a shared calibration effect (its forecasts already sit at the right level and scale) that several other foundation models also show, while at the monthly horizon TTM retains a genuine informational advantage. Because this edge is thin and even TTM is not best on every asset, a simple equal-weight average of TTM and Log-HAR matches the best single model and enters the MCS for 98 to 100% of assets across horizons, so a forecaster need not identify the best model for each asset in advance. Our most durable finding is that performance varies so widely across TSFM architectures that which TSFM one chooses matters more than whether to use a TSFM or an econometric model at all.
The remainder of this paper is organized as follows. Sec. (ref) reviews the related literature. Sec. (ref) describes the VOLARE dataset. Sec. (ref) presents the econometric and foundation model specifications, along with the forecast evaluation framework. Sec. (ref) reports the empirical results, Sec. (ref) assesses statistical significance through formal forecast-comparison tests, and Sec. (ref) contains robustness checks, with Sec. (ref) concluding.
This section reviews three strands of work. Subsec. (ref) covers realized volatility modeling and the HAR benchmark; Subsec. (ref) introduces TSFMs and their architectures; and Subsec. (ref) surveys their still-limited application to volatility forecasting.
The theory of realized volatility originates with andersen1998 and andersen2001exchange, who showed that the sum of squared intraday returns provides a consistent, nonparametric estimator of the integrated variance of asset prices. barndorffnielsen2002 developed the asymptotic distribution theory for realized variance (RV) in a stochastic volatility framework, deriving a central limit theorem and rate of convergence for the RV error around integrated variance. Subsequent work identified the key stylized facts of realized volatility: approximate log-normality, long-memory dependence with a fractional integration parameter $d \approx 0.4$, and slow mean reversion andersen2003.
At ultra-high sampling frequencies, microstructure noise from bid-ask bounce and discrete price changes biases the realized variance estimator upward. barndorffnielsen2008 introduced the realized kernel as a noise-consistent alternative that remains valid in the presence of market microstructure effects.
The HAR model of corsi2009 became the standard forecasting benchmark for realized volatility. By including daily, weekly, and monthly realized volatility components as regressors, the model approximates the long-memory decay of volatility autocorrelations. The specification is parsimonious (three regressors plus a constant) yet achieves accuracy comparable to long-memory models and to the more heavily parameterized machine-learning forecasters studied by christensen2023ml.
A large body of work has extended the HAR framework. The HAR-J model andersen2007roughing separates realized variance into continuous and jump components using bipower variation barndorffnielsen2004bipower. The HAR-RS model barndorffnielsen2010semivariance, patton2015 decomposes realized variance into positive and negative semivariance to capture asymmetric responses to upside and downside moves. The HARQ model bollerslev2016harq interacts the daily RV regressor with realized quarticity to account for time-varying measurement error. clements2021practical provide a practical guide to implementing and comparing these extensions.
Long-memory models offer a complementary approach. ARFIMA models granger1980, hosking1981, applied to realized volatility by andersen2003, explicitly parameterize the fractional integration order and remain competitive at longer forecast horizons where the slow mean reversion of realized volatility becomes the dominant dynamic. Machine learning methods have produced mixed results in realized volatility forecasting. bucci2020 reports that recurrent networks outperform ARFIMA-type benchmarks for S&P 500 realized volatility, and luong2018 find gains from random forests. In contrast, branco2024 conclude that simple linear models are difficult to beat after correcting for multiple testing. Surveys by gunnarsson2024 and leushuis2026 cover this literature.
Time series foundation models are large pretrained models that produce forecasts for arbitrary time series without task-specific training, analogous to large language models for text liang2024tsfmsurvey. They differ from task-specific deep learning models (multilayer perceptrons, recurrent networks, convolutional architectures, graph neural networks, and transformers), which must be trained from scratch on the target series before producing any forecast. TSFMs are instead pretrained on large corpora drawn from diverse domains (weather, energy, retail, transport) and forecast immediately via zero-shot inference. This eliminates the need for domain-specific training data, hyperparameter tuning, and GPU-intensive estimation, a practical advantage when deploying to new domains where historical data may be limited dooley2023forecastpfn.
Three architectural families dominate the first generation of TSFMs. Chronos ansari2024chronos converts continuous values into discrete tokens via uniform binning and forecasts recursively with a T5 encoder-decoder architecture raffel2020. Moirai woo2024moirai handles an arbitrary number of input series through an “any-variate” attention mechanism and uses mixture distribution outputs for uncertainty quantification. Lag-Llama rasul2024lagllama adapts the LLaMA decoder-only architecture for probabilistic time series forecasting using lag-based tokenization. Finance-specific foundation models have also been proposed, including Kronos shi2025kronos, which is pretrained on candlestick (open-high-low-close-volume, OHLCV) data from over 45 global exchanges using a learned tokenizer that maps price patterns into discrete tokens.
These first-generation architectures were quickly followed by improved successors and a wider range of models. ansari2025chronos2 introduced Chronos-2, which adds multivariate support and covariate handling; the Chronos team also released Chronos-Bolt, a faster variant of the original Chronos that produces direct quantile forecasts in a single forward pass and runs up to 250 times faster. Moirai 2.0 liu2025moirai2 demonstrated that smaller, better-trained models can match or exceed their larger predecessors using multi-token prediction and improved tokenization. Other models include TimesFM das2024timesfm, a patch-based decoder model from Google; Toto cohen2024toto, a model from Datadog pretrained on infrastructure monitoring metrics (e.g., server CPU usage, request latency); Moirai-MoE liu2024moiraimoe, a sparse Mixture of Experts (MoE) extension of Moirai\footnote{A Mixture of Experts replaces a single dense network with several specialized sub-networks (“experts”) and a gating mechanism that routes each input to a small subset of them, so the model's total capacity can grow while the computation per forecast stays low.}; Sundial liu2025sundial, which uses flow matching for generative forecasting; MOMENT goswami2024moment, a masked-encoder architecture designed for multiple time series tasks (forecasting, classification, anomaly detection)\footnote{We do not evaluate MOMENT because it is pretrained with a masked reconstruction objective rather than an autoregressive forecasting objective. Its forecasting head, a linear projection layer that maps patch embeddings to the forecast horizon, is not pretrained and must be trained on the target series before the model can produce any predictions goswami2024moment. This makes it incompatible with our zero-shot evaluation protocol.}; Timer liu2024timer; and TTM ekambaram2024ttm, IBM's lightweight TSMixer-based model and the smallest in our evaluation.
Several standardized benchmarks now evaluate TSFMs across domains. GIFT-Eval aksu2024gifteval provides a unified evaluation protocol across multiple datasets and forecast horizons. FEV-Bench shchur2025fevbench emphasizes realistic tasks with covariates and principled aggregation across tasks. TSFM-Bench li2024tsfmbench compares models across zero-shot and fine-tuned settings. A consistent finding is that no single model dominates across all domains and horizons, which motivates domain-specific evaluations of the kind we undertake here. tan2024llmuseful question whether language-model-based forecasters add value over simpler baselines at all, a skepticism that motivates the head-to-head design against econometric benchmarks that we adopt here.
The application of TSFMs to volatility forecasting is limited. The closest prior work is goel2025rv, who tested TimesFM 2.0 on realized volatility for 21 global equity indices. They found that zero-shot TimesFM did not consistently beat HAR and that fine-tuning was necessary to achieve competitive accuracy. Our paper differs in several respects. We evaluate nine foundation models across eight distinct architectures (Chronos-Bolt-Small, Chronos-Bolt-Base, Moirai 2.0, Moirai-MoE, Lag-Llama, TimesFM 2.5, Toto, Sundial, and TTM) and 50 assets in individual equities, foreign exchange, and futures, and we assess forecast significance and robustness through formal statistical testing.
rahimikia2025revisiting evaluated TSFMs from the Chronos and TimesFM families on daily excess returns and found that zero-shot TSFMs consistently underperformed strong machine-learning ensembles such as CatBoost and LightGBM. Because daily returns are near-white-noise while realized volatility is strictly positive, mean-reverting, and long-memory, closer to the macroeconomic and physical series on which TSFMs were pretrained, negative results on returns need not carry over to realized measures.
Other financial applications include stock price forecasting laniewski2025chronos, valeyre2024chronosstocks, foreign exchange volatility modeling nguyen2025volabert, Value-at-Risk forecasting goel2024var, and finance-specific pretrained models such as FinCast zhu2025fincast and Kronos shi2025kronos. Most directly related to our result, marconi2025 reports that small TTM models are competitive on financial forecasting tasks, including foreign-exchange volatility, with the strongest gains obtained through fine-tuning rather than zero-shot use. That study uses neither formal forecast-comparison tests nor a multi-model panel. Our results show that a small model can edge a well-specified Log-HAR on realized volatility across 50 assets in the zero-shot setting, by a narrow margin and without displacing the econometric benchmarks from the Model Confidence Set.
The gap in this literature, the absence of a multi-model, multi-asset evaluation with formal statistical testing for realized volatility, motivates the empirical design we describe next.
We use the VOLARE (VOLatility Archive for Realized Estimates) dataset of volare2026, which provides daily realized variance and related realized measures for a broad cross-section of financial assets. VOLARE is constructed from ultra-high-frequency tick data sourced from Kibot, covering 40 U.S.\ equities, 5 major currency pairs, and 5 commodity and index futures contracts. The equity sample begins on January 2, 2015 and runs through January 30, 2026, yielding 2,786 trading days per stock. The FX sample begins on September 25, 2009 and the futures sample on September 28, 2009 (up to 4,242 and 4,224 trading days, respectively).
For each asset-day, VOLARE provides realized measures computed at multiple sampling frequencies. We use all realized measures at the 5-minute sampling frequency, the standard bias-variance compromise against microstructure noise in the realized-volatility literature. liu2015fiveminute compare realized measures across asset classes and find 5-minute sampling hard to beat. Finer (1-minute) and noise-robust (realized-kernel) alternatives are available in VOLARE but are not used in this work. The measures are realized variance ($RV$), bipower variation ($BPV$), positive and negative realized semivariance ($RS^{+}$, $RS^{-}$), and realized quarticity ($RQ$). All values are expressed in decimal squared returns; for instance, a typical daily $RV$ for a U.S.\ equity is approximately $2.5 \times 10^{-4}$, corresponding to an annualized volatility of roughly 25%. VOLARE is well suited for our multi-asset comparison: it constructs all realized measures with a uniform methodology across assets. Cross-asset differences in model rankings then reflect genuine forecasting performance rather than measurement inconsistencies. The sample spans multiple volatility regimes, including the low-volatility period of 2017 to 2019, the COVID-19 shock of 2020, and the subsequent recovery. The equity sample comprises the 40 VOLARE stocks with complete coverage over the full 2015 to 2026 window, a balanced-panel restriction that requires survival over the sample and trades breadth for a common evaluation period. The stocks are AAPL, ADBE, AMD, AMGN, AMZN, AXP, BA, CAT, CRM, CSCO, CVX, DIS, GE, GOOGL, GS, HD, HON, IBM, JNJ, JPM, KO, MCD, META, MMM, MRK, MSFT, NFLX, NKE, NVDA, ORCL, PG, PM, SHW, TRV, TSLA, UNH, V, VZ, WMT, and XOM, covering all 11 Global Industry Classification Standard sectors. The FX sample consists of five major currency pairs (AUDUSD, EURUSD, GBPUSD, USDCAD, USDJPY) and the futures sample covers five contracts (Corn, Crude Oil, E-mini S&P 500, Gold, Natural Gas). Both FX and futures samples provide longer histories than the equity sample, because VOLARE's intraday coverage for these asset classes begins earlier.
Tab. (ref) reports descriptive statistics for the 5-minute realized variance across the three asset classes. As expected for a strictly positive quantity, realized variance is right-skewed across all asset classes, with heavy tails (kurtosis ranging from 35 to over 3,000) confirming that extreme volatility episodes are a persistent feature of the data. The first-order autocorrelation of daily $RV$, denoted $\rho_1$, averages 0.598 for equities, consistent with the well-documented persistence of volatility; the 22-day autocorrelation $\rho_{22}$ averages 0.098, so some dependence remains at the monthly horizon. The FX pairs display $RV$ smaller than equities by a factor of about eight ($\bar{RV} \approx 3 \times 10^{-5}$ for FX vs.\ $\approx 2.5 \times 10^{-4}$ for equities), with comparable persistence (cross-sectional average $\bar{\rho}_1 = 0.52$, where the bar denotes averaging across assets). Futures exhibit the widest cross-asset heterogeneity: the E-mini S&P 500 (ES) has $\rho_1 = 0.779$, while Gold (GC) shows near-zero autocorrelation ($\rho_1 \approx 0$), presenting a natural stress test for forecasting models. These moments describe realized variance as distributed in VOLARE; all forecasting and evaluation in the paper are conducted on the realized volatility scale $\sigma_t = \sqrt{RV_t}$.
This section describes the 17 forecasting models in our comparison: eight econometric benchmarks (Sec. (ref)), nine TSFMs (Sec. (ref)), the pretraining-data and contamination assessment (Sec. (ref)), and the evaluation framework (Sec. (ref)).
We consider eight econometric specifications that span the main approaches to realized volatility forecasting: the HAR family and its extensions, which aggregate lagged realized volatility at daily, weekly, and monthly horizons; ARFIMA, which models long memory through fractional integration; an ARMA model on log realized volatility; and the MEM, which enforces positivity through a multiplicative structure. We forecast realized volatility $\sigma_t \equiv \sqrt{RV_t}$ throughout, where $RV_t$ is the realized variance stored in VOLARE. We estimate the pure-RV models (HAR, Log-HAR, ARFIMA, ARMA, and MEM) directly on the volatility series $\sigma_t$. The augmented HAR variants (HAR-J, HAR-RS, and HARQ) instead carry variance-scale regressors (jumps, semivariances, and quarticity); we estimate these on the variance scale $RV$ and map their forecasts to volatility. We restrict the comparison to models that forecast the realized volatility series directly. Realized GARCH hansen2012realgarch and Realized EGARCH hansen2016realegarch instead jointly model daily returns and a realized measure to forecast the return conditional variance; mapping them onto our univariate realized-volatility target would require the return series and an auxiliary measurement equation, placing them on a different information set, so we leave them out of the comparison.
\paragraph{HAR Model corsi2009.} The HAR model captures the multi-horizon persistence of realized volatility by aggregating past realized volatility at daily, weekly, and monthly frequencies:
where $\sigma_{t-k:t} \equiv k^{-1}\sum_{i=0}^{k-1} \sigma_{t-i}$ denotes the average realized volatility over the previous $k$ days. We use the original specification, where the weekly and monthly components include overlapping lags (i.e., the weekly component averages days $t$ through $t-4$, not the non-overlapping “rotated” version that separates lags 2 to 5 from 6 to 22). A realized volatility forecast is non-negative by construction, but an unconstrained least-squares fit does not impose this and can return negative or near-zero values. We therefore estimate Eq. (ref) by non-negativity-constrained least squares, requiring the intercept and all lag coefficients to be non-negative, a sufficient positivity condition analogous to the GARCH non-negativity constraints of nelson1992. The constraint guarantees non-negative forecasts for any future input configuration and removes the need for any post-hoc adjustment of the predicted values.
\paragraph{HAR-J Model andersen2007roughing.} The HAR-J model augments the baseline HAR with a jump component to separate continuous and discontinuous variation:
where $J_t = \max(RV_t - BPV_t,\, 0)$ measures the jump component as the positive part of the difference between realized variance and bipower variation, a jump-robust estimator of integrated variance constructed from products of adjacent absolute returns barndorffnielsen2004bipower. A negative coefficient on $J_t$ indicates that large jumps reduce, rather than increase, future volatility.
\paragraph{HAR-RS Model patton2015.} The HAR-RS model decomposes realized variance into positive and negative semivariance components to capture asymmetric volatility responses:
where $RS_t^+$ and $RS_t^-$ denote good and bad realized semivariance, respectively, and $RS_t^+ + RS_t^- = RV_t$ by construction barndorffnielsen2010semivariance. Formally, $RS_t^{-}=\sum_{i} r_{t,i}^2\,\mathbb{1}\{r_{t,i}\le 0\}$ and $RS_t^{+}=\sum_{i} r_{t,i}^2\,\mathbb{1}\{r_{t,i}>0\}$, where $r_{t,i}$ is the $i$-th intraday return on day $t$. The six-regressor structure allows upside and downside risk to follow separate dynamics at each aggregation frequency.
\paragraph{HARQ Model bollerslev2016harq.} The HARQ model accounts for time-varying measurement error in realized volatility by interacting the daily RV regressor with the square root of realized quarticity:
where $RQ_t$ is the realized quarticity, $RQ_t=\tfrac{n}{3}\sum_{i} r_{t,i}^4$ with $n$ the number of intraday returns on day $t$. The interaction term attenuates the daily RV signal when measurement noise is high, as indicated by large values of $RQ_t$.
\paragraph{Log-HAR Model.} The Log-HAR model applies the HAR specification of corsi2009 to log-transformed realized volatility:
The logarithmic transformation maps volatility to the real line, so forecasts in levels are positive by construction after exponentiation, and the residual distribution is closer to Gaussian, the condition under which least squares is efficient taylor2017. Point forecasts in levels are recovered via bias-corrected retransformation: $\widehat{\sigma}_{t+h} = \exp(\widehat{\log \sigma}_{t+h} + s^2/2)$, where $s^2$ is the estimated residual variance of the log-volatility regression, following the standard log-normal adjustment. Log-HAR is our headline econometric benchmark: it is the most widely used log-space specification, it requires no positivity constraint, and we compute the relative loss ratios of Sec. (ref) against it.
\paragraph{ARFIMA Model granger1980, hosking1981.} The ARFIMA model captures the long-memory property of realized volatility through fractional differencing:
where $d \in [0, 0.5)$ is the fractional integration parameter governing the rate at which autocorrelations decay, $\Phi(L)$ and $\Theta(L)$ are autoregressive and moving average lag polynomials of orders $p$ and $q$, and $L$ is the lag operator. The boundary $d = 0$ nests the short-memory ARMA benchmark introduced below, so the value of fractional integration is testable; values near 0.4, typical for realized volatility, imply slow hyperbolic decay rather than the exponential decay of a standard ARMA model. We estimate $d$ by local-Whittle (Gaussian semiparametric) maximum likelihood robinson1995, which estimates the long-memory parameter directly from the periodogram near the zero frequency rather than through the log-periodogram regression of geweke1983. Given the estimated $\hat{d}$, we form the fractionally differenced series $w_t = (1-L)^{\hat{d}}(\log \sigma_t - \hat{\mu})$, select the short-memory orders $(p,q)$ over $\{0,1,2\}^2$ by the Bayesian information criterion (BIC), and fit an ARMA$(p,q)$ to $w_t$. We then re-integrate the forecasts of $w_t$ by applying the inverse fractional difference operator $(1-L)^{-\hat{d}}$.
\paragraph{ARMA Model.} We add an ARMA model fit directly to log realized volatility, $\Phi(L)(\log \sigma_t - \mu) = \Theta(L)\varepsilon_t$, with the orders $(p,q)$ selected over a grid $\{0,1,2\}^2$ by BIC at each estimation origin box1970. This is the short-memory counterpart of ARFIMA: it shares the log specification and the Gaussian-error fit but omits the fractional-integration term, so the comparison isolates the forecasting value of explicitly modeling long memory. Forecasts in levels use the same log-normal retransformation as Log-HAR.
\paragraph{MEM Model engle2002.} The MEM specifies realized volatility as the product of a conditional mean and a non-negative multiplicative innovation, $\sigma_t = \mu_t\,\varepsilon_t$ with $\mathrm{E}[\varepsilon_t \mid \mathcal{F}_{t-1}] = 1$, and a GARCH-type recursion for the conditional mean,
with $\omega, \alpha, \beta \ge 0$, so $\mu_t > 0$ by construction. We estimate $(\omega,\alpha,\beta)$ by exponential quasi-maximum likelihood, which is consistent for the conditional-mean parameters under correct specification of $\mu_t$ irrespective of the innovation density. The MEM is strictly positive by design, providing a second positivity-guaranteed benchmark alongside Log-HAR.
The five pure models (HAR, Log-HAR, ARFIMA, ARMA, and MEM) are used in iterated multistep mode, the configuration for which these specifications are designed and the one typically more accurate than direct projection when the model is not badly misspecified marcellino2006; for the volatility-specific comparison of direct versus iterated multiperiod forecasts, see ghysels2019. HAR and Log-HAR iterate by recursive plug-in, feeding each one-step forecast back as the most recent observation; ARFIMA, ARMA, and MEM iterate natively through their recursive structure. The three augmented HAR variants (HAR-J, HAR-RS, and HARQ) are estimated directly at each horizon $h$, because their auxiliary regressors (jumps, realized semivariances, and the quarticity interaction) cannot be projected forward without an auxiliary model for each one; estimating these specifications directly at the target horizon is the treatment adopted in the literature for horizon-specific HAR extensions bollerslev2016harq.
We re-estimate every econometric model at every origin of the rolling window defined in Sec. (ref), standard practice for the HAR family. We do not tabulate estimated coefficients or their standard errors; the reported quantities are out-of-sample forecast losses.
We evaluate nine TSFMs in a zero-shot setting, applied directly to realized volatility series without any domain-specific training or fine-tuning. The nine models span eight distinct architectures: Chronos-Bolt (small and base checkpoints), Moirai 2.0, Moirai-MoE, Lag-Llama, TimesFM 2.5, Toto, Sundial, and TTM.\footnote{carriero2024macro additionally evaluate TimeGPT garza2024timegpt for macroeconomic forecasting. We exclude TimeGPT because it is a proprietary, closed-source API that does not permit inspection of model weights or training data, precluding reproducibility.} Where a model is offered in multiple sizes we evaluate the small checkpoint, which the results tables denote with an “-S” suffix (for example, Moirai-2.0-S and Moirai-MoE-S); the surrounding text refers to each model by its architecture name. The zero-shot approach is the deployment mode most relevant in practice: it requires no labeled financial data and no retraining, and applies directly to any new asset. We leave fine-tuning strategies for future work.
A TSFM can be expressed as a parametric mapping from an observed history to a forecast distribution. We denote by $\sigma_{1:T} = (\sigma_1, \ldots, \sigma_T)$ the context window of $T$ past observations (in our case, $T = 1000$ daily realized volatility values). A TSFM with pretrained parameters $\hat{\theta}$ produces a forecast of the next $H$ values:
where $f_{\hat{\theta}}$ maps the input sequence to a predictive distribution over future values.\footnote{More precisely, $f_{\hat{\theta}}$ outputs a distribution $\hat{p}(\sigma_{T+1:T+H} \mid \sigma_{1:T}; \hat{\theta})$ from which we extract the conditional mean as the point forecast.} The model learns $\hat{\theta}$ during a pretraining phase on large external corpora of time series spanning weather, energy, retail, transport, and macroeconomic domains. At inference time, the model receives only the context window $\sigma_{1:T}$ and produces forecasts without any task-specific parameter updates; this is the zero-shot setting.
Every TSFM implements the mapping in Eq. (ref) through three stages: tokenization (converting $\sigma_{1:T}$ into the model's internal tokens; the strategies differ across models and are summarized in Tab. (ref)), encoding by a transformer vaswani2017 that attends across positions without imposing HAR's fixed daily, weekly, and monthly structure, and decoding by a head that maps representations back to forecasts. Transformers are either encoder-only (reading the whole input at once) or decoder-only (left-to-right, autoregressive); decoding heads either output quantiles directly, parameterize an explicit distribution (e.g., Student-$t$) that is sampled, or, for Sundial, generate trajectories via continuous normalizing flows. In all cases we extract the conditional mean as the point forecast, the summary aligned with QLIKE (Sec. (ref)): models with a mean head (Chronos-Bolt, TimesFM 2.5) expose it directly, sampling models (Lag-Llama, Sundial, Moirai-MoE) use the sample mean, and quantile-only models (Moirai 2.0) integrate the predictive quantile function; Toto is the one exception, discussed below.
All TSFMs receive a rolling context window of 1{,}000 daily realized volatility observations as input, with no additional covariates. We set the context length to 1{,}000 days to match the 1{,}000-day estimation window used for the econometric models, so that both classes of model condition on the same span of history.\footnote{Two architectures cannot accommodate a 1{,}000-day context and are run at their native 512-token limit: Moirai-MoE, whose positional encoding is fixed at 512 tokens, and TTM, whose r2.1 branches max out at a 512-day context. For these two models the context window is 512 days; all other TSFMs use 1000.}
For multi-step horizons ($h > 1$), the TSFM produces a trajectory of $h$ individual-step forecasts $(\widehat{\sigma}_{T+1}, \ldots, \widehat{\sigma}_{T+h})$. Our primary target is point-in-time realized volatility $\sigma_{T+h}$, the value $h$ steps ahead, so the point forecast at horizon $h$ is the $h$-th element of this trajectory, $\widehat{\sigma}_{T+h}$, not the average over the trajectory. Each trajectory element is the conditional mean of the model's predictive distribution at that step, obtained as described above. We use the conditional mean rather than the median because it is the point summary aligned with QLIKE, our primary evaluation loss;\footnote{The QLIKE-optimal forecast is the conditional mean of the variance, $\mathrm{E}[RV_{t+h}\mid\mathcal{F}_t]$. Because we model the volatility series and square the conditional-mean volatility forecast back to a variance for QLIKE, the two differ by a Jensen term equal to the conditional variance of the volatility forecast. The conditional mean nonetheless remains preferable to the conditional median, which carries an additional bias. We apply the same point-forecast construction to all 17 models and flag this Jensen term as a caveat, since recovering $\mathrm{E}[RV]$ exactly would require model-specific second-moment extraction that several sampling- and quantile-based TSFMs do not expose.} for the one heavy-tailed case (Toto) we use the analytic mean of the parameterized distribution for numerical stability.
A natural concern with zero-shot evaluation is whether TSFM performance is inflated because the models' pretraining corpora overlap with the test series. Tab. (ref) summarizes each model's training data and its financial content. No TSFM in our study was trained on realized volatility. VOLARE's realized variance series are second-moment statistics derived from ultra-high-frequency intraday returns at 5-minute sampling volare2026; this quantity does not appear in any known training corpus.
We treat contamination as a genuine concern. The models we evaluate were released in 2024 and 2025, and their pretraining windows can overlap our evaluation period in calendar time. Exact pretraining data cutoffs are not published for most of these models, so we cannot establish temporal disjointness between training and evaluation by date alone, and the evidence below is therefore indirect. For the models that do not disclose their corpora (Moirai 2.0, Moirai-MoE, Sundial), we cannot verify what they contain, and text or auxiliary contexts seen during pretraining could carry market-volatility information indirectly even without realized variance series being present. We concede this as a limitation that we cannot fully rule out for the undisclosed-corpus models.
Several pieces of evidence nonetheless make contamination an unlikely full explanation of our findings. First, the pattern of results is the opposite of what memorization or leakage would produce. Our findings are not a broad TSFM win: under robust loss ratios, most TSFMs do not beat Log-HAR, and only TTM does so consistently. Widespread contamination would inflate many models at once, not leave the bulk of models at or below the econometric benchmark. Second, the single model that wins, TTM, has the smallest capacity in our study (fewer than 1M parameters) and therefore the least room to memorize specific series; leakage benefiting the lowest-capacity model is hard to reconcile with a memorization account. Third, realized variance is a specific intraday-derived second moment computed at 5-minute sampling, absent from the known public corpora (e.g., Monash,\footnote{\url{https://forecastingdata.org/}} LOTSA\footnote{Large-scale Open Time Series Archive; see woo2024moirai.}), which contain raw series rather than this derived statistic. For models with disclosed corpora, financial data constitutes less than 1% of training observations (Tab. (ref)), limited to daily exchange rates and macroeconomic indicators.
We carry this into the conclusion as a limitation rather than treating it as resolved, but the more parsimonious reading of the pattern is that the modest advantage we document reflects transfer of general temporal structure (mean reversion, long memory, regime persistence) rather than memorization of specific financial dynamics.
We employ a walk-forward evaluation scheme with a rolling origin. Econometric models use a fixed estimation window of 1{,}000 trading days (approximately four years): at each origin, we estimate the model on the most recent 1{,}000 observations, and the window slides forward one day at a time, producing daily re-estimated forecasts. Daily re-estimation is the most demanding refresh cadence and avoids any look-ahead from stale coefficients, at no cost to the comparison because the same timing applies to every model. A window of roughly four years is standard in the realized volatility forecasting literature and is long enough to limit the sensitivity of least-squares estimates to individual volatility spikes bollerslev2016harq, clements2021practical. This matches the 1{,}000-observation context window supplied to the foundation models, so the two model classes condition on the same amount of history. Foundation models follow the same walk-forward timing but require no estimation step; only the context window slides forward. All model comparisons are conducted on the common out-of-sample period where both econometric and TSFM forecasts are available. For the equity sample (2,786 days), a 1,000-day window leaves 1,786 daily forecasts per model; because the three augmented HAR variants (HAR-J, HAR-RS, HARQ) need an extra 22-day monthly lag to construct their auxiliary regressors and therefore begin 22 days later, the common out-of-sample window shared by all 17 models is 1,764 forecasts per asset at $h = 1$. The FX and futures samples (over 4,000 days) yield substantially more.
We evaluate three forecast horizons: $h = 1$ (one day), $h = 5$ (one week), and $h = 22$ (one month). The forecast target is the point-in-time realized volatility $h$ days ahead, $\sigma_{t+h} = \sqrt{RV_{t+h}}$, not an average over the intervening days. A multi-day average overlaps with itself for $(h-1)/h$ of its content from one origin to the next, which induces serial correlation in the target and overstates its persistence relative to the underlying daily series, so it is not the quantity a forecaster conditioning on day-$t$ information wants to predict at horizon $h$. We produce the horizon-$h$ forecast by iterating the one-step recursion forward for the pure-RV models and by direct $h$-step estimation for the augmented HAR variants, as described in Sec. (ref); each TSFM returns the $h$-th element of its forecast trajectory. We report results for the $h$-day-average target as an additional robustness arm.
We report mean squared error (MSE) on the volatility scale, computed on $\sigma_{t+h}$ and $\widehat{\sigma}_{t+h}$ directly, but we focus on the QLIKE loss of patton2011, which is robust to noise in the realized-variance proxy and less sensitive to extreme volatility observations than MSE:
where $\widehat{RV}_t$ is the variance forecast and $RV_t$ is the realized variance. QLIKE is a loss on the variance scale, and its proxy-robustness property in patton2011, consistency of the forecast ranking under a conditionally unbiased variance proxy, holds for the variance, not its square root. We therefore evaluate QLIKE by squaring the volatility forecast back to a variance, $\widehat{RV}_t = \widehat{\sigma}_t^2$, and using realized variance $RV_t = \sigma_t^2$ as the proxy. This keeps MSE on the interpretable volatility scale while preserving QLIKE's proxy-robustness guarantee.
We winsorize each forecast to the in-sample support of realized volatility for that asset, $[\sqrt{\min RV},\ \sqrt{\max RV}]$, where the minimum and maximum are taken over the asset's full sample of realized variances. This is a wide guardrail rather than a tight bound. Because QLIKE diverges as the forecast approaches zero, an unbounded or floor-clipped forecast can distort it; bounding to the data's own support removes that distortion without an arbitrary constant. The lower bound keeps forecasts within the range of volatility the model was estimated on, avoiding the distortion an arbitrary numerical floor would introduce; the upper bound guards against the occasional extreme spike that the heavy-tailed predictive distributions of some TSFMs can produce. We apply the bound symmetrically to all 17 models, eight econometric and nine TSFM specifications, so that no model class is treated differently.
Pairwise forecast comparisons use the DM test on QLIKE loss differentials, with a Newey--West heteroskedasticity-and-autocorrelation-consistent variance newey1987 using a Bartlett kernel and $h-1$ lags to absorb the autocorrelation that multi-step loss differentials carry by construction. To address the multiple comparison problem inherent in evaluating 17 models, we compute the MCS, which identifies the subset of models whose forecasting ability cannot be statistically distinguished from the best model at a given significance level. We use the $T_{\max}$ statistic of hansen2011mcs with a moving-block bootstrap (block length 22 days, one trading month and the longest forecast horizon; $B = 10{,}000$ replications) at the $\alpha = 0.10$ level. Because pooled averages across assets can be dominated by a few high-loss series, we also report average loss ratios relative to Log-HAR: the per-asset QLIKE divided by Log-HAR's QLIKE on that asset, averaged across assets. Normalizing each asset by its own benchmark prevents a few high-volatility assets from dominating, as they do in the pooled mean of raw losses. These appear in Tab. (ref).
Additionally, we assess forecast efficiency using the MZ regression mincer1969:
where $\widehat{\sigma}_t$ is the model's volatility forecast and $\sigma_t$ the realized volatility. We run the regression on the volatility scale. Under forecast optimality, $\alpha = 0$ and $\beta = 1$, meaning the forecast is unbiased and captures the correct scale of variation. We test the joint null $H_0\!: \alpha = 0, \beta = 1$ using a Wald test with Newey--West standard errors and report cross-asset averages and rejection rates at the 5% level. We apply the MZ-based affine correction symmetrically to all models, econometric and TSFM alike, because the unbiasedness of least squares is an in-sample property that need not carry over to the out-of-sample forecasts produced under a rolling window.
To examine whether relative forecast performance is stable over time, we apply the giacomini2010 GR fluctuation test. For each model paired against the Log-HAR benchmark, we compute a rolling DM statistic over a window of size $m = \lfloor 0.3 \times T \rfloor$, producing a time path of relative QLIKE performance; the window fraction and the associated critical values follow giacomini2010. The test statistic is $\sup_t |S_t|$, compared against critical values from the distribution of the supremum of a standardized Brownian bridge. Rejection of the null indicates that the two models do not have equal predictive ability at every point in the sample, with one model significantly more accurate over some subperiod.
This section presents the realized volatility forecasting results.\footnote{Replication code is available at \url{https://github.com/Alessiobrini/tsfm-rv}.} We first examine aggregate forecast accuracy across loss functions and horizons, then analyze cross-asset and cross-market heterogeneity.
Tab. (ref) reports the cross-sectional average of per-asset loss functions across the 40 equities, and Tab. (ref) pools the same per-asset losses across all 50 assets (equities, FX, and futures). Fig. (ref) illustrates forecasts against realized values for four assets spanning equities and FX.
On the pooled cross-sectional average across all 50 assets (Tab. (ref)), several foundation models record low QLIKE. TTM achieves the lowest pooled QLIKE at every horizon (0.190 at $h = 1$), but the HAR family and a wide tier of foundation models sit within a narrow band just behind it, with the gaps among the leaders rarely exceeding 0.02 to 0.03 in QLIKE (Tab. (ref)). On its own, this pooled average would suggest a broad tier of foundation models matching or beating the HAR family.
However, the pooled mean is dominated by a few high-volatility assets: a model that does well on those assets can post a low average even if it loses to Log-HAR on most assets. To correct for this, Tab. (ref) reports the average across assets of each model's per-asset QLIKE ratio to Log-HAR, an aggregation that weights every asset equally and is not driven by outliers. Under this measure, only one foundation model beats Log-HAR at every horizon: TTM, with ratios of 0.982, 0.986, and 0.987 at $h = 1$, $5$, and $22$, an improvement of roughly 1.3 to 1.8%. TTM has fewer than one million parameters, the smallest model in the evaluation.
No other foundation model matches TTM's consistency. Sundial matches Log-HAR at the daily horizon (ratio 0.998) and falls behind at the two longer horizons. Moirai 2.0, Moirai-MoE-S, TimesFM 2.5, Chronos-Bolt, and Toto have loss ratios above one at all three horizons, so they lose to Log-HAR on the typical asset. The competitive econometric benchmarks, by contrast, cluster near Log-HAR: HAR matches it at $h = 1$ (0.998) before deteriorating slightly at longer horizons, and ARFIMA, ARMA, and the MEM all sit within a few percent of parity across horizons (Tab. (ref)). The contrast between the pooled means in Tab. (ref) and the loss ratios in Tab. (ref) is the central result of this section. The pooled average makes the foundation-model class look better than it is; once each asset is weighted equally, only TTM delivers a consistent improvement over the strongest econometric benchmark. This ranking does not depend on how the per-asset ratios are averaged: under a scale-symmetric geometric mean, TTM remains the only model below one at every horizon, so the arithmetic average is not what produces the result.
Two features of the loss-ratio table merit comment. HARQ is genuinely poor, with loss ratios as high as 5.132 at $h = 1$: its realized-quarticity correction does not help on this data and instead amplifies noise. The inflated QLIKE reflects HARQ's actual forecasts, not any flooring or clipping. The long-memory and multiplicative benchmarks, by contrast, are close to Log-HAR: ARFIMA, ARMA, and MEM all carry loss ratios near parity at $h = 1$ (1.012, 1.000, and 1.014) and stay within just over ten percent of Log-HAR at the longer horizons (Tab. (ref)), so fractional differencing and the multiplicative error structure track the log-HAR benchmark closely rather than falling well behind it on this sample.
Toto is competitive on pooled QLIKE (0.234 at $h = 1$) but its QLIKE spikes on a small number of commodity futures with extreme 2020 realized volatility, principally Gold (GC) and Crude (CL). These few contracts dominate the pooled cross-sectional average. The loss-ratio aggregation, which down-weights such outliers, places Toto near the middle of the foundation-model class rather than at the bottom. We return to this distinction in the cross-asset analysis.
Tab. (ref) extends the analysis to foreign exchange and futures markets.
The FX sample covers five major currency pairs with lower volatility levels and moderate persistence. On the cross-sectional average (Tab. (ref), Panel A), the spread among the leading models is narrow. At $h = 1$, TTM and the leading econometric and foundation models cluster within about 0.005 of each other on QLIKE, with Log-HAR at or near the top. At $h = 5$, Log-HAR ties the best foundation models on QLIKE, so no foundation model separates from the best econometric specification. At $h = 22$, Log-HAR has the lowest QLIKE, with TTM next; Moirai-MoE-S degrades sharply on FX at the monthly horizon, indicating that its forecasts diverge on the lower-amplitude currency series. On FX, Log-HAR is at or near the top at every horizon, and the foundation-model advantage, where it exists, is small.
The futures sample (Tab. (ref), Panel B) displays the widest cross-asset heterogeneity. At $h = 1$, Sundial and TTM tie for the lowest QLIKE, with HAR and Log-HAR just behind. At $h = 5$, Log-HAR has the lowest QLIKE, just ahead of HAR and TTM; at $h = 22$, Log-HAR leads more clearly, with MEM, HAR, and TTM following. The level HAR variants HAR-RS and HARQ produce inflated QLIKE on futures, and Toto's QLIKE spikes on the commodity contracts, consistent with the outliers noted above.
Three patterns emerge. First, TTM is the most consistent foundation model across classes: it is at or near the lowest QLIKE on equities, FX, and futures at most horizons, and it is the only foundation model to beat Log-HAR under the equal-weighted loss ratio (Tab. (ref)). No other foundation model achieves this consistency. Second, Log-HAR is the best or near-best specification on FX and futures across horizons, and second best on equities, confirming the practical value of the log transformation for forecast positivity and alignment with QLIKE. Third, dispersion within the foundation-model class is wide: apart from TTM, only Sundial reaches parity, and only at the daily horizon, while Moirai-MoE-S, TimesFM 2.5, Chronos-Bolt, and Toto lose to Log-HAR on the typical asset. This cross-asset variation reflects the match between each model and the distribution of realized volatility. Realized volatility, though persistent, is stationary and mean-reverting, closer to the series these models encounter in pretraining than the highly persistent macroeconomic series on which carriero2024macro find foundation models struggle.
Fig. (ref) provides a distributional view, plotting QLIKE ratios (model / Log-HAR) across all 50 assets. Values below one indicate the model outperforms Log-HAR. The competitive econometric benchmarks (HAR, ARFIMA, ARMA, MEM) cluster just above parity, and TTM is the only model whose distribution sits predominantly below one at all three horizons.
The rankings in Sec. (ref) show differences across models, but mean loss comparisons can be misleading when distributions are skewed and sample sizes vary across assets. We now apply four formal statistical tests: the MCS hansen2011mcs to identify the subset of models that cannot be distinguished from the best, pairwise DM tests diebold1995 to quantify directional win rates, MZ regressions mincer1969 to assess forecast efficiency, and GR fluctuation tests giacomini2010 to detect time variation in relative performance. Tab. (ref) reports MCS inclusion rates and pairwise DM win rates for all 50 assets.
The MCS sharpens the loss-ratio finding. TTM is the single dominant specification, with an all-horizon average inclusion rate of 0.96 that no other model approaches (Tab. (ref)). Behind it sits a broad set of econometric benchmarks rather than a single close competitor: ARMA, Log-HAR, HAR, ARFIMA, and the MEM all enter for three-quarters or more of assets at the daily horizon, with Log-HAR the most consistently admitted across horizons (86 to 90%). Among the remaining foundation models, Sundial enters frequently at the daily horizon (92%) but its inclusion fades at the longer horizons, and the rest of the class enters only rarely on average. Lag-Llama is the only reversal: it improves sharply at the monthly horizon (84%), mirroring its low pooled QLIKE there. The MCS therefore identifies TTM as the single strongest model but admits a wide set of econometric specifications alongside it, especially at the daily horizon, rather than a narrow two-model frontier.
We compute the DM test pairwise: for each of the 17 models in Tab. (ref), we test it against each of the remaining 16 on each of the 50 assets. Tab. (ref) reports the fraction of those tests in which each model achieves significantly lower QLIKE at the 5% level.\footnote{These win rates are unadjusted for multiple comparisons. Applying a Benjamini--Hochberg false-discovery-rate correction benjamini1995 at 5% within each model's 800 tests shrinks every win rate, most at the monthly horizon where pairwise differences are weak, but preserves the ranking: TTM retains the highest adjusted win rate at $h = 1$ and $h = 5$ (49% and 53%) and the small unadjusted gap at $h = 22$ closes to a tie (11% versus 11%).} TTM wins the highest fraction at the short and medium horizons, with Log-HAR the most consistent econometric model and edging TTM on the pairwise measure at $h = 22$. Win rates compress across all models at $h = 22$, consistent with forecast differences narrowing as horizons lengthen. The Chronos-Bolt checkpoints, TimesFM 2.5, and Moirai-MoE-S win a small share of comparisons beyond the daily horizon, confirming that the loss-ratio shortfall translates into pairwise losses.
Tab. (ref) reports average MZ regression results across 50 assets for all three horizons. The MZ regression, $\sigma_t = \alpha + \beta \, \widehat{\sigma}_t + \varepsilon_t$ on the volatility scale, tests whether forecasts are efficient: under the null, $\alpha = 0$ and $\beta = 1$. We apply the affine correction underlying these regressions symmetrically to all models, not to the foundation models alone (Subsec. (ref)). The horizon pattern reveals a calibration split. At the daily horizon the econometric benchmarks are the more efficient forecasts: Log-HAR has a slope of 1.043 and rejects the joint efficiency null for only 2% of assets, and ARMA and HAR for 6%. TTM, despite a slope close to unity (1.010), rejects for 88% of assets, and the other foundation models reject for essentially all of them. The high rejection rates of the foundation models at $h = 1$ reflect small but systematic departures from $(\alpha,\beta)=(0,1)$ rather than poor point accuracy: because TTM's slope is near unity, the rejections reflect precision in detecting small biases rather than a miscalibrated scale; Log-HAR's slope of 1.043 lies farther from one yet rejects for only 2% of assets, because its least-squares fit leaves the forecast efficient in sample by construction. As the horizon lengthens to $h = 22$, the econometric benchmarks undershoot more sharply, with slopes well below one (Log-HAR 0.532, HAR 0.364) and MZ $R^2$ falling toward zero, while the foundation models retain slopes closer to one (TTM 0.670, Moirai 2.0 0.679, TimesFM 2.5 0.665), a longer-horizon pattern consistent with christensen2023ml, who attribute the stronger relative performance of flexible models at longer horizons to their higher persistence approximating the long memory of realized volatility. At the daily horizon, then, Log-HAR is the more MZ-efficient forecast, and the foundation-model advantage in calibration appears only at the longest horizon.
{
}
Fig. (ref) plots the rolling DM statistic (QLIKE loss) of each foundation model against Log-HAR across the three forecast horizons, averaged across 50 assets. A positive rolling DM indicates that the comparison model produces lower QLIKE than Log-HAR during that window, so the foundation model outperforms the benchmark; a negative value indicates that Log-HAR is the more accurate. Values beyond the $\pm 2.80$ critical-value bands indicate rejection of the equal-predictive-ability null at the 5% level.
The GR tests show that relative performance against Log-HAR is time-varying rather than constant. TTM's rolling paths are the most stable: at the daily horizon it tracks close to parity with Log-HAR, and at the monthly horizon it rises toward the upper 5% band around the 2020 to 2021 volatility episode. The other foundation models sit below parity for most of the sample, consistent with their loss ratios above one. Most rolling statistics nonetheless remain within the $\pm 2.80$ bands, so for much of the sample the difference from Log-HAR is not statistically significant; the weakest of them breach the lower band over sustained windows, marking periods in which Log-HAR significantly outperforms them. The practical implication is that unconditional DM and MCS results, while informative about average performance, mask substantial temporal variation, and that the TTM versus Log-HAR ordering, though stable on average, is not constant period by period.
The four tests point to the same conclusion as the loss ratios: the foundation-model class is not uniformly strong, and only one model separates as the single best, though a competitive set of econometric benchmarks stays close behind it. The MCS places TTM at the frontier (0.96 all-horizon inclusion; Tab. (ref)) with Log-HAR the most consistently admitted econometric model (0.89) and HAR, ARMA, ARFIMA, and the MEM also entering frequently, especially at the daily horizon. The DM tests confirm that TTM wins the largest share of pairwise comparisons at the short and medium horizons (Tab. (ref)). The MZ regressions show that Log-HAR is the more efficient forecast at the daily horizon, while foundation models hold their calibration better at $h = 22$, where the HAR family undershoots (Tab. (ref)). The GR test shows that relative performance against Log-HAR is time-varying (Fig. (ref)). The wide heterogeneity across foundation-model architectures, from TTM's consistent frontier position to the parity-or-worse loss ratios of Chronos-Bolt, TimesFM 2.5, and Moirai-MoE-S, indicates that pretraining-data composition and output mechanism matter far more than membership in the foundation-model class.
The results in Secs. (ref) and (ref) establish model rankings and their statistical significance. A natural question is whether these rankings are stable across market regimes, estimation choices, and evaluation parameters. We assess this stability, then ask how much of the headline advantage is calibration rather than information, and whether combining the leading models improves on either alone.
Tab. (ref) splits the full sample at March 1, 2020, the onset of the COVID-19 volatility spike in U.S.\ markets, and reports forecast accuracy separately for the pre-COVID and post-COVID periods across all 50 assets. The qualitative ordering of the main analysis carries over to both regimes: TTM remains among the strongest specifications, with Log-HAR and the rest of the competitive econometric benchmarks clustered alongside it, and the foundation models that lose to Log-HAR on the full sample continue to do so within each sub-period. The gap between the leading models and the rest of the models widens in the post-COVID regime, where realized volatility is higher and more variable, but the relative ranking is preserved across the split. The level HAR variants are the least stable, with their QLIKE rising most in the high-volatility post-COVID sample, while Log-HAR and the log and multiplicative benchmarks (ARFIMA, ARMA, MEM) are comparatively stable. At the monthly horizon in the calm pre-COVID sample, every model's QLIKE exceeds one (daggers in Panel A), a level effect shared across all specifications that leaves the cross-model ordering unchanged. The model ranking is therefore stable across the COVID structural break.
Our headline results forecast the point-in-time realized volatility. Re-running the full pipeline under an alternative, $h$-day-average target reproduces the ranking and does not change the main result; we report the details in Appendix (ref) (Tab. (ref)).
We test whether the rankings depend on the winsorization bounds. They do not, because the bound rarely binds: across the 50 assets and 17 models (5.2 million forecast-date pairs in total), a forecast is clipped in only 0.05% of cases, and the upper cap binds in under 0.01%. Because the cap equals each asset's sample maximum realized volatility, it does not truncate genuine high-volatility forecasts, including those around the 2020 spike; it removes only the occasional extreme draw that lies beyond any realized value.
We also examine how foundation-model accuracy varies with the TSFM context length. Context effects are model-specific and horizon-dependent, but the conclusion that TTM is the only foundation model to beat Log-HAR under the equal-weighted loss ratio is unaffected; the full sensitivity analysis is in Appendix (ref) (Tab. (ref)).
A robustness concern is that a model's low loss could reflect good calibration, that is, alignment between the scale of forecasts and realizations, rather than genuine predictive information about future volatility dynamics. A Mincer--Zarnowitz recalibration shows that this concern refines rather than overturns the central message: at the shorter horizons TTM's edge over Log-HAR is largely a calibration effect that several other foundation models also enjoy, while at the monthly horizon TTM retains a genuine informational advantage. The MZ regression separates the two: regressing realized values on forecasts yields intercept and slope parameters that remove any linear bias, so that accuracy surviving the correction reflects information beyond an affine rescaling. As in the efficiency regressions, we apply the correction symmetrically to all models. Specifically, we estimate $\hat{\alpha}_t$ and $\hat{\beta}_t$ recursively from daily-origin forecasts up to $t - 1$ and form the corrected forecast $\widehat{\sigma}_t^{\text{MZ}} = \hat{\alpha}_t + \hat{\beta}_t \widehat{\sigma}_t$ using an expanding estimation window, at each of the three horizons. Tab. (ref) compares original and corrected QLIKE across $h = 1, 5, 22$.\footnote{We begin scoring the recursively corrected forecast only after a 252-day (one trading year) expanding-window warm-up, so the affine correction is estimated on a full year of forecasts before it is used, and the “Orig.” column is scored on this same post-warm-up window. Its QLIKE therefore differs slightly from the full-window pooled means in Tab. (ref) (for example, TTM at $h = 1$ is 0.192 here versus 0.190 there); the two are not in conflict. We verified that the qualitative reordering, with TTM in the middle of the ranking at the daily and weekly horizons and best at the monthly horizon, is robust to halving the warm-up to 126 days.} The effect of the correction depends on the horizon. At the daily and weekly horizons the correction reorders the leading models. The econometric benchmarks are already close to affine-efficient, so it leaves them almost unchanged: at $h = 1$, HAR and Log-HAR each move by at most about $0.002$ in QLIKE. Several foundation models improve sharply once the affine bias is removed. At $h = 1$, TimesFM 2.5 falls from $0.214$ to $0.190$ in QLIKE, Chronos-Bolt (base) from $0.221$ to $0.192$, and Moirai 2.0 from $0.204$ to $0.193$. Their raw forecasts therefore carry predictive information that a level-and-scale bias had masked. TTM moves the opposite way (0.192 to 0.198 at $h = 1$), dropping to eighth on the corrected QLIKE at both short horizons, where the eight best models lie within $0.01$ of one another. At these horizons part of TTM's raw advantage is calibration rather than superior information about future dynamics; good calibration is itself useful, but here it is not exclusive to TTM, since the same affine recalibration confers it on several econometric and foundation models alike. The monthly horizon is different. At $h = 22$, TTM remains the single most accurate model after the correction, with a corrected QLIKE of $0.484$ against $0.506$ for the recalibrated Log-HAR, ahead of ARFIMA (0.493) and the Chronos-Bolt checkpoints. At the longest horizon, TTM's edge is therefore not an artifact of calibration: it survives a recursive affine recalibration applied symmetrically to every model.
TTM and Log-HAR draw on different information, a foundation model's pretrained temporal structure and the HAR's explicit multi-horizon memory, so combining them may improve on either alone. We form two combinations of the TTM and Log-HAR volatility forecasts bates1969, timmermann2006: an equal-weight average, and a recursive Bates--Granger combination whose variance-minimizing weight is estimated from forecast errors observed strictly before each date (expanding window, clipped to $[0,1]$, with an equal-weight warm-up). Tab. (ref) reports the results. Both two-way combinations beat Log-HAR on the average loss ratio at every horizon and are at or below TTM's loss ratio as well: the equal-weight combination attains ratios of 0.977, 0.984, and 0.982, and the Bates--Granger combination 0.981, 0.982, and 0.978, each at or below TTM's 0.982, 0.986, and 0.987 at all three horizons. By Diebold--Mariano test the equal-weight combination has significantly lower QLIKE than Log-HAR on 36 of 50 assets at $h = 1$, but significantly beats TTM itself on at most 12, so it matches rather than dominates the best single model, the pattern expected from the forecast-combination puzzle timmermann2006 in which equal weights are hard to beat. Adding a third member, ARMA (the best-performing pure time-series benchmark, with a loss ratio of 1.000 at $h = 1$ in Tab. (ref)), to form a three-way combination helps only at the daily horizon: the equal-weight three-way average attains 0.981 at $h = 1$ but 0.994 and 1.013 at $h = 5$ and $h = 22$, where it adds nothing beyond the two-way combination and slips above parity at the monthly horizon. The combination is most valuable in the MCS: with the combinations added to the candidate set and the MCS recomputed across all 50 assets, the equal-weight two-way combination enters the set for 98 to 100% of assets across horizons, above TTM (82 to 94%) and far above Log-HAR (38 to 88%) and every other single foundation model. The practical reading reinforces our central message. A forecaster cannot know in advance which foundation model will work, the model-selection problem that, as this paper shows, dominates the foundation-versus-econometric choice. By averaging Log-HAR with TTM, the one small foundation model that beats it, that forecaster obtains accuracy matching the best single model while improving on the benchmark for most assets.
We evaluate nine zero-shot TSFMs for realized volatility forecasting across 50 assets spanning U.S.\ equities, foreign exchange, and futures, and compare them to eight econometric benchmarks with formal pairwise and multi-model forecast-comparison tests.
Three main findings emerge. First, only one foundation model beats a well-specified Log-HAR once each asset is weighted equally rather than pooled: TTM, the smallest model in the evaluation, is the only TSFM with a QLIKE loss ratio below one at every horizon (Tab. (ref)), by a narrow margin, and the MCS agrees. A uniform MZ recalibration (Sec. (ref)) shows this edge is largely a calibration effect at the shorter horizons, where several other TSFMs match TTM, but a genuine informational advantage at the monthly horizon. Second, the broad TSFM advantage in pooled means is largely an artifact of a few outlier assets; under loss-ratio aggregation the advantage shrinks to a single foundation model, while a competitive set of econometric benchmarks (HAR, ARFIMA, ARMA, and the MEM) clusters near parity with Log-HAR. Third, performance is so heterogeneous across architectures that model selection within the TSFM class matters more than the choice between TSFMs and econometric models.
For practitioners, the implication is measured. A general-purpose foundation model is not a drop-in improvement over HAR for realized volatility: most of the models we test do not beat Log-HAR on a typical asset. TTM is the exception, edging Log-HAR at all three horizons while running efficiently on CPU, which makes it a reasonable alternative to consider rather than a default to adopt. Because the edge is thin and even TTM is not best on every asset, a simple equal-weight average of TTM and Log-HAR matches the best single model and enters the MCS for 98 to 100% of assets across horizons, so a forecaster need not identify the best model for each asset in advance. For underperforming TSFMs, MZ bias correction can recover part of the underlying signal. The wide dispersion across architectures cautions against treating TSFMs as a uniform class; evaluating multiple architectures before deployment remains essential.
Several limitations apply. We evaluate only zero-shot performance; fine-tuning on realized volatility data may yield further gains, and a follow-up study examining parameter-efficient fine-tuning of the stronger architectures is a natural next step given how narrowly even the best zero-shot model beats the benchmark. We evaluate only point forecasts, specifically the conditional mean of each model's predictive distribution. As documented in Sec. (ref) (Tab. (ref)), no TSFM in our study was trained on realized volatility or any intraday-derived statistic, and financial series constitute less than 1% of training observations for all models with disclosed corpora. For the three models with undisclosed composition (Moirai 2.0, Moirai-MoE, Sundial), we cannot fully rule out indirect exposure to related daily financial series. Three extensions follow naturally: density evaluation under proper scoring rules such as the continuous ranked probability score, realized covariance forecasting and portfolio construction, and temporal holdout designs that further isolate the pretraining-data channel. The broader lesson is that for realized volatility the consequential choice is not foundation model versus econometric benchmark but which model within the foundation class, and a thin, recoverable edge is best captured by combining the leading models rather than selecting among them.
\paragraph{Conflict of interest.} The author declares no conflict of interest.
\paragraph{Funding.} This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
\paragraph{Data availability.} The realized measures analyzed in this study are derived from the VOLARE dataset, which is constructed from proprietary high-frequency tick data licensed from Kibot. The underlying tick data are not redistributable by the author, and the VOLARE-derived realized measures are available from the VOLARE project subject to its terms of use. The code that reproduces all results in this paper is openly available at \url{https://github.com/Alessiobrini/tsfm-rv}.