EconBase
← Back to paper

ReSGA: A Large Tail Risk Model for Learning Value-at-Risk and Expected Shortfall

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

99,165 characters · 13 sections · 58 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

ReSGA: A Large Tail Risk Model for Learning Value-at-Risk and Expected Shortfall

frontmatter{\thankstext{t2}{A user-oriented companion website for this paper, including stock-level tail risk forecasts, documentation, and links to the data and source code, is available at \url{https://tailrisk-resga.github.io}.}} \thankstext{t3}{Correspondence to Zhoufan Zhu (Email: [email removed])} \begin{aug} , \and \thanksref{t3} \end{aug} \begin{abstract} Learning Value-at-Risk (VaR) and Expected Shortfall (ES) is important for managing financial risks effectively. Existing approaches with limited parameters are vulnerable to model misspecification in the era of big data. To address this limitation, we propose a large tail risk model, the retrieval-enhanced self-grouping autoencoder (ReSGA), which is designed with millions of parameters to exploit the rich cross-sectional dependence and long-term temporal dynamics of assets using their characteristics. Applied to monthly US equity returns from 1926 to 2023 with 153 firm characteristics, ReSGA outperforms twelve econometric and machine learning competitors in terms of out-of-sample loss and statistical backtesting. In addition, its forecast advantages can translate into significant economic gains from long-short decile portfolios that are constructed by a new size-enhanced left-side momentum strategy. To clarify the role of complexity, we further conduct a systematic scaling analysis and demonstrate that improvements in joint VaR-ES forecasting are primarily driven by data complexity rather than model complexity. Finally, our analyses of group-importance and transfer-learning exhibit the interpretability and cross-market generalizability of ReSGA. \end{abstract} \begin{keyword} \kwd{AI for Finance, Expected Shortfall, Large Tail Risk Model, Value-at-Risk, Virtue of Complexity} \end{keyword}

Introduction

Value-at-Risk (VaR) and Expected Shortfall (ES) are central measures of financial tail risk in both academic research and regulatory practice. Considering the loss of a portfolio or investment, VaR characterizes only a quantile of the distribution of this loss, while ES summarizes expected losses beyond that quantile and therefore captures the severity of extreme downside outcomes. Moreover, ES is a coherent risk measure, satisfying properties like subadditivity, monotonicity, positive homogeneity, and translation invariance artzner1999. In contrast, VaR is generally not subadditive, implying that portfolio risk measured by VaR may exceed the sum of individual risks. These advantages make ES more consistent with the diversification principles of portfolio theory and capture the tail risk in a more comprehensive manner, explaining its growing role in modern risk management BCBS2019FRTB\footnote{For more studies on VaR and ES, one can refer to Acerbi2002on, Rockafellar2002Conditional, AcerbiSzekely2014, du2017backtesting, Li2023PELVE, and Chronopoulos2024Forecasting.}.

Although conceptually attractive, ES remains challenging to forecast in practice. A fundamental obstacle is that ES on its own is not elicitable, whereas VaR is gneiting2011. In other words, there exists no loss function for which ES is the unique minimizer of expected loss. This lack of direct elicitability partly explains why the development of ES forecasting methods has lagged behind that of VaR. A major breakthrough is provided by fissler2016, which introduce a class of consistent scoring rules, known as Fissler-Ziegel (FZ) loss functions, to utilize the joint elicitability of VaR and ES. This breakthrough leads to a growing literature on joint VaR-ES forecasting. For example, patton2019 employ FZ loss functions in semi-parametric settings based on the generalized autoregressive score (GAS) model and the generalized autoregressive conditional heteroskedasticity (GARCH) model; taylor2019 exploits the link between the asymmetric Laplace likelihood and the FZ loss for joint VaR-ES estimation; merlo2021 further extend the Laplace-based approach to multivariate regression settings. Until now, most existing work on joint VaR-ES forecasting remains confined to parametric or semi-parametric models. In such frameworks, the number of parameters is kept modest to ensure interpretability, and each asset is often modeled separately with its own set of parameters. While tractable, their performance depends critically on correct model specification, where the true data-generating process for tail risks may be far more complex than a GARCH or GAS model can capture.

Fortunately, the era of big data brings new opportunities for tail risk modeling. Risk forecasters now have access to decades of financial data, covering tens of thousands of assets, along with hundreds of firm characteristics. This abundance of information naturally raises a key question: Can we develop a large tail risk model, analogous to large language models (e.g., ChatGPT, Gemini, and Claude) in natural language processing, that learns from rich and diverse data to predict VaR and ES more accurately? In other words, instead of relying on asset-specific (semi-)parametric models, could a single high-capacity universal model trained on massive data serve as a general engine for joint VaR-ES forecasting?

To shed light on large tail risk models, we first develop a unified learning framework for joint VaR-ES forecasting. This framework assumes a universal function indexed by unknown parameters, mapping from predictors (e.g., firm characteristics) to two latent but unconstrained risk scores. Then, it transforms these two risk scores into valid VaR and ES forecasts through a fixed one-to-one mapping. This data-driven framework enables linear regression models, machine learning methods, and deep learning architectures to be treated in the same manner, with parameters estimated by minimizing the FZ loss.

Under this framework, we propose a new large tail risk model, the retrieval-enhanced self-grouping autoencoder (ReSGA), designed to fully exploit complex financial information for joint VaR-ES forecasting. ReSGA first learns latent group structures across assets, allowing tail-risk information to be shared among assets with similar financial behavior. It then incorporates a retrieval mechanism that enables each asset to draw on relevant long-term historical “experiences” from both its own past and other assets within the same group, thereby enriching usable temporal information. By leveraging cross-sectional dependence and long-term temporal dynamics in a flexible data-driven manner, ReSGA exploits spatial-temporal information more effectively than existing approaches.

Empirically, we study the monthly VaR and ES of United States (US) equity returns using a comprehensive dataset provided by jensen2023, covering more than 40,000 stocks from January 1926 to December 2023. Each stock-month observation is associated with 153 firm characteristics that are widely used in the asset pricing literature. We evaluate the out-of-sample VaR and ES forecasting performance of ReSGA and its 12 competing models through the average loss. From this analysis, we find that ReSGA consistently attains the lowest out-of-sample average loss across all stock universes considered. Besides the loss analysis, we further assess the statistical performance of all models via several statistical hypothesis tests, including the Diebold--Mariano (DM; diebold1995) test, the model confidence set (MCS; Hansen2011model) test, the conditional coverage (CC; Christoffersen1998EvaluatingIF) test, and the auxiliary expected shortfall regression (AESR; bayer2022) test. The out-of-sample testing results also provide consistent evidence in favor of ReSGA.

Beyond statistical assessments, we are interested in whether the tail risk forecasts generated by ReSGA can lead to economically meaningful gains. Following AtilganYigit2020, we construct a left-side momentum strategy by sorting stocks into deciles based on the predicted VaR or ES. The resulting VaR- and ES-sorted portfolios exhibit sharp nonlinearities concentrated in the extreme deciles: Stocks with the highest predicted tail risk (i.e., the lowest VaR or ES) deliver substantially lower returns and markedly worse downside risk, while a long-short strategy that buys the lowest-risk decile and sells the highest-risk decile, yields economically sizable Sharpe ratios. These patterns indicate that ReSGA’s tail risk forecasts capture cross-sectional variation in downside risk, which is also priced in the cross-section. Building on this insight, we propose a new size-enhanced left-side momentum signal that combines predicted ES with firm size. This signal is motivated by the “too big to fail” phenomenon in the US market: Tail risk in large firms is more likely to reflect compensated systematic risk, whereas similar tail risk in small firms often stems from idiosyncratic fragility. Based on this signal, we propose long-short portfolios for each considered model and find that all proposed non-econometric portfolios exhibit significant alphas relative to the Fama--French five-factor model. Hence, it indicates that our size-enhanced left-side momentum signal could broaden the mean-variance frontier. More importantly, our portfolio performances confirm that ReSGA delivers the best economic performance among all competing models, with the highest Sharpe ratio for the long-short decile portfolio and a clear monotonic return pattern across deciles. All of the above findings show that the forecasting advantage from ReSGA is not only statistically relevant but also economically meaningful.

Taken together, our statistical and economic findings above highlight that models can deliver better tail risk forecasts by exploiting richer information in the input. We view the informativeness of model input as one aspect of data complexity. This phenomenon connects to a broader debate on the virtue of complexity. In artificial intelligence, a growing body of evidence shows that the out-of-sample predictive performance of models often improves systematically with increases in either data complexity (with respect to in-sample data availability) or model complexity (with respect to model parameter size); see, for example, kaplan2020scalinglawsneurallanguage. This empirical regularity, known as the scaling law, underpins the success of large language models in the last five years. In finance, there has been an increasing focus on the investigation of model complexity, while the examination of data complexity remains very rare. Bryan2024virture provide early theoretical and empirical support for the virtue of model complexity in asset pricing, demonstrating that simple models with few parameters can severely understate return predictability relative to more complex alternatives. See more empirical evidence for this view in APTDidisheim2024 and Artificial2025Kelly. However, at the same time, other studies also question whether such gains from model complexity are economically meaningful or robust in asset pricing applications Berk2023CommentVirtueComplexity,Seemingly2025Stefan,Buncic2025Complexity,CarteaJinShi2025Complexity. In terms of risk forecasting, evidence on the role of model complexity is far more limited; li2025 show that machine learning models outperform linear models in realized volatility forecasting by capturing nonlinear temporal dynamics. However, to date, the role of model complexity and data complexity in tail risk forecasting remains largely unexplored.

To fill the gap, we further investigate the scaling performance of ReSGA and its competitors by varying model parameter size or in-sample data size to assess the role of model complexity or data complexity, respectively. First, our experiment results show that, holding the in-sample data fixed, simply increasing parameter size alone does not yield consistent performance improvements: Larger models do not reliably outperform smaller ones, and the best results in terms of out-of-sample average loss and portfolio Sharpe ratio typically occur at intermediate scales (around or below the in-sample sample size). In particular, models that exploit spatial–temporal inputs, such as ReSGA, benefit more from additional parameters and consistently outperform simpler architectures once sufficient parameters are available. Second, our complementary experiments vary the available in-sample data size while keeping the model architectures fixed. The corresponding results demonstrate that, in terms of out-of-sample average loss and portfolio Sharpe ratio, the performance of each model exhibits a clear improvement trend with the size of the in-sample data used. Together, the aforementioned scaling results indicate that our out-of-sample performance gains are primarily driven by richer data rather than by parameter growth alone. Hence, they provide limited support for a general virtue of model complexity, instead emphasizing the central role of data complexity in tail risk forecasting.

Lastly, we illustrate the usefulness of ReSGA through a group-importance analysis and a transfer-learning analysis. In the group-importance analysis, we explore economic model interpretability by investigating which types of firm characteristics drive tail risk predictability in the ReSGA model. It turns out that the predictive power for VaR and ES is concentrated in a small set of economic themes (or groups): Value, Low Risk, Momentum, Quality, and Short-Term Reversal. Additionally, the importance of these groups varies systematically across firm size and over time, reflecting shifts in the underlying economic sources of tail risk. In the transfer-learning analysis, we deploy the ReSGA model trained only on US equity data, without any re-estimation or fine-tuning, to five major international equity markets: China, Japan, the United Kingdom (UK), Australia and Canada. Despite substantial cross-market heterogeneity, ReSGA consistently presents reliable out-of-sample VaR-ES forecasting performance in every market, indicating that its predictive advantage is not US-specific and remains robust under market shifts. Hence, ReSGA, as a universal large tail risk model, has great generalizability in learning VaR and ES. At the portfolio level, however, return patterns differ substantially across countries. In China and Japan, higher size-enhanced left-side momentum is associated with higher returns, while in the UK, Australia, and Canada, large return dispersion concentrates in the extreme deciles. These findings suggest that tail-risk pricing varies across markets, likely reflecting differences in market structure, investor behavior, and risk premia.

The remaining paper proceeds as follows. (ref) introduces a unified VaR-ES learning framework and the ReSGA model. (ref) presents the empirical results for the US equity market, while (ref) examines the transferability of the ReSGA model trained on US data to other markets. (ref) concludes. Technical details are provided in Appendix (ref) and other appendices in the supplementary materials.

Methodology

Let $r_{i,t}$ denote the return of asset $i$ at time $t$, where $t = 1, \ldots, T$ and $i = 1, \ldots, N_t$, with $N_t$ representing the number of available assets at time $t$. The variables of interest in this paper are the conditional $\tau$-th VaR and ES of $r_{i,t}$, defined respectively as

align*[align* omitted — 184 chars of source]

where $\tau\in(0,1)$ is the quantile level, $F_{i,t}(\cdot)$ is the conditional distribution of $r_{i,t}$ given $\mathcal{F}_{t-1}$, and $\mathcal{F}_{t-1}$ is the information set containing all available information up to time $t-1$.

In the following subsections, (ref) introduces a unified learning framework for $\mathrm{VaR}_{i,t}$ and $\mathrm{ES}_{i,t}$. (ref) presents our new ReSGA model. (ref) describes a set of representative benchmark tail risk models that are used for comparison in our empirical study.

A Unified Learning Framework

Following the seminal work of koenker1978regression, a common way to learn VaR is based on the check loss function. However, as shown in gneiting2011, ES alone is not elicitable, meaning that it cannot be the unique minimizer of an expected loss function. fissler2016 solve this issue by showing that VaR and ES are jointly elicitable: $(\mathrm{VaR}, \mathrm{ES})$ is the unique minimizer of the general FZ loss function. Specifically, to jointly learn the VaR and ES of a univariate random variable $Y$, fissler2016 define the general FZ loss function (also termed as the scoring rule) as follows:

align[align omitted — 279 chars of source]

where $\tau\in(0, 1)$ is the quantile level, $\bm{1}\{\cdot\}$ is the indicator function, and $g_1,g_2:\mathbb{R} \to \mathbb{R}$ are two pre-determined functions, with $g_2$ being strictly increasing and convex, and $\tilde{g}_2$ being its primitive. In particular, when $g_1(x) = 0$ and $g_2(x) = -1/x$, the general FZ loss function in ((ref)) becomes the commonly used degree-0 FZ loss function:

align[align omitted — 133 chars of source]

where $(v, e)\in \Gamma$ for an admissible region $\Gamma=\{(v, e): e < v < 0\}$. Using the degree-0 FZ loss function, a joint learning of VaR and ES can be achieved by noting

align[align omitted — 191 chars of source]

As shown in (ref), the learning of $\mathrm{VaR}_{i,t}$ and $\mathrm{ES}_{i,t}$ under $\ell_{\mathrm{FZ0}}$ needs to meet a key admissible condition: $\mathrm{ES}_{i,t} < \mathrm{VaR}_{i,t} < 0$. To achieve this goal, we propose a new unified learning framework for $\mathrm{VaR}_{i,t}$ and $\mathrm{ES}_{i,t}$. Specifically, let $\bm{y}_{i,t} = (y_{i,t,1}, y_{i,t,2})'\in\mathbb{R}^2$ represent a vector of two latent risk scores $y_{i,t,1}$ and $y_{i,t,2}$. Then, we define a Softplus-based one-to-one mapping $p$: $\bm{y}_{i,t} \rightarrow (\mathrm{VaR}_{i,t}, \mathrm{ES}_{i,t})$, satisfying

align[align omitted — 232 chars of source]

where $\mathrm{Softplus}(\cdot)$ is a function Softplus defined as

align*[align* omitted — 118 chars of source]

Since both $\mathrm{Softplus}(y_{i,t,1})$ and $\mathrm{Softplus}(y_{i,t,2})$ are strictly positive, the mapping in (ref) guarantees that $\mathrm{ES}_{i,t} < \mathrm{VaR}_{i,t} < 0$, regardless of the values of $y_{i,t,1}$ and $y_{i,t,2}$. Consequently, we are able to learn $\mathrm{VaR}_{i,t}$ and $\mathrm{ES}_{i,t}$ through a model for $\bm{y}_{i,t}$, the form of which can be specified without imposing restrictions on $y_{i,t,1}$ and $y_{i,t,2}$.

Following the above idea, our learning framework only needs to specify a model for $\bm{y}_{i,t}$ indexed by unknown parameters $\bm{\theta}$. Therefore, unless otherwise stated, all tail risk models in this paper are designed for $\bm{y}_{i,t}$. By writing $\bm{y}_{i,t} \equiv \bm{y}_{i,t}(\bm{\theta})$, we have $(\mathrm{VaR}_{i,t}, \mathrm{ES}_{i,t})=p(\bm{y}_{i,t}(\bm{\theta}))$ according to ((ref)), so we can make use of result ((ref)) to estimate $\bm{\theta}$ by minimizing the empirical degree-0 FZ loss:

align[align omitted — 393 chars of source]

In practice, due to the massive data volume, we use the adaptive moment estimation (Adam) algorithm in Kingma2015AdamAM to solve optimization problem (ref). Given $\widehat{\bm{\theta}}$, $\mathrm{VaR}_{i,t}$ and $\mathrm{ES}_{i,t}$ are then learned by $\widehat{\mathrm{VaR}}_{i,t}$ and $\widehat{\mathrm{ES}}_{i,t}$, respectively, where $(\widehat{\mathrm{VaR}}_{i,t}, \widehat{\mathrm{ES}}_{i,t})=p(\bm{y}_{i,t}(\widehat{\bm{\theta}}))$.

The ReSGA Model

Let $\bm{Y}_{t} \in \mathbb{R}^{N_t \times 2}$ denote the matrix of cross-sectional risk scores at time $t$, where its $i$-th row is $\bm{y}_{i,t}'$. Under the learning framework in Section (ref), we design a new ReSGA model for $\bm{Y}_{t}$, which aims to capture the nonlinear relationships between asset characteristics and risk scores, while simultaneously accounting for spatial and temporal dependencies of assets. To facilitate the construction of ReSGA, we let $\bm{X}_{i,t-1} \in \mathbb{R}^{S \times P}$ be the feature matrix of asset $i$ up to time $t-1$, with its $s$-th row $\bm{x}_{i,t-1-S+s}'$, where $\bm{x}_{i,t} = (x_{i,t,1}, \ldots, x_{i,t,P})'\in\mathbb{R}^{P}$ contains $P$ characteristics of this asset at time $t$. Here, $P$ is the number of characteristics and $S$ is the number of time lags. Collecting all individual feature matrices, we obtain the tensor of asset characteristics $\mathcal{X}_{t-1} = [\bm{X}_{1,t-1}, \dots, \bm{X}_{N_t,t-1}] \in \mathbb{R}^{N_t \times S \times P}$.

Using $\mathcal{X}_{t-1}$ as the input, ReSGA, parameterized by $\bm{\theta}=(\bm{\phi}',\bm{\psi}')'$, learns $\bm{Y}_{t}$ through three core modules, encoder, retriever, and decoder, which are defined as follows:

align[align omitted — 324 chars of source]

where $\mathrm{Encoder}(\cdot; \bm{\phi})$, $\mathrm{Retriever}(\cdot; \mathbb{M}_t)$, and $\mathrm{Decoder}(\cdot, \cdot; \mathbb{M}_t, \bm{\psi})$ denote the functional forms of encoder, retriever, and decoder, respectively (see their detailed architectures in Sections (ref)--(ref)). Generally speaking,

itemize$\mathrm{Encoder}(\cdot; \bm{\phi})$, parameterized by $\bm{\phi}$, extracts a temporal feature $\mathcal{H}_t\in \mathbb{R}^{N_t \times S \times D}$ for all assets across lags in an autoregressive manner, along with a set $\mathbb{M}_t$ that captures latent group structures among assets according to their similarities in characteristics, where the dimension $D$ serves as a user-specific hyperparameter controlling model complexity; • $\mathrm{Retriever}(\cdot;\mathbb{M}_t)$ leverages the learned group structure $\mathbb{M}_t$ to retrieve long-term historical information not only from each asset’s own temporal history but also from other assets within the same group, yielding the retrieval feature $\bm{Z}_t \in \mathbb{R}^{N_t \times D}$; • $\mathrm{Decoder}(\cdot, \cdot; \mathbb{M}_t, \bm{\psi})$, parameterized by $\bm{\psi}$, hierarchically aggregates the group-, retrieval-, and asset-level information to generate the risk score $\bm{Y}_t$.

It is worth noting that $\mathbb{M}_t$ depends on the specific choice of $\mathcal{X}_{t-1}$ and loss function. In this study, under the degree-0 FZ loss in ((ref)), $\mathbb{M}_t$ reveals the dynamic topological relationships among assets for tail risk analysis, guided by the similarities in asset characteristics. When ReSGA is applied to other studies, different inputs and loss functions can lead to different interpretations for $\mathbb{M}_t$.

Competing Models

To assess the performance of ReSGA, we consider its 12 competing models. These competitors are classified into four categories: Point-wise, Temporal, Spatial–temporal and Econometric models, depending on whether they exploit temporal or cross-sectional information and whether they employ firm characteristics. Models in the first three categories are based on our learning framework in Section (ref), while models in the last category learn VaR and ES using existing econometric approaches. For more details about these 12 competing models, one can refer to Appendix (ref) in the supplementary materials.

\paragraph{Point-wise Models}

The point-wise models aim to capture the relationship between asset characteristics and tail risks, but ignore both temporal and cross-sectional dependencies. They directly map the most recent asset characteristics to risk scores as follows:

align[align omitted — 81 chars of source]

where $\bm{x}_{i,t-1} \in\mathbb{R}^{P}$ is the vector of $P$ characteristics for asset $i$ at time $t-1$, and $f(\cdot;\bm{\theta})$ is a parametric function indexed by $\bm{\theta}$. Particularly, we focus on two models in ((ref)): (i) the classical linear regression model, labeled as “Linear”; (ii) a three-hidden-layer neural network model proposed by gu2020empirical, labeled as “NN”.

\paragraph{Temporal Models}

Unlike point-wise models, the temporal models explicitly account for the dynamic evolution of asset characteristics over time by assuming

align[align omitted — 85 chars of source]

where $\bm{X}_{i, t - 1}\in\mathbb{R}^{S\times P}$ is the feature matrix of asset $i$ at time $t-1$, and $g(\cdot;\bm{\theta})$ is a parametric function indexed by $\bm{\theta}$. Note that the hyperparameter $S$ specifies the look-back time period of characteristics. When $S=1$, $\bm{X}_{i,t-1}'$ reduces to $\bm{x}_{i,t-1}$, which is the vector of the most recent characteristics (i.e., the input to point-wise models in ((ref))). Compared with point-wise models, the temporal models in ((ref)) exploit not only the predictive power of the latest characteristics, but also the historical evolution of characteristics, allowing them to capture dynamic patterns that may enhance predictive performance.

The temporal models require specific designs to effectively capture temporal dynamics. In this paper, we employ several representative temporal deep learning models in (ref), including:

enumerate• two extended models: a lag-augmented neural network model that flattens $\bm{X}_{i,t-1}$ into a vector as input (labeled as “LANN”), and a decomposition-based neural network model proposed by zeng2023transformers that separately builds trend and seasonal components (labeled as “DLinear”); • two recurrent neural network models: long short-term memory network model from hochreiter1997long (labeled as “LSTM”) and gated recurrent unit network model from GRU (labeled as “GRU”); • three transformer-type neural network models: the “Informer” model proposed by zhou2021informer, which leverages an attention-based encoder and decoder to capture temporal dependencies, along with its encoder-only and decoder-only variants (labeled as “EInformer” and “DInformer”, respectively).

\paragraph{Spatial-temporal Models}

For the spatial-temporal models, they simultaneously capture both the temporal evolution of asset characteristics and the cross-sectional interactions among assets using advanced deep learning architectures. Specifically, they are defined as

align[align omitted — 94 chars of source]

where $\mathcal{X}_{t-1} \in \mathbb{R}^{N_t \times S \times P}$ is the tensor of asset characteristics, and $m(\cdot;\bm{\theta})$ is a parametric function indexed by $\bm{\theta}$. Compared to the temporal models, which handle assets separately, the spatial–temporal models allow each asset’s risk scores to depend not only on its own historical information but also on that of other assets. By explicitly modeling these inter-asset dependencies, the spatial–temporal models are capable of capturing common shocks, contagion effects, and network spillovers, which are particularly relevant in financial markets. This spatial information sharing is expected to lead to improved robustness and predictive accuracy, especially during periods of heightened market co-movement or systemic stress.

Given the rich information contained in $\mathcal{X}_{t-1}$, the spatial–temporal models must be carefully designed to efficiently exploit both temporal and cross-sectional dependencies. Besides the proposed ReSGA model, we consider another spatial–temporal model in (ref): the self-grouping autoencoder model from GRAND (labeled as “SGA”). In short, the SGA model can be viewed as a simplified version of ReSGA. Both models share a common encoder–decoder architecture that learns time-varying group structures across assets. However, SGA removes all retrieval-related components and focuses solely on contemporaneous spatial dependencies. As a result, it could perform well when the temporal length $S$ is relatively short. In contrast, ReSGA extends this framework by incorporating a retrieval mechanism to capture long-term spatial–temporal interactions, making it suitable for longer sequences and more complex systemic risk patterns.

\paragraph{Econometric Models}

In the econometrics literature, patton2019 apply two widely used semi-parametric models, GAS and GARCH, to study the dynamics of VaR and ES. These models do not rely on machine learning or deep learning techniques to utilize the information of asset characteristics, yet they need to impose the constraint of $\mathrm{ES}_{i,t}<\mathrm{VaR}_{i,t}<0$ when implementing optimization in model estimation. Including them allows us to assess whether more complex models deliver meaningful improvements over well-established econometric approaches. Notably, unlike the aforementioned non-econometric models, which employ a unified learning framework with a single parameter set shared across assets, the GAS and GARCH models are estimated separately for each asset to learn VaR and ES directly.

(ref) summarizes the features of all considered tail risk models, such as nonlinearity from deep learning, temporal dynamics from lagged characteristics, and spatial information sharing within cross-sectional input. In addition, it also distinguishes models by whether they support long-history inputs (i.e., capturing long memory effect). By clarifying the incremental modeling capabilities across point-wise, temporal, spatial–temporal, and econometric models, this table clearly illustrates how ReSGA integrates all these features within a single model to exploit available information efficiently.

table[table omitted — 1,929 chars of source]

Empirical Study of US Equity

Data

This section analyzes a comprehensive US equity dataset compiled by jensen2023. This dataset covers more than 40,000 stocks, providing a broad and representative cross-section of the US equity market for large-scale risk forecasting. Meanwhile, it spans the period from January 1926 to December 2023 at a monthly frequency, offering nearly a century of observations across different market regimes, including economic expansions and major crisis episodes. In this dataset, each stock in every month is associated with 153 firm characteristics\footnote{See \url{https://jkpfactors.s3.amazonaws.com/documents/Documentation.pdf} for more details.}. As in gu2020empirical, we rank-normalize all characteristics cross-sectionally to mitigate the impact of outliers, and then impute missing characteristic values using contemporaneous cross-sectional medians. After this preprocessing treatment, we let $r_{i,t}$ denote the excess return of stock $i$ in month $t$, and construct the characteristics tensor $\mathcal{X}_{t-1}$ from all 153 characteristics.

Throughout the analysis, we apply the ReSGA model to learn $\mathrm{VaR}_{i,t}$ and $\mathrm{ES}_{i,t}$ at the quantile level $\tau=0.05$, with a comparison to the 12 competing models outlined in Section (ref). To evaluate out-of-sample forecasting performance for all models, we follow gu2020empirical to adopt an expanding-window process, with the last ten years of the sample (2014--2023) reserved for out-of-sample evaluation. Specifically, each model is trained on data from 1926 to 1995, validated on data from 1996 to 2013, and tested on data from 2014; this process then shifts forward by one year (trained on 1926--1996, validated on 1997--2014, and tested on 2015), and it is repeated until reaching the end of the out-of-sample period\footnote{As an exception, two econometric models, GAS and GARCH, have no need for validation, so they are estimated using all training and validation data.}. For the detailed implementation of the above training, validation, and testing process, one can refer to (ref).

Statistical Performance Evaluation

We first evaluate the out-of-sample forecasting performance of ReSGA and its 12 competitors from a statistical perspective. As in patton2019, we consider the following out-of-sample average loss as a natural assessment criterion:

align[align omitted — 264 chars of source]

where “oos” denotes the out-of-sample period, $\mathbb{S}_t$ denotes a group of stocks at time $t$, with $|\mathbb{S}_t|$ being its cardinality, $\ell_{\mathrm{FZ0}}$ is the degree-0 FZ loss function defined in (ref), and $\widehat{\mathrm{VaR}}_{i,t}$ and $\widehat{\mathrm{ES}}_{i,t}$ are forecasts of $\mathrm{VaR}_{i,t}$ and $\mathrm{ES}_{i,t}$ from each learned model.

table[table omitted — 2,466 chars of source]

(ref) reports the results of $\ell_{\text{oos}}$ across models and different stock groups. From this table, we can draw the following interesting findings:

itemize• ReSGA achieves the lowest loss across the full stock group, as well as within five size-based stock groups: mega, large, small, micro, and nano. This consistent dominance indicates that ReSGA provides a broadly effective engine for VaR and ES forecasting, with its combination of learned cross-sectional group structure and retrieval-based temporal augmentation. • Two spatial-temporal models tend to outperform seven temporal models, which in turn generally outperform two point-wise models. A notable exception is LANN, which performs worse than NN in most cases despite having richer inputs. By flattening the $S \times P$ characteristics matrix into a single vector, LANN discards temporal ordering and therefore fails to model sequential dependence directly. As a result, it fails to capture dynamic patterns such as persistence or volatility clustering, illustrating that exploiting temporal information requires appropriate model architecture. Nevertheless, the overall ranking among all eleven universal models provides strong empirical support for the virtue of data complexity: Improvements from point-wise to temporal models reflect the value of incorporating temporal information, while further gains from temporal to spatial-temporal models highlight the importance of cross-sectional information. • All eleven universal models substantially outperform two econometric models: GAS and GARCH. A natural explanation is that GAS and GARCH are primarily designed to capture return dynamics and largely overlook rich firm-level information, which is crucial for VaR and ES forecasting. Note that GAS and GARCH are estimated separately for each asset, whereas all universal models are trained on the pooled samples across assets. Hence, the above advantage of universal models over econometric models also points to a clear data scaling effect, highlighting the advantage of larger effective sample sizes and richer information content. • Within the universal model family, moving beyond linear specifications yields clear gains, since nonlinear models (e.g., NN, LSTM, SGA, and ReSGA) deliver lower losses than the linear model in most cases. This finding suggests that nonlinearities and interaction effects among firm characteristics play an important role in tail risk forecasting. • In terms of size-based stock groups, a clear monotonic pattern emerges: Larger firms systematically exhibit lower losses. This pattern is expected, as larger firms have greater liquidity, more standard published financial statements, and more stable trading environments, all of which contribute to their better predictability of tail risk.

In addition to the out-of-sample average loss, we further evaluate the out-of-sample performance of all models using four statistical hypothesis tests: the DM in diebold1995, MCS in Hansen2011model, CC in Christoffersen1998EvaluatingIF, and AESR in bayer2022. See their detailed descriptions in (ref) of the supplementary materials. In terms of the values of $\ell_{\text{oos}}$ with $\mathbb{S}_t$ being the full stock group, the DM test aims to compare the prediction performance of any two models, and the MCS test searches for a batch of models that perform significantly better than the others. The CC and AESR tests provide statistical evidence of the validity of VaR and ES forecasts, respectively. Notably, we do not adopt the e-backtesting approach of wang2025backtest to examine the validity of ES forecasts, since that framework is mainly designed for sequential monitoring and non-fixed-sample sizes, whereas our empirical exercise is based on fixed-sample out-of-sample evaluation.

(ref) reports the pairwise DM test statistics for all considered models. From this table, we have two clear findings. First, except for SGA, ReSGA dominates all other models with significantly lower values of $\ell_{\text{oos}}$. Second, all universal models deliver statistically significant improvements over econometric models.

Next, starting from the full set of models, our MCS test iteratively eliminates inferior models at the 90% confidence level, until only those statistically indistinguishable models remain to form the final model set. It turns out that only SGA and ReSGA are retained in the final model set. This outcome indicates that no other forecasting models perform as well as SGA and ReSGA in a statistical sense, underscoring the importance of spatial-temporal modeling for tail risk forecasting.

Moreover, we apply the CC and AESR tests to check the validity of VaR and ES forecasts for each stock at the significance level $\alpha \in \{0.01, 0.05, 0.10\}$, respectively, and then denote the proportion of stocks having valid VaR and ES forecasts as the pass rate. (ref) reports pass rates with respect to VaR and ES forecasts across all models. From this table, we find that, except for ES at the level $\alpha=0.01$, the two spatial-temporal models, SGA and ReSGA have the largest values of pass rates. In particular, ReSGA shows a substantially larger pass rate than SGA for the prediction of ES at the level $\alpha=0.05$ or $0.10$.

Overall, the above statistical performance evaluations consistently demonstrate the state-of-the-art performance of ReSGA in tail risk forecasting, which in turn implies the virtue of data complexity.

sidewaystable[p] \caption{Pairwise DM test statistics.} \begin{threeparttable} {2pt} \begin{tabular}{lccccccccccccc} \toprule & GAS & GARCH & Linear & NN & LANN & DLinear & LSTM & GRU & Informer & EInformer & DInformer & SGA & ReSGA \\ \midrule GAS & -- & & & & & & & & & & & & \\ GARCH & 1.04 & -- & & & & & & & & & & & \\ Linear & -3.35$^{***}$ & -2.39$^{**}$ & -- & & & & & & & & & & \\ NN & -3.39$^{***}$ & -2.41$^{**}$ & -0.85 & -- & & & & & & & & & \\ LANN & -3.30$^{***}$ & -2.34$^{**}$ & 1.02 & 8.37$^{***}$ & -- & & & & & & & & \\ DLinear & -3.42$^{***}$ & -2.39$^{**}$ & -1.73$^{*}$ & -0.11 & -2.86$^{***}$ & -- & & & & & & & \\ LSTM & -3.43$^{***}$ & -2.40$^{**}$ & -2.21$^{**}$ & -1.32 & -5.24$^{***}$ & -2.11$^{**}$ & -- & & & & & & \\ GRU & -3.45$^{***}$ & -2.41$^{**}$ & -2.46$^{**}$ & -1.80$^{*}$ & -5.65$^{***}$ & -2.89$^{***}$ & -2.23$^{**}$ & -- & & & & & \\ Informer & -3.44$^{***}$ & -2.41$^{**}$ & -3.80$^{***}$ & -1.62 & -4.46$^{***}$ & -3.13$^{***}$ & -1.31 & -0.56 & -- & & & & \\ EInformer & -3.36$^{***}$ & -2.37$^{**}$ & -0.23 & 0.42 & -0.81 & 0.76 & 1.22 & 1.45 & 2.27$^{**}$ & -- & & & \\ DInformer & -3.41$^{***}$ & -2.39$^{**}$ & -2.66$^{***}$ & -0.21 & -1.83$^{*}$ & -0.32 & 0.58 & 0.91 & 1.62 & -1.39 & -- & & \\ SGA & -3.51$^{***}$ & -2.43$^{**}$ & -4.62$^{***}$ & -3.56$^{***}$ & -8.01$^{***}$ & -4.40$^{***}$ & -3.05$^{***}$ & -2.17$^{**}$ & -1.40 & -2.25$^{**}$ & -2.13$^{**}$ & -- & \\ ReSGA & -3.51$^{***}$ & -2.44$^{**}$ & -4.69$^{***}$ & -4.49$^{***}$ & -7.63$^{***}$ & -4.82$^{***}$ & -4.13$^{***}$ & -3.26$^{***}$ & -2.33$^{**}$ & -2.60$^{***}$ & -2.51$^{**}$ & -1.12 & -- \\ \bottomrule \end{tabular} Notes. A lower-triangular matrix that reports the pairwise DM test statistic. Each entry shows the value of DM statistic by comparing the model in the row against the one in the column. Its negative (positive) value indicates that the row model yields lower (higher) loss, where the symbols ${***}$, ${**}$, and ${*}$ denote the rejection of the null hypothesis of equal prediction performance at the 1%, 5%, and 10% levels, respectively. \end{threeparttable}
table[table omitted — 2,029 chars of source]

Economic Performance Evaluation

To explore the economic implications of the VaR and ES, it is important to evaluate all tail risk models from an economic perspective. Following AtilganYigit2020, we first assess economic gains using decile portfolios sorted on the predicted VaR or ES. This trading strategy is referred to as “left-side momentum”. It guides us to monthly rebalance ten decile portfolios (P1--P10), after sorting all stocks in descending order according to their predicted VaR or ES at the end of each month. Note that our predicted VaR and ES are negative in this paper, so P1 corresponds to the lowest-risk portfolio and P10 represents the highest-risk portfolio.

To keep the presentation concise, (ref) reports portfolio results only for ReSGA, as the performance for other models is qualitatively similar. As shown in this table, under VaR-based sorting, average portfolio returns remain relatively stable up to P8 ($1.090\%$), but drop sharply to $0.619\%$ in P9 and turn negative in P10 ($-0.284\%$), while the Sharpe ratio falls to $-0.085$ in P10. This decline in returns is accompanied by a marked increase in downside risk: Maximum drawdown worsens from $0.499$ in P8 to $0.730$ in P9 and $0.807$ in P10. A similar pattern appears under ES-based sorting, with average returns decreasing from $1.196\%$ in P8 to $0.512\%$ in P9 and $-0.024\%$ in P10, and maximum drawdown deepening to $0.710$ in P9 and $0.810$ in P10.

The above findings indicate P9 and P10, which include stocks with high predicted tail risk, exhibit a substantially worse performance in terms of return, Sharpe ratio, and drawdown. AtilganYigit2020 attribute this pattern to behavioral underreaction to bad news and the persistence of left-tail risks, which concentrate losses in the most risk-exposed portfolios. Our findings mirror this stylized fact under both VaR- and ES-based sorting, supporting the view that tail-risk forecasts from the ReSGA can capture economically meaningful variation in downside risk across the cross-section.

table[table omitted — 1,736 chars of source]

Beyond the above left-side momentum, we further design a new size-enhanced left-side momentum strategy, which utilizes the trading signal motivated by the “too big to fail” phenomenon in the US stock market. Our key trading idea is to exploit the different economic implications of tail risk across firms of different sizes. Specifically, we assign strong buying signal to large-cap stocks with small values of ES, reflecting that downside risk in large firms often attracts market attention, policy support, or investor demand, and it thus can limit extreme losses and help preserve long-term value. In contrast, we assign strong selling signal to small-cap stocks with significantly negative ES, which often reflects idiosyncratic fragility, limited liquidity, or elevated default risk, rather than priced systematic risk.

Motivated by this idea, we construct the following trading signal:

align[align omitted — 176 chars of source]

where $\mathrm{Cap}_{i,t-1}$ denotes the log market capitalization of asset $i$ at time $t-1$, and $\overline{\mathrm{Cap}}_{t-1}$ is the corresponding cross-sectional mean. Here, the term $[1 - \exp(\widehat{\mathrm{ES}}_{i,t})]$ lies between $0$ and $1$, with its value decreasing as $\widehat{\mathrm{ES}}_{i,t}$ increases, serving as a proxy for tail-loss severity; the term $(\mathrm{Cap}_{i,t-1} - \overline{\mathrm{Cap}}_{t-1})$ tends to be positive for large-cap stocks and negative for small-cap stocks, acting as a classifier for firm size. Combining these two terms, the signal $\alpha_{i,t}$ in ((ref)) assigns high values to large-cap stocks with severe predicted tail risk, and low values to small-cap stocks with similar tail risk.

To assess the economic performance of our size-enhanced left-side momentum strategy, we adopt the signal $\alpha_{i,t}$ to propose monthly rebalanced decile portfolios as before, with $\widehat{\mathrm{ES}}_{i,t}$ computed from different models. Moreover, we regress the monthly returns of the resulting H--L portfolios on the Fama–French five factors (fama2015five) to examine whether there is an abnormal return (alpha) unexplained by standard factors. (ref) reports the results of this analysis, and it delivers the following findings:

itemize• The ReSGA-based strategy delivers the highest value of monthly average return, annualized Sharpe ratio and alpha for the H--L portfolio, outperforming all competing strategies. This result indicates that the best ES forecasting accuracy from ReSGA can translate into the strongest economic gains when combined with firm size information. • Regardless of the choice of models, decile portfolios exhibit a clear and robust monotonic pattern: Moving from P1 to P10, average portfolio returns and annualized Sharpe ratios decline steadily and become strongly negative in the lowest deciles (P9 and P10). This pattern confirms that the signal $\alpha_{i,t}$ can effectively rank stocks that have economically meaningful size-enhanced tail risk exposure. • For all strategies based on the universal models, the $p$-values of the estimated alphas are below $0.05$, whereas the corresponding $p$-values for the GAS- and GARCH-based strategies are $0.051$ and $0.128$, respectively. This contrast indicates that the econometric models exhibit relatively weak predictive power, highlighting the advantage of our proposed unified learning framework.

Taken together, (ref) show that the superior advantage of ReSGA in learning tail risks can translate into economically meaningful gains.

table[table omitted — 4,565 chars of source]

Scaling Performance Evaluation

To examine whether there is virtue of model complexity in learning VaR and ES, we begin by evaluating scaling performance with respect to parameter size for all considered machine-learning-based models (i.e., all universal models excluding Linear). Specifically, we adopt a within-architecture scaling strategy: For each model, we vary the number of model parameters by adjusting width/depth-related hyperparameters, while keeping all other mechanisms unchanged\footnote{For example, in the NN model, parameter size is controlled by the hidden-state dimension and the number of hidden layers. See (ref) for more details.}. This approach produces multiple variants of each model with substantially different numbers of learnable parameters. The connection between hyperparameter choices and parameter size for each model is reported in (ref) of the supplementary materials.

(ref) details the out-of-sample average loss of different models when their parameter size varies over several orders of magnitude. It includes exact loss values, while its accompanying figure visualizes the corresponding scaling patterns\footnote{For each model, we fit a polynomial function linking loss to log parameter size, with the polynomial order selected by the Akaike information criterion AIC.}. Combining both pieces of evidence, we obtain three main interesting findings:

itemize• ReSGA consistently outperforms all competitors once the parameter size reaches $10^5$ or above, with SGA ranking second, whereas Informer achieves the lowest loss at smaller scales. This pattern is evident from the bold entries in the table and the crossing behavior in the figure: ReSGA and SGA improve sharply from small to intermediate scales, while Informer deteriorates once its parameter size becomes large. These findings indicate that models with richer spatial-temporal inputs, such as SGA and ReSGA, need sufficient capacity in parameters to exploit useful information from inputs, while temporal models can perform well with relatively few parameters. In contrast, the point-wise model of NN benefits little from expanded capacity, likely because its input information is inherently limited. • There is no systematic evidence for a general virtue of model complexity. Except for LSTM, the parameter-richest specification (around $10^7$ parameters) does not deliver the best performance within each model family. The figure makes this non-monotonicity especially visible: Most loss curves flatten or turn upward after intermediate scales. Even for LSTM, where loss performance improves with scale, the gains are modest. Meanwhile, LSTM with around $10^7$ parameters still underperforms ReSGA with $10^5$ parameters. This absence of monotonic and substantial improvement suggests that simply increasing model size alone does not guarantee better tail-risk forecasts. • As shown by kaplan2020scalinglawsneurallanguage, in natural language processing, optimal performance is often achieved when model size is of the same order as the in-sample data size. This principle is only partially supported in our financial task. To be specific, ReSGA and SGA achieve their lowest out-of-sample losses at around $10^6$ parameters, which is comparable to the effective training sample size (approximately $2\times 10^6$); however, this alignment is not universal, since other models show no clear match between optimal parameter size and sample size. These observations point to the heterogeneity in the virtue of data complexity: Spatial-temporal models could exploit richer inputs, allowing for more capacity in parameters to be used effectively, whereas temporal and point-wise models rely on less information and consequently obtain fewer benefits from increased parameterization.
table[table omitted — 1,975 chars of source]

Meanwhile, (ref) reports the out-of-sample Sharpe ratios of H--L portfolios as parameter size varies, with an accompanying figure to visualize the trend of Sharpe ratio across parameter scales. From (ref), we find that Sharpe ratio fails to exhibit a clear monotonic pattern as the parameter size increases. First, the best-performing parameter size differs across models: Most models peak at intermediate scales, while LSTM and Informer attain their highest Sharpe ratios at the largest scale. Second, the leading model also changes across parameter scales. These mixed patterns indicate that increasing model complexity alone does not reliably improve portfolio-level gains.

table[table omitted — 1,754 chars of source]

To further examine the virtue of data complexity in learning VaR and ES, we next evaluate scaling performance with respect to in-sample data size. Specifically, we hold each model fixed and vary the fraction of available in-sample data used for training and validation, from $1\%$ to $100\%$ of the full in-sample dataset. For each fraction, we draw the in-sample data uniformly at random from the full in-sample dataset, and keep the hyperparameter configuration fixed at the values selected using the full in-sample dataset. This design holds model complexity fixed, and it isolates the role of in-sample data size that is one key dimension of data complexity.

(ref) reports the out-of-sample average loss across different in-sample data sizes, while its accompanying figure visualizes how these losses evolve as more in-sample data become available. From this table and its accompanying figure, we can reach the following findings:

itemize• All models benefit from more in-sample data (from $1\%$ to $100\%$). With the exception of EInformer, each model attains its best performance when the full in-sample dataset is available. Meanwhile, the loss curves in the figure mostly present monotonic pattern and decline as the sample fraction increases. • The best-performing model changes with the amount of in-sample data. When in-sample data are scarce (only $1\%$, $5\%$, or $10\%$), temporal models such as GRU and Informer outperform the other models. Once the fraction of used in-sample data reaches $25\%$ or above, the spatial-temporal models (SGA and especially ReSGA) become dominant, indicating that these models require sufficient data to fully exploit cross-sectional sharing and long-term temporal dependencies. • Point-wise model, NN, exhibits the largest improvement as the amount of in-sample data grows, consistent with its reliance on pooled cross-sectional information to compensate for weaker inductive structure. Nevertheless, even with the full in-sample dataset, NN remains clearly outperformed by the best temporal and spatial-temporal models. Thus, more data help broadly, with the greatest statistical gains when the model can use the richest in-sample information effectively.
table[table omitted — 1,782 chars of source]

Beyond evaluation from loss, (ref) reports the out-of-sample Sharpe ratios of H--L portfolios when the in-sample data size varies, while its accompanying figure visualizes the scaling patterns of Sharpe ratio. Being broadly consistent with the loss results in (ref), this table delivers the following notable results from the economic perspective:

itemize• ReSGA delivers the highest Sharpe ratio across different models once the in-sample data size reaches at least $25\%$ of the full dataset. • The fitted curves show that portfolio performance generally improves as more in-sample data become available, although the pattern is not strictly monotonic for every model.

Taken together, the above results show that a larger in-sample dataset can lead to stronger economic performance, especially for models that can effectively exploit rich information.

table[table omitted — 1,665 chars of source]

Overall, our analysis results provide little support for the virtue of model complexity. Instead, they highlight the importance of data complexity: Performance gains arise mainly from models that can exploit richer information with the use of a larger in-sample data size, rather than a larger parameter size alone. While the available sample size is always limited in finance, this motivates the next scaling dimension: expanding the effective information content of the input with appropriate model specification and enough parameters, which appears to be a more effective path to improving tail risk forecasting accuracy than simply increasing model size.

Group Importance

A further empirical question concerns which firm characteristics are most important for forecasting tail risk. Due to the superior performance of ReSGA from both statistical and economic perspectives, we utilize this model to tackle the above question. Assessing importance at the level of individual characteristics, however, is challenging due to the well-known dilution problem: When many characteristics are highly correlated, their marginal contributions become spread across related variables, obscuring the true economic drivers of predictive performance li2025. To address this issue, we assess importance at the group level rather than at the level of individual characteristics. Specifically, we adopt the 13-category taxonomy of jensen2023, which organizes the 153 firm characteristics into economically interpretable groups, including Low Risk, Value, Quality, Low Leverage, Momentum, Size, Profit Growth, Short-Term Reversal, Seasonality, Investment, Profitability, Debt Issuance, and Accruals. As argued by li2025, grouping characteristics reduces redundancy among highly correlated signals and provides a clearer decomposition of predictive content. This approach allows us to evaluate the incremental value of entire economic themes, rather than attributing importance to individual characteristics that may proxy for similar information.

We quantify group importance via a drop-group procedure. To be specific, for each group (denoted as $\mathbb{G}$), we follow gu2020empirical to construct perturbed testing samples in which all characteristics in group $\mathbb{G}$ are set to zero, while all remaining characteristics are left unchanged. By passing these perturbed samples through the trained model to generate new forecasts, we recompute the out-of-sample average loss in (ref) on these new forecasts, with the set $\mathbb{S}_t$ containing all examined stocks. Then, we measure the importance of group $\mathbb{G}$ by the increase in loss of perturbed samples relative to that obtained using the original testing samples. After computing importance for all 13 groups, we set negative values to zero to avoid spurious attribution and normalize the remaining values to sum to one.

figure[figure omitted — 231 chars of source]

(ref) exhibits the ten most important groups from ReSGA over the out-of-sample period (2014.01--2023.12). From this figure, two clear and economically intuitive patterns emerge.

itemize• The top five groups (Value, Low Risk, Momentum, Quality, and Short-Term Reversal) together account for more than 90% of importance for the full dataset. This concentration indicates that tail risk predictability is driven by a small set of groups rather than being evenly distributed across all characteristics. Similar findings can be found in the asset pricing literature (gu2020empirical, Gu2021AutoencoderAP; yang2024asset). Note that the concentration is even stronger for mega- and large-cap stocks, with top five groups receiving over $99\%$ importance, whereas the top five groups in small-, micro-, and nano-stocks contribute $91\%$, $89\%$, and $94\%$, respectively. • The ranking of the leading groups shows both stability and systematic variation across different levels of market capitalization. Generally speaking, Low Risk and Momentum are consistently important regardless of the level of market capitalization. This is intuitive, as volatility- and beta-related characteristics in Low Risk are closely linked to tail risk, while Momentum captures persistent return dynamics that affect tail outcomes. Apart from these two common drivers, for mega- and large-cap stocks, Quality and Low Leverage are more influential, suggesting that tail risk variation among large firms is closely tied to balance-sheet strength and financial stability; in contrast, for small-, micro-, and nano-cap stocks, Value and Short-Term Reversal become more crucial, where valuation performance and short-horizon price reversals are expected to closely relate to distress risk, illiquidity, and left-tail outcomes of those stocks.

Moreover, (ref) traces the month-by-month evolution of group importance for ReSGA, where we compute group importance using the increase of loss from each month. Its heatmap reinforces that Low Risk remains important almost throughout the entire out-of-sample period, consistent with its role documented from (ref). As one may expect, the importance of Low Risk was especially pronounced from April to September 2020, since the COVID-19 shock had severely disrupted global financial markets. In contrast, Value and Quality display a visible “complementary” pattern over time, where the elevated period of Quality importance often coincides with muted Value, and vice versa. This finding suggests that the dominant predictors of tail risk shift between valuation-driven repricing forces and balance-sheet resilience/financial strength. Additionally, Quality and Low Leverage exhibit similar dynamics in certain months, a similarity that is expected due to their overlapping economic content related to profitability, safety, and leverage constraints.

figure[figure omitted — 243 chars of source]

Transfer Learning

Fundamentally, our universal models are driven by a principle that tail risk is governed by a mapping structure from asset characteristics to risk scores. As long as relevant asset characteristics are available, the model, by construction, can be applied to different assets without re-estimation. This contrasts sharply with econometric models, such as GARCH and GAS, which are inherently asset-specific. Hence, it raises a natural question: Can a universal model, trained on the largest and most information-rich US equity market, generalize to other equity markets without re-estimation?

To address this question, we conduct a transfer learning exercise. To be specific, we train each universal model exclusively on US equity dataset as in Section (ref) and then apply it, without any re-training or fine-tuning, to five major international equity markets: China, Japan, UK, Australia, and Canada. This exercise allows us to assess whether the universal models capture fundamental relationships between firm characteristics and future tail risk across different institutional, regulatory, and liquidity environments.

(ref) reports the out-of-sample average loss across universal models and international equity markets. From this table, we have the following noteworthy findings:

itemize• Despite substantial cross-market heterogeneity, ReSGA delivers the lowest out-of-sample loss in Japan, UK, and Australia, and remains highly competitive in China (ranked fourth) and Canada (ranked second). This strong and stable performance indicates that the predictive structure learned by ReSGA from US data generalizes well across international equity markets, even without re-estimation. • In contrast, SGA performs poorly in most non-US markets, except for Canada. In particular, it is often dominated by simpler temporal models. This pattern suggests that cross-sectional dependence learned in the source market does not necessarily transfer well to other markets, potentially diluting asset-level tail-risk signals and introducing noise. The relatively strong performance of SGA in Canada is plausibly attributable to the close economic integration and high similarity between the US and Canadian equity markets, which makes cross-sectional structures learned from US data more transferable in this case. Note that ReSGA mitigates the above limitation of SGA by augmenting cross-sectional grouping with a retrieval mechanism that extends the effective temporal memory. The ability to exploit long historical information appears to compensate for potential misspecification in the cross-sectional structure, resulting in more robust transfer performance across markets. • Temporal models, such as GRU, LSTM, and DLinear, also exhibit relatively stable out-of-sample performance across markets and often rank among the top predictors. This robustness further supports the view that temporal dynamics constitute a more transferable source of predictive power than cross-sectional dependence when models are deployed across markets without re-estimation.
table[table omitted — 1,741 chars of source]

Besides the loss analysis, (ref) further displays the models in the final model set from the MCS test for the transfer learning across markets. For reference, we also include the earlier results for the US market, where only SGA and ReSGA remain in the final model set. The cross-market evidence from (ref) delivers a clear message: ReSGA is the only model selected in all final model sets. This result indicates that the advantage of ReSGA is not confined to the US market, but generalizes robustly across international equity universes. Moreover, the CC and AESR testing results in (ref) provide further evidence of the effectiveness of ReSGA in transfer learning: Across the five target markets, ReSGA delivers consistently strong pass rates for both VaR and ES forecasts and ranks among the top performers in nearly all cases.

table[table omitted — 1,675 chars of source]
table[table omitted — 2,220 chars of source]

Lastly, we examine the economic gains of ReSGA in the transfer learning exercise. (ref) reports cross-market portfolio performance based on the ES-driven trading signal defined in (ref). Its results show that portfolio return patterns differ substantially between US and non-US markets. In China and Japan, the decile portfolios exhibit a reversed pattern, with higher predicted signal associated with higher returns, leading to negative H--L portfolio performance. In the UK, Australia, and Canada, the results are more heterogeneous. Although monotonicity is weaker, the extreme deciles display substantial return dispersion, especially in the lowest-signal portfolios, which drive large negative H--L returns. These findings suggest that the pricing of tail risk varies across markets, likely reflecting differences in market structure, investor behavior, and risk premia.

Taken together, the transfer learning results reinforce our central conclusions: ReSGA functions as a general engine for tail risk forecasting.

table[table omitted — 2,388 chars of source]

Conclusion

This paper studies tail risk forecasting, with a particular focus on VaR and ES, in the era of big data. Methodologically, we develop a unified joint VaR-ES learning framework supervised by the FZ loss. Based on this framework, we propose a large tail risk model, ReSGA, which captures the nonlinear relationships between asset characteristics and VaR-ES by accounting for spatial and temporal dependencies among assets. Empirically, using nearly a century of US equity data covering over 40,000 stocks and 153 firm characteristics, we find that ReSGA consistently delivers the best out-of-sample forecasting performance among a broad set of competing models. These forecasting gains are economically meaningful: Trading strategies based on ReSGA forecasts exhibit pronounced left-tail momentum and deliver superior long-short decile portfolio performances using a newly proposed size-enhanced left-side momentum signal.

Moreover, we provide systematic evidence on the scaling behavior of ReSGA and other machine-learning-based universal models. Our results offer little support for a general virtue of model complexity measured by parameter size alone. Instead, they favor the virtue of data complexity in learning tail risk, as improvements in forecasting VaR and ES can be achieved through larger in-sample datasets or more informative model inputs. In other words, the success of ReSGA is not only due to its access to rich available information but also to its ability to effectively extract, share, and utilize this information. Finally, we conduct group-importance analysis and transfer-learning experiments to illustrate the interpretability and cross-market generalizability of ReSGA.