Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,896 characters · 17 sections · 39 citation commands
Expected Shortfall LASSO
\affil{University of Amsterdam}
Expected shortfall (ES) has quickly become the standard regulatory measure for market risk, as the Basel III framework of the Basel Committee on Banking Supervision has required banks to base their internal risk models on it since 2016 basel2016. Defined as the expectation of the returns that exceed Value-at-Risk (VaR), another risk measure which corresponds to a tail quantile, the global regulatory body considers ES to provide a more prudent capture of tail risk compared to VaR, partly due to theoretical considerations outlined in artzner1999coherent among others. Accurate forecasting of ES is therefore of great importance to financial institutions.
Models of ES are often parsimonious in nature due to the low signal-to-noise ratio in financial returns, which complicates estimation of large models. The call for parsimony is often amplified by the low frequency of return observations close to or in exceedance of the VaR, such that the researcher has few observations in their data set that are highly informative in the estimation of ES models. And yet, we know that the risk in returns is connected to various fundamental variables (see Chapter 4 in andersen2013financial and references therein). Moreover, dependence on these fundamental variables may be nonlinear, which, in turn, increase the amount of parameters in the model. For instance, the leverage effect states there is nonlinear relationship between stock market volatility and future returns, see, e.g. CAMPBELL1992. The development of estimation methods that can handle large ES models is therefore important.
In this paper we develop a semiparametric estimator of linear ES models with many regressors, with the express purpose of estimating models that may rely on many (transformations of) fundamental variables. The estimator is an $\ell_1$-penalized least squares estimator (LASSO) that utilizes pre-estimated VaR predictions and leverages a sparsity condition in the model. With the VaR predictions, an auxiliary dependent variable is created that may be regressed on the set of explanatory variables to obtain an ES estimator, as shown by barendse2020. We derive a nonasymptotic bound on the prediction and estimation errors, and provide conditions under which consistency is obtained. These conditions allow the amount of regressors $p$ to grow with and potentially be much larger than the sample size $T$. Although theoretical properties of LASSO estimators for time-series data have been developed in the literature, our estimator and results are novel because we explicitly account for the prediction error in the conditional variable that appears due to its dependence on the pre-estimated quantile predictions.
We use a penalty function that accounts for differences in the scale of the regressors which has not been studied in the literature before for least squares estimators, to our knowledge, and which is borrowed from the penalized quantile regression literature, see belloni2011. Although accounting for scale differences is an appealing property on its own, we also prefer this penalty function as it means the penalization function of the coefficients in the VaR and ES estimation steps is equivalent. Indeed, we rely on the quantile regression estimator of belloni2023high (from hereon BCMPW) to obtain the VaR predictions in the first step, although our results apply to any quantile estimator that satisfies the conditions.
As financial data is often dependent over time and heavy-tailed, we derive our results under conditions that allow for these properties. Specifically, we allow the data to be $\beta$-mixing sequences with finite moments of a certain order. To obtain our results, we develop a Fuk-Nagaev inequality for $\beta$-mixing sequences, which may be of independent interest. By using the Fuk-Nagaev inequality in conjuction with the same blocking strategy as utilized in BCMPW to obtain approximately independent blocks of data, we obtain rates on the model parameters that are directly comparable to those for the penalized quantile regression estimator in BCMPW. This result is closely related to the Fuk-Nagaev inequality for $\tau$-mixing sequences in babii2022machine, which relies on an alternative blocking strategy, and the tail inequality for sequences that satisfy a sub-Weibull property in Wong2020.
We apply our method in a systemic risk analysis. In this analysis we generate predictions of the Conditional Expected Shortfall (CoES) to measure the risk spillover from the financial sector to the entire stock market. The CoES measure was introduced in adriancovar as an extension of CoVaR, an alternative systemic risk measure which is commonly used but is less prudent than CoES from a similar argument to the one that favors ES over VaR. The estimation of CoES relies on three stages of quantile regression and one stage of ES regression. Like adriancovar we condition on seven lagged fundamental variables, but we extend their set of fundamental variables by also considering nonlinear transformations. Specifically, we use each the Chebyshev polynomials to transform each of the fundamental variables, where the degree of polynomials used may increase with the sample size. Our results show that the penalized VaR and ES estimators outperform the unpenalized benchmark estimator that strictly uses the untransformed set of fundamental variables by a considerable margin in terms of out-of-sample prediction error. We observe that a moderate degree of Chebyshev polynomials is optimal in our sample and generates CoES measurements that are more conservative than the benchmark and more responsive to news. Finally, we observe that the inclusion of nonlinear transformations of the fundamental variables results in a kind of leverage effect: risk predictions are less impacted by positive returns.
In related literature, there are several papers that study the properties of LASSO estimators for linear mean models in a time-series setting. These papers consider non-estimated conditional variables, and therefore do not apply to the ES estimation method proposed in this paper. Closely related are babii2022machine, who study the properties of the group-LASSO estimator for $\tau$-mixing data, and Wong2020, who consider the LASSO estimator for $\beta$-mixing sequences that satisfy a sub-Weibull property, which is stronger than the moment conditions we impose. medeiros20161 consider the adaptive LASSO estimator for models with martingale difference sequence errors and adamek2023lasso extend their results to the near-epoch dependence error case. masini2022regularized study the case with mixingales. wu2016, uematsu2019high, and chernozhukov2021lasso consider LASSO estimators for Bernoulli shift data. Finally, nardi2011autoregressive, kock2015oracle, and basu2015 develop LASSO estimators under more restrictive conditions on the data. hsu2008subset and wang2007regression develop LASSO estimators for scenarios in which the number of variables is smaller than the sample size. In contemporaneous research, xuming2023 also develop an estimator of high-dimensional ES models based on the least-squares procedure in barendse2020. Instead of a penalized estimator, they consider a Huberized robust estimator and develop their results under independence and sub-Gaussianity of the data.
The rest of the paper is organized as follows. Section (ref) develops the estimator and derives a nonasymptotic bound and and consistency results. Section (ref) contains the empirical study on CoES estimation. Section (ref) includes a Monte Carlo simulation study. Section (ref) develops the Fuk-Nagaev inequality. The Appendix includes proofs and additional results for the empirical analysis and Monte Carlo study.
Notation: Define, for $p\in \mathbb{N}$, $[p] = \{1,\ldots,p\}$. For $b \in \mathbb{R}^p$, we denote the $\ell_q$-norm as $\|b\|_q = (\sum_{i\in [p]} |b_i|^q)^{1/q}$, for $q \geq 1$, and $\|b\|_\infty = \max_{i\in[p]} |b_i|$ for $q=\infty$. For $a,b\in \mathbb{R}$, we denote $a\vee b = \max(a,b)$ and $a\wedge b = \min(a,b)$. For some vector $b \in \mathbb{R}^p$ and $V\subset [p]$ some index set, we denote by $b_V \in \mathbb{R}^p$ the vector for which ${b_V}_i = b_i$ if $i\in V$ and ${b_V}_i= 0 $ if $i \notin V$. Finally, for sequences $a_T$ and $b_T$, we write $a_T \lesssim b_T$ if there exists a constant $C > 0$ such that $a_T \leq C b_T$, for all $T\geq 1$, and $a_T \asymp b_T$ if $a_T \lesssim b_T$ and $b_T \lesssim a_T$.
We consider the following data generating process for the conditional variable $Y_t$, given the regressor vector $X_t$:
where $\{U_t\}$ is a sequence of random errors satisfying $U_t | X_t \sim \text{Unif}(0,1)$, $\{X_t\}$ is a sequence of random vectors satisfying $X_t \in \mathcal{X}\subseteq\mathbb{R}^p$, and $\alpha^0(\cdot)$ is some measurable $p$-dimensional vector functional on $(0,1)$. To obtain increasing quantiles we impose, for any $x \in \mathcal{X}$, the function $x'\alpha^0(u)$ is strictly increasing in $u\in(0,1)$. In forecasting scenarios, $X_t$ contains lagged variables. Finally, we note that this model nests the location-scale model, see Section (ref).
In this model the $\tau$-quantile of $Y_t$ conditional on $X_t$, denoted $Q_t(\tau)$, has the functional form \[Q_t(\tau) = X_t' \alpha^0(\tau),\] for some quantile level $\tau\in(0,1)$. Moreover, the conditional ES, denoted $ES_t(\tau)$, has the functional form \[ES_t(\tau) := E\left[Y_t \ | \ Y_t \leq Q_t(\tau), X_t \right] = X_t' \gamma^0(\tau),\] where $\gamma^0(\tau) = \int_0^\tau \alpha^0(u) du$ denotes the ES coefficient vector.
From the above specifications, we note that the DGP in ((ref)) allows the regressors $X_t$ to influence the conditional variables $Y_t$ differently depending on the level of $Y_t$, since $\alpha^0(\tau)$ and $\gamma^0(\tau)$ are dependent on $\tau$. Moreover, the regressor vector $X_t$ may include nonlinear transformations of some subset of regressors $Z_t$. Hence, the DGP allows the regressors $Z_t$ to influence $Y_t$ differently depending on the magnitude of (the elements of) regressor vector $Z_t$. Indeed, conditional on $U_t = \tau$, $\frac{\partial Y_t}{\partial Z_t} = \frac{\partial X_t'}{\partial Z_t} \alpha^0(\tau)$, which may depend on $Z_t$ if $X_t$ contains nonlinear transformations of $Z_t$. As an example, the credit spread may influence the market return differently depending on the level of the credit spread and the level of the market return.
From now on we fix $\tau$ and remove reference to it for notational convenience. Specifically, we use $\alpha^0 = \alpha^0(\tau)$,$\gamma^0 = \gamma^0(\tau)$, $Q_t = Q_t(\tau)$, and $ES_t(\tau) = ES_t$.
Before we introduce the estimator of the ES coefficient vector $\gamma^0$, we introduce some additional notation to describe the data set. For a sample of $T$ observations, let $Y = (Y_1,\ldots,Y_T)'$ denote the vector of conditional variables and $X = [X_1,\ldots,X_p]$ the regressor matrix, where $X_i = (X_{i1},\ldots,X_{iT})'$. Moreover, define the auxiliary conditional variable $\tilde{Y} = (\tilde{Y}_1,\ldots,\tilde{Y}_T)'$, where $\tilde{Y}_t = Q_t + \frac{1}{\tau} \mathds{1}(Y_t < Q_t) (Y_t - Q_t)$. A feasible counterpart to $\tilde{Y}$ is given by $\hat{Y} = (\hat{Y}_1,\ldots,\hat{Y}_T)'$ , where $\hat{Y}_t = \hat{Q}_t + \frac{1}{\tau} \mathds{1}(Y_t < \hat{Q}_t) (Y_t - \hat{Q}_t)$, with $\hat{Q}_t = X_t'\hat{\alpha}$ and $\hat{\alpha}$ some estimator of $\alpha^0$.
It can be shown that the ES parameters $\gamma^0$ optimize a (population) least squares problem with conditional variable $\tilde{Y}_t$ and regressor vector $X_t$. This follows from the regression errors $\varepsilon = (\varepsilon_1,\ldots,\varepsilon_T)' := \tilde{Y} - X'\gamma^0$ satisfying the condition $E[\varepsilon_t|X_t] = 0$ by definition of the expected shortfall above. See barendse2020 for details.
The above argument naturally leads to an estimator of $\gamma^0$ that solves a sample counterpart to the least squares problem using the feasible auxiliary conditional variables $\hat{Y}$:
with $\|\gamma\|_{1,T} := \sum_{i=1}^p \hat{\sigma}_i |\gamma_i|$, $\hat{\sigma}_i^2 := \frac{1}{T}\sum_{t=1}^T X_{it}^2$. The second term in the optimization problem in ((ref)) is a penalty term that depends on the magnitude of the elements of the coefficient vector $\gamma$ and the penalty level $\lambda \geq 0$. The norm $\|\cdot\|_{1,T}$ is introduced in belloni2011 and used in their quantile regression problem with LASSO penalization for iid data. belloni2023high also use it in their LASSO quantile regression for time-series data. It effectively standardizes the data by giving the coefficients corresponding to regressors with high dispersion a larger penalty than coefficients of regressors with low dispersion. We utilize the equivalent norm, such that the LASSO estimator of the quantile coefficients $\alpha^0$ uses the same penalization function as the LASSO estimator of the ES coefficients $\gamma^0$.
We derive a nonasymptotic bound on the prediction error $\frac{1}{T}\|X(\hat{\gamma}-\gamma^0) \|_2^2$ and the estimation error $\|\hat{\gamma}-\gamma^0\|_1$ under the following assumptions on the data.
The first part of Condition (ref) is a simplifying assumption. In alternative cases, the prediction error $\|\hat{Y} - X\gamma\|_2^2$ does not change. The estimation error does change, but is bounded by the rate for the standardized case multiplied by $\widebar{\sigma^{-1}} := \max_{i\in[p]} \sigma_i^{-1}$, see Appendix (ref).
Let $S_0 \subset \{1,\ldots,p\}$ denotes the sparsity set of $\gamma^0$, such that $\gamma_i = 0$ if $i\notin S_0$. The cardinality of $S_0$ is denoted by $s_0$. Also define the cone $A = \{\gamma : \|\gamma_{S_0^c}\|_1 \leq C_0 \|\gamma_{S_0}\|_1\}$, for some constant $C_0 \geq 5/3$.
Condition (ref) is a restricted eigenvalue condition using the population covariance matrix $\Sigma$, see bickel2009simultaneous. This condition is standard in the literature.
Lemma (ref) below provides nonasymptotic bounds on the estimation error $\|\hat{\gamma}-\gamma^0\|_1$ and the prediction error $\frac{1}{T}\|X(\hat{\gamma}-\gamma^0) \|_2^2$ that depend on the sparsity of the ES model and the penalty level, per usual, as well as on the prediction error contained in the feasible auxiliary conditional variables $\hat{Y}$. The bound holds on the intersection of the following events: $\mathcal{S} := \{\max_{i\in[p]}\left|\hat{\sigma}_i - \sigma_i\right| \leq \frac{1}{4}\}$; $\mathcal{T} := \{ \max_{i\in[p]} \frac{2}{T}\left|\varepsilon'X_j \right| \leq \lambda_0\}$; and $\mathcal{U} := \{ \max_{i,j\in [p]} \left| \frac{1}{T}\sum_{t=1}^TX_{it} X_{jt} - E[X_{it} X_{jt}] \right| \leq \lambda_1\}$, for positive constants $\lambda_0$ and $\lambda_1$.
In this section we provide conditions under which the prediction error converges to zero in probability and the estimator is consistent, i.e. $\|X(\hat{\gamma}-\gamma^0) \|_2^2/T = o_P(1)$ and $\|\hat{\gamma}-\gamma^0\|_1 = o_P(1)$. The result is obtained by showing that the event $\mathcal{T} \cap \mathcal{S} \cap \mathcal{U}$ appearing in Lemma (ref) occurs with high probability and that the bound in Lemma (ref) converges to zero under the imposed conditions.
Condition (ref) allows for time-series dependence in the data. A formal definition of $\beta$-mixing is provided in Section (ref). In Condition (ref) below we impose a bound on $\bar\beta(t)$.
Condition (ref) is imposed to allow for $X_{i1}\varepsilon_1$ and $X_{i1}X_{j1}$, $i,j\in[p]$, to not be sub-Gaussian, sub-exponential random variables. In financial settings, the state variables and idiosyncratic errors may be heavy-tailed, such that the sub-Gaussianity or sub-exponentiality property is restrictive. Condition (ref) may be rewritten to allow for different order of moment bounds on the random variables $X_{i1}\varepsilon_1$ and $X_{i1}X_{j1}$, $i,j\in[p]$, at the cost of more involved notation. This condition is also weaker than imposing the sub-Weibull property of Wong2020.
Finally, Condition (ref) imposes rates on the parameters $s_0$, $p$, $T$, and the prediction error $\|\hat{Y} - \tilde{Y}\|_2^2/T$. Conditions (ref).1 is a consequence of the time-series features of the data and need not be imposed for iid data. Condition (ref).2 reflects the polynomial tail part in the Fuk-Nagaev inequality for the maxima of high-dimensional sums and is required if $X_{i1}\varepsilon_1$ and $X_{i1}X_{j1}$ are not sub-Gaussian or sub-exponential random variables, for any $i,j\in[p]$. From Condition (ref) we observe that, if the value of $q$ increases, such that the random variables satisfy stricter moment conditions, Condition (ref).2 becomes less restrictive on the rates of $s_0$ and $p$. Condition (ref).3 is the standard exponential rate imposed in the LASSO problem if $X_{i1}\varepsilon_1$ and $X_{i1}X_{j1}$ are iid sub-Gaussian or sub-exponential random variables, for each $i,j\in[p]$. Finally, Condition (ref).4 requires that the estimation error in the auxiliary conditional variable converges to zero in probability.
Condition (ref) introduces the terms $a_T$ and $d_T$, which depend on $T$ and a parameter $\mu'$. Moreover, the parameter $\mu$ describes the dependence in the data and is large if dependence is low. The value of $\mu'$ may be chosen close to $\mu$, such that $d_T$ may be set approximately equal to $T$ if dependence is low. Given the choice of penalty parameter $\lambda$ below, there is a trade-off between the convergence rates of the estimator and predictions (see Lemma (ref)) and the rates of the parameters. The convergence rate improves for large $d_T$ and small $a_T$, such that we like to choose $\mu'$ large. On the other hand, Condition (ref).1 restricts $p$ more for small $d_T$ and large $a_T$, such that we like to choose $\mu'$ small. This trade-off is mediated by large $\mu$ (low independence), since it reduces the need for small $\mu'$ in Condition (ref).1.
As an example, take the high dependence case $\mu = 3$. If we choose $\mu' = 1$, then Condition (ref).1 requires $p = o(T^{\mu-\mu'}) = o(T^2)$. Moreover, Condition (ref).2 requires $s_0^q p^2 T^{2-\mu'(q+1)/(1+\mu')} = s_0^q p^2 T^{2-(q+1)/2} = o(1)$. On the other hand, if we choose $\mu' = 2$, we require $p = o(T^{\mu-\mu'}) = o(T)$ and $s_0^q p^2 T^{2-2(q+1)/3} = o(1)$, such that the trade-off between parameter rates and convergence rates is clearly visible. From the example, it is also clear that less dependence ($\mu \gg 2$) or less heavy-tailedness ($q \gg 2$) in the data will improve parameter rates and convergence rates.
Lemma (ref) below gives a consistency result on the estimation and prediction errors, for penalty levels that satisify \[\lambda \asymp \frac{(pa_T)^{2/q}}{d_T^{(q-1)/q}}\vee \frac{\sqrt{\log(pa_T)}}{\sqrt{d_T}}.\]
The results in Lemma (ref) are established under the condition that the prediction error in the auxiliary variables converges to zero in probability, i.e. $\|\hat{Y} - \tilde{Y}\|_2^2/T = o_P(1)$. In this section we first show that a nonasymptotic bound of $\|\hat{Y} - \tilde{Y}\|_2^2/T$ depends on the quantile level $\tau$ and the prediction error of the quantile predictions $\|\hat{Q} - Q\|_2^2/T$, where $\hat{Q} = (\hat{Q}_1,\ldots,\hat{Q}_T)'$ denotes a sample of quantile predictions and $Q = (Q_1,\ldots,Q_T)'$ denotes the sample of true quantiles. Second, we provide sufficient conditions under which $\|\hat{Y} - \tilde{Y}\|_2^2/T = o_P(1)$, if the quantile predictions $\hat{Q}$ are generated by the quantile regression estimator of BCMPW.
The upper-bound on $\|\hat{Y} - \tilde{Y}\|_2^2/T$ given in Lemma (ref) below increases as we consider values of the quantile level $\tau$ that are closer to zero. Such values are relevant in risk management, where $\tau$ is often set to $2.5\%$ or $5\%$. It also shows that the bound is dependent on the choice of quantile predictions $\hat{Q}$. Given that the prediction error $\|\hat{Q} - Q\|_2^2/T$ usually depends on the amount of regressors $p$, it follows that a trade-off exists between $p$ and $\tau$ in terms of the bounds on the estimation and prediction errors in Lemma (ref).
We consider quantile predictions that are generated using the penalized quantile regression estimator of BCMPW:
with $\nu \asymp \sqrt{\log(pa_T)/d_T}$. The quantile predictions follow as $\hat{Q}_t = X_t'\hat\alpha$.
Lemma (ref) below states that the prediction error $\|\hat{Q} - Q\|_2^2/T$ converges to zero in probability under Condition (ref). Condition (ref) contains the relevant conditions under which $\hat{\alpha}$ is shown to be consistent in BCMPW. Lemmas (ref) and (ref) together imply $\|\hat{Y} - \tilde{Y}\|_2^2/T = o_P(1)$.
We consider an application in systemic risk, in which we study the risk spillover from the banking sector to the entire market. To measure this relationship we make use of the CoES measure of adriancovar. This measure was originally introduced to quantify risk spillover from one financial firm to the financial sector, but it is also suitable to measure the spillover from the banking sector to the entire market, as we do here.
In adriancovar CoES was introduced alongside the CoVaR measure. CoVaR has been widely used for systemic risk assessment, and yet (Co)ES is preferred over (Co)VaR from a theoretical perspective: VaR does not satisfy the coherence property for risk measures artzner1999coherent, whereas ES does satisfy it. It is due to this argument and others that ES is now the preferred risk measure for financial institutions by the Basel Committee on Banking Supervision (see, e.g., basel2016).
We formally define CoES and related quantities. Given a universe of assets, consider the market return $R^{M}_t$ and the banking industry return $R^{I}_t$. The definition of CoES of the market relative to the banking industry relies on two Value-at-Risk measurements, $\text{VaR}^{I}_{t}$ and $\text{VaR}^{M}_{t}$, which are defined (implicitly) as the following tail quantiles:
with $Z_{t-1}$ denoting a vector of lagged state variables (and where we ignore the sign convention for Value-at-Risk).
From the definitions it is clear that $\text{VaR}^{I}_{t}$ denotes the $\tau-$quantile of $R^{I}_t$ given the lagged state variables $Z_{t-1}$. The $\text{VaR}^{M}_{t}$ denotes the $\tau-$quantile of $R^M_t$ given the lagged state variables $Z_{t-1} $and the contemporaneous industry return $R^{I}_t$. The quantile level $\tau\in(0,1)$ should be taken small, e.g. $\tau = 2.5\%$.
The $\text{CoES}_{t}$ is defined as follows
and therefore denotes the ES of $R^M_{t}$ given the regressor vector $(R^I_{t}, Z_{t-1}')'$ and the event $\left\{R^I_t = \text{VaR}^{I}_{t}\right\}$. From the conditioning on the latter event, it becomes clear that the distress state in the definition of CoES is quantified by the banking industry return being at its Value-at-Risk conditional on the state variables $Z_{t-1}$.
Finally, $\Delta$CoES describes the marginal change in CoES moving from a normal state of the banking sector to a distress state, i.e. from event $\{R^{I}_t = \text{Median}^{I}_t\}$ to $\{R^{I}_t = \text{VaR}^{I}_{t}\}$, where the former event stipulates the banking return is at its median value given state variables $Z_{t-1}$. Formally,
where $\text{CoES}^{\text{median}}_{t} := E\left[R^M_{t} \ \left | \ \right. R^M_{t} \leq \text{VaR}^{M}_{t}, \ R^I_{t} = \text{Median}^{I}_{t}, \ Z_{t-1} \right]$. We recall that the median is defined as the quantile at quantile level $\tau=50\%$, such that its estimation proceeds similarly to that of $\text{VaR}^{I}_{t}$, but at $\tau = 50\%$ instead of 2.5%.
We model the market return $R^{M}_t$ and industry return $R^{I}_t$ as in $(\ref{eq:linearmodel})$:
where the regressor vectors $\phi(\cdot)$ and $\psi(\cdot)$ are known vector functionals that (nonlinearly) transform their arguments. This model is highly flexible and allows for a nonlinear relationship between state variables, market return, and banking industry return. The relationship may also vary with quantile level $\tau\in(0,1)$. We elaborate on the choice of nonlinear transformations of the regressors in Section (ref). Finally, we recall that ((ref)) generalizes the location-scale model, which is commonly used in risk management.
This model implies the following VaR and ES models:
where we allow for implicit definitions of the coefficients for brevity. Finally, the CoES is given by
We obtain weekly banking industry return and market return for the US stock market from the website of Kenneth French (\url{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}). The weekly returns are respectively created from the daily 49 Industry Portfolio data set and the daily Factor data set.
We also collect several state variables, which are lagged. The choice of state variables choice is based on adriancovar and includes the variables: (1) equity volatility; (2) real estate return; (3) TED spread; (4) Treasury bill rate (change); (5) slope of the yield curve (change); (6) credit spread (change); (7) market return. Specifically, the equity volatility is calculated as the moving average of the most recent 22 daily squared market returns; the weekly real estate return and weekly market return are taken from the 49 Industry Portfolio and Factor data sets of Ken French, respectively; the TED spread is taken as the TEDRATE series from the FRED database; the Treasury bill rate is taken as the secondary market rate on 3-month Treasury bills from the H.15 series from the FRED database; the slope of the yield curve is taken as the difference between the secondary market rate on 10-year Treasury securities and 3-month Treasury bills, both obtained from the H.15 series; the credit spread is taken as the difference between the Moody's Baa rate, taken as the BAA10Y series from the FRED database, and the secondary market rate on 10-year Treasury securities from the H.15 series.
Our sample runs from 11 January 2002 to 21 January 2022 for a total of 1,000 observations. We split the data into a training and test set, both consisting of 500 observations. We estimate the parameters of the models on the training set using the penalized methods outlined in this paper. The penalty parameters are selected using 5-fold cross validation. The folds are chosen as subsequent blocks of observations to account for the time-series features of the data. Our test set therefore runs from 2012 to 2022 and includes several crises, including the early stages of the Covid pandemic in 2020.
To allow for a nonlinear relationships between the returns and the state variables, we create transformed state variables and returns using Chebyshev polynomials up to some degree $K$, which may be increasing with sample size $T$. In this case $p$ increases with the sample size $T$, such that we are in a high-dimensional setting. The theoretical results in this paper outline conditions under which the estimator is consistent and the prediction error is small in such settings.
We outline the procedure to obtain $\phi(R^{I}_t,Z_{t-1})$, denoting the regressor vector in the model of $R^M_t$. A similar procedure is applied to the regressor vector $\psi(Z_{t-1})$ in the model of $R^I_t$. Specifically, $\phi(R^{I}_t,Z_{t-1}) = (1,\phi_{1,t},\ldots,\phi_{7K,t})'$, with $\phi_{(i-1)K+k,t} = \phi_k(S_{it})$ corresponding to the $k$§ degree Chebyshev polynomial of $S_{it}$, where $S_{it}$ is the $i$th element of $S_t = (R^I_t,Z_{t-1}')'$, and
with $\tilde{S}_{it} = (2S_{it}-a_i-b_i)/(b_i-a_i)$, where $a_i$ and $b_i$ are suitably chosen endpoints of the approximation interval $[a_i,b_i]$ for $S_{i}$. For simplicity, we choose $a_i = \min_{t\in[T]}(S_{it})$ and $b_i = \max_{t\in[T]}(S_{it})$.
We generate VaR, ES, and CoES predictions for several choices of $K$, denoting the degree of Chebyshev polynomial transformations we consider. For $K=1$ we use the unpenalized estimator which is considered the benchmark. Indeed, adriancovar use the unpenalized quantile regression estimator with similar state variables to obtain CoVaR estimates. We also consider $K= 2$, 3, 5, and 10 and use the penalized estimators for these values of $K$.
It is suspected that $K$ should be moderate for two reasons. First, for tail quantile levels like $\tau = 2.5\%$, Lemma (ref) shows that the estimation error in the quantile models, which depends on $K$, have a pronounced effect on the convergence rates of the ES predictions and estimator. Second, the time-series dependence in and heavy-tailedness of the data implies that the rate of $K$ is determined by the rates of $s_0$ and $p$ (see Condition (ref)) and worsens, ceteris paribus, for higher dependence and heavy-tailedness. Financial data often are dependent over time and heavy-tailed, such that $K$ should be relatively small.
To compare VaR and ES predictions for different $K$, we use the following out-of-sample mean prediction errors. For VaR predictions the mean prediction error is captured by the mean tick-loss (MTL): \[\text{MTL} = \sum_{t} \left(\tau - \mathds{1}(Y_t - \hat{Q}_t)\right) (Y_t - \hat{Q}_t),\] where $\hat{Q}_t$ denotes a quantile prediction for period $t$ with parameters estimated over the training set and the sum runs over the periods in the test set.
For ES predictions the mean prediction error is captures by the mean squared prediction error for ES (ES-MSE): \[\text{ES-MSE} = \sum_{t} \left(\hat{Y}_t - \widehat{ES}_t\right)^2,\] where we recall $\hat{Y}_t = \hat{Q}_t + \frac{1}{\tau} \mathds{1}(Y_t < \hat{Q}_t) (Y_t - \hat{Q}_t)$, $\widehat{ES}_t$ denotes an ES prediction for period $t$ with parameters estimated over the training set, and the sum runs over the periods in the test set. The ES-MSE function is closely related to the objective function of the penalized ES estimator in this paper, see Equation ((ref)). The mean prediction errors MTL and ES-MSE will have different scales and should not be compared to each other.
Panel A in Table (ref) contains the mean prediction errors for the VaR and ES predictions of the weekly market return $R^M_t$ conditional on the weekly banking industry return $R^I_t$ and state variables $Z_{t-1}$. The results are shown for varying $K$, which denotes the degree of Chebyshev polynomials we consider.
We observe that $K=3$ is the optimal choice, giving the smallest prediction errors for the VaR and ES predictions at quantile level $\tau =2.5\%$ and performing much better than the unpenalized benchmark ($K=1$). This provides good evidence in favor of models of market risk that are nonlinear in the industry return and state variables. Moreover, this value of $K$ may be considered moderate, such that it is reasonable the asymptotic convergence results of Lemma (ref) apply.
At the median ($\tau=50\%$), the prediction error of the VaR prediction (median prediction) is less sensitive to $K$. The same holds to a lesser extent for the ES predictions. This can be explained by the mean-blur property of high-frequency returns: the mean (median) of such returns is indistinguishable from zero due to low signal-to-noise ratio (see, e.g. christoffersen2011elements).
Panel B in Table (ref) contains the mean prediction errors for the VaR predictions of the banking industry return $R^I_t$ conditional on the state variables $Z_{t-1}$. The choice $K=2$ is optimal, although its results are quite similar to $K=3$ at quantile level $\tau=2.5\%$. The mean-blur property, again, makes outcomes relatively insensitive to $K$ for the median ($\tau=50\%$).
Finally, Panel C presents average $\Delta$CoES estimates over the test set. The $\Delta$CoES estimates rely on each of the ES and VaR predictions we have discussed above, as evident from its definition in Section (ref). As discussed above, the mean prediction error results in Panels A and B point towards $K=3$ as the optimal choice. At the optimal choice ($K=3$) we observe that average $\Delta$CoES is largest in absolute value. This suggests that the benchmark ($K=1$) underestimates systemic risk.
Finally, Figure (ref) show $\Delta$CoES predictions over the test set for $K=1$ and $K=3$, respectively denoted the small and large model. We observe that the large model responds quicker to news and predicts much greater risk spillover in several periods, including the first half of 2020, which corresponds to the onset of the Covid pandemic in the US.
In Appendix (ref) we include more plots of the ES and VaR predictions generated by the models discussed above for $K=1$ and 3 and $\tau=2.5\%$ and 50%, as well as a plot of the market and industry returns. An interesting observation coming from these plots is that the larger ES and VaR models ($K=3$) for the market return are less sensitive to positive banking industry returns than the models that use $K=1$. This behavior is most pronounced for the large positive industry returns observed in 2020 and suggests that the inclusion of nonlinear transformations of the state variables can accommodate a leverage effect for the choice $K=3$, i.e. that negative industry returns increase market risk more than positive industry returns (see, e.g., christoffersen2011elements).
Finally, at $\tau = 50\%$ we note that the penalization of the VaR estimator for banking industry returns, gives predictions that are very close to zero for all periods. The unpenalized estimator ($K=1$) gives predictions that are much more volatile. This suggests that the penalized estimator is able to accommodate the low signal-to-noise ratio in the data as concerns the conditional median.
We study the small sample properties of our estimator in a Monte Carlo simulation experiment. In each Monte Carlo iteration we estimate the parameters of the quantile and ES models on simulated data using our penalized ES estimator and the penalized QR estimator of BCMPW. We then compare the estimation and prediction errors of the estimators to those of their unpenalized counterparts.
The simulation DGP we consider is a special case of ((ref)):
where $X_t$ denotes the regressor vector and $\nu_t\overset{iid}{\sim} N(0,\sigma_\nu^2)$ denotes the idiosyncratic error. In this model, $X_t'\xi$ and $X_t'\zeta$ describe the conditional mean and conditional volatility of $Y_t$, respectively. The elements of the regressor vector $X_t$ are generated as the Chebyshev polynomial transformations up to the $K$th degree of each of the elements of a regressor vector $Z_t$ to be defined below. We add one to each of Chebyshev polynomials and subsequently standardize by their respective standard deviations. The first manipulation ensures that the elements of $X_t$ are positive, such that $X_t'\zeta$ satisfies the positivity property of the conditional volatility if all elements of $\zeta$ are positive. The second manipulation ensures that the impact on $Y_t$ of the nonzero elements in the coefficient vectors $\xi$ and $\zeta$ is similar across Monte Carlo iterations. The regressor vector $X_t$ contains $p = 1 + dK$ elements, where $d$ denotes the dimension of the regressor vector $Z_t$, as it contains a constant and $K$ transformations of each of the elements in $Z_t$. That ((ref)) is a special case of ((ref)) follows since we can rewrite $Y_t = X_t'\alpha(U_{t})$, where $\alpha(u) = \xi + \zeta Q_\eta(u)$ and $Q_\eta(u)$ denotes the quantile function corresponding to the distribution of $\nu_t$.
The regressor vector $Z_t$ is correlated in the cross-section and over time. Specifically, $Z_t = (Z_{1,t},\ldots,Z_{d,t})'$ and
where $F_t$ and $\psi_{i,t}$ are iid $N(0,1-\rho^2)$ distributed, such that $\text{Corr}\left(Z_{i,t},Z_{j,t}\right) = \theta$, $\text{Corr}\left(Z_{i,t},Z_{i,t-1}\right) = \rho$, and $\text{Var}(Z_{i,t}) = 1$, for all $i,j\in[d]$.
The coefficient vectors $\xi$ and $\zeta$ are chosen such that the vectors are sparse and contain $s_0$ nonzero elements each. The nonzero elements will differ between $\xi$ and $\zeta$. We consider $Z_{1,t}$, $Z_{2,t}$, and $Z_{3,t}$ to be relevant to the prediction of $Y_t$. Let the first $1+3K$ elements of $X_t$ be the constant and the transformations of $Z_{1,t}$, $Z_{2,t}$, and $Z_{3,t}$. In each Monte Carlo iteration we randomly draw the relevant transformations of $Z_{1,t}$, $Z_{2,t}$, and $Z_{3,t}$. Specifically, let $S_\xi \subset \{2,\ldots,1+3K\}$ and $S_\zeta \subset \{2,\ldots,1+3K\}$ denote the respective sparsity sets, with cardinality $s_0$ each, where the $s_0$ elements in each set are randomly drawn without replacement from the index set $\{2,\ldots,1+3K\}$. The sequence of draws is important in our design. Hence, we let $S_\xi^i$ and $S_\zeta^i$ denote the index that was drawn as the $i$th draw in the sequence. We then choose $\xi_{S_\xi^i} = 1/(1 + i)$ and $\xi_i = 0$ if $i \in [p] \setminus S_\xi$. Similarly, we choose $\zeta_{S_\zeta^i} = 1/(1 + i)$ and $\zeta_i = 0$ if $i \in [p] \setminus S_\zeta$.
We study the following parameter choices: cross-sectional correlation $\alpha = 0.15$; time-series correlation $\rho=0.5$; sample sizes $T=\{250,500\}$; amount of untransformed regressors $d = 7$; degree of Chebyshev transformations $K = \{3,5\}$; sparsity coefficient $s_0 = 2$; quantile levels $\tau = \{2.5\%,10\%\}$; and error variances $\sigma^2_\nu = \{1/4,1\}$. As an example, in the scenario $d=7$, $K = 5$, and $T = 500$, the regression contains a constant and $dK = 35$ regressors, of which $2s_0 = 4$ are relevant. For both samples sizes and tail quantiles under consideration, we would not expect the unpenalized estimator to perform well in this scenario for weekly financial data.
For each parameter setting and each iteration we compute the estimation and prediction errors. The estimation error is measured by $\|\hat\gamma - \gamma^0\|_1$ for the ES estimator and by $\|\hat\alpha - \alpha^0\|_1$ for the quantile estimator. The prediction error is computed out-of-sample and given by the ES-MSE and MTL functions for the ES and quantile estimator, respectively, see Section (ref). To obtain the out-of-sample prediction errors we draw a sample of $2T$ observations, reserve the first $T$ observations for estimation, and compute the out-of-sample prediction error on the final $T$ observations. For the penalized estimators, we choose the penalty parameter by five-fold cross-validation for time-series data and refer to Section (ref) for details. We report the simulation average of the estimation and prediction errors over 1,000 Monte Carlo iterations.
Table (ref) contains results for the ES estimator. We first discuss the simulation average of the estimation error in Panel A. In all scenarios we observe the penalized estimator outperforms the unpenalized counterpart. The penalized estimator improves if we increase the sample size from $T = 250$ to 500, which suggests our estimator is consistent. The unpenalized estimator improves with the sample size if $K=3$, but if $K=5$ its performance is extremely bad. This indicates that the performance of the unpenalized estimator already breaks down at a moderate amount of regressors. Finally, the performance of the estimator improves if the signal-to-noise ratio, as measured by $\sigma_\nu$, decreases or if the quantile level $\tau$ is increased.
A similar picture arises for the prediction error in Panel B in Table (ref). We specifically note that the simulation average of the prediction error of the unpenalized estimator is about an order larger than that of the penalized estimator if $K = 5$. This suggests the out-of-sample performance of the predictions will be much better for the penalized estimator in comparison to the unpenalized estimator for scenarios with a moderate amount of regressors. We note that we should not compare the prediction errors at different quantile levels $\tau$, ceteris paribus, since the functional form of the prediction errors changes with $\tau$. Table (ref) in the Appendix contains the simulation results for the quantile estimator. We generally find similar results, although at $K=5$ the difference between the penalized and unpenalized quantile estimator in terms of average prediction error is not as stark as for the ES estimator.
Finally, we derive a Fuk-Nagaev inequality for the maxima of high-dimensional sums of $\beta$-mixing sequences. Before stating the Fuk-Nagaev inequality, we define the $\beta$-mixing property and the blocking strategy $(a_T,d_T)$.
We are now ready to state the Fuk-Nagaev inequality.