EconBase
← Back to paper

Regressions under Adverse Conditions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,746 characters · 17 sections · 96 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Regressions under Adverse Conditions

\baselineskip18pt \setcounter{totalnumber}{50} \setcounter{topnumber}{50} \setcounter{bottomnumber}{50} \abovedisplayskip1.5ex plus1ex minus1ex \belowdisplayskip1.5ex plus1ex minus1ex \abovedisplayshortskip1.5ex plus1ex minus1ex \belowdisplayshortskip1.5ex plus1ex minus1ex

abstract\singlespacing We introduce a new regression method that relates the mean of an outcome variable to covariates, under the “adverse condition” that a distress variable falls in its tail. This allows to tailor classical mean regressions to adverse scenarios, which receive increasing interest in economics and finance, among many others. In the terminology of the systemic risk literature, our method can be interpreted as a regression for the Marginal Expected Shortfall. We propose a two-step procedure to estimate the new models, show consistency and asymptotic normality of the estimator, and propose feasible inference under weak conditions that allow for cross-sectional and time series applications. Simulations verify the accuracy of the asymptotic approximations of the two-step estimator. Two empirical applications show that our regressions under adverse conditions are a valuable tool in such diverse fields as the study of the relation between systemic risk and asset price bubbles, and dissecting macroeconomic growth vulnerabilities into individual components.\\ Keywords: Estimation, Growth-at-Risk, Marginal Expected Shortfall, Regression Modeling, Systemic Risk \\ JEL classification: C18, C51, C58, E44

Introduction

\onehalfspacing

In this paper, we add to the econometric toolbox by introducing regressions under adverse conditions. While the classical ordinary least squares (OLS) regression relates the mean of an outcome variable to covariates, our new regression relates the mean of an outcome variable to covariates only when some other (distress) variable falls in its tail. Here, a variable falling in its tail means an exceedance of a large conditional quantile, which is a natural statistical interpretation of an adverse condition going back to KB78. Technically, our methodology can be interpreted as a regression for Aea17's Aea17 systemic risk measure Marginal Expected Shortfall (MES), which is precisely defined as the conditional mean of some outcome variable given that a distress variable exceeds a large quantile. Therefore, we use the terms regression under adverse conditions and MES regression interchangeably.

There are various applications where the relation between an outcome variable and explanatory variables is of particular interest if some distress variable experiences an extreme outcome. First, in the systemic risk literature, the relationship between a bank's stock performance and bank characteristics is mainly of interest if the financial system is in distress AB16, Aea17, BRS20, BDP20. In this case, the bank's sensitivity to financial turmoil (i.e., its systemic riskiness) is related to bank-level data, yielding insights into the determinants of systemic risk. Our regressions under adverse conditions can deliver such insights. Second, in the recent “Growth-at-Risk” literature ABG19,BS21,Aea22, interest focuses on the risk to a country's gross domestic product (GDP) growth. Our MES regressions allow growth risk---measured with the Expected Shortfall (ES)---to be dissected into different parts that can be attributed to specific sub-components (such as sub-regions), therefore contributing to the study of the interconnections of macroeconomic risks. While we apply our regressions under adverse conditions to real data in these two fields in Section (ref), we also explore some further uses in portfolio optimization in Appendix (ref) of the Online Supplement.

The three main contributions of this paper are as follows. First, we propose a general modeling framework for the---typically dynamic---conditional MES. Our models are flexible enough to incorporate possible non-linearities in the relationship between the MES and the regressors. Our framework also allows for contemporaneous and/or lagged covariates, which may be serially dependent. Second, we provide feasible inference tools for MES regressions, including parameter and asymptotic variance estimation. This allows practitioners to identify the regressors with significant explanatory content. Third, we illustrate the usefulness of our regressions under adverse conditions in the three diverse contexts mentioned above.

From the technical side, standard (one-step) M-estimation of MES regression models is infeasible as there exists no scalar consistent loss function that is uniquely minimized in expectation by the MES FH24, DFZ_CharMest. This contrasts with classical mean regressions, which can be estimated based on the squared error loss, as this loss is minimized (in expectation) by the mean.

To make up for this defect of the MES, we draw on bivariate “multi-objective” consistent loss functions for the MES (together with the quantile) proposed by FH24. Following tradition in the risk management literature, we denote quantiles as the Value-at-Risk (VaR). These bivariate loss functions lead to a two-step M-estimator: In the first step, we estimate the parameters of the VaR model for the distress variable by minimizing the quantile loss, exactly as in quantile regressions. In the second step, we obtain parameter estimates of the MES model by minimizing the squared error loss only for those observations with a distress variable larger than the (estimated) VaR. This event is typically called a “VaR exceedance”.

In developing asymptotic theory for our two-step estimator, we draw on classic results of NeweyMcFadden1994 for consistency and on invariance principles of DMR95 for asymptotic normality. Note that for asymptotic normality we cannot rely on standard results for two-step M-estimation NeweyMcFadden1994 because our second-step objective function is discontinuous due to the truncation at the VaR. As expected for two-step estimators, the estimation precision in the second step is influenced by the first-step estimator. We further propose methods for feasible inference based on consistent estimation of the asymptotic variance-covariance matrix.

We illustrate the good finite-sample performance the two-step M-estimator and the associated inference methods in simulations, covering typical sample sizes and probability levels (for the quantile model) that we employ in our empirical applications. We further compare our regression method with an ad hoc estimator that is used in BRS20, BDP20, Berger2020 and Karolyi2023 to estimate MES models and show that this method leads to inconsistent parameter estimates together with unrealistically small standard errors.

In Section (ref), we apply our regressions under adverse conditions in the two above sketched fields of finding covariates associated with systemic risk, and in dissecting an economic region's GDP growth vulnerability into contributions of individual countries. First, we reconsider an analysis similar to BRS20 to analyze the contemporaneous relation of asset price bubbles and the systemic riskiness (as measured by their MES) of the three systemically most relevant US banks. In doing so, we compare our MES regression estimator with the ad hoc estimation method used in BRS20. Overall, we find that the MES is negatively affected by bust periods, but the effect of boom periods is not statistically significant, yielding---as in our simulations---different results compared to BRS20.

Second, the recent “Growth-at-Risk” (GaR) literature, started by ABG19, analyzes downside risks to future GDP growth as a function of current economic and financial conditions. We apply the GaR methodology to the economic region consisting of Germany, France and the United Kingdom (UK)---the three largest economies in Europe. Following ABG19, we find that financial and economic conditions pose similar risks to future growth of the whole economic region. Our regressions under adverse conditions now allow us to investigate how these total (economic and financial) risk factors can be attributed to the individual countries. We obtain the striking finding that---in contrast to France and Germany---the risk to growth in the UK is mainly determined by financial conditions, which may be explained by the importance of the financial marketplace London. Similarly, we find that---in contrast to France and the UK---current economic conditions in the joint economic region are the main risk to future German growth, which may be due to Germany's strong export-oriented manufacturing base.

Our regressions under adverse conditions are related to quantile regressions of KB78 in that they model a tail functional, and as the (preliminary) first part VaR regression is simply a (possibly nonlinear) quantile regression. In that context we also refer to EM04 and CL23 for time series applications of quantile regressions, to Hog24 for MES forecasting models based on CCC--GARCH-type filters, and to DH24 for dynamic models for the related systemic risk measure CoVaR. Our method is more closely related to recently developed regressions for the ES of DimiBayer2019, PZC19, GBP21, Bar23+ and FMW23. Our and their methods share the property of relying on an ancillary quantile regression. However, while ES regressions can be estimated jointly with VaR/quantile regressions (essentially due to the joint elicitability of the pair (VaR, ES) due to FZ16a), our MES regressions require a two-step M-estimator (due to (VaR, MES) only being multi-objective elicitable in the sense of FH24). As the MES simplifies to the ES when the distress and outcome variables coincide, our MES regressions may be seen as a multivariate extension of ES regressions, which opens up many new fields of application, as outlined above and as demonstrated in more detail in the empirical applications.

Our MES regressions are further related to threshold models (or also: sample split models), which consider (possibly distinct) regression models for the two subsamples where some threshold variable is above or below a deterministic threshold parameter Han99,Han00,CH04. The main differences between our MES regressions and threshold models are as follows. First, in the latter the threshold is a fixed scalar, whereas our models allow the threshold (i.e., the quantile of the distress variable) to be influenced by the user via the quantile level. This flexibility is important, e.g., in the portfolio application, where the quantile level (most often 97.5%) is predetermined by the regulator BCBS19. Second, while the threshold is deterministic in sample split models, we allow the stochastic threshold (i.e., the VaR) to be modeled dynamically. Third, the models we consider are more general in that our framework covers non-linear relationships, whereas threshold regressions are mostly confined to linear models.

The remainder of the paper is structured as follows. Section (ref) introduces MES regression models and our proposed two-step estimator. Its asymptotic properties and feasible inference methods are presented in Section (ref). The simulations in Section (ref) verify the good finite-sample performance of our estimator. Section (ref) demonstrates the usefulness of MES regressions in two empirical applications and the final Section (ref) concludes. The proofs of all technical results are relegated to the Supplementary Material that contains the Appendices (ref)--(ref).

Modeling and Estimation

The Definition of MES

Consider some time series $\big\{\bm V_t=(\bm W_t^\prime, X_t,Y_t)^\prime\big\}_{t\in\mathbb{N}}$, where $\bm W_t$ are exogenous variables, $X_t$ denotes the distress or conditioning variable (e.g., a measure of macroeconomic or financial conditions) and $Y_t$ stands for the variable of interest (e.g., economic or financial losses). Let $\mathcal{F}_{t}=\sigma(\bm W_{t}, \bm V_{t-1}, \bm V_{t-2}, \ldots)$ be the information set generated by the contemporaneous $\bm W_t$ and past $\bm V_t$. The inclusion of $\bm W_t$ in $\mathcal{F}_t$ allows us to incorporate contemporaneous variables in MES regressions. For instance, in studying the relationship between asset price bubbles and systemic risk, BRS20 use a contemporaneous bubble indicator in an MES-type regression; see also the empirical application in Section (ref).

For $\beta\in(0,1)$ we define the VaR as the generalized inverse $\operatorname{VaR}_{t,\beta}:=\operatorname{VaR}_{\beta}(X_t\mid\mathcal{F}_t):=F_{X_t\mid\mathcal{F}_{t}}^{\leftarrow}(\beta)$, where $F_{X_t\mid\mathcal{F}_{t}}$ denotes the conditional cumulative distribution function (c.d.f.) of $X_t$ given $\mathcal{F}_t$. The distress event that quantifies the “adverse conditions” in the definition of the MES is that the stress variable $X_t$ exceeds its VaR, i.e., $\{X_t\geq\operatorname{VaR}_{t,\beta}\}$. With our orientation of $X_t$ denoting (economic or financial) losses, we commonly consider values for $\beta$ close to one. Left-tail truncations in the sense of $\{X_t\leq\operatorname{VaR}_{t,\beta}\}$ can be analyzed within our framework by simply considering $-X_t$. We formally define our regression target under adverse conditions as $\operatorname{MES}_{t,\beta}:=\operatorname{MES}_{\beta}(Y_t\mid\mathcal{F}_{t}):=\mathbb{E}_t\big[Y_t \mid X_t\geq\operatorname{VaR}_{\beta}(X_t\mid\mathcal{F}_t)\big]$, where $\mathbb{E}_t[\cdot]:=\mathbb{E}[\ \cdot\mid\mathcal{F}_t]$ and we suppress the dependence of the MES on $X_t$ for brevity.

Joint Models for VaR and MES

Let $\bm Z_t=(Z_{1,t},\ldots,Z_{k,t})^\prime$ be the observable covariates affecting the VaR and the MES (i.e., $\operatorname{VaR}_{t,\beta}$ and $\operatorname{MES}_{t,\beta}$). We assume that $\bm Z_t$ is $\mathcal{F}_{t}$-measurable, such that it can, e.g., be a subset of the contemporaneous $\bm W_t$ and past $\bm V_t$. The empirical applications in Section (ref) provide some concrete examples for $\bm Z_t$. We introduce the functions $v_t(\bm \theta^{v})=v(\bm Z_t;\, \bm \theta^v)$ and $m_t(\bm \theta^{m})=m(\bm Z_t;\, \bm \theta^m)$, where $\bm \theta^v$ ($\bm \theta^m$) is a generic vector from some parameter space $\bm \varTheta^v\subset\mathbb{R}^p$ ($\bm \varTheta^m\subset\mathbb{R}^q$), and $v(\,\cdot\,;\bm \theta^v)$ and $m(\,\cdot\,;\bm \theta^m)$ are measurable functions for all $\bm \theta^v\in\bm \varTheta^v$ and $\bm \theta^m\in\bm \varTheta^m$, respectively. We assume that there exist unique true parameters $\bm \theta^v_0 \in \bm \varTheta^{v}$ and $\bm \theta^m_0 \in \bm \varTheta^{m}$, such that almost surely (a.s.)

align[align omitted — 294 chars of source]

Thus, $v_t(\bm \theta_{0}^{v})$ is simply a (possibly non-linear) quantile regression model that relates the $\beta$-quantile of the distress variable $X_t$ to covariates $\bm Z_t$. This may be seen as an auxiliary regression step that allows to identify the “adverse condition” $\{X_t\geq\operatorname{VaR}_{\beta}(X_t\mid\mathcal{F}_t)\}$ in the actual quantity of interest $\operatorname{MES}_{\beta}(Y_t\mid\mathcal{F}_{t})=\mathbb{E}_t\big[Y_t \mid X_t\geq\operatorname{VaR}_{t,\beta}\big]$, which is modeled via $m_t(\bm \theta_{0}^{m})$ and serves to explain the interrelation between $Y_t$ and $X_t$ (in the form of the MES) with covariates. Therefore, we call the models $v_t(\bm \theta_{0}^{v})$ and $m_t(\bm \theta_{0}^{m})$ in (ref) a regression under adverse conditions or also an MES regression model. The leading special case of (ref) is the linear model shown in Example (ref) below, where $v_t(\bm \theta^{v})$ and $m_t(\bm \theta^{m})$ are linear in the parameters $\bm \theta^v$ and $\bm \theta^m$, respectively.

The fact that the conditional-on-$\mathcal{F}_t$ quantities $\operatorname{VaR}_{t,\beta}$ and $\operatorname{MES}_{t,\beta}$ only depend on $\bm Z_t$ is, of course, a modeling assumption. Except for the restriction that $\bm Z_t$ be $\mathbb{R}^{k}$-valued, this is quite a flexible assumption. It covers purely predictive regressions (where $\bm Z_t$ only contains information available at time $t-h$, say) and also static regressions (where $\bm Z_t$ only contains contemporaneous variables) and any combination of the two.

Note that we consider the case of separated parameters in (ref), where the parameters pertain only to the VaR model or the MES model. This is vital for our two-step M-estimator introduced below.

We also mention that our framework inherently models the interconnectedness between $X_t$ and $Y_t$ through the (parameter-free) definition of the MES. This contrasts with, e.g., AB16, whose estimation method (for the closely related CoVaR) can only capture linear effects of the VaR of $X_t$ onto $Y_t$.

exampleHere, we provide an example of a model that gives rise to a linear regression under adverse conditions in (ref). Consider the model \begin{align*} X_t &= \bm Z_t^{v\prime}\bm \theta_0^{v}+\varepsilon_{X,t},\\ Y_t &= \bm Z_t^{m\prime}\bm \theta_0^{m}+\varepsilon_{Y,t}, \end{align*} where $\bm Z_t^{v}$ and $\bm Z_t^{m}$ are ($p$ and $q$-dimensional) vectors containing some (or all) components of $\bm Z_t$, and we assume that $\operatorname{VaR}_{\beta}(\varepsilon_{X,t}\mid\mathcal{F}_t)=F_{\varepsilon_{X,t}\mid\mathcal{F}_t}^{\leftarrow}(\beta)=0$, and $\operatorname{MES}_{\beta}(\varepsilon_{Y,t} \mid \mathcal{F}_t) = \mathbb{E}_t\big[\varepsilon_{Y,t}\mid \varepsilon_{X,t}\geq \operatorname{VaR}_{\beta}(\varepsilon_{X,t}\mid\mathcal{F}_t)\big]=0$. Then, exploiting the $\mathcal{F}_t$-measurability of $\bm Z_t$, it is easy to check that (ref) holds with linear $v_t(\bm \theta_{0}^{v})=\bm Z_t^{v\prime}\bm \theta_0^{v}$ and linear $m_t(\bm \theta_{0}^{m})=\bm Z_t^{m\prime}\bm \theta_0^{m}$. Recall that $\operatorname{VaR}_\beta(\varepsilon_{X,t}\mid\mathcal{F}_t)=0$ is the standard assumption on the errors in quantile regressions. The requirement that $\operatorname{MES}_{\beta}(\varepsilon_{Y,t} \mid \mathcal{F}_t) = 0$ may then be viewed as the analog assumption in our MES regressions. Figure (ref) graphically illustrates a linear regression under adverse conditions with one covariate. The top panel shows the quantile regression line $v_t(\bm \theta_0^v)$ in blue and the bottom panel the MES regression line $m_t(\bm \theta_0^m)$ in red. The black points in both panels indicate the observations $(X_t,Y_t)^\prime$ with a VaR exceedance, such that $X_t\geq v_t(\bm \theta_0^v)$. It is obvious from this figure that the auxiliary quantile regression in the top plot is key in identifying the relevant points for the MES regression in the bottom plot.
figure[figure omitted — 372 chars of source]

While “usual” regressions simply relate some outcome $Y_t$ to covariates, our MES regressions allow us to investigate such a relationship conditional on an extreme realization of $X_t$. This is of interest, e.g., when relating a bank's losses $Y_t$ to institution characteristics, where the precise relationship is mainly of interest during times of large system-wide losses $X_t$. We use the model of Example (ref) to investigate this in the empirical application in Section (ref), but also in the other applications in Section (ref) and Appendix (ref). We discuss the relation of our MES regressions to OLS regressions that use $X_t$ as an explanatory variable in Appendix (ref).

remUsing the notation $\bm Z_t^m = (Z^m_{1,t},\dots, Z^m_{q,t})^\prime$ and $\bm \theta^m = (\theta_1^m, \dots, \theta_q^m)^\prime$, our models in (ref) allow for the definition of a marginal effect under adverse conditions with respect to $Z^m_{j,t}$ for any $j \in \{1,\dots,q\}$ via $\frac{\partial}{\partial Z^m_{j,t}} m_t(\bm \theta^m)$. This captures the impact of a unit change in $Z^m_{j,t}$ on the conditional MES given all other variables remain constant, which is analogous to the classical marginal effect for mean or quantile regressions. For the linear models of Example (ref), the marginal effect under adverse conditions is simply $\theta^m_j$.

Parameter Estimation

We now introduce our estimators of the unknown model parameters $\bm \theta_{0}^{v}$ and $\bm \theta_{0}^{m}$. Typically, (one-step) M-estimation requires a to-be-minimized loss function (or also: objective function). Yet, as pointed out in the Introduction, there is no real-valued loss function associated with the pair (VaR, MES). The existence of such a (strictly consistent) loss function is, however, necessary for consistent M-estimation of semiparametric models such as ours DFZ_CharMest.

To overcome this lack, we draw on FH24. Denoting by $\mathds{1}_{\{\cdot\}}$ the indicator function, they show that (under some smoothness conditions) the $\mathbb{R}^2$-valued loss function

equation[equation omitted — 353 chars of source]

is minimized in expectation by the true VaR and MES with respect to the lexicographic order. That is, for all $v,m\in\mathbb{R}$, \[ \mathbb{E}\Bigg[\bm S\Bigg(

pmatrix[pmatrix omitted — 64 chars of source]

,

pmatrix[pmatrix omitted — 18 chars of source]

\Bigg)\Bigg]\preceq_{\operatorname{lex}}\mathbb{E}\Bigg[\bm S\Bigg(

pmatrix[pmatrix omitted — 18 chars of source]

,

pmatrix[pmatrix omitted — 18 chars of source]

\Bigg)\Bigg], \] where $(x_1,x_2)^\prime\preceq_{\operatorname{lex}} (y_1,y_2)^\prime$ if $x_1<y_1$ or ($x_1=y_1$ and $x_2\leq y_2$), and $\operatorname{VaR}_\beta$ and $\operatorname{MES}_\beta$ denote the true VaR and MES of the joint distribution of $(X,Y)^\prime$. Note that $S^{\operatorname{VaR}}(\cdot,\cdot)$ in (ref) is the standard pinball-loss function known from quantile regression. Clearly, $S^{\operatorname{MES}}(\cdot,\cdot)$ resembles the squared error loss familiar from usual mean regressions, except that the indicator $\mathds{1}_{\{x>v\}}$ restricts the evaluation to observations with VaR exceedances in the first component. The pre-factor $1/2$ is simply a convenient normalization, enabling an easier representation of some subsequent quantities. In the related literature, loss functions are often also called scoring functions Gne11 and we use these two terms interchangeably.

To introduce our two-step estimator, suppose that we have a sample $\big\{(X_t,Y_t,\bm Z_t^\prime)^\prime\big\}_{t=1,\ldots,n}$ of size $n$. In a first step, the lexicographic order suggests to minimize the expected score of the VaR component. Approximating the expected score by the sample average of the scores, the first-step estimator is \[ \widehat{\bm \theta}_{n}^{v} = \operatorname*{arg\,min}_{\bm \theta^{v}\in\bm \varTheta^{v}}\frac{1}{n}\sum_{t=1}^{n}S^{\operatorname{VaR}}\big(v_t(\bm \theta^{v}), X_t\big). \] OH16 study this estimator in their nonlinear quantile regressions for deterministic regressors.

Given $\widehat{\bm \theta}_n^v$, the lexicographic order similarly suggests in a second step to minimize the sample average of the MES losses via

align[align omitted — 288 chars of source]

For this two-step estimator to be feasible, the requirement that the VaR model does not depend on the MES model is essential. Section (ref) shows that the estimation precision of $\widehat{\bm \theta}_{n}^{m}$ is affected by relying on $\widehat{\bm \theta}_n^v$ in the second step, which is the usual case for two-step estimators NeweyMcFadden1994.

Asymptotics for the Two-Step Estimators

We now show the joint asymptotic normality of our two-step M-estimators $\widehat{\bm \theta}_{n}^{v}$ and $\widehat{\bm \theta}_{n}^{m}$ under the regularity conditions in Assumptions (ref)--(ref) given below. Before presenting Assumptions (ref)--(ref), we have to introduce some notation. The joint c.d.f. of $(X_{t}, Y_t)^\prime\mid\mathcal{F}_t$ is denoted by $F_t(\cdot,\cdot)$, and its Lebesgue density (which we assume exists) by $f_t(\cdot,\cdot)$. Similarly, $F_t^{W}(\cdot)$ ($f_t^{W}(\cdot)$) denotes the distribution (density) function of $W_t\mid\mathcal{F}_t$ for $W\in\{X,Y\}$. For sufficiently smooth functions $\mathbb{R}^{p}\ni\bm \theta\mapsto f(\bm \theta) \in \mathbb{R}$, we denote the $(p\times1)$-gradient by $\nabla f(\bm \theta)$, its transpose by $\nabla' f(\bm \theta)$ and the $(p\times p)$-Hessian by $\nabla^2 f(\bm \theta)$. All (in-)equalities involving random quantities are meant to hold almost surely.

The asymptotic variance-covariance matrices of our estimators depend on the following matrices:

align*[align* omitted — 931 chars of source]

All these matrices exist by virtue of Assumptions (ref)--(ref), which we state now. For this, let $K<\infty$ be some large universal positive constant, and denote by $\norm{\bm x}$ the Euclidean norm when $\bm x$ is a vector, and the Frobenius norm when $\bm x$ is matrix-valued.

assumption\begin{enumerate} • There exist unique parameters $\bm \theta_0^v$ and $\bm \theta_0^m$, such that (ref) holds a.s. for all $t\in\mathbb{N}$. • The parameter space $\bm \varTheta=\bm \varTheta^{v}\times\bm \varTheta^{m}$ is compact, where $\bm \varTheta^{v}\subset\mathbb{R}^p$ and $\bm \varTheta^{m}\subset\mathbb{R}^q$. • $\bm \theta_0^v\in\operatorname{int}(\bm \varTheta^v)$ and $\bm \theta_{0}^{m}\in\operatorname{int}(\bm \varTheta^m)$, where $\operatorname{int}(\cdot)$ denotes the interior of a set. \end{enumerate}
assumption$\big\{(X_t, Y_t, \bm Z_t^\prime)^\prime\big\}_{t\in\mathbb{N}}$ is strictly stationary and $\beta$-mixing with coefficients $\beta(\cdot)$ of size $-r/(r-1)$ for some $r>1$.
assumption\begin{enumerate} • For all $t\in\mathbb{N}$, $F_t(\cdot,\cdot)$ belongs to a class of distributions on $\mathbb{R}^2$ that possess a positive Lebesgue density $f_t(x,y)$ for all $(x,y)^\prime\in\mathbb{R}^2$ such that $F_t(x,y)\in(0,1)$. • $\big|f_t^{X}(x)-f_t^{X}(x^\prime)\big|\leq K|x-x^\prime|$ and $\sup_{x\in\mathbb{R}}f_t^{X}(x)\leq K$. • The conditional density $f_t(x,y)$ is differentiable in its first argument, with partial derivative denoted by $\partial_1 f_t(x,y)$. • $\sup_{x\in\mathbb{R}}\big|\int_{-\infty}^{\infty}y\partial_1 f_t(x,y)\,\mathrm{d} y\big|\leq F(\mathcal{F}_t)$ and $\sup_{x\in\mathbb{R}}\int_{-\infty}^{\infty}|y| f_t(x,y)\,\mathrm{d} y\leq F_1(\mathcal{F}_t)$ for $\mathcal{F}_t$-measurable random variables $F(\mathcal{F}_t)$ and $F_1(\mathcal{F}_t)$. \end{enumerate}
assumption\begin{enumerate} • For all $t\in\mathbb{N}$, $v_t(\cdot)$ and $m_t(\cdot)$ are twice differentiable on $\operatorname{int}(\bm \varTheta^v)$ and $\operatorname{int}(\bm \varTheta^m)$ with gradients $\nabla v_t(\cdot)$ and $\nabla m_t(\cdot)$, and Hessians $\nabla^2 v_t(\cdot)$ and $\nabla^2 m_t(\cdot)$. • $\sup_{\bm \theta^v\in\bm \varTheta^v}\big|v_t(\bm \theta^v)\big|\leq V(\bm Z_t)$ and $\sup_{\bm \theta^m\in\bm \varTheta^m}\big|m_t(\bm \theta^m)\big|\leq M(\bm Z_t)$ for some measurable functions $V(\cdot)$ and $M(\cdot)$. • $\sup_{\bm \theta^v\in\bm \varTheta^v}\big\Vert\nabla v_t(\bm \theta^v)\big\Vert\leq V_1(\bm Z_t)$ and $\sup_{\bm \theta^m\in\bm \varTheta^m}\big\Vert\nabla m_t(\bm \theta^m)\big\Vert\leq M_1(\bm Z_t)$ for some measurable functions $V_1(\cdot)$ and $M_1(\cdot)$. • $\sup_{\bm \theta^v\in\bm \varTheta^v}\norm{\nabla^2v_t(\bm \theta^v)}\leq V_2(\bm Z_t)$ and $\sup_{\bm \theta^m\in\bm \varTheta^m}\norm{\nabla^2m_t(\bm \theta^m)}\leq M_2(\bm Z_t)$ for some measurable functions $V_2(\cdot)$ and $M_2(\cdot)$. • There exist neighborhoods of $\bm \theta_0^v$ and $\bm \theta_0^m$, such that $\norm{\nabla^2 v_t(\bm \tau^v) - \nabla^2 v_t(\bm \theta^v)}\leq V_3(\bm Z_t)\norm{\bm \tau^v-\bm \theta^v}$ and $\norm{\nabla^2 m_t(\bm \tau^m) - \nabla^2 m_t(\bm \theta^m)}\leq M_3(\bm Z_t)\norm{\bm \tau^m-\bm \theta^m}$ for all elements $\bm \theta^v$, $\bm \tau^v$ and $\bm \theta^m$, $\bm \tau^m$ of the respective neighborhoods, and measurable functions $V_3(\cdot)$ and $M_3(\cdot)$. \end{enumerate}
assumptionFor $r>1$ from Assumption (ref) and some $\iota>0$, it holds that $\mathbb{E}\big[V(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[V_1^{4r}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[V_2^{2r}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[V_3(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[M^{4r+\iota}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[M_1^{4r+\iota}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[M_2^{4r}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[M_3^{4r/(4r-1)}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[F^{4r/(4r-3)}(\mathcal{F}_t)\big]\leq K$, $\mathbb{E}\big[F_1^{4r/(4r-3)}(\mathcal{F}_t)\big]\leq K$, $\mathbb{E}|X_t|\leq K$, $\mathbb{E}|Y_t|^{4r+\iota}\leq K$.
assumptionThe matrices $\bm \varLambda$ and $\bm \varLambda_{(1)}$ are positive definite.
assumption$\sup_{\bm \theta^v\in\bm \varTheta^v}\sum_{t=1}^{n}\mathds{1}_{\{X_t=v_t(\bm \theta^v)\}}\leq K$ a.s. for all $n\in\mathbb{N}$.

Assumption (ref) (ref) ensures identification of the true parameters. Compactness in Assumption (ref) (ref) is a standard requirement in extremum estimation; see NeweyMcFadden1994. Assumption (ref) (ref) forces the true parameters to be interior to the respective parameter spaces, which is a key requirement for proving asymptotic normality. Assumption (ref) is a standard stationarity and mixing condition ensuring that suitable laws of large numbers and central limit theorems apply. Often, the $\beta$-mixing coefficients decrease exponentially, such that Assumption (ref) holds for any $r>1$ arbitrarily close to unity; see, e.g., CC02 and Lie05 for ARMA and GARCH processes. Assumption (ref) (ref) ensures strict (multi-objective) consistency of the scoring function given in (ref); see FH24. The final items (ref)--(ref) of Assumption (ref) may be interpreted as smoothness conditions for the conditional density. Since the VaR and MES only depend on $\bm Z_t$ under our modeling framework, it will most often be the case that the conditional density $f_t(\cdot,\cdot)$ will also be $\bm Z_t$-measurable; see, e.g., Appendix (ref). Nonetheless, the random variables $F(\mathcal{F}_{t})$ and $F_1(\mathcal{F}_{t})$ appearing in Assumption (ref) (ref) are in principle allowed to be $\mathcal{F}_{t}$-measurable. Assumption (ref) provides smoothness conditions on the VaR and MES model, and Assumption (ref) collects moment bounds. When mixing is exponential (corresponding to $r=1$ in Assumption (ref)), the moment bounds of Assumption (ref) are weakest, such that there is the usual moment-memory tradeoff. Note that the moment conditions for the VaR model are weaker than those for the MES model, because the latter is estimated based on a squared error-type loss. Assumption (ref) ensures that the asymptotic variance-covariance matrix in Theorem (ref) is well-defined, which can however be degenerate if $\bm V$ or $\bm M^\ast$ are singular; also see Hansen:82. Assumption (ref) is identical to Assumption 2 (G) in PZC19. It prevents an infinite number of $X_t$'s to lie on any possible “quantile regression line” $v_t(\bm \theta^v)$. Appendix (ref) shows that $K=\dim(\bm \theta^v)$ for linear models.

thmSuppose Assumptions (ref)--(ref) hold. Then, as $n\to\infty$, \begin{align*} \sqrt{n}\begin{pmatrix} \widehat{\bm \theta}_n^{v}-\bm \theta_0^v\\ \widehat{\bm \theta}_n^{m}-\bm \theta_0^m \end{pmatrix} \overset{d}{\longrightarrow}N\big(\boldsymbol{0}, \bm \varGamma \bm M \bm \varGamma^\prime), \end{align*} where \[ \bm \varGamma =\begin{pmatrix} \bm \varLambda^{-1} & \boldsymbol{0}\\ -\bm \varLambda_{(1)}^{-1}\bm \varLambda_{(2)}\bm \varLambda^{-1} & \bm \varLambda_{(1)}^{-1}\end{pmatrix}\in\mathbb{R}^{(p+q) \times (p+q)}\qquad\text{and}\qquad \bm M=\begin{pmatrix}\bm V &\boldsymbol{0} \\ \boldsymbol{0} & \bm M^\ast\end{pmatrix}\in\mathbb{R}^{(p+q)\times(p+q)}. \]

\sloppy The proof of Theorem (ref) is in Appendix (ref) and draws on DLS23. In the special case $\bm \varLambda_{(2)}=\boldsymbol{0}$, the first-step estimator does not have an impact on the asymptotic variance of the second-step estimator NeweyMcFadden1994. To gain some intuition for this, rewrite the curly bracket in $\bm \varLambda_{(2)}$ as $f_t^X(v_t(\bm \theta_0^v)) \left( \mathbb{E} \big[ Y_t \big| X_t = v_t(\bm \theta_0^v) \big] - \mathbb{E} \big[ Y_t \big| X_t \ge v_t(\bm \theta_0^v) \big] \right)$, which measures how the expectation of $Y_t$ reacts upon moving from $\{X_t = v_t(\bm \theta_0^v)\}$ further to the tail through $\{ X_t \ge v_t(\bm \theta_0^v)\}$. Hence, $\bm \varLambda_{(2)}=\boldsymbol{0}$ can be interpreted as a form of “conditional tail uncorrelatedness” of $X_t$ and $Y_t$, as $\bm \varLambda_{(2)}$ is a weighted expectation of the above quantity.

rem[Expected Shortfall Regressions] For $X_t = Y_t$, the definition of the MES simplifies to the Expected Shortfall, $\operatorname{ES}_{t,\beta} := \operatorname{ES}_{\beta}(X_t\mid\mathcal{F}_t) := \mathbb{E}_t\big[X_t \mid X_t \geq \operatorname{VaR}_{t,\beta}\big]$. For the ES, associated regression models (jointly with the conditional VaR) have been proposed recently. E.g., DimiBayer2019 and PZC19 consider joint M-estimation with a VaR model, which is feasible as (VaR, ES) are jointly elicitable, as opposed to the pair (VaR, MES), which is merely multi-objective elicitable. One can also use a two-step estimator for VaR and ES models, whose second step is essentially given by setting $Y_t = X_t$ in the right-hand side of (ref). Its asymptotic variance-covariance matrix resembles the one presented in Theorem (ref) by again setting $Y_t = X_t$ in all matrices but $\bm \varLambda_{(2)}$, where an ES-specific version is \begin{align*} \bm \varLambda_{(2)}^ES =\mathbb{E}\Big[ \big\{v_t(\bm \theta_0^v)- m_t(\bm \theta_0^m) \big\} f_{t}^{X}\big(v_t(\bm \theta_0^v)\big) \nabla m_t(\bm \theta_0^m)\nabla^\prime v_t(\bm \theta_0^v)\Big]. \end{align*} This arises as $\int_{-\infty}^{\infty} y f_t\big(v_t(\bm \theta_0^v),y\big) \,\mathrm{d} y = f_{t}^{X}\big(v_t(\bm \theta_0^v)\big) \mathbb{E}_t \big[ Y_t \mid X_t = v_t(\bm \theta^v_0) \big]$, and the latter conditional expectation equals $v_t(\bm \theta^v_0)$ if $Y_t = X_t$. Notice, however, that the joint distribution of $(Y_t, X_t)^\prime$ is degenerate if $Y_t = X_t$, such that our Assumption (ref) is violated. Therefore, Assumption (ref) is replaced in ES models by the requirement that the (univariate) conditional distribution of $X_t$ is sufficiently smooth around its $\beta$-quantile; see, e.g., PZC19.

For feasible inference, we have to estimate the asymptotic variance-covariance matrix of Theorem (ref) consistently. In Appendix (ref), we propose suitable estimators and show their consistency in Theorem (ref) under the additional Assumptions (ref)--(ref).

Appendix (ref) verifies Assumptions (ref)--(ref) and also Assumptions (ref)--(ref) for the linear MES regression model of Example (ref) subject to some regularity conditions.

Simulations

The Data-Generating Process

We consider a time series regression as a data-generating process (DGP), which exhibits the desirable features of a time-varying conditional mean and covariance matrix and at the same time results in a correctly specified regression for the MES (and the VaR) that allows for a comparison with the estimation used in BRS20, BDP20 in Section (ref).

Using the notation of Example (ref), the covariates $\bm Z^v_{t} = \bm Z^m_{t} =(1, Z_{1,t}, Z_{2,t})^\prime$ follow the (transformed) auto-regressions

align*[align* omitted — 171 chars of source]

where $\varepsilon_{Z,t}, \varepsilon_{\xi,t} \stackrel{\text{i.i.d.}}{\sim} {N}(0,1)$, independently of each other for all $t=1,\ldots,n$. The exponential transformation guarantees positivity of $Z_{1,t}$, which is required to introduce heteroskedasticity in the process

align[align omitted — 265 chars of source]

The multivariate $t$-distributed innovations $\boldsymbol{\varepsilon}_t =(\varepsilon_{1,t}, \varepsilon_{2,t})^\prime\stackrel{i.i.d.}{\sim} t_{6} \big( \boldsymbol{0}, \boldsymbol{\Sigma} =

psmallmatrix1 & 1.2 \\ 1.2 & 4

\big)$ are independent of the $\{\varepsilon_{Z,t}\}$ and $\{\varepsilon_{\xi,t}\}$. We choose $\bm \gamma=(\gamma_1, \ldots, \gamma_5)^\prime =(1,\ 1.5,\ 2,\ 0.25,\ 0.5)^\prime$ as the true values. For simplicity, our DGP in \eqref{eqn:DGPCrossSectional} implies that $X_t$ and $Y_t$ are driven by the same factors and only the heteroskedastic shock differs between the variables. This DGP results in the linear (VaR, MES) model

align*[align* omitted — 194 chars of source]

The true parameter values $\bm \theta_0^v=(\theta_{1,0}^v,\, \theta_{2,0}^v,\, \theta_{3,0}^v)^\prime$ and $\bm \theta_0^m=(\theta_{1,0}^m,\, \theta_{2,0}^m,\, \theta_{3,0}^m)^\prime$ are given by

align*[align* omitted — 298 chars of source]

where $\tilde m_\beta$ is the $\beta$-MES of the $t_{6} \big( \boldsymbol{0}, \boldsymbol{\Sigma} \big)$-distribution, and $\tilde q_\beta$ is the $\beta$-quantile of its first component. While the true parameters for $\theta_{3,0}^v$ and $\theta_{3,0}^m$ coincide, this specification does not violate the separated parameter condition in (ref) as we do not impose the condition $\theta_{3}^v = \theta_{3}^m$ in the estimation.

Simulation Results

We estimate the correctly specified six-parameter (VaR, MES) regression model based on the joint covariates $\bm Z^v_{t} = \bm Z^m_{t}$ given above. Table (ref) shows simulation results for the estimated parameters and their standard deviations. Results are based on $5000$ Monte Carlo replications of the DGP in (ref) for the probability levels $\beta \in \{0.9,\ 0.95,\ 0.975\}$, which we also use in our two applications in Sections (ref)--(ref). We consider the typical sample sizes $n \in \{500,\ 1000,\ 2000,\ 4000\}$. The asymptotic variances are estimated as detailed in Appendix (ref). A formal description of the table columns is given in the table caption.

table[table omitted — 4,752 chars of source]

We find that both the VaR and MES parameters are estimated consistently as the empirical bias and the empirical standard deviation shrink as $n$ grows in all settings. As expected, the estimates are more accurate for moderate (smaller) probability levels $\beta$, but even the largest choice $\beta = 0.975$ leads to fairly accurate estimates in large samples. Overall, the MES parameters are estimated with similar accuracy as the corresponding VaR parameters, and so are their asymptotic variances. These results are confirmed by the coverage rates of the resulting confidence intervals, where the empirical coverage rates approach the nominal rate of $95\%$. We observe some undercoverage in finite samples that, however, vanishes with decreasing $\beta$ and increasing sample size $n$.

Comparison with Existing Approaches to MES Regressions

BRS20, BDP20, Berger2020 and Karolyi2023 among others use the following approach to MES regressions. Define the empirical MES estimate over a rolling window of $S \in \mathbb{N}$ (typically, $S=250$) past observations, i.e.,

align[align omitted — 227 chars of source]

where $\widehat{Q}_\beta(X_{(t-S):t})$ denotes the sample quantile of $\{X_{t-S},\dots,X_t\}$ in the rolling window. BRS20, BDP20 then relate the transformed response $Y_t^\ast$ to covariates $\bm Z_t^m$ in a standard OLS mean regression; see BRS20 and BDP20 for details. We henceforth call this method the “ad hoc MES regression”. A mean regression of $Y_t^\ast$ on $\bm Z_t^m$ estimates the conditional expectation $ \mathbb{E} \big[ Y_t^\ast \mid \bm Z_{t}^m \big]$, whose interpretation is unclear and which is in general different from the de facto target of an MES regression, that is, $\operatorname{MES}_{t,\beta}=\mathbb{E}\big[Y_t \mid X_t\geq\operatorname{VaR}_{t,\beta}, \bm Z_{t}^m \big]$. A further drawback of the ad hoc MES regressions is that---in contrast to our MES regression---the covariates $\bm Z_t^m$ cannot contain past values of $X_t$ or $Y_t$ as these would influence both the left-hand and right-hand sides of the associated OLS regression, as $Y_t^\ast$ is a rolling average over past values of $Y_t$ (truncated by using lagged values of $X_t$).

To illustrate in a simpler context, such an ad hoc regression for the $\beta$-quantile would relate the moving window sample quantile $Y_t^{\dagger} = \widehat{Q}_\beta(X_{(t-S):t})$ to covariates. This contrasts with quantile regression of KB78, which relates $Y_t$ directly to the regressors by using the quantile-specific check loss function given in the first row of (ref).

figure[figure omitted — 392 chars of source]

Figure (ref) illustrates how severe the discrepancy between our and the ad hoc version of the MES regression is for the DGP in (ref) for $\beta = 0.95$. We do so by estimating a mean regression of the target $Y_t^\ast$ on the same covariates $(1,Z_{1,t}, Z_{2,t})'$ and by computing our MES estimate $\widehat{\bm \theta}_n^m$. While the parameter estimates of our MES regression are normally distributed around the true parameter values (given by the black lines), the estimates based on the ad hoc estimation method are far from the true values. Most strikingly, the ad hoc estimates of the two slope parameters associated with the covariates $Z_{1,t}$ and $Z_{2,t}$ vary around zero because the contemporaneous effect of the covariates on the rolling window quantity $Y_t^\ast$ is negligible. This illustrates that employing the ad hoc MES regression procedure does not capture the effect of covariates on the MES functional, but rather for some functional that is inherently hard to interpret; see (ref).

Empirical Applications

We illustrate the versatility of our MES regressions in three applications. These concern explanatory variables of systemic risk in the banking sector in Section (ref) and dissecting GDP growth vulnerabilities among the three biggest European economies in Section (ref). A further application concerning (equal risk contribution) portfolios is deferred to Appendix (ref).

Systemic Risk Regressions

With the introduction of RiskMetrics RM96, financial risk management---based mainly on the VaR---was beginning to be firmly established in the 1990s. Subsequently, the importance of adequately managing individual financial risks was also reflected in official regulations by the BCBS96. As these regulations did not prevent the financial crisis of 2007--08, more attention (of regulators as well as industry) has been paid to the systemic nature of financial risks. This means that more scrutiny was applied in studying the interconnectedness of individual institutions (in the context of the financial system as a whole) or the interconnectedness of individual trading desks (in the context of managing bank-wide risks). As a consequence of this development, a huge literature on systemic risk and its determinants has emerged ABT12,AB16,Aea17,BE17,BRS20,BDP20.

For instance, BRS20 investigate the contemporaneous relation of asset price bubbles and systemic risk by, among others, considering the (conditional) MES in their Section 4.2. However, our simulations in Section (ref) show that the ad hoc estimation method for MES regressions---as applied by the aforementioned authors---does not necessarily provide consistent parameter estimates nor allows for valid inference. Therefore, we now consider a simplified form of their analysis for the conditional MES by using our regressions under adverse conditions. Here, we mainly aim at identifying which covariates drive the conditional MES, and as in BRS20, we particularly focus on stock market boom and bust indicators.

Specifically, we focus on the three US banks that are classified as systemically most risky according to the FSB22 such that $Y_t$ equals the daily log-losses (i.e., the negative log-returns) of either the Bank of America Corporation (BAC), Citigroup (C) or JPMorgan Chase (JPM). We provide corresponding results for the other five US banks listed as a global systemically important bank (G-SIB) by the FSB22 in Appendix (ref). Throughout, we use the daily log-losses of the S&P 500 Financials for $X_t$ and set $\beta = 0.95$. Hence, the MES measures the mean loss of the bank conditional on the financial system being in distress.

We estimate the joint linear model $ \big( v_t(\bm \theta^v), m_t(\bm \theta^m) \big)' = \big( \bm Z_{t}^{v\,\prime} \bm \theta^v, \bm Z_{t}^{m\,\prime} \bm \theta^m\big)'$, where $ \bm \theta^v, \bm \theta^m \in \mathbb{R}^6$. For reasons of data availability and concerns of non-stationarity of some regressors, we use a restricted set of variables in $\bm Z_{t}^{v} = \bm Z_{t}^{m} = \bm Z_t = (1, Z_{t,1},\ldots,Z_{t,5})^\prime$, which contain an intercept, the change in spread between Moody's Baa-rated bonds and the 10-year Treasury bill rate (Change Spread), the spread between the 3-month LIBOR and 3-month Treasury bills (TED Spread), the VIX index as a forward-looking measure of market volatility, and separate indicator variables for stock market booms (SM Boom) and busts (SM Bust) from BRS20. To study the contemporaneous relation between systemic risk and bubbles, we follow BRS20 by lagging all explanatory variables by one time period, except for the boom and bust indicators for which we use contemporaneous values.

To allow for an immediate comparison of the estimation methods, we also estimate the “ad hoc MES regression” described in Section (ref), which is a classical OLS regression of the transformed target variable defined in (ref). We estimate the models based on daily data and use the largest common available sample ranging from May 5, 1993 until December 31, 2015, yielding a total of $n=5,555$ trading days.

table[table omitted — 3,068 chars of source]

Table (ref) displays the parameter estimates and the appertaining standard errors and $t$-test $p$-values for our MES regression using the inference methods developed in Section (ref). Also shown are the corresponding results of the ad hoc MES regression using standard inference methods for the OLS estimator. As in our simulations, we see that the parameter estimates and the standard errors of our MES regression and the ad hoc method based on a mean regression of the “empirically filtered” MES differ substantially. While the latter suggests a clearly significant influence of the bubble indicators in five out of six cases, our MES regression shows a different picture indicating that only the bust indicators have a (negative) significant effect for BAC and C, but not for JPM. Hence, given the functional of interest is the conditional MES, the ad hoc estimation method (likely) falsely classifies the boom indicators as being significant. Closely related, the substantially smaller standard errors of the ad hoc MES regression (caused by the OLS regression that omits the necessary “truncation” in the lower row of (ref)) are especially concerning in providing overconfident and possibly false conclusions on the drivers of systemic risk.

figure[figure omitted — 511 chars of source]

While MES regressions estimate the conditional MES as quantified in Theorem (ref), the ad hoc MES regression estimates the unconventional target functional of a conditional mean (through an OLS regression) of a rolling window MES estimate. A comparison of the respective model predictions in Figure (ref) shows that these can differ quite drastically, especially in turbulent times as in the year 2008: In contrast to our method, the ad hoc MES regression fails to quickly adapt to the rising levels of (systemic) risk due to its inherent dependence on the distant past through the rolling-window MES estimate in (ref). Hence, we suggest to clearly motivate the functional of interest and choose the appropriate methodological framework, which---when effects on the conditional MES are of interest---is provided by our regressions under adverse conditions.

The appropriateness of our (linear) regression under adverse conditions for the data is reinforced by in-sample model diagnostics. In classical correctly specified mean regressions, the expectation of the classical model residuals is zero conditional on known information, such as time or the model predictions. This is often analyzed by non-parametrically regressing the residuals on these quantities and plotting the results. In standard mean regressions, the classical model residuals $V(m,y)=m-y$ can be obtained as the derivative with respect to the model predictions $m$ of the (scaled) squared error loss $S(m,y)=\frac{1}{2}(m-y)^2$.

For our more complicated target functional (VaR, MES), we follow Pohle_GenResid and replace the model residuals by so-called “generalized model residuals”, whose conditional expectation---as above---equals zero for correctly specified models. The generalized model residuals correspond to the identification functions (known from forecast evaluation) evaluated at the model predictions and the outcomes. An identification function for the pair (VaR, MES) is given by $V\big( (v,m)', (x,y)' \big) = \big(\mathds{1}_{\{x\leq v\}}-\beta, \, \mathds{1}_{\{x>v\}}(m-y) \big)'$, which arises (almost everywhere) as the derivative of the lexicographical loss in (ref) with respect to $v$ and $m$ FH24. This construction parallels the derivative of the squared error mentioned above. In the following, we non-parametrically regress the generalized model residuals on time and the model predictions. Plots of the regression curve that deviate from the zero line are indicative of model misspecification (because then some combination of the covariates could predict the generalized residual---in violation of correct specification).

figure[figure omitted — 657 chars of source]

Figure (ref) shows generalized residuals of the VaR and MES parts of our regression, plotted against time and the model predictions, respectively. Also shown are the corresponding generalized MES residuals from the ad hoc MES regression. The blue line and its pointwise 95%-confidence band stem from a polynomial (mean) regression with automated parameter choices in the loess function of the statistical software R. Further details on the implementation are provided in Appendix (ref). The zero line is almost always contained in the blue band for the predictions from our MES regression, implying no evidence against model misspecification. In contrast, the ad hoc MES regression clearly violates this condition. Most troubling from a risk management perspective, the violations are most apparent during the financial crisis and for large model predictions, i.e., when accuracy of the method is most important.

Dissecting GDP Growth Vulnerabilities

In recent years, many researchers have quantified macroeconomic risks via the ES ABG19,Pea20,Aea21,DDP24. For instance, in their influential paper, ABG19 use the ES to measure US GDP growth vulnerability. Using an associated regression method, they then relate future ES to current macroeconomic and financial conditions by using US GDP growth and the US national financial conditions index (NFCI) as covariates.

In this section, we first introduce a slight (methodological) variation of ABG19's ABG19 procedure by estimating a joint VaR and ES regression. Then, we detail how our MES regressions can be used to gain additional insights on the decomposition of growth vulnerabilities in Europe. Here, our interest lies in both, understanding the driving forces of the conditional MES (see Table (ref)) and in providing an accurate model fit (see Figure (ref)).

Specifically, we denote by $X_t$ the negative quarterly year-on-year GDP growth of the economic area consisting of the three biggest European economies, Germany, France and the United Kingdom (UK). Formally, we estimate the parameters $\bm \theta^v = (\theta_1^v, \theta_2^v, \theta_3^v)^\prime$ and $\bm \theta^e = (\theta_1^e, \theta_2^e, \theta_3^e)^\prime$ of the linear (VaR, ES) regression

align[align omitted — 368 chars of source]

where $\operatorname{FCI}_{t-1}$ denotes the lagged “equally weighted” financial conditions index for advanced economies of Arrigoni2022 and $\mathcal{F}_t$ only contains time-$(t-1)$ information in the form of past $\operatorname{FCI}_{t-1}$ and $X_{t-1}$. ABG19 estimate the ES regression in (ref) by averaging over quantile regressions for multiple levels. In contrast, we employ the two-step M-estimator for the VaR and ES regression parameters that arises by setting $Y_t = X_t$ in (ref); cf. Remark (ref). In the numerical optimization, we additionally employ the non-crossing constraint that $\operatorname{ES}_{\beta}(X_t\mid \mathcal{F}_{t}) \ge \operatorname{VaR}_{\beta}(X_t\mid \mathcal{F}_{t})$ for all $t$.

table[table omitted — 1,444 chars of source]

The upper panel of Table (ref) (labeled “Joint Region”) shows the parameter estimates together with their standard errors. Notice that we do not report standard errors for the ES parameters in this panel, because no valid inference methods exist for a constrained two-step estimation of the conditional ES (as opposed to the unconstrained case discussed in Remark (ref)). The estimates are based on a sample of $n=99$ quarterly observations from 1995 Q1 until 2019 Q4, which (almost) corresponds to the available sample of the financial conditions index of Arrigoni2022. We find that the slope coefficients for both the VaR and ES model are positive and (for the VaR) statistically significant at any commonly used significance level. This confirms the result of ABG19 that future growth risk loads roughly equally on both financial and macroeconomic conditions (as captured by $\operatorname{FCI}_{t-1}$ and $X_{t-1}$ in (ref)--(ref)).

We now demonstrate how our MES regressions can be used to provide a more refined picture of growth vulnerabilities. Specifically, as we consider an entire economic region, the question naturally arises how the driving forces of growth vulnerabilities are distributed among the individual countries.

Formally, the common negative GDP growth $X_t = \sum_{d=1}^D w_{t,d} Y_{t,d}$ can be dissected into negative GDP growth, $Y_{t,d}$, of the $D=3$ individual countries. The time-varying weights $\bm w_t = (w_{t,1}, \dots, w_{t,D})^\prime$ are given by the countries' relative share of GDP in the entire economic region, such that $w_{t,d}\in(0,1)$ and $\sum_{d=1}^{D}w_{t,d}=1$. We partition the ES risk of $X_t$ from (ref) into its systemic MES components by writing

equation[equation omitted — 184 chars of source]

Then, we model the individual components of the right-hand side with our MES regression

align[align omitted — 184 chars of source]

jointly with the VaR regression in (ref). The conditional MES $\mathbb{E}_t[Y_{t,d}\mid X_t\geq\operatorname{VaR}_{t,\beta}]$ quantifies the expected (negative) GDP growth of country $d$ given that the economic region $X_t$ is in distress and, hence, measures the systemic vulnerability of economy $d$. The MES regression in (ref) therefore assesses how this systemic vulnerability is affected by past economic and financial conditions. By virtue of (ref), the three MES regressions in (ref) provide a complete dissection of the aggregated growth vulnerabilities captured by the ABG19 regression in (ref).

figure[figure omitted — 587 chars of source]

The lower panel of Table (ref) displays the MES parameter estimates together with their standard errors. Four out of six MES slope parameters, which equal the marginal effect under adverse conditions (see Remark (ref)), are significant at the $5\%$-level and we find the striking result that the explanatory power of the different variables varies strongly between the countries. This is in stark contrast to the classic (VaR, ES) regression of ABG19, where financial and economic conditions are equally important determinants of growth risk. For instance, the MES of the UK mainly loads on the financial conditions indicator, which may be explained by a high dependence on the financial sector of the UK with London as one of the most important financial hubs in the world. In contrast, the systemic vulnerability of the German economy is primarily driven by past economic conditions, which may be due to its strong export-oriented manufacturing sector.

For weights $\bm w_t $ being constant over time, the true ES regression parameters are given by the weighted average of the true MES regression parameters; see (ref). While this relationship is blurred in finite samples and due to the time-varying weights in this application (the weights for Germany vary between 0.4 and 0.45 and the ones for France and the UK between 0.26 and 0.31 in our sample), it can still be recognized in the parameter estimates in Table (ref). E.g., for the slope coefficient pertaining to $\operatorname{FCI}$, we have $0.663 \approx 0.7757 = 0.426 \cdot 0.352 + 0.291 \cdot 0.818 + 0.283 \cdot 1.370$, where 0.426, 0.291, and 0.283 are the average weights for Germany, France, and the UK.

Figure (ref) plots the implied VaR, ES and MES estimates for the regressions in (ref)--(ref) and (ref) for the joint economic region together with the three individual economies with solid lines. For comparison, the dashed lines show the implied VaR, ES and MES estimates of a four-dimensional DCC--GARCH model with Gaussian innovations Eng02. The display shows very plausible model fits of both methods, which is especially remarkable given that we estimate risk measure models in the tails with only $n=99$ observations.

figure[figure omitted — 594 chars of source]

Figure (ref) further illustrates the good in-sample model fit by plotting (VaR and MES) generalized model residuals (see Section (ref)) against time and the model predictions themselves. As the zero line is barely excluded from the confidence bands, we obtain (almost) no evidence against model misspecification for all models. However, due to the limited sample size, this was to be expected.

To summarize, while in the standard GaR literature growth vulnerabilities are the main object of study, our MES regressions allow to decompose these into systemic vulnerabilities of the different countries, therefore allowing a more fine-grained picture to emerge. We mention that our MES regressions may also be fruitfully used for single countries, such as the US. In such cases, one may decompose total growth into different economic sectors (e.g., manufacturing, service and agricultural) to shed more light on growth vulnerabilities in these different parts of the economy. We leave such explorations for future study.

Conclusion

Regressions typically model some functional of the scalar outcome $Y_t$ (say, the mean in mean regressions or quantiles in quantile regressions) as a function of covariates. In contrast, in this paper we model a functional of the bivariate outcome $(X_t,Y_t)^\prime$ (viz. the MES) as a function of covariates. This leads to what we term regressions under adverse conditions, where the mean of $Y_t$, given the adverse condition that a distress variable $X_t$ exceeds its conditional quantile, is related to covariates. We develop a two-step estimator and prove its asymptotic normality under general mixing conditions allowing for cross-sectional and time series applications, where we overcome the technical difficulty of having a discontinuous second-step objective function. We also propose feasible inference that works well as shown in our simulations. Our estimator particularly improves upon a recently proposed ad hoc estimation method of BRS20, BDP20, which we illustrate in simulations and an empirical application.

While our regression target---the conditional expectation of $Y_t$ given that $X_t$ exceeds its conditional quantile---is originally known as the MES in the (financial) systemic risk literature, we illustrate its value in three diverse applications focusing on the relation between systemic risk and asset price bubbles BRS20, dissecting GDP growth risks among countries ABG19, and in Appendix (ref) by analyzing portfolio risk allocation Bea18.

Moreover, we envision a multitude of further possible fields of applications such as in sensitivity analysis Asimit2019 or micro-econometric applications, where the MES could, e.g., be interpreted as the mean (personal or firm) income given the stress event of high inflation rates, or other challenging economic settings. We also refer to HTZ23, WTH24 and the references therein for the growing interest in quantile-truncated expectations in the micro-econometric literature.

A particular advantage of the MES as our regression target is its additivity CIM13 that---similar to our Section (ref)---allows for ES risk decompositions into different (additive) sub-components such as economic sectors. In a similar fashion, recently analyzed inflation risks LL24 could be decomposed into product categories. As such disaggregations are only possible by the additivity of the ES and MES risk measures as conditional expectations, these considerations provide further arguments for use of the ES over the VaR (and the MES over the CoVaR) in the recent debate of which risk measure is most suitable in practice EKT15, BCBS19, WangZitikis2021.

Supplementary Materials

The Supplementary Material contains the proof of Theorem (ref), and shows how to estimate the asymptotic variance-covariance matrix of Theorem (ref) consistently. It also verifies our main assumptions for the linear model of Example (ref) and provides an additional empirical application to portfolio optimization.

Replication material for the simulations and the applications in Section (ref) (together with the data) is available under \href{https://github.com/TimoDimi/replication_MES}{https://github.com/TimoDimi/replication_MES}. It draws on the corresponding open source package SystemicRisk package_SystemicRisk implemented in the statistical software R R2022.

Acknowledgments

We would like to thank the AE, two anonymous referees, Markus Brunnermeier, Tobias Fissler, Christoph Hanck, Simon Rother and Isabel Schnabel for their insightful comments that significantly improved the quality of the paper. We further thank Simon Rother for providing us with the bubble indicators used in BRS20.

Funding

The authors gratefully acknowledge support of the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) through grants 502572912 (Timo Dimitriadis), 460479886 and 531866675 (Yannick Hoga).