EconBase
← Back to paper

Regressions under Adverse Conditions

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

82,624 characters

Regressions under Adverse Conditions



\baselineskip18pt
\setcounter{totalnumber}{50}
\setcounter{topnumber}{50}
\setcounter{bottomnumber}{50}
\abovedisplayskip1.5ex plus1ex minus1ex
\belowdisplayskip1.5ex plus1ex minus1ex
\abovedisplayshortskip1.5ex plus1ex minus1ex
\belowdisplayshortskip1.5ex plus1ex minus1ex



\title{Regressions under Adverse Conditions
}


\author{
	Timo Dimitriadis\thanks{Faculty of Economics and Business, Goethe University Frankfurt, 60629 Frankfurt am Main, Germany, and Heidelberg Institute for Theoretical Studies (HITS), [email removed].}
\and
	Yannick Hoga\thanks{Faculty of Economics and Business Administration, University of Duisburg-Essen, Universit\"atsstra\ss e 12, D--45117 Essen, Germany, [email removed].}
}


 \date{\today}
\maketitle

\begin{abstract}
	\noindent
	\singlespacing
	We introduce a new regression method that relates the mean of an outcome variable to covariates, under the ``adverse condition'' that a distress variable falls in its tail.
	This allows to tailor classical mean regressions to adverse scenarios, which receive increasing interest in economics and finance, among many others.
	In the terminology of the systemic risk literature, our method can be interpreted as a regression for the Marginal Expected Shortfall.
	We propose a two-step procedure to estimate the new models, show consistency and asymptotic normality of the estimator, and propose feasible inference under weak conditions that allow for cross-sectional and time series applications.
	Simulations verify the accuracy of the asymptotic approximations of the two-step estimator.
	Two empirical applications show that our regressions under adverse conditions are a valuable tool in such diverse fields as the study of the relation between systemic risk and asset price bubbles, and dissecting macroeconomic growth vulnerabilities into individual components.\\


	\noindent \textbf{Keywords:} Estimation, Growth-at-Risk, Marginal Expected Shortfall, Regression Modeling, Systemic Risk \\
	\noindent \textbf{JEL classification:} C18, C51, C58, E44
\end{abstract}






\pagebreak
\section{Introduction}
\label{Motivation}

\onehalfspacing

In this paper, we add to the econometric toolbox by introducing \textit{regressions under adverse conditions}.
While the classical ordinary least squares (OLS) regression relates the mean of an outcome variable to covariates, our new regression relates the mean of an outcome variable to covariates only when some other (distress) variable falls in its tail.
Here, a variable falling in its tail means an exceedance of a large conditional quantile, which is a natural statistical interpretation of an adverse condition going back to \citet{KB78}.
Technically, our methodology can be interpreted as a regression for \citeauthor{Aea17}'s \citeyearpar{Aea17} systemic risk measure Marginal Expected Shortfall (MES), which is precisely defined as the conditional mean of some outcome variable given that a distress variable exceeds a large quantile.
Therefore, we use the terms \textit{regression under adverse conditions} and \textit{MES regression} interchangeably.

There are various applications where the relation between an outcome variable and explanatory variables is of particular interest if some distress variable experiences an extreme outcome.
First, in the systemic risk literature, the relationship between a bank's stock performance and bank characteristics is mainly of interest if the financial system is in distress \citep{AB16, Aea17, BRS20, BDP20}.
In this case, the bank's sensitivity to financial turmoil  (i.e., its systemic riskiness) is related to bank-level data, yielding insights into the determinants of systemic risk.
Our regressions under adverse conditions can deliver such insights.
Second, in the recent ``Growth-at-Risk'' literature \citep{ABG19,BS21,Aea22}, interest focuses on the risk to a country's gross domestic product (GDP) growth.
Our MES regressions allow growth risk---measured with the Expected Shortfall (ES)---to be dissected into different parts that can be attributed to specific sub-components (such as sub-regions), therefore contributing to the study of the interconnections of macroeconomic risks.
While we apply our regressions under adverse conditions to real data in these two fields in Section~\ref{sec:EmpiricalApplication}, we also explore some further uses in portfolio optimization in Appendix~\ref{sec:add appl} of the Online Supplement.

The three main contributions of this paper are as follows.
First, we propose a general modeling framework for the---typically dynamic---conditional MES.
Our models are flexible enough to incorporate possible non-linearities in the relationship between the MES and the regressors.
Our framework also allows for contemporaneous and/or lagged covariates, which may be serially dependent.
Second, we provide feasible inference tools for MES regressions, including parameter and asymptotic variance estimation.
This allows practitioners to identify the regressors with significant explanatory content. Third, we illustrate the usefulness of our regressions under adverse conditions in the three diverse contexts mentioned above.


From the technical side, standard (one-step) M-estimation of MES regression models is infeasible as there exists no \textit{scalar} consistent loss function that is uniquely minimized in expectation by the MES \citep{FH24, DFZ_CharMest}.
This contrasts with classical mean regressions, which can be estimated based on the squared error loss, as this loss is minimized (in expectation) by the mean.


To make up for this defect of the MES, we draw on \textit{bivariate} ``multi-objective'' consistent loss functions for the MES (together with the quantile) proposed by \citet[Theorem~4.2]{FH24}.
Following tradition in the risk management literature, we denote quantiles as the Value-at-Risk (VaR).
These bivariate loss functions lead to a \textit{two-step} M-estimator:
In the first step, we estimate the parameters of the VaR model for the distress variable by minimizing the quantile loss, exactly as in quantile regressions.
In the second step, we obtain parameter estimates of the MES model by minimizing the squared error loss only for those observations with a distress variable larger than the (estimated) VaR.
This event is typically called a ``VaR exceedance''.


In developing asymptotic theory for our two-step estimator, we draw on classic results of \citet{NeweyMcFadden1994} for consistency and on invariance principles of \citet{DMR95} for asymptotic normality.
Note that for asymptotic normality we cannot rely on standard results for two-step M-estimation \citep[as in][Sec.~6]{NeweyMcFadden1994} because our second-step objective function is discontinuous due to the truncation at the VaR.
As expected for two-step estimators, the estimation precision in the second step is influenced by the first-step estimator.
We further propose methods for feasible inference based on consistent estimation of the asymptotic variance-covariance matrix.

We illustrate the good finite-sample performance the two-step M-estimator and the associated inference methods in simulations, covering typical sample sizes and probability levels (for the quantile model) that we employ in our empirical applications.
We further compare our regression method with an \textit{ad hoc} estimator that is used in \citet{BRS20, BDP20}, \citet{Berger2020} and \citet{Karolyi2023} to estimate MES models and show that this method leads to inconsistent parameter estimates together with unrealistically small standard errors.


In Section~\ref{sec:EmpiricalApplication}, we apply our regressions under adverse conditions in the two above sketched fields of finding covariates associated with systemic risk, and in dissecting an economic region's GDP growth vulnerability into contributions of individual countries.
First, we reconsider an analysis similar to \citet[Table~7]{BRS20} to analyze the contemporaneous relation of asset price bubbles and the systemic riskiness (as measured by their MES) of the three systemically most relevant US banks.
In doing so, we compare our MES regression estimator with the ad hoc estimation method used in \citet{BRS20}.
Overall, we find that the MES is negatively affected by bust periods, but the effect of boom periods is not statistically significant, yielding---as in our simulations---different results compared to \citet{BRS20}.

Second, the recent ``Growth-at-Risk'' (GaR) literature, started by \citet{ABG19}, analyzes downside risks to future GDP growth as a function of current economic and financial conditions.
We apply the GaR methodology to the economic region consisting of Germany, France and the United Kingdom (UK)---the three largest economies in Europe.
Following \citet{ABG19}, we find that financial and economic conditions pose similar risks to future growth of the whole economic region.
Our regressions under adverse conditions now allow us to investigate how these total (economic and financial) risk factors can be attributed to the individual countries.
We obtain the striking finding that---in contrast to France and Germany---the risk to growth in the UK is mainly determined by \textit{financial} conditions, which may be explained by the importance of the financial marketplace London.
Similarly, we find that---in contrast to France and the UK---current \textit{economic} conditions in the joint economic region are the main risk to future German growth, which may be due to Germany's strong export-oriented manufacturing base.


Our regressions under adverse conditions are related to quantile regressions of \citet{KB78} in that they model a tail functional, and as the (preliminary) first part VaR regression is simply a (possibly nonlinear) quantile regression.
In that context we also refer to \citet{EM04} and \citet{CL23} for time series applications of quantile regressions, to \citet{Hog24} for MES forecasting models based on CCC--GARCH-type filters, and to \cite{DH24} for dynamic models for the related systemic risk measure CoVaR.
Our method is more closely related to recently developed regressions for the ES of \citet{DimiBayer2019}, \citet{PZC19}, \citet{GBP21}, \citet{Bar23+} and \citet{FMW23}.
Our and their methods share the property of relying on an ancillary quantile regression.
However, while ES regressions can be estimated jointly with VaR/quantile regressions (essentially due to the joint elicitability of the pair (VaR, ES) due to \citet{FZ16a}), our MES regressions require a two-step M-estimator (due to (VaR, MES) only being multi-objective elicitable in the sense of \citet{FH24}).
As the MES simplifies to the ES when the distress and outcome variables coincide, our MES regressions may be seen as a multivariate extension of ES regressions, which opens up many new fields of application, as outlined above and as demonstrated in more detail in the empirical applications.


Our MES regressions are further related to threshold models (or also: sample split models), which consider (possibly distinct) regression models for the two subsamples where some threshold variable is above or below a deterministic threshold parameter \citep{Han99,Han00,CH04}.
The main differences between our MES regressions and threshold models are as follows.
First, in the latter the threshold is a fixed scalar, whereas our models allow the threshold (i.e., the quantile of the distress variable) to be influenced by the user via the quantile level.
This flexibility is important, e.g., in the portfolio application, where the quantile level (most often 97.5\%) is predetermined by the regulator \citep{BCBS19}.
Second, while the threshold is deterministic in sample split models, we allow the stochastic threshold (i.e., the VaR) to be modeled dynamically.
Third, the models we consider are more general in that our framework covers non-linear relationships, whereas threshold regressions are mostly confined to linear models.


The remainder of the paper is structured as follows. Section~\ref{Modeling and Estimation} introduces MES regression models and our proposed two-step estimator. Its asymptotic properties and feasible inference methods are presented in Section~\ref{Asymptotic Properties}. The simulations in Section~\ref{sec:Simulations} verify the good finite-sample performance of our estimator.
Section~\ref{sec:EmpiricalApplication} demonstrates the usefulness of MES regressions in two empirical applications and the final Section~\ref{sec:Conclusion} concludes.
The proofs of all technical results are relegated to the Supplementary Material that contains the Appendices~\ref{sec:thm1}--\ref{sec:AddEmpResults}.





\section{Modeling and Estimation}
\label{Modeling and Estimation}

\subsection{The Definition of MES}

Consider some time series $\big\{\bm V_t=(\bm W_t^\prime, X_t,Y_t)^\prime\big\}_{t\in\mathbb{N}}$, where $\bm W_t$ are exogenous variables, $X_t$ denotes the distress or conditioning variable (e.g., a measure of macroeconomic or financial conditions) and $Y_t$ stands for the variable of interest (e.g., economic or financial losses).
Let $\mathcal{F}_{t}=\sigma(\bm W_{t}, \bm V_{t-1}, \bm V_{t-2}, \ldots)$ be the information set generated by the contemporaneous $\bm W_t$ and past $\bm V_t$.
The inclusion of $\bm W_t$ in $\mathcal{F}_t$ allows us to incorporate contemporaneous variables in MES regressions.
For instance, in studying the relationship between asset price bubbles and systemic risk, \citet{BRS20} use a contemporaneous bubble indicator in an MES-type regression; see also the empirical application in Section~\ref{sec:ApplMESReg}.


For $\beta\in(0,1)$ we define the VaR as the generalized inverse $\operatorname{VaR}_{t,\beta}:=\operatorname{VaR}_{\beta}(X_t\mid\mathcal{F}_t):=F_{X_t\mid\mathcal{F}_{t}}^{\leftarrow}(\beta)$, where $F_{X_t\mid\mathcal{F}_{t}}$ denotes the conditional cumulative distribution function (c.d.f.) of $X_t$ given $\mathcal{F}_t$.
The distress event that quantifies the ``adverse conditions'' in the definition of the MES is that the stress variable $X_t$ exceeds its VaR, i.e., $\{X_t\geq\operatorname{VaR}_{t,\beta}\}$.
With our orientation of $X_t$ denoting (economic or financial) losses, we commonly consider values for $\beta$ close to one.
Left-tail truncations in the sense of  $\{X_t\leq\operatorname{VaR}_{t,\beta}\}$ can be analyzed within our framework by simply considering $-X_t$.
We formally define our regression target under adverse conditions as $\operatorname{MES}_{t,\beta}:=\operatorname{MES}_{\beta}(Y_t\mid\mathcal{F}_{t}):=\mathbb{E}_t\big[Y_t \mid X_t\geq\operatorname{VaR}_{\beta}(X_t\mid\mathcal{F}_t)\big]$, where $\mathbb{E}_t[\cdot]:=\mathbb{E}[\ \cdot\mid\mathcal{F}_t]$ and we suppress the dependence of the MES on $X_t$ for brevity.




\subsection{Joint Models for VaR and MES}

Let $\bm Z_t=(Z_{1,t},\ldots,Z_{k,t})^\prime$ be the observable covariates affecting the VaR and the MES (i.e., $\operatorname{VaR}_{t,\beta}$ and $\operatorname{MES}_{t,\beta}$).
We assume that $\bm Z_t$ is $\mathcal{F}_{t}$-measurable, such that it can, e.g., be a subset of the contemporaneous $\bm W_t$ and past $\bm V_t$.
The empirical applications in Section~\ref{sec:EmpiricalApplication} provide some concrete examples for $\bm Z_t$.
We introduce the functions $v_t(\bm \theta^{v})=v(\bm Z_t;\, \bm \theta^v)$ and $m_t(\bm \theta^{m})=m(\bm Z_t;\, \bm \theta^m)$, where $\bm \theta^v$ ($\bm \theta^m$) is a generic vector from some parameter space $\bm \varTheta^v\subset\mathbb{R}^p$ ($\bm \varTheta^m\subset\mathbb{R}^q$), and $v(\,\cdot\,;\bm \theta^v)$ and $m(\,\cdot\,;\bm \theta^m)$ are measurable functions for all $\bm \theta^v\in\bm \varTheta^v$ and $\bm \theta^m\in\bm \varTheta^m$, respectively.
We assume that there exist unique true parameters $\bm \theta^v_0 \in \bm \varTheta^{v}$ and $\bm \theta^m_0 \in \bm \varTheta^{m}$, such that almost surely (a.s.)
\begin{align}
	\label{eqn:TrueModelParameters}
	\begin{pmatrix}
		\operatorname{VaR}_{\beta}(X_t\mid\mathcal{F}_t)\\
		\operatorname{MES}_{\beta}(Y_t\mid\mathcal{F}_{t})
	\end{pmatrix}
	= \begin{pmatrix}
		v_t(\bm \theta_{0}^{v})\\
		m_t(\bm \theta_{0}^{m})
	\end{pmatrix},\qquad t\in\mathbb{N}.
\end{align}
Thus, $v_t(\bm \theta_{0}^{v})$ is simply a (possibly non-linear) quantile regression model that relates the $\beta$-quantile of the distress variable $X_t$ to covariates $\bm Z_t$.
This may be seen as an auxiliary regression step that allows to identify the ``adverse condition'' $\{X_t\geq\operatorname{VaR}_{\beta}(X_t\mid\mathcal{F}_t)\}$ in the actual quantity of interest $\operatorname{MES}_{\beta}(Y_t\mid\mathcal{F}_{t})=\mathbb{E}_t\big[Y_t \mid X_t\geq\operatorname{VaR}_{t,\beta}\big]$, which is modeled via $m_t(\bm \theta_{0}^{m})$ and serves to explain the interrelation between $Y_t$ and $X_t$ (in the form of the MES) with covariates.
Therefore, we call the models $v_t(\bm \theta_{0}^{v})$ and $m_t(\bm \theta_{0}^{m})$ in \eqref{eqn:TrueModelParameters} a \textit{regression under adverse conditions} or also an \textit{MES regression model}.
The leading special case of \eqref{eqn:TrueModelParameters} is the linear model shown in  Example~\ref{ex:1} below, where $v_t(\bm \theta^{v})$ and $m_t(\bm \theta^{m})$ are linear in the parameters $\bm \theta^v$ and $\bm \theta^m$, respectively.



The fact that the conditional-on-$\mathcal{F}_t$ quantities $\operatorname{VaR}_{t,\beta}$ and $\operatorname{MES}_{t,\beta}$ only depend on $\bm Z_t$ is, of course, a modeling assumption.
Except for the restriction that $\bm Z_t$ be $\mathbb{R}^{k}$-valued, this is quite a flexible assumption. It covers purely predictive regressions (where $\bm Z_t$ only contains information available at time $t-h$, say) and also static regressions (where $\bm Z_t$ only contains contemporaneous variables) and any combination of the two.


Note that we consider the case of separated parameters in \eqref{eqn:TrueModelParameters}, where the parameters pertain only to the VaR model or the MES model.
This is vital for our two-step M-estimator introduced below.


We also mention that our framework inherently models the interconnectedness between $X_t$ and $Y_t$ through the (parameter-free) definition of the MES.
This contrasts with, e.g., \citet[Section III.B]{AB16}, whose estimation method (for the closely related CoVaR) can only capture linear effects of the VaR of $X_t$ onto $Y_t$.





\begin{example}\label{ex:1}
Here, we provide an example of a model that gives rise to a linear regression under adverse conditions in \eqref{eqn:TrueModelParameters}. Consider the model
\begin{align*}
	X_t &= \bm Z_t^{v\prime}\bm \theta_0^{v}+\varepsilon_{X,t},\\
	Y_t &= \bm Z_t^{m\prime}\bm \theta_0^{m}+\varepsilon_{Y,t},
\end{align*}
where $\bm Z_t^{v}$ and $\bm Z_t^{m}$ are ($p$ and $q$-dimensional) vectors containing some (or all) components of $\bm Z_t$,
and we assume that
$\operatorname{VaR}_{\beta}(\varepsilon_{X,t}\mid\mathcal{F}_t)=F_{\varepsilon_{X,t}\mid\mathcal{F}_t}^{\leftarrow}(\beta)=0$, and $\operatorname{MES}_{\beta}(\varepsilon_{Y,t}  \mid \mathcal{F}_t) = \mathbb{E}_t\big[\varepsilon_{Y,t}\mid \varepsilon_{X,t}\geq \operatorname{VaR}_{\beta}(\varepsilon_{X,t}\mid\mathcal{F}_t)\big]=0$.
Then, exploiting the $\mathcal{F}_t$-measurability of $\bm Z_t$, it is easy to check that \eqref{eqn:TrueModelParameters} holds with linear $v_t(\bm \theta_{0}^{v})=\bm Z_t^{v\prime}\bm \theta_0^{v}$ and linear $m_t(\bm \theta_{0}^{m})=\bm Z_t^{m\prime}\bm \theta_0^{m}$.
Recall that $\operatorname{VaR}_\beta(\varepsilon_{X,t}\mid\mathcal{F}_t)=0$ is the standard assumption on the errors in quantile regressions. The requirement that
$\operatorname{MES}_{\beta}(\varepsilon_{Y,t}  \mid \mathcal{F}_t) = 0$ may then be viewed as the analog assumption in our MES regressions.
Figure~\ref{fig:MESillu} graphically illustrates a linear regression under adverse conditions with one covariate.
The top panel shows the quantile regression line $v_t(\bm \theta_0^v)$ in blue and the bottom panel the MES regression line $m_t(\bm \theta_0^m)$ in red.
The black points in both panels indicate the observations $(X_t,Y_t)^\prime$ with a VaR exceedance, such that $X_t\geq v_t(\bm \theta_0^v)$.
It is obvious from this figure that the auxiliary quantile regression in the top plot is key in identifying the relevant points for the MES regression in the bottom plot.
\end{example}


\begin{figure}
	\centering
	\includegraphics[width=\linewidth]{input/illustration_MESreg}
	\caption{Top panel: Quantile regression line $v_t(\bm \theta_0^v)$ in blue.
	Bottom panel: MES regression line $m_t(\bm \theta_0^m)$ in red.
	Points in black in both panels correspond to observations with a VaR exceedance (i.e., $X_t\geq v_t(\bm \theta_0^v)$).}
	\label{fig:MESillu}
\end{figure}


While ``usual'' regressions simply relate some outcome $Y_t$ to covariates, our MES regressions allow us to investigate such a relationship \textit{conditional} on an extreme realization of $X_t$. This is of interest, e.g., when relating a bank's losses $Y_t$ to institution characteristics, where the precise relationship is mainly of interest during times of large system-wide losses $X_t$. We use the model of Example~\ref{ex:1} to investigate this in the empirical application in Section~\ref{sec:ApplMESReg}, but also in the other applications in Section~\ref{sec:GDP_MES} and Appendix~\ref{sec:add appl}.
We discuss the relation of our MES regressions to OLS regressions that use $X_t$ as an explanatory variable in Appendix~\ref{sec:RelationOLSXCovariate}.


\begin{rem}
	\label{rem:MarginalEffect}
	Using the notation $\bm Z_t^m = (Z^m_{1,t},\dots, Z^m_{q,t})^\prime$ and $\bm \theta^m = (\theta_1^m, \dots, \theta_q^m)^\prime$,
	our models in \eqref{eqn:TrueModelParameters} allow for the definition of a \emph{marginal effect under adverse conditions} with respect to $Z^m_{j,t}$ for any $j \in \{1,\dots,q\}$ via $\frac{\partial}{\partial Z^m_{j,t}} m_t(\bm \theta^m)$.
	This captures the impact of a unit change in $Z^m_{j,t}$ on the conditional MES given all other variables remain constant, which is analogous to the classical marginal effect for mean or quantile regressions.
	For the linear models of Example~\ref{ex:1}, the marginal effect under adverse conditions is simply $\theta^m_j$.
\end{rem}






\subsection{Parameter Estimation}
\label{sec:ParameterEstimation}

We now introduce our estimators of the unknown model parameters $\bm \theta_{0}^{v}$ and $\bm \theta_{0}^{m}$.
Typically, (one-step) M-estimation requires a to-be-minimized loss function (or also: objective function). Yet, as pointed out in the Introduction, there is no real-valued loss function associated with the pair (VaR, MES). The existence of such a (strictly consistent) loss function is, however, necessary for consistent M-estimation of semiparametric models such as ours \citep[Theorem 1]{DFZ_CharMest}.


To overcome this lack, we draw on \citet{FH24}.
Denoting by $\mathds{1}_{\{\cdot\}}$ the indicator function, they show that (under some smoothness conditions) the $\mathbb{R}^2$-valued loss function
\begin{equation}\label{eq:loss}
	\bm S\Bigg(\begin{pmatrix}v\\ m\end{pmatrix}, \begin{pmatrix}x\\ y\end{pmatrix}\Bigg) = \begin{pmatrix}S^{\operatorname{VaR}}(v,x)\\ S^{\operatorname{MES}}\big((v,m)^\prime, (x,y)^\prime\big) \end{pmatrix}=
	\begin{pmatrix}
		[\mathds{1}_{\{x\leq v\}}-\beta][v-x]\\
		\frac{1}{2}\mathds{1}_{\{x>v\}}[y-m]^2
	\end{pmatrix}
\end{equation}
is minimized in expectation by the true VaR and MES with respect to the lexicographic order. That is, for all $v,m\in\mathbb{R}$,
\[
	\mathbb{E}\Bigg[\bm S\Bigg(\begin{pmatrix}\operatorname{VaR}_\beta\\ \operatorname{MES}_\beta\end{pmatrix}, \begin{pmatrix}X\\ Y\end{pmatrix}\Bigg)\Bigg]\preceq_{\operatorname{lex}}\mathbb{E}\Bigg[\bm S\Bigg(\begin{pmatrix}v\\ m\end{pmatrix}, \begin{pmatrix}X\\ Y\end{pmatrix}\Bigg)\Bigg],
\]
where $(x_1,x_2)^\prime\preceq_{\operatorname{lex}} (y_1,y_2)^\prime$ if $x_1<y_1$ or ($x_1=y_1$ and $x_2\leq y_2$), and $\operatorname{VaR}_\beta$ and $\operatorname{MES}_\beta$ denote the true VaR and MES of the joint distribution of $(X,Y)^\prime$. Note that $S^{\operatorname{VaR}}(\cdot,\cdot)$ in \eqref{eq:loss} is the standard pinball-loss function known from quantile regression. Clearly, $S^{\operatorname{MES}}(\cdot,\cdot)$ resembles the squared error loss familiar from usual mean regressions, except that the indicator $\mathds{1}_{\{x>v\}}$ restricts the evaluation to observations with VaR exceedances in the first component.
The pre-factor $1/2$ is simply a convenient normalization, enabling an easier representation of some subsequent quantities.
In the related literature, loss functions are often also called \textit{scoring functions} \citep{Gne11} and we use these two terms interchangeably.


To introduce our two-step estimator, suppose that we have a sample $\big\{(X_t,Y_t,\bm Z_t^\prime)^\prime\big\}_{t=1,\ldots,n}$ of size $n$.
In a first step, the lexicographic order suggests to minimize the expected score of the VaR component. Approximating the expected score by the sample average of the scores, the first-step estimator is
\[
	\widehat{\bm \theta}_{n}^{v} = \operatorname*{arg\,min}_{\bm \theta^{v}\in\bm \varTheta^{v}}\frac{1}{n}\sum_{t=1}^{n}S^{\operatorname{VaR}}\big(v_t(\bm \theta^{v}), X_t\big).
\]
\citet{OH16} study this estimator in their nonlinear quantile regressions for deterministic regressors.


Given $\widehat{\bm \theta}_n^v$, the lexicographic order similarly suggests in a second step to minimize the sample average of the MES losses via
\begin{align}
	\label{eq:MES_estimator}
	\widehat{\bm \theta}_{n}^{m} = \operatorname*{arg\,min}_{\bm \theta^{m}\in\bm \varTheta^{m}}\frac{1}{n}\sum_{t=1}^{n}S^{\operatorname{MES}}\Big(\big(v_t(\widehat{\bm \theta}_{n}^{v}), m_t(\bm \theta^{m})\big)^\prime, \big(X_t, Y_t\big)^\prime\Big).
\end{align}
For this two-step estimator to be feasible, the requirement that the VaR model does not depend on the MES model is essential. Section~\ref{Asymptotic Properties} shows that the estimation precision of $\widehat{\bm \theta}_{n}^{m}$ is affected by relying on $\widehat{\bm \theta}_n^v$ in the second step, which is the usual case for two-step estimators \citep[Section 6]{NeweyMcFadden1994}.


\section{Asymptotics for the Two-Step Estimators}\label{Asymptotic Properties}


We now show the joint asymptotic normality of our two-step M-estimators $\widehat{\bm \theta}_{n}^{v}$ and $\widehat{\bm \theta}_{n}^{m}$ under the regularity conditions in Assumptions~\ref{ass:1}--\ref{ass:7} given below. Before presenting Assumptions~\ref{ass:1}--\ref{ass:7}, we have to introduce some notation. The joint c.d.f.~of $(X_{t}, Y_t)^\prime\mid\mathcal{F}_t$ is denoted by $F_t(\cdot,\cdot)$, and its Lebesgue density (which we assume exists) by $f_t(\cdot,\cdot)$. Similarly, $F_t^{W}(\cdot)$ ($f_t^{W}(\cdot)$) denotes the distribution (density) function of $W_t\mid\mathcal{F}_t$ for $W\in\{X,Y\}$. For sufficiently smooth functions $\mathbb{R}^{p}\ni\bm \theta\mapsto f(\bm \theta) \in \mathbb{R}$, we denote the $(p\times1)$-gradient by $\nabla f(\bm \theta)$, its transpose by $\nabla' f(\bm \theta)$ and the $(p\times p)$-Hessian by $\nabla^2 f(\bm \theta)$.
All (in-)equalities involving random quantities are meant to hold almost surely.




The asymptotic variance-covariance matrices of our estimators depend on the following matrices:
\begin{align*}
	\bm V&=\beta(1-\beta)\mathbb{E}\big[\nabla v_t(\bm \theta_0^v)\nabla^\prime v_t(\bm \theta_0^v)\big]\in\mathbb{R}^{p\times p},\\
	\bm \varLambda & =\mathbb{E}\Big[f_t^{X}\big(v_t(\bm \theta_0^v)\big)\nabla v_t(\bm \theta_0^v)\nabla^\prime v_t(\bm \theta_0^v)\Big]\in\mathbb{R}^{p\times p},\\
	\bm M^{\ast}&= \mathbb{E}\Big[\big\{Y_t -m_t(\bm \theta_0^m)\big\}^2 \mathds{1}_{\{X_t>v_t(\bm \theta_0^v)\}} \nabla m_t(\bm \theta_0^m) \nabla^\prime m_t(\bm \theta_0^m)\Big]\in\mathbb{R}^{q\times q},\\
	\bm \varLambda_{(1)} &= (1-\beta)\mathbb{E}\big[\nabla m_t(\bm \theta_0^m)\nabla^\prime m_t(\bm \theta_0^m)\big]\in\mathbb{R}^{q\times q},\\
	\bm \varLambda_{(2)} &=\mathbb{E}\bigg[\Big\{\int_{-\infty}^{\infty} y f_t\big(v_t(\bm \theta_0^v),y\big)\,\mathrm{d} y - m_t(\bm \theta_0^m)f_{t}^{X}\big(v_t(\bm \theta_0^v)\big)\Big\}\nabla m_t(\bm \theta_0^m)\nabla^\prime v_t(\bm \theta_0^v)\bigg]\in\mathbb{R}^{q\times p}.
\end{align*}
All these matrices exist by virtue of Assumptions~\ref{ass:1}--\ref{ass:7}, which we state now. For this, let $K<\infty$ be some large universal positive constant, and denote by $\norm{\bm x}$ the Euclidean norm when $\bm x$ is a vector, and the Frobenius norm when $\bm x$ is matrix-valued.




\begin{assumption}\label{ass:1}
\begin{enumerate}
\item\label{it:1i} There exist unique parameters $\bm \theta_0^v$ and $\bm \theta_0^m$, such that \eqref{eqn:TrueModelParameters} holds a.s.~for all $t\in\mathbb{N}$.

\item\label{it:1ii}
The parameter space $\bm \varTheta=\bm \varTheta^{v}\times\bm \varTheta^{m}$ is compact, where $\bm \varTheta^{v}\subset\mathbb{R}^p$ and $\bm \varTheta^{m}\subset\mathbb{R}^q$.

\item\label{it:1iii}
$\bm \theta_0^v\in\operatorname{int}(\bm \varTheta^v)$ and $\bm \theta_{0}^{m}\in\operatorname{int}(\bm \varTheta^m)$, where $\operatorname{int}(\cdot)$ denotes the interior of a set.

\end{enumerate}
\end{assumption}


\begin{assumption}\label{ass:2}
$\big\{(X_t, Y_t, \bm Z_t^\prime)^\prime\big\}_{t\in\mathbb{N}}$ is strictly stationary and $\beta$-mixing with coefficients $\beta(\cdot)$ of size $-r/(r-1)$ for some $r>1$.
\end{assumption}

\begin{assumption}\label{ass:3}
\begin{enumerate}
	\item\label{it:3i} For all $t\in\mathbb{N}$, $F_t(\cdot,\cdot)$ belongs to a class of distributions on $\mathbb{R}^2$ that possess a positive Lebesgue density $f_t(x,y)$ for all $(x,y)^\prime\in\mathbb{R}^2$ such that $F_t(x,y)\in(0,1)$.


	\item\label{it:3ii} $\big|f_t^{X}(x)-f_t^{X}(x^\prime)\big|\leq K|x-x^\prime|$ and $\sup_{x\in\mathbb{R}}f_t^{X}(x)\leq K$.

	\item\label{it:3iii} The conditional density $f_t(x,y)$ is differentiable in its first argument, with partial derivative denoted by $\partial_1 f_t(x,y)$.

	\item\label{it:3iv} $\sup_{x\in\mathbb{R}}\big|\int_{-\infty}^{\infty}y\partial_1 f_t(x,y)\,\mathrm{d} y\big|\leq F(\mathcal{F}_t)$ and $\sup_{x\in\mathbb{R}}\int_{-\infty}^{\infty}|y| f_t(x,y)\,\mathrm{d} y\leq F_1(\mathcal{F}_t)$ for $\mathcal{F}_t$-measurable random variables $F(\mathcal{F}_t)$ and $F_1(\mathcal{F}_t)$.

\end{enumerate}
\end{assumption}


\begin{assumption}\label{ass:4}
\begin{enumerate}
\item\label{it:4i} For all $t\in\mathbb{N}$, $v_t(\cdot)$ and $m_t(\cdot)$ are twice differentiable on $\operatorname{int}(\bm \varTheta^v)$ and $\operatorname{int}(\bm \varTheta^m)$ with gradients $\nabla v_t(\cdot)$ and $\nabla m_t(\cdot)$, and Hessians $\nabla^2 v_t(\cdot)$ and $\nabla^2 m_t(\cdot)$.

\item\label{it:4ii} $\sup_{\bm \theta^v\in\bm \varTheta^v}\big|v_t(\bm \theta^v)\big|\leq V(\bm Z_t)$ and $\sup_{\bm \theta^m\in\bm \varTheta^m}\big|m_t(\bm \theta^m)\big|\leq M(\bm Z_t)$ for some measurable functions $V(\cdot)$ and $M(\cdot)$.

\item\label{it:4iii} $\sup_{\bm \theta^v\in\bm \varTheta^v}\big\Vert\nabla v_t(\bm \theta^v)\big\Vert\leq V_1(\bm Z_t)$ and $\sup_{\bm \theta^m\in\bm \varTheta^m}\big\Vert\nabla m_t(\bm \theta^m)\big\Vert\leq M_1(\bm Z_t)$ for some measurable functions $V_1(\cdot)$ and $M_1(\cdot)$.

\item\label{it:4iv} $\sup_{\bm \theta^v\in\bm \varTheta^v}\norm{\nabla^2v_t(\bm \theta^v)}\leq V_2(\bm Z_t)$ and $\sup_{\bm \theta^m\in\bm \varTheta^m}\norm{\nabla^2m_t(\bm \theta^m)}\leq M_2(\bm Z_t)$ for some measurable functions $V_2(\cdot)$ and $M_2(\cdot)$.

\item\label{it:4v} There exist neighborhoods of $\bm \theta_0^v$ and $\bm \theta_0^m$, such that $\norm{\nabla^2 v_t(\bm \tau^v) - \nabla^2 v_t(\bm \theta^v)}\leq V_3(\bm Z_t)\norm{\bm \tau^v-\bm \theta^v}$ and $\norm{\nabla^2 m_t(\bm \tau^m) - \nabla^2 m_t(\bm \theta^m)}\leq M_3(\bm Z_t)\norm{\bm \tau^m-\bm \theta^m}$ for all elements $\bm \theta^v$, $\bm \tau^v$ and $\bm \theta^m$, $\bm \tau^m$ of the respective neighborhoods, and measurable functions $V_3(\cdot)$ and $M_3(\cdot)$.

\end{enumerate}
\end{assumption}

\begin{assumption}\label{ass:5}
 For $r>1$ from Assumption~\ref{ass:2} and some $\iota>0$, it holds that $\mathbb{E}\big[V(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[V_1^{4r}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[V_2^{2r}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[V_3(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[M^{4r+\iota}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[M_1^{4r+\iota}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[M_2^{4r}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[M_3^{4r/(4r-1)}(\bm Z_t)\big]\leq K$, $\mathbb{E}\big[F^{4r/(4r-3)}(\mathcal{F}_t)\big]\leq K$, $\mathbb{E}\big[F_1^{4r/(4r-3)}(\mathcal{F}_t)\big]\leq K$, $\mathbb{E}|X_t|\leq K$, $\mathbb{E}|Y_t|^{4r+\iota}\leq K$.
\end{assumption}

\begin{assumption}\label{ass:6}
The matrices  $\bm \varLambda$ and $\bm \varLambda_{(1)}$ are positive definite.
\end{assumption}


\begin{assumption}\label{ass:7}
$\sup_{\bm \theta^v\in\bm \varTheta^v}\sum_{t=1}^{n}\mathds{1}_{\{X_t=v_t(\bm \theta^v)\}}\leq K$ a.s.~for all $n\in\mathbb{N}$.
\end{assumption}

Assumption~\ref{ass:1}~\ref{it:1i} ensures identification of the true parameters.
Compactness in Assumption~\ref{ass:1}~\ref{it:1ii} is a standard requirement in extremum estimation; see \citet{NeweyMcFadden1994}.
Assumption~\ref{ass:1}~\ref{it:1iii} forces the true parameters to be interior to the respective parameter spaces, which is a key requirement for proving asymptotic normality. Assumption~\ref{ass:2} is a standard stationarity and mixing condition ensuring that suitable laws of large numbers and central limit theorems apply.
Often, the $\beta$-mixing coefficients decrease exponentially, such that Assumption~\ref{ass:2} holds for \textit{any} $r>1$ arbitrarily close to unity; see, e.g., \citet{CC02} and \citet{Lie05} for ARMA and GARCH processes.
Assumption~\ref{ass:3}~\ref{it:3i} ensures strict (multi-objective) consistency of the scoring function given in \eqref{eq:loss}; see \citet[Theorem~4.2]{FH24}.
The final items \ref{it:3ii}--\ref{it:3iv} of Assumption~\ref{ass:3} may be interpreted as smoothness conditions for the conditional density.
Since the VaR and MES only depend on $\bm Z_t$ under our modeling framework, it will most often be the case that the conditional density $f_t(\cdot,\cdot)$ will also be $\bm Z_t$-measurable; see, e.g., Appendix~\ref{sec:ex verif}. Nonetheless, the random variables $F(\mathcal{F}_{t})$ and $F_1(\mathcal{F}_{t})$ appearing in Assumption~\ref{ass:3}~\ref{it:3iv} are in principle allowed to be $\mathcal{F}_{t}$-measurable.
Assumption~\ref{ass:4} provides smoothness conditions on the VaR and MES model, and
Assumption~\ref{ass:5} collects moment bounds.
When mixing is exponential (corresponding to $r=1$ in Assumption~\ref{ass:2}), the moment bounds of Assumption~\ref{ass:5} are weakest, such that there is the usual moment-memory tradeoff. Note that the moment conditions for the VaR model are weaker than those for the MES model, because the latter is estimated based on a \textit{squared} error-type loss.
Assumption~\ref{ass:6} ensures that the asymptotic variance-covariance matrix in Theorem~\ref{thm:an} is well-defined, which can however be degenerate if $\bm V$ or $\bm M^\ast$ are singular; also see \citet[footnote 20]{Hansen:82}.
Assumption~\ref{ass:7} is identical to Assumption~2~(G) in \citet{PZC19}. It prevents an infinite number of $X_t$'s to lie on any possible ``quantile regression line'' $v_t(\bm \theta^v)$. Appendix~\ref{sec:ex verif} shows that $K=\dim(\bm \theta^v)$ for linear models.


\begin{thm}\label{thm:an}
Suppose Assumptions~\ref{ass:1}--\ref{ass:7} hold. Then, as $n\to\infty$,
\begin{align*}
\sqrt{n}\begin{pmatrix}
\widehat{\bm \theta}_n^{v}-\bm \theta_0^v\\
\widehat{\bm \theta}_n^{m}-\bm \theta_0^m
\end{pmatrix}
\overset{d}{\longrightarrow}N\big(\boldsymbol{0}, \bm \varGamma \bm M \bm \varGamma^\prime),
\end{align*}
where
\[
\bm \varGamma =\begin{pmatrix}
	\bm \varLambda^{-1} & \boldsymbol{0}\\
	-\bm \varLambda_{(1)}^{-1}\bm \varLambda_{(2)}\bm \varLambda^{-1} & \bm \varLambda_{(1)}^{-1}\end{pmatrix}\in\mathbb{R}^{(p+q) \times (p+q)}\qquad\text{and}\qquad \bm M=\begin{pmatrix}\bm V &\boldsymbol{0} \\ \boldsymbol{0} & \bm M^\ast\end{pmatrix}\in\mathbb{R}^{(p+q)\times(p+q)}.
\]
\end{thm}

\sloppy
The proof of Theorem~\ref{thm:an} is in Appendix~\ref{sec:thm1} and draws on \citet{DLS23}.
In the special case $\bm \varLambda_{(2)}=\boldsymbol{0}$, the first-step estimator does not have an impact on the asymptotic variance of the second-step estimator \citep[Section 6.2]{NeweyMcFadden1994}.
To gain some intuition for this, rewrite the curly bracket in  $\bm \varLambda_{(2)}$ as $f_t^X(v_t(\bm \theta_0^v))  \left( \mathbb{E} \big[ Y_t \big| X_t = v_t(\bm \theta_0^v) \big] - \mathbb{E} \big[ Y_t \big| X_t \ge v_t(\bm \theta_0^v) \big] \right)$, which measures how the expectation of $Y_t$ reacts upon moving from $\{X_t = v_t(\bm \theta_0^v)\}$ further to the tail through $\{ X_t \ge v_t(\bm \theta_0^v)\}$.
Hence, $\bm \varLambda_{(2)}=\boldsymbol{0}$ can be interpreted as a form of ``conditional tail uncorrelatedness'' of $X_t$ and $Y_t$, as $\bm \varLambda_{(2)}$ is a weighted expectation of the above quantity.



\begin{rem}[Expected Shortfall Regressions]
	\label{rem:ESreg}
	For $X_t = Y_t$, the definition of the MES simplifies to the Expected Shortfall, $\operatorname{ES}_{t,\beta} := \operatorname{ES}_{\beta}(X_t\mid\mathcal{F}_t) := \mathbb{E}_t\big[X_t \mid X_t \geq \operatorname{VaR}_{t,\beta}\big]$.
	For the ES, associated regression models (jointly with the conditional VaR) have been proposed recently.
	E.g., \citet{DimiBayer2019} and \citet{PZC19} consider \emph{joint} M-estimation with a VaR model, which is feasible as (VaR, ES) are \emph{jointly elicitable}, as opposed to the pair (VaR, MES), which is merely multi-objective elicitable.
	One can also use a two-step estimator for VaR and ES models, whose second step is essentially given by setting $Y_t = X_t$ in the right-hand side of \eqref{eq:MES_estimator}.
	Its asymptotic variance-covariance matrix resembles the one presented in Theorem \ref{thm:an} by again setting $Y_t = X_t$ in all matrices but $\bm \varLambda_{(2)}$, where an ES-specific version is
	\begin{align*}
		\bm \varLambda_{(2)}^\text{ES} =\mathbb{E}\Big[ \big\{v_t(\bm \theta_0^v)- m_t(\bm \theta_0^m) \big\} f_{t}^{X}\big(v_t(\bm \theta_0^v)\big) \nabla m_t(\bm \theta_0^m)\nabla^\prime v_t(\bm \theta_0^v)\Big].
	\end{align*}
	This arises as $\int_{-\infty}^{\infty} y f_t\big(v_t(\bm \theta_0^v),y\big) \,\mathrm{d} y =  f_{t}^{X}\big(v_t(\bm \theta_0^v)\big) \mathbb{E}_t \big[ Y_t \mid X_t = v_t(\bm \theta^v_0) \big]$, and the latter conditional expectation equals $v_t(\bm \theta^v_0)$ if $Y_t = X_t$.
	Notice, however, that the joint distribution of $(Y_t, X_t)^\prime$ is degenerate if $Y_t = X_t$, such that our Assumption \ref{ass:3} is violated.
	Therefore, Assumption \ref{ass:3} is replaced in ES models by the requirement that the (univariate) conditional distribution of $X_t$ is sufficiently smooth around its $\beta$-quantile; see, e.g., \citet[Assumption 2 (B)]{PZC19}.
\end{rem}


For feasible inference, we have to estimate the asymptotic variance-covariance matrix of Theorem~\ref{thm:an} consistently.
In Appendix~\ref{sec:thm3}, we propose suitable estimators and show their consistency in Theorem~\ref{thm:lrv} under the additional Assumptions~\ref{ass:9}--\ref{ass:8}.

Appendix~\ref{sec:ex verif} verifies Assumptions~\ref{ass:1}--\ref{ass:7} and also Assumptions~\ref{ass:9}--\ref{ass:11} for the linear MES regression model of Example~\ref{ex:1} subject to some regularity conditions.





\section{Simulations}
\label{sec:Simulations}


\subsection{The Data-Generating Process}

We consider a time series regression as a data-generating process (DGP), which exhibits the desirable features of a  time-varying conditional mean and covariance matrix and at the same time results in a correctly specified regression for the MES (and the VaR) that allows for a comparison with the estimation used in \cite{BRS20, BDP20} in Section \ref{sec:BrunnermeierMESReg}.

Using the notation of Example \ref{ex:1}, the covariates $\bm Z^v_{t} = \bm Z^m_{t} =(1, Z_{1,t}, Z_{2,t})^\prime$ follow the (transformed) auto-regressions
\begin{align*}
	Z_{1,t} &= 0.3 + 0.4 \exp(\xi_t)  \qquad \text{with } \qquad \xi_t = 0.6 \xi_{t-1} + \varepsilon_{\xi,t}, \\
	Z_{2,t} &= 0.75 Z_{2,t-1} + \varepsilon_{Z,t},
\end{align*}
where $\varepsilon_{Z,t}, \varepsilon_{\xi,t} \stackrel{\text{i.i.d.}}{\sim} {N}(0,1)$, independently of each other for all $t=1,\ldots,n$.
The exponential transformation guarantees positivity of $Z_{1,t}$, which is required to introduce heteroskedasticity in the process
\begin{align}
	\label{eqn:DGPCrossSectional}
	\big(X_t, Y_t \big)'
	&= (\gamma_1,\gamma_1)^\prime + (\gamma_2, \gamma_2)^\prime  Z_{1,t} + (\gamma_3,\gamma_3)^\prime  Z_{2,t} + \big(\gamma_4 + \gamma_5  Z_{1,t}  \big) \boldsymbol{\varepsilon}_t, \qquad t=1,\ldots,n.
\end{align}
The multivariate $t$-distributed innovations $\boldsymbol{\varepsilon}_t =(\varepsilon_{1,t}, \varepsilon_{2,t})^\prime\stackrel{\text{i.i.d.}}{\sim} t_{6} \big( \boldsymbol{0}, \boldsymbol{\Sigma} =  \begin{psmallmatrix}
	1 & 1.2 \\
	1.2 & 4
\end{psmallmatrix}\big)$
are independent of the $\{\varepsilon_{Z,t}\}$ and $\{\varepsilon_{\xi,t}\}$.
We choose $\bm \gamma=(\gamma_1, \ldots, \gamma_5)^\prime =(1,\ 1.5,\ 2,\ 0.25,\ 0.5)^\prime$ as the true values.
For simplicity, our DGP in \eqref{eqn:DGPCrossSectional} implies that $X_t$ and $Y_t$ are driven by the same factors and only the heteroskedastic shock differs between the variables.
This DGP results in the linear (VaR, MES) model
\begin{align*}
	v_t(\bm \theta^v) = \theta^v_1  + \theta^v_2 Z_{1,t} +  \theta^v_3 Z_{2,t}
	\qquad \text{ and } \qquad
	m_t(\bm \theta^m) = \theta^m_1  + \theta^m_2 Z_{1,t} +  \theta^m_3 Z_{2,t}.
\end{align*}
The true parameter values $\bm \theta_0^v=(\theta_{1,0}^v,\, \theta_{2,0}^v,\, \theta_{3,0}^v)^\prime$ and $\bm \theta_0^m=(\theta_{1,0}^m,\, \theta_{2,0}^m,\, \theta_{3,0}^m)^\prime$ are given by
\begin{align*}
	\bm \theta^v_{0} =
	\big(\gamma_1 + \gamma_4 \tilde q_\beta, \;
	\gamma_2 + \gamma_5 \tilde q_\beta,  \;
	\gamma_3 \big)^\prime
	\qquad \text{and} \qquad
	\bm \theta^m_{0} =
	\big(\gamma_1 + \gamma_4 \tilde m_\beta, \;
	\gamma_2 + \gamma_5 \tilde m_\beta, \;
	\gamma_3  \big)^\prime,
\end{align*}
where $\tilde m_\beta$ is the $\beta$-MES of the $t_{6} \big( \boldsymbol{0}, \boldsymbol{\Sigma} \big)$-distribution, and $\tilde q_\beta$ is the $\beta$-quantile of its first component.
While the \emph{true} parameters for $\theta_{3,0}^v$ and  $\theta_{3,0}^m$ coincide, this specification does not violate the separated parameter condition in \eqref{eqn:TrueModelParameters} as we do not impose the condition $\theta_{3}^v = \theta_{3}^m$ in the estimation.







\subsection{Simulation Results}

We estimate the correctly specified six-parameter (VaR, MES) regression model based on the joint covariates $\bm Z^v_{t} = \bm Z^m_{t}$ given above.
Table~\ref{tab:SimResultsCrossSectionalReg} shows simulation results for the estimated parameters and their standard deviations.
Results are based on $5000$ Monte Carlo replications of the DGP in \eqref{eqn:DGPCrossSectional} for the probability levels  $\beta \in \{0.9,\ 0.95,\ 0.975\}$, which we also use in our two applications in Sections \ref{sec:ApplMESReg}--\ref{sec:GDP_MES}.
We consider the typical sample sizes $n \in \{500,\ 1000,\ 2000,\ 4000\}$.
The asymptotic variances are estimated as detailed in Appendix~\ref{sec:thm3}.
A formal description of the table columns is given in the table caption.


\begin{table}[tb]
	\centering
	\scriptsize
	\begin{tabular}{rr lrrrr lrrrr lrrrr}
		\toprule
		\multicolumn{2}{l}{\textbf{VaR}} & &  \multicolumn{4}{c}{$\theta_1^v$} & & \multicolumn{4}{c}{$\theta_2^v$} & & \multicolumn{4}{c}{$\theta_3^v$}  \\
		\cmidrule{4-7}  	\cmidrule{9-12} 	 \cmidrule{14-17}
		$\beta$ & $n$ & &
		Bias & $\hat \sigma_\text{emp}$ & $\hat \sigma_\text{asy}$ & CI  & &
		Bias & $\hat \sigma_\text{emp}$ & $\hat \sigma_\text{asy}$ & CI  & &
		Bias & $\hat \sigma_\text{emp}$ & $\hat \sigma_\text{asy}$ & CI   \\
		\midrule
		\multirow{4}{*}{0.9}  & 500 &  & 0.049 & 0.581 & 0.585 & 0.92 &  & $-$0.033 & 0.489 & 0.503 & 0.91 &  & $-$0.003 & 0.163 & 0.159 & 0.93 \\
		& 1000 &  & 0.023 & 0.410 & 0.409 & 0.93 &  & $-$0.015 & 0.346 & 0.352 & 0.92 &  & $-$0.000 & 0.114 & 0.112 & 0.94 \\
		& 2000 &  & 0.012 & 0.288 & 0.287 & 0.93 &  & $-$0.007 & 0.244 & 0.247 & 0.93 &  & $-$0.000 & 0.081 & 0.079 & 0.94 \\
		& 4000 &  & 0.009 & 0.208 & 0.203 & 0.93 &  & $-$0.007 & 0.176 & 0.174 & 0.94 &  & $-$0.001 & 0.057 & 0.056 & 0.94 \\
		\addlinespace
		\multirow{4}{*}{0.95} & 500 &  & 0.104 & 0.824 & 0.850 & 0.89 &  & $-$0.073 & 0.687 & 0.730 & 0.87 &  & $-$0.001 & 0.232 & 0.230 & 0.91 \\
		& 1000 &  & 0.051 & 0.593 & 0.586 & 0.90 &  & $-$0.031 & 0.500 & 0.504 & 0.88 &  & 0.000 & 0.164 & 0.159 & 0.91 \\
		& 2000 &  & 0.034 & 0.412 & 0.410 & 0.92 &  & $-$0.024 & 0.346 & 0.353 & 0.91 &  & $-$0.001 & 0.115 & 0.112 & 0.93 \\
		& 4000 &  & 0.015 & 0.294 & 0.290 & 0.93 &  & $-$0.008 & 0.248 & 0.249 & 0.93 &  & 0.000 & 0.079 & 0.079 & 0.94 \\
		\addlinespace
		\multirow{4}{*}{0.975}& 500 &  & 0.214 & 1.210 & 1.425 & 0.87 &  & $-$0.130 & 1.007 & 1.179 & 0.82 &  & $-$0.001 & 0.338 & 0.377 & 0.89 \\
		& 1000 &  & 0.108 & 0.860 & 0.936 & 0.88 &  & $-$0.064 & 0.720 & 0.797 & 0.85 &  & 0.003 & 0.236 & 0.249 & 0.90 \\
		& 2000 &  & 0.061 & 0.604 & 0.626 & 0.90 &  & $-$0.035 & 0.516 & 0.536 & 0.88 &  & $-$0.003 & 0.169 & 0.169 & 0.92 \\
		& 4000 &  & 0.025 & 0.427 & 0.437 & 0.91 &  & $-$0.015 & 0.361 & 0.375 & 0.90 &  & 0.001 & 0.117 & 0.118 & 0.93 \\
		\midrule
		\midrule
		\\
		\multicolumn{2}{l}{\textbf{MES}} & &  \multicolumn{4}{c}{$\theta_1^m$} & & \multicolumn{4}{c}{$\theta_2^m$} & & \multicolumn{4}{c}{$\theta_3^m$}  \\
		\cmidrule{4-7}  	\cmidrule{9-12} 	 \cmidrule{14-17}
		$\beta$ & $n$ & &
		Bias & $\hat \sigma_\text{emp}$ & $\hat \sigma_\text{asy}$ & CI  & &
		Bias & $\hat \sigma_\text{emp}$ & $\hat \sigma_\text{asy}$ & CI  & &
		Bias & $\hat \sigma_\text{emp}$ & $\hat \sigma_\text{asy}$ & CI   \\
		\midrule
		\multirow{4}{*}{0.9}& 500 &  & 0.030 & 0.822 & 0.763 & 0.89 &  & $-$0.031 & 0.629 & 0.575 & 0.87 &  & 0.003 & 0.207 & 0.201 & 0.93 \\
		& 1000 &  & 0.018 & 0.614 & 0.573 & 0.90 &  & $-$0.018 & 0.466 & 0.433 & 0.88 &  & 0.005 & 0.144 & 0.141 & 0.94 \\
		& 2000 &  & 0.010 & 0.454 & 0.434 & 0.92 &  & $-$0.009 & 0.343 & 0.327 & 0.91 &  & $-$0.001 & 0.099 & 0.100 & 0.95 \\
		& 4000 &  & 0.006 & 0.325 & 0.317 & 0.93 &  & $-$0.006 & 0.243 & 0.239 & 0.93 &  & $-$0.000 & 0.070 & 0.071 & 0.95 \\
		\addlinespace
		\multirow{4}{*}{0.95}& 500 &  & 0.045 & 1.211 & 1.100 & 0.86 &  & $-$0.060 & 0.935 & 0.815 & 0.81 &  & $-$0.002 & 0.323 & 0.303 & 0.91 \\
		& 1000 &  & 0.039 & 0.906 & 0.849 & 0.88 &  & $-$0.035 & 0.698 & 0.640 & 0.86 &  & 0.001 & 0.220 & 0.218 & 0.94 \\
		& 2000 &  & 0.029 & 0.657 & 0.656 & 0.91 &  & $-$0.027 & 0.500 & 0.494 & 0.90 &  & $-$0.002 & 0.158 & 0.155 & 0.94 \\
		& 4000 &  & 0.010 & 0.493 & 0.487 & 0.92 &  & $-$0.010 & 0.373 & 0.366 & 0.92 &  & 0.001 & 0.112 & 0.111 & 0.94 \\
		\addlinespace
		\multirow{4}{*}{0.975} & 500 &  & 0.073 & 1.891 & 1.762 & 0.84 &  & $-$0.099 & 1.443 & 1.261 & 0.76 &  & $-$0.001 & 0.543 & 0.478 & 0.88 \\
		& 1000 &  & 0.062 & 1.334 & 1.256 & 0.87 &  & $-$0.061 & 1.019 & 0.931 & 0.82 &  & $-$0.005 & 0.360 & 0.336 & 0.90 \\
		& 2000 &  & 0.016 & 0.967 & 0.930 & 0.89 &  & $-$0.026 & 0.742 & 0.702 & 0.87 &  & $-$0.003 & 0.239 & 0.241 & 0.94 \\
		& 4000 &  & 0.028 & 0.741 & 0.734 & 0.91 &  & $-$0.026 & 0.565 & 0.552 & 0.90 &  & $-$0.004 & 0.174 & 0.173 & 0.94 \\
		\bottomrule
	\end{tabular}
	\caption{Simulation results for the parameter estimates of a (correctly specified) linear MES regression model based on the DGP in \eqref{eqn:DGPCrossSectional} and $5000$ simulation replications.
		The columns ``Bias'' show the average bias of the parameter estimates,  ``$\hat \sigma_\text{emp}$'' report the empirical standard deviation of the parameter estimates, ``$\hat \sigma_\text{asy}$'' the mean of the estimated standard deviations (computed from the diagonal elements of our estimate of $\bm \varGamma\bm M\bm \varGamma^\prime$) and the columns ``CI'' show the coverage rates of $95\%$-confidence intervals.}
	\label{tab:SimResultsCrossSectionalReg}
\end{table}




We find that both the VaR and MES parameters are estimated consistently as the empirical bias and the empirical standard deviation shrink as $n$ grows in all settings.
As expected, the estimates are more accurate for moderate (smaller) probability levels $\beta$, but even the largest choice $\beta = 0.975$ leads to fairly accurate estimates in large samples.
Overall, the MES parameters are estimated with similar accuracy as the corresponding VaR parameters, and so are their asymptotic variances.
These results are confirmed by the coverage rates of the resulting confidence intervals, where the empirical coverage rates approach the nominal rate of $95\%$.
We observe some undercoverage in finite samples that, however, vanishes with decreasing $\beta$ and increasing sample size $n$.




\subsection{Comparison with Existing Approaches to MES Regressions}
\label{sec:BrunnermeierMESReg}

\cite{BRS20, BDP20}, \citet{Berger2020} and \citet{Karolyi2023} among others use the following approach to MES regressions.
Define the empirical MES estimate over a rolling window of $S \in \mathbb{N}$ (typically, $S=250$) past observations, i.e.,
\begin{align}
	\label{eq:Y_MES_transform}
	Y_t^\ast = 	\bigg(\sum_{s=t-S}^{t} \mathds{1}_{\{ X_s \ge \widehat{Q}_\beta(X_{(t-S):t})\}} \bigg)^{-1}  \sum_{s=t-S}^{t}  Y_s \mathds{1}_{\{ X_s \ge \widehat{Q}_\beta(X_{(t-S):t}) \}},
\end{align}
where $\widehat{Q}_\beta(X_{(t-S):t})$ denotes the sample quantile of $\{X_{t-S},\dots,X_t\}$ in the rolling window.
\cite{BRS20, BDP20} then relate the transformed response $Y_t^\ast$ to covariates $\bm Z_t^m$ in a standard OLS mean regression; see \citet[equations (2) and (3)]{BRS20} and \citet[Section 2.3]{BDP20} for details.
We henceforth call this method the ``ad hoc MES regression''.
A mean regression of $Y_t^\ast$ on $\bm Z_t^m$ estimates the conditional expectation $ \mathbb{E} \big[ Y_t^\ast  \mid \bm Z_{t}^m \big]$, whose interpretation is unclear and which is in general different from the \textit{de facto} target of an MES regression, that is, $\operatorname{MES}_{t,\beta}=\mathbb{E}\big[Y_t \mid X_t\geq\operatorname{VaR}_{t,\beta}, \bm Z_{t}^m \big]$.
A further drawback of the ad hoc MES regressions is that---in contrast to our MES regression---the covariates $\bm Z_t^m$ cannot contain past values of $X_t$ or $Y_t$ as these would influence both the left-hand and right-hand sides of the associated OLS regression, as $Y_t^\ast$ is a rolling average over past values of $Y_t$ (truncated by using lagged values of $X_t$).

To illustrate in a simpler context, such an ad hoc regression for the $\beta$-quantile would relate the moving window sample quantile $Y_t^{\dagger}  = \widehat{Q}_\beta(X_{(t-S):t})$ to covariates.
This contrasts with quantile regression of \citet{KB78}, which relates $Y_t$ directly to the regressors by using the quantile-specific check loss function given in the first row of \eqref{eq:loss}.


\begin{figure}
	\centering
	\includegraphics[width=\linewidth]{input/ParamHist}
	\caption{Histograms (over the $M=5000$ simulation replications) of our MES regression parameter estimates in blue and the estimates using the technique of \cite{BRS20,BDP20} in red.
	We fix $\beta = 0.95$ and $n=4000$.
	The true values are marked by the vertical black lines.}
	\label{fig:ParamHist_Brunnermeier}
\end{figure}

Figure \ref{fig:ParamHist_Brunnermeier} illustrates how severe the discrepancy between our and the ad hoc version of the MES regression is for the DGP in \eqref{eqn:DGPCrossSectional} for $\beta = 0.95$.
We do so by estimating a mean regression of the target $Y_t^\ast$ on the same covariates $(1,Z_{1,t}, Z_{2,t})'$ and by computing our MES estimate $\widehat{\bm \theta}_n^m$.
While the parameter estimates of our MES regression are normally distributed around the true parameter values (given by the black lines), the estimates based on the ad hoc estimation method are far from the true values.
Most strikingly, the ad hoc estimates of the two slope parameters associated with the covariates $Z_{1,t}$ and $Z_{2,t}$ vary around zero because the contemporaneous effect of the covariates on the \textit{rolling window} quantity $Y_t^\ast$ is negligible.
This illustrates that employing the ad hoc MES regression procedure does not capture the effect of covariates on the MES functional, but rather for some functional that is inherently hard to interpret; see \eqref{eq:Y_MES_transform}.



\section{Empirical Applications}
\label{sec:EmpiricalApplication}

We illustrate the versatility of our MES regressions in three applications.
These concern explanatory variables of systemic risk in the banking sector in Section~\ref{sec:ApplMESReg} and dissecting GDP growth vulnerabilities among the three biggest European economies in Section~\ref{sec:GDP_MES}.
A further application concerning (equal risk contribution) portfolios is deferred to Appendix~\ref{sec:add appl}.




\subsection{Systemic Risk Regressions}
\label{sec:ApplMESReg}


With the introduction of RiskMetrics \citep{RM96}, financial risk management---based mainly on the VaR---was beginning to be firmly established in the 1990s.
Subsequently, the importance of adequately managing individual financial risks was also reflected in official regulations by the \citet{BCBS96}.
As these regulations did not prevent the financial crisis of 2007--08, more attention (of regulators as well as industry) has been paid to the \textit{systemic} nature of financial risks. This means that more scrutiny was applied in studying the interconnectedness of individual institutions (in the context of the financial system as a whole) or the interconnectedness of individual trading desks (in the context of managing bank-wide risks).
As a consequence of this development, a huge literature on systemic risk and its determinants has emerged \citep[e.g.,][]{ABT12,AB16,Aea17,BE17,BRS20,BDP20}.

For instance, \citet{BRS20} investigate the contemporaneous relation of asset price bubbles and systemic risk by, among others, considering the (conditional) MES in their Section 4.2.
However, our simulations in Section \ref{sec:BrunnermeierMESReg} show that the ad hoc estimation method for MES regressions---as applied by the aforementioned authors---does not necessarily provide consistent parameter estimates nor allows for valid inference.
Therefore, we now consider a simplified form of their analysis for the conditional MES by using our regressions under adverse conditions.
Here, we mainly aim at identifying which covariates drive the conditional MES, and as in  \citet{BRS20}, we particularly focus on stock market boom and bust indicators.


Specifically, we focus on the three US banks that are classified as systemically most risky  according to the \citet{FSB22} such that $Y_t$ equals the daily log-losses  (i.e., the negative log-returns) of either the Bank of America Corporation (BAC), Citigroup (C) or JPMorgan Chase (JPM).
We provide corresponding results for the other five US banks listed as a global systemically important bank (G-SIB) by the \citet{FSB22} in Appendix~\ref{sec:AddEmpResults}.
Throughout, we use the daily log-losses of the S\&P~500 Financials for $X_t$ and set $\beta = 0.95$.
Hence, the MES measures the mean loss of the bank \textit{conditional} on the financial system being in distress.


We estimate the joint linear model  $ \big( v_t(\bm \theta^v), m_t(\bm \theta^m) \big)' = \big( \bm Z_{t}^{v\,\prime} \bm \theta^v, \bm Z_{t}^{m\,\prime} \bm \theta^m\big)'$, where $ \bm \theta^v,  \bm \theta^m \in \mathbb{R}^6$. For reasons of data availability and concerns of non-stationarity of some regressors, we use a restricted set of variables in $\bm Z_{t}^{v} = \bm Z_{t}^{m} = \bm Z_t = (1, Z_{t,1},\ldots,Z_{t,5})^\prime$,
which contain an intercept, the change in spread between Moody's Baa-rated bonds and the 10-year Treasury bill rate (\textit{Change Spread}), the spread between the 3-month LIBOR and 3-month Treasury bills (\textit{TED Spread}), the VIX index as a forward-looking measure of market volatility, and separate indicator variables for stock market booms (\textit{SM Boom}) and busts (\textit{SM Bust}) from \citet{BRS20}.
To study the contemporaneous relation between systemic risk and bubbles, we follow \citet[Eq.~(3)]{BRS20} by lagging all explanatory variables by one time period, except for the boom and bust indicators for which we use contemporaneous values.



To allow for an immediate comparison of the estimation methods, we also estimate the ``ad hoc MES regression'' described in Section~\ref{sec:BrunnermeierMESReg}, which is a classical OLS  regression of the transformed target variable defined in \eqref{eq:Y_MES_transform}.
We estimate the models based on daily data and use the largest common available sample ranging from May 5, 1993 until December 31, 2015, yielding a total of $n=5,555$ trading days.




\begin{table}[tb]
	\centering
	\footnotesize
	\resizebox{\columnwidth}{!}{
	\begin{tabular}{c c l c rrr c rrr c rrr}
		\toprule
		&&&& \multicolumn{3}{c}{VaR Model} &&  \multicolumn{3}{c}{MES Model} &&  \multicolumn{3}{c}{Ad Hoc MES Regression} \\
		\cmidrule{5-7}   	\cmidrule{9-11}   	\cmidrule{13-15}
		Bank & & Covariate && Est. & SE & $p$-val  &&  Est. & SE & $p$-val &&  Est. & SE & $p$-val   \\
		\midrule
		\addlinespace
		\multirow{6}{*}{BAC}  &  & Intercept &  & $-$1.734 & 0.185 & 0.000 &  & $-$5.535 & 1.438 & 0.000 &  & $-$1.325 & 0.156 & 0.000 \\
		 &  & Change Spread &  & 0.591 & 0.107 & 0.000 &  & 2.184 & 0.752 & 0.004 &  & 1.965 & 0.090 & 0.000 \\
		 &  & TED Spread &  & 1.751 & 0.182 & 0.000 &  & 2.985 & 1.027 & 0.004 &  & $-$1.865 & 0.124 & 0.000 \\
		 &  & VIX &  & 0.092 & 0.015 & 0.000 &  & 0.147 & 0.058 & 0.011 &  & 0.068 & 0.009 & 0.000 \\
		 &  & ST Boom &  & $-$0.194 & 0.125 & 0.120 &  & 0.140 & 0.611 & 0.819 &  & 0.566 & 0.131 & 0.000 \\
		 &  & ST Bust &  & $-$0.571 & 0.169 & 0.001 &  & $-$1.483 & 0.604 & 0.014 &  & $-$0.280 & 0.181 & 0.121 \\
		\midrule
		\addlinespace
		\multirow{6}{*}{C}  &  & Intercept &  & $-$1.734 & 0.185 & 0.000 &  & $-$4.654 & 1.880 & 0.013 &  & $-$1.386 & 0.120 & 0.000 \\
		 &  & Change Spread &  & 0.591 & 0.107 & 0.000 &  & 1.141 & 0.920 & 0.215 &  & 1.791 & 0.070 & 0.000 \\
		 &  & TED Spread &  & 1.751 & 0.182 & 0.000 &  & 2.110 & 1.106 & 0.057 &  & $-$1.642 & 0.095 & 0.000 \\
		 &  & VIX &  & 0.092 & 0.015 & 0.000 &  & 0.263 & 0.063 & 0.000 &  & 0.092 & 0.007 & 0.000 \\
		 &  & ST Boom &  & $-$0.194 & 0.125 & 0.120 &  & $-$0.321 & 0.635 & 0.613 &  & 0.687 & 0.101 & 0.000 \\
		 &  & ST Bust &  & $-$0.571 & 0.169 & 0.001 &  & $-$1.476 & 0.608 & 0.015 &  & $-$0.328 & 0.139 & 0.018 \\
		\midrule
		\addlinespace
		\multirow{6}{*}{JPM} &  & Intercept &  & $-$1.734 & 0.185 & 0.000 &  & $-$2.601 & 0.873 & 0.003 &  & $-$0.234 & 0.084 & 0.005 \\
		 &  & Change Spread &  & 0.591 & 0.107 & 0.000 &  & 0.868 & 0.431 & 0.044 &  & 1.208 & 0.049 & 0.000 \\
		 &  & TED Spread &  & 1.751 & 0.182 & 0.000 &  & 2.455 & 0.927 & 0.008 &  & $-$1.668 & 0.067 & 0.000 \\
		 &  & VIX &  & 0.092 & 0.015 & 0.000 &  & 0.151 & 0.049 & 0.002 &  & 0.090 & 0.005 & 0.000 \\
		 &  & ST Boom &  & $-$0.194 & 0.125 & 0.120 &  & $-$0.151 & 0.444 & 0.734 &  & 0.422 & 0.070 & 0.000 \\
		 &  & ST Bust &  & $-$0.571 & 0.169 & 0.001 &  & $-$0.025 & 0.713 & 0.972 &  & 0.821 & 0.097 & 0.000 \\
		\bottomrule
	\end{tabular}
	}
	\caption{VaR and MES regression results for $\beta = 0.95$ for $X_t$ being the log-losses of the S\&P~500 Financials and $Y_t$ being the log-losses of Bank of America Corporation (BAC), Citigroup (C) and JPMorgan Chase (JPM) in the respective panels. The vertical panel entitled ``Ad Hoc MES Regression'' reports the results of the method described around  \eqref{eq:Y_MES_transform}. The respective parameter estimates are given in the columns ``Est.'', their standard errors in the columns ``SE'', and associated $t$-test $p$-values in the columns ``$p$-val''.}
	\label{tab:PredictiveRegression}
\end{table}



Table~\ref{tab:PredictiveRegression} displays the parameter estimates and the appertaining standard errors and $t$-test $p$-values for our MES regression using the inference methods developed in Section \ref{Asymptotic Properties}. Also shown are the corresponding results of the ad hoc MES regression using standard inference methods for the OLS estimator.
As in our simulations, we see that the parameter estimates and the standard errors of our MES regression and the ad hoc method based on a mean regression of the ``empirically filtered'' MES differ substantially.
While the latter suggests a clearly significant influence of the bubble indicators in five out of six cases, our MES regression shows a different picture indicating that only the bust indicators have a (negative) significant effect for BAC and C, but not for JPM.
Hence, given the functional of interest is the conditional MES, the ad hoc estimation method (likely) falsely classifies the boom indicators as being significant.
Closely related, the substantially smaller standard errors of the ad hoc MES regression (caused by the OLS regression that omits the necessary ``truncation'' in the lower row of \eqref{eq:loss}) are especially concerning in providing overconfident and possibly false conclusions on the drivers of systemic risk.



\begin{figure}[tb]
	\centering
	\includegraphics[width=\linewidth]{input/MES_BRS_regressions}
	\caption{Estimated sample paths of our MES regression and the ``ad hoc MES regression'' employed in \citet{BRS20} for $Y_t$ being negative returns of the Bank of America (BAC), $X_t$  negative returns of the S\&P\,500 Financials and $\beta = 0.95$.
	Returns on days with an exceedance of the estimated conditional VaR are displayed in black while all other returns are shown in gray.}
	\label{fig:MES_BRS_regressions}
\end{figure}


While MES regressions estimate the conditional MES as quantified in Theorem~\ref{thm:an}, the ad hoc MES regression estimates the unconventional target functional of a conditional mean (through an OLS regression) of a rolling window MES estimate.
A comparison of the respective model predictions in Figure~\ref{fig:MES_BRS_regressions} shows that these can differ quite drastically, especially in turbulent times as in the year 2008:
In contrast to our method, the ad hoc MES regression fails to quickly adapt to the rising levels of (systemic) risk due to its inherent dependence on the distant past through the rolling-window MES estimate in \eqref{eq:Y_MES_transform}.
Hence, we suggest to clearly motivate the functional of interest and choose the appropriate methodological framework, which---when effects on the conditional MES are of interest---is provided by our regressions under adverse conditions.


The appropriateness of our (linear) regression under adverse conditions for the data is reinforced by in-sample model diagnostics.
In classical correctly specified \emph{mean} regressions, the expectation of the classical model residuals is zero conditional on known information, such as time or the model predictions.
This is often analyzed by non-parametrically regressing the residuals on these quantities and plotting the results.
In standard mean regressions, the classical model residuals $V(m,y)=m-y$ can be obtained as the derivative with respect to the model predictions $m$ of the (scaled) squared error loss $S(m,y)=\frac{1}{2}(m-y)^2$.



For our more complicated target functional (VaR, MES), we follow \citet{Pohle_GenResid} and replace the model residuals by so-called ``generalized model residuals'', whose conditional expectation---as above---equals zero for correctly specified models.
The generalized model residuals correspond to the identification functions (known from forecast evaluation) evaluated at the model predictions and the outcomes.
An identification function for the pair (VaR, MES) is given by $V\big( (v,m)', (x,y)' \big) = \big(\mathds{1}_{\{x\leq v\}}-\beta, \, \mathds{1}_{\{x>v\}}(m-y) \big)'$, which arises (almost everywhere) as the derivative of the lexicographical loss in \eqref{eq:loss} with respect to $v$ and $m$ \citep{FH24}.
This construction parallels the derivative of the squared error mentioned above.
In the following, we non-parametrically regress the generalized model residuals on time and the model predictions.
Plots of the regression curve that deviate from the zero line are indicative of model misspecification (because then some combination of the covariates could predict the generalized residual---in violation of correct specification).


\begin{figure}[tb]
	\centering
	\includegraphics[width=\linewidth]{input/SysRisk_Diagnostics_Date.pdf}
	\includegraphics[width=\linewidth]{input/SysRisk_Diagnostics_Predictions.pdf}
	\caption{In-sample model diagnostics of the VaR and MES predictions of our regression and the ``ad hoc MES regression'' of \citet{BRS20, BDP20} for $Y_t$ being negative returns of the Bank of America (BAC), $X_t$  negative returns of the S\&P\,500 Financials and $\beta = 0.95$.
	The generalized residuals are plotted and non-parametrically regressed (in blue) against time (upper row) and the model predictions (lower row).}
	\label{fig:ModelDiagnostics_MES_BRS_regressions}
\end{figure}


Figure~\ref{fig:ModelDiagnostics_MES_BRS_regressions} shows generalized residuals of the VaR and MES parts of our regression, plotted against time and the model predictions, respectively.
Also shown are the corresponding generalized MES residuals from the ad hoc MES regression.
The blue line and its pointwise 95\%-confidence band stem from a polynomial (mean) regression with automated parameter choices in the \texttt{loess} function of the statistical software \texttt{R}.
Further details on the implementation are provided in Appendix~\ref{sec:AddEmpResults}.
The zero line is almost always contained in the blue band for the predictions from our MES regression, implying no evidence against model misspecification.
In contrast, the ad hoc MES regression clearly violates this condition.
Most troubling from a risk management perspective, the violations are most apparent during the financial crisis and for large model predictions, i.e., when accuracy of the method is most important.




\subsection{Dissecting GDP Growth Vulnerabilities}
\label{sec:GDP_MES}

In recent years, many researchers have quantified macroeconomic risks via the ES \citep{ABG19,Pea20,Aea21,DDP24}.
For instance, in their influential paper, \citet[Section II.B]{ABG19} use the ES to measure US GDP growth vulnerability.
Using an associated regression method, they then relate future ES to current macroeconomic and financial conditions by using US GDP growth and the US national financial conditions index (NFCI) as covariates.

In this section, we first introduce a slight (methodological) variation of \citeauthor{ABG19}'s \citeyearpar{ABG19} procedure by estimating a joint VaR and ES regression.
Then, we detail how our MES regressions can be used to gain additional insights on the decomposition of growth vulnerabilities in Europe.
Here, our interest lies in both, understanding the driving forces of the conditional MES (see Table \ref{tab:GDP_MES}) and in providing an accurate model fit (see Figure~\ref{fig:MES_GDP}).

Specifically, we denote by $X_t$ the negative  quarterly year-on-year GDP growth of the economic area consisting of the three biggest European economies, Germany, France and the United Kingdom (UK).
Formally, we estimate the parameters  $\bm \theta^v = (\theta_1^v, \theta_2^v, \theta_3^v)^\prime$ and  $\bm \theta^e = (\theta_1^e, \theta_2^e, \theta_3^e)^\prime$ of the linear (VaR, ES) regression
\begin{align}
		\operatorname{VaR}_{\beta}(X_t \mid \mathcal{F}_{t}) &= \theta_1^v + \theta_2^v \operatorname{FCI}_{t-1} + \theta_3^v X_{t-1}, \label{eq:VaR GaR}\\
		\operatorname{ES}_{\beta}(X_t \mid \mathcal{F}_{t}) = \mathbb{E}_t [X_t\mid X_t\geq\operatorname{VaR}_{t,\beta}] &= \theta_1^e + \theta_2^e \operatorname{FCI}_{t-1} + \theta_3^e X_{t-1},\label{eq:ES GaR}
\end{align}
where $\operatorname{FCI}_{t-1}$ denotes the lagged ``equally weighted'' financial conditions index for advanced economies of \citet{Arrigoni2022} and $\mathcal{F}_t$ only contains time-$(t-1)$ information in the form of past $\operatorname{FCI}_{t-1}$ and $X_{t-1}$.
\citet[Section II.B]{ABG19} estimate the ES regression in \eqref{eq:ES GaR} by averaging over quantile regressions for multiple levels.
In contrast, we employ the two-step M-estimator for the VaR and ES regression parameters that arises by setting $Y_t = X_t$ in \eqref{eq:MES_estimator}; cf.~Remark \ref{rem:ESreg}.
In the numerical optimization, we additionally employ the \emph{non-crossing constraint} that $\operatorname{ES}_{\beta}(X_t\mid \mathcal{F}_{t}) \ge \operatorname{VaR}_{\beta}(X_t\mid \mathcal{F}_{t})$ for all $t$.


\begin{table}[tb]
	\centering
	\begin{tabular}{ll c rr c rr c rr}
		\toprule
		&&& \multicolumn{2}{c}{$\theta_{1}$} && \multicolumn{2}{c}{$\theta_{2}$} && \multicolumn{2}{c}{$\theta_{3}$} \\
		\cmidrule{4-5}    \cmidrule{7-8}    \cmidrule{10-11}
		Country & Risk Measure &&   Est. & SE &&  Est. & SE  &&  Est. & SE  \\
		\midrule
		\multirow{2}{*}{Joint Region} & VaR && 0.602 & 0.212  &&   0.673 & 	0.097  && 0.884 & 0.100 \\
		& ES	&& 0.794 &  && 0.663 &  && 0.856 &  \\
		\midrule
		Germany 	& MES	 &&  2.160 & 0.261 && 0.352 & 0.160 && 1.350     &   0.077 \\
		France 			& MES	&&  $-0.557$ & 0.222 && 0.818  & 0.226  && 0.315    &  0.211 \\
		United Kingdom 	 & MES	&&   $-0.723$ & 0.335 && 1.370  & 0.312 && 0.122     &  0.255 \\
		\bottomrule
	\end{tabular}
	\caption{VaR, ES and MES regression parameter estimates at the level $\beta=0.9$ for the models in \eqref{eq:VaR GaR}--\eqref{eq:ES GaR} and \eqref{eq:GDPmodel} for $n=99$ observations of negative quarterly GDP growth between  1995 Q1 and 2019 Q4 for Germany, France, the United Kingdom (as $Y_{t,d}$) and their joint economic region (as $X_t$). The columns ``Est.'' report the parameter estimates and ``SE'' the estimated standard errors.
	The generic parameter names $\theta_{j}$, $j=1,2,3$ in the column headings refer to $\theta^v_{j}$, $\theta^e_{j}$ and $\theta^m_{j,d}$ from \eqref{eq:VaR GaR}, \eqref{eq:ES GaR} and \eqref{eq:GDPmodel}, respectively.}
	\label{tab:GDP_MES}
\end{table}

The upper panel of Table~\ref{tab:GDP_MES} (labeled ``Joint Region'') shows the parameter estimates together with their standard errors.
Notice that we do not report standard errors for the ES parameters in this panel, because no valid inference methods exist for a constrained two-step estimation of the conditional ES (as opposed to the unconstrained case discussed in Remark~\ref{rem:ESreg}).
The estimates are based on a sample of $n=99$ quarterly observations from 1995 Q1 until 2019 Q4, which (almost) corresponds to the available sample of the financial conditions index of \citet{Arrigoni2022}.
We find that the slope coefficients for both the VaR and ES model are positive and (for the VaR) statistically significant  at any commonly used significance level.
This confirms the result of \citet{ABG19} that future growth risk loads roughly equally on both financial and macroeconomic conditions (as captured by $\operatorname{FCI}_{t-1}$ and $X_{t-1}$ in \eqref{eq:VaR GaR}--\eqref{eq:ES GaR}).



We now demonstrate how our MES regressions can be used to provide a more refined picture of growth vulnerabilities.
Specifically, as we consider an entire economic region, the question naturally arises how the driving forces of growth vulnerabilities are distributed among the individual countries.


Formally, the common negative GDP growth $X_t = \sum_{d=1}^D w_{t,d} Y_{t,d}$ can be dissected into negative GDP growth, $Y_{t,d}$, of the $D=3$ individual countries.
The time-varying weights $\bm w_t = (w_{t,1}, \dots, w_{t,D})^\prime$ are given by the countries' relative share of GDP in the entire economic region, such that $w_{t,d}\in(0,1)$ and $\sum_{d=1}^{D}w_{t,d}=1$.
We partition the ES risk of $X_t$ from \eqref{eq:ES GaR} into its systemic MES components by writing
\begin{equation}\label{eq:ES decomp}
\mathbb{E}_t[X_t\mid X_t\geq\operatorname{VaR}_{t,\beta}] = \sum_{d=1}^{D} w_{t,d} \, \mathbb{E}_t[Y_{t,d}\mid X_t\geq\operatorname{VaR}_{t,\beta}].
\end{equation}
Then, we model the individual components of the right-hand side with our MES regression
\begin{align}
	\label{eq:GDPmodel}
	\mathbb{E}_t [Y_{t,d}\mid X_t\geq\operatorname{VaR}_{t,\beta}] &= \theta_{1,d}^m + \theta_{2,d}^m  \operatorname{FCI}_{t-1} + \theta_{3,d}^m X_{t-1},
\end{align}
jointly with the VaR regression in \eqref{eq:VaR GaR}.
The conditional MES $\mathbb{E}_t[Y_{t,d}\mid X_t\geq\operatorname{VaR}_{t,\beta}]$ quantifies the expected (negative) GDP growth of country $d$ given that the economic region $X_t$ is in distress and, hence, measures the systemic vulnerability of economy $d$.
The MES regression in \eqref{eq:GDPmodel} therefore assesses how this systemic vulnerability is affected by past economic and financial conditions.
By virtue of \eqref{eq:ES decomp}, the three MES regressions in \eqref{eq:GDPmodel} provide a complete dissection of the aggregated growth vulnerabilities captured by the \citet{ABG19} regression in \eqref{eq:ES GaR}.


\begin{figure}[tb]
	\centering
	\includegraphics[width=\linewidth]{input/MES_DCC_GDPgrowth}
	\caption{Estimated VaR and ES sample paths at level $\beta=0.9$ for our regression (solid lines) and a Gaussian DCC--GARCH model (dashed lines)  for the negative GDP growth of the joint economic region together with the estimated MES for Germany, France and the United Kingdom in the respective panels.
	Returns on quarters with a exceedance of the estimated conditional VaR (of our regression) are displayed as black points while all other returns are shown in gray color.}
	\label{fig:MES_GDP}
\end{figure}


The lower panel of Table~\ref{tab:GDP_MES} displays the MES parameter estimates together with their standard errors.
Four out of six MES slope parameters, which equal the marginal effect under adverse conditions (see Remark~\ref{rem:MarginalEffect}), are significant at the $5\%$-level and we find the striking result that the explanatory power of the different variables varies strongly between the countries.
This is in stark contrast to the classic (VaR, ES) regression of \citet{ABG19}, where financial and economic conditions are equally important determinants of growth risk.
For instance, the MES of the UK mainly loads on the financial conditions indicator, which may be explained by a high dependence on the financial sector of the UK with London as one of the most important financial hubs in the world.
In contrast, the systemic vulnerability of the German economy is primarily driven by past economic conditions, which may be due to its strong export-oriented manufacturing sector.


For weights $\bm w_t $ being constant over time, the true ES regression parameters are given by the weighted average of the true MES regression parameters; see \eqref{eq:ES decomp}.
While this relationship is blurred in finite samples and due to the time-varying weights in this application (the weights for Germany vary between 0.4 and 0.45 and the ones for France and the UK between 0.26 and 0.31 in our sample),
it can still be recognized in the parameter estimates in Table \ref{tab:GDP_MES}.
E.g., for the slope coefficient pertaining to $\operatorname{FCI}$, we have $0.663 \approx 0.7757 = 0.426 \cdot 0.352 + 0.291 \cdot 0.818 + 0.283 \cdot 1.370$, where 0.426, 0.291, and 0.283 are the average weights for Germany, France, and the UK.

Figure~\ref{fig:MES_GDP} plots the implied VaR, ES and MES estimates for the regressions in \eqref{eq:VaR GaR}--\eqref{eq:ES GaR} and \eqref{eq:GDPmodel} for the joint economic region together with the three individual economies with solid lines.
For comparison, the dashed lines show the implied VaR, ES and MES estimates of a four-dimensional DCC--GARCH model with Gaussian innovations \citep{Eng02}.
The display shows very plausible model fits of both methods, which is especially remarkable given that we estimate risk measure models in the tails with only $n=99$ observations.

\begin{figure}[tb]
	\centering
	\includegraphics[width=\linewidth]{input/GDP_Diagnostics_Date_DEU.pdf}
	\includegraphics[width=\linewidth]{input/GDP_Diagnostics_Predictions_DEU.pdf}
	\caption{In-sample model diagnostics of the VaR and MES predictions of our regression and a Gaussian DCC--GARCH model for $\beta = 0.9$. Here, $Y_t$ is the negative GDP growth of Germany and $X_t$ the joint economic region.
	The generalized residuals are plotted and non-parametrically regressed (in blue) against time (upper row) and the model predictions (lower row).}
	\label{fig:ModelDiagnostics_GDP_MES_DCC}
\end{figure}

Figure~\ref{fig:ModelDiagnostics_GDP_MES_DCC} further illustrates the good in-sample model fit by plotting (VaR and MES) generalized model residuals (see Section~\ref{sec:ApplMESReg}) against time and the model predictions themselves.
As the zero line is barely excluded from the confidence bands, we obtain (almost) no evidence against model misspecification for all models.
However, due to the limited sample size, this was to be expected.

To summarize, while in the standard GaR literature growth vulnerabilities are the main object of study, our MES regressions allow to decompose these into \textit{systemic} vulnerabilities of the different countries, therefore allowing a more fine-grained picture to emerge.
We mention that our MES regressions may also be fruitfully used for single countries, such as the US.
In such cases, one may decompose total growth into different economic sectors (e.g., manufacturing, service and agricultural) to shed more light on growth vulnerabilities in these different parts of the economy.
We leave such explorations for future study.








\section{Conclusion}
\label{sec:Conclusion}

Regressions typically model some functional of the \textit{scalar} outcome $Y_t$ (say, the mean in mean regressions or quantiles in quantile regressions) as a function of covariates.
In contrast, in this paper we model a functional of the \textit{bivariate} outcome $(X_t,Y_t)^\prime$ (viz.~the MES) as a function of covariates.
This leads to what we term \emph{regressions under adverse conditions}, where the mean of $Y_t$, given the adverse condition that a distress variable $X_t$ exceeds its conditional quantile, is related to covariates.
We develop a two-step estimator and prove its asymptotic normality under general mixing conditions allowing for cross-sectional and time series applications, where we overcome the technical difficulty of having a discontinuous second-step objective function.
We also propose feasible inference that works well as shown in our simulations.
Our estimator particularly improves upon a recently proposed \textit{ad hoc} estimation method of \citet{BRS20, BDP20}, which we illustrate in simulations and an empirical application.


While our regression target---the conditional expectation of $Y_t$ given that $X_t$ exceeds its conditional quantile---is originally known as the MES in the (financial) systemic risk literature, we illustrate its value in three diverse applications focusing on the relation between systemic risk and asset price bubbles \citep{BRS20}, dissecting GDP growth risks among countries \citep{ABG19}, and in Appendix \ref{sec:add appl} by analyzing portfolio risk allocation \citep{Bea18}.


Moreover, we envision a multitude of further possible fields of applications such as in sensitivity analysis \citep[Sec.~1.2]{Asimit2019} or micro-econometric applications, where the MES could, e.g., be interpreted as the mean (personal or firm) income given the stress event of high inflation rates, or other challenging economic settings.
We also refer to \citet{HTZ23}, \citet{WTH24} and the references therein for the growing interest in quantile-truncated expectations in the micro-econometric literature.



A particular advantage of the MES as our regression target is its additivity \citep{CIM13} that---similar to our Section \ref{sec:GDP_MES}---allows for ES risk decompositions into different (additive) sub-components such as economic sectors.
In a similar fashion, recently analyzed inflation risks \citep{LL24} could be decomposed into product categories.
As such disaggregations are only possible by the additivity of the ES and MES risk measures as conditional expectations, these considerations provide further arguments for use of the ES over the VaR (and the MES over the CoVaR) in the recent debate of which risk measure is most suitable in practice \citep{EKT15, BCBS19, WangZitikis2021}.







\section*{Supplementary Materials}

The Supplementary Material contains the proof of Theorem~\ref{thm:an}, and shows how to estimate the asymptotic variance-covariance matrix of Theorem~\ref{thm:an} consistently.
It also verifies our main assumptions for the linear model of Example~\ref{ex:1} and provides an additional empirical application to portfolio optimization.

Replication material for the simulations and the applications in Section \ref{sec:EmpiricalApplication} (together with the data) is available under \href{https://github.com/TimoDimi/replication_MES}{https://github.com/TimoDimi/replication\_MES}.
It draws on the corresponding open source package \texttt{SystemicRisk} \citep{package_SystemicRisk} implemented in the statistical software \texttt{R} \citep{R2022}.





\section*{Acknowledgments}

We would like to thank the AE, two anonymous referees, Markus Brunnermeier, Tobias Fissler, Christoph Hanck, Simon Rother and Isabel Schnabel for their insightful comments that significantly improved the quality of the paper.
We further thank Simon Rother for providing us with the bubble indicators used in \citet{BRS20}.





\section*{Funding}

The authors gratefully acknowledge support of the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) through grants 502572912 (Timo Dimitriadis), 460479886 and 531866675 (Yannick Hoga).



\pagebreak