EconBase
← Back to paper

Backtesting Systemic Risk Forecasts using Multi-Objective Elicitability

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

112,584 characters · 0 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Backtesting Systemic Risk Forecasts using Multi-Objective Elicitability

abstractSystemic risk measures such as CoVaR, CoES and MES are widely-used in finance, macroeconomics and by regulatory bodies. Despite their importance, we show that they fail to be elicitable and identifiable. This renders forecast comparison and validation, commonly summarised as `backtesting', impossible. The novel notion of multi-objective elicitability solves this problem. Specifically, we propose Diebold--Mariano type tests utilising two-dimensional scores equipped with the lexicographic order. We illustrate the test decisions by an easy-to-apply traffic-light approach. We apply our traffic-light approach to DAX 30 and S&P 500 returns, and infer some recommendations for regulators.\\ Keywords: Backtest; (Conditional) Elicitability; Forecasting; Identifiability; Lexicographic Order; Multi-objective Optimisation; Systemic Risk\\ JEL classification: C18 (Methodological Issues), C52 (Model Evaluation, Validation, and Selection), C58 (Financial Econometrics)\\
bibunit\section{Motivation} Regulating financial institutions in isolation is often not sufficient to prevent financial crises due to the interdependent risks these institutions face. In particular, their losses commonly exhibit a pronounced comonotonic behaviour in the extreme tails: When one financial institution, or the market as a whole, is in distress, other institutions are much more prone to being at risk as well. The U.S. subprime mortgage crisis of 2008--2009, the European sovereign debt crisis of 2010--2011 and the Covid-19 crash of 2020 have forcefully demonstrated this fact and also the need to assess the systemic nature of risk. As a consequence of these crises, a huge literature on measuring systemic risk has emerged over the last decade GK11,ChenIyengarMoallemi2013,AB16, Aea17, BE17, FeinsteinRudloffWeber2017. Systemic risk measures are important in various contexts. First, they are important in banking regulation under the Basel framework of the BCBSBF19, where they are vital in determining which banks are among the globally systemically important banks (G-SIBs). Such G-SIBs are then subjected to higher capital requirements. Second, in finance, systemic risk measures may be used to study spillover effects in the financial system AB16 or the build-up of asset price bubbles BRS20. Third, they may be used to study the linkage between the financial sector and the real economy. Among others, GKP16 and BE17 show that an increase in systemic risk is predictive of future declines in real economic activity. All these examples underscore the importance of accurately measuring and predicting systemic risk. In this paper, we revisit three influential systemic risk measures. First, we consider AB16's AB16 conditional value-at-risk (CoVaR) and conditional expected shortfall (CoES) as extensions of the well-known value-at-risk (VaR) and expected shortfall (ES) to the realm of systemic risk. If $Y$ are the losses of interest and $X$ the losses of a reference position, $\operatorname{CoVaR}_{\alpha|\beta}(Y|X)$ ($\operatorname{CoES}_{\alpha|\beta}(Y|X)$) is the VaR (ES) of $Y$ at level $\alpha$, given that $X$ is “in distress”. Here, we interpret the event that $X$ is in distress as $X$ being larger or equal than its $\beta$-quantile, i.e., $\{X\ge \operatorname{VaR}_\beta(X)\}$. Finally, we consider Aea17's Aea17 marginal expected shortfall, $\operatorname{MES}_{\beta}(Y|X)$, as the conditional mean of $Y$ given $\{X\ge \operatorname{VaR}_\beta(X)\}$. Section (ref) introduces the exact definitions. Bea17 distinguish between the “source-specific approach” and the “global approach” to systemic risk measurement. The source-specific approach considers individual sources of systemic risk, such as contagion risk or liquidity crises. In contrast, global measures of systemic risk potentially incorporate all mechanisms studied in the source-specific approach. Bea17 categorize CoVaR, CoES and MES under the global approach. In practice, forecasting systemic risk measures---such as CoVaR, CoES and MES---requires adequate models for the marginals $X$ and $Y$, and for their dependence structure. The literature has developed numerous different modelling approaches for this; see GT13 and BC19 for forecasting models for CoVaR and CoES, and BE17 and Eck18 for MES models. Due to the importance of systemic risk measures outlined above, it is vital to develop statistical quality assessments of the various models' predictive performances. It is the main aim of this paper to provide such tools, which are referred to as `backtests' in finance. Backtests have two main goals. On the one hand, one may wish to assess the absolute quality of forecasting models, also called the calibration, akin to model validation in statistics. Following the terminology of FZG16, we call such procedures “traditional backtests”. Roughly speaking, they check how well a sequence of risk measure forecasts aligns with corresponding observations of losses. Traditional backtests rely on the identifiability of the underlying risk measure, which ensures the existence of a (possible multivariate) function $\bm V$ that uniquely “identifies” the true report (see Definition (ref)). On the other hand, the presence of several alternative prediction models for a risk measure necessitates “comparative backtests” FZG16 to assess their predictive accuracy relative to each other. This is akin to statistical model selection procedures. Comparative backtests exploit the elicitability of the underlying risk measure. This implies the existence of a \textit{real-valued} loss (or also: scoring) function $S$, which is minimised in expectation by the optimal forecast (see Definition (ref)). Our contributions in this paper are the following: First, we show that $\operatorname{CoVaR}_{\alpha|\beta}$, $\operatorname{CoES}_{\alpha|\beta}$ and $\operatorname{MES}_\beta$ are not identifiable and elicitable as standalone risk measures (Proposition (ref)). The practical implication of this is that neither traditional nor comparative backtests can be carried out. In particular, any regulation based solely on these systemic risk measures is pointless, because neither the adequacy of the forecasts can be determined nor can different systemic risk forecasts be sensibly compared (say to a regulatory standard model). We provide a partial remedy for this drawback by giving joint identification functions for $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X))$, $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X), \operatorname{CoES}_{\alpha|\beta}(Y|X))$, and $(\operatorname{VaR}_\beta(X), \operatorname{MES}_\beta(Y|X))$ (Theorem (ref)). These identification functions can be used for (conditional) calibration tests in the spirit of NZ17. To the best of our knowledge, this entails the first traditional backtest for these systemic risk measures apart from Banulescu-RaduETAL2019. We contrast our approach with theirs in detail in Remark (ref). In particular, they use one-dimensional identification functions for $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X))$ and $(\operatorname{VaR}_\beta(X), \operatorname{MES}_\beta(Y|X))$, which fail to be strict in contrast to our two-dimensional identification functions. For the backtest of Banulescu-RaduETAL2019, this non-strictness leads to a complete loss of power in identifying certain misspecified systemic risk forecasts (Section (ref) in the Supplement). Theoretically, our results are akin to the fact that $\operatorname{ES}_\alpha(Y)$ is not identifiable on its own, but the pair $(\operatorname{VaR}_\alpha(Y), \operatorname{ES}_\alpha(Y))$ is identifiable FZ16a. In stark contrast to the joint elicitability of the pair $(\operatorname{VaR}_\alpha, \operatorname{ES}_\alpha)$, however, we show that the pairs $(\operatorname{VaR}_\beta,$ $\operatorname{CoVaR}_{\alpha|\beta})$, $(\operatorname{VaR}_\beta, \operatorname{MES}_\beta)$ and the triplet $(\operatorname{VaR}_\beta, \operatorname{CoVaR}_{\alpha|\beta}, \operatorname{CoES}_{\alpha|\beta})$ \emph{fail} to be elicitable (Section (ref)). So while traditional backtests for the above pairs and the triplet may be constructed by virtue of their identifiability, classical comparative backtests exploiting elicitability are not feasible. As a remedy to this negative result, we propose the novel concept of \emph{multi-objective elicitability}, which works with \textit{multivariate} scores $\bm S$ mapping to $\mathbb{R}^m$ equipped with a certain (partial) order relation $\preceq$. This contrasts sharply with classical $\mathbb{R}$-valued losses $S$. Their prevalence to date is grounded in tradition Gne11 and the fact that $\mathbb{R}$ is equipped with the canonical (total) order relation $\le$, which allows for straightforward comparisons of losses. Subsection (ref) introduces \emph{multi-objective scores} $\bm S$ and the corresponding concepts of \emph{multi-objective consistency and elicitability}. The terminology stems from the field of \emph{multi-objective optimisation}: According to Ehrgott2005 it is “a mathematical theory of optimization under multiple objectives”, and can be encountered in various fields of science, economics, logistics and engineering. Since this novel concept to forecast evaluation may open up the avenue to a whole field of applications and research (which is underpinned by further instances; see Example (ref)), we give a concise general outline of the theory, using partial orders on $\mathbb{R}^m$ or even infinite-dimensional real vector spaces. For the systemic risk forecasts we consider here, scores mapping to $\mathbb{R}^2$ equipped with the lexicographic (total) order---described in Subsection (ref)---are sufficient. In particular, the performance of different systemic risk forecasts must be ranked with regard to the lexicographic order. Specifically, Theorem (ref) shows that $(\operatorname{VaR}_\beta, \operatorname{CoVaR}_{\alpha|\beta})$, $(\operatorname{VaR}_\beta, \operatorname{CoVaR}_{\alpha|\beta}, \operatorname{CoES}_{\alpha|\beta})$ and $(\operatorname{VaR}_\beta, \operatorname{MES}_\beta)$ are multi-objective elicitable, and it provides classes of strictly multi-objective consistent scores. We outline in Section (ref) how these scores can be used for comparative backtests of Diebold--Mariano type. These comparative backtests are different from---in our case infeasible---“standard” comparative backtests in that they build on the newly introduced notion of multi-objective elicitability (with scores mapping to $\mathbb{R}^2$ equipped with the lexicographic order) instead of the “standard” notion of elicitability (with scores mapping to $\mathbb{R}$ equipped with the canonical order $\leq$). In particular, some systemic risk forecast is now preferable to some other forecast when the $\mathbb{R}^2$-valued score of the former is smaller (with regard to the lexicographic order) than that of the latter. Thus, financial institutions may build on this result to improve their prediction models for $(\operatorname{VaR}_\beta, \operatorname{CoVaR}_{\alpha|\beta})$, $(\operatorname{VaR}_\beta, \operatorname{CoVaR}_{\alpha|\beta}, \operatorname{CoES}_{\alpha|\beta})$ and $(\operatorname{VaR}_\beta, \operatorname{MES}_\beta)$, which is crucial for an adequate assessment of the diverse risks faced by these institutions. Since the multi-objective scores of Theorem (ref) take values in $\mathbb{R}^2$ equipped with the lexicographic order, some particularities arise for statistical hypothesis tests. While simple “two-sided” null hypotheses of equal predictive performance can be tested with a classical Wald-test, particular caution must be taken when testing for superior predictive ability. Due to the particularities of the lexicographic order, a straightforward “one-sided” composite null hypothesis would be insensitive to the systemic risk measure forecast, ignoring the primary goal of the backtesting procedure. Therefore, we suggest to use “one and a half”-sided composite null hypotheses, testing for superior predictive ability in the systemic risk component and equal performance in the auxiliary $\operatorname{VaR}_\beta(X)$ component. Section (ref) provides details, including an adaptation of the Basel framework's traffic-light approach to systemic risk backtests. An empirical application in Section (ref) demonstrates the viability of the comparative backtest. There, we consider daily log-losses of the DAX 30 with daily log-losses of the S&P 500 as a reference quantity. We compare systemic risk forecasts derived from a benchmark Gaussian copula model with those produced by a $t$-copula model, where in both models the correlation parameter of the copula is driven by generalised autoregressive score (GAS) dynamics CKL13. We find that the predictive performance of the $t$-copula is superior with $p$-values close to $3\%$, which is consistent with its popularity in empirical work. One conclusion from our empirical analysis is that fairly long samples are required to validly distinguish between different forecasts, because the effective sample sizes in comparing systemic risk forecasts are (almost by definition) reduced. Thus, the one year evaluation period for (univariate) VaR and ES forecasts in the Basel framework of the BCBSBF19 is, in our view, insufficient for systemic risk forecasts. The paper closes with a discussion and outlook (Section (ref)). Besides the parts already mentioned above, the Supplement provides proofs for the results of Section (ref) (Section (ref)) and further background material on multi-objective elicitability (Section (ref)). All other proofs are relegated to Section (ref). Section (ref) investigates the finite-sample properties of our comparative backtests in simulations. The \texttt{R} code to reproduce all numerical experiments is available online. Throughout the paper, we indicate vectors with bold letters. We highlight the distinction between row and column vectors only when it is essential, and use the symbol $'$ to indicate the transpose of a vector or matrix. \section{Formal definition of CoVaR, CoES and MES} Fix some non-atomic probability space $(\Omega, \mathfrak A, \operatorname{P})$ where all random objects are defined. Using standard notation, let $L^0(\mathbb{R}^d)$, $d=1,2$, be the space of all $\mathbb{R}^d$-valued random vectors on $(\Omega, \mathfrak A, \operatorname{P})$. Furthermore, for $p\in[1,\infty)$, let $L^p(\mathbb{R}^d)\subseteq L^0(\mathbb{R}^d)$ be the collection of random vectors whose components possess a finite $p$th moment. For $\bm X\in L^0(\mathbb{R}^d)$ let $F_{\bm X}$ be its joint distribution. Then define for $p\in\{0\}\cup[1, \infty)$ the collection $\mathcal{F}^p(\mathbb{R}^d):= \{F_{\bm X}\colon \bm X \in L^p(\mathbb{R}^d)\}$. We overload notation and identify any $F\in \mathcal{F}^0(\mathbb{R}^d)$ with its cumulative distribution function (cdf) $\mathbb{R}^d \to [0,1]$. Our systemic risk measures of interest---$\operatorname{CoVaR}$, $\operatorname{CoES}$ and $\operatorname{MES}$---are maps from $L^0(\mathbb{R}^2)$ (or $L^1(\mathbb{R}^2)$ for MES) to $ \mathbb{R}^* := (-\infty,\infty]$. They are law-determined, meaning that their values for $(X,Y)$ and $(\tilde X, \tilde Y)$ coincide if $F_{X,Y} = F_{\tilde X, \tilde Y}$. Hence, we can consider them as risk-functionals on $\mathcal{F}^0(\mathbb{R}^2)$ (or $\mathcal{F}^1(\mathbb{R}^2)$). Similarly, the popular univariate risk measure $\operatorname{VaR}_\beta$, $\beta\in[0,1]$, is a law-determined map $L^0(\mathbb{R})\to[-\infty,\infty]$. In the rest of the paper, we will frequently overload notation and identify these law-determined risk measures with their induced risk functionals. As such, we will use the terms `risk measure' and `risk functional' interchangeably. Let $(X,Y)\in L^0(\mathbb{R}^2)$ be a two-dimensional random vector. Here, $Y$ stands for the losses of a position of interest (with the sign convention that positive values are losses and negative values are gains) and $X$ is a univariate reference position or aggregate of a reference system, having the same sign convention. Denote by $F_{X,Y}$ their joint distribution function and by $F_X$ and $F_Y$ their marginals, respectively. Recall that for $\beta\in[0,1]$ the $\beta$-quantile of $F_X$ is the closed interval $q_\beta(F_X) = \{x\in\mathbb{R}\colon F_X(x-)\le \beta \le F_X(x)\}$, where $F_X(x-) := \lim_{t\uparrow x}F_X(x-)$. Then, $\operatorname{VaR}_\beta(X)$ is the lower $\beta$-quantile of $F_X$, i.e., $\operatorname{VaR}_\beta(X):= \operatorname{VaR}_\beta(F_X):= \inf q_\beta(F_X)$. For $\beta\in(0,1)$, $\operatorname{VaR}_\beta$ is always finite. Our sign convention is such that the larger the risk measure of a position, the riskier it is deemed. Hence, we typically choose a probability level of $\beta$ close to 1 for $\operatorname{VaR}_\beta$, such as $\beta=0.95$ or $\beta=0.99$. AB16 define $\operatorname{CoVaR}_\beta$ as the $\beta$-quantile of the conditional distribution function $F_Y(\ \cdot\mid X= \operatorname{VaR}_{\beta}(X)) = \operatorname{P}\{Y\leq\cdot\mid X = \operatorname{VaR}_{\beta}(X)\}$. The conditioning event in this definition is problematic for several reasons. First, it may have probability zero (which is the case when $F_X$ is continuous). Second, it does not fully capture the tail-risk of $F_X$. Third, since the roles of $Y$ and $X$ are asymmetric by construction, one may want to consider different probability levels to specify the event of `being in distress'. Thus, we follow GT13 and Banulescu-RaduETAL2019 in redefining $\operatorname{CoVaR}_{\alpha|\beta}\colon L^0(\mathbb{R}^2)\to\mathbb{R}$ for $\alpha\in(0,1)$, $\beta\in[0,1)$ as \begin{equation} \operatorname{CoVaR}_{\alpha|\beta}(Y| X):=\operatorname{CoVaR}_{\alpha|\beta}(F_{X,Y}):=\operatorname{VaR}_\alpha(F_{Y\mid X \geq \operatorname{VaR}_\beta(X)}), \end{equation} where $F_{Y\mid X \geq \operatorname{VaR}_\beta(X)}=\operatorname{P}\{Y\leq\cdot\mid X \geq \operatorname{VaR}_{\beta}(X)\}$. For $\beta=0$, we simply have $\operatorname{CoVaR}_{\alpha|0}(Y|X) := \operatorname{VaR}_\alpha(Y)$, and if $\beta=\alpha$ we simply write $\operatorname{CoVaR}_{\alpha}(Y| X) = \operatorname{CoVaR}_{\alpha|\alpha}(Y| X)$. Since $\operatorname{CoVaR}_{\alpha|\beta}(Y| X)$ is merely a quantile of the distribution $F_{Y\mid X \geq \operatorname{VaR}_\beta(X)}$, it inherits the same defects as $\operatorname{VaR}_\alpha(Y)$. That is, it ignores tail risks beyond the quantile level $\alpha$ and it fails to be coherent, particularly defying the rationale of advantageous diversification effects Aea99,MS14. The well-known \textit{Expected Shortfall} at level $\alpha\in(0,1)$, $\operatorname{ES}_\alpha(Y):= \operatorname{ES}_\alpha(F_Y):=\frac{1}{1-\alpha}\int_\alpha^1 \operatorname{VaR}_\gamma(Y) \mathrm{d} \gamma \in \mathbb{R}^*$, does not suffer from these defects. This motivates AB16 to introduce the conditional Expected Shortfall (CoES). As for $\operatorname{CoVaR}_{\alpha|\beta}$, we modify the conditioning event and formally introduce for $\alpha\in(0,1)$, $\beta\in[0,1)$, $\operatorname{CoES}_{\alpha|\beta}\colon L^0(\mathbb{R}^2)\to \mathbb{R}^*$ via \begin{align} &\operatorname{CoES}_{\alpha|\beta}(Y|X):= \operatorname{CoES}_{\alpha|\beta}(F_{X,Y}) := \frac{1}{1-\alpha}\int_\alpha^1 \operatorname{CoVaR}_{\gamma | \beta}(Y|X)\mathrm{d} \gamma . \end{align} If $F_{X,Y}$ is continuous, then $\operatorname{CoES}_{\alpha|\beta}(Y|X) = \operatorname{E}[Y | Y\ge \operatorname{CoVaR}_{\alpha|\beta}(Y|X) ,\ X\ge \operatorname{VaR}_\beta(X)]$. Again, $\operatorname{CoES}_{\alpha|0}(Y|X) = \operatorname{ES}_\alpha(Y)$, and we write $\operatorname{CoES}_\alpha(Y|X) := \operatorname{CoES}_{\alpha|\alpha}(Y|X)$. Finally, we consider the marginal Expected Shortfall (MES) of Aea17, which measures the expectation of $Y$ when $X$ is in distress, i.e., when $X$ is in its right tail. Specifically, we introduce for $\beta\in[0,1)$ the map $\operatorname{MES}_\beta\colon L^1(\mathbb{R}^2)\to\mathbb{R}$, \begin{equation} \operatorname{MES}_{\beta}(Y|X) := \operatorname{MES}_{\beta}(F_{X,Y}) := \operatorname{CoES}_{0|\beta}(Y|X) = \int_0^1 \operatorname{CoVaR}_{\gamma|\beta}(Y|X) \mathrm{d} \gamma . \end{equation} Again, $\operatorname{MES}_{0}(Y|X) = \operatorname{E}[Y]$, and $\operatorname{MES}_{\beta}(Y|X) = \operatorname{E}[Y | X\ge \operatorname{VaR}_\beta(X)]$ if $F_{X,Y}$ is continuous. If $F_X$ is discontinuous, Remark (ref) proposes a novel correction term which generalises the three measures considered in this paper. \section{(Conditional) identifiability and multi-objective elicitability} We present the theory in this section in all generality to serve as a basis for future research on multi-objective scores. Therefore, we work with general functionals which do not necessarily have the interpretation of risk functionals. All proofs are in Section (ref). \subsection{Notation, basic definitions and results} Adopting the decision-theoretic terminology of Gne11, we denote by $\mathsf{A}$ an action domain. This is the space of plausible forecasts, which can be finite for categorial forecasts, $\mathbb{R}$ or $\mathbb{R}^k$ for point forecasts, or a set of distributions for probabilistic forecasts. Moreover, let $\mathsf{O}$ be an observation domain---a set where verifying observations materialise---with $\mathsf{O} = \mathbb{R}^d$ as a leading example. Denote by $\mathcal{F}^0(\mathsf{O})$ the set of all probability distributions on $\mathsf{O}$. Let $\mathcal{F}\subseteq \mathcal{F}'$ be subclasses of $\mathcal{F}^0(\mathsf{O})$. We consider a general, possibly set-valued functional $\bm T\colon\mathcal{F}'\to\mathcal P(\mathsf{A})$ with $\mathcal P(\mathsf{A})$ the power set of $\mathsf{A}$. Later on, $\bm T$ will have the interpretation of a risk functional. Note that $\mathcal{F}'$ is the class of distributions, where our functional $\bm T$ is defined on, and $\mathcal{F}\subseteq \mathcal{F}'$ is the subclass on which $\bm T$ will be identifiable/elicitable (see Definitions (ref) and (ref)). For instance, for $\bm T(F_{X,Y})=(\operatorname{VaR}_\beta(F_X),$ $\operatorname{CoVaR}_{\alpha|\beta}(F_{X,Y}) )$, we have $\mathcal{F}'=\mathcal{F}^0(\mathbb{R}^2)$ and $\mathcal{F}$ is given in Theorem (ref) (i)/Theorem (ref) (ii). We adopt the \emph{selective} notion of forecasts discussed in FFHR2021 where one is content with correctly specifying a single element $\bm t \in \bm T(F) \subseteq \mathsf{A}$ as opposed to specifying the entire set $\bm T(F)$. (If one is interested in \emph{exhaustive} forecasts---i.e., in forecasts for the whole set $\bm T(F)$---one can change the action domain to $\mathcal P(\mathsf{A})$.) If $\bm T$ attains singletons only, we identify the value of $\bm T(F)$ with its unique element. This identification allows us to treat point-valued functionals as set-valued functionals without loss of generality. We mention that the risk functionals to be considered in Section (ref) are all point-valued. A function $\bm g\colon \mathsf{A}\times \mathsf{O}\to\mathbb{R}^{\mathcal I}$, where $\mathcal I$ is an index set, is called $\mathcal{F}$-\textit{integrable} if for all components $g_i$, $i\in\mathcal I$, it holds that $\int |g_i(\bm r,\bm y)|\,\mathrm{d} F(\bm y)<\infty$ for all $\bm r\in \mathsf{A}$, $F\in\mathcal{F}$. If $\bm g$ is $\mathcal{F}$-integrable, we define the map $\bar \bm g\colon \mathsf{A} \times \mathcal{F} \to\mathbb{R}^{\mathcal I}$, $\bar \bm g(\bm r, F) := \int g_i(\bm r,\bm y)\,\mathrm{d} F(\bm y)$ for $\bm r\in \mathsf{A}$, $F\in\mathcal{F}$. A similar convention and notation is used for maps $\bm a\colon\mathsf{O}\to\mathbb{R}^{\mathcal I}$. We start with the classical definition of elicitability and consistent scoring functions, mapping to $\mathbb{R}$. \begin{defn} An $\mathcal{F}$-integrable map $S\colon \mathsf{A}\times \mathsf{O}\to\mathbb{R}$ is an \emph{$\mathcal{F}$-consistent scoring function} for $\bm T$ if $\bar S(\bm t,F)\le \bar S(\bm r,F)$ for all $\bm t\in \bm T(F)$, all $\bm r\in\mathsf{A}$ and for all $F\in\mathcal{F}$. It is a \emph{strictly} $\mathcal{F}$-consistent scoring function for $\bm T$ if, additionally, $\bar S(\bm t,F)= \bar S(\bm r,F)$ implies that $\bm r\in \bm T(F)$. $\bm T$ is \emph{elicitable on $\mathcal{F}$} if there is a strictly $\mathcal{F}$-consistent scoring function for $\bm T$. \end{defn} \begin{defn} An $\mathcal{F}$-integrable map $\bm V\colon \mathsf{A}\times \mathsf{O}\to\mathbb{R}^m$ is an \emph{$\mathcal{F}$-identification function} for $\bm T$ if $\bar \bm V(\bm t,F)=\bm 0$ for all $\bm t\in \bm T(F)$ and for all $F\in\mathcal{F}$. It is a \emph{strict} $\mathcal{F}$-identification function for $\bm T$ if, additionally, for all $F\in\mathcal{F}$ and for all $\bm r\in \mathsf{A}$, $\bar \bm V(\bm r,F)=\bm 0$ implies that $\bm r\in \bm T(F)$. $\bm T$ is \emph{identifiable on $\mathcal{F}$} if there is a strict $\mathcal{F}$-identification function for $\bm T$. \end{defn} Suppose in a risk management context that $\bm T$ corresponds to the $\operatorname{VaR}_{\alpha}$-functional and $y_t$ are the observed losses. Subject to mild conditions on $\mathcal{F}$, $\operatorname{VaR}_\alpha$ is elicitable, where the `pinball loss' $S(r,y) = (\mathds{1}\{y>r\} - 1+\alpha)(r-y)$ is strictly $\mathcal{F}$-consistent. This allows to compare competing VaR forecasts $(r_{t,(1)})_{t=1, \ldots, n}$ and $(r_{t,(2)})_{t=1, \ldots, n}$ via their empirical average score differences $\overline{d}_n = \overline{S}_{1n}-\overline{S}_{2n} = \frac{1}{n} \sum_{t=1}^n S(r_{t,(1)},y_t) - S(r_{t,(2)},y_t)$. A negative (positive) sign of $\overline{d}_n$ indicates superiority (inferiority) of $(r_{t,(1)})_{t=1, \ldots, n}$ over $(r_{t,(2)})_{t=1, \ldots, n}$. Identifiability, on the other hand, opens the way to test for calibration by checking, e.g., if the test statistic $\frac{1}{n}\sum_{t=1}^n \bm V(r_{t,(i)},y_t)$ is sufficiently close to $\bm 0$ or not. E.g., when $\bm T$ corresponds to $\operatorname{VaR}_\alpha$ checking calibration amounts to checking if the empirical VaR-violation rate is roughly $1-\alpha$, which can be done in terms of $V(r,y) = \mathds{1}\{y>r\} - (1-\alpha)$. These examples demonstrate the importance of elicitability and identifiability for comparing and evaluating (risk) forecasts in practice. Under regularity conditions, the notions of elicitability and identifiability are equivalent for point-valued functionals mapping to $\mathbb{R}$ SteinwartPasinETAL2014. There are also important functionals which fail to be elicitable and identifiable, most prominently the variance and expected shortfall Gne11. In such situations, the notions of \emph{conditional} elicitability and \emph{conditional} identifiability can be helpful. Following the concept presented in FZ16a, we slightly adapt EKT15's EKT15 original definition of conditional elicitability. We also introduce the corresponding counterpart of conditional identifiability. \begin{defn} Consider two functionals $\bm T_j:\mathcal{F}'\to \mathcal P(\mathsf{A}_j)$, $j=1,2$, and let $\mathcal{F}\subseteq \mathcal{F}'$. \begin{enumerate}[\normalfont (i)] • $\bm T_{2}$ is \textit{conditionally elicitable with} $\bm T_{1}$ on $\mathcal{F}$, if $\bm T_1$ is elicitable on $\mathcal{F}$ and $\bm T_2$ is elicitable on \( \mathcal{F}_{\bm r_1}:=\left\{F\in\mathcal{F}\colon\ \bm r_1\in \bm T_1(F)\right\} \) for any $\bm r_1\in \mathsf{A}_1$. • $\bm T_{2}$ is \textit{conditionally identifiable with} $\bm T_{1}$ on $\mathcal{F}$, if $\bm T_1$ is identifiable on $\mathcal{F}$ and $\bm T_2$ is identifiable on \( \mathcal{F}_{\bm r_1}:=\left\{F\in\mathcal{F}\colon\ \bm r_1\in \bm T_1(F)\right\} \) for any $\bm r_1\in \mathsf{A}_1$. \end{enumerate} \end{defn} It is easy to see that the variance is conditionally elicitable and conditionally identifiable with the mean, and that $\operatorname{ES}_\alpha$ is conditionally elicitable and conditionally identifiable with $\operatorname{VaR}_\alpha$ on appropriate classes of distributions, respectively EKT15. The pairs (mean, variance) and $(\operatorname{VaR}_\alpha, \operatorname{ES}_\alpha)$ even turn out to be elicitable and identifiable FZ16a. For identifiability, this is an instance of the following proposition, which is already stated without a proof in the discussion of FZ16a. \begin{prop} If $\bm T_2$ is conditionally identifiable with $\bm T_1$ on $\mathcal{F}$, then the pair $(\bm T_1,\bm T_2)$ is identifiable on $\mathcal{F}$. \end{prop} Of course, an analogue to Proposition (ref) for (conditional) elicitability would be desirable, and it has been stated as an open problem in the discussion of FZ16a. Unfortunately, the answer is negative: While Section (ref) establishes the conditional elicitability of $\operatorname{CoVaR}_{\alpha|\beta}$, $(\operatorname{CoVaR}_{\alpha|\beta}, \operatorname{CoES}_{\alpha|\beta})$ and $\operatorname{MES}_\beta$ all with $\operatorname{VaR}_\beta$, the corresponding pairs and the triplet generally fail to be elicitable; see Section (ref). \subsection{Multi-objective scores, consistency and elicitability} To overcome the structural drawback of elicitability in comparison to identifiability, in particular the lack of an analogue to Proposition (ref), we introduce the novel notion of \emph{multi-objective scoring functions} and the corresponding concepts of \emph{multi-objective consistency} and \emph{multi-objective elicitability}. It is inspired by the fundamental observation that identification functions are generally multivariate. To be more precise, the dimension $m$ of the identification function $\bm V$ usually coincides with the dimension $k$ of the functional. If $k=m=1$ an identification function is often induced by the derivative of a consistent scoring function $S$. Also, the antiderivative of an (oriented) identification function yields a consistent score, thus roughly establishing a one-to-one correspondence between the class of identification functions and the class of consistent scoring functions. For a $k$-dimensional functional, the gradient of a consistent score is $\mathbb{R}^k$-valued and naturally induces an identification function, subject to smoothness conditions. However, not every $k$-dimensional identification function possesses an antiderivative for $k\ge2$. This is due to integrability conditions asserting that if it was integrable, the corresponding Hessian of the stipulated antiderivative would need to be symmetric (see also Example (ref) for an illustration). This rules out the one-to-one relation in the multivariate setting, giving rise to a \emph{gap} between the class of consistent scoring functions and the one of identification functions. This gap and its consequences on estimation are discussed in detail by DFZ2020. We sidestep this integrability constraint by introducing the concept of \textit{multivariate} scoring functions. Indeed, it is this multivariate structure of identification functions which facilitates the straightforward proof of Proposition (ref). Therefore, we mimic this multi-dimensionality for scores, letting them map to $\mathbb{R}^m$, where usually $m=k$, or even more generally to some real vector space $\mathbb{R}^{\mathcal I}$, where the index set ${\mathcal I}$ may be finite, countable or even uncountable. In the application of the general theory developed here to the systemic risk measures, it suffices to consider $\mathbb{R}^2$-valued scores; see Theorem (ref) in Section (ref). The motivation for defining a classical score as a univariate map to $\mathbb{R}$ (see Definition (ref)) is grounded in tradition on the one hand. On the other hand, $\mathbb{R}$ is equipped with the canonical (total) order relation $\le$, which allows for straightforward comparisons of forecasts by checking whether the empirical average score differences satisfy $\overline{d}_n\geq0$ or $\overline{d}_n\leq0$. Consistency in Definition (ref) ultimately relies on the existence of an order. Hence, we need to equip $\mathbb{R}^{\mathcal I}$ with a vector partial order $\preceq$ and write $(\mathbb{R}^{\mathcal I}, \preceq)$. A binary relation $\preceq$ on $\mathbb{R}^{\mathcal I}$ is a \emph{partial order} if it is reflexive ($\forall \bm x \in \mathbb{R}^{\mathcal I}$, $\bm x\preceq \bm x$), antisymmetric ($\forall \bm x,\bm y\in \mathbb{R}^{\mathcal I}$ if $\bm x\preceq \bm y$ and $\bm y\preceq \bm x$, then $\bm x=\bm y$), transitive ($\forall \bm x,\bm y, \bm z \in \mathbb{R}^{\mathcal I}$ if $\bm x\preceq \bm y$ and $\bm y\preceq \bm z$, then $\bm x\preceq \bm z$). Two elements $\bm x,\bm y\in \mathbb{R}^{\mathcal I}$ are \emph{comparable} if $\bm x\preceq \bm y$ or $\bm y\preceq \bm x$. A \emph{total} order is a partial order where all elements are comparable. A partial order $\preceq$ on $\mathbb{R}^{\mathcal I}$ induces a \emph{strict} partial order $\prec$ on $\mathbb{R}^{\mathcal I}$ as follows. For all $\bm x,\bm y\in \mathbb{R}^{\mathcal I}$ it holds that $\bm x\prec \bm y$ if and only if $\bm x\preceq \bm y$ and $\bm x\neq \bm y$. A partial order $\preceq$ on $\mathbb{R}^{\mathcal I}$ is a \emph{vector partial order} if it is compatible with addition and positive scaling. That is, if for all $\bm x,\bm y, \bm z\in\mathbb{R}^{\mathcal I}$ and for all $\lambda \in (0,\infty)$ it holds that $\bm x \preceq \bm y$ implies that $\bm x + \bm z \preceq \bm y + \bm z$, and $\bm x \preceq \bm y$ implies that $\lambda \bm x \preceq \lambda \bm y$. (In the sequel, we always mean a vector partial order whenever we write “partial order”.) The canonical choice is the componentwise order defined for $\bm x=(x_i)_{i \in \mathcal I},\bm y = (y_i)_{i \in \mathcal I}\in\mathbb{R}^{\mathcal I}$ as $\bm x \preceq \bm y$ if and only if $x_i\le y_i$ for all $i\in \mathcal I$. The use of a partial order $\preceq$ on $\mathbb{R}^{\mathcal I}$ leads to two different generalisations of Definition (ref). These are due to the fact that in a partial order $\bm x \preceq \bm y$ implies that $\bm y \nprec \bm x$, but the reverse implication fails if $\bm x$ and $\bm y$ are not comparable. \begin{defn} \begin{enumerate}[(i)] • An $\mathcal{F}$-integrable map $\bm S\colon \mathsf{A}\times \mathsf{O}\to(\mathbb{R}^{\mathcal I}, \preceq)$ is a \emph{strongly} multi-objective $\mathcal{F}$-consistent scoring function for $\bm T\colon \mathcal{F}\to\mathcal P(\mathsf{A})$ if $\bar \bm S(\bm t,F)\preceq \bar \bm S(\bm r,F)$ for all $\bm t\in \bm T(F)$, all $\bm r\in\mathsf{A}$ and all $F\in\mathcal{F}$. It is a strictly strongly multi-objective $\mathcal{F}$-consistent scoring function for $\bm T$ if, additionally, $\bar \bm S(\bm t,F)= \bar \bm S(\bm r,F)$ implies that $\bm r\in \bm T(F)$. $\bm T$ is \emph{strongly multi-objective-elicitable on $\mathcal{F}$} (with respect to the order $\preceq$ on $\mathbb{R}^{\mathcal I}$) if there is a strictly strongly multi-objective $\mathcal{F}$-consistent scoring function for $\bm T$. • An $\mathcal{F}$-integrable map $\bm S\colon \mathsf{A}\times \mathsf{O} \to(\mathbb{R}^{\mathcal I}, \preceq)$ is a \emph{weakly} multi-objective $\mathcal{F}$-consistent scoring function for $\bm T\colon \mathcal{F}\to\mathcal P(\mathsf{A})$ if $\bar \bm S(\bm r,F)\nprec \bar \bm S(\bm t,F)$ for all $\bm t\in \bm T(F)$, all $\bm r\in\mathsf{A}$ and all $F\in\mathcal{F}$. It is a strictly weakly multi-objective $\mathcal{F}$-consistent scoring function for $\bm T$ if, additionally, for all $F\in\mathcal{F}$ and for any $\bm r_0\in\mathsf{A}$ it holds that if $\bar \bm S(\bm r,F)\nprec \bar \bm S(\bm r_0,F)$ for all $\bm r \in\mathsf{A}$, then $\bm r_0\in \bm T(F)$. $\bm T$ is \emph{weakly multi-objective-elicitable on $\mathcal{F}$} (with respect to the order $\preceq$ on $\mathbb{R}^{\mathcal I}$) if there is a strictly weakly multi-objective $\mathcal{F}$-consistent scoring function for $\bm T$. \end{enumerate} \end{defn} As discussed, a partial order on $\mathbb{R}^{\mathcal I}$ is sufficient to define multi-objective consistency. Using the notation $\bar \bm S(B,F):= \{\bar \bm S(\bm r,F)\colon \bm r\in B\}$ for $F\in\mathcal{F}$ and $B\subseteq \mathsf{A}$, the definition of multi-objective consistency does not require comparability of all elements $\bar \bm S(\mathsf{A},F)$ for some fixed $F\in\mathcal{F}$. Weak multi-objective consistency solely implies that all elements of $\bar \bm S(\bm T(F),F)$ are minimal in $\bar \bm S(\mathsf{A},F)$. The strict version additionally ensures that all minimal elements of $\bar \bm S(\mathsf{A},F)$ are in $\bar \bm S(\bm T(F),F)$. Strong multi-objective consistency additionally means that the elements of $\bar \bm S(\bm T(F),F)$ are not only minimal in $\bar \bm S(\mathsf{A},F)$, but that they are the (unique) least element of $\bar \bm S(\mathsf{A},F)$. Due to the uniqueness of a least element, $\bar \bm S(\bm T(F),F)$ is a singleton. The strict version of strong multi-objective consistency additionally means that $\bar \bm S(\bm T(F),F)\not\subset\bar \bm S(\mathsf{A}\setminus \bm T(F),F)$. See Figure (ref) for an illustration of the two situations. Clearly, (strict) strong consistency implies (strict) weak consistency. If we omit the qualifiers “weak” or “strong”, we refer to the strong version. \begin{figure}[t!] \caption{In both panels, the shaded areas correspond to $\bar \bm S(\mathsf{A},F)$ of a multi-objective score $\bm S$ mapping to $\mathbb{R}^2$ equipped with the componentwise order. The red sets correspond to $\bar \bm S(\bm T(F),F)$. In panel (a), this is the unique least element, illustrating the situation of strict strong multi-objective consistency. In panel (b), this set corresponds to all minimal elements, depicting the situation of strict weak multi-objective elicitability.} \end{figure} Multi-objective consistency translates the essence of consistency as a `truth serum', or being incentive compatible, to the multivariate realm. No matter what forecast $\bm r\in\mathsf{A}$ an agent issues, they would be better off (or at least not worse off) in expectation when issuing a correctly specified functional value $\bm t\in \bm T(F)$. (In the weak version, if an agent issued some $\bm t\in \bm T(F)$, they would not be better off with some other forecast.) Strictness means that any action $\bm r\in \mathsf{A}\setminus \bm T(F)$ leads to a worse outcome in expectation than the truth $\bm t\in \bm T(F)$. So this minimal requirement of honouring truthful forecasts is preserved. \begin{rem} To the best of our knowledge, the notion of multi-objective scoring functions with the related concepts is novel to the forecast evaluation literature. However, it has some connections to the concept of \emph{forecast dominance} introduced by EhmETAL2016. Let $(S_i)_{i\in\mathcal I}$ be some class of univariate (strictly) $\mathcal{F}$-consistent scoring functions $S_i\colon \mathsf{A}\times \mathsf{O}\to\mathbb{R}$ for some functional $\bm T\colon\mathcal{F}\to\mathcal P(\mathsf{A})$. Adapting their definition slightly, we say that a forecast $\bm r\in\mathsf{A}$ dominates $\tilde \bm r\in\mathsf{A}$ relative to $(S_i)_{i\in\mathcal I}$ if $\bar S_i(\bm r,F) \le \bar S_i(\tilde \bm r,F)$ for all $F\in\mathcal{F}$ and for all $i\in\mathcal I$. In our terminology, we could phrase this forecast ranking in terms of a multivariate score $\bm S\colon \mathsf{A}\times \mathsf{O} \to (\mathbb{R}^{\mathcal I}, \preceq)$, $\bm S(\bm r,\bm y) = \big(S_i(\bm r,\bm y)\big)_{i\in\mathcal I}$, where $\preceq$ is the usual componentwise partial order. By virtue of mixture representations, EhmETAL2016 impressively demonstrate that for quantiles and expectiles, one can rephrase forecast dominance with respect to (nearly) all consistent scoring functions equivalently in terms of extremal or elementary scores $(S_\theta)_{\theta\in \mathbb{R}}\subseteq (S_i)_{i\in\mathcal I}$. For this situation, we can easily construct a multi-objective score $\bm S = (S_\theta)_{\theta\in \mathbb{R}}$, mapping to the function space $\mathbb{R}^\mathbb{R}$ equipped with the componentwise order. EhmETAL2016 provide numerous instances of comparable and incomparable forecast rankings. This example provides an easy construction for a (strictly) multi-objective $\mathcal{F}$-consistent score, $\bm S = (S_\theta)_{\theta\in \mathbb{R}}$, if $\bm T$ is elicitable. \end{rem} For verifying observations $(\bm y_t)_{t=1, \ldots, n}$, our Definition (ref) suggests to compare two sequences of forecasts $(\bm r_{t,(i)})_{t=1, \ldots, n}$ ($i=1,2$) via their average scores $\overline{\bm S}_{1n}=\frac{1}{n} \sum_{t=1}^n \bm S(\bm r_{t,(1)},\bm y_t)$ and $\overline{\bm S}_{2n}=\frac{1}{n} \sum_{t=1}^n \bm S( \bm r_{t,(2)},\bm y_t)$. However, while using classical univariate scores always leads to a conclusive forecast ranking (ignoring questions of statistical significance for a moment), the presence of a \emph{partial} order may lead to inconclusive rankings, particularly if neither of the two forecasts is correctly specified. For instance, if $\overline{\bm S}_{1n}=(1, 4)$ and $\overline{\bm S}_{2n}=(4, 1)$ (e.g., for the bivariate systemic risk scores of Theorem (ref)), then neither $\overline{\bm S}_{1n}\preceq\overline{\bm S}_{2n}$ nor $\overline{\bm S}_{2n}\preceq\overline{\bm S}_{1n}$ in the componentwise order on $\mathbb{R}^2$. \subsection{Multi-objective elicitability with respect to the lexicographic order \\ and conditional elicitability} To overcome the issue of inconclusive forecast rankings, we equip $\mathbb{R}^{\mathcal I}$ with a total order. In a total order, all elements are comparable. Hence, the notions of weak and strong multi-objective consistency / elicitability coincide and a distinction is redundant. A promising choice is the lexicographic order. For our purposes, and to ease the exposition, it suffices to consider $\mathbb{R}^2$ equipped with the lexicographic order $\preceq_{\mathrm{lex}}$. On $\mathbb{R}^2$ it holds that $(x_1,x_2)\preceq_{\mathrm{lex}} (y_1,y_2)$ if $x_1<y_1$ or if ($x_1=y_1$ and $x_2\le y_2$). The lexicographic order is widely used for preference relations in microeconomics. In particular, it is well-known for the fact that it cannot be represented by a real-valued utility function MWG1995. There are at least two reasons for our choice of the lexicographic order. First, as a total order, it allows for conclusive forecast rankings, which we exploit in our comparative backtests in Section (ref). For instance, if as above $\overline{\bm S}_{1n}=(1, 4)$ and $\overline{\bm S}_{2n}=(4, 1)$ for the systemic risk scores, then the inconclusive ranking is resolved because $\overline{\bm S}_{1n}\preceq_{\mathrm{lex}} \overline{\bm S}_{2n}$. Second, the lexicographic order opens the way to the following analogue of Proposition (ref). \begin{thm} If $\bm T_2$ is conditionally elicitable with $\bm T_1$ on $\mathcal{F}$ and $\bm T_1(F)$ is a singleton for all $F\in\mathcal{F}$, then the pair $(\bm T_1,\bm T_2)$ is multi-objective elicitable on $\mathcal{F}$ with respect to the lexicographic order $\preceq_{\mathrm{lex}}$ on $\mathbb{R}^2$. \end{thm} Theorem (ref) is almost a direct analogue of Proposition (ref), with the intriguing exception that $\bm T_1$ is assumed to be a singleton on $\mathcal{F}$ in Theorem (ref); see Remark (ref) for further details. We illustrate how the construction of Theorem (ref) leads to strictly consistent multi-objective scores mapping to $(\mathbb{R}^2, \preceq_{\mathrm{lex}})$ for the cases of (mean, variance) and $(\operatorname{VaR}_\alpha, \operatorname{ES}_\alpha)$ in Subsection (ref). There, we also provide further examples of conditionally elicitable functionals failing to be elicitable in the traditional sense, but which are multi-objective elicitable with respect to $(\mathbb{R}^2, \preceq_{\mathrm{lex}})$. Section (ref) elaborates on how the multi-objective scores can be considered a “generalised antiderivative” of a $k$-dimensional identification function with relaxed symmetry conditions. Section (ref) provides further details on comparing misspecified forecasts under multi-objective elicitability, and on the sensitivity of multi-objective scores with respect to increasing information sets. Finally, Section (ref) revisits a powerful necessary condition for identifiability and elicitability, namely the Convex Level Sets (CxLS) property, for multi-objective scores. \begin{figure} \begin{tikzpicture}[every edge/.style={imp}] \matrix[nodes={inner sep=2mm}, row sep=0.7cm,column sep=0.6cm] { \node (c-id) [rectangle, draw] {Conditional Identifiability}; & & \node (c-el) [rectangle, draw] {Conditional Elicitability}; \\ \\ \node (id) [rectangle, draw] {Identifiability}; & \node (el) [rectangle, draw] {Elicitability}; & \node (mo-el) [rectangle, draw, text width=5cm, align=center] {Multi-Objective Elicitability with respect to $(\mathbb{R}^2, \preceq_{\mathrm{lex}})$}; \\ }; \path (c-id) edge[implies-implies,double equal sign distance] node[above] {1)} (c-el) (c-el) edge[imp] node[right] {Theorem (ref)} (mo-el) (c-id) edge[imp] node[left] {Proposition (ref)} (id) (el) edge[imp] (mo-el) ; \end{tikzpicture} \caption{Illustration of the most important structural results for real-valued (risk) functionals $\bm T_1$ and $\bm T_2$. The equivalence 1) follows from SteinwartPasinETAL2014 under some regularity conditions.} \end{figure} The proof of Theorem (ref) explicitly exploits the asymmetric structure of the lexicographic order which fits well with the asymmetric notion of conditional elicitability, where the roles of $\bm T_1$ and $\bm T_2$ may not be changed. In particular, in the setup of Theorem (ref), the pair $(\bm T_2,\bm T_1)$ is generally not multi-objective elicitable on $\mathcal{F}$ with respect to the lexicographic order $\preceq_{\mathrm{lex}}$ on $\mathbb{R}^2$. More generally, we suspect that under the conditions of Theorem (ref) the lexicographic order is the only order on $\mathbb{R}^2$ which renders $(\bm T_2,\bm T_1)$ multi-objective elicitable. In the rest of the paper, we focus on multi-objective scores mapping to $(\mathbb{R}^2, \preceq_{\mathrm{lex}})$. We use the results of Section (ref), summarised in Figure (ref), extensively to prove our structural results for the systemic risk measures in the next section. \section{Structural results for CoVaR, CoES and MES} \subsection{CoVaR, CoES and MES fail to be identifiable or elicitable} The following proposition shows that the three risk measures $\operatorname{CoVaR}_{\alpha|\beta}$, $\operatorname{CoES}_{\alpha|\beta}$ and $\operatorname{MES}_{\beta}$ generally fail to be identifiable or elicitable on sufficiently rich classes of bivariate distributions $\mathcal{F}\subseteq \mathcal{F}^0(\mathbb{R}^2)$. It is proven by showing that the CxLS property (which remains necessary for multi-objective scores; see Proposition (ref)) is violated in each case. \begin{prop} For $\alpha, \beta\in(0,1)$, $\operatorname{CoVaR}_{\alpha|\beta}$, $\operatorname{CoES}_{\alpha|\beta}$ and $\operatorname{MES}_{\beta}$ are neither identifiable nor elicitable on any class $\mathcal{F}\subseteq \mathcal{F}^0(\mathbb{R}^2)$ containing all bivariate normal distributions along with their mixtures. \end{prop} Proposition (ref) casts doubt on traditional and comparative backtesting approaches for CoVaR, CoES and MES as standalone systemic risk measures. Thus, these measures should not be used for regulatory purposes on their own, because forecasts for them can neither be verified for their adequacy nor can they be sensibly compared to improve their modeling. Proposition (ref) can also be shown on other classes $\mathcal{F}\subseteq \mathcal{F}^0(\mathbb{R}^2)$ that are sufficiently rich, e.g., the classes containing all measures with finite support (see Remark (ref)). Importantly, $\mathcal{F}$ must be convex and must not only consist of distributions with independent marginals. \subsection{Joint identifiability results} Section (ref) establishes the conditional identifiability and conditional elicitability of $\operatorname{CoVaR}_{\alpha|\beta}(Y|X)$, $(\operatorname{CoVaR}_{\alpha|\beta}(Y|X), \operatorname{CoES}_{\alpha|\beta}(Y|X))$, and $\operatorname{MES}_{\beta}(Y|X)$ with $\operatorname{VaR}_\beta(X)$, respectively. This, in combination with Proposition (ref), immediately yields the following joint identifiability results, where for $\alpha,\beta\in(0,1)$ and $p\in\{0\}\cup[1,\infty]$ we use the notation \begin{align}\nonumber \mathcal{F}_{\alpha}^p(\mathbb{R}) &:=\big\{F\in \mathcal{F}^p(\mathbb{R}) \colon F\big(\operatorname{VaR}_\alpha(F)\big)=\alpha\big\},\\ \mathcal{F}_{(\alpha)}^p(\mathbb{R}) &:=\big\{F\in \mathcal{F}_{\alpha}^p(\mathbb{R})\colon F\big(\operatorname{VaR}_\alpha(F) + \varepsilon\big)>\alpha \ \text{for all } \varepsilon>0\big\}, \\ \nonumber \mathcal{F}_{(\beta)}^0(\mathbb{R}^2)&:=\big\{F_{X,Y}\in\mathcal{F}^0(\mathbb{R}^2)\colon F_X \in \mathcal{F}_{(\beta)}^0(\mathbb{R})\big\}. \end{align} \begin{thm} Let $\alpha,\beta\in(0,1)$. Consider the strict $\mathcal{F}_{(\beta)}^0(\mathbb{R}^2)$-identification function for $F_{X,Y}\mapsto \operatorname{VaR}_{\beta}(F_X)$, $V^{\operatorname{VaR}}\colon \mathbb{R}\times\mathbb{R}^2\to\mathbb{R}$, $V^{\operatorname{VaR}}\big(v,(x,y)\big) = \mathds{1}\{x\le v\} - \beta$. \begin{enumerate}[\normalfont (i)] • On $\{F_{X,Y}\in \mathcal{F}^0(\mathbb{R}^2)\colon F_X\in \mathcal{F}^0_{(\beta)}(\mathbb{R}), \ F_{Y|X\ge\operatorname{VaR}_\beta(X)}\in \mathcal{F}^0_{(\alpha)}(\mathbb{R})\}$, the pair $F_{X,Y}\mapsto (\operatorname{VaR}_\beta(F_X),$ $\operatorname{CoVaR}_{\alpha|\beta}(F_{X,Y}) )$ is identifiable with a strict identification function $\bm V^{(\operatorname{VaR}, \operatorname{CoVaR})}\colon \mathbb{R}^2\times \mathbb{R}^2\to\mathbb{R}^2$, \begin{align} \bm V^{(\operatorname{VaR}, \operatorname{CoVaR})}\big((v,c),(x,y)\big) &=\begin{pmatrix} \mathds{1}\{x\le v\} - \beta\\ \mathds{1}\{x>v\} \big[\mathds{1}\{y\le c\} - \alpha\big] \end{pmatrix}. \end{align} • On $\{F_{X,Y}\in \mathcal{F}^0(\mathbb{R}^2)\colon F_X\in \mathcal{F}^0_{(\beta)}(\mathbb{R}), \ F_{Y|X\ge\operatorname{VaR}_\beta(X)}\in \mathcal{F}^1_{(\alpha)}(\mathbb{R})\}$, the triplet $F_{X,Y}\mapsto (\operatorname{VaR}_\beta(F_X), \operatorname{CoVaR}_{\alpha|\beta}(F_{X,Y}), \operatorname{CoES}_{\alpha|\beta}(F_{X,Y}) )$ is identifiable with a strict identification function $\bm V^{(\operatorname{VaR}, \operatorname{CoVaR}, \operatorname{CoES})}\colon \mathbb{R}^3\times \mathbb{R}^2\to\mathbb{R}^3$, \begin{equation} \big((v,c,e),(x,y)\big) \mapsto\begin{pmatrix} \mathds{1}\{x\le v\} - \beta\\ \mathds{1}\{x>v\} \big[\mathds{1}\{y\le c\} - \alpha\big]\\ \mathds{1}\{x>v\} \Big[e - \frac{1}{1-\alpha}\big(y\mathds{1}\{y>c\} + c(\mathds{1}\{y\le c\} - \alpha)\big) \Big] \end{pmatrix}. \end{equation} • On $\{F_{X,Y}\in \mathcal{F}^0(\mathbb{R}^2)\colon F_X\in \mathcal{F}^0_{(\beta)}(\mathbb{R}), \ F_{Y|X\ge\operatorname{VaR}_\beta(X)}\in \mathcal{F}^1(\mathbb{R})\}$, the pair $F_{X,Y}\mapsto (\operatorname{VaR}_\beta(F_X),$ $\operatorname{MES}_{\beta}(F_{X,Y}) )$ is identifiable with a strict identification function $\bm V^{(\operatorname{VaR}, \operatorname{MES})}\colon \mathbb{R}^2\times \mathbb{R}^2\to\mathbb{R}^2$, \begin{align} \bm V^{(\operatorname{VaR}, \operatorname{MES})}\big((v,\mu),(x,y)\big) &=\begin{pmatrix} \mathds{1}\{x\le v\} - \beta\\ \mathds{1}\{x>v\} \big[\mu-y\big] \end{pmatrix}. \end{align} \end{enumerate} \end{thm} Following NZ17, Theorem (ref) can readily be used to assess the absolute forecast quality via joint (\emph{Wald}-type) calibration tests. These either test the null hypothesis of \textit{unconditional} calibration, $\operatorname{E}[\bm V(\bm r_t, (X_t,Y_t))]=\bm 0$ for all $t\in \mathbb N$, or the more informative null of \emph{conditional} calibration, $\operatorname{E}[\bm V(\bm r_t, (X_t,Y_t))\,|\,\mathfrak F_{t-1}]=\bm 0$ for all $t\in \mathbb N$. Here, the $\sigma$-algebra $\mathfrak F_{t-1}$ represents the information available to the forecaster at time $t-1$. Recall that the null of conditional calibration is equivalent to $\operatorname{E}[\varphi_{t-1}\bm V(\bm r_t, (X_t,Y_t))]=\bm 0$ for all $\mathfrak F_{t-1}$-measurable random vectors $\varphi_{t-1}$ for all $t\in \mathbb N$. Unless $\mathfrak F_{t-1}$ is particularly simple, this approach is statistically not feasible, because a finite selection $\varphi_{1,t-1}, \ldots, \varphi_{\ell, t-1}$ of $\mathfrak F_{t-1}$-measurable random vectors is required in practice. Note that using only a finite collection amounts to testing a broader null than the original one of conditional calibration. \begin{rem} It is worth pointing out the improvements upon the substantial work of Banulescu-RaduETAL2019, who develop---inter alia---traditional backtests of systemic risk measures. Translating their notation into ours, they propose in their equation (4) a one-dimensional identification function for the pair $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X))$ of the form $V\colon \mathbb{R}^2\times \mathbb{R}^2\to\mathbb{R}$, \begin{equation} V\big((v,c),(x,y)\big) = \mathds{1}\{x> v\} \mathds{1}\{y> c\} - (1-\alpha)(1-\beta). \end{equation} Banulescu-RaduETAL2019 point out that this an identification function since $ \bar V\big((\operatorname{VaR}_\beta(X),$ $\operatorname{CoVaR}_{\alpha|\beta}(X|Y)),F_{X,Y}\big) =0. $ However, it fails to be strict, since, so long as $(1-\alpha')(1-\beta') = (1-\alpha)(1-\beta)$, we have $\bar V\big((\operatorname{VaR}_{\beta'}(X), \operatorname{CoVaR}_{\alpha'|\beta'}(X|Y)),F_{X,Y}\big)=0$. Thus, the specific null of `correct' unconditional calibration $\operatorname{E}[V(\bm r_t, (X_t, Y_t))]=0$ with $V$ from (ref) is too broad, since it is satisfied for all $\bm r_t=(\operatorname{VaR}_{\beta'}(X),\operatorname{CoVaR}_{\alpha'|\beta'}(X|Y))$ with $(1-\alpha')(1-\beta') = (1-\alpha)(1-\beta)$. This implies only trivial power against such alternatives for the corresponding Wald-test for calibration. This is in stark contrast to the Wald-test employing our two-dimensional identification function in (ref), where $\operatorname{E}[\bm V^{(\operatorname{VaR}, \operatorname{CoVaR})}(\bm r_t, (X_t, Y_t))]=\bm 0$ \textit{if and only if} $\bm r_t=(\operatorname{VaR}_\beta(X),\operatorname{CoVaR}_{\alpha|\beta}(X|Y))$. We refer to Section (ref) for a simulation study illustrating the possible detrimental effects of the non-strictness of (ref). We stress that the non-strictness is not only problematic in the uncountably many cases where $(1-\alpha')(1-\beta') = (1-\alpha)(1-\beta)$, but in \textit{all} cases where VaR is overpredicted and CoVaR underpredicted or the other way around. This is because by overpredicting (underpredicting) VaR and underpredicting (overpredicting) CoVaR, the different biases cancel out in the one-dimensional identification function in (ref), thus leading to a loss of power in identifying misspecified forecasts; see Figure (ref). Intuitively speaking, two moment conditions via a two-dimensional identification function are necessary to uniquely identify the two-dimensional pair $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X))$. This is in line with other two-dimensional strict identification functions for two-dimensional functionals in the literature, e.g. for Value at Risk and Expected Shortfall, see FZ16a, NZ17. \end{rem} \subsection{Multi-objective elicitability results} Recall that $\operatorname{VaR}_{\beta}(Y)$, $(\operatorname{VaR}_\beta(Y), \operatorname{ES}_\beta(Y) )$ and $\operatorname{E}(Y)$ are all elicitable. Surprisingly, their conditional counterparts $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X) )$, $(\operatorname{VaR}_\beta(X),$ $\operatorname{CoVaR}_{\alpha|\beta}(Y|X), \operatorname{CoES}_{\alpha|\beta}(Y|X) )$ and $(\operatorname{VaR}_\beta(X), \operatorname{MES}_\beta(Y|X) )$ fail to be elicitable despite being identifiable (Section (ref)). This is due to integrability conditions, causing an extreme gap between the class of strict identification functions and the class of strictly consistent scoring functions, which turns out to be empty. Thus, comparative backtests cannot be implemented using a scalar strictly consistent scoring function. Furthermore, the conditional elicitability results of Section (ref) can hardly be used for forecast comparisons, unless the $\operatorname{VaR}_\beta(X)$ forecasts are the same and correctly specified. However, the conditional elicitability in combination with Theorem (ref) immediately yields the following novel joint multi-objective elicitability results with respect to the lexicographic order $\preceq_{\mathrm{lex}}$ on $\mathbb{R}^2$. These results can readily be used for comparing systemic risk forecasts as detailed in Section (ref). We again use the notation introduced in (ref). \begin{thm} Let $\alpha,\beta\in(0,1)$. \begin{enumerate}[\normalfont (i)] • On $\mathcal{F} \subseteq \mathcal{F}_{\beta}^0(\mathbb{R}^2)$, the score $S^{\operatorname{VaR}}\colon \mathbb{R}\times\mathbb{R}^2\to\mathbb{R}$, \begin{equation} S^{\operatorname{VaR}}\big(v,(x,y)\big) = \big(\mathds{1}\{x\le v\} - \beta\big)h(v) - \mathds{1}\{x\le v\}h(x) +a^{\operatorname{VaR}}(x,y) \end{equation} is strictly $\mathcal{F}$-consistent for $\mathcal{F}\ni F_{X,Y}\mapsto \operatorname{VaR}_{\beta}(F_X)$, if $h\colon\mathbb{R}\to\mathbb{R}$ is strictly increasing and for all $v\in\mathbb{R}$, the function $(x,y)\mapsto a^{\operatorname{VaR}}(x,y)-\mathds{1}\{x\le v\}h(x)$ is $\mathcal{F}$-integrable. • On $\mathcal{F}\subseteq \{F_{X,Y}\in \mathcal{F}^0(\mathbb{R}^2)\colon F_X\in \mathcal{F}^0_{(\beta)}(\mathbb{R}), \ F_{Y|X\ge\operatorname{VaR}_\beta(X)}\in \mathcal{F}^0_{\alpha}(\mathbb{R})\}$, the pair $\mathcal{F}\ni F_{X,Y}\mapsto (\operatorname{VaR}_\beta(F_X), \operatorname{CoVaR}_{\alpha|\beta}(F_{X,Y}))$ is multi-objective elicitable with respect to $(\mathbb{R}^2, \preceq_{\mathrm{lex}})$. A strictly $\mathcal{F}$-consistent multi-objective scoring function $\bm S^{(\operatorname{VaR}, \operatorname{CoVaR})}\colon \mathbb{R}^2\times \mathbb{R}^2\to(\mathbb{R}^2, \preceq_{\mathrm{lex}})$ is given by \begin{align} \bm S^{(\operatorname{VaR}, \operatorname{CoVaR})}\big((v,c),(x,y)\big) &=\begin{pmatrix} S^{\operatorname{VaR}}\big(v,(x,y)\big)\\ S_v^{\operatorname{CoVaR}}\big(c,(x,y)\big) \end{pmatrix}, \\[0.5em] \nonumber S_v^{\operatorname{CoVaR}}\big(c,(x,y)\big) &= \mathds{1}\{x>v\} \Big[\big(\mathds{1}\{y\le c\} - \alpha\big) g(c) - \mathds{1}\{y\le c\}g(y) + a(y)\Big] \\ \nonumber &\quad + a^{\operatorname{CoVaR}}(x,y), \end{align} where $g\colon\mathbb{R}\to\mathbb{R}$ is strictly increasing and $S_v^{\operatorname{CoVaR}}$ is $\mathcal{F}$-integrable for all $v\in\mathbb{R}$. \item On $\mathcal{F}\subseteq \{F_{X,Y}\in \mathcal{F}^0(\mathbb{R}^2)\colon F_X\in \mathcal{F}^0_{(\beta)}(\mathbb{R}), \ F_{Y|X\ge\operatorname{VaR}_\beta(X)}\in \mathcal{F}^1_{\alpha}(\mathbb{R})\}$, the triplet $\mathcal{F}\ni F_{X,Y}\mapsto (\operatorname{VaR}_\beta(F_X), \operatorname{CoVaR}_{\alpha|\beta}(F_{X,Y}), \operatorname{CoES}_{\alpha|\beta}(F_{X,Y}))$ is multi-objective elicitable with respect to $(\mathbb{R}^2, \preceq_{\mathrm{lex}})$. A strictly $\mathcal{F}$-consistent multi-objective scoring function $\bm S^{(\operatorname{VaR}, \operatorname{CoVaR}, \operatorname{CoES})}\colon \mathbb{R}^3\times \mathbb{R}^2\to(\mathbb{R}^2, \preceq_{\mathrm{lex}})$ is given by \begin{align}\label{eq:S_CoES} &\bm S^{(\operatorname{VaR}, \operatorname{CoVaR}, \operatorname{CoES})}\big((v,c,e),(x,y)\big) =\begin{pmatrix} S^{\operatorname{VaR}}\big(v,(x,y)\big)\\ S_v^{(\operatorname{CoVaR}, \operatorname{CoES})} \big((c,e),(x,y)\big) \end{pmatrix},\\[0.5em] \nonumber &S_v^{(\operatorname{CoVaR}, \operatorname{CoES})}\big((c,e),(x,y)\big) = \mathds{1}\{x>v\}\Big[\big(\mathds{1}\{y\le c\} - \alpha\big) g(c) - \mathds{1}\{y\le c\}g(y) \\ \nonumber & \quad +\phi'(e)\Big(e - \tfrac{1}{1-\alpha}\big(y\mathds{1}\{y>c\} + c(\mathds{1}\{y\le c\} - \alpha)\big)\Big) - \phi(e)+ a(y) \Big] + a^{\operatorname{CoES}}(x,y), \end{align} where $g\colon\mathbb{R}\to\mathbb{R}$ is increasing, $\phi\colon\mathbb{R}\to\mathbb{R}$ is strictly convex with subgradient $\phi'<0$, and $S_v^{(\operatorname{CoVaR}, \operatorname{CoES})}$ is $\mathcal{F}$-integrable for all $v\in\mathbb{R}$. \item On $\mathcal{F}\subseteq\{F_{X,Y}\in \mathcal{F}^0(\mathbb{R}^2)\colon F_X\in \mathcal{F}^0_{(\beta)}(\mathbb{R}), \ F_{Y|X\ge\operatorname{VaR}_\beta(X)}\in \mathcal{F}^1(\mathbb{R})\}$, the pair $F_{X,Y}\mapsto (\operatorname{VaR}_\beta(F_X), \operatorname{MES}_{\beta}(F_{X,Y}))$ is multi-objective elicitable with respect to ${(\mathbb{R}^2, \preceq_{\mathrm{lex}})}$. A strictly $\mathcal{F}$-consistent multi-objective scoring function $\bm S^{(\operatorname{VaR}, \operatorname{MES})}\colon \mathbb{R}^2\times \mathbb{R}^2\to(\mathbb{R}^2, \preceq_{\mathrm{lex}})$ is given by \begin{align}\label{eq:S_MES} \bm S^{(\operatorname{VaR}, \operatorname{MES})}\big((v,\mu),(x,y)\big) &=\begin{pmatrix} S^{\operatorname{VaR}}\big(v,(x,y)\big)\\ S_v^{\operatorname{MES}} \big(\mu,(x,y)\big) \end{pmatrix}, \\[0.5em] \nonumber S_v^{\operatorname{MES}}\big(\mu,(x,y)\big) &= \mathds{1}\{x>v\} \Big[\phi'(\mu)(\mu - y) - \phi(\mu) + a(y)\Big] + a^{\operatorname{MES}}(x,y), \end{align} where $\phi\colon\mathbb{R}\to\mathbb{R}$ is strictly convex with subgradient $\phi'$ and $S_v^{\operatorname{MES}}$ is $\mathcal{F}$-integrable for all $v\in\mathbb{R}$. \end{enumerate} \end{thm} The choice of the functions $a, a^{\operatorname{VaR}}, a^{\operatorname{CoVaR}}, a^{\operatorname{CoES}}, a^{\operatorname{MES}}$ in Theorem \ref{thm:mo el} is inessential. It may only influence the integrability of the scoring functions on the one hand and it can control the sign of the scores on the other hand. For the remaining functions in Theorem \ref{thm:mo el}, one could choose the standard choices, e.g., the identity for $h$ and $g$ in \eqref{eq:Score VaR}, \eqref{eq:S_CoVaR} and \eqref{eq:S_CoES}, giving rise to the common `pinball loss', or $\phi(y) = y^2$ in \eqref{eq:S_MES} leading to the square loss. For further possible choices, especially for $\phi$ in \eqref{eq:S_CoES}, we refer to Subsection \ref{Description of Tests}. Table~\ref{tab:overview} summarises the results of Section~\ref{Structural results for CoVaR, CoES and MES}. For purposes of comparison, the final three rows display the properties of $\operatorname{VaR}_{\beta}(Y)$ and $\operatorname{ES}_{\beta}(Y)$. The multi-objective elicitability of the pair $(\operatorname{VaR}_\beta(Y), \operatorname{ES}_\beta(Y) )$ follows from Example \ref{example: lex functionals b}. Table~\ref{tab:overview} highlights the structural difference between $\operatorname{VaR}_{\beta}(Y)$, $\operatorname{E}(Y)$, $(\operatorname{VaR}_\beta(Y), \operatorname{ES}_\beta(Y) )$ on the one hand and their conditional counterparts $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X) )$, $(\operatorname{VaR}_\beta(X), \operatorname{MES}_\beta(Y|X) )$, and $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X),$ $\operatorname{CoES}_{\alpha|\beta}(Y|X))$ on the other hand. While elicitability in the usual sense holds for the former, it fails for the latter. This highlights the importance of the newly introduced concept of multi-objective elicitability, which allows for comparative backtests. Section~\ref{sec:tests} shows how to implement comparative backtests with the multi-objective scores of Theorem~\ref{thm:mo el}. \begin{table} \caption{\label{tab:overview}Overview of properties of (systemic) risk measures. \ding{51} (\ding{55}) indicates that property does (does not) apply.} \centering \small \begin{tabular}{lccc} \toprule Risk Measure & Identifiability & Elicitability & Multi-objective \\ & & &Elicitability \\ \midrule $\operatorname{CoVaR}_{\alpha|\beta}(Y|X)$ & \ding{55} & \ding{55} & \ding{55} \\ $\operatorname{CoES}_{\alpha|\beta}(Y|X)$ & \ding{55} & \ding{55} & \ding{55} \\ $\operatorname{MES}_{\alpha|\beta}(Y|X)$ & \ding{55} & \ding{55} & \ding{55} \\ \midrule $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X) )$ & \ding{51} & \ding{55} & \ding{51} \\ $(\operatorname{VaR}_\beta(X), \operatorname{CoVaR}_{\alpha|\beta}(Y|X), \operatorname{CoES}_{\alpha|\beta}(Y|X) )$ & \ding{51} & \ding{55} & \ding{51} \\ $(\operatorname{VaR}_\beta(X), \operatorname{MES}_\beta(Y|X) )$ & \ding{51} & \ding{55} & \ding{51} \\ \midrule $\operatorname{VaR}_{\alpha}(Y)$ & \ding{51} & \ding{51} & \ding{51} \\ $\operatorname{ES}_{\alpha}(Y)$ & \ding{55} & \ding{55} & \ding{55} \\ $(\operatorname{VaR}_\alpha(Y), \operatorname{ES}_\alpha(Y) )$ & \ding{51} & \ding{51} & \ding{51} \\ \bottomrule \end{tabular}\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\qquad\ \end{table} \section{Diebold--Mariano tests for multi-objective scores} \label{sec:tests} \subsection{Two-sided tests}\label{Two-Sided Tests} \citet{DM95} propose to use formal hypothesis tests to account for sampling uncertainty in forecast comparisons. These so-called Diebold--Mariano (DM) tests are widely used in empirical forecast comparisons and continue to be studied in the theoretical literature. However, consistent with the extant notion of strict consistency, DM tests have hitherto relied on scalar scoring functions. Thus, here we show how to use our two-dimensional multi-objective scores from Theorem~\ref{thm:mo el} in DM tests, with a special focus on the implications caused by the lexicographic order. To that end, denote by $\bm S=(S_1, S_2)^\prime$ one of the multi-objective scores of Theorem~\ref{thm:mo el}. Let $\{\bm r_{t,(1)}\}_{t=1,\ldots,n}$ and $\{\bm r_{t,(2)}\}_{t=1,\ldots,n}$ be the appertaining competing sequences of forecasts (e.g., if $\bm S=\bm S^{(\operatorname{VaR},\operatorname{CoVaR})}$, then $\bm r_{t,(i)}=(\widehat{\operatorname{VaR}}_{t,(i)},\ \widehat{\operatorname{CoVaR}}_{t,(i)})$ for $i=1,2$). The verifying observations are $\{(X_t, Y_t)\}_{t=1,\ldots,n}$. We compare the two forecasts via the (bivariate) score differences $\bm d_t:=(d_{1t}, d_{2t})^\prime:=\bm S(\bm r_{t,(1)}, (X_t, Y_t))-\bm S(\bm r_{t,(2)}, (X_t, Y_t))$. The two-sided null hypothesis is that both forecasts predict equally well on average, i.e., $H_0^{=} \colon \operatorname{E}[\overline{\bm d}_n]=\bm 0$ for all $n=1,2,\ldots$, where $\overline{\bm d}_n:=(\overline{d}_{1n}, \overline{d}_{2n})^\prime:=\frac{1}{n}\sum_{t=1}^{n}\bm d_t$. (Along the lines of \cite{GW06}, one can also test the conditional null hypothesis $H_0^{*=} \colon \operatorname{E}[ \bm d_t\mid \mathfrak F_{t-1}]=\bm 0$ for all $t=1,2,\ldots$, where the $\sigma$-algebra $\mathfrak F_{t-1}$ contains all information available at time $t-1$.) We test $H_0^{=}$ using the Wald-type test statistic \begin{equation} \label{eq:T_n} \mathcal{T}_n=n \overline{\bm d}_n^\prime\widehat{\bm \varOmega}_n^{-1} \overline{\bm d}_n, \end{equation} where $\widehat{\bm \varOmega}_n$ is some consistent estimator of the variance-covariance matrix $\bm \varOmega_n=\operatorname{Var}(\sqrt{n}\overline{\bm d}_n)$ under the null hypothesis (i.e., in the componentwise norm, $\|\widehat{\bm \varOmega}_n-\bm \varOmega_n\|\overset{\operatorname{P}}{\longrightarrow}0$, as $n\to\infty$). To account for possible autocorrelation in the sequence $\{\bm d_t\}_{t=1,\ldots,n}$, one can use \begin{multline}\label{eq:hat Omega_n} \widehat{\bm \varOmega}_n=\begin{pmatrix}\widehat{\sigma}_{11,n} & \widehat{\sigma}_{12,n} \\ \widehat{\sigma}_{12,n} & \widehat{\sigma}_{22,n}\end{pmatrix}=\frac{1}{n}\sum_{t=1}^{n}(\bm d_t-\overline{\bm d}_n)(\bm d_t-\overline{\bm d}_n)^\prime \\ + \frac{1}{n}\sum_{h=1}^{m_n}w_{n,h}\sum_{t=h+1}^{n}\Big[(\bm d_t-\overline{\bm d}_n)(\bm d_{t-h}-\overline{\bm d}_n)^\prime + (\bm d_{t-h}-\overline{\bm d}_n)(\bm d_t-\overline{\bm d}_n)^\prime\Big], \end{multline} where $m_n\rightarrow\infty$ is a sequence of integers satisfying $m_n=o(n^{1/4})$, and $w_{n,h}$ is a uniformly bounded scalar triangular array with $w_{n,h}\rightarrow1$, as $n\to\infty$, for all $h=1,\ldots,m_n$; see \citet{Whi01} for detail. Under the assumption that $\{\bm d_t\}_{t=1,\ldots,n}$ does not exhibit autocorrelation under the null, $m_n$ can be set to 0 such that $\widehat{\bm \varOmega}_n$ is simply the sample variance-covariance matrix. \begin{thm}\label{thm:DM} Suppose that $\bm \varOmega_n\longrightarrow\bm \varOmega$, as $n\to\infty$, where $\bm \varOmega$ is positive definite. Then, under technical Assumption~\ref{ass:DM} (see Supplement Section~\ref{app:Proofs}), it holds under $H_0^{=}$ that \[ \sqrt{n}\widehat{\bm \varOmega}_n^{-1/2}\overline{\bm d}_{n}\overset{d}{\longrightarrow}N(\bm 0, \bm I_{2\times2}),\qquad\text{as }n\to\infty, \] where $\bm I_{2\times2}$ denotes the $(2\times2)$-identity matrix. In particular, $\mathcal{T}_n\overset{d}{\longrightarrow}\chi_{2}^{2}$, as $n\to\infty$, where $\chi_2^{2}$ denotes a $\chi^2$-distribution with 2 degrees of freedom. \end{thm} Thus, we reject $H_0^{=}$ at significance level $\nu$, if $\mathcal{T}_n>\chi_{2,1-\nu}^2$, where $\chi_{2,1-\nu}^2$ is the $(1-\nu)$-quantile of the $\chi_2^2$-distribution. A typical non-rejection region in terms of $\overline{d}_{1n}$ and $\overline{d}_{2n}$ is sketched in Figure (ref) (a), and has the well-known ellipse shape. For brevity, we leave out a formal investigation of our test under the alternative. We mention, however, that consistency of our test can be established along the lines of GW06. \begin{figure} \caption{Non-rejection regions for two-sided DM test in (a), lexicographic DM test in (b), and one and a half-sided DM test in (c).} \end{figure} \subsection{One and a half-sided tests} In several contexts, it may be desirable to perform a one-sided comparative backtest to establish the superiority of risk forecasts $\{\bm r_{t,(2)}\}_{t=1,\ldots,n}$ over some benchmark forecasts $\{\bm r_{t,(1)}\}_{t=1,\ldots,n}$. These different forecasts could stem from two different internal models of a financial institution (say, a legacy model and and extension thereof). It could also be the case that the benchmark forecasts originate from a regulatory standard model, and---in line with the conservative backtesting approach of FZG16---the financial institution has the onus of proof to show the superiority of its internal model over the standard model. In such situations, it is tempting to test the null hypothesis $\operatorname{E}[\overline{\bm d}_n]\preceq_{\mathrm{lex}}\bm 0$, which is equivalent to \begin{equation} \operatorname{E}[\overline{d}_{1n}]<0 \qquad\text{or}\qquad\Big(\operatorname{E}[\overline{d}_{1n}]=0\quad\text{and}\quad\operatorname{E}[\overline{d}_{2n}]\leq0 \Big). \end{equation} Here, the goal would be to reject the null of better or, at least, equally good benchmark forecasts as evidence of the superiority of $\{\bm r_{t,(2)}\}_{t=1,\ldots,n}$. However, as under standard conditions (see Theorem (ref)), $\sqrt{n}(\overline{d}_{1n}, \overline{d}_{2n})^\prime$ has an asymptotic bivariate normal distribution, the probability that $\overline{d}_{1n}=0$ and $\overline{d}_{2n}\leq0$ vanishes for large sample sizes. Thus, testing (ref) amounts to a test of the null $\operatorname{E}[\overline{d}_{1n}]<0$, i.e., that the benchmark VaR forecasts are superior. This null can be tested via $\mathcal{T}_{1n}=\sqrt{n}\overline{d}_{1n}/\widehat{\sigma}_{11,n}^{1/2}$, where we reject (ref) at significance level $\nu\in(0,1)$ when $\mathcal{T}_{1n}>\Phi^{-1}(1-\nu)$. The corresponding non-rejection region is sketched in Figure (ref) (b). In particular, we would reject $\operatorname{E}[\overline{\bm d}_n]\preceq_{\mathrm{lex}}\bm 0$ solely based on the predictive performance of the $\operatorname{VaR}(X_t)$ component. In other words, once $\operatorname{E}[\overline{d}_{1n}]<0$ is rejected such that the internal VaR forecasts are superior, the internal forecasts $\bm r_{t,(1)}$ are preferable in the lexicographic order \textit{irrespective} of the quality of the systemic risk component. Since this would entirely ignore the motivation of the backtest, we disregard this one-sided backtesting approach. As a compromise, we suggest to test the following \textit{`one and a half-sided'} null hypothesis. Since the two systemic risk forecasts only play a role for the ranking in the lexicographic order when $\operatorname{E}[\overline{d}_{1n}]=0$, our suggested null hypothesis takes the form \begin{equation*} H_0^{\preceq_{\mathrm{lex}}} \colon \operatorname{E}[\overline{d}_{1n}]=0\quad\text{and}\quad\operatorname{E}[\overline{d}_{2n}]\leq0\qquad\text{for all }n=1,2,\ldots. \end{equation*} We demonstrate below how to interpret a rejection of this null. The region in $\mathbb{R}^2$ pertaining to $H_0^{\preceq_{\mathrm{lex}}}$ is the lower part of the vertical axis in Figure (ref) (c). Obviously, $H_0^{\preceq_{\mathrm{lex}}}$ is the union of all $H_0^{(c)}$ with $c\le0$, where \[ H_0^{(c)} \colon \operatorname{E}[\overline{d}_{1n}]=0\quad\text{and}\quad \operatorname{E}[\overline{d}_{2n}]=c\qquad\text{for all }n=1,2,\ldots \] For each individual $c\le0$, this can be tested using the Wald-type test statistic $\mathcal{T}_{n}^{(c)}$, where $\mathcal{T}_{n}^{(c)}$ is defined similarly as $\mathcal{T}_n$, only with $\bm d_t$ replaced by $\bm d_t^{(c)}=(d_{1t}, d_{2t}-c)^\prime$. (Note that this substitution leaves $\widehat{\bm \varOmega}_n$ unaffected.) Thus, we reject $H_0^{\preceq_{\mathrm{lex}}}$ if and only if \begin{equation} \mathcal{T}_n^{(c)}> \chi_{2,1-\widetilde{\nu}}^2\qquad\text{for all }c\leq0, \end{equation} where $\widetilde{\nu}\in(0,1)$. The area associated with the appertaining non-rejection region is shaded in pink in Figure (ref) (c). The rejection condition (ref) is of course equivalent to \( \mathcal{T}_n^{\operatorname{OS}} := \inf_{c\le 0} \mathcal{T}_n^{(c)} = \mathcal{T}_n^{(c^*)} >\chi_{2,1-\widetilde{\nu}}^2, \) where the solution $c^*:= \min\big\{0, \overline{d}_{2n} - (\widehat{\sigma}_{12,n}/\widehat{\sigma}_{11,n})\overline{d}_{1n}\big\}$ follows from a simple quadratic minimisation problem. Hence, \begin{equation} \mathcal{T}_n^{\operatorname{OS}}=n\Big(\overline{d}_{1n}, \max\big\{\overline{d}_{2n}, (\widehat{\sigma}_{12,n}/\widehat{\sigma}_{11,n})\overline{d}_{1n}\big\}\Big)\widehat{\bm \varOmega}_n^{-1}\begin{pmatrix}\overline{d}_{1n}\\ \max\big\{\overline{d}_{2n}, (\widehat{\sigma}_{12,n}/\widehat{\sigma}_{11,n})\overline{d}_{1n}\big\}\end{pmatrix}. \end{equation} To illustrate this graphically, note that the line defined by $z_{2}=(\widehat{\sigma}_{12,n}/\widehat{\sigma}_{11,n})z_{1}$---indicated by the dashed line in Figure (ref) (c)---passes through the extremal (negative and positive) horizontal points of the ellipse. Thus, if $\overline{d}_{2n}>(\widehat{\sigma}_{12,n}/\widehat{\sigma}_{11,n})\overline{d}_{1n}$, (ref) parametrizes the upper half of the tilted ellipse, and if $\overline{d}_{2n}\leq(\widehat{\sigma}_{12,n}/\widehat{\sigma}_{11,n})\overline{d}_{1n}$ the area below. In our numerical experiments, we use $\mathcal{T}_n^{\operatorname{OS}}$ to test $H_0^{\preceq_{\mathrm{lex}}}$. The next proposition shows that rejecting $H_0^{\preceq_{\mathrm{lex}}}$ when $\mathcal{T}_n^{\operatorname{OS}}>\chi_{2,1-\widetilde{\nu}}^2$ leads to a test of level $\nu=1/2\big[1+\widetilde{\nu}-F_{\chi_1^2}(\chi_{2,1-\widetilde{\nu}}^{2})\big]$, where $F_{\chi_1^2}$ denotes the cdf of a $\chi_{1}^{2}$-distribution. \begin{prop} Under the conditions of Theorem (ref), rejecting $H_0^{\preceq_{\mathrm{lex}}}$ if $\mathcal{T}_n^{\operatorname{OS}}>\chi_{2,1-\widetilde{\nu}}^2$ leads to an asymptotic size $\nu$-test with $\nu=1/2\big[1+\widetilde{\nu}-F_{\chi_1^2}(\chi_{2,1-\widetilde{\nu}}^{2})\big]$. That is, \[ \sup_{c\leq0}\lim_{n\to\infty}\operatorname{P}\Big\{H_0^{\preceq_{\mathrm{lex}}}\ \text{is rejected based on}\ \mathcal{T}_{n}^{\operatorname{OS}}\ \Big\vert\ H_0^{(c)}\ \text{holds}\Big\}=\nu. \] \end{prop} \begin{rem} For a test of $H_0^{\preceq_{\mathrm{lex}}}$ with desired significance level of $\nu$, one can determine the required level $\widetilde{\nu}$ from $\nu=1/2\big[1+\widetilde{\nu}-F_{\chi_1^2}(\chi_{2,1-\widetilde{\nu}}^{2})\big]$ using standard root-finding algorithms. E.g., if $\nu=1\%$ / $\nu=5\%$ / $\nu=10\%$, then $\widetilde{\nu}=1.60\%$ / $\widetilde{\nu}=7.66\%$ / $\widetilde{\nu}=14.9\%$. \end{rem} \begin{rem} To compare systemic risk forecasts on an equal footing, it may be desirable to compare them based on the same marginal model being used and, hence, based on the same VaR forecasts $\widehat{\operatorname{VaR}}_t=\widehat{\operatorname{VaR}}_{t,(1)}=\widehat{\operatorname{VaR}}_{t,(2)}$. In this case, differences in predictive ability can be attributed solely to the different dependence models; see, e.g., NZ19+ and Hog20a+. (Note that Theorem (ref) no longer applies for identical VaR forecasts, since the limit of the covariance matrix $\bm \varOmega$ is only positive semi-definite.) When $\widehat{\operatorname{VaR}}_t=\widehat{\operatorname{VaR}}_{t,(1)}=\widehat{\operatorname{VaR}}_{t,(2)}$, testing for $\operatorname{E}[\overline{d}_{1n}]=0$ is redundant, and $H_0^{=}$ and $H_0^{\preceq_{\mathrm{lex}}}$ reduce to $\operatorname{E}[\overline{d}_{2n}]=0$ and $\operatorname{E}[\overline{d}_{2n}]\leq0$, respectively. These hypotheses can be tested using $\mathcal{T}_{2n}=\sqrt{n}\overline{d}_{2n}/\widehat{\sigma}_{22,n}^{1/2}$. Under $\operatorname{E}[\overline{d}_{2n}]=0$, $\mathcal{T}_{2n}$ converges to an $N(0,1)$-limit when restricting the assumptions of Theorem (ref) to the sequence $d_{2t}$. Hence, we would reject $\operatorname{E}[\overline{d}_{2n}]=0$ (or $\operatorname{E}[\overline{d}_{2n}]\leq0$) at significance level $\nu\in(0,1)$, if $|\mathcal{T}_{2n}|>\Phi^{-1}(1-\nu/2)$ (or $\mathcal{T}_{2n}>\Phi^{-1}(1-\nu)$). \end{rem} \begin{figure} \caption{Interpretation of test decisions. Yellow zone corresponds to ellipse. Red (grey) zone corresponds to the rectangle to the left (right) of the ellipse. Green (orange) zone corresponds to area immediately above (below) the ellipse.} \end{figure} Figure (ref) presents a schematic for interpreting test decisions on $H_0^{\preceq_{\mathrm{lex}}}$ depending on the values of $\overline{d}_{1n}$ and $\overline{d}_{2n}$. In a regulatory context, it extends the three-zone traffic-light classification of the BCBSBF19: \begin{quote} { “The green zone corresponds to backtesting results that do not themselves suggest a problem with the quality or accuracy of a bank's model. The yellow zone encompasses results that do raise questions in this regard, but where such a conclusion is not definitive. The red zone indicates a backtesting result that almost certainly indicates a problem with a bank's risk model.”} \end{quote} The union of the yellow and the orange areas in Figure (ref) corresponds to the non-rejection region depicted in Figure (ref) (c). By symmetry, the union of the yellow and the green areas is the non-rejection region associated with the null \[ H_0^{\succeq_{\mathrm{lex}}} \colon \operatorname{E}[\overline{d}_{1n}]=0\quad\text{and}\quad\operatorname{E}[\overline{d}_{2n}]\geq0\qquad\text{for all }n=1,2,\ldots. \] Thus, for $(\overline{d}_{1n},\overline{d}_{2n})'$ in the yellow area we can neither reject $H_0^{\preceq_{\mathrm{lex}}}$ nor $H_0^{\succeq_{\mathrm{lex}}}$ (at significance level $\nu=1/2\big[1+\widetilde{\nu}-F_{\chi_1^2}(\chi_{2,1-\widetilde{\nu}}^{2})\big]$, respectively). This implies that there is no evidence of differences in predictive ability between the internal and the benchmark model (at level $\widetilde{\nu}$). From a regulatory perspective, the bank's internal model is `at the boundary' of what can be deemed acceptable. Hence, we suggest heightened attention and close monitoring by the regulator, as indicated by the yellow colour. When $(\overline{d}_{1n},\overline{d}_{2n})'$ falls into the orange area, then---while the VaR forecasts are of comparable quality---the benchmark model provides superior systemic risk forecasts. Here, as indicated by the orange colour, a revision of the internal systemic risk forecasts is called for, while the internal marginal model is in order. In contrast, the green area indicates superior systemic risk forecasts of the internal model, with VaR forecasts being comparably accurate. In this case, the financial institution should be allowed to use its internal model for risk forecasting purposes. The green colour indicates a pass of the backtest. In the red region, the VaR forecasts of the internal model are deemed inferior to the benchmark predictions, whence there is no basis to compare the systemic risk forecasts. (The red region corresponds to the rejection region of the null $\operatorname{E}[\overline{d}_{1n}]\ge0$ at level $\nu'=\frac12 - \frac12 F_{\chi_1^2}(\chi_{2,1-\widetilde{\nu}}^{2})$. E.g., for $\nu = 5\%$ we get $\widetilde \nu = 7.66\%$ and $\nu'= 1.17\%$.) Here, the bank should not be allowed to use its own marginal model (the traffic light is red), but instead should be required to use the benchmark model for the marginals to ensure a fair assessment of the systemic risk forecasts. For this comparison, where the VaR forecasts are identical, one can focus solely on the systemic risk component by using $\mathcal{T}_{2n}$; see Remark (ref). For the comparison via $\mathcal{T}_{2n}$, we suggest to adopt a similar decision heuristic as in FZG16: Neither $H_0^{\preceq_{\mathrm{lex}}}$ nor $H_0^{\succeq_{\mathrm{lex}}}$ can be rejected when $|\mathcal{T}_{2n}|\leq \Phi^{-1}(1-\nu)$ (corresponding to our yellow zone). When $\mathcal{T}_{2n}> \Phi^{-1}(1-\nu)$ ($\mathcal{T}_{2n}< \Phi^{-1}(\nu)$) such that $H_0^{\preceq_{\mathrm{lex}}}$ ($H_0^{\succeq_{\mathrm{lex}}}$) can be rejected, the internal systemic risk forecasts are superior (inferior), corresponding to our green (red) zone. Similarly as for the red region, there are no grounds for meaningful systemic risk forecast comparisons in the grey area, since the internal model's VaR forecasts are superior. (Formally, it corresponds to the rejection region of the null $\operatorname{E}[\overline{d}_{1n}]\le0$ at level $\nu'=\frac12 - \frac12 F_{\chi_1^2}(\chi_{2,1-\widetilde{\nu}}^{2})$.) In this case, there is no cause for action on the end of the bank (hence the grey colour), but the regulator should rather adopt the bank's VaR model as a basis for comparing the systemic risk forecasts. As before, the subsequent comparison of systemic risk forecasts should be carried out by using $\mathcal{T}_{2n}$, since the VaR forecasts are identical. Section (ref) investigates the finite-sample properties of $\mathcal{T}_{n}$, $\mathcal{T}_n^{\operatorname{OS}}$, and $\mathcal{T}_{2n}$ under $H_0^{=}$ and $H_0^{\preceq_{\mathrm{lex}}}$ in detail. Here, we only summarize the main findings. First, size is adequate already for $n=500$, which is encouraging since effective sample sizes in risk forecast comparisons are small. Second, power increases markedly in $n$. Third, comparisons for (CoVaR, CoES) are slightly more powerful than those for CoVaR alone, most likely due to the increased informational content of the CoES component. Fourth, as expected for one-sided tests, departures from $H_0^{\preceq_{\mathrm{lex}}}$ and $H_0^{=}$ are easier to detect for the former. Fifth, it is in general easier to detect differences in predictive ability of the systemic risk component, when the VaR component is identical (instead of only comparable) across forecasts. Intuitively, the inclusion of comparable VaR forecasts dilutes the power of the test in the systemic risk component. While the schematic in Figure (ref) is motivated by a regulatory framework, we stress that it can be used in the context of any comparative backtest between different models. Such a general example is provided in Section (ref). \section{Empirical application} Consider daily log-losses $X_{-r+1},\ldots,X_n$ on the S&P 500 and log-losses $Y_{-r+1},\ldots,Y_n$ on the DAX 30 from 2000--2020, where the data are taken from \href{https://www.wsj.com/market-data/quotes}{www.wsj.com/market-data/quotes} (ticker symbols: SPX and DX:DAX). So if $P_{Z,t}$ denotes the stock index value at time $t$, then $Z_t=-\log(P_{Z,t}/P_{Z,t-1})$ ($Z\in\{X,Y\}$). We only keep those observations where data on both indexes are available, giving us $n+r=5\,193$ observations. Here, for $\alpha=\beta=0.95$, we compare rolling-window (VaR, CoVaR, CoES) forecasts for the series $\{(X_t,Y_t) \}_{t=1,\ldots,n}$, where $r=1\,000$ denotes the moving window length. The choice of $X$ and $Y$ amounts to considering the risk for large losses of the DAX 30, given that the world's leading stock index---the S&P 500---is in distress. To promote flow in this section, we often refer to the simulation setup in Section (ref) for details on the time series models and the risk forecast computation. For short-term risk management purposes, \emph{conditional} (systemic) risk measure forecasts are more informative than unconditional ones. Conditional risk measures are based on the conditional distribution of $(X_t, Y_t)$, that is, $F_{(X_t,Y_t) \mid \mathfrak F_{t-1}}(x,y) = \operatorname{P}\{X_t\le x, Y_t\le y \mid \mathfrak F_{t-1}\} =: \operatorname{P}_{t-1}\{X_t\le x, Y_t\le y\}$ for $x,y\in\mathbb{R}$. Here, the filtration $\mathfrak F_{t-1}$ is generated by the information available to a forecaster at time $t-1$. These are usually past observations $(X_{t-1},Y_{t-1}), (X_{t-2}, Y_{t-2}), \ldots$, and possibly additional exogenous information. Here, we assume $\mathfrak{F}_{t-1} = \sigma\big((X_{t-1},Y_{t-1}), (X_{t-2}, Y_{t-2}), \ldots\big)$, such that we forecast the \emph{conditional} risk measures $\operatorname{VaR}_{t}(X_t) = \operatorname{VaR}_{\beta}(F_{X_t\mid\mathfrak F_{t-1}})$, $\operatorname{CoVaR}_{t}(Y_t\vert X_t) = \operatorname{CoVaR}_{\alpha|\beta}(F_{(X_t,Y_t) \mid \mathfrak F_{t-1}})$, and $\operatorname{CoES}_{t}(Y_t\vert X_t) = \operatorname{CoES}_{\alpha|\beta}(F_{(X_t,Y_t) \mid \mathfrak F_{t-1}})$. For notational brevity, we suppress the dependence of the risk measures on the risk levels $\alpha$ and $\beta$, which we fix at $\alpha=\beta=0.95$. We consider two different methods for $(\operatorname{VaR}_{t}(X_t), \operatorname{CoVaR}_{t}(Y_t\vert X_t), \operatorname{CoES}_{t}(Y_t\vert X_t))$ forecasting. The first method uses a simple GARCH(1,1) model for $X_t$ and $Y_t$ and---as a dependence model for the respective innovations $\varepsilon_{x,t}$ and $\varepsilon_{y,t}$---a Gaussian copula driven by GAS dynamics. That is, we assume $(\varepsilon_{x,t},\varepsilon_{y,t})\mid\mathfrak{F}_{t-1}$ to have Gaussian copula density $c(\,\cdot\,; \rho_t)$ with time-varying correlation parameter $\rho_t\in(-1,1)$ following GAS dynamics. Details on the precise specification are in Subsection (ref), where the same model (with $\vartheta=\infty$ in Equation (ref)) is used in the simulations. Both the marginal and the dependence model are regularly used as benchmark models in forecast comparisons. The second method uses the GJR--GARCH(1,1) model of GJR93. The GJR--GARCH model possesses an additional parameter that allows positive and negative shocks of equal magnitude to have a different effect on volatility. As a dependence model, we now use a $t$-copula driven by GAS dynamics, similarly as in Equation (ref). For both models we remain agnostic regarding the specific distribution of the $\varepsilon_{x,t}$ and $\varepsilon_{y,t}$---both in the model estimation (by using the robust Gaussian quasi-maximum likelihood estimator for the GARCH-type marginal models) and the risk forecasting stage (by using their empirical cdfs $\widehat{F}_x$ and $\widehat{F}_y$ in computing the risk measures). For details on how the risk predictions are calculated, we refer to Subsection (ref). We now compare two sequences of rolling-window predictions. For each rolling window of length $r=1\,000$, we re-fit the two models on a daily basis. This gives us $n=4\,193$ forecasts $\{\bm r_{t,(1)}\}_{t=1,\ldots,n}$ from the GARCH model with Gaussian copula, and $\{\bm r_{t,(2)}\}_{t=1,\ldots,n}$ from the GJR--GARCH with $t$-copula. We interpret the $\bm r_{t,(1)}$ as generic benchmark forecasts, which are to be improved upon by the $\bm r_{t,(2)}$. In an oversight context, the $\bm r_{t,(1)}$ may be some regulatory benchmark forecasts and the $\bm r_{t,(2)}$ forecasts from the bank's internal model. Or from a banking perspective, the $\bm r_{t,(1)}$ may be forecasts issued from a trading desk's legacy model and the $\bm r_{t,(2)}$ are forecasts from a refined version thereof. The latter context is more realistic here, because in the regulatory framework the focus is more often on the relation of between an individual bank's returns and the market as a whole. In any case, the methodology remains the same. We compare forecasts based on the verifying observations $\{(X_t,Y_t)\}_{t=1,\ldots,n}$. \begin{figure}[t!] \caption{Top: DAX log-losses on days of VaR violation of S&P 500 (black). CoVaR (CoES) forecasts are shown as the blue (red) line. Violation of CoVaR (CoES) forecast is indicated by a blue (red) dot. All forecasts from GARCH with Gaussian copula. Bottom: Same as top, only with forecasts from GJR--GARCH with $t$-copula.} \end{figure} Figure (ref) shows CoVaR and CoES forecasts from both models, where the top panel corresponds to the GARCH model with Gaussian copula and the bottom panel to the GJR--GARCH with $t$-copula. Specifically, the panels show the CoVaR (blue) and CoES (red) forecasts for the DAX log-losses (black) on days where the S&P 500 exceeds its VaR forecast. Note that due to the different marginal models (and, hence, the different VaR forecasts), the black lines differ slightly in the upper and lower panel of Figure (ref). By definition, the S&P 500 should only exceed its VaR forecast on 5% of all trading days, i.e., on $0.05 \cdot 4193=209.65$ days in our out-of-sample period. With 213 (218) VaR violations, our marginal GARCH(1,1) model (GJR--GARCH(1,1) model) is close to the ideal frequency. By definition, we expect our CoVaR forecasts to be not exceeded on 95% of these 213 days (218 days) with a VaR violation. With 15 and 16 exceedances (blue dots), which correspond to non-exceedance frequencies of 93.0% and 92.7%, the Gaussian copula and the $t$-copula model are reasonably close to the 95%-benchmark. However, for the Gaussian copula, the CoVaR exceedances seem to cluster more, such as during the beginning of the Covid-19 pandemic in early 2020 (top panel of Figure (ref)). Such violation clusters are undesirable from a risk management perspective, providing some informal evidence in favour of the $t$-copula model. We investigate this more formally in the following. Note that both panels of Figure (ref) indicate marked spikes in systemic risk during the financial crisis of 2008--2009, during the European sovereign debt crisis in the first half of the 2010's, and---most markedly---in 2020 as a consequence of the Covid-19 pandemic. As pointed out above, we regard the GARCH model with Gaussian copula as our benchmark model. So we now want to test $H_0^{\preceq_{\mathrm{lex}}}$ for (VaR, CoVaR) and (VaR, CoVaR, CoES). Due to the less clustered CoVaR exceedances and the compelling empirical evidence in favour of GJR--GARCH models GJR93,BEK11 and GAS-$t$-copula models CKL13,BC19, we expect the GJR--GARCH model with $t$-copula to produce lower scores, i.e., better risk forecasts, possibly leading to a rejection of $H_0^{\preceq_{\mathrm{lex}}}$. We carry out the tests using $\mathcal{T}_n^{\operatorname{OS}}$ with the scores of Equations (ref) and (ref) having 0-homogeneous score differences, and with $\widehat{\bm \varOmega}_n$ from (ref) (with $m_n=0$); see Supplement Subsection (ref) for details. For VaR and ES forecasts, scoring functions giving 0-homogeneous score differences are typically recommended, since they allow for `unit-consistent' and powerful comparisons NZ17. We confirm the latter in our simulations for systemic risk forecasts as well, thus justifying our choice. Let $\overline{\bm d}_n=(\overline{d}_{1n}, \overline{d}_{2n})^\prime$ be defined as in Subsection (ref). Indeed, computing $\overline{\bm d}_n$, we find that the GJR--GARCH with $t$-copula produces lower scores for both the (VaR, CoVaR) and the (VaR, CoVaR, CoES) forecasts. The score differences are even statistically significant at the 5%-level: The $p$-values for the $\mathcal{T}_n^{\operatorname{OS}}$-based Wald test are 2.9% for (VaR, CoVaR) and 3.0% for (VaR, CoVaR, CoES). \begin{figure}[t!] \caption{Panel (a): Values of $\overline{d}_{1n}$ and $\overline{d}_{2n}$ for (VaR, CoVaR) forecasts, indicated by $\times$. Panel (b): Values of $\overline{d}_{1n}$ and $\overline{d}_{2n}$ for (VaR, CoVaR, CoES) forecasts, indicated by $\times$.} \end{figure} Figure (ref) illustrates the test decisions. Panel (a) shows the results for the (VaR, CoVaR) comparison. Both $\overline{d}_{1n}$ and $\overline{d}_{2n}$ are positive, favouring the forecasts $\bm r_{t,(2)}$ from the GJR--GARCH model with $t$-copula. Additionally, the pair $\overline{\bm d}_n=(\overline{d}_{1n}, \overline{d}_{2n})^\prime$ lies above the yellow ellipse, which traces the contour of a bivariate normal distribution with probability content $(100-7.66)\%=92.34\%$. This confirms the significance of the score difference at the $5\%$-level; recall Remark (ref). The results for the (VaR, CoVaR, CoES) forecasts in panel (b) are qualitatively similar. Adopting our traffic-light interpretation, the GJR--GARCH model with $t$-copula would be deemed an adequate risk forecasting model. Nonetheless, the borderline significance of this example shows that discriminating between (systemic) risk forecasts requires long samples for the given parameter choice of $\beta=0.95$. This is as expected, because by only considering observations with one extreme component, the effective sample size is massively reduced to roughly $n(1-\beta)$. Indeed, for both models, the effective out-of-sample period for comparing systemic risk forecasts is reduced from a length of 4193 to just over 200. The practical implication is that one should allow for sufficiently large samples to prove the superiority of the internal model over the benchmark. Alternatively, one could adapt a slightly looser definition of distress in the reference asset $X$, e.g., by setting $\beta=0.9$. The current evaluation period for VaR forecasts specified in the Basel framework by the BCBSBF19 is one year, amounting to roughly 250 daily returns. Even for evaluating VaR and ES forecasts, the horizon of 250 days has been called into question for being too short DLS20, and this is only magnified for systemic risk forecasts. So whatever evaluation period for systemic risk measures is eventually settled on in a regulatory context, it likely needs to be far in excess of one year. \section{Discussion and outlook} To our knowledge, this is the first paper to come up with comparative backtests for the systemic risk measures CoVaR, CoES and MES, which are crucial inputs in financial, macroeconomic and regulatory applications. Model selection procedures based on our results may enhance modelling attempts of these quantities in financial institutions. Moreover, the fact that our notion of multi-objective elicitability serves as a `truth serum' implies that the regulatory framework can be improved by enticing financial institutions to accurately model systemic risk. In terms of calibration tests, we are only aware of one more paper: Banulescu-RaduETAL2019 introduce such tests, which unfortunately hinge on non-strict identification functions. This may lead to a severe loss of power under the alternative as demonstrated in Remark (ref) and Section (ref). Due to the strictness of our identification functions, such a phenomenon cannot happen in our context. The novel concept of multi-objective elicitability is likely to be fruitful also in applications beyond the proposed DM-backtests for systemic risk measures. In Subsection (ref), we provide more examples of interesting situations where conditional elicitability holds, yet classical joint elicitability fails. By virtue of Theorem (ref), one can now construct incentive compatible elicitation mechanisms for these functionals, or can come up with $M$-estimation procedures in a regression context. Beyond the confines of finance, we anticipate many other interesting applications of our backtests. For instance, in economics, ABG19 and Aea21 have recently drawn attention to tail risks and their interconnections by popularizing the Growth-at-Risk, which is simply the VaR of GDP growth. This literature has led to an increase in the use of risk forecast evaluation methods in economics BS21. Hence, the backtests developed in this paper should be relevant for future macroeconomic applications as well. \begin{thebibliography}{39} \bibitem[{Acharya \emph{et al.}(2017)Acharya, Pedersen, Philippon and Richardson}]{Aea17} Acharya VV, Pedersen LH, Philippon T, Richardson M. 2017. Measuring systemic risk. \emph{The Review of Financial Studies} \textbf{30}: 2--47. \bibitem[{Adrian \emph{et al.}(2019)Adrian, Boyarchenko and Giannone}]{ABG19} Adrian T, Boyarchenko N, Giannone D. 2019. Vulnerable growth. \emph{The American Economic Review} \textbf{109}: 1263--1289. \bibitem[{Adrian and Brunnermeier(2016)}]{AB16} Adrian T, Brunnermeier MK. 2016. {CoVaR}. \emph{The American Economic Review} \textbf{106}: 1705--1741. \bibitem[{Adrian \emph{et al.}(2021)Adrian, Grinberg, Liang, Malik and Yu}]{Aea21} Adrian T, Grinberg F, Liang N, Malik S, Yu J. 2021. The term structure of {G}rowth-at-{R}isk. \emph{\textup{Forthcoming in} American Economic Journal: Macroeconomics} \url{https://www.aeaweb.org/articles/pdf/doi/10.1257/mac.20180428}. \bibitem[{Artzner \emph{et al.}(1999)Artzner, Delbaen, Eber and Heath}]{Aea99} Artzner P, Delbaen F, Eber JM, Heath D. 1999. Coherent measures of risk. \emph{Mathematical Finance} \textbf{9}: 203--228. \bibitem[{{Bank for International Settlements}(2019)}]{BCBSBF19} {Bank for International Settlements}. 2019. \emph{Basel Framework}. Basel, \url{http://www.bis.org/basel_framework/index.htm?export=pdf}. \bibitem[{Banulescu-Radu \emph{et al.}(2021)Banulescu-Radu, Hurlin, Leymarie and Scaillet}]{Banulescu-RaduETAL2019} Banulescu-Radu D, Hurlin C, Leymarie J, Scaillet O. 2021. Backtesting marginal expected shortfall and related systemic risk measures. \emph{Management Science} \textbf{67}: 5730--5754. \bibitem[{Benoit \emph{et al.}(2017)Benoit, Colliard, Hurlin and P{\'e}rignon}]{Bea17} Benoit S, Colliard JE, Hurlin C, P{\'e}rignon C. 2017. Where the risks lie: A survey on systemic risk. \emph{Review of Finance} \textbf{21}: 109--152. \bibitem[{Bernardi and Catania(2019)}]{BC19} Bernardi M, Catania L. 2019. Switching generalized autoregressive score copula models with application to systemic risk. \emph{Journal of Applied Econometrics} \textbf{34}: 43--65. \bibitem[{Brownlees \emph{et al.}(2011)Brownlees, Engle and Kelly}]{BEK11} Brownlees C, Engle R, Kelly B. 2011. A practical guide to volatility forecasting through calm and storm. \emph{Journal of Risk} \textbf{14}: 3--22. \bibitem[{Brownlees and Engle(2017)}]{BE17} Brownlees C, Engle RF. 2017. {SRISK}: A conditional capital shortfall measure of systemic risk. \emph{The Review of Financial Studies} \textbf{30}: 48--79. \bibitem[{Brownlees and Souza(2021)}]{BS21} Brownlees C, Souza ABM. 2021. Backtesting global {G}rowth-at-{R}isk. \emph{Journal of Monetary Economics} \textbf{118}: 312--330. \bibitem[{Brunnermeier \emph{et al.}(2020)Brunnermeier, Rother and Schnabel}]{BRS20} Brunnermeier M, Rother S, Schnabel I. 2020. Asset price bubbles and systemic risk. \emph{The Review of Financial Studies} \textbf{33}: 4272--4317. \bibitem[{Chen \emph{et al.}(2013)Chen, Iyengar and Moallemi}]{ChenIyengarMoallemi2013} Chen C, Iyengar G, Moallemi CC. 2013. {An Axiomatic Approach to Systemic Risk}. \emph{Management Science} \textbf{59}: 1373--1388. \bibitem[{Creal \emph{et al.}(2013)Creal, Koopman and Lucas}]{CKL13} Creal D, Koopman SJ, Lucas A. 2013. Generalized autoregressive score models with applications. \emph{Journal of Applied Econometrics} \textbf{28}: 777--795. \bibitem[{Diebold and Mariano(1995)}]{DM95} Diebold FX, Mariano RS. 1995. Comparing predictive accuracy. \emph{Journal of Business & Economic Statistics} \textbf{13}: 253--263. \bibitem[{Dimitriadis \emph{et al.}(2020{a})Dimitriadis, Fissler and Ziegel}]{DFZ2020} Dimitriadis T, Fissler T, Ziegel JF. 2020{a}. {The Efficiency Gap}. \emph{Preprint.} \url{https://arxiv.org/abs/2010.14146}. \bibitem[{Dimitriadis \emph{et al.}(2020{b})Dimitriadis, Liu and Schnaitmann}]{DLS20} Dimitriadis T, Liu X, Schnaitmann J. 2020{b}. Encompassing tests for value at risk and expected shortfall multi-step forecasts based on inference on the boundary. \emph{Preprint.} \url{https://arxiv.org/abs/2009.07341}. \bibitem[{Eckernkemper(2018)}]{Eck18} Eckernkemper T. 2018. Modeling systemic risk: Time-varying tail dependence when forecasting marginal expected shortfall. \emph{Journal of Financial Econometrics} \textbf{16}: 63--117. \bibitem[{Ehm \emph{et al.}(2016)Ehm, Gneiting, Jordan and Kr{\"u}ger}]{EhmETAL2016} Ehm W, Gneiting T, Jordan A, Kr{\"u}ger F. 2016. {Of quantiles and expectiles: consistent scoring functions, Choquet representations and forecast rankings}. \emph{Journal of the Royal Statistical Society: Series B (Statistical Methodology)} \textbf{78}: 505--562. \bibitem[{Ehrgott(2005)}]{Ehrgott2005} Ehrgott M. 2005. \emph{Multicriteria Optimization}. Berlin, Heidelberg: Springer. \bibitem[{Emmer \emph{et al.}(2015)Emmer, Kratz and Tasche}]{EKT15} Emmer S, Kratz M, Tasche D. 2015. What is the best risk measure in practice? a comparison of standard measures. \emph{Journal of Risk} \textbf{18}: 31--60. \bibitem[{Feinstein \emph{et al.}(2017)Feinstein, Rudloff and Weber}]{FeinsteinRudloffWeber2017} Feinstein Z, Rudloff B, Weber S. 2017. Measures of systemic risk. \emph{SIAM Journal on Financial Mathematics} \textbf{8}: 672--708. \bibitem[{Fissler \emph{et al.}(2021)Fissler, Frongillo, Hlavinov\'a and Rudloff}]{FFHR2021} Fissler T, Frongillo R, Hlavinov\'a J, Rudloff B. 2021. Forecast evaluation of quantiles, prediction intervals, and other set-valued functionals. \emph{Electronic Journal of Statistics} \textbf{15}: 1034--1084. \bibitem[{Fissler and Ziegel(2016)}]{FZ16a} Fissler T, Ziegel JF. 2016. Higher order elicitability and {O}sband's principle. \emph{The Annals of Statistics} \textbf{44}: 1680--1707. \bibitem[{Fissler \emph{et al.}(2016)Fissler, Ziegel and Gneiting}]{FZG16} Fissler T, Ziegel JF, Gneiting T. 2016. Expected shortfall is jointly elicitable with value-at-risk: Implications for backtesting. \emph{Risk Magazine} : 58--61. \bibitem[{Giacomini and White(2006)}]{GW06} Giacomini R, White H. 2006. Tests of conditional predictive ability. \emph{Econometrica} \textbf{74}: 1545--1578. \bibitem[{Giesecke and Kim(2011)}]{GK11} Giesecke K, Kim B. 2011. Systemic risk: What defaults are telling us. \emph{Management Science} \textbf{57}: 1387--1405. \bibitem[{Giglio \emph{et al.}(2016)Giglio, Kelly and Pruitt}]{GKP16} Giglio S, Kelly B, Pruitt S. 2016. Systemic risk and the macroeconomy: An empirical evaluation. \emph{Journal of Financial Economics} \textbf{119}: 457--471. \bibitem[{Girardi and Tolga Erg{\"u}n(2013)}]{GT13} Girardi G, Tolga Erg{\"u}n A. 2013. Systemic risk measurement: Multivariate {GARCH} estimation of {CoVaR}. \emph{Journal of Banking & Finance} \textbf{37}: 3169--3180. \bibitem[{Glosten \emph{et al.}(1993)Glosten, Jagannathan and Runkle}]{GJR93} Glosten LR, Jagannathan R, Runkle DE. 1993. On the relation between the expected value and the volatility of the nominal excess return on stocks. \emph{The Journal of Finance} \textbf{48}: 1779--1801. \bibitem[{Gneiting(2011)}]{Gne11} Gneiting T. 2011. Making and evaluating point forecasts. \emph{Journal of the American Statistical Association} \textbf{106}: 746--762. \bibitem[{Hoga(2021)}]{Hog20a+} Hoga Y. 2021. Modeling time-varying tail dependence, with application to systemic risk forecasting. \emph{\textup{Forthcoming in} Journal of Financial Econometrics} : 1--31. \bibitem[{Mainik and Schaanning(2014)}]{MS14} Mainik G, Schaanning E. 2014. On dependence consistency of {CoVaR} and some other systemic risk measures. \emph{Statistics & Risk Modeling} \textbf{31}: 49--77. \bibitem[{Mas-Colell \emph{et al.}(1995)Mas-Colell, Whinston and Green}]{MWG1995} Mas-Colell A, Whinston MD, Green JR. 1995. \emph{Microeconomic Theory}. Oxford University Press. \bibitem[{Nolde and Zhang(2020)}]{NZ19+} Nolde N, Zhang J. 2020. Conditional extremes in asymmetric financial markets. \emph{Journal of Business & Economic Statistics} \textbf{38}: 201--213. \bibitem[{Nolde and Ziegel(2017)}]{NZ17} Nolde N, Ziegel JF. 2017. Elicitability and backtesting: Perspectives for banking regulation. \emph{The Annals of Applied Statistics} \textbf{11}: 1833--1874. \bibitem[{Steinwart \emph{et al.}(2014)Steinwart, Pasin, Williamson and Zhang}]{SteinwartPasinETAL2014} Steinwart I, Pasin C, Williamson R, Zhang S. 2014. {Elicitation and Identification of Properties}. \emph{JMLR Workshop Conf. Proc.} \textbf{35}: 1--45. \bibitem[{White(2001)}]{Whi01} White H. 2001. \emph{Asymptotic Theory for Econometricians}. San Diego: Academic Press, {F}irst edn. \end{thebibliography}
bibunit