Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,616 characters · 12 sections · 56 citation commands
Systemic Risk Surveillance
\baselineskip18pt \setcounter{totalnumber}{50} \setcounter{topnumber}{50} \setcounter{bottomnumber}{50} \abovedisplayskip1.5ex plus1ex minus1ex \belowdisplayskip1.5ex plus1ex minus1ex \abovedisplayshortskip1.5ex plus1ex minus1ex \belowdisplayshortskip1.5ex plus1ex minus1ex
\doublespacing
The numerous financial crises of recent times and their severe economic reverberations have raised awareness of the importance of systemic risk AB16,Aea17,VZ19. To better appreciate the difference between risk and systemic risk, consider the returns on shares of banks. While bank returns are often subject to higher volatility during crises (implying larger risk), this does not necessarily translate into increased commonality (implying larger systemic risk). However, it is precisely the increased commonality that regulators have come to be most concerned about, as this may entail system-wide distress with potentially severe economic costs GKP16.
Therefore, it is important to evaluate systemic risk forecasts in the financial system. While classical financial “one-shot” backtests are designed for such forecast evaluations, under repeated application their statistical type I errors (false test rejections) accumulate over time---up to the point that a true null will be rejected with probability approaching one. As systemic risks are to be monitored continuously (e.g., daily), it is necessary to control accumulated false rejections, aligning with recent interest in safe anytime-valid statistical inference shafer2021testing, vovk2021values, ramdas2023game, WWZ23. In this paper, we interchangeably call methods with “time-uniform” false rejection guarantees monitoring procedures, surveillance schemes or online tests.
Despite the apparent need for systemic risk surveillance, there is a lack of statistically valid tools for this in the literature. It is the main aim of this paper to fill this gap by providing such tools and to show their validity. To the best of our knowledge, there only exist “one-shot” backtests that assess the adequacy of systemic risk forecasts Bea21,FH21. However, their rejection rates would accumulate under repeated application. We propose monitoring procedures for the most popular systemic risk measure, the conditional Value-at-Risk (CoVaR) of AB16, and the reverse CoVaR (RCoVaR). In Appendix (ref), we also develop monitoring schemes for the conditional expected shortfall (CoES) and marginal expected shortfall (MES) of Aea17.
For the construction of the monitoring procedure for the CoVaR as the most important systemic risk measure, we draw on recent work of FH21. They provide so-called {identification functions} that uniquely identify both the stand-alone risk and systemic risk. Under the null hypothesis that the CoVaR forecasts are correctly specified (i.e., calibrated), the probabilistic structure of the corresponding identification functions is fully known: they are independent and identically distributed (IID) Bernoulli random variables. This remains true without making any additional assumptions about the dynamic structure of the return series. As the probabilistic structure under the null is fully known, it can be used to simulate time-uniform and non-asymptotic critical values. The simulation-based construction is particularly useful as systemic risk events are rare by definition, such that standard asymptotic theory may not be relied upon to provide accurate approximations in finite samples Hog19a+. We also extend our monitoring procedure to the CoVaR with reversed conditioning.
Because the probabilistic structure of the identification functions for CoES and MES is not fully characterized under the null, a direct extension of our CoVaR monitoring scheme is infeasible. Instead, in Appendix (ref), we exploit a systemic-risk analogue of the classical cumulative violation sequence used in Acerbi2002spectral, DE17, Du2024powerful, among others. We show that, under the null, the full probabilistic structure of this sequence---namely its uniform-type distribution and independence---is known, which again enables the computation of non-asymptotic, time-uniform critical values via simulation.
Our approach of constructing the CoVaR monitoring procedure is closely related to HD22a+, who propose online detection schemes for single univariate risk forecasts, viz. Value-at-Risk (VaR) and expected shortfall (ES). However, we improve upon their procedure by using normalized detectors, and most importantly, extend their theory to systemic risk measures. The latter necessitates monitoring multiple time series at once, whereas HD22a+ only consider a single series. This multivariate monitoring necessitates a Bonferroni-type correction to obtain critical values with finite-sample validity, which are not too conservative in practice. A further benefit of the Bonferroni correction is that it allows to attribute a monitoring alarm to a specific financial institution.
Additional work that is related to ours are the monitoring procedures proposed by WG13, NLL14 and MSW20. These authors do, however, not focus on systemic risk itself, but only on quantities that can at best be described as crude approximations to it, such as correlation WG13 and the copula NLL14,MSW20. Moreover, these procedures rely on an initial period of non-contamination, which our approach does not require.
Regarding calibration (backtesting) under repeated use, our work relates to the literature on safe and anytime-valid inference (SAVI) based on e-values. For VaR and ES, several e-value--based monitoring schemes have been proposed, but they do not extend to systemic risk measures. While e-value tests are typically valid indefinitely, they are often overly conservative. In contrast, our approach fixes the monitoring horizon and delivers non-asymptotic, simulation-based, and hence (almost) exact size control via (almost) exact critical values. horvath2025sequential take yet another approach by proposing a sequential monitoring procedure for VaR and ES models that is, however, highly model-specific and based on asymptotic theory.
Our Monte Carlo simulations use a realistic DCC--GARCH model Eng02 for the asset returns to demonstrate the good finite-sample properties of our monitoring procedures. Despite the use of a potentially conservative Bonferroni-type correction, the empirical size remains close to the nominal level---even in moderately high-dimensional settings, where numerous individual tests could, in principle, amplify such conservative tendencies. We also show that, under the null, the detector to first raise a false alarm pertains to the VaR or the CoVaR with equal probability. The simulations further demonstrate the high power of our procedures in identifying misspecified forecasts, again exhibiting a balanced pattern regarding which detector first signals a suboptimal predictive performance.
In the empirical application, we apply our monitoring procedure to the stock returns of systemically relevant US banks. In doing so, we evaluate systemic risk forecasts from two competing multivariate GARCH models over a calm period and three turbulent episodes: the global financial crisis, the COVID-19 pandemic, and the recent US tariff period under President Trump. The monitoring results indicate that the constant covariance structure of the simpler CCC–GARCH model is inadequate for capturing the systemic component of financial risk during such volatile periods, whereas the forecasts of the more flexible DCC–GARCH cannot be rejected.
The remainder of the paper is structured as follows. Section (ref) formally introduces the CoVaR and its reverse variant and our corresponding monitoring procedures. A related monitoring scheme for the CoES and MES is constructed in Appendix (ref). Simulations in Section (ref) explore the finite-sample performance of our methods, and Section (ref) presents the empirical application. Finally, Section (ref) concludes. Next to our CoES and MES surveillance schemes in Section (ref), the appendix contains Section (ref) with the proofs for our CoVaR procedures, and Section (ref) that contains the algorithms to compute critical values.
First, we define the systemic risk measures we consider in Section (ref). Then, we introduce monitoring procedures for the CoVaR in Section (ref), which we extend to the reverse CoVaR in Section (ref). Surveillance schemes for the CoES and MES are presented in Appendix (ref).
\sloppy Consider the to-be-monitored sequence of random variables $\big\{(X_t,Y_{1t},\ldots,Y_{Kt})^\prime\big\}_{t \in \mathbb{N}}$, where $K \in \mathbb{N}$ is the number of monitored variables. Here, $Y_{kt}$ ($k=1,\ldots,K$) stands for the log-losses of interest (e.g., the losses of a bank's shares / losses of a business unit) and $X_t$ are the log-losses of some reference position (e.g., system-wide losses in the financial system / bank-wide losses of all business units). In a concrete monitoring situation, the variables at time $t=1,2\ldots$ are still to be (sequentially) observed and are not yet available when setting up the surveillance scheme at time $t=0$.
Let $\mathcal{F}_{t}=\sigma\big((X_t,Y_{1t},\ldots,Y_{Kt})^\prime,\bm Z_t,(X_{t-1},Y_{1,t-1},\ldots,Y_{K,t-1})^\prime,\bm Z_{t-1},\ldots\big)$ denote the time-$t$ information set with (possibly multivariate) $\bm Z_t$ containing additional covariates. Further, we denote by $F_{(X_t,Y_{kt})^\prime\mid\mathcal{F}_{t-1}}(\cdot,\cdot)$ the cumulative distribution function (CDF) of $(X_t,Y_{kt})^\prime\mid\mathcal{F}_{t-1}$ for $k\in[K]$, where we use the shorthand notation $[K] := \{1,\dots,K\}$ for any $K \in \mathbb{N}$. Throughout this paper, we assume that $(X_t,Y_{kt})^\prime\mid\mathcal{F}_{t-1}$ has a strictly positive Lebesgue density for all $(x,y)\in\mathbb{R}^2$ such that $F_{(X_t,Y_{kt})^\prime\mid\mathcal{F}_{t-1}}(x,y)\in(0,1)$. This assumption ensures that we do not have to rely on generalized inverses to compute our (quantile-based) risk measures, which we introduce next.
Define the VaR as the quantile at level $\beta \in (0,1)$ of the conditional distribution $F_{X_t\mid\mathcal{F}_{t-1}}$, i.e., $\operatorname{VaR}_{t,\beta}=F_{X_t\mid\mathcal{F}_{t-1}}^{-1}(\beta)$. The VaR is a popular univariate risk measure. Note that the inverse $F_{X_t\mid\mathcal{F}_{t-1}}^{-1}(\cdot)$ exists thanks to our assumption on the distribution of $(X_t, Y_{kt})^\prime\mid\mathcal{F}_{t-1}$.
We now introduce the systemic risk measures CoVaR and its reverse variant that we denote by RCoVaR. The stress event in the definition of all these measures is that the loss of the reference position exceeds its VaR, i.e., $\{X_t\geq\operatorname{VaR}_{t,\beta}\}$. Then, we define the CoVaR of the $k$-th series $Y_{kt}$ as
where $F_{Y_{kt}\mid X_t \geq \operatorname{VaR}_{t,\beta}, \mathcal{F}_{t-1}}(\cdot)=\mathbb{P}\{Y_{kt}\leq \cdot\mid X_t \geq \operatorname{VaR}_{t,\beta},\ \mathcal{F}_{t-1}\}$, and $\alpha \in (0,1)$. As for the VaR, our assumption on the CDF of $(X_t, Y_{kt})^\prime\mid\mathcal{F}_{t-1}$ ensures that the inverse in (ref) exists. Note that the CoVaR is simply the $\alpha$-VaR of the distribution of $Y_{kt}$ conditional on $\{X_t\geq\operatorname{VaR}_{t,\beta}\}$, i.e., conditional on $X_t$ being in distress. Without the conditioning on the distress event $\{X_t\geq\operatorname{VaR}_{t,\beta}\}$ (or, equivalently, with $\beta=0$), the CoVaR reduces to the VaR of $Y_{kt}$, such that $\operatorname{CoVaR}_{t,\alpha|0} = F_{Y_{kt}\mid\mathcal{F}_{t-1}}^{-1}(\alpha)$. Since $X_t$ denotes financial losses, we typically consider values for $\alpha$ and $\beta$ close to one for the CoVaR to capture (downside) systemic risk (e.g., $\alpha=\beta=0.95$). When $\alpha=\beta$ we simply write $\operatorname{CoVaR}_{kt,\alpha} = \operatorname{CoVaR}_{kt,\alpha|\alpha}$.
Following among others GT13, NZ20 and Bea21, our CoVaR definition deviates from the original one of AB16, who use $\{X_t=\operatorname{VaR}_{t,\beta}\}$ as the stress event. This latter choice is problematic because it often has probability zero and it does not fully incorporate all tail events of $X_t$. We refer to DH24 for a detailed account of the advantages of the definition in (ref).
The above interpretation of the CoVaR in (ref) with $Y_{kt}$ as losses of financial institutions and $X_t$ as market losses coincides with what AB16 call the Exposure CoVaR, and it measures how institution $k$ is exposed to market risk. A regulator might however also be interested in the reverse direction, i.e., how an individual bank affects the market. Therefore, we additionally define the CoVaR with reverse conditioning as
where $\operatorname{VaR}_{kt,\beta} = F_{Y_{kt}\mid\mathcal{F}_{t-1}}^{-1}(\beta)$ is now specific to institution $k$. The reverse CoVaR in (ref) measures the impact of distress in bank $k$ on the general market, which corresponds to the notion of banks as transmitters of systemic risk. In contrast, our initial CoVaR definition views banks as receivers of systemic risk. Of course, both definitions reveal useful information, and which definition is more suitable depends on the context.
We now propose procedures for monitoring the correct specification of CoVaR forecasts. Importantly, the monitoring should be carried out in an online fashion, i.e., sequentially as new observations become available. Classical backtests with asymptotic validity based on a test statistic $Z_T$ with sample size $T$ and critical value $c_\iota$ for significance level $\iota \in (0,1)$ require that under the null hypothesis of correct specification, $H_0$, the asymptotic rejection probability $\lim_{T \to \infty} \mathbb{P}_{H_0}\{Z_T > c_\iota\} \le \iota$ is below $\iota$. In contrast, sequential monitoring procedures satisfy the stronger non-asymptotic notion
This implies that even though the test decision is monitored every trading day---not unusual in a risk management context---the accumulated probability of a false rejection (type I error) remains controlled. We will achieve such finite sample type I error guarantees in (ref) by constructing detectors (essentially: test statistics), whose probabilistic structure is (almost) fully known under the null hypothesis, such that exact critical values can be calculated.
In a practical monitoring situation, we are about to sequentially observe $(X_1,Y_{11},\ldots,Y_{K1})^\prime, \ldots, (X_n,Y_{1n},\ldots,Y_{Kn})^\prime$ at times (trading days) $1,\dots,n$. Our monitoring procedure defined below requires a rolling monitoring window of fixed length $m \le n$, as illustrated by the blue bars in Figure (ref). Thus, at each monitoring time $T \in \{m,\dots,n\}$---the endpoint of a given monitoring window---the procedure relies on the window spanning the times $\{T-m +1 ,\dots,T\}$. We only start monitoring when the first full window of $m$ trading days is available at time $T=m$. While starting earlier with an initially expanding window would be possible, this would come at the cost of an explosion of technicalities, which we avoid for better readability.
We now formalize the null hypothesis of risk measure forecast adequacy, which is closely related to the statistical notion of conditional forecast calibration; see GR_2023 for a recent and detailed treatment of forecast calibration. The strongest notion of ideal forecasts $\widehat{\operatorname{VaR}}_{t,\beta}$ and $\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}$ for the VaR and CoVaR is given by
We synonymously refer to this null as testing ideal forecasts, correct specification or calibration.
Clearly, $H_0^{\operatorname{CoVaR}}$ is not directly testable, because the true $\operatorname{VaR}_{t,\beta}$ and $\operatorname{CoVaR}_{kt,\alpha|\beta}$ are not observable even ex post. To circumvent this, we exploit the existence of a strict identification function for the CoVaR (jointly with the VaR), given in FH21,
Identification functions are at the heart of many classical backtests NZ17. The property that renders them useful in testing calibration is that their (conditional) expectation is zero if and only if the true (conditional) VaR and CoVaR are inserted. More precisely, FH21 show that
under our conditions on the CDF of $(X_t, Y_{kt})^\prime\mid\mathcal{F}_{t-1}$. Hence, the relevant implication in monitoring calibration becomes
The following result shows that the full, joint probabilistic properties of the indicators
is known under the null hypothesis $H_0^{\operatorname{CoVaR}}$.
As the full probabilistic structure of $\big\{(I_t,I_{kt})^\prime\big\}_{t\in\mathbb{N}}$ is known for each $k \in [K]$ individually, the result of Proposition (ref) can be used for our monitoring procedure to simulate critical values that are valid in a monitoring sense (ref) without the need for asymptotic approximations. Importantly, the binary nature of the CoVaR identification function in (ref) facilitates such a treatment, while remaining agnostic about the (conditional) distributions of $X_t$ and $Y_{kt}$. Since for $K > 1$, Proposition (ref) stays silent on the dependence structure between the contemporaneous $I_{kt}$ and $I_{k't}$ with $k \neq k'$, we will exploit a Bonferroni-type correction in Theorem (ref) below.
For a (prospective) monitoring sample of size $n \in \mathbb{N}$, Proposition (ref) leads to the testable implications of $H_0^{\operatorname{CoVaR}}$ that
where $\operatorname{Bern}(\gamma)$ denotes a Bernoulli distribution with success probability $\gamma\in(0,1)$.
We now explain the construction of a VaR detector that monitors $I_{t} \overset{\text{IID}}{\sim}\operatorname{Bern}(1-\beta)$ from (ref), and continue with a CoVaR detector below. For monitoring the VaR-specific sequence $I_t$, we adapt the method of HD22a+ and use a moving sum (MOSUM) detector inspired by the one-shot backtest of KW15 of the form
where $T \in \{m,\dots,n\}$ and $a\in(0,1)$ with canonical choice $a = 0.5$. The constant $a$ is a user-specified parameter that allows to vary the sensitivity of the procedure towards violations of correct unconditional calibration (i.e., $\mathbb{E}[I_t]=1-\beta$) and violations of IIDness of the $I_t$. These two properties are assessed by $\operatorname{VaR}^{uc}(T)$ and $\operatorname{VaR}^{iid}(T)$, respectively. The length of the rolling monitoring window is denoted by $m\in\mathbb{N}$, but the dependence of $\operatorname{VaR}(T)$ on $m$ is not made explicit for notational brevity.
For $\operatorname{VaR}^{uc}(T)$ in (ref), we choose the standardized detector
The mean, $\mathbb{E}_{H_0^{\operatorname{CoVaR}}}[V_T]$, and variance, $\operatorname{Var}_{H_0^{\operatorname{CoVaR}}}(V_T)$, are obtained from simulations as the full probabilistic structure of $I_t$ is known under $H_0^{\operatorname{CoVaR}}$; see (ref). The standardization in $\operatorname{VaR}^{uc}(T)$ differs from HD22a+ and is used to better balance the sensitivity of the two detectors in (ref). In contrast to $\mathbb{E}_{H_0^{\operatorname{CoVaR}}}[V_T]$ and $\operatorname{Var}_{H_0^{\operatorname{CoVaR}}}(V_T)$, the sequence $V_T$ has to be computed from the data (i.e., observations and appertaining forecasts) and is responsible for giving the procedure its power. Specifically, $V_T$ is designed to uncover deviations from $\mathbb{E}[I_t]=1-\beta$, i.e., correct unconditional calibration. By construction, large values of $\operatorname{VaR}^{uc}(T)$ provide evidence against the null, where deviations in both directions (i.e., $\mathbb{E}[I_t]>1-\beta$ and $\mathbb{E}[I_t]<1-\beta$) raise an alarm.\footnote{ One could also consider one-sided detectors here. However, we refrain from doing so, because it remains unclear how a too conservative (and, hence, classically acceptable) misspecification of VaR forecasts interact with the validity of the CoVaR forecasts in light of the joint identification function in (ref); also see WWZ23. }
For $\operatorname{VaR}^{iid}(T)$ in (ref), we consider the durations $d_i:=t_i-t_{i-1}$ ($t_0:=T-m$) between VaR violation times in the window $\{T-m+1,\ldots,T\}$. These are defined as
where $S_{T} := \sum_{t=T-m+1}^{T}I_t$ counts the number of VaR violations in the window $\{T-m+1,\ldots,T\}$. As a measure of the “inequality” between the durations $d_i$, we follow KW15 and use the Gini coefficient
The detector for the second part of (ref) then is the standardized Gini coefficient
where $\mathbb{E}_{H_0^{\operatorname{CoVaR}}}[g_T]$ and $\operatorname{Var}_{H_0^{\operatorname{CoVaR}}}(g_T)$ are the mean and variance, respectively, of $g_T$ under the null. As above, these moments are derived from Monte Carlo simulations, and it is only $g_T$ itself that is computed from the data.
The idea underlying the use of $\operatorname{VaR}^{iid}(T)$ is to test the IID property of $I_t$ by considering the spacing between the VaR exceedances (i.e., those $t$ for which $I_t=1$). Unevenly spaced exceedances with large values for the Gini coefficient (and, hence, large values of $\operatorname{VaR}^{iid}(T)$) weigh against IIDness. Note that in principle also too evenly spaced exceedances (with small values for the Gini coefficient) provide evidence against $H_0^{\operatorname{CoVaR}}$. However, such a violation of the null is of no concern in risk management, where only a clustering of exceedances is seen as problematic KW15.
To also monitor the CoVaR forecasts, i.e., to sequentially test $I_{kt}\overset{\text{IID}}{\sim}\operatorname{Bern}\big((1-\alpha)(1-\beta)\big)$ from (ref), we use the same strategy as for monitoring VaR predictions. Therefore, we define the CoVaR detectors $\operatorname{CoVaR}_{k}(T)$ for any $k \in [K]$ similarly as the VaR detector $\operatorname{VaR}(T)$:
Here, the right-hand side quantities are defined in analogy to those in (ref), with the only difference that $I_{kt}$ replaces $I_t$, and $(1-\alpha)(1-\beta)$ replaces $1-\beta$ at every occurrence.
The following theorem shows how the probability of a false detection in the sense of (ref) can be bounded uniformly over time and over different institutions $k \in [K]$, for which the CoVaR is forecasted.
In other words, Theorem (ref) shows that size at level $\iota\in(0,1)$ is controlled in finite samples if we reject $H_0^{\operatorname{CoVaR}}$ as soon as either
The two probabilities in the upper row of (ref) are associated with a standard Bonferroni correction for testing the $K+1$ hypotheses of calibration of $\widehat{\operatorname{VaR}}_{t,\beta}$ and $\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}$ for $k \in [K]$. However, the presence of the negative third term on the right-hand side of (ref) allows for a less conservative---and, hence, potentially more powerful---procedure. We mention that there may not exist $v$ and $c_k$, such that the equality in (ref) holds exactly, because of the binary nature of the detectors. Therefore, in practice $v$ and $c_k$ should be chosen to lead to a sum that is smaller but as close as possible to $\iota$; see also step (ref) in Algorithm (ref) below.
For the practical computation of the critical values $v$ and $c_k$, we use the following Algorithm (ref), which approximates the probabilities in (ref) by using the probabilistic structure in (ref) to sample under the null.
The above algorithm also overcomes the theoretical ambiguity that there exists a continuum of critical values $v$ and $c_k$ that ensure the probabilities in (ref) sum to $\iota$. In principle, such a continuum allows to weight the importance of the VaR and CoVaR hypotheses. However, the fourth step in Algorithm (ref) advocates a natural approach that balances the importance of VaR and CoVaR by choosing $v$ and $c_k$ to correspond to some $(1-\nu)$-quantile of the variables $\sup_{T=m,\ldots,n}\operatorname{VaR}(T)$ and $\sup_{T=m,\ldots,n}\operatorname{CoVaR}_{k}(T)$, respectively. Then, $\nu$ in step 5 simply has to be chosen to ensure that the probabilities in (ref) sum to $\iota$. Note that this entails the computation of two critical values only---one for the VaR (written $v^{\operatorname{CoVaR}}$) and one for the CoVaR (written $c^{\operatorname{CoVaR}}$), where the latter one is the same for all $k \in [K]$, because the distribution of $I_{kt}\overset{\text{IID}}{\sim}\operatorname{Bern}\big((1-\alpha)(1-\beta)\big)$ is independent of $k$. These critical values have the superscript “${\operatorname{CoVaR}}$” to distinguish them from critical values used for monitoring different systemic risk measures.
Clearly, the use of Boole's inequality for $K > 1$ (in the proof of Theorem (ref)) implies that our monitoring procedure is conservative, indicated by the inequality in (ref). Only for $K=1$, this becomes an equality such that size is kept exactly. While a conservative test reduces the probability of actually making a type I error, this usually has the drawback of increasing the probability of a type II error, such that the ability to identify (systemic) risk changes is reduced. However, we show in simulations in Section (ref) that our procedure is not too conservative for values of $K$ representative of practical applications.
A possible reason for the good performance of the Bonferroni correction is explained by Hol79, who states that: “The power gain obtained by using a sequentially rejective Bonferroni test [based on ordering $p$-values] instead of a classical Bonferroni test depends very much upon the alternative. It is small if all the hypotheses are `almost true', but it may be considerable if a number of hypotheses are `completely wrong'.” In our case, systemic risk builds up slowly for all banks, rendering all hypotheses 'almost true' at first, thus leading to good detection properties of our simple Bonferroni-type corrections.
The use of the Bonferroni-type correction has two further advantages. First, it ensures controlled size in finite samples. In particular, we do not have to rely on large-sample asymptotics, which may be unreliable in backtesting contexts Hog19a+. Second, the Bonferroni correction allows us to attribute a rejection of the null to a specific institution. For instance, if the detector $\operatorname{CoVaR}_{k^{\ast}}$ for some $k^\ast \in [K]$ is the first to raise an alarm, the rejection of the null can be pinpointed to bank $k^\ast$. Of course, precisely identifying the most vulnerable institution is of utmost importance in systemic risk analysis.\footnote{ A further implication of $H_0^{\operatorname{CoVaR}}$ is that $I_s$ and $I_{kt}$ are independent for $s\neq t$. We do not explicitly reflect this implication in the construction of our detectors. While doing so may increase the power of the detector (depending, of course, on the specific type of alternative), this would again render infeasible the attribution of a rejection of $H_0^{\operatorname{CoVaR}}$ to a specific institution. }
Section (ref) focuses on a monitoring procedure for the CoVaR in (ref), which measures the individual exposure of (e.g., financial institution) $Y_{kt}$ onto the joint (market) $X_t$. Now, we consider monitoring of the RCoVaR in (ref), which measures the effect the individual (e.g., financial institution) $Y_{kt}$ has on the joint (market) $X_t$.
Similarly as above, for given VaR forecasts $\widehat{\operatorname{VaR}}_{kt,\beta}$ and RCoVaR forecasts $\widehat{\operatorname{RCoVaR}}_{kt,\alpha|\beta}$, the null is
Employing analogous arguments as in Section (ref), we rely on a testable implication of $H_0^{\operatorname{RCoVaR}}$ that is based on the indicators $I_{kt}^{v} := \mathds{1}_{\{Y_{kt}> \widehat{\operatorname{VaR}}_{kt,\beta}\}}$ and $I_{kt}^{c} := \mathds{1}_{\{Y_{kt}> \widehat{\operatorname{VaR}}_{kt,\beta},\ X_t> \widehat{\operatorname{RCoVaR}}_{kt,\alpha|\beta}\}}$. Specifically, we have the following analog to Proposition (ref).
The proof is similar to that of Proposition (ref) and, hence, omitted. Proposition (ref) suggests the following surveillance approach. VaR changes may be monitored via the detector $\operatorname{VaR}_{k}(T)$, which is defined in analogy to $\operatorname{VaR}(T)$ at (ref) with $I_t$ replaced by $I_{kt}^{v}$. Similarly, CoVaR surveillance now uses the detector $\operatorname{RCoVaR}_{k}(T)$, where the definition is identical to that of $\operatorname{CoVaR}_{k}(T)$ at (ref) except that $I_{kt}^{c}$ is used in place of $I_{kt}$. \color{black}
As in Section (ref), almost the full probabilistic structure of the detectors is known under $H_0^{\operatorname{RCoVaR}}$, such that the probabilities in (ref) and, thereby, the critical values $v_k$ and $c_k$ can be computed explicitly via simulations. For this, we employ Algorithm (ref) in Appendix (ref), which is very similar to Algorithm (ref), but takes into account the differences between (ref) and (ref). Again, Algorithm (ref) only computes two critical values---one for the VaR component ($v^{\operatorname{RCoVaR}}$), and one for the RCoVaR component ($c^{\operatorname{RCoVaR}}$).
As in Theorem (ref), our procedure is less conservative than a standard Bonferroni correction. Recall that a Bonferroni correction suggests monitoring each of the $2K$ hypotheses at a significance level of $\iota/(2K)$, such that the total probability of rejection under the null is bounded from above by $\iota$. In contrast, our monitoring scheme based on Theorem (ref) is less conservative because the sum in (ref) is less than the “Bonferroni sum”, i.e.,
Here, we investigate size and power of our monitoring procedures. We place a particular emphasis on studying size, because Bonferroni-type corrections---such as those underlying our monitoring schemes---are known to be conservative when many hypotheses are tested.
We simulate the variables $\big\{(X_t,Y_{1t},\ldots,Y_{Kt})^\prime\big\}_{t=-E+1,\ldots,n}$ from a DCC--GARCH model of Eng02. The sequence consists of $E \in \mathbb{N}$ time points for model estimation (used in one simulation setting), and $n \in \mathbb{N}$ time points for generating forecasts and their evaluation. The data-generating process (DGP) is
where $\bm D_t^2$ is the $\mathcal{F}_{t-1}$-measurable diagonal matrix containing the componentwise conditional variances of $\bm W_t\mid\mathcal{F}_{t-1}$, the matrix $\bm R_t$ is the conditional correlation of $\bm W_t\mid\mathcal{F}_{t-1}$, and $t_{\nu}(\bm R_t)$ denotes the multivariate $t$-distribution with degrees of freedom equal to $\nu>2$, correlation matrix $\bm R_t$ and marginals standardized to have zero mean and unit variance. Therefore, the conditional variance-covariance matrix of $\bm W_t\mid\mathcal{F}_{t-1}$ is $\bm H_t=\bm D_t\bm R_t\bm D_t$. The matrices $\bm D_t$ and $\bm R_t$ are modeled in the DCC--GARCH fashion of Eng02 as
where $\overline{\bm Q}$ is the unconditional correlation matrix of the “devolatilized” series $\bm \varepsilon_{t-1}$, and $\operatorname{diag}(\bm a)$ and $\operatorname{diag}(\bm A)$ denote the diagonal matrices containing the elements of the vector $\bm a$ and the diagonal elements of the matrix $\bm A$, respectively. Note that both the conditional variance matrix $\bm D_t^2$ and the conditional correlation matrix $\bm R_t$ evolve in a GARCH-type fashion (the latter through the recursion for $\bm Q_t$).
We parametrize the process as follows: We choose $\overline{\bm Q}$ to be the matrix with ones on the main diagonal and 0.5 elsewhere (i.e., an equicorrelation matrix), such that the unconditional correlation between any two elements of $\bm \varepsilon_t$ equals 0.5. Furthermore, we fix the parameters $\nu=5$, $\bm \omega_G=(0.1,\ldots, 0.1)^\prime$, $\bm \alpha_G=(0.1,\ldots,0.1)^\prime$, and $\alpha_{Q}=0.1$. In contrast, the autoregressive parameters $\bm \beta_G=(\beta_{G,1},\ldots,\beta_{G,K+1})^\prime$ and $\beta_{Q}$ vary over time. Specifically,
such that there is an upward change in the persistences at time $t^{\ast} \in [n]$. In the baseline simulation setup, we fix $\beta_\text{post} = 0.85$.
Through (ref), we misspecify the dynamics of the process, which will then be picked up by the monitoring procedure. Of course, if $t^{\ast}=n$, then there is no structural change in the sample, and the systemic risk forecasts (issued from the estimates of the initial fixed window) are correctly specified. Hence, except for the estimation error, the forecasts are correctly specified, such that a rejection probability close to the nominal level should be expected.
For generating the forecasts, we either use fixed DCC--GARCH parameters or (in one setup) estimate the parameters based on a sample $\big\{(X_t,Y_{1t},\ldots,Y_{Kt})^\prime\big\}_{t=-E+1,\ldots,0}$ of length $E$ by using the R package rmgarch rmgarch. In doing so, we estimate the marginals via standard Gaussian quasi-maximum likelihood estimation (to obtain $K+1$ estimates of $\omega_{G,i}$, $\alpha_{G,i}$ and $\beta_{G,i}$). The dependence parameters ($\nu$, $\alpha_Q$ and $\beta_Q$) are estimated in a second step using a multivariate $t$-assumption for the $\bm \varepsilon_t$.
Given DCC-forecasts (either based on fixed or estimated model parameters) for $\widehat{\bm D}_t$, $\widehat{\bm R}_t$ and $\widehat{\bm H}_t$, we obtain the VaR forecasts at time $t$ as $\widehat{\operatorname{VaR}}_{t,\beta} = \widehat{\bm D}_{t,11}^{-1} q_{t_\nu}(\beta)$, where $q_{t_\nu}(\beta)$ denotes the $\beta$-quantile of the univariate $t_{{\nu}}$-distribution with unit variance and (either fixed or estimated) degrees of freedom $\nu$. For the CoVaR forecasts, we are not aware of a closed-form solution. Instead, we apply a root finding algorithm to approximate $\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}$ via
The probability in (ref) is taken with respect to $(X_t, Y_{kt})\mid\mathcal{F}_{t-1} \sim t_\nu(\widehat{\bm H}_{t,k})$, where $\widehat{\bm H}_{t,k} = \big( ( \widehat{\bm H}_{t, 11}, \widehat{\bm H}_{t, 1 (k+1)})' \mid ( \widehat{\bm H}_{t, 1 (k+1)}, \widehat{\bm H}_{t, (k+1) (k+1)})' \big)$ consists of the respective entries of the forecasted variance-covariance matrix $\widehat{\bm H}_{t}$. Based on the forecasts $\widehat{\operatorname{VaR}}_{t,\beta}$ and $\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}$, we compute the VaR and CoVaR indicators as described in Sections (ref)--(ref).
All following results are based on 5000 simulation replications. We first fix $n=1000$, $t^\ast=n$, $m=250$, and generate the systemic risk forecasts based on the correct DCC--GARCH parameters from the “pre-break” period such that we omit parameter estimation noise. We deliberately do not compare against one-shot systemic risk backtests of, e.g., Bea21 and FH21 since these---opposed to our monitoring procedure---accumulate type I errors when applied repeatedly.
The columns “Joint” in Tables (ref) and (ref) show the rejection rates of the CoVaR and RCoVaR monitoring procedures under the no-break null hypotheses---i.e., $t^\ast=n$ in (ref). We consider $\alpha = \beta \in \{0.9, 0.95\}$ together with $K \in \{1,2,5,10\}$ for the CoVaR and $K \in \{2,5,10\}$. For $K=1$, the rejection rates of the CoVaR and RCoVaR coincide due to the symmetry of the DGP. We find that all combined empirical rejection rates in the columns “Joint” are below $10.28\%$, which corresponds to our theoretical finding that the monitoring procedures hold size exactly.\footnote{Given that these empirical rejection rates are based on 5000 Monte Carlo replications, a $99\%$-confidence interval of IID Bernoulli-distributed variables with success probability 0.1 is given by approximately $[8.9\%, 11.1\%]$, indicating that our procedure is able to hold size exactly, while the deviations from $10\%$ are explained by the finite number of Monte Carlo replications.} Even for $K=10$, the rejection rates are not too conservative and lie between $6\%$ and $9\%$, which illustrates that the application of Boole's inequality in the proof of Theorem (ref) is relatively tight here.
As the columns “Joint” merely indicate whether any of the detectors raises an incorrect alarm under the null, we also analyze which of the detectors is responsible for the rejection. For this, the right-hand side columns (denoted “VaR” and with the series numbers “1”--“10”) of the tables show the frequencies how often a given detector is responsible for first raising a false alarm. Notice here that the joint rejection frequency is somewhat lower than the sum of the individual frequencies as multiple detectors occasionally reject simultaneously. Overall, we find that for both, CoVaR and RCoVaR, the spurious alarms of the individual detectors are relatively balanced between the VaR and the systemic risk detectors as well as between the different assets $k=1,\dots,K$. Especially the former is noteworthy, and as desired by the choice of the level $\nu$ in steps 4 and 5 of the Algorithms (ref) and (ref).
We now analyze our procedures' power to detect misspecified forecasts. For this, Figure (ref) plots rejection frequencies under the alternative ($t^\ast<n$ in (ref)). The plots in the left panel display power against the break point $t^\ast \in \{0,1, \dots, 1000\}$, with $\beta_{\text{post}} = 0.85$ in (ref). The right panel shows power for a fixed break point at $t^\ast = 0$ for a varying degree of post-break parameter misspecification $\beta_\text{post} \in \{0.7, 0.75, 0.8, 0.85, 0.87, 0.89, 0.899\}$. We again consider $K\in\{1,2,5,10\}$ and $\alpha = \beta \in \{0.9, 0.95\}$. Notice that monitoring only starts at $t=m=250$. In the left panel, the plots recover the “joint” size from Tables (ref)--(ref) for $t^\ast = n = 1000$, and in the right panel, for $\beta_\text{post} = 0.7$.
As expected, power increases monotonically for earlier break points as well as for higher degrees of parameter misspecification in all plots, and the procedure works equally well for both probability levels. Whether a larger number $K$ of to-be-monitored stocks increases or decreases power depends on the specific situation. In general, monitoring more sequences simultaneously is subject to a trade-off between a stricter correction of the significance level and more series in which misspecifications can be detected.
When comparing the monitoring procedures across systemic risk measures, we find that the reverse CoVaR procedure is the most powerful, which can be explained by the fact that more ($2K$ instead of $K+1$) forecast series are monitored, and that more ($K$ instead of 1) VaR series are monitored simultaneously, which are not “as far in the tail” as the CoVaR.
Figure (ref) shows the joint rejection rates in the same setting as in the left panel of Figure (ref) in black, keeping $K=5$ and $\alpha = \beta = 0.9$ fixed. Additionally, the colored dashed (VaR) and dot-dashed (Systemic Risk) lines indicate how often the respective detector was the first to raise an alarm, akin to the numbered columns in Tables (ref)--(ref). We find that the systemic risk detectors for different $k \in [K]$ have identical rates of rejecting first, which is sensible given the symmetry of the DCC--GARCH DGP. Moreover, the VaR detector is more powerful than the systemic risk detectors, most likely as the VaR is not as far in the tail.
We continue to analyze the effect parameter estimation noise within the forecasts has on the rejection frequencies. Recall that, as argued in Remark (ref), we view model estimation error as part of a misspecified forecast sequence. Figure (ref) compares the rejection rates when the forecasts are based on (correctly specified) DCC--GARCH models that are estimated on an in-sample period of length $E \in \{1000, 2000, 5000, \infty\}$ with $K \in \{1,2,5,10\}$, while fixing $n=1000$, $m=250$, $t^\ast = 0$, $\beta_\text{post} = 0.85$, and $\alpha = \beta = 0.9$. We find that small(er) estimation window sizes distort the systemic risk forecasts and, hence, deliver increased rejection rates of up to $20\%$ even in the “no-break case” with $t^\ast = 1000$. This effect is somewhat more pronounced for the RCoVaR than for the CoVaR. A further effect of the model estimation is that power increases naturally for $t^\ast < 1000$ when the length of the estimation period $E$ is decreased.
Overall, our simulations reinforce the theoretical finding that the surveillance schemes hold size exactly, even when the tests are employed repeatedly at every time point. Furthermore, the power of our tests behaves naturally for varying break times, different break magnitudes, and a varying dimensionality of the monitored sequences. Finally, we find that VaR and systemic risk rejections occur relatively balanced under the null, whereas the VaR detector naturally has more power (i.e., it detects earlier), due to the systemic risk measures being further out in the tail.
We apply the systemic risk surveillance procedures to real financial data to analyze their sensitivity to sub-optimal risk forecasts in practice. For the market indicator $X_t$, we use the negative returns of the S&P\,500 Financials index (SPF) and we use the $K=5$ systemically important US banks Bank of America Corp (BAC), Citigroup Inc (C), Goldman Sachs Group Inc (GS), JPMorgan Chase & Co (JPM) and Wells Fargo & Co (WFC) as $(Y_{1t}, \dots, Y_{Kt})^\prime$. The four considered monitoring time periods each span $n=1000$ days and are chosen as (i) the global financial crisis: 23 September 2005 -- 14 September 2009, containing the bankruptcy of Lehman Brothers on 15 September 2008; (ii) a calm period: 14 January 2013 -- 30 December 2016; (iii) the COVID period: 23 March 2017 -- 12 March 2021, containing the outbreak of the COVID pandemic; and (iv) Trump's tariffs: 29 October 2021 -- 23 October 2025, containing the introduction of Trump's tariffs on 2 April 2025 (and the collapse of the Silicon Valley Bank on 10 March 2023). In all four settings, we use a rolling monitoring window of $m=250$ trading days, consider the systemic risk measures at levels $\alpha = \beta = 0.95$ and use the monitoring level of $\iota = 0.1$.
We generate systemic risk forecasts by modeling the $K+1=6$ returns by a CCC--GARCH model with Gaussian innovations and a DCC--GARCH model with Student's $t$ innovations, introduced by Eng02 and further described in Section (ref). We deliberately choose a Gaussian distribution for the CCC--GARCH model's innovations to analyze a suboptimal model. The latter DCC--GARCH--$t$ model performs relatively well in horse races of multivariate volatility models LRV12,LRV13 and is, hence, still a standard forecasting model in financial risk management. From these two models, we generate CoVaR and RCoVaR forecasts as described in Section (ref) based on the presumed distributions of the model innovations $\bm \varepsilon_t\mid\mathcal{F}_{t-1}\sim t_{\nu}(\bm R_t)$; using $\nu = \infty$ for the Gaussian case. The models are estimated in all four settings using $E = 1500$ observations, i.e., data starting approximately six years before the beginning of the monitoring period.
Figure (ref) presents the VaR and CoVaR detector values, normalized by their respective critical values, for both multivariate GARCH models across the four considered time periods. Under this normalization, values exceeding unity indicate rejections of the null hypothesis. The detector values may take negative values as a consequence of their standardization, as defined, for example, in (ref).
While we do not find rejections during the calm time for either forecasting model, the detectors raise an alarm for the CCC--GARCH--$\mathcal{N}$ forecasts in all three other considered periods. Importantly, these alarms are raised by the systemic CoVaR detector, hence implying that specifically the systemic component of the risk forecasts is misspecified. This highlights the importance of monitoring systemic risk as opposed to monitoring only the risk component (i.e., the VaR) via the procedures of HD22a+ and WWZ23. In contrast, for the DCC--GARCH--$t$ forecasts, only the VaR component raises an alarm after the start of the COVID pandemic, showing that the dynamic correlation structure implies by the DCC is much better able to capture the changing co-movements in turbulent times.
Figure (ref) provides corresponding results for the RCoVaR forecasts with overall similar findings. While there are no rejections during the calm time in either model, especially the RCoVaR detectors raise alarms for the CCC--GARCH--$\mathcal{N}$ forecasts around the times where (systemic) market risks have increased markedly. In contrast, for the Student's $t$ DCC--GARCH forecasts, mostly the VaR detectors yield rejections, implying once more that this model with its dynamic correlation structure is much better suited to model systemic risks.
As for the attribution of forecast failure to a specific institution, consider the financial crisis, which had its origins in the housing market Mis11. During this time, Wells Fargo's increased systemic riskiness was responsible for the forecast failure of the CCC--GARCH--$\mathcal{N}$ model. This holds true when Wells Fargo is viewed both as a systemic risk receiver and transmitter (see Figures (ref) and (ref), respectively). The preeminent role of Wells Fargo early on may be explained by its heavy reliance on mortgage lending (later on increased through the acquisition of Wachovia in early 2008), which contrasts with the other considered retail and investment banks.
Our empirical results are consistent with LRV12,LRV13. In their horse race of multivariate volatility models, they find that DCC and CCC specifications perform equally well during calm times, yet the former are more adequate for crises periods. We find their conclusion for volatility forecasting to also hold true for systemic risk forecasting. However, in contrast to the (repeated) “one-shot analysis” of LRV12,LRV13, our monitoring procedures also allow us to date the time point when a forecast breakdown of a specific model occurs and, equally important, also to pinpoint the model failure to a specific institution.
Regulators as well as financial institutions are particularly concerned about the commonalities in risk factors, i.e., systemic risk. To effectively take preventive measures, it becomes vital to detect changes in systemic risk assessments as soon as possible. To that end, this paper proposes formal monitoring tools for systemic risk. These are shown to work well in simulations and are useful in practice, as the empirical application to the US banking sector in Section (ref) demonstrates.
The advantages of our proposed procedures are fourfold. First, unlike classical “one-shot” backtests, our monitoring schemes control size under repeated application over a fixed horizon, as required for, e.g., daily risk monitoring in financial markets. Second, size control holds in finite samples by construction, in contrast to asymptotic one-shot backtests such as FH21. Third, our procedures accommodate multiple time series simultaneously, unlike the one-shot backtests of Bea21, which are restricted to bivariate settings (i.e., $K=1$ in our notation). Fourth, a Bonferroni-type correction enables us to attribute detected deficiencies in systemic risk forecasts to specific institutions, a key feature in regulatory applications where identification matters as much as detection.
We stress that while our empirical application deals with systemic risk in the financial system, our monitoring procedures may be used more widely in other contexts. For instance, it could be used by individual banks to monitor systemic risk forecasts in their financial positions. Again, the key concern of the institutions is not necessarily the risk inherent in their positions (as risk is associated with reward in financial markets), but the commonality in their exposures. This is because it is precisely during times of extreme co-movements that diversification benefits vanish, which---as the saying goes---is the only free lunch around. For this reason, it may also be useful for banks to monitor the systemic risk in their positions.
\singlespacing {4pt}