EconBase
← Back to paper

Systemic Risk Surveillance

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

74,494 characters

Systemic Risk Surveillance



\baselineskip18pt
\setcounter{totalnumber}{50}
\setcounter{topnumber}{50}
\setcounter{bottomnumber}{50}
\abovedisplayskip1.5ex plus1ex minus1ex
\belowdisplayskip1.5ex plus1ex minus1ex
\abovedisplayshortskip1.5ex plus1ex minus1ex
\belowdisplayshortskip1.5ex plus1ex minus1ex


\title{Systemic Risk Surveillance\thanks{
		The first author gratefully acknowledges support of the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) through grant 502572912 and the second author through grants 460479886, 531866675 and 568876076.
		Replication material for the simulations and application is available on Github under \href{https://github.com/TimoDimi/replication_CoVaR_Monitoring}{https://github.com/TimoDimi/replication\_CoVaR\_Monitoring}.
	}
}



\author{
	Timo Dimitriadis\thanks{Faculty of Economics and Business, Goethe University Frankfurt, 60629 Frankfurt am Main, Germany, and Heidelberg Institute for Theoretical Studies, [email removed].}
\and
	Yannick Hoga\thanks{Faculty of Economics and Business Administration, University of Duisburg-Essen, Universit\"atsstra\ss e 12, D--45117 Essen, Germany, [email removed].}
}


\date{\today}
\maketitle

\begin{abstract}
	\noindent
	Following several episodes of financial market turmoil in recent decades, changes in systemic risk have drawn growing attention.
	Therefore, we propose surveillance schemes for systemic risk, which allow to detect misspecified systemic risk forecasts in an ``online'' fashion.
	This enables daily monitoring of the forecasts while controlling for the accumulation of false test rejections.
	Such online schemes are vital in taking timely countermeasures to avoid financial distress.
	Our monitoring procedures allow multiple series at once to be monitored, thus increasing the likelihood and the speed at which early signs of trouble may be picked up.
	The tests hold size by construction, such that the null of correct systemic risk assessments is only rejected during the monitoring period with (at most) a pre-specified probability.
	Monte Carlo simulations illustrate the good finite-sample properties of our procedures.	An empirical application to US banks during multiple crises demonstrates the usefulness of our surveillance schemes for both regulators and financial institutions. \\

	\noindent \textbf{Keywords:} CoVaR, Forecasting, Monitoring, Multiple Testing, Systemic Risk \\
	\noindent \textbf{JEL classification:} C52 (Model Evaluation, Validation, and Selection); G17 (Financial Forecasting and Simulation); G32 (Financial Risk and Risk Management)
\end{abstract}




\section{Motivation}

\doublespacing


The numerous financial crises of recent times and their severe economic reverberations have raised awareness of the importance of systemic risk \citep{AB16,Aea17,VZ19}.
To better appreciate the difference between risk and \textit{systemic} risk, consider the returns on shares of banks.
While bank returns are often subject to higher volatility during crises (implying larger risk), this does not necessarily translate into increased commonality (implying larger \emph{systemic} risk).
However, it is precisely the increased commonality that regulators have come to be most concerned about, as this may entail system-wide distress with potentially severe economic costs \citep{GKP16}.


Therefore, it is important to evaluate systemic risk forecasts in the financial system.
While classical financial ``one-shot'' backtests are designed for such forecast evaluations, under repeated application their statistical type I errors (false test rejections) accumulate over time---up to the point that a true null will be rejected with probability approaching one.
As systemic risks are to be monitored continuously (e.g., daily), it is necessary to control \emph{accumulated} false rejections, aligning with recent interest in safe anytime-valid statistical inference \citep{shafer2021testing, vovk2021values, ramdas2023game, WWZ23}.
In this paper, we interchangeably call  methods with ``time-uniform'' false rejection guarantees monitoring procedures, surveillance schemes or online tests.

Despite the apparent need for systemic risk surveillance, there is a lack of statistically valid tools for this in the literature.
It is the main aim of this paper to fill this gap by providing such  tools and to show their validity.
To the best of our knowledge, there only exist ``one-shot'' backtests that assess the adequacy of systemic risk forecasts \citep{Bea21,FH21}.
However, their rejection rates would accumulate under repeated application.
We propose monitoring procedures for the most popular systemic risk measure, the conditional Value-at-Risk (CoVaR) of \citet{AB16}, and the reverse CoVaR (RCoVaR).
In Appendix~\ref{sex:CoESandMES}, we also develop monitoring schemes for the conditional expected shortfall (CoES) and marginal expected shortfall (MES) of  \citet{Aea17}.


For the construction of the monitoring procedure for the CoVaR as the most important systemic risk measure, we draw on recent work of \citet{FH21}.
They provide so-called {identification functions} that uniquely identify both the stand-alone risk and systemic risk.
Under the null hypothesis that the CoVaR forecasts are correctly specified (i.e., calibrated), the probabilistic structure of the corresponding identification functions is fully known: they are independent and identically distributed (IID) Bernoulli random variables. This remains true without making any additional assumptions about the dynamic structure of the return series.
As the probabilistic structure under the null is fully known, it can be used to simulate time-uniform and non-asymptotic critical values.
The simulation-based construction is particularly useful as systemic risk events are rare by definition, such that standard asymptotic theory may not be relied upon to provide accurate approximations in finite samples \citep{Hog19a+}.
We also extend our monitoring procedure to the CoVaR with reversed conditioning.


Because the probabilistic structure of the identification functions for CoES and MES is not fully characterized under the null, a direct extension of our CoVaR monitoring scheme is infeasible. Instead, in Appendix~\ref{sex:CoESandMES}, we exploit a systemic-risk analogue of the classical \textit{cumulative violation sequence} used in \citet{Acerbi2002spectral}, \citet{DE17}, \citet{Du2024powerful}, among others.
We show that, under the null, the full probabilistic structure of this sequence---namely its uniform-type distribution and independence---is known, which again enables the computation of non-asymptotic, time-uniform critical values via simulation.



Our approach of constructing the CoVaR monitoring procedure is closely related to \citet{HD22a+}, who propose online detection schemes for single \textit{univariate} risk forecasts, viz.~Value-at-Risk (VaR) and expected shortfall (ES).
However, we improve upon their procedure by using normalized detectors, and most importantly, extend their theory to \emph{systemic} risk measures.
The latter necessitates monitoring multiple time series at once, whereas \citet{HD22a+} only consider a single series.
This multivariate monitoring necessitates a Bonferroni-type correction to obtain critical values with finite-sample validity, which are not too conservative in practice.
A further benefit of the Bonferroni correction is that it allows to attribute a monitoring alarm to a specific financial institution.



Additional work that is related to ours are the monitoring procedures proposed by \citet{WG13}, \citet{NLL14} and \citet{MSW20}.
These authors do, however, not focus on systemic risk itself, but only on quantities that can at best be described as crude approximations to it, such as correlation \citep{WG13} and the copula \citep{NLL14,MSW20}.
Moreover, these procedures rely on an initial period of non-contamination, which our approach does not require.

Regarding calibration (backtesting) under repeated use, our work relates to the literature on safe and anytime-valid inference (SAVI) based on e-values. For VaR and ES, several e-value--based monitoring schemes have been proposed, but they do not extend to \emph{systemic} risk measures. While e-value tests are typically valid indefinitely, they are often overly conservative. In contrast, our approach fixes the monitoring horizon and delivers non-asymptotic, simulation-based, and hence (almost) exact size control via (almost) exact critical values.
\citet{horvath2025sequential} take yet another approach by proposing a sequential monitoring procedure for VaR and ES models that is, however, highly model-specific and based on asymptotic theory.


Our Monte Carlo simulations use a realistic DCC--GARCH model \citep{Eng02} for the asset returns to demonstrate the good finite-sample properties of our  monitoring procedures.
Despite the use of a potentially conservative Bonferroni-type correction, the empirical size remains close to the nominal level---even in moderately high-dimensional settings, where numerous individual tests could, in principle, amplify such conservative tendencies.
We also show that, under the null, the detector to first raise a false alarm pertains to the VaR or the CoVaR with equal probability. The simulations further demonstrate the high power of our procedures in identifying misspecified forecasts, again exhibiting a balanced pattern regarding which detector first signals a suboptimal predictive performance.

In the empirical application, we apply our monitoring procedure to the stock returns of systemically relevant US banks.
In doing so, we evaluate systemic risk forecasts from two competing multivariate GARCH models over a calm period and three turbulent episodes: the global financial crisis, the COVID-19 pandemic, and the recent US~tariff period under President Trump.
The monitoring results indicate that the constant covariance structure of the simpler CCC–GARCH model is inadequate for capturing the \emph{systemic} component of financial risk during such volatile periods, whereas the forecasts of the more flexible DCC–GARCH cannot be rejected.





The remainder of the paper is structured as follows.
Section~\ref{Systemic Risk Monitoring} formally introduces the CoVaR and its reverse variant and our corresponding monitoring procedures.
A related monitoring scheme for the CoES and MES is constructed in Appendix~\ref{sex:CoESandMES}.
Simulations in Section~\ref{Simulations} explore the finite-sample performance of our methods, and Section~\ref{Empirical Application} presents the empirical application.
Finally, Section~\ref{Conclusion} concludes.
Next to our CoES and MES surveillance schemes in Section~\ref{sex:CoESandMES}, the appendix contains Section~\ref{sec:Proofs of the main paper} with the proofs for our CoVaR procedures, and Section~\ref{sec:Algorithms} that contains the algorithms to compute critical values.







\section{Systemic Risk Monitoring}
\label{Systemic Risk Monitoring}

First, we define the systemic risk measures we consider in Section~\ref{Defining Systemic Risk}.
Then, we introduce monitoring procedures for the CoVaR in Section~\ref{CoVaR Monitoring}, which we extend to the reverse CoVaR in Section~\ref{ReversedCoVaRMonitoring}.
Surveillance schemes for the CoES and MES are presented in Appendix~\ref{sex:CoESandMES}.


\subsection{Defining Systemic Risk Measures}
\label{Defining Systemic Risk}

\sloppy
Consider the to-be-monitored sequence of random variables $\big\{(X_t,Y_{1t},\ldots,Y_{Kt})^\prime\big\}_{t \in \mathbb{N}}$, where $K \in \mathbb{N}$ is the number of monitored variables.
Here, $Y_{kt}$ ($k=1,\ldots,K$) stands for the log-losses of interest (e.g., the losses of a bank's shares / losses of a business unit) and $X_t$ are the log-losses of some reference position (e.g., system-wide losses in the financial system / bank-wide losses of all business units).
In a concrete monitoring situation, the variables at time $t=1,2\ldots$ are still to be (sequentially) observed and are not yet available when setting up the surveillance scheme at time $t=0$.

Let $\mathcal{F}_{t}=\sigma\big((X_t,Y_{1t},\ldots,Y_{Kt})^\prime,\bm Z_t,(X_{t-1},Y_{1,t-1},\ldots,Y_{K,t-1})^\prime,\bm Z_{t-1},\ldots\big)$ denote the time-$t$ information set with (possibly multivariate) $\bm Z_t$ containing additional covariates.
Further, we denote by $F_{(X_t,Y_{kt})^\prime\mid\mathcal{F}_{t-1}}(\cdot,\cdot)$ the cumulative distribution function (CDF) of $(X_t,Y_{kt})^\prime\mid\mathcal{F}_{t-1}$ for $k\in[K]$, where we use the shorthand notation $[K] := \{1,\dots,K\}$ for any $K \in \mathbb{N}$.
Throughout this paper, we assume that $(X_t,Y_{kt})^\prime\mid\mathcal{F}_{t-1}$ has a strictly positive Lebesgue density for all $(x,y)\in\mathbb{R}^2$ such that $F_{(X_t,Y_{kt})^\prime\mid\mathcal{F}_{t-1}}(x,y)\in(0,1)$.
This assumption ensures that we do not have to rely on generalized inverses to compute our (quantile-based) risk measures, which we introduce next.


Define the VaR as the quantile at level $\beta \in (0,1)$ of the conditional distribution $F_{X_t\mid\mathcal{F}_{t-1}}$, i.e., $\operatorname{VaR}_{t,\beta}=F_{X_t\mid\mathcal{F}_{t-1}}^{-1}(\beta)$.
The VaR is a popular \textit{univariate} risk measure.
Note that the inverse $F_{X_t\mid\mathcal{F}_{t-1}}^{-1}(\cdot)$ exists thanks to our assumption on the distribution of $(X_t, Y_{kt})^\prime\mid\mathcal{F}_{t-1}$.

We now introduce the systemic risk measures CoVaR and its reverse variant that we denote by RCoVaR.
The stress event in the definition of all these measures is that the loss of the reference position exceeds its VaR, i.e., $\{X_t\geq\operatorname{VaR}_{t,\beta}\}$.
Then, we define the CoVaR of the $k$-th series $Y_{kt}$ as
\begin{align}
	\label{eqn:DefCoVaR}
	\operatorname{CoVaR}_{kt,\alpha|\beta}=F_{Y_{kt}\mid X_t \geq \operatorname{VaR}_{t,\beta},\mathcal{F}_{t-1}}^{-1}(\alpha),
\end{align}
where $F_{Y_{kt}\mid X_t \geq \operatorname{VaR}_{t,\beta}, \mathcal{F}_{t-1}}(\cdot)=\mathbb{P}\{Y_{kt}\leq \cdot\mid X_t \geq \operatorname{VaR}_{t,\beta},\ \mathcal{F}_{t-1}\}$, and $\alpha \in (0,1)$.
As for the VaR, our assumption on the CDF of $(X_t, Y_{kt})^\prime\mid\mathcal{F}_{t-1}$ ensures that the inverse in \eqref{eqn:DefCoVaR} exists.
Note that the CoVaR is simply the $\alpha$-VaR of the distribution of $Y_{kt}$ \textit{conditional} on $\{X_t\geq\operatorname{VaR}_{t,\beta}\}$, i.e., conditional on $X_t$ being \emph{in distress}.
Without the conditioning on the distress event $\{X_t\geq\operatorname{VaR}_{t,\beta}\}$ (or, equivalently, with $\beta=0$), the CoVaR reduces to the VaR of $Y_{kt}$, such that
$\operatorname{CoVaR}_{t,\alpha|0} = F_{Y_{kt}\mid\mathcal{F}_{t-1}}^{-1}(\alpha)$.
Since $X_t$ denotes financial losses, we typically consider values for $\alpha$ and $\beta$ close to one for the CoVaR to capture (downside) systemic risk (e.g., $\alpha=\beta=0.95$).
When $\alpha=\beta$ we simply write $\operatorname{CoVaR}_{kt,\alpha} = \operatorname{CoVaR}_{kt,\alpha|\alpha}$.

Following among others \citet{GT13}, \citet{NZ20} and \citet{Bea21}, our CoVaR definition deviates from the original one of \citet{AB16}, who use $\{X_t=\operatorname{VaR}_{t,\beta}\}$ as the stress event. This latter choice is problematic because it often has probability zero and it does not fully incorporate all tail events of $X_t$.
We refer to \citet[Section~2.1]{DH24} for a detailed account of the advantages of the definition in \eqref{eqn:DefCoVaR}.

The above interpretation of the CoVaR in \eqref{eqn:DefCoVaR} with $Y_{kt}$ as losses of financial institutions and $X_t$ as market losses coincides with what \citet[Sec.~II.D]{AB16} call the \textit{Exposure CoVaR}, and it measures how institution $k$ is \emph{exposed} to market risk.
A regulator might however also be interested in the reverse direction, i.e., how an individual bank \textit{affects} the market.
Therefore, we additionally define the CoVaR with \emph{reverse} conditioning as
\begin{align}
	\label{eqn:DefRevCoVaR}
	\operatorname{RCoVaR}_{kt,\alpha|\beta} = F_{X_t\mid Y_{kt} \geq \operatorname{VaR}_{kt,\beta},\mathcal{F}_{t-1}}^{-1}(\alpha),\qquad k \in [K],
\end{align}
where $\operatorname{VaR}_{kt,\beta} = F_{Y_{kt}\mid\mathcal{F}_{t-1}}^{-1}(\beta)$ is now specific to institution $k$.
The reverse CoVaR in \eqref{eqn:DefRevCoVaR} measures the impact of distress in bank $k$ on the general market, which corresponds to the notion of banks as transmitters of systemic risk.
In contrast, our initial CoVaR definition views banks as receivers of systemic risk.
Of course, both definitions reveal useful information, and which definition is more suitable depends on the context.




\subsection{CoVaR Monitoring}
\label{CoVaR Monitoring}

We now propose procedures for monitoring the correct specification of CoVaR forecasts.
Importantly, the monitoring should be carried out in an \textit{online} fashion, i.e., sequentially as new observations become available.
Classical backtests with asymptotic validity based on a test statistic $Z_T$ with sample size $T$ and critical value $c_\iota$ for significance level $\iota \in (0,1)$ require that under the null hypothesis of correct specification, $H_0$, the asymptotic rejection probability $\lim_{T \to \infty} \mathbb{P}_{H_0}\{Z_T > c_\iota\} \le \iota$ is below $\iota$.
In contrast, sequential \emph{monitoring procedures} satisfy the stronger non-asymptotic notion
\begin{align}
	\label{eqn:MonitoringGuarantee}
	\mathbb{P}_{H_0} \big\{ \exists T \in \mathcal{I}: \; Z_T > c_\iota \big\} \le \iota.
\end{align}
This implies that even though the test decision is monitored every trading day---not unusual in a risk management context---the \textit{accumulated} probability of a  false rejection (type I error) remains controlled.
We will achieve such \textit{finite sample} type I error guarantees in \eqref{eqn:MonitoringGuarantee} by constructing detectors (essentially: test statistics), whose probabilistic structure is (almost) fully known under the null hypothesis, such that exact critical values can be calculated.





\begin{figure}[tb]
	\centering
	\includegraphics[width=\textwidth]{input/Illustration_Windows.pdf}
	\caption{Illustration of the Monitoring Windows.}
	\label{fig:MonitoringWindows}
\end{figure}

In a practical monitoring situation, we are about to sequentially observe $(X_1,Y_{11},\ldots,Y_{K1})^\prime, \ldots, (X_n,Y_{1n},\ldots,Y_{Kn})^\prime$ at times (trading days) $1,\dots,n$.
Our monitoring procedure defined below requires a \emph{rolling monitoring window} of fixed length $m \le n$, as illustrated by the blue bars in Figure~\ref{fig:MonitoringWindows}.
Thus, at each \textit{monitoring time} $T \in \{m,\dots,n\}$---the endpoint of a given monitoring window---the procedure relies on the window spanning the times $\{T-m +1 ,\dots,T\}$.
We only start monitoring when the first full window of $m$ trading days is available at time $T=m$.
While starting earlier with an initially expanding window would be possible, this would come at the cost of an explosion of technicalities, which we avoid for better readability.




We now formalize the null hypothesis of \emph{risk measure forecast adequacy}, which is closely related to the statistical notion of conditional forecast calibration; see \citet{GR_2023} for a recent and detailed treatment of forecast calibration.
The strongest notion of \textit{ideal} forecasts $\widehat{\operatorname{VaR}}_{t,\beta}$ and $\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}$ for the VaR and CoVaR is given by
\begin{equation*}
	H_0^{\operatorname{CoVaR}} \colon \widehat{\operatorname{VaR}}_{t,\beta} = \operatorname{VaR}_{t,\beta}
	\qquad \text{and} \qquad
	\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta} = \operatorname{CoVaR}_{kt,\alpha\mid\beta}, \quad k \in [K],\quad t \in \mathbb{N}.
\end{equation*}
We synonymously refer to this null as testing ideal forecasts, correct specification or calibration.

Clearly, $H_0^{\operatorname{CoVaR}}$ is not directly testable, because the true $\operatorname{VaR}_{t,\beta}$ and $\operatorname{CoVaR}_{kt,\alpha|\beta}$ are not observable even \textit{ex post}.
To circumvent this, we exploit the existence of a strict identification function for the CoVaR (jointly with the VaR), given in \citet[Theorem~S.3.1]{FH21},
\begin{align}
	\label{eqn:CoVaRJointID}
	\bm V \big((v,c)', (x,y)' \big) =
	\begin{pmatrix}\mathds{1}_{\{x\leq v\}}-\beta\\ \mathds{1}_{\{x> v\}}\big[\mathds{1}_{\{y\leq c\}}-\alpha\big]\end{pmatrix}.
\end{align}
Identification functions are at the heart of many classical backtests \citep{NZ17}.
The property that renders them useful in testing calibration is that their (conditional) expectation is zero \textit{if and only if} the true (conditional) VaR and CoVaR are inserted.
More precisely, \citet[Theorem~S.3.1]{FH21} show that
\begin{align*}
	\mathbb{E} \Big[\bm V\big( ( v,c)' , (X_t, Y_{kt})' \big) \;\Big|\; \mathcal{F}_{t-1} \Big] = \boldsymbol{0} \qquad\Longleftrightarrow\qquad
	v = \operatorname{VaR}_{t,\beta}  \;  \text{ and } \;   c = \operatorname{CoVaR}_{kt,\alpha|\beta}
\end{align*}
under our conditions on the CDF of $(X_t, Y_{kt})^\prime\mid\mathcal{F}_{t-1}$.
Hence, the relevant implication in monitoring calibration
becomes
\begin{align}\label{eq:5p}
	 \mathbb{E} \left[\bm V\left(\begin{pmatrix}\widehat{\operatorname{VaR}}_{t,\beta}\\\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}\end{pmatrix},\begin{pmatrix} X_t\\Y_{kt}\end{pmatrix}\right) \; \Bigg| \; \mathcal{F}_{t-1} \right]=\boldsymbol{0} \qquad\text{for all}\ k \in [K], \; \, t\in \mathbb{N}.
\end{align}



The following result shows that the full, joint probabilistic properties of the indicators
\begin{align}
	\label{eqn:Indicators}
	I_{t} := \mathds{1}_{\{X_t > \widehat{\operatorname{VaR}}_{t,\beta}\}}
	\qquad \text{and} \qquad
	I_{kt} := \mathds{1}_{\{X_t> \widehat{\operatorname{VaR}}_{t,\beta},\ Y_{kt}> \widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}\}}
\end{align}
is known under the null hypothesis $H_0^{\operatorname{CoVaR}}$.

\begin{prop}
	\label{prop:CoVaR null}
	Under $H_0^{\operatorname{CoVaR}}$, it holds for all $k \in [K]$ that
	\begin{align}
		\label{eqn:CoVaRH0Law}
		\big\{(I_t,I_{kt})^\prime\big\}_{t\in\mathbb{N}} \overset{d}{=} \big\{(\mathds{1}_{\{U_{1t}>\beta\}},\mathds{1}_{\{U_{1t}>\beta,\ U_{2t}>\alpha\}})^\prime\big\}_{t\in \mathbb{N}},
	\end{align}
	where $U_{it}\overset{IID}{\sim}\mathcal{U}[0,1]$ are independent of each other for  $i=1,2$ and all $t \in \mathbb{N}$.
\end{prop}


As the full probabilistic structure of $\big\{(I_t,I_{kt})^\prime\big\}_{t\in\mathbb{N}}$ is known for each $k \in [K]$ individually, the result of Proposition~\ref{prop:CoVaR null} can be used for our monitoring procedure to simulate critical values that are valid in a monitoring sense~\eqref{eqn:MonitoringGuarantee} without the need for asymptotic approximations.
Importantly, the binary nature of the CoVaR identification function in \eqref{eqn:CoVaRJointID} facilitates such a treatment, while remaining agnostic about the (conditional) distributions of $X_t$ and $Y_{kt}$.
Since for $K > 1$, Proposition~\ref{prop:CoVaR null} stays silent on the dependence structure between the contemporaneous $I_{kt}$ and $I_{k't}$ with $k \neq k'$, we will exploit a Bonferroni-type correction in Theorem~\ref{thm:CoVaR monitor} below.

For a (prospective) monitoring sample of size $n \in \mathbb{N}$, Proposition~\ref{prop:CoVaR null} leads to the testable implications of $H_0^{\operatorname{CoVaR}}$ that
\begin{equation}
	\label{eqn:Bernoulli}
	I_{t} \overset{\text{IID}}{\sim}\operatorname{Bern}(1-\beta) \quad \text{and} \quad I_{kt} \overset{\text{IID}}{\sim}\operatorname{Bern}\big((1-\alpha)(1-\beta)\big), \quad k \in [K],\ t \in [n],
\end{equation}
where $\operatorname{Bern}(\gamma)$ denotes a Bernoulli distribution with success probability $\gamma\in(0,1)$.

\begin{rem}[VaR monitoring for the CoVaR]
	\label{rem:VaRMonitoring}
	In CoVaR monitoring, it is crucial \emph{not} to disregard the VaR indicator $I_t$ in \eqref{eqn:Bernoulli}, even though superficially it only concerns the VaR while not being directly related to the CoVaR.
	To see why, consider probability levels $\alpha^\prime\neq\alpha$ and $\beta^\prime\neq\beta$ satisfying $(1-\alpha^\prime)(1-\beta^\prime)=(1-\alpha)(1-\beta)$.
	Then, by the same arguments leading to \eqref{eqn:Bernoulli}, $I^\prime_{kt}:=\mathds{1}_{\{X_t> \operatorname{VaR}_{t,\beta^\prime},\ Y_{kt}> \operatorname{CoVaR}_{kt,\alpha^\prime|\beta^\prime}\}}\overset{\text{IID}}{\sim}\operatorname{Bern}\big((1-\alpha^\prime)(1-\beta^\prime)\big) = \operatorname{Bern}\big((1-\alpha)(1-\beta)\big)$.
	Therefore, the true $\operatorname{CoVaR}_{kt,\alpha|\beta}$ are indistinguishable from the incorrect $\operatorname{CoVaR}_{kt,\alpha^\prime|\beta^\prime}$, because $I_{kt}$ and $I_{kt}^\prime$ are both IID with the same $\operatorname{Bern}\big((1-\alpha)(1-\beta)\big)$-distribution. This would allow the possibility of completely misspecified CoVaR forecasts not being detected by a monitoring procedure that disregards $I_t$.
	Thus, it is vital to also monitor for VaR calibration via $I_{t} \overset{\text{IID}}{\sim}\operatorname{Bern}(1-\beta)\neq\operatorname{Bern}(1-\beta^\prime)$.
\end{rem}

We now explain the construction of a VaR detector that monitors $I_{t} \overset{\text{IID}}{\sim}\operatorname{Bern}(1-\beta)$ from \eqref{eqn:Bernoulli}, and continue with a CoVaR detector below.
For monitoring the VaR-specific sequence $I_t$, we adapt the method of \citet{HD22a+}
and use a moving sum (MOSUM) detector inspired by the one-shot backtest of \citet{KW15} of the form
\begin{equation}
	\label{eq:VaR det}
	\operatorname{VaR}(T)=a\operatorname{VaR}^{uc}(T) + (1-a)\operatorname{VaR}^{iid}(T),
\end{equation}
where $T \in \{m,\dots,n\}$ and $a\in(0,1)$ with canonical choice $a = 0.5$.
The constant $a$ is a user-specified parameter that allows to vary the sensitivity of the procedure towards violations of correct unconditional calibration (i.e., $\mathbb{E}[I_t]=1-\beta$) and violations of IIDness of the $I_t$. These two properties are assessed by $\operatorname{VaR}^{uc}(T)$ and $\operatorname{VaR}^{iid}(T)$, respectively.
The length of the rolling monitoring window is denoted by $m\in\mathbb{N}$, but the dependence of $\operatorname{VaR}(T)$ on $m$ is not made explicit for notational brevity.


For $\operatorname{VaR}^{uc}(T)$ in \eqref{eq:VaR det}, we choose the standardized detector
\begin{align}
	\label{eqn:VaRDetStandardization}
	\operatorname{VaR}^{uc}(T) = \frac{V_T- \mathbb{E}_{H_0^{\operatorname{CoVaR}}}[V_T] }{ \sqrt{\operatorname{Var}_{H_0^{\operatorname{CoVaR}}}(V_T) }}
	\quad \text{with} \quad
	V_T = \bigg|\frac{1}{m}\sum_{t=T-m+1}^{T}I_{t}-(1-\beta)\bigg|,\qquad T\geq m.
\end{align}
The mean, $\mathbb{E}_{H_0^{\operatorname{CoVaR}}}[V_T]$, and variance, $\operatorname{Var}_{H_0^{\operatorname{CoVaR}}}(V_T)$, are obtained from simulations as the full probabilistic structure of $I_t$ is known under $H_0^{\operatorname{CoVaR}}$; see \eqref{eqn:Bernoulli}.
The standardization in $\operatorname{VaR}^{uc}(T)$ differs from \citet{HD22a+} and is used to better balance the sensitivity of the two detectors in \eqref{eq:VaR det}.
In contrast to $\mathbb{E}_{H_0^{\operatorname{CoVaR}}}[V_T]$ and $\operatorname{Var}_{H_0^{\operatorname{CoVaR}}}(V_T)$, the sequence $V_T$ has to be computed from the data (i.e., observations and appertaining forecasts) and is responsible for giving the procedure its power.
Specifically, $V_T$ is designed to uncover deviations from $\mathbb{E}[I_t]=1-\beta$, i.e., correct unconditional calibration.
By construction, large values of $\operatorname{VaR}^{uc}(T)$ provide evidence against the null, where deviations in both directions (i.e., $\mathbb{E}[I_t]>1-\beta$ and $\mathbb{E}[I_t]<1-\beta$) raise an alarm.\footnote{
	One could also consider one-sided detectors here. However, we refrain from doing so, because it remains unclear how a too conservative (and, hence, classically acceptable) misspecification of VaR forecasts interact with the validity of the CoVaR forecasts in light of the joint identification function in \eqref{eqn:CoVaRJointID}; also see \citet[Remark~2]{WWZ23}.
}




For $\operatorname{VaR}^{iid}(T)$ in \eqref{eq:VaR det}, we consider the durations $d_i:=t_i-t_{i-1}$ ($t_0:=T-m$) between VaR violation times in the window $\{T-m+1,\ldots,T\}$.
These are defined as
\begin{align*}
	\{t_1,\ldots,t_{S_{T}}\} =\big\{t\in\{T-m+1,\ldots,T\}\ :\ I_{t}=1\big\},
\end{align*}
where $S_{T} := \sum_{t=T-m+1}^{T}I_t$ counts the number of VaR violations in the window $\{T-m+1,\ldots,T\}$.
As a measure of the ``inequality'' between the durations $d_i$, we follow \citet{KW15} and use the Gini coefficient
\begin{align*}
	g_{T} &= \frac{\frac{1}{S_{T}^2}\sum_{i,j=1}^{S_{T}}|d_i-d_j|}{2\overline{d}},
	\qquad \text{with} \qquad
	 \overline{d}=\frac{1}{S_{T}}\sum_{i=1}^{S_{T}}d_i.
\end{align*}
The detector for the second part of \eqref{eq:VaR det} then is the standardized Gini coefficient
\begin{align*}
	\operatorname{VaR}^{iid}(T) = \frac{g_{T} -  \mathbb{E}_{H_0^{\operatorname{CoVaR}}}[g_T]}{ \sqrt{\operatorname{Var}_{H_0^{\operatorname{CoVaR}}}(g_T) }},
\end{align*}
where $\mathbb{E}_{H_0^{\operatorname{CoVaR}}}[g_T]$ and $\operatorname{Var}_{H_0^{\operatorname{CoVaR}}}(g_T)$ are the mean and variance, respectively, of $g_T$ under the null.
As above, these moments are derived from Monte Carlo simulations, and it is only $g_T$ itself that is computed from the data.


The idea underlying the use of $\operatorname{VaR}^{iid}(T)$ is to test the IID~property of $I_t$ by considering the spacing between the VaR exceedances (i.e., those $t$ for which $I_t=1$). Unevenly spaced exceedances with large values for the Gini coefficient (and, hence, large values of $\operatorname{VaR}^{iid}(T)$) weigh against IIDness.
Note that in principle also too evenly spaced exceedances (with small values for the Gini coefficient) provide evidence against $H_0^{\operatorname{CoVaR}}$.
However, such a violation of the null is of no concern in risk management, where only a clustering of exceedances is seen as problematic \citep[see, e.g.,][Fig.~1]{KW15}.


To also monitor the CoVaR forecasts, i.e., to sequentially test $I_{kt}\overset{\text{IID}}{\sim}\operatorname{Bern}\big((1-\alpha)(1-\beta)\big)$ from \eqref{eqn:Bernoulli}, we use the same strategy as for monitoring VaR predictions.
Therefore, we define the CoVaR detectors $\operatorname{CoVaR}_{k}(T)$ for any $k \in [K]$ similarly as the VaR detector $\operatorname{VaR}(T)$:
\begin{align}
	\label{eq:CoVaR det}
	\operatorname{CoVaR}_{k}(T)=a\operatorname{CoVaR}^{uc}_{k}(T) + (1-a)\operatorname{CoVaR}^{iid}_{k}(T).
\end{align}
Here, the right-hand side quantities are defined in analogy to those in \eqref{eq:VaR det}, with the only difference that $I_{kt}$ replaces $I_t$, and $(1-\alpha)(1-\beta)$ replaces $1-\beta$ at every occurrence.

The following theorem shows how the probability of a false detection in the sense of \eqref{eqn:MonitoringGuarantee} can be bounded uniformly over time and over different institutions $k \in [K]$, for which the CoVaR is forecasted.

\begin{thm}
	\label{thm:CoVaR monitor}
	For any significance level $\iota \in (0,1)$, it holds that
	\begin{multline}\label{eqn:CoVaRMonitoringTypeIError}
		\mathbb{P}_{H_0^{\operatorname{CoVaR}}} \Big\{ \exists T \in \{m, \dots, n\}: \quad \operatorname{VaR}(T) \ge v  \\ \text{  or  }  \;   \operatorname{CoVaR}_{k}(T) \ge c_k \quad \text{for some $k \in [K]$}  \Big\} \le \iota,
	\end{multline}
	if the critical values $v$ and the $c_k$'s are chosen such, that
	\begin{multline}
		\mathbb{P}_{H_0^{\operatorname{CoVaR}}}\Big\{\sup_{T=m,\ldots,n}\operatorname{VaR}(T) \geq v \Big\} + \sum_{k=1}^{K}\mathbb{P}_{H_0^{\operatorname{CoVaR}}}\Big\{\sup_{T=m,\ldots,n}\operatorname{CoVaR}_{k}(T) \geq c_k\Big\} \\
		- \sum_{k=1}^{K}\mathbb{P}_{H_0^{\operatorname{CoVaR}}}\Big\{\sup_{T=m,\ldots,n}\operatorname{VaR}(T) \geq v,\ \sup_{T=m,\ldots,n}\operatorname{CoVaR}_{k}(T) \geq c_k\Big\}=\iota.
		\label{eq:size ineq}
	\end{multline}
\end{thm}

In other words, Theorem~\ref{thm:CoVaR monitor} shows that size at level $\iota\in(0,1)$ is controlled \textit{in finite samples} if we reject $H_0^{\operatorname{CoVaR}}$ as soon as either
	\begin{equation*}
		\operatorname{VaR}(T) \geq v \quad\text{or}\quad \operatorname{CoVaR}_{k}(T) \geq c_k\qquad\text{for some}\ T=m,\ldots,n,\quad k \in [K].
	\end{equation*}

The two probabilities in the upper row of \eqref{eq:size ineq} are associated with a standard Bonferroni correction for testing the $K+1$ hypotheses of calibration of $\widehat{\operatorname{VaR}}_{t,\beta}$ and $\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}$ for $k \in [K]$.
However, the presence of the \textit{negative} third term on the right-hand side of \eqref{eq:size ineq} allows for a less conservative---and, hence, potentially more powerful---procedure.
We mention that there may not exist $v$ and $c_k$, such that the equality in \eqref{eq:size ineq} holds exactly, because of the binary nature of the detectors.
Therefore, in practice $v$ and $c_k$ should be chosen to lead to a sum that is smaller but as close as possible to $\iota$; see also step~\ref{it:final} in Algorithm~\ref{algo:1} below.


For the practical computation of the critical values $v$ and $c_k$, we use the following Algorithm~\ref{algo:1}, which approximates the probabilities in \eqref{eq:size ineq} by using the probabilistic structure in \eqref{eqn:CoVaRH0Law} to sample under the null.



\begin{algorithm}
	\label{algo:1}
	To compute critical values for Theorem~\ref{thm:CoVaR monitor} proceed as follows:
	\begin{enumerate}
		\item
		\label{it:1}
		Generate a large number $B$ of mutually independent samples $\big\{U_{1t}\overset{\text{IID}}{\sim}\mathcal{U}[0,1]\big\}_{t \in [n]}$ and $\big\{U_{2t}\overset{\text{IID}}{\sim}\mathcal{U}[0,1]\big\}_{t \in [n]}$.

		\item
		\label{it:2}
		\sloppy
		Compute the $B$ sequences $\big\{I_{t}^\ast=\mathds{1}_{\{U_{1t} > \beta\}}\big\}_{t \in [n]}$ and $\big\{I_{1t}^{\ast}=\mathds{1}_{\{U_{1t} > \beta,\ U_{2t}>\alpha \}}\big\}_{t \in [n]}$.

		\item\label{it:3}
		For $b=1,\ldots,B$, calculate:
		\begin{itemize}
			\item[i.]
			$\sup_{T=m,\ldots,n}\operatorname{VaR}^{b}(T)$, where $\operatorname{VaR}^{b}(T)$ is defined as $\operatorname{VaR}(T)$, except that $\{I_t\}$ is replaced by the $b$-th sample from $\{I_t^{\ast}\}$ from step \ref{it:2} of this algorithm;

			\item[ii.]
			$\sup_{T=m,\ldots,n}\operatorname{CoVaR}^{b}(T)$, where $\operatorname{CoVaR}^{b}(T)$ is defined as $\operatorname{CoVaR}_{1}(T)$, except that $\{I_{1t}\}$ is replaced by the $b$-th sample from $\{I_{1t}^{\ast}\}$ from step \ref{it:2} of this algorithm.
		\end{itemize}

		\item On a fine grid for $\nu\in[0,1]$, compute the empirical $(1-\nu)$-quantiles of
		\begin{itemize}
			\item[i.] $\big\{\sup_{T=m,\ldots,n}\operatorname{VaR}^{b}(T)\big\}_{b \in [B]}$, which we denote by $v(\nu)$, and
			\item[ii.] $\big\{\sup_{T=m,\ldots,n}\operatorname{CoVaR}^{b}(T)\big\}_{b \in [B]}$, which we denote by $c(\nu)$.
		\end{itemize}

		\item\label{it:final} Find the value $\nu \in [0,1]$ for which
		\begin{multline*}
			\frac{1}{B}\sum_{b=1}^{B}\mathds{1}_{\big\{\sup_{T=m,\ldots,n} \operatorname{VaR}^{b}(T)\geq v(\nu) \big\}} + \frac{K}{B}\sum_{b=1}^{B}\mathds{1}_{\big\{\sup_{T=m,\ldots,n} \operatorname{CoVaR}^{b}(T)\geq  c(\nu)\big\}} \\
			- \frac{K}{B}\sum_{b=1}^{B} \mathds{1}_{\big\{\sup_{T=m,\ldots,n}\operatorname{VaR}^{b}(T)\geq v(\nu),\ \sup_{T=m,\ldots,n}\operatorname{CoVaR}^{b}(T)\geq c(\nu)\big\}}
		\end{multline*}
		is equal to (or smaller than) $\iota$; cf.~\eqref{eq:size ineq}. The appertaining critical values will be denoted by
		$v^{\operatorname{CoVaR}}$ and $c^{\operatorname{CoVaR}}$.

	\end{enumerate}
\end{algorithm}

The above algorithm also overcomes the theoretical ambiguity that there exists a continuum of critical values $v$ and $c_k$ that ensure the probabilities in \eqref{eq:size ineq} sum to $\iota$.
In principle, such a continuum allows to weight the importance of the VaR and CoVaR hypotheses.
However, the fourth step in Algorithm~\ref{algo:1} advocates a natural approach that balances the importance of VaR and CoVaR by choosing $v$ and $c_k$ to correspond to some $(1-\nu)$-quantile of the variables $\sup_{T=m,\ldots,n}\operatorname{VaR}(T)$ and $\sup_{T=m,\ldots,n}\operatorname{CoVaR}_{k}(T)$, respectively.
Then, $\nu$ in step~5 simply has to be chosen to ensure that the probabilities in \eqref{eq:size ineq} sum to $\iota$.
Note that this entails the computation of two critical values only---one for the VaR (written $v^{\operatorname{CoVaR}}$) and one for the CoVaR (written $c^{\operatorname{CoVaR}}$), where the latter one is the same for all $k \in [K]$, because the distribution of $I_{kt}\overset{\text{IID}}{\sim}\operatorname{Bern}\big((1-\alpha)(1-\beta)\big)$ is independent of $k$.
These critical values have the superscript ``${\operatorname{CoVaR}}$'' to distinguish them from critical values used for monitoring different systemic risk measures.


Clearly, the use of Boole's inequality for $K > 1$ (in the proof of Theorem~\ref{thm:CoVaR monitor}) implies that our monitoring procedure is conservative, indicated by the \emph{inequality} in  \eqref{eqn:CoVaRMonitoringTypeIError}.
Only for $K=1$, this becomes an equality such that size is kept exactly.
While a conservative test reduces the probability of actually making a type I error, this usually has the drawback of increasing the probability of a type II error, such that the ability to identify (systemic) risk changes is reduced.
However, we show in simulations in Section~\ref{Simulations} that our procedure is not too conservative for values of $K$ representative of practical applications.


A possible reason for the good performance of the Bonferroni correction is explained by \citet{Hol79}, who states that: ``The power gain obtained by using a sequentially rejective Bonferroni test [based on ordering $p$-values] instead of a classical Bonferroni test depends very much upon the alternative.
It is small if all the hypotheses are `almost true', but it may be considerable if a number of hypotheses are `completely wrong'.''
In our case, systemic risk builds up slowly for all banks, rendering all hypotheses 'almost true' at first, thus leading to good detection properties of our simple Bonferroni-type corrections.

The use of the Bonferroni-type correction has two further advantages.
First, it ensures controlled size in finite samples.
In particular, we do not have to rely on large-sample asymptotics, which may be unreliable in backtesting contexts \citep{Hog19a+}.
Second, the Bonferroni correction allows us to attribute a rejection of the null to a specific institution.
For instance, if the detector $\operatorname{CoVaR}_{k^{\ast}}$ for some $k^\ast \in [K]$ is the first to raise an alarm, the rejection of the null can be pinpointed to bank $k^\ast$.
Of course, precisely identifying the most vulnerable institution is of utmost importance in systemic risk analysis.\footnote{
	A further implication of $H_0^{\operatorname{CoVaR}}$ is that $I_s$ and $I_{kt}$ are independent for $s\neq t$. We do not explicitly reflect this implication in the construction of our detectors. While doing so may increase the power of the detector (depending, of course, on the specific type of alternative), this would again render infeasible the attribution of a rejection of $H_0^{\operatorname{CoVaR}}$ to a specific institution.
}



\begin{rem}[Flexibility of detector specification]
	Other CoVaR monitoring detectors $\operatorname{VaR}(T)$ and $\operatorname{CoVaR}_k(T)$ depending solely on $\big\{(I_t,I_{kt})^\prime\big\}_{t\in [n]}$ can be constructed that preserve the accumulated type I error control in \eqref{eqn:MonitoringGuarantee}; compare Algorithm~\ref{algo:1}.
	This flexibility enables the construction of detectors specifically tailored to the particular form of misspecification one aims to detect.
	The same observation applies to the monitoring procedures for the RCoVaR in Section~\ref{ReversedCoVaRMonitoring}, and CoES and MES in Appendix~\ref{sex:CoESandMES}.
\end{rem}






\subsection{CoVaR Monitoring with Reversed Conditioning}
\label{ReversedCoVaRMonitoring}


Section~\ref{CoVaR Monitoring} focuses on a monitoring procedure for the CoVaR in \eqref{eqn:DefCoVaR}, which measures the individual exposure of (e.g., financial institution) $Y_{kt}$ onto the joint (market) $X_t$.
Now, we consider monitoring of the RCoVaR in \eqref{eqn:DefRevCoVaR}, which measures the effect the individual (e.g., financial institution) $Y_{kt}$ has on the joint (market) $X_t$.


Similarly as above, for given VaR forecasts $\widehat{\operatorname{VaR}}_{kt,\beta}$ and RCoVaR forecasts $\widehat{\operatorname{RCoVaR}}_{kt,\alpha|\beta}$, the null is
\begin{equation*}
	H_0^{\operatorname{RCoVaR}} \colon \widehat{\operatorname{VaR}}_{kt,\beta} = \operatorname{VaR}_{kt,\beta}\qquad	\text{and} \qquad
	\widehat{\operatorname{RCoVaR}}_{kt,\alpha|\beta} = \operatorname{RCoVaR}_{kt,\alpha|\beta}, \quad k \in [K], \quad t \in \mathbb{N}.
\end{equation*}

Employing analogous arguments as in Section~\ref{CoVaR Monitoring}, we rely on a testable implication of $H_0^{\operatorname{RCoVaR}}$ that is based on the indicators $I_{kt}^{v} := \mathds{1}_{\{Y_{kt}> \widehat{\operatorname{VaR}}_{kt,\beta}\}}$ and $I_{kt}^{c} := \mathds{1}_{\{Y_{kt}> \widehat{\operatorname{VaR}}_{kt,\beta},\ X_t> \widehat{\operatorname{RCoVaR}}_{kt,\alpha|\beta}\}}$.
Specifically, we have the following analog to Proposition~\ref{prop:CoVaR null}.



\begin{prop}
	\label{prop:CoVaR rev null}
	Under $H_0^{\operatorname{RCoVaR}}$, it holds for all $k \in [K]$ that
	\[
		\big\{(I_{kt}^v,I_{kt}^c)^\prime\big\}_{t\in\mathbb{N}} \overset{d}{=} \big\{(\mathds{1}_{\{U_{1t}>\beta\}},\mathds{1}_{\{U_{1t}>\beta,\ U_{2t}>\alpha\}})^\prime\big\}_{t\in\mathbb{N}},
	\]
	where $U_{it}\overset{IID}{\sim}\mathcal{U}[0,1]$ are independent of each other for  $i=1,2$ and all $t \in \mathbb{N}$.
\end{prop}

The proof is similar to that of Proposition~\ref{prop:CoVaR null} and, hence, omitted.
Proposition~\ref{prop:CoVaR rev null} suggests the following surveillance approach.
VaR changes may be monitored via the detector $\operatorname{VaR}_{k}(T)$, which is defined in analogy to $\operatorname{VaR}(T)$ at \eqref{eq:VaR det} with $I_t$ replaced by $I_{kt}^{v}$.
Similarly, CoVaR surveillance now uses the detector $\operatorname{RCoVaR}_{k}(T)$, where the definition is identical to that of $\operatorname{CoVaR}_{k}(T)$ at \eqref{eq:CoVaR det} except that $I_{kt}^{c}$ is used in place of $I_{kt}$.
\color{black}

\begin{thm}
	\label{thm:CoVaR rev monitor}
	For any significance level $\iota \in (0,1)$, it holds that
	\begin{multline}
		\label{eqn:RevCoVaRMonitoringTypeIError}
		\mathbb{P}_{H_0^{\operatorname{RCoVaR}}} \Big\{ \exists T \in \{m,\ldots,n\}: \quad \operatorname{VaR}_{k}(T) \ge v_k  \\
		\text{  or  }   \operatorname{RCoVaR}_{k}(T) \ge c_k \quad \text{for some $k \in [K]$}  \Big\} \le \iota,
	\end{multline}
	if the critical values $v_k$ and $ c_k$ are chosen such, that
	\begin{equation}
		\label{eq:size ineq rev}
		\sum_{k=1}^{K}\mathbb{P}_{H_0^{\operatorname{RCoVaR}}}\bigg\{\Big\{\sup_{T=m,\ldots,n}\operatorname{VaR}_{k}(T) \geq v_k \Big\}\cup\Big\{\sup_{T=m,\ldots,n}\operatorname{RCoVaR}_{k}(T) \geq c_k \Big\}\bigg\}=\iota.
	\end{equation}
\end{thm}

As in Section~\ref{CoVaR Monitoring}, almost the full probabilistic structure of the detectors is known under $H_0^{\operatorname{RCoVaR}}$, such that the probabilities in \eqref{eq:size ineq rev} and, thereby, the critical values $v_k$ and $c_k$ can be computed explicitly via simulations.
For this, we employ Algorithm~\ref{algo:RevCoVaR} in Appendix~\ref{sec:Algorithms}, which is very similar to Algorithm~\ref{algo:1}, but takes into account the differences between \eqref{eq:size ineq} and \eqref{eq:size ineq rev}.
Again, Algorithm~\ref{algo:RevCoVaR} only computes two critical values---one for the VaR component ($v^{\operatorname{RCoVaR}}$), and one for the RCoVaR component ($c^{\operatorname{RCoVaR}}$).


As in Theorem~\ref{thm:CoVaR monitor}, our procedure is less conservative than a standard Bonferroni correction.
Recall that a Bonferroni correction suggests monitoring each of the $2K$ hypotheses at a significance level of $\iota/(2K)$, such that the total probability of rejection under the null is bounded from above by $\iota$.
In contrast, our monitoring scheme based on Theorem~\ref{thm:CoVaR rev monitor} is less conservative because the sum in \eqref{eq:size ineq rev} is less than the ``Bonferroni sum'', i.e.,
\begin{multline*}
	\sum_{k=1}^{K}\mathbb{P}_{H_0^{\operatorname{RCoVaR}}}\bigg\{\Big\{\sup_{T=m,\ldots,n}\operatorname{VaR}_{k}(T) \geq v_k \Big\}\cup\Big\{\sup_{T=m,\ldots,n}\operatorname{RCoVaR}_{k}(T) \geq c_k \Big\}\bigg\}\\
	\leq \sum_{k=1}^{K}\bigg[\mathbb{P}_{H_0^{\operatorname{RCoVaR}}}\Big\{\sup_{T=m,\ldots,n}\operatorname{VaR}_{k}(T) \geq v_k\Big\}+\mathbb{P}_{H_0^{\operatorname{RCoVaR}}}\Big\{\sup_{T=m,\ldots,n}\operatorname{RCoVaR}_{k}(T) \geq c_k \Big\}\bigg].
\end{multline*}




\begin{rem}[Model-estimation error]
	\label{rem:ModelEstimation}
	We intentionally exclude model-estimation error from the uncertainty quantification of our detectors. Consequently, we treat the forecasts as the outcome of a particular modeling choice---including its estimation---and therefore evaluate the forecasting model and the estimation method jointly as in \citet{DM95}, \citet{GW06} and \citet{RS19}. In this sense, the risk forecasts are regarded as the final product that must satisfy the respective null hypotheses, irrespective of the underlying model and its estimation. This approach parallels \citet{HD22a+}, who provide several arguments in favor of this practice for monitoring risk forecasts, all of which carry over directly to the present setting.
	See Figure~\ref{fig:PowerEstWindowCoVaR} for an analysis of our procedures' sensitivity to model-estimation error.
\end{rem}



\section{Simulations}
\label{Simulations}

Here, we investigate size and power of our monitoring procedures.
We place a particular emphasis on studying size, because Bonferroni-type corrections---such as those underlying our monitoring schemes---are known to be conservative when many hypotheses are tested.



\subsection{The DCC--GARCH Data-Generating Process}
\label{Data-Generating Process}

We simulate the variables $\big\{(X_t,Y_{1t},\ldots,Y_{Kt})^\prime\big\}_{t=-E+1,\ldots,n}$ from a DCC--GARCH model of \citet{Eng02}.
The sequence consists of $E \in \mathbb{N}$ time points for model estimation (used in one simulation setting), and $n \in \mathbb{N}$ time points for generating forecasts and their evaluation.
The data-generating process (DGP) is
\begin{equation*}
	\bm W_t:=(X_t,Y_{1t},\ldots,Y_{Kt})^\prime=\bm D_t \bm \varepsilon_t,\qquad\bm \varepsilon_t\mid\mathcal{F}_{t-1}\sim t_{\nu}(\bm R_t),
\end{equation*}
where $\bm D_t^2$ is the $\mathcal{F}_{t-1}$-measurable diagonal matrix containing the componentwise conditional variances of $\bm W_t\mid\mathcal{F}_{t-1}$, the matrix $\bm R_t$ is the conditional correlation of $\bm W_t\mid\mathcal{F}_{t-1}$, and $t_{\nu}(\bm R_t)$ denotes the multivariate $t$-distribution with degrees of freedom equal to $\nu>2$, correlation matrix $\bm R_t$ and marginals standardized to have zero mean and unit variance.
Therefore, the conditional variance-covariance matrix of $\bm W_t\mid\mathcal{F}_{t-1}$ is $\bm H_t=\bm D_t\bm R_t\bm D_t$. The matrices $\bm D_t$ and $\bm R_t$ are modeled in the DCC--GARCH fashion of \citet{Eng02} as
\begin{align*}
	\bm D_t^2 &= \operatorname{diag}(\bm \omega_{G}) + \operatorname{diag}(\bm \alpha_{G})\circ \bm W_{t-1}\bm W_{t-1}^{\prime} + \operatorname{diag}(\bm \beta_{G})\circ\bm D_{t-1}^2,\\
	\bm Q_t &= \overline{\bm Q}(1-\alpha_Q-\beta_Q) + \alpha_Q \bm \varepsilon_{t-1}\bm \varepsilon_{t-1}^\prime + \beta_Q \bm Q_{t-1},\\
	\bm R_t &= \operatorname{diag}(\bm Q_t)^{-1/2}\bm Q_t\operatorname{diag}(\bm Q_t)^{-1/2},
\end{align*}
where $\overline{\bm Q}$ is the unconditional correlation matrix of the ``devolatilized'' series $\bm \varepsilon_{t-1}$, and  $\operatorname{diag}(\bm a)$ and  $\operatorname{diag}(\bm A)$ denote the diagonal matrices containing the elements of the vector $\bm a$ and the diagonal elements of the matrix $\bm A$, respectively.
Note that both the conditional variance matrix $\bm D_t^2$ and the conditional correlation matrix $\bm R_t$ evolve in a GARCH-type fashion (the latter through the recursion for $\bm Q_t$).


We parametrize the process as follows:
We choose $\overline{\bm Q}$ to be the matrix with ones on the main diagonal and 0.5 elsewhere (i.e., an equicorrelation matrix), such that the unconditional correlation between any two elements of $\bm \varepsilon_t$ equals 0.5.
Furthermore, we fix the parameters $\nu=5$, $\bm \omega_G=(0.1,\ldots, 0.1)^\prime$, $\bm \alpha_G=(0.1,\ldots,0.1)^\prime$, and $\alpha_{Q}=0.1$.
		In contrast, the autoregressive parameters $\bm \beta_G=(\beta_{G,1},\ldots,\beta_{G,K+1})^\prime$ and $\beta_{Q}$ vary over time.
		Specifically,
		\begin{align}
			\label{eqn:SimParamMisspec}
			\beta_{G,i} = \beta_{G,i,t} =\begin{cases}  0.7, & t\leq t^{\ast},\\
				\beta_\text{post}, & t> t^{\ast},
			\end{cases}
			\qquad\text{and}\qquad \beta_Q = \beta_{Q,t} =\begin{cases}  0.7, & t\leq t^{\ast},\\
				\beta_\text{post}, & t> t^{\ast},
			\end{cases}
		\end{align}
		such that there is an upward change in the persistences at time $t^{\ast} \in [n]$.
		In the baseline simulation setup, we fix $\beta_\text{post} = 0.85$.

		Through \eqref{eqn:SimParamMisspec}, we misspecify the dynamics of the process, which will then be picked up by the monitoring procedure.
		Of course, if $t^{\ast}=n$, then there is no structural change in the sample, and the systemic risk forecasts (issued from the estimates of the initial fixed window) are correctly specified.
		Hence, except for the estimation error, the forecasts are correctly specified, such that a rejection probability close to the nominal level should be expected.







		\subsection{Systemic Risk Forecasts from DCC--GARCH Models}
		\label{sec:SystemicRiskForecasts}

		For generating the forecasts, we either use fixed DCC--GARCH  parameters or (in one setup) estimate the parameters based on a sample $\big\{(X_t,Y_{1t},\ldots,Y_{Kt})^\prime\big\}_{t=-E+1,\ldots,0}$ of length $E$ by using  the \texttt{R} package \texttt{rmgarch} \citep{rmgarch}.
		In doing so, we estimate the marginals via standard Gaussian quasi-maximum likelihood estimation (to obtain $K+1$ estimates of $\omega_{G,i}$, $\alpha_{G,i}$ and $\beta_{G,i}$).
		The dependence parameters ($\nu$, $\alpha_Q$ and $\beta_Q$) are estimated in a second step using a multivariate $t$-assumption for the $\bm \varepsilon_t$.


		Given DCC-forecasts (either based on fixed or estimated model parameters) for $\widehat{\bm D}_t$, $\widehat{\bm R}_t$ and $\widehat{\bm H}_t$, we obtain the VaR forecasts at time $t$ as $\widehat{\operatorname{VaR}}_{t,\beta} = \widehat{\bm D}_{t,11}^{-1} q_{t_\nu}(\beta)$, where $q_{t_\nu}(\beta)$ denotes the $\beta$-quantile of the univariate $t_{{\nu}}$-distribution with unit variance and (either fixed or estimated) degrees of freedom $\nu$.
		For the CoVaR forecasts, we are not aware of a closed-form solution.
		Instead, we apply a root finding algorithm to approximate $\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}$ via
		\begin{align}
			\label{eqn:CoVaRFCApprox}
			\mathbb{P}
			\Big\{ Y_{kt} \ge \widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}, \; X_t \ge \widehat{\operatorname{VaR}}_{t,\beta} \;\Big| \; \mathcal{F}_{t-1} \Big\} -  (1-\alpha) (1-\beta) = 0.
		\end{align}
		The probability in  \eqref{eqn:CoVaRFCApprox} is taken with  respect to $(X_t, Y_{kt})\mid\mathcal{F}_{t-1} \sim t_\nu(\widehat{\bm H}_{t,k})$, where $\widehat{\bm H}_{t,k} = \big( ( \widehat{\bm H}_{t, 11}, \widehat{\bm H}_{t, 1 (k+1)})'  \mid ( \widehat{\bm H}_{t, 1 (k+1)}, \widehat{\bm H}_{t, (k+1) (k+1)})' \big)$
		consists of the respective  entries of the forecasted variance-covariance matrix $\widehat{\bm H}_{t}$.
		Based on the forecasts $\widehat{\operatorname{VaR}}_{t,\beta}$ and $\widehat{\operatorname{CoVaR}}_{kt,\alpha|\beta}$, we compute the VaR and CoVaR indicators as described in Sections~\ref{CoVaR Monitoring}--\ref{ReversedCoVaRMonitoring}.





		\subsection{Simulation Results under the Null Hypothesis}
		\label{sec:Simulation_Results_Null}


		All following results are based on 5000 simulation replications.
		We first fix $n=1000$, $t^\ast=n$, $m=250$, and generate the systemic risk forecasts based on the correct DCC--GARCH parameters from the ``pre-break'' period such that we omit parameter estimation noise.
		We deliberately do not compare against one-shot systemic risk backtests of, e.g., \citet{Bea21} and \citet{FH21} since these---opposed to our monitoring procedure---accumulate type I errors when applied repeatedly.


		\begin{table}[tb]
			\centering
			\resizebox{\linewidth}{!}{
			\small
			\begin{tabular}{llllrlrrrrrrrrrrrr}
				\toprule
				$\alpha = \beta$ &  & $K$ &  & Joint &  & VaR & & \multicolumn{10}{c}{CoVaR for series number $k$} \\
				\cmidrule{9-18}
				&&&&&&&& 1 & 2 & 3 & 4 & 5 & 6 & 7 & 8 & 9 & 10 \\
				\midrule
				\multirow{4}{*}{0.9} &  & 1 &  & 10.00 &  & 5.30 && 4.70 &  &  &  &  &  &  &  &  &  \\
				&  & 2 &  & 10.28 &  & 3.30 && 3.46 & 3.56 &  &  &  &  &  &  &  &  \\
				&  & 5 &  & 7.58 &  & 1.46 && 1.14 & 1.32 & 1.40 & 1.10 & 1.32 &  &  &  &  &  \\
				&  & 10 &  & 8.36 &  & 0.78 && 0.92 & 0.74 & 0.96 & 0.80 & 0.64 & 0.54 & 0.88 & 0.86 & 0.78 & 0.58 \\
				\midrule
				\multirow{4}{*}{0.95} &  & 1 &  & 9.62 &  & 4.28 && 5.36 &  &  &  &  &  &  &  &  &  \\
				&  & 2 &  & 9.26 &  & 3.46 && 3.04 & 2.82 &  &  &  &  &  &  &  &  \\
				&  & 5 &  & 8.24 &  & 1.56 && 1.22 & 1.24 & 1.38 & 1.50 & 1.52 &  &  &  &  &  \\
				&  & 10 &  & 6.38 &  & 0.68 && 0.66 & 0.60 & 0.52 & 0.74 & 0.60 & 0.52 & 0.46 & 0.68 & 0.70 & 0.50 \\
				\bottomrule
			\end{tabular}
		}
			\caption{Joint and disaggregated \textit{first} rejection rates (in percent) of our CoVaR surveillance procedure under the null hypothesis of correctly specified forecasts.
			Results are reported for a nominal level of $\iota = 10\%$, for $K \in \{1,2,5,10\}$ and $\alpha = \beta \in \{0.9, 0.95\}$.
			The column ``VaR'' reports the percentage of cases in which the VaR detector raised the first alarm, while the columns labeled ``1''--``10'' report the corresponding first-alarm rates of the respective CoVaR detectors.}
			\label{tab:SizeCoVaR}
		\end{table}



		\begin{table}[tb]
			\centering
			\resizebox{\linewidth}{!}{
			\small
			\begin{tabular}{llllrlllrrrrrrrrrr}
				\toprule
				&&&&&&&& \multicolumn{10}{c}{Series number $k$} \\
				\cmidrule{9-18}
				$\alpha = \beta$ &  & $K$ &  & Joint &  & Measure &  & 1 & 2 & 3 & 4 & 5 & 6 & 7 & 8 & 9 & 10 \\
				\midrule
				&  & \multirow{2}{*}{2} &  & \multirow{2}{*}{9.32} &  & VaR &  & 2.22 & 2.48 &  &  &  &  &  &  &  &  \\
				&  & &  &  &  & RCoVaR &  & 2.32 & 2.44 &  &  &  &  &  &  &  &  \\[0.3em]
				\multirow{2}{*}{0.9} &  & \multirow{2}{*}{5} &  & \multirow{2}{*}{6.96} &  & VaR &  & 0.70 & 0.88 & 0.62 & 0.92 & 1.06 &  &  &  &  &  \\
				&  &  &  &  &  & RCoVaR &  & 0.66 & 0.56 & 0.56 & 0.68 & 0.62 &  &  &  &  &  \\[0.3em]
				&  & \multirow{2}{*}{10} &  & \multirow{2}{*}{7.70} &  & VaR &  & 0.48 & 0.40 & 0.50 & 0.44 & 0.36 & 0.48 & 0.64 & 0.62 & 0.54 & 0.50 \\
				&  &  &  & &  & RCoVaR &  & 0.30 & 0.34 & 0.28 & 0.52 & 0.34 & 0.30 & 0.30 & 0.30 & 0.36 & 0.26 \\
				\midrule
				&  & \multirow{2}{*}{2} &  & \multirow{2}{*}{9.64} &  & VaR &  & 2.80 & 2.92 &  &  &  &  &  &  &  &  \\
				&  & &  & &  & RCoVaR &  & 2.12 & 2.22 &  &  &  &  &  &  &  &  \\[0.3em]
				\multirow{2}{*}{0.95} &  &  \multirow{2}{*}{5} &  & \multirow{2}{*}{8.12} &  & VaR &  & 1.00 & 1.02 & 1.20 & 0.88 & 1.18 &  &  &  &  &  \\
				&  &  &  &  &  & RCoVaR &  & 0.80 & 0.70 & 0.78 & 0.74 & 0.82 &  &  &  &  &  \\[0.3em]
				&  &  \multirow{2}{*}{10} &  & \multirow{2}{*}{6.70} &  & VaR &  & 0.66 & 0.56 & 0.44 & 0.44 & 0.40 & 0.50 & 0.48 & 0.52 & 0.54 & 0.50 \\
				&  & &  & &  & RCoVaR &  & 0.26 & 0.30 & 0.22 & 0.22 & 0.20 & 0.22 & 0.28 & 0.22 & 0.42 & 0.12 \\
				\bottomrule
			\end{tabular}
		}
			\caption{Joint and disaggregated \textit{first} rejection rates (in percent) of our RCoVaR surveillance procedure under the null hypothesis of correctly specified forecasts.
			Results are reported for a nominal level of $\iota = 10\%$, for $K \in \{2,5,10\}$ and $\alpha = \beta \in \{0.9, 0.95\}$.
			The columns labeled ``1''--``10'' report the percentage of cases in which the respective VaR or CoVaR detector---indicated in the column ``Measure''---raised the first alarm.}
			\label{tab:SizeRCoVaR}
		\end{table}



		The columns ``Joint'' in Tables~\ref{tab:SizeCoVaR} and \ref{tab:SizeRCoVaR} show the rejection rates of the CoVaR and RCoVaR monitoring procedures under the no-break null hypotheses---i.e., $t^\ast=n$ in \eqref{eqn:SimParamMisspec}.
		We consider $\alpha = \beta \in \{0.9, 0.95\}$ together with $K \in \{1,2,5,10\}$ for the CoVaR and  $K \in \{2,5,10\}$.
		For $K=1$, the rejection rates of the CoVaR and RCoVaR coincide due to the symmetry of the DGP.
		We find that all combined empirical rejection rates in the columns ``Joint'' are below $10.28\%$, which corresponds to our theoretical finding that the monitoring procedures hold size exactly.\footnote{Given that these empirical rejection rates are based on 5000 Monte Carlo replications, a $99\%$-confidence interval of IID Bernoulli-distributed variables with success probability 0.1 is given by approximately $[8.9\%, 11.1\%]$, indicating that our procedure is able to hold size \emph{exactly}, while the deviations from $10\%$ are explained by the finite number of Monte Carlo replications.}
		Even for $K=10$, the rejection rates are not too conservative and lie between $6\%$ and $9\%$, which illustrates that the application of Boole's inequality in the proof of Theorem~\ref{thm:CoVaR monitor} is relatively tight here.


		As the columns ``Joint'' merely indicate whether \emph{any} of the detectors raises an incorrect alarm under the null, we also analyze \emph{which} of the detectors is responsible for the rejection.
		For this, the right-hand side columns (denoted ``VaR'' and with the series numbers ``1''--``10'') of the tables show the frequencies how often a given detector is responsible for \emph{first} raising a false alarm.
		Notice here that the joint rejection frequency is somewhat lower than the sum of the individual frequencies as multiple detectors occasionally reject simultaneously.
		Overall, we find that for both, CoVaR and RCoVaR, the spurious alarms of the individual detectors are relatively balanced between the VaR and the systemic risk detectors as well as between the different assets $k=1,\dots,K$.
		Especially the former is noteworthy, and as desired by the choice of the level $\nu$ in steps 4 and 5 of the Algorithms~\ref{algo:1} and \ref{algo:RevCoVaR}.









\subsection{Simulation Results under the Alternative Hypothesis}
\label{sec:Simulation_Results_Alternative}

We now analyze our procedures' power to detect misspecified forecasts.
For this, Figure~\ref{fig:PowerCoVaR} plots rejection frequencies under the alternative ($t^\ast<n$ in \eqref{eqn:SimParamMisspec}).
The plots in the left panel display power against the break point $t^\ast \in \{0,1, \dots, 1000\}$, with $\beta_{\text{post}} = 0.85$ in \eqref{eqn:SimParamMisspec}.
The right panel shows power for a fixed break point at  $t^\ast = 0$ for a varying degree of post-break parameter misspecification $\beta_\text{post} \in \{0.7, 0.75, 0.8, 0.85, 0.87, 0.89, 0.899\}$.
We again consider $K\in\{1,2,5,10\}$ and $\alpha = \beta \in \{0.9, 0.95\}$.
Notice that monitoring only starts at $t=m=250$.
In the left panel, the plots recover the ``joint'' size from Tables~\ref{tab:SizeCoVaR}--\ref{tab:SizeRCoVaR} for $t^\ast = n = 1000$, and in the right panel, for $\beta_\text{post} = 0.7$.


\begin{figure}[tb]
	\centering
	\includegraphics[width=\textwidth]{input/power_CoVaR.pdf}
	\caption{Rejection rates of our CoVaR and RCoVaR surveillance methods plotted against the break point in \eqref{eqn:SimParamMisspec} in the left panel and against the post-break DCC parameter in the right panel.
	In both plots, we consider $\alpha = \beta \in \{0.9, 0.95\}$ as well as $K \in \{1,2,5,10\}$ for the CoVaR and  $K \in \{2,5,10\}$ for the RCoVaR.}
	\label{fig:PowerCoVaR}
\end{figure}

As expected, power increases monotonically for earlier break points as well as for higher degrees of parameter misspecification in all plots, and the procedure works equally well for both probability levels.
Whether a larger number $K$ of to-be-monitored stocks increases or decreases power depends on the specific situation.
In general, monitoring more sequences simultaneously is subject to a trade-off between a stricter correction of the significance level and more series in which misspecifications can be detected.


When comparing the monitoring procedures \emph{across} systemic risk measures, we find that the reverse CoVaR procedure is the most powerful, which can be explained by the fact that more ($2K$ instead of $K+1$) forecast series are monitored, and that more ($K$ instead of 1) VaR series are monitored simultaneously, which are not ``as far in the tail'' as the CoVaR.


\begin{figure}[tb]
	\centering
	\includegraphics[width=0.75\textwidth]{input/power_first_detectors_CoVaR.pdf}
	\caption{Rejection rates of the joint CoVaR and RCoVaR procedure in black, and frequencies how often the individual detectors are the \emph{first} to generate a detection in colors in the setting of the left panel of Figure~\ref{fig:PowerCoVaR}.
	Dashed lines depict ``first'' rejection rates for the VaR detector, and dot-dashed lines for the systemic risk detectors. The colors indicate the respective component of $\bm W_t$.
	The nominal level of $\iota = 10\%$ is indicated by the dashed horizontal line.}
	\label{fig:PowerCoVaRIndividual}
\end{figure}


Figure~\ref{fig:PowerCoVaRIndividual} shows the joint rejection rates in the same setting as in the left panel of Figure~\ref{fig:PowerCoVaR} in black, keeping $K=5$ and $\alpha = \beta = 0.9$ fixed.
Additionally, the colored dashed (VaR) and dot-dashed (Systemic Risk) lines indicate how often the respective detector was the first to raise an alarm, akin to the numbered columns in Tables~\ref{tab:SizeCoVaR}--\ref{tab:SizeRCoVaR}.
We find that the systemic risk detectors for different $k \in [K]$ have identical rates of rejecting first, which is sensible given the symmetry of the DCC--GARCH DGP.
Moreover, the VaR detector is more powerful than the systemic risk detectors, most likely as the VaR is not as far in the tail.


\begin{figure}[tb]
	\centering
	\includegraphics[width=\textwidth]{input/power_by_EstimationEffect_CoVaR.pdf}
	\caption{Rejection rates of our CoVaR and RCoVaR surveillance methods plotted against the break point for different estimation window lengths $E$, where $E = \infty$ implies the use of the correct parameters.
	We consider $K \in \{1,2,5,10\}$ and note that the rejection rates of CoVaR and RCoVaR coincide for $K=1$.
	We further keep $n=1000$, $t^\ast = 0$,  $\beta_\text{post} = 0.85$ and $\alpha = \beta = 0.9$ fixed.}
	\label{fig:PowerEstWindowCoVaR}
\end{figure}

We continue to analyze the effect parameter estimation noise within the forecasts has on the rejection frequencies.
Recall that, as argued in Remark~\ref{rem:ModelEstimation}, we view model estimation error as part of a misspecified forecast sequence.
Figure~\ref{fig:PowerEstWindowCoVaR} compares the rejection rates when the forecasts are based on (correctly specified) DCC--GARCH models that are estimated on an in-sample period of length $E \in \{1000, 2000, 5000, \infty\}$ with $K \in \{1,2,5,10\}$, while fixing $n=1000$, $m=250$, $t^\ast = 0$, $\beta_\text{post} = 0.85$, and $\alpha = \beta = 0.9$.
We find that small(er) estimation window sizes distort the systemic risk forecasts and, hence, deliver increased rejection rates of up to $20\%$ even in the ``no-break case'' with $t^\ast = 1000$.
This effect is somewhat more pronounced for the RCoVaR than for the CoVaR.
A further effect of the model estimation is that power increases naturally for  $t^\ast < 1000$ when the length of the estimation period $E$ is decreased.

Overall, our simulations reinforce the theoretical finding that the surveillance schemes hold size exactly, even when the tests are employed repeatedly at every time point.
Furthermore, the power of our tests behaves naturally for varying break times, different break magnitudes, and a varying dimensionality of the monitored sequences.
Finally, we find that VaR and systemic risk rejections occur relatively balanced under the null, whereas the VaR detector naturally has more power (i.e., it detects earlier), due to the systemic risk measures being further out in the tail.








\section{Empirical Application}
\label{Empirical Application}

		We apply the systemic risk surveillance procedures to real financial data to analyze their sensitivity to sub-optimal risk forecasts in practice.
		For the market indicator $X_t$, we use the negative returns of the S\&P\,500 Financials index (SPF) and we use the $K=5$ systemically important US banks Bank of America Corp (BAC), Citigroup Inc (C), Goldman Sachs Group Inc (GS), JPMorgan Chase \& Co (JPM) and Wells Fargo \& Co (WFC) as $(Y_{1t}, \dots, Y_{Kt})^\prime$.
		The four considered monitoring time periods each span $n=1000$ days and are chosen as
		(i) the global financial crisis: 23 September 2005 -- 14 September 2009, containing the bankruptcy of Lehman Brothers on 15 September 2008;
		(ii) a calm period: 14 January 2013 -- 30 December 2016;
		(iii) the COVID period: 23 March 2017 -- 12 March 2021, containing the  outbreak of the COVID pandemic; and
		(iv) Trump's tariffs: 29 October 2021 -- 23 October 2025, containing the introduction of Trump's tariffs on 2 April 2025 (and the collapse of the Silicon Valley Bank on 10 March 2023).
		In all four settings, we use a rolling monitoring window of $m=250$ trading days, consider the systemic risk measures at levels $\alpha = \beta = 0.95$ and use the monitoring level of $\iota = 0.1$.

		\begin{figure}[tb]
			\centering
			\includegraphics[width=\textwidth]{input/Detectors_GARCH_CoVaR.pdf}
			\caption{
				Normalized VaR and CoVaR detector values for the SPF, and the financial institutions BAC, C, GS, JPM and WFC for forecasts from a Gaussian CCC--GARCH and a Student's $t$ DCC--GARCH model.
				We use the specifications $\iota=0.1$, $\alpha=\beta=0.95$,  $E = 1500$, $n=1000$ and $m=250$.
				The displayed detector values are normalized by their respective critical values, such that a  detector exceeding the black horizontal unit line implies a detection.
				The vertical dotted lines represent the bankruptcy of Lehman Brothers on 15 September 2008,
				the beginning of the COVID crisis on 13 March 2020 (when the US declared a national emergency),
				the collapse of the Silicon Valley Bank on 10 March 2023, and
				the tariff announcement of Donald Trump on 2 April 2025.
			}
			\label{fig:Detectors_CoVaR}
		\end{figure}


		We generate systemic risk forecasts by modeling the $K+1=6$ returns by a CCC--GARCH model with Gaussian innovations and a DCC--GARCH model with Student's $t$ innovations, introduced by \citet{Eng02} and further described in Section~\ref{Data-Generating Process}.
		We deliberately choose a Gaussian distribution for the CCC--GARCH model's innovations to analyze a suboptimal model.
		The latter DCC--GARCH--$t$ model performs relatively well in horse races of multivariate volatility models \citep{LRV12,LRV13} and is, hence, still a standard forecasting model in financial risk management.
		From these two models, we generate CoVaR and RCoVaR forecasts as described in Section~\ref{sec:SystemicRiskForecasts}  based on the presumed distributions of the model innovations $\bm \varepsilon_t\mid\mathcal{F}_{t-1}\sim t_{\nu}(\bm R_t)$; using $\nu = \infty$ for the Gaussian case.
		The models are estimated in all four settings using $E = 1500$ observations, i.e., data starting approximately six years before the beginning of the monitoring period.

		Figure~\ref{fig:Detectors_CoVaR} presents the VaR and CoVaR detector values, normalized by their respective critical values, for both multivariate GARCH models across the four considered time periods.
		Under this normalization, values exceeding unity indicate rejections of the null hypothesis.
		The detector values may take negative values as a consequence of their standardization, as defined, for example, in \eqref{eqn:VaRDetStandardization}.

		While we do not find rejections during the calm time for either forecasting model, the detectors raise an alarm for the CCC--GARCH--$\mathcal{N}$ forecasts in all three other considered periods.
		Importantly, these alarms are raised by the systemic CoVaR detector, hence implying that specifically the systemic component of the risk forecasts is misspecified.
		This highlights the importance of monitoring \textit{systemic} risk as opposed to monitoring only the risk component (i.e., the VaR) via the procedures of \citet{HD22a+} and \citet{WWZ23}.
		In contrast, for the DCC--GARCH--$t$ forecasts, only the VaR component raises an alarm after the start of the COVID pandemic, showing that the dynamic correlation structure implies by the DCC is much better able to capture the changing co-movements in turbulent times.


	\begin{figure}[tb]
			\centering
			\includegraphics[width=\textwidth]{input/Detectors_GARCH_RevCoVaR.pdf}
			\caption{
				Normalized VaR and reverse CoVaR detector values. For details, see the caption of Figure~\ref{fig:Detectors_CoVaR}.
			}
			\label{fig:Detectors_RevCoVaR}
	\end{figure}


		Figure~\ref{fig:Detectors_RevCoVaR} provides corresponding results for the RCoVaR forecasts with overall similar findings.
		While there are no rejections during the calm time in either model, especially the RCoVaR detectors raise alarms for the CCC--GARCH--$\mathcal{N}$ forecasts around the times where (systemic) market risks have increased markedly.
		In contrast, for the Student's $t$ DCC--GARCH forecasts, mostly the VaR detectors yield rejections, implying once more that this model with its dynamic correlation structure is much better suited to model systemic risks.

		As for the attribution of forecast failure to a specific institution, consider the financial crisis, which had its origins in the housing market \citep{Mis11}.
		During this time, Wells Fargo's increased systemic riskiness was responsible for the forecast failure of the CCC--GARCH--$\mathcal{N}$ model.
		This holds true when Wells Fargo is viewed both as a systemic risk \textit{receiver} and \textit{transmitter} (see Figures~\ref{fig:Detectors_CoVaR} and \ref{fig:Detectors_RevCoVaR}, respectively).
		The preeminent role of Wells Fargo early on may be explained by its heavy reliance on mortgage lending (later on increased through the acquisition of Wachovia in early 2008), which contrasts with the other considered retail and investment banks.


		Our empirical results are consistent with \citet{LRV12,LRV13}.
		In their horse race of multivariate volatility models, they find that DCC and CCC specifications perform equally well during calm times, yet the former are more adequate for crises periods.
		We find their conclusion for \textit{volatility} forecasting to also hold true for \textit{systemic risk} forecasting.
		However, in contrast to the (repeated) ``one-shot analysis'' of \citet{LRV12,LRV13}, our monitoring procedures also allow us to date the time point when a forecast breakdown of a specific model occurs and, equally important, also to pinpoint the model failure to a specific institution.


\section{Conclusion}
\label{Conclusion}

Regulators as well as financial institutions are particularly concerned about the commonalities in risk factors, i.e., systemic risk. To effectively take preventive measures, it becomes vital to detect changes in systemic risk assessments as soon as possible. To that end, this paper proposes formal monitoring tools for systemic risk. These are shown to work well in simulations and are useful in practice, as the empirical application to the US~banking sector in Section~\ref{Empirical Application} demonstrates.

The advantages of our proposed procedures are fourfold.
First, unlike classical ``one-shot'' backtests, our monitoring schemes control size under \emph{repeated application} over a fixed horizon, as required for, e.g., daily risk monitoring in financial markets.
Second, size control holds \emph{in finite samples} by construction, in contrast to asymptotic one-shot backtests such as \citet{FH21}.
Third, our procedures accommodate \emph{multiple time series} simultaneously, unlike the one-shot backtests of \citet{Bea21}, which are restricted to bivariate settings (i.e., $K=1$ in our notation).
Fourth, a Bonferroni-type correction enables us to \emph{attribute} detected deficiencies in systemic risk forecasts to specific institutions, a key feature in regulatory applications where identification matters as much as detection.

We stress that while our empirical application deals with systemic risk in the financial system, our monitoring procedures may be used more widely in other contexts.
For instance, it could be used by individual banks to monitor systemic risk forecasts in their financial positions.
Again, the key concern of the institutions is not necessarily the risk inherent in their positions (as risk is associated with reward in financial markets), but the commonality in their exposures.
This is because it is precisely during times of extreme co-movements that diversification benefits vanish, which---as the saying goes---is the only free lunch around. For this reason, it may also be useful for banks to monitor the systemic risk in their positions.




\singlespacing
\small
\setlength{\bibsep}{4pt}
\bibliographystyle{jaestyle2}
\bibliography{bib_CoVaRMonitoring}