Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,893 characters · 13 sections · 107 citation commands
Dynamic CoVaR Modeling and Estimation
\baselineskip18pt \setcounter{totalnumber}{50} \setcounter{topnumber}{50} \setcounter{bottomnumber}{50} \abovedisplayskip1.5ex plus1ex minus1ex \belowdisplayskip1.5ex plus1ex minus1ex \abovedisplayshortskip1.5ex plus1ex minus1ex \belowdisplayshortskip1.5ex plus1ex minus1ex
\onehalfspacing
Since the introduction of the Value-at-Risk (VaR), risk forecasts have become a key input in financial decision making. For instance, VaR and Expected Shortfall (ES) forecasts are now routinely used for setting capital requirements of financial institutions under the Basel framework LS21. Consequently, a huge literature on forecasting VaR and ES has emerged MF00,EM04,NR16,Mas17,PZC19, DimiHalbleib2021. By definition, these measures are primarily designed to assess the risks faced by banks in isolation. Thus, these measures are well-suited to address microprudential objectives in banking regulation, that is, to limit the risk taking of individual institutions.
However, in the aftermath of the financial crisis of 2007--09, macroprudential objectives have gained importance on the regulatory agenda AER12,Aea17. While also attempting to curb the risk taking of individual financial institutions, the macroprudential approach additionally takes into account the commonality of risk exposures among banks. To do so, a measure of interconnectedness of banks is now used under the Basel framework of the BCBSBF19 to determine the global systemically important banks (G-SIBs), which are subjected to higher capital requirements. This new focus has spurred the development of systemic risk measures to accurately measure the interlinkages for which VaR and ES are unsuitable. By now, a plethora of systemic risk measures is available GK11,CIM13,AB16,Aea17.
One of the most popular systemic risk measures is the conditional VaR (CoVaR) of AB16. It is defined as a quantile of a financial loss (e.g., of an entire market), given that a reference asset (e.g., a systemically important bank) is in distress, where the latter is taken to mean an exceedance of the reference asset's VaR. Note that this definition corresponds to the slightly redefined version of AB16's AB16 CoVaR, which is due to GT13 and relies on the broader “stress event” of a VaR exceedance AB16. Section (ref) outlines our reasons for doing so.
While forecasting models and asymptotic properties of forecasts are well explored for VaR and ES Cea07,GS08,WZ16,Hog18+, the models for CoVaR forecasting are hitherto rather ad-hoc and little is known about the consistency of the forecasts. For instance, DCC--GARCH models of Eng02---which are popular due to their ability to accurately forecast conditional variance-covariance matrices LRV12,CM14---may be used to generate CoVaR forecasts. However, as FZ16b point out, “[n]o formally established asymptotic results exist for the full estimation of the DCC [...] models”; see also DFL18. This renders their application in forecasting systemic risk measures questionable. Moreover, there are additional problems associated with their use in systemic risk forecasting that stem from the non-uniqueness of the decomposition of the variance-covariance matrix (see Section (ref) and Appendix (ref)). Recently, francq2025inference and Hog25 derive asymptotic properties for forecasts of systemic risk, as measured by the CoVaR and the marginal expected shortfall (MES). However, the multivariate GARCH-type models they consider for prediction do not allow for a dynamic conditional correlation structure. In sum, there is a need for dynamic multivariate models that can deliver accurate systemic risk forecasts with strong theoretical underpinnings.
This paper fills this gap. Specifically, the first main contribution of this paper is to introduce such dynamic models for the CoVaR, which we call CoCAViaR models, as they nest the classical CAViaR models of EM04. In the spirit of EM04 and PZC19, we model the pair (VaR, CoVaR) depending on past financial losses, lagged (VaR, CoVaR) model values, and (possibly) external covariates. Our models are semiparametric in the sense that the quantities of interest (VaR and CoVaR) are modeled parametrically, yet no further (parametric or otherwise) assumptions are placed on the conditional distribution.
There are new challenges in modeling systemic risk vis-\`{a}-vis modeling univariate quantities such as VaR EM04,WhiteKimManganelli2015,Catania2022 and ES PZC19. Modeling VaR or (VaR, ES) requires to specify the univariate dynamics only. Yet, since CoVaR measures the interlinkage of bivariate random variables, it becomes necessary to model the co-movements as well. We explore different models for doing so, and identify several competitive performers in our empirical application.
Our second main contribution is to propose an estimator for the model parameters and derive its large sample properties. The main technical hurdle to overcome is that---unlike VaR and (VaR, ES)---the pair (VaR, CoVaR) fails to be elicitable, such that no real-valued scoring function exists that is uniquely minimized by the true report FH24. This renders standard M-estimation---adopted for VaR and ES models by EM04, DimiBayer2019 and PZC19---infeasible DFZ_CharMest. Instead, we exploit the multi-objective elicitability of (VaR, CoVaR) FH24. This property suggests a two-step M-estimator. In the first step, the score in the VaR component is minimized, and then the CoVaR score is minimized in the second step. We show that this leads to a consistent estimator. In Appendices (ref) and (ref), we also establish asymptotic normality of our two-step M-estimator and propose valid inference based on consistent estimation of the asymptotic variance-covariance matrix. Our proofs show how to deal with non-smooth and discontinuous objective functions in the context of two-step M-estimation for dynamic models (based on past model values). We speculate that our proof strategy may also be used elsewhere and, therefore, may be of independent interest.
In Appendix (ref), we also show consistency of our estimator for parameters from (a fairly general class of) models that are based on misspecified initial values (which arise, e.g., when the true model is initialized in the infinite past). While this issue is well-explored for estimators of GARCH-type models FZ10, it is not treated for current semiparametric VaR and ES models that face the additional difficulty of non-smooth objective functions EM04,WhiteKimManganelli2015,PZC19,Catania2022.
Simulations confirm the good finite-sample properties of our two-step estimator. While the parameters of the CoCAViaR models can be estimated accurately in realistic settings, reliable inference for the model parameters requires large sample sizes or moderate probability levels. However, as the main use of our CoCAViaR models is in forecasting, inference for the model parameters is of lesser importance than consistent estimates.
In the empirical application, we compare CoVaR forecasts issued from our CoCAViaR models with those from benchmark DCC--GARCH models. We do so for the four most systemically risky US banks according to the FSB22, whose impact on a broader market index is assessed by CoVaR. Our various CoCAViaR model specifications tend to outperform the DCC--GARCH benchmarks, particularly in the CoVaR forecasts but also in the VaR component. Often, these differences in predictive ability are also statistically significant, as judged by the DM95-type comparative backtest of FH24. The superiority of our proposals may be explained by two reasons. First, our CoCAViaR specifications are specifically designed to model the quantities of interest (VaR and CoVaR), whereas multivariate GARCH processes focus on modeling the complete predictive distribution. Second, our estimation technique is tailored to provide an accurate description of the (VaR, CoVaR) evolution. In particular, the estimator is not too strongly influenced by center-of-the-distribution observations as, e.g., standard estimators of multivariate GARCH models.
The rest of the paper is structured as follows. Section (ref) formally introduces the CoVaR, our modeling framework and the appertaining parameter estimator. Section (ref) gives large sample results for our estimator. We illustrate the finite-sample properties of our estimator in Section (ref). Section (ref) presents the empirical application and the final Section (ref) concludes. The main proofs as well as additional details are given in the Online Appendix (containing Sections (ref)--(ref)), where we also verify our main assumptions for CCC--GARCH models and vector autoregressive (VAR) models.
Throughout the paper, we consider a sample of size $n\in\mathbb{N}$ of the bivariate series $\big\{(X_t,Y_t)^\prime\big\}_{t\in\mathbb{N}}$, which is defined on the probability space $\big(\Omega, \mathcal{F}, \mathbb{P}\big)$. Specifically, $Y_t$ stands for the log-losses of interest (e.g., system-wide losses in the financial system) and $X_t$ are the log-losses of some reference position (e.g., the losses of a bank's shares). Here, log-losses are simply the negated log-returns. The information set that the forecaster is interested in conditioning on at time $(t-1)$ is $\mathcal{F}_{t-1}=\sigma\big((X_{t-1},Y_{t-1})^\prime,\bm Z_{t-1},\ldots,(X_{1},Y_{1})^\prime,\bm Z_{1},\boldsymbol{\mathcal{I}}_0 \big)$. Here, the variable $\bm Z_t$ contains some (possibly multivariate) exogenous covariates, and $\boldsymbol{\mathcal{I}}_0$ represents known and (possibly multivariate) pre-sample information that is used for model initialization at time $t=1$ (see below for details).
For $\beta\in[0,1)$, we define $\operatorname{VaR}_{\beta}(F)=F^{\leftarrow}(\beta)$ to be the $\beta$-quantile of the distribution $F$. Then, the conditional VaR is simply $\operatorname{VaR}_{t,\beta}=\operatorname{VaR}_{\beta}(F_{X_t\mid\mathcal{F}_{t-1}})$, where $F_{X_t\mid\mathcal{F}_{t-1}}$ denotes the conditional distribution of $X_t$ given $\mathcal{F}_{t-1}$. The stress event considered in the definition of the CoVaR is that the loss of the reference position exceeds its VaR, i.e., $\{X_t\geq\operatorname{VaR}_{t,\beta}\}$. With our orientation of $X_t$ denoting financial losses, we commonly consider values for $\beta$ close to one, such as $\beta=0.95$. For $\alpha\in(0,1)$ we define $\operatorname{CoVaR}_{t,\alpha|\beta}=\operatorname{CoVaR}_{\alpha|\beta}(F_{X_t,Y_t\mid\mathcal{F}_{t-1}})$, where $F_{X_t,Y_t\mid\mathcal{F}_{t-1}}$ denotes the distribution of $(X_t,Y_t)^\prime\mid\mathcal{F}_{t-1}$ and \[ \operatorname{CoVaR}_{\alpha|\beta}(F_{X,Y})=\operatorname{VaR}_\alpha(F_{Y\mid X \geq \operatorname{VaR}_\beta(F_{X})}) \] for a joint distribution function $F_{X,Y}$ with marginals $F_{X}$ and $F_{Y}$. Again, we usually consider values of $\alpha$ close to one. If $\alpha=\beta$ we simply write $\operatorname{CoVaR}_{t,\alpha} = \operatorname{CoVaR}_{t,\alpha|\alpha}$, and for $\beta=0$ we simply have $\operatorname{CoVaR}_{t,\alpha|\beta} = \operatorname{VaR}_{\alpha}(F_{Y_t\mid\mathcal{F}_{t-1}})$.
As pointed out in the Motivation, our CoVaR definition follows GT13 and deviates from the original one of AB16. The latter authors use $\{X=\operatorname{VaR}_{\beta}(F_X)\}$ as the stress event instead of $\{X\geq\operatorname{VaR}_{\beta}(F_X)\}$, such that $\operatorname{CoVaR}_{\alpha|\beta}^{=}(F_{X,Y})=\operatorname{VaR}_{\alpha}(F_{Y\mid X=\operatorname{VaR}_{\beta}(F_X)})$. Our reasons for adopting the “inequality version” of the CoVaR are as follows. First, on an intuitive level, conditioning on $\{X\geq\operatorname{VaR}_{\beta}(F_X)\}$ instead of $\{X=\operatorname{VaR}_{\beta}(F_X)\}$ better captures “tail risks”. Second, $\operatorname{CoVaR}_{\alpha|\beta}(F_{X,Y})$ is dependence consistent in the sense of MS14, while $\operatorname{CoVaR}_{\alpha|\beta}^{=}(F_{X,Y})$ is not BDJ17. Third, a further advantage of the redefined CoVaR is its multi-objective elicitability shown by FH24. This property allows CoVaR forecasts to be meaningfully evaluated via so-called backtests, which are, however, not available for AB16's AB16 $\operatorname{CoVaR}_{\alpha|\beta}^{=}(F_{X,Y})$. Fourth, (dynamic models for) $\operatorname{CoVaR}_{\alpha|\beta}^{=}(F_{X,Y})$ would be harder to estimate than $\operatorname{CoVaR}_{\alpha|\beta}(F_{X,Y})$ without imposing further parametric assumptions. This is because the $\operatorname{CoVaR}_{\alpha|\beta}^{=}(F_{X,Y})$ is the $\alpha$-quantile of the distribution of $Y\mid X=\operatorname{VaR}_{\beta}(F_X)$, which requires non-parametric methods with their usually slower convergence rates. For all these reasons, we work with the definition of GT13, which---next to the authors mentioned above---is also used by AFS13, BC19, NZ20 and CR22.
An attractive class of models for financial data---and, hence, also for CoVaR modeling---are multivariate GARCH processes, such as the DCC--GARCH of Eng02 or its corrected version by Aie13. These model the conditional covariance matrix $\bm H_t := \operatorname{Var}\big((X_t,Y_t)^\prime\mid\mathcal{F}_{t-1}\big)$ dynamically and are regularly used for volatility forecasting. However, in order to generate CoVaR forecasts, one requires the “square-root matrix” $\bm \varSigma_t$ of $\bm H_t$, satisfying $\bm \varSigma_t\bm \varSigma_t^\prime = \bm H_t$. For this matrix decomposition there exist infinitely many possibilities, such as the one based on the symmetric eigenvalue decomposition ($\bm \varSigma_t^{s}$) or the lower triangular matrix of the Cholesky decomposition ($\bm \varSigma_t^{l}$). The problem is that for non-spherically distributed shocks, each possibility, while implying the same variance-covariance dynamics $\bm H_t$, may imply different values for the CoVaR (forecasts). Thus, in principle, a given single GARCH-type model for $\bm H_t$ may be consistent with an infinite number of CoVaR forecasts (depending on the choice of $\bm \varSigma_t$), thus creating an unsatisfactory ambiguity when applied to CoVaR forecasting. We provide additional details on this ambiguity in Appendix (ref), where we also show that this identification problem for the CoVaR can only arise for non-spherically distributed shocks.
This deficiency of multivariate GARCH models underlines the necessity to construct explicit CoVaR models as we do in the following. In the spirit of EM04 and PZC19, we consider semiparametric models for the pair (VaR, CoVaR) of the general form
where $\bm \theta^v$ and $\bm \theta^c$ are generic parameters from some parameter spaces $\bm \varTheta^v\subset\mathbb{R}^{p}$ and $\bm \varTheta^c\subset\mathbb{R}^q$, respectively. For $t=1$, equation (ref) is taken to mean that the model only depends on the initialization information $\boldsymbol{\mathcal{I}}_0$ (introduced in Section (ref)) and (possibly) the respective parameter $\bm \theta^v$ or $\bm \theta^c$. Throughout the paper, we assume that the underlying data-generating process is such that the model in (ref) is correctly specified for (VaR, CoVaR). That is, there exist true parameters $\bm \theta^v_0\in\bm \varTheta^v$ and $\bm \theta^c_0\in\bm \varTheta^c$, such that
almost surely (a.s.).
Before offering some comments on our general framework, we give a specific instance of a model satisfying (ref) in Example (ref) now. A discussion of the conditions under which (ref) holds for the models of Example (ref) follows in Remark (ref).
A few comments on our general modeling approach given in (ref) and (ref) are in order. First, while we assume the model in (ref) to be correctly specified, we make no assumption on its stationarity. Indeed, because the information set $\mathcal{F}_{t-1}$ increases with $t$, $v_t(\bm \theta^v)$ and $c_t(\bm \theta^c)$ will in general be non-stationary, even for the true parameters (i.e., for $\bm \theta^v=\bm \theta_0^v$ and $\bm \theta^c=\bm \theta_0^c$).
Second, our model in (ref) is semiparametric in the sense that---while the model dynamics are governed by the parameters $\bm \theta^v$ and $\bm \theta^c$---we impose no additional assumptions on the conditional distribution of $(X_t,Y_t)^\prime\mid\mathcal{F}_{t-1}$.
Third, as the $v_t(\cdot)$ and $c_t(\cdot)$ functions may vary with time $t$, our modeling framework is sufficiently flexible to allow for the inclusion of lagged (VaR, CoVaR) as in Example (ref). This is important because including lags of model values often leads to a better predictive performance, as is the case for GARCH models, which improve upon ARCH models by including lags of volatility in the variance equation.
Fourth, the VaR model in (ref) generalizes classical (time series) quantile regressions based on lagged values of $X_t$, $Y_t$ and the covariate vector $\bm Z_{t}$ EM04,OH16. Also, the “VAR for VaR” models (of dimension two) of WhiteKimManganelli2015 are nested by choosing $\beta=0$ such that the CoVaR simply becomes the VaR of $Y_t$.
Fifth, the fact that the VaR model $v_t(\cdot)$ does not depend on the CoVaR parameter $\bm \theta^c$ renders our two-step estimator (introduced in Section (ref) below) feasible. Likewise, we assume that the CoVaR model $c_t(\cdot)$ does not include the VaR parameter $\bm \theta^v$. Our two-step estimator would remain feasible if we dropped this requirement. This would, however, come at the cost of a further explosion of technicality in the proofs, without any offsetting improvements (in terms of predictive accuracy) of the resulting models; see the empirical application in Section (ref) for evidence on this.
Sixth, besides using the full history of lagged values of $X_t$, $Y_t$ and $\bm Z_t$, our models must be initialized at time $t=1$. For this, we use the initialization information $\boldsymbol{\mathcal{I}}_0$ that may be thought of as capturing the conditions at which the observable history started. E.g., when $X_t$ or $Y_t$ denote returns on JPMorgan Chase shares from 2001 onwards, then $\boldsymbol{\mathcal{I}}_0$ may capture general market conditions on the first day the merged bank JPMorgan Chase started trading in 2001. We give specific possibilities for the initialization of the SAV CoCAViaR models in the following example.
\setcounter{example}{0}
The SAV CoCAViaR models of Example (ref) are, among others, correctly specified (in the sense of (ref)) under a type of Jea98's Jea98 extended constant conditional correlation (ECCC) GARCH process, which has been studied extensively in the literature HT04,CK10:
where $\widetilde{\bm \omega}\in\mathbb{R}^2$, $\widetilde{\bm A}, \widetilde{\bm B}\in\mathbb{R}^{2\times2}$. In (ref), the $\mathcal{F}_{t-1}$-measurable conditional volatilities $\sigma_{X,t}$ and $\sigma_{Y,t}$ are independent of the independent and identically distributed (i.i.d.) innovations $\big(\varepsilon_{X,t}, \varepsilon_{Y,t} \big)' \overset{\text{i.i.d.}}{\sim} F \big( \boldsymbol{0}, \bm \varSigma \big)$, where $F$ denotes a generic absolutely continuous, bivariate distribution with zero mean and covariance matrix $\bm \varSigma=
$, $\rho\in(-1,1)$. Equation \eqref{eqn:ECCCmodel} deviates from \citet{Jea98} by using \textit{absolute} (instead of \textit{squared}) returns as drivers of volatility, since these have more predictive content for volatility \citep{FG07}. Model~\eqref{eqn:ECCCmodel} implies the (VaR, CoVaR) dynamics in \eqref{eqn:SAVCoCAViaRModelClass}, where the true parameters $\bm \omega_0, \bm A_0, \bm B_0$ of the model in \eqref{eqn:SAVCoCAViaRModelClass} arise as transformations of $\widetilde{\bm \omega}, \widetilde{\bm A}, \widetilde{\bm B}$.\footnote{ Multiplying the rows of the volatility equation in \eqref{eqn:ECCCmodel} by the true VaR and CoVaR of the innovations respectively gives that $\bm \theta_0^v = \big( \omega_{1,0}, A_{11,0} , A_{12,0}, B_{11,0}, B_{12,0} \big)^\prime = \big( v_\varepsilon \widetilde{\omega}_1, v_\varepsilon \widetilde{A}_{11}, v_\varepsilon \widetilde{A}_{12}, \widetilde{B}_{11}, v_\varepsilon/ c_\varepsilon \widetilde{B}_{12} \big)^\prime$ and $\bm \theta_0^c = \big( \omega_{2,0}, A_{21,0} , A_{22,0}, B_{21,0}, B_{22,0} \big)^\prime = \big( c_\varepsilon \widetilde{\omega}_2, c_\varepsilon \widetilde{A}_{12}, c_\varepsilon \widetilde{A}_{22}, c_\varepsilon / v_\varepsilon \widetilde{B}_{21}, \widetilde{B}_{22} \big)^\prime$, where $v_\varepsilon$ is the $\beta$-VaR of $\varepsilon_{X,t}$ and $c_\varepsilon$ the $\alpha|\beta$-CoVaR of the pair $(\varepsilon_{X,t}, \varepsilon_{Y,t})^\prime$, whose analytical form is given in MS14. }
While ECCC--GARCH models are primarily designed to model the complete predictive distribution of $(X_t,Y_t)^\prime\mid\mathcal{F}_{t-1}$, they can also be used to predict VaR and CoVaR (via (ref) and the parameter transformation given in footnote (ref)). When (ref) describes the true data-generating process for $(X_t, Y_t)^\prime$, then there will be little difference between the (VaR, CoVaR) forecasts by our model in Example (ref) and those issued by the ECCC--GARCH model (assuming consistent parameter estimates in both cases). However, when (ref) does not describe the underlying dynamics of $(X_t, Y_t)^\prime$, the two approaches to (VaR, CoVaR) modeling may yield widely different predictions for reasons discussed in more detail in Section (ref).
In our simulations in Section (ref), we use (a restricted version of) model (ref) to generate data to assess how well our two-step M-estimator performs. In our empirical application, we consider various SAV CoCAViaR specifications based on zero restrictions in $\bm A$ and $\bm B$, as well as generalizations to so-called “asymmetric slope” models based on the positive and negative components of $X_t$ and $Y_t$ instead of on their absolute values.
We now introduce estimators of the unknown model parameters $\bm \theta_{0}^{v}$ and $\bm \theta_{0}^{c}$. As we consider M-estimation in this paper, we require a to-be-minimized objective (or also: scoring) function. However, as pointed out in the Motivation, there is no real-valued scoring function associated with the pair (VaR, CoVaR). DFZ_CharMest show that the existence of such a (consistent) scoring function is a necessary condition for consistent M-estimation of semiparametric models.
To overcome this drawback, FH24 propose a $\mathbb{R}^2$-valued scoring function in the closely related context of forecast evaluation. To be able to compare forecasts, $\mathbb{R}^2$ has to be equipped with an order, and FH24 show that the lexicographic order is suitable for that purpose. Specifically, they prove that (under some regularity conditions) the expectation of the $\mathbb{R}^2$-valued scoring function
is minimized by the true VaR and CoVaR with respect to the lexicographic order. That is, for all $v,c\in\mathbb{R}$, \[ \mathbb{E}\Bigg[\bm S\Bigg(
,
\Bigg)\Bigg]\preceq_{\operatorname{lex}}\mathbb{E}\Bigg[\bm S\Bigg(
,
\Bigg)\Bigg], \] where $(x_1,x_2)^\prime\preceq_{\operatorname{lex}} (y_1,y_2)^\prime$ if $x_1<y_1$ or ($x_1=y_1$ and $x_2\leq y_2$). Note that $S^{\operatorname{VaR}}(\cdot,\cdot)$ in (ref) is the standard tick-loss function known from quantile regression. Clearly, $S^{\operatorname{CoVaR}}(\cdot,\cdot)$ is similar in structure to $S^{\operatorname{VaR}}(\cdot,\cdot)$, with the only difference being the indicator $\mathds{1}_{\{x>v\}}$ that restricts the evaluation to observations with VaR exceedances in the first component. In the related literature, these scoring functions are often also called “loss functions” Gne11. To avoid confusion with financial “losses”, we adhere to the term “scoring function”.
The definition of the lexicographic order suggests the following two-step M-estimator of $\bm \theta_{0}^{v}$ and $\bm \theta_{0}^{c}$. In the first step, we estimate $\bm \theta_0^v$ via
such that the parameters are chosen to minimize the average empirical score in the VaR component over the VaR parameter space $\bm \varTheta^{v}$. For the VaR model $v_t(\cdot)$, this is the quantile regression estimator of EM04.
With the estimate $\widehat{\bm \theta}_n^v$ at hand, the lexicographic order then suggests to minimize the average empirical score in the second component to estimate $\bm \theta_0^c$ via
For this two-step estimator to be feasible, the requirement that the VaR evolution does not depend on $\bm \theta^c$ is essential. Appendix (ref) shows that the presence of $\widehat{\bm \theta}_n^v$ in the second-stage minimization impacts the asymptotic variance of $\widehat{\bm \theta}_{n}^{c}$. Of course, this is usually the case for two-step estimators NeweyMcFadden1994. A two-step estimator similar in spirit to (ref)--(ref) is employed in DimiHogaMES2024 for their static regressions under adverse conditions.
Note that the estimators $\widehat{\bm \theta}_{n}^{v}$ and $\widehat{\bm \theta}_{n}^{c}$ depend on the assumed initialization information in $\boldsymbol{\mathcal{I}}_0$ through the model functions $v_t(\bm \theta^{v})$ and $c_t(\bm \theta^{c})$. For instance, the parameter estimators $\widehat{\bm \theta}_n^v$ and $\widehat{\bm \theta}_n^c$ for the model in Example (ref) rely on initial values for $v_1(\bm \theta^v)$ and $c_1(\bm \theta^c)$, which depend on whether the model was initialized according to (ref) or (ref). We mention that in case (ref)---when $\boldsymbol{\mathcal{I}}_0$ contains the infinite past---the estimators cannot be computed as $v_1(\bm \theta^v)$ and $c_1(\bm \theta^c)$ depend on infinitely many lags. In this case, one has to rely on a truncated information set for estimation.
Next, Section (ref) shows consistency of the parameter estimators in the case where the full $\boldsymbol{\mathcal{I}}_0$ can be used for estimation (such as (ref)--(ref)). The case where $\boldsymbol{\mathcal{I}}_0$ needs to be truncated (such as (ref)) is treated in Appendix (ref).
Here, we show consistency of our two-step M-estimators $\widehat{\bm \theta}_{n}^{v}$ and $\widehat{\bm \theta}_{n}^{c}$. To do so, we introduce several regularity conditions in Assumption (ref). Throughout the paper, we let $K\in(0,\infty)$ denote some large universal constant not depending on $n$, $t$ or any other values. Further, $\Vert\bm x\Vert$ denotes the Euclidean norm when $\bm x$ is a vector, and the Frobenius norm when $\bm x$ is matrix-valued. The joint cumulative distribution function (c.d.f.) of $(X_{t}, Y_t)^\prime\mid\mathcal{F}_{t-1}$ is denoted by $F_t(\cdot,\cdot)$, and its Lebesgue density (which we assume exists) by $f_t(\cdot,\cdot)$. Similarly, $F_t^{W}(\cdot)$ ($f_t^{W}(\cdot)$) denotes the distribution (density) function of $W_t\mid\mathcal{F}_{t-1}$ for $W\in\{X,Y\}$. For a sufficiently smooth function $\mathbb{R}^{p}\ni\bm \theta\mapsto f(\bm \theta) \in \mathbb{R}$, we denote the $(p\times1)$-gradient by $\nabla f(\bm \theta)$, its transpose by $\nabla' f(\bm \theta)$ and the $(p\times p)$-Hessian by $\nabla^2 f(\bm \theta)$.
Since some of the conditions in Assumption (ref) are high-level, we verify these for CCC--GARCH and VAR models in Appendices (ref) and (ref), respectively. In Assumption (ref), item (ref) ensures a correctly specified VaR and CoVaR model, including a correct model initialization at $t=1$ through $\mathcal{I}_0$, which can, e.g., be achieved by one of the possibilities (ref)--(ref) discussed in Example (ref) and Remark (ref). We treat arbitrarily initialized models in Appendix (ref). Item (ref) is a standard stationarity condition for the observables. Yet notice that the (VaR, CoVaR) process $\big(v_t(\cdot), c_t(\cdot)\big)^\prime$ does not have to be stationary and---in general---is non-stationary due to the expanding information set $\mathcal{F}_{t-1}$ for processes initialized in the finite past. Item (ref) ensures---among other things---strict (multi-objective) consistency of the scoring function given in (ref); see FH24 for details. Uniform Lipschitz conditions on derivatives of the joint conditional c.d.f. are imposed in item (ref). The next item (ref) provides certain boundedness conditions on the conditional probability density function of $X_t$. Compactness of the parameter space in (ref) is a standard requirement in extremum estimation; see NeweyMcFadden1994. Assumption (ref) (ref) is also a standard condition, imposed in similar form by EM04, PZC19 and Catania2022. In Appendices (ref) and (ref), we show how this condition can be verified based on Theorem 21.9 in Dav94 for an underlying CCC--GARCH and a linear VAR process, respectively. Item (ref) for the VaR model is identical to condition (ID) in Wei91 and, loosely speaking, ensures that $v_t(\bm \theta^v)$ differs sufficiently from $v_t(\bm \theta_0^v)$ with positive (conditional) probability for any $\bm \theta^v \not= \bm \theta_0^v$. This condition is required to establish the unique identifiability of the parameter $\bm \theta_0^v$ in the sense of Whi80. Condition (b) of item (ref) is the analogous condition for the CoVaR model, required for the unique identifiability of $\bm \theta_0^c$. The final items (ref)--(ref) are smoothness and moment conditions. We mention that the moment bounds in item (ref) on the observables, $\mathbb{E}|X_t|\leq K$ and $\mathbb{E}|Y_t|^{1+\iota}\leq K$, are sufficiently weak to be practically always satisfied in financial and economic applications.
Our first main theoretical result establishes the consistency of $\widehat{\bm \theta}_{n}^{v}$ and $\widehat{\bm \theta}_{n}^{c}$.
The proof of Theorem (ref) exploits an “in probability” version of Lemma 2.2 in Whi80 (also see Whi96 and the proof of Theorem 1 in Wei91), which does not require stationarity of the true model values $v_t(\bm \theta_0^v)$ and $c_t(\bm \theta_0^c)$. The first result that $\widehat{\bm \theta}_{n}^{v}\overset{\mathbb{P}}{\longrightarrow}\bm \theta_0^{v}$ is essentially a version of Theorem 1 in EM04; the only difference being that our regularity conditions are more involved, since we also show that $\widehat{\bm \theta}_{n}^{c}\overset{\mathbb{P}}{\longrightarrow}\bm \theta_0^{c}$. The proof of this latter result is complicated by the fact that the CoVaR score \[ S^{\operatorname{CoVaR}}\big((v_t(\bm \theta^v),c_t(\bm \theta^c))^\prime, (X_t,Y_t)^\prime\big)=\mathds{1}_{\{X_t>v_t(\bm \theta^v)\}}\big[\mathds{1}_{\{Y_t\leq c_t(\bm \theta^c)\}}-\alpha\big]\big[c_t(\bm \theta^c)-Y_t\big] \] is discontinuous in the VaR parameter $\bm \theta^v$. This fact necessitates many of the regularity conditions imposed in Assumption (ref).
We provide additional asymptotic theory in the Online Appendix, which we only briefly outline here. First, following EM04, WhiteKimManganelli2015, PZC19 and Catania2022, Theorem (ref) requires correct use of initialization information in the computation of the estimator. When using arbitrary (and, hence, possibly incorrect) initial values, Theorem (ref) in Appendix (ref) shows that consistency of the estimator is retained under some weak contraction condition on the models. This condition is trivially satisfied for the models we use in the simulations and the application; see (ref) and (ref).
Second, we also establish asymptotic normality of $(\widehat{\bm \theta}_n^v, \widehat{\bm \theta}_n^c)$ in Theorem (ref) in Appendix (ref) under the additional Assumption (ref), which we verify for CCC--GARCH and VAR models in Appendices (ref) and (ref), respectively.
Third, we propose estimators of the asymptotic variance of the limiting distribution of $(\widehat{\bm \theta}_n^v, \widehat{\bm \theta}_n^c)$ and prove their consistency in Theorem (ref) in Appendix (ref). This allows for feasible inference on the model parameters.
Here, we consider estimation of a dynamic SAV CoCAViaR model given in (ref). For this, we simulate $\big\{(X_t, Y_t)^\prime\big\}_{t=1,\ldots,n}$ from the absolute value ECCC--GARCH model in (ref) with diagonal $\widetilde{\bm A}$ and $\widetilde{\bm B}$, such that Assumptions (ref) and (ref) are met (see Appendix (ref) for the verification). We choose the parameter values $\widetilde{\bm \omega} = (0.04,\ 0.02)^\prime$, $\widetilde{\bm A} =
$, $\widetilde{\bm B} =
$ and let $F$ be the multivariate (marginally standardized) $t$-distribution with $\nu = 8$ degrees of freedom and a residual correlation of $\rho = 0.5$. We initialize the volatility process at $t=1$ with the starting values $\big(\sigma_{X,1}, \sigma_{Y,1}\big)^\prime = \widetilde{\bm \omega}$, hence complying with item (ref) in Remark (ref).
The off-diagonal zero-restrictions in $\widetilde{\bm A}$ and $\widetilde{\bm B}$ result in (more or less) the classic CCC--GARCH model of Bol90. Recall that $\widetilde{B}_{12} = 0$ is essential for our two-step M-estimator, whereas $\widetilde{B}_{21} = 0$ just facilitates the derivation of the asymptotic theory. We estimate the SAV-diag CoCAViaR model, which arises for diagonal $\bm A$ and $\bm B$ in (ref). (Table (ref) below provides a complete nomenclature of CoCAViaR models considered in this paper.) Therefore, the to-be-estimated parameters are $\bm \theta^v = (\omega_1, A_{11}, B_{11})^\prime$ and $\bm \theta^c = (\omega_2, A_{22}, B_{22})^\prime$, whose true values can be obtained from $\widetilde{\bm \omega}$, $\widetilde{\bm A}$ and $\widetilde{\bm B}$ as in footnote (ref).
Table (ref) shows simulation results for the dynamic CoCAViaR model based on $M=5000$ replications for the choices $\alpha = \beta \in \{0.9,\ 0.95\}$ and sample sizes $n \in \{500,\ 1000,\ 2000,\ 4000\}$. We correctly initialize the model recursion in the numerical estimation process by using the model parameters $\bm \omega$ as described in item (ref) of Example (ref). We show in Online Supplement (ref) that the results are almost unchanged when using an (incorrect) initialization with a constant value as discussed in item (ref) of Example (ref). The asymptotic variance-covariance matrices are estimated as detailed in Appendix (ref); see in particular Remark (ref) of that appendix. A formal description of the table columns is given in the table caption.
The columns reporting the (average and median) bias and the standard deviations confirm the consistency of the estimator from Theorem (ref). In general, the VaR parameters are estimated with smaller empirical bias than the CoVaR parameters, which is not surprising given that CoVaR is further out in the tail and, hence, subject to larger estimation uncertainty. Furthermore, the average bias is often larger than the median bias, indicating that the empirical distributions of the parameter estimates are still subject to some skewness or outliers. Notice that even for the largest sample size of $n=4000$ for our choice of $\alpha = \beta = 0.95$, the CoVaR model is essentially estimated as a $95\%$-quantile based on an effective sample size of only $\tilde n = (1-\beta) n = 200$ observations, which is an inherently difficult task. We further see that sample sizes of around 2000 days are required to reliably estimate the models for these extreme levels. This is especially true for the CoVaR parameters.
Next, we assess the precision of the standard errors and coverage of the confidence intervals for the parameters derived from Theorems (ref) and (ref) (see Appendix (ref) and (ref)). The results for the estimated standard deviations and the confidence interval coverage rates in Table (ref) show that asymptotic variance-covariance estimation is a very difficult task for (Co)CAViaR models. The empirical standard errors are somewhat overestimated for the VaR parameters in Table (ref), whereas they are interestingly more accurate for the CoVaR parameters. While the confidence intervals for the VaR parameters are rather conservative, the ones for CoVaR display some undercoverage for the extreme probability levels of $\alpha = \beta = 0.95$, and exhibit almost correct coverage for $\alpha = \beta = 0.9$. Appendix (ref) illustrates that simple adjustments of the bandwidth choices do not result in meaningful improvements, which demonstrates the need for future research on improving the estimation accuracy of the asymptotic variance-covariance matrix for dynamic (Co)CAViaR models. We mention however that for us, the main purpose of these models lies in prediction (see the next section), where inference is of lesser importance.
For the empirical forecasting application, we use daily close-to-close log-losses from January 4, 2000 until December 31, 2021 with a total of $n=5535$ trading days, obtained from the financial data provider Refinitiv. We use data for the S&P 500 index as our $Y_t$, and for $X_t$, we use Bank of America (BAC), Citigroup (C), Goldman Sachs (GS), JPMorgan Chase (JPM) and the S&P 500 Financials (SPF) that represents the financial sector of the S&P 500. The individual financial institutions are the four systemically most risky US banks according to the FSB22. The information set of interest is $\mathcal{F}_{t-1}=\sigma(X_{t-1}, Y_{t-1}, \ldots, X_1, Y_1)$ (such that $\boldsymbol{\mathcal{I}}_0$ is the empty set). Then, the quantity to be forecasted, i.e., $\operatorname{CoVaR}_{t,\alpha\mid\beta}$, measures the spillover risk of the financial system/institution to the overall economy conditional on the current state of the market (embodied by $\mathcal{F}_{t-1}$). We focus on the probability levels $\alpha = \beta = 0.95$ and estimate all models using a rolling window with estimation samples of length 3000 days. To reduce the computational burden, we only re-estimate all entertained models (introduced below) every 100 days.
In Section (ref), we present predictions and parameter estimates for six different CoCAViaR models, while Section (ref) compares these specifications with six distinct DCC--GARCH CoVaR forecasts.
We consider six forecasting models from the CoCAViaR class in this section. The first three candidate models are from the general class of SAV CoCAViaR models given in (ref). The acronym SAV indicates that the driving forces of the models are the absolute values of the log-losses, $|X_{t-1}|$ and $|Y_{t-1}|$. The top panel of Table (ref) summarizes which covariates are included in each of the employed SAV CoCAViaR model specifications. The suffix “diag” in the first model indicates that the off-diagonal elements of the parameter matrices $\bm A$ and $\bm B$ are set to zero. Similarly, “full” indicates that the full specification of (ref) is used (only restricting $B_{12} = 0$ to facilitate two-step M-estimation) and the suffix “fullA” indicates that the full matrix $\bm A$ is considered, while $\bm B$ is restricted to be a diagonal matrix.
Notice that the SAV-full model is not covered by our modeling framework (ref), because the CoVaR model depends on lags of $v_t(\bm \theta^v)$, such that $c_t(\bm \theta^c,\bm \theta^v)$ is a function of the VaR parameters $\bm \theta^v$ as well. Nonetheless, we include this specification here to show that the gains in forecast accuracy of this extension may be non-existent in practice.
Along the lines of EM04, we extend the CoCAViaR model class by using signed values of $X_t$ and $Y_t$ to the Asymmetric Slope (AS) CoCAViaR models,
where $\bm \omega \in \mathbb{R}^2, \bm A^+, \bm A^-, \bm B \in \mathbb{R}^{2 \times 2}$ are collected in the parameter vectors $\bm \theta^v$ and $\bm \theta^c$. Here, we define $x^+ = \max(x,0)$ and $x^- = -\min(x,0)$ for $x \in \mathbb{R}$. Parameter equalities in $\bm A^+$ and $\bm A^-$ can be used to generate absolute values $|X_{t-1}|$ and $|Y_{t-1}|$ in (ref). Intuitively, the positive values of $X_{t-1}$ and $Y_{t-1}$ (i.e., financial losses in our orientation) are expected to contribute more to the future VaR and CoVaR than their negative values. This is much like large losses often have more predictive content for volatility than equally large gains in the GJR--GARCH models of GJR93. The three model suffixes “pos”, “signs” and “mixed” in the bottom panel of Table (ref) imply that for “pos” only the positive components are included, for “signs” both positive and negative components are used, and “mixed” includes a mix of positive, negative and absolute value losses (see Table (ref) for details).
Table (ref) reports the estimated parameters together with their standard errors for the six CoCAViaR model specifications. The results are based on the initial estimation window of 3000 observations. All CoCAViaR models in this section are estimated using the initialization with the model parameters $\bm \omega$ as described in item (ref) of Example (ref). Except for the SAV-full model, the autoregressive coefficients are all between 0.798 and 0.958, and are highly significant. The other coefficients are all barely significant, but their direction is very reasonable. The “cross-terms” seem to be less important in general. In the SAV-full model, including the lagged $v_{t-1}(\cdot)$ and $c_{t-1}(\cdot)$ in the CoVaR model results in unintuitive parameter estimates, possibly resulting from a multicollinearity problem. This is also reflected by the model's inferior forecasting performance; see Table (ref) below. Hence, the additional generality of the SAV-full model (where $c_t(\bm \theta^c,\bm \theta^v)$ also depends on $\bm \theta^v$) does not seem to be relevant empirically. Consistent with the idea of GJR93 that losses have a larger impact on volatility than do equally large gains, we find that losses are more important than gains in predicting systemic risk in the AS models; see especially $Y_{t-1}^+$ and $Y_{t-1}^-$ in the CoVaR equation of the AS-mixed model.
Figure (ref) illustrates the rolling window forecasts from the SAV-fullA CoCAViaR model for the evaluation window ranging from December 6, 2011 until December 31, 2021. The upper panel shows the log-losses of JPMorgan Chase---the systemically most important bank according to the FSB22---together with its VaR forecasts. Losses exceeding the VaR forecasts are highlighted in black, which correspond to days with an (out-of-sample) stress event $\big\{X_t \ge \widehat \operatorname{VaR}_{t,\beta}\big\}$ in the definition of the CoVaR in Section (ref). The lower panel shows the log-losses of the S&P\,500 together with the model's CoVaR forecasts. There, the log-losses of days with a VaR exceedance (of JPMorgan Chase) are displayed in black, such that the CoVaR forecasts can be interpreted as $\alpha = 95\%$ quantile forecasts among those days with VaR exceedances. We defer a more formal investigation of the (VaR, CoVaR) forecasts to the following section.
As competitors for our six CoCAViaR models, we use six different DCC--GARCH specifications. DCC--GARCH models have attained benchmark status, because of their accurate variance-covariance matrix predictions LRV12,Laurent2013,CM14. Particularly, we use three DCC--GARCH specifications, containing two standard DCC--GARCH(1,1) models with multivariate Gaussian and $t$-distributed innovations, respectively, and a DCC specification based on a univariate GJR--GARCH(1,1) model. The models are abbreviated as “DCC-n”, “DCC-t” and “DCC-gjr”, respectively. We estimate all DCC--GARCH models by maximum likelihood using the rmgarch package of Galanos2022 for the statistical software R.
As discussed in Section (ref), different “square roots” $\bm \varSigma_t$ of the conditional covariance matrix $\bm H_t=\bm \varSigma_t\bm \varSigma_t^\prime$ may lead to different CoVaR forecasts of DCC--GARCH models (see also Appendix (ref)). Here, we consider VaR and CoVaR forecasts that are obtained by combining the above three DCC model specifications with a symmetric and a Cholesky decomposition (suffix “sym” respectively “Chol” after the model abbreviation) of the forecasted variance-covariance matrices, yielding a total of six sets of forecasts. For instance, the forecasts abbreviated “DCC-t-Chol” are computed from a DCC--GARCH with $t$-innovations based on the Cholesky decomposition of $\bm H_t$. We obtain the final VaR and CoVaR forecasts by multiplying $\bm \varSigma_t$ with a nonparametric estimate of the VaR and CoVaR of the model residuals.\footnote{The forecast evaluation results in Table (ref) differ between the DCC-n-sym and the DCC-n-Chol models (despite the normal distribution being spherical; see Section (ref) and Appendix (ref)) as the empirical distribution, which is used in the nonparametric estimate, is in general not spherical.} We refer to Appendix (ref) for details on the computation of VaR and CoVaR forecasts from DCC--GARCH models.
The multi-objective elicitability of (VaR, CoVaR) complicates inference on the predictive ability of the models, since the scoring function in (ref) is bivariate such that standard DM95 tests---based on scalar scores---cannot be applied. We follow FH24 in applying their “one and a half-sided” tests, which they illustrate via an extended traffic light system. Figure (ref) exemplarily displays their extended traffic light approach for two models. Doing so requires us to fix a baseline model, and we arbitrarily choose the “DCC-t-Chol” model for all evaluations. Such a baseline model is necessary as classical extensions to multi-model forecast comparison methods---such as the model confidence set of Hansen2011---are not available in the case of bivariate (multi-objective) scoring functions.
Multi-objective elicitability implies that the scores of two competing sequences of CoVaR forecasts can only be compared if their underlying VaR forecasts perform equally well. To obtain a reasonable finite-sample counterpart, FH24 interpret this as an insignificant score difference in the standard DM95 test for the VaR forecasts; that is, the null that the expected VaR score differences (calculated based on the first component in (ref)) are equal to zero cannot be rejected.
The red zone in Figure (ref) indicates that the baseline VaR is significantly superior, and the comparison model is rejected without consideration of its CoVaR forecasts. The grey zone indicates the reverse, while the three remaining zones in the intermediary corridor imply insignificant score differences of the VaR forecasts. Here, the orange zone implies that the baseline CoVaR is significantly superior (i.e., the CoVaR score differences based on the second component of (ref) are smaller than zero in expectation), the green zone that the alternative model is superior and the yellow zone represents insignificant score differences. As the baseline is the “DCC-t-Chol” model, a comparison with our CoCAViaR models should ideally yield results in the green zone (which indicates superior CoVaR forecasts and comparable VaR forecasts) or in the grey zone (which implies superior VaR forecasts) for the new models to have merit in practice.
Table (ref) presents results on the forecasting performance of all 12 models. For the VaR, we report the average VaR score (multiplied by 10 for better readability) using the first component of (ref), its rank among the different models, and the “hits” as the percentage of days where the loss is larger than the VaR forecast, $X_t \ge \widehat \operatorname{VaR}_{t,\beta}$. For the CoVaR, we report the average CoVaR score using the second component in (ref) (multiplied by 1000 for better readability), the corresponding rank, and the CoVaR hits defined as the percentage of days where $Y_t \ge \widehat \operatorname{CoVaR}_{t,\alpha\mid\beta}$ among all days with a VaR hit, i.e., the $t$ where $X_t \ge \widehat \operatorname{VaR}_{t,\beta}$. The VaR forecast should ideally be exceeded with probability $1-\beta=5\%$, and---on those days---the CoVaR forecast should be exceeded with probability $1-\alpha = 5\%$. Hence, for reasonable VaR and CoVaR forecasts we expect the numbers in the two “hits” columns of Table (ref) to be close to 5. The last two columns of Table (ref) report results for the previously described one and a half-sided test of FH24. There we use “DCC-t-Chol” as the baseline model, and report the test's $p$-value together with the resulting zone of their extended traffic light system. For each considered asset $X_t$, we sort the table rows (i.e., the models) according to their CoVaR score, as the CoVaR forecasts are of main interest in this section.
Among the CoCAViaR models, we find a better performance of the AS than the SAV models for the VaR, but a relatively comparable performance for CoVaR forecasting. A reason for this may be that the additional VaR parameters in the AS models are estimated with more precision and, hence, the predictive content of the positive/negative parts emerges more clearly than for CoVaR, where the effective sample size is much reduced. It is further noteworthy that the SAV-full CoCAViaR model, which includes the lagged VaR in the CoVaR equation, performs below average. This might be caused by a high collinearity of the VaR and CoVaR, and also shows that the restriction $B_{12}=0$---which we have to impose for our two-step M-estimator---is likely to be unproblematic in practice.
Overall, we find a superior forecasting performance of the CoCAViaR models for all five employed assets for $X_t$ compared to the DCC models. While the rankings of the VaR scores vary over the different assets, the superior forecasting performance of the CoCAViaR models is more substantial for the CoVaR. This is supported by the CoVaR hits (corresponding to unconditional forecast calibration) being much closer to the nominal level of $5\%$ than for the DCC models, whose hit frequencies are almost all above $10\%$. While none of the CoCAViaR models are significantly outperformed by being in the red or orange zone, many significantly outperform the baseline DCC model and are located in the green and grey zones. Also notice that among the CoCAViaR models, the SAV-full specification (not covered by our theoretical framework) performs below average. Therefore, our modeling framework (ref) seems to be sufficiently general to capture the main features of (systemic) risk evolution through time.
In comparing the (VaR, CoVaR) forecasts of our CoCAViaR models with those issued by the DCC--GARCH models, we have used the same scoring function (ref) in the comparative backtest as we did for estimating the CoCAViaR models. Therefore, one may be concerned that using the same scoring function to estimate our CoCAViaR models and to evaluate the predictions, lends an unfair advantage in the forecast comparison to the CoCAViaR models. To alleviate such concerns, we also use a scoring function different from (ref) in the comparative backtest of FH24 as a robustness check. For brevity, the results are reported in Appendix (ref) and they show that the dominance of our CoCAViaR models still holds.
The superiority of our CoCAViaR models may be surprising, because the second-stage estimator is based on the reduced estimation sample where $\{X_t \geq \widehat{\operatorname{VaR}}_{t,\beta} \}$. Yet, doing so allows our models to focus on the salient features of the data in the tail, thus offsetting the drawback of the lower effective sample size. This focus means, however, that we only model a very narrow---although for many purposes important---aspect of the conditional distribution of $(X_t,Y_t)^\prime\mid\mathcal{F}_{t-1}$ (namely the VaR and CoVaR). In contrast, multivariate GARCH models are capable of modeling the complete predictive distribution of $(X_t,Y_t)^\prime\mid\mathcal{F}_{t-1}$, but they may not excel at modeling each aspect of it equally well. This echoes EM04, who conclude in the context of univariate VaR forecasting that “[a]lthough GARCH might be a useful model for describing the evolution of volatility, the results in this article show that it might provide an unsatisfactory approximation when applied to tail estimation.”
Our first main contribution is to propose a flexible, semiparametric approach for modeling (VaR, CoVaR) over time. Since we only model (VaR, CoVaR) and not the full predictive distribution, we “let the tails speak for themselves” DuM83. As we find in an empirical application on the systemic riskiness of US financial institutions, this yields models that improve upon benchmark DCC--GARCH processes in terms of predictive accuracy. As a second main contribution, we propose a two-step M-estimator for the model parameters and derive its large sample properties. From an econometric perspective, our proofs are non-standard since the loss function that has to be used for the CoVaR is not only non-smooth, but also discontinuous in the VaR model (parameter).
We expect our modeling framework to have applications beyond the one considered here, for instance in predictive CoVaR regressions as in AB16. Furthermore, in macroeconomics, ABG19,Aea21 have recently drawn attention to tail risks and their interconnections by popularizing the Growth-at-Risk. Hence, the models proposed in this paper could become relevant for studying interconnections of macroeconomic risks. Much like BS21 compared the predictive accuracy of different Growth-at-Risk models, our models may be used as competitors in Co-Growth-at-Risk comparisons.