EconBase
← Back to paper

Persistence-Robust Break Detection in Predictive Quantile and CoVaR Regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

70,132 characters · 15 sections · 87 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Persistence-Robust Break Detection in Predictive CoVaR Regressions

\baselineskip18pt \setcounter{totalnumber}{50} \setcounter{topnumber}{50} \setcounter{bottomnumber}{50} \abovedisplayskip1.5ex plus1ex minus1ex \belowdisplayskip1.5ex plus1ex minus1ex \abovedisplayshortskip1.5ex plus1ex minus1ex \belowdisplayshortskip1.5ex plus1ex minus1ex

abstractForecasting systemic risk (as measured by AB16's AB16 CoVaR) is important in economics and finance. However, predictive relationships may be unstable over time. Therefore, this paper develops structural break tests in predictive CoVaR regressions. These tests can detect changes in the forecasting power of covariates, and are based on the principle of self-normalization. We show that our tests are valid irrespective of whether the predictors are stationary or near-stationary, rendering the inference procedures suitable for a wide range of practical applications. Also, we derive the power of our tests under local alternatives, which shows that power is higher when predictors are near-stationary instead of stationary. As a final theoretical contribution, we propose unsupervised change point tests that are consistent even in the presence of arbitrarily many breaks in the predictive relationship. Simulations illustrate the good finite-sample properties of our methods. An empirical application to systemic risk forecasting models for the US banking system shows the usefulness of our tests by uncovering changes in the predictive content of the VIX.\\ Keywords: Change Points, CoVaR, Mild Integration, Predictive Regressions, Quantiles \\ JEL classification: C12 (Hypothesis Testing); C52 (Model Evaluation, Validation, and Selection); C58 (Financial Econometrics)

Motivation

Recent financial crises and their widespread impacts have heightened awareness of the systemic nature of risk AB16,Aea17,VZ19. While markets often experience higher levels of volatility during such turbulent periods---indicating greater overall risk---this does not automatically translate to increased systemic risk. For example, banks may exhibit more volatile returns, signaling higher individual risk, yet without a corresponding rise in their co-movements (i.e., no increase in systemic risk). However, it is this rise in commonality, that has become a key concern for regulators, as it suggests the potential for widespread financial distress and significant economic consequences GKP16.

Therefore, researchers have investigated the predictive content of numerous variables for future levels of systemic risk. The perhaps most important systemic risk measure in this context is the conditional Value-at-Risk (CoVaR) of AB16. Many researchers have related this measure to lagged covariates via a so-called CoVaR regression. Among the studied predictors are bank-level variables, such as (the log of) total assets BDP20,BRS20, and financial variables, such as equity volatility AB16,BDP20,BRS20 or the VIX Hea16. But also macroeconomic variables have been investigated, such as inflation BRS20, the TED spread as a measure of liquidity risk AB16,BDP20, and 10-year government bond rates BRS20.

To this date, there are no methods that allow to test for the stability of these predictive relations. Such change point tests for predictive CoVaR regressions are, however, important for at least two reasons. First, unstable models should be used with caution (if at all) for forecasting purposes PPP13. Second, knowledge of structural breaks may improve modeling attempts and generate further research into the underlying causes of the instabilities. For instance, ST21 carry out this program for predictive mean regressions. Specifically, they show that suitably accounting for breaks improves the forecasting performance of several predictors of stock returns. It is the main purpose of this paper to fill this gap in the literature by proposing suitable change point tests for predictive CoVaR regressions.

Indeed, the theoretical literature on CoVaR regressions seems underdeveloped. Currently, there exist only theoretical results for standard inference on (constant) parameters in CoVaR regressions with stationary regressors Lea24,DH26. In particular, these papers do not consider the problem of structural break testing. We stress that, while the theoretical literature on CoVaR regressions is limited, they are frequently applied in practice AB16,BRS20,BDP20.

One key feature of our change point tests is their persistence-robustness in the sense that they allow predictors to be stationary or near-stationary. This is useful because there are many applications of CoVaR regressions, where covariates exhibit different degrees of serial dependence. For instance, among the predictors mentioned above, the short-term TED spread (i.e., the difference between the 3-month LIBOR rate and the 3-month secondary market Treasury bill rate) and equity volatility, are known to be reasonably persistent. Likewise, inflation as a predictor of systemic risk is strongly serially dependent. Moreover, 10-year government bond rates, that account for the connection between sovereigns and banks, are a classic example of a highly persistent time series. In light of this variety of predictors displaying different degrees of serial dependence, change point tests possessing some persistence-robustness---such as ours---seem to be called for.

We mention that structural break tests in regression contexts typically require knowledge of the persistence of the predictors. For instance, in their change point tests for predictive mean regressions, Zea23a show that asymptotic distributions differ between stationary and non-stationary covariates, such that the correct choice of critical values depends on whether the predictors are stationary or non-stationary. Conveniently, this is not the case for our test.

Although our main emphasis in this paper is on developing change point tests for CoVaR regressions, as a corollary we also obtain change point tests for simple predictive quantile regressions (QRs). The reason for this is that CoVaR regressions can only be estimated jointly with quantile regressions (FH24). Of course, testing for changes in predictive QRs is equally important as doing so in CoVaR regressions, due to the significant interest that the former have garnered in the past (Lee16,FL19,cai2022new,FLS23,Lea24a,MSK24,MWT25). The works most closely related to our persistence-robust change point tests for predictive QRs are those of Qu08 and OQ11, who develop structural break tests for quantile regressions with exclusively stationary covariates. Our test differs from theirs in the following respects. First, our structural change test is robust to whether predictors are stationary or near-stationary. What is more, under the alternative, we show that local power is higher when the predictors are highly persistent. Second, our tests are based on the principle of self-normalization (SN), pioneered in a change point context by SZ10. Third, we also derive “unsupervised” tests in the spirit of ZL18, which do not require a priori knowledge of the number of change points.

All our persistence-robust change point tests work by comparing estimates of regression coefficients based on suitably chosen subsamples. We show that these subsample estimates converge (in a functional sense) to Brownian motion both when predictors are stationary and near-stationary. Because these Brownian motions only differ in their covariance matrices, our self-normalized structural break tests following SZ10 are valid for stationary and near-stationary predictors in the sense that no a priori knowledge of their persistence is required. The technical development relies on convexity results of Kat09 and recent results for stationary and near-stationary variables in MP20. Note that we only draw on MP20 to utilize certain properties of stationary and near-stationary predictors. We use these results for a distinctly different end, viz. (quantile, CoVaR) regressions---much unlike MP20, who consider (cointegrating) mean regressions. Moreover, we go beyond MP20 by developing functional central limit theory for the regression estimators.

Apart from delivering persistence-robustness, SN also offers further advantages. First, it is straightforward to implement, as it only requires recursively estimated coefficients as inputs. Second, our functional results underlying the SN-based change point tests also facilitate “off-the-shelf” inference for the full-sample estimates when no break was detected; see Section (ref) for an application. Third, compared with approaches that directly estimate asymptotic variances, SN-based tests often offer more adequate size control in finite samples Sha10,Sha15. Fourth, “standard” change point tests that involve consistent estimation of asymptotic (long-run) variances often suffer from the problem of non-monotonic power, where power may even decrease in the distance from the null Vog99. By a clever construction of the normalizer, SN sidesteps this problem and, hence, does not exhibit this non-monotonicity SZ10.

Simulations show that our tests possess excellent finite-sample properties. For both stationary and near-stationary predictors, size is close to the nominal level for sample sizes typically encountered in practice. This not only holds for the quantile regression but also for the CoVaR regression, where the effective sample size is severely reduced by the very definition of CoVaR as a systemic risk measure. We also show that our tests have high power in detecting deviations from the null of a stable predictive relationship.

Our second main contribution is to apply our test to a systemic risk prediction model for the US financial sector. As a predictor, we use the volatility index (VIX), which measures the short-term future volatility in the S&P 500 expected by market participants Wha09. The VIX and other measures of market volatility are widely used as predictors of systemic risk AB16,Hea16,Ned20. We measure systemic risk for all global systemically important US banks with respect to the S&P 500 Financials. While throughout our sample a high VIX signals higher future systemic risk, our test shows that the precise predictive power of the VIX is liable to change over time.

The remainder of this paper is structured as follows. We introduce basic notation and definitions in Section (ref). Then, Section (ref) presents our change point tests for (quantile, CoVaR) regressions. Section (ref) summarizes the results of extensive Monte Carlo simulations, and Section (ref) contains our systemic risk application. The final Section (ref) concludes. To promote flow in the main paper, we relegate the introduction of some assumptions and the asymptotic variance-covariance matrices of the parameter estimators to Appendices (ref)--(ref). The appendix further contains detailed proofs and an additional application in Sections (ref)--(ref).

Preliminaries

Defining Quantiles and the CoVaR

Let $Z_t$ denote the variable of interest and let $Y_t$ stand for the “distress” variable. The systemic importance of $Y_t$ for $Z_t$ will then be measured by the CoVaR. In the empirical application in Section (ref), $Z_t$ will denote the log-losses of a large US bank and $Y_t$ those of the S&P 500 Financials. The $(k+1)\times 1$-vector of predictors is $\bm X_{t-1}=(1,\bm x_{t-1}^\prime)^\prime$, where $\bm x_{t-1}$ contains the $\mathbb{R}^{k}$-valued, non-constant predictors. (We use bold font to denote vector- or matrix-valued quantities and normal font to denote scalar-valued objects.) As mentioned in the Motivation, several possibly persistent variables have been used in CoVaR regressions as predictors.

To formally introduce quantiles and the CoVaR, let $\mathcal{F}_{t}=\sigma(Z_t, Y_t, \bm X_t,Z_{t-1}, Y_{t-1}, \bm X_{t-1},\ldots)$ denote the time-$t$ information set, and let $\alpha,\beta\in(0,1)$ be two probability levels. The quantity to be modeled in the auxiliary QR is the conditional $\alpha$-quantile $Q_{\alpha}(Y_t\mid\mathcal{F}_{t-1}):=Q_{\alpha}(F_{Y_t\mid\mathcal{F}_{t-1}}):=F_{Y_t\mid\mathcal{F}_{t-1}}^{\leftarrow}(\alpha)$ of the cumulative distribution function $F_{Y_t\mid\mathcal{F}_{t-1}}(\cdot):=\mathbb{P}\{Y_t\leq\cdot\mid\mathcal{F}_{t-1}\}$.

The CoVaR is the $\beta$-quantile of the distribution of $Z_t\mid \big\{Y_t\geq Q_{\alpha}(Y_t\mid\mathcal{F}_{t-1}),\mathcal{F}_{t-1}\big\}$, i.e., \[ \operatorname{CoVaR}_{\beta\mid\alpha}\big((Z_t,Y_t)^\prime\mid\mathcal{F}_{t-1}\big)=Q_{\beta}\big(F_{Z_t\mid Y_t\geq Q_{\alpha}(Y_t\mid\mathcal{F}_{t-1}),\, \mathcal{F}_{t-1}}\big). \] Thus, the CoVaR is only distinct from a standard quantile by the conditioning event $\{Y_t\geq Q_{\alpha}(Y_t\mid\mathcal{F}_{t-1})\}$, which we interpret as a “distress event” for large $\alpha$. Therefore, as $\alpha\uparrow1$, the CoVaR measures the risk in $Z_t$ (via its $\beta$-quantile) conditional on $Y_t$ being in distress, such that the systemic risk interpretation of the CoVaR becomes apparent. For $\alpha=0$, the distress event occurs with probability one, and the CoVaR degenerates to a quantile, i.e., $\operatorname{CoVaR}_{\beta\mid0}\big((Z_t,Y_t)^\prime\mid\mathcal{F}_{t-1}\big)=Q_{\beta}(Z_t\mid\mathcal{F}_{t-1})$.

Predictive (Quantile, CoVaR) Regressions

We model quantiles and the CoVaR via the linear (quantile, CoVaR) regression (henceforth simply called a CoVaR regression)

align[align omitted — 320 chars of source]

The model in (ref) is a standard linear predictive quantile regression. The assumption on the QR error $\epsilon_t$ in (ref) ensures that $Q_{\alpha}(Y_t\mid\mathcal{F}_{t-1})=\bm X_{t-1}^\prime\bm \alpha_0$, whence the coefficient $\bm \alpha_0$ measures the strength of the predictive content of $\bm X_{t-1}$ for the (conditional) $\alpha$-quantile of $Y_t$. Analogously, the assumption on the (quantile, CoVaR) errors $(\epsilon_t,\delta_t)^\prime$ in (ref) implies that $\operatorname{CoVaR}_{\beta\mid\alpha}\big((Z_t,Y_t)^\prime\mid\mathcal{F}_{t-1}\big)=\bm X_{t-1}^\prime\bm \beta_0$, such that $\bm \beta_0$ quantifies the forecasting power of $\bm X_{t-1}$ for future levels of systemic risk. In Appendix (ref) of the simulations, we present a concrete data-generating process (DGP) for $(Y_t,Z_t)^\prime$ that gives rise to linear (quantile, CoVaR) regressions as in (ref)--(ref).

To estimate the predictive CoVaR regression in (ref)--(ref), we adapt the estimator of DH26 to our context. Given a sample $\big\{(Z_t,Y_t,\bm X_{t-1}^\prime)^\prime\big\}_{t=1,\ldots,n}$, we estimate the parameter vector $\bm \alpha_0$ via

equation*[equation* omitted — 164 chars of source]

where $\rho_{\alpha}(u)=u(\alpha - \mathds{1}_{\{u\leq0\}})$ is the standard pinball loss known from quantile regressions KB78. To estimate $\bm \beta_0$, we use

equation*[equation* omitted — 221 chars of source]

This estimator is similar in spirit to the QR estimator $\widehat{\bm \alpha}_n$, except that in minimizing the loss, only those observations $(Z_t,Y_t,\bm X_{t-1}^\prime)^\prime$ are used for which $Y_t$ is larger than its predicted conditional quantile, i.e., $Y_t>\bm X_{t-1}^\prime\widehat{\bm \alpha}_n$. We mention that this restriction to observations with a quantile exceedance is unavoidable by definition of the CoVaR as a quantile of the distribution $F_{Z_t\mid Y_t\geq Q_{\alpha}(Y_t\mid\mathcal{F}_{t-1}),\mathcal{F}_{t-1}}$ that conditions on $Y_t\geq Q_{\alpha}(Y_t\mid\mathcal{F}_{t-1})$. For a more technical argument, we refer to Theorem 4.2 (ii) of FH24. Underlying the estimators $\widehat{\bm \alpha}_n$ and $\widehat{\bm \beta}_n$ is the assumption that the predictive relationships in (ref)--(ref) are unchanged throughout the sample.

CoVaR Regressions With Instabilities

Since the predictive content of $\bm X_{t-1}$ may change over time, the coefficients $\bm \alpha_0$ and $\bm \beta_0$ in (ref)--(ref) may vary with $t$, such that

align[align omitted — 320 chars of source]

Our goal is to construct a test of the no-break hypothesis for model (ref)--(ref), i.e., \[ \mathcal{H}_0^{\operatorname{CoVaR}}\colon

pmatrix[pmatrix omitted — 47 chars of source]

=\ldots=

pmatrix[pmatrix omitted — 47 chars of source]

\equiv

pmatrix[pmatrix omitted — 43 chars of source]

. \] To do so, we follow SZ10 and rely on estimates based on certain subsamples $\big\{(Z_t, Y_t,\bm X_{t-1}^\prime)^\prime\big\}_{t=\lfloor nr\rfloor + 1,\ldots,\lfloor ns\rfloor}$, where $0\leq r<s\leq1$ and $\lfloor \cdot\rfloor$ denotes the floor function. Specifically, the subsample estimates for the CoVaR regression are

align*[align* omitted — 454 chars of source]

Define the stacked estimator $\widehat{\bm \gamma}_n(r,s):=\big(\widehat{\bm \alpha}_n^\prime(r,s),\widehat{\bm \beta}_n^\prime(r,s)\big)^\prime$. Then, our self-normalized test statistic for $\mathcal{H}_0^{\operatorname{CoVaR}}$, inspired by SZ10, is given by \[ \mathcal{U}_{n,\bm \gamma}:=\sup_{s\in[\iota,1-\iota]} s^2(1-s)^2\big[\widehat{\bm \gamma}_n(0,s) - \widehat{\bm \gamma}_n(s,1)\big]^\prime \bm{\mathcal{N}}_{n,\bm \gamma}^{-1}(s)\big[\widehat{\bm \gamma}_n(0,s) - \widehat{\bm \gamma}_n(s,1)\big] \] with cut-off $\iota\in(0,1/2)$ and normalizer

multline*[multline* omitted — 415 chars of source]

The cut-off $\iota$ ensures that estimates are based on at least a $\iota$-fraction of the available sample, that is, on $\lfloor n\iota\rfloor$ many observations. For technical reasons, the choice $\iota=0$ is not allowed, which is similar as in the fixed-regressor QRs considered in ZS13.

The test statistic $\mathcal{U}_{n,\bm \gamma}$ has two different components. First, the outer terms involving $\big[\widehat{\bm \gamma}_n(0,s) - \widehat{\bm \gamma}_n(s,1)\big]$ indicate a structural instability in (ref)--(ref) if they are far from zero and, thus, give the test its power. When $s$ is close to 0 or 1, this difference is weighted down by the factor $s^2(1-s)^2$, because then either $\widehat{\bm \gamma}_n(0,s)$ or $\widehat{\bm \gamma}_n(s,1)$ is based on very few observations, such that larger differences between the two estimates are not uncommon. The second component of $\mathcal{U}_{n,\bm \gamma}$ is the normalizer $\bm{\mathcal{N}}_{n,\bm \gamma}(s)$, which serves to give a nuisance parameter-free limiting distribution (see the proof of Corollary (ref)).

To derive the asymptotic limit of our test statistic $\mathcal{U}_{n,\bm \gamma}$, we require some assumptions on the predictors $\bm x_{t-1}$. We introduce these in the following Section (ref) (see Assumptions (ref)--(ref)). Additional required assumptions on the predictive QR model in (ref) (see Assumptions (ref)--(ref) in Appendix (ref)), and the predictive CoVaR regression in (ref) (see Assumptions (ref)--(ref) in Appendix (ref)) are relegated to the appendix to promote flow. Importantly, Assumptions (ref)--(ref) in Appendices (ref)--(ref) are sufficiently general to allow for conditional heteroskedasticity of the regression errors $(\epsilon_t,\delta_t)^\prime$.

Main Results

Assumptions on the Predictors

In line with most of the predictive regression literature, we assume the stochastic predictors $\bm x_{t}$ in $\bm X_{t}=(1,\bm x_{t}^\prime)^\prime$ to be generated by the additive components model \[ \bm x_{t}=\bm \mu_{x} + \bm \xi_t, \] where $\bm \mu_{x}$ is an $\mathbb{R}^{k}$-valued constant and the (zero-mean) stochastic component $\bm \xi_t$ obeys the following autoregression:

assumption[Predictors] The stochastic $\bm \xi_t$ are generated by \begin{equation} \bm \xi_t = \bm R_n\bm \xi_{t-1} + \bm u_{t},\qquad t\in\mathbb{N}, \end{equation} where $\bm u_{t}$ is a linear process defined in Assumption (ref) below and the autoregressive matrix $\bm R_n$ satisfies \begin{equation} \bm C_n:=n^{\kappa}(\bm R_n-\bm I_k)\underset{(n\to\infty)}{\longrightarrow}\bm C, \end{equation} where $\bm I_{k}$ denotes the $(k\times k)$-identity matrix and $\kappa\geq0$. The variables $\bm \xi_t$ belong to one of the following classes: \begin{enumerate} • Stationary predictors: Equation (ref) holds with $\kappa=0$ and $\bm R:=\bm R_n=\bm I_k + \bm C$ has spectral radius $\rho(\bm R)<1$. • Near-stationary predictors: Equation (ref) holds with $\kappa\in(0,1)$ and $\bm C$ is a negative stable matrix, i.e., all its eigenvalues have negative real part. \end{enumerate} The process $\bm \xi_t$ in (ref) is initialized at $\bm \xi_0=O_{\mathbb{P}}(1)$ under (I0), and $\bm \xi_0=o_{\mathbb{P}}(n^{\kappa/2})$ under (NS).

A similar condition to Assumption (ref) is entertained by KMS23. Our Assumption (ref) covers the cases of stationary and near-stationary regressors in Assumption N from MP20 with, in their notation, $\kappa_n=n^{\kappa}$. Variables generated according to (NS) in Assumption (ref) have also been termed mildly integrated PM09,Lee16 or moderately integrated MP09. However, we adopt the terminology of MP20 here, as it is closest to our framework. Allowing $\kappa=1$ in (NS) would lead to what MP20 call near-nonstationary regressors. These are closely related to nearly integrated variables, considered in the earlier literature CES95,JM06a.

Of course, it would be desirable for additional robustness to also allow for near-nonstationary covariates with $\kappa=1$ in Assumption (ref). However, in that case, limit theory is already non-Gaussian for the quantile regression Lee16, such that our self-normalized approach cannot be expected to work. Covering near-nonstationary regressors requires use of other methods, such as IVX-based approaches. Yet, investigating these is beyond the scope of the present paper, and is left for future research.

It will turn out that the estimators $\widehat{\bm \alpha}_n(r,s)$ and $\widehat{\bm \beta}_n(r,s)$ have a different convergence rate depending on whether the predictors are stationary or near-stationary in Assumption (ref). To unify notation across these two cases, we introduce the normalizing matrix \[ \bm D_n=

cases\sqrt{n}\bm I_{k+1} & for (I0),\\ \operatorname{diag}(\sqrt{n},n^{\frac{1+\kappa}{2}}\bm I_{k}) & for (NS).

\]

For the predictor innovations $\bm u_t$, we follow the linear framework of MP20. To that end, let $\Vert\cdot\Vert$ denote the spectral norm.

assumption[Predictor innovations] For each $t\in\mathbb{N}$, $\bm u_t$ has linear process representation \[ \bm u_t=\sum_{j=0}^{\infty}\bm F_j\bm \varepsilon_{t-j},\qquad \sum_{j=0}^{\infty}\Vert\bm F_j\Vert<\infty,\qquad \sum_{j=1}^{\infty}j\Vert\bm F_j\Vert^2<\infty, \] where $\bm F_0:=\bm I_{k}$, $\bm F(1):=\sum_{j=1}^{\infty}\bm F_j$ has full rank, and $\bm \varepsilon_t$ is a $\mathbb{R}^{k}$-valued martingale difference sequence with respect to $\widetilde{\mathcal{F}}_{t}=\sigma(\bm \varepsilon_t,\bm \varepsilon_{t-1},\ldots)$, such that $\mathbb{E}_{\widetilde{\mathcal{F}}_{t-1}}[\bm \varepsilon_t\bm \varepsilon_t^\prime]=\bm \varSigma_{\varepsilon}>0$ and the sequence $\{\Vert\bm \varepsilon_t\Vert^2\}_{t\in\mathbb{Z}}$ is uniformly integrable.

Often in the predictive regression literature, it is assumed that the regression disturbances are uncorrelated with the predictor increments $\bm u_t$ Dea22. We do not require a similar condition in our linear process Assumption (ref) for the $\bm u_t$. Also note that while the requirement $\mathbb{E}_{\widetilde{\mathcal{F}}_{t-1}}[\bm \varepsilon_t\bm \varepsilon_t^\prime]=\bm \varSigma_{\varepsilon}>0$ rules out conditional heteroskedasticity of the $\bm \varepsilon_t$ with respect to their own past, conditional on the full information set $\mathcal{F}_{t-1}$, variances of the $\bm \varepsilon_t$ are allowed to vary.

The sole purpose of Assumptions (ref)--(ref) is to ensure that we can draw on certain results of MP20 for stationary and near-stationary predictors $\bm x_{t-1}$. The remaining development for our CoVaR regressions in (ref)--(ref) instead relies on results for estimators derived from convex minimization problems Kat09. To be able to apply those, the appendix imposes some regularity conditions on (ref) (Assumptions (ref)--(ref) in Appendix (ref)) and (ref) (Assumptions (ref)--(ref) in Appendix (ref)).

Change Point Test for Predictive CoVaR Regressions

Before presenting our first main result, we have to introduce some additional notation. Define $\mathcal{D}_{\iota}=\big\{(r,s)\in[0,1]^2\colon \iota\leq r<s\leq 1-\iota,\ s-r\geq\iota\big\}$. Let $\ell^{\infty}(\mathcal{D}_{\iota})$ denote the space of real-valued, bounded functions on $\mathcal{D}_{\iota}$ endowed with the uniform topology VW96. The $d$-fold product of this space is denoted by $(\ell^{\infty}(\mathcal{D}_{\iota}))^d$, which comes equipped with the product topology.

thmSuppose $\mathcal{H}_0^{\operatorname{CoVaR}}$ holds true for the model (ref)--(ref). If Assumptions (ref)--(ref) are satisfied, then, as $n\to\infty$, \begin{equation*} (s-r)\begin{pmatrix}\bm D_n\big[\widehat{\bm \alpha}_n(r,s) - \bm \alpha_0\big]\\ \bm D_n\big[\widehat{\bm \beta}_n(r,s) - \bm \beta_0\big] \end{pmatrix} \overset{d}{\longrightarrow}\overline{\bm \varSigma}^{1/2}\big[\overline{\bm W}(s)-\overline{\bm W}(r)\big]\qquadin (\ell^{\infty}(\mathcal{D}_{\iota}))^{2k+2}, \end{equation*} where $\overline{\bm W}(\cdot)$ is a $(2k+2)$-variate standard Brownian motion, and $\overline{\bm \varSigma}$ is defined in Appendix (ref).
proofSee Appendix (ref) for a sketch of the proof. Appendices (ref)--(ref) provide full detail.

The proof of Theorem (ref) draws on several sources. The general outline of the proof is similar to that of Theorem 2 in HS25, where they show the functional convergence of parameter estimators in linear (quantile, expected shortfall) regressions. Apart from considering a completely different functional, the main technical differences are as follows. First, we go beyond HS25 by considering two-parameter convergence in Theorem (ref). Doing so allows us to derive an “unsupervised” break test in the spirit of ZL18 later on (see Section (ref)), which does not require the number of breaks to be pre-specified. Second, in contrast to the stationary setup in HS25, the limit theory in Theorem (ref) unifies the cases of stationary and near-stationary regressors. We are able to do so by leveraging several results of MP20 for (I0) and (NS) variables.

The precise form of the asymptotic variance-covariance matrix $\overline{\bm \varSigma}$ is irrelevant for our test, because---by virtue of self-normalization---it simply cancels out in the limiting distribution of our test statistic $\mathcal{U}_{n,\bm \gamma}$ (see Corollary (ref)). The (non-functional) convergence for fixed $r=0$ and $s=1$ of Theorem (ref) for $\widehat{\bm \alpha}_n(r,s)$ is similar to that in Lee16 and FL19 (who, however, go beyond the (I0) and (NS) case of Assumption (ref) by also considering near-nonstationary and mildly explosive regressors). Yet, on account of the functional convergence, Theorem (ref) provides a much stronger conclusion in the (I0) and (NS) case, which may be used to derive uniformly valid inference on structural breaks in the predictive relationship, as we show next.

corUnder the conditions of Theorem (ref) it holds that, as $n\to\infty$, \begin{equation*} \mathcal{U}_{n,\bm \gamma} \overset{d}{\longrightarrow}\sup_{s\in[\iota,1-\iota]}\big[\overline{\bm W}(s)-s\overline{\bm W}(1)\big]^\prime \overline{\bm{\mathcal{\bm W}}}^{-1}(s)\big[\overline{\bm W}(s)-s\overline{\bm W}(1)\big]=:\mathcal{U}_{2k+2}, \end{equation*} where $\overline{\bm W}(\cdot)$ is a $(2k+2)$-variate standard Brownian motion and the normalizer is \begin{align*} \overline{\bm{\mathcal{\bm W}}}(s) &= \int_{\iota}^{s}\big\{\overline{\bm W}(r)-(r/s)\overline{\bm W}(s)\big\}\big\{\overline{\bm W}(r)-(r/s)\overline{\bm W}(s)\big\}^\prime\,\mathrm{d} r \\ & + \int_{s}^{1-\iota}\Big\{\big[\overline{\bm W}(1)-\overline{\bm W}(r)\big] - \big((1-r)/(1-s)\big)\big[\overline{\bm W}(1)-\overline{\bm W}(s)\big]\Big\}\times\\ &\times\Big\{\big[\overline{\bm W}(1)-\overline{\bm W}(r)\big] - \big((1-r)/(1-s)\big)\big[\overline{\bm W}(1)-\overline{\bm W}(s)\big]\Big\}^\prime\,\mathrm{d} r. \end{align*}
proofThe result follows by a straightforward application of the continuous mapping theorem to Theorem (ref); see Appendix (ref) for full detail.

Denote by $\mathcal{U}_{\ell,1-\nu}$ the $(1-\nu)$-quantile of the distribution of $\mathcal{U}_{\ell}$ $(\ell\in\mathbb{N})$ for $\nu\in(0,1)$. Then, Corollary (ref) shows that rejecting $\mathcal{H}_{0}^{\operatorname{CoVaR}}$ if $\mathcal{U}_{n,\bm \gamma} > \mathcal{U}_{2k+2,1-\nu}$ leads to an asymptotic level-$\nu$ test.

Table (ref) shows some selected $(1-\nu)$-quantiles of the limiting distribution $\mathcal{U}_{\ell}$ from Corollary (ref) for different $\ell$'s. The critical values have been computed based on 500,000 draws from the limiting distribution, where the Brownian motion was approximated on a grid of 5,000 equally spaced points in $[0,1]$.

table[table omitted — 909 chars of source]

Importantly, these critical values are valid irrespective of whether predictors are (I0) or (NS), which delivers the persistence-robustness of our test. The reason we obtain such a unified approach lies in Theorem (ref), which shows that the functional limit of the subsample estimates is Brownian motion with only the variance-covariance matrix differing between the (I0) and (NS) case. These covariance matrices, however, cancel out in the limit due to SN. The only quantity appearing in $\mathcal{U}_{2k+2}$ that is related to the data-generating process is the dimension $k$ of the predictors. In our case, $k$ is a fixed integer.

rem[Computation of test statistic] It is often computationally easier to approximate the test statistic $\mathcal{U}_{n,\bm \gamma}$ in the following way. First, for $j\in\mathbb{N}$ define $\bm T_n(j)=(j/n)(1-j/n)\big[\widehat{\bm \gamma}(0,j/n) - \widehat{\bm \gamma}(j/n, 1)\big]$, such that $\widehat{\bm \gamma}(0,j/n)$ ($\widehat{\bm \gamma}(j/n,1)$) is the regression estimate based on observations from $t=1,\ldots,j$ ($t=j+1,\ldots,n$). Then, $\mathcal{U}_{n,\bm \gamma}$ can be approximated via \begin{equation*} \widetilde{\mathcal{U}}_{n,\bm \gamma}=\sup_{j=\lfloor\iota n\rfloor+1,\ldots,\lfloor(1-\iota) n\rfloor}\bm T_n^{\prime}(j)\bm V_n^{-1}(j)\bm T_n(j), \end{equation*} where \begin{multline*} \bm V_n(j) = \frac{1}{n}\bigg\{\sum_{i=\lfloor\iota n\rfloor+1}^{j}(i/n)^2\big[\widehat{\bm \gamma}(0,i/n) - \widehat{\bm \gamma}(0, j/n)\big]\big[\widehat{\bm \gamma}(0,i/n) - \widehat{\bm \gamma}(0,j/n)\big]^\prime\\ \sum_{i=j}^{\lfloor(1-\iota) n\rfloor}(1-i/n)^2\big[\widehat{\bm \gamma}(i/n,1) - \widehat{\bm \gamma}(j/n, 1)\big]\big[\widehat{\bm \gamma}(i/n,1) - \widehat{\bm \gamma}(j/n, 1)\big]^\prime\bigg\}. \end{multline*} Lengthy, but straightforward, calculations show that $\mathcal{U}_{n,\bm \gamma}-\widetilde{\mathcal{U}}_{n,\bm \gamma}=o_{\mathbb{P}}(1)$, such that a test of $\mathcal{H}_{0}^{\operatorname{CoVaR}}$ may also be based on $\widetilde{\mathcal{U}}_{n,\bm \gamma}$.
rem[Mixed-persistence predictors] A careful reading of the proofs shows that Corollary (ref) also holds for “mixed”-persistence regressors, where $k_{1}$ predictors (say $\bm X_{t-1,(I0)}$) are stationary and $k-k_1$ are near-stationary (say $\bm X_{t-1,(NS)}$), as long as Lemma (ref) in Appendix (ref) continues to hold for the $\bm X_{t-1}=(\bm X_{t-1,(I0)}^\prime, \bm X_{t-1,(NS)}^\prime)^\prime$ and some limiting matrix $\bm \varOmega$ (then with normalizing matrix $\bm D_n=\operatorname{diag}(\sqrt{n}\bm I_{k_1},\, n^{\frac{1+\kappa}{2}}\bm I_{k-k_1})$).

Finally, we show how Theorem (ref) can be specialized to construct a change point test for predictive QRs.

corSuppose $\mathcal{H}_0^{Q}\colon\bm \alpha_{0,1}=\ldots=\bm \alpha_{0,n}\equiv\bm \alpha_0$ holds true for model (ref). If Assumptions (ref)--(ref) are satisfied, then, as $n\to\infty$, \begin{equation*} \mathcal{U}_{n,\bm \alpha}:=\sup_{s\in[\iota,1-\iota]} s^2(1-s)^2\big[\widehat{\bm \alpha}_n(0,s) - \widehat{\bm \alpha}_n(s,1)\big]^\prime \bm{\mathcal{N}}_{n,\bm \alpha}^{-1}(s)\big[\widehat{\bm \alpha}_n(0,s) - \widehat{\bm \alpha}_n(s,1)\big]\overset{d}{\longrightarrow} \mathcal{U}_{k+1}, \end{equation*} where the normalizer $\bm{\mathcal{N}}_{n,\bm \alpha}(s)$ is defined as $\bm{\mathcal{N}}_{n,\bm \gamma}(s)$ with $\widehat{\bm \gamma}_n(\cdot,\cdot)$ replaced by $\widehat{\bm \alpha}_n(\cdot,\cdot)$ at every occurrence.
proofAnalogous to the proof of Corollary (ref), but using Theorem (ref) in Appendix (ref) in place of Theorem (ref). Therefore, Corollary (ref) holds under a subset of the conditions of Corollary (ref).

Power Analysis

Here, we show that our test based on $\mathcal{U}_{n,\bm \gamma}$ has power against a one-break local alternative. To do so, consider the model (ref)--(ref) under \[ \mathcal{H}_1^{\operatorname{CoVaR}}\colon

pmatrix[pmatrix omitted — 48 chars of source]

=

pmatrix[pmatrix omitted — 90 chars of source]

,\qquad t=1,\ldots,n, \] where $\bm a(\cdot)$ and $\bm b(\cdot)$ are $\mathbb{R}^{k+1}$-valued, componentwise step functions on $[0,1]$. This local alternative is inspired by KPA88. Extensions to certain smooth $\bm a(\cdot)$ and $\bm b(\cdot)$ are also possible along the lines of KPA88 and Qu08. However, in view of later results on multiple discrete breaks (see Section (ref)) and to keep the exposition as simple as possible, we confine ourselves to step functions here.

thmSuppose $\mathcal{H}_1^{\operatorname{CoVaR}}$ holds true for the model (ref)--(ref). If Assumptions (ref)--(ref) are satisfied, then, as $n\to\infty$, \begin{multline*} (s-r)\begin{pmatrix}\bm D_n\big[\widehat{\bm \alpha}_n(r,s) - \bm \alpha_0\big]\\ \bm D_n\big[\widehat{\bm \beta}_n(r,s) - \bm \beta_0\big] \end{pmatrix} \overset{d}{\longrightarrow}\overline{\bm \varSigma}^{1/2}\big[\overline{\bm W}(s)-\overline{\bm W}(r)\big] +\begin{pmatrix}\int_{r}^{s}\bm a(x)\,\mathrm{d} x\\ \int_{r}^{s}\bm b(x)\,\mathrm{d} x\end{pmatrix}\quad in (\ell^{\infty}(\mathcal{D}_{\iota}))^{2k+2}, \end{multline*} where $\overline{\bm W}(\cdot)$ and $\overline{\bm \varSigma}$ are as in Theorem (ref).
proofSee Appendix (ref).

Our test statistic $\mathcal{U}_{n,\bm \gamma}$ is specifically designed to have power under one-break alternatives. Formally, we show that when the single break point $s^\ast$ lies in the interval $(\iota,1-\iota)$, then our test is consistent against certain local alternatives.

corSuppose $\mathcal{H}_1^{\operatorname{CoVaR}}$ holds true for the model (ref)--(ref), where $\bm c(x):=\big(\bm a^\prime(x),\bm b^\prime(x)\big)^\prime=\bm c_1\mathds{1}_{\{x\leq s^{\ast}\}} + \bm c_2\mathds{1}_{\{x>s^{\ast}\}}$ for some break point $s^\ast\in(\iota,1-\iota)$ and $\bm c_1,\bm c_2\in\mathbb{R}^{2k+2}$ with $\bm c_1\neq\bm c_2$. If Assumptions (ref)--(ref) are satisfied, then for any $\nu\in(0,1)$, \[ \lim_{\Vert\bm c_2-\bm c_1\Vert\to\infty}\lim_{n\to\infty}\mathbb{P}\big\{\mathcal{U}_{n,\bm \gamma}>\mathcal{U}_{2k+2,1-\nu}\big\}=1. \]
proofSee Appendix (ref).

We omit the corresponding local power result for the QR, because it follows analogously. Importantly, Corollary (ref) shows that local power is higher for near-stationary than for stationary predictors. This is because under stationarity the test is consistent against alternatives in a $n^{-1/2}$-neighborhood of the null, whereas for near-stationary variables consistency is obtained even in “more local” $n^{-(1+\kappa)/2}$-neighborhoods of the slope coefficients (see $\mathcal{H}_{1}^{\operatorname{CoVaR}}$). The intuition behind this result is that the signal strength in regressions with highly persistent variables is much larger, allowing the parameters to be estimated much more precisely, which---in turn---aids in detecting potential breaks.

Multiple Breaks

The $\mathcal{U}_{n,\bm \gamma}$ statistic is specifically designed for one-break alternatives. Therefore, in this section, we briefly propose a test statistic $\mathcal{V}_{n,\bm \gamma}$ tailored for multiple possible breaks along the lines of the “unsupervised” approach of ZL18. We derive the asymptotic limit of $\mathcal{V}_{n,\bm \gamma}$ under the null and show consistency under local alternatives with arbitrarily many breaks. Crucially, the number of break points does not have to be pre-specified ex ante in this test and, in this sense, it is “unsupervised”.

To introduce the test statistic, let $\bm G=\bm G(r,s)$ be some arbitrary function on $\mathcal{D}=\big\{(r,s)\in[0,1]^2\colon 0\leq r\leq s\leq 1\big\}$ and define

align*[align* omitted — 401 chars of source]

Then, our test statistic is defined as

multline[multline omitted — 394 chars of source]

where $\iota\in(0,1/2)$ and $\widehat{\bm S}_n(r,s):= (s-r)\widehat{\bm \gamma}_n(r,s)$. To gain some intuition, note that straightforward algebra implies that

equation[equation omitted — 189 chars of source]

such that the $\bm d$-terms in $\mathcal{V}_{n,\bm \gamma}$ compare subsample estimates, giving the test its power. In contrast, the “meat” matrices $\bm \varXi^{-1}$ in (ref) simply serve as (cleverly constructed) normalizers.

corSuppose $\mathcal{H}_0^{\operatorname{CoVaR}}$ holds true for the model (ref)--(ref). If Assumptions (ref)--(ref) are satisfied, then, as $n\to\infty$, \begin{multline*} \mathcal{V}_{n,\bm \gamma} \overset{d}{\longrightarrow}\sup_{(r_1,r_2)\in\mathcal{D}_{\iota}}\bm d^\prime(\Delta\overline{\bm W}, 0,r_1,r_2)\bm \varXi^{-1}(\Delta\overline{\bm W}, 0,r_1,r_2)\bm d(\Delta\overline{\bm W}, 0,r_1,r_2)\\ +\sup_{(s_1,s_2)\in\mathcal{D}_{\iota}}\bm d^\prime(\Delta\overline{\bm W}, s_1,s_2,1)\bm \varXi^{-1}(\Delta\overline{\bm W},s_1,s_2,1)\bm d(\Delta\overline{\bm W}, s_1,s_2,1)=:\mathcal{V}_{2k+2}, \end{multline*} where $\Delta\overline{\bm W}(r,s)=\overline{\bm W}(s)-\overline{\bm W}(r)$ for the $(2k+2)$-variate standard Brownian motion $\overline{\bm W}(\cdot)$.
proofSee Appendix (ref).

The limiting distribution $\mathcal{V}_{2k+2}$ coincides with $T^{\ast}(\mathbb{B})$ from Theorem 3.2 in ZL18.

Corollary (ref) shows consistency of our test even for the case of $M^{\ast}$ breaks.

corSuppose $\mathcal{H}_1^{\operatorname{CoVaR}}$ holds true for the model (ref)--(ref), where $\bm c(x):=\big(\bm a^\prime(x), \bm b^\prime(x)\big)=\bm c_i$ for $x\in(s_{i-1}^{\ast},s_{i}^{\ast}]$ are the break magnitudes of the $M^{\ast}\in\mathbb{N}$ break points occurring at times $0=s_{0}^{\ast}<s_{1}^{\ast}<\ldots<s_{M^{\ast}}^{\ast}<s_{M^{\ast}+1}^{\ast}=1$. Moreover, assume for the break magnitudes that $\bm c_{i}\neq\bm c_{i+1}$ for all $i=1,\ldots,M^{\ast}$ and the break times obey $\min_{0\leq i\leq M^{\ast}}|s_{i+1}^{\ast}-s_{i}^{\ast}|>\iota$. If Assumptions (ref)--(ref) are satisfied, then, for any $\nu\in(0,1)$, \[ \lim_{\min_{1\leq i\leq M^{\ast}}\Vert\bm c_{i+1}-\bm c_{i}\Vert\to\infty}\lim_{n\to\infty}\mathbb{P}\big\{\mathcal{V}_{n,\bm \gamma}>\mathcal{V}_{2k+2,1-\nu}\big\}=1, \] where $\mathcal{V}_{2k+2,1-\nu}$ denotes the $(1-\nu)$-quantile of $\mathcal{V}_{2k+2,1-\nu}$.
proofSee Appendix (ref).

Simulations

We provide simulation evidence on the finite-sample properties of our tests (size and power) in Appendix (ref). We draw the following main conclusions from the simulation study. First, the change point tests for the predictive CoVaR regressions have reasonable size even for samples as small as $n=1000$. Second, for extremely persistent predictors, the tests display some liberal tendencies. Yet, under near-stationarity, the size distortions vanish for large $n$, as predicted by our theory. Third, the power to identify structural breaks is higher for more persistent predictors---consistent with the higher local power for near-stationary covariates in Corollary (ref) (recall Section (ref)).

Stability of the VIX as a Systemic Risk Predictor

Up until the global financial crisis (GFC) of 2007--9, banking regulation focused almost exclusively on microprudential objectives; that is, it focused on limiting the risks of each financial institution in isolation. However, such individual regulations were insufficient to prevent the crisis. Thus, in the aftermath, attention has shifted towards macroprudential objectives, where---next to ensuring the financial viability of each bank in isolation---the goal is to improve the stability of the financial system as a whole. Achieving this requires a gauge of the interconnectedness of financial institutions, which is provided by systemic risk measures AB16,Aea17.

As a consequence, researchers have aimed for a better understanding of systemic risk. For instance, AB16 and Hea16 investigate the predictive content of volatility for future levels of systemic risk. However, whether such predictive relationships remain stable over time has not been investigated before due to a lack of appropriate statistical methodology, which this paper provides.

Particular interest attaches to the VIX as a predictor for future systemic riskiness, because of what BS14 term the “volatility paradox”. This phenomenon describes systemic risk as building up in times of low volatility in equities, only to materialize during crises. This occurs endogenously because financial institutions take excessive risk when volatility is low. Due to this tight relationship, the VIX as a forward-looking volatility measure can be seen as a natural predictor of systemic risk.

The VIX, whose value at time $t$ we denote by $\operatorname{VIX}_t$, measures the market participants' risk-neutral expectation of the variance in stock market returns on the S&P 500. It is computed by the Chicago Board Options Exchange (CBOE) based on option prices. Figure (ref) plots the VIX over our 10 year sample period from 2005--2014, which is chosen to include the GFC and the European sovereign debt crisis. The most marked spike in expected volatility occurs after the collapse of Lehman Brothers in late 2008 during the GFC. One can see clearly that the VIX took a long time to return to its pre-crisis level, suggesting a reasonably high persistence. Further spikes of the VIX, such as during the European sovereign debt crisis of the early 2010s, are also noticeable.

figure[figure omitted — 138 chars of source]

More formally, consistent with HZ12, we find that the VIX is somewhere on the borderline of stationarity. As in CWW15 and in line with Assumption (ref), we report the AR(1) coefficient estimate for the VIX, which is 0.9824. The KPSS test of Kea92 rejects the null of stationarity at any conventional significance level. However, the ADF--GLS test of ERS96 rejects the null of a unit root in the VIX, with a $p$-value well below $1\%$.\footnote{A similar result is obtained when applying the Dickey--Fuller-type test of CT09, which is robust to unconditional heteroskedasticity. Specifically, the wild bootstrap-based implementation of this test yields a $p$-value of 1.2%. In computing this result, we have adopted the standard settings implemented in the bootUR package of SW23.} In light of this conflicting evidence where both stationarity and non-stationarity are rejected, the exact degree of persistence in the VIX seems difficult to determine. Therefore, the persistence-robustness of our test becomes an empirically desirable feature when testing for changes next.

Specifically, we now investigate the stability of the VIX as a predictor of systemic risk in the US financial system. We do so for the “standard” CoVaR in Section (ref), where the conditioning is on the individual financial institution, and for the Exposure CoVaR in Section (ref), where we condition on the financial system.

“Standard” CoVaR

We now apply our test from Corollary (ref) to check the constancy of the predictive relationship between the VIX and future levels of systemic risk in the US banking sector. To do so, we use daily log-losses of the S&P 500 Financials (SPF) index as our $Z_t$, and the log-losses on a particular US bank as $Y_t$. We focus on those US banks that, at the end of our sample, were ranked as a global systemically important bank (G-SIB) by the Financial Stability Board FSB15. Then, $\operatorname{CoVaR}_{\beta\mid\alpha}\big((Z_t,Y_t)^\prime\mid\mathcal{F}_{t-1}\big)$ measures the impact that an extreme loss of the bank has on the financial sector, i.e., its riskiness to the system. We examine the structural stability of the following predictive model for the CoVaR:

align[align omitted — 344 chars of source]

Therefore, we investigate whether the predictive power of volatility for systemic risk is subject to change over time. To estimate the model, we use daily data from 2005 to 2014, all obtained from finance.yahoo.com, resulting in a sample of size $n=2515$.

Hea16 use the change in the VIX as a systemic risk predictor, citing concerns of non-stationarity of the VIX in levels. This practice was also followed in the subsequent literature Ned20. While this step may be necessary to achieve stationarity (as required by the theory developed in Hea16), going from levels to differences reduces the discriminatory power of subsequent tests, as is well-known from the predictive regression literature BD15a. Due to the persistence-robustness of our change point test, we can work directly with the levels of the VIX, which optimizes power; see the discussion below Corollary (ref).

table[table omitted — 1,291 chars of source]

For all US G-SIBs, Table (ref) shows realizations of the test statistic $\mathcal{U}_{n,\bm \gamma}$ for $\alpha=\beta\in\{0.9,\ 0.95\}$ to truly capture systemic risk. The test statistics for all banks exceed the 10%-critical value of 151.4 (see Table (ref)) at least once, such that there is evidence for structural breaks in the predictive content of the VIX of varying degree for each institution.

figure[figure omitted — 980 chars of source]

To provide additional insight on a specific rejection in Table (ref), we focus on Citigroup. Figure (ref) provides an overview for the case $\alpha=\beta=0.95$. Consistent with Table (ref), the 5%-critical value of our CoVaR test is exceeded at some point in time by the function $s\mapsto s^2(1-s)^2\big[\widehat{\bm \gamma}_n(0,s) - \widehat{\bm \gamma}_n(s,1)\big]^\prime \bm{\mathcal{N}}_{n,\bm \gamma}^{-1}(s)\big[\widehat{\bm \gamma}_n(0,s) - \widehat{\bm \gamma}_n(s,1)\big]$, which forms the basis of our test statistic $\mathcal{U}_{n,\bm \gamma}$. However, the top panel allows us to see more clearly the period responsible for the break. It is obvious that the evidence against a stable predictive relationship is strongest during the global financial crisis, with the maximum value of the test statistic attained on October 15th, 2007.

In principle, this break could be due to an instability in the quantile model (ref) or the CoVaR model (ref). However, when plotting the analogous function $s\mapsto s^2(1-s)^2\big[\widehat{\bm \alpha}_n(0,s) - \widehat{\bm \alpha}_n(s,1)\big]^\prime \bm{\mathcal{N}}_{n,\bm \alpha}^{-1}(s)\big[\widehat{\bm \alpha}_n(0,s) - \widehat{\bm \alpha}_n(s,1)\big]$ for only the QR in (ref), we see that it solidly stays below the 5%-critical value (see the middle panel of Figure (ref)). This points to an instability exclusively in the predictive relationship for future systemic risk in (ref).

The bottom panel of Figure (ref) sheds additional light on how the predictive content of the VIX varies over time for systemic risk. It does so by plotting rolling-window estimates of the slope coefficient $\beta_{1,t}$ of the VIX in the CoVaR regression (ref), where each window consists of 500 observations. During the GFC we observe elevated levels of predictability by the VIX, which subsequently declines (but stays positive). Therefore, as expected, during times where systemic risk materializes, the VIX as a “fear index” provides a good indication of high future systemic risk in the financial sector. Yet even during times of lower predictability, the VIX and systemic risk remain tightly positively linked, consistent with the notion of AB16 that “low volatility environments breed systemic risk”.

Exposure CoVaR

Our above definition of the CoVaR corresponds to the standard definition of AB16, where the system (here: S&P 500 Financials) is considered conditional on some institution (here: US G-SIB) being in distress. However, as pointed out by AB16, the conditioning may be reversed to obtain the Exposure CoVaR, which measures the risk of an institution given system failure. The two CoVaR definitions give complementary information. For instance, a small, but highly connected, bank may have a large Exposure CoVaR, due to its many links with other institutions. However, its “standard” CoVaR---measuring its impact on the system---would be low because of its small size, that would allow other banks to quickly pick up its positions in case of failure. Since the Exposure CoVaR is also of interest in its own right, we repeat part of the above analysis now for the Exposure CoVaR.

Specifically, we investigate the structural stability of the Exposure CoVaR regression

align[align omitted — 346 chars of source]

which essentially corresponds to (ref)--(ref). However, here we interchange the roles of $Y_t$ and $Z_t$, such that $Y_t$ now denotes the S&P 500 Financials log-losses and $Z_t$ the log-losses on one of the eight US G-SIBs. Doing so allows us to study the Exposure CoVaR for each G-SIB.

table[table omitted — 1,014 chars of source]

Table (ref), which is the analog of Table (ref), shows the results of our stability tests for the Exposure CoVaR regression in (ref)--(ref). Here, we find little evidence of instability in the predictive relationship. Only for three banks (for $\alpha=\beta=0.95$) do we find some indications for change at the 10%-level. However, with 16 tests being carried out at the 10%-level, finding three rejections is not uncommon even when the null holds true in all cases. Overall, full-sample estimates of the Exposure CoVaR regression in (ref)--(ref) seem credible.

To obtain these full-sample estimates, we only have to run the single quantile regression for the S&P 500 Financials in (ref), which gives \[ \widehat{Q}_{\alpha}(Y_t\mid\mathcal{F}_{t-1}) = -0.021 + 0.0024\cdot \operatorname{VIX}_{t-1}. \] The positive sign of the estimated slope coefficient $\alpha_1$ indicates that a high expectation of future volatility (as measured by the VIX) is indicative of higher levels of risk for the US financial system. We display the slope coefficient estimates $\widehat{\beta}_1$ of the appertaining Exposure CoVaR regression (ref) in Table (ref). The VIX has a strong positive influence on the future exposure of banks to trouble in the financial system, as the positive and (mostly) significant estimates of $\beta_1$ suggest.

The $p$-values in Table (ref) (displayed in parentheses below the estimates) are calculated based on the functional convergence in Theorem (ref). To see precisely how, write $\widehat{\bm \beta}_{n}(0,s)=\big(\widehat{\beta}_{0,n}(0,s),\ldots,\widehat{\beta}_{k,n}(0,s)\big)^\prime$ and $\bm \beta_0=(\beta_{0},\ldots,\beta_{k})^\prime$, and define the self-normalizer $\mathcal{S}_{n,\beta_i}=\int_{\iota}^{1}s^2\big[\widehat{\beta}_{i,n}(0,s) - \widehat{\beta}_{i,n}(0,1)\big]^2\,\mathrm{d} s$. Then, following ideas from Sha10, Theorem (ref) and the continuous mapping theorem imply under $\mathcal{H}_0^{\operatorname{CoVaR}}$ and $\beta_{i}=0$ that, as $n\to\infty$, \[ \mathcal{T}_n:=n\widehat{\beta}_{i,n}^2(0,1)/\mathcal{S}_{n,\beta_i}\overset{d}{\longrightarrow}W^2(1)/\int_{\iota}^{1}\big[W(s)-sW(1)\big]^2\,\mathrm{d} s=:\mathcal{T}, \] where $W(\cdot)$ denotes a standard Brownian motion. (A similar result also holds for the QR coefficients.) Then, if $\mathcal{T}_n$ exceeds the $(1-\nu)$-quantile of $\mathcal{T}$, we conclude that the $i$-th CoVaR coefficient is significantly different from zero at level $\nu$. Importantly, this “off-the-shelf” inference method afforded by our functional convergence result remains valid regardless of whether the predictors are stationary or near-stationary. Therefore, the $p$-values of Table (ref) are persistence-robust in the same sense as our structural break tests.

table[table omitted — 1,056 chars of source]

Overall, this section's results suggest that the forecasting power of the VIX for risk and exposure systemic risk is rather stable. Moreover, the VIX is a statistically significant covariate, with high expected volatility predicting higher levels of risk and systemic risk in the US financial system. We stress that this significance result for the VIX is robust to whether or not the VIX is stationary or near-stationary.

Comparing the results of Sections (ref) and (ref), we find that the VIX has somewhat instable forecasting power for predicting the riskiness of individual banks to the system, yet is a rather stable predictor of the riskiness the financial system poses for a single bank. This suggests that the channel linking individual institutions to the system is more prone to changes than that connecting the system with a specific bank. In turn, this could be due to idiosyncratic effects on the level of the bank, such as---in the most extreme case---the collapse of an entire institution (e.g., Lehman Brothers) that destabilizes the financial system. During its demise, the risk to the system is significantly increased relative to the pre-collapse time, leading naturally to time-variation in the bank's risk to the system.

Conclusion

This paper is the first to propose structural break tests for CoVaR regressions. Importantly, our tests are valid irrespective of whether predictors are stationary or near-stationary. In his review of self-normalization, Sha15 writes on the use of SN that “when the dependence in the time series is too strong (say, near-integrated time series), inference of certain parameter becomes difficult because the information aggregated over time does not accumulate quickly due to strong dependence.” We show in this paper that when the dependence in the predictors is not too strong (at most near-stationary), then self-normalized change point tests still work in the sense of providing unified inference on possible breaks in the predictive relationship. Crucially, no pre-tests for the serial dependence properties of the predictors are required, which would otherwise impair the validity of the subsequent break test. As a corollary, we also obtain change point tests for QRs, where the results are novel for near-stationary predictors.

An empirical application highlights the importance of persistence-robust tests, where the VIX offers an example of a predictor somewhere on the boundary between stationarity and non-stationarity. Our test can be applied validly (in the sense of holding size) regardless of whether the VIX is stationary or near-stationary. We find the predictive relationship between volatility and future systemic risk in the US financial system to be stable or instable depending on the direction of conditioning in the CoVaR. We also provide an application of our persistence-robust QR stability test to some putative equity premium predictors (Appendix (ref)) that are also borderline stationary. Our test reveals that there are significant instabilities in quantile prediction models for the US equity premium.

Future work could consider robustifying inference on change points also with respect to (I1) predictors. In this case, we no longer expect SN to work, but the IVX approach of MP09 and PM09 may be a feasible option. Indeed, IVX has been used successfully in developing unified inference in predictive QR Lee16. One drawback, however, is that IVX leads to slower rates of convergence for the (NS) case and, thus, possibly lower power. Investigations such as these are left for future research.