EconBase
← Back to paper

Change-Point Testing for Risk Measures in Time Series

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

83,677 characters · 17 sections · 50 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Change-Point Testing for Risk Measures in Time Series

\onehalfspacing

titlepage\thispagestyle{empty} \begin{abstract} We propose novel methods for change-point testing for nonparametric estimators of expected shortfall and related risk measures in weakly dependent time series. We can detect general multiple structural changes in the tails of marginal distributions of time series under general assumptions. Self-normalization allows us to avoid the issues of standard error estimation. The theoretical foundations for our methods are functional central limit theorems, which we develop under weak assumptions. An empirical study of S&P 500 and US Treasury bond returns illustrates the practical use of our methods in detecting and quantifying instability in the tails of financial time series. \noindentKeywords: Time Series, Risk Measure, Change-Point Test, Confidence Interval, Self-Normalization, Sectioning, Expected Shortfall, Unsupervised Change Point Detection \noindentJEL classification: C14, C58, G32 \end{abstract}

Introduction

The quantification of risk is a central topic of study in finance. The need to guard against unforeseeable events with adverse and potentially catastrophic consequences has led to an extensive and burgeoning literature on risk measures. The two most popular financial risk measures, the value-at-risk (VaR) and the expected shortfall (ES), are extensively used for risk and portfolio management as well as regulation in the finance industry mcneil2015quantitative. Parallel to the developments in risk estimation, there has been longstanding interest in detecting changes in financial time series, especially changes in the tail structure, which is essential for effective risk and portfolio management. Indeed, empirical findings strongly suggest that financial time series frequently exhibit changes in their underlying statistical structure due to changes in economic conditions, e.g. monetary policies, or critical social events. Although there is an established literature on structural change detection for parametric time series models, including monitoring of proxies for risk such as tail index, there is little work concerning monitoring of general tail structure, or of risk measures in particular. To underscore this point, the existing literature on risk measure estimation assumes stationarity of time series observations over a time period of interest, with the stationarity of the risk measure being key for the estimation to make sense. However, it is a challenge to verify this assumption.

We provide tools to detect general and potentially multiple changes in the tail structure of time series, and in particular, tools for monitoring for changes in risk measures such as ES and related measures such as conditional tail moments (CTM) methni_etal2014 over time periods of interest. Specifically, we develop retrospective change-point tests to detect changes in ES and related risk measures. Additionally, we offer new ways of constructing confidence intervals for these risk measures. Our methods are applicable to a wide variety of time series models as they depend on functional central limit theorems for the risk measures, which we develop under weak assumptions. As described below, our methods complement and extend the existing literature in several ways.

As mentioned previously, the literature lacks tools to monitor general tail structure or risk measures such as ES or CTM over time. This deficiency appears to be two-fold. (1) Although there are studies on VaR change-point testing, for example, qu2008, often one is interested in characterizations of tail structure that are more informative than simple location measures. Indeed, the introduction of ES as an alternative risk measure to VaR was, to a great extent, driven by the need to quantify tail structure, particularly, the expected magnitude of losses conditional on losses being in the tail. Aligned with this goal, a popular measure of tail structure is the tail index, which describes tail thickness and governs distributional moments. In tail index estimation, an extreme value theory approach is typically taken with the assumption of so-called regularly-varying Pareto-type tails. However, tail index estimation is very sensitive to the choice of which fraction of sample observations is classified to be “in the tail”.\footnote{The Hill estimator hill1975 is widely used and requires the user to choose the fraction of sample observations deemed to be “in the tail” to use for estimation. However, generally, there appears to be no best way to select this fraction. In the specific setting of change-point testing for tail index, recommendations for this fraction in the literature range from the top 20th percentile to the top 5th percentile of observations kim_etal2009,kim_etal2011,hoga2017. Such choice heavily influences the quality of tail index estimation and change detection and must be made on a case-by-case basis rocco2014. It is a delicate matter as choosing too small of a fraction results in high estimation variance and choosing too large of a fraction often results in high bias due to misspecification of where the tail begins.} Moreover, the typical regularly-varying Pareto-type tail assumptions may not even be valid in some situations. And even if they are valid, the tail index is invariant to changes of the location and scale types, and thus these types of structural changes in the tail would remain undetected using tail index-based change-point tests. (2) As mentioned before, all previous studies on risk measure estimation implicitly assume the risk measure is constant over some time period of interest---otherwise, risk measure estimation and confidence interval construction could behave erratically. For instance, if there is a sudden change in VaR (at some level) in the middle of a time series, naively estimating VaR using the entire time series could result in a wrong estimate of VaR or ES. Hence, given the importance of ES and related risk measures, a simple test for ES change over a time period of interest is a useful first step prior to follow-up statistical analysis.

To simultaneously address both deficiencies, we propose retrospective change-point testing for ES and related risk measures such as CTM. We introduce, in particular, a consistent test for a potential ES change at an unknown time based on a variant of the widely used cumulative sum (CUSUM) process.\footnote{The cumulative sum (CUSUM) process was first introduced by page1954 and is discussed in detail in csorgo_etal2011.} We subsequently generalize this test to the case of multiple potential ES change points, leveraging recent work by zhang_etal2018, and unlike existing change-point testing methodologies in the literature, our test does not require the number of potential change points to be specified in advance. Our use of a risk measure such as ES is attractive in many ways. First, the fraction of observations used in ES estimation, for example, the upper 5th percentile of observations, directly has meaning, and is often comparable to the fraction of observations used in tail index estimation, as discussed previously. Second, the use of ES does not require parametric-type tail assumptions, which is not the case with use of the tail index. Third, ES can detect much more general structural changes in the tail such as location or scale changes, while such changes go undetected when using tail index. Moreover, our change-point tests can be used to check in a statistically principled way whether or not ES and related risk measures are constant over a time period of interest, and provide additional validity when applying existing estimation methods for these risk measures.

A key detail that has been largely ignored in previous studies is standard error estimation. Here, the issue is two-fold. (1) In the construction of confidence intervals for risk measures, consistent estimation of standard errors is nontrivial due to standard errors involving the entire time series correlation structure at all integer lags, and thus being infinite-dimensional in nature. A few studies in which standard error estimation has been addressed (or bypassed) are chen_etal2005,chen2008,wang_etal2008,xu2016. However, confidence interval construction in these studies all require user-specified tuning parameters such as the bandwidth in periodogram kernel smoothing or window width in the blockwise bootstrap and empirical likelihood methods. Although the choice of these tuning parameters significantly influences the quality of the resulting confidence intervals, it is not always clear how to best to select them. (2) Choosing the tuning parameters is not only difficult, but data-driven approaches may result in non-monotonic test power, as pointed out in numerous studies vogelsang1999,crainiceanu_etal2007,deng_etal2008,shao_etal2010.\footnote{For general change-point tests with time series observations, consistent estimation of standard errors is typically done using periodogram kernel smoothing, and the performance of such tests is heavily influenced by the choice of kernel bandwidth.}

We offer the following solutions to the above issues. (1) To address the issue of often problematic selection of tuning parameters for confidence interval construction with time series data, we make use of ratio statistics to cancel out unknown standard errors and form pivotal quantities, thereby avoiding the often difficult estimation of such nuisance parameters. We examine confidence interval construction using a technique originally referred to in the simulation literature as sectioning asmussen_etal2007, which involves splitting the data into equal-size non-overlapping sections, separately evaluating the estimator of interest using the data in each section, and relying on a normal approximation to form an asymptotically pivotal t-statistic. We also examine its generalization, referred to in the simulation literature as standardized time series schruben1983,glynn_etal1990 and in the time series literature as self-normalization lobato2001,shao2010, which uses functionals different from the t-statistic to create asymptotically pivotal quantities. (2) In the context of change-point testing using ES and related risk measures, to avoid potentially troublesome standard error estimation, we follow shao_etal2010 and zhang_etal2018 and apply the method of self-normalization for change-point testing by dividing CUSUM-type processes by corresponding processes designed to cancel out the unknown standard error. The processes we divide by take into account potential change point(s) and thus avoids the problem of non-monotonic power which often plagues change-point tests that rely on consistent standard error estimation, as discussed previously.

The outline of our paper is as follows. We describe the setting and assumptions in Section (ref). In Section (ref), we develop the asymptotic theory for VaR and ES, specifically, functional central limit theorems, which provide the theoretical basis for the proposed confidence interval construction and change-point testing methodologies. In introducing our statistical methods, we first discuss the simpler task of confidence interval construction for risk measures in Section (ref). Then, with several fundamental ideas in place, we take up testing for a single potential change point in VaR and ES in Section (ref). We extend the change-point testing methodology to an unknown, possibly multiple, number of change points in Section (ref). In Section (ref), we demonstrate the good finite-sample performance of our proposed methods through simulations. We conclude with an empirical study in Section (ref) of returns data for the S&P 500 index and US Treasury bond returns. The proofs of our theoretical results as well as additional simulation and empirical results can be found in the Appendix.

Asymptotic Theory

Model Setup

Suppose the random variable $X$ and the stationary sequence of random variables $X_1, \dots, X_n$ have the marginal distribution function $F$. For some level $p \in (0,1)$, we want to estimate the risk measures VaR (value-at-risk) and ES (expected shortfall) defined by:\footnote{It is reasonable to define VaR and ES as upper tail characteristics when considering a loss process.}

eqnarray*[eqnarray* omitted — 99 chars of source]

Let $\widehat{F}_n(\cdot) = n^{-1} \sum_{i=1}^n \indic{X_i \le \cdot}$ be the sample distribution function. We consider the following plug-in sample-based nonparametric estimators:

eqnarray[eqnarray omitted — 206 chars of source]

For the change-point testing, we estimate these risk measures on subsets of the data. Hence, for $m > l$, we will also consider the following nonparametric estimators based on samples $X_l,\dots, X_m$, with $\widehat{F}_{l:m}(\cdot) = (m-l+1)^{-1} \sum_{i=l}^m \indic{X_i \le \cdot}$:

eqnarray[eqnarray omitted — 226 chars of source]

We derive the asymptotic theory under general assumptions for the stochastic process. Let $\mathcal{F}_l^m$ denote the $\sigma$-algebra generated by $X_l, \dots, X_m$, and let $\mathcal{F}_l^\infty$ denote the $\sigma$-algebra generated by $X_l, X_{l+1}, \dots$. The $\alpha$-mixing coefficient, as introduced by rosenblatt1956, is defined as $$\alpha(k) = \sup_{A \in \mathcal{F}_1^j, \, B \in \mathcal{F}_{j+k}^\infty, \, j \ge 1} \mathop\mathrm{abs}{\Parg{A}\Parg{B} - \Parg{A,B}},$$ and a sequence is said to be $\alpha$-mixing if $\lim_{k \to \infty} \alpha(k) = 0$. The dependence described by $\alpha$-mixing is the least restrictive as it is implied by the other types of mixing; see doukhan1994 for a comprehensive discussion. In what follows, $D[0,1]$ denotes the space of real-valued functions on $[0,1]$ that are right-continuous and have left limits, and convergence in distribution on this space is defined with respect to the Skorohod topology billingsley1999. Also, for some index set $\mathcal{I}$, $\ell^\infty(\mathcal{I})$ denotes the space of real-valued bounded functions on $\mathcal{I}$, and convergence in distribution on this space is defined with respect to the uniform topology. We denote the integer part of a real number $x$ by $[x]$ and the positive part by $[x]_+$. A standard Brownian motion on the real line is denoted by $W$. Our theoretical results are developed using the following assumption.

assumptionFor some constants $\gamma > 0$ and $a > \max(3,(2+\gamma)/\gamma)$: \begin{enumerate} • $X_1, X_2, \dots$ is $\alpha$-mixing with $\alpha(k) = O(k^{-a})$$\Earg{\mathop\mathrm{abs}{X}^{2+\gamma}} < \infty$$X$ has a positive and differentiable density $f$ in a neighborhood of $VaR$, and for each $k \ge 2$, $(X_1,X_k)$ has joint density in a neighborhood of $(VaR,VaR)$. \end{enumerate}

Assumption (ref) condition (i) is a form of asymptotic independence, which ensures that the time series is not too serially dependent so that non-degenerate limit distributions are possible. A very wide range of commonly used financial time series models such as GARCH models and diffusion models satisfy this condition. Condition (ii) indicates a trade-off between the strength of the moment condition of the underlying marginal distribution and the rates of $\alpha$-mixing, with weaker $\alpha$-mixing conditions requiring stronger moment conditions and vice versa. Condition (iii) is a standard assumption in VaR and ES estimation, and rules out pathological situations in which there are many ties among the observations near $VaR$.

We point out, in particular, that while our results are illustrated for the most widely used risk measures, VaR and ES, they can be adapted to many other important functionals of the underlying marginal distribution. One straightforward adaptation (by assuming stronger $\alpha$-mixing and moment conditions) is to CTM: $\Earg{X^\beta \mid X > VaR}$ for some level $p \in (0,1)$ and some $\beta > 0$ methni_etal2014. Our results also easily extend to multivariate time series, but for simplicity of illustration, we focus on univariate time series.

Functional Central Limit Theorems

We develop functional central limit theorems jointly for VaR and ES under weak assumptions. These functional central limit theorems allow the construction of general change point statistics.

thmUnder Assumption (ref), the process \begin{align} \left\{ \sqrt{n} t\begin{bmatrix}\widehat{VaR}_{[nt]} - VaR \\ \widehat{ES}_{[nt]} - ES \end{bmatrix} : t \in [0,1]\right\} \end{align} converges in distribution in $D[0,1] \times D[0,1]$ to $\Sigma^{1/2} (W_1,W_2)$, where $W_1$ and $W_2$ are independent standard Brownian motions, and $\Sigma \in \mathbb{R}^{2 \times 2}$ is symmetric positive-semidefinite with components \begin{align} & \Sigma_{11} = \frac{1}{f^2(VaR)}\left(\Earg{g^2(X_1)} + 2 \sum_{i=2}^\infty \Earg{g(X_1)g(X_i)}\right) \\ & \Sigma_{12} = \frac{1}{f(VaR)(1-p)}\left(\Earg{g(X_1)h(X_1)} + \sum_{i=2}^\infty \Bigl(\Earg{g(X_1)h(X_i)} + \Earg{h(X_1)g(X_i)}\Bigr)\right) \\ & \Sigma_{22} = \frac{1}{(1-p)^2}\left(\Earg{h^2(X_1)} + 2 \sum_{i=2}^\infty \Earg{h(X_1)h(X_i)}\right), \end{align} where $g(X) = \indic{X \le VaR} - p$ and $h(X) = [X-VaR]_+ - \Earg{[X-VaR]_+}$.

Our Assumption $\ref{A3}$ conditions (i)-(ii) are weaker than those of existing central limit theorems in the literature (c.f. chen_etal2005,chen2008), which require $\alpha$-mixing coefficients to decay exponentially fast, as well as additional regularity of the marginal and pairwise joint densities of the observations.

In deriving the ES functional central limit theorem, we used the following Bahadur representation for $\widehat{ES}_n$.

propSuppose Assumption (ref) holds for some $\gamma > 0$ and $a > (2+\gamma)/\gamma$. Then, for any $\gamma' > 0$ satisfying $-1/2 + 1/(2a) + \gamma' < 0$, $$\widehat{ES}_n - \biggl(VaR + \frac{1}{1-p} \frac{1}{n} \sum_{i=1}^n [X_i - VaR]_+ \biggr) = o_{a.s.}(n^{-1 + 1/(2a) + \gamma'} \log n).$$

sun_etal2010 developed such a Bahadur representation in the setting of independent, identically-distributed data, but to the best of our knowledge, no such representation exists in the stationary, $\alpha$-mixing setting. Such a Bahadur representation is generally useful for developing limit theorems in many different contexts.

We also have the following extension of the functional central limit theorems for VaR and ES in Theorem (ref), where now there are two “time” indices instead of just one. The standard functional central limit theorems are useful for detecting a single change point in a time series, but the following extension will allow us to detect an unknown, possibly multiple, number of change points, as we will discuss later. We point out that this result does not follow automatically from Theorem (ref) and an application of the continuous mapping theorem, because the estimators in ((ref))-((ref)) are not additive, for instance, for $m > l$, $\widehat{ES}_{l:m} \ne \widehat{ES}_{1:m} - \widehat{ES}_{1:l-1}$.

thmFix any $\delta > 0$ and consider the index set $\Delta = \{ (s,t) \in [0,1]^2 : t-s \ge \delta \}$. Under Assumption (ref), with the modification that $(X_1, X_k)$ has a joint density for all $k \ge 2$, the process \begin{align} \left\{ \sqrt{n} (t-s)\begin{bmatrix}\widehat{VaR}_{[ns]:[nt]} - VaR \\ \widehat{ES}_{[ns]:[nt]} - ES \end{bmatrix} : (s,t) \in \Delta \right\} \end{align} converges in distribution in $\ell^\infty(\Delta) \times \ell^\infty(\Delta)$ to the process \begin{align} \left\{\Sigma^{1/2} \begin{bmatrix} W_1(t) - W_1(s) \\ W_2(t) - W_2(s) \end{bmatrix} : (s,t) \in \Delta \right\}, \end{align} where $\Sigma$ is from Theorem (ref), and $W_1$ and $W_2$ are independent standard Brownian motions.

Statistical Methods

Confidence Intervals

In time series analysis, confidence interval construction for an unknown quantity is often difficult, due to dependence. Indeed, from Theorem (ref), we see that the standard errors appearing in the normal limiting distributions depend on the time series autocovariance at all integer lags. To construct confidence intervals using Theorem (ref), these standard errors must be estimated. One approach, taken in chen_etal2005,chen2008, is to estimate using kernel smoothing the spectral density at zero frequency of the transformed time series $g(X_1),g(X_2),\dots$ and $h(X_1),h(X_2),\dots$, where $g$ and $h$ are from Theorem (ref). Although it is known that under certain moment and correlation assumptions, spectral density estimators are consistent for stationary processes brockwell_etal1991,anderson1971, in practice it is nontrivial to obtain robust and precise estimates due to the need to select tuning parameters for the kernel smoothing-based approach. As for other approaches, under certain conditions, resampling methods such as the moving block bootstrap kunsch1989,liu_etal1992 and the subsampling method for time series politis_etal1999 bypass direct standard error estimation and yield confidence intervals that asymptotically have the correct coverage probability. However, these approaches also require user-chosen tuning parameters such as the block length in the moving block bootstrap and the window width in subsampling. We introduce an alternative way to construct confidence intervals with asymptotically correct coverage, where the quality of the resulting confidence intervals is less sensitive to the choice of tuning parameters and is therefore more robust.

We first examine a technique called sectioning from the simulation literature asmussen_etal2007, which can be used to construct confidence intervals for general estimators. The method is as follows. Let $Y_1(\cdot), Y_2(\cdot), \dots$ be a sequence of random bounded real-valued functions on $[0,1]$. For some user-specified integer $m \ge 2$, suppose we have the joint convergence in distribution:

align[align omitted — 215 chars of source]

where $\sigma > 0$ and $\mathcal{N}_1, \dots, \mathcal{N}_m$ are independent standard normal random variables. Taking $Y_n(\cdot)$ to be the processes in ((ref)) for VaR or ES, this is guaranteed by Theorem (ref). Then, with

align*[align* omitted — 179 chars of source]

by the Continuous Mapping Theorem, as $n \to \infty$ with $m$ fixed, $m^{1/2}\overline{Y}_n/S_n$ converges in distribution to the Student's $t$-distribution with $m-1$ degrees of freedom. So using the $t$-distribution critical values, we can obtain confidence intervals for VaR and ES with asymptotically correct coverage. Note that this method requires the number $n/m$ to be sufficiently large for the asymptotic distribution in ((ref)) to hold. Following asmussen_etal2007, we recommend selecting the integer $m$ within the range of $[2,10]$ in practice. In Section (ref), we demonstrate through simulations that the sectioning method is very robust to choices of $m$ within this range.

Next, we examine a generalization of sectioning, called self-normalization. This technique has been studied recently in the time series literature lobato2001,shao2010, as well as earlier in the simulation literature, where it is known as standardization schruben1983,glynn_etal1990. As with sectioning, the method applies to confidence interval construction for general estimators. With self-normalization, the idea is to use a ratio-type statistic where the unknown standard error appears in both the numerator and the denominator and thus cancels, resulting in a pivotal limiting distribution. As before, let $Y_1(\cdot), Y_2(\cdot), \dots$ be a sequence of random bounded real-valued functions on $[0,1]$. Suppose we have the distributional convergence in $D[0,1]$: $Y_n(\cdot) \overset{d}{\to} \sigma W(\cdot)$, where $\sigma > 0$. Again, taking $Y_n(\cdot)$ to be the processes in ((ref)) for VaR or ES, this distributional convergence is guaranteed by Theorem (ref). Consider some positive homogeneous functional $T : D[0,1] \mapsto \mathbb{R}$, i.e., satisfying $T(\sigma Y) = \sigma T(Y)$ for $\sigma > 0$ and $Y \in D[0,1]$. If the Continuous Mapping Theorem can applied, then

eqnarray[eqnarray omitted — 84 chars of source]

The right side of ((ref)) is a pivotal limiting distribution. So using its critical values, which may be computed via simulation, we can obtain confidence intervals for VaR and ES with asymptotically correct coverage.

The choice of form for the functional $T$ is up to the user. One convenient choice is $T(Y) = \left(\int_0^1(Y(t) - tY(1))^2 dt\right)^{1/2}$, using which we obtain the following result for ES:

eqnarray[eqnarray omitted — 222 chars of source]

This form for $T$, as advocated for by lobato2001, shao2010 and shao_etal2010 among others, ensures that the numerator and denominator of the right side of ((ref)) are quantities based on independent normal random variables, which makes quantiles for their ratio particularly easy to compute.

Testing for a Single Change Point

As is the case with confidence interval construction with time series data, change-point testing in time series based on statistics constructed from functional central limit theorems and the continuous mapping theorem is often nontrivial due to the need to estimate standard errors. Motivated by the maximum likelihood method in the parametric setting, variants of the page1954 CUSUM statistic are commonly used for nonparametric change-point tests csorgo_etal2011, and generally rely on asymptotic approximations (via functional central limit theorems and the continuous mapping theorem) to supply critical values of pivotal limiting distributions under the null hypothesis of no change. As discussed in vogelsang1999,shao_etal2010,zhang_etal2018, testing procedures where standard errors are estimated directly, for example, by estimating the spectral density of transformed time series via a kernel-smoothing approach, can be biased under the change-point alternative. Such bias can result in nonmonotonic power, i.e., power can decrease in some ranges as the alternative deviates from the null. To avoid this issue, shao_etal2010 and zhang_etal2018 propose using self-normalization techniques to general change-point testing. We adopt this idea to our specific problem of detecting changes in tail risk measures.

As motivated in the Introduction (Section (ref)), it is important to perform hypothesis tests for abrupt changes of risk measures in the time series setting. We introduce the methodology for joint testing of VaR and ES. In this section, we consider the case of at most one change point. We consider the case of an unknown number of change points in the next section, using the approach in zhang_etal2018. For a time series $X_1, \dots, X_n$, let $(VaR_{X_i}, ES_{X_i})$ be the VaR and ES at a fixed level $p$ for the marginal distribution of $X_i$. We test the following null and alternative hypotheses.

align*[align* omitted — 674 chars of source]

We base our change-point test on the following variant of the CUSUM process:

align[align omitted — 201 chars of source]

Note that we split the above process for all possible break points $t \in (0,1)$ into a difference of two estimators, an estimator using $X_1, \dots, X_{[nt]}$ and an estimator using $X_{[nt]+1}, \dots, X_n$. Moreover, splitting the process as in ((ref)) avoids potential VaR estimation using a sequence containing the change point, which could have undesirable behavior. In Proposition (ref) below, we self-normalize the process in ((ref)) using the approach of shao_etal2010. Note that the self-normalizer of ((ref)) below takes into account the potential change point and is split into two separate integrals involving $X_1, \dots, X_{[nt]}$ and $X_{[nt]+1}, \dots, X_n$.

propSuppose Assumption (ref) holds. Under the null hypothesis $\mathcal{H}_0$, \begin{align} G_n := \sup_{t \in [0,1]} C_n(t)^T D_n(t)^{-1} C_n(t), \end{align} with \begin{align*} & C_n(t) := t(1-t) \begin{bmatrix} \widehat{VaR}_{1:[nt]} - \widehat{VaR}_{[nt]+1:n} \\ \widehat{ES}_{1:[nt]} - \widehat{ES}_{[nt]+1:n} \end{bmatrix} \\ & D_n(t) := n^{-1}\sum_{i=1}^{[nt]} \left(\frac{i}{n}\right)^2 \begin{bmatrix} \widehat{VaR}_{1:i} - \widehat{VaR}_{1:[nt]} \\ \widehat{ES}_{1:i} - \widehat{ES}_{1:[nt]} \end{bmatrix}^{\otimes 2} + n^{-1}\sum_{i=[nt]+1}^n \left(\frac{n-i+1}{n}\right)^2 \begin{bmatrix} \widehat{VaR}_{i:n} - \widehat{VaR}_{[nt]+1:n} \\ \widehat{ES}_{i:n} - \widehat{ES}_{[nt]+1:n} \end{bmatrix}^{\otimes 2}, \end{align*} converges in distribution to \begin{align} G := \sup_{t \in [0,1]} C(t)^T D(t)^{-1} C(t), \end{align} with \begin{align*} & C(t) := \begin{bmatrix} W_1(t) - tW_1(1) \\ W_2(t) - tW_2(1) \end{bmatrix} \\ & D(t) := \int_0^t \begin{bmatrix} W_1(s) - \frac{s}{t}W_1(t) \\ W_2(s) - \frac{s}{t}W_2(t) \end{bmatrix}^{\otimes 2} ds + \int_t^1 \begin{bmatrix} W_1(1) - W_1(s) - \frac{1-s}{1-t}(W_1(1) - W_1(t)) \\ W_2(1) - W_2(s) - \frac{1-s}{1-t}(W_2(1) - W_2(t)) \end{bmatrix}^{\otimes 2} ds. \end{align*} Assume the alternative hypothesis $\mathcal{H}_1$ is true with the change point occurring at some fixed (but unknown) $t^* \in (0,1)$. For any fixed difference $VaR_{X_{[nt^*]}} = c_1 \ne c_2 = VaR_{X_{[nt^*]+1}}$ or $ES_{X_{[nt^*]}} = d_1 \ne d_2 = ES_{X_{[nt^*]+1}}$, we have $G_n \overset{P}{\to} \infty$ as $n \to \infty$. Furthermore, if the difference varies with $n$ according to $c_1 - c_2 = n^{-1/2 + \epsilon} L$ or $d_1 - d_2 = n^{-1/2 + \epsilon} L$ for some $L \ne 0$ and $\epsilon \in (0,1/2)$, then $G_n \overset{P}{\to} \infty$ as $n \to \infty$.

The distribution of $G$ is pivotal, and its critical values may be obtained via simulation. For testing $\mathcal{H}_0$ versus $\mathcal{H}_1$ at some level, we reject $\mathcal{H}_0$ if the test statistic $G_n$ exceeds some corresponding critical value of $G$. To obtain critical values, we simulate 5,000 replications, with each replication consisting of 2,000 independent standard normal random variables to approximate standard Brownian motions on $[0,1]$.

In subsequent discussions concerning ((ref)), we will refer to the process appearing in the numerator as the “CUSUM process”, the process appearing in the denominator as the “self-normalizer process”, and the entire ratio process as the “self-normalized CUSUM process”.\footnote{To theoretically evaluate the efficiency of statistical tests, an analysis based on sequences of so-called local limiting alternatives (for example, sequences $c_1 - c_2 = O(n^{-1/2})$ in Proposition (ref) above) can be considered (see, for example, vandervaart1998). However, such an analysis would be considerably involved, and we leave it for future study.}

Extension to Multiple Change Points

We extend our single change-point testing methodology to the case of multiple change points. Typically, the number of potential change points in the alternative hypothesis must be prespecified. However, we leverage the recent work of zhang_etal2018 and introduce joint change-point tests for VaR and ES, that can accommodate an unknown, possibly multiple, number of change points in the alternative hypothesis. We fix some small $\delta > 0$ and consider the following null and alternative hypotheses (following the notation from Section (ref)).

align*[align* omitted — 669 chars of source]

Consider the index set $\Delta = \{ (s,t) \in [\delta,1-\delta]^2 : t-s \ge \delta \}$ and the test statistic

align*[align* omitted — 184 chars of source]

where

align*[align* omitted — 1,451 chars of source]

Then, under $\mathcal{H}_0$, applying Theorem 3 of zhang_etal2018, our Theorem (ref) above yields the following result.

corSuppose Assumption (ref) holds. Under the null hypothesis $\mathcal{H}_0$, \begin{align*} H_n \overset{d}{\to} & \sup_{(s_1,s_2) \in \Delta} E(0,s_1,s_2)^T F(0,s_1,s_2)^{-1} E(0,s_1,s_2) + \sup_{(t_1,t_2) \in \Delta} E(t_1,t_2,1)^T F(t_1,t_2,1)^{-1} E(t_1,t_2,1) \\ & := H, \end{align*} where \begin{align*} E(r_1,r_2,r_3) = & \begin{bmatrix} W_1(r_2) - W_1(r_1) - \frac{r_2 - r_1}{r_3 - r_1}(W_1(r_3) - W_1(r_1)) \\ W_2(r_2) - W_2(r_1) - \frac{r_2 - r_1}{r_3 - r_1}(W_2(r_3) - W_2(r_1)) \end{bmatrix} \\ F(r_1,r_2,r_3) = & \int_{r_1}^{r_2} \begin{bmatrix} W_1(s) - W_1(r_1) - \frac{s - r_1}{r_2 - r_1}(W_1(r_2) - W_1(_1)) \\ W_2(s) - W_2(r_1) - \frac{s - r_1}{r_2 - r_1}(W_2(r_2) - W_2(r_1)) \end{bmatrix}^{\otimes 2} ds \\ & \quad + \int_{r_2}^{r_3} \begin{bmatrix} W_1(r_3) - W_1(s) - \frac{r_3 - s}{r_3 - r_2}(W_1(r_3) - W_1(r_2)) \\ W_2(r_3) - W_2(s) - \frac{r_3 - s}{r_3 - r_2}(W_2(r_3) - W_2(r_2)) \end{bmatrix}^{\otimes 2} ds. \end{align*} Under the alternative $\mathcal{H}_1$, $H_n \overset{P}{\to} \infty$ as $n \to \infty$.

To reduce the computational burden of the method, we use a grid approximation suggested by zhang_etal2018, where in the doubly-indexed set $\Delta$, one index is reduced to a coarser grid. Specifically, let $\mathcal{G}_\delta = \{(1+k\delta)/2 : k \in \mathbb{Z}\} \cap [0,1]$ and consider the modified statistic

align[align omitted — 307 chars of source]

As before, under $\mathcal{H}_0$, we have

align[align omitted — 341 chars of source]

Note that simply using the original doubly-indexed set $\Delta$, for a sample of size $n$ and an arbitrary number of change points, we would need to search for maxima over $O(n^2)$ points. However, using the grid approximation, we need only search for maxima over $O(n)$ points. In contrast, if we were to use a direct extension of the single-change point detection methodology of Section (ref), with $m$ change points (which needs to be specified in advance), we would need to search for maxima over $O(n^m)$ points. Hence, the methodology introduced in this section offers significant computational savings.

We reject $\mathcal{H}_0$ if $\widetilde{H}_n$ exceeds the critical value corresponding to a desired test level of the pivotal quantity $\widetilde{H}$. To obtain critical values, we simulate 10,000 replications, with each replication consisting of 5,000 independent standard normal random variables to approximate standard Brownian motion on $[0,1]$.

Simulations

We perform an extensive simulation study to investigate the finite-sample performance of ES confidence interval construction using the sectioning and self-normalization methods (Section (ref)) as well as upper tail change detection (Sections (ref) and (ref)) using ES. Our simulations are based on a widely used and practically relevant data-generating process for modeling financial time series, namely Generalized Autoregressive Conditional Heteroskedasticity (GARCH) models. Specifically, in the main text we consider \[ X_i=\sigma_i\epsilon_i, \] \[ \sigma_i^2 = \omega + \lambda_1 X_{i-1}^2+\lambda_2\sigma_{i-1}^2, \] where the conditional variance $\sigma_i^2$ follows a GARCH(1,1) process. The innovations $\epsilon_i$ are assumed to be i.i.d. standard normal, and the model parameters are set as $\omega=0.01$, $\lambda_1=0.1$, and $\lambda_2=0.8$. This parameterization yields a highly persistent yet stationary process, as $\lambda_1+\lambda_2$ approaches 1. Our choice of parameters follows bollerslev1986generalized, where these parameters are estimated for a GARCH(1,1) model on lower-frequency financial time series data.\footnote{bollerslev1986generalized estimates the parameters $(\omega,\lambda_1,\lambda_2)=(0.007,0.135,0.829)$ {using quarterly U.S. inflation data from February 1948 to April 1983. While these parameters are not calibrated based on high-frequency asset returns, they are widely cited in the literature and serve as a benchmark for evaluating model performance.}} The Appendix collects a comprehensive set of robustness results with different parameters for the GARCH model and $t$-distributed innovations. In particular, in Figure (ref) we also show the results for a more persistent GARCH(1,1) process with parameters $\omega=0.01, \lambda_1=0.1$, and $\lambda_2=0.9$, which better reflects the dynamics of higher-frequency financial return series. This choice of $\lambda_1$ and $\lambda_2$ is also adopted in the recent paper horvath2024sequential. All our results are robust to these parameter choices.

Confidence Intervals

We begin by studying how the empirical coverage probability of confidence intervals depends on the sample size. In Figure (ref), we vary the time series sample size from 200 to 3,000 and examine the widths and empirical coverage probabilities of the 95% confidence intervals for ES in the upper 5th percentile. These confidence intervals are constructed using the sectioning and self-normalization methods (Section (ref)) for the GARCH(1,1) process introduced above. For comparison, we also include confidence intervals constructed using the block bootstrap method, as described in buhlmann2002bootstraps.

For the sectioning method, we show that the results are robust to the sectioning parameter $m$. We study the coverage probability and confidence intervals for the sectioning parameter $m\in \{2,5,8,10\}$. For the self-normalization method, we use the functional $T(Y)$ defined as $T(Y)=( \int_0^1 (Y(t)-tY(1))^2dt)^{1/2}$, as recommended in Section (ref). The block bootstrap method is applied with block lengths $l\in \{20,50\}$, and for this method, confidence intervals are constructed using the 0.025 and 0.975 quantiles of its empirical distribution. We average the estimates run over 10,000 replications. For each replication, we use a burn-in period of 5,000 to ensure the GARCH(1,1) process approximately reaches stationarity. Appendix B collects a comprehensive set of additional robustness results that confirm our findings. We repeat the simulation study for GARCH (1,1) models with parameters $\lambda_2 \in \{0.5, 0.6, 0.7, 0.8\}$ and $t$-distributed innovations $\epsilon_i$ with degrees of freedom $\nu=15$.

Figure (ref) shows that both the self-normalization method and sectioning method with different parameters $m$ achieve good coverage when the sample size is large. In particular, the coverage probability of the confidence intervals constructed by the sectioning method is robust to different choices of $m \in [2,10]$. In contrast, the block bootstrap method yields lower coverage probabilities and performs worse than our proposed methods{, requiring larger sample sizes to approximate the nominal coverage rate}. Additionally, the block bootstrap is also sensitive to the choice of the block length. Furthermore, the sectioning method with a larger $m\in \{5,8,10\}$ produces smaller confidence intervals, but generally yields lower empirical coverage probabilities compared to the self-normalization method. This is especially pronounced for small sample sizes such as 200 or 500.\footnote{{While a sample size of 200 or 500 is statistically small, obtaining such samples at lower frequencies (e.g., quarterly data) requires a long time span.}} As the sample size increases, the performance of the two methods becomes more similar, and the desired coverage probability of 0.95 is approximately reached. While choosing $m=2$ in the sectioning method can lead to wider confidence intervals, this can be easily identified in practice, allowing for the selection of a larger $m$ that leads to a tighter interval width.

figure[figure omitted — 1,466 chars of source]

Detection of Location Change in Tail

We investigate through simulations the detection of abrupt location changes using ES in the upper 5th percentile, as discussed in Section (ref). To assess the robustness of our method in detecting change points at different locations within the time series, we consider changes occurring at $t=rT$ with $r\in \{0.5,0.6,0.7,0.8,0.9\}$, where $T$ is the time horizon.

We consider the GARCH(1,1) process introduced previosuly before the potential distribution change at time $t=rT$. After the change point, we consider the following process, where the magnitude $\mu$ of the location change varies between 0 (the null hypothesis of no change) and 1, that is, $\mu \in [0,1]$: \[ X_i=\mu+\sigma_i\epsilon_i, \] \[ \sigma_i^2 = 0.01 + 0.1 (X_{i-1}-\mu)^2+0.8\sigma_{i-1}^2. \]

Figure (ref) shows the approximate power of change-point tests using the self-normalized CUSUM statistic (see ((ref))-((ref))) at the 0.05 significance level for varying magnitudes of the abrupt location change. For each data point, we perform 100 replications of the change-point testing using times series sequences of length 2,000. For each replication, we use a burn-in period of 5,000 to ensure the GARCH(1,1) process approximately reaches stationarity.

As expected, the power of the test increases monotonically with the magnitude of the location change. In accordance with the desired 0.05 significance level of our procedure, the probability of false positive detection is approximately 0.05, as indicated by the points with zero magnitude of location change in Figure (ref). While the location of the change point does affect the detection probability, the test is robust when the change occurs between the $r=0.5$ and $0.8$ fractions of the series. For changes occurring at the very end of the series (at $r=0.9$), the detection probability decreases. However, the test can still achieve high detection power as the magnitude of the change increases.

While our method can detect changes in the mean, it is specifically designed to capture more general changes in the distribution. Therefore, when the change occurs solely in the mean, our method may require a relatively larger magnitude of shift for detection compared to methods designed specifically for mean change detection. In the next section, we focus on more complex structural breaks, where our method is particularly useful and demonstrates strong performance.

figure[figure omitted — 1,062 chars of source]

Detection of General Change in Tail

We also investigate the detection of general structural changes in the upper tail of the underlying marginal distribution occurring at different points within the time interval. Although the relationship between power and the “magnitude" of the change in the upper tail is not as simple as in the case of pure location change, nevertheless, with Proposition (ref) we will detect the change with high probability as our sample size increases. In our simulations, we study the detection of general structural changes in the tail using ES in the upper 5th percentile, as discussed in Section (ref).

As in the previous section, we consider the GARCH(1,1) process with i.i.d. normal innovations and parameters $\omega=0.01$, $\lambda_1=0.1$, and $\lambda_2=0.8$ before the distribution change at time $t=rT$, with $r=\{0.5,0.6,0.7,0.8,0.9\}$. After the change point, we consider the following processes:

enumerate[label=(\arabic*)] • $X_i=\sigma_i\epsilon_i$, $\sigma_i^2 = 0.01 + \lambda_1 X_{i-1}^2+0.8\sigma_{i-1}^2$, $\epsilon_i \overset{i.i.d.}{\sim} \mathcal{N}(0,1)$; • $X_i=\sigma_i\epsilon_i$, $\sigma_i^2 = 0.01 + 0.1 X_{i-1}^2+\lambda_2\sigma_{i-1}^2$, $\epsilon_i \overset{i.i.d.}{\sim} \mathcal{N}(0,1)$; • $X_i=\sigma_i\epsilon_i$, $\sigma_i^2 = 0.01 + 0.1 X_{i-1}^2+0.8\sigma_{i-1}^2$, $\epsilon_i \overset{i.i.d.}{\sim} t(\nu)$.

In the first setup, we vary the persistence parameter $\lambda_1$ from $0.1$ to $0.19$. In the second setup, we vary the parameter $\lambda_2$ from $0.8$ to $0.89$. {These ranges of parameters ensure that the process remains stationary.} In the last setup, we consider a change in the heavy-tailedness of the underlying distribution, transitioning from Gaussian innovations to $t$-distributed innovations with {degrees of freedom $\nu\in \{2,4,6,8,10,12,14,16,100,1000\}$. As $\nu$ increases, the $t$-distribution approaches the Gaussian distribution, and hence we expect the detection probability to decline accordingly}. Figure (ref) shows the approximate power of change-point tests using the self-normalized CUSUM statistic (see ((ref))) at the 0.05 significance level. For each data point, we perform 100 replications using times series sequences of length 2,000 and use an initial burn-in period of 5,000 to approximately reach stationarity.

figure[figure omitted — 1,860 chars of source]

As expected, for all types of changes, the power is approximately a monotonic function of the magnitude of change. Furthermore, our method is generally robust to the location of the change point, provided that it is not located near the very end of the time period. When the change point occurs close to the end of the series, the detection probability decreases slightly, but our method is still able to detect the change with high probability as the magnitude of the change increases.

Detection of Multiple Changes in Tail

We additionally investigate the detection of multiple structural changes in the upper tail of the underlying marginal distribution. Here, we again use ES in the upper 5th percentile and compare the practical performance of the single change-point methodology discussed in Section (ref) with the unsupervised multiple change-point methodology discussed in Section (ref). We consider the following variant of the GARCH(1,1) process introduced previously:

align[align omitted — 352 chars of source]

\[ \qquad \qquad \sigma_i^2 = 0.01 + 0.1 X_{i-1}^2+0.8\sigma_{i-1}^2. \] In this GARCH(1,1) process, the innovations initially have relatively light tails, then change to heavier tails with $t(\nu)$-distributed innovations, and finally revert back to the original lighter tails.

figure[figure omitted — 970 chars of source]

Figure (ref) shows the approximate detection power of at least one change point as $\nu$ varies using the single change-point methodology (see ((ref))-((ref))) versus the unsupervised multiple change-point methodology (see ((ref))-((ref))) at the 0.05 significance level. For each data point, we perform 100 replications using times series sequences of length 1,200. For each replication, we use an initial burn-in period of 5,000 to approximately reach stationarity in the GARCH(1,1) process.

We observe that the single change-point testing method is unable to detect a change in the process in ((ref)), even for extremely strong deviations from the null such as the case $\nu = 2$. In fact, its power decays to zero as the magnitude of the change increases. This occurs because the single change-point test assumes that only one change point occurs and seeks to divide the time series into two sections separated by a single change point. Consequently, the effects of multiple changes can be “canceled out”. On the other hand, the unsupervised multiple change-point test exhibits the desired performance with increasing power as the magnitude of the change increases. Hence, it is a promising candidate for detecting more complex patterns of changes in the tails of time series.

Empirical Applications

Data and Expected Shortfall Estimation

We study the changes in tail risk for two of the most important macroeconomic indicators that reflect the conditions in the U.S. stock market and for interest rates. First, we consider the daily log returns of the S&P 500 Index, which is a proxy for a market risk factor, for the years 1928-2023.\footnote{We obtain the time-series from Yahoo Finance.} Second, we analyze the weekly log returns of U.S. 1-year and 10-year zero-coupon bonds for the years 1961-2023. We obtain the data from the U.S. Treasury Discount Bond Database provided by filipovic2022shrinking, which is shown to provide the most reliable estimates. These bonds are tradable portfolios of U.S. Treasuries and reflect information in interest rates.

figure[figure omitted — 1,045 chars of source]
figure[figure omitted — 1,064 chars of source]

We begin our study with an analysis of the expected shortfall (ES) over time. We first apply the two methods of confidence interval construction from Section (ref) to the daily log returns of S&P 500 for the period from January 03, 1928 to December 29, 2023. In Figure (ref), we show ES estimates of the lower 5th percentile of log returns throughout this time period along with 95% confidence bands computed using the sectioning and self-normalization methods. We use the sectioning method with $m=10$. (Our simulation study demonstrates that the coverage results are robust to the choices of the tuning parameter $m$.) We use a rolling window of 100 days with 90 days of overlap between successive windows. The self-normalization method appears to be more conservative and yields a wider confidence band compared to the sectioning method, which agrees with the results presented in Figure (ref). Overall, the ES estimates and confidence bands appear to capture well the increased volatility of returns during periods of financial instability, such as the 1929 Wall Street Crash, World War II in 1940, the 1987 Black Monday Crash, the 2008 Financial Crisis, and the COVID-19 Pandemic in 2020.

We next apply the two methods of confidence interval construction from Section (ref) to weekly log returns of U.S. 1-year and 10-year zero-coupon bonds for the period from July 03, 1961 to August 31, 2023. In Figure (ref), we show ES estimates of the lower 10th percentile of log returns throughout this time period along with 95% confidence bands computed using the sectioning and self-normalization methods. We use a rolling window of 40 weeks with 30 weeks of overlap between successive windows. Again, the self-normalization method appears to be more conservative and yields a wider confidence band compared to the sectioning method. Although time series for bond returns have a greater degree of autocorrelation compared to S&P 500 returns, our methods are still applicable.\footnote{For comparison, we also estimate the ES with a parametric GARCH(1,1) model. Figures (ref) and (ref) in the Appendix collect the results. We show that for both the S&P 500 log returns and U.S. 10-year bond log returns, the ES estimates are broadly similar using the parametric GARCH model and our method. However, the parametric GARCH model can lead to unstable and spurious results. For example, for the U.S. 1-year bond log returns, the ES estimates show some noticeable differences, indicating that our method and the parametric model can yield different results. The estimation of the GARCH model on a rolling window can also lead to issues, and on approximately 10% of the rolling windows the optimizer failed to converge. For example, for the period from March 3, 1930 to July 23, 1930, the optimization method failed for both GARCH(1,1) models with either $t$-distributed innovations or normal innovations. We only plot the ES estimates from the GARCH model when the fitting process succeeds. Overall, our method appears to be more robust and reliable.}

These rolling window estimates of ES with confidence intervals provide a valuable tool to get insight into the variation of the tail behavior of these time series. However, they are not a formal test for change points. Such tests are based on our novel statistics, which we study next.

Change-Point Detection

We now apply our change-point testing methodology to the equity and bond return time-series. We use our single and multiple change-point testing methodology for ES of the lower 5th percentile of S&P 500 daily log returns and of the lower 10th percentile of U.S. 1-year and 10-year zero-coupon bonds weekly log returns. For a systematic study, we run the tests on rolling windows over the full return time series. We compare the behavior of the tests and also visualize the return time-series on windows of interest.\footnote{Recall that our change-point testing methodology is meant to test for change(s) in a given time window, and help assess whether or not risk measures are stationary over the window. For the rolling window demonstrations in this section, if one is interested in performing statistical inference simultaneously for all windows, then one could run a multiple hypothesis testing procedure that controls the False Discovery Rate (FDR). A robust (albeit conservative) choice of procedure, which is valid for arbitrary dependence among the tests, is the Benjamini-Yekutieli (B-Y) procedure benjamini_yekutieli_2001. A description of the B-Y procedure is provided at the end of Appendix C.}

In more detail, we apply our single (based on ((ref))-((ref))) and multiple (based on ((ref))-((ref))) change-point tests to S&P 500 returns in successive 6-month and 1-year time windows, with 1 month of shift between successive windows. In Figures (ref)-(ref) in Appendix C, we plot the test statistics for single and multiple change-point tests for ES in the lower 5th percentile of S&P 500 log returns for different time windows. We also indicate the critical value for the 0.05 significance level. Then, we apply the two change-point tests to weekly log returns of the zero-coupon bonds. We use successive 5-year and 10-year time windows for both bonds, with 3 months of shift between successive windows. In Figures (ref)-(ref) in Appendix C, we plot the test statistics for single and multiple change-point tests for ES in the lower 10th percentile of U.S. 1-year and 10-year zero-coupon bond log returns for different time windows.

In Tables (ref) and (ref), for S&P 500 returns and bond returns, respectively, we report examples of the single and multiple change-point test results. For each window, we report the test statistics and the corresponding p-values.\footnote{ In Table (ref), we also report results from applying the B-Y multiple hypothesis testing adjustment for single and multiple change-point tests over rolling windows for the bond returns. Despite the conservativeness of the B-Y procedure, the bond returns generally exhibit less volatility relative to the S&P 500 returns, and allow for discoveries with FDR control. Unfortunately, the B-Y procedure is too conservative to reliably make discoveries with FDR control for S&P 500 returns.}

table[table omitted — 2,513 chars of source]
table[table omitted — 3,597 chars of source]

There are examples of time windows where both the single and multiple change-point tests detect change at the 0.05 significance level. We collect some of these examples in the top block of Tables (ref) and (ref). To gain better intuition for the results, we also plot the return time series for some of these examples. In Figure (ref), we plot the return time series for the S&P 500 log returns for periods including major events such as the 1929 Wall Street Crash, World War II in 1940, the 1987 Black Monday Crash, the 2008 Global Financial Crisis, the 2011 August Market Decline, and the 2020 Covid-19 Pandemic. Similarly, for the 1-year and 10-year bond returns, the top block of Table (ref) shows examples of time windows where both the single and multiple change-point tests detect change at the 0.05 significance level. In Figure (ref), we plot the return time series for some of these time windows, including the 1980-1982 US recession (due in part to government restrictive monetary policy, and to a lesser degree the Iranian Revolution of 1979, which resulted in significant oil price increases) and the 2008 Financial Crisis. These findings suggest that structural breaks in ES for stock returns closely align with major financial crises. This highlights the sensitivity of tail risk to systemic market disruptions. For bond returns, the detected breaks often coincide with significant shifts in monetary policy, which implies that changes in risk may reflect expectations about interest rate and policy uncertainty.

figure[figure omitted — 626 chars of source]
figure[figure omitted — 714 chars of source]

Interestingly, the change points detected for U.S. 1-year bond returns can differ from those for the U.S. 10-year bond returns. For example, during the period from August 05, 2012 to July 23, 2017, the p-values of both single and multiple-change tests for U.S. 1-year bond returns are less than 0.001, indicating the presence of change points. This is illustrated in the third plot in Figure (ref). However, the p-values of both tests for U.S. 10-year bond returns exceed 0.7, suggesting no significant change points during the same period. This discrepancy might be due to the Federal Reserve's decision in 2015 to raise the federal funds rate, initiating a monetary tightening cycle after the Great Recession. Short-term yields, such as those on 1-year bonds, are directly tied to the federal funds rate and quickly reflect such policy changes. In contrast, long-term yields, such as those on 10-year bonds, are less sensitive to short-term policy changes and more influenced by long-term expectations of future economic growth and inflation.

On the other hand, different window sizes and change-point testing methods can yield different results. There are time windows where the single change-point test fails, but the multiple change-point test succeeds; see the bottom block of Table (ref) for the S&P 500 daily returns. One example is the period between May 01, 2008 and April 30, 2009. Here, the p-value is 0.018 for the unsupervised multiple change-point test, with the test statistic (in ((ref))) having value 171.7, which indicates that there are one or more change points during this time window. However, the single change-point test performs poorly on this time window, with a p-value of 0.487 and the test statistic (in ((ref))) having a value of 8.0. As seen in Figure (ref), changes in the tail of the distribution of the S&P 500 log returns are clearly present at the end of 2008, corresponding to the 2008 Financial Crisis. Because the single change-point test assumes that only one change point occurs and seeks to divide the time series into two sections separated by a single change point, the effects of multiple changes can be “canceled out”. The same issue arises for the U.S. 1-year bond log returns in the period from September 19, 1976 to August 31, 1986, and for U.S. 10-year bond log returns in the period from December 18, 1977 to November 29, 1987. This problem is addressed by the multiple change-point test.

However, there is also a trade-off associated with using the multiple change-point test, as it can suffer from low power when there is a single change point and the time window is short. This is supported by the examples in the middle block in Table (ref) for the S&P 500 log returns. In these instances, the single change-point test has p-values that are all less than 0.025, but the multiple change-point test has p-values that are between 0.059 and 0.156. This suggests that the multiple change-point test detects some evidence of a change point, but not enough to achieve statistical significance at the 0.05 level.

In light of the above discussion, we recommend the following use of single and multiple change-point tests. In general, the multiple change-point test can be run first. If the test fails to reach statistical significance, and it is plausible that there is only one change point, then to increase power, the single change-point test can be run (perhaps within a local window). Alternatively, the multiple change-point test can be used to check the result of the single change-point test, and help rule out the possibility that the effects of multiple changes are canceled out. In summary, the single and multiple change-point tests can be deployed in complementary ways.

Conclusion

We propose a novel methodology to perform confidence interval construction and change-point testing for fundamental nonparametric estimators of risk such as ES. This allows for evaluation of the homogeneity of ES and related measures such as the conditional tail moments, and in particular allows the researcher to detect general tail structural changes in time series observations. While current approaches to tail structural change testing typically involve quantities such as the tail index and thus require parametric modeling of the tail, our approach does not require such assumptions. Moreover, we are able to detect more general structural changes in the tail using ES, for example, location and scale changes, which are undetectable using tail index. Hence, we advocate for the use of ES for general purpose monitoring for tail structural change.

We view our methods as being more robust compared to extant approaches which involve consistent estimation of standard errors or blockwise versions of the bootstrap or empirical likelihood, which can be more sensitive to tuning parameters. We note that our proposed sectioning and self-normalization methods for confidence interval construction and change-point testing still require some user choice, for example, the number of sections to use in sectioning or which particular functional to use for self-normalization. Our simulations demonstrate that our methods are robust to these user choices.

Our simulations show the promising finite-sample performance of our procedures. Our empirical study of S&P 500 and US Treasury bond returns illustrates the practical use of our methods in detecting and quantifying market instability via the tails of financial time series.

\setcounter{table}{0} \setcounter{figure}{0}