Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,452 characters · 9 sections · 74 citation commands
V1 A Residual Bootstrap for Conditional Expected Shortfall
\doublespacing
The assessment of market risk is a key challenge that financial market participants face on a daily basis. To evaluate the risk, financial institutions primarily employ the risk measure Value-at-Risk (VaR) to meet the capital requirements enforced by the Basel Committee on Banking Supervision. Despite its popularity, the VaR is not a coherent risk measure as it fails to fulfill the subadditivity property artzneretal1999. A coherent alternative is the related risk measure Expected Shortfall (ES). For a given level $\alpha$, it is defined as the expected return in the worst $100\alpha \%$ cases and is therefore sometimes called Expected Tail Loss.\footnote{In the literature, ES is also sometimes referred to as conditional VaR since it is defined as the expected loss given a VaR exceedence. Since conditional refers to temporal dependence (i.e.\ conditional on past returns) in this paper, we refrain from using this term to prevent any confusion.} In contrast to the VaR, the ES provides valuable information on the severity of an incurred loss, which makes it the preferred risk measure in practice (c.f.\ acerbitasche2002a, acerbitasche2002a; acerbitasche2002b). Consequently, the Basel Committee published revised standards in January 2016 resembling a shift from VaR towards ES as the underlying risk measure osmundsen2018.
In the literature there is an increasing interest in conditional risk measures, which take into account the temporal dependence of asset returns. Frequently, the volatility dynamics are specified by a (semi-) parametric model such that the conditional ES can be expressed as the product of the conditional volatility and the ES of the innovations’ distribution (c.f.\ francq2015risk, francq2015risk, Example 2). The latter can be treated as additional parameter, which is generally unknown just like the parameters of the conditional volatility model. Inferring the parameters from data leads to an evaluation of the conditional ES that is prone to estimation risk. As argued in beutner2018residual this estimation uncertainty can be substantial for risk measures related to extreme events.
The uncertainty around point estimates is typically determined by asymptotic theory, in which one replaces unknown quantities in the limiting distribution by consistent estimates. For example, caiwang2008 and martinsetal2018 study the behavior of proposed nonparametric estimators for conditional VaR and ES, based on asymptotics and simulation studies. An alternative approach is based on bootstrap approximations. Regarding the estimators of the volatility model’s parameters, several bootstrap methods have been examined, among which the sub-sample bootstrap hall2003inference, the block bootstrap corradi2008bootstrap, the wild bootstrap shimizu2009bootstrapping and the residual bootstrap, both with recursive pascual2006bootstrap,hidalgo2007goodness and fixed design (shimizu2009bootstrapping, shimizu2009bootstrapping; cavaliere2018fixed, cavaliere2018fixed, beutner2018residual, beutner2018residual). However, the estimation of the conditional ES has received only limited attention in the bootstrap literature. christoffersen2005estimation construct intervals for conditional ES based on a recursive-design residual bootstrap method. gao2008estimation compare coverage probabilities for conditional ES based on this bootstrap method and asymptotic normality results in their simulation study.
In this paper, we extend results of beutner2018residual derived for conditional VaR to the conditional ES estimator. In particular, we follow the two-step procedure of francq2015risk for the estimation of the underlying parameters. In a first step, we obtain estimates of the parameters of the stochastic volatility model by quasi-maximum-likelihood (QML) estimation. Based on the model's residuals, an estimate for the innovations’ ES is obtained in the second step. We propose a fixed-design residual bootstrap method to mimic the finite sample distribution of this two-step estimator for a general class of volatility models. Moreover, an algorithm is provided for the construction of bootstrap intervals for the conditional ES.
The remainder of the paper is organized as follows. Section (ref) introduces a general class of volatility models and derives the conditional ES. The two-step estimation procedure is described in Section (ref) and corresponding asymptotic results are provided under the assumptions imposed by beutner2018residual. In Section (ref), a fixed-design residual bootstrap method is proposed and proven to be consistent. In addition, bootstrap intervals are constructed for the conditional ES. Section (ref) consists of a Monte Carlo study. Section (ref) concludes. Auxiliary results and proofs of the main results are gathered in the Appendix.
We consider conditional volatility models of the form
with $t\in \mathbb{Z}$, where $\epsilon_t$ denotes the log-return, $\{\sigma_t\}$ is a volatility process and $\{\eta_t\}$ is a sequence of independent and identically distributed (i.i.d.) variables. The volatility is assumed to be a measurable function of past observations
with $\sigma:\mathbb{R}^\infty\times \Theta\to(0,\infty)$ and $\theta_0$ denotes the true parameter vector belonging to the parameter space $\Theta \subset \mathbb{R}^r$, $r \in \mathbb{N}$. Various commonly used volatility models satisfy (ref)--(ref); for examples see francq2015risk (francq2015risk, Table 1). Consider an arbitrary real-valued random variable $X$ (e.g.\ stock return) with cdf $F_X$. If $\mathbb{E}_X[X^-]<\infty$ with $X^{-} = \max\{-X,0\}$, then the ES at level $\alpha \in (0,1)$ is finite and given by $ES_\alpha(X) = -\mathbb{E}_X\big[X|X<F_X^{-1}(\alpha)\big]$. Let $\mathcal{F}_{t-1}$ denote the $\sigma$-algebra generated by $\{\epsilon_u, u<t\}$. It follows that the conditional ES of $\epsilon_t$ given $\mathcal{F}_{t-1}$ at level $\alpha \in (0,1)$ is
As $\eta_t$ are i.i.d., the ES at level $\alpha$ of $\eta_t$ is constant for a given $\alpha$ and can be treated as a parameter. Setting $\mu_\alpha = -\mathbb{E}\big[\eta_t|\eta_t<\xi_\alpha\big]$ with $\xi_\alpha = F^{-1}(\alpha)$ and $F$ denoting the cdf of $\eta_t$, (ref) reduces to
Typically, $\alpha$ is chosen small (e.g.\ 5% and 1%) such that $\xi_\alpha<0$ and hence $\mu_\alpha>0$. Except for special cases\footnote{We derive the analytical expressions for $\mu_\alpha$ for the cases in which $\eta_t$ are normally as well as Student-$t$ distributed in Appendix B.}, $\mu_\alpha$ is unknown and needs to estimated just like $\theta_0$.
For the estimation of the parameters $\theta_0$ and $\mu_\alpha$ we employ the two-step procedure of francq2015risk (francq2015risk, Remark 3). First, the vector of the conditional volatility parameters $\theta_0$ is estimated by quasi-maximum-likelihood (QML). Since
can generally not be determined completely given a sample $\epsilon_1, ,\dots, \epsilon_n$, we replace the unknown presample observations by arbitrary values, say $\tilde{\epsilon}_t$, $t\leq 0$, yielding
Then the QML estimator of $\theta_0$ is defined as a measurable solution $\hat{\theta}_n$ of
with the criterion function specified by
In the second step, $\mu_\alpha$ can be estimated on the basis of the first-step residuals, i.e. $\hat{\eta}_t=\epsilon_t/\tilde{\sigma}_t(\hat{\theta}_n)$. A reasonable estimator of $\mu_\alpha$ (c.f.\ gao2008estimation, gao2008estimation) is given by
where $\hat{\xi}_{n,\alpha}$ is the empirical $\alpha$-quantile of $\hat{\eta}_1,\dots,\hat{\eta}_n$, i.e.\ $\hat{\xi}_{n,\alpha}=\hat{\mathbbm{F}}_n^{-1}(\alpha)$ with empirical distribution function $\hat{\mathbbm{F}}_n(x)=\frac{1}{n}\sum_{t=1}^n \mathbbm{1}_{\{\hat{\eta}_t\leq x\}}$.
Having obtained estimators for $\theta_0$ and $\mu_\alpha$, we turn to the estimation of the conditional ES of the one-period ahead observation at level $\alpha$. For notational convenience, we use the abbreviation $ES_{n,\alpha}$ to denote $ES_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$. Employing (ref)--(ref) we can estimate $ES_{n,\alpha}$ by
For the asymptotic analysis of (ref)--(ref) we assume the conditions of beutner2018residual, which we restate for completeness.
For a discussion of the conditions we refer to francq2015risk (francq2015risk, Section 2 and 3) and beutner2018residual (beutner2018residual, Section 3). On the basis of the previous assumptions we extend the strong consistency result of francq2015risk (francq2015risk, Theorem 1) to the estimator of the ES at level $\alpha$ of $\eta_t$.
To lighten notation, we henceforth write $D_t(\theta) =\frac{1}{\sigma_t(\theta)}\frac{\partial\sigma_t(\theta)}{\partial \theta}$ and drop the argument when evaluated at the true parameter, i.e.\ $D_t=D_t(\theta_0)$. The next result provides the joint asymptotic distribution of $\hat{\theta}_n$ and $\hat{\mu}_{n,\alpha}$ and is due to francq2012risk.
In order to evaluate $\sigma^{2}_{\alpha}$ and $x_{\alpha}$ in Theorem (ref), we need expressions for the variance and covariance term respectively. After basic manipulation we find
with $p_{\alpha} = \mathbb{E}[\eta_{t}^{2}\mathbbm{1}_{\{\eta_t<\xi_\alpha\}}] - \alpha$ and $q_\alpha = \mathbb{E}[\eta_t^{3}\mathbbm{1}_{\{\eta_t< \xi_\alpha\}}]$. In a GARCH($p,q$) setting, gao2008estimation quantify the uncertainty around $\hat{\theta}_n$ and $\hat{\mu}_{n,\alpha}$ using (ref) while replacing the unknown quantities in $\Gamma_\alpha$ by estimates. In the same spirit, $\xi_\alpha$ and $\mu_\alpha$ can be substituted by $\hat{\xi}_{n,\alpha}$ and $\hat{\mu}_{n,\alpha}$ while $\Omega$, $J$, $q_\alpha$, $p_\alpha$ and $\kappa$ can be replaced by
with $\hat{D}_t = \tilde{D}_t(\hat{\theta}_n)$ and $\tilde{D}_t(\theta) = \frac{1}{\tilde{\sigma}_t(\theta)}\frac{\partial\tilde{\sigma}_t(\theta)}{\partial \theta}$. The strong consistency of the estimators in (ref) follows from beutner2018residual (beutner2018residual, Lemma 2 and Theorem 1). Based on (ref) we obtain a consistent estimator for $\Gamma_\alpha$ denoted by $\hat{\Gamma}_{n,\alpha}$. Note that in the joint asymptotic distribution in Theorem (ref) the pdf of $\eta_{t}$ does not occur. This is in contrast to the limiting distribution of the parameters that comprise the conditional VaR estimator beutner2018residual. Hence, no density estimation (by e.g.\ kernel smoothing) is required here.
The asymptotic behavior of the conditional ES estimator can be studied by employing Theorem (ref). Since the conditional volatility varies over time, a limiting distribution cannot exist and therefore the concept of weak convergence is not applicable in this context. beutner2017justification (beutner2017justification, Section 4) advocate a merging concept that generalizes the notion of weak convergence, i.e.\ two sequences of (random) probability measures $\{P_n\},\{Q_n\}$ merge (in probability) if and only if their bounded Lipschitz distance $d_{BL}(P_n,Q_n)$ converges to zero (in probability). Assuming two independent samples, one for parameter estimation and one for conditioning, the delta method suggests that the ES estimator, centered at $ES_{n,\alpha}$ and inflated by $\sqrt{n}$, and
given $\mathcal{F}_n$ merge in probability. Equation (ref) highlights once more the relevance of the merging concept since the conditional variance still depends on $n$ and does not converge as $n \to \infty$. In combination with Theorem (ref) and $\hat{\Gamma}_{n,\alpha} \overset{a.s.}{\to} \Gamma_\alpha$, $100(1-\gamma)\%$ confidence intervals for $ES_{n,\alpha}$ can be constructed with bounds given by
where $\Phi$ denotes the standard normal cdf. It has to be mentioned that researchers rarely have a replicate, independent of the original series, to their disposal.\footnote{Exceptions would include some experimental settings.} An asymptotic justification for the interval on the basis of a single sample is given in beutner2017justification. Bootstrap methods offer an alternative way to quantify the uncertainty around the estimators.
We propose a fixed-design residual bootstrap procedure, described in Algorithm (ref), to approximate the distribution of the estimators in (ref)-(ref).
In the following subsection we show the asymptotic validity of the fixed-design bootstrap procedure described in Algorithm (ref).
Subsequently, we employ the usual notation for bootstrap asymptotics, i.e.\ “$\overset{p^*}{\to}$" and “$\overset{d^*}{\to}$", as well as the standard bootstrap stochastic order symbol “$o_{p^*}(1)$" (c.f.\ chang2003sieve, chang2003sieve). The asymptotic validity of the bootstrap corresponding to the stochastic volatility part is shown in beutner2018residual (beutner2018residual, Proposition 1). Therefore, we focus only on $\hat{\mu}_{n,\alpha}^*$. By construction, we have $\sum_{t=1}^n \mathbbm{1}_{\{\hat{\eta}_t^*<\hat{\xi}_{n,\alpha}^*\}}=\floor{\alpha n}+1$, where $\floor{x}$ denotes the largest integer not exceeding $x$. Defining $\alpha_n=\frac{\floor{\alpha n}+1}{n}$, we standardize (ref) such that the bootstrap estimator satisfies
where the scaling factor $-1/\alpha_n$ in (ref) converges to $-1/\alpha$ since $\alpha\leq \alpha_n \leq \alpha +\frac{1}{n}$. The different terms in brackets are given by {\allowdisplaybreaks
} Employing arguments of chen2007nonparametric (chen2007nonparametric, Lemma 2) Lemma (ref) in Appendix (ref) states the asymptotic negligibility of $A_n^*$, i.e.\ $A_n^*\overset{p^*}{\to}0$ in probability. The term $B_n^* = 0$ since $\frac{1}{n}\sum_{t=1}^n \mathbbm{1}_{\{\hat{\eta}_t^*<\hat{\xi}_{n,\alpha}^*\}}=\alpha_n$ by construction. Further, Lemma (ref) in Appendix (ref) states that $C_n^*=\alpha \mu_\alpha \Omega \sqrt{n}\big(\hat{\theta}_n^*-\hat{\theta}_n\big)+o_{p^*}(1)$ in probability. Last, we have $D_n^*\overset{d^*}{\to}N(0,\nu_\alpha)$ almost surely by Lemma (ref) in Appendix (ref). The previous discussion together with the asymptotic expansion of $\sqrt{n}\big(\hat{\theta}_n^*-\hat{\theta}_n\big)$ in beutner2018residual (beutner2018residual, Equation 4.4) yields
in probability. Employing Lemma (ref) once more leads to the paper's main result.
Theorem (ref) is useful to validate the bootstrap for the conditional ES estimator. For the asymptotic behavior of the conditional ES estimator we refer to (ref) and the text preceding it. The following corollary is established.
Having proven first-order asymptotic validity of the bootstrap procedure described in Section (ref), we turn to constructing bootstrap confidence intervals for ES.
Clearly, the ES evaluation in (ref) is subject to estimation risk that needs to be quantified. We propose the following algorithm to obtain approximately $100(1-\gamma)\%$ confidence intervals.
For a discussion of the three interval types in Algorithm (ref), we refer to beutner2018residual (beutner2018residual, Section 4.3). In the next section, features of the fixed-design bootstrap confidence intervals for the conditional ES are studied by means of simulations.
To assess the proposed bootstrap procedure in finite samples, we consider a simulation setup similar to beutner2018residual. The Data Generating Process (DGP) is a GARCH($1,1$), which falls in the class of conditional volatility models defined in (ref)--(ref). More specifically, we consider
with $\theta_0=(\omega_0,\alpha_0,\beta_0)^\prime$. Regarding the GARCH parameters we study two scenarios:
The innovations $\{\eta_t\}$ are drawn from two different distributions: the Student-$t$ distribution with $\nu=6$ degrees of freedom and the standard normal distribution (which corresponds to the case $\nu=\infty$). Whereas in the latter case the innovations are appropriately standardized, in the former we draw from the normalized density $f(x)=\frac{1}{\sigma_\nu}f_\nu(x/\sigma_\nu)$ such that $\mathbb{E}[\eta_t^2]=1$, where $\sigma_\nu^2=\frac{\nu-2}{\nu}$ and $f_\nu(x)=\frac{\Gamma(\frac{\nu+1}{2})}{\sqrt{\nu \pi}\Gamma(\frac{\nu}{2})}\big(1+\frac{x^2}{\nu}\big)^{-\frac{\nu+1}{2}}$. In this setting, the ES of the innovations’ distribution reduces to $\mu_\alpha=\frac{f_{\nu-2}(\xi_\alpha)}{\alpha}$ with $\xi_\alpha=\sigma_\nu F_\nu^{-1}(\alpha)$ and $F_\nu(x)=\int_{-\infty}^x f_\nu(y) dy$; we refer to Appendix (ref) for details. For the experiment, the ES level takes two values: $\alpha \in \{0.01, 0.05 \}$. We consider four different sample sizes $n \in \{ 500; 1{,}000; 5{,}000; 10{,}000 \}$ and the number of bootstrap replicates is fixed at $B = 2{,}000$. For each model, we simulate $S = 2{,}000$ independent Monte Carlo trajectories. All simulations are carried out on a HP Z640 workstation with 16 cores using Matlab R2016a. The numerical optimization of the log-likelihood function is performed using the built-in function fmincon. Parallel computing by means of parfor is employed to reduce running time significantly.
beutner2018residual (beutner2018residual) demonstrate that the bootstrap distribution mimics adequately the finite sample distribution of the estimator of the volatility parameters. In a similar fashion, we assess whether the bootstrap distribution (given a particular sample) mimics the distribution of the ES parameter estimator.
Figure (ref) displays the density estimates for the distribution of $\sqrt{n}(\hat{\mu}_{n,\alpha}-\mu_\alpha)$ and $\sqrt{n}(\hat{\mu}_{n,\alpha}^*-\hat{\mu}_{n,\alpha})$ in the high persistence case for $n=5{,}000$ with $\alpha \in \{0.01,0.05\}$. In both cases, we observe that the density plots are bell curves around the value zero, which supports the theoretical results of Theorem (ref) and (ref). Since the density graphs for the other scenarios are very similar, they are not reported in order to conserve space. We continue by studying the coverage probabilities of the three bootstrap intervals introduced in Section (ref).
Table (ref) reports the results of the three $90\%$--bootstrap intervals for the $5\%$--ES with Student-$t$ distributed innovations (which we consider as benchmark). For moderate sample sizes, we observe satisfactory coverage probabilities that lie relatively close to the nominal level of $90\%$. For small sample size ($n=500)$, the intervals exhibit small under-coverage with values ranging from $4.00$ to $5.85$ percentage points ($pp$) below the nominal value. For all three intervals, we find that the average rate of the conditional ES being below the interval is considerably less than it being above the interval. This phenomenon is most pronounced in smaller sample size. Concerning the average length of the intervals, we can make two important observations. Firstly, the SY interval is generally larger than the EP/RT interval.\footnote{By construction, the EP and the RT interval are of equal length.} As sample size increases, this gap disappears and all intervals' average lengths shrink. Secondly, the average length of intervals is larger in the high persistence case, as the conditional volatility varies more compared to the lower persistence case. In the following, we study deviations from the benchmark specification. Table (ref) considers a change in the innovation distribution $F$, while Table (ref) and (ref) take into account a change in the ES level $\alpha$ and a change in the nominal coverage probability $100(1-\gamma)\%$, respectively.
Table (ref) considers the case where the innovations follow a standard normal distribution. Results are qualitatively similar to the benchmark. In particular, coverage rates are generally close to the nominal level for $n\geq 1{,}000$ yet the under-coverage in smaller sample sizes is less in this scenario. For example, the average coverage is at most $3.70pp$ below the $90\%$ level even when $n=500$. In general, results seem to be “less extreme" compared to the benchmark: $(i)$ the average length of all intervals is smaller for all sample sizes and $(ii)$ the average rate of the conditional ES being above the interval lies closer to the corresponding rate below the interval. Moreover, we observe that there is no interval that outperforms the others in the case of $\eta_{t}$ being standard normally distributed.
Table (ref) provides simulation results for the conditional ES at level $\alpha = 0.01$, where the DGP is a GARCH($1,1$) with Student-$t$ innovations (6 degrees of freedom). Unsurprisingly, we find that the average length of all intervals is considerably larger compared to the benchmark. More strikingly, we observe that the phenomenon of under-coverage appears across sample sizes. For the lowest sample size considered, i.e. $n=500$, average coverage rates are between $15pp$ and $20pp$ below nominal value. This problem is still severe for the case $n=1{,}000$, as rates are still approximately $10pp$ too low. Results are more satisfactory for the two highest sample sizes. An explanation for this result can be found in gao2008estimation (gao2008estimation, Remark 3.3) who assert that the effective sample size for the estimation of ES is solely $n\alpha$. All in all, we conclude that larger sample sizes are needed to obtain acceptable coverage probabilities.
Table (ref) considers an increase in the interval's nominal value from $90\%$ to $95\%$. Once again we conclude that the results are qualitatively similar to the benchmark. The average lengths of the intervals are larger for every sample size considered. These results are to be expected for bootstrap intervals with higher nominal value.
It might be of interest to compare the results for the conditional ES with those reported in beutner2018residual for the conditional VaR. They find that the EP interval performs worse than the RT interval in small samples, which is in line with the theoretical findings in falk1991coverage. This result does not carry over to the conditional ES, since in most instances (except Table (ref)) the EP interval even outperforms the RT interval. To make a full comparison, we also computed the results where the DGP is a T-GARCH($1,1$). Results are vastly comparable and available upon request.
To summarize, the simulation study suggests that the fixed-design bootstrap works well in terms of average coverage. In comparison to the conditional VaR, higher sample sizes are necessary to obtain coverage rates close to the nominal value. There is no clear evidence to have a preference for any of the three intervals based on the simulation results in all different settings. This directly contrasts results for the conditional VaR in beutner2018residual for which the RT bootstrap interval is found superior.
This paper studies the two-step estimation procedure of francq2015risk associated with the conditional ES. In the first step, the conditional volatility parameters are estimated by QMLE, while the second step corresponds to the estimation of conditional ES based on the first-step residuals. We find that the estimators of the parameters that comprise the conditional ES have a joint asymptotic distribution that does not depend on any density. This is in direct contrast with the conditional VaR estimator for which density estimation is required. A fixed-design residual bootstrap method is proposed to mimic the finite sample distribution of the two-step estimator and its consistency is proven under mild assumptions. In addition, an algorithm is provided for the construction of bootstrap intervals for the conditional ES to take into account the uncertainty induced by estimation. Three interval types are suggested and a simulation study is conducted to investigate their performance in finite samples. Firstly, we find that average coverage rates of all intervals are close to nominal value, except when sample size is low. Secondly, we find that there is no clear evidence that any of the proposed intervals outperforms the others. This contrasts the results in beutner2018residual for the conditional VaR, who find superiority of the reversed-tails bootstrap interval.
The present work can be extended by developing a bootstrap procedure for the one-step approach estimator of francq2015risk. This suggestion is left for future research.