Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
113,992 characters · 12 sections · 128 citation commands
Journal of Econometrics
\begingroup \footnote{ $^{\textcolor{white}{**}*}$Department of Econometrics and Data Science, Vrije Universiteit Amsterdam, De Boelelaan 1105 1081 HV Amsterdam, Netherlands. E-mail address: [email removed]\\ $^{\textcolor{white}{*}**}$a.s.r. Archimedeslaan 10, 3584 BA Utrecht. E-mail address: [email removed]\\ $^{***}$Department of Quantitative Economics, Maastricht University, Tongersestraat 53, 6211 LM Maastricht, Netherlands. E-mail address: [email removed] (corresponding author) } \addtocounter{footnote}{-1} \endgroup
\doublespacing
Risk management has tremendously developed in past decades becoming an increasing practice. With minimum capital requirements being enforced by current legislation (Basel III and Solvency II), financial institutions and insurance companies monitor risk by using conventional measures such as Value-at-Risk (VaR). Typically, the volatility dynamics are specified by a (semi-)parametric model leading to conditional risk measure versions. For GARCH-type models the conditional VaR reduces to the conditional volatility scaled by a quantile of the innovations' distribution. The latter is conventionally treated as additional parameter and forms together with the others the risk parameter francq2015risk. The true parameters are generally unknown and need to be estimated to obtain an estimate for the conditional VaR. Clearly, this VaR evaluation is subject to estimation risk that needs to be quantified for appropriate risk management.
Whereas an estimator based on a single step is available after re-parameterization francq2015risk, a widely used approach is the following two-step estimation procedure. First, the parameters of the stochastic volatility model are estimated. Arguably the most popular estimation method in a GARCH-type setting is the Gaussian quasi-maximum-likelihood (QML) method. Based on the model's residuals the quantile is estimated by its empirical counterpart in a second step. For realistic sample sizes (e.g.\ $500$ or $1{,}000$ daily observations) the estimators are subject to considerable estimation risk. In particular, the estimation uncertainty associated with the quantile estimator is substantial for extreme quantiles (e.g.\ $\leq 5\%$).
To quantify the uncertainty around the point estimators, one traditionally relies on asymptotic theory while replacing the unknown quantities in the limiting distribution by consistent estimates. An alternative approach -- frequently employed in practice -- is based on a bootstrap approximation. Regarding the estimators of the GARCH parameters, various bootstrap methods have been studied to approximate the estimators' finite sample distribution including the subsample bootstrap hall2003inference, the block bootstrap corradi2008bootstrap, the wild bootstrap shimizu2009bootstrapping and the residual bootstrap. The residual bootstrap method is particularly popular and can be further divided into recursive pascual2006bootstrap,hidalgo2007goodness,jeong2017residual and fixed shimizu2009bootstrapping,cavaliere2018fixed design. Whereas in the former the bootstrap observations are generated recursively using the estimated volatility dynamics, the latter design keeps the dynamics of the bootstrap samples fixed at the value of the original series. Further recent applications of fixed or recursive bootstrap designs or variants thereof to conditional volatility models can be found in HETLAND2021, FRANCQ202247 and cavaliere2020bootstrap.
The estimation of the quantile and the conditional VaR have received only selected attention in the bootstrap literature and proposed bootstrap methods have been, to the best of our knowledge, exclusively investigated by means of simulation. christoffersen2005estimation examine various quantile estimators and construct intervals for the conditional VaR using a recursive-design residual bootstrap method. In addition, hartz2006accurate presume the innovation distribution to be standard normal such that the quantile parameter is known; they propose a resampling method based on a residual bootstrap and a bias-correction step to account for deviations from the normality assumption. In contrast, spierdijk2016confidence develops an $m$-out-of-$n$ without-replacement bootstrap to construct confidence intervals for ARMA-GARCH VaR.
This paper proposes a fixed-design residual bootstrap method to mimic the finite sample distribution of the two-step estimator and provides an algorithm for the construction of bootstrap intervals for the conditional VaR. The proposed bootstrap method is proven to be consistent for a general class of volatility models. In particular, our framework does not only encompass GARCH but also several GARCH extensions such as the threshold GARCH (T-GARCH) of zakoian1994threshold and the GJR-GARCH named after Glosten, Jagannathan and Runkle (glosten1993relation). The bootstrap consistency is established under a set of mild assumptions, which relaxes moment conditions on the innovations imposed in the GARCH bootstrap literature. To the best of our knowledge this paper is the first to theoretically validate the residual bootstrap for the quantile and the conditional VaR.
The remainder of the paper is organized as follows. Section (ref) specifies the model and the conditional VaR is derived. The two-step estimation procedure is described in Section (ref) and asymptotic theory is provided under mild assumptions. In Section (ref), a fixed-design residual bootstrap method is proposed and proven to be consistent. Further, bootstrap intervals are constructed for the conditional VaR and extensions to the bootstrap methods presented here are discussed. A simulation study is conducted in Section (ref) and an empirical application illustrates the interval estimation based on the fixed-design residual bootstrap. Section (ref) concludes. Appendix (ref) contains proof of the main results, whereas Supplementary Appendix (ref) contains auxiliary results and their proofs. Finally, Supplementary Appendix (ref) is devoted to the related recursive-design residual bootstrap and Supplementary Appendix (ref) contains additional simulation results.
We consider a conditional volatility model of the form
with $t\in \mathbb{Z}$, where $\{\epsilon_t\}$ denotes the sequence of log-returns, $\{\sigma_t\}$ is a volatility process and $\{\eta_t\}$ is a sequence of independent and identically distributed (iid) variables satisfying $\mathbb{E}\big[\eta_t^2\big]=1$. The volatility is presumed to be a measurable function of past observations
with $\sigma:\mathbb{R}^\infty\times \Theta\to(0,\infty)$ and $\theta_0$ denotes the true parameter vector belonging to the parameter space $\Theta \subset \mathbb{R}^r$, $r \in \mathbb{N}$. Subsequently, we consider two examples for the functional form of (ref): the well-known GARCH model (engle1982autoregressive, engle1982autoregressive; bollerslev1986generalized, bollerslev1986generalized) and the T-GARCH model of zakoian1994threshold. Whereas the first is frequently applied in practice, the second is motivated by our empirical application (see Section (ref)).
Throughout the paper, for any cumulative distribution function (cdf), say $G$, we define the generalized inverse by $G^{-1}(u)=\inf\big\{\tau \in \mathbb{R} : G(\tau)\geq u\big\}$ and write $G(\cdot-)$ to denote its left limit. Generally, for an arbitrary real-valued random variable $X$ (e.g. stock return) with cdf $F_X$, the VaR at level $\alpha \in (0,1)$, is given by $VaR_\alpha(X)=-F_X^{-1}(\alpha)$.
Let $\mathcal{F}_n$ denote the $\sigma$-algebra generated by $\{\epsilon_t, t\leq n\}$. It follows that the conditional VaR of $\epsilon_{n+1}$ given $\mathcal{F}_n$ at level $\alpha \in (0,1)$ is $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n) = \sigma (\epsilon_n,\epsilon_{n-1},\dots;\theta_0) VaR_\alpha(\eta_{n+1})$. For given $\alpha$, the quantile of $\eta_{n+1}$ is constant and can be treated as a parameter. Thus, denoting the cdf of the $\eta_t$'s by $F$ and setting $\xi_\alpha = F^{-1}(\alpha)$, the conditional VaR of $\epsilon_{n+1}$ given $\mathcal{F}_n$ at level $\alpha$ reduces to
Typically, $\alpha$ is fixed at a sufficiently small level such that $\xi_\alpha<0$. Except for special cases (e.g. normality of $\eta_t$), $\xi_\alpha$ is unknown and needs to estimated just like $\theta_0$.
We estimate the parameters $\theta_0$ and $\xi_\alpha$ following the two-step procedure of francq2015risk (francq2015risk, Section 4.2). In the first step, we estimate the conditional volatility parameter $\theta_0$ by Gaussian QML. This approach is motivated as follows: if the innovations $\{\eta_t\}$ were Gaussian, the variables $\eta_t(\theta) = \epsilon_t/\sigma_t(\theta)$ would be iid $N(0,1)$ whenever $\theta=\theta_0$, where $\sigma_{t}(\theta) = \sigma(\epsilon_{t-1},\dots,\epsilon_{1}, \epsilon_{0},\epsilon_{-1},\dots;\theta)$. The 'Q' in QML stands for 'quasi' and refers to the fact that $F$ does not need to be the standard normal distribution function. Obviously, given a sample $\epsilon_1, \dots, \epsilon_n$, we generally cannot determine $\sigma_{t}(\theta)$ completely. Replacing the unknown presample observations by arbitrary values, say $\tilde{\epsilon}_t$, $t\leq 0$, we obtain $\tilde{\sigma}_{t}(\theta) = \sigma(\epsilon_{t-1},\dots,\epsilon_{1}, \tilde{\epsilon}_{0},\tilde{\epsilon}_{-1},\dots;\theta)$, which serves as an approximation for $\sigma_{t}(\theta)$. The QML estimator of $\theta_0$ is defined by
In the second step, we estimate $\xi_\alpha$ on the basis of the first-step residuals, i.e. $\hat{\eta}_t=\epsilon_t/\tilde{\sigma}_t(\hat{\theta}_n)$. The empirical $\alpha$-quantile of $\hat{\eta}_1,\dots,\hat{\eta}_n$ is given by
where $\rho_\alpha(u)=u(\alpha-\mathbbm{1}_{\{u< 0\}})$ is the usual asymmetric absolute loss function (cf.\ koenker2006quantile, koenker2006quantile). Equivalently, we can write $\hat{\xi}_{n,\alpha}=\hat{\mathbbm{F}}_n^{-1}(\alpha)$ with $\hat{\mathbbm{F}}_n(x)=\frac{1}{n}\sum_{t=1}^n \mathbbm{1}_{\{\hat{\eta}_t\leq x\}}$ being the empirical distribution function (edf) of the residuals.
Having obtained estimators for $\theta_0$ and $\xi_\alpha$, we turn to the estimation of the conditional VaR of the one-period ahead observation at level $\alpha$. Whereas the notation $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ stresses the object's conditional nature, we henceforth proceed with the abbreviation $VaR_{n,\alpha}$ for notational convenience. Employing (ref) -- (ref) we can estimate $VaR_{n,\alpha}$ by
Clearly, the estimator's large sample properties cannot be studied using traditional tools such as consistency since (ref) does not permit a limit.
For the subsequent asymptotic analysis, we introduce the following assumptions.
The previous set of assumptions is comparable to the conditions imposed by francq2015risk. Assumption (ref) calls for a correct specification of the volatility structure. If the researcher incorrectly specifies a volatility function $\varsigma(\dots;\vartheta)$ instead, the estimator of the misspecified conditional volatility model $\hat{\vartheta}_n$ will converge to a pseudo-true value, i.e.\ $\vartheta_0=\arg \min_\vartheta \mathbb{E}\big[\frac{1}{2}\frac{\epsilon_t^2}{\varsigma_t^2(\vartheta)}+\log \varsigma_t(\vartheta)\big]$. The corresponding edf of the residuals $\frac{1}{n}\sum_{t=1}^n \mathbbm{1}_{\{\epsilon_t/\varsigma_t(\hat{\vartheta}_n)\leq x\}}$ converges to $\bar{F}(x)=\mathbb{E}\big[F\big(x\frac{\varsigma_t(\vartheta_0)}{\sigma_t(\theta_0)}\big)\big]$ in view of Lemma (ref) while the $\alpha$-quantile estimator converges to $\bar{F}^{-1}(\alpha)$, which is generally different from $F^{-1}(\alpha)$. Thus, the correct specification of the volatility function is crucial and one can test for it using the recently developed test by Maria2019; for further recent results on goodness-of-fit testing for GARCH models see Bardet2020. Regarding the bootstrap method to be developed below, it is plausible to expect that in the presence of a misspecified conditional volatility model it will be consistent for the pseudo-true values, although we do not provide a rigorous proof.
Regarding the innovation process we do not need to assume $\mathbb{E}[\eta_t]=0$ (cf.\ francq2004maximum, francq2004maximum, Remark 2.5). The iid condition in Assumption (ref)(ref) is vital for (ref) to hold and is the basis of the residual bootstrap in Section (ref). Under correct specification of the volatility process the iid assumption imposed on the innovations can be tested for by considering the errors $\epsilon_t/ \sigma(\epsilon_{t-1},\epsilon_{t-2},\dots;\hat{\theta}_n)$, $t=1,..,n$, and applying the test of cho2011generalized.
Whereas cavaliere2018fixed assume the existence of the sixth moment of $\eta_t$ for the fixed-design bootstrap in ARCH($q$) models, we only require the fourth moment to be finite in Assumption (ref)((ref)). We assume $\theta_0$ belongs to the interior of the parameter space in Assumption (ref). Parameters on the boundary yield non-standard problems, which require special treatment. cavaliere2020bootstrap provide bootstrap inference on the boundary of the parameter space with application to conditional volatility models. We return to this issue in our empirical application in Remark (ref). In Assumption (ref) the function $\sigma(x_1,x_2,\dots;\theta)$ is presumed to be monotonically increasing in $\theta$, which is used to establish the strong consistency of the quantile estimator. While the monotonicity condition is a feature shared by various stochastic volatility models (cf.\ berkes2003limit, berkes2003limit, Lemma 4.1), it excludes the exponential GARCH nelson1991conditional and the log-GARCH geweke1986comment,pantula1986modeling. Further, we require higher order of moments in Assumption (ref) for the bootstrap, which does not seem to be restrictive for the classical GARCH-type models (cf.\ francq2011garch, francq2011garch, p.\ 165; hamadeh2011asymptotic, hamadeh2011asymptotic, p.\ 501). In particular, Assumption (ref) is presumed to hold with $a=\pm 12$, $b=12$ and $c=6$ for establishing the convergence of the bootstrap information matrix.
On the basis of the previous assumptions we extend the strong consistency result of francq2015risk (francq2015risk, Theorem 1) to the quantile estimator.
To lighten notation, we henceforth write $D_t(\theta) =\frac{1}{\sigma_t(\theta)}\frac{\partial\sigma_t(\theta)}{\partial \theta}$ and drop the argument when evaluated at the true parameter, i.e.\ $D_t=D_t(\theta_0)$. The next result provides the joint asymptotic distribution of $\hat{\theta}_n$ and $\hat{\xi}_{n,\alpha}$ and is due to francq2015risk.
In applied work, to use Theorem (ref) in order to conduct inference for $(\theta_0, \xi_{\alpha})$ one needs a consistent estimator for $\Sigma_{\alpha}$. It follows from Lemma (ref) that $\kappa$, $\Omega$, and $J$ can be consistently estimated by
respectively, with $\hat{D}_t = \tilde{D}_t(\hat{\theta}_n)$ and $\tilde{D}_t(\theta) = \frac{1}{\tilde{\sigma}_t(\theta)}\frac{\partial\tilde{\sigma}_t(\theta)}{\partial \theta}$. To estimate $\lambda_{\alpha}$ and $\zeta_{\alpha}$, which also appear in $\Sigma_{\alpha}$, one can use $\hat{\xi}_{n,\alpha}$ for $\xi_{\alpha}$ and
for $p_{\alpha}$ (see also Lemma (ref) for the properties of $\hat{p}_{n,\alpha}$). To estimate the density $f$, which also appears in $\lambda_{\alpha}$ and $\zeta_{\alpha}$, kernel smoothing is commonly employed, i.e.\
with kernel function $k$ and bandwidth $h_n>0$. Hence, whenever we use a consistent $\hat{\mathbbm{f}}_n^S$ for $f$ we obtain, combined with Equations (ref) and (ref), a consistent estimator for $\Sigma_\alpha$ denoted by $\hat{\Sigma}_{n,\alpha}$. For the case that $\epsilon_t$ in Equation ((ref)) follows a GARCH$(p,q)$ process, gao2008estimation considered estimating $f$ by a kernel density estimator with Lipschitz-continuous kernels such as $k(x)=\phi(x)$, where $\phi$ is the standard normal density function. An alternative estimator is based on the uniform kernel $k(x)=\frac{1}{2}\mathbbm{1}_{\{|x|\leq 1\}}$ yielding $\hat{\mathbbm{f}}_n^S(\hat{\xi}_{n,\alpha})\overset{p}{\to}f(\xi_\alpha)$ whenever $h_n \sim n^{-\varrho}$ for some $\varrho \in (0,1/2]$.
At this point it is worth recalling that interest lies in $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ (cf. Equation ((ref))). More precisely, one is not so much interested in the distribution of this random variable or its moments but in inference for its possible realizations. Note also that $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ varies with $n$ and does not converge. This illustrates that inference for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ is different from inference for $(\xi_{\alpha},\theta_0)$ which is just a real-valued vector of dimension $r+1$. Nevertheless, it is intuitively clear how to conduct inference for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$, yet this difference results in two technical issues that need to be addressed to provide a thorough theoretical justification of inferential procedures for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$. These technical issues are known for quite some time; see, for instance, phillips1979sampling, kreiss2015discussion and the textbook pesaran2015time[p. 389]. Several examples illustrating these two issues can be found in beutner2017justification. In a nutshell the first issue stems from the fact that, as mentioned, $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ varies over time (cf. Eq. ((ref))) implying that a limiting distribution cannot exist. The second issue is a result of $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ being a conditional quantity which on the one hand requires conditioning on the sample observed so far, yet on the other hand this conditioning eliminates all randomness, making it impossible to establish useful distributional results. We now briefly illustrate this here for the case that $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ is the object of interest; a very detailed illustration of this second issue by means of a GARCH(1,1) can be found in Example 2.1 in beutner2017justification. Fixing some arbitrary starting values $\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots$ and making the dependency of $\sigma_{n+1}$ on the realized values of the random variables $\epsilon_n,\epsilon_{n-1},\ldots,\epsilon_1$ and the starting values $\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots$ explicit a feasible version of the quantity of interest is
where $\epsilon_n^r,\epsilon_{n-1}^r,\ldots,\epsilon_{1}^r$ denote the realized values of $\epsilon_n,\epsilon_{n-1},\ldots,\epsilon_1$. Here feasible is in the sense that only parameter uncertainty remains. Now replacing, as at the beginning of Section 3, the unknown $\theta_0$ by the estimator $\hat{\theta}_{n}$ and doing the same with $\xi_{\alpha}$ we see that these estimators must enter Equation (ref) given the realizations, i.e. as $\hat{\theta}_{n}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$ and $\hat{\xi}_{\alpha,n}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$, respectively, because otherwise we would end up with an inconsistency. Indeed, if they entered as $\hat{\theta}_n(\epsilon_1,\epsilon_{2},\ldots,\epsilon_n)$ and $\hat{\xi}_{n,\alpha}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_n)$, respectively, we would treat $\epsilon_1,\epsilon_{2},\ldots,\epsilon_n$ as random and observed at the same time. Upon replacing $\theta_0$ and $\xi_{\alpha}$ in Equation (ref) by $\hat{\theta}_{n}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$ and $\hat{\xi}_{n,\alpha}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$, respectively, no random variable appears in
which is just the estimator of Equation (3.3) with all dependencies made explicit. As just observed no random variable appears in $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$ which therefore does not have a distribution which could be used to construct confidence intervals for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$. In the article beutner2017justification merging is proposed as a solution to overcome the first issue and sample splitting as a means to solve the second. The results of this article can be used justify, through their Theorem 1, approximating the distribution of the VaR estimator, centered at $VaR_{n,\alpha}$ and inflated by $\sqrt{n}$, by
and consequently to justify the intuitive (conditional) confidence interval
for ${VaR}_{n,\alpha}$. Here $\Phi$ denotes the standard normal cdf, and we used again the short-hand notation for $\tilde{\sigma}_{n+1}(\hat{\theta})$ and $\hat{\xi}_{n,\alpha}$, i.e. did not make the dependency on the starting values etc. explicit. Note that the interval (ref) lacks a proper theoretical justification as it treats the non-random $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$ as random because it assigns an asymptotic distribution to it. In order to use Theorem 3 and Corollary 2 of beutner2017justification to provide a sound theoretical justification of (ref) let $n_E: \mathbb{N} \to \mathbb{N}$ and $n_P: \mathbb{N} \to \mathbb{N}$ be such that for all $n \in \mathbb{N}$ we have $n_E(n) < n_P(n)$. First, in the sample split approach we will only use $\epsilon_1,\ldots,\epsilon_{n_E}$ to estimate $\xi_{\alpha}$ and $\theta_0$ which we denote by $\hat{\xi}_{n_E,\alpha}^{SPL}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_{n_E})$ and $\hat{\theta}_{n_E}^{SPL}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_{n_E})$, respectively. To lighten the notation a bit we will also use the shorter notations $\hat{\xi}_{n_E,\alpha}^{SPL}$ and $\hat{\theta}_{n_E}^{SPL}$, respectively. Second, in the sample split approach conditioning on $\mathcal{F}_n$ on the left-hand side of (ref) is replaced by conditioning on $\epsilon_{n_P},\ldots,\epsilon_n$ only. The sample split version of (3.3) which we denote by $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{SPL}$ becomes
where $\epsilon_n^r,\epsilon_{n-1}^r,\ldots,\epsilon_{n_P}^r$ and $\tilde{\epsilon}_{0},\tilde{\epsilon}_{-1},\ldots$ are as before and $c_{n_P-1},\ldots,c_{1}$ are constants. These constants could be viewed as starting values but since they are not replacing unobserved value as $\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots$ do we prefer to use a different letter for them. Note also that $\epsilon_1,\ldots,\epsilon_{n_E}$ enter the sample split version $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{SPL}$ as random variables which is in contrast to $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$. Therefore, $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{SPL}$ does have a non-degenerate distribution which can be used to construct confidence intervals. The sample split (conditional) confidence interval is defined as
Here $\hat{\Sigma}_{n_E,\alpha}^{SPL}$ is as $\hat{\Sigma}_{n,\alpha}$ but based on the first $n_E$ observations only. Note that this interval is meaningful because $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{SPL}$ is random and converges in distribution after centering and scaling. Intuitively, we would say that the statistically meaningful interval in (ref) provides a theoretical justification for the interval (ref) if
and if
converges to
This is also the concept employed in beutner2017justification. Under the conditions of Theorem 3 of this article the convergence in (ref) holds and (ref) converges to (ref) by Corollary 2 of this article. It is worth pointing out that convergence is in probability and that for this concept to be applicable $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$ and (ref) are treated as random. Theorem 2 above ensures that one of the assumptions of Corollary 3 of beutner2017justification holds so that it provides the basis for a sound theoretical justification of (ref) as a (conditional) confidence interval. Formally, we can state
Although the interval in (ref) is, as just outlined, well-justified, it may perform poorly since the density estimation appears rather sensitive regarding the choice of bandwidth (see gao2008estimation, gao2008estimation, Section 4). Bootstrap methods offer an alternative way to quantify the uncertainty around the estimators.
Bootstrap approximations frequently provide better insight into the actual distribution than the asymptotic approximation, yet they require a careful set-up. hall2003inference show that conventional bootstrap methods are inconsistent in a GARCH model lacking finite fourth moment in the case of the squared innovations' distribution not being in the domain of attraction of the normal distribution. They consider a subsample bootstrap instead and study its asymptotic properties. In correspondence, an $m$-out-of-$n$ without-replacement bootstrap is proposed by spierdijk2016confidence to construct confidence intervals for ARMA-GARCH VaR.
pascual2006bootstrap present a residual bootstrap in a GARCH($1,1$) setting and assess its finite sample properties by means of simulation. Their bootstrap scheme follows a recursive design in which the bootstrap observations are generated iteratively using the estimated volatility dynamics. Building upon their results, christoffersen2005estimation construct bootstrap confidence intervals for (conditional) VaR and Expected Shortfall and compare them to competitive methods within the GARCH($1,1$) model. Theoretical results on the recursive-design residual bootstrap are provided by hidalgo2007goodness and jeong2017residual for the ARCH($\infty$) and GARCH($p,q$) model, respectively.
In contrast, shimizu2009bootstrapping considers fixed-design variants of the wild and the residual bootstrap in which the ARMA-GARCH dynamics of the bootstrap samples are kept fixed at the values of the original series. The bootstrap estimators are based on a single Newton-Raphson iteration simplifying the proofs of first-order asymptotic validity. shimizu2009bootstrapping's approach for the residual bootstrap is also employed in a multivariate GARCH setting by francq2016variance. Recently, cavaliere2018fixed study the fixed-design residual bootstrap in the context of ARCH($q$) models and propose a bootstrap Wald statistic based on a QML bootstrap estimator. While their theory has been developed independently to ours, their simulation study indicates that the fixed-design bootstrap performs as well as the recursive-design bootstrap.
We propose a fixed-design residual bootstrap procedure, described in Algorithm (ref), to approximate the distribution of the estimators in (ref) -- (ref).
Proposition (ref) establishes the asymptotic validity of the bootstrap for the volatility parameters.
Establishing the asymptotic validity of the bootstrap for the second part appears challenging since the bootstrap innovations are drawn from the discrete distribution $\hat{\mathbbm{F}}_n$. To overcome this issue we rely on arguments employed by bahadur1966note and berkes2003limit. The following theorem states the paper's main result.
Theorem (ref) is useful to validate the bootstrap for the conditional VaR estimator. For the asymptotic behavior of the conditional VaR estimator we refer to (ref) and the text around it. The following corollary is established.
Clearly, the VaR evaluation in (ref) is subject to estimation risk that needs to be quantified. We propose the following algorithm to obtain approximately $100(1-\gamma)\%$ confidence intervals.
The interval in (ref) is obtained by the EP method, that is frequently encountered in the bootstrap literature. It is obtained from the (typically) infeasible equal-tailed confidence interval
where $G_{n}^{-1}$ is the (unknown) quantile function of $\sqrt{n}( \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}-{VaR}_{n,\alpha})$, which is replaced by its bootstrap analogue $\hat{G}_{n,B}^{*-1}$. The same reasoning leads to the SY interval but with test statistic $\sqrt{n}|( \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}-{VaR}_{n,\alpha})|$ instead of $\sqrt{n}( \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}-{VaR}_{n,\alpha})$ which makes it also clear that the interval in (ref) presumes symmetry for rationalizing its construction. “Flipping around" the tails of the ET interval leads to the RT interval given in (ref). Clearly, the RT and the EP have equal length. Whereas (ref) in its current form emphasizes the interval's name, RT type intervals are frequently reported in their reduced form, i.e. the lower and upper bound of (ref) simplify to the $\gamma/2$ and $1-\gamma/2$ quantiles of $\frac{1}{B}\sum_{b=1}^B \mathbbm{1}_{\big\{ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{*(b)}\leq x\big\}}$, respectively. RT intervals can either be motivated by the results of falk1991coverage\footnote{In a random sample setting falk1991coverage prove that the RT bootstrap interval for quantiles has asymptotically greater coverage than the corresponding EP bootstrap interval. For additional insights we refer to hall1988bootstrap.} or as the bootstrap analogue of the (uncentered) statistic $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$. It is worth mentioning that RT type bootstrap intervals for the VaR are also constructed in reduced form by christoffersen2005estimation. Regardless of whether we use an EP, RT or SY interval the meaning is always the same: Given the past up to and including time $n$ the probability that the conditional VaR for period $n+1$ is contained in the intervals is approximately equal to $100(1-\gamma)\%$.
The asymptotic normality result in Theorem (ref) as well as the bootstrap consistency in Theorem (ref) are derived, inter alia, under the assumption that the innovations are iid. In case this is not believed to be true -- e.g. if the suggested specification tests mentioned in Section (ref) indicate otherwise -- asymptotic normality of $\sqrt{n}(\hat{\theta}_n-\theta_0)$ can still be established under regularity assumptions. Escanciano2009 studies the QML estimator under some dependence among the $\eta_t$'s while imposing slightly stronger (moment) conditions, whereas the related paper of Linton2010 investigates estimators in a GARCH(1,1) with dependent errors but under weaker moment conditions. A multivariate version of the dependence condition in Escanciano2009 can be found in FrancqZakoian2016.
The bootstrap method presented in Algorithm (ref) is contingent on the iid assumption. In general, one cannot expect bootstrap procedures that are based on an iid assumption to be robust against deviations from this assumption; for a well-known example, one may refer to GONCALVES2004. Alternative bootstrap techniques may be used if the iid condition is thought to be unrealistic. A variety of bootstrap methods exist that can capture dependence and non-identical random variables; see e.g. Lahiri03 for a broad overview. The wild or multiplier bootstrap Mammen93, DavidsonFlachaire08 is particularly suited for dealing with non-identical variables, but does not capture dependence, unless it is properly modified Shao10, FSU20. However, it remains an open question which bootstrap method combined with the fixed design approach leads to valid bootstrap procedures.
In order to evaluate the finite sample performance of the proposed bootstrap procedure a Monte Carlo experiment is conducted. We confine ourselves to four conditional volatility specifications related to Examples (ref) and (ref) in Section (ref). The first two are GARCH($1,1$) parameterizations with
which are similar to the specifications of gao2008estimation (gao2008estimation, Section 4) and spierdijk2016confidence (spierdijk2016confidence, Section 4.2). In addition, we study two T-GARCH($1,1$) scenarios likewise associated with high and low persistence:
Within the experiment the VaR level takes value $\alpha =0.05$ and there are two possible innovation distributions: a Student-$t$ distribution with $6$ degrees of freedom (df) and the standard normal distribution.\footnote{The Student-t innovations are appropriately standardized to satisfy $\mathbb{E}\eta_t^2=1$.} We consider four estimation sample sizes, $n \in \{250; 500; 1{,}000; 5{,}000\}$, whereas the number of bootstrap replicates is fixed and equal to $B=999$. For each model version we simulate $S=10{,}000$ independent Monte Carlo trajectories.
The numerical optimization of the log-likelihood function is carried out employing the build-in function fmincon and running time is reduced by parallel computing using parfor. The code is available on the website of the third author, and simulation results for a VaR level of $\alpha=0.01$ can be found in the working paper version.
\captionsetup[table]{labelformat=simple, labelsep=space}
\bgroup
\egroup
\captionsetup[table]{labelformat=simple, labelsep=space}
\bgroup
\egroup
Table (ref) and (ref) report the results of the three $90\%$--bootstrap intervals for the $5\%$--VaR when the innovation distribution is Student-t (henceforth referred to as baseline) and the model is a GARCH(1,1) and a T-GARCH(1,1), respectively. In both tables the results of the interval (ref) based on asymptotic (AS) theory are included for comparison, where a Gaussian kernel is utilized together with a bandwidth following silverman1986density's (silverman1986density) rule-of-thumb. In the GARCH($1,1$) high persistence case (right part of Table 1), we see that the average coverage varies around $90\%$ across all sample sizes for the RT and the SY interval. In contrast, the EP and the AS interval fall short of the nominal $90\%$ by $10.43$ and $4.15$ percentage points (pp), respectively, for small sample size ($n=250$). Nevertheless, their average coverage approaches the nominal value as the sample size increases. Remarkably, for all four intervals the average rate of the conditional VaR being below the interval is considerably less than the average rate of the conditional VaR being above the interval when the sample size is rather small ($n \leq 500$). Regarding the intervals' length, we observe that the SY interval is on average larger than the EP/RT interval. As the sample size increases this gap diminishes and the intervals' average lengths shrink. Considering the low persistent case (left part of Table (ref)) we find similar results regarding the intervals' average coverage, yet their average lengths turn out to be smaller compared to the high persistent case. This is intuitive as the conditional volatility tends to vary less in the low persistent case. Regarding the T-GARCH($1,1$) in Table (ref), the overall picture is similar as in the GARCH(1,1) case, however the under-coverage in small and medium-sized samples appears to be more extreme for the EP and reduced for the AS interval.
Simulation results for the scenario when the $\eta_t$'s follow a standard normal distribution and when the model is a GARCH(1,1) and a T-GARCH(1,1), respectively, are tabulated in Tables (ref) and (ref) which are given in Appendix (ref). Here we only note that, although the error distribution underlying the QMLE is correctly specified in this case, the qualitative results stated above with regard to Table (ref) persist: the RT and the SY intervals possess accurate coverage rates across sample sizes, whereas the EP and the AS interval exhibit under-coverage in samples of rather small size with different extent. Moreover, we observe that the intervals are on average shorter in the Gaussian case than in the baseline case. This seems partially driven by a smaller variance of $\hat{\xi}_{n,\alpha}$; for $\alpha=0.05$ the asymptotic variance $\zeta_\alpha$ in (ref) is equal to $3.11$ in the Gaussian case compared to $5.72$ in the Student-t case with $6$ degrees of freedom.
While the small-sample-performance of the AS interval can be explained by its embodied density estimation, the question arises why the EP interval performs worse than the other bootstrap intervals, which seems counter-intuitive at first. Howbeit the results are in line with the theoretical findings of falk1991coverage (falk1991coverage, unnumbered Corollary, p.\ 488). In a random sample setting they prove that the RT bootstrap interval for quantiles has asymptotically greater coverage than the corresponding EP bootstrap interval. The emerging gap\footnote{We neglect their $o(n^{-1/2})$ term. Take note that the theoretical results of falk1991coverage are not directly applicable in our setting due to GARCH-type effects.}
Table (ref) presents the average coverage gap between the EP and the RT bootstrap interval of the conditional VaR for the baseline specification as well as for three deviations from the baseline (change in $F$, $\alpha$ and $\gamma$). For example, in the low persistence GARCH($1,1$) case of the baseline with $n=250$, the average coverage gap amounts to $90.73\%-81.20\%=9.53$pp (see also Table (ref)). It is striking that all values are positive within Table (ref), which highlights the superiority of the RT bootstrap interval over the EP bootstrap interval. Further, it is eminent that average coverage gap tends to decrease with increasing sample size, which supports (i). Comparing columns (1) and (3) we also find that the average coverage gap tends to be larger for the $1\%$--VaR than for the $5\%$--VaR, which gives rise to (ii). Regarding (iii), the result of falk1991coverage (falk1991coverage) suggests that the gap slightly decreases when increasing the nominal coverage from $90\%$ to $95\%$. Such tendency is precisely observed when comparing columns (1) and (4) of Table (ref).
\bgroup
\egroup
\bgroup
\egroup
With regard to Remark (ref) in Section (ref), Tables (ref) and (ref) report the simulation results for the recursive-design bootstrap for the DGPs of Tables (ref) and (ref), respectively. We refer to Appendix (ref) for computational details. In comparison to the fixed-design approach (see Tables (ref) and (ref)) we find that the recursive-design method performs similarly in terms of average coverage for each interval type, which corresponds to the simulation results of cavaliere2018fixed. It is striking, however, that the intervals' average lengths are larger in the recursive-design than in the fixed-design set-up. For example, in the high persistence GARCH(1,1) case (right part of Table (ref)) for $n=500$ the average length in the recursive-design approach is $0.604$ for the EP/RT interval compared to $0.581$ in the fixed-design. As the sample size increases this difference disappears.
In summary, the simulations suggest that the RT and the SY bootstrap interval work well for both bootstrap designs and that they outperform in smaller samples the AS interval in terms of average coverage even though their tails are unequally represented. In contrast, for both bootstrap designs the EP interval falls short of its nominal coverage, which is in line with the theoretical findings of falk1991coverage. Since the fixed RT method leads on average to shorter intervals than the corresponding SY method and its recursive-design counterpart, this suggests to favor the fixed-design RT bootstrap interval in (ref).
We analyze the French stock market index CAC 40 for the period January 1, 2015 -- January 1, 2020. The index values for the period are retrieved from Yahoo Finance and daily (log-) returns (expressed in $\%$) are computed using $\epsilon_t = 100\log(p_t/p_{t-1})$, where $p_t$ denotes the closing value of the index at trading day $t$.
Figure (ref)(a) displays the resulting series of returns. We disregard the observations from July 1, 2019 onwards, which we leave for the out-of-sample evaluation, yielding $n=1{,}146$ remaining observations (i.e.\ January 1, 2015 - July 1, 2019). For the volatility process we consider the T-GARCH($1,1$) model specified in Example (ref).\footnote{We also consider an Asymmetric Power GARCH model ding1993long, i.e.\ $\sigma_{t+1}^\delta= \omega_0 + \alpha_0^+ (\epsilon_{t}^+)^\delta +\alpha_0^-(\epsilon_{t}^-)^\delta + \beta_0 \sigma_{t}^\delta$ with $\delta>0$, which nests the GARCH($1,1$) model ($\delta=2$, $\alpha_0^+=\alpha_0^-$) and the T-GARCH($1,1$) model ($\delta=1$) of Examples (ref) and (ref). In practice, the impact of the power $\delta$ on the volatility is minor and the QML approach of hamadeh2011asymptotic suggests a $\delta$ close to $1$ in favor for the T-GARCH specification.} Table (ref) reports the corresponding point estimates with standard errors obtained by bootstrapping based on Algorithm (ref). As documented in numerous studies we find that the volatility persistence is close to unity. Further, we observe that $\hat{\alpha}^-_n$ is considerably larger than $\hat{\alpha}^+_n$ indicating a strong leverage effect, i.e.\ negative returns tend to increase volatility by more than positive returns of the same magnitude.
Figure (ref)(b) plots the histogram of the residuals with the normal distribution superimposed. Further, we test the condition that the innovations are iid (see Assumption (ref)(ref)) with the generalized run tests of cho2011generalized.\footnote{The implementation of the tests is available on the website of the first author.} These tests are particularly suitable in this case since they can be based on the residuals and are sensitive against a wide range of alternatives. The test statistic of the sup-norm based test is $0.40$, which corresponds to a p-value of $0.27$. consequently, one cannot reject the null hypothesis of iid innovations at any common significance level. Similarly, the generalized run test based on the $L_1$-norm cannot be rejected at a $10\%$ significance level.
Next, we perform a rolling window analysis starting with subperiod January 1, 2015 -- July 1, 2019 and ending with subperiod July 8, 2015 -- January 1, 2020. We have $130$ subperiods each consisting of $1{,}146$ observations. For each rolling window period we fit a T-GARCH($1,1$) model and estimate the one-period-ahead conditional VaR associated with level $\alpha=0.05$. For example, for the first window the T-GARCH($1,1$) estimates are reported in Table (ref) and the conditional $5\%$-VaR of the one-period ahead (i.e.\ July 1, 2019) is estimated by $1.11$. Further, we obtain the associated $95\%$-confidence intervals based on bootstrap and asymptotic normality. In addition to the RT intervals of the fixed- and residual-design bootstrap, we also computed an interval based on the asymptotic distribution. The corresponding intervals are $[0.850,1.136]$ (fixed-design), $[0.834,1.115]$ (recursive-design), and $[0.828,1.106]$ (asymp.\ normality). Although the intervals are fairly similar, the asymptotic and recursive bootstrap intervals are shorter than the fixed-design interval. Given its tendency to underestimate variability in finite samples, this result is unsurprising for the asymptotic interval, although for the recursive bootstrap this contrasts the simulation findings. Note that the fixed-design iid and block bootstraps produce very similar interval, which is not surprising as our conducted specification tests did not indicate any violation of the iid assumption on the innovations.
The results of the rolling window analysis are visualized in Figure (ref). It plots the realized return together with (the opposite of) the estimated conditional VaR. For clarity we only indicate the lower and upper bound of the $95\%$ RT fixed-design bootstrap interval.
We observe that in more turbulent times (e.g.\ August, 2019), the estimated VaR amplifies. In such volatile periods we expect the estimation risk to increase and, accordingly, we find wider bootstrap confidence intervals. In this regard, although not directly related to the bootstrap procedures considered it is worth mentioning that a proposal for monitoring VaR estimates over time can be found in HogaDemetrescu.
In this paper we study the two-step estimation procedure of francq2015risk associated with the conditional VaR. In the first step, the conditional volatility parameters are estimated by QMLE, while the second step corresponds to approximating the quantile of the innovations' distribution by the empirical quantile of the residuals. A fixed-design residual bootstrap method is proposed to mimic the finite sample distribution of the two-step estimator and its consistency is proven under mild assumptions. In addition, an algorithm is provided for the construction of bootstrap intervals for the conditional VaR to take into account the uncertainty induced by estimation. Three interval types are suggested and a large-scale simulation study is conducted to investigate their performance in finite samples. We find that the equal-tailed percentile interval based on the fixed-design residual bootstrap tends to fall short of its nominal value, whereas the corresponding interval based on reversed tails yields accurate average coverage combined with the shortest average length. Although the result seems counter-intuitive at first, it is in line with the theoretical findings of falk1991coverage. In the simulation study we also consider the recursive-design residual bootstrap. It turns out that the recursive-design and the fixed-design bootstrap perform similar in terms of average coverage. Yet in smaller samples the fixed-design scheme leads on average to shorter intervals. Further, the interval estimation by means of the fixed-design residual bootstrap is illustrated in an empirical application to daily returns of the French stock index CAC 40.
Natural extensions of this work are encompassing other risk measures such as Expected Shortfall heinemann2018residual and developing a bootstrap procedure for the one-step estimator of francq2015risk. Further, it is worthwhile to consider a smoothed bootstrap version in the spirit of hall1989smoothing, which offers potential gains in accuracy. The latter two extensions are left for future research, as is the question how to extend the fixed-design bootstrap in order to give valid bootstrap inference when the iid assumption does not hold.
The authors thank Franz Palm, Hanno Reuvers, Jean-Michel Zako\"ian and Christian Francq for useful comments and suggestions as well as Dewi Peerlings and Benoit Duvocelle for computational support during the revision stage. In addition, the authors are grateful to the editors Oliver Linton and Torben Andersen, one of the associate editors and to two anonymous referees for their constructive remarks.
This research was financially supported by the Netherlands Organisation for Scientific Research (NWO).