EconBase
← Back to paper

A Residual Bootstrap for Conditional Value-at-Risk

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

113,992 characters · 12 sections · 128 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Journal of Econometrics

center[center omitted — 103 chars of source]
center[center omitted — 207 chars of source]
abstractA fixed-design residual bootstrap method is proposed for the two-step estimator of francq2015risk associated with the conditional Value-at-Risk. The bootstrap's consistency is proven for a general class of volatility models and intervals are constructed for the conditional Value-at-Risk. A simulation study reveals that the equal-tailed percentile bootstrap interval tends to fall short of its nominal value. In contrast, the reversed-tails bootstrap interval yields accurate coverage. We also compare the theoretically analyzed fixed-design bootstrap with the recursive-design bootstrap. It turns out that the fixed-design bootstrap performs equally well in terms of average coverage, yet leads on average to shorter intervals in smaller samples. An empirical application illustrates the interval estimation.\\ \\ Key words: Residual bootstrap; Value-at-Risk; GARCH\\ JEL codes: C14; C15; C58

\begingroup \footnote{ $^{\textcolor{white}{**}*}$Department of Econometrics and Data Science, Vrije Universiteit Amsterdam, De Boelelaan 1105 1081 HV Amsterdam, Netherlands. E-mail address: [email removed]\\ $^{\textcolor{white}{*}**}$a.s.r. Archimedeslaan 10, 3584 BA Utrecht. E-mail address: [email removed]\\ $^{***}$Department of Quantitative Economics, Maastricht University, Tongersestraat 53, 6211 LM Maastricht, Netherlands. E-mail address: [email removed] (corresponding author) } \addtocounter{footnote}{-1} \endgroup

\doublespacing

Introduction

Risk management has tremendously developed in past decades becoming an increasing practice. With minimum capital requirements being enforced by current legislation (Basel III and Solvency II), financial institutions and insurance companies monitor risk by using conventional measures such as Value-at-Risk (VaR). Typically, the volatility dynamics are specified by a (semi-)parametric model leading to conditional risk measure versions. For GARCH-type models the conditional VaR reduces to the conditional volatility scaled by a quantile of the innovations' distribution. The latter is conventionally treated as additional parameter and forms together with the others the risk parameter francq2015risk. The true parameters are generally unknown and need to be estimated to obtain an estimate for the conditional VaR. Clearly, this VaR evaluation is subject to estimation risk that needs to be quantified for appropriate risk management.

Whereas an estimator based on a single step is available after re-parameterization francq2015risk, a widely used approach is the following two-step estimation procedure. First, the parameters of the stochastic volatility model are estimated. Arguably the most popular estimation method in a GARCH-type setting is the Gaussian quasi-maximum-likelihood (QML) method. Based on the model's residuals the quantile is estimated by its empirical counterpart in a second step. For realistic sample sizes (e.g.\ $500$ or $1{,}000$ daily observations) the estimators are subject to considerable estimation risk. In particular, the estimation uncertainty associated with the quantile estimator is substantial for extreme quantiles (e.g.\ $\leq 5\%$).

To quantify the uncertainty around the point estimators, one traditionally relies on asymptotic theory while replacing the unknown quantities in the limiting distribution by consistent estimates. An alternative approach -- frequently employed in practice -- is based on a bootstrap approximation. Regarding the estimators of the GARCH parameters, various bootstrap methods have been studied to approximate the estimators' finite sample distribution including the subsample bootstrap hall2003inference, the block bootstrap corradi2008bootstrap, the wild bootstrap shimizu2009bootstrapping and the residual bootstrap. The residual bootstrap method is particularly popular and can be further divided into recursive pascual2006bootstrap,hidalgo2007goodness,jeong2017residual and fixed shimizu2009bootstrapping,cavaliere2018fixed design. Whereas in the former the bootstrap observations are generated recursively using the estimated volatility dynamics, the latter design keeps the dynamics of the bootstrap samples fixed at the value of the original series. Further recent applications of fixed or recursive bootstrap designs or variants thereof to conditional volatility models can be found in HETLAND2021, FRANCQ202247 and cavaliere2020bootstrap.

The estimation of the quantile and the conditional VaR have received only selected attention in the bootstrap literature and proposed bootstrap methods have been, to the best of our knowledge, exclusively investigated by means of simulation. christoffersen2005estimation examine various quantile estimators and construct intervals for the conditional VaR using a recursive-design residual bootstrap method. In addition, hartz2006accurate presume the innovation distribution to be standard normal such that the quantile parameter is known; they propose a resampling method based on a residual bootstrap and a bias-correction step to account for deviations from the normality assumption. In contrast, spierdijk2016confidence develops an $m$-out-of-$n$ without-replacement bootstrap to construct confidence intervals for ARMA-GARCH VaR.

This paper proposes a fixed-design residual bootstrap method to mimic the finite sample distribution of the two-step estimator and provides an algorithm for the construction of bootstrap intervals for the conditional VaR. The proposed bootstrap method is proven to be consistent for a general class of volatility models. In particular, our framework does not only encompass GARCH but also several GARCH extensions such as the threshold GARCH (T-GARCH) of zakoian1994threshold and the GJR-GARCH named after Glosten, Jagannathan and Runkle (glosten1993relation). The bootstrap consistency is established under a set of mild assumptions, which relaxes moment conditions on the innovations imposed in the GARCH bootstrap literature. To the best of our knowledge this paper is the first to theoretically validate the residual bootstrap for the quantile and the conditional VaR.

The remainder of the paper is organized as follows. Section (ref) specifies the model and the conditional VaR is derived. The two-step estimation procedure is described in Section (ref) and asymptotic theory is provided under mild assumptions. In Section (ref), a fixed-design residual bootstrap method is proposed and proven to be consistent. Further, bootstrap intervals are constructed for the conditional VaR and extensions to the bootstrap methods presented here are discussed. A simulation study is conducted in Section (ref) and an empirical application illustrates the interval estimation based on the fixed-design residual bootstrap. Section (ref) concludes. Appendix (ref) contains proof of the main results, whereas Supplementary Appendix (ref) contains auxiliary results and their proofs. Finally, Supplementary Appendix (ref) is devoted to the related recursive-design residual bootstrap and Supplementary Appendix (ref) contains additional simulation results.

Model

We consider a conditional volatility model of the form

align[align omitted — 57 chars of source]

with $t\in \mathbb{Z}$, where $\{\epsilon_t\}$ denotes the sequence of log-returns, $\{\sigma_t\}$ is a volatility process and $\{\eta_t\}$ is a sequence of independent and identically distributed (iid) variables satisfying $\mathbb{E}\big[\eta_t^2\big]=1$. The volatility is presumed to be a measurable function of past observations

align[align omitted — 114 chars of source]

with $\sigma:\mathbb{R}^\infty\times \Theta\to(0,\infty)$ and $\theta_0$ denotes the true parameter vector belonging to the parameter space $\Theta \subset \mathbb{R}^r$, $r \in \mathbb{N}$. Subsequently, we consider two examples for the functional form of (ref): the well-known GARCH model (engle1982autoregressive, engle1982autoregressive; bollerslev1986generalized, bollerslev1986generalized) and the T-GARCH model of zakoian1994threshold. Whereas the first is frequently applied in practice, the second is motivated by our empirical application (see Section (ref)).

exampleSuppose $\{\epsilon_t\}$ follows a GARCH$(1,1)$ process given by (ref) and $\sigma_{t}^2 = \omega_0 + \alpha_0 \epsilon_{t-1}^2+ \beta_0 \sigma_{t-1}^2$, where $\theta_0 = (\omega_0,\alpha_0,\beta_0)'\in (0,\infty)\times[0,\infty)\times[0,1)$. The recursive structure implies $\sigma_{t}=\sigma(\epsilon_{t-1},\epsilon_{t-2},\dots;\theta_0) = \sqrt{\sum_{k=1}^\infty \beta_0^{k-1} \big(\omega_0 + \alpha_0 \epsilon_{t-k}^2\big)}$.
exampleSuppose $\{\epsilon_t\}$ follows a T-GARCH$(1,1)$ process given by (ref) and $\sigma_{t} =\: \omega_0 + \alpha_0^+ \epsilon_{t-1}^+ +\alpha_0^- \epsilon_{t-1}^- + \beta_0 \sigma_{t-1}$ with parameters $\theta_0 =(\omega_0,\alpha_0^+,\alpha_0^- ,\beta_0)'\in (0,\infty)\times[0,\infty)\times[0,\infty)\times[0,1)$ and $ \epsilon_{t}^+=\max\{\epsilon_{t},0\}$ and $ \epsilon_{t}^-=\max\{-\epsilon_{t},0\}$. The model's recursive structure yields $\sigma_{t}=\sigma(\epsilon_{t-1},\epsilon_{t-2},\dots;\theta_0) =\! \sum_{k=1}^\infty \beta_0^{k-1} \big(\omega_0 + \alpha_0^+ \epsilon_{t-k}^+ +\alpha_0^-\epsilon_{t-k}^-\big)$.

Throughout the paper, for any cumulative distribution function (cdf), say $G$, we define the generalized inverse by $G^{-1}(u)=\inf\big\{\tau \in \mathbb{R} : G(\tau)\geq u\big\}$ and write $G(\cdot-)$ to denote its left limit. Generally, for an arbitrary real-valued random variable $X$ (e.g. stock return) with cdf $F_X$, the VaR at level $\alpha \in (0,1)$, is given by $VaR_\alpha(X)=-F_X^{-1}(\alpha)$.

Let $\mathcal{F}_n$ denote the $\sigma$-algebra generated by $\{\epsilon_t, t\leq n\}$. It follows that the conditional VaR of $\epsilon_{n+1}$ given $\mathcal{F}_n$ at level $\alpha \in (0,1)$ is $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n) = \sigma (\epsilon_n,\epsilon_{n-1},\dots;\theta_0) VaR_\alpha(\eta_{n+1})$. For given $\alpha$, the quantile of $\eta_{n+1}$ is constant and can be treated as a parameter. Thus, denoting the cdf of the $\eta_t$'s by $F$ and setting $\xi_\alpha = F^{-1}(\alpha)$, the conditional VaR of $\epsilon_{n+1}$ given $\mathcal{F}_n$ at level $\alpha$ reduces to

align[align omitted — 112 chars of source]

Typically, $\alpha$ is fixed at a sufficiently small level such that $\xi_\alpha<0$. Except for special cases (e.g. normality of $\eta_t$), $\xi_\alpha$ is unknown and needs to estimated just like $\theta_0$.

Estimation

We estimate the parameters $\theta_0$ and $\xi_\alpha$ following the two-step procedure of francq2015risk (francq2015risk, Section 4.2). In the first step, we estimate the conditional volatility parameter $\theta_0$ by Gaussian QML. This approach is motivated as follows: if the innovations $\{\eta_t\}$ were Gaussian, the variables $\eta_t(\theta) = \epsilon_t/\sigma_t(\theta)$ would be iid $N(0,1)$ whenever $\theta=\theta_0$, where $\sigma_{t}(\theta) = \sigma(\epsilon_{t-1},\dots,\epsilon_{1}, \epsilon_{0},\epsilon_{-1},\dots;\theta)$. The 'Q' in QML stands for 'quasi' and refers to the fact that $F$ does not need to be the standard normal distribution function. Obviously, given a sample $\epsilon_1, \dots, \epsilon_n$, we generally cannot determine $\sigma_{t}(\theta)$ completely. Replacing the unknown presample observations by arbitrary values, say $\tilde{\epsilon}_t$, $t\leq 0$, we obtain $\tilde{\sigma}_{t}(\theta) = \sigma(\epsilon_{t-1},\dots,\epsilon_{1}, \tilde{\epsilon}_{0},\tilde{\epsilon}_{-1},\dots;\theta)$, which serves as an approximation for $\sigma_{t}(\theta)$. The QML estimator of $\theta_0$ is defined by

align[align omitted — 270 chars of source]

In the second step, we estimate $\xi_\alpha$ on the basis of the first-step residuals, i.e. $\hat{\eta}_t=\epsilon_t/\tilde{\sigma}_t(\hat{\theta}_n)$. The empirical $\alpha$-quantile of $\hat{\eta}_1,\dots,\hat{\eta}_n$ is given by

align[align omitted — 134 chars of source]

where $\rho_\alpha(u)=u(\alpha-\mathbbm{1}_{\{u< 0\}})$ is the usual asymmetric absolute loss function (cf.\ koenker2006quantile, koenker2006quantile). Equivalently, we can write $\hat{\xi}_{n,\alpha}=\hat{\mathbbm{F}}_n^{-1}(\alpha)$ with $\hat{\mathbbm{F}}_n(x)=\frac{1}{n}\sum_{t=1}^n \mathbbm{1}_{\{\hat{\eta}_t\leq x\}}$ being the empirical distribution function (edf) of the residuals.

Having obtained estimators for $\theta_0$ and $\xi_\alpha$, we turn to the estimation of the conditional VaR of the one-period ahead observation at level $\alpha$. Whereas the notation $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ stresses the object's conditional nature, we henceforth proceed with the abbreviation $VaR_{n,\alpha}$ for notational convenience. Employing (ref) -- (ref) we can estimate $VaR_{n,\alpha}$ by

align[align omitted — 331 chars of source]

Clearly, the estimator's large sample properties cannot be studied using traditional tools such as consistency since (ref) does not permit a limit.

For the subsequent asymptotic analysis, we introduce the following assumptions.

assumption{(Compactness)} $\Theta$ is a compact subset of $\mathbb{R}^r$.
assumption{(Stationarity & Ergodicity)} $\{\epsilon_t\}$ is a strictly stationary and ergodic solution of (ref) with (ref).
assumption{(Volatility process)} The function $\sigma:\mathbb{R}^\infty\times \Theta\to(0,\infty)$ is known and for any real sequence $\{x_i\}$, the function $\theta\to\sigma(x_1,x_2,\dots;\theta)$ is continuous. Almost surely, $\sigma_t(\theta)>\underline{\omega}$ for any $\theta \in \Theta$ and some $\underline{\omega}>0$ and $\mathbb{E}[\sigma_t^s(\theta_0)]<\infty$ for some $s>0$. Moreover, for any $\theta \in \Theta$, we assume $\sigma_t(\theta_0)/\sigma_t(\theta)=1$ almost surely (a.s.) if and only if $\theta=\theta_0$.
assumption{(Initial conditions)} There exists a constant $\rho \in (0,1)$ and a random variable $C_1$ measurable with respect to $\mathcal{F}_0$ and $\mathbb{E}[|C_1|^s]<\infty$ for some $s>0$ such that \begin{enumerate}[(i)] • $\sup_{\theta \in \Theta}|\sigma_t(\theta)-\tilde{\sigma}_t(\theta)|\leq C_1 \rho^t$; • $\theta\to \sigma(x_1, x_2, \dots;\theta)$ has continuous second-order derivatives satisfying \begin{align*} \sup_{\theta \in \Theta}\bigg|\bigg|\frac{\partial \sigma_t(\theta)}{\partial \theta}-\frac{\partial \tilde{\sigma}_t(\theta)}{\partial \theta}\bigg|\bigg|\leq C_1 \rho^t, \qquad \quad \sup_{\theta \in \Theta}\bigg|\bigg|\frac{\partial^2 \sigma_t(\theta)}{\partial \theta \partial \theta'}-\frac{\partial^2 \tilde{\sigma}_t(\theta)}{\partial \theta\partial \theta'}\bigg|\bigg|\leq C_1 \rho^t, \end{align*} where $||\cdot||$ denotes the Euclidean norm. \end{enumerate}
assumption{(Innovation process)} The innovations $\{\eta_t\}$ satisfy \begin{enumerate}[(i)] • $\eta_t\overset{iid}{\sim}F$ with $F$ being continuous, $\mathbb{E}\big[\eta_t^2\big]=1$ and $\eta_t$ is independent of $\{\epsilon_u:u<t\}$; • $\eta_t$ admits a density $f$ which is continuous and strictly positive around $\xi_\alpha<0$; • $\mathbb{E}\big[\eta_t^4\big]<\infty$. \end{enumerate}
assumption{(Interior)} $\theta_0$ belongs to the interior of $\Theta$ denoted by $\mathring{\Theta}$.
assumption{(Non-degeneracy)} There does not exist a non-zero $\lambda\in \mathbb{R}^r$ such that $\lambda'\frac{\partial \sigma_t(\theta_0)}{\partial \theta}=0$ a.s.
assumption{(Monotonicity)} For any real sequence $\{x_i\}$ and for any $\theta_1,\theta_2 \in \Theta$ satisfying $\theta_1\leq \theta_2$ componentwise, we have $\sigma(x_1,x_2,\dots;\theta_1)\leq \sigma(x_1,x_2,\dots;\theta_2)$.
assumption{(Moments)} There exists a neighborhood $\mathscr{V}(\theta_0)$ of $\theta_0$ such that the following variables have finite expectation \begin{align*} (i)\sup_{\theta \in \mathscr{V}(\theta_0)}\bigg|\frac{ \sigma_t(\theta_0)}{\sigma_t(\theta)}\bigg|^a, \qquad \;\;\: (ii) \sup_{\theta \in \mathscr{V}(\theta_0)}\bigg|\bigg|\frac{1}{\sigma_t(\theta)}\frac{\partial \sigma_t(\theta)}{\partial \theta}\bigg|\bigg|^{b}, \qquad \;\;\: (iii) \sup_{\theta \in \mathscr{V}(\theta_0)}\bigg|\bigg|\frac{1}{\sigma_t(\theta)}\frac{\partial^2 \sigma_t(\theta)}{\partial \theta \partial \theta'}\bigg|\bigg|^c \end{align*} for some $a$, $b$, $c$ (to be specified).\footnote{Note that the variables in (i)--(iii) are strictly stationary (francq2011garch, francq2011garch, p.\ 181/406).}
assumption{(Scaling Stability)} There exists a function $g$ such that for any $\theta \in \Theta$, for any $\lambda>0$, and any real sequence $\{x_i\}$ \begin{align*} \lambda \sigma(x_1,x_2,\dots;\theta)=\sigma(x_1,x_2,\dots;\theta_\lambda), \end{align*} where $\theta_\lambda=g(\theta,\lambda)$ and $g$ is differentiable in $\lambda$.

The previous set of assumptions is comparable to the conditions imposed by francq2015risk. Assumption (ref) calls for a correct specification of the volatility structure. If the researcher incorrectly specifies a volatility function $\varsigma(\dots;\vartheta)$ instead, the estimator of the misspecified conditional volatility model $\hat{\vartheta}_n$ will converge to a pseudo-true value, i.e.\ $\vartheta_0=\arg \min_\vartheta \mathbb{E}\big[\frac{1}{2}\frac{\epsilon_t^2}{\varsigma_t^2(\vartheta)}+\log \varsigma_t(\vartheta)\big]$. The corresponding edf of the residuals $\frac{1}{n}\sum_{t=1}^n \mathbbm{1}_{\{\epsilon_t/\varsigma_t(\hat{\vartheta}_n)\leq x\}}$ converges to $\bar{F}(x)=\mathbb{E}\big[F\big(x\frac{\varsigma_t(\vartheta_0)}{\sigma_t(\theta_0)}\big)\big]$ in view of Lemma (ref) while the $\alpha$-quantile estimator converges to $\bar{F}^{-1}(\alpha)$, which is generally different from $F^{-1}(\alpha)$. Thus, the correct specification of the volatility function is crucial and one can test for it using the recently developed test by Maria2019; for further recent results on goodness-of-fit testing for GARCH models see Bardet2020. Regarding the bootstrap method to be developed below, it is plausible to expect that in the presence of a misspecified conditional volatility model it will be consistent for the pseudo-true values, although we do not provide a rigorous proof.

Regarding the innovation process we do not need to assume $\mathbb{E}[\eta_t]=0$ (cf.\ francq2004maximum, francq2004maximum, Remark 2.5). The iid condition in Assumption (ref)(ref) is vital for (ref) to hold and is the basis of the residual bootstrap in Section (ref). Under correct specification of the volatility process the iid assumption imposed on the innovations can be tested for by considering the errors $\epsilon_t/ \sigma(\epsilon_{t-1},\epsilon_{t-2},\dots;\hat{\theta}_n)$, $t=1,..,n$, and applying the test of cho2011generalized.

Whereas cavaliere2018fixed assume the existence of the sixth moment of $\eta_t$ for the fixed-design bootstrap in ARCH($q$) models, we only require the fourth moment to be finite in Assumption (ref)((ref)). We assume $\theta_0$ belongs to the interior of the parameter space in Assumption (ref). Parameters on the boundary yield non-standard problems, which require special treatment. cavaliere2020bootstrap provide bootstrap inference on the boundary of the parameter space with application to conditional volatility models. We return to this issue in our empirical application in Remark (ref). In Assumption (ref) the function $\sigma(x_1,x_2,\dots;\theta)$ is presumed to be monotonically increasing in $\theta$, which is used to establish the strong consistency of the quantile estimator. While the monotonicity condition is a feature shared by various stochastic volatility models (cf.\ berkes2003limit, berkes2003limit, Lemma 4.1), it excludes the exponential GARCH nelson1991conditional and the log-GARCH geweke1986comment,pantula1986modeling. Further, we require higher order of moments in Assumption (ref) for the bootstrap, which does not seem to be restrictive for the classical GARCH-type models (cf.\ francq2011garch, francq2011garch, p.\ 165; hamadeh2011asymptotic, hamadeh2011asymptotic, p.\ 501). In particular, Assumption (ref) is presumed to hold with $a=\pm 12$, $b=12$ and $c=6$ for establishing the convergence of the bootstrap information matrix.

On the basis of the previous assumptions we extend the strong consistency result of francq2015risk (francq2015risk, Theorem 1) to the quantile estimator.

theorem(Strong Consistency) \begin{itemize} • francq2015risk Under Assumptions (ref)--(ref), (ref)((ref)) and (ref)((ref)) the estimator in (ref) is strongly consistent, i.e. $\hat{\theta}_n \overset{a.s.}{\to} \theta_0$. • If in addition Assumptions (ref) and (ref)(i) hold with $a=-1$, then the estimator in (ref) satisfies $\hat{\xi}_{n,\alpha} \overset{a.s.}{\to} \xi_\alpha$. \end{itemize}

To lighten notation, we henceforth write $D_t(\theta) =\frac{1}{\sigma_t(\theta)}\frac{\partial\sigma_t(\theta)}{\partial \theta}$ and drop the argument when evaluated at the true parameter, i.e.\ $D_t=D_t(\theta_0)$. The next result provides the joint asymptotic distribution of $\hat{\theta}_n$ and $\hat{\xi}_{n,\alpha}$ and is due to francq2015risk.

theorem(Asymptotic Distribution, francq2015risk) Suppose Assumptions (ref)--(ref), (ref) and (ref) hold with $a=b=4$ and $c=2$. Then, we have \begin{align} \begin{pmatrix} \sqrt{n}(\hat{\theta}_n-\theta_0)\\ \sqrt{n}(\xi_{\alpha} - \hat{\xi}_{n,\alpha}) \end{pmatrix} \overset{d}{\to}N\big(0, \Sigma_\alpha\big) \qquad with\qquad \Sigma_\alpha= \begin{pmatrix} \frac{\kappa-1}{4}J^{-1} & \lambda_\alpha J^{-1}\Omega\\ \lambda_\alpha \Omega'J^{-1} & \zeta_\alpha \end{pmatrix}, \end{align} where $\kappa = \mathbb{E}[\eta_t^4]$, $\Omega = \mathbb{E}[D_t]$, $J=\mathbb{E}[D_tD_t']$, $\lambda_\alpha = \xi_\alpha\frac{\kappa-1}{4}+\frac{p_\alpha}{2f(\xi_{\alpha})}$, $\zeta_\alpha = \xi_\alpha^2\frac{\kappa-1}{4}+\frac{\xi_{\alpha} p_\alpha}{f(\xi_\alpha)}+\frac{\alpha(1-\alpha)}{f^2(\xi_\alpha)}$ and $p_\alpha =\mathbb{E}[\eta_t^2 \mathbbm{1}_{\{\eta_t<\xi_\alpha\}}]-\alpha$.
remarkIt is worth mentioning that the asymptotics in this theorem for $\hat{\xi}_{n,\alpha}$ are for $\alpha$ fixed while $n$ goes to infinity. If, for instance, $\alpha$ is very small for moderate $n$ the distribution in the following theorem might not provide a good approximation. For such cases, approximations based on extreme value theory may provide better approximations to the unknown finite sample distribution. See, for example, McNeil2002 and the recent li_peng_song_2022 for an application of GARCH models with extreme value theory for VaR estimation. Note that the latter authors consider simultaneous inference for VaR and expected shortfall.

In applied work, to use Theorem (ref) in order to conduct inference for $(\theta_0, \xi_{\alpha})$ one needs a consistent estimator for $\Sigma_{\alpha}$. It follows from Lemma (ref) that $\kappa$, $\Omega$, and $J$ can be consistently estimated by

align[align omitted — 247 chars of source]

respectively, with $\hat{D}_t = \tilde{D}_t(\hat{\theta}_n)$ and $\tilde{D}_t(\theta) = \frac{1}{\tilde{\sigma}_t(\theta)}\frac{\partial\tilde{\sigma}_t(\theta)}{\partial \theta}$. To estimate $\lambda_{\alpha}$ and $\zeta_{\alpha}$, which also appear in $\Sigma_{\alpha}$, one can use $\hat{\xi}_{n,\alpha}$ for $\xi_{\alpha}$ and

align[align omitted — 162 chars of source]

for $p_{\alpha}$ (see also Lemma (ref) for the properties of $\hat{p}_{n,\alpha}$). To estimate the density $f$, which also appears in $\lambda_{\alpha}$ and $\zeta_{\alpha}$, kernel smoothing is commonly employed, i.e.\

align[align omitted — 122 chars of source]

with kernel function $k$ and bandwidth $h_n>0$. Hence, whenever we use a consistent $\hat{\mathbbm{f}}_n^S$ for $f$ we obtain, combined with Equations (ref) and (ref), a consistent estimator for $\Sigma_\alpha$ denoted by $\hat{\Sigma}_{n,\alpha}$. For the case that $\epsilon_t$ in Equation ((ref)) follows a GARCH$(p,q)$ process, gao2008estimation considered estimating $f$ by a kernel density estimator with Lipschitz-continuous kernels such as $k(x)=\phi(x)$, where $\phi$ is the standard normal density function. An alternative estimator is based on the uniform kernel $k(x)=\frac{1}{2}\mathbbm{1}_{\{|x|\leq 1\}}$ yielding $\hat{\mathbbm{f}}_n^S(\hat{\xi}_{n,\alpha})\overset{p}{\to}f(\xi_\alpha)$ whenever $h_n \sim n^{-\varrho}$ for some $\varrho \in (0,1/2]$.

At this point it is worth recalling that interest lies in $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ (cf. Equation ((ref))). More precisely, one is not so much interested in the distribution of this random variable or its moments but in inference for its possible realizations. Note also that $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ varies with $n$ and does not converge. This illustrates that inference for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ is different from inference for $(\xi_{\alpha},\theta_0)$ which is just a real-valued vector of dimension $r+1$. Nevertheless, it is intuitively clear how to conduct inference for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$, yet this difference results in two technical issues that need to be addressed to provide a thorough theoretical justification of inferential procedures for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$. These technical issues are known for quite some time; see, for instance, phillips1979sampling, kreiss2015discussion and the textbook pesaran2015time[p. 389]. Several examples illustrating these two issues can be found in beutner2017justification. In a nutshell the first issue stems from the fact that, as mentioned, $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ varies over time (cf. Eq. ((ref))) implying that a limiting distribution cannot exist. The second issue is a result of $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ being a conditional quantity which on the one hand requires conditioning on the sample observed so far, yet on the other hand this conditioning eliminates all randomness, making it impossible to establish useful distributional results. We now briefly illustrate this here for the case that $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ is the object of interest; a very detailed illustration of this second issue by means of a GARCH(1,1) can be found in Example 2.1 in beutner2017justification. Fixing some arbitrary starting values $\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots$ and making the dependency of $\sigma_{n+1}$ on the realized values of the random variables $\epsilon_n,\epsilon_{n-1},\ldots,\epsilon_1$ and the starting values $\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots$ explicit a feasible version of the quantity of interest is

equation[equation omitted — 241 chars of source]

where $\epsilon_n^r,\epsilon_{n-1}^r,\ldots,\epsilon_{1}^r$ denote the realized values of $\epsilon_n,\epsilon_{n-1},\ldots,\epsilon_1$. Here feasible is in the sense that only parameter uncertainty remains. Now replacing, as at the beginning of Section 3, the unknown $\theta_0$ by the estimator $\hat{\theta}_{n}$ and doing the same with $\xi_{\alpha}$ we see that these estimators must enter Equation (ref) given the realizations, i.e. as $\hat{\theta}_{n}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$ and $\hat{\xi}_{\alpha,n}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$, respectively, because otherwise we would end up with an inconsistency. Indeed, if they entered as $\hat{\theta}_n(\epsilon_1,\epsilon_{2},\ldots,\epsilon_n)$ and $\hat{\xi}_{n,\alpha}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_n)$, respectively, we would treat $\epsilon_1,\epsilon_{2},\ldots,\epsilon_n$ as random and observed at the same time. Upon replacing $\theta_0$ and $\xi_{\alpha}$ in Equation (ref) by $\hat{\theta}_{n}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$ and $\hat{\xi}_{n,\alpha}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$, respectively, no random variable appears in

align*[align* omitted — 489 chars of source]

which is just the estimator of Equation (3.3) with all dependencies made explicit. As just observed no random variable appears in $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$ which therefore does not have a distribution which could be used to construct confidence intervals for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$. In the article beutner2017justification merging is proposed as a solution to overcome the first issue and sample splitting as a means to solve the second. The results of this article can be used justify, through their Theorem 1, approximating the distribution of the VaR estimator, centered at $VaR_{n,\alpha}$ and inflated by $\sqrt{n}$, by

align[align omitted — 321 chars of source]

and consequently to justify the intuitive (conditional) confidence interval

align[align omitted — 702 chars of source]

for ${VaR}_{n,\alpha}$. Here $\Phi$ denotes the standard normal cdf, and we used again the short-hand notation for $\tilde{\sigma}_{n+1}(\hat{\theta})$ and $\hat{\xi}_{n,\alpha}$, i.e. did not make the dependency on the starting values etc. explicit. Note that the interval (ref) lacks a proper theoretical justification as it treats the non-random $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$ as random because it assigns an asymptotic distribution to it. In order to use Theorem 3 and Corollary 2 of beutner2017justification to provide a sound theoretical justification of (ref) let $n_E: \mathbb{N} \to \mathbb{N}$ and $n_P: \mathbb{N} \to \mathbb{N}$ be such that for all $n \in \mathbb{N}$ we have $n_E(n) < n_P(n)$. First, in the sample split approach we will only use $\epsilon_1,\ldots,\epsilon_{n_E}$ to estimate $\xi_{\alpha}$ and $\theta_0$ which we denote by $\hat{\xi}_{n_E,\alpha}^{SPL}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_{n_E})$ and $\hat{\theta}_{n_E}^{SPL}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_{n_E})$, respectively. To lighten the notation a bit we will also use the shorter notations $\hat{\xi}_{n_E,\alpha}^{SPL}$ and $\hat{\theta}_{n_E}^{SPL}$, respectively. Second, in the sample split approach conditioning on $\mathcal{F}_n$ on the left-hand side of (ref) is replaced by conditioning on $\epsilon_{n_P},\ldots,\epsilon_n$ only. The sample split version of (3.3) which we denote by $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{SPL}$ becomes

align*[align* omitted — 723 chars of source]

where $\epsilon_n^r,\epsilon_{n-1}^r,\ldots,\epsilon_{n_P}^r$ and $\tilde{\epsilon}_{0},\tilde{\epsilon}_{-1},\ldots$ are as before and $c_{n_P-1},\ldots,c_{1}$ are constants. These constants could be viewed as starting values but since they are not replacing unobserved value as $\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots$ do we prefer to use a different letter for them. Note also that $\epsilon_1,\ldots,\epsilon_{n_E}$ enter the sample split version $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{SPL}$ as random variables which is in contrast to $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$. Therefore, $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{SPL}$ does have a non-degenerate distribution which can be used to construct confidence intervals. The sample split (conditional) confidence interval is defined as

align[align omitted — 809 chars of source]

Here $\hat{\Sigma}_{n_E,\alpha}^{SPL}$ is as $\hat{\Sigma}_{n,\alpha}$ but based on the first $n_E$ observations only. Note that this interval is meaningful because $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{SPL}$ is random and converges in distribution after centering and scaling. Intuitively, we would say that the statistically meaningful interval in (ref) provides a theoretical justification for the interval (ref) if

equation[equation omitted — 535 chars of source]

and if

equation[equation omitted — 538 chars of source]

converges to

equation[equation omitted — 451 chars of source]

This is also the concept employed in beutner2017justification. Under the conditions of Theorem 3 of this article the convergence in (ref) holds and (ref) converges to (ref) by Corollary 2 of this article. It is worth pointing out that convergence is in probability and that for this concept to be applicable $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$ and (ref) are treated as random. Theorem 2 above ensures that one of the assumptions of Corollary 3 of beutner2017justification holds so that it provides the basis for a sound theoretical justification of (ref) as a (conditional) confidence interval. Formally, we can state

corollaryAssume that the assumptions of Theorem (ref) are fulfilled. Define the function $\psi: \mathbb{R}^{\infty} \times \Upsilon$ with $\Upsilon=\mathbb{R} \times \Theta$ by $$\psi(x_1,x_2,\ldots;\upsilon)=\psi(x_1,x_2,\ldots;(\xi,\theta))=-\xi \sigma(x_1,x_2,\ldots;\theta),$$ and put $\upsilon_0=(\xi_{\alpha}, \theta_0)$. Assume further that \begin{enumerate} • $\Big|\Big|\frac{\partial \sigma (\epsilon_n, \epsilon_{n-1}, \ldots ;\theta_0)}{\partial \theta}\Big|\Big|=O_{p}(1)$; • $\sup_{\upsilon \in \mathscr{V}(\upsilon_0)}\Big|\Big|\frac{\partial^2 \psi (\epsilon_n, \epsilon_{n-1}, \ldots ;\upsilon)}{\partial \upsilon \partial \upsilon'}\Big|\Big|=O_{p}(1)$ for some open neighborhood $\mathscr{V}(\upsilon_0)$ around $\upsilon_0$; • Given sequences $\{\tilde{\epsilon}_t\}$ and $\{c_t\}$, we have \begin{align*} & \sqrt{n}\big(\psi(\epsilon_n,\ldots,\epsilon_{t_1},c_{t_1-1},\ldots,c_1,\tilde{\epsilon}_0,\ldots; \upsilon_0) - \psi (\epsilon_n, \epsilon_{n-1}, \ldots; \upsilon_0)\big)=o_{p}(1),\\ & \bigg|\bigg|\frac{\partial \psi(\epsilon_n,\ldots,\epsilon_{t_1},c_{t_1-1},\ldots,c_1,\tilde{\epsilon}_0,\ldots; \upsilon_0) }{\partial \upsilon} - \frac{\partial \psi (\epsilon_n, \epsilon_{n-1}, \ldots; \upsilon_0)}{\partial \upsilon}\bigg|\bigg|=o_{p}(1),\\ & \sup_{\upsilon \in \mathscr{V}(\upsilon_0)} \bigg|\bigg|\frac{\partial^2 \psi(\epsilon_n,\ldots,\epsilon_{t_1},c_{t_1-1},\ldots,c_1,\tilde{\epsilon}_0,\ldots; \upsilon_0)}{\partial \upsilon \partial \upsilon'} - \frac{\partial^2 \psi (\epsilon_n, \epsilon_{n-1}, \ldots; \upsilon_0)}{\partial \upsilon \partial \upsilon'}\bigg|\bigg| \\ & =o_{p}(1) \end{align*} for any $t_1 \geq 1$ such that $(n-t_1) / l_n \rightarrow \infty$ as $n \to \infty$ and for some model-specific $l_n$ with $l_n \rightarrow \infty$. \end{enumerate} Moreover, let $n_E$ and $n_P$ fulfill Assumption 3.a of beutner2017justification and $\{\epsilon_t\}$ Assumption 3.c of the same article. Then the difference between $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{SPL}$ and $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$ converges to zero in probability and the same holds for the difference between (ref) and (ref).
proofThe claims made are proved if the assumptions of Theorem 3 and Corollary 2 in beutner2017justification are met. Assumption 1.a of beutner2017justification holds by Theorem (ref). Note that the scaling sequence $m_T$ in beutner2017justification equals $\sqrt{n}$ here. Moreover according to Theorem (ref) $\hat{\upsilon}_n=(\hat{\xi}_{n,\alpha},\hat{\theta}_n)$ scaled by $\sqrt{n}$ and centered at $\upsilon_0$ converges to a multivariate normal distribution so that Assumption 5 of beutner2017justification also holds. We now turn to Assumption 1.b of beutner2017justification. Clearly $\upsilon \to \psi(\cdot;\upsilon$) is continuous because by Assumption 3 above $\theta \to \sigma(\cdot,;\theta)$ is continuous. Furthermore, because $\upsilon \to \psi(\cdot;\upsilon)$ is the product of the functions $\xi \to \xi$ and $\theta \to \sigma(\cdot;\theta)$ which are both twice differentiable (for $\theta \to \sigma(\cdot;\theta)$ this holds by Assumption 4 (ii) above) the same holds for $\upsilon \to \psi(\cdot;\upsilon)$. Therefore, Assumption 3.b of beutner2017justification holds. The gradient of $\psi(\epsilon_n,\epsilon_{n-1},\ldots;\upsilon_0)$ is given by $$ \left(-\sigma(\epsilon_n,\epsilon_{n-1},\ldots;\theta_0), -\xi_{\alpha}\frac{\partial \sigma (\epsilon_n, \epsilon_{n-1}, \ldots ;\theta_0)}{\partial \theta}\right). $$ By Assumption 3 above and Markov's inequality $\sigma(\epsilon_n,\epsilon_{n-1},\ldots;\theta_0)$ is bounded in probability. Hence, together with Assumption 1. in the statement of the corollary it follows that Assumption 1.c in beutner2017justification holds. Assumption 1.d and 1.e in beutner2017justification are just the Assumptions 2. and 3. of the corollary. Obviously, Assumptions 3.a and 3.c of beutner2017justification hold under the assumptions of the corollary. Assumption 3.b of this article is met by Assumption 2 above which is nothing else than Assumption 3.b of beutner2017justification for Value-at-Risk. As explained in the proof of Theorem 3 in beutner2017justification the Assumption 2.b of that article is irrelevant for the quantity on the right-hand side of (ref) and for (ref). Assumption 2.a of that article is obviously met. This finishes the proof.
remarkLike some of the Assumptions 1-10 will hold or not depending on the specification of $\sigma$, the same is true for the assumptions of Corollary (ref) that ensure applicability of the results of beutner2017justification. For popular time series models like a GARCH(1,1) they have been verified in beutner2019technical.

Although the interval in (ref) is, as just outlined, well-justified, it may perform poorly since the density estimation appears rather sensitive regarding the choice of bandwidth (see gao2008estimation, gao2008estimation, Section 4). Bootstrap methods offer an alternative way to quantify the uncertainty around the estimators.

Bootstrap

Bootstrap approximations frequently provide better insight into the actual distribution than the asymptotic approximation, yet they require a careful set-up. hall2003inference show that conventional bootstrap methods are inconsistent in a GARCH model lacking finite fourth moment in the case of the squared innovations' distribution not being in the domain of attraction of the normal distribution. They consider a subsample bootstrap instead and study its asymptotic properties. In correspondence, an $m$-out-of-$n$ without-replacement bootstrap is proposed by spierdijk2016confidence to construct confidence intervals for ARMA-GARCH VaR.

pascual2006bootstrap present a residual bootstrap in a GARCH($1,1$) setting and assess its finite sample properties by means of simulation. Their bootstrap scheme follows a recursive design in which the bootstrap observations are generated iteratively using the estimated volatility dynamics. Building upon their results, christoffersen2005estimation construct bootstrap confidence intervals for (conditional) VaR and Expected Shortfall and compare them to competitive methods within the GARCH($1,1$) model. Theoretical results on the recursive-design residual bootstrap are provided by hidalgo2007goodness and jeong2017residual for the ARCH($\infty$) and GARCH($p,q$) model, respectively.

In contrast, shimizu2009bootstrapping considers fixed-design variants of the wild and the residual bootstrap in which the ARMA-GARCH dynamics of the bootstrap samples are kept fixed at the values of the original series. The bootstrap estimators are based on a single Newton-Raphson iteration simplifying the proofs of first-order asymptotic validity. shimizu2009bootstrapping's approach for the residual bootstrap is also employed in a multivariate GARCH setting by francq2016variance. Recently, cavaliere2018fixed study the fixed-design residual bootstrap in the context of ARCH($q$) models and propose a bootstrap Wald statistic based on a QML bootstrap estimator. While their theory has been developed independently to ours, their simulation study indicates that the fixed-design bootstrap performs as well as the recursive-design bootstrap.

Fixed-design Residual Bootstrap

We propose a fixed-design residual bootstrap procedure, described in Algorithm (ref), to approximate the distribution of the estimators in (ref) -- (ref).

algorithm[algorithm omitted — 1,309 chars of source]
remarkIn contrast to the literature, the bootstrap errors are drawn with replacement from the residuals rather than the standardized residuals. In fact, re-centering would be inappropriate in the case of $\mathbb{E}[\eta_t]\neq 0$. In addition, re-scaling of the residuals is typically redundant as $\frac{1}{n}\sum_{t=1}^n \hat{\eta}_t^2=1$ is implied by $\hat{\theta}_n \in \mathring{\Theta}$ under Assumption (ref); see francq2011garch, francq2011garch, p.\ 182/406 and note that the solution requires $\hat{\theta}_n$ belonging to the interior (francq2011garch, Oct.\ 2018, personal communication).
remarkThe term `fixed-design' refers to the fact that the bootstrap observations are generated using $\tilde{\sigma}_t(\hat{\theta}_n)=\sigma(\epsilon_{t-1},\dots,\epsilon_1,\tilde{\epsilon}_0,\tilde{\epsilon}_{-1},\dots;\hat{\theta}_n)$. In contrast, a recursive-design scheme replicates the model's dynamic structure, i.e.\ $\epsilon_t^\star = \sigma_t^\star \eta_t^\star$ with $\sigma_t^\star = \sigma(\epsilon_{t-1}^\star,\dots,\epsilon_1^\star,\tilde{\epsilon}_0,\tilde{\epsilon}_{-1},\dots;\hat{\theta}_n)$ and $\eta_t^\star \overset{iid}{\sim} \hat{\mathbbm{F}}_n$, which is computationally more demanding. We refer to Appendix (ref) for a complete description. See also cavaliere2018fixed for more theoretical insights on the difference in the design in an ARCH($q$).
remarkWhereas (ref) involves a nonlinear optimization, shimizu2009bootstrapping proposes a Newton-Raphson type bootstrap estimator instead. The Newton-Raphson bootstrap estimator corresponding to (ref) is given by \begin{align*} \hat{\theta}_n^{*NR} = \hat{\theta}_n+\hat{J}_n^{-1}\frac{1}{2n}\sum_{t=1}^n \hat{D}_t \big(\eta_t^{*2}-1\big), \end{align*} which can considerably speed up computations.

Proposition (ref) establishes the asymptotic validity of the bootstrap for the volatility parameters.

propositionSuppose Assumptions (ref)--(ref), (ref)((ref)), (ref)((ref)), (ref), (ref), (ref) and (ref) hold with $a=\pm 12$, $b=12$ and $c=6$. Then, we have \begin{align*} \sqrt{n}\big(\hat{\theta}_n^* - \hat{\theta}_n\big) \overset{d^*}{\to}N\bigg(0,\frac{\kappa-1}{4}J^{-1}\bigg) \end{align*} almost surely.

Establishing the asymptotic validity of the bootstrap for the second part appears challenging since the bootstrap innovations are drawn from the discrete distribution $\hat{\mathbbm{F}}_n$. To overcome this issue we rely on arguments employed by bahadur1966note and berkes2003limit. The following theorem states the paper's main result.

theorem(Bootstrap consistency) Suppose Assumptions (ref)--(ref) hold with $a=\pm 12$, $b=12$ and $c=6$. Then, we have \begin{align*} \begin{pmatrix} \sqrt{n}(\hat{\theta}_n^*-\hat{\theta}_n)\\ \sqrt{n}(\hat{\xi}_{n,\alpha} - \hat{\xi}_{n,\alpha}^*) \end{pmatrix} \overset{d^*}{\to}N\big(0, \Sigma_\alpha\big) \end{align*} in probability.

Theorem (ref) is useful to validate the bootstrap for the conditional VaR estimator. For the asymptotic behavior of the conditional VaR estimator we refer to (ref) and the text around it. The following corollary is established.

corollaryUnder the assumptions of Theorem (ref) the conditional distribution of $\sqrt{n}\big( \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{*}- \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}\big)$ given $\mathcal{F}_n$ and (ref) given $\mathcal{F}_n$ merge in probability.

Bootstrap Confidence Intervals for VaR

Clearly, the VaR evaluation in (ref) is subject to estimation risk that needs to be quantified. We propose the following algorithm to obtain approximately $100(1-\gamma)\%$ confidence intervals.

algorithm[algorithm omitted — 3,838 chars of source]

The interval in (ref) is obtained by the EP method, that is frequently encountered in the bootstrap literature. It is obtained from the (typically) infeasible equal-tailed confidence interval

align*[align* omitted — 572 chars of source]

where $G_{n}^{-1}$ is the (unknown) quantile function of $\sqrt{n}( \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}-{VaR}_{n,\alpha})$, which is replaced by its bootstrap analogue $\hat{G}_{n,B}^{*-1}$. The same reasoning leads to the SY interval but with test statistic $\sqrt{n}|( \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}-{VaR}_{n,\alpha})|$ instead of $\sqrt{n}( \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}-{VaR}_{n,\alpha})$ which makes it also clear that the interval in (ref) presumes symmetry for rationalizing its construction. “Flipping around" the tails of the ET interval leads to the RT interval given in (ref). Clearly, the RT and the EP have equal length. Whereas (ref) in its current form emphasizes the interval's name, RT type intervals are frequently reported in their reduced form, i.e. the lower and upper bound of (ref) simplify to the $\gamma/2$ and $1-\gamma/2$ quantiles of $\frac{1}{B}\sum_{b=1}^B \mathbbm{1}_{\big\{ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}^{*(b)}\leq x\big\}}$, respectively. RT intervals can either be motivated by the results of falk1991coverage\footnote{In a random sample setting falk1991coverage prove that the RT bootstrap interval for quantiles has asymptotically greater coverage than the corresponding EP bootstrap interval. For additional insights we refer to hall1988bootstrap.} or as the bootstrap analogue of the (uncentered) statistic $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.5ex}} \stackon[1pt]{VaR}{\tmpbox} _{n,\alpha}$. It is worth mentioning that RT type bootstrap intervals for the VaR are also constructed in reduced form by christoffersen2005estimation. Regardless of whether we use an EP, RT or SY interval the meaning is always the same: Given the past up to and including time $n$ the probability that the conditional VaR for period $n+1$ is contained in the intervals is approximately equal to $100(1-\gamma)\%$.

Bootstrap Extensions

The asymptotic normality result in Theorem (ref) as well as the bootstrap consistency in Theorem (ref) are derived, inter alia, under the assumption that the innovations are iid. In case this is not believed to be true -- e.g. if the suggested specification tests mentioned in Section (ref) indicate otherwise -- asymptotic normality of $\sqrt{n}(\hat{\theta}_n-\theta_0)$ can still be established under regularity assumptions. Escanciano2009 studies the QML estimator under some dependence among the $\eta_t$'s while imposing slightly stronger (moment) conditions, whereas the related paper of Linton2010 investigates estimators in a GARCH(1,1) with dependent errors but under weaker moment conditions. A multivariate version of the dependence condition in Escanciano2009 can be found in FrancqZakoian2016.

The bootstrap method presented in Algorithm (ref) is contingent on the iid assumption. In general, one cannot expect bootstrap procedures that are based on an iid assumption to be robust against deviations from this assumption; for a well-known example, one may refer to GONCALVES2004. Alternative bootstrap techniques may be used if the iid condition is thought to be unrealistic. A variety of bootstrap methods exist that can capture dependence and non-identical random variables; see e.g. Lahiri03 for a broad overview. The wild or multiplier bootstrap Mammen93, DavidsonFlachaire08 is particularly suited for dealing with non-identical variables, but does not capture dependence, unless it is properly modified Shao10, FSU20. However, it remains an open question which bootstrap method combined with the fixed design approach leads to valid bootstrap procedures.

Numerical Illustration

Monte Carlo Experiment

In order to evaluate the finite sample performance of the proposed bootstrap procedure a Monte Carlo experiment is conducted. We confine ourselves to four conditional volatility specifications related to Examples (ref) and (ref) in Section (ref). The first two are GARCH($1,1$) parameterizations with

enumerate[(i)] • high persistence: $\theta_0 = (\omega_0,\alpha_0,\beta_0)’= \big(0.05\times 20^2/252,0.15,0.8\big)'$; • low persistence: $\theta_0 = (\omega_0,\alpha_0,\beta_0)’ = \big(0.05\times 20^2/252,0.4,0.55\big)'$,

which are similar to the specifications of gao2008estimation (gao2008estimation, Section 4) and spierdijk2016confidence (spierdijk2016confidence, Section 4.2). In addition, we study two T-GARCH($1,1$) scenarios likewise associated with high and low persistence:

enumerate[(i)] \setcounter{enumi}{2} • high persistence: $\theta_0 = (\omega_0,\alpha_0^+, \alpha_0^-, \beta_0)’ = \big(0.05\times 20/\sqrt{252},0.05,0.10,0.8\big)'$; • low persistence: $\theta_0 = (\omega_0,\alpha_0^+, \alpha_0^-, \beta_0)’ = \big(0.05\times 20/\sqrt{252},0.1,0.3,0.55\big)'$.

Within the experiment the VaR level takes value $\alpha =0.05$ and there are two possible innovation distributions: a Student-$t$ distribution with $6$ degrees of freedom (df) and the standard normal distribution.\footnote{The Student-t innovations are appropriately standardized to satisfy $\mathbb{E}\eta_t^2=1$.} We consider four estimation sample sizes, $n \in \{250; 500; 1{,}000; 5{,}000\}$, whereas the number of bootstrap replicates is fixed and equal to $B=999$. For each model version we simulate $S=10{,}000$ independent Monte Carlo trajectories.

The numerical optimization of the log-likelihood function is carried out employing the build-in function fmincon and running time is reduced by parallel computing using parfor. The code is available on the website of the third author, and simulation results for a VaR level of $\alpha=0.01$ can be found in the working paper version.

\captionsetup[table]{labelformat=simple, labelsep=space}

\bgroup

table[table omitted — 2,888 chars of source]

\egroup

\captionsetup[table]{labelformat=simple, labelsep=space}

\bgroup

table[table omitted — 2,310 chars of source]

\egroup

Table (ref) and (ref) report the results of the three $90\%$--bootstrap intervals for the $5\%$--VaR when the innovation distribution is Student-t (henceforth referred to as baseline) and the model is a GARCH(1,1) and a T-GARCH(1,1), respectively. In both tables the results of the interval (ref) based on asymptotic (AS) theory are included for comparison, where a Gaussian kernel is utilized together with a bandwidth following silverman1986density's (silverman1986density) rule-of-thumb. In the GARCH($1,1$) high persistence case (right part of Table 1), we see that the average coverage varies around $90\%$ across all sample sizes for the RT and the SY interval. In contrast, the EP and the AS interval fall short of the nominal $90\%$ by $10.43$ and $4.15$ percentage points (pp), respectively, for small sample size ($n=250$). Nevertheless, their average coverage approaches the nominal value as the sample size increases. Remarkably, for all four intervals the average rate of the conditional VaR being below the interval is considerably less than the average rate of the conditional VaR being above the interval when the sample size is rather small ($n \leq 500$). Regarding the intervals' length, we observe that the SY interval is on average larger than the EP/RT interval. As the sample size increases this gap diminishes and the intervals' average lengths shrink. Considering the low persistent case (left part of Table (ref)) we find similar results regarding the intervals' average coverage, yet their average lengths turn out to be smaller compared to the high persistent case. This is intuitive as the conditional volatility tends to vary less in the low persistent case. Regarding the T-GARCH($1,1$) in Table (ref), the overall picture is similar as in the GARCH(1,1) case, however the under-coverage in small and medium-sized samples appears to be more extreme for the EP and reduced for the AS interval.

Simulation results for the scenario when the $\eta_t$'s follow a standard normal distribution and when the model is a GARCH(1,1) and a T-GARCH(1,1), respectively, are tabulated in Tables (ref) and (ref) which are given in Appendix (ref). Here we only note that, although the error distribution underlying the QMLE is correctly specified in this case, the qualitative results stated above with regard to Table (ref) persist: the RT and the SY intervals possess accurate coverage rates across sample sizes, whereas the EP and the AS interval exhibit under-coverage in samples of rather small size with different extent. Moreover, we observe that the intervals are on average shorter in the Gaussian case than in the baseline case. This seems partially driven by a smaller variance of $\hat{\xi}_{n,\alpha}$; for $\alpha=0.05$ the asymptotic variance $\zeta_\alpha$ in (ref) is equal to $3.11$ in the Gaussian case compared to $5.72$ in the Student-t case with $6$ degrees of freedom.

While the small-sample-performance of the AS interval can be explained by its embodied density estimation, the question arises why the EP interval performs worse than the other bootstrap intervals, which seems counter-intuitive at first. Howbeit the results are in line with the theoretical findings of falk1991coverage (falk1991coverage, unnumbered Corollary, p.\ 488). In a random sample setting they prove that the RT bootstrap interval for quantiles has asymptotically greater coverage than the corresponding EP bootstrap interval. The emerging gap\footnote{We neglect their $o(n^{-1/2})$ term. Take note that the theoretical results of falk1991coverage are not directly applicable in our setting due to GARCH-type effects.}

enumerate• tends to be smaller for larger sample sizes, • tends to be larger for more extreme quantiles, and • tends to vary with the nominal coverage rate in a non-monotonic way.

Table (ref) presents the average coverage gap between the EP and the RT bootstrap interval of the conditional VaR for the baseline specification as well as for three deviations from the baseline (change in $F$, $\alpha$ and $\gamma$). For example, in the low persistence GARCH($1,1$) case of the baseline with $n=250$, the average coverage gap amounts to $90.73\%-81.20\%=9.53$pp (see also Table (ref)). It is striking that all values are positive within Table (ref), which highlights the superiority of the RT bootstrap interval over the EP bootstrap interval. Further, it is eminent that average coverage gap tends to decrease with increasing sample size, which supports (i). Comparing columns (1) and (3) we also find that the average coverage gap tends to be larger for the $1\%$--VaR than for the $5\%$--VaR, which gives rise to (ii). Regarding (iii), the result of falk1991coverage (falk1991coverage) suggests that the gap slightly decreases when increasing the nominal coverage from $90\%$ to $95\%$. Such tendency is precisely observed when comparing columns (1) and (4) of Table (ref).

table[table omitted — 1,836 chars of source]

\bgroup

table[table omitted — 2,495 chars of source]

\egroup

\bgroup

table[table omitted — 1,878 chars of source]

\egroup

With regard to Remark (ref) in Section (ref), Tables (ref) and (ref) report the simulation results for the recursive-design bootstrap for the DGPs of Tables (ref) and (ref), respectively. We refer to Appendix (ref) for computational details. In comparison to the fixed-design approach (see Tables (ref) and (ref)) we find that the recursive-design method performs similarly in terms of average coverage for each interval type, which corresponds to the simulation results of cavaliere2018fixed. It is striking, however, that the intervals' average lengths are larger in the recursive-design than in the fixed-design set-up. For example, in the high persistence GARCH(1,1) case (right part of Table (ref)) for $n=500$ the average length in the recursive-design approach is $0.604$ for the EP/RT interval compared to $0.581$ in the fixed-design. As the sample size increases this difference disappears.

In summary, the simulations suggest that the RT and the SY bootstrap interval work well for both bootstrap designs and that they outperform in smaller samples the AS interval in terms of average coverage even though their tails are unequally represented. In contrast, for both bootstrap designs the EP interval falls short of its nominal coverage, which is in line with the theoretical findings of falk1991coverage. Since the fixed RT method leads on average to shorter intervals than the corresponding SY method and its recursive-design counterpart, this suggests to favor the fixed-design RT bootstrap interval in (ref).

Empirical Application

We analyze the French stock market index CAC 40 for the period January 1, 2015 -- January 1, 2020. The index values for the period are retrieved from Yahoo Finance and daily (log-) returns (expressed in $\%$) are computed using $\epsilon_t = 100\log(p_t/p_{t-1})$, where $p_t$ denotes the closing value of the index at trading day $t$.

figure[figure omitted — 742 chars of source]

Figure (ref)(a) displays the resulting series of returns. We disregard the observations from July 1, 2019 onwards, which we leave for the out-of-sample evaluation, yielding $n=1{,}146$ remaining observations (i.e.\ January 1, 2015 - July 1, 2019). For the volatility process we consider the T-GARCH($1,1$) model specified in Example (ref).\footnote{We also consider an Asymmetric Power GARCH model ding1993long, i.e.\ $\sigma_{t+1}^\delta= \omega_0 + \alpha_0^+ (\epsilon_{t}^+)^\delta +\alpha_0^-(\epsilon_{t}^-)^\delta + \beta_0 \sigma_{t}^\delta$ with $\delta>0$, which nests the GARCH($1,1$) model ($\delta=2$, $\alpha_0^+=\alpha_0^-$) and the T-GARCH($1,1$) model ($\delta=1$) of Examples (ref) and (ref). In practice, the impact of the power $\delta$ on the volatility is minor and the QML approach of hamadeh2011asymptotic suggests a $\delta$ close to $1$ in favor for the T-GARCH specification.} Table (ref) reports the corresponding point estimates with standard errors obtained by bootstrapping based on Algorithm (ref). As documented in numerous studies we find that the volatility persistence is close to unity. Further, we observe that $\hat{\alpha}^-_n$ is considerably larger than $\hat{\alpha}^+_n$ indicating a strong leverage effect, i.e.\ negative returns tend to increase volatility by more than positive returns of the same magnitude.

table[table omitted — 722 chars of source]

Figure (ref)(b) plots the histogram of the residuals with the normal distribution superimposed. Further, we test the condition that the innovations are iid (see Assumption (ref)(ref)) with the generalized run tests of cho2011generalized.\footnote{The implementation of the tests is available on the website of the first author.} These tests are particularly suitable in this case since they can be based on the residuals and are sensitive against a wide range of alternatives. The test statistic of the sup-norm based test is $0.40$, which corresponds to a p-value of $0.27$. consequently, one cannot reject the null hypothesis of iid innovations at any common significance level. Similarly, the generalized run test based on the $L_1$-norm cannot be rejected at a $10\%$ significance level.

Next, we perform a rolling window analysis starting with subperiod January 1, 2015 -- July 1, 2019 and ending with subperiod July 8, 2015 -- January 1, 2020. We have $130$ subperiods each consisting of $1{,}146$ observations. For each rolling window period we fit a T-GARCH($1,1$) model and estimate the one-period-ahead conditional VaR associated with level $\alpha=0.05$. For example, for the first window the T-GARCH($1,1$) estimates are reported in Table (ref) and the conditional $5\%$-VaR of the one-period ahead (i.e.\ July 1, 2019) is estimated by $1.11$. Further, we obtain the associated $95\%$-confidence intervals based on bootstrap and asymptotic normality. In addition to the RT intervals of the fixed- and residual-design bootstrap, we also computed an interval based on the asymptotic distribution. The corresponding intervals are $[0.850,1.136]$ (fixed-design), $[0.834,1.115]$ (recursive-design), and $[0.828,1.106]$ (asymp.\ normality). Although the intervals are fairly similar, the asymptotic and recursive bootstrap intervals are shorter than the fixed-design interval. Given its tendency to underestimate variability in finite samples, this result is unsurprising for the asymptotic interval, although for the recursive bootstrap this contrasts the simulation findings. Note that the fixed-design iid and block bootstraps produce very similar interval, which is not surprising as our conducted specification tests did not indicate any violation of the iid assumption on the innovations.

The results of the rolling window analysis are visualized in Figure (ref). It plots the realized return together with (the opposite of) the estimated conditional VaR. For clarity we only indicate the lower and upper bound of the $95\%$ RT fixed-design bootstrap interval.

figure[figure omitted — 424 chars of source]

We observe that in more turbulent times (e.g.\ August, 2019), the estimated VaR amplifies. In such volatile periods we expect the estimation risk to increase and, accordingly, we find wider bootstrap confidence intervals. In this regard, although not directly related to the bootstrap procedures considered it is worth mentioning that a proposal for monitoring VaR estimates over time can be found in HogaDemetrescu.

remarkThe point estimate $\hat{\alpha}^+_n$ in Table (ref) is close to zero, indicating that the parameter $\alpha^+_0$ may lie on the boundary. In view of Assumption (ref), the proposed bootstrap method becomes invalid when a parameter lies on the boundary of the parameter space. cavaliere2020bootstrap modify the fixed volatility bootstrap to account for nuisance parameters on the boundary by shrinking the estimates for bootstrap sampling toward the boundary at an appropriate rate. Following their suggestion and using instead $\theta^*_n=(\hat{\omega}_n,0,\hat{\alpha}^-_n,\hat{\beta}_n)'$ to generate the bootstrap sample, does not notably alter the bootstrap results presented in this subsection.

Concluding Remarks

In this paper we study the two-step estimation procedure of francq2015risk associated with the conditional VaR. In the first step, the conditional volatility parameters are estimated by QMLE, while the second step corresponds to approximating the quantile of the innovations' distribution by the empirical quantile of the residuals. A fixed-design residual bootstrap method is proposed to mimic the finite sample distribution of the two-step estimator and its consistency is proven under mild assumptions. In addition, an algorithm is provided for the construction of bootstrap intervals for the conditional VaR to take into account the uncertainty induced by estimation. Three interval types are suggested and a large-scale simulation study is conducted to investigate their performance in finite samples. We find that the equal-tailed percentile interval based on the fixed-design residual bootstrap tends to fall short of its nominal value, whereas the corresponding interval based on reversed tails yields accurate average coverage combined with the shortest average length. Although the result seems counter-intuitive at first, it is in line with the theoretical findings of falk1991coverage. In the simulation study we also consider the recursive-design residual bootstrap. It turns out that the recursive-design and the fixed-design bootstrap perform similar in terms of average coverage. Yet in smaller samples the fixed-design scheme leads on average to shorter intervals. Further, the interval estimation by means of the fixed-design residual bootstrap is illustrated in an empirical application to daily returns of the French stock index CAC 40.

Natural extensions of this work are encompassing other risk measures such as Expected Shortfall heinemann2018residual and developing a bootstrap procedure for the one-step estimator of francq2015risk. Further, it is worthwhile to consider a smoothed bootstrap version in the spirit of hall1989smoothing, which offers potential gains in accuracy. The latter two extensions are left for future research, as is the question how to extend the fixed-design bootstrap in order to give valid bootstrap inference when the iid assumption does not hold.

Acknowledgements

The authors thank Franz Palm, Hanno Reuvers, Jean-Michel Zako\"ian and Christian Francq for useful comments and suggestions as well as Dewi Peerlings and Benoit Duvocelle for computational support during the revision stage. In addition, the authors are grateful to the editors Oliver Linton and Torben Andersen, one of the associate editors and to two anonymous referees for their constructive remarks.

This research was financially supported by the Netherlands Organisation for Scientific Research (NWO).

thebibliography\bibitem[\citeauthoryear{Bahadur}{Bahadur}{1966}]{bahadur1966note} Bahadur, R.R. (1966). \newblock A note on quantiles in large samples. \newblock {\em The Annals of Mathematical Statistics\/} {\em 37\/}(3), 577--580. \bibitem[\citeauthoryear{Bardet, Kamila, and Kengne}{Bardet et al.}{2020}]{Bardet2020} Bardet, J.M., K. Kamila, and W. Kengne (2020). \newblock Consistent model selection criteria and goodness-of-fit test for common time series models. \newblock {\em Electronic Journal of Statistics\/} {\em 14\/}(1), 2009--2052. \bibitem[\citeauthoryear{Berkes and Horv{\'a}th}{Berkes and Horv{\'a}th}{2003}]{berkes2003limit} Berkes, I. and L. Horv{\'a}th (2003). \newblock Limit results for the empirical process of squared residuals in GARCH models. \newblock {\em Stochastic Processes and their Applications\/} {\em 105\/}(2), 271--298. \bibitem[\citeauthoryear{Beutner, Heinemann, and Smeekes}{Beutner et al.}{2019}]{beutner2019technical} Beutner, E., A. Heinemann, and S. Smeekes (2019). \newblock A general framework for prediction in time series models. \newblock Working paper, Maastricht University, \url{https://arxiv.org/pdf/1902.01622.pdf}. \bibitem[\citeauthoryear{Beutner, Heinemann, and Smeekes}{Beutner et al.}{2021}]{beutner2017justification} Beutner, E., A. Heinemann, and S. Smeekes (2021). \newblock A justification of conditional confidence intervals. \newblock {\em Electronic Journal of Statistics\/} {\em 15\/}(1), 2517--2565. \bibitem[\citeauthoryear{Billingsley}{Billingsley}{1986}]{billingsley1986probability} Billingsley, P. (1986). \newblock {\em Probability and Measure\/} (2nd ed.). \newblock New York: John Wiley & Sons. \bibitem[\citeauthoryear{Bollerslev}{Bollerslev}{1986}]{bollerslev1986generalized} Bollerslev, T. (1986). \newblock Generalized autoregressive conditional heteroskedasticity. \newblock {\em Journal of Econometrics\/} {\em 31\/}(3), 307--327. \bibitem[\citeauthoryear{Cavaliere, Nielsen, Pedersen, and Rahbek}{Cavaliere et al.}{2022}]{cavaliere2020bootstrap} Cavaliere, G., H.B. Nielsen, R.S. Pedersen, and A. Rahbek (2022). \newblock Bootstrap inference on the boundary of the parameter space, with application to conditional volatility models. \newblock {\em Journal of Econometrics\/} {\em 227\/}(1), 241--263. \bibitem[\citeauthoryear{Cavaliere, Pedersen, and Rahbek}{Cavaliere et al.}{2018}]{cavaliere2018fixed} Cavaliere, G., R.S. Pedersen, and A. Rahbek (2018). \newblock The fixed volatility bootstrap for a class of ARCH($q$) models. \newblock {\em Journal of Time Series Analysis\/} {\em 39}, 920--941. \bibitem[\citeauthoryear{Cho and White}{Cho and White}{2011}]{cho2011generalized} Cho, J.S. and H. White (2011). \newblock Generalized runs tests for the iid hypothesis. \newblock {\em Journal of Econometrics\/} {\em 162\/}(2), 326--344. \bibitem[\citeauthoryear{Christoffersen and Gon{\c{c}}alves}{Christoffersen and Gon{\c{c}}alves}{2005}]{christoffersen2005estimation} Christoffersen, P. and S. Gon{\c{c}}alves (2005). \newblock Estimation risk in financial risk management. \newblock {\em The Journal of Risk\/} {\em 7\/}(3), 1--28. \bibitem[\citeauthoryear{Corradi and Iglesias}{Corradi and Iglesias}{2008}]{corradi2008bootstrap} Corradi, V. and E.M. Iglesias (2008). \newblock Bootstrap refinements for QML estimators of the GARCH(1,1) parameters. \newblock {\em Journal of Econometrics\/} {\em 144\/}(2), 500--510. \bibitem[\citeauthoryear{Cs{\"o}rg{\H o} and R{\'e}v{\'e}sz}{Cs{\"o}rg{\H o} and R{\'e}v{\'e}sz}{1981}]{csorgo1981strong} Cs{\"o}rg{\H o}, M. and P. R{\'e}v{\'e}sz (1981). \newblock {\em Strong Approximations in Probability and Statistics}. \newblock Budapest: Akad{\'e}miai Kiad{\'o}. \bibitem[\citeauthoryear{Davidson and Flachaire}{Davidson and Flachaire}{2008}]{DavidsonFlachaire08} Davidson, R. and E. Flachaire (2008). \newblock The wild bootstrap, tamed at last. \newblock {\em Journal of Econometrics\/} {\em 146}, 162--169. \bibitem[\citeauthoryear{Ding, Granger, and Engle}{Ding et al.}{1993}]{ding1993long} Ding, Z., C.W. Granger, and R.F. Engle (1993). \newblock A long memory property of stock market returns and a new model. \newblock {\em Journal of Empirical Finance\/} {\em 1\/}(1), 83--106. \bibitem[\citeauthoryear{Engle}{Engle}{1982}]{engle1982autoregressive} Engle, R.F. (1982). \newblock Autoregressive conditional heteroscedasticity with estimates of the variance of {U}nited {K}ingdom inflation. \newblock {\em Econometrica\/} {\em 50\/}(4), 987--1007. \bibitem[\citeauthoryear{Escanciano}{Escanciano}{2009}]{Escanciano2009} Escanciano, J.C. (2009). \newblock Quasi-maximum likelihood estimation of semi-strong GARCH model. \newblock {\em Econometric Theory\/} {\em 25\/}(2), 561--570. \bibitem[\citeauthoryear{Falk and Kaufmann}{Falk and Kaufmann}{1991}]{falk1991coverage} Falk, M. and E. Kaufmann (1991). \newblock Coverage probabilities of bootstrap-confidence intervals for quantiles. \newblock {\em The Annals of Statistics\/} {\em 19\/}(1), 485--495. \bibitem[\citeauthoryear{Francq, Horv{\'a}th, and Zako{\"\i}an}{Francq et al.}{2016}]{francq2016variance} Francq, C., L. Horv{\'a}th, and J.M. Zako{\"\i}an (2016). \newblock Variance targeting estimation of multivariate GARCH models. \newblock {\em Journal of Financial Econometrics\/} {\em 14\/}(2), 353--382. \bibitem[\citeauthoryear{Francq and Zako{\"\i}an}{Francq and Zako{\"\i}an}{2004}]{francq2004maximum} Francq, C. and J.M. Zako{\"\i}an (2004). \newblock Maximum likelihood estimation of pure \uppercase{Garch} and \uppercase{Arma}-\uppercase{Garch} processes. \newblock {\em Bernoulli\/} {\em 10\/}(4), 605--637. \bibitem[\citeauthoryear{Francq and Zako\"ian}{Francq and Zako\"ian}{2011}]{francq2011garch} Francq, C. and J.M. Zako\"ian (2011). \newblock {\em \uppercase{GARCH} Models: Structure, Statistical Inference and Financial Applications}. \newblock Chichester: John Wiley & Sons. \bibitem[\citeauthoryear{Francq and Zako{\"\i}an}{Francq and Zako{\"\i}an}{2015}]{francq2015risk} Francq, C. and J.M. Zako{\"\i}an (2015). \newblock Risk-parameter estimation in volatility models. \newblock {\em Journal of Econometrics\/} {\em 184\/}(1), 158--173. \bibitem[\citeauthoryear{Francq and Zako{\"\i}an}{Francq and Zako{\"\i}an}{2016}]{FrancqZakoian2016} Francq, C. and J.M. Zako{\"\i}an (2016). \newblock Estimating multivariate volatility models equation by equation. \newblock {\em Journal of the Royal Statistical Society, Series B\/} {\em 76\/}(3), 613--635. \bibitem[\citeauthoryear{Francq and Zakoïan}{Francq and Zakoïan}{2022}]{FRANCQ202247} Francq, C. and J.M. Zakoïan (2022). \newblock Testing the existence of moments for garch processes. \newblock {\em Journal of Econometrics\/} {\em 227\/}(1), 47--64. \bibitem[\citeauthoryear{Friedrich, Smeekes, and Urbain}{Friedrich et al.}{2020}]{FSU20} Friedrich, M., S. Smeekes, and J.P. Urbain (2020). \newblock Autoregressive wild bootstrap inference for nonparametric trends. \newblock {\em Journal of Econometrics\/} {\em 214}, 81--109. \bibitem[\citeauthoryear{Gao and Song}{Gao and Song}{2008}]{gao2008estimation} Gao, F. and F. Song (2008). \newblock Estimation risk in \uppercase{GARCH} \uppercase{V}a\uppercase{R} and \uppercase{ES} estimates. \newblock {\em Econometric Theory\/} {\em 24\/}(5), 1404--1424. \bibitem[\citeauthoryear{Geweke}{Geweke}{1986}]{geweke1986comment} Geweke, J. (1986). \newblock Comment on: modelling the persistence of conditional variances. \newblock {\em Econometric Reviews\/} {\em 5}, 57--61. \bibitem[\citeauthoryear{Glosten, Jagannathan, and Runkle}{Glosten et al.}{1993}]{glosten1993relation} Glosten, L.R., R. Jagannathan, and D.E. Runkle (1993). \newblock On the relation between the expected value and the volatility of the nominal excess return on stocks. \newblock {\em The Journal of Finance\/} {\em 48\/}(5), 1779--1801. \bibitem[\citeauthoryear{Gonçalves and Kilian}{Gonçalves and Kilian}{2004}]{GONCALVES2004} Gonçalves, S. and L. Kilian (2004). \newblock Bootstrapping autoregressions with conditional heteroskedasticity of unknown form. \newblock {\em Journal of Econometrics\/} {\em 123\/}(1), 89--120. \bibitem[\citeauthoryear{Hall, DiCiccio, and Romano}{Hall et al.}{1989}]{hall1989smoothing} Hall, P., T.J. DiCiccio, and J.P. Romano (1989). \newblock On smoothing and the bootstrap. \newblock {\em The Annals of Statistics\/} {\em 17\/}(2), 692--704. \bibitem[\citeauthoryear{Hall and Heyde}{Hall and Heyde}{1980}]{hall1980martingale} Hall, P. and C.C. Heyde (1980). \newblock {\em Martingale Limit Theory and its Application}. \newblock New York: Academic Press. \bibitem[\citeauthoryear{Hall and Martin}{Hall and Martin}{1988}]{hall1988bootstrap} Hall, P. and M.A. Martin (1988). \newblock On bootstrap resampling and iteration. \newblock {\em Biometrika\/} {\em 75\/}(4), 661--671. \bibitem[\citeauthoryear{Hall and Yao}{Hall and Yao}{2003}]{hall2003inference} Hall, P. and Q. Yao (2003). \newblock Inference in \uppercase{ARCH} and \uppercase{GARCH} models with heavy--tailed errors. \newblock {\em Econometrica\/} {\em 71\/}(1), 285--317. \bibitem[\citeauthoryear{Hamadeh and Zako{\"\i}an}{Hamadeh and Zako{\"\i}an}{2011}]{hamadeh2011asymptotic} Hamadeh, T. and J.M. Zako{\"\i}an (2011). \newblock Asymptotic properties of \uppercase{LS} and \uppercase{QML} estimators for a class of nonlinear \uppercase{GARCH} processes. \newblock {\em Journal of Statistical Planning and Inference\/} {\em 141\/}(1), 488--507. \bibitem[\citeauthoryear{Hartz, Mittnik, and Paolella}{Hartz et al.}{2006}]{hartz2006accurate} Hartz, C., S. Mittnik, and M. Paolella (2006). \newblock Accurate value-at-risk forecasting based on the normal-\uppercase{GARCH} model. \newblock {\em Computational Statistics & Data Analysis\/} {\em 51\/}(4), 2295--2312. \bibitem[\citeauthoryear{Heinemann and Telg}{Heinemann and Telg}{2018}]{heinemann2018residual} Heinemann, A. and S. Telg (2018). \newblock A residual bootstrap for conditional expected shortfall. \newblock arXiv Preprint 1811.11557. \bibitem[\citeauthoryear{Hetland, Pedersen, and Rahbek}{Hetland et al.}{ress}]{HETLAND2021} Hetland, S., R.S. Pedersen, and A. Rahbek (in press). \newblock Dynamic conditional eigenvalue garch. \newblock {\em Journal of Econometrics\/}. \bibitem[\citeauthoryear{Hidalgo and Zaffaroni}{Hidalgo and Zaffaroni}{2007}]{hidalgo2007goodness} Hidalgo, J. and P. Zaffaroni (2007). \newblock A goodness-of-fit test for \uppercase{ARCH}($\infty$) models. \newblock {\em Journal of Econometrics\/} {\em 141\/}(2), 835--875. \bibitem[\citeauthoryear{Hjort and Pollard}{Hjort and Pollard}{2011}]{hjort2011asymptotics} Hjort, N.L. and D. Pollard (2011). \newblock Asymptotics for minimisers of convex processes. \newblock {\em Preprint arXiv:1107.3806v1\/}. \bibitem[\citeauthoryear{Hoga and Demetrescu}{Hoga and Demetrescu}{2023}]{HogaDemetrescu} Hoga, Y. and M. Demetrescu (2023). \newblock Monitoring value-at-risk and expected shortfall forecasts. \newblock {\em Management Science\/} {\em 69\/}(5), 2954--2971. \bibitem[\citeauthoryear{Jeong}{Jeong}{2017}]{jeong2017residual} Jeong, M. (2017). \newblock Residual-based \uppercase{GARCH} bootstrap and second order asymptotic refinement. \newblock {\em Econometric Theory\/} {\em 33\/}(3), 779--790. \bibitem[\citeauthoryear{Jim\'{e}nez-Gamero, Lee, and Meintanis}{Jim\'{e}nez-Gamero et al.}{2020}]{Maria2019} Jim\'{e}nez-Gamero, M.D., S. Lee, and S.G. Meintanis (2020). \newblock Goodness-of-fit tests for parametric specifications of conditionally heteroscedastic models. \newblock {\em TEST\/} {\em 29}, 682--703. \bibitem[\citeauthoryear{Koenker and Xiao}{Koenker and Xiao}{2006}]{koenker2006quantile} Koenker, R. and Z. Xiao (2006). \newblock Quantile autoregression. \newblock {\em Journal of the American Statistical Association\/} {\em 101\/}(475), 980--990. \bibitem[\citeauthoryear{Kreiss}{Kreiss}{2016}]{kreiss2015discussion} Kreiss, J.P. (2016). \newblock Discussion: bootstrap prediction intervals for linear, nonlinear and nonparametric autoregressions. \newblock {\em Journal of Statistical Planning and Inference\/} {\em 177}, 28--30. \bibitem[\citeauthoryear{Lahiri}{Lahiri}{2003}]{Lahiri03} Lahiri, S.N. (2003). \newblock {\em Resampling Methods for Dependent Data}. \newblock New York: Springer-Verlag. \bibitem[\citeauthoryear{Li, Peng, and Song}{Li et al.}{ress}]{li_peng_song_2022} Li, S., L. Peng, and X. Song (in press). \newblock Simultaneous confidence bands for conditional value-at-risk and expected shortfall. \newblock {\em Econometric Theory\/}. \bibitem[\citeauthoryear{Linton, Pan, and Wang}{Linton et al.}{2010}]{Linton2010} Linton, O., J. Pan, and H. Wang (2010). \newblock Estimation for a nonstationary semi-strong {GARCH}(1,1) model with heavy-tailed errors. \newblock {\em Econometric Theory\/} {\em 26\/}(1), 1--28. \bibitem[\citeauthoryear{Mammen}{Mammen}{1993}]{Mammen93} Mammen, E. (1993). \newblock Bootstrap and wild bootstrap for high dimensional linear models. \newblock {\em Annals of Statistics\/} {\em 21}, 255--285. \bibitem[\citeauthoryear{McNeil and Frey}{McNeil and Frey}{2000}]{McNeil2002} McNeil, A.J. and R. Frey (2000). \newblock Estimation of tail-related risk measures for heteroscedastic financial time series: an extreme value approach. \newblock {\em Journal of Empirical Finance\/} {\em 7\/}(3), 271 -- 300. \bibitem[\citeauthoryear{Nelson}{Nelson}{1991}]{nelson1991conditional} Nelson, D.B. (1991). \newblock Conditional heteroskedasticity in asset returns: A new approach. \newblock {\em Econometrica: Journal of the Econometric Society\/} {\em 59\/}(2), 347--370. \bibitem[\citeauthoryear{Pantula}{Pantula}{1986}]{pantula1986modeling} Pantula, S.G. (1986). \newblock Modeling the persistence of conditional variances: a comment. \newblock {\em Econometric Reviews\/} {\em 5}, 79--97. \bibitem[\citeauthoryear{Pascual, Romo, and Ruiz}{Pascual et al.}{2006}]{pascual2006bootstrap} Pascual, L., J. Romo, and E. Ruiz (2006). \newblock Bootstrap prediction for returns and volatilities in \uppercase{garch} models. \newblock {\em Computational Statistics & Data Analysis\/} {\em 50\/}(9), 2293--2312. \bibitem[\citeauthoryear{Pesaran}{Pesaran}{2015}]{pesaran2015time} Pesaran, M.H. (2015). \newblock {\em Time Series and Panel Data Econometrics}. \newblock Oxford: Oxford University Press. \bibitem[\citeauthoryear{Phillips}{Phillips}{1979}]{phillips1979sampling} Phillips, P.C.B. (1979). \newblock The sampling distribution of forecasts from a first-order autoregression. \newblock {\em Journal of Econometrics\/} {\em 9\/}(3), 241--261. \bibitem[\citeauthoryear{Roussas}{Roussas}{1997}]{roussas1997course} Roussas, G.G. (1997). \newblock {\em A Course in Mathematical Statistics\/} (2nd ed.). \newblock San Diego: Academic Press. \bibitem[\citeauthoryear{Shao}{Shao}{2010}]{Shao10} Shao, X. (2010). \newblock The dependent wild bootstrap. \newblock {\em Journal of the American Statistical Association\/} {\em 105}, 218--235. \bibitem[\citeauthoryear{Shimizu}{Shimizu}{2009}]{shimizu2009bootstrapping} Shimizu, K. (2009). \newblock {\em Bootstrapping Stationary \uppercase{ARMA}--\uppercase{GARCH} Models}. \newblock Springer. \bibitem[\citeauthoryear{Silverman}{Silverman}{1986}]{silverman1986density} Silverman, B. (1986). \newblock {\em Density Estimation for Statistics and Data Analysis Estimation Density}. \newblock London: Chapman and Hall. \bibitem[\citeauthoryear{Spierdijk}{Spierdijk}{2016}]{spierdijk2016confidence} Spierdijk, L. (2016). \newblock Confidence intervals for \uppercase{ARMA}--\uppercase{GARCH} value-at-risk: the case of heavy tails and skewness. \newblock {\em Computational Statistics & Data Analysis\/} {\em 100}, 545--559. \bibitem[\citeauthoryear{Xiong and Li}{Xiong and Li}{2008}]{xiong2008some} Xiong, S. and G. Li (2008). \newblock Some results on the convergence of conditional distributions. \newblock {\em Statistics & Probability Letters\/} {\em 78\/}(18), 3249--3253. \bibitem[\citeauthoryear{Zako{\"\i}an}{Zako{\"\i}an}{1994}]{zakoian1994threshold} Zako{\"\i}an, J.M. (1994). \newblock Threshold heteroskedastic models. \newblock {\em Journal of Economic Dynamics and Control\/} {\em 18\/}(5), 931--955.