The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
113,837 characters
Journal of Econometrics
\begin{center}
\scshape
\LARGE{\textbf{A Residual Bootstrap for Conditional Value-at-Risk}}
\normalfont
\end{center}
\begin{center}
\par\vspace{0.5cm}
\text{Eric Beutner$^*$}
\hspace{1.5cm}
\text{Alexander Heinemann$^{**}$}
\hspace{1.4cm}
\text{Stephan Smeekes$^{***}$}
\par\vspace{2cm}
\par
\par
\text{\today}
\vspace{1cm}
\end{center}
\begin{abstract}
A fixed-design residual bootstrap method is proposed for the two-step estimator of \cite{francq2015risk} associated with the conditional Value-at-Risk. The bootstrap's consistency is proven for a general class of volatility models and intervals are constructed for the conditional Value-at-Risk. A simulation study reveals that the equal-tailed percentile bootstrap interval tends to fall short of its nominal value. In contrast, the reversed-tails bootstrap interval yields accurate coverage. We also compare the theoretically analyzed fixed-design bootstrap with the recursive-design bootstrap. It turns out that the fixed-design bootstrap performs equally well in terms of average coverage, yet leads on average to shorter intervals in smaller samples. An empirical application illustrates the interval estimation.\\ \\
\textbf{Key words:} Residual bootstrap; Value-at-Risk; GARCH\\
\textbf{JEL codes:} C14; C15; C58
\end{abstract}
\begingroup
\renewcommand{}\footnote{\hspace{-0.7cm}
$^{\textcolor{white}{**}*}$Department of Econometrics and Data Science, Vrije Universiteit Amsterdam, De Boelelaan 1105
1081 HV Amsterdam, Netherlands. E-mail address: [email removed]\\
$^{\textcolor{white}{*}**}$a.s.r. Archimedeslaan 10, 3584 BA Utrecht. E-mail address: [email removed]\\
$^{***}$Department of Quantitative Economics, Maastricht University, Tongersestraat 53, 6211 LM Maastricht, Netherlands. E-mail address: [email removed] (corresponding author)
}
\addtocounter{footnote}{-1}
\endgroup
\newpage
\doublespacing
\section{Introduction}
\label{sec:4.1}
Risk management has tremendously developed in past decades becoming an increasing practice. With minimum capital requirements being enforced by current legislation (Basel III and Solvency II), financial institutions and insurance companies monitor risk by using conventional measures such as Value-at-Risk (VaR). Typically, the volatility dynamics are specified by a (semi-)parametric model leading to conditional risk measure versions. For GARCH-type models the conditional VaR reduces to the conditional volatility scaled by a quantile of the innovations' distribution. The latter is conventionally treated as additional parameter and forms together with the others the \textit{risk parameter} \citep{francq2015risk}. The true parameters are generally unknown and need to be estimated to obtain an estimate for the conditional VaR.
Clearly, this VaR evaluation is subject to estimation risk that needs to be quantified for appropriate risk management.
Whereas an estimator based on a single step is available after re-parameterization \citep{francq2015risk}, a widely used approach is the following two-step estimation procedure. First, the parameters of the stochastic volatility model are estimated. Arguably the most popular estimation method in a GARCH-type setting is the Gaussian quasi-maximum-likelihood (QML) method. Based on the model's residuals the quantile is estimated by its empirical counterpart in a second step. For realistic sample sizes (e.g.\ $500$ or $1{,}000$ daily observations) the estimators are subject to considerable estimation risk. In particular, the estimation uncertainty associated with the quantile estimator is substantial for extreme quantiles (e.g.\ $\leq 5\%$).
To quantify the uncertainty around the point estimators, one traditionally relies on asymptotic theory while replacing the unknown quantities in the limiting distribution by consistent estimates. An alternative approach -- frequently employed in practice -- is based on a bootstrap approximation. Regarding the estimators of the GARCH parameters, various bootstrap methods have been studied to approximate the estimators' finite sample distribution including the subsample bootstrap \citep{hall2003inference}, the block bootstrap \citep{corradi2008bootstrap}, the wild bootstrap \citep{shimizu2009bootstrapping} and the residual bootstrap. The residual bootstrap method is particularly popular and can be further divided into recursive \citep{pascual2006bootstrap,hidalgo2007goodness,jeong2017residual} and fixed \citep{shimizu2009bootstrapping,cavaliere2018fixed} design. Whereas in the former the bootstrap observations are generated recursively using the estimated volatility dynamics, the latter design keeps the dynamics of the bootstrap samples fixed at the value of the original series. Further recent applications of fixed or recursive bootstrap designs or variants thereof to conditional volatility models can be found in \cite{HETLAND2021}, \cite{FRANCQ202247} and \cite{cavaliere2020bootstrap}.
The estimation of the quantile and the conditional VaR have received only selected attention in the bootstrap literature and proposed bootstrap methods have been, to the best of our knowledge, exclusively investigated by means of simulation. \cite{christoffersen2005estimation} examine various quantile estimators and construct intervals for the conditional VaR using a recursive-design residual bootstrap method. In addition, \cite{hartz2006accurate} presume the innovation distribution to be standard normal such that the quantile parameter is known; they propose a resampling method based on a residual bootstrap and a bias-correction step to account for deviations from the normality assumption. In contrast, \cite{spierdijk2016confidence} develops an $m$-out-of-$n$ without-replacement bootstrap to construct confidence intervals for ARMA-GARCH VaR.
This paper proposes a fixed-design residual bootstrap method to mimic the finite sample distribution of the two-step estimator and provides an algorithm for the construction of bootstrap intervals for the conditional VaR. The proposed bootstrap method is proven to be consistent for a general class of volatility models. In particular, our framework does not only encompass GARCH but also several GARCH extensions such as the threshold GARCH (T-GARCH) of \cite{zakoian1994threshold} and the GJR-GARCH named after Glosten, Jagannathan and Runkle (\citeyear{glosten1993relation}). The bootstrap consistency is established under a set of mild assumptions, which relaxes moment conditions on the innovations imposed in the GARCH bootstrap literature. To the best of our knowledge this paper is the first to theoretically validate the residual bootstrap for the quantile and the conditional VaR.
The remainder of the paper is organized as follows. Section \ref{sec:4.2} specifies the model and the conditional VaR is derived. The two-step estimation procedure is described in Section \ref{sec:4.3} and asymptotic theory is provided under mild assumptions. In Section \ref{sec:4.4}, a fixed-design residual bootstrap method is proposed and proven to be consistent. Further, bootstrap intervals are constructed for the conditional VaR and extensions to the bootstrap methods presented here are discussed. A simulation study is conducted in Section \ref{sec:4.5} and an empirical application illustrates the interval estimation based on the fixed-design residual bootstrap. Section \ref{sec:4.6} concludes. Appendix \ref{app:1} contains proof of the main results, whereas Supplementary Appendix \ref{app:4.A} contains auxiliary results and their proofs. Finally, Supplementary Appendix \ref{app:4.B} is devoted to the related recursive-design residual bootstrap and Supplementary Appendix \ref{sec add simulation} contains additional simulation results.
\section{Model}
\label{sec:4.2}
We consider a conditional volatility model of the form
\begin{align}
\label{eq:4.2.1}
\epsilon_t = \sigma_t\eta_t
\end{align}
with $t\in \mathbb{Z}$, where $\{\epsilon_t\}$ denotes the sequence of log-returns, $\{\sigma_t\}$ is a volatility process and $\{\eta_t\}$ is a sequence of independent and identically distributed (iid) variables satisfying $\mathbb{E}\big[\eta_t^2\big]=1$. The volatility is presumed to be a measurable function of past observations
\begin{align}
\label{eq:4.2.2}
\sigma_{t}=\sigma_{t}(\theta_0)=\sigma(\epsilon_{t-1},\epsilon_{t-2},\dots;\theta_0)
\end{align}
with $\sigma:\mathbb{R}^\infty\times \Theta\to(0,\infty)$ and $\theta_0$ denotes the true parameter vector belonging to the parameter space $\Theta \subset \mathbb{R}^r$, $r \in \mathbb{N}$. Subsequently, we consider two examples for the functional form of \eqref{eq:4.2.2}: the well-known GARCH model (\citeauthor{engle1982autoregressive}, \citeyear{engle1982autoregressive}; \citeauthor{bollerslev1986generalized}, \citeyear{bollerslev1986generalized}) and the T-GARCH model of \cite{zakoian1994threshold}. Whereas the first is frequently applied in practice, the second is motivated by our empirical application (see Section \ref{sec:4.5.2}).
\begin{example}
\label{ex:4.1}
Suppose $\{\epsilon_t\}$ follows a GARCH$(1,1)$ process given by \eqref{eq:4.2.1} and $\sigma_{t}^2 = \omega_0 + \alpha_0 \epsilon_{t-1}^2+ \beta_0 \sigma_{t-1}^2$, where $\theta_0 = (\omega_0,\alpha_0,\beta_0)'\in (0,\infty)\times[0,\infty)\times[0,1)$. The recursive structure implies $\sigma_{t}=\sigma(\epsilon_{t-1},\epsilon_{t-2},\dots;\theta_0) = \sqrt{\sum_{k=1}^\infty \beta_0^{k-1} \big(\omega_0 + \alpha_0 \epsilon_{t-k}^2\big)}$.
\end{example}
\begin{example}
\label{ex:4.2}
Suppose $\{\epsilon_t\}$ follows a T-GARCH$(1,1)$ process given by \eqref{eq:4.2.1} and $\sigma_{t} =\: \omega_0 + \alpha_0^+ \epsilon_{t-1}^+ +\alpha_0^- \epsilon_{t-1}^- + \beta_0 \sigma_{t-1}$ with parameters $\theta_0 =(\omega_0,\alpha_0^+,\alpha_0^- ,\beta_0)'\in (0,\infty)\times[0,\infty)\times[0,\infty)\times[0,1)$ and $ \epsilon_{t}^+=\max\{\epsilon_{t},0\}$ and $ \epsilon_{t}^-=\max\{-\epsilon_{t},0\}$. The model's recursive structure yields $\sigma_{t}=\sigma(\epsilon_{t-1},\epsilon_{t-2},\dots;\theta_0) =\! \sum_{k=1}^\infty \beta_0^{k-1} \big(\omega_0 + \alpha_0^+ \epsilon_{t-k}^+ +\alpha_0^-\epsilon_{t-k}^-\big)$.
\end{example}
Throughout the paper, for any cumulative distribution function (cdf), say $G$, we define the generalized inverse by $G^{-1}(u)=\inf\big\{\tau \in \mathbb{R} : G(\tau)\geq u\big\}$ and write $G(\cdot-)$ to denote its left limit. Generally, for an arbitrary real-valued random variable $X$
(e.g.~stock return) with cdf $F_X$, the VaR at level $\alpha \in (0,1)$, is given by $VaR_\alpha(X)=-F_X^{-1}(\alpha)$.
Let $\mathcal{F}_n$ denote the $\sigma$-algebra generated by $\{\epsilon_t, t\leq n\}$. It follows that the conditional VaR of $\epsilon_{n+1}$ given $\mathcal{F}_n$ at level $\alpha \in (0,1)$ is $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n) = \sigma (\epsilon_n,\epsilon_{n-1},\dots;\theta_0) VaR_\alpha(\eta_{n+1})$.
For given $\alpha$, the quantile of $\eta_{n+1}$ is constant and can be treated as a parameter. Thus, denoting the cdf of the $\eta_t$'s by $F$ and setting $\xi_\alpha = F^{-1}(\alpha)$, the conditional VaR of $\epsilon_{n+1}$ given $\mathcal{F}_n$ at level $\alpha$ reduces to
\begin{align}
\label{eq:4.2.4}
VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n) = -\xi_\alpha\: \sigma_{n+1}(\theta_0).
\end{align}
Typically, $\alpha$ is fixed at a sufficiently small level such that $\xi_\alpha<0$. Except for special cases (e.g. normality of $\eta_t$), $\xi_\alpha$ is unknown and needs to estimated just like $\theta_0$.
\section{Estimation}
\label{sec:4.3}
We estimate the parameters $\theta_0$ and $\xi_\alpha$ following the two-step procedure of \citeauthor{francq2015risk} (\citeyear{francq2015risk}, Section 4.2).
In the first step, we estimate the conditional volatility parameter $\theta_0$ by Gaussian QML. This approach is motivated as follows: if the innovations $\{\eta_t\}$ were Gaussian, the variables $\eta_t(\theta) = \epsilon_t/\sigma_t(\theta)$ would be iid $N(0,1)$ whenever $\theta=\theta_0$, where $\sigma_{t}(\theta) = \sigma(\epsilon_{t-1},\dots,\epsilon_{1}, \epsilon_{0},\epsilon_{-1},\dots;\theta)$.
The 'Q' in QML stands for 'quasi' and refers to the fact that $F$ does not need to be the standard normal distribution function. Obviously, given a sample $\epsilon_1, \dots, \epsilon_n$, we generally cannot determine $\sigma_{t}(\theta)$ completely. Replacing the unknown presample observations by arbitrary values, say $\tilde{\epsilon}_t$, $t\leq 0$, we obtain $\tilde{\sigma}_{t}(\theta) = \sigma(\epsilon_{t-1},\dots,\epsilon_{1}, \tilde{\epsilon}_{0},\tilde{\epsilon}_{-1},\dots;\theta)$, which serves as an approximation for $\sigma_{t}(\theta)$. The QML estimator of $\theta_0$ is defined by
\begin{align}
\label{eq:4.3.3}
\hat{\theta}_n=\arg\max_{\theta \in \Theta} \frac{1}{n}\sum_{t=1}^n \tilde{\ell}_t(\theta) \qquad \text{with} \qquad \tilde{\ell}_t(\theta)=-\frac{1}{2}\bigg(\frac{\epsilon_t}{\tilde{\sigma}_t(\theta)}\bigg)^2-\log \tilde{\sigma}_t(\theta).
\end{align}
In the second step, we estimate $\xi_\alpha$ on the basis of the first-step residuals, i.e.~$\hat{\eta}_t=\epsilon_t/\tilde{\sigma}_t(\hat{\theta}_n)$. The empirical $\alpha$-quantile of $\hat{\eta}_1,\dots,\hat{\eta}_n$ is given by
\begin{align}
\label{eq:4.3.4}
\hat{\xi}_{n,\alpha}= \arg \min_{z \in \mathbb{R}} \frac{1}{n} \sum_{t=1}^n \rho_\alpha(\hat{\eta}_t-z),
\end{align}
where $\rho_\alpha(u)=u(\alpha-\mathbbm{1}_{\{u< 0\}})$ is the usual asymmetric absolute loss function (cf.\ \citeauthor{koenker2006quantile}, \citeyear{koenker2006quantile}). Equivalently, we can write $\hat{\xi}_{n,\alpha}=\hat{\mathbbm{F}}_n^{-1}(\alpha)$ with $\hat{\mathbbm{F}}_n(x)=\frac{1}{n}\sum_{t=1}^n \mathbbm{1}_{\{\hat{\eta}_t\leq x\}}$ being the empirical distribution function (edf) of the residuals.
Having obtained estimators for $\theta_0$ and $\xi_\alpha$, we turn to the estimation of the conditional VaR of the one-period ahead observation at level $\alpha$. Whereas the notation $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ stresses the object's conditional nature, we henceforth proceed with the abbreviation $VaR_{n,\alpha}$ for notational convenience. Employing \eqref{eq:4.3.3} -- \eqref{eq:4.3.4} we can estimate $VaR_{n,\alpha}$ by
\begin{align}
\label{eq:4.3.5}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}=-\hat{\xi}_{n,\alpha}\: \tilde{\sigma}_{n+1}\big(\hat{\theta}_n\big).
\end{align}
Clearly, the estimator's large sample properties cannot be studied using traditional tools such as consistency since \eqref{eq:4.3.5} does not permit a limit.
For the subsequent asymptotic analysis, we introduce the following assumptions.
\begin{assumption}{(Compactness)}
\label{as:4.1}
$\Theta$ is a compact subset of $\mathbb{R}^r$.
\end{assumption}
\begin{assumption}{(Stationarity \& Ergodicity)}
\label{as:4.2}
$\{\epsilon_t\}$ is a strictly stationary and ergodic solution of \eqref{eq:4.2.1} with \eqref{eq:4.2.2}.
\end{assumption}
\begin{assumption}{(Volatility process)}
\label{as:4.3}
The function $\sigma:\mathbb{R}^\infty\times \Theta\to(0,\infty)$ is known and for
any real sequence $\{x_i\}$, the function $\theta\to\sigma(x_1,x_2,\dots;\theta)$ is continuous. Almost surely, $\sigma_t(\theta)>\underline{\omega}$ for any $\theta \in \Theta$ and some $\underline{\omega}>0$ and $\mathbb{E}[\sigma_t^s(\theta_0)]<\infty$ for some $s>0$. Moreover, for any $\theta \in \Theta$, we assume $\sigma_t(\theta_0)/\sigma_t(\theta)=1$ almost surely (a.s.) if and only if $\theta=\theta_0$.
\end{assumption}
\begin{assumption}{(Initial conditions)}
\label{as:4.4}
There exists a constant $\rho \in (0,1)$ and a random variable $C_1$ measurable with respect to $\mathcal{F}_0$ and $\mathbb{E}[|C_1|^s]<\infty$ for some $s>0$ such that
\begin{enumerate}[(i)]
\item \label{as:4.4.1} $\sup_{\theta \in \Theta}|\sigma_t(\theta)-\tilde{\sigma}_t(\theta)|\leq C_1 \rho^t$;
\item \label{as:4.4.2} $\theta\to \sigma(x_1, x_2, \dots;\theta)$ has continuous second-order derivatives satisfying
\begin{align*}
\sup_{\theta \in \Theta}\bigg|\bigg|\frac{\partial \sigma_t(\theta)}{\partial \theta}-\frac{\partial \tilde{\sigma}_t(\theta)}{\partial \theta}\bigg|\bigg|\leq C_1 \rho^t, \qquad \quad \sup_{\theta \in \Theta}\bigg|\bigg|\frac{\partial^2 \sigma_t(\theta)}{\partial \theta \partial \theta'}-\frac{\partial^2 \tilde{\sigma}_t(\theta)}{\partial \theta\partial \theta'}\bigg|\bigg|\leq C_1 \rho^t,
\end{align*}
where $||\cdot||$ denotes the Euclidean norm.
\end{enumerate}
\end{assumption}
\begin{assumption}{(Innovation process)}
\label{as:4.5}
The innovations $\{\eta_t\}$ satisfy
\begin{enumerate}[(i)]
\item \label{as:4.5.1} $\eta_t\overset{iid}{\sim}F$ with $F$ being continuous, $\mathbb{E}\big[\eta_t^2\big]=1$ and $\eta_t$ is independent of $\{\epsilon_u:u<t\}$;
\item \label{as:4.5.2} $\eta_t$ admits a density $f$ which is continuous and strictly positive around $\xi_\alpha<0$;
\item \label{as:4.5.3} $\mathbb{E}\big[\eta_t^4\big]<\infty$.
\end{enumerate}
\end{assumption}
\begin{assumption}{(Interior)}
\label{as:4.6}
$\theta_0$ belongs to the interior of $\Theta$ denoted by $\mathring{\Theta}$.
\end{assumption}
\begin{assumption}{(Non-degeneracy)}
\label{as:4.7}
There does not exist a non-zero $\lambda\in \mathbb{R}^r$ such that $\lambda'\frac{\partial \sigma_t(\theta_0)}{\partial \theta}=0$ a.s.
\end{assumption}
\begin{assumption}{(Monotonicity)}
\label{as:4.8}
For any real sequence $\{x_i\}$ and for any $\theta_1,\theta_2 \in \Theta$ satisfying $\theta_1\leq \theta_2$ componentwise, we have $\sigma(x_1,x_2,\dots;\theta_1)\leq \sigma(x_1,x_2,\dots;\theta_2)$.
\end{assumption}
\begin{assumption}{(Moments)}
\label{as:4.9}
There exists a neighborhood $\mathscr{V}(\theta_0)$ of $\theta_0$ such that the following variables have finite expectation
\begin{align*}
\text{(i)}\sup_{\theta \in \mathscr{V}(\theta_0)}\bigg|\frac{ \sigma_t(\theta_0)}{\sigma_t(\theta)}\bigg|^a, \qquad \;\;\: \text{(ii)} \sup_{\theta \in \mathscr{V}(\theta_0)}\bigg|\bigg|\frac{1}{\sigma_t(\theta)}\frac{\partial \sigma_t(\theta)}{\partial \theta}\bigg|\bigg|^{b}, \qquad \;\;\: \text{(iii)} \sup_{\theta \in \mathscr{V}(\theta_0)}\bigg|\bigg|\frac{1}{\sigma_t(\theta)}\frac{\partial^2 \sigma_t(\theta)}{\partial \theta \partial \theta'}\bigg|\bigg|^c
\end{align*}
for some $a$, $b$, $c$ (to be specified).\footnote{Note that the variables in (i)--(iii) are strictly stationary (\citeauthor{francq2011garch}, \citeyear{francq2011garch}, p.\ 181/406).}
\end{assumption}
\begin{assumption}{(Scaling Stability)}
\label{as:4.10}
There exists a function $g$ such that for any $\theta \in \Theta$, for any $\lambda>0$, and any real sequence $\{x_i\}$
\begin{align*}
\lambda \sigma(x_1,x_2,\dots;\theta)=\sigma(x_1,x_2,\dots;\theta_\lambda),
\end{align*}
where $\theta_\lambda=g(\theta,\lambda)$ and $g$ is differentiable in $\lambda$.
\end{assumption}
The previous set of assumptions is comparable to the conditions imposed by \cite{francq2015risk}. Assumption \ref{as:4.3} calls for a correct specification of the volatility structure. If the researcher incorrectly specifies a volatility function $\varsigma(\dots;\vartheta)$ instead, the estimator of the misspecified conditional volatility model $\hat{\vartheta}_n$ will converge to a pseudo-true value, i.e.\ $\vartheta_0=\arg \min_\vartheta \mathbb{E}\big[\frac{1}{2}\frac{\epsilon_t^2}{\varsigma_t^2(\vartheta)}+\log \varsigma_t(\vartheta)\big]$. The corresponding edf of the residuals $\frac{1}{n}\sum_{t=1}^n \mathbbm{1}_{\{\epsilon_t/\varsigma_t(\hat{\vartheta}_n)\leq x\}}$ converges to $\bar{F}(x)=\mathbb{E}\big[F\big(x\frac{\varsigma_t(\vartheta_0)}{\sigma_t(\theta_0)}\big)\big]$ in view of Lemma \ref{lem:4.1} while the $\alpha$-quantile estimator converges to $\bar{F}^{-1}(\alpha)$, which is generally different from $F^{-1}(\alpha)$. Thus, the correct specification of the volatility function is crucial and one can test for it using the recently developed test by \cite{Maria2019}; for further recent results on goodness-of-fit testing for GARCH models see \cite{Bardet2020}. Regarding the bootstrap method to be developed below, it is plausible to expect that in the presence of a misspecified conditional volatility model it will be consistent for the pseudo-true values, although we do not provide a rigorous proof.
Regarding the innovation process we do not need to assume $\mathbb{E}[\eta_t]=0$ (cf.\ \citeauthor{francq2004maximum}, \citeyear{francq2004maximum}, Remark 2.5).
The iid condition in Assumption \ref{as:4.5}\eqref{as:4.5.1} is vital for \eqref{eq:4.2.4} to hold and is the basis of the residual bootstrap in Section \ref{sec:4.4.1}.
Under correct specification of the volatility process the iid assumption imposed on the innovations can be tested for by considering the errors $\epsilon_t/ \sigma(\epsilon_{t-1},\epsilon_{t-2},\dots;\hat{\theta}_n)$, $t=1,..,n$, and applying the test of \cite{cho2011generalized}.
Whereas \cite{cavaliere2018fixed} assume the existence of the sixth moment of $\eta_t$ for the fixed-design bootstrap in ARCH($q$) models, we only require the fourth moment to be finite in Assumption \ref{as:4.5}(\ref{as:4.5.3}).
We assume $\theta_0$ belongs to the interior of the parameter space in Assumption \ref{as:4.6}. Parameters on the boundary yield non-standard problems, which require special treatment. \cite{cavaliere2020bootstrap} provide bootstrap inference on the boundary of the parameter space with application to conditional volatility models. We return to this issue in our empirical application in Remark \ref{rem:5.1}.
In Assumption \ref{as:4.8} the function $\sigma(x_1,x_2,\dots;\theta)$ is presumed to be monotonically increasing in $\theta$, which is used to establish the strong consistency of the quantile estimator. While the monotonicity condition is a feature shared by various stochastic volatility models (cf.\ \citeauthor{berkes2003limit}, \citeyear{berkes2003limit}, Lemma 4.1), it excludes the exponential GARCH \citep{nelson1991conditional} and the log-GARCH \citep{geweke1986comment,pantula1986modeling}. Further, we require higher order of moments in Assumption \ref{as:4.9} for the bootstrap, which does not seem to be restrictive for the classical GARCH-type models (cf.\ \citeauthor{francq2011garch}, \citeyear{francq2011garch}, p.\ 165; \citeauthor{hamadeh2011asymptotic}, \citeyear{hamadeh2011asymptotic}, p.\ 501). In particular, Assumption \ref{as:4.9} is presumed to hold with $a=\pm 12$, $b=12$ and $c=6$ for establishing the convergence of the bootstrap information matrix.
On the basis of the previous assumptions we extend the strong consistency result of \citeauthor{francq2015risk} (\citeyear{francq2015risk}, Theorem 1) to the quantile estimator.
\begin{theorem}\textit{(Strong Consistency)}\label{thm:4.1}
\begin{itemize}
\item[(i)] \citep{francq2015risk} Under Assumptions \ref{as:4.1}--\ref{as:4.3}, \ref{as:4.4}(\ref{as:4.4.1}) and \ref{as:4.5}(\ref{as:4.5.1}) the estimator in \eqref{eq:4.3.3} is strongly consistent, i.e. $\hat{\theta}_n \overset{a.s.}{\to} \theta_0$.
\item[(ii)] If in addition Assumptions \ref{as:4.6} and \ref{as:4.9}(i) hold with $a=-1$, then the estimator in \eqref{eq:4.3.4} satisfies $\hat{\xi}_{n,\alpha} \overset{a.s.}{\to} \xi_\alpha$.
\end{itemize}
\end{theorem}
To lighten notation, we henceforth write $D_t(\theta) =\frac{1}{\sigma_t(\theta)}\frac{\partial\sigma_t(\theta)}{\partial \theta}$ and drop the argument when evaluated at the true parameter, i.e.\ $D_t=D_t(\theta_0)$. The next result provides the joint asymptotic distribution of $\hat{\theta}_n$ and $\hat{\xi}_{n,\alpha}$ and is due to \cite{francq2015risk}.
\begin{theorem} \label{thm:4.2}
(Asymptotic Distribution, \citealp{francq2015risk}) Suppose Assumptions \ref{as:4.1}--\ref{as:4.7}, \ref{as:4.9} and \ref{as:4.10} hold with $a=b=4$ and $c=2$. Then, we have
\begin{align}
\label{eq:4.3.6}
\begin{pmatrix}
\sqrt{n}(\hat{\theta}_n-\theta_0)\\
\sqrt{n}(\xi_{\alpha} - \hat{\xi}_{n,\alpha})
\end{pmatrix}
\overset{d}{\to}N\big(0, \Sigma_\alpha\big) \qquad \mbox{with}\qquad
\Sigma_\alpha=
\begin{pmatrix}
\frac{\kappa-1}{4}J^{-1} & \lambda_\alpha J^{-1}\Omega\\
\lambda_\alpha \Omega'J^{-1} & \zeta_\alpha
\end{pmatrix},
\end{align}
where $\kappa = \mathbb{E}[\eta_t^4]$, $\Omega = \mathbb{E}[D_t]$, $J=\mathbb{E}[D_tD_t']$, $\lambda_\alpha = \xi_\alpha\frac{\kappa-1}{4}+\frac{p_\alpha}{2f(\xi_{\alpha})}$, $\zeta_\alpha = \xi_\alpha^2\frac{\kappa-1}{4}+\frac{\xi_{\alpha} p_\alpha}{f(\xi_\alpha)}+\frac{\alpha(1-\alpha)}{f^2(\xi_\alpha)}$ and $p_\alpha =\mathbb{E}[\eta_t^2 \mathbbm{1}_{\{\eta_t<\xi_\alpha\}}]-\alpha$.
\end{theorem}
\begin{remark}
It is worth mentioning that the asymptotics in this theorem for $\hat{\xi}_{n,\alpha}$ are for $\alpha$ fixed while $n$ goes to infinity. If, for instance, $\alpha$ is very small for moderate $n$ the distribution in the following theorem might not provide a good approximation. For such cases, approximations based on extreme value theory may provide better approximations to the unknown finite sample distribution. See, for example, \cite{McNeil2002} and the recent \cite{li_peng_song_2022} for an application of GARCH models with extreme value theory for VaR estimation. Note that the latter authors consider simultaneous inference for VaR and expected shortfall.
\end{remark}
In applied work, to use Theorem \ref{thm:4.2} in order to conduct inference for $(\theta_0, \xi_{\alpha})$ one needs a consistent estimator for $\Sigma_{\alpha}$. It follows from Lemma \ref{lem:4.2} that $\kappa$, $\Omega$, and $J$ can be consistently estimated by
\begin{align}
\label{eq:4.3.7}
\begin{split}
\hat{\kappa}_n=\frac{1}{n}\sum_{t=1}^n\hat{\eta}_t^4, \qquad
\hat{\Omega}_n=\frac{1}{n}\sum_{t=1}^n\hat{D}_t, \qquad \mbox{ and } \quad \hat{J}_n=\frac{1}{n}\sum_{t=1}^n\hat{D}_t\hat{D}_t',
\end{split}
\end{align}
respectively, with $\hat{D}_t = \tilde{D}_t(\hat{\theta}_n)$ and $\tilde{D}_t(\theta) = \frac{1}{\tilde{\sigma}_t(\theta)}\frac{\partial\tilde{\sigma}_t(\theta)}{\partial \theta}$.
To estimate $\lambda_{\alpha}$ and $\zeta_{\alpha}$, which also appear in $\Sigma_{\alpha}$, one can use $\hat{\xi}_{n,\alpha}$ for $\xi_{\alpha}$ and
\begin{align}\label{eq consistency matrix own}
\hat{p}_{n,\alpha} = (1 \slash n) \sum_{t=1}^n \hat{\eta}_t^2 \mathbbm{1}_{\{\eta_t < \hat{\xi}_{n,\alpha}\}}-\alpha
\end{align}
for $p_{\alpha}$ (see also Lemma \ref{alg:4.2} for the properties of $\hat{p}_{n,\alpha}$). To estimate the density $f$, which also appears in $\lambda_{\alpha}$ and $\zeta_{\alpha}$, kernel smoothing is commonly employed, i.e.\
\begin{align}\label{eq:4.3.8}
\hat{\mathbbm{f}}_n^S(x) =\frac{1}{nh_n}\sum_{t=1}^n k\bigg(\frac{x-\hat{\eta}_t}{h_n}\bigg)
\end{align}
with kernel function $k$ and bandwidth $h_n>0$. Hence, whenever we use a consistent $\hat{\mathbbm{f}}_n^S$ for $f$ we obtain, combined with Equations \eqref{eq:4.3.7} and \eqref{eq:4.3.8}, a consistent estimator for $\Sigma_\alpha$ denoted by $\hat{\Sigma}_{n,\alpha}$. For the case that $\epsilon_t$ in Equation (\ref{eq:4.2.1}) follows a GARCH$(p,q)$ process, \cite{gao2008estimation} considered estimating $f$ by a kernel density estimator with Lipschitz-continuous kernels such as $k(x)=\phi(x)$, where $\phi$ is the standard normal density function. An alternative estimator is based on the uniform kernel $k(x)=\frac{1}{2}\mathbbm{1}_{\{|x|\leq 1\}}$ yielding $\hat{\mathbbm{f}}_n^S(\hat{\xi}_{n,\alpha})\overset{p}{\to}f(\xi_\alpha)$ whenever $h_n \sim n^{-\varrho}$ for some $\varrho \in (0,1/2]$.
At this point it is worth recalling that interest lies in $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ (cf.~Equation (\ref{eq:4.2.4})). More precisely, one is not so much interested in the distribution of this random variable or its moments but in inference for its possible realizations. Note also that $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ varies with $n$ and does not converge. This illustrates that inference for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ is different from inference for $(\xi_{\alpha},\theta_0)$ which is just a real-valued vector of dimension $r+1$. Nevertheless, it is intuitively clear how to conduct inference for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$, yet this difference results in two technical issues that need to be addressed to provide a thorough theoretical justification of inferential procedures for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$. These technical issues are known for quite some time; see, for instance, \cite{phillips1979sampling}, \citet{kreiss2015discussion} and the textbook \cite{pesaran2015time}[p.~389]. Several examples illustrating these two issues can be found in \citet{beutner2017justification}. In a nutshell the first issue stems from the fact that, as mentioned, $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ varies over time (cf.~Eq.~(\ref{eq:4.2.4})) implying that a limiting distribution cannot exist.
The second issue is a result of $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ being a conditional quantity which on the one hand requires conditioning on the sample observed so far, yet on the other hand this conditioning eliminates all randomness, making it impossible to establish useful distributional results. We now briefly illustrate this here for the case that $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$ is the object of interest; a very detailed illustration of this second issue by means of a GARCH(1,1) can be found in Example 2.1 in \cite{beutner2017justification}. Fixing some arbitrary starting values $\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots$ and making the dependency of $\sigma_{n+1}$ on the realized values of the random variables $\epsilon_n,\epsilon_{n-1},\ldots,\epsilon_1$ and the starting values $\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots$ explicit a feasible version of the quantity of interest is
\begin{equation}\label{eq VaR for sample splitting}
VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)= -\xi_{\alpha}\, \tilde{\sigma}_{n+1}(\epsilon_n^r,\epsilon_{n-1}^r,\ldots,\epsilon_1^r,\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots;\theta_0),
\end{equation}
where $\epsilon_n^r,\epsilon_{n-1}^r,\ldots,\epsilon_{1}^r$ denote the realized values of $\epsilon_n,\epsilon_{n-1},\ldots,\epsilon_1$. Here feasible is in the sense that only parameter uncertainty remains. Now replacing, as at the beginning of Section 3, the unknown $\theta_0$ by the estimator $\hat{\theta}_{n}$ and doing the same with $\xi_{\alpha}$ we see that these estimators must enter Equation \eqref{eq VaR for sample splitting} \textit{given the realizations}, i.e.~as $\hat{\theta}_{n}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$ and $\hat{\xi}_{\alpha,n}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$, respectively, because otherwise we would end up with an inconsistency. Indeed, if they entered as $\hat{\theta}_n(\epsilon_1,\epsilon_{2},\ldots,\epsilon_n)$ and $\hat{\xi}_{n,\alpha}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_n)$, respectively, we would treat $\epsilon_1,\epsilon_{2},\ldots,\epsilon_n$ as random and observed at the same time. Upon replacing $\theta_0$ and $\xi_{\alpha}$ in Equation \eqref{eq VaR for sample splitting} by $\hat{\theta}_{n}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$ and $\hat{\xi}_{n,\alpha}(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)$, respectively, no random variable appears in
\begin{align*}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}= & -\hat{\xi}_{n,\alpha}(\epsilon_1^r,\epsilon_{2}^r,\ldots, \epsilon_n^r) \nonumber \\ & \times \sigma_{n+1}(\epsilon_n^r,\epsilon_{n-1}^r,\ldots,\epsilon_1^r,s_0,s_{-1},\ldots;\hat{\theta}_n(\epsilon_1^r,\epsilon_{2}^r,\ldots,\epsilon_n^r)),
\end{align*}
which is just the estimator of Equation (3.3) with all dependencies made explicit. As just observed no random variable appears in $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}$ which therefore does not have a distribution which could be used to construct confidence intervals for $VaR_{\alpha}(\epsilon_{n+1}|\mathcal{F}_n)$. In the article \citet{beutner2017justification} merging is proposed as a solution to overcome the first issue and sample splitting as a means to solve the second. The results of this article can be used justify, through their Theorem 1, approximating the distribution of the VaR estimator, centered at $VaR_{n,\alpha}$ and inflated by $\sqrt{n}$, by
\begin{align}
\label{eq:4.3.9}
N\left(0, \begin{pmatrix}
-\xi_\alpha \frac{\partial \sigma_{n+1}(\theta_0)}{\partial \theta}\\
\sigma_{n+1}
\end{pmatrix}' \Sigma_\alpha \begin{pmatrix}
-\xi_\alpha \frac{\partial \sigma_{n+1}(\theta_0)}{\partial \theta}\\
\sigma_{n+1}
\end{pmatrix}\right),
\end{align}
and consequently to justify the intuitive (conditional) confidence interval
\begin{align}\label{eq:4.3.10}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}\pm \frac{\Phi^{-1}(\gamma/2)}{\sqrt{n}}
\left\{\begin{pmatrix}
-\hat{\xi}_{n,\alpha} \frac{\partial \tilde{\sigma}_{n+1}(\hat{\theta}_n)}{\partial \theta}\\
\tilde{\sigma}_{n+1}(\hat{\theta}_n)
\end{pmatrix}' \hat{\Sigma}_{n,\alpha} \begin{pmatrix}
-\hat{\xi}_{n,\alpha} \frac{\partial \tilde{\sigma}_{n+1}(\hat{\theta}_n)}{\partial \theta}\\
\tilde{\sigma}_{n+1}(\hat{\theta}_n)
\end{pmatrix}\right\}^{1/2}
\end{align}
for ${VaR}_{n,\alpha}$. Here $\Phi$ denotes the standard normal cdf, and we used again the short-hand notation for $\tilde{\sigma}_{n+1}(\hat{\theta})$ and $\hat{\xi}_{n,\alpha}$, i.e.~did not make the dependency on the starting values etc.~explicit. Note that the interval \eqref{eq:4.3.10} lacks a proper theoretical justification as it treats the non-random $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}$ as random because it assigns an asymptotic distribution to it. In order to use Theorem 3 and Corollary 2 of \citet{beutner2017justification} to provide a sound theoretical justification of \eqref{eq:4.3.10} let $n_E: \mathbb{N} \to \mathbb{N}$ and $n_P: \mathbb{N} \to \mathbb{N}$ be such that for all $n \in \mathbb{N}$ we have $n_E(n) < n_P(n)$. First, in the sample split approach we will only use $\epsilon_1,\ldots,\epsilon_{n_E}$ to estimate $\xi_{\alpha}$ and $\theta_0$ which we denote by $\hat{\xi}_{n_E,\alpha}^{SPL}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_{n_E})$ and $\hat{\theta}_{n_E}^{SPL}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_{n_E})$, respectively. To lighten the notation a bit we will also use the shorter notations
$\hat{\xi}_{n_E,\alpha}^{SPL}$ and $\hat{\theta}_{n_E}^{SPL}$, respectively. Second, in the sample split approach conditioning on $\mathcal{F}_n$ on the left-hand side of \eqref{eq VaR for sample splitting} is replaced by conditioning on $\epsilon_{n_P},\ldots,\epsilon_n$ only. The sample split version of (3.3) which we denote by $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{SPL}$ becomes
\begin{align*}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{SPL}=
& -\hat{\xi}_{n_E,\alpha}^{SPL}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_{n_E}) \tilde{\sigma}_{n+1}^{SPL}(\hat{\theta}^{SPL}_{n_E})
\nonumber \\
= & -\hat{\xi}_{n_E,\alpha}^{SPL}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_{n_E}) \nonumber \\
& \times \tilde{\sigma}_{n+1}(\epsilon_n^r,\epsilon_{n-1}^r,\ldots,\epsilon_{n_P}^r,c_{n_P-1},\ldots,c_{1},\tilde{\epsilon}_{0},\tilde{\epsilon}_{-1}, \ldots;\hat{\theta}_{n_E}^{SPL}(\epsilon_1,\epsilon_{2},\ldots,\epsilon_{n_E})),
\end{align*}
where $\epsilon_n^r,\epsilon_{n-1}^r,\ldots,\epsilon_{n_P}^r$ and $\tilde{\epsilon}_{0},\tilde{\epsilon}_{-1},\ldots$ are as before and $c_{n_P-1},\ldots,c_{1}$ are constants. These constants could be viewed as starting values but since they are not replacing unobserved value as $\tilde{\epsilon}_0, \tilde{\epsilon}_{-1},\ldots$ do we prefer to use a different letter for them. Note
also that $\epsilon_1,\ldots,\epsilon_{n_E}$ enter the sample split version $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{SPL}$ as random variables which is in contrast to $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}$. Therefore, $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{SPL}$ does have a non-degenerate distribution which can be used to construct confidence intervals. The sample split (conditional) confidence interval is defined as
\begin{align}\label{eq var spl interval}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{SPL}\pm \frac{\Phi^{-1}(\gamma/2)}{\sqrt{n_E}}
\left\{\begin{pmatrix}
-\hat{\xi}_{n_E,\alpha}^{SPL} \frac{\partial \tilde{\sigma}_{n+1}^{SPL}(\hat{\theta}_{n_E}^{SPL})}{\partial \theta}\\
\tilde{\sigma}_{n+1}^{SPL}(\hat{\theta}_{n_E}^{SPL})
\end{pmatrix}' \hat{\Sigma}_{n_E,\alpha}^{SPL} \begin{pmatrix}
-\hat{\xi}_{n_E,\alpha}^{SPL} \frac{\partial \tilde{\sigma}_{n+1}^{SPL}(\hat{\theta}_{n_E}^{SPL})}{\partial \theta}\\
\tilde{\sigma}_{n+1}^{SPL}(\hat{\theta}_{n_E}^{SPL})
\end{pmatrix}\right\}^{1/2}.
\end{align}
Here $\hat{\Sigma}_{n_E,\alpha}^{SPL}$ is as $\hat{\Sigma}_{n,\alpha}$ but based on the first $n_E$ observations only. Note that this interval is meaningful because $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{SPL}$ is random and converges in distribution after centering and scaling. Intuitively, we would say that the statistically meaningful interval in \eqref{eq var spl interval} provides a theoretical justification for the interval \eqref{eq:4.3.10} if
\begin{equation}\label{eq convergence var's}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{SPL} \mbox{ converges to }
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}
\end{equation}
and if
\begin{equation}\label{eq convergence matrices I}
\left\{\begin{pmatrix}
-\hat{\xi}_{n_E,\alpha}^{SPL} \frac{\partial \tilde{\sigma}_{n+1}^{SPL}(\hat{\theta}_{n_E}^{SPL})}{\partial \theta}\\
\tilde{\sigma}_{n+1}^{SPL}(\hat{\theta}_{n_E}^{SPL})
\end{pmatrix}' \hat{\Sigma}_{n_E,\alpha}^{SPL} \begin{pmatrix}
-\hat{\xi}_{n_E,\alpha}^{SPL} \frac{\partial \tilde{\sigma}_{n+1}^{SPL}(\hat{\theta}_{n_E}^{SPL})}{\partial \theta}\\
\tilde{\sigma}_{n+1}^{SPL}(\hat{\theta}_{n_E}^{SPL})
\end{pmatrix}\right\}^{1/2}
\end{equation}
converges to
\begin{equation}\label{eq convergence matrices II}
\left\{\begin{pmatrix}
-\hat{\xi}_{n,\alpha} \frac{\partial \tilde{\sigma}_{n+1}(\hat{\theta}_n)}{\partial \theta}\\
\tilde{\sigma}_{n+1}(\hat{\theta}_n)
\end{pmatrix}' \hat{\Sigma}_{n,\alpha} \begin{pmatrix}
-\hat{\xi}_{n,\alpha} \frac{\partial \tilde{\sigma}_{n+1}(\hat{\theta}_n)}{\partial \theta}\\
\tilde{\sigma}_{n+1}(\hat{\theta}_n)
\end{pmatrix}\right\}^{1/2}.
\end{equation}
This is also the concept employed in \citet{beutner2017justification}. Under the conditions of Theorem 3 of this article the convergence in \eqref{eq convergence var's} holds and \eqref{eq convergence matrices I} converges to
\eqref{eq convergence matrices II} by Corollary 2 of this article. It is worth pointing out that convergence is in probability and that for this concept to be applicable $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}$ and \eqref{eq convergence matrices II} are treated as random. Theorem 2 above ensures that one of the assumptions of Corollary 3 of \citet{beutner2017justification} holds so that it provides the basis for a sound theoretical justification of \eqref{eq:4.3.10} as a (conditional) confidence interval. Formally, we can state
\begin{corollary}\label{corollary supp} Assume that the assumptions of Theorem \ref{thm:4.1} are fulfilled. Define the function $\psi: \mathbb{R}^{\infty} \times \Upsilon$ with $\Upsilon=\mathbb{R} \times \Theta$ by
$$\psi(x_1,x_2,\ldots;\upsilon)=\psi(x_1,x_2,\ldots;(\xi,\theta))=-\xi \sigma(x_1,x_2,\ldots;\theta),$$
and put $\upsilon_0=(\xi_{\alpha}, \theta_0)$. Assume further that
\begin{enumerate}
\item $\Big|\Big|\frac{\partial \sigma (\epsilon_n, \epsilon_{n-1}, \ldots ;\theta_0)}{\partial \theta}\Big|\Big|=O_{p}(1)$;
\item $\sup_{\upsilon \in \mathscr{V}(\upsilon_0)}\Big|\Big|\frac{\partial^2 \psi (\epsilon_n, \epsilon_{n-1}, \ldots ;\upsilon)}{\partial \upsilon \partial \upsilon'}\Big|\Big|=O_{p}(1)$ for some open neighborhood $\mathscr{V}(\upsilon_0)$ around $\upsilon_0$;
\item Given sequences $\{\tilde{\epsilon}_t\}$ and $\{c_t\}$, we have
\begin{align*}
& \sqrt{n}\big(\psi(\epsilon_n,\ldots,\epsilon_{t_1},c_{t_1-1},\ldots,c_1,\tilde{\epsilon}_0,\ldots; \upsilon_0) - \psi (\epsilon_n, \epsilon_{n-1}, \ldots; \upsilon_0)\big)=o_{p}(1),\\
& \bigg|\bigg|\frac{\partial \psi(\epsilon_n,\ldots,\epsilon_{t_1},c_{t_1-1},\ldots,c_1,\tilde{\epsilon}_0,\ldots; \upsilon_0) }{\partial \upsilon} - \frac{\partial \psi (\epsilon_n, \epsilon_{n-1}, \ldots; \upsilon_0)}{\partial \upsilon}\bigg|\bigg|=o_{p}(1),\\
& \sup_{\upsilon \in \mathscr{V}(\upsilon_0)} \bigg|\bigg|\frac{\partial^2 \psi(\epsilon_n,\ldots,\epsilon_{t_1},c_{t_1-1},\ldots,c_1,\tilde{\epsilon}_0,\ldots; \upsilon_0)}{\partial \upsilon \partial \upsilon'} - \frac{\partial^2 \psi (\epsilon_n, \epsilon_{n-1}, \ldots; \upsilon_0)}{\partial \upsilon \partial \upsilon'}\bigg|\bigg| \\
& =o_{p}(1)
\end{align*}
for any $t_1 \geq 1$ such that $(n-t_1) / l_n \rightarrow \infty$ as $n \to \infty$ and for some model-specific $l_n$ with $l_n \rightarrow \infty$.
\end{enumerate}
Moreover, let $n_E$ and $n_P$ fulfill Assumption 3.a of \citet{beutner2017justification} and $\{\epsilon_t\}$ Assumption 3.c of the same article. Then the difference between $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{SPL}$ and $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}$ converges to zero in probability and the same holds for the difference between \eqref{eq convergence matrices I} and \eqref{eq convergence matrices II}.
\end{corollary}
\begin{proof}
The claims made are proved if the assumptions of Theorem 3 and Corollary 2 in \citet{beutner2017justification} are met. Assumption 1.a of \citet{beutner2017justification} holds by Theorem \ref{thm:4.2}. Note that the scaling sequence $m_T$ in \cite{beutner2017justification} equals $\sqrt{n}$ here. Moreover according to Theorem \ref{thm:4.2} $\hat{\upsilon}_n=(\hat{\xi}_{n,\alpha},\hat{\theta}_n)$ scaled by $\sqrt{n}$ and centered at $\upsilon_0$ converges to a multivariate normal distribution so that Assumption 5 of \citet{beutner2017justification} also holds. We now turn to Assumption 1.b of \citet{beutner2017justification}. Clearly $\upsilon \to \psi(\cdot;\upsilon$) is continuous because by Assumption 3 above $\theta \to \sigma(\cdot,;\theta)$ is continuous. Furthermore, because $\upsilon \to \psi(\cdot;\upsilon)$ is the product of the functions $\xi \to \xi$ and $\theta \to \sigma(\cdot;\theta)$ which are both twice differentiable (for $\theta \to \sigma(\cdot;\theta)$ this holds by Assumption 4 (ii) above) the same holds for $\upsilon \to \psi(\cdot;\upsilon)$. Therefore, Assumption 3.b of \citet{beutner2017justification} holds. The gradient of $\psi(\epsilon_n,\epsilon_{n-1},\ldots;\upsilon_0)$ is given by
$$
\left(-\sigma(\epsilon_n,\epsilon_{n-1},\ldots;\theta_0), -\xi_{\alpha}\frac{\partial \sigma (\epsilon_n, \epsilon_{n-1}, \ldots ;\theta_0)}{\partial \theta}\right).
$$
By Assumption 3 above and Markov's inequality $\sigma(\epsilon_n,\epsilon_{n-1},\ldots;\theta_0)$ is bounded in probability. Hence, together with Assumption 1.~in the statement of the corollary it follows that Assumption 1.c in \citet{beutner2017justification} holds.
Assumption 1.d and 1.e in \citet{beutner2017justification} are just the Assumptions 2.~and 3.~of the corollary. Obviously, Assumptions 3.a and 3.c of \cite{beutner2017justification} hold under the assumptions of the corollary. Assumption 3.b of this article is met by Assumption 2 above which is nothing else than Assumption 3.b of \citet{beutner2017justification} for Value-at-Risk. As explained in the proof of Theorem 3 in \citet{beutner2017justification} the Assumption 2.b of that article is irrelevant for the quantity on the right-hand side of \eqref{eq convergence var's} and for \eqref{eq convergence matrices II}. Assumption 2.a of that article is obviously met. This finishes the proof.
\end{proof}
\begin{remark}
Like some of the Assumptions 1-10 will hold or not depending on the specification of $\sigma$, the same is true for the assumptions of Corollary \ref{corollary supp} that ensure applicability of the results of \citet{beutner2017justification}. For popular time series models like a GARCH(1,1) they have been verified in \citet{beutner2019technical}.
\end{remark}
Although the interval in \eqref{eq:4.3.10} is, as just outlined, well-justified, it may perform poorly since the density estimation appears rather sensitive regarding the choice of bandwidth (see \citeauthor{gao2008estimation}, \citeyear{gao2008estimation}, Section 4). Bootstrap methods offer an alternative way to quantify the uncertainty around the estimators.
\section{Bootstrap}
\label{sec:4.4}
Bootstrap approximations frequently provide better insight into the actual distribution than the asymptotic approximation, yet they require a careful set-up. \cite{hall2003inference} show that conventional bootstrap methods are inconsistent in a GARCH model lacking finite fourth moment in the case of the squared innovations' distribution not being in the domain of attraction of the normal distribution. They consider a subsample bootstrap instead and study its asymptotic properties. In correspondence, an $m$-out-of-$n$ without-replacement bootstrap is proposed by \cite{spierdijk2016confidence} to construct confidence intervals for ARMA-GARCH VaR.
\cite{pascual2006bootstrap} present a residual bootstrap in a GARCH($1,1$) setting and assess its finite sample properties by means of simulation. Their bootstrap scheme follows a recursive design in which the bootstrap observations are generated iteratively using the estimated volatility dynamics. Building upon their results, \cite{christoffersen2005estimation} construct bootstrap confidence intervals for (conditional) VaR and Expected Shortfall and compare them to competitive methods within the GARCH($1,1$) model. Theoretical results on the recursive-design residual bootstrap are provided by \cite{hidalgo2007goodness} and \cite{jeong2017residual} for the ARCH($\infty$) and GARCH($p,q$) model, respectively.
In contrast, \cite{shimizu2009bootstrapping} considers fixed-design variants of the wild and the residual bootstrap in which the ARMA-GARCH dynamics of the bootstrap samples are kept fixed at the values of the original series. The bootstrap estimators are based on a single Newton-Raphson iteration simplifying the proofs of first-order asymptotic validity. \citeauthor{shimizu2009bootstrapping}'s approach for the residual bootstrap is also employed in a multivariate GARCH setting by \cite{francq2016variance}. Recently, \cite{cavaliere2018fixed} study the fixed-design residual bootstrap in the context of ARCH($q$) models and propose a bootstrap Wald statistic based on a QML bootstrap estimator. While their theory has been developed independently to ours, their simulation study indicates that the fixed-design bootstrap performs as well as the recursive-design bootstrap.
\subsection{Fixed-design Residual Bootstrap}
\label{sec:4.4.1}
We propose a fixed-design residual bootstrap procedure, described in Algorithm \ref{alg:4.1}, to approximate the distribution of the estimators in \eqref{eq:4.3.3} -- \eqref{eq:4.3.5}.
\begin{algorithm}\textit{(Fixed-design residual bootstrap)}
\label{alg:4.1}
\begin{enumerate}
\item For $t=1,\dots, n$, generate $\eta_t^* \overset{iid}{\sim} \hat{\mathbbm{F}}_n$ and the bootstrap observation $\epsilon_t^* = \tilde{\sigma}_t(\hat{\theta}_n) \eta_t^*$.
\item Calculate the bootstrap estimator
\begin{align}
\label{eq:4.4.1}
\hat{\theta}_n^* = \arg \max_{\theta \in \Theta} \frac{1}{n}\sum_{t=1}^n \ell_t^*(\theta) \qquad \text{with} \qquad \ell_t^*(\theta)=-\frac{1}{2}\bigg(\frac{\epsilon_t^{*}}{\tilde{\sigma}_t(\theta)}\bigg)^2-\log \tilde{\sigma}_t(\theta).
\end{align}
\item For $t=1,\dots,n$ compute the bootstrap residual $\hat{\eta}_t^* = \epsilon_t^*/\tilde{\sigma}_t(\hat{\theta}_n^*)$
and obtain
\begin{align}
\label{eq:4.4.2}
\hat{\xi}_{n,\alpha}^* = \arg\min_{z \in \mathbb{R}} \frac{1}{n}\sum_{t=1}^n\rho_\alpha(\hat{\eta}_t^*-z).
\end{align}
\item Obtain the bootstrap estimator of the conditional VaR
\begin{align}
\label{eq:4.4.3}
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{*}=-\hat{\xi}_{n,\alpha}^{*}\: \tilde{\sigma}_{n+1}\big(\hat{\theta}_n^{*}\big).
\end{align}
\end{enumerate}
\end{algorithm}
\begin{remark}
\label{rem:4.1}
In contrast to the literature, the bootstrap errors are drawn with replacement from the residuals rather than the standardized residuals. In fact, re-centering would be inappropriate in the case of $\mathbb{E}[\eta_t]\neq 0$.
In addition, re-scaling of the residuals is typically redundant as $\frac{1}{n}\sum_{t=1}^n \hat{\eta}_t^2=1$
is implied by $\hat{\theta}_n \in \mathring{\Theta}$
under Assumption \ref{as:4.10}; see \citeauthor{francq2011garch}, \citeyear{francq2011garch}, p.\ 182/406 and note that the solution requires $\hat{\theta}_n$ belonging to the interior (\citeauthor{francq2011garch}, Oct.\ 2018, personal communication).
\end{remark}
\begin{remark}
\label{rem:4.2}
The term `fixed-design' refers to the fact that the bootstrap observations are generated using $\tilde{\sigma}_t(\hat{\theta}_n)=\sigma(\epsilon_{t-1},\dots,\epsilon_1,\tilde{\epsilon}_0,\tilde{\epsilon}_{-1},\dots;\hat{\theta}_n)$. In contrast, a recursive-design scheme replicates the model's dynamic structure, i.e.\ $\epsilon_t^\star = \sigma_t^\star \eta_t^\star$ with $\sigma_t^\star = \sigma(\epsilon_{t-1}^\star,\dots,\epsilon_1^\star,\tilde{\epsilon}_0,\tilde{\epsilon}_{-1},\dots;\hat{\theta}_n)$ and $\eta_t^\star \overset{iid}{\sim} \hat{\mathbbm{F}}_n$, which is computationally more demanding. We refer to Appendix \ref{app:4.B} for a complete description. See also \cite{cavaliere2018fixed} for more theoretical insights on the difference in the design in an ARCH($q$).
\end{remark}
\begin{remark}
\label{rem:4.3}
Whereas \eqref{eq:4.4.1} involves a nonlinear optimization, \cite{shimizu2009bootstrapping} proposes a Newton-Raphson type bootstrap estimator instead. The Newton-Raphson bootstrap estimator corresponding to \eqref{eq:4.4.1} is given by
\begin{align*}
\hat{\theta}_n^{*NR} = \hat{\theta}_n+\hat{J}_n^{-1}\frac{1}{2n}\sum_{t=1}^n \hat{D}_t \big(\eta_t^{*2}-1\big),
\end{align*}
which can considerably speed up computations.
\end{remark}
Proposition \ref{prop:4.1} establishes the asymptotic validity of the bootstrap for the volatility parameters.
\begin{proposition}
\label{prop:4.1}
Suppose Assumptions \ref{as:4.1}--\ref{as:4.4}, \ref{as:4.5}(\ref{as:4.5.1}), \ref{as:4.5}(\ref{as:4.5.3}), \ref{as:4.6}, \ref{as:4.7}, \ref{as:4.9} and \ref{as:4.10} hold with $a=\pm 12$, $b=12$ and $c=6$. Then, we have
\begin{align*}
\sqrt{n}\big(\hat{\theta}_n^* - \hat{\theta}_n\big)
\overset{d^*}{\to}N\bigg(0,\frac{\kappa-1}{4}J^{-1}\bigg)
\end{align*}
almost surely.
\end{proposition}
Establishing the asymptotic validity of the bootstrap for the second part appears challenging since the bootstrap innovations are drawn from the discrete distribution $\hat{\mathbbm{F}}_n$. To overcome this issue we rely on arguments employed by \cite{bahadur1966note} and \cite{berkes2003limit}. The following theorem states the paper's main result.
\begin{theorem}\textit{(Bootstrap consistency)}
\label{thm:4.3}
Suppose Assumptions \ref{as:4.1}--\ref{as:4.10} hold with $a=\pm 12$, $b=12$ and $c=6$. Then, we have
\begin{align*}
\begin{pmatrix}
\sqrt{n}(\hat{\theta}_n^*-\hat{\theta}_n)\\
\sqrt{n}(\hat{\xi}_{n,\alpha} - \hat{\xi}_{n,\alpha}^*)
\end{pmatrix}
\overset{d^*}{\to}N\big(0, \Sigma_\alpha\big)
\end{align*}
in probability.
\end{theorem}
Theorem \ref{thm:4.3} is useful to validate the bootstrap for the conditional VaR estimator. For the asymptotic behavior of the conditional VaR estimator we refer to \eqref{eq:4.3.9} and the text around it. The following corollary is established.
\begin{corollary}
\label{cor:4.1}
Under the assumptions of Theorem \ref{thm:4.3} the conditional distribution of $\sqrt{n}\big(
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{*}-
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}\big)$ given $\mathcal{F}_n$ and \eqref{eq:4.3.9} given $\mathcal{F}_n$ merge in probability.
\end{corollary}
\subsection{Bootstrap Confidence Intervals for VaR}
\label{sec:4.4.3}
Clearly, the VaR evaluation in \eqref{eq:4.3.5} is subject to estimation risk that needs to be quantified. We propose the following algorithm to obtain approximately $100(1-\gamma)\%$ confidence intervals.
\begin{algorithm}{\textit{(Fixed-design Bootstrap Confidence Intervals for VaR)}}
\label{alg:4.2}
\begin{enumerate}
\item Acquire a set of $B$ bootstrap replicates, i.e. $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{*(b)}$ for $b=1,\dots,B$, by repeating Algorithm \ref{alg:4.1}.
\item[2.1.] Obtain the \textit{equal-tailed percentile} (EP) interval
\begin{align}
\label{eq:4.4.8}
\bigg[
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}-\frac{1}{\sqrt{n}}\hat{G}_{n,B}^{*-1}(1-\gamma/2),\:
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}-\frac{1}{\sqrt{n}}\hat{G}_{n,B}^{*-1}(\gamma/2)\bigg]
\end{align}
with $\hat{G}_{n,B}^{*-1}(\cdot)$ being the quantile function (generalized inverse) of $\hat{G}_{n,B}^*(x)=\frac{1}{B}\sum_{b=1}^B \mathbbm{1}_{\big\{\sqrt{n}\big(
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{*(b)}-
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}\big)\leq x\big\}}$.
\item[2.2.] Calculate the \textit{reversed-tails} (RT) interval
\begin{align}
\label{eq:4.4.9}
\bigg[
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}+\frac{1}{\sqrt{n}}\hat{G}_{n,B}^{*-1}(\gamma/2),
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}+\frac{1}{\sqrt{n}}\hat{G}_{n,B}^{*-1}(1-\gamma/2)\bigg].
\end{align}
\item[2.3.] Compute the \textit{symmetric} (SY) interval
\begin{align}
\label{eq:4.4.10}
\bigg[
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}-\frac{1}{\sqrt{n}}\hat{H}_{n,B}^{*-1}(1-\gamma),\:
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}+\frac{1}{\sqrt{n}}\hat{H}_{n,B}^{*-1}(1-\gamma)\bigg]
\end{align}
with $\hat{H}_{n,B}^{*-1}(\cdot)$ being the quantile function (generalized inverse) of $\hat{H}_{n,B}^*(x)=\frac{1}{B}\sum_{b=1}^B \mathbbm{1}_{\big\{\sqrt{n}\big|
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{*(b)}-
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}\big|\leq x\big\}}$.
\end{enumerate}
\end{algorithm}
The interval in \eqref{eq:4.4.8} is obtained by the EP method, that is frequently encountered in the bootstrap literature. It is obtained from the (typically) infeasible equal-tailed confidence interval
\begin{align*}
\bigg[
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}-\frac{1}{\sqrt{n}}G_{n}^{-1}(1-\gamma/2),\:
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}-\frac{1}{\sqrt{n}}G_{n}^{-1}(\gamma/2)\bigg],
\end{align*}
where $G_{n}^{-1}$ is the (unknown) quantile function of $\sqrt{n}(
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}-{VaR}_{n,\alpha})$, which is replaced by its bootstrap analogue $\hat{G}_{n,B}^{*-1}$. The same reasoning leads to the SY interval but with test statistic $\sqrt{n}|(
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}-{VaR}_{n,\alpha})|$ instead of $\sqrt{n}(
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}-{VaR}_{n,\alpha})$ which makes it also clear that the interval in \eqref{eq:4.4.10} presumes symmetry for rationalizing its construction.
``Flipping around" the tails of the ET interval leads to the RT interval given in \eqref{eq:4.4.9}. Clearly, the RT and the EP have equal length. Whereas \eqref{eq:4.4.9} in its current form emphasizes the interval's name, RT type intervals are frequently reported in their reduced form, i.e.~the lower and upper bound of \eqref{eq:4.4.9} simplify to the $\gamma/2$ and $1-\gamma/2$ quantiles of $\frac{1}{B}\sum_{b=1}^B \mathbbm{1}_{\big\{
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}^{*(b)}\leq x\big\}}$, respectively. RT intervals can either be motivated by the results of \cite{falk1991coverage}\footnote{In a random sample setting \cite{falk1991coverage} prove that the RT bootstrap interval for quantiles has asymptotically greater coverage than the corresponding EP bootstrap interval. For additional insights we refer to \cite{hall1988bootstrap}.} or as the bootstrap analogue of
the (uncentered) statistic $
\savestack{\tmpbox}{\stretchto{
\scaleto{
\scalerel*[\widthof{\ensuremath{VaR}}]{\kern-.6pt\bigwedge\kern-.6pt}
{\rule[-\textheight/2]{1ex}{\textheight}}
}{\textheight}
}{0.5ex}}
\stackon[1pt]{VaR}{\tmpbox}
_{n,\alpha}$.
It is worth mentioning that RT type bootstrap intervals for the VaR are also constructed in reduced form by \cite{christoffersen2005estimation}. Regardless of whether we use an EP, RT or SY interval the meaning is always the same: Given the past up to and including time $n$ the probability that the conditional VaR for period $n+1$ is contained in the intervals is approximately equal to $100(1-\gamma)\%$.
\subsection{Bootstrap Extensions}
\label{sec:4.4.4}
The asymptotic normality result in Theorem \ref{thm:4.2} as well as the bootstrap consistency in Theorem \ref{thm:4.3} are derived, inter alia, under the assumption that the innovations are iid. In case this is not believed to be true -- e.g.~if the suggested specification tests mentioned in Section \ref{sec:4.3} indicate otherwise -- asymptotic normality of $\sqrt{n}(\hat{\theta}_n-\theta_0)$ can still be established under regularity assumptions. \cite{Escanciano2009} studies the QML estimator under some dependence among the $\eta_t$'s while imposing slightly stronger (moment) conditions, whereas the related paper of \cite{Linton2010} investigates estimators in a GARCH(1,1) with dependent errors but under weaker moment conditions. A multivariate version of the dependence condition in \cite{Escanciano2009} can be found in \cite{FrancqZakoian2016}.
The bootstrap method presented in Algorithm \ref{alg:4.1} is contingent on the iid assumption. In general, one cannot expect bootstrap procedures that are based on an iid assumption to be robust against deviations from this assumption; for a well-known example, one may refer to \cite{GONCALVES2004}. Alternative bootstrap techniques may be used if the iid condition is thought to be unrealistic. A variety of bootstrap methods exist that can capture dependence and non-identical random variables; see e.g.~\cite{Lahiri03} for a broad overview. The wild or multiplier bootstrap \citep{Mammen93, DavidsonFlachaire08} is particularly suited for dealing with non-identical variables, but does not capture dependence, unless it is properly modified \citep{Shao10, FSU20}. However, it remains an open question which bootstrap method combined with the fixed design approach leads to valid bootstrap procedures.
\section{Numerical Illustration}
\label{sec:4.5}
\subsection{Monte Carlo Experiment}
\label{sec:4.5.1}
In order to evaluate the finite sample performance of the proposed bootstrap procedure a Monte Carlo experiment is conducted. We confine ourselves to four conditional volatility specifications related to Examples \ref{ex:4.1} and \ref{ex:4.2} in Section \ref{sec:4.2}.
The first two are GARCH($1,1$) parameterizations with
\begin{enumerate}[(i)]
\item high persistence: $\theta_0 = (\omega_0,\alpha_0,\beta_0)’= \big(0.05\times 20^2/252,0.15,0.8\big)'$;
\item low persistence: $\theta_0 = (\omega_0,\alpha_0,\beta_0)’ = \big(0.05\times 20^2/252,0.4,0.55\big)'$,
\end{enumerate}
which are similar to the specifications of \citeauthor{gao2008estimation} (\citeyear{gao2008estimation}, Section 4) and \citeauthor{spierdijk2016confidence} (\citeyear{spierdijk2016confidence}, Section 4.2). In addition, we study two T-GARCH($1,1$) scenarios likewise associated with high and low persistence:
\begin{enumerate}[(i)]
\setcounter{enumi}{2}
\item high persistence: $\theta_0 = (\omega_0,\alpha_0^+, \alpha_0^-, \beta_0)’ = \big(0.05\times 20/\sqrt{252},0.05,0.10,0.8\big)'$;
\item low persistence: $\theta_0 = (\omega_0,\alpha_0^+, \alpha_0^-, \beta_0)’ = \big(0.05\times 20/\sqrt{252},0.1,0.3,0.55\big)'$.
\end{enumerate}
Within the experiment the VaR level takes value $\alpha =0.05$ and there are two possible innovation distributions: a Student-$t$ distribution with $6$ degrees of freedom (df) and the standard normal distribution.\footnote{The Student-t innovations are appropriately standardized to satisfy $\mathbb{E}\eta_t^2=1$.} We consider four estimation sample sizes, $n \in \{250; 500; 1{,}000; 5{,}000\}$, whereas the number of bootstrap replicates is fixed and equal to $B=999$. For each model version we simulate $S=10{,}000$ independent Monte Carlo trajectories.
The numerical optimization of the log-likelihood function is carried out employing the build-in function \textit{fmincon} and running time is reduced by parallel computing using \textit{parfor}. The code is available on the website of the third author, and simulation results for a VaR level of $\alpha=0.01$ can be found in the working paper version.
\captionsetup[table]{labelformat=simple, labelsep=space}
\bgroup
\begin{table}[tbp]
\begin{center}
\caption{\textbf{Fixed-design} bootstrap confidence intervals and asymptotic interval for \textbf{GARCH($1$,$1$)} with Student-t innovations}
\label{tab:4.1}
\end{center}
\centering
\resizebox{\textwidth}{!}{\begin{tabular}{rc:ccc:ccc}
\hline \hline
\multicolumn{1}{c}{\begin{tabular}[c]{@{}c@{}}Sample\\ Size\end{tabular}} & & \begin{tabular}[c]{@{}c@{}}Average \\ coverage\end{tabular} & \begin{tabular}[c]{@{}c@{}}Av. coverage\\ below/above\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ length\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ coverage\end{tabular} & \begin{tabular}[c]{@{}c@{}}Av. coverage\\ below/above\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ length\end{tabular} \\ \hline
\multicolumn{1}{c}{} & & & & & & & \\[-9pt]
\multicolumn{1}{c}{} & & \multicolumn{6}{c}{} \\[-8pt]
\multicolumn{1}{c}{} & & \multicolumn{3}{c}{low persistence} & \multicolumn{3}{c}{high persistence} \\
250 & EP & 81.2 & 6.81/11.99 & 0.595 & 79.57 & 7.28/13.15 & 0.797 \\
& RT & 90.73 & 2.98/6.29 & 0.595 & 90.22 & 3.28/6.50 & 0.797 \\
& SY & 88.87 & 3.38/7.75 & 0.620 & 88.28 & 3.48/8.24 & 0.829 \\
& AS & 86.35 & 3.37/10.28 & 0.617 & 85.85 & 3.93/10.22 & 0.803 \\
\hdashline
500 & EP & 84.15 & 6.13/9.72 & 0.431 & 84.07 & 6.20/9.73 & 0.581 \\
& RT & 90.65 & 3.83/5.52 & 0.431 & 90.82 & 3.86/5.32 & 0.581 \\
& SY & 89.71 & 3.84/6.45 & 0.443 & 89.52 & 3.98/6.50 & 0.595 \\
& AS & 88.17 & 3.80/8.03 & 0.445 & 87.4 & 4.34/8.26 & 0.573 \\
\hdashline
1,000 & EP & 85.53 & 5.88/8.59 & 0.306 & 85.73 & 5.64/8.63 & 0.419 \\
& RT & 90.56 & 4.10/5.34 & 0.306 & 90.41 & 4.07/5.52 & 0.419 \\
& SY & 89.68 & 4.10/6.22 & 0.311 & 89.72 & 4.08/6.20 & 0.425 \\
& AS & 88.57 & 4.11/7.32 & 0.314 & 88.09 & 4.43/7.48 & 0.412 \\
\hdashline
5,000 & EP &87.59 & 5.72/6.69 & 0.144 & 87.97 & 5.35/6.68 & 0.191 \\
& RT & 90.43 & 4.58/4.99 & 0.144 & 89.77 & 4.82/5.41 & 0.191 \\
& SY & 89.69 & 4.74/5.57 & 0.145 & 89.69 & 4.56/5.75 & 0.192 \\
& AS & 89.62 & 4.57/5.81 & 0.146 & 88.86 & 4.84/6.30 & 0.188 \\
\hline \hline
\end{tabular}
}
\vspace{0.15cm}
\caption*{Table \ref{tab:4.1} reports distinct features of the \textbf{fixed-design} bootstrap confidence intervals and the asymptotic interval for the conditional VaR at \textbf{level} $\bm{\alpha=0.05}$ with \textbf{nominal coverage} $\bm{1-\gamma=90\%}$. For each interval type and different sample sizes ($n$), the interval's average coverage rates (in $\%$), the average rate of the conditional VaR being below/above the interval (in $\%$) and the interval's average length are tabulated. The bootstrap intervals are based on $B=999$ bootstrap replications and the averages are computed using $S=10{,}000$ simulations. The results are for the low (left part) and high (right part) persistence parametrization of a \textbf{GARCH(1,1)} with (normalized) Student-t innovations ($6$ df).
}
\end{table}
\egroup
\captionsetup[table]{labelformat=simple, labelsep=space}
\bgroup
\begin{table}[tbp]
\begin{center}
\caption{\textbf{Fixed-design} bootstrap confidence intervals and asymptotic interval for \textbf{T-GARCH($1$,$1$)} with Student-t innovations}
\label{tab:4.1a}
\end{center}
\centering
\resizebox{\textwidth}{!}{\begin{tabular}{rc:ccc:ccc}
\hline \hline
\multicolumn{1}{c}{\begin{tabular}[c]{@{}c@{}}Sample\\ Size\end{tabular}} & & \begin{tabular}[c]{@{}c@{}}Average \\ coverage\end{tabular} & \begin{tabular}[c]{@{}c@{}}Av. coverage\\ below/above\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ length\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ coverage\end{tabular} & \begin{tabular}[c]{@{}c@{}}Av. coverage\\ below/above\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ length\end{tabular} \\ \hline
\multicolumn{1}{c}{} & & & & & & & \\[-9pt]
\multicolumn{1}{c}{} & & \multicolumn{6}{c}{} \\[-8pt]
& \multicolumn{1}{l}{} & \multicolumn{3}{c}{low persistence} & \multicolumn{3}{c}{high persistence} \\
250 & EP &
79.74 & 7.16/13.10 & 0.140 & 78.93 & 7.64/13.43 & 0.289 \\
& RT &
90.22 & 3.42/6.36 & 0.140 & 90.34 & 2.98/6.68 & 0.289 \\
& SY &
88.37 & 3.64/7.99 & 0.145 & 88.53 & 3.51/7.96 & 0.302 \\
& AS &
87.59 & 3.27/9.14 & 0.146 & 87.87 & 3.14/8.99 & 0.304 \\
\hdashline
500 & EP &
82.41 & 6.51/11.08 & 0.104 & 82.12 & 6.66/11.22 & 0.211 \\
& RT &
90.13 & 4.12/5.75 & 0.104 & 90.83 & 3.40/5.77 & 0.211 \\
& SY &
88.83 & 4.26/6.91 & 0.106 & 89.35 & 3.50/7.15 & 0.218 \\
& AS &
89.08 & 3.57/7.35 & 0.108 & 89.39 & 3.28/7.33 & 0.221 \\
\hdashline
1,000 & EP &
84.98 & 6.09/8.93 & 0.076 & 83.46 & 6.84/9.70 & 0.155 \\
& RT &
90.16 & 4.63/5.21 & 0.076 & 90.29 & 4.41/5.30 & 0.155 \\
& SY &
89.37 & 4.33/6.30 & 0.077 & 89.22 & 4.34/6.44 & 0.159 \\
& AS &
89.53 & 4.06/6.41 & 0.079 & 89.35 & 4.09/6.56 & 0.161 \\
\hdashline
5,000 & EP &
88.37 & 5.06/6.57 & 0.036 & 87.95 & 5.08/6.97 & 0.074 \\
& RT &
90.72 & 4.72/4.56 & 0.036 & 90.16 & 4.84/5.0 & 0.074 \\
& SY &
90.52 & 4.44/5.04 & 0.036 & 89.95 & 4.41/5.64 & 0.074 \\
& AS &
91.23 & 4.03/4.74 & 0.037 & 90.54 & 4.11/5.35 & 0.076 \\
\hline \hline
\end{tabular}
}
\vspace{0.15cm}
\caption*{Table \ref{tab:4.1a} is exactly as Table \ref{tab:4.1} but for a \textbf{T-GARCH(1,1)} with (normalized) Student-t innovations ($6$ df) instead of a GARCH(1,1) with Student-t innovations ($6$ df).}
\end{table}
\egroup
Table \ref{tab:4.1} and \ref{tab:4.1a} report the results of the three $90\%$--bootstrap intervals for the $5\%$--VaR when the innovation distribution is Student-t (henceforth referred to as baseline) and the model is a GARCH(1,1) and a T-GARCH(1,1), respectively. In both tables the results of the interval \eqref{eq:4.3.10} based on asymptotic (AS) theory are included for comparison, where a Gaussian kernel is utilized together with a bandwidth following \citeauthor{silverman1986density}'s (\citeyear{silverman1986density}) rule-of-thumb. In the GARCH($1,1$) high persistence case (right part of Table 1), we see that the average coverage varies around $90\%$ across all sample sizes for the RT and the SY interval. In contrast, the EP and the AS interval fall short of the nominal $90\%$ by $10.43$ and $4.15$ percentage points (pp), respectively, for small sample size ($n=250$). Nevertheless, their average coverage approaches the nominal value as the sample size increases. Remarkably, for all four intervals the average rate of the conditional VaR being below the interval is considerably less than the average rate of the conditional VaR being above the interval when the sample size is rather small ($n \leq 500$). Regarding the intervals' length, we observe that the SY interval is on average larger than the EP/RT interval. As the sample size increases this gap diminishes and the intervals' average lengths shrink. Considering the low persistent case (left part of Table \ref{tab:4.1}) we find similar results regarding the intervals' average coverage, yet their average lengths turn out to be smaller compared to the high persistent case. This is intuitive as the conditional volatility tends to vary less in the low persistent case. Regarding the T-GARCH($1,1$) in Table \ref{tab:4.1a}, the overall picture is similar as in the GARCH(1,1) case, however the under-coverage in small and medium-sized samples appears to be more extreme for the EP and reduced for the AS interval.
Simulation results for the scenario when the $\eta_t$'s follow a standard normal distribution and when the model is a GARCH(1,1) and a T-GARCH(1,1), respectively, are tabulated in Tables \ref{tab:4.2} and \ref{tab:4.2a} which are given in Appendix \ref{sec add simulation}. Here we only note that, although the error distribution underlying the QMLE is correctly specified in this case, the qualitative results stated above with regard to Table \ref{tab:4.1} persist: the RT and the SY intervals possess accurate coverage rates across sample sizes, whereas the EP and the AS interval exhibit under-coverage in samples of rather small size with different extent.
Moreover, we observe that the intervals are on average shorter in the Gaussian case than in the baseline case. This seems partially driven by a smaller variance of $\hat{\xi}_{n,\alpha}$; for $\alpha=0.05$ the asymptotic variance $\zeta_\alpha$ in \eqref{eq:4.3.6} is equal to $3.11$ in the Gaussian case compared to $5.72$ in the Student-t case with $6$ degrees of freedom.
While the small-sample-performance of the AS interval can be explained by its embodied density estimation, the question arises why the EP interval performs worse than the other bootstrap intervals, which seems counter-intuitive at first. Howbeit the results are in line with the theoretical findings of \citeauthor{falk1991coverage} (\citeyear{falk1991coverage}, unnumbered Corollary, p.\ 488). In a random sample setting they prove that
the RT bootstrap interval for quantiles has asymptotically greater coverage than the corresponding EP bootstrap interval. The emerging gap\footnote{We neglect their $o(n^{-1/2})$ term. Take note that the theoretical results of \cite{falk1991coverage} are not directly applicable in our setting due to GARCH-type effects.}
\begin{enumerate}
\item[(i)] tends to be smaller for larger sample sizes,
\item[(ii)] tends to be larger for more extreme quantiles, and
\item[(iii)] tends to vary with the nominal coverage rate in a non-monotonic way.
\end{enumerate}
Table \ref{tab:4.5} presents the average coverage gap between the EP and the RT bootstrap interval of the conditional VaR for the baseline specification as well as for three deviations from the baseline (change in $F$, $\alpha$ and $\gamma$). For example, in the low persistence GARCH($1,1$) case of the baseline with $n=250$, the average coverage gap amounts to $90.73\%-81.20\%=9.53$pp (see also Table \ref{tab:4.1}). It is striking that all values are positive within Table \ref{tab:4.5}, which highlights the superiority of the RT bootstrap interval over the EP bootstrap interval. Further, it is eminent that average coverage gap tends to decrease with increasing sample size, which supports (i). Comparing columns (1) and (3) we also find that the average coverage gap tends to be larger for the $1\%$--VaR than for the $5\%$--VaR, which gives rise to (ii). Regarding (iii), the result of \citeauthor{falk1991coverage} (\citeyear{falk1991coverage}) suggests that the gap slightly decreases when increasing the nominal coverage from $90\%$ to $95\%$. Such tendency is precisely observed when comparing columns (1) and (4) of Table \ref{tab:4.5}.
\begin{table}[h]
\caption{\textbf{Average gap} between the \textbf{RT and the EP} fixed-design bootstrap intervals for different settings}\label{tab:4.5}
\centering
\begin{tabular}{rcccccccc}
\hline \hline
\begin{tabular}[c]{@{}c@{}}Sample\\ size\end{tabular} &
(1) & (2) & (3) & (4) & (1) & (2) & (3) & (4) \\
\hline
\\
& \multicolumn{8}{c}{Panel I: GARCH($1,1$)} \\
& \multicolumn{4}{c}{low persistence} & \multicolumn{4}{c}{high persistence} \\
250 & 9.53 & 9.07 & 13.26 & 8.41 & 10.65 & 9.63 & 14.79 & 9.30 \\
500 & 6.50 & 6.86 & 10.09 & 5.55 & 6.75 & 6.88 & 9.87 & 5.68 \\
1,000 & 5.03 & 5.02 & 8.87 & 4.38 & 4.68 & 4.85 & 8.60 & 4.24 \\
5,000 & 2.84 & 2.67 & 5.67 & 2.32 & 1.80 & 1.48 & 5.07 & 1.74 \\
& \multicolumn{8}{c}{Panel II: T-GARCH($1,1$)} \\
& \multicolumn{4}{c}{low persistence} & \multicolumn{4}{c}{high persistence} \\
250 & 10.48 & 9.39 & 15.22 & 9.07 & 11.41 & 10.41 & 15.03 & 10.23 \\
500 & 7.72 & 6.11 & 10.84 & 6.73 & 8.71 & 7.94 & 11.55 & 7.86 \\
1,000 & 5.18 & 4.23 & 8.55 & 4.40 & 6.83 & 5.54 & 9.20 & 5.82 \\
5,000 & 2.35 & 1.63 & 4.63 & 1.89 & 2.21 & 2.02 & 5.10 & 1.73 \\
\hline \hline
\end{tabular}
\vspace{0.15cm}
\caption*{Table \ref{tab:4.5} reports the average coverage gap between the RT and the EP fixed-design bootstrap interval in percentage points for different settings and sample sizes. For varying sample sizes ($n$) Panel I presents the results for the low and high persistence parameterization of a GARCH($1,1$), whereas Panel II displays the results for the corresponding T-GARCH($1,1$) processes.\\
($1$) $5\%$--VaR, Student-t innovations and $90\%$ nominal coverage (baseline)\\
($2$) $5\%$--VaR, Gaussian innovations and $90\%$ nominal coverage\\
($3$) $1\%$--VaR, Student-t innovations and $90\%$ nominal coverage\\
($4$) $5\%$--VaR, Student-t innovations and $95\%$ nominal coverage}
\end{table}
\bgroup
\begin{table}[tbp]
\caption{\textbf{Recursive-design} bootstrap confidence intervals for \textbf{GARCH(1,1)} with Student-t innovations}\label{tab:4.6}
\centering
\resizebox{\textwidth}{!}{\begin{tabular}{rc:ccc:ccc}
\hline \hline
\multicolumn{1}{c}{\begin{tabular}[c]{@{}c@{}}Sample\\ Size\end{tabular}} & & \begin{tabular}[c]{@{}c@{}}Average \\ coverage\end{tabular} & \begin{tabular}[c]{@{}c@{}}Av. coverage\\ below/above\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ length\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ coverage\end{tabular} & \begin{tabular}[c]{@{}c@{}}Av. coverage\\ below/above\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ length\end{tabular} \\ \hline
\multicolumn{1}{c}{} & & & & & & & \\[-9pt]
\multicolumn{1}{c}{} & & \multicolumn{3}{c}{low persistence} & \multicolumn{3}{c}{high persistence} \\
250 & EP &
81.65 & 5.99/12.36 & 0.619 & 80.54 & 5.87/13.59 & 0.859 \\
& RT &
90.49 & 3.78/5.73 & 0.619 & 90.09 & 4.08/5.83 & 0.859 \\
& SY &
90.20 & 2.78/7.02 & 0.653 & 90.49 & 2.93/6.58 & 0.918 \\
\hdashline
500 & EP &
84.66 & 5.55/9.79 & 0.441 & 84.50 & 5.38/10.12 & 0.604 \\
& RT &
90.39 & 4.29/5.32 & 0.441 & 90.25 & 4.60/5.15 & 0.604 \\
& SY &
90.62 & 3.36/6.02 & 0.459 & 90.88 & 3.27/5.85 & 0.628 \\
\hdashline
1,000 & EP &
85.78 & 5.40/8.82 & 0.310 & 86.19 & 5.08/8.73 & 0.427 \\
& RT &
90.24 & 4.47/5.29 & 0.310 & 90.33 & 4.28/5.39 & 0.427 \\
& SY &
90.30 & 3.76/5.94 & 0.318 & 90.69 & 3.57/5.74 & 0.438 \\
\hdashline
5,000 & EP &
87.68 & 5.60/6.72 & 0.144 & 87.75 & 5.35/6.90 & 0.191 \\
& RT &
90.28 & 4.72/5.00 & 0.144 & 89.96 & 4.84/5.20 & 0.191 \\
& SY &
90.14 & 4.49/5.37 & 0.146 & 89.69 & 4.59/5.72 & 0.193 \\
\hline \hline
\end{tabular}
}
\vspace{0.15cm}
\caption*{Table \ref{tab:4.6} reports distinct features of the \textbf{recursive-design} bootstrap confidence intervals for the conditional VaR at \textbf{level} $\bm{\alpha=0.05}$ with \textbf{nominal coverage} $\bm{1-\gamma=90\%}$. For each interval type and different sample sizes ($n$), the interval's average coverage rates (in $\%$), the average rate of the conditional VaR being below/above the interval (in $\%$) and the interval's average length are tabulated. The bootstrap intervals are based on $B=999$ bootstrap replications and the averages are computed using $S=10{,}000$ simulations. The results are for the low (left part) and the high (right part) persistence parametrization of a \textbf{GARCH(1,1)} with Student-$t$ innovations ($6$ df).
}
\end{table}
\egroup
\bgroup
\begin{table}[tbp]
\caption{\textbf{Recursive-design} bootstrap confidence intervals for \textbf{T-GARCH(1,1)} with Student-t innovations}\label{tab:4.6a}
\centering
\resizebox{\textwidth}{!}{\begin{tabular}{rc:ccc:ccc}
\hline \hline
\multicolumn{1}{c}{\begin{tabular}[c]{@{}c@{}}Sample\\ Size\end{tabular}} & & \begin{tabular}[c]{@{}c@{}}Average \\ coverage\end{tabular} & \begin{tabular}[c]{@{}c@{}}Av. coverage\\ below/above\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ length\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ coverage\end{tabular} & \begin{tabular}[c]{@{}c@{}}Av. coverage\\ below/above\end{tabular} & \begin{tabular}[c]{@{}c@{}}Average\\ length\end{tabular} \\ \hline
\multicolumn{1}{c}{} & & & & & & & \\[-9pt]
& \multicolumn{1}{l}{} & \multicolumn{3}{c}{low persistence} & \multicolumn{3}{c}{high persistence} \\
250 & EP &
79.62 & 6.79/13.59 & 0.143 & 79.12 & 7.35/13.53 & 0.294 \\
& RT &
90.54 & 3.49/5.97 & 0.143 & 90.98 & 2.67/6.35 & 0.294 \\
& SY &
89.31 & 3.33/7.36 & 0.150 & 89.17 & 3.10/7.73 & 0.309 \\
\hdashline
500 & EP &
82.33 & 6.25/11.42 & 0.105 & 82.28 & 6.56/11.16 & 0.214 \\
& RT &
90.27 & 4.26/5.47 & 0.105 & 91.29 & 3.37/5.34 & 0.214 \\
& SY &
89.39 & 3.97/6.64 & 0.109 & 89.90 & 3.36/6.74 & 0.223 \\
\hdashline
1,000 & EP &
85.22 & 5.81/8.97 & 0.077 & 83.79 & 6.58/9.63 & 0.157 \\
& RT &
89.97 & 4.88/5.15 & 0.077 & 90.29 & 4.37/5.34 & 0.157 \\
& SY &
90.09 & 4.12/5.79 & 0.079 & 89.69 & 4.13/6.18 & 0.162 \\
\hdashline
5,000 & EP &
88.53 & 4.82/6.65 & 0.036 & 88.10 & 4.93/6.97 & 0.075 \\
& RT &
90.40 & 4.95/4.65 & 0.036 & 90.12 & 4.99/4.89 & 0.075 \\
& SY &
90.91 & 4.23/4.86 & 0.037 & 90.61 & 4.13/5.26 & 0.076 \\
\hline \hline
\end{tabular}
}
\vspace{0.15cm}
\caption*{Table \ref{tab:4.6a} is exactly as Table \ref{tab:4.6} but for a \textbf{T-GARCH(1,1)} with Student-$t$ innovations ($6$ df).
}
\end{table}
\egroup
With regard to Remark \ref{rem:4.2} in Section \ref{sec:4.4.1}, Tables \ref{tab:4.6} and \ref{tab:4.6a} report the simulation results for the recursive-design bootstrap for the DGPs of Tables \ref{tab:4.1} and \ref{tab:4.1a}, respectively. We refer to Appendix \ref{app:4.B} for computational details. In comparison to the fixed-design approach (see Tables \ref{tab:4.1} and \ref{tab:4.1a}) we find that the recursive-design method performs similarly in terms of average coverage for each interval type, which corresponds to the simulation results of \cite{cavaliere2018fixed}. It is striking, however, that the intervals' average lengths are larger in the recursive-design than in the fixed-design set-up. For example, in the high persistence GARCH(1,1) case (right part of Table \ref{tab:4.6}) for $n=500$ the average length in the recursive-design approach is $0.604$ for the EP/RT interval compared to $0.581$ in the fixed-design. As the sample size increases this difference disappears.
In summary, the simulations suggest that the RT and the SY bootstrap interval work well for both bootstrap designs and that they outperform in smaller samples the AS interval in terms of average coverage even though their tails are unequally represented. In contrast, for both bootstrap designs the EP interval falls short of its nominal coverage, which is in line with the theoretical findings of \cite{falk1991coverage}. Since the fixed RT method leads on average to shorter intervals than the corresponding SY method and its recursive-design counterpart, this suggests to favor the fixed-design RT bootstrap interval in \eqref{eq:4.4.9}.
\subsection{Empirical Application}
\label{sec:4.5.2}
We analyze the French stock market index CAC 40 for the period January 1, 2015 -- January 1, 2020. The index values for the period are retrieved from Yahoo Finance and daily (log-) returns (expressed in $\%$) are computed using $\epsilon_t = 100\log(p_t/p_{t-1})$, where $p_t$ denotes the closing value of the index at trading day $t$.
\begin{figure}[tbp]
\centering
\begin{subfigure}[b]{0.45\textwidth}
\label{fig:2a}
\centering
\includegraphics[width=\textwidth]{fig_4_2_a_revision_bw.jpg}
\caption{Returns of CAC 40 }
\end{subfigure}
\quad
\begin{subfigure}[b]{0.45\textwidth}
\label{fig:2b}
\centering
\includegraphics[width=\textwidth]{fig_4_2_b_revision_bw.jpg}
\caption{Histogram of the residuals $\hat{\eta}_t$'s}
\end{subfigure}
\caption{The returns of the French stock market index CAC 40 are plotted in (a) for the period January 1, 2015 -- January 1, 2020. The histogram of the residuals is plotted in (b) after fitting a T-GARCH($1,1$) model to the subperiod January 1, 2015 -- July 1, 2019. A scaled normal density is superimposed.}
\label{fig:4.2}
\end{figure}
Figure \ref{fig:4.2}(a) displays the resulting series of returns. We disregard the observations from July 1, 2019 onwards, which we leave for the out-of-sample evaluation, yielding $n=1{,}146$ remaining observations (i.e.\ January 1, 2015 - July 1, 2019). For the volatility process we consider the T-GARCH($1,1$) model specified in Example \ref{ex:4.2}.\footnote{We also
consider an Asymmetric Power GARCH model \citep{ding1993long}, i.e.\ $\sigma_{t+1}^\delta= \omega_0 + \alpha_0^+ (\epsilon_{t}^+)^\delta +\alpha_0^-(\epsilon_{t}^-)^\delta + \beta_0 \sigma_{t}^\delta$ with $\delta>0$, which nests the GARCH($1,1$) model ($\delta=2$, $\alpha_0^+=\alpha_0^-$) and the T-GARCH($1,1$) model ($\delta=1$) of Examples \ref{ex:4.1} and \ref{ex:4.2}. In practice, the impact of the power $\delta$ on the volatility is minor and the QML approach of \cite{hamadeh2011asymptotic} suggests a $\delta$ close to $1$ in favor for the T-GARCH specification.}
Table \ref{tab:4.7} reports the corresponding point estimates with standard errors obtained by bootstrapping based on Algorithm \ref{alg:4.1}. As documented in numerous studies we find that the volatility persistence is close to unity. Further, we observe that $\hat{\alpha}^-_n$ is considerably larger than $\hat{\alpha}^+_n$ indicating a strong leverage effect, i.e.\ negative returns tend to increase volatility by more than positive returns of the same magnitude.
\begin{table}[tbp]
\caption{T-GARCH($1,1$) estimates for CAC 40}\label{tab:4.7}
\centering
\begin{tabular}{lcccc}
\hline \hline
& $\hat{\omega}_n$ & $\hat{\alpha}^+_n$ & $\hat{\alpha}^-_n$ & $\hat{\beta}_n$ \\ \hline
point estimate & $0.0292$ & $0.0046$ & $0.1798$ & $0.9026$ \\
std. error & $0.0109$ & $0.0215$ & $0.0339$ & $0.0234$ \\ \hline \hline
\end{tabular}
\vspace{0.15cm}
\caption*{T-GARCH($1,1$) estimates for the subperiod January 1, 1998 -- December 31, 2017. The standard errors are obtained by applying the fixed-design residual bootstrap with $B=2{,}000$ bootstrap replications.}
\end{table}
Figure \ref{fig:4.2}(b) plots the histogram of the residuals with the normal distribution superimposed. Further, we test the condition that the innovations are iid (see Assumption \ref{as:4.5}\eqref{as:4.5.1}) with the generalized run tests of \cite{cho2011generalized}.\footnote{The implementation of the tests is available on the website of the first author.} These tests are particularly suitable in this case since they can be based on the residuals and are sensitive against a wide range of alternatives. The test statistic of the sup-norm based test is $0.40$, which corresponds to a p-value of $0.27$. consequently, one cannot reject the null hypothesis of iid innovations at any common significance level. Similarly, the generalized run test based on the $L_1$-norm cannot be rejected at a $10\%$ significance level.
Next, we perform a rolling window analysis starting with subperiod January 1, 2015 -- July 1, 2019 and ending with subperiod July 8, 2015 -- January 1, 2020. We have $130$ subperiods each consisting of $1{,}146$ observations. For each rolling window period we fit a T-GARCH($1,1$) model and estimate the one-period-ahead conditional VaR associated with level $\alpha=0.05$.
For example, for the first window the T-GARCH($1,1$) estimates are reported in Table \ref{tab:4.7} and the conditional $5\%$-VaR of the one-period ahead (i.e.\ July 1, 2019) is estimated by $1.11$.
Further, we obtain the associated $95\%$-confidence intervals based on bootstrap and asymptotic normality. In addition to the RT intervals of the fixed- and residual-design bootstrap, we also computed an interval based on the asymptotic distribution.
The corresponding intervals are $[0.850,1.136]$ (fixed-design), $[0.834,1.115]$ (recursive-design), and $[0.828,1.106]$ (asymp.\ normality). Although the intervals are fairly similar, the asymptotic and recursive bootstrap intervals are shorter than the fixed-design interval. Given its tendency to underestimate variability in finite samples, this result is unsurprising for the asymptotic interval, although for the recursive bootstrap this contrasts the simulation findings. Note that the fixed-design iid and block bootstraps produce very similar interval, which is not surprising as our conducted specification tests did not indicate any violation of the iid assumption on the innovations.
The results of the rolling window analysis are visualized in Figure \ref{fig:4.3}. It plots the realized return together with (the opposite of) the estimated conditional VaR. For clarity we only indicate the lower and upper bound of the $95\%$ RT fixed-design bootstrap interval.
\begin{figure}[tbp]
\centering
\includegraphics[width=\textwidth]{fig_4_3_revision_bw.jpg}
\caption{Returns and the estimated conditional VaR (solid) for the period June 2, 2019 -- December 31, 2019. The estimation rests on the $1{,}146$ preceding observations. Lower and upper bounds for the conditional VaR (dashed) are based on the fixed-design bootstrap scheme using the RT method with $1-\gamma = 95\%$.}
\label{fig:4.3}
\end{figure}
We observe that in more turbulent times (e.g.\ August, 2019), the estimated VaR amplifies. In such volatile periods we expect the estimation risk to increase and, accordingly, we find wider bootstrap confidence intervals. In this regard, although not directly related to the bootstrap procedures considered it is worth mentioning that a proposal for monitoring VaR estimates over time can be found in \citet{HogaDemetrescu}.
\begin{remark}
\label{rem:5.1}
The point estimate $\hat{\alpha}^+_n$ in Table \ref{tab:4.7} is close to zero, indicating that the parameter $\alpha^+_0$ may lie on the boundary. In view of Assumption \ref{as:4.6}, the proposed bootstrap method becomes invalid when a parameter lies on the boundary of the parameter space. \cite{cavaliere2020bootstrap} modify the fixed volatility bootstrap to account for nuisance parameters on the boundary by shrinking the estimates for bootstrap sampling toward the boundary at an appropriate rate. Following their suggestion and using instead $\theta^*_n=(\hat{\omega}_n,0,\hat{\alpha}^-_n,\hat{\beta}_n)'$ to generate the bootstrap sample, does not notably alter the bootstrap results presented in this subsection.
\end{remark}
\section{Concluding Remarks}
\label{sec:4.6}
In this paper we study the two-step estimation procedure of \cite{francq2015risk} associated with the conditional VaR. In the first step, the conditional volatility parameters are estimated by QMLE, while the second step corresponds to approximating the quantile of the innovations' distribution by the empirical quantile of the residuals. A fixed-design residual bootstrap method is proposed to mimic the finite sample distribution of the two-step estimator and its consistency is proven under mild assumptions. In addition, an algorithm is provided for the construction of bootstrap intervals for the conditional VaR to take into account the uncertainty induced by estimation. Three interval types are suggested and a large-scale simulation study is conducted to investigate their performance in finite samples. We find that the equal-tailed percentile interval based on the fixed-design residual bootstrap tends to fall short of its nominal value, whereas the corresponding interval based on reversed tails yields accurate average coverage combined with the shortest average length. Although the result seems counter-intuitive at first, it is in line with the theoretical findings of \cite{falk1991coverage}. In the simulation study we also consider the recursive-design residual bootstrap. It turns out that the recursive-design and the fixed-design bootstrap perform similar in terms of average coverage. Yet in smaller samples the fixed-design scheme leads on average to shorter intervals. Further, the interval estimation by means of the fixed-design residual bootstrap is illustrated in an empirical application to daily returns of the French stock index CAC 40.
Natural extensions of this work are encompassing other risk measures such as Expected Shortfall \citep{heinemann2018residual} and developing a bootstrap procedure for the one-step estimator of \cite{francq2015risk}. Further, it is worthwhile to consider a smoothed bootstrap version in the spirit of \cite{hall1989smoothing}, which offers potential gains in accuracy. The latter two extensions are left for future research, as is the question how to extend the fixed-design bootstrap in order to give valid bootstrap inference when the iid assumption does not hold.
\section*{Acknowledgements}
The authors thank Franz Palm, Hanno Reuvers, Jean-Michel Zako\"ian and Christian Francq for useful comments and suggestions
as well as Dewi Peerlings and Benoit Duvocelle for computational support during the revision stage. In addition, the authors are grateful to the editors Oliver Linton and Torben Andersen, one of the associate editors and to two anonymous referees for their constructive remarks.
This research was financially supported by the Netherlands Organisation for Scientific Research (NWO).
\begin{thebibliography}{}
\bibitem[\protect\citeauthoryear{Bahadur}{Bahadur}{1966}]{bahadur1966note}
Bahadur, R.R. (1966).
\newblock A note on quantiles in large samples.
\newblock {\em The Annals of Mathematical Statistics\/}~{\em 37\/}(3),
577--580.
\bibitem[\protect\citeauthoryear{Bardet, Kamila, and Kengne}{Bardet
et~al.}{2020}]{Bardet2020}
Bardet, J.M., K.~Kamila, and W.~Kengne (2020).
\newblock Consistent model selection criteria and goodness-of-fit test for
common time series models.
\newblock {\em Electronic Journal of Statistics\/}~{\em 14\/}(1), 2009--2052.
\bibitem[\protect\citeauthoryear{Berkes and Horv{\'a}th}{Berkes and
Horv{\'a}th}{2003}]{berkes2003limit}
Berkes, I. and L.~Horv{\'a}th (2003).
\newblock Limit results for the empirical process of squared residuals in
\uppercase{GARCH} models.
\newblock {\em Stochastic Processes and their Applications\/}~{\em 105\/}(2),
271--298.
\bibitem[\protect\citeauthoryear{Beutner, Heinemann, and Smeekes}{Beutner
et~al.}{2019}]{beutner2019technical}
Beutner, E., A.~Heinemann, and S.~Smeekes (2019).
\newblock A general framework for prediction in time series models.
\newblock Working paper, Maastricht University,
\url{https://arxiv.org/pdf/1902.01622.pdf}.
\bibitem[\protect\citeauthoryear{Beutner, Heinemann, and Smeekes}{Beutner
et~al.}{2021}]{beutner2017justification}
Beutner, E., A.~Heinemann, and S.~Smeekes (2021).
\newblock A justification of conditional confidence intervals.
\newblock {\em Electronic Journal of Statistics\/}~{\em 15\/}(1), 2517--2565.
\bibitem[\protect\citeauthoryear{Billingsley}{Billingsley}{1986}]{billingsley1986probability}
Billingsley, P. (1986).
\newblock {\em Probability and Measure\/} (2nd ed.).
\newblock New York: John Wiley \& Sons.
\bibitem[\protect\citeauthoryear{Bollerslev}{Bollerslev}{1986}]{bollerslev1986generalized}
Bollerslev, T. (1986).
\newblock Generalized autoregressive conditional heteroskedasticity.
\newblock {\em Journal of Econometrics\/}~{\em 31\/}(3), 307--327.
\bibitem[\protect\citeauthoryear{Cavaliere, Nielsen, Pedersen, and
Rahbek}{Cavaliere et~al.}{2022}]{cavaliere2020bootstrap}
Cavaliere, G., H.B. Nielsen, R.S. Pedersen, and A.~Rahbek (2022).
\newblock Bootstrap inference on the boundary of the parameter space, with
application to conditional volatility models.
\newblock {\em Journal of Econometrics\/}~{\em 227\/}(1), 241--263.
\bibitem[\protect\citeauthoryear{Cavaliere, Pedersen, and Rahbek}{Cavaliere
et~al.}{2018}]{cavaliere2018fixed}
Cavaliere, G., R.S. Pedersen, and A.~Rahbek (2018).
\newblock The fixed volatility bootstrap for a class of \uppercase{ARCH}($q$)
models.
\newblock {\em Journal of Time Series Analysis\/}~{\em 39}, 920--941.
\bibitem[\protect\citeauthoryear{Cho and White}{Cho and
White}{2011}]{cho2011generalized}
Cho, J.S. and H.~White (2011).
\newblock Generalized runs tests for the iid hypothesis.
\newblock {\em Journal of Econometrics\/}~{\em 162\/}(2), 326--344.
\bibitem[\protect\citeauthoryear{Christoffersen and
Gon{\c{c}}alves}{Christoffersen and
Gon{\c{c}}alves}{2005}]{christoffersen2005estimation}
Christoffersen, P. and S.~Gon{\c{c}}alves (2005).
\newblock Estimation risk in financial risk management.
\newblock {\em The Journal of Risk\/}~{\em 7\/}(3), 1--28.
\bibitem[\protect\citeauthoryear{Corradi and Iglesias}{Corradi and
Iglesias}{2008}]{corradi2008bootstrap}
Corradi, V. and E.M. Iglesias (2008).
\newblock Bootstrap refinements for \uppercase{QML} estimators of the
\uppercase{GARCH}(1,1) parameters.
\newblock {\em Journal of Econometrics\/}~{\em 144\/}(2), 500--510.
\bibitem[\protect\citeauthoryear{Cs{\"o}rg{\H o} and
R{\'e}v{\'e}sz}{Cs{\"o}rg{\H o} and R{\'e}v{\'e}sz}{1981}]{csorgo1981strong}
Cs{\"o}rg{\H o}, M. and P.~R{\'e}v{\'e}sz (1981).
\newblock {\em Strong Approximations in Probability and Statistics}.
\newblock Budapest: Akad{\'e}miai Kiad{\'o}.
\bibitem[\protect\citeauthoryear{Davidson and Flachaire}{Davidson and
Flachaire}{2008}]{DavidsonFlachaire08}
Davidson, R. and E.~Flachaire (2008).
\newblock The wild bootstrap, tamed at last.
\newblock {\em Journal of Econometrics\/}~{\em 146}, 162--169.
\bibitem[\protect\citeauthoryear{Ding, Granger, and Engle}{Ding
et~al.}{1993}]{ding1993long}
Ding, Z., C.W. Granger, and R.F. Engle (1993).
\newblock A long memory property of stock market returns and a new model.
\newblock {\em Journal of Empirical Finance\/}~{\em 1\/}(1), 83--106.
\bibitem[\protect\citeauthoryear{Engle}{Engle}{1982}]{engle1982autoregressive}
Engle, R.F. (1982).
\newblock Autoregressive conditional heteroscedasticity with estimates of the
variance of {U}nited {K}ingdom inflation.
\newblock {\em Econometrica\/}~{\em 50\/}(4), 987--1007.
\bibitem[\protect\citeauthoryear{Escanciano}{Escanciano}{2009}]{Escanciano2009}
Escanciano, J.C. (2009).
\newblock Quasi-maximum likelihood estimation of semi-strong \uppercase{GARCH}
model.
\newblock {\em Econometric Theory\/}~{\em 25\/}(2), 561--570.
\bibitem[\protect\citeauthoryear{Falk and Kaufmann}{Falk and
Kaufmann}{1991}]{falk1991coverage}
Falk, M. and E.~Kaufmann (1991).
\newblock Coverage probabilities of bootstrap-confidence intervals for
quantiles.
\newblock {\em The Annals of Statistics\/}~{\em 19\/}(1), 485--495.
\bibitem[\protect\citeauthoryear{Francq, Horv{\'a}th, and Zako{\"\i}an}{Francq
et~al.}{2016}]{francq2016variance}
Francq, C., L.~Horv{\'a}th, and J.M. Zako{\"\i}an (2016).
\newblock Variance targeting estimation of multivariate \uppercase{GARCH}
models.
\newblock {\em Journal of Financial Econometrics\/}~{\em 14\/}(2), 353--382.
\bibitem[\protect\citeauthoryear{Francq and Zako{\"\i}an}{Francq and
Zako{\"\i}an}{2004}]{francq2004maximum}
Francq, C. and J.M. Zako{\"\i}an (2004).
\newblock Maximum likelihood estimation of pure \uppercase{Garch} and
\uppercase{Arma}-\uppercase{Garch} processes.
\newblock {\em Bernoulli\/}~{\em 10\/}(4), 605--637.
\bibitem[\protect\citeauthoryear{Francq and Zako\"ian}{Francq and
Zako\"ian}{2011}]{francq2011garch}
Francq, C. and J.M. Zako\"ian (2011).
\newblock {\em \uppercase{GARCH} Models: Structure, Statistical Inference and
Financial Applications}.
\newblock Chichester: John Wiley \& Sons.
\bibitem[\protect\citeauthoryear{Francq and Zako{\"\i}an}{Francq and
Zako{\"\i}an}{2015}]{francq2015risk}
Francq, C. and J.M. Zako{\"\i}an (2015).
\newblock Risk-parameter estimation in volatility models.
\newblock {\em Journal of Econometrics\/}~{\em 184\/}(1), 158--173.
\bibitem[\protect\citeauthoryear{Francq and Zako{\"\i}an}{Francq and
Zako{\"\i}an}{2016}]{FrancqZakoian2016}
Francq, C. and J.M. Zako{\"\i}an (2016).
\newblock Estimating multivariate volatility models equation by equation.
\newblock {\em Journal of the Royal Statistical Society, Series B\/}~{\em
76\/}(3), 613--635.
\bibitem[\protect\citeauthoryear{Francq and Zakoïan}{Francq and
Zakoïan}{2022}]{FRANCQ202247}
Francq, C. and J.M. Zakoïan (2022).
\newblock Testing the existence of moments for garch processes.
\newblock {\em Journal of Econometrics\/}~{\em 227\/}(1), 47--64.
\bibitem[\protect\citeauthoryear{Friedrich, Smeekes, and Urbain}{Friedrich
et~al.}{2020}]{FSU20}
Friedrich, M., S.~Smeekes, and J.P. Urbain (2020).
\newblock Autoregressive wild bootstrap inference for nonparametric trends.
\newblock {\em Journal of Econometrics\/}~{\em 214}, 81--109.
\bibitem[\protect\citeauthoryear{Gao and Song}{Gao and
Song}{2008}]{gao2008estimation}
Gao, F. and F.~Song (2008).
\newblock Estimation risk in \uppercase{GARCH} \uppercase{V}a\uppercase{R} and
\uppercase{ES} estimates.
\newblock {\em Econometric Theory\/}~{\em 24\/}(5), 1404--1424.
\bibitem[\protect\citeauthoryear{Geweke}{Geweke}{1986}]{geweke1986comment}
Geweke, J. (1986).
\newblock Comment on: modelling the persistence of conditional variances.
\newblock {\em Econometric Reviews\/}~{\em 5}, 57--61.
\bibitem[\protect\citeauthoryear{Glosten, Jagannathan, and Runkle}{Glosten
et~al.}{1993}]{glosten1993relation}
Glosten, L.R., R.~Jagannathan, and D.E. Runkle (1993).
\newblock On the relation between the expected value and the volatility of the
nominal excess return on stocks.
\newblock {\em The Journal of Finance\/}~{\em 48\/}(5), 1779--1801.
\bibitem[\protect\citeauthoryear{Gonçalves and Kilian}{Gonçalves and
Kilian}{2004}]{GONCALVES2004}
Gonçalves, S. and L.~Kilian (2004).
\newblock Bootstrapping autoregressions with conditional heteroskedasticity of
unknown form.
\newblock {\em Journal of Econometrics\/}~{\em 123\/}(1), 89--120.
\bibitem[\protect\citeauthoryear{Hall, DiCiccio, and Romano}{Hall
et~al.}{1989}]{hall1989smoothing}
Hall, P., T.J. DiCiccio, and J.P. Romano (1989).
\newblock On smoothing and the bootstrap.
\newblock {\em The Annals of Statistics\/}~{\em 17\/}(2), 692--704.
\bibitem[\protect\citeauthoryear{Hall and Heyde}{Hall and
Heyde}{1980}]{hall1980martingale}
Hall, P. and C.C. Heyde (1980).
\newblock {\em Martingale Limit Theory and its Application}.
\newblock New York: Academic Press.
\bibitem[\protect\citeauthoryear{Hall and Martin}{Hall and
Martin}{1988}]{hall1988bootstrap}
Hall, P. and M.A. Martin (1988).
\newblock On bootstrap resampling and iteration.
\newblock {\em Biometrika\/}~{\em 75\/}(4), 661--671.
\bibitem[\protect\citeauthoryear{Hall and Yao}{Hall and
Yao}{2003}]{hall2003inference}
Hall, P. and Q.~Yao (2003).
\newblock Inference in \uppercase{ARCH} and \uppercase{GARCH} models with
heavy--tailed errors.
\newblock {\em Econometrica\/}~{\em 71\/}(1), 285--317.
\bibitem[\protect\citeauthoryear{Hamadeh and Zako{\"\i}an}{Hamadeh and
Zako{\"\i}an}{2011}]{hamadeh2011asymptotic}
Hamadeh, T. and J.M. Zako{\"\i}an (2011).
\newblock Asymptotic properties of \uppercase{LS} and \uppercase{QML}
estimators for a class of nonlinear \uppercase{GARCH} processes.
\newblock {\em Journal of Statistical Planning and Inference\/}~{\em 141\/}(1),
488--507.
\bibitem[\protect\citeauthoryear{Hartz, Mittnik, and Paolella}{Hartz
et~al.}{2006}]{hartz2006accurate}
Hartz, C., S.~Mittnik, and M.~Paolella (2006).
\newblock Accurate value-at-risk forecasting based on the
normal-\uppercase{GARCH} model.
\newblock {\em Computational Statistics \& Data Analysis\/}~{\em 51\/}(4),
2295--2312.
\bibitem[\protect\citeauthoryear{Heinemann and Telg}{Heinemann and
Telg}{2018}]{heinemann2018residual}
Heinemann, A. and S.~Telg (2018).
\newblock A residual bootstrap for conditional expected shortfall.
\newblock arXiv Preprint 1811.11557.
\bibitem[\protect\citeauthoryear{Hetland, Pedersen, and Rahbek}{Hetland
et~al.}{ress}]{HETLAND2021}
Hetland, S., R.S. Pedersen, and A.~Rahbek (in press).
\newblock Dynamic conditional eigenvalue garch.
\newblock {\em Journal of Econometrics\/}.
\bibitem[\protect\citeauthoryear{Hidalgo and Zaffaroni}{Hidalgo and
Zaffaroni}{2007}]{hidalgo2007goodness}
Hidalgo, J. and P.~Zaffaroni (2007).
\newblock A goodness-of-fit test for \uppercase{ARCH}($\infty$) models.
\newblock {\em Journal of Econometrics\/}~{\em 141\/}(2), 835--875.
\bibitem[\protect\citeauthoryear{Hjort and Pollard}{Hjort and
Pollard}{2011}]{hjort2011asymptotics}
Hjort, N.L. and D.~Pollard (2011).
\newblock Asymptotics for minimisers of convex processes.
\newblock {\em Preprint arXiv:1107.3806v1\/}.
\bibitem[\protect\citeauthoryear{Hoga and Demetrescu}{Hoga and
Demetrescu}{2023}]{HogaDemetrescu}
Hoga, Y. and M.~Demetrescu (2023).
\newblock Monitoring value-at-risk and expected shortfall forecasts.
\newblock {\em Management Science\/}~{\em 69\/}(5), 2954--2971.
\bibitem[\protect\citeauthoryear{Jeong}{Jeong}{2017}]{jeong2017residual}
Jeong, M. (2017).
\newblock Residual-based \uppercase{GARCH} bootstrap and second order
asymptotic refinement.
\newblock {\em Econometric Theory\/}~{\em 33\/}(3), 779--790.
\bibitem[\protect\citeauthoryear{Jim\'{e}nez-Gamero, Lee, and
Meintanis}{Jim\'{e}nez-Gamero et~al.}{2020}]{Maria2019}
Jim\'{e}nez-Gamero, M.D., S.~Lee, and S.G. Meintanis (2020).
\newblock Goodness-of-fit tests for parametric specifications of conditionally
heteroscedastic models.
\newblock {\em TEST\/}~{\em 29}, 682--703.
\bibitem[\protect\citeauthoryear{Koenker and Xiao}{Koenker and
Xiao}{2006}]{koenker2006quantile}
Koenker, R. and Z.~Xiao (2006).
\newblock Quantile autoregression.
\newblock {\em Journal of the American Statistical Association\/}~{\em
101\/}(475), 980--990.
\bibitem[\protect\citeauthoryear{Kreiss}{Kreiss}{2016}]{kreiss2015discussion}
Kreiss, J.P. (2016).
\newblock Discussion: bootstrap prediction intervals for linear, nonlinear and
nonparametric autoregressions.
\newblock {\em Journal of Statistical Planning and Inference\/}~{\em 177},
28--30.
\bibitem[\protect\citeauthoryear{Lahiri}{Lahiri}{2003}]{Lahiri03}
Lahiri, S.N. (2003).
\newblock {\em Resampling Methods for Dependent Data}.
\newblock New York: Springer-Verlag.
\bibitem[\protect\citeauthoryear{Li, Peng, and Song}{Li
et~al.}{ress}]{li_peng_song_2022}
Li, S., L.~Peng, and X.~Song (in press).
\newblock Simultaneous confidence bands for conditional value-at-risk and
expected shortfall.
\newblock {\em Econometric Theory\/}.
\bibitem[\protect\citeauthoryear{Linton, Pan, and Wang}{Linton
et~al.}{2010}]{Linton2010}
Linton, O., J.~Pan, and H.~Wang (2010).
\newblock Estimation for a nonstationary semi-strong {GARCH}(1,1) model with
heavy-tailed errors.
\newblock {\em Econometric Theory\/}~{\em 26\/}(1), 1--28.
\bibitem[\protect\citeauthoryear{Mammen}{Mammen}{1993}]{Mammen93}
Mammen, E. (1993).
\newblock Bootstrap and wild bootstrap for high dimensional linear models.
\newblock {\em Annals of Statistics\/}~{\em 21}, 255--285.
\bibitem[\protect\citeauthoryear{McNeil and Frey}{McNeil and
Frey}{2000}]{McNeil2002}
McNeil, A.J. and R.~Frey (2000).
\newblock Estimation of tail-related risk measures for heteroscedastic
financial time series: an extreme value approach.
\newblock {\em Journal of Empirical Finance\/}~{\em 7\/}(3), 271 -- 300.
\bibitem[\protect\citeauthoryear{Nelson}{Nelson}{1991}]{nelson1991conditional}
Nelson, D.B. (1991).
\newblock Conditional heteroskedasticity in asset returns: A new approach.
\newblock {\em Econometrica: Journal of the Econometric Society\/}~{\em
59\/}(2), 347--370.
\bibitem[\protect\citeauthoryear{Pantula}{Pantula}{1986}]{pantula1986modeling}
Pantula, S.G. (1986).
\newblock Modeling the persistence of conditional variances: a comment.
\newblock {\em Econometric Reviews\/}~{\em 5}, 79--97.
\bibitem[\protect\citeauthoryear{Pascual, Romo, and Ruiz}{Pascual
et~al.}{2006}]{pascual2006bootstrap}
Pascual, L., J.~Romo, and E.~Ruiz (2006).
\newblock Bootstrap prediction for returns and volatilities in
\uppercase{garch} models.
\newblock {\em Computational Statistics \& Data Analysis\/}~{\em 50\/}(9),
2293--2312.
\bibitem[\protect\citeauthoryear{Pesaran}{Pesaran}{2015}]{pesaran2015time}
Pesaran, M.H. (2015).
\newblock {\em Time Series and Panel Data Econometrics}.
\newblock Oxford: Oxford University Press.
\bibitem[\protect\citeauthoryear{Phillips}{Phillips}{1979}]{phillips1979sampling}
Phillips, P.C.B. (1979).
\newblock The sampling distribution of forecasts from a first-order
autoregression.
\newblock {\em Journal of Econometrics\/}~{\em 9\/}(3), 241--261.
\bibitem[\protect\citeauthoryear{Roussas}{Roussas}{1997}]{roussas1997course}
Roussas, G.G. (1997).
\newblock {\em A Course in Mathematical Statistics\/} (2nd ed.).
\newblock San Diego: Academic Press.
\bibitem[\protect\citeauthoryear{Shao}{Shao}{2010}]{Shao10}
Shao, X. (2010).
\newblock The dependent wild bootstrap.
\newblock {\em Journal of the American Statistical Association\/}~{\em 105},
218--235.
\bibitem[\protect\citeauthoryear{Shimizu}{Shimizu}{2009}]{shimizu2009bootstrapping}
Shimizu, K. (2009).
\newblock {\em Bootstrapping Stationary \uppercase{ARMA}--\uppercase{GARCH}
Models}.
\newblock Springer.
\bibitem[\protect\citeauthoryear{Silverman}{Silverman}{1986}]{silverman1986density}
Silverman, B. (1986).
\newblock {\em Density Estimation for Statistics and Data Analysis Estimation
Density}.
\newblock London: Chapman and Hall.
\bibitem[\protect\citeauthoryear{Spierdijk}{Spierdijk}{2016}]{spierdijk2016confidence}
Spierdijk, L. (2016).
\newblock Confidence intervals for \uppercase{ARMA}--\uppercase{GARCH}
value-at-risk: the case of heavy tails and skewness.
\newblock {\em Computational Statistics \& Data Analysis\/}~{\em 100},
545--559.
\bibitem[\protect\citeauthoryear{Xiong and Li}{Xiong and
Li}{2008}]{xiong2008some}
Xiong, S. and G.~Li (2008).
\newblock Some results on the convergence of conditional distributions.
\newblock {\em Statistics \& Probability Letters\/}~{\em 78\/}(18), 3249--3253.
\bibitem[\protect\citeauthoryear{Zako{\"\i}an}{Zako{\"\i}an}{1994}]{zakoian1994threshold}
Zako{\"\i}an, J.M. (1994).
\newblock Threshold heteroskedastic models.
\newblock {\em Journal of Economic Dynamics and Control\/}~{\em 18\/}(5),
931--955.
\end{thebibliography}