Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,970 characters · 7 sections · 66 citation commands
Estimation of Integrated Volatility Functionals with Kernel Spot Volatility Estimators
In this work, we are concerned with the estimation of integrated volatility functionals of the form $V_{T}(g):=\int_{0}^{T}g(c_s)ds$, where $g$ is a smooth function and $c_t:= \sigma_t\sigma_t^*$ is the spot {\color{Blue} variance-covariance} process of a $d$-dimensional It\^o semimartingale given by
Here, $\{W_{t}\}_{t\geq{}0}$ is $d$-dimensional standard Brownian Motion (BM) and $\{J_{t}\}_{t\geq{}0}$ is a pure-jump process of bounded variation. Such functionals have many applications in financial econometrics. In the simplest case, when $d=1$ and $g(x) =x$, this is the well-studied integrated volatility (IV) or variance $IV_T=\int_0^{T}\sigma^2_sds$, which measures the overall variability or {\color{Blue} uncertainty} latent in the process $X$ during $[0,T]$. Again, when $d=1$ and $g(x) = x^2$, $V_T(g)$ corresponds to the integrated quarticity, which appears in the limiting distribution of the realized variance estimator $\widehat{IV}_T=\sum_{i=1}^n (\Delta_i^nX)^2$ and, hence, it is crucial to construct feasible confidence intervals for IV. Another function appearing in the literature is $g(x)=\log(x)$, when researching the statistical features of a factor $Y$ driving the volatility in a model where $\sigma$ is assumed to be the exponential of $Y$. Volatility functionals have also been used in principal component analysis by ait2019principal and in estimating integrated moments for option pricing purposes by li2016generalized.
One of the fundamental methods to estimate $V_{T}(g)=\int_{0}^{T}g(c_s)ds$ consists of approximating the integral by a Riemman sum and {\color{Blue} plugging} in a suitable spot volatility estimator $\hat{c}$ in place of $c$:
This method was pioneered by jacod2013quarticity, where, as the spot volatility estimator, a local rolling window estimator with uniform weights was used:
Two interesting facts emerged from their theory. {\color{Red} For the ease of exposition, we focus on the $d=1$-case in the introduction, though our main results are stated for general dimension $d\geq 1$.} First of all, the rate of convergence of the estimation error is unexpectedly of order $\Delta_n^{\frac{1}{2}}$ even though $\hat{c}_{i}^{n}$ can only attain the rate $\Delta_n^{\frac{1}{4}}$\footnote{{\color{Red} Intuitively, by {\color{Blue} Taylor} expansion of the estimator $g(\hat{c})$, we can write $\Delta_n\sum_{i=1}^n g(\hat{c}_i^n) - \Delta_n\sum_{i=1}^{n}g(c_{t_{i-1}})= \Delta_n\sum_{i=1}^{n}g'(c_{t_{i-1}}) \left(\hat{c}_{i}^n - c_{t_{i-1}}\right) + \frac{1}{2}\Delta_n\sum_{i=1}^{n}g''(c_{t_{i-1}}) \left(\hat{c}_{i}^n - c_{t_{i-1}}\right)^2+\text{higher order terms}$. The estimation errors in the first term of the right-hand side are “averaged out" so that the convergence rate changes from $n^{-1/4}$ to $n^{-1/2}$.}}. Secondly, when $k_n \sim \theta/\sqrt{\Delta_n} $ for some $\theta\in(0,\infty)$, which yields the best convergence rate for the spot volatility estimator $\hat{c}_{i}^{n}$, $\frac{1}{\sqrt{\Delta_n}}\left( \widehat{V}_{T}(g) - V_{T}(g) \right) $ converges to a mixed Gaussian distribution with four bias terms:
In principle, the bias terms can be estimated and eliminated from $\widehat{V}_{T}(g)$ to obtain a feasible CLT with centered Gaussian limiting distribution. This strategy was explored by jacod2015estimation. An alternative solution, however, was adopted in jacod2013quarticity. Specifically, the window length $k_n$ was made to converge to $\infty$ at a slower rate (under-smoothing) so that $k_n \sqrt{\Delta_n}\to{}0$. This corresponds to the case $\theta= 0 $ and, hence, the first, third, and fourth bias terms would vanish, while the second term becomes dominant. To keep the convergence rate $\sqrt{\Delta_n} $ of the estimator $\widehat{V}_T(g) $, this second term needs to be estimated and the estimator needs to be de-biased. The resulting estimator takes the form
which indeed achieves the rate $\Delta_n^{\frac{1}{2}}$ with semi-parametrically optimal asymptotic variance. {\color{Red} A similar approach to deal with {\color{Blue} edge} effects and biases was also considered in kristensen2010nonparametric, who also proposed kernel estimators and obtained a central limit theorem (though the model considered therein lacks leverage effects).} Alternatively, as mentioned before, in the optimal case $\theta \in (0,\infty)$ and employing forward uniform kernels, estimators for all the bias terms were proposed in jacod2015estimation, where an unbiased estimator with centered limiting distribution was developed.
li2019efficient took a different approach towards bias correction. They proposed a type of `jackknife' estimators, which are formed as a linear combination of a few uncorrected estimators of the form (ref) associated with different windows $k_n$. Since these estimators have similar bias terms but different coefficients depending on the {\color{Blue} window} sizes, the biases can be eliminated by a proper linear combination. Both the `two-scale' jackknife (i.e., when combining only two uncorrected {\color{Blue} estimators}) and jacod2013quarticity's estimator use under-smoothing, so that the biases due to boundary effects, volatility of volatility, and volatility jumps are rendered asymptotically negligible. One drawback of this approach is that one needs to determine the degree of undersmoothing. In other words, if, for instance, we take $k_n \sim \beta\Delta_n^{\alpha}$ with $\alpha>-1/2$, then one needs to tune both parameters $\beta$ and $\alpha$, which is nontrivial. In contrast, a multiscale jackknife estimator, formed as a linear combination of three (or more) estimators, can in principle cancel all the biases regardless of the speed of convergence of the window size. However, again, this involves more tuning parameters that need to be carefully calibrated.
Other related work on this topic includes mykland2009inference, who proposed using the same method as in jacod2013quarticity, but {\color{Red} keeping $k_n = k$ constant}; for the quarticity or more generally for estimating $\int_0^T g(c_s) ds $ with a power functional $g(x) = x^r$, they obtained a CLT with rate $1/\sqrt{\Delta_n}$ and an asymptotic variance that is not efficient, {\color{Red} but can arbitrarily approach the optimal variance} when $k$ is large. {\color{Red} Glotter established a Locally Asymptotically Mixed Normality (LAMN) condition for the estimation of integrated volatility functionals in a continuous diffusion process where the volatility is driven by an independent It\^o process. The result therein shows that the optimal rate of convergence is $n^{-1/2}$. Another related result in the same direction is Renault}. {\color{Red} More recently, LiLiu2021 generalized jacod2013quarticity 's result to allow a long-memory component such as {\color{Blue} an} fBM with Hurst parameter $H>1/2$. They also established a semiparametric efficiency lower bound, generalizing Renault's results.} chen2018inference developed an estimator for $\widehat{V}_{T}(g)$ in the presence of microstructure noise, based on a forward finite difference approximation of the standard pre-averaging estimator of the integrated volatility.
Our work aims to apply a general kernel spot volatility estimator to the plug-in estimator of integrated volatility functionals. The general kernel estimator (fan2008spot, kristensen2010nonparametric) is defined as
where $K_{{b}}(x):=K(x/{b})/{b}$ {\color{Red} for some bandwidth $b>0$}, and $k_n$ is an appropriate sequence that controls the bandwidth $b_n=k_n\Delta_n$ and that needs to be calibrated. Of course, one can see (ref) as a particular case of (ref) with $K(x)={\bf 1}_{[0,1]}(x)$. {\color{Red} kristensen2010nonparametric developed an asymptotic theory for (ref) when $t_i$ is a fixed time $\tau\in (0,T)$ and showed asymptotic normality to the spot variance $c_{\tau}$ under certain path-wise H\"older continuity and non-leverage conditions on $c$. See also fan2008spot and mancini2015estimation for other related results. All these results applied undersmoothing, which allowed them to neglect the “target error" coming from approximating the spot volatility by a kernel weighted integrated volatility.}
In this paper, we consider the following version:
with $\bar{K}_{n}\left(t_i\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right)$, which is known to be more robust than (ref) {\color{Blue} against} edge effects {\color{Red} (see also kristensen2010nonparametric for other edge-robust alternatives)}. As a result, we can, in principle, consider the whole range $i=1,\dots, n$ in (ref). Furthermore, we find it more convenient for our proofs to consider the estimator
The coefficient $\bar{K}_n(t_i)$ attenuates the contribution of $g\left(\hat{c}_{i}^{n}\right)$ for $t_i$ near $0$ and $T$ and, thus, serves as an additional edge effect correction. {\color{Red} kristensen2010nonparametric also considered a kernel-based estimator similar to (ref), but applied an edge effect correction, similar to that in jacod2015estimation, by eliminating the terms corresponding to $i$'s close to $0$ and $n$, and did not consider the adjustments $\bar{K}_n(t_i)$. Their asymptotic theory also only considered the undersmoothing asymptotic regime mentioned above.}
The motivation for considering general kernels comes from FigLi and FigWu, where it has been shown, both theoretically and by Monte Carlo experiments, that a general two-sided kernel with unbounded support could significantly reduce the variance of the spot volatility estimator $\hat{c}$. More specifically, FigLi studied the infill asymptotic behavior of the mean-square error (MSE) of (ref) {for continuous It\^o semimartingales} and showed that, when setting $k_n$ `optimally' (i.e., in order to minimize the leading order terms of the MSE), the double exponential kernel $K(x)=.5 e^{-|x|}$ minimizes the resulting MSE over all kernels. This result was established under the absence of leverage effects. FigWu extended this result, by first proving a CLT for a general {two-sided} kernel, with optimal convergence rate and in the presence of jumps and leverage {effects}. It was then shown that exponential kernels minimize the asymptotic variance of the CLT. It is worth noting that, even when applying undersmoothing, the asymptotic variance could be reduced when choosing a kernel of unbounded support. Specifically, when applying undersmoothing, the asymptotic variance of the spot volatility estimator is proportional to $\|K\|^2=\int_{-\infty}^{\infty} K^2(x) dx$. When $K$ is constrained to have support on $[0,1]$, the optimal kernel is indeed the uniform kernel $K_{unif}(x)={\bf 1}_{[0,1]}(x)$ with $\|K_{unif}\|^2=1$\footnote{{\color{Red} Indeed, by Jensen's inequality, $K_{unif}(x)={\bf 1}_{[0,1]}(x)$ satisfies $\int K^2_{unif}(x)dx=1=\left(\int K(x)dx\right)^2\leq \int K^2(x)dx$, which shows that a uniform kernel $K_{unif}(x)$ minimizes the asymptotic variance when applying undersmoothing.}}. However, picking the kernels $.5e^{-|x|}$ {\color{Red} and $e^{-x}{\bf 1}_{x>0}$ will result in estimators whose asymptotic variances are respectively a quarter and a half of} that obtained by a uniform kernel $K_{unif}(x)={\bf 1}_{[0,1]}(x)$. {\color{Red} Similar comments apply when considering only backward looking kernels such as ${\bf 1}_{[-1,0)}(x)$ and $e^{-|x|}{\bf 1}_{x<0}$.}
A natural question is whether the estimation accuracy gained when estimating the spot volatility by a general kernel transfers into gains in estimating integrated functionals. Intuitively, by the Taylor expansion of the estimator $g(\hat{c})$, we can write $g(\hat{c}) - g(c)= g'(c) \left(\hat{c} - c\right) + \frac{1}{2}g''(c)\left(\hat{c} - c\right)^2 + O_p\left(|\hat{c} - c|^3 \right)$. After taking expectations, the variance of the spot volatility estimator, which appears in the second term, becomes a bias term of the integrated volatility estimator. Therefore, the bias of the estimator $\widehat{V}_{T}(g)$ could in principle be reduced by adopting a general kernel.
As it turns out, the CLT of (ref) again exhibits several biases, which depend on a rather intricate and nontrivial manner on the kernel. The bias terms compared to the uniform kernel case involves a highly nontrivial transformation of the kernel function. The second bias term, involving $\frac{1}{\theta} \int_0^T g^{\prime \prime}(c_s) c_s^2ds$, contains the $L_2$ norm of the kernel and, thus, can be reduced because certain kernels of unbounded support {\color{Blue} possess} smaller $L_2$ norm. We then propose bias-corrected estimator similar to those in jacod2015estimation, and derive the corresponding CLT. We show by Monte Carlo experiments that our general kernel estimator typically has superior performance compared with the jackknife estimators of li2019efficient. Of course, our estimator is simpler as it does not involve as many tuning parameters, which are hard to tune-up. In fact, our Monte Carlo experiments suggest our estimators are significantly more stable in the sense that they exhibit better performance for most values of the bandwidth, while the jackknife estimator requires accurate tuning of the `bandwidth' to match our estimator's performance. We also observe that even without bias corrections, under certain conditions, the plug-in estimator with general kernel spot volatility estimator has good performance, which could be a result of the increased accuracy of {\color{Blue} the} spot volatility estimator and/or the reduction in bias attained by using a general kernel.
While working on the final stages of writing the present manuscript, we became aware of a recent work by benvenutifunctionals, who also studied a type of plug-in estimator with a general two-sided kernel spot volatility estimator. However, unlike our work, only a suboptimal bandwidth asymptotic regime (i.e., assuming under-smoothing) was considered and only consistency was established.
The rest of this paper is organized as follows. Section (ref) introduces the framework, assumptions, and main results. Section (ref) illustrates the performance of our method via Monte Carlo simulations and compares it to the jackknife estimator of li2019efficient. The proofs are deferred to an appendix section.
Notation: We shall use the following notation:
Throughout, we consider a $d$-dimensional It\^o semimartingale $X=[X^1,\dots,X^d]^*\in\mathbb{R}^{d\times{}1}$ of the form:
where all stochastic processes ($\mu :=\left\{\mu_{t}\right\}_{t \geq 0}$, $\sigma :=\left\{\sigma_{t}\right\}_{t \geq 0}$, $W :=\left\{W_{t}\right\}_{t \geq 0}$, $\delta:=\left\{\delta(t, z):t\geq 0, z\in E\right\}$, $\mathfrak{p}:=\{\mathfrak{p}(B):B\in\mathcal{B}(\mathbb{R}_{+}\times E)\}$, etc.) are defined on a complete filtered probability space $\left(\Omega, \mathcal{F}, \mathbb{F}, \mathbb{P}\right)$ with filtration $ \mathbb{F}=(\mathcal{F}_{t})_{t \geq 0}$. Here, {\color{DG} $W=[W^1,\dots,W^d]^*$} is a $d$-dimensional standard Brownian Motion (BM) adapted to the filtration \(\mathbb{F}\), $\delta$ is a predictable $\mathbb{R}^d$-valued function on $\Omega \times \mathbb{R}_{+} \times E$, and $\mathfrak{p}$ is a Poisson random measure on $\mathbb{R}_{+} \times E$ for some arbitrary Polish space $E$ with compensator $\mathfrak{q}(\mathrm{d} u, \mathrm{~d} x)=\mathrm{d} u \otimes \lambda(\mathrm{d} x)$, where $\lambda$ is a $\sigma-$ finite measure on $E$ having no atom. We also denote the spot variance-covariance process as $c_t := \sigma_t \sigma_t^*$, which takes values in the set $\mathcal{M}_d^+$ of nonnegative definite {\color{DG} $d\times d$} {\color{DR} symmetric} matrices. {\color{Blue} As common} in the literature, $c$ is called the spot volatility of the process. We assume $c$ is also an It\^o semimartingale following the dynamics:
where {$W:=\{W_t\}_{t\geq0}$} is the same $d$-dimensional Brownian Motion driving the dynamics of $X$. {\color{Red} The assumption (ref) is common in the literature (see, e.g., the monographs JacodProtter,jacodaitsahalia), and was also assumed in the work of jacod2013quarticity{\color{Blue} )}.} {\color{DG} For each $1\leq \ell, m\leq d$, $\{\tilde{\mu}^{\ell m}_t\}_{t\geq{}0}$ is adapted locally bounded {\color{DR} and} $\{\tilde{\sigma}^{\ell m}_t\}_{t\geq{}0}$ is adapted c\`adl\`ag.}
We now state the main assumption on the process $X$:
The parameter $r$ determines the jump activity of the process: the larger $r$ is, the more active or frequent are the small jumps of the process. When $r=0$, the process {\color{Blue} exhibits} {\color{DR} finitely} many jumps in any bounded time interval (in that case, we say that the jumps are of finite activity). When $r<1$, the jump component of the process is of bounded variation, which is a standard {\color{Blue} constraint} for the truncated realized quadratic variation to achieve efficiency (JacodProtter).
Next, we give the conditions on the kernel function $K$ {\color{DG} for which we need to define the function} \[ {\color{DG} L(t) := \int_{t}^{\infty} K(u) d u \mathbf{1}_{\{t>0\}}-\int_{-\infty}^{t} K(u) d u \mathbf{1}_{\{t \leq 0\}}.} \] As explained in the introduction, we aim to consider two-sided kernels of unbounded support. It is important to remark that this type of {\color{Blue} kernel} presents some challenges to our analysis (see comments before (ref)).
For an arbitrary process $\{U_{t}\}_{t\geq{}0}$, a filtration $\mathbb{F}=(\mathcal{F}_{t})_{t \geq 0}$, and a given time span $\Delta_{n}>0$, we shall use the notation \[ U_{i}^{n}:=U_{i\Delta_{n}},\qquad \Delta_{i}^{n} U:=U_{i}^{n}-U_{i-1}^{n}, \qquad \mathcal{F}^n_{i} := \mathcal{F}_{i \Delta_n}. \] Stable convergence in law is denoted by \( \stackrel{st}{\longrightarrow}\). See (2.2.4) in JacodProtter for the definition of this type of convergence. As usual, $a_{n}\sim c_{n}$ means that $a_{n}/c_{n}\to{}1$ as $n\to\infty$.
Throughout, we assume that we sample the process $X$ in ((ref)) at discrete times $t_{i} :=t_{i, n} :=i \Delta_{n}$, where $\Delta_{n} :=T / n$ and $T\in(0,\infty)$ is a given fixed time horizon. To estimate the spot volatility $c_{t_i}$, we adopt a Nadaraya-Watson type of kernel estimator of the form:
where \(K_{{b}}(x):=K(x / b) / b\), $k_{n}\in\mathbb{N}$, and ${b_n}:={k}_{n} \Delta_n$ is the bandwidth of the kernel function. This is a variation of the estimator
first proposed in fan2008spot and kristensen2010nonparametric, and studied in recent papers such as FigLi, FigWu. The denominator \[ \bar{K}_{n}\left(t_i\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right), \] {\color{DR} in (ref) is known to} improve the performance (ref) for values of $\tau=t_i$ near the edge of the estimation interval $[0,T]$. The variation (ref) was also considered in kristensen2010nonparametric and FigLi {\color{DR} when $t_i$ is {replaced} with a fix time $\tau\in(0,T)$}. In that case, $\bar{K}_{n}\left(\tau\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-\tau\right)\to1$ and the asymptotic behavior of ${\tilde{c}^{n,X}_{\tau}}$ and \[ {\hat{c}^{n,X}_{\tau}} :=\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-\tau\right)\left(\Delta_{j}^{n} X\right)\left(\Delta_{j}^{n} X\right)^*/\bar{K}_{n}\left(\tau\right) \] are equivalent. However, for our estimator (see (ref) below), we need to consider $t_i$ for all $i=1,\dots, n$ and, thus, it is not direct that we can simply omit $\bar{K}_{n}\left(t_i\right)$ (in fact, its presence helps with the analysis).
To handle jumps in the process $X$, we also consider a truncated version {\color{DR} of (ref):}
with a suitable sequence of truncation levels $v_n\in(0, \infty]$. {\color{DR} More specifically, we consider two types of conditions on $v_n$. In the case that $X$ is continuous, we take
while if $X$ is discontinuous, we take
where $\ell\geq{}4$ is a positive integer that will be specified below. Note that when $\alpha=\infty$ in (ref), we recover the non-truncated version ((ref)) and, in that case, the value of $\varpi$ is irrelevant.} For simplicity, we {\color{DR} often omit the dependence on $X$ and $v_n$ on (ref) and (ref) and simply write $\hat{c}_{i}^{n}$}.
{Our estimation target is the integrated volatility functional, defined as
for a continuous function $g:\mathcal{M}_d^+\to\mathbb{R}$. As proposed in jacod2013quarticity, a natural estimator for this integral is given by {\color{DR} $\Delta_{n} \sum_{i=1}^{\left[t / \Delta_{n}\right]} g\left(\hat{c}_{i}^{n}\right)$, which can be interpreted as a Riemann sum approximation of (ref).} However, as we shall see, it will be easier to derive the asymptotic properties of the estimator
Furthermore, we can see (ref) as a weighted average version of $\Delta_{n} \sum_{i=1}^{\left[t / \Delta_{n}\right]} g\left(\hat{c}_{i}^{n}\right)$, which gives less weight to $g\left(\hat{c}_{i}^{n}\right)$ for values of $t_i$ near $0$ and $T$, where the estimator $\hat{c}_{i}^{n}$ is less accurate due to edge effects.}
{Our first result characterizes the asymptotic behavior of the estimation error of $V(g)^{n}_{t}$} when the bandwidth converges to $0$ at {an} "optimal" rate {\color{DR} (c.f. Theorem (ref) below and the remarks before).}
To make this CLT “feasible” {\color{Blue} in practical work}, the bias terms need to be estimated. {For $A^1$, we use that
For $A^2$, we can apply Theorem (ref) to the functional $$ h(x):=\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g(x)\left(x^{pu}x^{qv}+x^{pv}x^{qu}\right), $$ for any $x\in \mathbb{R}^{d\times d}$, to get
The term $A^3$ is the hardest to estimate since it involves the {\color{Blue} `volvol'} $\tilde{c}$. In the case of a uniform {\color{DR} right-sided} kernel, {\color{DR} $K(x)={\bf 1}_{[0,1)}(\chi)$}, jacod2015estimation proposed the estimator
with $\hat{c}$ given by ((ref)) and showed that ${\color{DR} \widehat{A}^{n, 3}}$ converges to a linear combination of the bias terms $A^{2}$ and $A^3$. The following theorem shows the corresponding result for a general kernel function. The proof of Theorem (ref) is {given} in Appendix (ref).}
Under the Assumption ((ref)), based on ((ref)), ((ref)), and ((ref)), we {\color{DR} then} propose {a bias-corrected} estimator of the form:
where $$ C_{K,1}:=\frac{\int_{-\infty}^{\infty} L(z)^2 d z+\left(\int_{-\infty}^{0}-\int_{0}^{\infty}\right) L(z) d z}{{\int_{-\infty}^{\infty} L(z)(L(z)-L(z-1))d z+\frac{1}{2}-\int_{-\infty}^{\infty}K(z)(|z|\wedge{}1)dz}}{}, $$ and $$ C_{K,2}:={\int^{\infty}_{-\infty} K(z)(K(z)-K(z-1))d zC_{K,1}}. $$ Next, we show that indeed ${\widetilde{V}(g)^{n}_T}$ enjoys a centered CLT. Its simple proof is given in Section (ref).
As mentioned above, the bias term $A^3 $ contains the `volvol' $\tilde{c} $ and the estimator of $A^3$ in ((ref)) could introduce extra variance to the {asymptotically} unbiased estimator ((ref)). {An alternative} way to eliminate {the} bias is {undersmoothing, i.e., through the selection of a bandwidth sequence $b_n$ converging to $0$ at a faster rate than $\sqrt{\Delta_n}$ ($b_n \ll \sqrt{\Delta_n}$) or, equivalently, picking $k_n$ such that $\sqrt{\Delta_n}k_n \stackrel{n \to \infty}{\longrightarrow} 0$ (meaning that $\theta = 0$ in ((ref)))}. This is the approach put forward in jacod2013quarticity. {In that case, it is expected that the bias terms $A^1$ and $A^3$ will vanish and the bias term $A^2 $ will dominate}. {{\color{DR} Following the same idea,} we can devise an asymptotically unbiased {\color{DR} estimator as follows}:
The following result establishes the asymptotic behavior of $\bar{V}(g)^n_T $. {Its proof is {given} in Appendix (ref).}
In this section we analyze the performance of {\color{DR} our} kernel-based estimators. To isolate the effect of the bandwidth on the estimators' performance, we only consider continuous processes (i.e., $\delta\equiv 0$ in (ref) and, thus, $X\equiv X'$) and focus on the untruncated estimator (ref) and the corresponding estimators $V(g)^{n}_{t}$ of (ref) and $\widetilde{V}(g)^{n}_T$ of (ref).
We consider the data generating model in li2019efficient:
with $\mathbb{E}\left[d W_t d B_t\right] = -0.75 dt $. Throughout, we use the function $g(c) = c^2$, which corresponds to the integrated quarticity{,} and also the function $g(c) = \log(c)$. We simulate data for $T= 5$ days and $T = 21$ days, and assume the process $X$ is observed once {\color{Red} every 1 minute or every 5 minutes, with 6.5 trading hours per day, for all the results of Subsection (ref), but not for Subsection (ref), where we assume a frequency of 1-second}.
We first show the finite sample behavior of the estimator is consistent with the Central Limit Theorem of Theorem (ref). Under the notation of Theorem (ref), we choose $\color{DR}{ \theta = 0.2}$, $g(x) = x^2 $, and estimate the functional volatility using the exponential kernel $K(x) = \frac{1}{2}e^{-|x|}$. The choice of $\theta$ was obtained from the formula $b^*=\theta^*\sqrt{\Delta_n}$, where $b^*$ was chosen using the iterative method in FigLi. The sample frequency of the data is 1 second. Based on 5,000 simulated paths from the model ((ref)), we {calculate the standardized estimation error for each path $i=1,\dots, 5000$} according to Theorem (ref):
where $c^i_s = (\sigma^i_s)^2 $, $\{\hat{c}^{n,i}_j\}_j $ are the kernel spot volatility estimates of the $i$th simulated path, and the variance ${\rm Var}(Z)^{i}=2 \int_0^T(g''(c^i_s)c^i_s)^2ds$ and the biases $A^{\ell,i}$ ($\ell=1,2,3$), as defined in (ref), are computed via Riemann sum approximations using the true simulated volatility $\{c^{n,i}_j\}_{j=0,\dots,n}$ and volvol observations $\{\tilde{c}^{n,i}_j\}_{j=0,\dots,n}$. We then compare the histogram of the standardized estimates $z_i$ with a standard normal distribution (see Figure (ref)). As shown therein, the distribution of the estimation error is consistent with the asymptotic result in Theorem (ref).\%, though the convergence rate may be slow.\\
Similarly, we verify the behavior of the bias-corrected estimator proposed in Corollary (ref). With the same value of $\theta$, $g(x)$, exponential kernel, and simulated data, the standardized bias-corrected estimator takes the form:
The corresponding histogram is shown in Figure (ref), which again suggest the CLT stated in Corollary (ref) is valid.
We now compare our bias-corrected estimator $\widetilde{V}(g)_T^{n}$ to the Jackknife estimator of li2019efficient. For the Jackknife method, we consider the two-scale estimator $TS_n$, its boundary-adjusted version $TS_n'${,} and the multiscale estimator $MS_n$ as defined in li2019efficient, Section 3.1. The parameters for the Jackknife estimators are chosen according to li2019efficient. {\color{DR} By running numerous simulations, we determine that the values chosen by li2019efficient broadly optimize their estimator's performance}. For the multiscale estimator $MS_n$, $\left(\psi_1, \psi_2, \psi_3\right)=(-2.5,8,-4.5)$, and $\left(k_{1, n}, k_{2, n}, k_{3, n}\right)=(15,30,45)$ for data sampled every 5 minutes, or $(40,80,120)$ for data sampled every minute. For the two-scale estimator, $\left(\psi_1, \psi_2\right)=(-1,2)$ along with the same $\left(k_{1, n}, k_{2, n}\right)$ as above. For our kernel estimator, we use 3 kernels: a two-sided exponential kernel $K(x)=\frac{1}{2}e^{-|x|}$, a two-sided uniform kernel $K(x)=\frac{1}{2}{\bf 1}_{[-1,1]}(x)$, and a right-sided uniform kernel $K(x)={\bf 1}_{[0,1]}(x)$. We analyze the performance of the simple biased estimator $V(g)^n_T$ in ((ref)) and the bias-corrected estimator $\widetilde{V}(g)^{n}_T$ in ((ref)). We also consider the simple biased estimator $V(g)^{n,{JR}}_T$ and the bias-corrected estimator $\tilde{V}(g)^{n,{JR}}_T$ of (ref) and (ref), respectively, which were proposed in jacod2015estimation.
The bandwidths of all the kernel estimators are chosen using the iterative method in FigLi, Section 5. We iterate the method only once. For a two-sided kernel, the integrated volatility of volatility used when updating the bandwidth is estimated by the Two-time Scale Realized Volatility (TSRV) estimator proposed in FigLi. For the right-sided uniform kernel, the TSRV estimator no longer applies. We then use the realized variance estimator on the estimated volatility path as a replacement. As finite-sample corrections, in ((ref)), at {each} $t_i$, we replace $\int_{-\infty}^{\infty}K^2(u) du $ with $ \sum_{j=1}^n K^2_h(t_i - t_j) \Delta_n h$, and replace $\int_{-\infty}^{\infty}L(s) ds $ and $\int_{-\infty}^{\infty}L^2(s)ds $ with $ \frac{\Delta_n}{b}\sum_{j=1}^n L\left(\frac{t_{j-1}-t_i}{b} \right)$ and $\frac{\Delta_n}{b}\sum_{j=1}^n L^2\left(\frac{t_j-t_i}{b} \right)$, respectively, where $ L\left(\frac{t_j - t_i}{b} \right) = -\sum_{l=1}^{j}K_{{b}}(t_{l-1}-t_i) $ if $j \leq i$ and $\sum_{l=j+1}^{n}K_{{b}}(t_{l-1}-t_i)$ if $ j > i$.
In Table (ref), we report the relative bias and relative root-mean-square-errors in percentage unit for the considered estimators. Our results of the Jackknife estimators are consistent with the ones in li2019efficient, Table 1 therein. As shown in Table (ref), the kernel estimators {\color{DR} exhibit a similar performance to} the Jackknife estimators. However, the Jackknife estimators have more tuning parameters, while ours requires only one tuning parameter (namely, $\theta$). As we will show in Section (ref), the {\color{Red} most important} advantage of our estimators {\color{Red} lies} on their stability relative to the choice of the bandwidth. That is, our estimators seem to exhibit satisfactory performance for a large range of the tuning parameters, while other estimators are more sensitive to choosing `good' values for their tuning parameters.
{\color{DR} Returning to} the results of Table (ref),} we also observe that in some cases the kernel estimators have smaller relative root mean square error (RMSE) compared with the Jackknife estimators, especially when the time span $T$ is short ($T = 5$ days). Intuitively, the superior performance of the spot volatility estimators with exponential kernel (shown in FigLi, Section 7) {\color{DR} is expected to be more evident} when the time span is relatively short, e.g. 5 days. For reference, we include the tables for the function $g(c) = log(c)$ in Table (ref). These tables confirm our previous conclusion. It is worth noting that in general, the bias-corrected kernel estimator $\widetilde{V}^{n}_T$ has {\color{DR} significantly} smaller RMSE compared with the simple biased estimator $V(g)^{n}_T$, which also happens among the Jacod & Rosenbaum’s estimators, $V(g)^{n,{JR}}_T$ and $\widetilde{V}^{n,JR}_T$.
In this subsection, we investigate how sensitive the estimators are relative to the bandwidth parameter. For the Jackknife estimators, we focus on the boundary adjusted multi-scale estimator $\text{MS}$ (which, as shown in Tables (ref) and (ref), typically has the best performance among all {\color{Red} the considered} Jackknife estimators), and for the kernel estimators, we {\color{DR} first consider an} exponential kernel. We {\color{DR} try} a sequence of bandwidths $k_n$ for the kernel estimator, and a sequence of triplets {$(k_{1,n}, k_{2,n}, k_{3,n}) = ([\frac{k_n}{2}], k_n, [\frac{3k_n}{2}])$} and $(\psi_1, \psi_2, \psi_3) = (-2.5,8,-4.5)$ for $MS$, which include the parameters chosen in li2019efficient {\color{DR} when $k_n=30$ (see Table 1 therein)}. The results are shown in the Figures (ref)-(ref), where $\text{S0 exp}$ denotes the simple biased exponential kernel estimator in ((ref)) and $\text{S3 exp}$ denotes the bias-corrected estimators with exponential kernel in ((ref)). {\color{DR} We can conclude that the estimator $\text{S3 exp}$ exhibits the most stable or consistent performance for} either 5 or 21 days data with 1 or 5 minutes sampling frequency, and both $g(c)=c^2$ or $\log (c)$. {\color{DR} In other words, the estimator's values are {\color{DG} in general} relatively invariant to changes in the bandwidth, {\color{DR} especially} when working with low frequency and $g(c)=c^2$. Under those conditions, the others estimators require more careful tuning of the bandwidth in order to show good performance. We can also observe that the bias-corrected exponential kernel estimator almost always outperforms the boundary-adjusted multi-scale Jackknife estimator $\text{MS}$ and the simple biased exponential kernel estimator $\text{S0 exp}$.}
{\color{DR} Finally,} in Figure (ref), we analyze the bandwidth stability of three estimators: the bias-corrected estimators (ref) with exponential and right-sided uniform kernels and the bias-corrected estimator (ref) proposed in jacod2015estimation. Again, since we are not introducing jumps in the process $X$, the estimator $\hat{c}_i^{n}$ used for (ref) and the estimator $\hat{c}_i^{n,{JR}}$ used for (ref) do not have any truncation (i.e., we fix $v_n=\infty$ in (ref) and (ref)). Figure (ref) shows that our proposed estimator $\widetilde{V}^{n}_T$ with either exponential or right-sided uniform kernel almost always outperforms $\widetilde{V}^{n, JR}_T$. The exponential kernel appears to be less sensitive to the choice of $k_n$ than the right-sided uniform kernel. The estimator $\widetilde{V}^{n, JR}_T$ also seems to require a careful tuning of the bandwidth in order to achieve a good performance.