EconBase
← Back to paper

Estimation of Integrated Volatility Functionals with Kernel Spot Volatility Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,970 characters · 7 sections · 66 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimation of Integrated Volatility Functionals with Kernel Spot Volatility Estimators

abstractFor a multidimensional It\^o semimartingale, we consider the problem of estimating integrated volatility functionals. jacod2013quarticity studied a plug-in type of estimator based on a Riemann sum approximation of the integrated functional and a spot volatility estimator with a forward uniform kernel. Motivated by recent results that show that spot volatility estimators with general {two-sided} kernels of unbounded support are more accurate, in this paper an estimator using a general kernel spot volatility estimator as the plug-in is considered. A biased central limit theorem for estimating the integrated functional is established with an optimal convergence rate. Central limit theorems for properly de-biased estimators are also obtained both at the optimal convergence regime for the bandwidth and when applying undersmoothing. Our results show that one can significantly reduce the estimator's bias by adopting a general kernel instead of the standard uniform kernel. Our proposed bias-corrected estimators are found to maintain remarkable robustness against bandwidth selection in a variety of sampling frequencies and functions.

Introduction

In this work, we are concerned with the estimation of integrated volatility functionals of the form $V_{T}(g):=\int_{0}^{T}g(c_s)ds$, where $g$ is a smooth function and $c_t:= \sigma_t\sigma_t^*$ is the spot {\color{Blue} variance-covariance} process of a $d$-dimensional It\^o semimartingale given by

equation[equation omitted — 55 chars of source]

Here, $\{W_{t}\}_{t\geq{}0}$ is $d$-dimensional standard Brownian Motion (BM) and $\{J_{t}\}_{t\geq{}0}$ is a pure-jump process of bounded variation. Such functionals have many applications in financial econometrics. In the simplest case, when $d=1$ and $g(x) =x$, this is the well-studied integrated volatility (IV) or variance $IV_T=\int_0^{T}\sigma^2_sds$, which measures the overall variability or {\color{Blue} uncertainty} latent in the process $X$ during $[0,T]$. Again, when $d=1$ and $g(x) = x^2$, $V_T(g)$ corresponds to the integrated quarticity, which appears in the limiting distribution of the realized variance estimator $\widehat{IV}_T=\sum_{i=1}^n (\Delta_i^nX)^2$ and, hence, it is crucial to construct feasible confidence intervals for IV. Another function appearing in the literature is $g(x)=\log(x)$, when researching the statistical features of a factor $Y$ driving the volatility in a model where $\sigma$ is assumed to be the exponential of $Y$. Volatility functionals have also been used in principal component analysis by ait2019principal and in estimating integrated moments for option pricing purposes by li2016generalized.

One of the fundamental methods to estimate $V_{T}(g)=\int_{0}^{T}g(c_s)ds$ consists of approximating the integral by a Riemman sum and {\color{Blue} plugging} in a suitable spot volatility estimator $\hat{c}$ in place of $c$:

equation[equation omitted — 118 chars of source]

This method was pioneered by jacod2013quarticity, where, as the spot volatility estimator, a local rolling window estimator with uniform weights was used:

equation[equation omitted — 139 chars of source]

Two interesting facts emerged from their theory. {\color{Red} For the ease of exposition, we focus on the $d=1$-case in the introduction, though our main results are stated for general dimension $d\geq 1$.} First of all, the rate of convergence of the estimation error is unexpectedly of order $\Delta_n^{\frac{1}{2}}$ even though $\hat{c}_{i}^{n}$ can only attain the rate $\Delta_n^{\frac{1}{4}}$\footnote{{\color{Red} Intuitively, by {\color{Blue} Taylor} expansion of the estimator $g(\hat{c})$, we can write $\Delta_n\sum_{i=1}^n g(\hat{c}_i^n) - \Delta_n\sum_{i=1}^{n}g(c_{t_{i-1}})= \Delta_n\sum_{i=1}^{n}g'(c_{t_{i-1}}) \left(\hat{c}_{i}^n - c_{t_{i-1}}\right) + \frac{1}{2}\Delta_n\sum_{i=1}^{n}g''(c_{t_{i-1}}) \left(\hat{c}_{i}^n - c_{t_{i-1}}\right)^2+\text{higher order terms}$. The estimation errors in the first term of the right-hand side are “averaged out" so that the convergence rate changes from $n^{-1/4}$ to $n^{-1/2}$.}}. Secondly, when $k_n \sim \theta/\sqrt{\Delta_n} $ for some $\theta\in(0,\infty)$, which yields the best convergence rate for the spot volatility estimator $\hat{c}_{i}^{n}$, $\frac{1}{\sqrt{\Delta_n}}\left( \widehat{V}_{T}(g) - V_{T}(g) \right) $ converges to a mixed Gaussian distribution with four bias terms:

enumerate• A term of the form $-\frac{\theta}{2}\left(g(c_0) + g(c_T) \right) $ due to border effects; • A {\color{Blue} second} term of the form $\frac{1}{\theta} \int_0^T g^{\prime \prime}(c_s) c_s^2ds $ due to the nonlinearity of $g$; • A third term of the form $-\frac{\theta}{12} \int_0^T g^{\prime \prime}(\tilde{c}_s) $, where $\tilde{c} $ is the volatility of volatility, due to target error $c_i - \bar{c}_{i}^n$, where $\bar{c}_{i}^{n}=\frac{1}{k_{n} \Delta_{n}} \sum_{j=0}^{k_{n}-1}c_{i+j}$; • A fourth term involving the jump component of the volatility process $\theta \sum_{s\le t} G(c_{s-}, c_s) $ for a suitable function of $G$ depending on g. This error disappears when $c$ is continuous.

In principle, the bias terms can be estimated and eliminated from $\widehat{V}_{T}(g)$ to obtain a feasible CLT with centered Gaussian limiting distribution. This strategy was explored by jacod2015estimation. An alternative solution, however, was adopted in jacod2013quarticity. Specifically, the window length $k_n$ was made to converge to $\infty$ at a slower rate (under-smoothing) so that $k_n \sqrt{\Delta_n}\to{}0$. This corresponds to the case $\theta= 0 $ and, hence, the first, third, and fourth bias terms would vanish, while the second term becomes dominant. To keep the convergence rate $\sqrt{\Delta_n} $ of the estimator $\widehat{V}_T(g) $, this second term needs to be estimated and the estimator needs to be de-biased. The resulting estimator takes the form

equation[equation omitted — 151 chars of source]

which indeed achieves the rate $\Delta_n^{\frac{1}{2}}$ with semi-parametrically optimal asymptotic variance. {\color{Red} A similar approach to deal with {\color{Blue} edge} effects and biases was also considered in kristensen2010nonparametric, who also proposed kernel estimators and obtained a central limit theorem (though the model considered therein lacks leverage effects).} Alternatively, as mentioned before, in the optimal case $\theta \in (0,\infty)$ and employing forward uniform kernels, estimators for all the bias terms were proposed in jacod2015estimation, where an unbiased estimator with centered limiting distribution was developed.

li2019efficient took a different approach towards bias correction. They proposed a type of `jackknife' estimators, which are formed as a linear combination of a few uncorrected estimators of the form (ref) associated with different windows $k_n$. Since these estimators have similar bias terms but different coefficients depending on the {\color{Blue} window} sizes, the biases can be eliminated by a proper linear combination. Both the `two-scale' jackknife (i.e., when combining only two uncorrected {\color{Blue} estimators}) and jacod2013quarticity's estimator use under-smoothing, so that the biases due to boundary effects, volatility of volatility, and volatility jumps are rendered asymptotically negligible. One drawback of this approach is that one needs to determine the degree of undersmoothing. In other words, if, for instance, we take $k_n \sim \beta\Delta_n^{\alpha}$ with $\alpha>-1/2$, then one needs to tune both parameters $\beta$ and $\alpha$, which is nontrivial. In contrast, a multiscale jackknife estimator, formed as a linear combination of three (or more) estimators, can in principle cancel all the biases regardless of the speed of convergence of the window size. However, again, this involves more tuning parameters that need to be carefully calibrated.

Other related work on this topic includes mykland2009inference, who proposed using the same method as in jacod2013quarticity, but {\color{Red} keeping $k_n = k$ constant}; for the quarticity or more generally for estimating $\int_0^T g(c_s) ds $ with a power functional $g(x) = x^r$, they obtained a CLT with rate $1/\sqrt{\Delta_n}$ and an asymptotic variance that is not efficient, {\color{Red} but can arbitrarily approach the optimal variance} when $k$ is large. {\color{Red} Glotter established a Locally Asymptotically Mixed Normality (LAMN) condition for the estimation of integrated volatility functionals in a continuous diffusion process where the volatility is driven by an independent It\^o process. The result therein shows that the optimal rate of convergence is $n^{-1/2}$. Another related result in the same direction is Renault}. {\color{Red} More recently, LiLiu2021 generalized jacod2013quarticity 's result to allow a long-memory component such as {\color{Blue} an} fBM with Hurst parameter $H>1/2$. They also established a semiparametric efficiency lower bound, generalizing Renault's results.} chen2018inference developed an estimator for $\widehat{V}_{T}(g)$ in the presence of microstructure noise, based on a forward finite difference approximation of the standard pre-averaging estimator of the integrated volatility.

Our work aims to apply a general kernel spot volatility estimator to the plug-in estimator of integrated volatility functionals. The general kernel estimator (fan2008spot, kristensen2010nonparametric) is defined as

align[align omitted — 143 chars of source]

where $K_{{b}}(x):=K(x/{b})/{b}$ {\color{Red} for some bandwidth $b>0$}, and $k_n$ is an appropriate sequence that controls the bandwidth $b_n=k_n\Delta_n$ and that needs to be calibrated. Of course, one can see (ref) as a particular case of (ref) with $K(x)={\bf 1}_{[0,1]}(x)$. {\color{Red} kristensen2010nonparametric developed an asymptotic theory for (ref) when $t_i$ is a fixed time $\tau\in (0,T)$ and showed asymptotic normality to the spot variance $c_{\tau}$ under certain path-wise H\"older continuity and non-leverage conditions on $c$. See also fan2008spot and mancini2015estimation for other related results. All these results applied undersmoothing, which allowed them to neglect the “target error" coming from approximating the spot volatility by a kernel weighted integrated volatility.}

In this paper, we consider the following version:

align[align omitted — 175 chars of source]

with $\bar{K}_{n}\left(t_i\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right)$, which is known to be more robust than (ref) {\color{Blue} against} edge effects {\color{Red} (see also kristensen2010nonparametric for other edge-robust alternatives)}. As a result, we can, in principle, consider the whole range $i=1,\dots, n$ in (ref). Furthermore, we find it more convenient for our proofs to consider the estimator

equation[equation omitted — 127 chars of source]

The coefficient $\bar{K}_n(t_i)$ attenuates the contribution of $g\left(\hat{c}_{i}^{n}\right)$ for $t_i$ near $0$ and $T$ and, thus, serves as an additional edge effect correction. {\color{Red} kristensen2010nonparametric also considered a kernel-based estimator similar to (ref), but applied an edge effect correction, similar to that in jacod2015estimation, by eliminating the terms corresponding to $i$'s close to $0$ and $n$, and did not consider the adjustments $\bar{K}_n(t_i)$. Their asymptotic theory also only considered the undersmoothing asymptotic regime mentioned above.}

The motivation for considering general kernels comes from FigLi and FigWu, where it has been shown, both theoretically and by Monte Carlo experiments, that a general two-sided kernel with unbounded support could significantly reduce the variance of the spot volatility estimator $\hat{c}$. More specifically, FigLi studied the infill asymptotic behavior of the mean-square error (MSE) of (ref) {for continuous It\^o semimartingales} and showed that, when setting $k_n$ `optimally' (i.e., in order to minimize the leading order terms of the MSE), the double exponential kernel $K(x)=.5 e^{-|x|}$ minimizes the resulting MSE over all kernels. This result was established under the absence of leverage effects. FigWu extended this result, by first proving a CLT for a general {two-sided} kernel, with optimal convergence rate and in the presence of jumps and leverage {effects}. It was then shown that exponential kernels minimize the asymptotic variance of the CLT. It is worth noting that, even when applying undersmoothing, the asymptotic variance could be reduced when choosing a kernel of unbounded support. Specifically, when applying undersmoothing, the asymptotic variance of the spot volatility estimator is proportional to $\|K\|^2=\int_{-\infty}^{\infty} K^2(x) dx$. When $K$ is constrained to have support on $[0,1]$, the optimal kernel is indeed the uniform kernel $K_{unif}(x)={\bf 1}_{[0,1]}(x)$ with $\|K_{unif}\|^2=1$\footnote{{\color{Red} Indeed, by Jensen's inequality, $K_{unif}(x)={\bf 1}_{[0,1]}(x)$ satisfies $\int K^2_{unif}(x)dx=1=\left(\int K(x)dx\right)^2\leq \int K^2(x)dx$, which shows that a uniform kernel $K_{unif}(x)$ minimizes the asymptotic variance when applying undersmoothing.}}. However, picking the kernels $.5e^{-|x|}$ {\color{Red} and $e^{-x}{\bf 1}_{x>0}$ will result in estimators whose asymptotic variances are respectively a quarter and a half of} that obtained by a uniform kernel $K_{unif}(x)={\bf 1}_{[0,1]}(x)$. {\color{Red} Similar comments apply when considering only backward looking kernels such as ${\bf 1}_{[-1,0)}(x)$ and $e^{-|x|}{\bf 1}_{x<0}$.}

A natural question is whether the estimation accuracy gained when estimating the spot volatility by a general kernel transfers into gains in estimating integrated functionals. Intuitively, by the Taylor expansion of the estimator $g(\hat{c})$, we can write $g(\hat{c}) - g(c)= g'(c) \left(\hat{c} - c\right) + \frac{1}{2}g''(c)\left(\hat{c} - c\right)^2 + O_p\left(|\hat{c} - c|^3 \right)$. After taking expectations, the variance of the spot volatility estimator, which appears in the second term, becomes a bias term of the integrated volatility estimator. Therefore, the bias of the estimator $\widehat{V}_{T}(g)$ could in principle be reduced by adopting a general kernel.

As it turns out, the CLT of (ref) again exhibits several biases, which depend on a rather intricate and nontrivial manner on the kernel. The bias terms compared to the uniform kernel case involves a highly nontrivial transformation of the kernel function. The second bias term, involving $\frac{1}{\theta} \int_0^T g^{\prime \prime}(c_s) c_s^2ds$, contains the $L_2$ norm of the kernel and, thus, can be reduced because certain kernels of unbounded support {\color{Blue} possess} smaller $L_2$ norm. We then propose bias-corrected estimator similar to those in jacod2015estimation, and derive the corresponding CLT. We show by Monte Carlo experiments that our general kernel estimator typically has superior performance compared with the jackknife estimators of li2019efficient. Of course, our estimator is simpler as it does not involve as many tuning parameters, which are hard to tune-up. In fact, our Monte Carlo experiments suggest our estimators are significantly more stable in the sense that they exhibit better performance for most values of the bandwidth, while the jackknife estimator requires accurate tuning of the `bandwidth' to match our estimator's performance. We also observe that even without bias corrections, under certain conditions, the plug-in estimator with general kernel spot volatility estimator has good performance, which could be a result of the increased accuracy of {\color{Blue} the} spot volatility estimator and/or the reduction in bias attained by using a general kernel.

While working on the final stages of writing the present manuscript, we became aware of a recent work by benvenutifunctionals, who also studied a type of plug-in estimator with a general two-sided kernel spot volatility estimator. However, unlike our work, only a suboptimal bandwidth asymptotic regime (i.e., assuming under-smoothing) was considered and only consistency was established.

The rest of this paper is organized as follows. Section (ref) introduces the framework, assumptions, and main results. Section (ref) illustrates the performance of our method via Monte Carlo simulations and compares it to the jackknife estimator of li2019efficient. The proofs are deferred to an appendix section.

Notation: We shall use the following notation:

itemize• The set of (nonnegative definite) $d\times{}d$ real matrices $x=[x^{\ell m}]_{\ell,m=1}^{d}$ is denoted ($\mathcal{M}_+^d$) $\mathbb{R}^{d\times d}$; • For a $C^1$ function $g:\mathbb{R}^{d\times d}\to\mathbb{R}$, its gradient {\color{Red} $\nabla g(x)$} is the $d\times d$ matrix with $(\ell,m)$ entry $\partial_{\ell m}g(x):=\frac{\partial g(x)}{\partial x^{\ell m}}$; • For $x,y\in\mathbb{R}^{d\times d}$, its inner product is defined as $x\cdot y=\sum_{\ell,m}x^{\ell m}y^{\ell m}$. Also, $x^{*}$ is the transpose of $x$ and $\|x\|=\sqrt{\sum_{\ell,m=1}^n (x^{\ell m})^2}$ is the Euclidian norm.

The Setting, Estimator, and Main Results

Throughout, we consider a $d$-dimensional It\^o semimartingale $X=[X^1,\dots,X^d]^*\in\mathbb{R}^{d\times{}1}$ of the form:

equation[equation omitted — 312 chars of source]

where all stochastic processes ($\mu :=\left\{\mu_{t}\right\}_{t \geq 0}$, $\sigma :=\left\{\sigma_{t}\right\}_{t \geq 0}$, $W :=\left\{W_{t}\right\}_{t \geq 0}$, $\delta:=\left\{\delta(t, z):t\geq 0, z\in E\right\}$, $\mathfrak{p}:=\{\mathfrak{p}(B):B\in\mathcal{B}(\mathbb{R}_{+}\times E)\}$, etc.) are defined on a complete filtered probability space $\left(\Omega, \mathcal{F}, \mathbb{F}, \mathbb{P}\right)$ with filtration $ \mathbb{F}=(\mathcal{F}_{t})_{t \geq 0}$. Here, {\color{DG} $W=[W^1,\dots,W^d]^*$} is a $d$-dimensional standard Brownian Motion (BM) adapted to the filtration \(\mathbb{F}\), $\delta$ is a predictable $\mathbb{R}^d$-valued function on $\Omega \times \mathbb{R}_{+} \times E$, and $\mathfrak{p}$ is a Poisson random measure on $\mathbb{R}_{+} \times E$ for some arbitrary Polish space $E$ with compensator $\mathfrak{q}(\mathrm{d} u, \mathrm{~d} x)=\mathrm{d} u \otimes \lambda(\mathrm{d} x)$, where $\lambda$ is a $\sigma-$ finite measure on $E$ having no atom. We also denote the spot variance-covariance process as $c_t := \sigma_t \sigma_t^*$, which takes values in the set $\mathcal{M}_d^+$ of nonnegative definite {\color{DG} $d\times d$} {\color{DR} symmetric} matrices. {\color{Blue} As common} in the literature, $c$ is called the spot volatility of the process. We assume $c$ is also an It\^o semimartingale following the dynamics:

equation[equation omitted — 210 chars of source]

where {$W:=\{W_t\}_{t\geq0}$} is the same $d$-dimensional Brownian Motion driving the dynamics of $X$. {\color{Red} The assumption (ref) is common in the literature (see, e.g., the monographs JacodProtter,jacodaitsahalia), and was also assumed in the work of jacod2013quarticity{\color{Blue} )}.} {\color{DG} For each $1\leq \ell, m\leq d$, $\{\tilde{\mu}^{\ell m}_t\}_{t\geq{}0}$ is adapted locally bounded {\color{DR} and} $\{\tilde{\sigma}^{\ell m}_t\}_{t\geq{}0}$ is adapted c\`adl\`ag.}

We now state the main assumption on the process $X$:

assumptionThe process $X$ follows the dynamics ((ref)) with $c_t=\sigma_t\sigma_t^*$ satisfying ((ref)) and, for some $r \in[0,1)$, a measurable function $\Gamma_m: E \rightarrow \mathbb{R}_{+}$, a constant $C_m<\infty$, and a localizing sequence of stopping times $\left(\tau_m\right)_{m \geq 1}$ such that $\tau_m \rightarrow \infty$, we have $$ t \in\left[0, \tau_m\right] \Longrightarrow\left\{\begin{array}{l} \left|\mu_t\right|+\left|\sigma_t\right|+\left|\tilde{\mu}_t\right|+\left|\tilde{\sigma}_t\right| \leq C_m, \\ |\delta(t, z)| \wedge 1 \leq \Gamma_m(z), \quad \text { where } \int \Gamma_m(z)^r \lambda(d z)<\infty. \end{array}\right. $$

The parameter $r$ determines the jump activity of the process: the larger $r$ is, the more active or frequent are the small jumps of the process. When $r=0$, the process {\color{Blue} exhibits} {\color{DR} finitely} many jumps in any bounded time interval (in that case, we say that the jumps are of finite activity). When $r<1$, the jump component of the process is of bounded variation, which is a standard {\color{Blue} constraint} for the truncated realized quadratic variation to achieve efficiency (JacodProtter).

Next, we give the conditions on the kernel function $K$ {\color{DG} for which we need to define the function} \[ {\color{DG} L(t) := \int_{t}^{\infty} K(u) d u \mathbf{1}_{\{t>0\}}-\int_{-\infty}^{t} K(u) d u \mathbf{1}_{\{t \leq 0\}}.} \] As explained in the introduction, we aim to consider two-sided kernels of unbounded support. It is important to remark that this type of {\color{Blue} kernel} presents some challenges to our analysis (see comments before (ref)).

assumptionThe kernel function $K : \mathbb{R} \rightarrow \mathbb{R}$ is bounded such that \begin{enumerate} • $\int K(x) d x=1$; • $K$ is Lipschitz and piecewise $C^2$ on its support $(A,B)$, where $A < B$, $A \in [-\infty, 0]$, and $B \in [0, \infty]$; • (i) $\int|K(x) x| d x<\infty$; (ii) $K(x) |x|^{2+\epsilon} \rightarrow 0$, $K^{\prime}(x) |x|^{2+\epsilon} \rightarrow 0$, with $\epsilon>1/2 $ as $|x| \rightarrow \infty$ ; (iii) $\int |K^{(j)}(x)|dx < \infty$, for $j=0,1,2,3$; (iv) $\int_0^\infty\int_{w}^\infty |K'(x)|dxdw<\infty$; (v) $\int K^{2}(x)dx<\infty$; {\color{DG} (vi) $\int_{-\infty}^{\infty}|L(u)u|du<\infty$}; • (i) $V_{-\infty}^{\infty}(K^{(j)}):=\lim_{m\to\infty}V_{-m}^{m}(K^{(j)})<\infty$, where $V_{-m}^{m}(K^{(j)})$ is the total variation of $K^{(j)}$ on the interval $[-m,m]$, $j=0,1,2,3$; (ii) $V_{-\infty}^{\infty}(K^2)<\infty$; • $K^{\prime}$ and $K^{\prime\prime}$ exist piecewise on $(A,B)\subset\mathbb{R}$ and are absolutely continuous on {\color{Red} their} corresponding domain $\operatorname{dom}(K^{\prime})$ and $\operatorname{dom}(K^{\prime\prime})\subset(A,B)$, respectively; • $K^{\prime\prime\prime}$ exists piecewise on $(A,B)\subset\mathbb{R}$. \end{enumerate}
remark{\color{Red} Even though the list of assumptions above seems long, it is important to remark that all kernels commonly used in practice satisfy these conditions. The assumptions on the regularity of the high-order derivatives $K^{(j)}$ for $j=2,3$ are only used to show the following limit: \begin{equation} \sqrt{\Delta_{n}} \sum_{i=1}^n\left[\bar{K}_{n}\left(t_i\right)-1\right] g\left(c_i^n\right)\longrightarrow \theta \int_{-\infty}^{0}L(u)du g\left(c_{0}\right)-\theta \int_{0}^{\infty}L(u)du g\left(c_{T}\right), \end{equation} where recall that $\bar{K}_{n}\left(t_i\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right)$. This limit is easy to deal with in the case of a uniform kernel because, in that case, $\bar{K}_{n}\left(t_i\right)\equiv 1$ for almost all $i$ except if $i$ is close to $0$ or $n$.}

For an arbitrary process $\{U_{t}\}_{t\geq{}0}$, a filtration $\mathbb{F}=(\mathcal{F}_{t})_{t \geq 0}$, and a given time span $\Delta_{n}>0$, we shall use the notation \[ U_{i}^{n}:=U_{i\Delta_{n}},\qquad \Delta_{i}^{n} U:=U_{i}^{n}-U_{i-1}^{n}, \qquad \mathcal{F}^n_{i} := \mathcal{F}_{i \Delta_n}. \] Stable convergence in law is denoted by \( \stackrel{st}{\longrightarrow}\). See (2.2.4) in JacodProtter for the definition of this type of convergence. As usual, $a_{n}\sim c_{n}$ means that $a_{n}/c_{n}\to{}1$ as $n\to\infty$.

Throughout, we assume that we sample the process $X$ in ((ref)) at discrete times $t_{i} :=t_{i, n} :=i \Delta_{n}$, where $\Delta_{n} :=T / n$ and $T\in(0,\infty)$ is a given fixed time horizon. To estimate the spot volatility $c_{t_i}$, we adopt a Nadaraya-Watson type of kernel estimator of the form:

align[align omitted — 245 chars of source]

where \(K_{{b}}(x):=K(x / b) / b\), $k_{n}\in\mathbb{N}$, and ${b_n}:={k}_{n} \Delta_n$ is the bandwidth of the kernel function. This is a variation of the estimator

align[align omitted — 175 chars of source]

first proposed in fan2008spot and kristensen2010nonparametric, and studied in recent papers such as FigLi, FigWu. The denominator \[ \bar{K}_{n}\left(t_i\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right), \] {\color{DR} in (ref) is known to} improve the performance (ref) for values of $\tau=t_i$ near the edge of the estimation interval $[0,T]$. The variation (ref) was also considered in kristensen2010nonparametric and FigLi {\color{DR} when $t_i$ is {replaced} with a fix time $\tau\in(0,T)$}. In that case, $\bar{K}_{n}\left(\tau\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-\tau\right)\to1$ and the asymptotic behavior of ${\tilde{c}^{n,X}_{\tau}}$ and \[ {\hat{c}^{n,X}_{\tau}} :=\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-\tau\right)\left(\Delta_{j}^{n} X\right)\left(\Delta_{j}^{n} X\right)^*/\bar{K}_{n}\left(\tau\right) \] are equivalent. However, for our estimator (see (ref) below), we need to consider $t_i$ for all $i=1,\dots, n$ and, thus, it is not direct that we can simply omit $\bar{K}_{n}\left(t_i\right)$ (in fact, its presence helps with the analysis).

To handle jumps in the process $X$, we also consider a truncated version {\color{DR} of (ref):}

equation[equation omitted — 299 chars of source]

with a suitable sequence of truncation levels $v_n\in(0, \infty]$. {\color{DR} More specifically, we consider two types of conditions on $v_n$. In the case that $X$ is continuous, we take

equation[equation omitted — 171 chars of source]

while if $X$ is discontinuous, we take

equation[equation omitted — 216 chars of source]

where $\ell\geq{}4$ is a positive integer that will be specified below. Note that when $\alpha=\infty$ in (ref), we recover the non-truncated version ((ref)) and, in that case, the value of $\varpi$ is irrelevant.} For simplicity, we {\color{DR} often omit the dependence on $X$ and $v_n$ on (ref) and (ref) and simply write $\hat{c}_{i}^{n}$}.

{Our estimation target is the integrated volatility functional, defined as

equation[equation omitted — 79 chars of source]

for a continuous function $g:\mathcal{M}_d^+\to\mathbb{R}$. As proposed in jacod2013quarticity, a natural estimator for this integral is given by {\color{DR} $\Delta_{n} \sum_{i=1}^{\left[t / \Delta_{n}\right]} g\left(\hat{c}_{i}^{n}\right)$, which can be interpreted as a Riemann sum approximation of (ref).} However, as we shall see, it will be easier to derive the asymptotic properties of the estimator

equation[equation omitted — 167 chars of source]

Furthermore, we can see (ref) as a weighted average version of $\Delta_{n} \sum_{i=1}^{\left[t / \Delta_{n}\right]} g\left(\hat{c}_{i}^{n}\right)$, which gives less weight to $g\left(\hat{c}_{i}^{n}\right)$ for values of $t_i$ near $0$ and $T$, where the estimator $\hat{c}_{i}^{n}$ is less accurate due to edge effects.}

{Our first result characterizes the asymptotic behavior of the estimation error of $V(g)^{n}_{t}$} when the bandwidth converges to $0$ at {an} "optimal" rate {\color{DR} (c.f. Theorem (ref) below and the remarks before).}

theoremLet $g$ be a $C^{3}$ function such that, for all $x \in \mathbb{R}^{d \times d}$, \begin{equation} |g(x)| \leq C\left(1+\|x\|^{\color{DG} \ell}\right),\quad\left| \partial _{p_1 q_1,\dots,p_j q_j}g(x)\right| \leq C\left(1+\|x\|^{{\color{DG} \ell}-j}\right), \quad j=1,2,3{,} \end{equation} for $p_1, q_1,\dots,p_j, q_j\in\{1, \ldots, d\}$ and for some constants $C > 0$ and ${\color{DG} \ell} \ge 4$. Assume $k_n$ satisfies \begin{equation} k_{n} \sim \frac{\theta}{\sqrt{\Delta_{n}}} \qquad (i.e.,\;\; the bandwidth\ {{b}_n}:=k_n\Delta_n\sim \theta \sqrt{\Delta_n}), \end{equation} for some $\theta \in(0, \infty)$. Then, under ((ref)) and Assumptions (ref) and (ref), the following assertions hold true: \begin{itemize} • Suppose that {\color{DR} $X$ is continuous} (i.e., $\delta\equiv 0$ in ((ref))). Then, the estimator (ref) of the integrated volatility functional $V(g)_T$ with {\color{DR} $\hat{c}_{i}^{n}$ given by ((ref)), {under} the condition (ref)\footnote{In particular, by taking $\alpha=\infty$, we recover the untruncated estimator (ref).},} satisfies the following stable convergence in law, as $n \rightarrow \infty$: \begin{equation} {\color{DR} \frac{1}{\sqrt{\Delta_{n}}}\Big(V(g)_T^{n}-V(g)_T\Big) \stackrel{st}{\longrightarrow} A^{1} + A^{2}+A^{3}+Z}, \end{equation} where $Z$ is {\color{Blue} an} r.v. defined on an extension $\left(\widetilde{\Omega}, \widetilde{\mathcal{F}},(\widetilde{\mathcal{F}}_{t})_{t \geq 0}, \widetilde{\mathbb{P}}\right)$ of $\left(\Omega, \mathcal{F}, (\mathcal{F}_{t})_{t \geq 0}, \mathbb{P}\right)$, which, conditionally on $\mathcal{F}$, is a centered Gaussian variable with variance \begin{equation} {\widetilde{\mathbb{E}}\left(Z^{2} \mid \mathcal{F}\right) =\sum_{j,k,l,m} \int_0^T\partial_{jk}g(c_s)\partial_{lm}g(c_s)(c_s^{jl}c_s^{km}+c_s^{jm}c_s^{kl})ds,} \end{equation} {and \begin{equation} \begin{split} A^{1}&:=-\theta \int_{0}^{\infty}L(u)dug\left(c_{0}\right)+\theta \int_{-\infty}^{0}L(u)dug\left(c_{T}\right), \\ A^{2}&:=\frac{1}{2\theta} \int_0^T \sum_{p,q,u,v=1}^{d}\partial^{2}_{p q, u v} g\left(c_s\right) \check{c}_s^{p q, u v} d s \int_{-\infty}^{\infty} K^2(u) d u,\\ A^{3}&:=\frac{\theta}{2} \int_0^T \sum_{p,q,u,v=1}^{d}\partial^{2}_{p q, u v} g\left(c_s\right)\tilde{c}_s^{p q, u v} d s\left[\int_{-\infty}^{\infty} L(u)^2 d u+\left(\int_{-\infty}^{0}-\int_{0}^{\infty}\right) L(u) d u\right]{\color{DR} ,} \end{split} \end{equation} where} $\check{c}_{s}^{pq,uv}:=c_{s}^{pu}c_{s}^{qv} +c_{s}^{pv}c_{s}^{qu}$ and $\tilde{c}^{pq,uv}_{s}:=\sum_{r=1}^d\tilde{\sigma}_s^{p q, r} \tilde{\sigma}_s^{u v, r}$. • {\color{DR} When $X$ is discontinuous, we {\color{DR} still} have ((ref)) for the estimator $V\left(g\right)^{n}_{T}$ with the truncated version ((ref)) provided that the stronger condition (ref) is satisfied}. \end{itemize}
remarkNote that the bias terms depend on the kernel in a rather intricate manner (especially, the term $A^3$). The {\color{Blue} signs} of the coefficients are always the same regardless of $K$. For instance, $\int_{-\infty}^{\infty} L(u)^2 d u+\left(\int_{-\infty}^{0}-\int_{0}^{\infty}\right) L(u) d u$ is negative because $|L(u)|\leq 1$ for all $u$, and $L(u)\leq{}0$ ($L(u)\geq 0$) for any $u<0$ ($u>0$). Thus, $A^3$ goes in {\color{Blue} the} opposite direction to $A^2$.
remarkIf we use the uniform {\color{DR} right-sided} kernel $K(x)={\bf 1}_{[0,1)}(x)$ in our estimators, {as} in jacod2013quarticity,jacod2015estimation, then \begin{equation} \bar{K}_n\left(t_{i}\right)=\Delta_{n}\sum_{j=1}^{n}{K_{b_n}}\left(t_{j-1}-t_{i}\right)=\Delta_{n}\sum_{j=i+1}^{{k_n}+i}\frac{1}{{b_n}}=1, \quad for \ 0\leq i\leq n-k_{n}. \end{equation} {\color{DR} In that case,} the estimator in jacod2015estimation, defined as \[ V\left(g\right)^{n,JR}_{T}:=\Delta_{n}\sum_{i=1}^{n-k_{n}-1}g\left(\hat{c}_{i}^{n,JR}\right), \] where \begin{equation} \hat{c}_i^{n,{JR}}:=\frac{1}{k_n \Delta_n} \sum_{j=0}^{k_n-1} \left(\Delta_{i+j}^n X\right)\left(\Delta_{i+j}^n X\right)^* \mathbbm{1}_{\left\{\left\|\Delta_{i+j}^n X\right\| \leq v_n\right\}}, \end{equation} can be written {\color{DR} in terms of our estimator (ref) as follows:} \begin{equation} V\left(g\right)^{n,JR}_{T}=V\left(g\right)^{n}_{T}-\Delta_{n}\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)+\Delta_n g\left(\hat{c}_{0}^{n}\right). \end{equation} {The formula (ref) holds because} ((ref)) ensures that $\hat{c}_{i}^{n}=\hat{c}_{i+1}^{n,JR}$, for $0\leq i\leq n-k_{n}$. {It is easy to see that, for {the uniform kernel $K(x)={\bf 1}_{[0,1)}(x)$,} $\Delta_{n}^{1/2}\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)\to \frac{\theta}{2}g\left(c_{T}\right)$ and the bias terms of Theorem (ref) reduce to \begin{equation} \begin{array}{l} A^{1}:=-\frac{\theta}{2}g\left(c_{0}\right){,} \\ A^{2}:=\frac{1}{2\theta} \int_0^T \sum_{p,q,u,v=1}^{d}\partial^{2}_{p q, u v} g\left(c_s\right) \check{c}_s^{p q, u v} d s,\\ A^{3}:=-\frac{\theta}{12} \int_0^T \sum_{p,q,u,v=1}^{d}\partial^{2}_{p q, u v} g\left(c_s\right)\tilde{c}_s^{p q, u v} d s. \end{array} \end{equation} Thus, Theorem (ref) implies that \begin{equation} \frac{1}{\sqrt{\Delta_{n}}}\left(V(g)_T^{n,JR}-V(g)_T\right) \stackrel{st}{\longrightarrow}\tilde{A}^{1} + A^{2}+A^{3}+Z, \end{equation} where $\tilde{A}^{1}=-\frac{\theta}{2}\left[g\left(c_{0}\right)+g\left(c_{T}\right)\right]$. We then recover the same result as jacod2015estimation.}

To make this CLT “feasible” {\color{Blue} in practical work}, the bias terms need to be estimated. {For $A^1$, we use that

equation[equation omitted — 224 chars of source]

For $A^2$, we can apply Theorem (ref) to the functional $$ h(x):=\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g(x)\left(x^{pu}x^{qv}+x^{pv}x^{qu}\right), $$ for any $x\in \mathbb{R}^{d\times d}$, to get

equation[equation omitted — 268 chars of source]

The term $A^3$ is the hardest to estimate since it involves the {\color{Blue} `volvol'} $\tilde{c}$. In the case of a uniform {\color{DR} right-sided} kernel, {\color{DR} $K(x)={\bf 1}_{[0,1)}(\chi)$}, jacod2015estimation proposed the estimator

equation[equation omitted — 276 chars of source]

with $\hat{c}$ given by ((ref)) and showed that ${\color{DR} \widehat{A}^{n, 3}}$ converges to a linear combination of the bias terms $A^{2}$ and $A^3$. The following theorem shows the corresponding result for a general kernel function. The proof of Theorem (ref) is {given} in Appendix (ref).}

theoremUnder the notations {\color{DR} and conditions} of Theorem (ref), we have, {as $n\to\infty$}, \begin{equation} \begin{aligned} {\color{DR} \widehat{A}^{n, 3}}\stackrel{\mathbb{P}}{\longrightarrow}&\,\frac{1}{4\theta} \int_0^T \sum_{j, k, l, m}\partial^2_{jk, lm} g\left(c_s\right) \check{c}_s^{jk, lm} d s {\int^{\infty}_{-\infty} K(z)(K(z)-K(z-1))d z}\\ &+\frac{\theta}{4} \int_0^T \sum_{j, k, l, m}\partial^2_{jk, lm} g\left(c_s\right) \tilde{c}_s^{jk, lm} d s\\ &\quad\times \Bigg({\int_{-\infty}^{\infty} L(z)(L(z)-L(z-1))d z+\frac{1}{2}-\int_{-\infty}^{\infty}K(z)(|z|\wedge1)dz\Bigg),} \end{aligned} \end{equation} when $\hat{c}$ is given by ((ref)) {\color{DR} with the condition (ref)}. Furthermore, {\color{DR} when $X$ is continuous (i.e., $\delta\equiv 0$ in ((ref))),} (ref) still holds {\color{DR} under the weaker condition (ref)} {(in particular, we can use the untruncated estimator (ref) by taking $\alpha=\infty$)}.

Under the Assumption ((ref)), based on ((ref)), ((ref)), and ((ref)), we {\color{DR} then} propose {a bias-corrected} estimator of the form:

equation[equation omitted — 582 chars of source]

where $$ C_{K,1}:=\frac{\int_{-\infty}^{\infty} L(z)^2 d z+\left(\int_{-\infty}^{0}-\int_{0}^{\infty}\right) L(z) d z}{{\int_{-\infty}^{\infty} L(z)(L(z)-L(z-1))d z+\frac{1}{2}-\int_{-\infty}^{\infty}K(z)(|z|\wedge{}1)dz}}{}, $$ and $$ C_{K,2}:={\int^{\infty}_{-\infty} K(z)(K(z)-K(z-1))d zC_{K,1}}. $$ Next, we show that indeed ${\widetilde{V}(g)^{n}_T}$ enjoys a centered CLT. Its simple proof is given in Section (ref).

corollaryUnder the assumptions {\color{DR} and notation} of Theorem (ref), {as $n\to\infty$}, we have \begin{equation} \frac{1}{\sqrt{\Delta_n}}\left( {\widetilde{V}(g)^{n}_T} - V(g)\right) \stackrel{st}{\to} Z, \end{equation} {\color{DR} where ${\widetilde{V}(g)^{n}_T}$ is defined as in (ref) with $\hat{c}$ given by ((ref)) {\color{DR} with the} condition (ref). Furthermore, {\color{DR} when $X$ is continuous, (ref) holds under the weaker condition {(ref)}.}}
remarkjacod2015estimation proposed the following bias-corrected estimator: \begin{align} \widetilde{V}(g)^{n,{JR}}_T&:=V(g)^{n,{JR}}_T+\frac{k_n \Delta_n}{2}\left(g\left(\hat{c}_1^{n,{JR}}\right)+g\left(\hat{c}^{n,{JR}}_{n-k_n+1}\right)\right)-\frac{3}{4k_n}V(h)_T^{n,{JR}}\\ &\quad+\frac{\Delta_n}{8} \sum_{i=1}^{n-2k_n+1}\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g\left(\hat{c}_i^{n,{JR}}\right)\left(\hat{c}_{i+k_n}^{n,{JR}, p q}-\hat{c}_i^{n,{JR}, p q}\right)\left(\hat{c}_{i+k_n}^{n,{JR}, u v}-\hat{c}_i^{n,{JR}, u v}\right), \nonumber \end{align} where $\hat{c}_i^{n,{JR}}$ and $V\left(g\right)^{n,JR}_{T}$ are defined as in ((ref)) and ((ref)), respectively. {\color{DR} There is a connection between (ref) and our estimator (ref) taking a right-sided uniform kernel $K(x)={\bf 1}_{[0,1)}(x)$. Indeed, note that, with that kernel,} ((ref)) reduces to \begin{equation} \begin{split} {\widetilde{V}(g)^{n}_T} &:= V\left(g\right)^{n}_{T} +\frac{k_n \Delta_n}{2} g\left(\hat{c}_1^n\right)-\frac{3}{4k_n}V\left(h\right)^{n}_{T}\\ &\quad+\frac{\Delta_n}{8} \sum_{i=1}^{n-2k_n+1}\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g\left(\hat{c}_i^n\right)\left(\hat{c}_{i+k_n}^{n, p q}-\hat{c}_i^{n, p q}\right)\left(\hat{c}_{i+k_n}^{n, u v}-\hat{c}_i^{n, u v}\right). \end{split} \end{equation} Then, by {\color{DR} (ref) and} ((ref)), we {\color{DR} conclude that} \begin{equation} \begin{split} { \widetilde{V}(g)^{n}_T} &=\widetilde{V}(g)^{n,{JR}}_T +\Delta_{n}\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)-\Delta_n g\left(\hat{c}_{0}^{n}\right)-\frac{k_n \Delta_n}{2} g\left(\hat{c}^{n,{JR}}_{n-k_n+1}\right)\\ &\quad-\frac{3}{4k_n}\Delta_n\sum_{i=n-k_{n}-1}^{n}h\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)+\frac{3\Delta_n}{4k_n} h\left(\hat{c}_{0}^{n}\right)\\ &\quad+\frac{\Delta_n}{8}\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g\left(\hat{c}_{n-2k_n+1}^{n}\right)\left(\hat{c}_{n-k_n+1}^{n, p q}-\hat{c}_{n-2k_n+1}^{n, p q}\right)\left(\hat{c}_{n-k_n+1}^{n, u v}-\hat{c}_{n-2k_n+1}^{n, u v}\right)\\ &\quad-\frac{\Delta_n}{8}\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g\left(\hat{c}_{0}^{n}\right)\left(\hat{c}_{k_n}^{n, p q}-\hat{c}_0^{n, p q}\right)\left(\hat{c}_{k_n}^{n, u v}-\hat{c}_0^{n, u v}\right). \end{split} \end{equation} Since $$ \sqrt{\Delta_n}\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)\stackrel{\mathbb{P}}{\longrightarrow} \frac{\theta}{2}g\left(c_T\right),\qquad \frac{k_n \sqrt{\Delta_n}}{2} g\left(\hat{c}^{n,{JR}}_{n-k_n+1}\right)\stackrel{\mathbb{P}}{\longrightarrow} \frac{\theta}{2}g\left(c_T\right), $$ we know that $\Delta_n\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)-\frac{k_n \Delta_n}{2} g\left(\hat{c}^{n,{JR}}_{n-k_n+1}\right)=o_{p}\left(\sqrt{\Delta_n}\right)$. Also note that $$\sqrt{\Delta_n}\sum_{i=n-k_{n}-1}^{n}h\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)\stackrel{\mathbb{P}}{\longrightarrow} \frac{\theta}{2}h\left(c_T\right), $$ which implies that $\frac{3}{4k_n}\Delta_n\sum_{i=n-k_{n}-1}^{n}h\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)=O_{p}\left(\frac{\sqrt{\Delta_n}}{k_n}\right)$. The remaining terms in ((ref)) can {\color{DR} directly} be proved {\color{DR} to be} $o_{p}\left(\sqrt{\Delta_n}\right)$. We then conclude that $$ \frac{1}{\sqrt{\Delta_n}}\left(\widetilde{V}(g)^{n}_T -\widetilde{V}(g)^{n,{JR}}_T\right)\stackrel{\mathbb{P}}{\longrightarrow} 0. $$ Hence, by Corollary (ref), we can recover the {\color{DR} stable} convergence in law for the bias-corrected estimator in jacod2015estimation: \begin{equation} \frac{1}{\sqrt{\Delta_n}}\left( \widetilde{V}(g)^{n,{JR}}_T - V(g)\right) \stackrel{st}{\to} Z. \end{equation}

As mentioned above, the bias term $A^3 $ contains the `volvol' $\tilde{c} $ and the estimator of $A^3$ in ((ref)) could introduce extra variance to the {asymptotically} unbiased estimator ((ref)). {An alternative} way to eliminate {the} bias is {undersmoothing, i.e., through the selection of a bandwidth sequence $b_n$ converging to $0$ at a faster rate than $\sqrt{\Delta_n}$ ($b_n \ll \sqrt{\Delta_n}$) or, equivalently, picking $k_n$ such that $\sqrt{\Delta_n}k_n \stackrel{n \to \infty}{\longrightarrow} 0$ (meaning that $\theta = 0$ in ((ref)))}. This is the approach put forward in jacod2013quarticity. {In that case, it is expected that the bias terms $A^1$ and $A^3$ will vanish and the bias term $A^2 $ will dominate}. {{\color{DR} Following the same idea,} we can devise an asymptotically unbiased {\color{DR} estimator as follows}:

equation[equation omitted — 177 chars of source]

The following result establishes the asymptotic behavior of $\bar{V}(g)^n_T $. {Its proof is {given} in Appendix (ref).}

theoremAssume $k_n$ satisfies \begin{equation} k_{n}^2 \Delta_n \to 0,\quad k_n^3 \Delta_n \to \infty{.} \end{equation} Then, under ((ref)) and Assumptions (ref) and (ref), we have the following stable convergence in law \begin{equation} \frac{1}{\sqrt{\Delta_{n}}}\left(\bar{V}(g)_T^{n}-V(g)_T\right) \stackrel{st}{\longrightarrow} Z, \end{equation}where $Z$ is as in Theorem (ref) {\color{DR} and ${\bar{V}(g)^{n}_T}$ is defined as in (ref) with $\hat{c}$ given by ((ref)) under condition (ref). Furthermore, {\color{DR} when $X$ is continuous, (ref) holds under the weaker condition (ref).}}

Simulation Study

In this section we analyze the performance of {\color{DR} our} kernel-based estimators. To isolate the effect of the bandwidth on the estimators' performance, we only consider continuous processes (i.e., $\delta\equiv 0$ in (ref) and, thus, $X\equiv X'$) and focus on the untruncated estimator (ref) and the corresponding estimators $V(g)^{n}_{t}$ of (ref) and $\widetilde{V}(g)^{n}_T$ of (ref).

Simulation design

We consider the data generating model in li2019efficient:

equation[equation omitted — 167 chars of source]

with $\mathbb{E}\left[d W_t d B_t\right] = -0.75 dt $. Throughout, we use the function $g(c) = c^2$, which corresponds to the integrated quarticity{,} and also the function $g(c) = \log(c)$. We simulate data for $T= 5$ days and $T = 21$ days, and assume the process $X$ is observed once {\color{Red} every 1 minute or every 5 minutes, with 6.5 trading hours per day, for all the results of Subsection (ref), but not for Subsection (ref), where we assume a frequency of 1-second}.

Validity of the asymptotic theory

We first show the finite sample behavior of the estimator is consistent with the Central Limit Theorem of Theorem (ref). Under the notation of Theorem (ref), we choose $\color{DR}{ \theta = 0.2}$, $g(x) = x^2 $, and estimate the functional volatility using the exponential kernel $K(x) = \frac{1}{2}e^{-|x|}$. The choice of $\theta$ was obtained from the formula $b^*=\theta^*\sqrt{\Delta_n}$, where $b^*$ was chosen using the iterative method in FigLi. The sample frequency of the data is 1 second. Based on 5,000 simulated paths from the model ((ref)), we {calculate the standardized estimation error for each path $i=1,\dots, 5000$} according to Theorem (ref):

equation[equation omitted — 200 chars of source]

where $c^i_s = (\sigma^i_s)^2 $, $\{\hat{c}^{n,i}_j\}_j $ are the kernel spot volatility estimates of the $i$th simulated path, and the variance ${\rm Var}(Z)^{i}=2 \int_0^T(g''(c^i_s)c^i_s)^2ds$ and the biases $A^{\ell,i}$ ($\ell=1,2,3$), as defined in (ref), are computed via Riemann sum approximations using the true simulated volatility $\{c^{n,i}_j\}_{j=0,\dots,n}$ and volvol observations $\{\tilde{c}^{n,i}_j\}_{j=0,\dots,n}$. We then compare the histogram of the standardized estimates $z_i$ with a standard normal distribution (see Figure (ref)). As shown therein, the distribution of the estimation error is consistent with the asymptotic result in Theorem (ref).\%, though the convergence rate may be slow.\\

figure[figure omitted — 207 chars of source]

Similarly, we verify the behavior of the bias-corrected estimator proposed in Corollary (ref). With the same value of $\theta$, $g(x)$, exponential kernel, and simulated data, the standardized bias-corrected estimator takes the form:

equation[equation omitted — 138 chars of source]

The corresponding histogram is shown in Figure (ref), which again suggest the CLT stated in Corollary (ref) is valid.

figure[figure omitted — 216 chars of source]
table[table omitted — 1,680 chars of source]
table[table omitted — 1,686 chars of source]

Comparison with Jackknife estimator

We now compare our bias-corrected estimator $\widetilde{V}(g)_T^{n}$ to the Jackknife estimator of li2019efficient. For the Jackknife method, we consider the two-scale estimator $TS_n$, its boundary-adjusted version $TS_n'${,} and the multiscale estimator $MS_n$ as defined in li2019efficient, Section 3.1. The parameters for the Jackknife estimators are chosen according to li2019efficient. {\color{DR} By running numerous simulations, we determine that the values chosen by li2019efficient broadly optimize their estimator's performance}. For the multiscale estimator $MS_n$, $\left(\psi_1, \psi_2, \psi_3\right)=(-2.5,8,-4.5)$, and $\left(k_{1, n}, k_{2, n}, k_{3, n}\right)=(15,30,45)$ for data sampled every 5 minutes, or $(40,80,120)$ for data sampled every minute. For the two-scale estimator, $\left(\psi_1, \psi_2\right)=(-1,2)$ along with the same $\left(k_{1, n}, k_{2, n}\right)$ as above. For our kernel estimator, we use 3 kernels: a two-sided exponential kernel $K(x)=\frac{1}{2}e^{-|x|}$, a two-sided uniform kernel $K(x)=\frac{1}{2}{\bf 1}_{[-1,1]}(x)$, and a right-sided uniform kernel $K(x)={\bf 1}_{[0,1]}(x)$. We analyze the performance of the simple biased estimator $V(g)^n_T$ in ((ref)) and the bias-corrected estimator $\widetilde{V}(g)^{n}_T$ in ((ref)). We also consider the simple biased estimator $V(g)^{n,{JR}}_T$ and the bias-corrected estimator $\tilde{V}(g)^{n,{JR}}_T$ of (ref) and (ref), respectively, which were proposed in jacod2015estimation.

The bandwidths of all the kernel estimators are chosen using the iterative method in FigLi, Section 5. We iterate the method only once. For a two-sided kernel, the integrated volatility of volatility used when updating the bandwidth is estimated by the Two-time Scale Realized Volatility (TSRV) estimator proposed in FigLi. For the right-sided uniform kernel, the TSRV estimator no longer applies. We then use the realized variance estimator on the estimated volatility path as a replacement. As finite-sample corrections, in ((ref)), at {each} $t_i$, we replace $\int_{-\infty}^{\infty}K^2(u) du $ with $ \sum_{j=1}^n K^2_h(t_i - t_j) \Delta_n h$, and replace $\int_{-\infty}^{\infty}L(s) ds $ and $\int_{-\infty}^{\infty}L^2(s)ds $ with $ \frac{\Delta_n}{b}\sum_{j=1}^n L\left(\frac{t_{j-1}-t_i}{b} \right)$ and $\frac{\Delta_n}{b}\sum_{j=1}^n L^2\left(\frac{t_j-t_i}{b} \right)$, respectively, where $ L\left(\frac{t_j - t_i}{b} \right) = -\sum_{l=1}^{j}K_{{b}}(t_{l-1}-t_i) $ if $j \leq i$ and $\sum_{l=j+1}^{n}K_{{b}}(t_{l-1}-t_i)$ if $ j > i$.

In Table (ref), we report the relative bias and relative root-mean-square-errors in percentage unit for the considered estimators. Our results of the Jackknife estimators are consistent with the ones in li2019efficient, Table 1 therein. As shown in Table (ref), the kernel estimators {\color{DR} exhibit a similar performance to} the Jackknife estimators. However, the Jackknife estimators have more tuning parameters, while ours requires only one tuning parameter (namely, $\theta$). As we will show in Section (ref), the {\color{Red} most important} advantage of our estimators {\color{Red} lies} on their stability relative to the choice of the bandwidth. That is, our estimators seem to exhibit satisfactory performance for a large range of the tuning parameters, while other estimators are more sensitive to choosing `good' values for their tuning parameters.

{\color{DR} Returning to} the results of Table (ref),} we also observe that in some cases the kernel estimators have smaller relative root mean square error (RMSE) compared with the Jackknife estimators, especially when the time span $T$ is short ($T = 5$ days). Intuitively, the superior performance of the spot volatility estimators with exponential kernel (shown in FigLi, Section 7) {\color{DR} is expected to be more evident} when the time span is relatively short, e.g. 5 days. For reference, we include the tables for the function $g(c) = log(c)$ in Table (ref). These tables confirm our previous conclusion. It is worth noting that in general, the bias-corrected kernel estimator $\widetilde{V}^{n}_T$ has {\color{DR} significantly} smaller RMSE compared with the simple biased estimator $V(g)^{n}_T$, which also happens among the Jacod & Rosenbaum’s estimators, $V(g)^{n,{JR}}_T$ and $\widetilde{V}^{n,JR}_T$.

{\color{DR} Bandwidth Sensitivity}

In this subsection, we investigate how sensitive the estimators are relative to the bandwidth parameter. For the Jackknife estimators, we focus on the boundary adjusted multi-scale estimator $\text{MS}$ (which, as shown in Tables (ref) and (ref), typically has the best performance among all {\color{Red} the considered} Jackknife estimators), and for the kernel estimators, we {\color{DR} first consider an} exponential kernel. We {\color{DR} try} a sequence of bandwidths $k_n$ for the kernel estimator, and a sequence of triplets {$(k_{1,n}, k_{2,n}, k_{3,n}) = ([\frac{k_n}{2}], k_n, [\frac{3k_n}{2}])$} and $(\psi_1, \psi_2, \psi_3) = (-2.5,8,-4.5)$ for $MS$, which include the parameters chosen in li2019efficient {\color{DR} when $k_n=30$ (see Table 1 therein)}. The results are shown in the Figures (ref)-(ref), where $\text{S0 exp}$ denotes the simple biased exponential kernel estimator in ((ref)) and $\text{S3 exp}$ denotes the bias-corrected estimators with exponential kernel in ((ref)). {\color{DR} We can conclude that the estimator $\text{S3 exp}$ exhibits the most stable or consistent performance for} either 5 or 21 days data with 1 or 5 minutes sampling frequency, and both $g(c)=c^2$ or $\log (c)$. {\color{DR} In other words, the estimator's values are {\color{DG} in general} relatively invariant to changes in the bandwidth, {\color{DR} especially} when working with low frequency and $g(c)=c^2$. Under those conditions, the others estimators require more careful tuning of the bandwidth in order to show good performance. We can also observe that the bias-corrected exponential kernel estimator almost always outperforms the boundary-adjusted multi-scale Jackknife estimator $\text{MS}$ and the simple biased exponential kernel estimator $\text{S0 exp}$.}

{\color{DR} Finally,} in Figure (ref), we analyze the bandwidth stability of three estimators: the bias-corrected estimators (ref) with exponential and right-sided uniform kernels and the bias-corrected estimator (ref) proposed in jacod2015estimation. Again, since we are not introducing jumps in the process $X$, the estimator $\hat{c}_i^{n}$ used for (ref) and the estimator $\hat{c}_i^{n,{JR}}$ used for (ref) do not have any truncation (i.e., we fix $v_n=\infty$ in (ref) and (ref)). Figure (ref) shows that our proposed estimator $\widetilde{V}^{n}_T$ with either exponential or right-sided uniform kernel almost always outperforms $\widetilde{V}^{n, JR}_T$. The exponential kernel appears to be less sensitive to the choice of $k_n$ than the right-sided uniform kernel. The estimator $\widetilde{V}^{n, JR}_T$ also seems to require a careful tuning of the bandwidth in order to achieve a good performance.

commentWe also notice $S0$ has similar behavior compared to the boundary adjusted two-scale estimator $TS'$, yet the estimator $S0$ doesn't require any bias correction and even outperforms the $TS'$ estimator under a wide range of bandwidth, except for 21 days data with 5 min sampling frequency, $g(c)=\log (c)$.
figure[figure omitted — 657 chars of source]
figure[figure omitted — 653 chars of source]
figure[figure omitted — 665 chars of source]
figure[figure omitted — 661 chars of source]
figure[figure omitted — 743 chars of source]