EconBase
← Back to paper

Estimation of Integrated Volatility Functionals with Kernel Spot Volatility Estimators

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

65,970 characters

Estimation of Integrated Volatility Functionals with Kernel Spot Volatility Estimators


\maketitle
\begin{abstract}
For a multidimensional It\^o semimartingale, we consider the problem of estimating integrated volatility functionals. \cite{jacod2013quarticity} studied a plug-in type of estimator based on a Riemann sum approximation of the integrated functional and a spot volatility estimator with a forward uniform kernel. Motivated by recent results that show that spot volatility estimators with general {two-sided} kernels of unbounded support are more accurate, in this paper an estimator using a general kernel spot volatility estimator as the plug-in is considered. A biased central limit theorem for estimating the integrated functional is established with an optimal convergence rate. Central limit theorems for properly de-biased estimators are also obtained both at the optimal convergence regime for the bandwidth and when applying undersmoothing. Our results show that one can significantly reduce the estimator's bias by adopting a general kernel instead of the standard uniform kernel. Our proposed bias-corrected estimators are found to maintain remarkable robustness against bandwidth selection in a variety of sampling frequencies and functions.
\smallskip
\end{abstract} \hspace{10pt}

\section{Introduction}
In this work, we are concerned with the  estimation of integrated volatility functionals of the form  $V_{T}(g):=\int_{0}^{T}g(c_s)ds$, where $g$ is a smooth function and $c_t:= \sigma_t\sigma_t^*$ is the spot {\color{Blue} variance-covariance} process of a $d$-dimensional It\^o semimartingale given by
\begin{equation}
dX_t = \mu_t dt + \sigma_t d W_t +dJ_t.
\end{equation}
Here, $\{W_{t}\}_{t\geq{}0}$ is $d$-dimensional standard Brownian Motion (BM) and $\{J_{t}\}_{t\geq{}0}$ is a pure-jump process of bounded variation.
Such functionals have many applications in financial econometrics. In the simplest case, when $d=1$ and $g(x) =x$, this is the well-studied integrated volatility (IV) or variance $IV_T=\int_0^{T}\sigma^2_sds$, which measures the overall variability or {\color{Blue} uncertainty} latent in the process $X$ during $[0,T]$.
Again, when $d=1$ and $g(x) = x^2$, $V_T(g)$ corresponds to the integrated quarticity, which appears in the limiting distribution of the realized variance estimator $\widehat{IV}_T=\sum_{i=1}^n (\Delta_i^nX)^2$ and, hence, it is crucial to construct feasible confidence intervals for IV. Another function appearing in the literature  is $g(x)=\log(x)$, when researching the statistical features of a factor $Y$ driving the volatility in a model where $\sigma$ is assumed to be the exponential of $Y$. Volatility functionals have also been used in principal component analysis by \cite{ait2019principal}  and in estimating integrated moments for option pricing purposes by \cite{li2016generalized}.

One of the fundamental methods to estimate $V_{T}(g)=\int_{0}^{T}g(c_s)ds$ consists of approximating the integral by a Riemman sum and {\color{Blue} plugging} in a suitable spot volatility estimator $\hat{c}$ in place of $c$:
\begin{equation} \label{eq:V_jacod}
\widehat{V}_{T}(g)=\Delta_{n} \sum_{i=1}^{n-k_{n}+1} g\left(\hat{c}_{i}^{n}\right).
\end{equation}
This method was pioneered by \cite{jacod2013quarticity}, where,  as the spot volatility estimator, a local rolling window estimator with uniform weights was used:
 \begin{equation} \label{eq:V_jacodb}
	\hat{c}_{i}^{n,unif}=\frac{1}{k_{n} \Delta_{n}} \sum_{j=0}^{k_{n}-1}\left( \Delta^n_{i+j}X\right)^{2}.
\end{equation}
Two interesting facts emerged from their theory. {\color{Red} For the ease of exposition, we focus on the $d=1$-case in the introduction, though our main results are stated for general dimension $d\geq 1$.} First of all, the rate of convergence of the estimation error is unexpectedly of order $\Delta_n^{\frac{1}{2}}$ even though $\hat{c}_{i}^{n}$ can only attain the rate $\Delta_n^{\frac{1}{4}}$\footnote{{\color{Red}  Intuitively, by {\color{Blue} Taylor} expansion of the estimator $g(\hat{c})$, we can write $\Delta_n\sum_{i=1}^n g(\hat{c}_i^n) - \Delta_n\sum_{i=1}^{n}g(c_{t_{i-1}})= \Delta_n\sum_{i=1}^{n}g'(c_{t_{i-1}}) \left(\hat{c}_{i}^n - c_{t_{i-1}}\right) + \frac{1}{2}\Delta_n\sum_{i=1}^{n}g''(c_{t_{i-1}}) \left(\hat{c}_{i}^n - c_{t_{i-1}}\right)^2+\text{higher order terms}$. The estimation errors in the first term of the right-hand side are ``averaged out" so that the convergence rate changes from $n^{-1/4}$ to $n^{-1/2}$.}}. Secondly, when $k_n \sim \theta/\sqrt{\Delta_n} $ for some $\theta\in(0,\infty)$, which yields the best convergence rate for the spot volatility estimator $\hat{c}_{i}^{n}$, $\frac{1}{\sqrt{\Delta_n}}\left( \widehat{V}_{T}(g) - V_{T}(g) \right) $ converges to a mixed Gaussian distribution with four bias terms:
\begin{enumerate}
\item A term of the form $-\frac{\theta}{2}\left(g(c_0) + g(c_T) \right) $ due to border effects;
\item A {\color{Blue} second} term of the form $\frac{1}{\theta} \int_0^T g^{\prime \prime}(c_s) c_s^2ds $ due to the nonlinearity of $g$;
\item A third term of the form $-\frac{\theta}{12} \int_0^T g^{\prime \prime}(\tilde{c}_s) $, where $\tilde{c} $ is the volatility of volatility, due to target error $c_i - \bar{c}_{i}^n$, where $\bar{c}_{i}^{n}=\frac{1}{k_{n} \Delta_{n}} \sum_{j=0}^{k_{n}-1}c_{i+j}$;
\item A fourth term involving the jump component of the volatility process $\theta \sum_{s\le t} G(c_{s-}, c_s) $ for a suitable function of $G$ depending on g. This error disappears when $c$ is continuous.
\end{enumerate}

In principle, the bias terms can be estimated and eliminated from $\widehat{V}_{T}(g)$ to obtain a feasible CLT with centered Gaussian limiting distribution. This strategy was explored by \cite{jacod2015estimation}.
An alternative solution, however, was adopted in \cite{jacod2013quarticity}.
Specifically, the window length $k_n$ was made to converge to $\infty$ at a slower rate (under-smoothing) so that $k_n \sqrt{\Delta_n}\to{}0$. This corresponds to the case $\theta= 0 $ and, hence, the first, third, and fourth bias terms would vanish, while the second term becomes dominant. To keep the convergence rate $\sqrt{\Delta_n} $ of the estimator $\widehat{V}_T(g) $, this second term needs to be estimated and the estimator needs to be de-biased. The resulting estimator takes the form
\begin{equation}
\Delta_n \sum_{i=1}^{n - k_n + 1}\left( g\left(\hat{c}_{i}^{n}\right) - \frac{1}{k_n} g^{\prime \prime }(\hat{c}_i)\hat{c}_i^2 \right),
\end{equation}
which indeed achieves the rate $\Delta_n^{\frac{1}{2}}$ with semi-parametrically optimal asymptotic variance. {\color{Red} A similar approach to deal with {\color{Blue} edge} effects and biases was also considered in \cite{kristensen2010nonparametric}, who also proposed kernel estimators and obtained a central limit theorem (though the model considered therein lacks leverage effects).} Alternatively, as mentioned before, in the optimal case $\theta \in (0,\infty)$ and employing forward uniform kernels, estimators for all the bias terms were proposed in \cite{jacod2015estimation}, where an unbiased estimator with centered limiting distribution was developed.

\cite{li2019efficient} took a different approach towards bias correction. They proposed a type of `jackknife' estimators, which are formed as a linear combination of a few uncorrected estimators of the form \eqref{eq:V_jacodb} associated with different windows $k_n$.
Since these estimators have similar bias terms but different coefficients depending on the {\color{Blue} window} sizes, the biases can be eliminated by a proper linear combination. Both the `two-scale' jackknife (i.e., when combining only two uncorrected {\color{Blue} estimators}) and \cite{jacod2013quarticity}'s estimator
use under-smoothing, so that the biases due to boundary effects, volatility of volatility, and volatility jumps are rendered asymptotically negligible. One drawback of this approach is that one needs to determine the degree of undersmoothing. In other words, if, for instance, we take $k_n \sim \beta\Delta_n^{\alpha}$ with $\alpha>-1/2$, then one needs to tune both parameters $\beta$ and $\alpha$, which is nontrivial.
In contrast, a multiscale jackknife estimator, formed as a linear combination of three (or more) estimators, can in principle cancel all the biases regardless of the speed of convergence of the window size. However, again, this involves more tuning parameters that need to be carefully calibrated.

Other related work on this topic includes \cite{mykland2009inference}, who proposed using the same method as in \cite{jacod2013quarticity}, but {\color{Red} keeping $k_n = k$ constant}; for the quarticity or more generally for estimating $\int_0^T g(c_s) ds  $ with  a power functional $g(x) = x^r$, they obtained a CLT with rate $1/\sqrt{\Delta_n}$ and an asymptotic variance that is not efficient, {\color{Red} but can arbitrarily approach the optimal variance} when $k$ is large.  {\color{Red} \cite{Glotter} established a Locally Asymptotically
Mixed Normality (LAMN) condition for the estimation of integrated volatility functionals in a continuous diffusion process where the volatility is driven by an independent It\^o process. The result therein shows that the optimal rate of convergence is $n^{-1/2}$. Another related result in the same direction is \cite{Renault}}. {\color{Red} More recently, \cite{LiLiu2021} generalized \cite{jacod2013quarticity} 's result to allow a long-memory component such as {\color{Blue} an} fBM with Hurst parameter $H>1/2$. They also established a semiparametric
efficiency lower bound, generalizing \cite{Renault}'s results.} \cite{chen2018inference} developed an estimator for $\widehat{V}_{T}(g)$ in the presence of microstructure noise,  based on a forward finite difference approximation of the standard pre-averaging estimator of the integrated volatility.

Our work aims to apply a general kernel spot volatility estimator to the plug-in estimator of integrated volatility functionals. The general kernel estimator (\cite{fan2008spot}, \cite{kristensen2010nonparametric}) is defined as
\begin{align}\label{MDKNN00}
	{\tilde{c}^{n}_{i}} :=\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right)\left(\Delta_{j}^{n} X\right)^{2},
\end{align}
where $K_{{b}}(x):=K(x/{b})/{b}$ {\color{Red} for some bandwidth $b>0$}, and $k_n$ is an appropriate sequence that controls the bandwidth $b_n=k_n\Delta_n$ and that needs to be calibrated.
Of course, one can see \eqref{eq:V_jacodb} as a particular case of \eqref{MDKNN00} with $K(x)={\bf 1}_{[0,1]}(x)$. {\color{Red} \cite{kristensen2010nonparametric} developed an asymptotic theory for \eqref{MDKNN00} when $t_i$ is a fixed time $\tau\in (0,T)$ and showed asymptotic normality to the spot variance $c_{\tau}$ under certain path-wise H\"older continuity and non-leverage conditions on $c$. See also \cite{fan2008spot} and \cite{mancini2015estimation} for other related results. All these results applied undersmoothing, which allowed them to neglect the ``target error" coming from approximating the spot volatility by a kernel weighted integrated volatility.}


In this paper, we consider the following version:
\begin{align}\label{MDKNNa0}
	{\hat{c}^{n}_{i}} :=\frac{\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right)\left(\Delta_{j}^{n} X\right)^2}{\bar{K}_{n}\left(t_i\right)},
\end{align}
with $\bar{K}_{n}\left(t_i\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right)$,
which is known to be more robust than \eqref{MDKNN00} {\color{Blue} against} edge effects {\color{Red} (see also \cite{kristensen2010nonparametric} for other edge-robust alternatives)}. As a result, we can, in principle, consider the whole range $i=1,\dots, n$ in \eqref{eq:V_jacod}. Furthermore, we find it more convenient for our proofs to consider the estimator
\begin{equation} \label{eq:V_jacodBis}
\widehat{V}_{T}(g)=\Delta_{n} \sum_{i=1}^{n} \bar{K}_n(t_i)g\left(\hat{c}_{i}^{n}\right).
\end{equation}
The coefficient  $\bar{K}_n(t_i)$ attenuates the contribution of $g\left(\hat{c}_{i}^{n}\right)$ for $t_i$ near $0$ and $T$ and, thus, serves as an additional edge effect correction. {\color{Red} \cite{kristensen2010nonparametric} also considered a kernel-based estimator similar to \eqref{eq:V_jacodBis}, but applied an edge effect correction, similar to that in \cite{jacod2015estimation}, by eliminating the terms corresponding to $i$'s close to $0$ and $n$,  and did not consider the adjustments $\bar{K}_n(t_i)$. Their asymptotic theory also only considered the undersmoothing asymptotic regime mentioned above.}

The motivation for considering general kernels comes from \cite{FigLi} and \cite{FigWu}, where it has been shown, both theoretically and by Monte Carlo experiments, that a general two-sided kernel with unbounded support could significantly reduce the variance of the spot volatility estimator $\hat{c}$. More specifically, \cite{FigLi} studied the infill asymptotic behavior of the mean-square error (MSE) of \eqref{MDKNN00} {for continuous It\^o semimartingales} and showed that, when setting $k_n$ `optimally' (i.e., in order to minimize the leading order terms of the MSE), the double exponential kernel $K(x)=.5 e^{-|x|}$ minimizes the resulting MSE over all kernels. This result was established under the absence of leverage effects. \cite{FigWu} extended this result, by first proving a CLT for a general {two-sided} kernel, with optimal convergence rate and in the presence of jumps and leverage {effects}. It was then shown that exponential kernels minimize the asymptotic variance of the CLT. It is worth noting that, even when applying undersmoothing, the asymptotic variance could be reduced when choosing a kernel of unbounded support. Specifically, when applying undersmoothing, the asymptotic variance of the spot volatility estimator is proportional to $\|K\|^2=\int_{-\infty}^{\infty} K^2(x) dx$. When $K$ is constrained to have support on $[0,1]$, the optimal kernel is indeed the uniform kernel $K_{unif}(x)={\bf 1}_{[0,1]}(x)$ with $\|K_{unif}\|^2=1$\footnote{{\color{Red} Indeed, by Jensen's inequality, $K_{unif}(x)={\bf 1}_{[0,1]}(x)$ satisfies $\int K^2_{unif}(x)dx=1=\left(\int K(x)dx\right)^2\leq \int K^2(x)dx$, which shows that a uniform kernel $K_{unif}(x)$ minimizes the asymptotic variance when applying undersmoothing.}}. However, picking the kernels $.5e^{-|x|}$ {\color{Red} and $e^{-x}{\bf 1}_{x>0}$ will result in estimators whose asymptotic variances are respectively a quarter and a half of} that obtained by a uniform kernel $K_{unif}(x)={\bf 1}_{[0,1]}(x)$. {\color{Red} Similar comments apply when considering only backward looking kernels such as ${\bf 1}_{[-1,0)}(x)$ and $e^{-|x|}{\bf 1}_{x<0}$.}

A natural question is whether the estimation accuracy gained when estimating the spot volatility by a general kernel transfers into gains in estimating integrated functionals. Intuitively, by the Taylor expansion of the estimator $g(\hat{c})$, we can write $g(\hat{c}) - g(c)= g'(c) \left(\hat{c} - c\right) + \frac{1}{2}g''(c)\left(\hat{c} - c\right)^2 + O_p\left(|\hat{c} - c|^3 \right)$. After taking expectations, the variance of the spot volatility estimator, which appears in the second term,  becomes a bias term of the integrated volatility estimator. Therefore, the bias of the estimator $\widehat{V}_{T}(g)$ could in principle be reduced by adopting a general kernel.

As it turns out, the CLT of \eqref{MDKNN00} again exhibits several biases, which depend on a rather intricate and nontrivial manner on the kernel. The bias terms compared to the uniform kernel case involves a highly nontrivial transformation of the kernel function. The second bias term, involving $\frac{1}{\theta} \int_0^T g^{\prime \prime}(c_s) c_s^2ds$, contains the $L_2$ norm of the kernel and, thus, can be reduced because certain kernels of unbounded support {\color{Blue} possess} smaller $L_2$ norm. We then propose bias-corrected estimator similar to those in \cite{jacod2015estimation}, and derive the corresponding CLT. We show by Monte Carlo experiments that our general kernel estimator typically has superior performance compared with the jackknife estimators of \cite{li2019efficient}. Of course, our estimator is simpler as it does not involve as many tuning parameters, which are hard to tune-up. In fact, our Monte Carlo experiments suggest our estimators are significantly more stable in the sense that they exhibit better performance for most values of the bandwidth, while the jackknife estimator requires accurate tuning of the `bandwidth' to match our estimator's performance. We also observe that even without bias corrections, under certain conditions, the plug-in estimator with general kernel spot volatility estimator has good performance, which could be a result of the increased accuracy of {\color{Blue} the} spot volatility estimator and/or the reduction in bias attained by using a general kernel.

While working on the final stages of writing the present manuscript, we became aware of a recent work by \cite{benvenutifunctionals}, who also studied a type of plug-in estimator with a general two-sided kernel spot volatility estimator. However, unlike our work, only a suboptimal bandwidth asymptotic regime (i.e., assuming under-smoothing) was considered and only consistency was established.

The rest of this paper is organized as follows. Section \ref{setting} introduces the framework, assumptions, and  main results.  Section \ref{SimulationSect} illustrates the performance of our method via Monte Carlo simulations and compares it to the jackknife estimator of \cite{li2019efficient}.  The proofs are deferred to an appendix section.

\textbf{Notation:} We shall use the following notation:
\begin{itemize}
\item The set of (nonnegative definite) $d\times{}d$ real matrices $x=[x^{\ell m}]_{\ell,m=1}^{d}$ is denoted ($\mathcal{M}_+^d$) $\mathbb{R}^{d\times d}$;
\item For a $C^1$ function $g:\mathbb{R}^{d\times d}\to\mathbb{R}$, its gradient {\color{Red} $\nabla g(x)$} is the $d\times d$ matrix with $(\ell,m)$ entry $\partial_{\ell m}g(x):=\frac{\partial g(x)}{\partial x^{\ell m}}$;
\item For $x,y\in\mathbb{R}^{d\times d}$, its inner product is defined as $x\cdot y=\sum_{\ell,m}x^{\ell m}y^{\ell m}$. Also, $x^{*}$ is the transpose of $x$ and $\|x\|=\sqrt{\sum_{\ell,m=1}^n (x^{\ell m})^2}$ is the Euclidian norm.

\end{itemize}


\section{The Setting, Estimator, and Main Results} \label{setting}
Throughout, we consider a $d$-dimensional It\^o semimartingale $X=[X^1,\dots,X^d]^*\in\mathbb{R}^{d\times{}1}$ of the form:
\begin{equation} \label{eq:X}
\begin{split}
X_{t}=&X_0 + \int_0^t\mu_{s} ds +\int_0^t\sigma_{s} d W_{s}\\
&+\int_0^t \int_E \delta(s, z) \mathbbm{1}_{\{|\delta(s, z)| \leq 1\}}(\mathfrak{p}-\mathfrak{q})(d s, d z)+\int_0^t \int_E \delta(s, z) \mathbbm{1}_{\{|\delta(s, z)|>1\}} \mathfrak{p}(d s, d z),
\end{split}
\end{equation}
where all stochastic processes ($\mu :=\left\{\mu_{t}\right\}_{t \geq 0}$, $\sigma :=\left\{\sigma_{t}\right\}_{t \geq 0}$,  $W :=\left\{W_{t}\right\}_{t \geq 0}$, $\delta:=\left\{\delta(t, z):t\geq 0, z\in E\right\}$, $\mathfrak{p}:=\{\mathfrak{p}(B):B\in\mathcal{B}(\mathbb{R}_{+}\times E)\}$, etc.) are defined on a complete filtered probability space  $\left(\Omega, \mathcal{F}, \mathbb{F}, \mathbb{P}\right)$  with filtration $ \mathbb{F}=(\mathcal{F}_{t})_{t \geq 0}$. Here, {\color{DG} $W=[W^1,\dots,W^d]^*$} is a $d$-dimensional standard Brownian Motion (BM)  adapted to the filtration \(\mathbb{F}\), $\delta$ is a predictable $\mathbb{R}^d$-valued function on $\Omega \times \mathbb{R}_{+} \times E$, and $\mathfrak{p}$ is a Poisson random measure on $\mathbb{R}_{+} \times E$ for some arbitrary Polish space $E$ with compensator $\mathfrak{q}(\mathrm{d} u, \mathrm{~d} x)=\mathrm{d} u \otimes \lambda(\mathrm{d} x)$, where $\lambda$ is a $\sigma-$ finite measure on $E$ having no atom.
We also denote the spot variance-covariance process as $c_t := \sigma_t \sigma_t^*$, which takes values in the set $\mathcal{M}_d^+$ of nonnegative definite {\color{DG} $d\times d$} {\color{DR} symmetric} matrices. {\color{Blue} As common} in the literature, $c$ is called the spot volatility of the process. We assume $c$ is also an It\^o semimartingale following the dynamics:
\begin{equation} \label{eq:sigma}
{\color{DG} c_{t}^{\ell m}=c_0^{\ell m} + {\int_0^t{\tilde{\mu}^{\ell m}_{s}} \mathrm{d} s+\int_{0}^{t}\tilde{\sigma}^{\ell m}_{s} \mathrm{d}W_{s} }, \quad 1\leq \ell, m\leq d,}
\end{equation}
where {$W:=\{W_t\}_{t\geq0}$} is the same $d$-dimensional Brownian Motion driving the dynamics of $X$. {\color{Red} The assumption \eqref{eq:sigma} is common in the literature (see, e.g., the monographs \cite{JacodProtter,jacodaitsahalia}), and was also assumed in the work of \cite{jacod2013quarticity}{\color{Blue} )}.} {\color{DG} For each $1\leq \ell, m\leq d$, $\{\tilde{\mu}^{\ell m}_t\}_{t\geq{}0}$ is adapted locally bounded {\color{DR} and}  $\{\tilde{\sigma}^{\ell m}_t\}_{t\geq{}0}$ is adapted c\`adl\`ag.}

We now state the main assumption on the process $X$:
\begin{assumption} \label{process}
The process $X$ follows the dynamics (\ref{eq:X}) with $c_t=\sigma_t\sigma_t^*$ satisfying (\ref{eq:sigma}) and, for some $r \in[0,1)$, a measurable function $\Gamma_m: E \rightarrow \mathbb{R}_{+}$, a constant $C_m<\infty$, and a localizing sequence of stopping times $\left(\tau_m\right)_{m \geq 1}$ such that $\tau_m \rightarrow \infty$, we have
$$
t \in\left[0, \tau_m\right] \Longrightarrow\left\{\begin{array}{l}
\left|\mu_t\right|+\left|\sigma_t\right|+\left|\tilde{\mu}_t\right|+\left|\tilde{\sigma}_t\right| \leq C_m, \\
|\delta(t, z)| \wedge 1 \leq \Gamma_m(z), \quad \text { where } \int \Gamma_m(z)^r \lambda(d z)<\infty.
\end{array}\right.
$$
\end{assumption}
The parameter $r$ determines the jump activity of the process: the larger $r$ is, the more active or frequent are the small jumps of the process. When $r=0$, the process {\color{Blue} exhibits} {\color{DR} finitely} many jumps in any bounded time interval (in that case, we say that the jumps are of finite activity). When $r<1$, the jump component of the process is of bounded variation, which is a standard {\color{Blue} constraint} for the truncated realized quadratic variation to achieve efficiency (\cite{JacodProtter}).

Next, we give the conditions on the kernel function $K$ {\color{DG} for which we need to define the function}
\[
	{\color{DG} L(t) := \int_{t}^{\infty} K(u) d u \mathbf{1}_{\{t>0\}}-\int_{-\infty}^{t} K(u) d u \mathbf{1}_{\{t \leq 0\}}.}
\]
As explained in the introduction, we aim to consider two-sided kernels of unbounded support.  It is important to remark that this type of {\color{Blue} kernel} presents some challenges to our analysis (see comments before \eqref{eq:left}).
\begin{assumption} \label{kernel}
The kernel function $K : \mathbb{R} \rightarrow \mathbb{R}$ is bounded such that
\begin{enumerate}
\item[{\rm a.}] $\int K(x) d x=1$;
\item[{\rm b.}] $K$ is Lipschitz and piecewise $C^2$  on its support $(A,B)$,  where $A < B$, $A \in [-\infty, 0]$, and $B \in [0, \infty]$;
\item[{\rm c.}] (i) $\int|K(x)  x| d x<\infty$; (ii) $K(x) |x|^{2+\epsilon} \rightarrow 0$, $K^{\prime}(x) |x|^{2+\epsilon} \rightarrow 0$, with $\epsilon>1/2 $ as $|x| \rightarrow \infty$ ; (iii) $\int |K^{(j)}(x)|dx < \infty$, for $j=0,1,2,3$; (iv) $\int_0^\infty\int_{w}^\infty |K'(x)|dxdw<\infty$; (v) $\int K^{2}(x)dx<\infty$; {\color{DG} (vi) $\int_{-\infty}^{\infty}|L(u)u|du<\infty$};
\item[{\rm d.}] (i) $V_{-\infty}^{\infty}(K^{(j)}):=\lim_{m\to\infty}V_{-m}^{m}(K^{(j)})<\infty$, where $V_{-m}^{m}(K^{(j)})$ is the total variation of $K^{(j)}$ on the interval $[-m,m]$, $j=0,1,2,3$; (ii) $V_{-\infty}^{\infty}(K^2)<\infty$;
\item[{\rm e.}] $K^{\prime}$ and $K^{\prime\prime}$ exist piecewise on $(A,B)\subset\mathbb{R}$ and are absolutely continuous on {\color{Red} their} corresponding domain $\operatorname{dom}(K^{\prime})$ and $\operatorname{dom}(K^{\prime\prime})\subset(A,B)$, respectively;
\item[{\rm f.}] $K^{\prime\prime\prime}$ exists piecewise on $(A,B)\subset\mathbb{R}$.
\end{enumerate}
\end{assumption}

\begin{remark}
	{\color{Red} Even though the list of assumptions above seems long, it is important to remark that all kernels commonly used in practice satisfy these conditions. The assumptions on the regularity of the high-order derivatives $K^{(j)}$ for $j=2,3$ are only used to show the following limit:
	\begin{equation}\label{eq:v5}
    \sqrt{\Delta_{n}} \sum_{i=1}^n\left[\bar{K}_{n}\left(t_i\right)-1\right] g\left(c_i^n\right)\longrightarrow \theta \int_{-\infty}^{0}L(u)du g\left(c_{0}\right)-\theta \int_{0}^{\infty}L(u)du g\left(c_{T}\right),
\end{equation}
where recall that $\bar{K}_{n}\left(t_i\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right)$. This limit is easy to deal with in the case of a uniform kernel because, in that case, $\bar{K}_{n}\left(t_i\right)\equiv 1$ for almost all $i$ except if $i$ is close to $0$ or $n$.}
\end{remark}

For an arbitrary process $\{U_{t}\}_{t\geq{}0}$,  a filtration  $\mathbb{F}=(\mathcal{F}_{t})_{t \geq 0}$,  and  a given time span $\Delta_{n}>0$, we shall use the notation
\[
	U_{i}^{n}:=U_{i\Delta_{n}},\qquad
	\Delta_{i}^{n} U:=U_{i}^{n}-U_{i-1}^{n}, \qquad \mathcal{F}^n_{i} := \mathcal{F}_{i \Delta_n}.
\]
Stable convergence in law is denoted by \( \stackrel{st}{\longrightarrow}\). See (2.2.4) in \cite{JacodProtter} for the definition of this type of convergence. As usual, $a_{n}\sim c_{n}$  means that $a_{n}/c_{n}\to{}1$ as $n\to\infty$.

Throughout, we assume that we sample the process $X$ in (\ref{eq:X}) at discrete times   $t_{i} :=t_{i, n} :=i \Delta_{n}$,  where $\Delta_{n} :=T / n$ and $T\in(0,\infty)$ is a given fixed time horizon.
To estimate the spot volatility $c_{t_i}$, we adopt a Nadaraya-Watson type of kernel estimator of the form:
\begin{align}\label{MDKNN}
	{\hat{c}^{n,X}_{i}} :=\frac{\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right)\left(\Delta_{j}^{n} X\right)\left(\Delta_{j}^{n} X\right)^*}{\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right)},
\end{align}
where \(K_{{b}}(x):=K(x / b) / b\), $k_{n}\in\mathbb{N}$, and ${b_n}:={k}_{n} \Delta_n$ is the bandwidth of the kernel function. This is a variation of the estimator
\begin{align}\label{MDKNNb}
	{\tilde{c}^{n,X}_{\tau}} :=\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-\tau\right)\left(\Delta_{j}^{n} X\right)\left(\Delta_{j}^{n} X\right)^*,
\end{align}
first proposed in \cite{fan2008spot} and \cite{kristensen2010nonparametric},  and studied in recent papers such as \cite{FigLi}, \cite{FigWu}. The denominator
\[
	\bar{K}_{n}\left(t_i\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-t_i\right),
\]
{\color{DR} in \eqref{MDKNN} is known to} improve the performance \eqref{MDKNNb} for values of $\tau=t_i$ near the edge of the estimation interval $[0,T]$. The variation \eqref{MDKNN} was also considered in \cite{kristensen2010nonparametric} and \cite{FigLi} {\color{DR} when $t_i$ is {replaced} with a fix time $\tau\in(0,T)$}. In that case, $\bar{K}_{n}\left(\tau\right):=\Delta_{n}\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-\tau\right)\to1$ and the asymptotic behavior of ${\tilde{c}^{n,X}_{\tau}}$ and
\[
	{\hat{c}^{n,X}_{\tau}} :=\sum_{j=1}^{n} K_{{k}_{n}\Delta_n}\left(t_{j-1}-\tau\right)\left(\Delta_{j}^{n} X\right)\left(\Delta_{j}^{n} X\right)^*/\bar{K}_{n}\left(\tau\right)
\]
are equivalent. However, for our estimator (see \eqref{Target0a} below), we need to consider $t_i$ for all $i=1,\dots, n$ and, thus, it is not direct that we can simply omit $\bar{K}_{n}\left(t_i\right)$ (in fact, its presence helps with the analysis).

To handle jumps in the process $X$, we also consider a truncated version {\color{DR} of \eqref{MDKNN}:}
\begin{equation}\label{estivoltrun}
\hat{c}_i^{n,X,v_n}:=\frac{\sum_{j=1}^n K_{k_n \Delta_n}\left(t_{j-1}-t_{i}\right)\left(\Delta_j X\right)\left(\Delta_j X\right)^*\mathbbm{1}_{\left\{\left\|\Delta_j X\right\| \leq v_n\right\}}}{\Delta_n \sum_{j=1}^n K_{k_{n}\Delta_n}\left(t_{j-1}-t_{i}\right)}{,}
\end{equation}
with a suitable sequence of truncation levels $v_n\in(0, \infty]$. {\color{DR} More specifically, we consider two types of conditions on $v_n$. In the case that $X$ is continuous, we take
    \begin{equation}\label{moreassum1}
        v_n=\alpha \Delta_n^{\varpi}, \quad \text { for some arbitrary fixed } \alpha\in(0,\infty]\text { and } \varpi < \frac{1}{2},
    \end{equation}
while if $X$ is discontinuous, we take
\begin{equation}\label{moreassum}
	v_n=\alpha \Delta_n^{\varpi}, \quad \text { for some fixed } \alpha\in(0,\infty)\text{ and } \varpi \in\left[\frac{2{\color{DG} \ell}-1}{2(2{\color{DG} \ell}-r)}, \frac{1}{2}\right),
\end{equation}
where $\ell\geq{}4$ is a positive integer that will be specified below. Note that when $\alpha=\infty$ in \eqref{moreassum1}, we recover the non-truncated version (\ref{MDKNN}) and, in that case, the value of $\varpi$ is irrelevant.}
For simplicity, we {\color{DR} often omit the dependence on $X$ and $v_n$ on \eqref{MDKNN} and \eqref{estivoltrun} and simply write $\hat{c}_{i}^{n}$}.


{Our estimation target is the integrated volatility functional, defined as
\begin{equation}\label{Target0a}
V(g)_{t}:=\int_{0}^{t} g\left(c_{s}\right) d s,
\end{equation}
for a continuous function $g:\mathcal{M}_d^+\to\mathbb{R}$.
As proposed in \cite{jacod2013quarticity}, a natural estimator for this integral is given by {\color{DR} $\Delta_{n} \sum_{i=1}^{\left[t / \Delta_{n}\right]} g\left(\hat{c}_{i}^{n}\right)$, which can be interpreted as a Riemann sum approximation of \eqref{Target0a}.} However, as we shall see, it will be easier to derive the asymptotic properties of the estimator
\begin{equation} \label{eq:simple_estimator}
V(g)^{n}_{t}:=\Delta_{n} \sum_{i=1}^{\left[t / \Delta_{n}\right]} g\left(\hat{c}_{i}^{n}\right){\bar{K}_n\left(t_i\right)}.
\end{equation}
Furthermore, we can see \eqref{eq:simple_estimator} as a weighted average version of $\Delta_{n} \sum_{i=1}^{\left[t / \Delta_{n}\right]} g\left(\hat{c}_{i}^{n}\right)$, which gives less weight to $g\left(\hat{c}_{i}^{n}\right)$ for values of $t_i$ near $0$ and $T$, where the estimator $\hat{c}_{i}^{n}$ is less accurate due to edge effects.}

{Our first result characterizes the asymptotic behavior of the estimation error of $V(g)^{n}_{t}$} when the bandwidth converges to $0$ at {an} "optimal" rate {\color{DR} (c.f. Theorem \eqref{clt_suboptimal} below and the remarks before).}

\begin{theorem} \label{clt}
Let $g$ be a $C^{3}$ function such that, for all $x \in \mathbb{R}^{d \times d}$,
\begin{equation} \label{eq:g_prime_bound}
|g(x)| \leq C\left(1+\|x\|^{\color{DG} \ell}\right),\quad\left|  \partial _{p_1 q_1,\dots,p_j q_j}g(x)\right| \leq C\left(1+\|x\|^{{\color{DG} \ell}-j}\right), \quad j=1,2,3{,}
\end{equation} for $p_1, q_1,\dots,p_j, q_j\in\{1, \ldots, d\}$ and for some constants $C > 0$ and ${\color{DG} \ell} \ge 4$. Assume $k_n$ satisfies
\begin{equation} \label{eq: k_n}
k_{n} \sim \frac{\theta}{\sqrt{\Delta_{n}}} \qquad (i.e.,\;\; \text{the bandwidth}\ {{b}_n}:=k_n\Delta_n\sim \theta \sqrt{\Delta_n}),
\end{equation}
for some $\theta \in(0, \infty)$. Then, under (\ref{eq:X}) and Assumptions \ref{process} and \ref{kernel}, the following assertions hold true:
\begin{itemize}
    \item[a)] Suppose that {\color{DR} $X$ is continuous} (i.e., $\delta\equiv 0$ in (\ref{eq:X})). Then, the estimator \eqref{eq:simple_estimator} of the integrated volatility functional $V(g)_T$ with {\color{DR} $\hat{c}_{i}^{n}$ given by (\ref{estivoltrun}), {under} the condition \eqref{moreassum1}\footnote{In particular, by taking $\alpha=\infty$, we recover the untruncated estimator \eqref{MDKNN}.},}
    satisfies the following stable convergence in law, as $n \rightarrow \infty$:
\begin{equation}\label{eq:convergence_in_law}
{\color{DR} \frac{1}{\sqrt{\Delta_{n}}}\Big(V(g)_T^{n}-V(g)_T\Big)
\stackrel{st}{\longrightarrow} A^{1} + A^{2}+A^{3}+Z},
\end{equation}
where $Z$ is {\color{Blue} an} r.v. defined on an extension $\left(\widetilde{\Omega}, \widetilde{\mathcal{F}},(\widetilde{\mathcal{F}}_{t})_{t \geq 0}, \widetilde{\mathbb{P}}\right)$ of $\left(\Omega, \mathcal{F}, (\mathcal{F}_{t})_{t \geq 0}, \mathbb{P}\right)$, which, conditionally on $\mathcal{F}$, is a centered Gaussian variable with variance
\begin{equation}\label{VarForNL}
{\widetilde{\mathbb{E}}\left(Z^{2} \mid \mathcal{F}\right) =\sum_{j,k,l,m} \int_0^T\partial_{jk}g(c_s)\partial_{lm}g(c_s)(c_s^{jl}c_s^{km}+c_s^{jm}c_s^{kl})ds,}
\end{equation}
{and
\begin{equation}\label{MDfnOfAs}
\begin{split}
A^{1}&:=-\theta \int_{0}^{\infty}L(u)dug\left(c_{0}\right)+\theta \int_{-\infty}^{0}L(u)dug\left(c_{T}\right), \\
A^{2}&:=\frac{1}{2\theta} \int_0^T \sum_{p,q,u,v=1}^{d}\partial^{2}_{p q, u v} g\left(c_s\right) \check{c}_s^{p q, u v} d s \int_{-\infty}^{\infty} K^2(u) d u,\\
A^{3}&:=\frac{\theta}{2} \int_0^T \sum_{p,q,u,v=1}^{d}\partial^{2}_{p q, u v} g\left(c_s\right)\tilde{c}_s^{p q, u v}  d s\left[\int_{-\infty}^{\infty} L(u)^2 d u+\left(\int_{-\infty}^{0}-\int_{0}^{\infty}\right) L(u) d u\right]{\color{DR} ,}
\end{split}
\end{equation}
where} $\check{c}_{s}^{pq,uv}:=c_{s}^{pu}c_{s}^{qv} +c_{s}^{pv}c_{s}^{qu}$ and $\tilde{c}^{pq,uv}_{s}:=\sum_{r=1}^d\tilde{\sigma}_s^{p q, r} \tilde{\sigma}_s^{u v, r}$.
\item[b)]
{\color{DR} When $X$ is discontinuous, we {\color{DR} still} have (\ref{eq:convergence_in_law}) for the estimator $V\left(g\right)^{n}_{T}$ with the truncated version (\ref{estivoltrun}) provided that the stronger condition \eqref{moreassum} is satisfied}.
\end{itemize}
\end{theorem}

\begin{remark}
Note that the bias terms depend on the kernel in a rather intricate manner (especially, the term $A^3$). The {\color{Blue} signs} of the coefficients are always the same regardless of $K$. For instance, $\int_{-\infty}^{\infty} L(u)^2 d u+\left(\int_{-\infty}^{0}-\int_{0}^{\infty}\right) L(u) d u$  is negative because $|L(u)|\leq 1$ for all $u$, and $L(u)\leq{}0$ ($L(u)\geq 0$) for any $u<0$ ($u>0$). Thus, $A^3$ goes in {\color{Blue} the} opposite direction to $A^2$.
\end{remark}
\begin{remark}
If we use the uniform {\color{DR} right-sided} kernel $K(x)={\bf 1}_{[0,1)}(x)$ in our estimators, {as} in \cite{jacod2013quarticity,jacod2015estimation}, then
\begin{equation}\label{eq:K_bar_equal_1}
\bar{K}_n\left(t_{i}\right)=\Delta_{n}\sum_{j=1}^{n}{K_{b_n}}\left(t_{j-1}-t_{i}\right)=\Delta_{n}\sum_{j=i+1}^{{k_n}+i}\frac{1}{{b_n}}=1, \quad \text{for} \ 0\leq i\leq n-k_{n}.
\end{equation}
{\color{DR} In that case,} the estimator in \cite{jacod2015estimation}, defined as
\[
	V\left(g\right)^{n,JR}_{T}:=\Delta_{n}\sum_{i=1}^{n-k_{n}-1}g\left(\hat{c}_{i}^{n,JR}\right),
\]
where
\begin{equation}\label{vola_JR}
    \hat{c}_i^{n,{JR}}:=\frac{1}{k_n \Delta_n} \sum_{j=0}^{k_n-1} \left(\Delta_{i+j}^n X\right)\left(\Delta_{i+j}^n X\right)^* \mathbbm{1}_{\left\{\left\|\Delta_{i+j}^n X\right\| \leq v_n\right\}},
\end{equation}
can be written {\color{DR} in terms of our estimator \eqref{eq:simple_estimator} as follows:}
\begin{equation}\label{eq:right_V_estimator}
    V\left(g\right)^{n,JR}_{T}=V\left(g\right)^{n}_{T}-\Delta_{n}\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)+\Delta_n g\left(\hat{c}_{0}^{n}\right).
\end{equation}
{The formula \eqref{eq:right_V_estimator} holds because} (\ref{eq:K_bar_equal_1}) ensures that $\hat{c}_{i}^{n}=\hat{c}_{i+1}^{n,JR}$, for $0\leq i\leq n-k_{n}$. {It is easy to see that, for {the uniform kernel $K(x)={\bf 1}_{[0,1)}(x)$,} $\Delta_{n}^{1/2}\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)\to \frac{\theta}{2}g\left(c_{T}\right)$ and the bias terms of Theorem \ref{clt} reduce to
\begin{equation}
\begin{array}{l}
A^{1}:=-\frac{\theta}{2}g\left(c_{0}\right){,} \\
A^{2}:=\frac{1}{2\theta} \int_0^T \sum_{p,q,u,v=1}^{d}\partial^{2}_{p q, u v} g\left(c_s\right) \check{c}_s^{p q, u v} d s,\\
A^{3}:=-\frac{\theta}{12} \int_0^T \sum_{p,q,u,v=1}^{d}\partial^{2}_{p q, u v} g\left(c_s\right)\tilde{c}_s^{p q, u v}  d s.
\end{array}
\end{equation}
Thus, Theorem \ref{clt} implies that
\begin{equation} \label{eq:clt_right}
\frac{1}{\sqrt{\Delta_{n}}}\left(V(g)_T^{n,JR}-V(g)_T\right) \stackrel{st}{\longrightarrow}\tilde{A}^{1} +  A^{2}+A^{3}+Z,
\end{equation} where $\tilde{A}^{1}=-\frac{\theta}{2}\left[g\left(c_{0}\right)+g\left(c_{T}\right)\right]$. We then recover the same result as \cite{jacod2015estimation}.}
\end{remark}


To make this CLT ``feasible'' {\color{Blue} in practical work}, the bias terms need to be estimated.
{For $A^1$, we use that
\begin{equation} \label{eq:2side_edge_bias}
k_n \sqrt{\Delta_n}g\left(\hat{c}^n_1 \right) \stackrel{\mathbb{P}}{\to} \theta g(c_0), \quad k_n \sqrt{\Delta_n}g\left(\hat{c}^n_n \right) \stackrel{\mathbb{P}}{\to} \theta g(c_T).
\end{equation}
For $A^2$, we can apply Theorem \ref{clt} to the functional
$$
    h(x):=\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g(x)\left(x^{pu}x^{qv}+x^{pv}x^{qu}\right),
$$
for any $x\in \mathbb{R}^{d\times d}$, to get
\begin{equation}\label{eq:h}
V\left(h\right)^{n}_{T}=\Delta_{n}\sum_{i=1}^{n}h\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)\stackrel{\mathbb{P}}{\longrightarrow}  \int_0^T \sum_{p,q,u,v=1}^{d}\partial^{2}_{p q, u v} g\left(c_s\right) \check{c}_s^{p q, u v} d s.
\end{equation}
The term $A^3$ is the hardest to estimate since it involves the {\color{Blue} `volvol'} $\tilde{c}$. In the case of a uniform {\color{DR} right-sided} kernel, {\color{DR} $K(x)={\bf 1}_{[0,1)}(\chi)$}, \cite{jacod2015estimation} proposed the estimator
\begin{equation}\label{An3}
    {\color{DR} \widehat{A}^{n, 3}}:=\frac{\sqrt{\Delta_n}}{8} \sum_{i=1}^{n-2 k_n+1} \sum_{j, k, l, m} \partial_{j k, l m}^2 g(\hat{c}_i^n)\left(\hat{c}_{i+k_n}^{n,j k}-\hat{c}_i^{n,j k}\right)\left(\hat{c}_{i+k_n}^{n,l m}-\hat{c}_i^{n,l m}\right),
\end{equation}
with $\hat{c}$ given by (\ref{vola_JR}) and showed that  ${\color{DR} \widehat{A}^{n, 3}}$ converges to a linear combination of the bias terms $A^{2}$ and $A^3$. The following theorem shows the corresponding result for a general kernel function. The proof of Theorem \ref{thm2.2} is {given} in Appendix \ref{proofthm2.2}.}
\begin{theorem}\label{thm2.2}
    Under the notations {\color{DR} and conditions} of Theorem \ref{clt}, we have, {as $n\to\infty$},
\begin{equation}\label{eq:bias_estimates}
    \begin{aligned}
    {\color{DR} \widehat{A}^{n, 3}}\stackrel{\mathbb{P}}{\longrightarrow}&\,\frac{1}{4\theta} \int_0^T \sum_{j, k, l, m}\partial^2_{jk, lm} g\left(c_s\right) \check{c}_s^{jk, lm} d s
    {\int^{\infty}_{-\infty} K(z)(K(z)-K(z-1))d z}\\
    &+\frac{\theta}{4} \int_0^T \sum_{j, k, l, m}\partial^2_{jk, lm} g\left(c_s\right) \tilde{c}_s^{jk, lm} d s\\
    &\quad\times \Bigg({\int_{-\infty}^{\infty} L(z)(L(z)-L(z-1))d z+\frac{1}{2}-\int_{-\infty}^{\infty}K(z)(|z|\wedge{}1)dz\Bigg),}
    \end{aligned}
\end{equation}
when $\hat{c}$ is given by (\ref{estivoltrun}) {\color{DR} with the condition \eqref{moreassum}}. Furthermore, {\color{DR} when $X$ is continuous (i.e., $\delta\equiv 0$ in (\ref{eq:X})),} \eqref{eq:bias_estimates} still holds {\color{DR} under the weaker condition \eqref{moreassum1}} {(in particular, we can use the untruncated estimator \eqref{MDKNNb} by taking $\alpha=\infty$)}.
\end{theorem}
Under the Assumption (\ref{eq: k_n}), based on  (\ref{eq:2side_edge_bias}), (\ref{eq:h}), and (\ref{eq:bias_estimates}), we {\color{DR} then} propose {a bias-corrected} estimator of the form:
\begin{equation} \label{eq:unbiased_V}
\begin{split}
{\widetilde{V}(g)^{n}_T} &:= V\left(g\right)^{n}_{T} +k_n \Delta_n g\left(\hat{c}_1^n\right) {\int_0^{\infty} L(u)d u}-k_n \Delta_n g\left(\hat{c}_n^n\right){\int_{-\infty}^0 L(u) d u}\\
&\quad-\frac{\Delta_n}{4} \sum_{i=1}^{n-2k_n+1}\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g\left(\hat{c}_i^n\right)\left(\hat{c}_{i+k_n}^{n, p q}-\hat{c}_i^{n, p q}\right)\left(\hat{c}_{i+k_n}^{n, u v}-\hat{c}_i^{n, u v}\right)C_{K,1}\\
&\quad+\frac{1}{2k_n}V\left(h\right)^{n}_{T}\left[C_{K,2}-\int K^2(u) du\right]{\color{DR} ,}
\end{split}
\end{equation}
where
$$
C_{K,1}:=\frac{\int_{-\infty}^{\infty} L(z)^2 d z+\left(\int_{-\infty}^{0}-\int_{0}^{\infty}\right) L(z) d z}{{\int_{-\infty}^{\infty} L(z)(L(z)-L(z-1))d z+\frac{1}{2}-\int_{-\infty}^{\infty}K(z)(|z|\wedge{}1)dz}}{},
$$
and
$$
C_{K,2}:={\int^{\infty}_{-\infty} K(z)(K(z)-K(z-1))d zC_{K,1}}.
$$
Next, we show that indeed ${\widetilde{V}(g)^{n}_T}$ enjoys a centered CLT. Its simple proof is given in Section \ref{SmplCorUnb}.
\begin{corollary}\label{BiasCorrectedCLT0}
Under the assumptions {\color{DR} and notation} of Theorem \ref{clt},  {as $n\to\infty$}, we have \begin{equation}\label{UnBsCLTb}
\frac{1}{\sqrt{\Delta_n}}\left( {\widetilde{V}(g)^{n}_T} - V(g)\right) \stackrel{st}{\to} Z,
\end{equation}
{\color{DR} where ${\widetilde{V}(g)^{n}_T}$ is defined as in \eqref{eq:unbiased_V} with $\hat{c}$ given by (\ref{estivoltrun}) {\color{DR} with the} condition \eqref{moreassum}. Furthermore, {\color{DR} when $X$ is continuous, \eqref{UnBsCLTb} holds under the weaker condition {\eqref{moreassum1}}.}}
\end{corollary}





\begin{remark}
\cite{jacod2015estimation} proposed the following bias-corrected estimator:
\begin{align}\label{CRJEN}
    \widetilde{V}(g)^{n,{JR}}_T&:=V(g)^{n,{JR}}_T+\frac{k_n \Delta_n}{2}\left(g\left(\hat{c}_1^{n,{JR}}\right)+g\left(\hat{c}^{n,{JR}}_{n-k_n+1}\right)\right)-\frac{3}{4k_n}V(h)_T^{n,{JR}}\\
    &\quad+\frac{\Delta_n}{8} \sum_{i=1}^{n-2k_n+1}\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g\left(\hat{c}_i^{n,{JR}}\right)\left(\hat{c}_{i+k_n}^{n,{JR}, p q}-\hat{c}_i^{n,{JR}, p q}\right)\left(\hat{c}_{i+k_n}^{n,{JR}, u v}-\hat{c}_i^{n,{JR}, u v}\right),
    \nonumber
\end{align}
where $\hat{c}_i^{n,{JR}}$ and     $V\left(g\right)^{n,JR}_{T}$
are defined as in (\ref{vola_JR}) and (\ref{eq:right_V_estimator}), respectively. {\color{DR} There is a connection between \eqref{CRJEN} and our estimator \eqref{eq:unbiased_V} taking a right-sided uniform kernel $K(x)={\bf 1}_{[0,1)}(x)$. Indeed, note that, with that kernel,} (\ref{eq:unbiased_V}) reduces to
\begin{equation}
\begin{split}
{\widetilde{V}(g)^{n}_T} &:= V\left(g\right)^{n}_{T} +\frac{k_n \Delta_n}{2} g\left(\hat{c}_1^n\right)-\frac{3}{4k_n}V\left(h\right)^{n}_{T}\\
&\quad+\frac{\Delta_n}{8} \sum_{i=1}^{n-2k_n+1}\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g\left(\hat{c}_i^n\right)\left(\hat{c}_{i+k_n}^{n, p q}-\hat{c}_i^{n, p q}\right)\left(\hat{c}_{i+k_n}^{n, u v}-\hat{c}_i^{n, u v}\right).
\end{split}
\end{equation}
Then, by {\color{DR} \eqref{eq:K_bar_equal_1} and} (\ref{eq:right_V_estimator}), we {\color{DR} conclude that}
\begin{equation}\label{relationtoJR}
\begin{split}
{ \widetilde{V}(g)^{n}_T} &=\widetilde{V}(g)^{n,{JR}}_T +\Delta_{n}\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)-\Delta_n g\left(\hat{c}_{0}^{n}\right)-\frac{k_n \Delta_n}{2} g\left(\hat{c}^{n,{JR}}_{n-k_n+1}\right)\\
&\quad-\frac{3}{4k_n}\Delta_n\sum_{i=n-k_{n}-1}^{n}h\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)+\frac{3\Delta_n}{4k_n} h\left(\hat{c}_{0}^{n}\right)\\
&\quad+\frac{\Delta_n}{8}\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g\left(\hat{c}_{n-2k_n+1}^{n}\right)\left(\hat{c}_{n-k_n+1}^{n, p q}-\hat{c}_{n-2k_n+1}^{n, p q}\right)\left(\hat{c}_{n-k_n+1}^{n, u v}-\hat{c}_{n-2k_n+1}^{n, u v}\right)\\
&\quad-\frac{\Delta_n}{8}\sum_{p, q, u, v=1}^d \partial_{p q, u v}^2 g\left(\hat{c}_{0}^{n}\right)\left(\hat{c}_{k_n}^{n, p q}-\hat{c}_0^{n, p q}\right)\left(\hat{c}_{k_n}^{n, u v}-\hat{c}_0^{n, u v}\right).
\end{split}
\end{equation}
Since
$$
\sqrt{\Delta_n}\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)\stackrel{\mathbb{P}}{\longrightarrow} \frac{\theta}{2}g\left(c_T\right),\qquad
\frac{k_n \sqrt{\Delta_n}}{2} g\left(\hat{c}^{n,{JR}}_{n-k_n+1}\right)\stackrel{\mathbb{P}}{\longrightarrow} \frac{\theta}{2}g\left(c_T\right),
$$
we know that $\Delta_n\sum_{i=n-k_{n}-1}^{n}g\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)-\frac{k_n \Delta_n}{2} g\left(\hat{c}^{n,{JR}}_{n-k_n+1}\right)=o_{p}\left(\sqrt{\Delta_n}\right)$. Also note that
$$\sqrt{\Delta_n}\sum_{i=n-k_{n}-1}^{n}h\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)\stackrel{\mathbb{P}}{\longrightarrow} \frac{\theta}{2}h\left(c_T\right),
$$
which implies that $\frac{3}{4k_n}\Delta_n\sum_{i=n-k_{n}-1}^{n}h\left(\hat{c}_{i}^{n}\right)\bar{K}\left(t_{i}\right)=O_{p}\left(\frac{\sqrt{\Delta_n}}{k_n}\right)$. The remaining terms in (\ref{relationtoJR}) can {\color{DR} directly} be proved {\color{DR} to be} $o_{p}\left(\sqrt{\Delta_n}\right)$. We then conclude that
$$
\frac{1}{\sqrt{\Delta_n}}\left(\widetilde{V}(g)^{n}_T -\widetilde{V}(g)^{n,{JR}}_T\right)\stackrel{\mathbb{P}}{\longrightarrow} 0.
$$
Hence, by Corollary \ref{BiasCorrectedCLT0}, we can recover the {\color{DR} stable} convergence in law for the bias-corrected estimator in \cite{jacod2015estimation}:
 \begin{equation}
\frac{1}{\sqrt{\Delta_n}}\left( \widetilde{V}(g)^{n,{JR}}_T - V(g)\right) \stackrel{st}{\to} Z.
\end{equation}
\end{remark}


As mentioned above, the bias term $A^3 $ contains the `volvol' $\tilde{c} $ and the estimator of $A^3$ in (\ref{eq:bias_estimates}) could introduce extra variance to the {asymptotically} unbiased estimator (\ref{eq:unbiased_V}). {An alternative} way to eliminate {the} bias is {undersmoothing, i.e., through the selection of a bandwidth sequence $b_n$ converging to $0$ at a faster rate than $\sqrt{\Delta_n}$ ($b_n \ll \sqrt{\Delta_n}$) or, equivalently, picking $k_n$ such that $\sqrt{\Delta_n}k_n  \stackrel{n \to \infty}{\longrightarrow} 0$ (meaning that $\theta = 0$ in (\ref{eq: k_n}))}. This is the approach put forward in \cite{jacod2013quarticity}.
{In that case, it is expected that the bias terms $A^1$ and $A^3$ will vanish and the bias term $A^2 $ will dominate}. {{\color{DR} Following the same idea,} we can devise an asymptotically unbiased {\color{DR} estimator as follows}:
\begin{equation}\label{v_bar}
    \bar{V}(g)^n_T := \Delta_n \sum_{i= 1}^{n} \left( g(\hat{c}^n_i) - \frac{1}{2k_n} h(\hat{c}^n_i) \int K^2(u) du\right)\bar{K}\left(t_{i}\right).
\end{equation}
The following result establishes the asymptotic behavior of $\bar{V}(g)^n_T $. {Its proof is {given} in Appendix \ref{PrfSubs23}.}
\begin{theorem} \label{clt_suboptimal}
Assume $k_n$ satisfies \begin{equation} \label{eq:k_n_suboptimal}
k_{n}^2 \Delta_n \to 0,\quad  k_n^3 \Delta_n \to \infty{.}
\end{equation}
Then, under (\ref{eq:X}) and Assumptions \ref{process} and \ref{kernel}, we have the following stable convergence in law
\begin{equation} \label{eq:convergence_in_law_unbiased}
\frac{1}{\sqrt{\Delta_{n}}}\left(\bar{V}(g)_T^{n}-V(g)_T\right) \stackrel{st}{\longrightarrow} Z,
\end{equation}where $Z$ is as in Theorem \ref{clt} {\color{DR} and ${\bar{V}(g)^{n}_T}$ is defined as in \eqref{v_bar} with $\hat{c}$ given by (\ref{estivoltrun}) under condition \eqref{moreassum}. Furthermore, {\color{DR} when $X$ is continuous, \eqref{eq:convergence_in_law_unbiased} holds under the weaker condition \eqref{moreassum1}.}}

\end{theorem}

\section{Simulation Study}\label{SimulationSect}
In this section we analyze the performance of {\color{DR} our} kernel-based estimators. To isolate the effect of the bandwidth on the estimators' performance, we only consider continuous processes (i.e., $\delta\equiv 0$ in \eqref{eq:X} and, thus, $X\equiv X'$) and focus on the untruncated estimator \eqref{MDKNN} and the corresponding estimators $V(g)^{n}_{t}$ of \eqref{eq:simple_estimator} and $\widetilde{V}(g)^{n}_T$ of \eqref{eq:unbiased_V}.

\subsection{Simulation design}
We consider the data generating model in \cite{li2019efficient}:
\begin{equation} \label{eq:expOU_model}
\left\{\begin{array}{l}
d X_t = \sigma_t d W_t  \\
\sigma_t = exp(-1.6+F_t), \quad d F_t = -5 F_t dt + 2dB_t,
\end{array}\right.
\end{equation} with $\mathbb{E}\left[d W_t d B_t\right]  = -0.75 dt $. Throughout, we use the function $g(c) = c^2$, which corresponds to the integrated quarticity{,} and also the function $g(c) = \log(c)$.  We simulate data for  $T= 5$ days and  $T = 21$ days, and assume the process $X$ is observed once {\color{Red} every 1 minute or every 5 minutes, with 6.5 trading hours per day, for all the results of Subsection \ref{Sec33}, but not for Subsection  \ref{Sec32}, where we assume a frequency of 1-second}.

\subsection{Validity of the asymptotic theory}\label{Sec32}
We first show the finite sample behavior of the estimator is consistent with the Central Limit Theorem of Theorem \ref{clt}. Under the notation of Theorem \ref{clt}, we choose $\color{DR}{ \theta = 0.2}$, $g(x) = x^2 $, and estimate the functional volatility using the exponential kernel $K(x) = \frac{1}{2}e^{-|x|}$. The choice of $\theta$ was obtained from the formula $b^*=\theta^*\sqrt{\Delta_n}$, where $b^*$ was chosen using the iterative method in \cite{FigLi}. The sample frequency of the data is 1 second. Based on 5,000 simulated paths from the model (\ref{eq:expOU_model}), we {calculate the
standardized estimation error for each path $i=1,\dots, 5000$} according to Theorem \ref{clt}: \begin{equation}
z_i = \frac{\left(\Delta_n\sum_{j=1}^n g(\hat{c}^{n,i}_j)\bar{K}\left(t_j\right) - \int_0^T g(c^{i}_s) ds\right)/\sqrt{\Delta_n} - (A^{1,i}+A^{2,i} +A^{3,i})}{\sqrt{{\rm Var}(Z)^{i}}},
\end{equation}
where $c^i_s = (\sigma^i_s)^2 $, $\{\hat{c}^{n,i}_j\}_j $ are the kernel spot volatility estimates of the $i$th simulated path, and the variance ${\rm Var}(Z)^{i}=2 \int_0^T(g''(c^i_s)c^i_s)^2ds$ and the biases $A^{\ell,i}$ ($\ell=1,2,3$), as defined in \eqref{MDfnOfAs}, are computed via Riemann sum approximations  using the true simulated volatility $\{c^{n,i}_j\}_{j=0,\dots,n}$ and volvol observations $\{\tilde{c}^{n,i}_j\}_{j=0,\dots,n}$.
We then compare the histogram of the standardized estimates $z_i$ with a standard normal distribution (see Figure \ref{fig:hist_exp}). As shown therein, the distribution of the  estimation error is consistent with the asymptotic result in Theorem \ref{clt}.\\%, though the convergence rate may be slow.\\
\begin{figure}[h]
    \centering
    \includegraphics[scale = 0.6]{thm2_1.png}
    \caption{Histogram of $z_i$ and the density of standard normal distribution in Theorem \ref{clt}.  }
    \label{fig:hist_exp}
\end{figure}
Similarly, we verify the behavior of the bias-corrected estimator proposed in Corollary \ref{BiasCorrectedCLT0}. With the same value of $\theta$, $g(x)$, exponential kernel, and simulated data, the standardized bias-corrected estimator takes the form:
\begin{equation}
    z_i = \frac{\frac{1}{\sqrt{\Delta_n}}\left(\widetilde{V}(g)^{n,i}_T - \int_0^T g(c^{i}_s) ds\right)}{\sqrt{Var(Z^i)}}.
\end{equation}
The corresponding histogram is shown in Figure \ref{fig:2}, which again suggest the CLT stated in Corollary \ref{BiasCorrectedCLT0} is valid.
\begin{figure}[h]
    \centering
    \includegraphics[scale = 0.6]{thm2_2.png}
    \caption{Histogram of $z_i$ and the density of standard normal distribution in Corollary \ref{BiasCorrectedCLT0}.  }
    \label{fig:2}
\end{figure}
\begin{table}[h!]
\setlength\tabcolsep{3pt}
\begin{center}
\begin{tabular}{  c c c c c c c c c }
\hline
& \multicolumn{4}{c}{$T=21$ days } &\multicolumn{4}{c}{$T=5$ days }\\
\hline
  & \multicolumn{2}{c}{$\Delta_n = 5 $ minutes } &\multicolumn{2}{c}{$\Delta_n = 1$ minute }& \multicolumn{2}{c}{$\Delta_n = 5 $ minutes } &\multicolumn{2}{c}{$\Delta_n = 1$ minute }\\
  & BIAS(\%) & RMSE(\%) & BIAS(\%) & RMSE(\%) & BIAS(\%) & RMSE(\%) & BIAS(\%) & RMSE(\%)\\
\hline
\multicolumn{9}{c}{exponential kernel}\\
$V ^n_T$ & 5.884 & 9.890 & 2.345 & 3.941 & 13.291 & 17.627 & 6.718 &8.537 \\
$\widetilde{V}^{n}_T$ & 1.557 & 8.214 & 0.042 & 3.336 & 1.531 & 16.096 & 0.949 &6.861 \\
\multicolumn{9}{c}{two-sided uniform kernel}\\
$V ^n_T$ & 4.824 & 9.497 & 1.821 & 3.729 & 11.332 & 16.557 & 5.783 & 7.903\\
$\widetilde{V}^{n}_T$ & 1.804 & 8.357 & 0.084 & 3.352 & 2.288 & 15.843 & 1.183 & 6.874\\
\multicolumn{9}{c}{right-sided uniform kernel}\\
$V ^{n}_T$ & 1.727 & 8.237 & 0.208 & 3.384 & 8.380 & 16.339 & 4.638 & 7.086\\
$\widetilde{V}^{n}_T$ & 1.083 & 8.334 & -0.267 & 3.361 & 1.393 & 16.509 & 0.665 & 6.965\\
\multicolumn{9}{c}{Jacod \& Rosenbaum's Estimators}\\
$V ^{n,JR}_T$ & 5.096 & 11.127 & 1.656 & 4.115 & 19.714 & 23.313 & 8.574 & 10.459\\
$\widetilde{V}^{n,JR}_T$ & 1.795 & 8.469 & 0.061 & 3.370 & 2.002 & 16.562 & 1.139 &7.067 \\
\multicolumn{9}{c}{Jackknife Estimators}\\
$TS_n$ & 4.365 & 8.223 & 2.195 & 3.435 & 3.802 & 14.834 & 2.991 &6.779 \\
$TS_n'$ & 1.323 & 8.162 & 0.213 & 3.308 & -8.019 & 20.895 & -3.855 &8.714 \\
$MS_n$ & 1.989 & 8.220 & 0.098 & 3.347 & 1.618 & 16.641 & 1.041 &6.930 \\
\hline
\end{tabular}
\caption{$g(c) = c^2$.}
\label{table:Jackknife}
\end{center}
\end{table}
\begin{table}[h!]
\setlength\tabcolsep{3pt}
\begin{center}
\begin{tabular}{  c c c c c c c c c }
\hline
& \multicolumn{4}{c}{$T=21$ days } &\multicolumn{4}{c}{$T=5$ days }\\
\hline
  & \multicolumn{2}{c}{$\Delta_n = 5 $ minutes } &\multicolumn{2}{c}{$\Delta_n = 1$ minute }& \multicolumn{2}{c}{$\Delta_n = 5 $ minutes } &\multicolumn{2}{c}{$\Delta_n = 1$ minute }\\
  & BIAS(\%) & RMSE(\%) & BIAS(\%) & RMSE(\%) & BIAS(\%) & RMSE(\%) & BIAS(\%) & RMSE(\%)\\
\hline
\multicolumn{9}{c}{exponential kernel}\\
$V ^n_T$ & 3.922 & 4.716 & 1.819 & 2.192 & 11.166 & 11.597 & 5.035 & 5.335\\
$\widetilde{V}^{n}_T$ & 0.486 & 2.218 & 0.120 & 0.908 & 1.007 & 4.427 & 0.169 & 2.111\\
\multicolumn{9}{c}{two-sided uniform kernel}\\
$V ^n_T$ & 3.313 & 4.016 & 1.463 & 1.818 & 9.620 & 10.219 & 4.309 & 4.673\\
$\widetilde{V}^{n}_T$ & 0.300 & 2.790 & -0.025 & 0.993 & 1.127 & 4.515 & 0.190 & 2.122\\
\multicolumn{9}{c}{right-sided uniform kernel}\\
$V ^{n}_T$ & 1.224 & 2.373 & 0.485 & 1.045 & 7.605 & 8.281 & 3.379 & 3.800\\
$\widetilde{V}^{n}_T$ & -0.753 & 3.440 & -0.432 & 1.277 & 0.091 & 4.463 & -0.216 & 2.122\\
\multicolumn{9}{c}{Jacod \& Rosenbaum's Estimators}\\
$V ^{n,{JR}}_T$ & 4.239 & 4.565 & 1.790 & 2.018 & 17.397 & 17.148 & 7.763 &7.785 \\
$\widetilde{V}^{n,JR}_T$ & -0.059 & 3.228 & -0.126 & 1.118 & 0.820 & 4.478 & 0.135 & 2.124\\
\multicolumn{9}{c}{Jackknife Estimators}\\
$TS_n$ & 2.269 & 3.962 & 1.296 & 2.152 & 1.913 & 5.099 & 0.957 &2.472 \\
$TS_n'$ & -0.425 & 2.931 & -0.142 & 1.472 & -9.518 & 10.686 & -5.169 &5.547 \\
$MS_n$ & 0.681 & 2.487 & 0.282 & 1.064 & 0.988 & 4.540 & 0.202 & 2.143\\
\hline
\end{tabular}
\caption{$g(c) = log(c)$.}
\label{table:Jackknife_log}
\end{center}
\end{table}
\subsection{Comparison with Jackknife estimator}\label{Sec33}
We now compare our bias-corrected estimator $\widetilde{V}(g)_T^{n}$ to the Jackknife estimator of \cite{li2019efficient}. For the Jackknife method, we consider the two-scale estimator $TS_n$, its boundary-adjusted version $TS_n'${,} and the multiscale estimator $MS_n$ as defined in \cite{li2019efficient}, Section 3.1. The parameters for the Jackknife estimators are chosen according to \cite{li2019efficient}. {\color{DR} By running numerous simulations, we determine that the values chosen by \cite{li2019efficient} broadly optimize their estimator's performance}. For the multiscale estimator $MS_n$, $\left(\psi_1, \psi_2, \psi_3\right)=(-2.5,8,-4.5)$, and $\left(k_{1, n}, k_{2, n}, k_{3, n}\right)=(15,30,45)$ for data sampled every 5 minutes, or $(40,80,120)$ for data sampled every minute. For the two-scale estimator, $\left(\psi_1, \psi_2\right)=(-1,2)$ along with the same $\left(k_{1, n}, k_{2, n}\right)$ as above. For our kernel estimator, we use 3 kernels: a two-sided exponential kernel $K(x)=\frac{1}{2}e^{-|x|}$, a two-sided uniform kernel $K(x)=\frac{1}{2}{\bf 1}_{[-1,1]}(x)$, and a right-sided uniform kernel $K(x)={\bf 1}_{[0,1]}(x)$. We analyze the performance of the simple biased estimator $V(g)^n_T$ in (\ref{eq:simple_estimator}) and the bias-corrected estimator $\widetilde{V}(g)^{n}_T$ in (\ref{eq:unbiased_V}). We also consider the simple biased estimator $V(g)^{n,{JR}}_T$ and the bias-corrected estimator $\tilde{V}(g)^{n,{JR}}_T$ of \eqref{eq:right_V_estimator} and \eqref{CRJEN}, respectively, which were proposed in \cite{jacod2015estimation}.

The bandwidths of all the kernel estimators are chosen using the iterative method in \cite{FigLi}, Section 5. We iterate the method only once. For a two-sided kernel, the integrated volatility of volatility used when updating the bandwidth is estimated by the Two-time Scale Realized Volatility (TSRV) estimator proposed in \cite{FigLi}. For the right-sided uniform kernel,  the TSRV estimator no longer applies. We then use the realized variance estimator on the estimated volatility path as a replacement.
As finite-sample corrections, in (\ref{eq:unbiased_V}), at {each} $t_i$, we replace $\int_{-\infty}^{\infty}K^2(u) du $ with $
\sum_{j=1}^n K^2_h(t_i - t_j) \Delta_n h$, and replace $\int_{-\infty}^{\infty}L(s) ds  $ and $\int_{-\infty}^{\infty}L^2(s)ds $ with $
 \frac{\Delta_n}{b}\sum_{j=1}^n L\left(\frac{t_{j-1}-t_i}{b} \right)$ and $\frac{\Delta_n}{b}\sum_{j=1}^n L^2\left(\frac{t_j-t_i}{b} \right)$, respectively, where
 $ L\left(\frac{t_j - t_i}{b}  \right) =
 -\sum_{l=1}^{j}K_{{b}}(t_{l-1}-t_i) $ if $j \leq i$ and $\sum_{l=j+1}^{n}K_{{b}}(t_{l-1}-t_i)$ if $ j > i$.


In Table \ref{table:Jackknife}, we report the relative bias and relative root-mean-square-errors in percentage unit for the considered estimators. Our results of the Jackknife estimators are consistent with the ones in \cite{li2019efficient}, Table 1 therein. As shown in Table \ref{table:Jackknife}, the kernel estimators {\color{DR} exhibit a similar performance to} the Jackknife estimators. However, the Jackknife estimators have more tuning parameters, while ours requires only one tuning parameter (namely, $\theta$).
As we will show in Section \ref{sensitivity}, the {\color{Red} most important} advantage of our estimators {\color{Red} lies} on their stability relative to the choice of the bandwidth. That is, our estimators seem to exhibit satisfactory performance for a large range of the tuning parameters, while other estimators are more sensitive to choosing `good' values for their tuning parameters.

{\color{DR} Returning to} the results of Table \ref{table:Jackknife},} we also observe that in some cases the kernel estimators have smaller relative root mean square error (RMSE) compared with the Jackknife estimators, especially when the  time span $T$ is short ($T = 5$ days). Intuitively, the superior performance of the spot volatility estimators with exponential  kernel (shown in \cite{FigLi}, Section 7) {\color{DR} is expected to be more evident} when the time span is relatively short, e.g. 5 days.  For reference, we include the tables for the function $g(c) = log(c)$ in Table \ref{table:Jackknife_log}. These tables confirm our previous conclusion.
It is worth noting that in general, the bias-corrected kernel estimator $\widetilde{V}^{n}_T$ has {\color{DR} significantly} smaller RMSE compared with the simple biased estimator $V(g)^{n}_T$, which also happens among the Jacod \& Rosenbaum’s estimators, $V(g)^{n,{JR}}_T$ and $\widetilde{V}^{n,JR}_T$.

\subsubsection{{\color{DR}  Bandwidth Sensitivity}\label{sensitivity}}
In this subsection, we investigate how sensitive the estimators are relative to the bandwidth parameter. For the Jackknife estimators, we focus on the boundary adjusted multi-scale estimator $\text{MS}$ (which, as shown in Tables \ref{table:Jackknife} and \ref{table:Jackknife_log}, typically has the best performance among all {\color{Red} the considered} Jackknife estimators), and for the kernel estimators, we {\color{DR} first consider an} exponential kernel. We {\color{DR} try} a sequence of bandwidths $k_n$ for the kernel estimator, and a sequence of triplets {$(k_{1,n}, k_{2,n}, k_{3,n}) = ([\frac{k_n}{2}], k_n, [\frac{3k_n}{2}])$} and $(\psi_1, \psi_2, \psi_3) = (-2.5,8,-4.5)$ for $MS$, which include the parameters chosen in \cite{li2019efficient} {\color{DR} when $k_n=30$ (see Table 1 therein)}. The results are shown in the Figures \ref{fig:bdw_square_21}-\ref{fig:bdw_log_5}, where $\text{S0 exp}$ denotes the simple biased exponential kernel estimator in (\ref{eq:simple_estimator})  and $\text{S3 exp}$ denotes the bias-corrected estimators with exponential kernel in (\ref{eq:unbiased_V}). {\color{DR} We can conclude that the estimator $\text{S3 exp}$ exhibits the most stable or consistent performance for} either 5 or 21 days data with 1 or 5 minutes sampling frequency, and both $g(c)=c^2$ or $\log (c)$. {\color{DR} In other words, the estimator's values are {\color{DG} in general} relatively invariant to changes in the bandwidth, {\color{DR} especially} when working with low frequency and $g(c)=c^2$. Under those conditions, the others estimators require more careful tuning of the bandwidth in order to show good performance.
We can also observe that the bias-corrected exponential kernel estimator almost always outperforms the boundary-adjusted multi-scale Jackknife estimator $\text{MS}$ and the simple biased exponential kernel estimator $\text{S0 exp}$.}

{\color{DR} Finally,} in Figure \ref{fig:mixed}, we analyze the bandwidth stability of three estimators: the bias-corrected estimators \eqref{eq:unbiased_V} with exponential and right-sided uniform kernels and the bias-corrected estimator \eqref{CRJEN} proposed in \cite{jacod2015estimation}. Again, since we are not introducing jumps in the process $X$, the estimator $\hat{c}_i^{n}$ used for \eqref{eq:unbiased_V} and the estimator $\hat{c}_i^{n,{JR}}$ used for \eqref{CRJEN} do not have any truncation (i.e., we fix $v_n=\infty$ in \eqref{estivoltrun} and \eqref{vola_JR}). Figure \ref{fig:mixed} shows that our proposed estimator $\widetilde{V}^{n}_T$ with either exponential or right-sided uniform kernel almost always outperforms $\widetilde{V}^{n, JR}_T$. The exponential kernel appears to be less sensitive to the choice of $k_n$ than the right-sided uniform kernel. The estimator $\widetilde{V}^{n, JR}_T$ also seems to require a careful tuning of the bandwidth in order to achieve a good performance.
\begin{comment}
We also notice $S0$ has similar behavior compared to the boundary adjusted two-scale estimator $TS'$, yet the estimator $S0$ doesn't require any bias correction and even outperforms the $TS'$ estimator under a wide range of bandwidth, except for 21 days data with 5 min sampling frequency, $g(c)=\log (c)$.
\end{comment}


\begin{figure}
\centering
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{21day5min.png}
  \caption{data with 5 minute sample frequency. }
  \label{fig:sub1}
\end{subfigure}
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{21day1min.png}
  \caption{data with 1 minute sample frequency.}
  \label{fig:sub2}
\end{subfigure}
\caption{ RMSE(\%) v.s. the bandwidth $k_n$, for 21 days data, $g(c) = c^2$ and 1000 simulated path. $S0 $ exp denotes the simple biased estimator with exponential kernel and $S3$ exp denotes the unbiased estimator with exponential kernel. }
\label{fig:bdw_square_21}
\end{figure}


\begin{figure}
\centering
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{5day5min.png}
  \caption{data with 5 minute sample frequency. }
  \label{fig:sub1}
\end{subfigure}
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{5day1min.png}
  \caption{data with 1 minute sample frequency.}
  \label{fig:sub2}
\end{subfigure}
\caption{ RMSE(\%) v.s. the bandwidth $k_n$, for 5 days data, $g(c) = c^2$ and 1000 simulated path. $S0 $ exp denotes the simple biased estimator with exponential kernel and $S3$ exp denotes the unbiased estimator with exponential kernel. }
\label{fig:bdw_square_5}
\end{figure}

\begin{figure}
\centering
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{21day5minlog.png}
  \caption{data with 5 minute sample frequency. }
  \label{fig:sub1}
\end{subfigure}
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{21day1minlog.png}
  \caption{data with 1 minute sample frequency.}
  \label{fig:sub2}
\end{subfigure}
\caption{ RMSE(\%) v.s. the bandwidth $k_n$, for 21 days data, $g(c) = log(c)$ and 1000 simulated path. $S0$   exp denotes the simple biased estimator with exponential kernel and $S3 $ exp denotes the unbiased estimator with exponential kernel. }
\label{fig:bdw_log_21}
\end{figure}
\begin{figure}
\centering
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{5day5minlog.png}
  \caption{data with 5 minute sample frequency. }
  \label{fig:sub1}
\end{subfigure}
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{5day1minlog.png}
  \caption{data with 1 minute sample frequency.}
  \label{fig:sub2}
\end{subfigure}
\caption{ RMSE(\%) v.s. the bandwidth $k_n$, for 5 days data, $g(c) = log(c)$ and 1000 simulated path. $S0$   exp denotes the simple biased estimator with exponential kernel and $S3 $ exp denotes the unbiased estimator with exponential kernel. }
\label{fig:bdw_log_5}
\end{figure}

\begin{figure}
\centering
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{5day5minmixednew.png}
  \caption{data with 5 minute sample frequency. }
  \label{fig:sub1}
\end{subfigure}
\begin{subfigure}{.5\textwidth}
  \centering
  \includegraphics[width=1\linewidth]{5day1minmixednew.png}
  \caption{data with 1 minute sample frequency.}
  \label{fig:sub2}
\end{subfigure}
\caption{ RMSE(\%) v.s. the bandwidth $k_n$, for 5 days data, $g(c) = c^2$ and 1000 simulated path. $S3$ exp and $S3$ unif denotes the unbiased estimator with exponential kernel and right-sided uniform kernel, respectively. $S3$ JR corresponds to the bias-corrected estimator (\ref{CRJEN}) in \cite{jacod2015estimation}.}
\label{fig:mixed}
\end{figure}



\newpage