The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
89,516 characters
Kernel Estimation of Spot Volatility with Microstructure Noise Using Pre-Averaging
\maketitle
\begin{abstract}
We first revisit the problem of estimating the spot volatility of an It\^o semimartingale using a kernel estimator. We prove a Central Limit Theorem with an optimal convergence rate for a general two-sided kernel. Next, we introduce a new pre-averaging/kernel estimator for spot volatility to handle the microstructure noise of ultra high-frequency observations. We prove a Central Limit Theorem for the estimation error with an optimal rate and study the optimal selection of the bandwidth and kernel functions. We show that the pre-averaging/kernel estimator's asymptotic variance is minimal for two-sided exponential kernels, hence, justifying the need of working with kernels of unbounded support as opposed to the most commonly used uniform kernel. We also develop a feasible implementation of the proposed estimators with optimal bandwidth. Monte Carlo experiments confirm the superior performance of the devised method.
\medskip
\noindent
\textbf{AMS 2000 subject classifications}: 62M09, 62G05.
\smallskip
\noindent
\textbf{Keywords and Phrases}: Spot volatility estimation; kernel estimation; pre-averaging; microstructure noise; bandwidth selection; kernel function selection.
\end{abstract} \hspace{10pt}
\section{Introduction}
{It\^o semimartingale} models for the dynamics of asset returns have been widely used in financial econometrics. Such a process takes the form
\begin{equation}\label{ItoModel00}
dX_t = \mu_t dt + \sigma_t d W_t+{dJ_{t}},
\end{equation}
where {\(\{W_t\}_{t\geq{}0}\)} is a standard Brownian motion and {$\{J_{t}\}_{t\geq{}0}$ is the jump component}. {The spot volatility \(\sigma_t\) is a key feature of the model as it} plays a crucial rule in option pricing, portfolio management, and financial risk management. Since last decade, there has been some growing interest in the estimation of volatility due to the wide availability of high frequency data. In this work, we are concerned with spot volatility estimation in an It\^o semimartingale model {\color{Blue} via kernel smoothing}. This is one of the most widely used nonparametric methods in statistics, dating back to the seminal work of \cite{rosenblatt1956remarks} and \cite{parzen} (see also the monograph \cite{wand1995monographs}).
One of the earliest works on kernel-based estimation of spot volatility dates back to \cite{foster1994continuous_optKernel}, where they studied a weighted rolling window estimator, which is essentially a kernel estimator with compact support. Asymptotic normality was established under abstract conditions that {\color{Blue} were not directly stated in terms of the coefficients of the It\^o semimartingale (\ref{ItoModel00})}. Concretely, they worked with a time series discretization of the model (\ref{ItoModel00}).
\cite{fan2008spot} established the asymptotic normality for a general kernel estimator, {\color{Blue} this time working directly with the model (\ref{ItoModel00}) {\color{Blue} under relatively mild conditions on the coefficients,} but without jumps. {\color{Blue} However, the} result therein also required a certain condition on the convergence rate of the bandwidth to $0$}, which allowed them to neglect the ``target error" coming from approximating the spot volatility by a kernel {\color{Blue} weighted volatility}.
As a result, {\color{Blue} the convergence rates of the estimators were suboptimal (see Section 6 in \cite{FigLi} for more details)}.
{\color{Blue} \cite{kristensen2010nonparametric} also proved a Central Limit Theorem (CLT) for kernel-based estimators under the absence of jumps and a non-leverage condition (i.e., $\sigma$ and $W$ were assumed to be independent). \cite{Yuetal2014} generalized Kristensen's result by allowing a jump component of finite activity (FA), but still assuming non-leverage effects. \cite{mancini2015estimation} considered more general It\^o semimartingales, but again FA jumps. All these works {\color{Blue} only considered CLTs with \emph{suboptimal convergence rates}.}}
\cite{alvarez2012estimation} proposed an estimator of \(\sigma^p_t \) by considering forward finite difference approximations of the realized power variation process of order \(p\), which is essentially a forward-looking kernel estimator with uniform kernel. \cite{JacodProtter} {\color{Blue} (Section 13.3 therein)} considered both backward and forward finite difference approximations of the realized quadratic variation. Both works obtained the best possible convergence rates for their CLTs {\color{Blue} for a rather} general It\^o semimartingale model ({\color{Blue} in the case \cite{JacodProtter}, also including jumps}). We also refer to \cite{jacodaitsahalia}, {\color{Blue} Chapter 8},
for an extensive review of the relevant literature.
More recently, \cite{FigLi} studied the leading order terms of the mean-square error (MSE) of kernel-based estimators {for continuous It\^o semimartingales} under a certain local condition on the covariance function of the spot variance $\sigma_{t}^{2}$, which covers not only Brownian driven volatilities but also those driven by fractional Brownian motion and other Gaussian processes. Using the asymptotics for the MSE, the optimal convergence rate was established and formulas for the optimal bandwidth and kernel functions were derived {\color{Blue} under a non-leverage condition}. CLTs for general {\color{Blue} right-sided} kernel estimators were also obtained (see also Remark 8.10 in \cite{jacodaitsahalia}, where a result for a general right-sided kernel with compact support was stated without proof).
One of the objectives of the present work is then to extend the results of \cite{FigLi} {\color{Blue} and \cite{Yuetal2014}\footnote{{\color{Blue} As explained above, \cite{Yuetal2014} established a CLT for a general two-sided kernel but with {\color{Blue} a} suboptimal converge rate, FA jumps, and non-leverage.}}}, and prove a CLT for a general \emph{two-sided} kernel {\color{Blue} of unbounded support,} {\color{Blue} with optimal convergence rate and in the presence of jumps and leverage {\color{Blue} effects}}. As {\color{Blue} proved} in this paper in greater generality, such kernels can have better performance than either one-sided or {compactly supported} kernels. Until now, this fact seems to have eluded the literature, which has almost exclusively focussed on uniform kernels.
While the results described in the previous paragraph are {important} for intermediate intraday frequencies (e.g., 1 to 5 minute), it is widely accepted that financial returns at ultra high-frequency are contaminated by market microstructure noise.
Specifically, high-frequency asset prices exhibit several stylized features, which cannot be accounted by It\^o semimartingales, such as clustering noises, bid/ask bounce effects, and roundoff errors {\color{Blue} (cf. \cite{Campbell}, Chapter 3, \cite{Zeng:2003}, \cite{jacodaitsahalia}, Chapter 2)}. Such discrepancies between macro and micro movements are typically modeled by an additive noise. The literature of statistical estimation methods under microstructure noise has grown extensively since last decade and is still a highly researched subject (see \cite{zhang2005tale}, \cite{HanLun}, \cite{Bandi}, \cite{MyklandZhang2012}, \cite{barndorff2008designing}, \cite{podolskij2009estimation}, and \cite{jacod2009microstructure} for a few seminal works in the area as well as the monograph \cite{jacodaitsahalia}).
Most of the existing literature on volatility estimation for high frequency data with microstructure noise has mainly focused on the estimation of {\color{Blue} the} integrated volatility or variance (IV), defined as $IV_{T}=\int_{0}^{T}\sigma_{t}^2dt$. \cite{zhang2005tale} showed that scaled by \((2n)^{-1}\), the realized variance estimator, the gold standard for IV estimation in the absence of microstructure noise, consistently estimates the variance of the microstructure noise, instead of the integrated volatility, as the sampling frequency \(n\) increases. There are several approaches to overcome this problem: the Two Scale Realized Variance (TSRV) estimator by \cite{zhang2005tale} and the efficient Multiscale Realized Variance by \cite{Zhang2006}; the Realized Kernel estimator by \cite{barndorff2008designing}; the pre-averaging method by \cite{podolskij2009estimation} and \cite{jacod2009microstructure}; and the Quasi-Maximun Likelihood Estimator (QMLE) by \cite{xiu2010quasi}.
Spot volatility estimation is often viewed as a byproduct of integrated volatility estimation {\color{Blue} since, in principle, we can recover the spot volatility $\sigma_t^2$ as a finite-difference approximation of an estimate of the integrated volatility}. Following this idea, \cite{zu2014estimating} {constructed} {\color{Blue} a} Two-Scale Realized Spot Variance (TSRSV) estimator based on the TSRV integrated variance estimator {\color{Blue} of \cite{zhang2005tale}}. They proved consistency and derived the asymptotic distribution of the estimation error with a convergence rate of \(n^{-1/12}\), which is suboptimal.
The second objective of our work is to construct a kernel based estimator of the spot volatility based on the pre-averaging integrated variance estimator of \cite{jacod2009microstructure}. The basic idea is simple and natural. If we denote $\widehat{IV}^{pre-av}_{t}$ the pre-averaging estimator of $IV_{t}=\int_{0}^{t}\sigma_{s}^{2}ds$, our estimators combines this with a kernel localization technique as follows:
\[
\hat{\sigma}_{t}^{2}=\int_{0}^{t}\frac{1}{b_n}K\left(\frac{s-t}{b_n}\right)d IV^{pre-av}_{s},
\]
where $K$ is a suitable kernel function and $b_n>0$ is the bandwidth, which should converge to $0$ at an appropriate rate. We establish the asymptotic mix normality of our estimator and identify two asymptotic regimes for two different bandwidth convergence regimes. One of those regimes yields the optimal convergence rate of \(n^{-1/8}\) for our estimator. It is important to point out that the asymptotic theory for the kernel/pre-averaging estimator cannot be derived from that for the pre-averaging integrated variance and also is substantially different and harder than that for kernel based estimators in the absence of microstructure noise.
{\color{Blue} Though combining pre-averaging and kernel {\color{Blue} smoothing} is a natural idea, to the best of knowledge, there are only two related results in the literature.} \cite{jacodaitsahalia} {\color{Blue} (Section 8.7 therein)}, stated, without proof, a stable convergence result of {\color{Blue} a pre-averaging estimator for the spot volatility of a continuous It\^o semimartingale}\footnote{{\color{Blue} The estimator therein is different from ours. Our estimator includes a debiasing term, which is {\color{Blue} omitted} in \cite{jacodaitsahalia}. Our Monte Carlo experiments show that such a correction is important in finite samples.}}, but only in the case of a one-sided uniform kernel $K(t)={\bf 1}_{[0,1]}(t)$ {\color{Blue} (see also \cite{chen2018inference} for a similar estimator)}. Here we consider {\color{Blue} a truncated version to handle the jumps and a general two-side kernel} (see below as to the need of considering such kernels). {\color{Blue} \cite{Yu2} also proposed {\color{Blue} a pre-averaging kernel estimator for the spot volatility, slightly} different from our estimator. They established asymptotic normality with suboptimal convergence rate for their untruncated estimator in the case of a continuous It\^o semimartingale, and for their truncated estimator in the presence of L\'evy jumps of bounded variation. {\color{Blue} In both situations, a non-leverage condition was adopted. In our case, we consider not only leverage effects, but also more general jump processes, not necessarily of L\'evy type and with no restriction in the index of jump activity, under both the suboptimal and optimal convergence rate regimes.}}
As an important application of our results, we study the problem of bandwidth and kernel function selection. Using our CLT, we first derive the optimal bandwidth and then the optimal kernel function (the one that minimizes the limiting variance) at the optimal rate.
We {then} show that the optimal kernel is a two-sided exponential or Laplace function $K(x)=\frac{1}{2}e^{-|x|}$. This fact justifies the necessity of developing the asymptotic theory for general kernels {\color{Blue} of unbounded support} over the more widely used uniform kernels. Again, we emphasize that our work is critical because it calls into question the indiscriminate use of uniform one-sided kernels in the literature. If we were constrained to compactly supported kernels {\color{Blue} in the suboptimal asymptotic regime}, a uniform kernel would be the best, but this is no longer the case if we allow kernels with unbounded support {\color{Blue} and/or consider an optimal convergence rate regime}. Similarly, two-sided kernels will perform better than one-sided, even if compactly supported.
The implementation of the optimum bandwidth (at the optimum rate) is more challenging because it involves the vol vol and the spot volatility itself. Hence, to implement it we develop a new method, which iteratively estimates the spot volatility, the vol vol, and the optimal bandwidth. Using Monte Carlo simulation, we compare our estimator with the TSRSV estimator of \cite{zu2014estimating} and show a significant improved accuracy. We also illustrate the improvement achieved by the optimal exponential kernel and the calibrated optimal bandwidth via our iterative method.
We finish the introduction by giving one more reason as to the importance of estimating the spot volatility.
As mentioned above, while spot volatility estimation can, at least conceptually, be seen as a byproduct of integrated variance estimation, interestingly enough, one can also use spot volatility estimation as an intermediate step toward the estimation of integrated volatility functionals of the form $I_{T}(g):=\int_{0}^{T}g(\sigma_{s}^{2})ds$. Specifically, once an estimator $\hat{\sigma}_{t}^{2}$ of $\sigma^{2}_{t}$ has been developed, one can naturally devise an estimator for $I_{T}(g)$ of the form $\hat{I}_{T}(g)=\Delta_{n}\sum_{i=1}^{n}g(\hat{\sigma}^{2}_{t_{i}})$, where $t_{i}=i\Delta_{n}$ and $\Delta_{n}=T/n$, followed by an appropriate bias correction adjustment. In the absence of noise, \cite{jacod2013quarticity}, \cite{li2019efficient}, and \cite{mykland2009inference} have developed methods for the estimation of these functionals (see also \cite{li2016generalized}, \cite{ait2019principal}, and \cite{li2017adaptive} for related methods and other applications thereof). Recently, \cite{chen2018inference} developed an estimator for $\hat{I}_{T}(g)$ based on a forward finite difference approximation of the standard pre-averaging estimator of the integrated variance.
The rest of the paper is organized as follows. Section \ref{setting} introduces the setting of the problem and the main result. Section \ref{optimal_para} shows an application of our main theorem: the optimal parameter and kernel selection. The simulations are provided in Section \ref{simulation}. {\color{Blue} Some conclusions are given in Section \ref{ConcludeSec}}. Proofs of our main results can be found in two appendices.
\section{The Setting, Estimator, and Main Results} \label{setting}
Throughout, {we consider an It\^o semimartingale of the form:}
\begin{equation} \label{eq:X}
\begin{split}
X_{t}=&X_0 + \int_0^t\mu_{s} ds +\int_0^t\sigma_{s} d W_{s} \\
&+\int_{0}^{t} \int_{E} \delta(s, z) \mathbbm{1}_{\{{|\delta(s, z)|} \leq 1\}}(\mathfrak{p}-\mathfrak{q})(d s, d z)
+\int_{0}^{t} \int_{E} \delta(s, z) \mathbbm{1}_{\{{|\delta(s, z)|}>1\}} \mathfrak{p}(d s, d z),
\end{split}
\end{equation}
where all stochastic processes ($\mu :=\left\{\mu_{t}\right\}_{t \geq 0}$, $\sigma :=\left\{\sigma_{t}\right\}_{t \geq 0}$, $W :=\left\{W_{t}\right\}_{t \geq 0}$, $\mathfrak{p}:=\{\mathfrak{p}(B):B\in\mathcal{B}(\mathbb{R}_{+}\times E)\}$) are defined on a complete filtered probability space $\left(\Omega^{(0)}, \mathcal{F}^{(0)}, \mathbb{F}^{(0)}, \mathbb{P}^{(0)}\right)$ with filtration $ \mathbb{F}^{(0)}=\big(\mathcal{F}^{(0)}_{t}\big)_{t \geq 0}$ and {are assumed to satisfy standard conditions for $X$ to be well-defined}. Here, $W$ is a standard Brownian Motion (BM) adapted to the filtration \(\mathbb{F}^{(0)}\), {and} $\mathfrak{p}$ is a Poisson random measure on $\mathbb{R}_{+} \times E$ for some arbitrary Polish space $E$ with compensator $\mathfrak{q}(\mathrm{d} u, \mathrm{d} x)=\mathrm{d} u \otimes \lambda(\mathrm{d} x)$, {where} \(\lambda\) is a \(\sigma-\) finite measure on \(E\) having no atom. {For further details regarding It\^o semimartingales, see Section 2.1.4 in \cite{JacodProtter}.}
We denote the spot variance process \(c_t = \sigma_t^2\) and assume it is also an It\^o semimartingale with the following {dynamics:}
\begin{equation} \label{eq:sigma}
c_{t}=c_0 + {\int_0^t{\tilde{\mu}_{s}} \mathrm{d} s+\int_{0}^{t}\tilde{\sigma}_{s} \mathrm{d}B_{s} + \int_{0}^{t}\int_{E} \tilde{\delta}(s, z)(\mathfrak{p}-\mathfrak{q})(\mathrm{d} s, \mathrm{d} z)},
\end{equation}
where {$B:=\{B_t\}_{t\geq0}$} is a standard Brownian Motion adapted to $\mathbb{F}^{(0)}$ {so that} $d\left<W,B\right>_{t}=\rho_{t} d t$. Here, $\{\tilde{\mu}_t\}_{t\geq{}0}$ is adapted locally bounded; \(\{\rho_t\}_{t\geq{}0}\) is adapted, locally bounded, c\`adl\`ag; $\{\tilde{\sigma}_t\}_{t\geq{}0}$ is adapted c\`adl\`ag and $\tilde{\delta}$ is a predictable function on ${\mathbb{R}_{+}} \times E$ satisfying standard conditions for the process above to be well-defined {(see \cite{JacodProtter})}.
We now state the {main assumption on the process \(X\)}:
\begin{assumption} \thlabel{X}
The process \(X\) satisfies (\ref{eq:X}) with \(c_t= \sigma^2_{t}\) satisfying (\ref{eq:sigma}) and, for some \(r \in [0,2]\), measurable functions $\Gamma_{m}, \Lambda_{m}:E\to\mathbb{R}_{+}$, {constants $C_{m}<\infty$,} and a localizing sequence of stopping times {$\left(\tau_{m}\right)_{m\geq{}1}$ such that $\tau_{m}\to\infty$}, we have
\begin{equation*}
t \in\left[0, \tau_{m}\right] \Longrightarrow\left\{\begin{array}{l}
\left|\mu_{t}\right|+\left|\sigma_{t}\right|+\left|\tilde{\mu_{t}}\right|+\left|\tilde{\sigma_{t}}\right| \leq C_m, \\
|\delta(t, z)| \wedge 1 \leq \Gamma_{m}(z), \quad \text{ where }\int \Gamma_{m}(z)^{r} \lambda(d z)<\infty,\\
\left|\tilde{\delta}(t, z)\right| \wedge 1 \leq \Lambda_{m}(z), \quad \text{ where }\int \Lambda_{m}(z)^{2} \lambda(d z)<\infty.
\end{array}\right.
\end{equation*}
\end{assumption}
The parameter $r$ plays a key role in our asymptotic results. In short, $r$ determines the jump activity of the process: the larger $r$ is, the more active or frequent are the small jumps of the process. When $r=1$, the process exhibit finite many jumps in any bounded time interval (in that case, we say that the jumps are of finite activity). When $r<1$, the jump component of the process is of bounded variation.
To establish {the} central limit theorem for the kernel estimator $\hat{c}_{t}$, we need some assumptions on the kernel.
\begin{assumption} \thlabel{kernel}
The kernel function $K : \mathbb{R} \rightarrow \mathbb{R}$ is bounded and
\begin{enumerate}
\item $\int K(x) d x=1$;
\item K is Lipschitz and piecewise $C^1$ on $(-\infty,\infty)$;
\item (i) $\int|K(x) x| d x<\infty$ ; (ii) $K(x) x^{2} \rightarrow 0$, as $|x| \rightarrow \infty$ ; (iii) $\int |K^{\prime}(x)|dx < \infty$.
\end{enumerate}
\end{assumption}
For an arbitrary process $\{U_{t}\}_{t\geq{}0}$ and a given time span $\Delta_{n}>0$, we shall use the notation
\[
U_{i}^{n}:=U_{i\Delta_{n}},\qquad
\Delta_{i}^{n} U:=U_{i}^{n}-U_{i-1}^{n}.
\]
Stable convergence in law is denoted by \( \stackrel{st}{\longrightarrow}\). See (2.2.4) in \cite{JacodProtter} for the definition of this type of convergence. As usual, $a_{n}\sim b_{n}$ means that $a_{n}/b_{n}\to{}1$ as $n\to\infty$.
Throughout the paper, we consider two settings: observations with and without market microstructure noise. In the absence of microstruture noise, we use standard kernel estimation, while to {handle} the noise we propose a type of pre-averaging kernel estimator. These two settings together with the main results are presented in the following two subsections.
\subsection{Observations without microstructure noise}
In this subsection, we assume that we can directly observe the process $X$ in (\ref{eq:X}) at discrete times $t_{i} :=t_{i, n} :=i \Delta_{n}$, where $\Delta_{n} :=T / n$ and $T\in(0,\infty)$ is a given fixed time horizon. We also consider a sequence of truncation levels $v_{n}$ satisfying
\begin{equation} \label{eq:trucation_para}
v_{n}=\alpha \Delta_{n}^{\varpi} \quad \text { for some } \alpha>0, \quad \varpi \in\left(0, \frac{1}{2}\right).
\end{equation}
To estimate the spot volatility $c_{\tau}$, at a given time $\tau \in (0,T)$, we adopt the kernel estimator, studied in \cite{fan2008spot,kristensen2010nonparametric} and {its truncated version}, studied in \cite{Yuetal2014} {\color{Blue} and \cite{mancini2015estimation}}:
\begin{align}\label{MDKNN}
{\hat{c}^{n}\left({m}_{n}\right)_{\tau}} &:=\sum_{i=1}^{n} K_{{m}_{n}\Delta_n}\left(t_{i-1}-\tau\right)\left(\Delta_{i}^{n} X\right)^{2},\\
\label{MDKNN_truncated}
{\hat{c}^{n}\left({m}_{n},v_n\right)_{\tau}} &:=\sum_{i=1}^{n} K_{{m}_{n}\Delta_n}\left(t_{i-1}-\tau\right)\left(\Delta_{i}^{n} X\right)^{2}\mathbbm{1}_{\left\{\left|\Delta_{i}^{n} X\right| \leq v_{n}\right\}},
\end{align}
where \(K_{b}(x):=K(x / b) / b\), $m_{n}\in\mathbb{N}$, and $b_n:={m}_{n} \Delta_n$ is the bandwidth of the kernel function\footnote{Here, ${m}_{n}$ is equivalent to $k_n$ in the Theorem 13.3.7 of \cite{JacodProtter}, while ${m}_{n}\Delta_n$ is equivalent to the bandwidth $h_n$ of \cite{FigLi}.}. The asymptotic behavior of this estimator with one-sided uniform kernels (i.e., $K(x)={\bf 1}_{[0,1]}(x)$ or $K(x)={\bf 1}_{[-1,0]}(x)$) was studied in \cite{JacodProtter}. {\cite{Yuetal2014} showed a CLT for (\ref{MDKNN_truncated}) at the suboptimal convergence rate ($\beta=0$) under a nonleverage condition (i.e., $d\left<W,B\right>_{t}=0$ in (\ref{eq:X})-(\ref{eq:sigma}) and compound Poisson jump component}.
In this part, we extend the results to general two-sided kernels with possibly unbounded support, {optimal rate, and more general type of It\^o semimartingale}. There is an important motivation for considering {general unbounded} kernels since, as proved in \cite{FigLi} {in the non-leverage case and without jumps}, exponential and some other nonuniform unbounded kernels can yield estimators with {significantly} better performance than those based on uniform kernels. {In Section \ref{optimal_para} below, we show that this is also true under the more general semimartingale model (\ref{eq:X})-(\ref{eq:sigma}).}
We now proceed to describe the limiting distribution of the estimation error of (\ref{MDKNN})-(\ref{MDKNN_truncated}). Let \(V, V^{\prime}\) be independent centered Gaussian variables, {independent of $\mathcal{F}^{(0)}$,} defined on a ``very good'' filtered extension {$\Big(\widetilde{\Omega}^{(0)}, \widetilde{\mathcal{F}}^{(0)},\big(\widetilde{\mathcal{F}}_{t}^{(0)}\big)_{t \geq 0}, \widetilde{\mathbb{P}}^{(0)}\Big)$
of $\Big(\Omega^{(0)}, \mathcal{F}^{(0)}, \big(\mathcal{F}_{t}^{(0)}\big)_{t \geq 0}, \mathbb{P}^{(0)}\Big)$} (see \cite{JacodProtter} for definition) such that
\begin{equation} \label{eq:Y_Yprime}
\operatorname{\mathbb{E}}\left( V^2\right) =2\int K^2(u) du, \quad \operatorname{\mathbb{E}}\left(V^{\prime 2} \right)= \int L^2(t)dt,
\end{equation} where \(L(t) = \int_{t}^{\infty} K(u) d u \mathbf{1}_{\{t>0\}}-\int_{-\infty}^{t} K(u) d u \mathbf{1}_{\{t \leq 0\}}\).
Next, let \(Z^{(0)}_{\tau}, Z^{\prime (0)}_{\tau}\) be defined as
\begin{equation} \label{eq:limiting_Z}
Z^{(0)}_{\tau} = c_{\tau} V, \quad Z^{\prime (0)}_{{\tau}} = {\tilde{\sigma}_{\tau} }V^{\prime}.
\end{equation}
Now we are ready to introduce our main theorem for a general kernel estimator in the absence of microstructure noise. The proof is given in Appendix \ref{PrfOfMnRsltTh1}.
\begin{theorem} \thlabel{thm_no_noise}
Let the sequence $\{{m}_{n}\}_{n\geq{}1}$ that controls the bandwidth of the kernel estimator be such that
${m}_{n} \rightarrow \infty$, ${m}_{n}\Delta_n\rightarrow 0$, and
\begin{equation}\label{IDBWN}
{\color{Blue} {m}_{n} \sqrt{\Delta_{n}} \rightarrow \beta, \quad\text{with}\quad \beta \in [0,\infty]}.
\end{equation} Then, under \thref{X,kernel} above, at a given time \(\tau \in [0,T],\)
we have: \begin{enumerate}[label=\alph*)]
\item If $X$ is continuous, both the truncated version (\ref{MDKNN_truncated}) and the non-truncated version (\ref{MDKNN}) satisfy the following stable convergence in law, as $n\to\infty$:
\begin{equation} \label{eq:stable_convergence}
\begin{aligned}
\rm (i) &\quad \sqrt{{m}_{n}}\left(\hat{c}^{{n}}_{\tau}-c_{\tau}\right)\stackrel{st}{\longrightarrow}Z^{(0)}_{\tau}+\beta Z_{\tau}^{\prime (0)},\quad \text{if}\quad \beta < \infty,\\
\rm (ii) &\quad \frac{1}{\sqrt{{m}_{n}\Delta_n}}\left(\hat{c}^{{n}}_{\tau}-c_{\tau}\right) \stackrel{st}{\longrightarrow}Z_{\tau}^{\prime (0)},\quad\text{if}\quad \beta =\infty,\end{aligned}
\end{equation} where \(Z^{(0)}_{\tau}, Z_{\tau}^{\prime (0)}\) are defined as in (\ref{eq:limiting_Z}).
\item {\color{Blue} Suppose
\begin{equation} \label{eq:m_n_range}
m_{n} \Delta_{n}^{a} \rightarrow \beta^{\prime} \in(0, \infty), \quad \text {where } a \in(0,1),
\end{equation}
so that \eqref{IDBWN} holds with $\beta=0$ when $a<1/2$, $\beta=\beta'\in(0,\infty)$ when $a=1/2$, {\color{Blue} or} $\beta=\infty$ when $a>1/2$. Then, when $X$ is discontinuous,} we have (\ref{eq:stable_convergence}) for the non-truncated version (\ref{MDKNN}), as soon as
\begin{equation*}
\text{either } r<\frac{4}{3}, \quad \text { or }\quad \frac{4}{3} \leq r<\frac{2}{1+a} \quad\left(\text{and then } a<\frac{1}{2}\right).
\end{equation*}
\item Under (\ref{eq:m_n_range}), when $X$ is discontinuous, we have (\ref{eq:stable_convergence}) for the truncated version, as soon as
\begin{equation} \label{eq:r_range_truncated}
r<\frac{2}{1+a \wedge(1-a)}, \quad \varpi>\frac{a\wedge(1-a)}{2(2-r)}.
\end{equation}
\end{enumerate}
\end{theorem}
\begin{remark}
{The CLTs above generalize the results in \cite{FigLi}, where only right-sided kernels were considered under the absence of jumps, in \cite{JacodProtter} and \cite{alvarez2012estimation}, where only one-sided uniform kernels (i.e., $K(x)={\bf 1}_{[0,1]}(x)$ or $K(x)={\bf 1}_{[-1,0]}(x)$) were studied, and in \cite{jacodaitsahalia}, where a CLT for a general right-sided kernel with compact support was stated without proof. The proof of Theorem \ref{thm_no_noise} is also different from that in} \cite{FigLi} and is based on the approach of \cite{JacodProtter}. {\color{Blue} The case with $\beta=0$ produces a CLT with convergence rate $m_n^{-1/2}$, which vanishes slower than $\Delta_n^{1/4}$, the optimal rate. In that case, our result generalizes \cite{fan2008spot}, \cite{kristensen2010nonparametric}, \cite{Yuetal2014}, and \cite{mancini2015estimation} by allowing jumps of both finite and infinite activity and {\color{Blue} dependence} between the volatility and the Brownian motion driving the log-return process $X$ ({\color{Blue} leverage effects}).}
\end{remark}
\begin{remark}
{\color{Blue} As stated by the points (b)-(c) above, in the presence of jumps, both estimators (\ref{MDKNN}) and \eqref{MDKNN_truncated} can attain the optimal convergence rate of $\Delta_n^{1/4}$, but only if the index of jump activity is less than $4/3$. In the presence of higher jump activity, the estimators can only achieved the suboptimal convergence rate of $m_n^{-1/2}\gg \Delta_n^{1/4}$ (case $a<1/2$ and $\beta=0$). It is worth noting the surprising fact that, even in the presence of jumps, the untruncated kernel estimator (\ref{MDKNN}) can still consistently estimate the spot volatility. In the case of finitely many jumps, we may explain this fact by noting that in a small local window, there could be at most a finite number of jumps, while, in the limit, there are increasingly more increments that do not contain jumps\footnote{{\color{Blue} We thank a referee for pointing out this interesting insight.}}. Nevertheless, in practice and for better finite sample performance, one typically would prefer the truncated version of the estimator.}
\end{remark}
\begin{remark}
{As explained in the introduction, it is critical to expand the results to general two-sided kernels of unbounded support since these kernels exhibit superior performance. For instance, in the suboptimal rate case ($\beta=0$), the kernel $K$ with support $[0,1]$ that minimizes the asymptotic variance $2\int K^{2}(u)du$ is the uniform kernel $K_{unif}(x)={\bf 1}_{[0,1]}(x)$ since, by Jensen's inequality, $\int_0^1 K^{2}(u)du\geq{}(\int_{0}^{1}K(x)dx)^{2}=1=\int_0^1 K_{unif}^{2}(u)du$. However, there are many other kernels that are two-sided or of unbounded support and that attain smaller variance, even in the suboptimal rate case $\beta=0$. For instance, both $K(x)=2^{-1}{\bf 1}_{[-1,1]}(x)$ and $K_{exp^+}(x)=e^{-x}{\bf 1}_{(0,\infty)}(x)$ are such that $\int K^{2}(u)du=1/2$. In the optimal rate case ($\beta\in(0,\infty)$), the optimal kernel is the two-sided exponential $K(x)=2^{-1}e^{-|x|}$ as shown in Subsection \ref{OptNoMicroSection} below.}
\end{remark}
\subsection{Observations in the presence of microstructure noise}
In this part, we assume that our observations of $X$ are contaminated by ``microstructure'' noise. That is, we assume we observe
\begin{equation} \label{SmplSchm0}
Y_{{t_i}}:=X_{t_i} + \epsilon_{t_i},
\end{equation}
where $\epsilon = \{\epsilon_t\}$ is the noise process and, as before, $t_{i} :=t_{i, n} :=i \Delta_{n}$, $0 \leq i \leq n$, with $\Delta_{n} :=T / n$ and a fixed time horizon $T\in(0,\infty)$.
We allow the noise $\epsilon $ to depend on $X$, but in such a way that, conditionally on the whole process $X$, $\{\epsilon_t\}_{t\geq 0} $ is a family of independent, centered random variables. More formally, following the framework of \cite{JacodProtter}, for each time $t$, we consider a transition probability $Q_{t}\left(\omega^{(0)}, d z\right)$ from $\left(\Omega^{(0)}, \mathcal{F}^{(0)}_{t}\right)$ into $(\mathbb{R},\mathcal{B}(\mathbb{R}))$, and the canonical process $\{\epsilon_{t}\}_{t\geq{}0}$ on $\mathbb{R}^{[0, \infty)}$ defined as $ \epsilon_{t}(\tilde{\omega})=\tilde{\omega}(t)$ for $t\geq{}0$ and $\tilde{\omega}\in\mathbb{R}^{[0, \infty)}$. Next, we construct a new probability space $\left(\mathbb{R}^{[0, \infty)}, \mathbb{B}, \sigma\left(\epsilon_{s} : s \in[0, t)\right), \mathbb{Q}\right) $, where $\mathbb{B}$ is the product Borel $\sigma$-field and $\mathbb{Q}=\otimes_{t \geq 0} Q_{t}$. We then define an enlarged filtered probability space $\left(\Omega, \mathcal{F},\left(\mathcal{F}_{t}\right)_{t \geq 0}, \mathbb{P}\right)$ and a filtration $(\mathcal{H}_t) $ as follows:
\begin{equation*}
\left\{\begin{array}{l}\Omega=\Omega^{(0)} \times {\mathbb{R}^{[0, \infty)}}, \\ {\mathcal{F}_{t}=\mathcal{F}^{(0)}_{t} \otimes \sigma\left(\epsilon_{s} : s \in[0,t)\right), \quad \mathcal{H}_{t}=\mathcal{F}^{(0)}\otimes \sigma\left(\epsilon_{s} : s \in[0,t)\right)} \\ {\mathbb{P}(d \omega^{(0)}, {d \tilde{\omega}} )=\mathbb{P}^{(0)}(d \omega^{(0)}) \mathbb{Q}(\omega^{(0)}, d {\tilde{\omega}}).}\end{array}\right.
\end{equation*}
Any variable or process in either $\Omega^{(0)} $ or $\mathbb{R}^{[0,\infty)} $ can be {extended} in the usual way to a variable or a process on $\Omega $.
We now state the assumptions on the $\mathcal{F}^{(0)} $-conditional law of the noise process as well as some slightly different assumptions on the spot variance process and kernel function.
\begin{assumption} \thlabel{noise}
All variables $\left(\epsilon_t: t \geq 0 \right)$ are independent conditionally on $\mathcal{F}^{(0)}$, and we have
\begin{itemize}
\item $\mathbb{E}\left(\left.\epsilon_{t} \right| \mathcal{F}^{(0)}\right)=0$,
\item For all $p>0$, the process $ \mathbb{E}\left(\left. \left|\epsilon_{t}\right|^{p} \right| \mathcal{F}^{(0)}\right)$ is $\left(\mathcal{F}^{(0)}_{t}\right)$-adapted and locally bounded,
\item The conditional variance process $\gamma_{t}= \mathbb{E}\left(\left. \left|\epsilon_{t}\right|^{2} \right| \mathcal{F}^{(0)}\right)$ is c\`adl\`ag.
\end{itemize}
\end{assumption}
Along the lines of \cite{JacodProtter} (originally proposed in \cite{jacod2009microstructure}), to construct the pre-averaging estimator, we need:
\begin{itemize}
\item[(i)] A sequence of positive integers $k_n$, which represent the length of the pre-averaging window, satisfying
\begin{equation}\label{AsympCndkn}
k_{n}=\frac{1}{\theta \sqrt{\Delta_{n}}}+\mathrm{o}\left(\frac{1}{\Delta_{n}^{1/4}}\right), \quad \mbox{ {for some} }\; \theta>0;
\end{equation}
\item[(ii)] A real-valued weight function $g$ on [0, 1], satisfying that g is continuous, piecewise $C^1$ with a piecewise Lipschitz derivative $g'$ such that\footnote{{It is enough to ask $g\in L^{2}([0,1])$, but, since the pre-averaging estimator is invariant to scalings of the weight function $g$, without loss of generality, we can impose the condition $|g|_{L^{2}}=1$.}}
\begin{equation*}
g(0)=g(1)=0, \quad \int_{0}^{1} g(s)^{2} d s=1{.}
\end{equation*}
\item[(iii)] A sequence $v_n$ representing the truncation level, satisfying
\begin{equation} \label{eq:v_n_noise}
v_{n}=\alpha \left(k_{n}\Delta_n \right)^{\varpi} \quad \text{ for some } \alpha > 0, \;\varpi \in (0,\frac{1}{2});
\end{equation}
\end{itemize}
Next, for an arbitrary process $U$, we define the sequences:
\begin{equation} \label{eq:notation_bar_hat}
\begin{array}{l}
{\overline{U}_{i}^{n}=\sum_{j=1}^{k_{n}-1} g\left(\frac{j}{k_{n}}\right) \Delta_{i+j-1}^{n} U} = -\sum_{j=1}^{k_{n}} \left(g\left(\frac{j}{k_{n}}\right) - g\left(\frac{j-1}{k_{n}}\right)\right) U_{i+j-2}^{n},\\
{\widehat{U}_{i}^{n}=\sum_{j=1}^{k_{n}}\left(g\left(\frac{j}{k_{n}}\right)-g\left(\frac{j-1}{k_{n}}\right)\right)^{2} \left(\Delta_{i+j-1}^{n} U\right)^2}.
\end{array}
\end{equation}
As seen from the definition, $\overline{U}_{i}^{n} $ is the weighted average of the increments $\Delta_{i+j-1}U, j= 1, \cdots, k_n -1 $, while $ \widehat{U}_{i}^{n}$ is a de-biasing term.
For a weight function $g$ as above, let
\begin{equation} \label{eq:phi}
\phi_{k_{n}}(g)=\sum_{i=1}^{k_{n}} g(\frac{i}{k_{n}})^{2};\quad
\phi_{k_{n}}^{\prime}(h)=\sum_{i=1}^{k_{n}}\left(g(\frac{i}{k_{n}})-g(\frac{i-1}{k_{n}})\right)^{2},
\end{equation}
and note that
\begin{equation} \label{eq:phi_g}
\begin{array}{l}
\phi_{k_{n}}(g) =k_{n} \int_0^1 g^2(s) ds+\mathrm{O}(1) = k_n +\mathrm{O}(1) ,\\
\phi_{k_{n}}^{\prime}(g) =\frac{1}{k_{n}} \int_0^1 {(g^{\prime}(s))^{2}} ds+\mathrm{O}\left(\frac{1}{k_{n}^{2}}\right). \end{array}
\end{equation}
Now we can define the pre-averaging estimator of the spot variance $c_{\tau} $ at $\tau\in(0,T)$. We consider a non-truncated version, defined as
\begin{equation} \label{eq:non_truncated_pre}
\hat{c}\left(k_{n}, {m}_{n}\right)_{\tau} =\frac{1}{\phi_{k_{n}}\left(g\right)} \sum_{j=1}^{n-k_n+1} K_{{m}_{n} \Delta_n}\left(t_{j-1} - \tau\right)\left(\left(\overline{Y}_{j}^{n}\right)^2 - \frac{1}{2}\widehat{Y}_{j}^{n}\right),
\end{equation}
as well as, two truncated versions:
\begin{equation}\label{PreAverEst0}
\begin{split}
\hat{c}\left(k_{n}, {m}_{n},v_n,1\right)_{\tau} &=\frac{1}{\phi_{k_{n}}\left(g\right)} \sum_{j=1}^{n-k_n+1} K_{{m}_{n} \Delta_n}\left(t_{j-1} - \tau\right)\left(\left(\overline{Y}_{j}^{n}\right)^2 \mathbbm{1}_{\{|\bar{Y}^n_j | \leq v_n\}} - \frac{1}{2}\widehat{Y}_{j}^{n}\right),\\
\hat{c}\left(k_{n}, {m}_{n},v_n,2\right)_{\tau} &=\frac{1}{\phi_{k_{n}}\left(g\right)} \sum_{j=1}^{n-k_n+1} K_{{m}_{n} \Delta_n}\left(t_{j-1} - \tau\right)\left(\left(\overline{Y}_{j}^{n}\right)^2 - \frac{1}{2}\widehat{Y}_{j}^{n}\right)\mathbbm{1}_{\{|\bar{Y}^n_j | \leq v_n\}} .
\end{split}
\end{equation}
{\color{Blue} The basic idea is the same as in the case where the efficient process $X$ is observed without noise. We see $\overline{Y}_{j}^{n}$ as a noise-free proxy of the increment $\Delta_{j}^n X$. By properly choosing the truncation {\color{Green} level} $v_n\to{}0$ (e.g., $v_n\gg \sqrt{u_n ln(1/u_n)}$ with $u_n:=k_n \Delta_n$), the event $|\bar{Y}^n_j | >v_n$ will indicate the occurrence of a ``big" jump happening during the time interval $[j\Delta_n,(j+k_n)\Delta_n]$ and, thus, we eliminate such a term from the summations in \eqref{PreAverEst0}. The estimator $\hat{c}\left(k_{n}, {m}_{n},v_n,2\right)_{\tau}$ is closer to the one defined in \cite{Yu2}, while $\hat{c}\left(k_{n}, {m}_{n},v_n,1\right)_{\tau}$ is similar to the one considered in \cite{chen2018inference}, though therein only the one-sided kernel $K(x)={\bf 1}_{[0,1]}(x)$ is studied. It will be interesting to compare their statistical properties and finite-sample performance.}
Before giving the asymptotic behavior of the pre-averaging estimators {\color{Blue} (\ref{eq:non_truncated_pre}) and (\ref{PreAverEst0})}, we introduced the limiting distributions. Below, $Z_{\tau}, Z^{\prime}_{\tau} $ are defined on a good extension $\Big(\widetilde{\Omega}, \widetilde{\mathcal{F}},\big(\widetilde{\mathcal{F}}_{t}\big)_{t>0}, \widetilde{\mathbb{P}}\Big)$ of the space $\left(\Omega, \mathcal{F},\left(\mathcal{F}_{t}\right)_{t \geq 0}, \mathbb{P}\right)$ {so that,} conditionally on $ \mathcal{F}$, {they} are independent Gaussian random variables with conditional variance \begin{equation} \label{eq:delta12}
\begin{aligned}
&\delta_1^2(\tau) := \widetilde{\mathbb{E}}\left(Z_{\tau}^2 | \mathcal{F}\right) = 4\left(\Phi_{22} c_{\tau}^{2}/\theta+2 \Phi_{12} c_{\tau} \gamma_{\tau}\theta+\Phi_{11} \gamma_{\tau}^{2}\theta^{3}\right) \int K^2(u) d u,\\
&\delta_2^2(\tau) := \widetilde{\mathbb{E}}\left(Z_{\tau}^{\prime 2} | \mathcal{F}\right)= \tilde{\sigma}_{\tau}^2\int L^2(t)dt,
\end{aligned}
\end{equation} with $\phi_{1}(s)=\int_{s}^{1} g^{\prime}(u) g^{\prime}(u-s) \mathrm{d} u$, $\phi_{2}(s)=\int_{s}^{1} g(u) g(u-s) \mathrm{d} u$, $ \Phi_{i j}=\int_{0}^{1} \phi_{i}(s) \phi_{j}(s) \mathrm{d} s$, and \(L(t) = \int_{t}^{\infty} K(u) d u \mathbf{1}_{\{t>0\}}-\int_{-\infty}^{t} K(u) d u \mathbf{1}_{\{t \leq 0\}}\).
The following result establishes the asymptotic behavior of the estimation error for the proposed estimators. The proof is given in Appendix \ref{PrfOfMnRslt}.
\begin{theorem} \label{thm}
Let $\{{m}_{n}\}_{n\geq{}1}$ be a sequence of positive integers such that $m_{n}\to\infty$, ${m}_{n} \Delta_n \rightarrow 0$, ${m}_{n}\sqrt{\Delta_n} \rightarrow \infty $, and ${m}_{n}\Delta_n^{3/4}\rightarrow \beta$ for some $\beta \in [0,\infty]$, and let $k_{n}$, $v_n$, and $g$ be as described in (i)-(iii) above. {Then,} under \thref{X,kernel,noise}, we have:
\begin{enumerate}
\item When $X$ is continuous, the pre-averaging estimators {\color{Blue} (\ref{eq:non_truncated_pre}) and (\ref{PreAverEst0})} are {\color{Blue} all} such that, as $n\to\infty$,
\begin{equation}\label{CLTTCs}
\begin{aligned}
\rm (i) &\quad {m}_{n}^{1 / 2} \Delta_{n}^{1 / 4} \left(\hat{c}_{\tau} -c_{\tau}\right) \stackrel{st}{\longrightarrow} Z_{\tau} + \beta Z^{\prime}_{\tau}, \mbox{ if } {\color{Blue} \beta\in[0,\infty)}, \\
\rm (ii) &\quad \frac{1}{\sqrt{{m}_{n} \Delta_n}} \left(\hat{c}_{\tau} -c_{\tau}\right) \stackrel{st}{\longrightarrow} Z^{\prime}_{\tau}, \mbox{ if } \beta = \infty;
\end{aligned}
\end{equation}
\item When $X$ is discontinuous and $r \in (0,2]$, with \begin{equation} \label{eq:m_n_range_noise}
m_{n} \Delta_{n}^{a} \rightarrow \beta^{\prime} \in(0, \infty), \quad \text { where } a \in(\frac{1}{2},1),
\end{equation} and $r, \varpi$ satisfying
\begin{equation} \label{eq:r_w_condition_noise}
r < \frac{5}{2} - 2\left[(a - \frac{1}{4})\wedge (1-a+\frac{1}{4})\right],\quad
\varpi \geq \frac{(a- \frac{1}{4})\wedge (1-(a- \frac{1}{4} ) ) - \frac{1}{4}}{2-r},
\end{equation}
the truncated pre-averaging estimator {\color{Blue} $\hat{c}(k_n, m_n, v_n, 1)_{\tau} $ in (\ref{PreAverEst0}) } satisfies (\ref{CLTTCs}) with $\beta=0$ when $a<3/4$, with $\beta=\beta'$ when $a=3/4$, {\color{Blue} or} with $\beta=\infty$ when $3/4<a<1$.
\item {\color{Blue} When $X$ is discontinuous and $r \in (0,2]$, with (\ref{eq:m_n_range_noise}) and $r, \varpi$ satisfying
\begin{equation} \label{eq:w_condition_noise_2}
{\color{Blue} r \le 4-\frac{2}{a \vee \left(3/2-a \right)},\quad \frac{(a- \frac{1}{4})\wedge (1-(a- \frac{1}{4} ) ) - \frac{1}{4}}{2-r}\leq \varpi \leq \frac{(2-2a) \vee (2a-1)}{r},}
\end{equation}
the truncated pre-averaging estimator $\hat{c}(k_n, m_n, v_n, 2)_{\tau} $ in (\ref{PreAverEst0}) satisfies (\ref{CLTTCs}) with $\beta=0$ when $a<3/4$, with $\beta=\beta'\in(0,\infty)$ when $a=3/4$, and with $\beta=\infty$ when $3/4<a<1$.}
\end{enumerate}
\end{theorem}
\begin{remark}
{\color{Blue} The second and third points in the above theorem show that the truncated pre-averaging estimators (\ref{PreAverEst0}) can achieve the optimal convergence rate of $\Delta_n^{1/8}$, but only if the index of jump activity is restricted to be $r<3/2$ for $\hat{c}(k_n, m_n, v_n, 1)$ and $r<4/3$ for $\hat{c}(k_n, m_n, v_n, 2)$. If the index of jump activity is larger than $3/2$ and $4/3$ respectively , the estimators can only achieved suboptimal convergence rates. When comparing their theoretical properties, $\hat{c}(k_n, m_n, v_n, 1) $ can, in principle, handle jumps with higher index $r$ than $\hat{c}(k_n, m_n, v_n, 2) $ at the optimal bandwidth. However, as we will see in Section \ref{simulation}, $\hat{c}(k_n, m_n, v_n, 2) $ appears to be more effective at eliminating jumps when the jump size is large in the presence of finite activity jumps. The two estimators have similar performance when the jump's size is relatively small.}
\end{remark}
\begin{remark}
{Let us give some intuition or heuristic explanation of the estimator (\ref{PreAverEst0}) and {\color{Blue} its asymptotic behavior established above.}}
For the estimation of the integrated variance (IV), \([X,X]_T = \int_0^T c_t dt \), \cite{jacod2009microstructure} proposed the {following pre-averaging estimator:}
\begin{equation*}
\widehat{[X,X]_s} := \frac{1}{\phi_{k_n}(g)} \frac{s}{s- k_n \Delta_n}\sum_{j=1}^{[s/\Delta_n]-k_n+1}\left(\left(\overline{Y}_{j}^{n}\right)^2 - \frac{1}{2}\widehat{Y}_{j}^{n}\right), \quad s \in (0, T],
\end{equation*}
for {\color{Blue} a continuous It\^o semimartingale} $X$. It was shown that:
\begin{equation*}
\frac{1}{\Delta_{n}^{1/4}}\left( \widehat{[X,X]_T} -[X,X]_T\right) \stackrel{st}{\longrightarrow} \mathcal{U}_{T}^{\text {noise }},
\end{equation*}
where $\mathcal{U}_{T}^{\text {noise }}$ is a centered Gaussian process with conditional variance
\[\delta_T:=\operatorname{\mathbb{E}}\left(\left(\mathcal{U}_{T}^{\text {noise }}\right)^{2} | \mathcal{F}\right)=
{\int_{0}^{T} \zeta_t d t}:=\int_{0}^{T} 4\left(\Phi_{22} c_{t}^{2}/\theta+2 \Phi_{12}c_t\gamma_{t}\theta+\Phi_{11} \gamma_{t}^{2}\theta^{3}\right) d t. \]
{In the no-thresholding case ($v_{n}=\infty$)}, the spot volatility estimator (\ref{PreAverEst0}) can be viewed as a localization of {the IV process} in that \begin{equation*}
\hat{c}_{t}\approx \int K_{m_n \Delta_n}(s-t)d\widehat{\left[X,X\right]}_{s}.
\end{equation*}
More specifically, the factor $\frac{s}{s-k_n \Delta_n} $ is omitted for the spot volatility estimator.
{If we use the representation $\mathcal{U}_{t}^{\text {noise }}=\int_{0}^{t}(\zeta_{s})^{1/2} dB^{U}_s$, where $B^U$ is a Wiener process, we can then heuristically argue that $\hat{c}_{t}-c_{t}\approx \int K_{m_n \Delta_n}(s-t)d(\widehat{\left[X,X\right]}_{s}-{\left[X,X\right]}_{s})=\Delta_{n}^{1/4}\int K_{m_n \Delta_n}(s-t)dU_{s}^{\text{noise}}=\Delta_{n}^{1/4}\int K_{m_n \Delta_n}(s-t)(\zeta_{s})^{1/2}dB^{U}_{s}$. Therefore, the variance of the estimation error at time \(t\) is expected to be close to}
\begin{equation*}
\sqrt{\Delta_n}\int K^2_{m_n\Delta_n}(s-t) {\zeta_{s}ds} \approx \frac{1}{m_n\sqrt{\Delta_n}}4\left(\Phi_{22} c_{t}^{2}/\theta+2 \Phi_{12} c_{t}\gamma_{t}\theta+\Phi_{11} \gamma_{t}^{2}\theta^{3}\right) \int K^2(u) du.
\end{equation*}
which is indeed the case, {but only} when \({m}_{n}\Delta_n^{3/4}\rightarrow \beta = 0\) as formally shown in Theorem \ref{thm}. {It is important to remark that the proof of Theorem \ref{thm} does not rely on the heuristic arguments above.}
\end{remark}
\section{An application: Optimal Parameter Tuning}\label{optimal_para}
In this section, as an application of our main {\color{Blue} Theorems \ref{thm_no_noise} and} \ref{thm}, we show how to tune {\color{Blue} the bandwidth parameter $\beta$ and the pre-averaging parameter $\theta$, as well as the kernel function $K$ of the estimator,} in order to minimize the asymptotic variance of the estimation error {\color{Blue} $\hat{c}_{\tau}-c_{\tau}$}. {Two possible approaches can be taken. Minimize the asymptotic variance of $\hat{c}_{\tau}$, say {\color{Blue} $\tilde{\delta}^2(\tau)$}, at each time $\tau$ or minimize the \emph{integrated asymptotic variance} {\color{Blue} $\int_{0}^{T}\tilde{\delta}^{2}(t)dt$} over the period $[0,T]$. In our simulations of Section \ref{simulation}, we implemented both methods and found out that the second method yields slightly better results. An explanation for this is given in Remark \ref{ExplLocvsGlob} below. Therefore, in this part, we focus on the second approach.}
By necessity, the optimal choices of $\theta$ and $\beta$ {under the criterion of the previous paragraph} will be expressed in terms of the integrated variance and quarticity, $IV_{T}:=\int_0^T c_{t}dt$ and $QrT_{T}:=\int_0^Tc_{t}^{2}dt$, respectively, the Integrated Volatility of Volatility (IVV), $\int_0^T \tilde{\sigma}_{t}^2 dt$, and the integrated variance of the noise $\epsilon_{t}$, $\int_0^T \gamma_{t}d t $. We can estimate $\int_0^T \tilde{\sigma}_{t}^2 dt$ and $\int_0^T \gamma_{t}d t $ separately, while for $IV_{T}$ and $QrT_{T}$, we {propose an iterative procedure} in which an initial rough estimate of $c_{t}$ on a grid of $[0,T]$ is used to determine estimates of {$IV_{T}$ and $QrT_{T}$}. These estimates are then used to find suitable estimates of the optimal values for $\theta$ and $\beta$. {Finally,} these estimated $\hat{\theta}$ and $\hat{\beta}$ are applied in the kernel pre-averaging estimator {(\ref{PreAverEst0})}
to refine our estimates of $c_{t}$ on the grid.
\begin{remark}
{\color{Blue} A related problem, that is not being considered here in detail, is that of tuning the truncation level $v_n$ in the truncated estimators \eqref{MDKNN_truncated} and \eqref{PreAverEst0}. Most of the literature about this problem {\color{Blue} has been} in the context of estimating the integrated variance $IV_T=\int_0^T \sigma_s^2ds$. It has been customary in econometric studies to adopt a power threshold of the form $v_n= c \Delta_n^\gamma$. The rule of thumb is to take a value of $\gamma$ close to $.5$ and $c$ that depends on an estimate of the volatility level. For instance, \cite{JacodTodorov:2014} took $\gamma=.49$ and $c=4\sqrt{BPV}$, where $BPV:= \frac{\pi}{2} \sum_{i=2}^n |\Delta_{i-1}^nX||\Delta_i^nX|$ is the Bipower variation. More recently, this issue has {\color{Blue} also been studied in the literature} using more objective and {\color{Blue} statistically valid approaches}, but only in the absence of microstructure noise. In the case of FA jumps and constant volatility $\sigma$, \cite{FigueroaLopezMancini:2019} showed that the optimal threshold (in terms of minimizing the conditional mean-square error) is asymptotically equivalent to $\sqrt{2\sigma^2 \Delta_n\ln(1/\Delta_n)}$ and proposed an itervative method to estimate $\sigma$. In the presence of small jumps that behave like those of an $\alpha$-stable L\'evy process, \cite{gong2021} showed that the optimal threshold is asymptotically equivalent to $\sqrt{(2-\alpha)\sigma^2 \Delta_n\ln(1/\Delta_n)}$. Again, these results are in the absence of microstructure noise and for the problem of estimating the integrated variance. However, given the local nature of spot volatility estimation, one can imagine that similar results may hold for the estimators \eqref{MDKNN_truncated} and \eqref{PreAverEst0}. We leave this problem for future research.}
\end{remark}
\subsection{Optimal selection of \texorpdfstring{$\theta$}{Lg} }
Recall we set $k_n = \frac{1}{\theta \sqrt{\Delta_n}} +\mathrm{o}\left(\frac{1}{\Delta_{n}^{1/4}}\right)$ and, thus, the parameter $\theta $ determines the length of the pre-averaging window $k_n$. The following corollary, which follows easily from Theorem \ref{thm}, {gives us a method to tune $\theta$ up.}
\begin{corollary}
{The optimal value $\theta^{*}$ of $\theta$, which is set to minimize the integrated asymptotic variance of the pre-averaging kernel estimator (\ref{PreAverEst0}), is such that}
\begin{equation} \label{eq:theta_optimal}
{(\theta^{\star})^{2} = \frac{\sqrt{\Phi_{12}^2 \left(\int_0^T c_t \gamma_t d t \right)^2 + 3\Phi_{11}\Phi_{22}\int_0^T \gamma_t^2 d t \int_0^T c_t^2 d t} - \Phi_{12}\int_0^T c_t \gamma_t d t}{3\Phi_{11}\int_0^T\gamma_{t}^2 d t}.}
\end{equation}
\end{corollary}
\begin{remark} {Note that the local version of (\ref{eq:theta_optimal}) (i.e., the value of $\theta$ that minimizes the spot asymptotic variance $\delta^{2}_{1}(t)$) is such that:
\begin{equation} \label{eq:theta_optimalLocal}
(\theta^{\star,local}_{t})^{2} = c_{t}\frac{\sqrt{\Phi_{12}^2 + 3\Phi_{11}\Phi_{22}} - \Phi_{12}}{3\Phi_{11}{\color{Blue} \gamma_{t}}}.
\end{equation}
In the context of integrated volatility estimation, \cite{jacod2015microstructure} obtained the same formula (see Eq.~(3.8) therein). It was also proposed a two-step procedure to implement it. However, in our simulation, we found out that the performance of the estimator is less sensitive to the choice of \(\theta\) than to that of the bandwidth.}
\end{remark}
\subsection{Optimal bandwidth selection} \label{band_section}
From Theorem \ref{thm}, we can deduce that when ${m}_{n} $ (the {bandwidth in} $\Delta_{n}$ units) is of the form ${m}_{n} = \beta \Delta_n^{-3/4} $ for some constant $\beta \in (0,\infty)$, the optimal convergence rate of $\Delta_{n}^{1/8}$ is attained and
we further have:
\begin{equation*}
\Delta_n^{-1/8} \left(\hat{c}\left(k_{n}, {m}_{n},v_n\right)_{\tau} -c_{\tau}\right)\stackrel{st}{\longrightarrow}
\beta^{-1/2} \left( Z_{\tau} + \beta Z^{\prime}_{\tau} \right).
\end{equation*} Therefore, the limiting distribution {\color{Blue} above} has conditional variance ${\bar{\delta}^{2}(\tau)}:=\frac{1}{\beta}\delta_1^2(\tau)+\beta \delta_2^2(\tau)$, where \(\delta_1^2(\tau)\) and \( \delta_2^2(\tau) \) are given as in (\ref{eq:delta12}).
{The following result gives the optimal value of $\beta$ that minimizes $\int_{0}^{T}\bar{\delta}^{2}(\tau)d\tau$.}
\begin{corollary} \thlabel{bandwidth}
Let
\begin{equation*}
\Theta(\theta):=\Theta(\theta;g):=\frac{\Phi_{22}}{\theta}\int_0^Tc_{t}^{2}dt+2 \Phi_{12} \theta \int_0^T\gamma_{t}c_{t}dt+\Phi_{11}\theta^{3} \int_0^T\gamma_{t}^{2}dt.
\end{equation*}
With the bandwidth $b_n = m_n \Delta_n=\beta \Delta_{n}^{1/4}$, the optimal value of $b_n$, which is set to minimize {$\int_{0}^{T}\bar{\delta}^{2}_{\tau}d\tau$},
is given by
\begin{equation} \label{eq:optimal_h}
{b}_{n}^{\star} = \sqrt{\frac{\int_0^T\delta_1^2(t) dt}{\int_0^T\delta_2^2(t) dt}}\Delta_n^{1/4}= \Delta_n^{1/4}\sqrt{\frac{4{\Theta(\theta)} \int K^2(u) d u}{\int_0^T\tilde{\sigma}_{t}^2 dt\int L^2(v)dv}}.
\end{equation}
With this optimal bandwidth choice, the integrated variance {$\int_{0}^{T}\bar{\delta}^{2}(\tau)d\tau$} of the limiting distribution for the {\color{Blue} scaled estimation error $\Delta_n^{-1/8} \left(\hat{c}\left(k_{n}, {m}_{n},v_n\right)_{\tau} -c_{\tau}\right)$} is given by
\begin{equation}\label{AsymptVarbb}
2\sqrt{\int_0^T \delta_1^2(t)dt \int_0^T \delta_2^2(t) dt} = 4 \sqrt{{\Theta(\theta)}\int_0^T\tilde{\sigma}_{t}^2 dt
\int K^2(u) d u \int L^2(v)dv}.
\end{equation}
\end{corollary}
Note that \(b_n^{\star}\) contains unknown theoretical quantities that need to be estimated in order to devise a plug in type estimator. Under the assumption of $\gamma_t \equiv \gamma $, the variance of the noise, \(\gamma \), can be estimated using the estimator in \cite{zhang2005tale}:
\begin{equation*}
\hat{\gamma} = \frac{1}{2n}\sum_{i=1}^n \left(Y^n_{i} - Y^n_{i-1}\right)^2.
\end{equation*}
For the estimation of the IVV, \(\int_0^T \tilde{\sigma}_{t}^2 dt\),
we start by obtaining a preliminary estimate of the spot variance $c$ on the grid $\tau\in\{t_{i}\}_{i=0,\dots,n}$, via the estimator (\ref{PreAverEst0}), staring with some sensible initial estimates of the tuning parameter values. For example, we can set \( b_n =m_{n}\Delta_{n}= \Delta_n^{1/4}\). Let us denote these initial {estimates as $\hat{c}_{t_{i},0}$}. We then {compute the sparse realized quadratic variation of the $\hat{c}_{t_{i}}$'s to estimate the Integrated Volatility of Volatility \({\rm IVV}=\int_0^T \tilde{\sigma}_{t}^2 dt\):}
\[
{\widehat{IVV}_{T,0}:=\sum_{i=0}^{[n/p]-1}(\hat{c}_{t_{(i+1)p},0}-\hat{c}_{t_{i p},0})^{2},}
\]
for some positive integer $p\ll n$. We also implemented a pre-averaging integrated variance estimator for the IVV based on the spot variance estimates. However, the choice of tuning parameters here could be tricky and the performance is similar to the {simpler} sparse Realized Variance estimator above. As for \(\int_0^T c^2_{{t}}dt\), we can simply compute the sum of squares of the preliminary estimates {$\hat{c}^{2}_{t_{i},0}$} and multiply by $\Delta_n$\footnote{{In the simulations, we also tried the preaveraged quarticity estimator of \cite{jacod2009microstructure} (Eq.~(3.14) therein) but the results were suboptimal.}}.
Now with these estimates, we can calculate an estimate of the optimal bandwidth \(b_n^{\star} \) using the result of Corollary \ref{bandwidth}. Such an approximate optimal bandwidth can then be used to refine our estimates of the spot variance grid. Continuing this procedure iteratively, we hope to obtain good estimates of the optimal bandwidth.
Note that (\ref{eq:optimal_h}) sets the same bandwidth for the entire path of $X$.
We can also consider a local or non-homogeneous bandwidth: for $\tau \in [0,T]$, the local bandwidth {is set} to minimize the {asymptotic} variance of the estimation error at time $\tau$. {Concretely, by setting ${m}_{n} = \beta \Delta_n^{-3/4} $ and minimizing the asymptotic spot variance ${\bar{\delta}^{2}(\tau)}=\beta^{-1}\delta_1^2(\tau)+\beta \delta_2^2(\tau)$}, the optimal bandwidth is given by
\begin{equation} \label{OptSpot00}
{b_n^{\star, local}(\tau)}= \frac{\delta_1(\tau)}{\delta_2(\tau)}\Delta_n^{1/4}
={\Delta_n^{1/4}}\sqrt{\frac{4{\Theta_{\tau}(\theta)} \int K^2(u) d u}{\tilde{\sigma}_{\tau}^2\int L^{2}(u)du}},
\end{equation}
with $\delta_{1}(\tau)$ and $\delta_{2}(\tau)$ defined as in Theorem \ref{thm} and $\Theta_{\tau}(\theta)$ defined as:
\begin{equation*}
\Theta_{\tau}(\theta):=\frac{\Phi_{22}}{\theta}c_{\tau}^{2}+2 \Phi_{12} \gamma_{\tau}\theta c_{\tau}+\Phi_{11} \gamma_{\tau}^{2}\theta^{3}.
\end{equation*}
With this optimal bandwidth, the variance of the limiting distribution for the estimation error is given by
\begin{equation}\label{AsymptVarbbb}
2\delta_1(\tau) \delta_2(\tau) = 4 \sqrt{{\Theta_{\tau}(\theta)}\tilde{\sigma}_{\tau}^2
\int K^2(u) d u \int L^{2}(u) du}.
\end{equation}
Since the local bandwidth has the flexibility to adapt to the volatility level, we may expect {that a data-driven estimate of the bandwidth $b_n^{\star, local}(\tau)$ in (\ref{OptSpot00}) should outperform a data-driven estimate of the homogeneous bandwidth $b_n^{\star}$ in (\ref{eq:optimal_h}). However, in our Monte Carlo simulations of Section \ref{simulation}, we found out this is not always the case. A possible explanation for this is given below (see also Remark \ref{ExplLocvsGlob} for further analysis).}
\begin{remark}\label{PossExpl00}
We can see the constant bandwidth (\ref{eq:optimal_h}) as an approximation of the optimal local bandwidth (\ref{OptSpot00}), where the {mean} values $\int_{0}^{T}\Theta_{t}(\theta)dt/T$ and $\int_{0}^{T}\tilde{\sigma}_{t}^2dt/T$ are used as proxies of the spot values $\Theta_{\tau}(\theta)$ and $\tilde{\sigma}_{\tau}^{2}$, respectively. These global proxies have the advantages of being easier and more accurate to estimate. {This may be one of the reasons why a data-driven estimate of the constant bandwidth $b_n^{\star}$ may be able to outperform a data-driven estimate of the local version $b_n^{\star, local}(\tau)$ in some situations.}
\end{remark}
\subsection{Optimal kernel function}\label{OptKernelSection}
With the optimal bandwidths of Section \ref{band_section}, we can now obtain a formula for the asymptotic variance, which enjoys an explicit dependence on the kernel function $K$. It is then natural to attempt to find the kernel that minimizes such a variance. As observed from (\ref{AsymptVarbb}) or (\ref{AsymptVarbbb}), we only need to minimize
\begin{equation*}
I(K) = \int K^2(u) d u \int L^2(u) du=
\int K^2(u) d u \iint_{x y\geq 0} K(x)K(y)(|x|\wedge{}|y|)d x d y,
\end{equation*}
over all kernels $K$ such that $\int K(u)du=1$, where for the second equality above we used that \(L(t) = \int_{t}^{\infty} K(u) d u \mathbf{1}_{\{t>0\}}-\int_{-\infty}^{t} K(u) d u \mathbf{1}_{\{t \leq 0\}}\).
It has been proved in \cite{FigLi}, Section 4.1, that, among all the kernel functions satisfying \thref{kernel}, the exponential kernel function $K^{\exp}(x)=\frac{1}{2} \exp (-|x|)$ is the one that minimizes the functional $I(K)$. {\color{Blue} \cite{FigLi} (see Remark 4.2 therein) showed that, compare to the two-sided uniform (resp., Epanechnikov) kernels, the integrated asymptotic variance can be reduced by about 14\% (resp., 6\%) when using exponential kernel.}
\cite{FigLi} also showed that exponential kernels have {\color{Blue} a computational advantage} since they {\color{Blue} enable us} to reduce the time complexity for estimating the volatility on all the grid points {\color{Blue} $t_{1}<\dots<t_n$}, from $O(n^{2})$ to $O(n)$. This property is particularly useful when working with high-frequency observations, {\color{Blue} where $n$ is quite large.}
\subsection{Tuning parameters under the absence of microstructure noise}\label{OptNoMicroSection}
{By following the same arguments as above, we can determine the optimal bandwidth parameter and kernel function for the estimators (\ref{MDKNN})-(\ref{MDKNN_truncated}) under the no-microstructure-noise model (\ref{eq:X})-(\ref{eq:sigma}). Specifically, we first take a bandwidth of the form $b_{n}=\beta\Delta_{n}^{1/2}$ ($\beta\in(0,\infty)$), which, from Theorem \ref{thm_no_noise}, leads {\color{Blue} to} the best possible rate of convergence $\Delta_{n}^{-1/4}$ of (\ref{MDKNN})-(\ref{MDKNN_truncated}). In that case, the asymptotic variance will take the form $\bar{\delta}_{\tau}^{2}=\beta^{-1}\delta_1^2(\tau)+\beta \delta_2^2(\tau)$, where
\[
\delta_{1}^{2}(\tau)=2 c_{\tau}^{2}\int K^2(u) du,\qquad
\delta_{2}^{2}(\tau)=2\tilde{\sigma}_{\tau}^{2} \int L^2(t)dt.
\]
Then, the optimal value of $\beta$ that minimizes the asymptotic variance is $\beta^{*}=\delta_{1}(\tau)/\delta_{2}(\tau)$, leading to the optimal bandwidth
\begin{equation} \label{OptSpot00NMN}
\tilde{b}_n^{\star, local} = \frac{\delta_1(\tau)}{\delta_2(\tau)}\Delta_n^{1/2}
=\Delta_n^{1/2}\sqrt{\frac{2 c_{\tau}^{2}\int K^2(u) d u}{\tilde{\sigma}_{\tau}^2\int L^{2}(u)du}}.
\end{equation}
Plugging $\beta^{*}$ into $\bar{\delta}_{\tau}^{2}$, leads to the optimal asymptotic variance of
\[
2\delta_1(\tau)\delta_2(\tau)= 4 \sqrt{c_{\tau}^{2}\tilde{\sigma}_{\tau}^2
\int K^2(u) d u \int L^{2}(u) du},
\]
which, as before, is minimized by the two-side exponential kernel $K(x)=2^{-1}e^{-|x|}$.}
\section{Simulation Study} \label{simulation}
In this section, we study the performance of the kernel pre-averaging estimators {\color{Blue} (\ref{eq:non_truncated_pre}) and (\ref{PreAverEst0}),} together with the implementation procedure described in Subsection \ref{band_section}, and compare the results with the Two Scale Realized Spot Variance (TSRSV) estimator proposed in \cite{zu2014estimating}.
\subsection{Simulation design and performance metrics} \label{heston_section}
We implemented two {\color{Blue} different data generating} models: {\color{Blue} a Heston model and a one-factor stochastic volatility (SV1F) model. More specifically, in Subsections \ref{ElmiJumpSec}-\ref{OptimalBdSec}, we consider the Heston model:}
\begin{equation} \label{HestonMld00}
\begin{split}
&{\color{Blue} Y_{t_i}=X_{t_i}+\varepsilon_{t_i}},\\
&\mathrm{d} X_{t}=\left(\mu-c_{t} / 2\right) \mathrm{d} t+{c_{t}^{1/2}} \mathrm{d} W_{t} + J_{t}^{X}\mathrm{d} N_{t}^{X},\\
&\mathrm{d} c_{t}=\kappa\left(\alpha-c_{t}\right) \mathrm{d} t+\gamma c_{t}^{1 / 2} \mathrm{d} B_{t} +\sqrt{c_{t-}} J_{t}^{c} \mathrm{d} N_{t}^{c},
\end{split}
\end{equation}
where we assume $B_t = \rho W_t + \sqrt{1 - \rho^2} {\tilde{W}_t}$, with $\tilde{W}$ being a Brownian motion independent with $W$.
We adopt the same parameter values as in \cite{zhang2005tale}, but properly normalized so that the time unit is one day:
\begin{equation} \label{ParValHston00}
{\mu=0.05 / 252,\quad \kappa=5 / 252, \quad \alpha=0.04 / 252,\quad \gamma=0.5 / 252, \quad \rho=-0.5.}
\end{equation}
We set the noise as $\epsilon_{i}^{n}:=\epsilon_{t_{i}} \stackrel{i.i.d.}{\sim} \mathcal{N}\left(0,0.0005^{2}\right)$, and the initial values to $X_0 = 1$ and $c_0 = 0.04/252$. The jump parameters are taken from \cite{chen2018inference} and set to be \(J^X_t\, {\color{Blue} \stackrel{{\rm iid}}{\sim}}\, N(-0.01,0.02^2)\), \(N_{t+\Delta}^{X}-N_{t}^{X} \sim \text {Poisson }(36 \Delta/252)\), \(\log \left(J_{t}^{c}\right) \,{\color{Blue} \stackrel{{\rm iid}}{\sim}}\, N(-5, 0.8)\), and \(N_{t+\Delta}^{c}-N_{t}^{c} \sim \frac{1}{\sqrt{252}}\operatorname{Poisson}(12 \Delta/252)\), {\color{Blue} with all these random processes being mutually independent.}
{\color{Blue} We also consider the One Factor Stochastic Volatility (SV1F) model (cf. \cite{zu2014estimating}, \cite{barndorff2008designing}, \cite{Yu2}):
\begin{equation} \label{eq:sv1f_model}
\begin{aligned}
& Y^n_i = X^n_i + \epsilon^n_i\\
&\mathrm{d} X_{t}=\mu \mathrm{d} t+\exp \left(\beta_{0}+\beta_{1} {\color{Green} \gamma_{t}}\right) \mathrm{~d} W_{t} + \mathrm{~d} J_t, \\
&\mathrm{d} {\color{Green} \gamma_{t}}=\alpha {\color{Green} \gamma_{t}} \mathrm{~d} t+\mathrm{d} B_{t}.
\end{aligned}
\end{equation}
The model above is adopted in Subsections \ref{ComSubH} and \ref{CmpTrNoTR} with different parameter values that will be specified therein.}
Throughout, we use the usual triangular weight function \(g(x) = 2 x \wedge(1-x) \). We simulate data for one day ($T=1$), and assume the data is observed once every second, with 6.5 trading hours per day. The number of observation is then \(n = 23 400\).
For the $j$th simulated path \(\{{X^{(j)}_{t_i}}: 0\leq i\leq n, t_i = i T/n\}\), we estimate the corresponding skeleton of the spot variance process, \(\{c_{t_i,j}\}_{i=1,\dots,n}\), for a given pre-averaging parameter \(\theta\) and a bandwidth parameter \({\tilde{\beta}}\) (the bandwidth is {\color{Blue} then given by} \(\tilde{\beta} \Delta_n^{1/4}\)). The estimated path is denoted as \(\{\hat c_{t_i,j}\}_{i=1,\dots,n}\). Next, we calculate the average of the squared errors (ASE),
\[
ASE_{j} = \frac{1}{n-2l+1}\sum_{i=l}^{n-l}\left(\hat{c}_{t_i,j} - c_{t_i,j} \right)^2.
\]
Here, \(l=[0.1n] \) is used to further alleviate boundary effects. Then, we take the square root of the average of the \(ASEs\) over all the simulated paths:
\[
\widehat{RMSE} = \sqrt{\frac{1}{m}\sum_{j=1}^{m} ASE_{j}},
\]
where $m$ is the number of simulations. This is an estimate of
\[
RMSE = \sqrt{\operatorname{\mathbb{E}}\left[\frac{1}{n-2l+1}\sum_{i=l}^{n-l}\left(\hat{c}_{t_i} - c_{t_i} \right)^2 \right]}.
\]
\subsection{Elimination of jumps and truncation}\label{ElmiJumpSec}
{In this subsection, we will show that} the truncation in the estimator (\ref{PreAverEst0}) {does a good job in eliminating} the jumps of the process (\ref{HestonMld00}). {To this end, we compare the performance of the truncated estimator \( \hat{c}\left(k_{n}, {m}_{n},v_n,1\right)\) in (\ref{PreAverEst0}), with that of the non-truncated estimator
\begin{equation}\label{PreAverEst0_cts}
\hat{c}\left(k_{n}, {m}_{n}, Y^*\right)_{\tau}:=\frac{1}{\phi_{k_{n}}\left(g\right)} \sum_{j=1}^{n-k_n+1} K_{{m}_{n} \Delta_n}\left(t_{j-1} - \tau\right)\left(\left(\overline{Y}_{j}^{*n}\right)^2 - \frac{1}{2}\widehat{Y}_{j}^{*n}\right),
\end{equation}
applied to the continuous} Heston model:
\begin{equation} \label{HestonMld00_cts}
\begin{split}
&Y_{i}^{*n}=X_{i}^{*n}+\varepsilon_{i}^{n},\\
&\mathrm{d} X^*_{t}=\left(\mu-c^*_{t} / 2\right) \mathrm{d} t+\sqrt{c^*_{t}} \mathrm{d} W_{t} ,\\
&\mathrm{d} c^*_{t}=\kappa\left(\alpha-c^*_{t}\right) \mathrm{d} t+\gamma\sqrt{c_{t}^{*}} \mathrm{d} B_{t} .
\end{split}
\end{equation}
We set \(\beta = 1\), \(\theta = 5\), and $v_n = 1.8\times \sqrt{BPV}(k_n \Delta_n)^{0.47}$, where $BPV = \frac{\pi}{2} \sum_{i=2}^{n}\left|\Delta_{i-1}^{n} X \| \Delta_{i}^{n} X\right|$\footnote{A similar threshold is applied in \cite{JacodTodorov:2014}.}. {\color{Blue} The results are based on 2000 simulated paths of both \(Y\) and \(Y^{*}\)}.
As proposed in \cite{kristensen2010nonparametric}, in order to alleviate the {edge effects, we replace $K_{{m}_{n} \Delta_n}\left(t_{i-1} - \tau\right)$ in (\ref{PreAverEst0}) and (\ref{PreAverEst0_cts}) with}
\begin{equation*}
K^{adj}_{m_n \Delta_n}\left( t_{i-1} - \tau\right) = \frac{K_{{m}_{n} \Delta_n}\left(t_{i-1} - \tau\right)}{\Delta_n\sum_{j=1}^{n-k_n+1} K_{{m}_{n} \Delta_n}\left(t_{j-1} - \tau\right)}.
\end{equation*}
The \(\widehat{RMSE}\) of the three estimators is reported in Table \ref{table:jump}.
\begin{table}[h!]
\begin{center}
\begin{tabular}{ | c | c | c | c | c |}
\hline
&{\begin{tabular}{c}
\( \hat{c}\left(k_{n}, {m}_{n}, Y^*\right) \)
\end{tabular}}&{\begin{tabular}{c}
\( \hat{c}\left(k_{n}, {m}_{n}, Y\right)\)
\end{tabular}} & {\begin{tabular}{c}
$\hat{c}\left(k_{n}, {m}_{n}, v_n\right)$
\end{tabular}}\\
\hline
$\widehat{RMSE}\times 10^{5}$ &5.483386 &16.83420
&5.419338 \\
\hline
\end{tabular}
\caption{Comparison between truncated and non-truncated estimators.}
\label{table:jump}
\end{center}
\end{table}
{\color{Blue} These results suggests that} the {truncation procedure} can effectively eliminate the jumps under this Heston model, since the estimated RMSE of the truncated estimator for the model (\ref{HestonMld00}) is even less than that of the non-truncated estimator based on the continuous model (\ref{HestonMld00_cts}).
\subsection{Validity of the asymptotic theory and necessity of de-biasing}
We first show that the asymptotic behavior of the estimation error is consistent with our theoretical result. By \thref{bandwidth}, the optimal rate of convergence of the estimation error is attained when the bandwidth takes the form \(m_n^{\star} \Delta_n = \beta \Delta_n^{1/4}\), for some \(\beta \in (0,\infty) \), and, thus, we only analyze the {\color{Blue} case 1(i) ($\beta\in(0,\infty)$)} of Theorem \ref{thm}. We aim to estimate the {\color{Blue} spot variance \(c_{0.5}\) in the Heston model \eqref{HestonMld00} without jumps. Accordingly and for simplicity, we use the untruncated pre-averaging kernel estimator \eqref{eq:non_truncated_pre}. We take \(\beta = 1\) and exponential kernel}. The histogram of the estimation errors, $\hat{c}_{0.5} - c_{0.5}$, based on 25,000 simulated paths, is shown in Figure \ref{fig:hist_exp}. We also plot the theoretical density of the estimation error as prescribed by Theorem \ref{thm} but with the true parameter values for $\gamma$ and $\theta $, and replacing $c_{0.5}$ with the {\color{Blue} average value} of $c_{0.5}$ over all 25,000 path. As it can be seen, the theoretical density is consistent with the empirical results.
\begin{figure}[h]
\centering
\includegraphics[scale = 0.5]{hist_exp.png}
\caption{Histogram of $\hat{c}_t - c_t$ at $t=0.5$ and the density of the theoretical limiting distribution.}
\label{fig:hist_exp}
\end{figure}
To investigate the need of the bias correction term \(\widehat{Y}_{j}^{n}\) in $\hat{c}\left(k_{n}, {m}_{n},v_n,1\right)_{\tau}$, let us consider a new estimator without the bias correction term, \(\tilde{c}_{\tau} = \sum_{j=1}^{n-k_n+1} K_{m_n \Delta_n}\left(t_{j-1} - \tau\right)\left(\overline{Y}_{j}^{n}\right)^2 \mathbbm{1}_{\{|\bar{Y}^n_j | \leq v_n\}} \). We show the histogram of the estimation errors $\tilde{c}_{0.5} - c_{0.5} $ for 25,000 simulated paths, and, for comparisons, also plot the same theoretical asymptotic density function of Figure \ref{fig:hist_exp}. As shown in {left panel of} Figure \ref{fig:debias}, the estimator {\color{Blue} $\tilde{c}_{0.5}$} significantly overestimates the spot variance, which shows the necessity of the bias correction term \(\widehat{Y}_{j}^{n}\) in (\ref{PreAverEst0}).
\begin{figure}
\centering
\includegraphics[scale = 0.45]{No_debias.png}
\includegraphics[scale = 0.45]{exp_unif.png}
\caption{Left Panel: The effect of bias correction term. Right Panel: Comparison of the asymptotic distribution between uniform and exponential kernel.}
\label{fig:debias}
\end{figure}
\subsection{Performance for different kernels}
Before analyzing the empirical performance of the estimators for different kernels, we compare the theoretical asymptotic densities of the estimation error for the exponential and uniform kernels. This is shown in {right panel of Figure \ref{fig:debias}}. We can see therein that, as predicted in Subsection \ref{OptKernelSection}, the exponential kernel estimator {\color{Blue} has smaller} asymptotic variance.
We now proceed to compare the finite sample performance of {\color{Blue} the untruncated pre-averaging kernel estimator \eqref{eq:non_truncated_pre}}
for different kernels in the Heston model \eqref{HestonMld00} {\color{Blue} without jumps}. We assume both a non-leverage setting (\(\rho =0\)) and a negative correlation setting (\(\rho = -0.5\)).
{\color{Blue} We} fix \(\theta = 5\) and apply the iterative homogeneous bandwidth selection method introduced in Subsection \ref{band_section} {\color{Blue} with different} kernels. We report the estimated $RMSE$ with the initial bandwidth $\beta =1 $ and the result of iterative bandwidth selection method after one iteration in Table \ref{table:diff_kernel} for the following {\color{Blue} four kernels}:
\begin{equation*}
\begin{split}
&K_{{\rm exp}}(x) = \frac{1}{2} e^{-|x|} , \quad
K_{unif}(x) = \frac{1}{2} \mathbbm{1}_{\{|x|<1\}} \\
&K_{1}(x) = |1-x| \mathbbm{1}_{\{|x|<1\}} , \quad
K_{2}(x) = \frac{3}{4} (1-x^2) \mathbbm{1}_{\{|x|<1\}} .
\end{split}
\end{equation*}
This shows that, indeed, the exponential kernel provides the best performance.
\begin{table}[h!]
\begin{center}
\begin{tabular}{ | l | l | c |}
\hline
\multicolumn{3}{|c|}{{$\widehat{RMSE}\times 10^{5}$ ($\rho = 0$)}} \\
\hline
Kernel & $\beta = 1$& \begin{tabular}{c}Optimal\\
Bandwidth Selection\end{tabular} \\
\hline
$K_{{\rm exp}}$ &1.400 & 1.068 \\
\hline
$K_{unif}$ &1.890 & 1.608 \\
\hline
$K_1$ &2.173 & 1.648 \\
\hline
$K_2$ &2.064 & 1.476 \\
\hline
\end{tabular}
\caption{Comparison of different kernel functions.}
\label{table:diff_kernel}
\end{center}
\end{table}
\subsection{Optimal bandwidth}\label{OptimalBdSec}
First, we show that the suboptimal bandwidth, which corresponds to $\beta = 0$ in Theorem \ref{thm}, indeed performs worse than the optimal bandwidth, even though its asymptotic variance is easier to estimate without the \(\beta Z^{\prime}_{\tau}\) term. {\color{Blue} For simplicity, we again only consider the Heston model \eqref{HestonMld00} without jumps and the untruncated pre-averaging kernel estimator \eqref{eq:non_truncated_pre}. We will compare the truncated and untruncated versions in more detail below in Subsection \ref{CmpTrNoTR}.}
In Table \ref{table:opt_subopt}, we compare the optimal bandwidth \(h_1 = \beta \Delta_n^{1/4} \) with the suboptimal bandwidths \( h_2 = \beta \Delta_n^{0.28} \) and \(h_3 = \beta \Delta_n^{0.3} \), using the exponential kernel with \(\beta = 1,2,3,4\) respectively, {\color{Blue} based} on 1000 simulated path. The results show the advantage in using the optimal bandwidth for the same level of the coefficient $\beta $.
\begin{table}[h!]
\begin{center}
\begin{tabular}{ | c | c | c |c|}
\hline
\multicolumn{4}{|c|}{{$\widehat{RMSE}\times 10^{5}(\rho = -0.5)$}} \\
\hline
Bandwidth & $h_1$(optimal) & $h_2$ (suboptimal) & $h_3$ (suboptimal)\\
\hline
$\beta = 1$ & 1.418 & 1.605 & 1.754 \\
\hline
$\beta = 2$ & 1.133 & 1.225 & 1.308\\
\hline
$\beta = 3$ & 1.077 & 1.121 & 1.678\\
\hline
$\beta = 4$ & 1.050 & 1.073 & 1.104 \\
\hline
\end{tabular}
\caption{Comparison between optimal bandwidth and suboptimal bandwidth}
\label{table:opt_subopt}
\end{center}
\end{table}
Next, we compare the results of the iterative homogeneous and local bandwidth selection methods, as discussed in Subsection \ref{band_section}. Based on some initial simulations, we observed that the parameter \(\theta \), which controls the length of the pre-averaging window \(k_n\) as \(k_n = \frac{1}{\theta\sqrt{\Delta_n}}\), has comparatively smaller effect on the performance of estimator than that of the bandwidth. Therefore, throughout this section, we fix \( \theta = 5\), which is computed by (\ref{eq:theta_optimal}) using true parameter values, and consider different bandwidth selection techniques\footnote{We also consider other values of $\theta$ and the results were similar.}.
In Table \ref{table:diff_band}, we report the estimated RMSE for different bandwidth selection methods. For the homogeneous bandwidth selection method (\ref{eq:optimal_h}), we apply the realized variance of sparsely sampled (5 min) spot variance estimates $\{\hat{c}_{t_{i}}\}$ to estimate the vol vol \(\int_0^T \tilde{\sigma}^2_t dt \) as described in Section \ref{band_section}. We fix the estimated vol vol after the first iteration to prevent the increased variance brought by the iterative method. The first two iterations are shown in the first two columns of the table and we can see that the second iteration does not improve the result significantly. Therefore, one iteration of the bandwidth selection method is sufficient in practice.
For the local bandwidth method, we use \(\int_0^T \tilde{\sigma}^2_t dt/T \) as a proxy of $\tilde{\sigma}^2_{\tau}$ in the formula (\ref{OptSpot00}).
As a reference, we also give the results of using an oracle optimal bandwidth, which is computed by the true parameter values and the simulated spot variance process with Eqs.~(\ref{eq:optimal_h}) and (\ref{OptSpot00}) for the optimal homogeneous and optimal local bandwidths, respectively. In the last column, we provide the result of a semi-oracle type of bandwidth, where we use the estimated spot variance ``skeleton" $\{\hat{c}_{t_{i}}\}$ to estimate $\int_0^Tc_t dt$ and $\int_0^T c^2_t dt$, via Riemann sums\footnote{We also apply the pre-averaging estimate of quarticity given in \cite{jacod2010limit}, but the results were less optimal.}, while using the true parameter of $\gamma$ given in (\ref{ParValHston00}) to estimate $\int_0^T \tilde{\sigma}^2_t dt=\gamma^{2} \int_{0}^{T}c_{t}dt$. The last simplification is possible due to the special structure of the diffusion coefficient of variance process in the Heston model (\ref{HestonMld00}). A similar approach can be applied to other popular volatility models such as CEV models. As we can see therein, the data-driven approaches (1st two columns) are quite close to the oracle and semi-oracle estimates.
\begin{table}[h!]
\begin{center}
\begin{tabular}{ | c | c | c | c | c |}
\hline
\multicolumn{5}{|c|}{ $\widehat{RMSE}\times 10^{5}(\rho = -0.5)$} \\
\hline
& 1st Iter. & 2nd Iter & Oracle & Semi-oracle\\
\hline
homogeneous &1.0530 &1.0529&1.0540 &1.0533 \\
\hline
local & 1.0571 & 1.0551 &1.0542 &1.0547 \\
\hline
\end{tabular}
\caption{Comparison of different bandwidth selection methods based on 1000 simulations. RMSE for initial bandwidth \(\beta = 1\) is $1.4086\times 10^{-5}$. Columns 2 and 3 show the results corresponding to the 1st and 2nd iterations of bandwidth selection methods. Column 4 and 5 show the result using oracle and semi-oracle bandwidths, respectively.}
\label{table:diff_band}
\end{center}
\end{table}
\begin{remark}\label{ExplLocvsGlob}
The estimator with local bandwidth has the flexibility to adjust its bandwidth at different times based on the data. Therefore, theoretically, this estimator should be able to achieve a lower {\color{Blue} value of the integrated asymptotic variance $\int_{0}^{T}\bar{\delta}^{2}(\tau)d\tau$, which, as defined in \thref{bandwidth}, is given by}:
\begin{equation*}
\Delta_n^{1/4} {\color{Blue} \int_0^T\left( \frac{1}{\beta_{t}}\delta_1^2(t)+\beta_{t} \delta_2^2(t)\right) d t}.
\end{equation*}
However, {\color{Blue} our} simulations show that the performance of the local bandwidth
is almost the same as that of the homogeneous bandwidth. To further investigate this phenomenon, in {\color{Blue} the} left panel of Figure \ref{fig:zeta1bdw}, we show the estimated RMSE for different {\color{Blue} times \( \tau\)} against the parameter $\beta$ in the bandwidth formula $b_{n}=\beta\Delta_{n}^{1/4}$. As before we simulate the Heston model (\ref{HestonMld00}) with the same parameters as in (\ref{ParValHston00}), but with the vol vol parameter $\gamma= 1/252$. We can conclude from the figure that the optimal {\color{Blue} $\beta$-value is almost the same for different $\tau$'s, and this value is also close to the theoretical optimal homogeneous} bandwidth based on the asymptotic variance of the estimator. Thus, an estimator with homogeneous bandwidth can achieve a similar result without extra computation cost. This trend is less obvious when the vol vol parameter $\gamma$ is relatively small. In the right panel of Figure \ref{fig:zeta1bdw} we show the estimated RMSE vs. $\beta$ when $\gamma = 0.5/252$. {\color{Blue} In that case, the} perceived almost flat trend as the bandwidth increases shows that the realized variance can serve as a good proxy of the spot volatility, at least for the purpose of tuning the parameters of the estimators, since the spot volatility estimator degenerates to the integrated volatility estimator when the bandwidth gets large. Note, however, that the MSE paths are slowly tickling up as $\beta$ increases and each of those paths exhibit an optimal bandwidth. These are again relatively close for different times $\tau$ and also close to the theoretical optimal {\color{Blue} homogeneous} bandwidth. In conclusion, when the vol vol parameter is not known, the theoretical optimal bandwidth can provide a good guideline for the empirical experiments and a homogeneous bandwidth is sufficient in achieving similar result as local bandwidth while reducing the estimation error and computation cost caused by {\color{Blue} the} latter.
\begin{figure}
\centering
\includegraphics[scale = 0.45]{zeta1bdw.png}
\includegraphics[scale = 0.45]{zeta0_5bdw.png}
\caption{Left Panel: MSE v.s. bandwidth when {$\gamma = 1/252$}. Right panel: MSE v.s. bandwidth when {$\gamma = 0.5/252$}}
\label{fig:zeta1bdw}
\end{figure}
\end{remark}
\newpage
\subsection{Comparison with TSRSV}\label{ComSubH}
{\color{Blue} In this section, we adopt the model \eqref{eq:sv1f_model} with the same parameters as \cite{zu2014estimating}}:
\begin{equation}
\mu = 0.03, \;\beta_1 = 0.125,\;\alpha=-0.025, \;\rho = -0.3,\; \beta_0 = \beta^2_1/(2\alpha).
\end{equation}
{\color{Blue} We also take $\gamma_0\sim\mathcal{N}\left(0, -\frac{1}{2\alpha}\right)$ and $J=0$.}
The microstructure noise $\epsilon^n_i$ is set to be $\epsilon_{i}^{n}:=\epsilon_{t_{i}} \stackrel{i.i.d.}{\sim} \mathcal{N}\left(0,\omega^2\right) $, {\color{Blue} where, as in \cite{zu2014estimating}, $\omega^2$ can take one of three possible levels: $0.0001$, $0.001$, and $0.01$.}
For the TSRSV estimator, we implement the smoothing version of the TSRSV (see \cite{zu2014estimating}, Section 3.1), denoted as {\color{Blue} $\hat{c}_{TSRSV}$}, and calculate the bandwidth and scale parameters according to Section 3.4 in \cite{zu2014estimating}. For our pre-averaging estimator, we implement the non-truncated version (denoted as {\color{Blue} $\hat{c}_{PA}$}) with exponential kernel and use the iterative method described in {\color{Blue} Subsection} \ref{band_section} for bandwidth selection. {\color{Blue} We consider two sampling frequencies: 1-sec or 5-sec.} In Table \ref{table:TSRSV}, we report the RMSE of the two estimators. {\color{Blue}
As shown in the table, the pre-averaging estimator has a superior performance, especially when the noise level is large.}
\begin{table}[h!]
\begin{center}
\begin{tabular}{ c |c c |c c |c c }
\hline
\multirow{2}{5em}{Frequency} &\multicolumn{2}{c}{$\omega^2 = 0.0001$} &
\multicolumn{2}{c}{$\omega = 0.001$} & \multicolumn{2}{c}{$\omega = 0.01$} \\
& $\hat{c}_{PA} $ &$\hat{c}_{TSRSV}$ & $\hat{c}_{PA} $ &$\hat{c}_{TSRSV}$ & $\hat{c}_{PA} $ & $\hat{c}_{TSRSV}$ \\
\hline
1 sec & 0.0411 & 0.0727 & 0.0546 & 0.1121 & 0.0634& 0.2345\\
5 sec & 0.0505 & 0.1487 & 0.0649 & 0.1392 & 0.1066 &0.3005
\end{tabular}
\caption{Comparison between TSRSV and kernel pre-averaging estimator. {\color{Blue} We set the parameter $\theta$ in \eqref{AsympCndkn} to be $5$, $5$, and $1.5$ for noise level $0.0001, 0.001,0.01$ respectively. For the TSRSV estimator, to reduce the computational cost, instead of choosing the initial bandwidth using the cross validation method proposed in \cite{kristensen2010nonparametric} as did in Section 4.2.1 \cite{zu2014estimating}, we use the bandwidth already selected in \cite{zu2014estimating}, Table 8-10, as the initial values to estimate the vol vol. Note that the obtained RMSE of the smoothing TSRSV estimator under the different noise levels (0.0727, 0.1121, and 0.2345) match the results in \cite{zu2014estimating}, who reported the values 0.094, 0.118, and 0.223, respectively.}}
\label{table:TSRSV}
\end{center}
\end{table}
\subsection{{\color{Blue} Comparison between the truncated and untruncated estimators}}\label{CmpTrNoTR}
{\color{Blue} In this subsection, we study the two versions of the truncated estimators (\ref{PreAverEst0}) and the non-truncated estimator (\ref{eq:non_truncated_pre}) under various levels of jump size and sample frequency, using the simulation setting in \cite{Yu2}. The parameters therein are chosen from \cite{Huang2005}}:
\begin{equation}
\mu = 0.03,\; \beta_1 = 0.125,\;\alpha=-0.1, \;\rho = 0, \;\beta_0 = 0.
\end{equation}
We also conduct our study in the same experiment design {\color{Blue} as} \cite{Yu2}. {\color{Blue} More specifically, we consider three levels of jump activity (no jumps, compound Poisson jumps, and Variance Gamma jumps); 3 noise levels ($\omega = 0025,0.035,0.05 $); and three different sample frequencies (1 observation every 10 sec, every 30 sec, and every 60 sec.). In the case of finite activity jumps, $J_{t}=\sum_{j=1}^{N_{t}} Z_{\tau_{j}}$ with $Z_{\tau_{j}} \sim N(0,\sigma_Y^2) $ and {\color{Blue} $\{N_t\}_{t\geq{}0} \sim Poisson(3)$}, while in the case of infinite activity jumps, $J_t = cG_t + \eta \tilde{W}_{G_t} $ with $G_t \sim Gamma(t/b,b)$, $b = 0.23$, $c = -0.2$, $\eta = 0.2$, and $\tilde{W} $ is an independent Brownian motion,} as in \cite{Mancini2009}.
{\color{Blue} In Table \ref{table:finite_jump}, we report the RMSE of our non-truncated estimators (\ref{eq:non_truncated_pre}) and truncated estimators (\ref{PreAverEst0}), denoted as $\hat{c}_{T,1}$, $\hat{c}_{T,2}$ and $\hat{c}_{Non}$, respectively.
We observe that when there are no jumps, all three estimators have similar performance, with the non-truncated estimator giving slightly better results. However, when jumps are present, the truncated estimators have a much superior performance, especially when the jump size is large, and $\hat{c}_{T,2} $ appears to be more effective at eliminating jumps compared with $\hat{c}_{T,1} $.}
{\color{Blue} \cite{Yu2} also proposed a pre-averaging kernel estimators {\color{Blue} for the spot volatility. As mentioned in the introduction, their estimator has a different de-biasing term, which could affect the finite-sample performance, and their asymptotic normality is only established with} suboptimal convergence rate.} Our results in Table \ref{table:finite_jump} are comparable with \cite{Yu2}, Section 5. For example, the RMSE 0.1421 under $\omega = 0.025, \sigma_Y = 0$ with 10 sec data is close to the RMSE $\sqrt{0.0201} = 0.1417$ in \cite{Yu2}.
\begin{table}[h!]
\begin{center}
\begin{tabular}{ c |c c c |c c c |c c c }
\hline
\multirow{2}{5em}{Frequency} &\multicolumn{3}{c}{$\omega = 0.025$} &
\multicolumn{3}{c}{$\omega = 0.035$} & \multicolumn{3}{c}{$\omega = 0.05$} \\
& $\hat{c}_{T,2}$ &$\hat{c}_{T,1}$ & $\hat{c}_{Non} $ & $\hat{c}_{T,2}$ & $\hat{c}_{T,1}$ & $\hat{c}_{Non} $ & $\hat{c}_{T,2}$ &$\hat{c}_{T,1}$ & $\hat{c}_{Non} $\\
\hline
\multicolumn{10}{l}{Scenario A: Diffusion with no jumps $\sigma_Y =0 $ }\\
10 sec &0.1421& 0.1421& 0.1420& 0.1474& 0.1474& 0.1472& 0.1689& 0.1690& 0.1677\\
30 sec &0.1724& 0.1724& 0.1719& 0.1763& 0.1763& 0.1758& 0.1846& 0.1846& 0.1837\\
60 sec &0.2057& 0.2057& 0.2054& 0.2085& 0.2085& 0.2083& 0.2161& 0.2161& 0.2160\\
\hline
\multicolumn{10}{l}{Scenario B: Diffusion with small jumps $\sigma_Y =0.5 $ }\\
10 sec &0.1311& 0.1630& 1.0120& 0.1347& 0.1641& 1.0004& 0.1592& 0.1805& 0.9943\\
30 sec & 0.1660& 0.1893& 0.9758& 0.1690& 0.1913& 0.9798& 0.1797& 0.1971& 0.9783\\
60 sec &0.2083& 0.2209& 0.9738& 0.2113& 0.2215& 0.9732& 0.2205& 0.2255& 0.9730\\
\hline
\multicolumn{10}{l}{Scenario C: Diffusion with large jumps $\sigma_Y =1.5 $ } \\
10 sec & 0.1952 &0.9402 &9.4648& 0.2013& 0.9397& 9.3509& 0.2090& 0.9013& 9.1702\\
30 sec & 0.2262 &0.9183 &9.0481& 0.2256& 0.9132& 9.0438& 0.2279& 0.8990& 8.9954\\
60 sec & 0.2640 &0.8949 &8.9794& 0.2635& 0.8892& 8.9731& 0.2590& 0.8707& 8.9514\\
\hline
\multicolumn{10}{l}{Scenario D: Diffusion with jumps of infinite activity } \\
10 sec& 0.1246& 0.1247& 0.1271& 0.1330& 0.1331& 0.1349& 0.1529& 0.1530& 0.1547\\
30 sec&0.1576& 0.1577& 0.1594& 0.1587& 0.1588& 0.1598& 0.1728& 0.1729& 0.1740\\
60 sec&0.1940& 0.1940& 0.1952& 0.1959& 0.1959& 0.1966& 0.2063& 0.2063& 0.2076\\
\end{tabular}
\caption{The RMSE of the pre-averaging estimators. {\color{Blue} We set the parameter $\theta$ in \eqref{AsympCndkn} to be $5$, $3$, and $2$ for 10-sec, 30-sec, and 60-sec data, respectively. The truncation level is set to be $v_n = \alpha \sqrt{BPV \phi_{k_n}(g)}(\Delta_n)^{0.49}$, with $BPV = \frac{\pi}{2} \sum_{i=2}^{n}\left|\Delta_{i-1}^{n} X \| \Delta_{i}^{n} X\right|$ and calculated on sparsely sampled data (5 min frequency). When $\sigma_Y = 0$, $\alpha = 5,4,4$ for 10-sec, 30-sec and 60-sec data, respectively; when $\sigma_Y = 0.5$, $\alpha = 5.5,3.5,3$ for 10-sec, 30-sec and 60-sec data, respectively; when $\sigma_Y = 1.5$, $\alpha = 5.2,3,2.5$ for the respective frequencies; and, finally, in the case of jump with infinite activity, we set $\alpha = 6,4,4$ for 10-sec, 30-sec and 60-sec data, respectively.}}
\label{table:finite_jump}
\end{center}
\end{table}
\section{{\color{Blue} Conclusions}}\label{ConcludeSec}
{\color{Blue} We introduce high-frequency-based kernel estimators of the spot volatility in both the absence and presence of microstructure noise. One of the key differences of our results from those of earlier literature is to consider a general kernel in an asymptotic regime for the bandwidth that leads to optimal convergence rates for the resulting kernel estimators. Under this regime, kernels of unbounded support offer improved performance compared to uniform or other kernels with bounded support. General two-sided kernels of unbounded support were already advocated in the work of \cite{FigLi}, where it was proved for the first time that exponential kernels are optimal, hence, formally validating an old conjecture of \cite{foster1994continuous_optKernel}. Unfortunately, \cite{FigLi} imposed strong assumptions for the validity of their results, the most important of which are the absence of leverage effects, microstructure noise, and jumps. These three effects are, of course, pervasive in real transaction data. In this work, we are able to relax all of those constraints and consider a rather general model. We further develop a feasible implementation of the proposed estimators. Via Monte Carlo experiments, we confirm the superior performance of the proposed estimators.}
\section*{Acknowledgements}
{\color{Blue} The authors are grateful to the Associate Editor and two anonymous referees for their multiple suggestions that help to significantly improve the original manuscript.}