EconBase
← Back to paper

Estimating spot volatility under infinite variation jumps with dependent market microstructure noise

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

99,242 characters · 14 sections · 89 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating spot volatility under infinite variation jumps with dependent market microstructure noise

frontmatter\address[SUFE]{School of Statistics and Management, Shanghai University of Finance and Economics} \address[UM]{Department of Mathematics, University of Macau} \cortext[cor1]{Corresponding author. Email addresses: [email removed] (Qiang LIU), [email removed] (Zhi LIU). } \begin{abstract} Jumps and market microstructure noise are stylized features of high-frequency financial data. It is well known that they introduce bias in the estimation of volatility (including integrated and spot volatilities) of assets, and many methods have been proposed to deal with this problem. When the jumps are intensive with infinite variation, the {efficient} estimation of spot volatility {under serially dependent noise} is not available and is thus in need. For this purpose, we propose a novel estimator of spot volatility with a hybrid use of the pre-averaging technique and the empirical characteristic function. Under mild assumptions, the results of consistency and asymptotic normality of our estimator are established. Furthermore, we show that our estimator achieves an almost efficient convergence rate with optimal variance {when the jumps are either less active or active with symmetric structure}. Simulation studies verify our theoretical conclusions. We apply our proposed estimator to empirical analyses, such as estimating the weekly volatility curve using second-by-second transaction price data. \\ \\ JEL Classification: C13, C14, G10, G12 \end{abstract} \begin{keyword} Empirical characteristic function \sep High-frequency data \sep Jumps \sep Jump activity \sep Kernel smoothing \sep {Dependent} market microstructure noise \sep Pre-averaging \sep Spot volatility \end{keyword}

Introduction

In the current information era, computational technology is developing rapidly and is widely used in the financial market. Hence, high-frequency data is becoming increasingly available. Consequently, vast research, both in statistics and econometrics, has been conducted to analyze high-frequency data. One of the most widely studied problems is quantifying the variational strength of assets, namely, volatility. The volatility plays a crucial role in many areas of financial economics, including asset and derivative pricing, portfolio allocation, risk management, and hedging (see, e.g., BS1973, S1964, M1952, J2006). We refer to AJ2014 for a comprehensive introduction to the topic.

In an arbitrage-free and frictionless financial market, the logarithmic price process of an asset, say $\{{X_t}\}_{0\leq t\leq T}$, must be represented as a semi-martingale (see DS1994). Mathematically, $X_t$ can be written as

equation[equation omitted — 123 chars of source]

where $B$ is a standard Brownian motion, $J$ is a jump process, and $b$ and $\sigma$ are adapted and locally bounded c$\grave{\text a}$dl$\grave{\text a}$g processes. The continuous Brownian motion part describes the normal fluctuation of the financial market, while the jump process models accidental changes. The former, the volatility, what we are concerned to study, is completely determined by the coefficient process $\sigma$. Until now, three quantities regarding the information of $\sigma$ have been defined and investigated. They are integrated volatility $\int_{0}^{T} \sigma_{t}^2dt$, spot volatility $\sigma_\tau^2$ for any given $\tau \in[0,T]$, and the Laplace transform of volatility $\int_{0}^{T} e^{-u\sigma_{t}^2}dt$ with $u\in \mathbb{R}. $\footnote{This study focuses on (integrated or spot) volatility, and the estimation of the Laplace transform of volatility is investigated by TT2012a, TT2012b, TTG2011, WLX2019, WLX2019b, HLV2020, and etc.} In practice, the entire path of $X_t$ with $t\in[0,T]$ is not available, and only a finite number of data were observed at some discrete time points. Throughout this paper, we assume that the observation time points are equidistantly distributed along the time interval $[0,T]$; namely, the observed data are $\{X_{t_i}: i=0,1,\cdots,n\}$, where $n$ is the number of observations and $t_i = i\Delta_n$ with $\Delta_n = T/n$. Taking $n \rightarrow \infty$ results in high-frequency data, and based on this infilled setting, our theoretical results are established.

To characterize the intensity of jumps of a semi-martingale $X$ over $[0,T]$, YJ2009 introduced the jump activity index, which is defined as

align*[align* omitted — 84 chars of source]

where $\Delta X_t = X_t - X_{t^{-}}$ is the size of the jump at time $t$. We see from the definition that as $\beta$ increases, the (small) jumps tend to be more frequent: $X$ contains (almost surely) finite number of jumps if and only if $\beta = 0$; if $\beta < 1$, the jump part of $X$ is of finite variation, and $\beta > 1$ implies that $X$ contains an infinite variation jump part. For the finite activity jump part $J$, we can write it as $ J_t = \sum_{i=1}^{N_t} Y_i$, where $N_t<\infty, a.s.$ is the number of jumps up to time $t$, and $Y_i$ is the size of the $i$th jump. For the infinite activity case, we have the following expression separating the jump process into two parts (one with absolute jump sizes smaller than 1 and the other one with absolute jump sizes larger than one):

align*[align* omitted — 133 chars of source]

where $\mu$ is a jump measure and $\nu$ is its predictable compensator. For It$\hat{\text{o}}$ semi-martingales, we have $0 \leq \beta < 2$. When $X$ is a L$\acute{\text{e}}$vy process, the jump activity index coincides with the Blumenthal–Getoor index; in particular, if $X$ is a stable process, then $\beta$ is the stable index. Without consideration of jumps, the realized variance, which is the sum of squared increments, has been extensively studied in estimating integrated volatility (see AB1998, ABDL2003, BNS2002a, and many others). When jumps are present, the realized variance is no longer a consistent estimator of the integrated volatility; hence, jump-robust methods are proposed. The available approaches include the thresholding technique suggested by M2009, M2011, and MR2011, and the multi-power estimator introduced in BN2004, BGJPS2006, W2006, J2008, and references therein. Unfortunately, the jump part is restricted to at most finite variations to obtain the limiting distribution for both estimators. This issue not only limits the theory of inference of volatility, but also limits real applications because empirical studies in YJ2009 and JKLM2012 shown that jumps could be very frequent, with the estimated jump activity index being even larger than 1.5. When jumps of infinite variation are present, for the first time in the literature, JT2014 considered the estimation of integrated volatility, and KLJ2015 studied the test of pure-jump processes. JT2014 proposed a nonparametric integrated volatility estimator by using the empirical characteristic function of the observation increments, and LLL2018 extended their method to estimate spot volatility via kernel smoothing. Based on the empirical characteristic function, K2019 presented a novel test for jumps of infinite variation. All these mentioned literatures do not consider the presence of market microstructure noise.

As another stylized feature of financial data, market microstructure noise may be caused by price discreteness, bid-ask spread bounce, or rounding error. Mathematically, we observe $Y_{t_i}$ with

align[align omitted — 103 chars of source]

where $\epsilon_{t_i}$ represents the noise term at time $t_i$. The noise term introduces a bias in the estimation of volatility, and the most widely used technique to tackle this problem is pre-averaging, as proposed by JLMPV2009. The idea is that averaging multiple original observations can reduce the variance of the noise and bring the data closer to the latent process. Thus, constructing estimators based on the pre-averaged data, instead of the original observations, results in estimates closer to those based on the actual underlying process. The detailed realization of pre-averaging for our use and related analysis is presented in Section (ref). Other alternative methods to remove the bias from the market microstructure noise are the two-time scale and multi-time scale realized volatility estimators in ZMA2005 and Z2006, the quasi-maximum likelihood estimator in X2010, the realized kernel method in BNHLS2008a, among others.

From a practical perspective, it is more realistic to consider the simultaneous presence of jumps and market microstructure noise. An intuitive idea is to deal with jumps and market microstructure noise sequentially using the aforementioned techniques (see, e.g., FW2007, PV2009b, JLK2014, COP2014, YP2014, CHP2018, BNS2020, FW2022, JT2018, and references therein). { This study considers the estimation of spot volatility under both infinite variation jumps and dependent market microstructure noise. We first dispose the original observational data using the pre-averaging technique and then construct an estimation procedure based on the empirical characteristic function of the pre-averaged increments. Then, application of kernel smoothing results in our spot volatility estimator. Subsequently, the bias terms caused by both jumps and market microstructure noise are estimated and removed. Under some commonly used assumptions, we establish the consistency and asymptotic normality results for our estimator and verify their finite sample performance by extensive simulation studies. } { The contributions of our study can be summarized as follows. First, although the estimation of integrated volatility under jumps and market microstructure noise has been widely studied, the estimation of spot volatility with such a consideration is relative rare. More importantly, estimating spot volatility enables one to quantify the variational strength of an asset at any given time, compared with a fixed time interval for integrated volatility. Thus, the consideration of the former is more meaningful from the perspective of practical application. Second, our proposed spot volatility estimator is robust to jumps without any restrictions on the jump activity index $\beta$, and it achieves an almost efficient convergence rate of $(\log(n))^{a}/n^{1/8}$ for any $a>0$ with optimal variance {when the jump process satisfies either $\beta \leq \frac{3}{2}$ or $\beta < 2$ with symmetric property}. A concurrent work of FW2022 also constructed a kernel based spot volatility estimator based on pre-averaged increments, but they applied thresholding method to deal with the jumps. Their spot volatility estimator can only obtain the efficient convergence rate when $\beta<3/2$. When $\beta \geq 3/2$, the convergence rate will get slower as $\beta$ increases, and it becomes extremely slow when $\beta$ is close to 2. Simulation studies in Section (ref) also show that our estimator outperforms theirs when $\beta$ is large. Third, most, if not all, of the aforementioned works require the market microstructure noise to be mutually independent and identically distributed. Empirically, AMZ2011, JLZ2017, DX2021,LL2022 and references therein have found evidence of dependent microstructure noise in financial markets. To the best of our knowledge, our work is the first one considering serially dependent market microstructure noise in the estimation of spot volatility. Moreover, our noise structure also accommodates heteroscedasticity, endogenousness, and heavy-tailed distribution. }

The remainder of this paper is organized as follows. In Section (ref), we introduce our model and describe the conditions of the underlying data generation process and the market microstructure noise. In Section (ref), we explain the construction of our spot volatility estimator, and present their asymptotic properties in Section (ref). The simulation studies conducted in Section (ref) verify our theoretical results. In Section (ref), we apply the proposed method to real financial datasets. Section (ref) concludes the paper. All theoretical proofs are presented in Appendix.

Model setup

In this section, we describe the detailed formulation of $X$ in (ref) and $\epsilon$ in (ref) and give related assumptions on them.

Latent log-price process

The log-price process, $\{{X_t}\}_{0\leq t\leq T}$ in (ref), is defined in the filtered probability space $(\Omega, \mathcal{F}, \{ \mathcal{F}_{t}\}_{0 \leq t\leq T}, \mathrm{P})$. Specifically, for the jump process $J$, we decompose it as

equation[equation omitted — 179 chars of source]

where $K \geq 1$ is a finite number; $\gamma^{(k)}$ with $k=1,\cdots,K$ are adapted and locally bounded c$\grave{\text a}$dl$\grave{\text a}$g processes; $L^{(k)}$ are mutually independent {pure jump} L$\acute{\text{e}}$vy processes, they are also independent of $B$, and the Blumenthal-Getoor index of $L^{(k)}$ is $\beta_k$ with $1\leq \beta_K \leq \beta_{K-1} \leq \cdots\leq \beta_2\leq \beta_1<2$; $\delta(s,x)$ is a predictable process on $\Omega \times [0,T] \times \mathbb{R}$; $\mu'(ds,dx)$ is a homogeneous Poisson random measure on $[0,T] \times \mathbb{R}$. { For ease of demonstration, we separate $L^{(k)}$ as a positive jump part $L^{(k),+}$ and a negetive jump part $-L^{(k),-}$, and we write

align*[align* omitted — 131 chars of source]

where $L^{(k),+}_s$ and $L^{(k),-}_s$ are two independent L$\acute{\text{e}}$vy processes with the same index $\beta_{k}$ and positive jumps. It can be seen that if $L^{(k)}$ is symmetric, it is equivalent to assuming $\gamma^{(k),+} = \gamma^{(k),-} = \gamma^{(k)}$ with $L^{(k)} = L^{(k),+} - L^{(k),-}$. } According to JS2003, we can write $L^{(k),\pm}$ as

equation[equation omitted — 198 chars of source]

where $\mu^{(k),\pm}(ds,dx)$ is a Poisson random measure on $[0,T] \times \mathbb{R}$ with the compensator $\nu^{(k),\pm}(ds,dx)=F^{(k),\pm}(dx)ds$. We assume that the volatility process $\{\sigma_t\}_{0 \leq t\leq T}$ is also an It$\hat{\text{o}}$ semi-martingale, and it can be represented as

align[align omitted — 233 chars of source]

where $\widetilde{B}$ is a standard Brownian motion independent of $B$, $\tilde{\mu}(ds,dx)$ is a compensated homogeneous Poisson measure with L$\acute{\text{e}}$vy measure $dt \times \tilde{\nu}(dx)$, which has an arbitrary dependence structure with $\mu^{(k),\pm}(ds,dx)$ and $\mu'(ds,dx)$; $\tilde{b}, \tilde{\sigma}, \tilde{\sigma}'$ are adapted and locally bounded c$\grave{\text a}$dl$\grave{\text a}$g processes, and $\tilde{\delta}(s,x)$ is a predictable process on $\Omega \times [0,T] \times \mathbb{R}$.

For our complete model consisting of (ref), (ref), (ref), and (ref), we only require that $B$ and $L^{(k),\pm}$ with $k=1,\cdots,K$ are mutually independent, whereas the other coefficient processes and driving processes can have any dependence structures. In the joint presence of $B$ in $X$ and $\sigma$, the dependent relationship between $\tilde{\mu}(ds,dx)$ in (ref) and $\mu^{(k),\pm}(ds,dx)$ in (ref), $\mu^{'}(ds,dx)$ in (ref) depicts the continuous leverage effect and co-jumps in the price and volatility processes, respectively.

asuFor a sequence of stopping times $\{\tau_n, n=1,2,\cdots\}$ increases to infinity, a sequence of real values $a_n$, non-negative Lebesgue integrable function $J$ on $R$, and a real number $0 \leq r < 1$, and for $0 \leq t<s \leq t+1$ and $k=1,\cdots,K$, we have \begin{align} & \mathbf{E}[(V_{s\wedge \tau_n}-V_{t \wedge \tau_n})^2] \leq a_n(s-t), for V= b', \sigma, \gamma^{(k),\pm}, \tilde{\sigma}, \tilde{\sigma}', \int_{\mathbb{R}}\tilde{\delta}(\cdot,x) \tilde{\nu}(dx), \\ &|V_t|\leq a_n, \hbox{for} V=b', \sigma, \gamma^{(k),\pm}, \delta, \tilde{b}, \tilde{\sigma}, \tilde{\sigma}', \int_{\mathbb{R}}\tilde{\delta}(\cdot,x) \tilde{\nu}(dx),\\ & |\delta(\omega,t,x)|^r\wedge 1\leq a_nJ(x). \end{align} Moreover, $\sigma_{\tau}^2>0$ holds almost surely for any $\tau\in [0,T]$.

Conditions (ref)--(ref) in Assumption (ref) impose boundedness and continuity conditions on the coefficient processes of $X$ and $\sigma$ to guarantee their availability. These are standard and widely used in the high-frequency literature. Equivalently, according to the localization procedure stated in Section 4.4.1 of JP2012, we can assume that all the coefficient processes of $X$ and $\sigma$ are bounded. Moreover, condition (ref) for $\sigma$ limits its fluctuations over a short period of time, which enables us to treat it locally as a “constant" and consider the estimation of spot volatility. {

asuFor $k=1,\cdots,K$, the positive jump processes $L^{(k),\pm}$ are independent L$\acute{e}$vy processes with characteristics $(0,0,F^{(k),\pm})$. Moreover, for $x\in (0,1]$, there exist a uniform constant $r \in [0,1)$ and a function $g$, such that the tail functions $\overline{F}^{(k),\pm}(x) = F^{(k),\pm}((x,+\infty))$ satisfy \begin{align} \left| \overline{F}^{(k),\pm}(x)- \frac{1}{x^{\beta_k}} \right| \leq g(x), \end{align} where $g$ is a decreasing function with $\int_{0}^{1} x^{r-1}g(x)dx < \infty$.

} {Both Assumptions (ref) and (ref) involve a number $r\in[0,1)$. The magnitude of $r$ in Assumption (ref) controls the activity of the finite variation jump $J^{(2)}$, while the one in Assumption (ref) restricts the deviation degree from stable processes around zero for $L^{(k),\pm}$, which drive the infinite variation jump component of $J$. The smaller $r$ is, the stronger the two assumptions are.}

{Similar structure of $L^{(k),\pm}$ in Assumption 2.2 is also adopted by JT2014 for single infinite variation jump part. They demonstrated with detailed analysis therein that such an assumption can incorporate temper stable processes, which include time changed Brownian motion of normal inverse Gaussian process and CGMY model. Besides, any jump process of finite variation such as Merton's model, Kou's model, Variance Gamma process, Inverse Gaussian process can be modeled by $J^{(2)}$. Thus, our composition of the jump process $J$ is general enough to cover a wide range of jump models used in applications.\footnote{Interested readers can refer to CT2004 for the complete introduction and description of commonly used jump processes in financial modeling.} }

{In fact, further relaxations to (ref) can be considered to incorporate more general jump models. First, it was commented in JT2014 that (ref) on $L^{(k),\pm}/(a^{(k),\pm})^{1/\beta_k}$ and $(a^{(k),\pm})^{1/\beta_k} \cdot \gamma^{(k),\pm}$ is equivalent to

align*[align* omitted — 101 chars of source]

with any positive constants $a^{(k),\pm}$ for $L^{(k),\pm}$ and $\gamma^{(k),\pm}$. More generally, JT2016 allowed $a^{(k),\pm}$ to be time-varying and random. Furthermore, JT2018 considered $\beta_ka^{(k),\pm}$ for the assumption, which includes the generalized hyperbolic model as a special case.}\\

Market microstructure noise

{ If market microstructure noise is involved during the observation procedure, then a contaminated version of the data, as in (ref), will be obtained. For different financial models, the statistical assumptions on the noise can range from the simplest one of i.i.d. case, as in most of related literature, to very complex one with serially dependence, which is considered in AMZ2011, JLZ2017, JLZ2019, LLV2020,DX2021,LL2022 and references therein. In this paper, we consider a unified setting with the following generalized structure:

asuThe noise $\{\epsilon_{t_i}: i=0,\cdots,n\}$ can be realized as \begin{align} \epsilon_{t_i} = w_{t_i} \cdot\chi_{i}. \end{align} The sequence $\{w_{t_i}: i=0,\cdots,n\}$ is nonnegative and satisfies Lipschitz condition, namely, for $i,j=0,\cdots,n$, it holds that \begin{align} \mathbf{E}[(w_{t_i} - w_{t_j})^2] \leq |t_i - t_j|. \end{align} The sequence $\{\chi_{i}\}_{i\in \mathbb{Z}}$ is independent with $X$ and $w$, and $\chi_{i}$ is standard normal random variable with $\rho(i-j) = \mathbf{E}[\chi_{i} \chi_{j}].$ Moreover, there exists $d_n$ such that $\chi_{i} $ and $\chi_{j}$ are independent when $|i-j| >d_n$, and \begin{align} \sum_{i=0}^{d_n} \rho(i) < \infty. \end{align}

} { Assumption (ref) accommodates several empirical features of the microstructure noise, including heavy tails, endogenousness, heteroscedasticity, autocorrelation, etc. The tail part of $\epsilon$ is heavier than the one of $\chi$, which can be obviously seen from (ref). One special case of $w$ is nonnegative It$\hat{\text{o}}$ semimartingale defined on $(\Omega, \mathcal{F}, \{ \mathcal{F}_{t}\}_{0 \leq t\leq T}, \mathrm{P})$, and its driving Brownian motion and Poisson random measure can have any dependence structure with the ones in $X$, allowing $\epsilon$ and $X$ to be correlated. Moreover, $w$ can accommodate possible diurnal features of the noise so that the size of the noise may change over time, while the autocorrelation remains roughly unchanged. Our assumption allows for infinite-order autocorrelation since $d_n$ can tend to infinity as $n \rightarrow \infty$ and arbitrary shrinking magnitude if only (ref) holds. Such a consideration can cover many settings in existing literature as special cases, such as i.i.d. noise and finite-order dependent noise, exponentially decaying dependence in AMZ2011, polynomially decaying dependence in JLZ2017, JLZ2019, LLV2020 and LL2022. Moreover, we do not specify any serial dependence structure of $\chi$ such as moving average model considered in DX2021. } {

rmkWe have several comments on Assumption (ref) where we assume that the noise sequence $\{\chi_i\}_{i \in \mathbb{Z}}$ is from the normal distribution. The reason is as follows. To establish our central limit theorems in Section (ref), we require the following result: For any $j \in \mathbb{Z}$ and sequence $p_n\rightarrow \infty$ as $n\rightarrow \infty$, it holds that \begin{align} \mathbf{E} \left[ \cos\left(\frac{1}{\sqrt{p_n}} \sum_{i=j}^{j+p_n} \chi_i \right) \right] - \mathbf{E} \left[ \cos\left(\frac{1}{\sqrt{p_n}} \sum_{i=j}^{j+p_n} \chi'_i \right) \right] = o\left(\frac{1}{p_n^{1/4}}\right), \end{align} where $\{ \chi'_i\}_{i \in \mathbb{Z}}$ is a sequence of standard normal random variables with $\mathbf{E}[\chi'_{i} \chi'_{j}] = \rho(i-j)$ and (ref), $\chi'_{i} $ and $\chi'_{j}$ are independent when $|i-j| >d_n$. Under the normal assumption of $\chi_i$, the approximation error in (ref) vanishes, hence we do not need to consider it in the proof. However, the normal assumption can be relaxed. When the noise terms $\chi_i$'s are mutually independent, an assumption of finite fourth order moment on $\chi_i$ will be enough to obtain (ref), this is proved in Lemma 1 of WLX2019. In general dependent cases, it can be relaxed to conditions of finite moments of any order and mixing correlation coefficients, by applying a similar technique used in JLZ2017 and LL2022.

}

Spot volatility estimator

With the observed data $\{Y_{t_i}, i=0,1,\cdots, n\}$, we are interested in estimating the spot volatility $\sigma^2_\tau$ at a fixed time $\tau\in[0, T]$. Throughout this paper, we define $\Delta_i^nV := V_{t_i} - V_{t_{i-1}}$ {and $\Delta_{i,j}^nV := V_{t_{i+j-1}} - V_{t_{i-1}} = \sum_{i'=i}^{i+j-1} \Delta_{i'}^nV$} for a general process $V$ and $t_{j}^{i}:=(jp_n+i)\Delta_n$, where $p_n$ is a positive integer. $\lfloor x \rfloor$ denotes the integer part of a real number $x$, and $\mathbf{R}\{y\}$ denotes the real part of a complex number $y$.

Pre-averaging

To remove the influence of the market microstructure noise, we apply the pre-averaging technique in JLMPV2009 to our raw observations before using them to construct the spot volatility estimator.

asuThe function $g(x)$, supported in the interval $[0,1]$, is nonnegative, continuous, and piecewise continuously differentiable with a piecewise Lipschitz derivative, and \begin{align*} g(0) = g(1) =0, \quad \int_{0}^{1} g^2(x)dx > 0. \end{align*}

For a process $\{V_t\}_{0\leq t\leq T}$ observed at time points $\{t_i: i=0,1,\cdots,n\}$, we define the non-overlapping pre-averaged observations as

align[align omitted — 246 chars of source]

where $g_i^n = g(i/p_n)$. By some simple calculations, we can obtain

align*[align* omitted — 178 chars of source]

Let $p_n = O(n^{\eta})$ with $0 \leq \eta < 1/2$, we see from above that $\Delta_{jp_n}^{n}\overline{Y}$ is dominated by the noise part $\Delta_{jp_n}^{n}\overline{\epsilon}$, this makes extracting volatility information from the logarithmic price part $\Delta_{jp_n}^{n}\overline{X}$ impossible. Thus, the pre-averaging method does not work in this scenario. A special case is the raw increment of $\Delta_i^nY=\Delta_i^nX + \Delta_i^n\epsilon = O_p(\sqrt{\Delta_n}) + O_p(1)$. As $p_n$ increases, the order of $\Delta_{jp_n}^{n}\overline{\epsilon}$ decreases, and that of $\Delta_{jp_n}^{n}\overline{X}$ increases. If $p_n = O(n^{1/2})$, then $\Delta_{jp_n}^{n}\overline{X}$ and $ \Delta_{jp_n}^{n}\overline{\epsilon}$ have the same order, it is possible to estimate the volatility of $X$ if we can remove the bias. Furthermore, if $p_n = O(n^{\eta})$ with $1/2<\eta < 1$, then $\Delta_{jp_n}^{n}\overline{Y}$ is dominated by $\Delta_{jp_n}^{n}\overline{X}$, and the influence of $\epsilon$ is negligible. The drawback of this choice of $p_n$ is that the total number of pre-averaged data $\lfloor n/p_n \rfloor$ is small.

We define the following quantities regarding the pre-averaging function $g(x)$ for later use. For a real number $0 \leq \theta \leq 2$,

align[align omitted — 109 chars of source]

{ and

align[align omitted — 199 chars of source]

} We note that $\phi_{\theta}^n, \psi^{n} = O(1)$, \footnote{ According to Mean Value Theorem, there exist $i/p_n \leq s_i \leq (i+1)/p_n$ for $i=0,\cdots,(p_n-1)$ such that

align*[align* omitted — 311 chars of source]

under Assumptions (ref) and (ref). }and they will be directly formulated into our spot volatility estimator introduced in the next section.

Spot volatility estimator

Now, we are ready to define our estimator $\widehat{\sigma^2}_{\tau,n}(u,h)$ of the spot volatility $\sigma^2_{\tau}$ at any given time $\tau \in[0,T]$ as follows:

align[align omitted — 200 chars of source]

where $K_h(\cdot):= K(\cdot/h)/h$, with $K(\cdot)$ being the kernel function and $h$ the bandwidth parameter, $ v_n = \frac{1}{\phi_2^np_n^2\Delta_n}$, and {

align[align omitted — 736 chars of source]

}

asuThe kernel function $K(x)$, supported in the interval $[a, b]$, is nonnegative and continuously differentiable with \begin{equation} \int_a^bK^2(x)dx < +\infty, \int_a^b K(x)dx =1. \end{equation}

{ To explain the intuition behind our estimator of spot volatility, we exemplify a special case as follows. We consider $b_t \equiv 0, \sigma_t\equiv\sigma$ in (ref), $\delta(s,x) \equiv 0$ for $J^{(2)}$ in (ref) and $w_{t_i} \equiv w$ in (ref).\footnote{In fact, our theoretical results demonstrates that the drift term $b$ and the finite variation jump process $J^{(2)}$ can always be neglected.} For the infinite variation jump $J^{(1)}$ in (ref), we consider $K=1$ and $L^{(1)}$ being a strictly symmetric stable L$\acute{\text{e}}$vy process with $\gamma^{(1),+}_t = \gamma^{(1),-}_t \equiv \gamma$. Moreover, we assume that the characteristic functions of $L^{(1)}$ at time point $t=1$, $L^{(1)}_1$, can be written as $\mathbf{E}[e^{iuL^{(1)}_1}]=e^{-C_1|u|^{\beta_1}}$, where $u \in \mathbb{R}$ and $C_1$ is a constant. The analysis of our estimation procedure can be separated into the following steps:

enumerate• The quantity $S_{\tau,n}(u,h)$ in (ref) is close to $\mathbf{E}\Big[ \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big]$. To illustrate, we let $K(x) = 1_{\{-1 \leq x \leq 0\} }$, for which we have \begin{align*} S_{\tau,n}(u,h) &= \frac{p_n\Delta_n}{h} \sum_{j=1}^{\lfloor n/p_n \rfloor } K\Big(\frac{jp_n\Delta_n -\tau}{h}\Big) \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \\ & \approx \frac{1}{k_n} \sum_{j=\lfloor \frac{\tau-h}{p_n\Delta_n} \rfloor}^{\lfloor \frac{\tau}{p_n\Delta_n} \rfloor} \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)}, \end{align*} with $k_n = \Big\lfloor \frac{h}{p_n\Delta_n} \Big\rfloor$. It is obvious that $S_{\tau,n}(u,h)$ is the empirical version of the expectation $\mathbf{E}\Big[ \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big]$. • The explicit form of $\mathbf{E}\Big[ \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big]$ can be obtained based on the characteristic function of $\Delta_{jp_n}^n\overline{Y}$. We notice that \begin{align*} \Delta_{jp_n}^{n}\overline{Y} = \Delta_{jp_n}^{n}\overline{Y} + \Delta_{jp_n}^{n}\overline{\epsilon} = \sigma \Delta_{jp_n}^{n}\overline{B} + \gamma \Delta_{jp_n}^{n}\overline{L^{(1)}}+ \Delta_{jp_n}^{n}\overline{\epsilon}, \end{align*} and that after some simple calculations, we obtain that $\Delta_{jp_n}^{n}\overline{B} \sim \mathcal{N}(0, \phi_2^np_n\Delta_n)$ and $\Delta_{jp_n}^{n}\overline{\epsilon} \sim \mathcal{N}\Big(0, \displaystyle\frac{\psi^{n}w^2}{p_n}\Big)$. Because $B, L^{(1)}, \epsilon$ are mutually independent, we then have\footnote{Regarding the strictly symmetric stable L$\acute{\text{e}}$vy process $L^{(1)}$, since $\Big(\displaystyle\frac{L_{ct}^{(1)}}{c^{1/{\beta_1}}}\Big)_{t \geq 0} =^{d} (L_t^{(1)})_{t \geq 0}$ holds for any constant $c>0$, the characteristic function of $L_t^{(1)}$ for any $t \in [0,T]$ can be written as $\mathbf{E}[e^{iuL^{(1)}_t}]=e^{-C_1|u|^{\beta_1} t}$. This yields \begin{align*} \mathbf{E}\Big[\exp{ \Big( \frac{i\cdot u\gamma\Delta_{jp_n}^n\overline{L^{(1)}}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] &= \mathbf{E}\Big[\exp{ \Big( \frac{i\cdot u\gamma \sum_{i=1}^{p_n-1}g_i^n \Delta_{jp_n+i}^{n}L^{(1)} }{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] \\ &= \prod_{i=1}^{p_n-1} \mathbf{E}\Big[\exp{ \Big( \frac{i\cdot u\gamma g_i^n \Delta_{jp_n+i}^{n}L^{(1)} }{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] = \prod_{i=1}^{p_n-1} \exp{\Big(-C_1 \Big| \frac{u\gamma g_i^n\Delta_n^{1/\beta_1}}{\sqrt{\phi_2^np_n\Delta_n}} \Big|^{\beta_1} } \Big) \\ & = \exp{\Big(-C_1\frac{\phi_{\beta_1}^{n}|u\gamma|^{\beta_1}}{(\sqrt{\phi_2^n})^{\beta_1}}(p_n\Delta_n)^{1-\frac{\beta_1}{2}}\Big) }. \end{align*}} \begin{align} \begin{split} & \mathbf{E}\Big[ \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] = \mathbf{R}\Big\{ \mathbf{E}\Big[\exp{ \Big(\frac{i \cdot u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] \Big\} \\ &= \mathbf{R}\Big \{ \mathbf{E}\Big[\exp{\Big(\frac{i\cdot u\sigma \Delta_{jp_n}^{n}\overline{B}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] \cdot \mathbf{E}\Big[\exp{ \Big( \frac{i\cdot u\gamma\Delta_{jp_n}^n\overline{L^{(1)}}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] \cdot \mathbf{E}\Big[\exp{\Big(\frac{\text{i}\cdot u \Delta_{jp_n}^{n}\overline{\epsilon}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] \Big \} \\ &= \exp\Big(\frac{-u^2\sigma^2}{2}\Big)\cdot \exp{\Big(-C_1\frac{\phi_{\beta_1}^{n}|u\gamma|^{\beta_1}}{(\sqrt{\phi_2^n})^{\beta_1}}(p_n\Delta_n)^{1-\frac{\beta_1}{2}}\Big) } \cdot \exp\Big(\frac{-u^2w^2v_n\psi^{n} }{2}\Big). \end{split} \end{align} • The spot volatility can be retrieved from the composition of $\mathbf{E}\Big[ \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big]$. Because $p_n\Delta_n \rightarrow 0$ and $0 \leq \beta_1< 2$, we have that, as $n\rightarrow \infty$, \begin{align} \mathbf{E}\Big[ \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] \rightarrow \exp\Big(\frac{-u^2\sigma^2}{2}\Big) \cdot \exp\Big(\frac{-u^2w^2v_n\psi^{n} }{2}\Big). \end{align} Then, taking logarithm to the estimator of the left-hand side term in (ref) and multiplying $\frac{-2}{u^2}$ sequentially results in an estimator of $\sigma^2 + v_n\psi^{n}w^2$. • The spot volatility estimator can be naturally obtained after removing a consistent estimator of $v_n\psi^{n}w^2$. The quantity $\widehat{w^2}_{\tau,n}(h)$ in (ref) is a kernel based estimator of $\psi^{n}w^2$ with a faster convergence rate than the one for spot volatility estimator (The rigorous proof is given by Lemma (ref) in Appendix). • Eliminating the influence from the jump process $L^{(1)}$ is necessary for central limit theorem. We see from (ref) and the analysis in Step 3 above that the presence of $L^{(1)}$ bring in an extra bias term with the form \begin{align*} \frac{2C_1}{u^2}\frac{\phi_{\beta_1}^{n}|u\gamma|^{\beta_1}}{(\sqrt{\phi_2^n})^{\beta_1}}(p_n\Delta_n)^{1-\frac{\beta_1}{2}}. \end{align*} This term converges to 0 as $n$ tends to infinity, which means it has no effect on the consistency of our spot volatility estimator. However, to establish the central limit theorem, the convergence rate must be faster than the main term. Thus, a further procedure for the purpose of estimating and removing the bias term has to be considered. We will present the detailed discussion in Section (ref) later.

} For our proposed spot volatility estimator, there are some other issues worth to be noted. From the above analysis, we know that, as $n \rightarrow \infty$, the limiting value of $S_{\tau,n}(u,h)$ lies within the interval $(0,1)$, and the thresholding step with scales of $1/n$ and $(n-1)/n$ in (ref) is to guarantee this and excludes the possibility of outliers. The threshold plays no role asymptotically. Furthermore, such disposal guarantees the validity of the logarithmic function. Recall that we provide two different choices for $p_n$ when discussing the pre-averaging method in Section (ref). One of which corresponds to $p_n^2\Delta_n= O(1)$, and the other one satisfies $p_n^2\Delta_n \rightarrow \infty$ as $n\rightarrow \infty$. For the latter case, we have $v_n \rightarrow 0$, thus in (ref), $\mathbf{E}\Big[ \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big] \rightarrow \exp\Big(\frac{-u^2\sigma^2}{2}\Big)$. Namely, the influence of the noise $\epsilon$ on the consistency of our estimator is asymptotically negligible. As a result, the term $- v_{n} \cdot \widehat{w^2}_{\tau,n}(h)$ in (ref) is no longer necessary under this scenario. But we do not separately discuss this case because involving such a procedure elevates the performance of the estimator, especially for finite samples.

rmkFor the implementation of $S_{\tau,n}(u,h)$ in (ref) and $\widehat{w^2}_{\tau,n}(h)$ in (ref), the following factors $F_S$ and $F_w$ are further multiplied, respectively, for finite sample adjustment: \begin{align*} F_S = \frac{ 1 }{ \frac{p_n\Delta_n}{h} \sum_{j= 1 \vee \lceil \frac{ah+\tau}{p_n\Delta_n} \rceil}^{\lfloor \frac{n}{p_n} \rfloor \wedge \lfloor \frac{bh+\tau}{p_n\Delta_n} \rfloor } K\Big(\frac{jp_n\Delta_n -\tau}{h} \Big) }, \quad F_w= \frac{ 1 }{\frac{\Delta_n}{h} \sum_{i= 1 \vee \lceil \frac{ah+\tau}{\Delta_n} \rceil}^{n \wedge \lfloor \frac{bh+\tau}{\Delta_n} \rfloor } K\Big(\frac{j\Delta_n -\tau}{h} \Big) }, \end{align*} since $K(x)$ is only defined for $x \in [a,b]$. On one hand, such an adjustment can always avoid unnecessary discretization errors in approximating the integral $\int_{a}^{b} K(x)dx = 1$. On the other hand, the adjustment is necessary when the time point $\tau$ is close to the endpoint $0$ ($T$) with $a<0$ ($b>0$). For convenience of presentation, one default premise for our later theoretical analysis is \begin{align*} \Big \lfloor \frac{n}{p_n} \Big \rfloor > \frac{bh+\tau}{k_n\Delta_n}, \quad 1 < \Big \lfloor \frac{ah+\tau}{k_n\Delta_n} \Big\rfloor, \quad n > \frac{bh+\tau}{\Delta_n}, \quad 1 < \Big \lfloor \frac{ah+\tau}{\Delta_n} \Big \rfloor. \end{align*} They are always well satisfied by choosing a proper domain $[a,b]$.

Asymptotic properties

In this section, we demonstrate the consistency and asymptotic normality properties of the spot volatility estimator $\widehat{\sigma^2}_{\tau,n}(u,h)$. Under Assumption (ref), the characteristic functions of $L^{(k),\pm}$ with $k=1,\cdots,K$ can be obtained. Based on that, the bias term due to jump can be estimated and removed. We use the notations $\rightarrow^{p}$ and $\rightarrow^{d}$ for convergence in probability and convergence in distribution, respectively.

thmUnder Assumptions (ref)-(ref), (ref)-(ref), and assume that, as $\Delta_n\rightarrow 0$, it holds that \begin{align} h\rightarrow 0, \quad p_n\rightarrow \infty, \quad \frac{h}{p_n\Delta_n} \rightarrow \infty, \quad \frac{p_n^2\Delta_n}{n^{1/3}} \rightarrow 0, \end{align} and meanwhile, \begin{align*} (d_n)^2 \sqrt{h} \rightarrow 0, \quad (d_n)^2 \sqrt{\frac{d_n \Delta_n}{h}} \rightarrow 0. \end{align*} Then, for any fixed $\tau \in [0,T]$ and $u\in \mathbb{R}$, we have \begin{equation} \widehat{\sigma^2}_{\tau,n}(u,h) \rightarrow^p \sigma^2_{\tau}. \end{equation}

{ From (ref), we see that owing to the infinite variation jump processes $L^{(k),\pm}$ with $k=1,\cdots,K$, the estimator $\widehat{\sigma^2}_{\tau,n}(u,h)$ suffers a bias term. The bias term $b_{\tau,n}(u)$ takes the following form\footnote{See Lemma (ref) in the Appendix.}

align[align omitted — 547 chars of source]

with $C_k= \int_{0}^{\infty} \frac{\sin x}{x^{\beta_k}}dx$, $D_k = \int_{0}^{\infty} \frac{1- \cos x}{x^{\beta_k}}dx$ and the notation $\langle x \rangle^{\beta} := \text{sign}(x) \cdot |x|^{\beta}$. Because $K,C_k, D_k, \phi_{\beta_k}^{n}, \phi_2^n, \gamma_{\tau}^{(k),\pm}$ are finite and $1\leq \beta_K\leq \beta_{K-1} \leq \cdots\leq \beta_2\leq \beta_1<2$, we have

align[align omitted — 109 chars of source]

}

To obtain the asymptotic normality property of our estimator, we require the parameter $u$ to be a series $u_n$ converging to 0 as $\Delta_n \rightarrow 0$.

thmUnder Assumptions (ref)-(ref), (ref)-(ref), \begin{enumerate} • If $\beta_1\leq 1.5$, and, as $\Delta_n \rightarrow 0$, it holds that, $u_n\rightarrow 0, ~h\rightarrow 0, ~p_n\rightarrow \infty$ and \begin{eqnarray} \sup{\frac{h}{u_n^4\sqrt{p_n\Delta_n}}} < \infty, \frac{\sqrt{p_n\Delta_n}}{u_n^2\sqrt{h}} \rightarrow 0, \quad \frac{p_n^2\Delta_n}{n^{1/5}} \rightarrow 0, \end{eqnarray} and meanwhile, \begin{align*} \frac{(d_n)^2h}{\sqrt{p_n \Delta_n}} \rightarrow 0, \quad \frac{(d_n)^5}{p_n} \rightarrow 0. \end{align*} Then, for any fixed $\tau \in [0,T]$, we have \begin{equation} \sqrt{\frac{h}{p_n\Delta_n}}\cdot\frac{\widehat{\sigma^2}_{\tau,n}(u_n,h)-\sigma_\tau^2}{\sqrt{2} (\sigma_{\tau}^2 + v_n\psi^n w_{\tau}^2)} \rightarrow^d\Big(\int_a^b K^2(x)dx \Big)^{1/2}\cdot {\cal N}(0,1). \end{equation} • If $ \beta_1 < 2$, and, for any $\delta>0$, as $\Delta_n \rightarrow 0$, it holds that, $u_n\rightarrow 0,~h\rightarrow 0,~p_n\rightarrow \infty$ and \begin{eqnarray} \sup{\frac{h}{u_n^6\sqrt{p_n\Delta_n}}} < \infty, \frac{h}{(p_n\Delta_n)^{1/2+\delta} } \rightarrow \infty, \frac{p_n^2\Delta_n}{n^{1/5}} \rightarrow 0, \end{eqnarray} and meanwhile, \begin{align*} \frac{(d_n)^2h}{\sqrt{p_n \Delta_n}} \rightarrow 0, \quad \frac{(d_n)^5}{p_n} \rightarrow 0. \end{align*} Then, for any fixed $\tau \in [0,T]$, we have \begin{equation} \sqrt{\frac{h}{p_n\Delta_n}}\cdot\frac{\widehat{\sigma^2}_{\tau,n}(u_n,h)-\sigma_\tau^2 - b_{\tau,n}(u_n) }{\sqrt{2} (\sigma_{\tau}^2 + v_n\psi^n w_{\tau}^2)} \rightarrow^d\Big(\int_a^b K^2(x)dx \Big)^{1/2}\cdot {\cal N}(0,1). \end{equation} \end{enumerate}

First, we discuss the convergence rate and asymptotic variance in (ref). For the pre-averaging procedure introduced in Section (ref), we fix our choice of $p_n$ with $p_n = O(n^{1/2})$, under which $\Delta_{jp_n}^{n}\overline{X}$ and $ \Delta_{jp_n}^{n}\overline{\epsilon}$ are of the same order. { For condition (ref), with any $a>0$, we set $u_n = O\left( \frac{1}{(\log(n))^{a/2}}\right), h = O\left(\frac{1}{n^{1/4}(\log(n))^{2a}}\right)$. Then $\sqrt{\frac{h}{p_n\Delta_n}}$ in (ref) equals $\frac{n^{1/8}}{(\log(n))^a}$, resulting in a convergence rate of $\frac{(\log(n))^a}{n^{1/8}}$ for our spot volatility estimator. } This rate is almost efficient for the estimation of spot volatility with presence of market microstructure noise, compared with the efficient rate of $n^{-1/8}$ discussed in YFZZ2013, FW2022. Comparing our conclusion (ref) with the ones in these studies, we can conclude that the limiting variance of our estimator is also optimal.

By comparing (ref) and (ref), we see that our spot volatility estimator $\widehat{\sigma^2}_{\tau,n}(u_n,h)$ is free from the bias term $b_{\tau,n}(u)$ in (ref) when $\beta_1 \leq 1.5$, but if $\beta_1 > 1.5$, the bias term is not asymptotically negligible. To see how this critical value of 1.5 is derived, let us consider $u_n =O(n^{-x_1}) $ and $ h = O(n^{-x_2})$ for illustration. Under this setting, condition ((ref)) is equivalent to $(x_1,x_2) \in R_1$ with

align*[align* omitted — 132 chars of source]

To make $b_{\tau,n}(u_n)$ in (ref) asymptotically negligible in the central limit theorem (ref), we require

align*[align* omitted — 143 chars of source]

which is equivalent to

align*[align* omitted — 100 chars of source]

Observing that $\min_{(x_1,x_2)\in R_1} \beta_0(x_1,x_2) > 1.5$ and $\min_{(x_1,x_2)\in R_1} \beta_0(x_1,x_2) \rightarrow 1.5$ as $(x_1,x_2) \rightarrow (0,\frac{1}{4})$, we get the condition $\beta_1 \leq 1.5$. In this case, the convergence rate of the spot volatility estimator is $n^{x_2/2 - 1/4}$, and a faster rate can be obtained if we consider a slower convergence rate for $u_n$ with order $O\left(\frac{1}{(\log (n))^{a/2}}\right)$, as discussed above.

For the general case of $\beta_1 < 2$, (ref) implies that a nearly optimal convergence rate can be achieved by our spot volatility estimator if $b_{\tau,n}(u_n)$ is known.\footnote{In practice, since $b_{\tau,n}(u_n)$ is unknown, we will construct estimator to remove this bias later. {Furthermore, we need to assume that the jumps are symmetric so as to guarantee the efficiency of both convergence rate and asymptotic variance.}} Under (ref),\footnote{Compared with (ref), condition (ref) is more restrictive, which can be seen from the following expressions:

align*[align* omitted — 283 chars of source]

} the convergence rate of $\frac{(\log(n))^a}{n^{1/8}}$ for any $a>0$ is obtained when we set $u_n = O( \frac{1}{(\log(n))^{a/3}}), h = O(\frac{1}{n^{1/4}(\log(n))^{2a}})$. {We stress that this convergence rate holds for any $\beta_1<2$. In FW2022, the authors employed the pre-averaging method to filter the noise and the thresholding method to remove the jumps. We compare the convergence rate of our estimator with the one proposed in FW2022\footnote{Given $\beta_1 < \widehat{\beta}$, namely $r < \widehat{\beta}$ in FW2022, and according to (2.23) therein, we can get the marginal value of $a$ by taking

align*[align* omitted — 132 chars of source]

Since (2.22) implies that $m_n = O(n^{a})$, combining this with the conclusions in (2.21) results in the fastest convergence rate.} in Table (ref). We see that their convergence rate decreases fast when $\beta_1$ is not smaller than $3/2$, and the rate gets extremely slow when the jump activity index is close to 2.

table[table omitted — 597 chars of source]

}

Although the presence of infinite variation jumps does not affect the limiting behavior of $\widehat{\sigma^2}_{\tau,n}(u_n,h)$ if $\beta_1 \leq 1.5$, prior knowledge of $\beta_1$ is not achievable. Furthermore, when $\beta_1 < 2$, the result (ref) is not applicable because $b_{\tau,n}(u_n)$ is unknown in practice. Thus, a further procedure removing the bias term $b_{\tau,n}(u_n)$ is by no means required. Moreover, eliminating $b_{\tau,n}(u_n)$ always improves the finite sample performance of our spot volatility estimator, even if the bias is negligible.

To this end, {we assume that $L^{(k)}$ are symmetric}. By following the idea in LLL2018, we firstly construct a de-biased estimator that works for $K=1$. We recall that $K$ is the total number of infinite variation jump processes, and define

align[align omitted — 152 chars of source]

with

align[align omitted — 289 chars of source]

where $\lambda > 1$, and the ratio above is set to zero when its denominator vanishes. { To explain the intuition, we note that if $L^{(k)}$ are symmetric, namely $\gamma^{(k),+} = \gamma^{(k),-} = \gamma$, the bias term $b_{\tau,n}(u)$ in (ref) then becomes

align[align omitted — 242 chars of source]

Specifically, with $K=1$, we notice that $ b_{\tau,n}(u)$ is linear with respect to $u$, in the sense that

align[align omitted — 102 chars of source]

Plugging (ref) and the conclusion $\widehat{\sigma^2}_{\tau,n}(\lambda u_n,h) \approx \sigma_{\tau}^2 + b_{\tau,n}(\lambda u_n)$ from Theorem (ref) into ((ref)) yields $\widehat{B}_{\tau,n}(\lambda,u_n,h) \approx b_{\tau,n}(u_n) $, which enables us to estimate the bias and remove it as done in (ref). } When $K >1 $, the estimator $\widehat{B}_{\tau,n}(\lambda,u_n,h)$ is generally not sufficient to remove the bias term $b_{\tau,n}(u_n)$ since the result (ref) does not hold for (ref). But (ref) still holds for each summanding term in (ref), this inspires us to remove the bias term in an iterative way as follows.

enumerate• Select real numbers $\lambda > 1$, $\xi>0$ (typically small), and initially set \begin{align} \widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda,\xi,0) = \widehat{\sigma^2}_{\tau,n}(u_n,h). \end{align} • With $\widehat{\sigma^2}_{\tau,n}(u_n, h, \lambda, \xi, i-1)$ for $i=1,\cdots, K-1$, iterate \begin{align} \begin{split} & \widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda,\xi,i) = \widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda,\xi,i-1) + u_n^2\sqrt{\frac{p_n\Delta_n}{h}}\xi \\ & - \frac{(\widehat{\sigma^2}_{\tau,n}(\lambda u_n,h,\lambda,\xi,i-1) - \widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda,\xi,i-1))^2}{\widehat{\sigma^2}_{\tau,n}(\lambda^2u_n,h,\lambda,\xi,i-1) - 2\widehat{\sigma^2}_{\tau,n}(\lambda u_n,h,\lambda,\xi,i-1) + \widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda,\xi,i-1)}, \end{split} \end{align} and take 0 for the above ratio when its denominator vanishes. • Take $\widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda,\xi,K)$ as the de-biased estimator.

The same idea is also considered in JT2016 and LL2020. In fact, the above iterative procedure further requires some specific structures on $\beta_1, \cdots, \beta_K$.

asu{For $k=1,\cdots,K$, the jump processes $L^{(k)}$ are symmetric,} their jump activity indices $\beta_1, \cdots, \beta_K$ are scattered on the discrete points $\{ 2-i\rho: i=1,\cdots,\lfloor \frac{1}{\rho} \rfloor \}$ for some unknown constant $\rho \in(0,1)$.

For the debiased estimators introduced above, we have the following version of central limit theorem for general situation $\beta_1 < 2$.

thmUnder Assumptions (ref)-(ref), (ref)-(ref), (ref), for any $\delta>0$, if $\Delta_n \rightarrow 0$, it holds that, $u_n\rightarrow 0,~h\rightarrow 0,~p_n\rightarrow \infty$ and \begin{eqnarray} \sup{\frac{h}{u_n^6\sqrt{p_n\Delta_n}}} < \infty, \frac{h}{(p_n\Delta_n)^{1/2+\delta} } \rightarrow \infty, \frac{p_n^2\Delta_n}{n^{1/5}} \rightarrow 0, \end{eqnarray} \begin{enumerate} • when $K=1$, for any fixed $\tau \in [0,T]$, we have \begin{equation} \sqrt{\frac{h}{p_n\Delta_n}}\cdot\frac{ \widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda)-\sigma_\tau^2}{\sqrt{2} (\sigma_{\tau}^2 + v_n\psi^n w_{\tau}^2)} \rightarrow^d\Big(\int_a^b K^2(x)dx \Big)^{1/2}\cdot {\cal N}(0,1). \end{equation} • when $K>1$, for any fixed $\tau \in [0,T]$, we have \begin{equation} \sqrt{\frac{h}{p_n\Delta_n}}\cdot\frac{\widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda,\xi,K)-\sigma_\tau^2 }{\sqrt{2} (\sigma_{\tau}^2 + v_n\psi^n w_{\tau}^2)} \rightarrow^d\Big(\int_a^b K^2(x)dx \Big)^{1/2}\cdot {\cal N}(0,1). \end{equation} \end{enumerate}

Note that the asymptotic condition (ref) in Theorem (ref) is the same as (ref) in Theorem (ref), following the similar analysis on the convergence rate and variance for the result (ref), we conclude that the estimators $\widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda)$ and $\widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda,\xi,K)$ in Theorem (ref) also have an almost efficient convergence rate of $\frac{(\log(n))^a}{n^{1/8}}$ for any $a>0$ and an optimal limiting variance.

We observe that the denominator on the left-hand side of (ref) and (ref) involve unknown quantities $\sigma_{\tau}^2$ and $\psi^n w_{\tau}^2$, replacing them with their respective consistent estimators $\widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda)$ (for $K=1$), $\widehat{\sigma^2}_{\tau,n}(u_n,h,\lambda,\xi,K)$ (for $K>1$), and $\widehat{w^2}_{\tau,n}(h)$ provides feasible versions of central limit theorem.

{ We now give some comments on the effect of the jump process $J$ in (ref) on the estimation of spot volatility. The finite variation jump part $J^{(2)}$ does not have influence on the consistency and central limit theorem of our spot volatility estimator. This is also true for $J^{(1)}$ when $\beta_1 \leq 1.5$, as seen from (ref) in Theorem (ref). The presence of $J^{(1)}$ brings in a bias term, $b_{\tau,n}(u)$ in (ref), and its explicit form can be obtained by using the characteristic function of $J^{(1)}$ under Assumption (ref). Moreover, to remove the bias term, symmetric condition on $L^{(k)}$ is necessary so that the linear relationship in (ref) can be applied. The symmetry assumption is also required in related literature (JT2014,JT2016,JT2018,LLL2018,LL2020). A possible way to avoid the assumption is by differencing the adjacent increments so that the jump component of the new ones are symmetric, but it is at a cost of increasing the asymptotic variance of the estimator by 2, which is documented and discussed in JT2014. {In a word, our spot volatility estimator achieves both efficient convergence rate and efficient variance when either $\beta_1 \leq 1.5$ or $\beta_1 <2$ with symmetric jumps.} }

rmkWe note that Assumption (ref) is not a prerequisite. It is more appropriate to say that if Assumption (ref) is satisfied, then the minimum iteration number of (ref) required for (ref) to hold is $K$. If Assumption (ref) is not satisfied, then we can always make it true by decreasing $\rho$ and adding “fictitious" indices so that the indices fill in the whole set $\{ 2-i\rho: i=1,\cdots,\lfloor \frac{1}{\rho} \rfloor \}$. Meanwhile, we set the associated coefficient processes $\gamma^{(k)}$ to be 0 for all those “fictitious" indices. This only increases the number of iterations and does not affect our de-biasing procedure (ref), so the result (ref) remains valid with these new indices. Our analysis also implies that, because we do not know the real value of $K$, a relatively larger choice of iteration times is preferred and is not harmful from the asymptotic viewpoint. However, for finite samples, the de-biasing procedure (ref) can make the whole estimation unstable; thus, the iteration times cannot be too large. Interested readers can refer to JT2016 for more discussion on choosing the proper iteration times. We must admit that the instability problem of the de-biasing procedure remains unsolved, which limits its application. Thus, we do not consider verifying Theorem (ref) via simulation studies in the next section.

Simulation studies

In this section, we demonstrate the finite sample performance of our estimator $\widehat{\sigma^2}_{\tau,n}(u_n,h)$ and verify its asymptotic properties developed in the last section via simulation studies.

Simulation design

We consider the following underlying data generating process, whose continuous part is the widely used Heston model, as the log price process of an asset:

align[align omitted — 285 chars of source]

where $B, W$ are standard Brownian motions with $\mathbf{E}[B_tW_t] =-0.3 t$; $L^{(1)}$ and $L^{(2)}$ are two mutually independent strictly symmetric stable L$\acute{\text{e}}$vy processes with Blumenthal-Getoor index $\beta_1, \beta_2$ respectively; $\sum_{i=1}^{N_t} \xi_{i}$ is a compound Poisson process where $N_t$ is a Poisson process with intensity $\lambda'=3$ and $\xi_i\stackrel{i.i.d}\sim \mathcal{N}(0, 1)$; the initial value of the volatility process, namely, the spot volatility $\sigma_0^2$, is randomly sampled from its stationary distribution $\Gamma(a,b)$ with scale parameter $b=\frac{2\cdot6}{(0.5)^2}$ and shape parameter $a= 0.25\cdot b$. For the parameters regarding the continuous part of (ref), we adopt the same setting as in LT2014 and LL2020. For the discontinuous part, we fix $\gamma^{(1)}= 0.15, \gamma^{(2)}=0.05$, $\beta_2 = 1$, and vary the magnitude of $\beta_1$. The generating method of $L^{(1)}, L^{(2)}$ is based on the built-in generator of random variables with stable distribution in the software R. We generate $L^{(1)}, L^{(2)}$ such that $\mathbf{E}[e^{iuL^{(1)}_1}]=e^{-|u|^{\beta_1}}$ and $\mathbf{E}[e^{iuL^{(2)}_1}]=e^{-|u|^{\beta_2}}$. {As to the noise process $\epsilon$ in Assumption (ref), we consider the stationary case with $w_{t_i} \equiv \sigma_{\epsilon}$ and

align[align omitted — 142 chars of source]

where $s$ is a constant within $(-0.5, 0.5)$, $\chi_{i} \sim^{i.i.d.} \mathcal{N}(0,1)$. It includes the special case of $\epsilon_{t_i} \sim^{i.i.d.} \mathcal{N}(0,\sigma_{\epsilon}^2)$ by taking $d_n = 0$; that is the setting in JLMPV2009, JLK2014, WLX2019. When $d_n \geq 1$, $\chi$ is a moving average (MA) model which approximates a fractionally difference process, and its autocorrelation decays slowly, which is inline with the empirical founding of JLZ2017. The model (ref) is also considered by JLZ2019.\footnote{We also note that, for finite $d_n$, the autocorrelation function first decays slowly till lag $d_n$ and vanishes after $d_n$; and for $d_n = \infty$, the autocorrelation function decays polynomially at the rate of $2s-1$ (see pp. 72-73 in T2002). } }

We take $T=1$ week and $n=5 \times 6.5 \times 3600$, corresponding to 1-second data in a 6.5-hour trading day for one week (five consecutive trading days) in practice. We take the weight function $g(x) = x \wedge (1-x)$ and $p_n = \lfloor \frac{1}{3} \sqrt{n} \rfloor = 114$ for pre-averaging the original observations, where the choice of constant $\frac{1}{3}$ is suggested by JLMPV2009 and widely adopted in the literature. For our estimator $\widehat{\sigma^2}_{\tau,n}(u_n,h)$, we set the tuning parameter $u_n$ as

align[align omitted — 275 chars of source]

where $(\log(n))^{-1/24}$ satisfies the condition (ref) when our spot volatility estimator achieves an almost efficient convergence rate with $h = O(\frac{1}{n^{1/4}(\log(n))^{1/6}})$. The denominator consistently estimates $\sqrt{\sigma_{\tau}^2+ v_n \psi^n w_{\tau}^2}$, as shown in (ref), Lemma (ref), and Theorem (ref). From (ref), we see that $S_{\tau,n}(u_n,h)$ is close to $\exp\{\frac{-u_n^2}{2}(\sigma_{\tau}^2+ v_n \psi^n w_{\tau}^2)\}$, which shows that the magnitude of $u_n$ is always deemed large or small with respect to that of $(\sigma_{\tau}^2+ v_n \psi^n w_{\tau}^2)$. Theoretically, we require $u_n\rightarrow 0$ as $n\rightarrow \infty$, but $u_n$ is a constant when $n$ is fixed in practice. We do such a scaling to guarantee that $u_n$ is sufficiently small, and is free from the influence of $(\sigma_{\tau}^2+ v_n \psi^n w_{\tau}^2)$. Similar adjustments were also made in JT2018, LLL2018, and LL2020.

Simulation results

First, we demonstrate the finite sample performance of our estimator $\widehat{\sigma^2}_{\tau,n}(u_n,h)$ for different choices of time point $\tau$, kernel function $K(x)$, bandwidth parameter $h$, noise variance $\sigma_{\epsilon}^2$ for $i.i.d.$ noise (namely $d_n = 0$), and jump activity index $\beta_1$ to see how these parameters impact our estimation procedure. For each set of parameters, we use Euler discretization to generate the sample path of (ref), and then obtain the value of $\widehat{\sigma^2}_{\tau,n}(u_n,h)$ with the finite sample adjustment introduced in Remark (ref). We repeat the procedure for 5000 times and record the average of the following relative biases (R.B.):

align[align omitted — 121 chars of source]

and their standard deviations (S.D.). We note that the theoretical relative bias (T.R.B.) and their theoretical standard deviations (T.S.D.) can be written as

align[align omitted — 246 chars of source]

with

align[align omitted — 450 chars of source]

which are obtained from (ref), (ref), (ref) with $d_n=0$, and Remark (ref). We let the parameters $\tau = 0, 0.5$, $h=n^{-0.26}, n^{-0.3}, n^{-0.35} $, $\sigma_{\epsilon}^2 = 0.01^2, 0.03^2, 0.05^2$, $\beta_1 = 1.2, 1.5, 1.8$, and consider different kernel functions of

align[align omitted — 296 chars of source]

and record the results in Table (ref). We see from the table that

enumerate• The values of S.D. at $\tau=0.5$ are slightly smaller than that at $\tau=0$. This is because fewer data (half, to be precise) are used at $\tau = 0$ compared with the estimation at $\tau=0.5$. Similarly, with fixed $\tau, \sigma_{\epsilon}^2, \beta_1, K(x)$, a relatively larger selection of $h$ always results in a smaller value of S.D., because more data are incorporated to calculate the estimator. • For fixed $\tau,h, \sigma_{\epsilon}^2, \beta_1$, the magnitude of S.D. increases as we change the kernel function from $K_1$ to $K_4$, since the value of $K^2$ in (ref) increases.\footnote{$K^2$ is the Riemann sum of the integral $\int_{a}^{b} (K(x))^2dx$, and we have $\int_{-1}^{1} (K_1(x))^2dx = \frac{1}{2}, \int_{-1}^{1} (K_2(x))^2dx = \frac{3}{5}, \int_{-1}^{1} (K_3(x))^2dx = \frac{5}{7}, \int_{-1}^{1} (K_4(x))^2dx = \frac{350}{429}$.} In addition, because $K_4$ puts an extremely small weight on the data far away from the time point $\tau$, fewer data points are used for the estimation compared with other kernel functions. • When other parameters stay the same, relatively larger values of $\sigma_{\epsilon}^2$ or $\beta_1$ always correspond to larger values of S.D. or R.B., respectively. This can be observed from their theoretical forms of T.S.D. or T.R.B. in (ref).

All these observations verify our asymptotic results provided in Section (ref).

table[table omitted — 7,501 chars of source]

{Next, we consider the dependent noise case of (ref) for different decreasing rate of autocorrelation and dependent span, by varying $s$ and $d_n$, respectively. We fix $K(x) = K_1(x), \sigma_{\epsilon}^2 = 0.01^2, \beta_1 = 1.2, \tau = 0.5$ and document the results of R.B. and S.D. in Table (ref). We note that T.R.B. remains the same as in (ref) while T.S.D. turns to be

align[align omitted — 166 chars of source]

where $\psi^n$ is given in (ref) with

align*[align* omitted — 185 chars of source]

Under the moving average model (ref), $\rho(0)=1$, and for $k\geq 1$, we have $\rho(k) = \rho(-k)$ with

align*[align* omitted — 216 chars of source]

Based on this and for different choices of $s$ and $d_n$, we get $\psi^n$ in Table (ref). Comparing the results in Table (ref) and Table (ref), we observe that, for any fixed bandwidth parameters $h=n^{-0.26}, n^{-0.30}, n^{-0.35}$, when we consider varying $\psi^n$, a larger $\psi^n$ always results in a larger value of S.D.. This is inline with our theoretical results in (ref). }

table[table omitted — 1,421 chars of source]
table[table omitted — 494 chars of source]

{We now compare the finite sample performance of our estimator with the method based on pre-averaging and thresholding in FW2022 for jumps with different intensity. FW2022 proposed two truncated estimators of spot volatility, which are denoted as $\widehat{\sigma^2_{\tau}}(p_n, h,s_n, 1)$ and $\widehat{\sigma^2_{\tau}}(p_n, h,s_n,2)$, respectively,\footnote{Their notations $k_n, m_n,v_n, \phi_{k_n}(g)$ are $p_n, \lfloor nh \rfloor, s_n, p_n\phi_2^n$ used in our paper, respectively.}

align[align omitted — 444 chars of source]

with $s_n= \alpha (p_n\Delta_n)^{\varpi}$ for some $\alpha>0, \varpi \in (0,\frac{1}{2})$ and

align*[align* omitted — 227 chars of source]

We take $s_n = 1.8 \sqrt{BPV} (p_n \Delta_n)^{0.47}$ with $BPV = \frac{\pi}{2} \sum_{i=2}^{n} |\Delta_{i-1}^n Y||\Delta_{i}^n Y|$. To alleviate the edge effect, similar to our adjustment in Remark (ref), they replace $K_h(t_j - \tau)$ by

align*[align* omitted — 112 chars of source]

For the comparison, we set $n=23400$, as used in FW2022, and let $\sigma_{\epsilon} = 0.01$ and $\tau= 0.5$. We consider the bounded uniform kernel function $K_{uni}(x) = \frac{1}{2}1_{\{ |x|\leq 1 \} }$ for all the estimators first, and record the R.B., S.D. and the mean squared error (M.S.E.) in Table (ref). From the results, we see that our spot volatility estimator performs the best in all cases. }

table[table omitted — 1,864 chars of source]

We then verify the asymptotic normality of the spot volatility estimator established in Theorem (ref), and we consider the $i.i.d.$ noise for simplicity. The parameters are set as $\tau = 0.5$, $h = \frac{1}{n^{1/4}(\log(n))^{1/6}}$,\footnote{A more elaborate selection can be considered by scaling it with a constant in a way described in Section 3.2 of FW2022.} $\sigma_{\epsilon}^2 = 0.05^2$, and the kernel function $K_1(x)$ is used. We consider the studentized statistics in (ref) and (ref):

equation[equation omitted — 216 chars of source]

and

equation[equation omitted — 239 chars of source]

where the latter removes the bias term $b_{\tau, n}(u_n)$ owing to the presence of an infinite variation jump process $L^{(1)}$. We generate a total of 5000 paths from model (ref), and calculate the statistics, and display their histograms in Figure (ref). The figure shows that the overall shape of the density function of these statistics is close to standard normal distribution. Moreover, as we increase $\beta_1$ from 1.2 to 1.8, the statistic Sta-1 suffers a positive bias with gradually increasing size, while the statistic Sta-2 performs very stable like a standard normal distribution. Furthermore, we record the mean values and the coverage percentages under given nominal levels ($90\%, 95\%$ and $99\%$) for the statistic Sta-2 with the setting of parameters $\tau=0,0.5$, $\sigma_{\epsilon}^2=0.01^2, 0.03^2, 0.05^2$, $\beta_1=1.2, 1.5, 1.8$ and kernel functions $K_1, K_2, K_3, K_4$ in Table (ref). We see that the absolute values of the means are smaller than 0.2, and the differences between the percentages and the standard values are controlled within 5$\%$, which verifies the central limit theorem of (ref).

figure[figure omitted — 723 chars of source]
table[table omitted — 5,063 chars of source]

{Recall that in previous experiments, we take $T=1$ week and $n=5 \times 6.5 \times 3600$, which corresponds to 1-second frequency data in the stock market of 6.5 trading hour per day within one week (five consecutive trading days). Now, we demonstrate the central limit theorem (ref) for possibly low-frequency data met in practice by choosing different values for $n$. Specifically, $n=23400, 3900, 1950, 650$ correspond to 5-second, 30-second, 1-minute and 3-minute real data respectively. We use the weight function $g(x) = x \wedge (1-x)$ and $p_n = \lfloor \frac{1}{3} \sqrt{n} \rfloor$, the kernel function $K_1(x)$ and bandwidth parameter $h = \frac{1}{n^{1/4}(\log(n))^{1/6}}$. Other related parameters are fixed as $\tau = 0.5, \beta_1 = 1.8, \sigma_{\epsilon}^2 = 0.05^2$. The histograms of Sta-2 are presented in Figure (ref), where each histogram is based on 5000 repetitions. We see that, as the frequency gets lower, the histogram gradually deviates from the standard normal distribution. This is natural since less data are available. In fact, for 5-second, 30-second and 1-minute frequencies (namely $n=23400,3900, 1950$), the difference between the distribution of the estimates and standard normal distribution is quite small. For $n=650$, $p_n=9$ observational data are used for pre-averaging, consequently only $\lfloor \frac{h}{p_n\Delta_n} \rfloor = 3$ pre-averaged data are used by $S_{\tau,n}(u,h)$ in (ref) to estimate the expectation $\mathbf{E}\Big[ \cos{ \Big(\frac{u\Delta_{jp_n}^n\overline{Y}}{\sqrt{\phi_2^np_n\Delta_n}} \Big)} \Big]$, which deviates the distribution of the estimate from the standard normal distribution. One possible way to improve our spot volatility estimator is to use overlapping pre-averaged data, instead of non-overlapping case in this paper, but the theoretical derivation could be rather complicated. Such an extension may be considered in our future studies. }

figure[figure omitted — 598 chars of source]

Empirical analysis

In this section, we apply our estimator of spot volatility to some real high-frequency data. Specifically, we estimate the volatility curve within a week using 1-second trading price data and demonstrate the distributional pattern of spot volatility estimates.

We select four giant corporations, Apple (APPL), Facebook (FB), Intel (INTC), and Microsoft (MSFT), from the LOBSTER\footnote{https://lobsterdata.com} database and extract the second-by-second transactional price time series by taking the last recording price in each one-second interval, namely, the previous tick strategy. The dataset under study spans from September 12, 2016, to November 04, 2016, containing eight consecutive weeks, where each week consists of five trading days with 6.5 trading hours.

As discussed in the section of simulation studies, we first perform a pre-averaging procedure to the second-by-second price data with $g(x) = x \wedge (1-x)$ and $p_n = 114$. To calculate the estimate $\widehat{\sigma^2}_{\tau,n}(u_n,h)$ at different time points, we let the kernel function $K(x) = 1_{\{ -1 \leq x < 0 \} }$ and bandwidth $h = \frac{1}{n^{1/4}(\log(n))^{1/6}}$ with $n= 5 \times 6.5 \times 3600$, which means that 36 historical pre-averaged data closest to $\tau$ are used to construct the spot volatility estimator at $\tau$, with equal weights. The parameter $u_n$ is set according to (ref). And the setting of $d_n = 10$ is considered for demonstration.

The volatility curves within each week for different stocks are displayed in Figure (ref), and the histograms of the spot volatility estimate are presented in Figure (ref). We see that most of the daily volatility curves are like “U" shape while there is no clear pattern for weekly volatility curves. In addition, we can observe some distinctive volatility jumps from the volatility curves. The histograms show that the spot volatility estimates are distributed as a chi-square-like distribution, and most of them cluster around 0.

figure[figure omitted — 388 chars of source]
figure[figure omitted — 310 chars of source]

Conclusion and future work

In this paper, we propose a jump-robust estimator of spot volatility by using noisy high-frequency data, where the jumps considered can be of infinite variation. Under mild assumptions, we establish the consistency and asymptotic normality results. Moreover, the constructed estimator is shown to be almost rate-efficient and variance-efficient {when the jumps are either less active or active with symmetric structure}. Simulation studies are implemented to verify our theoretical results. We also apply the estimator to analyze the variation of spot volatility curve for several stocks in the financial market.

This study inspires us to extend our work in the future. First, as demonstrated in Theorem (ref), whether or not to de-bias depends on whether the largest jump activity index $\beta_1$ is larger than 1.5. It is necessary to propose a rigorous hypothesis testing procedure regarding this before applying our spot volatility estimator. Second, we see from Theorem (ref) that the de-biasing procedures differ for the cases of $K=1$ and $K>1$; thus, estimating the number of infinite variation jump processes can help us decide which one to use and the number of iterations to run for the iterative algorithm. Third, it is mentioned that the de-biasing procedures are unstable, thus finding a better procedure is necessary. Finally, from the perspective of practical application, the parameters $p_n, u_n, h,d_n$ should be tuned in a subtler data-driven way so that our spot volatility estimator can have the best finite sample performance.

Acknowledgements

Qiang Liu's work is supported by Fundamental Research Funds for the Central Universities, Shanghai University of Finance and Economics (No. 2021110482, No. 2022110007) and Shanghai Pujiang Program (22PJC046), Zhi Liu gratefully acknowledges financial support from NSFC (No. 11971507) and FDCT (0041/2021/ITP).