EconBase
← Back to paper

Optimal Iterative Threshold-Kernel Estimation of Jump Diffusion Processes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

96,541 characters · 15 sections · 55 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Optimal Iterative Threshold-Kernel Estimation of Jump Diffusion Processes

\abstract{ In this paper, we {\color{DR} propose a new threshold-kernel jump-detection} method for jump-diffusion processes, which iteratively applies thresholding and kernel methods in an approximately optimal way to achieve improved finite-sample performance. As in FLN2013, we use the expected number of jump misclassifications as the objective function to optimally select the threshold parameter of the jump detection scheme. We prove that the objective function is quasi-convex and obtain a new second-order infill approximation of the optimal threshold {\color{DR} in closed form}. The approximate optimal threshold depends not only on the spot volatility $\sigma_t$, but also the jump intensity and the value of the jump density at the origin. {Estimation} methods for these quantities are then developed, where the spot volatility is estimated by a kernel estimator with thresholding and the value of the jump density at the origin is estimated by a density kernel estimator applied to those increments deemed to {contain} jumps by the chosen thresholding criterion. Due to the interdependency between the model parameters and the approximate optimal estimators built to estimate them, a type of iterative fixed-point {\color{Red} estimation} algorithm is developed to implement them. Simulation studies {for a prototypical stochastic volatility model,} show that it is not only feasible to implement the higher-order local optimal threshold scheme but also that this is superior to those based only on the first order approximation and/or on average values of the parameters over the estimation time period.}

Introduction

In this work, we study {a} jump diffusion process of the form \[ X_{t} := \int_{0}^{t} \gamma_u du + \int_{0}^{t} \sigma_u dW_{u} + \sum_{j=1}^{N_{t}} \zeta_{j}, \] where {$W$ is a {Wiener} process, $N$ is an independent Poisson process with local intensity $\{\lambda_{t}\}_{t\geq{}0}$, and $\{\zeta_{j}\}_{j\geq{}1}$ are i.i.d. variables independent {of} $W$ and $N$. With the presence of jumps}, several statistical inference problems, including volatility estimation and jump detection, can be addressed by the {thresholding approach {developed} by Mancini (2001, 2004, 2009)}. The basic idea is to {introduce a threshold tuning parameter $B$ so that} whenever the absolute value of {an} increment $\Delta X := X_{t_i} - X_{t_{i - 1}}$ exceeds $B$, we conclude that an unusual event (aka a “jump") has happened during {\color{Red} the interval} $(t_{i-1},t_{i}]$, based on which we can {then} proceed to estimate the volatility and other parameters. Many works have been conducted to further {extend} the threshold method to various statistical inference problems. For {an} It\^{o} semimartingale with finite or infinite jump activity, jump detection and integrated volatility estimation was studied by Mancini:2009 and Jacod2007,Jac08MPVperSM. We also refer to Corsietal2010, ait2009testing,ait2009estimating,ait2010brownian, CLTIA, figueroa2012statistical, jing2012jump, and others for further applications of the threshold method.

One of the key issues that we have to address in order to have {a} good performance of the {jump detection procedure} is {the selection of the threshold $B$}. {Ideally, we hope to select the best possible threshold {under a suitable criterion}. Such a problem was studied by Figueroa-L\'{o}pez and Nisen (2013) using the expected number of jump {\color{Red} misclassifications} {as the estimation loss function and, more recently, by FLM2017 using the mean-square error of the threshold realized quadratic variation}. Under the assumption of zero drift, constant volatility {$\sigma$}, and constant jump intensity, FLN2013 showed that} the {first-order} approximation of the optimal threshold {is given by} $\sqrt{3\,{\sigma^{2}}{h} \log(1/ {h})}$ (cf. {Theorems 4.2 and 4.3 therein), when ${h}$, the time {span} between observations, shrinks to $0$ (i.e., infill or high-frequency asymptotics)}. {\color{Blue} Based on this result, FLN2013 proposed a method to estimate time-dependent deterministic volatilities and, by simulation, showed that its performance is good for smooth volatilities.} In this work, we generalize this framework {in three directions}. {We first prove that the loss function is {quasi-convex} and admits a global minimum in the more general case of non-homogeneous drift, volatility, and jump intensity}. {\color{Blue} A simpler version of this result was stated without proof in FLN2013}. We then proceed to obtain a second-order asymptotic approximation of the optimal localized threshold, {\color{DR} in closed form,} which {depends on the spot volatility $\sigma_t$, the local jump intensity $\lambda_{t}$, and the value of the jump density at the origin. We find out that, as expected, if} the spot volatility is high, then it is more preferable to have a larger threshold. {However, when} the jump intensity or the jump density at the origin is large, the possibility of having {smaller} jumps is higher, which favors a smaller threshold {to {detect} such jumps}. Although {an explicit formula for the second-order approximation is derived}, the method is not feasible unless we are able to estimate all the unknown parameters {appearing in this formula}: the spot volatility, the jump intensity, and the jump density at the origin. To this end, we {apply} kernel estimation techniques, as described below, to devise feasible plug-in type estimators for the optimal threshold.

{Kernel estimation} has a long history and has been applied to a large range of statistical problems. {In our work, we use it} to estimate the jump density at the origin. The problem we are facing differs from the usual {density kernel estimation} in several ways. Firstly, the data we have {is} contaminated by {noise}, and to make things even worse, part of the data may not contain {any information at all about} the density we {want} to estimate. Moreover, due to the usage of {a} threshold, the data we have {is at best {drawn} from a truncated distribution and, the point at which we hope to estimate the density, is not even} inside the support of the truncated data. Due to these reasons, we have to adjust the {standard method of kernel density estimation} and select the threshold appropriately so that we can get a satisfactory estimation of the jump density {at the origin}. It turns out that the optimal threshold that we should use in such a situation is larger than the one we use for optimal jump detection {(see} Section (ref) for the intuition behind this).

Another quantity we have to estimate is the spot volatility, which can also be estimated by the kernel estimator. One earlier research on this topic is foster1994continuous, where {a} rolling window estimator is analyzed, which is similar to the idea of the kernel estimation {with a uniform kernel}. The kernel-based estimation of the spot volatility, with general kernel, was studied by fan2008spot, kristensen2010nonparametric, {mancini2015estimation and,} more recently, Figueroa-L\'{o}pez and Li (2017). See also the excellent monographs of JacodProtter and JacodAitSahaliaBook for a {general treatment of the problem of spot volatility estimation of It\^o semimartigales via} uniform kernels (though Remark 8.10 in JacodAitSahaliaBook also briefly mentions the case of a general kernel with support on $[0,1]$). One of the key issues related to kernel estimators of spot volatility is how to select the bandwidth. kristensen2010nonparametric proposed a leave-one-out cross-validation method, which is a general method, but suffers from the loss of accuracy and computational inefficiency. In {this} work, we {adapt and extend the approach of Figueroa-L\'{o}pez and Li (2017) by applying a threshold-kernel estimator of the spot volatility rather than just kernel estimation.} The leading order {terms} of the MSE of the estimator {are explicitly derived}, based on which {we} {propose a procedure for optimal} bandwidth and kernel selection. The {CLT of the estimation error is also given.}

{As explained above,} the approximated optimal threshold depends on the spot volatility, jump intensity, and the {value of the jump density at the origin}, while {the approximated optimal} estimators of these three quantities depend on the threshold. Such {an interdependency} immediately suggests an iterative algorithm that starts with an initial guess of these parameters and gradually {converges} to a fixed point result. Due to the {nature} of the threshold estimator, the result is purely determined by whether the absolute value of each data increment exceeds the threshold, so we can conclude convergence without any ambiguity based on whether each data increment is included by the threshold or not.

The rest of the paper is organized as follows. {Section (ref) introduces the framework and {assumptions}. In Section (ref)}, we analyze the optimal threshold and obtain the second order approximation {thereof}. {The {bias} and variance of the estimator are derived in Section (ref)}. In Section (ref), we consider the kernel estimation of the jump density at the origin. The threshold-kernel estimation of the spot volatility is studied in Section (ref). The three estimators are {then} combined into an iterative algorithm {presented} in Section (ref). Finally, {the performance of the proposed methods are analyzed through several {simulations} in Section (ref). {Conclusions and some thoughts about} future work are provided in Section (ref). The proofs of the main results are deferred to an Appendix section.}

The Optimal Threshold {of TRV}

In this section we extend the modelling framework and optimal thresholding results {of FLN2013}. Specifically, we will allow non-constant drift, volatility, and {\color{Red} jump} intensity, though we keep the jump density {constant through time}. In the first subsection, we {introduce} all the assumptions that we need for the optimal threshold results. However, we temporarily set {the} drift, volatility, and intensity to be deterministic, {which would subsequently be relaxed} when we discuss the kernel threshold estimation of spot volatility. All the results can be generalized to {\color{Blue} stochastic drift and volatility, and doubly stochastic Poisson process $N$, as long as we assume that the Brownian motion and jumps of the semimartingale are independent from all these processes, since we can always condition on the paths of the drift, the volatility, and the jump intensity of $N$}. {\color{DR} It is also important to point out that,} {though our results in this section are derived under the just mentioned {\color{Blue} independence} assumption, our simulation experiments show that this is not essential as the proposed estimators perform well under prototypical stochastic volatility {\color{Blue} models} with leverage.}

The Framework and Assumptions

{Throughout}, {we consider an It\^{o} semimartingale of the form:}

equation[equation omitted — 178 chars of source]

where {$W=\{W_t\}_{t\geq{}0}$ is a Wiener process,} $\{\zeta_{j}\}_{j\geq{}1}$ are i.i.d. variables with density $f$, ${N=\{ N_{t}\}_{t \geq 0}}$ is a non-homogeneous Poisson process with {intensity} function $\{\lambda_t\}_{t \geq 0}$, and the continuous component $\{X^c_t\}_{t \geq 0}$ and jump component $\{J_t\}_{t \geq 0}$ are independent. {\color{Blue} The processes $\gamma$ and $\sigma$ satisfy standard conditions for the integrals in ((ref)) to be well-defined. In this section, we shall additionally assume the following conditions on $\gamma$, $\sigma$, and $\lambda$\footnote{{\color{Blue} In Section (ref), we will consider stochastic processes $\gamma$ and $\sigma$.}}}:

assumptionThe functions $\gamma:[0,\infty)\to\mathbb{R}$, $\sigma:[0,\infty)\to\mathbb{R}^{+}$, and $\lambda:[0,\infty)\to\mathbb{R}^{+}$ are deterministic such that, for any given fixed {\color{Blue} $t > 0$, \begin{equation} \begin{split} &\sigma_{t}:={\inf_{0 \leq s \leq t}} \sigma_s > 0, \quad\bar{\sigma}_{t}:={\sup_{0 \leq s \leq t}} \sigma_s <\infty,\\ &\gamma_{t}:={\inf_{0 \leq s \leq t}} \gamma_{s} > 0, \quad\bar{\gamma}_{t}:={\sup_{0 \leq s \leq t}} \gamma_{s} <\infty,\\ &\lambda_{t}:={\inf_{0 \leq s \leq t}} \lambda_{s}> 0, \quad\bar{\lambda}_{t}:={\sup_{0 \leq s \leq t}} \lambda_{s}<\infty. \end{split} \end{equation} Furthermore,} we assume that $t\mapsto\sigma_{t}$ is continuous.

The following notation will be needed:

equation[equation omitted — 267 chars of source]

Note that with these notations, our model assumptions imply that, for any $t,h\geq{}0$ and $k\in\mathbb{N}$, $$ X^{c}_{t+h}-X^{c}_{t} =_{D} N \left( h\overline{\gamma}_{t,h}, h\overline{\sigma}^{2}_{t,h} \right),\qquad \mathbb{P} \left( X_{t+h} - X_{t} \in dx | N_{t+h} - N_{t} = k \right) = \phi_{t,h}*f^{*k}(x)dx, $$ where $\phi_{t,h}$ is the density of $X^{c}_{t+h}-X^{c}_{t}$, i.e. $ \phi_{t,h}(x):= \frac{1}{\overline{\sigma}_{t,h}\sqrt{h}}\phi \left( \frac{x- h\overline{\gamma}_{t,h}}{\overline{\sigma}_{t,h}\sqrt{h}} \right) $. For these types of processes, the associated local characteristics are of the form $(\gamma,\sigma,\nu)$, where the density of the local L\'{e}vy measure is given by $\nu_{t}(x) = \lambda_{t} f(x)$.

assumption{ The jump density $f$ {\color{Blue} has} the form \begin{equation} f(x) = pf_{+}(x){\bf 1}_{[x \geq 0]} + qf_{-}(x){\bf 1}_{[x < 0]}, \end{equation} where $p \in [0,1]$ and $q:=1-p$, and $f_{+}:[0,\infty)\to[0,\infty)$ and $f_{-}:(-\infty,0]\to[0,\infty)$ are bounded functions such that $ \int_0^\infty f_+(x) dx = \int_{-\infty}^0 f_-(x) dx = 1 $. Furthermore, we assume that \[ f_{\pm}(0) = \lim_{x\to{}0^{\pm}}f_{\pm}(x) \in (0, \infty). \]}

{The following notations will also be needed:}

equation[equation omitted — 318 chars of source]

Note that $\mathcal{C}_{0}(f) = f(0)$ and $\mathcal{C}_d(f)=0$ if $f$ is continuous at the origin. For some results, we also need the following assumption:

assumption$f_{+} \in {\rm C}^1([0, b))$, $f_{-} \in {\rm C}^1((a, 0])$, for some {$a \in (-\infty, 0)$, $b \in (0, \infty)$} and $f_{\pm}'(0) := \lim_{x \to 0^{\pm}} f'_{\pm}(x)$ exists.

Throughout, we assume that we observe the process $X$ {\color{Blue} at evenly spaced times,

equation[equation omitted — 62 chars of source]

where $h_{n}$ is the time span between observations and $T:=T_{n}:=nh_{n}$ is the time horizon. We will also} use {\color{Blue} $\Delta^n_i X := X_{t_{i}} - X_{t_{i - 1}}$ to denote the increment of the underlying process over $[t_{i-1},t_{i})$}, and when no ambiguity can be brought, we will drop the superscript $n$. Finally, we introduce the {\color{Blue} jump detection procedure we consider in this work. We first specify} a vector of {thresholds} $[\vec{B}]_{T}^{n} = (B^n_1, ..., B^n_n)$, where we often drop the superscript $n$ when no confusion can be generated. {\color{Blue} Given} $[\vec{B}]_{T}^{n} $, we {\color{Blue} would} conclude that a jump {\color{Blue} had occurred} during $[t_{i - 1}, t_i)$ {\color{Blue} whenever} $|\Delta_i X| > B_i$. {\color{Blue} As a byproduct of this jump detection criterion, we can then devise the following natural} estimators {of $N_{T}$, $J_{T}$, and the integrated variance $IV_T:=\int_0^T \sigma_s^2 ds$:}

equation[equation omitted — 312 chars of source]

{These estimators were first studied in mancini2001disentangling, mancini2004estimation}. The estimator $\widehat{IV}_T$ has extensively been studied in the literature and is commonly called the truncated or thresholded realized quadratic variation (TRV) of $X$.

Optimal Threshold and Its Approximation

In this subsection, we {\color{Blue} formulate} the problem of optimal threshold selection. {We adopt the approach in FLN2013, which we now {briefly review} for completeness. We} seek to find a threshold $[\vec{B}]_{T} = (B_1, ..., B_n) \in \mathbb{R}_{+}^{n}$ to minimize the {loss function}:

equation[equation omitted — 242 chars of source]

The above loss function {represents} the expected number of {“jump"} {mis-classifications} (i.e., subintervals erroneously classified as having {jumps when in fact they do not, or not having jumps when in fact they do)}. The previous formulation gives the same weight to both types of error, while a more general loss function is given by:

equation[equation omitted — 255 chars of source]

For our purpose, (ref) is enough, but in {certain applications, (ref) may} be useful. For instance, it is more likely that {market participants} become more conservative when they {erroneously} identify a price {change} as an unusual event, i.e., a jump. In this case, {one may prefer to} take $w <1$.

In both (ref) and (ref), the loss function is additive. Therefore, we can optimize each $B_i$ separately. Indeed, we define the following loss function for given $t$ and $h$:

align[align omitted — 199 chars of source]

If we {were} able to devise a method to find $B^* = \mbox{argmin}_B L_{t,h}(B; w)$ for any $t$ and $h$, then, by setting $t = t_{i - 1}$ and $h = t_{i} - t_{i - 1}$, we {would be} able to {\color{Blue} specify} the whole optimal $[\vec{B}]_T$. Obviously, the first issue that we have to address is whether or not there is a global minimum point {\color{Blue} $B^{*}$}. As it turns out, the loss function (ref) is {quasi-convex}\footnote{{A mapping $g: D \rightarrow \mathbb{R}$, for convex $D$, is quasi-convex if for any $\lambda \in [0,1]$ and $x,y \in D$, $g(x \lambda + y (1-\lambda)) \leq \max \lbrace g(x), g(y) \rbrace$.}} in $B$, when $h$ is small enough. This property was established in FLN2013 for a driftless L\'evy processes (i.e., $\gamma\equiv{}0$ and $\sigma$ and $\lambda$ are constants). Nonzero drifts create some {nontrivial} subtleties that are resolved in the following {\color{Blue} theorem, which was stated without proof in FLN2013.}

theorem[Uniform {Quasi-Convexity} of the Loss Functions] Assume that we have model (ref), and {Assumptions} {(ref)-(ref)} are satisfied. {Then, for} any fixed $T > 0$, there exists $h_{0}:=h_{0}(T) > 0$, such that, for all $t \in [0,T]$, $h \in (0,h_{0}]$, and $w > 0$, the function $L_{t,h}(B;w)$ is quasi-convex in $B$, and possesses a unique global minimum point $B^{*}_{t,h}$.

We proceed to give a fixed-point formulation of the optimal threshold $B^{*}_{t,h}$, which in turn {enables us to find a {second-order} asymptotic expansion for $B^{*}_{t,h}$ in a high-frequency asymptotic regime ($h\to{}0$).} {This} characterization will {equip us with the} theoretical basis for developing feasible estimation algorithms later. {\color{Blue} In what follows, we focus on the case of $w=1$ and for easiness of notation, we drop the variable $w$ in $L_{t,h}(B; w)$.}

theorem[Characterizations of the Optimal Threshold] Assume that we have model (ref), and {Assumptions (ref)-(ref)} are satisfied. For each fixed $T>0$, there exists $h_{0}:=h_{0}(T)>0$ such that, for any $t\in [0,T]$ and $h\in (0,h_{0})$, the optimal threshold $B_{t,h}^{*}$, based on the increment $X_{t+h}-X_{t}$, is such that, \begin{align}\nonumber B_{t,h}^{*} &= h\overline{\gamma}_{t,h} + \sqrt{2 h\overline{\sigma}^{2}_{t,h}} \left[\ln \left(1+\exp \left( \frac{-2B_{t,h}^{*}\overline{\gamma}_{t,h}}{\overline{\sigma}^{2}_{t,h}}\right)\right)\right.\\ &\qquad\qquad\qquad\quad\qquad \left.-\ln\left(\sqrt{2\pi h\overline{\sigma}^{2}_{t,h}}\sum_{k=1}^{\infty} \frac{\left( h\overline{\lambda}_{t,h}\right)^{k}}{k!} \left[\phi_{t,h}*f^{*k}(B_{t,h}^{*})+\phi_{t,h}*f^{*k}(-B_{t,h}^{*})\right] \right)\right]^{1/2}. \end{align} Furthermore, as $h \to{} 0$, we have {the asymptotics:} \begin{equation} B_{t,h}^{*} = \sqrt{h}\, {\bar\sigma_{t,h}} \left[ 3\log \left( 1/h \right) - 2 \log \left( \sqrt{2\pi}\, \mathcal{C}_0(f) {\bar\sigma_{t,h} \bar\lambda_{t,h}} \right) \right]^{1/2} + o(h^{\frac{1}{2} + \alpha} ), \end{equation} for any $\alpha \in (0, 1/2)$. {\color{Blue} If, furthermore, $\sigma_{t}^{2},\lambda_{t}\in C^{1}((0,T))$ and continuous on $[0,T]$, then} the asymptotics in (ref) remains true if we replace $\bar\sigma_{t,h}$ and $\bar\lambda_{t,h}$ with $\sigma_t$ and $\lambda_t$, respectively.
remarkThe last assertion of Theorem (ref) remains true if $t\to\sigma_{t}^{2}$ is H\"older continuous for any exponent $\chi\in (0,1/2)$. In particular, this is the case for any volatility model driven by a Brownian motion (see RevuzYor). Recall that if a function is H\"older continuous with exponent $\chi$, then it is H\"older continuous for any exponent $\chi'<1/2$.

{The previous result extends the first-order approximation $\sqrt{ 3\sigma_{t,h}^2 h \log \left( 1/h \right)}$ of FLN2013, whose remainder is just of order $O\left( h^{1/2} \log^{-1/2} \left( 1/h \right) \right)$. However, with} the second order approximation, the remainder is $o(h^{1 - \epsilon})$ for any {$\epsilon \in (1/2, 1)$}. It is convenient to introduce the following notations for the first- and second-order optimal threshold approximations, respectively:

equation[equation omitted — 314 chars of source]

{\color{Blue} These tell us that, in a high-frequency sampling setting, the single most important parameter to determine a suitable threshold level $B$ is the spot volatility $\sigma_{t}$,} followed by the parameter $\nu_{t}(0):=\lambda_{t}\mathcal{C}_{0}(f)$, which broadly determines the likelihood of a small jump occurrence around time $t$. It is interesting to note that the optimal threshold ${B}^{*2}_{t,h}$ can differ substantially from $B^{*1}_{t, h}$ when $\sigma_{t} \lambda_{t}{\mathcal{C}_{0}(f)}$ is large. This is intuitive since, for instance, if $\sigma_{t}$ and ${C_{0}(f)}$ are fixed, as the jump rate $\lambda_{t}$ increases, the optimal threshold {\color{Blue} should decrease} in order to account for an {\color{Blue} increase in the appearance of “small" jumps. If the threshold is not adjusted, there would be more “false-negatives", i.e., missed jumps. On the other hand, as $\lambda_{t}$ decreases, the optimal threshold should be larger} in order to offset an increment in the likelihood of {false-positives} (namely, wrongly concluding the occurrence of a jump during the small interva $[t,t+h]$). {\color{Red} Similarly, for fixed $\sigma_{t}$ and $\lambda_{t}$ as the likelihood for small jumps, approximately parameterized by ${C_{0}(f)}$, increases (decreases) the optimal threshold decreases (increases) accordingly.}

Although we have proved the asymptotic properties of (ref), these optimal thresholds are not yet {feasible,} since we still need to estimate the spot volatility $\sigma_t^2$, {\color{Red} jump} intensity $\lambda_t$, and the {mass concentration of the jump density at the origin, $\mathcal{C}_0(f)$}. We will introduce estimators to these quantities in {Subsection (ref) and Section (ref)}, {respectively.}

remarkAlthough {the criterion ((ref))} provides a reasonable approach {for threshold selection,} there is no guarantee that {the resulting} optimal threshold is the one that minimizes the {mean-square error of the truncated realized quadratic variation $\widehat{IV}_{T}$ introduced in ((ref)).} We refer to FLM2017 for some results regarding {the latter problem.}

{Bias and Variance}

We conclude with the following asymptotic result of the estimation error of the TRV, which generalizes {a result of} FLN2016 to non-homogeneous drift, volatility, and jump intensities. {\color{Blue} As usual, the notation $a_{h}\sim b_{h}$, as $h\to{}0$, means that $\lim_{h\to{}0}a_{h}/b_{h}=1$.}

propositionSuppose that the assumptions of Theorem (ref) are enforced {and that $\vec{B}=(B_{n})_{n\geq{}1}$ is set to be $B_{n}^{*1} = \sqrt{3 {{\sigma}_{t_i}^2} h_n \log(1/h_n)}$. Then, as $n\to\infty$,} \begin{align*} \mathbb{E} \left[TRV(X)[\vec{B}]_{T}^{n} \right] - \int_0^T \sigma_s^2 ds \sim h_{n} \int_0^T (\gamma_s^2 - \lambda_s \sigma_s^{2}) ds, \quad {\rm Var}\left( TRV(X)[\vec{B}]_{T}^{n} \right) \sim 2 h_{n} \int_0^T\sigma_s^{4} ds. \end{align*} Furthermore, the asymptotic {\color{Blue} behavior} above {also} holds with any {threshold sequence} of the form \[ \check{B}_{n,i} = \sqrt{{c_{n,i} {\sigma}_{t_{i}}^2} {h_{n}} \log(1/{h_{n}})} + o(\sqrt{{h_{n}} \log(1/{h_{n}})}), \] provided that {$c:=\liminf_{n\to\infty}\inf_{i}c_{n,i} \in (2, \infty)$}.
proof{Let us write $B_{n,i}$ of the form $\sqrt{3 {{\sigma}^{2}_{{t_{i}}}} h_n \log(1/h_n)}$}. {The bias of the} TRV estimator can be decomposed as the following: \begin{align}\nonumber &\quad TRV(X)[\vec{B}]_{T_{n}}^{n} - \int_0^T \sigma_s^2 ds \\ & = \sum_{i = 1}^n \left( | \Delta^n_i X |^2 1_{[\Delta^n_i N = 0]} - h_n \overline{\sigma}^{2}_{t_{i - 1},h_n} \right) + \sum_{i = 1}^n | \Delta^n_i X |^2 1_{[ | \Delta^n_i X | \leq B_n, \Delta^n_i N \neq 0]} - \sum_{i = 1}^n | \Delta^n_i X |^2 1_{[ | \Delta^n_i X | > B_n, \Delta^n_i N = 0]}. \end{align} {\color{Blue} Using Lemmas C.1 and C.2 in FLN2016} as well as Assumption (ref), for any $0<\epsilon<1/2$, we have: \begin{align} \mathbb{E} \left[ | \Delta^n_i X |^2 \mathbf{1}_{[ | \Delta^n_i X | \leq B_{n,i}, \Delta^n_i N \neq 0]} \right] & = O(B_{n,i}^{3} h_{n})= O(h_{n}^{\frac{5}{2}} \left[ \log \left( 1/h_{n} \right) \right]^{\frac{3}{2}}), \\ \mathbb{E} \left[ |\Delta^n_i X |^2 \mathbf{1}_{[ | \Delta^n_i X | > B_{n,i}, \Delta^n_i N = 0]} \right] & = O(\sqrt{h_{n}}B_{n,i}\phi(B_{n,i}/\bar{\sigma}_{t_{i},h_{n}}\sqrt{h_{n}}))= O \left( h_{n}^{\frac{5}{2}-{\epsilon}} \left[ \log \left( 1/h_{n} \right) \right]^{\frac{1}{2}} \right), \end{align} where the $O(\cdot)$ terms are uniform in $i$. These would imply {that the second and third terms of ((ref)) are of orders $O_P(h^{3/2} \left[ \log \left( 1/h \right) \right]^{3/2})$ and $O_P( h^{3/2-{\epsilon}} \left[ \log \left( 1/h \right) \right]^{1/2} )$, respectively. Both of these terms are then $o(h_{n})$.} For the first term therein, note that \begin{align*} \mathbb{E} \left[ | \Delta^n_i X |^2 \mathbf{1}_{[\Delta^n_i N = 0]} - h_n \overline{\sigma}^{2}_{t_{i - 1},h_n} \right] & = \mathbb{P} (\Delta^n_i N \neq 0) h_n \overline{\sigma}^{2}_{t_{i - 1},h_n} + \mathbb{P} (\Delta^n_i N = 0) h_n^2 \overline{\gamma}^{2}_{t_{i - 1},h_n} \\ & = h_n^2 (\overline{\gamma}^{2}_{t_{i - 1},h_n} - \overline{\lambda}_{t_{i - 1},h_n} \overline{\sigma}^{2}_{t_{i - 1},h_n} ) + O(h_n^3) , \\ \mathbb{E} \left[ \left( | \Delta^n_i X |^2 \mathbf{1}_{[\Delta^n_i N = 0]} - h_n \overline{\sigma}^{2}_{t_{i - 1},h_n} \right)^2 \right] & = \mathbb{P} (\Delta^n_i N \neq 0) h_n^2 \overline{\sigma}^{4}_{t_{i - 1},h_n} + \mathbb{P} (\Delta^n_i N = 0) 2 h_n^2 \overline{\sigma}^{4}_{t_{i - 1},h_n} \&= 2 h_n^2 \overline{\sigma}^{4}_{t_{i - 1},h_n} + O(h_n^3). \end{align*} Calculating the summation of the above and noticing the independence of different terms, we conclude {the} first part of the desired result. {For {$\check{B}_{n,i}$, the term ((ref)) will instead be of order $O_P( h^{1+c/2-{\epsilon}} \left[ \log \left( 1/h \right) \right]^{1/2})$.} Therefore,} as long as {$c > 2$}, the asymptotic behavior does not change. This proves the second part of the desired result.
remarkThe motivation for considering the threshold {$\breve{B}_{n,i}$} in {Proposition (ref)} {comes from} the fact that the true value {of $\sigma^{2}$} is not available and, in practice, we have to use an estimate $\hat{\sigma}^{2}$ of it. Suppose we have an estimator of $\sigma_{t_i}^2$ denoted by $\hat{\sigma}_{t_i}^2$, and we use the corresponding estimated threshold $\hat{B}_{n}^{*1} = \sqrt{3 \hat{\sigma}_{t_i}^2 h_n \log(1/h_n)}$. The second part of Proposition (ref) tells us that if, {for instance, the estimator is such that} $\liminf_{n \to \infty} \hat{\sigma}_{t_{i}}^2 / \sigma_{t_{i}}^2 = c > 2/3$, we would have $ \hat{B}_{n}^{*1} = \sqrt{3 \hat{\sigma}_{t_i}^2 h_n \log(1/h_n)} \geq \sqrt{3 (c - \epsilon) \sigma_{t_i}^2 h_n \log(1/h_n)}, $ for $n$ large enough and $\epsilon \in (0, c - 2/3)$. This {will result} in an estimator such that the asymptotics of {the} expectation and variance of Proposition (ref) hold.

A Threshold-Kernel Estimation of the Jump Density at $0$

In this section, we investigate the estimation of the jump density at the origin, which is needed {in order to implement the second order optimal threshold $B^{*2}_{t, h}$ given by (ref)}. We propose a method based on kernel estimators. {\color{Blue} For a related method, but for a more general class of It\^o semimartingales, see Ueltzhofer. The main difference between the method proposed below and the one proposed in that paper is the thresholding technique.}

{\color{Blue} We impose} the following regularity conditions, which in particular imply that $\mathcal{C}_{0}(f)=f(0)$.

assumption$f \in C^2 \left( [a, b] \right)$ for {some} $a < 0 < b$. Also, $f(0) \neq 0$ and $f^{\prime\prime} (0) \neq 0$.
remark{It is possible to relax the previous assumption. {For instance, if the density $f$ merely satisfies} Assumption (ref), the} estimation of $f(0^+)$ and $f(0^-)$ would have to be done separately using one-sided kernel estimators. The basic idea is the same {as} what we present below, but the convergence rate and the choice of bandwidth will be different.

{As mentioned above, we wish to construct a consistent estimator for $\mathcal{C}_0(f)=f(0)$, which is not feasible during a fixed time interval $[0, T]$. Hence, in this part, we consider a high-frequency/long-run sampling setting, where {simultaneously} \[ {h}_{n}=t_{i}-t_{i-1} \to 0,\qquad T_{n}=t_{n} \to \infty, \] as $n\to\infty$. Throughout, we also assume that $\gamma$, $\sigma$, and $\lambda$ are constant so that the distribution of $\Delta_i X$ does not depend on $i$.

In the spirit of threshold {estimation, the basic idea is to treat the {“large"} increments $\Delta_i X$, whose absolute values exceed an appropriate threshold, as {proxies of} the process' jumps. These large increments can then be plugged into a {standard} kernel estimator of $ f(0)$. {Concretely, we consider the estimator:}

equation[equation omitted — 182 chars of source]

under the convention that $0/0=0$ in the case that $\{{i:} | \Delta_i X | > B \} =\emptyset$. As usual, {$K_{\delta}(x):= K(x/{\delta})/{\delta}$, where {$K:[0,\infty)\to [0,\infty)$} is a right-sided kernel function such that $\int_0^\infty K(x) dx = 1$ and ${\delta}$} is the bandwidth parameter. We also use $|A|$ to denote the number of elements in {a} set $A$. We expect that the estimator (ref) will have poor performance if $|\{i: |\Delta X_{i}| > B\}|$ is small, but, since we assume that $T \to \infty$ and $f(x)\neq{}0$ in a neighborhood of $\{0\}$, {for large-enough $n$, we have $P(\{|\Delta_{i} X| > B\} = \emptyset) \approx e^{-\lambda T} \to 0$. For} our implementation of ((ref)) {in the Monte Carlo studies of Section (ref)}, we will set $\hat{f}(0)=0$ if $|\{{i}: |\Delta_i X | > B\}|\leq{}5${, which simply makes the second order threshold to be the first order threshold}.

In what follows, $f^*$ {stands} for the density of {$|\Delta_{i} X|$}, {which depends on $n$}, while $f^*_{|\Delta X| | |\Delta X| > B}$ {stands} for the density of {$|\Delta_{i} X|$} conditioning on $|{\Delta_{i} X}| > B$. {To analyze the performance of the estimator ((ref)) and choose a suitable thresholding level $B$ {and bandwidth ${\delta}$}, we decompose the estimation error into the following two terms:}

enumerate[(i)] • $E_1 = \frac{1}{|\{ |\Delta_{i} X| > B\}|} \sum_{ | \Delta_{i} X | > B} K_{{\delta}}( |\Delta_{i} X| - B) - f^*_{ |\Delta X| | |\Delta X| > B }(B)$, • $E_2 = f^*_{ |\Delta X| | |\Delta X| > B }(B) - 2f(0)$.

{Next, we {follow} a “greedy" strategy to {determine suitable values for the} threshold $B$ and bandwidth ${\delta}$. Specifically, we minimize $E_2$ to obtain an “optimal" threshold $B$, and with that given, we minimize $E_1$ to obtain {an} “optimal" bandwidth ${\delta}$. Minimizing $E_1 + E_2$ directly will be a much more involved problem, and requires more assumptions. However, we believe solving such a problem does not significantly improve the performance of the proposed estimator. Therefore, we leave it as an open problem.

{Minimizing $E_1$ over ${\delta}$} given $B$ is {closely related to the standard} theory of kernel density estimation, so we can directly {apply the general theory for such a problem}. We only need to ensure that $|\{ |\Delta_{i} X| > B\}| \to \infty$, which follows from Proposition (ref) {\color{Blue} below} with {the} additional assumption that $T\to\infty$. {Two} widely used methods are plug-in method and cross-validation, which both have pros and cons. These methods are beyond the scope of this paper {and, for simplicity, we instead use the well-known Silverman's (1986) rule of thumb for bandwidth selection}:

equation[equation omitted — 98 chars of source]

where “$ \mbox{sd}$" is the standard deviation of $\{\Delta_i X : |\Delta_i X| > B \}$ and $L$ is the number of observations, i.e. $| \{\Delta_i X : |\Delta_i X| > B \} |$. Such a rule of thumb works the best with Gaussian kernel function and Gaussian density function. However, the method {is known to be} robust for other kernel and density functions.

We now proceed to show that $B^{*} = \sqrt{4 {h} \sigma^2 \log(1/{h})}$ minimizes the leading order terms of the second error $E_2$. The proof of the following two results are given in Appendix (ref).

propositionSuppose that Assumption (ref) is satisfied and $\gamma$, $\sigma$, and $\lambda$ are constant. {Further assume that $B \to 0$ and $B/\sqrt{{h}} \to \infty$.} Then, $E_2$ converges to $0$ as ${h}\to{}0$ if and only if ${h}^{-3/2} \exp\left( -\frac{B^2}{2 {h} \sigma^2 } \right) \to 0$. {Under this condition, we have} \begin{equation} E_2 = \frac{ 2 }{\lambda \sqrt{2 \pi {h}^3 \sigma^2}} \exp\left( -\frac{B^2}{2 {h} \sigma^2 }\right) + 2f(0) B + o(B)+o(h^{-3/2}e^{-\frac{B^{2}}{2h\sigma^{2}}}). \end{equation} Furthermore, if $E_2$ converges to $0$, then $\mathbb{P} ( |{\Delta_{i} X}| > B ) = \lambda {h} + o({h})$, {as ${h}\to{}0$.}

In addition to {providing us conditions} for the error $E_{2}$ to vanish, Proposition (ref) implies that, {in that case,} $\mathbb{E} \left[ |\{ i : |\Delta_i X| > B\}| \right] = \lambda T + o(T)$, {as $T\to\infty$ and ${h}\to{}0$}. Therefore, the average sample size that can be used for the estimation of $f(0)$ is approximately constant with respect to $B$. {Heuristically, this suggests that the selection of $B$ will not affect significantly the selection of {\color{Blue} $\delta$} that minimizes $E_{1}$.} We {are now ready to} obtain an approximate optimal threshold $B$, which minimizes the leading order terms of $E_2$.

corollaryThe approximate optimal threshold {\color{Blue} $\widetilde{B}^{*}$} that minimizes the leading order term of $E_2$ given by (ref) is {such that} \begin{equation} {\color{Blue} \widetilde{B}^*} = \sqrt{4 {h} \sigma^2 \log(1/{h})} + O(\sqrt{{h} \log\log(1/{h})}), \end{equation}

It is interesting to notice that the “optimal" threshold here is not the same as the one identified in the previous section. Indeed, if we do use the optimal threshold $B^{*1}$ or $B^{*2}$ in (ref), $E_2$ would {diverge}. It is interesting and important to get some sense why the optimal thresholds differ from each other. Indeed, in the previous section, we optimize the expected number of jump misclassification. In that case, we are minimizing the sum of unconditional false positive (mistakenly claim a jump) and unconditional false negative (miss a jump). However, since the probability that a jump occurs is so small, proportional to the length of the time increments, the probability of having a false negative, {by nature}, cannot be too large. Therefore, by having the expected number of misclassification as the objective function, we would choose a threshold in favour of having a much smaller unconditional false positive rate. {As it turns out, if we choose $B^{*1}$ or $B^{*2}$, conditioning on $|\Delta X| > B$, the probability} that no jump occur is comparable to the probability that a jump occurs, both $O({h})$. That is, the conditional false negative rate does not vanish. Such a situation would minimize the expected number of {\color{Red} misclassifications}, but would not enable us to distinguish the distribution of the jump from the noise. Using $\sqrt{4 {h} \sigma^2 \log(1 / {h})}$, on the other hand, makes the conditional false negative {vanishing and, thus, enables us to get consistent} estimation of jump density.}

Threshold-Kernel Estimation of Spot Volatility

{In this section, we consider the estimation of the spot volatility of a jump-diffusion process, which is needed to implement the approximate optimal threshold formula ((ref)), but is also an important problem on its own. Unlike Section (ref), here we {\color{Blue} also work with certain} stochastic volatility models. The precise conditions are given below.}

The {idea} of kernel estimation of spot volatility is to take a weighted average of the squared increments (see, e.g., foster1994continuous and fan2008spot):

equation[equation omitted — 161 chars of source]

{Here,} $K(\cdot)$ is {a} kernel function with $\int K(x) dx = 1$, $K_{\delta}(x)=K(x/\delta)/\delta$, and ${\delta} > 0$ is the bandwidth. However, when jumps do occur, the estimator above becomes inaccurate. A natural idea is to combine (ref) with the threshold method. Concretely, {given a} threshold vector $[\vec{B}]^{n}_{T}=(B_{1}^{n},\dots,B_{n}^{n})$, we consider the local threshold-kernel estimator:

equation[equation omitted — 203 chars of source]

{In what follows} {we will} investigate the properties of (ref). In order to do this, we will have to deal with the randomness of the volatility, {for which we extend some of the results} in FLL2016. We will mention the assumptions on $\{\sigma_t\}_{t \geq 0}$ and $K$ in Subsection (ref), and then discuss the asymptotic properties of (ref) in subsequent subsections.

Assumptions on the Volatility Process

The first {\color{Blue} assumptions are some non-leverage and boundedness conditions, which enable us to condition on the whole path of the volatility and drift and use estimates from FLN2016}:

assumptionIn (ref), {$(\gamma, \sigma)$ are locally bounded c\'adl\'ag} independent of the Brownian motion $W$ and {\color{Blue} the jump component $J$. Furthermore, there exists a deterministic $M_T <\infty$ for which $\bar{\gamma}_{T}$ and $\bar{\sigma}_{T}^{2}$ defined in ((ref)) satisfy $\bar{\gamma}_{T}<M_{T}$ and $\bar{\sigma}_{T}^{2}<M_{T}$. The intensity $\lambda$ is still assumed to be deterministic such that $\underline{\lambda}_{t}:=\inf_{0 \leq s \leq t} \lambda_{s}> 0$ and $\bar{\lambda}_{t}:=\sup_{0 \leq s \leq t} \lambda_{s}<\infty$.}

{We now introduce the} key assumption on the volatility process.

assumptionSuppose that for $\varpi > 0$ and certain functions $L: \mathbb{R}_+ \rightarrow \mathbb{R}_+$, $C_{\varpi}: \mathbb{R} \times \mathbb{R} \rightarrow \mathbb{R}$, such that $C_{\varpi}$ is not identically zero and \begin{equation} \begin{split} C_{\varpi}(hr,hs) & = h^{\varpi}C_{\varpi}(r,s), \quad for r, s \in \mathbb{R}, h \in \mathbb{R}_+, \end{split} \end{equation} the variance process $V := \{V_t = \sigma_t^2 : t\geq 0\}$ satisfies \begin{equation} \mathbb{E} [ (V_{t+r} - V_{t})(V_{t+s} - V_{t}) ] = L(t)C_{\varpi}(r,s) + o ((r^2 + s^2)^{\varpi/2}) , \quad r,s \rightarrow 0. \end{equation}

An additional assumption on the kernel function $K$ is the following:

assumptionGiven $\varpi > 0$ and $C_{\varpi}$ as defined in Assumption (ref), {the kernel} function $K:\mathbb{R} \rightarrow \mathbb{R}$ satisfies the following conditions: \begin{enumerate}[(1)] • $\int K(x) dx = 1${;} • $K$ is Lipschitz and piecewise $C^1$ on its support $(A,B)$, where $-\infty \leq A < 0 < B \leq \infty${;} • {(i)} $\int |K(x)||x|^{\varpi} dx < \infty$; {(ii) $K(x)x^{\varpi + 1} \rightarrow 0$, as} $|x| \rightarrow \infty$; {(iii)} $\int |K^{\prime}(x)| dx < \infty$, {(iv)} $V_{-\infty}^{\infty} (|K^{\prime}|) < \infty$, where $V_{-\infty}^{\infty}(\cdot)$ is the total variation{;} • $\iint K(x)K(y)C_{\varpi}(x,y)dxdy > 0$. \end{enumerate}

We refer to FLL2016 for more details on {\color{Blue} Assumptions} (ref), (ref) and (ref). {We just mention here that Assumption (ref) covers a wide range of} frameworks such as deterministic and smooth volatility, Brownian motion and fractional Brownian motion driven volatility, etc. In the following subsection, we will establish asymptotic properties of (ref) based on Assumption (ref), (ref) and (ref).

Asymptotic Properties of Threshold-Kernel Estimator

FLL2016 proves the following result under Assumption (ref), (ref), and (ref) (c.f. Section 3 therein):

equation[equation omitted — 412 chars of source]

where $X^c$ is the continuous part of $X$ {defined in ((ref))}. The key result to extend the {theory of kernel estimators, as developed in FLL2016,} to the threshold-kernel estimators ((ref)) is the following.

propositionSuppose that {\color{Blue} Assumptions (ref),} (ref), (ref), and (ref) are satisfied, {and take a bandwidth sequence ${\delta}_{n}$ such that ${h}_{n}/{\delta}_{n}\to{}0$. Let ${B_{i}}:=B_{n,i}(c) := \sqrt{c {\bar{\sigma}_{t_{i},h}^2} {{h}} \log(1/{{h}})} + o({\sqrt{{h} \log(1/{h})}}) $, with $c > 0$.} Then, we have: \begin{equation} \bar{\mathcal{E}}_{n}:=\sum _{i=1}^n K_{{\delta}}(t_{i-1} - \tau) \left[(\Delta_i X^c)^2- (\Delta_i X)^2 1_{\{|\Delta_i X| \leq B_i \}}\right] = O_P\left( \max\{ {h}, {h}^{c/2} \log^{{1/2}}(1 / {h}) \} \right). \end{equation} {Furthermore,} \begin{align} {\mathbb{E}\left(\bar{\mathcal{E}}_{n}^{2}\right)= {O}\left(\frac{{h}^2}{{\delta}}\right) + {O}\left( \frac{{h}^{1 + \frac{c}{2}}}{{\delta}} [\log(1/{{h}})]^{\frac{3}{2}} \right) + {O}\left( {h}^{c} \log(1/{{h}}) \right).} \end{align}
proof{Let $\mathcal{E}_{i}:=(\Delta_i X)^2 \textbf{1}_{\{|\Delta_i X| \leq B_i \}} - (\Delta_i X^c)^2$ and} observe that \begin{equation} \begin{split} {\mathcal{E}_{i}} & = - ( \Delta_i X^c )^2 \mathbf{1}_{[ \Delta_i N \neq 0]} + ( \Delta_i X )^2 \mathbf{1}_{[ | \Delta_i X | \leq {B_{i}}, \Delta_i N \neq 0]} - ( \Delta_i X )^2 \mathbf{1}_{[ | \Delta_i X | > {B_{i}}, \Delta_i N = 0]} =:\mathcal{E}_{i,1}+\mathcal{E}_{i,2}+\mathcal{E}_{i,3}. \end{split} \end{equation} Now, {\color{Blue} conditioning on the paths of $\sigma$ and $\gamma$ and applying Lemmas C.1-C.2 in FLN2016}, the following holds: \begin{equation} \begin{split} & {( \Delta_i X )^2 \mathbf{1}_{[ | \Delta_i X | \leq B_{i}, \Delta_i N \neq 0]} = {O_P\left( B_{i}^{3}{h}\right)}=O_P\left( {h}^{5/2} [\log(1/{{h}})]^{3/2} \right)}, \\ & {( \Delta_i X )^2 \mathbf{1}_{[ | \Delta_i X | > B_{i}, \Delta_i N = 0]} = {\color{Blue} O_{P}(\sqrt{h_{n}}B_{n,i}\phi(B_{n,i}/\bar{\sigma}_{t_{i},h_{n}}\sqrt{h_{n}}))}= O_P\left( {h}^{1 + c/2} [\log(1/{{h}})]^{1/2} \right)},\\ &{ ( \Delta_i X^c )^2 \mathbf{1}_{[ \Delta_i N \neq 0]} = O_{P}({h}^2)}. \end{split} \end{equation} From Assumption (ref), the above holds uniformly over $1 \leq i \leq n$. Therefore, by Assumption (ref), we have: $$ \sum _{i=1}^n K_{{\delta}}(t_{i-1} - \tau) \left[ (\Delta_i X)^2 \textbf{1}_{\{|\Delta_i X| \leq B_i \}} - (\Delta_i X^c)^2 \right] = O_P\left( \max\{ {h}, {h}^{c/2} \log^{1/2}(1 / {h}) \} \right). $$ {For the second assertion of the theorem, first note that \begin{align*} \mathbb{E}\left(\bar{\mathcal{E}}_{n}\right)&=\sum _{i=1}^n K_{{\delta}}(t_{i-1} - \tau) \mathbb{E}\left[\mathcal{E}_{i,1}+\mathcal{E}_{i,2}+\mathcal{E}_{i,3}\right]\\ &=\sum _{i=1}^n K_{{\delta}}(t_{i-1} - \tau) \left[{O}({h}^2)+ {O}\left( {h}^{\frac{5}{2}} [\log(1/{{h}})]^{\frac{3}{2}} \right) + {O}\left( {h}^{1 + \frac{c}{2}} [\log(1/{{h}})]^{\frac{1}{2}} \right)\right]\\ &= {O}\left({h}\right) + {O} \left( {h}^{\frac{c}{2}} [\log(1/{{h}})]^{\frac{1}{2}} \right). \end{align*} Similarly, \begin{align*} {\rm Var}\left(\bar{\mathcal{E}}_{n}\right)&=\sum _{i=1}^n K^{2}_{{\delta}}(t_{i-1} - \tau) {\rm Var}\left((\Delta_i X^c)^2 - (\Delta_i X)^2 1_{\{|\Delta_i X| \leq B_i \}}\right)\\ &\leq4\sum _{i=1}^n K^{2}_{{\delta}}(t_{i-1} - \tau) \left[\mathbb{E}(\mathcal{E}^{2}_{i,1})+\mathbb{E}(\mathcal{E}^{2}_{i,2})+\mathbb{E}(\mathcal{E}^{2}_{i,3})\right]\\ &=\sum _{i=1}^n K^{2}_{{\delta}}(t_{i-1} - \tau) \left[{O}({h}^3) +{O}\left( {h}^{2 + \frac{c}{2}} [\log(1/{{h}})]^{\frac{3}{2}} \right)\right]\\ &= {O}\left(\frac{{h}^2}{{\delta}}\right) +{O}\left( \frac{{h}^{1 + \frac{c}{2}}}{{\delta}} [\log(1/{{h}})]^{\frac{3}{2}} \right). \end{align*} We then conclude the result.}

With {Proposition (ref)}, {\color{Blue} we get} the following proposition, which characterizes the leading order terms of the MSE of the threshold-kernel estimator ((ref)). This allows us to {perform} bandwidth and kernel function selection.

propositionAssume that {Assumptions} {\color{Blue} (ref),} (ref), (ref), and (ref) are satisfied, and take the threshold vector to be $B_{{n,i}}(c) = \sqrt{c \bar{\sigma}_{t_i,h}^2 {h} \log(1/{h})} + o(\sqrt{{h} \log(1 / {h})})$ for any $c\in (\frac{\varpi}{\varpi + 1}, \infty)$. Then, we have that, for each $\tau \in (0,T)$, \begin{equation} \mathbb{E} \left[ \left( TKW(\tau, n, {\delta}) - \sigma_\tau^2 \right)^2 \right] = 2\frac{{h}}{{\delta}} \mathbb{E}[\sigma_{\tau}^4] \int K^2(x)dx + {\delta}^{\varpi} L(\tau) \iint K(x)K(y) C_{\varpi} (x,y) dxdy + o\left(\frac{{h}}{{\delta}}\right) + o \left({\delta}^{\varpi}\right). \end{equation}
proofWe consider the following decomposition: \begin{equation} \begin{split} TKW(\tau, n, {\delta}) - \sigma_\tau^2 & = \sum _{i=1}^n K_{{\delta}}(t_{i-1} - \tau) \left[ (\Delta_i X)^2 1_{\{|\Delta_i X| \leq B_i \}} - (\Delta_i X^c)^2 \right] + \left[ \sum _{i=1}^n K_{{\delta}}(t_{i-1} - \tau) (\Delta_i X^c)^2 - \sigma_\tau^2 \right] \\ & =: (I) + (II). \end{split} \end{equation} From (ref), we have that the second moment of (II) above converges with rate $O\left(\frac{{h}}{{\delta}}\right) + O\left({\delta}^{\varpi}\right)$. The optimal rate {of (II) is given by ${h}^{\varpi / (1 + \varpi)}$ and is attained with ${\delta} \sim {h}^{1/(\varpi + 1)}$}. Therefore, {by Proposition (ref), as long as $c > \varpi / (1 + \varpi)$}, (I) is of higher order than (II), in which case, (I) will be either of $o\left(\frac{{h}}{{\delta}}\right)$ or {\color{Blue} $o\left({\delta}^{\varpi}\right)$}. This completes the proof.
remarkThe leading order term {of the MSE of ((ref))} does not depend on the {threshold}. However, by selecting the optimal threshold or its approximations, we are able to optimize the {sub-order} part of the error, which enhances the performance of the estimator in practice. Also, {since taking $c\in(2,\infty)$ does not change the asymptotic rate of convergence}, we have {\color{Red} a} certain degree of robustness of this method.

With some further assumptions, we are {also} able to obtain the {CLT} of the threshold-kernel estimator. The proof of the following {result} is similar to that Proposition (ref), {\color{Blue} but taking advantage of Theorems 6.1 and 6.2 in FLL2016 which deal with the analogous results without jumps.}

theoremAssume that {A}ssumption (ref), (ref), (ref), (ref) and (ref) are satisfied, and take the threshold vector to be ${\color{Blue} B_{n,i}(c)} = \sqrt{c \bar{\sigma}_{t_i,h}^2 {h} \log(1/{h})} + o(\sqrt{{h} \log(1 / {h})})$ for any $c\in (\frac{\varpi}{\varpi + 1}, \infty) $. Then, for each $\tau \in (0,T)$, \begin{equation} \left( \frac{{h}}{{\delta}} \right)^{-1/2} \left[ TKW(\tau, n, {\delta}) - \int_0^T K_{{\delta}}(t - \tau) \sigma_t^2 dt \right] \rightarrow _D \delta_1 N(0,1), \end{equation} {where $\delta^{2}_{1}=2\sigma_{\tau}^{4}\int K^{2}(x)dx$.} Furthermore, suppose that either one of the following conditions holds: \begin{enumerate}[(1)] • $\{\sigma^2_t\}_{t\geq{}0}$ is an It\^{o} process given by $\sigma_t^2 = \sigma_{0}^{2}+\int_0^t {f_{s}} ds + \int_0^t {g_{s}} d{\color{Blue} B_s}$, {where {\color{Blue} $B$ is a Brownian motion independent of $W$ and} we {further assume that} $\sup_{t \in [0, T]}\mathbb{E}[|f_t|] < \infty$, $\sup_{t \in [0, T] } \mathbb{E} [ g_{t}^{2} ] < \infty$, and $\mathbb{E} [ ( g_{\tau + h} - g_\tau )^2 ] \to 0$ as $h \to 0$;} • $\sigma^2_t = f(t, Z_t)$, for a deterministic function $f:\mathbb{R} \times \mathbb{R} \to \mathbb{R}$ such that $f \in C^{1,2}(\mathbb{R})$, and a Gaussian process $\{Z_t\}_{t\geq{}0}$ satisfying Assumption (ref) and some mild additional conditions\footnote{We refer the reader to FLL2016 for more details. In FLL2016, we assume $\sigma_t^2 = f(Z_t)$, but it is actually trivial to generalize to the case that $\sigma_t^2 = f(t, Z_t)$ for $f \in C^{1,2}(\mathbb{R})$.}. \end{enumerate} Then, on an extension $(\bar{\Omega}, \bar{\mathscr{F}}, \bar{\mathbb{P}})$ of the probability space $(\Omega, \mathscr{F},\mathbb{P})$, equipped with a standard normal variable $\xi$ independent of {\color{Blue} $\{\sigma_{t}\}_{t\geq{}0}$}, we have, for each $\tau \in (0,T)$, \begin{equation} {\delta}^{-\varpi/2} \left( \int_0^T K_{{\delta}}(t - \tau) (\sigma_t^2 - \sigma_\tau^2) dt \right) \rightarrow_{D}\, \delta_2 \xi, \end{equation} where, under the condition (1) above, $\delta_2^2 = g(\tau, \omega)^2 \iint K(x) K(y) C(x, y) dxdy$, while, under the condition (2), $ \delta_2^2 = [f_{2}(\tau, Z_\tau)]^2 L^{(Z)}(\tau) \iint K(x)K(y) C^{(Z)}_\varpi (x,y) dxdy $. Here, $f_{2}(t, z) = \frac{\partial f}{\partial z}(t, z)$.

It is interesting to realize the {difference} between the range of $c$ allowed here and the one allowed for the integrated volatility. Indeed, for $\varpi \in (0, \infty)$, the range for spot volatility estimation is strictly larger than the range for the integrated volatility estimation. The reason is that the estimation of spot volatility is much less accurate than the integrated volatility. Therefore, we may conclude that even with a bad estimation of spot volatility, we are still able to get a threshold that is accurate enough for us to apply the threshold estimation and obtain another estimation of the spot volatility.

Bandwidth and Kernel Selection

With the leading order approximation we obtained from the previous subsection, we are now able to develop a feasible plug-in type bandwidth selection method. Furthermore, we can derive the optimal kernel function when the volatility is driven by Brownian motion. In this subsection, we describe all related results, which are direct {consequences} of Proposition (ref), and are parallel to results given by FLL2016. {We refer to FLL2016 for the details of the proofs.}

The first result is the theoretical approximated optimal bandwidth, which can be obtained by taking {the derivatives} of the leading order terms in (ref) with respect to the bandwidth ${\delta}$.

propositionWith the same assumptions as Proposition (ref), the approximated optimal bandwidth, denoted by ${\delta}^{a, opt}_n$, which is defined to minimize the leading order term of MSE in (ref), is given by \begin{equation} \begin{split} {\delta}^{a, opt}_n = & n^{-1/(\varpi + 1)} \left[ \frac{2 T \mathbb{E}[\sigma_{\tau}^4] \int K^2(x)dx} {\varpi L(\tau) \iint K(x)K(y) C_{\varpi} (x,y) dxdy} \right] ^{1/(\varpi + 1)}, \end{split} \end{equation} while the {attained global minimum of the} approximated MSE is given by \begin{equation} \begin{split} MSE^{a, opt}_n & = n^{-\varpi/(1+\varpi)} \frac{1 + \varpi}{\varpi} \left( 2 T \mathbb{E}[\sigma_{\tau}^4] \int K^2(x)dx \right) ^{\varpi/(1+\varpi)} \left( \varpi L(\tau) \iint K(x)K(y) C_{\varpi} (x,y) dxdy \right) ^{1/(1+\varpi)}. \end{split} \end{equation}

As shown in FLL2016, {the resulting bandwidth obtained by replacing $ \mathbb{E}[\sigma_{\tau}^4]$ and $L(\tau)$ in the formula ((ref)) with their integrated versions, $\int_0^T \mathbb{E}[\sigma_{\tau}^4] d\tau$ and $\int_0^T L(\tau) d\tau$,} is asymptotically equivalent to the optimal bandwidth that minimizes the integrated MSE, $\int_{0}^{T}\mathbb{E}\left[(\hat{\sigma}_{t}^{2}-\sigma_{t}^{2})^{2}\right]dt$. {\color{Blue} In the case of a volatility process driven by Brownian motion, as in the setup (1) of Theorem (ref), FLL2016 showed that $\varpi=1$, $C_{1}(x,y)=\min\{|x|,|y|\}{\bf 1}_{xy\geq{}0}$, and $L(t)=\mathbb{E}(g^{2}_{t})$, which leads to the formula:

equation[equation omitted — 265 chars of source]

Furthermore, since, at best, we only have one realization of the path of $\sigma$ and we are working with a nonparametric setting for $\sigma$, it is natural to use $\int_0^T \sigma_{t}^4dt$ and $\int_{0}^{T}g_{t}^{2}dt$ as proxies of $\mathbb{E}[\int_{0}^{T}\sigma_{\tau}^4d\tau]$ and $ \mathbb{E}[\int_{0}^{T}g_{t}^{2}dt]$, respectively. These considerations suggest the following bandwidth selection method:

equation[equation omitted — 240 chars of source]

Alternatively, by virtue of the independence condition in Assumption (ref), we can see ((ref)) as an approximation of the optimal bandwidth that minimizes the conditional integrated MSE, $\mathbb{E}\left[\int_{0}^{T}(\hat{\sigma}_{t}^{2}-\sigma_{t}^{2})^{2}dt|\sigma_{s},\gamma_{s}:0\leq{}s\leq{}T\right]$.

However, the bandwidth ((ref)) is not yet {feasible}, since it depends on {the} unknown random quantities $\int_{0}^{T}\sigma_{t}^4dt$ and $\int_{0}^{T}g^{2}_{t}dt$. A well-known estimator of $\int_0^T \sigma_{t}^4 dt$ is the truncated realized quarticity, which is defined by $\widehat{IQ} =(3{h})^{-1} \sum_{i = 1}^n (\Delta_i X)^4{\bf 1}_{\{|\Delta_{i}X|<B_{n,i}\}}$ (see Proposition 1 in Mancini:2009 for consistency). The estimation of $\int_0^T g_{t}^{2}dt$ is more involved. This quantity is sometimes called the integrated {\color{Red} vol of vol (or vol vol for short)} and is essentially the quadratic variation of the volatility process. FLL2016 introduced an estimator based on the Two-time Scale Realized Quadratic Variation introduced in Two_Time_Scale. Concretely, let} $\hat{\sigma}_{l, t_i}^2$ and $\hat{\sigma}_{r, t_i}^2$ be the left and right side estimator of $\sigma_{t_i}^2$, respectively, defined as the following:

equation[equation omitted — 518 chars of source]

{Next,} we define the following two {\color{Blue} finite differences:} $\Delta_i \hat{\sigma}^2 = \hat{\sigma}_{r, t_{i + 1}}^2 - \hat{\sigma}_{l, t_i}^2$, $\Delta_i^{(k)} \hat{\sigma}^2 = \hat{\sigma}_{r, t_{i + k}}^2 - \hat{\sigma}_{l, t_i}^2$. {Finally,} we can construct the following estimator:

equation[equation omitted — 225 chars of source]

Here, $b$ is a small enough integer, when compared to $n$. The purpose of introducing such a number $b$ is to alleviate the boundary effect of the {one sided} estimators, since, for instance, it is expected that $\hat{\sigma}_{l, t_i}^2$ will be more inaccurate as $i$ {gets} smaller. The consistency of the TSRVV estimator can be proved by {\color{Blue} Proposition} (ref) and the corresponding results from FLL2016.

The final result that we will mention in this subsection is about the optimal kernel function. Indeed, as was proved in FLL2016, when the volatility is driven by Brownian motion, the optimal kernel function is given by the double exponential function.

theoremWith the same assumptions as Proposition (ref) and assuming $C_{\varpi}(r, s) = \min\{|r|, |s|\} \textbf{1}_{\{rs > 0\}}$, we have that the optimal kernel function that minimizes the approximated optimal MSE given by (ref) is the double exponential kernel function: $$K^{opt}(x) = \frac{1}{2} e^{-|x|}, \quad x \in \mathbb{R} .$$

Full Implementation Scheme of The Threshold-Kernel Estimation

In this section, we {propose} a complete data-driven threshold-kernel estimation scheme. {We consider several versions, depending} on whether we treat the volatility to be constant or not and whether we use {the first- or second-}order approximation formula. {One of our main interests is to investigate whether or not local and/or second-order thresholding {\color{Blue} can improve} the performance of threshold estimation.}

{Let us recall that the key problem at hand is jump detection; i.e., we hope to determine whether $\Delta_i N = 0$ or not. We are, of course, also interested in estimating the volatility, jump intensity, and jump density, but we are operating under the premise that effective jump detection leads to good estimation of the other model features.} In {Section (ref)}, we {introduced} the expected number of jump misclassification as the objective function and obtained the {theoretical first and second order infill approximations} of the optimal threshold, respectively given by

equation[equation omitted — 285 chars of source]

where, {with certain abuse of notation, we denote $\sigma_i^2 := \sigma^{2}_{t_{i}}$ and $\lambda_i := \lambda_{t_{i}}$.} Although we have assumed that $\mathcal{C}_0(f)$ remains constant as the time evolves, we do allow non-constant volatility $\sigma_t$ and jump intensity $\lambda_t$.

{\color{Blue} Since estimating spot values is typically less accurate than estimating average values, {a simple first approach to implement ((ref)) is to substitute $\sigma^{2}_{i}$ and $\lambda_{i}$ by their average values, $\bar{\sigma}^{2}:=\int_{0}^{T}\sigma_{s}^{2}ds/T$ and $\bar{\lambda}:=\int_{0}^{T}\lambda_{s}ds/T$, respectively. This simplification leads us to consider the following threshold sequences:}

equation[equation omitted — 301 chars of source]

where the superscript $c$ above is used to denote “constant" volatility and jump intensity. In light of (ref), natural estimates of $\bar{\lambda}$ and $\bar{\sigma}^{2}$ are given by

equation[equation omitted — 226 chars of source]

respectively. The estimator of $\mathcal{C}_0(f)=f(0)$, as developed in Section (ref), is given by

equation[equation omitted — 204 chars of source]

where the bandwidth $\delta$ is set according to Silverman's rule of thumb ((ref)) and, for the threshold ${B}_i$, we could use the same threshold as in ((ref)) or an estimate of $\widetilde{B}_{i}=\sqrt{4 {h} \sigma^2 \log(1/{h})}$ as suggested in Corollary (ref). In the algorithms below and in the simulations of Section (ref), we use the former threshold. Putting all together, the Algorithms (ref) and (ref) below detail the implementation of the 1st and 2nd order constant thresholds ((ref)). Algorithm 1 is the same as that proposed in FLN2013 and, because it generates a nonincreasing sequence of thresholds and volatility estimates, is guaranteed to finish in finitely many steps. See the end of this section for more information about the stopping criteria for Algorithm (ref).

algorithm[algorithm omitted — 895 chars of source]
algorithm[algorithm omitted — 941 chars of source]

We now consider the implementation of the local or non-constant thresholds ((ref)). First of all, since Theorem (ref) establishes that $\sigma_i^2$ has a much greater effect on the approximated optimal threshold than that of $\lambda_i$, we simplify the problem by estimating $\lambda_{i}$ with $\hat{\lambda}$ as defined in ((ref)). The estimation of $\sigma_i^2$, per our discussion in Section (ref), is given by the kernel estimator:

equation[equation omitted — 177 chars of source]

Above, we could try to calibrate the bandwidth ${\delta}$ using an approach similar to that described in Section (ref). However, for simplicity, in the simulations we set $\delta=h_{n}^{1/2}$, which, per ((ref)), is rate optimal at first order. Based on the $\hat{\sigma}_{i}^{2}$'s, $\hat{\lambda}$, and $\widehat{\mathcal{C}_0(f)}$ as defined in ((ref)), we can then compute estimates of the first and second order approximation of the optimal thresholds as follows:

equation[equation omitted — 320 chars of source]

where above the superscript $n$ stands for non-constant volatility estimation. Algorithm (ref) below gives the details of the implementation of the non-constant thresholds ((ref)). Therein, the initial threshold is taken as $B^{n1}_i=\left[ 3\hat{\bar{\sigma}}^{2}_{0} {h} \log \left( 1/ {h} \right) \right]^{1/2}$, where $\hat{\bar{\sigma}}^{2}_{0}$ is an initial estimate of $\bar{\sigma}^{2}:=\int_{0}^{T}\sigma_{s}^{2}ds/T$ such as those obtained from the previous Algorithms. In the simulations of Section (ref), we take that from Algorithm (ref). See also below for more details about the “stopping criteria" of the algorithm.

algorithm[algorithm omitted — 1,072 chars of source]

Note} that, in (ref) and (ref), $B^{c2}$, and $B^{n2}$ may not be well defined, under a finite sample setting. {Indeed, for} a fixed time period and a fixed sample size, it is possible to have $3\log \left( 1/{h} \right) < 2 \log \left( \sqrt{2\pi} \mathcal{C}_0(f) \hat\sigma_{i} \hat\lambda \right)$, in which case the square root {in ((ref))} is not well defined. Of course, asymptotically this is never an issue since we only need to consider a small enough ${h}$. As to implementation, {however, it is natural to use} $B^{*2}$ whenever $3\log \left( 1/{h} \right) > 2 \log \left( \sqrt{2\pi} \mathcal{C}_0(f) \sigma_{i} \lambda_{i} \right)$, and use $B^{*1}$, otherwise.

{We now briefly discuss {some} stopping criteria for the Algorithms (ref) and (ref). {Typically, most iterative algorithms are stopped when} the updated value is “close" enough to the old value. However, for the threshold estimator, we note that there are only $2^{n}$ possible threshold vectors after the initial set up. Therefore, there are only two possible situations for the “while" loop in Algorithm (ref):

enumerate• After a few iterations, the algorithm comes to a fixed threshold vector $[\vec{B}_T]$. • After a few iterations, the algorithm comes to a loop of threshold vectors given by $[\vec{B}_T^1]$, ..., $[\vec{B}_T^k]$.

As we will see at the end of Subsection (ref), generally the threshold vector converges within 2 iterations.}

Monte Carlo Study

In this section, {\color{Blue} we investigate} the performance of our proposed methods. Specifically, in Section (ref), we will compare the four different threshold methods given by (ref) and (ref) {\color{Blue} and detailed in Algorithms (ref)-(ref)}. In Section (ref), we investigate the performance of the threshold-kernel estimation of the jump density at the origin.

Throughout, we consider the jump-diffusion model given by (ref), {with the continuous part $\{X_t^c\}_{t \geq 0}$ following a} Heston model:

equation[equation omitted — 155 chars of source]

Here, $V_t = \sigma_t^2$ is the variance process. The parameters of (ref) are selected according to the following setting {also used in Two_Time_Scale}:

equation[equation omitted — 110 chars of source]

As to the initial values, we use $X^c_0 = 1$ and {$V_{0}= \sigma_0^2 = 0.04$}. {\color{DR} The unit of time in this study is 1 year and, thus, the parameter values above are annualized.} Although {\color{Blue} the properties} of the threshold-kernel estimators studied in this work were derived under a non-leverage setting ({\color{Blue} i.e.,} {$\rho = 0$}, where $\rho$ is the correlation between $B_t$ and $W_t$), we run simulations on both the non-leverage setting and a negative leverage setting ($\rho = -0.5$) in order to check the robustness of the method against the leverage effect.

As to the jump component, we consider {\color{Blue} Merton type of jumps}:

equation[equation omitted — 146 chars of source]

The intensity of the jump component is set to be {a} constant value, i.e., $\lambda_t \equiv \lambda$ for all $t \geq 0$. For the values of $\lambda$ and $\vartheta$, we consider the following {scenarios:}

enumerate$\lambda = 50$ {and} ${\vartheta}=0.03$, which gives an average annualized volatility of about $\sqrt{0.04+50(0.03)^{2}}\approx 0.29$; • $\lambda = 100$ {and} ${\vartheta}=0.03$, which gives an average annualized volatility of about $\sqrt{0.04+100(0.03)^{2}}\approx 0.36$; • $\lambda = 200$ {and} ${\vartheta}=0.03$, which gives an average annualized volatility of about $\sqrt{0.04+200(0.03)^{2}}\approx 0.46$; • $\lambda = 1000$ {and} ${\vartheta}=0.01$, which gives an annualized volatility of about $\sqrt{0.04+1000(0.01)^{2}}\approx0.37$.

The reason for choosing these $\lambda$'s is to investigate how {\color{Blue} large levels of} jump intensity can affect the {\color{Blue} performance of the estimators}, {while} ${\vartheta}$ is selected accordingly {so that} the annualized volatility {\color{Blue} is reasonable}.

{We assume} that there are 252 trading days in a year and 6.5 trading hours in each day. We focus on 5-minute data, which is standard in the literature to avoid microstructure noise effects. Furthermore, the length of the data is set to be 1 month (21 trading days), 3 month (63 trading days), and 1/2 year.

Comparison of Different Thresholds

We now proceed to examine how the different “optimal" threshold approximation methods introduced in Section (ref) {\color{Blue} affect} the number of jump {misclassifications}. {\color{Blue} In Tables (ref), we report the average total number of jump mis-classifications corresponding to the four threshold approximation methods $B^{c1}$, $B^{c2}$, $B^{n1}$ and $B^{n2}$, as well as an oracle threshold, {where we use the second order approximation $B_{i}^{*2}$ in ((ref))} with all the true parameter values plugged in. In each case, we} compute {\color{Blue} the average number of jump misclassifications:}

equation[equation omitted — 386 chars of source]

where $m$ is the number of simulations, $X^{(j)}_{\cdot}$ and $N^{(j)}$ are the jth simulated paths of $X$ and $N$, {\color{Blue} respectively,} and $a\in\{c1,c2,n1,n2,*2\}$, depending on the used thresholding method. {\color{Blue} For the non-constant methods, we use an exponential kernel $K(x)=e^{-|x|}/2$ to estimate the spot volatility, which, as shown in Theorem (ref), is optimal. We ran only 4 iterations of the iterative algorithms described in Section (ref). As shown below, this typically suffices to reach convergence.}

{\color{Blue} The conclusion is that the jump detection method based on the second order approximation with non-constant volatility estimation (“n2" method) performs the best among all the four methods. Although, {as it should be expected, this is slightly} worse than the oracle one, it is remarkably close to the latter. Even for a relatively low value of $\lambda=50$, where is typically hard to estimate $\lambda$ and $C_{0}(f)$ because of relatively few jumps, the 2nd order local method is still a bit better than those based on constant threshold. For instance, for a time horizon of 1 month, the n2 method only misses about 1 jump out of the expected 4 jumps during the month. The difference between the constant and local thresholds becomes more crucial as the intensity of jumps increases. For an intensity of $200$, the method will only miss about 3 of the expected 16 jumps.}

{\color{Blue} As mentioned above, the results of Table (ref) were based on 4 iterations of the Algorithms of Section (ref). To assess the convergence of the algorithm, in Table (ref), we show the results of the average number of jump misclassifications $\bar{\mathcal{L}}^{a}$, as defined in ((ref)), for each of the first 4 iterations of Algorithm (ref) based on the 2nd order approximation. As it can be seen, convergence is typically reached after the 2nd iteration.}

\begin {table}

center[center omitted — 2,030 chars of source]

\end {table}

\begin {table}

center[center omitted — 1,990 chars of source]

\end {table}

Estimation of Jump Density at the Origin {\color{Blue} and Spot Volatility}

{We now study} the performance of the kernel estimator of the jump density at the origin that we proposed in Section (ref). Since we have already confirmed that the second order approximation of the optimal threshold with non-constant volatility estimation outperforms other thresholds, we will only consider {this threshold in this} and later subsections.

{\color{Blue} The results are shown in {\color{Blue} Table (ref)}. These basically confirm what we expect that the performance of the estimator improves as the time-horizon and intensity become larger (for the same level of jump variance). It is hard to compare the performance of the estimators when $\vartheta=0.03$ to those when $\vartheta=0.01$ and $\lambda=1000$ because, though we expect more jumps in the latter case, those will also be much harder to detect since $\vartheta$ is smaller.} Finally, an interesting phenomenon} is that we usually underestimate the jump density at the origin. This is acceptable for our purpose. Indeed, if we denote $\widehat{B^2}$ as the estimated second order threshold, we generally have $B^1 > \widehat{B^2} > B^2$. This is better than having $\widehat{B^2} < B^2$, in which case we might suffer significantly from {false positives (i.e., mis-classifying the increments of the continuous component as jumps)}.

\begin {table}

center[center omitted — 1,645 chars of source]

\end {table}

{\color{Blue} We finally give some illustrations about the performance of the the kernel/threshold spot volatility estimator ((ref)). We apply 4 iterations of the local Algorithm (ref) based on the 2nd order approximation of the optimal threshold. In Figure (ref), we show a prototypical realization of the variance process $\{V_{t}\}_{t\geq{}0}$ defined in ((ref)) together with the estimated spot variance process resulting from the 1st iteration (red dotted), from the final iteration 4 (long-dashed blue), and from the oracle (green double-dashed), which uses $B_{i}^{*1}$ in ((ref)) with the true values of $\sigma_{i}^{2}=V_{t_{i}}$, $C_{0}(f)$, and $\lambda$. We take $\lambda=200$, $\vartheta=0.03$, $T=6$ months, and $h=5$ minutes. The three spot variance estimates are close to each other and are able to fit well the overall level of the volatility through time. The Sum Of Square Errors, \[ {\rm SSE}=\sum_{i=1}^{n}(\hat{\sigma}_{t_{i}}^{2}-\sigma^{2}_{t_{i}})^{2}, \] for the 1st, 4th, and oracle estimates are respectively given by $1.6525$, $1.4457$, and $1.4450$. Figure (ref) shows the same results corresponding to $\lambda=1000$ and $\vartheta=0.01$. The SSE are in this case $2.0015$, $1.5069$, and $1.4006$ for the 1st, 4th, and oracle estimates, respectively.}

figure[figure omitted — 877 chars of source]
figure[figure omitted — 932 chars of source]

Conclusion and Future Work

In this paper, we study the problem of jump detection via the thresholding method, which {is obviously closely} related to the problem of spot volatility estimation. We extend the approximated optimal threshold of FLN2013 by considering a second-order approximation and a non-homogeneous parameter setting. The result is of theoretical interest since the remainder of the second order approximation is much smaller and, at the same time, the resulting threshold estimator is time-invariant, which makes more sense in reality. Monte Carlo studies also demonstrate the superior performance of the second-order approximation.

The higher accuracy comes with the price of more parameters to estimate. We first managed to build a threshold-kernel estimator of the jump density at the origin. We propose a different “optimal" threshold for this purpose and demonstrate the reason why this should be different from the original “optimal" threshold. The intuition is that we have to be more accurate when claiming that an increment contains a jump in order to have a good estimation of its density at the origin. We also put forward a modified version of the threshold-kernel estimator of spot volatility where increments that exceed the threshold are filtered out.

In order to implement the proposed methods, we need to resolve some key {obstacles}. Concretely, {estimates of} the optimal threshold, the jump density at the origin, and the spot volatility depend on each other. To resolve the issue, we propose an iterative threshold-kernel estimation scheme. Although we are not guaranteed that the iterative algorithm always converges, Monte Carlo studies show that this rarely creates any problem in reality.

{The spirit of jump detection by threshold method is to claim that a jump occurs whenever the absolute value of the increment of the process exceeds the threshold, which, by definition, is a binary outcome. In this case, when an increment is {close} to the threshold, a small difference in the increment can lead to totally different results. One way to alleviate such a problem is to estimate the probability that a jump happens during a specific time interval, which is similar to the idea of Logistic regression. This {\color{Red} suggests} an alternative approach to threshold-based classification. Given a non-decreasing function $F:[0, \infty) \to [0, 1]$ and an increment $|\Delta_{i}X|$, we can postulate that the probability that a jump occurs during $[t_{i-1},t_{i}]$ is $F(|\Delta_{i} X|)$. We can then adopt the following loss function, that is frequently {used in} classification problems:

align*[align* omitted — 200 chars of source]

Indeed, it could be cumbersome to optimize over all continuous functions $F$. However, we can try to limit ourself to a suitable, relatively small, class of possible functions $F$. One possible direction is to consider $F_B(x) = F(x / B)$, which is a generalization of what we have done in this paper. Another possible direction is to consider $F(x) = F_{1n}(x) \mathbf{1}_{\{ 0 < x < B\}} + F_{2n}(x) \mathbf{1}_{\{ x \geq B\}}$, where $F_{1n}$ and $F_{2n}$ are two functions that can depend on $n$. This can potentially provide insight on how the shape of $F$ should look like around the “optimal" threshold.}