EconBase
← Back to paper

Bayesian Inference on Volatility in the Presence of Infinite Jump Activity and Microstructure Noise

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

86,509 characters · 20 sections · 78 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bayesian Inference on Volatility in the Presence of Infinite Jump Activity and Microstructure Noise

abstractVolatility estimation based on high-frequency data is key to accurately measure and control the risk of financial assets. A L\'{e}vy process with infinite jump activity and microstructure noise is considered one of the simplest, yet accurate enough, models for financial data at high-frequency. Utilizing this model, we propose a “purposely misspecified" posterior of the volatility obtained by ignoring the jump-component of the process. The misspecified posterior is further corrected by a simple estimate of the location shift and re-scaling of the log likelihood. Our main result establishes a Bernstein-von Mises (BvM) theorem, which states that the proposed adjusted posterior is asymptotically Gaussian, centered at a consistent estimator, and with variance equal to the inverse of the Fisher information. In the absence of microstructure noise, our approach can be extended to inferences of the integrated variance of a general It\^o semimartingale. Simulations are provided to demonstrate the accuracy of the resulting credible intervals, and the frequentist properties of the approximate Bayesian inference based on the adjusted posterior.

MSC 2010 subject classifications: Primary 62M09; secondary 62F15.

Keywords and phrases: Bernstein-von Mises theorem, Semiparametric inference, It\^o semimartingales, high-frequency based inference, Microstructure noise

Introduction

In the past decade, jumps have played an increasingly important role in asset price modeling. The necessity of jumps is supported by both empirical and realistic considerations such as (i) sudden and relatively large changes observed in real stock prices; (ii) the implied volatility smile phenomenon, which is more pronounced for short maturity options; and (iii) the proper management of risk peterintro, Cont:2003. Though early on (e.g., Merton's model) the attention was centered on finite-jump activity models (i.e., those exhibiting finite jumps in finite time intervals), infinite-activity models are considered more realistic as suggested by many studies based on real asset returns infjump, Li2008Bayesian, Szerszen2009Baye, Yu2011MCMC,Yang2017jump. Here we consider a one-dimensional L\'{e}vy process $X=\{X_{t}\}_{t\geq{}0}$ defined on some probability space $(\Omega, \mathcal{F}, (\mathcal{F}_t)_{t\ge 0}, P)$ over a fixed time horizon $t \in [0, T]$ , which is a fundamental and widely-used tool to model jump processes with infinite activity. Concretely,

equation[equation omitted — 66 chars of source]

where $\mu\in\mathbb{R}$ and $\theta\in[0,\infty)$ are the drift and the variance parameters, respectively, $W=\{W_t\}_{t\ge 0}$ is a Wiener process, and $J=\{J_t\}_{t\ge 0}$ is an independent pure-jump L\'evy process. In financial applications, $X_{t}$ typically represents the log-return or log-price process $\log(S_{t}/S_{0})$ of an asset with price process $\{S_{t}\}_{t\geq{}0}$. In that case, the parameter $\sigma=\theta^{1/2}$ is called the volatility of the process and contributes to the total “variability" of the process $X$. Further details about the model and its components are given in \S (ref).

With improvements in computational power and the advent of electronic-based financial markets, high-frequency data (every minute, second, or even nanosecond) has become widely available. While exploiting the convenience of massive data, we also suffer from market microstructure frictions (e.g., serial autocorrelation, price discreteness, and temporary demand-supply imbalance) caused by the nature of trading at high frequency. In an attempt to explain the nature of tick-by-tick data, zhou1996noise and TSRV introduced the concept of microstructure noise, in which the observed transaction log-price $Y_{t}$ at time $t$ is a noisy measure of an underlying “efficient" log-price $X_t$:

equation[equation omitted — 113 chars of source]

Our purpose is to estimate the variance parameter $\theta$ based on high-frequency sampling observations $Y_{t_{0}},Y_{t_{1}},\dots,Y_{t_{n}}$ ($0=t_{0}<\dots<t_{n}=T$) of the process over a fixed period of time $[0,T]$. From the perspective of frequentist point estimation, when there is no microstructure noise, mancini2009 proposed a consistent estimator by eliminating those increments of the process, $\Delta_{i}Y:=Y_{t_{i}}-Y_{t_{i-1}}$, which are larger in absolute value than a suitably defined threshold. The asymptotic efficiency of the estimator with the restriction of a bounded variation jump process $J$ is proved later in cont2011. When the microstructure noise is taken into account but jumps are not present, several estimators have been proposed. The two-scale estimator in TSRV considered two different estimation scales of the process to estimate and eliminate the effect of the noise. The preaveraging approach in JACOD2009preaveraging replaced the increments $\Delta_{i}Y$ with a weighted summation over a small window. The kernel method in Barndorff-NielsenKernel utilized the weighted realized autocovariances. When both noise and jumps are present, PODOLSKIJ2009IA, podolskij2009 introduced the modulated bipower variation estimator using the bipower variation of the weighted average of the increments. The estimator is consistent, but cannot achieve the efficient convergence rate $n^{-1/4}$, which represents the best rate that can be achieved for the estimation problem in presence of noise and jumps. CHRISTENSEN2010 proposed two quantile-based realized volatility estimators by employing empirical quantiles of the averaged returns. The estimators are both consistent and asymptotically efficient, but only applicable for processes with finite jumps. More recently, Jing2014jumpest combined the preaveraging method of JACOD2009preaveraging and the thresholding ideas of mancini2009 to construct a consistent estimator of the integrated variance that is robust to both noise and infinite jump activity. The details of this estimator are explained in \S (ref).

Whereas there are numerous frequentist estimators available, the development of an explicit and efficient Bayesian approach which can accommodate high-frequency data remains a largely open problem. For a fully Bayesian approach, the joint posterior of the parameters must be derived based on the full likelihood function and a joint prior distribution for all the parameters in the model. Then, integrating the joint posterior over the nuisance parameters (in our case the parameters related to the jump component $J$ and microstructure noise $\varepsilon$) yields a marginal posterior distribution for the parameter of interest. Since the posterior is often intractable, Markov Chain Monte Carlo (MCMC) methods are typically used to sample from the joint posterior, and then numerical integration over the nuisance parameters is achieved by simply ignoring the corresponding MCMC output for those parameters. MCMC-based Bayesian methods have been applied to the volatility estimation problem by several studies. Jones99bayesianestimation and Eraker2003impact used MCMC for a diffusion process augmented by a Poisson jump process. More recently, additional model complexity has been accommodated by taking infinite activity into consideration. Yu2011MCMC proposed an MCMC estimation method using both spot and option prices. Their jumps are assumed to follow either a variance gamma process or an $\alpha$-stable process. jasra2011a developed an automated sequential Monte Carlo algorithm by adding an additional re-sampling step for variance gamma jumps. he2014efficient applied a slice sampling approach with a similar variance gamma assumption. Griffin2016 incorporated realized variation and realized power variation into a MCMC procedure, and analyzed a generalized variance gamma process. Yang2017jump considered both returns and the Chicago Board of Options Exchange (CBOE) Volatility Index (VIX) to obtain the posterior for the jump part. The variance gamma process and normal inverse gamma process were considered.

Although the papers mentioned above considered Bayesian inference derived from the joint posterior, they all require strong assumptions about the structure of the jumps, which severely limit the practical value of these methods. Without these simplifying assumptions, it is quite challenging to write down the full likelihood function under the semi-parametric setting (ref), which means that it is also difficult to obtain the full joint posterior without such assumptions. One of these assumptions is the choice of a particular specification of the jump process $J$, among many possible jump processes. However, empirical results in Li2008Bayesian, Yu2011MCMC, and Kou2017 suggested that different jump assumptions lead to different estimation results for volatility. The posterior depends heavily on the structure of the jumps. Thus, sticking to just one jump type increases the possibility of misspecification and, therefore, can lead to inaccurate estimation and inference.

Moreover, specifying and calculating the distribution of the jump component may incur heavy computational costs, especially when working with high-frequency data. For this reason, nearly all of the aforementioned studies consider only daily returns data. Some literature like jasra2011a and Griffin2016 did apply their methods to hourly data and 5-minute data, respectively. However, they both fixed one of the parameters of the jump process as constant, in order to reduce the computational load.

The difficulties of deriving the posterior and the associated heavy computational costs are mainly caused by the jumps, which are only related to the nuisance parameters. Our target of estimation, the variance or volatility, is not affected by the jumps, and is modeled by a simple Gaussian process, for which Bayesian inference can be more easily obtained. Based on this observation, one plausible idea to tackle the problem is to ignore the nuisance parameters in the nonparametric part of the process, replace the nuisance parameters in the parametric part by their consistent estimators, and construct a posterior only for the parameter of interest. The advantages of such an approach are that one need not specify a prior on the jump process, and it is not necessary to obtain samples from the full joint posterior. By contrast, we will directly obtain an approximation to the marginal posterior for the volatility, which we will show can be used for accurate Bayesian inference. This approach was recently used by Martin. They derived a `purposely misspecified' posterior for a jump-diffusion model with constant volatility, finite jump activity and without microstructure noise, which targets the parameter of interest, the volatility, directly. Using a misspecified model on purpose, the inherent difficulty of specifying the likelihood function in a nonparametric model is tackled by omitting the complicated nuisance component of the model. The bias and the inaccurate variance caused by the misspecification are later corrected by applying a location shift and rescaling the likelihood using a Gibbs posterior. They showed that the adjusted posterior possesses good asymptotic properties, as guaranteed by a Bernstein-von Mises theorem.

In this paper, we study a `purposely misspecified’ posterior for the variance $\theta$ of the model (ref) either with or without microstructure noise, which is a considerably more difficult and realistic setting in comparison to the finite jump activity model without microstructure noise that was studied by Martin. Our main result is a Bernstein-von Mises Theorem for the adjusted posterior for the volatility parameter, which shows that the proposed posterior is asymptotically normal and centered at a consistent estimator, and with variance shrinking at rates $n^{-1/2}$ and $n^{-1}$, respectively, depending on whether a microstructure noise is incorporated or not in the model.

The novel contributions of this paper can be summarized as follows. First, we allow the jump process to be any L\'{e}vy process with bounded variation, i.e. there is no parametric assumption about the nuisance component, and no assumption of finite jump activity. We also allow for an additive microstructure noise in the data. These relaxations of the stronger assumptions made in the existing literature help to alleviate inaccuracies introduced by model misinterpretation, and also avoid expensive computational costs. In fact, we also show that in the situations when the microstructure noise can be ignored (e.g., when working with medium range frequencies), our approach can be extended to the estimation of the integrated variance of a general It\^o semimartingale $X$. In particular, we allow stochastic volatility and a general pure-jump semimartingale component $J$.

It is important to remark that our proposed inference procedure is among the first Bayesian approaches that can accommodate truly high-frequency data; due to high computational costs and lack of theoretical performance guarantees, most of the existing literature involves methods which are only applicable to low frequency data, such as daily observations. Finally, our results suggest that, under certain circumstances, misspecification on purpose can serve as a vehicle for accurate approximate Bayesian inference about low-dimensional interest parameters in complex, possibly infinite-dimensional models.

The paper is organized as follows. A detailed description of the setting and model are provided in \S (ref). Differences between finite and infinite activity when deriving the `purposely misspecified' posterior are highlighted in \S (ref). This analysis reveals the importance of proposing a modified version of the Bernstein-von Mises theorem, which is stated in \S (ref). The misspecified model is presented in \S (ref), and further extended in \S (ref). The main results are stated in \S (ref) and \S (ref). Simulation results given in \S (ref) illustrate the performance of our procedures. Discussion and concluding remarks are in \S (ref). The proofs and further technical details appear in the Appendix.

Model setup

As mentioned in \S (ref), we consider a one-dimensional continuous-time process defined on some probability space $(\Omega, \mathcal{F}, (\mathcal{F}_t)_{t\ge 0}, P)$ over a fixed and finite time horizon, $X=\{X_t; t\in [0,T]\}$, which is assumed to follow model (ref). It consists of a drift part with constant coefficient $\mu\in \mathbb{R}$, a diffusion part with constant coefficient $\theta\in \mathbb{R}_{+}$, which represents the volatility or variance, and a pure jump part $J = \{J_t\}_{t\ge 0}$. The parameter space for $\theta$, denoted as $\Theta$, is assumed to be a bounded and open subset of $(0,+\infty)$ such that $0\notin \bar{\Theta}$.

The jump component $J$ is assumed to be a pure jump L\'{e}vy process, which is used in many fields of science. In mathematical finance, a L\'{e}vy process is widely recognized to provide a better fit to intraday returns than plain Brownian motion. A comprehensive overview of the applications of L\'{e}vy processes can be found in levyprocessbook and Cont:2003. A L\'evy processes is defined as a c\`{a}dl\`{a}g, real valued stochastic process, which has independent and stationary increments, and is stochastically continuous. It is known that a L\'evy process $X$ takes the general form ((ref)) with $J$ defined as

equation[equation omitted — 231 chars of source]

where $\mu$ is a Poisson random measure on $\mathbb{R}_{+}\times \mathbb{R}\backslash\{0\}$ with mean measure $\nu(dx)dt$ such that $\int_{\mathbb{R}\backslash\{0\}}(|x|^{2}\wedge{}1)\nu(dx)<\infty$. This is the so-called L\'{e}vy-It\^{o} decomposition of $X$ and $\nu$ is called the L\'evy measure of $X$.

The contaminated process, which equals to $X_t$ plus a noise component $\varepsilon_t$, is observed at equally-spaced discrete times $\{0=t_0< t_1<\ldots<t_n=T\}$, $t_j - t_{j-1} = \Delta_n = T/n$. More specifically, we observe

equation[equation omitted — 109 chars of source]

The data is assumed to be generated by the model ((ref))-((ref)) with true volatility value $\theta^{*}$, which is the target to be estimated. The L\'{e}vy model with microstructure noise ((ref)) is considered one of the simplest, yet accurate enough, models for financial data at high-frequency. For an assessment of its empirical accuracy, we refer to FLKiseop2015.

The process $Y$ satisfies the following assumptions:

customass{(N)} \begin{enumerate} • The microstructure noise components, $\varepsilon = \{\varepsilon_{t_j}\}_{j=1}^n$, are independent and identically distributed (i.i.d.), and follow a $\mathcal{N}(0,\sigma_{\varepsilon}^2)$ distribution. In Bayesian framework, we assume that the i.i.d. holds true conditionally on the unknown parameter $\sigma_{\varepsilon}$. • The processes $\varepsilon$ and $X$ are independent. \end{enumerate}
customass{(JD)} The Blumenthal-Getoor index $\alpha$ of $J$ is less than $1$: \begin{equation} \alpha = \inf\left\{p>0:\int_{|x|\le 1} |x|^p\nu(dx)<\infty\right\} < 1. \end{equation} In particular, this implies that the paths of the process $J$ are of bounded variation, almost surely.
customass{(JF)} The process $J$ has a finite $16$th moment. Combined with (ref)-1, this assumption is equivalent to $$\int_{|x|\ge 1} x^{16}\,\nu(dx)<\infty.$$
remark\begin{enumerate} • hansen2006moderatefreq suggested that the independence assumption for $\varepsilon$ and $X$ is reasonable for moderate intraday frequency (e.g. 1 minute). • For a L\'evy process, the Blumenthal-Getoor index $\alpha$ controls the small jump activity of the process: it becomes larger as the small jumps are more persistent. The assumption of $\alpha<1$ is inspired by cont2011 and JacodSum, and used later in \S (ref) to apply a central limit theorem (CLT) for a threshold estimator of the volatility. JacodSum concluded that when $\alpha\ge 1$, there is no CLT in general for a realized quadratic threshold estimator of the integrated variance. Its rate of convergence to the integrated variance is much slower than $n^{-1/2}$. A detailed proof of both the CLT and lack-of-CLT can be found in cont2011. A similar bounded variation assumption also appear in previous studies, such as CHRISTENSEN2010, cont2011, JacodSum, and Jing2014jumpest. • It is important to remark that in the absence of microstructure noise, we can take a stochastic volatility model and much more general pure-jump semimartingales $J$ of bounded variation (see \S (ref)). We also don't require the condition (ref). \end{enumerate}

For future reference, let us recall the following common notation for the increments and jumps of an arbitrary continuous-time c\`adl\`ag process $\{U_{t}\}_{t\geq{}0}$: \[ \Delta_{i}U=\Delta_{i}^{n}U=U_{t_{i}}-U_{t_{i-1}},\quad \Delta U_{t}=U_{t}-U_{t^{-}}. \]

Comparison with finite jump activity models

In this section, we present a motivating example using a simpler finite jump activity model, in order to illustrate the usefulness of the approximate Bayesian inference obtained via purposeful misspecification. Martin proposed this approach, but did not make comparisons to the true marginal posterior for the volatility parameter. In the next subsection, we provide this comparison through a simulation experiment. We then summarize the theoretical results in Martin and explain what issues arise when considering the more complicated and realistic setting of infinite jump activity.

An illustration through simulation

We first empirically compare the “purposely misspecified" posterior from Martin with a marginalized full Bayesian posterior. The goal of the comparison is to assess the accuracy of the former method and to motivate our approach. A simple jump diffusion model without noise is considered. Concretely, model (ref) is used with a compound Poisson jump process: \[ J_{t}=\sum_{i=0}^{N_{t}}\xi_{i}. \] Here, $N=\{N_{t}\}_{t\geq{}0}$ is a Poisson process with rate $\lambda$, and $\{\xi_i\}_{i\geq{}1}$ are i.i.d. random variables independent of $N$ and $W$. We assume that $\{\xi_i\}_{i\geq{}1}$, which represent the jump sizes, follow a uniform distribution $U(-1,1)$. This assumption enables us to derive a joint posterior and perform Gibbs sampling for the parameters $ \Theta = (\mu, \theta, \lambda)$. The other parameters and settings are inherited from Martin: $$\lambda=5,\quad \mu = 1,\quad \theta = 10,\quad n = 5000,\quad T = 1.$$ For simplicity, in what follows, we approximate the Poisson process by a Bernoulli process; namely, ${N}$ is assumed to be a point process such that $P[{N}_{t_{i}}-{N}_{t_{i-1}}=1]=\lambda \Delta_{n}$ and $P[{N}_{t_{i}}-{N}_{t_{i-1}}=0]=1-\lambda \Delta_{n}$.

The joint posterior density based on the data $\pmb{X}^{(n)}=(\Delta_1 X,\dots, \Delta_n X)$ can be written as

align*[align* omitted — 395 chars of source]

The priors chosen for $\mu, \theta, \lambda$ are a standard Gaussian distribution, an inverse gamma distribution, and a beta distribution, respectively. The posterior for $\theta$ is estimated by two methods: (i) Gibbs sampling from the full joint posterior, followed by numerical integration to yield the marginal posterior (i.e., we simply ignore the MCMC output for the nuisance parameters $\mu$ and $\lambda$); and (ii) a direct posterior for $\theta$ obtained by purposeful misspecification. We emphasize that the Gibbs sampling approach, which is exact modulo finite simulation error, is only available here because of the very strong assumptions made regarding the jump process. This method is not available for the most complicated and realistic settings we consider in this paper. The second method is an approximation using a misspecified model to directly obtain a posterior for $\theta$ without the need to first obtain the full joint posterior and marginalize. The latter method, as shown in this paper, works quite well even in much more complicated and realistic settings than those considered in this section. Figure (ref)-(ref) compares the two approaches. Figure (ref) shows the posteriors for 10 different simulations. The `purposely misspecified' posterior typically resembles quite well the Gibbs distribution of the samples simulated from the joint posterior, which is supposed to recover the correct posterior for the volatility through marginalization of the joint one. Both posteriors center around the true volatility. The 95% highest posterior density intervals are shown in Figure (ref). The similarities of the two empirical posteriors as well as their credible intervals demonstrate the accuracy of the `purposely misspecified' posterior, and therefore, the validity of the inference based on it.

figure[figure omitted — 908 chars of source]

In general, it is quite complicated to perform fully Bayesian analysis for infinite jump activity models based on high-frequency data because of the lack of tractable joint posteriors. To perform MCMC sampling from those joint posteriors, some studies (e.g. Li2008Bayesian) consider the unobserved jump increments $\Delta_i J$ as a latent parameter. However, with high frequency data, this may cause numerical difficulties. Taking model 3 in mancini2009, for example, a variance gamma jump component $J$ is utilized with drift $-0.2$, variance $0.2$, and variance of the subordinator $0.23$ (see more details in \S (ref)). The sample size is $1000$, and the time interval $\Delta_n=0.001$. Under these model settings, the jumps have extremely small sizes. More than $90\%$ of the jump sizes are less than $10^{-7}$, and, hence, they are difficult to be recovered in the MCMC sampling. As mentioned in \S (ref), in the previous studies that incorporate infinite jump activity for high-frequency data, models are simplified in order to conduct MCMC sampling.

Theoretical challenges

Martin applied their purposely misspecified approach to the simpler model setting of an uncontaminated jump-diffusion model with constant volatility and finite jump activity. They first constructed a misspecified model by omitting the jump part $J$. Under this misspecified model, the resulting misspecified posterior was shown to be asymptotically normal conditionally on a given path of $J$. Since the result works for all possible $J$, it can be generalized to a version which does not depend on $J$. Even though such an asymptotic normality does hold for a suitably centered and scaled misspecified posterior for the volatility, the misspecification of the model has the adverse effect of causing this misspecified posterior to center in the wrong place and to have an incorrect and inefficient variance compared to the true marginal posterior obtained by marginalizing the full joint posterior over the drift and jump parts of the model. Therefore, Martin proposed to correct for the bias and inefficiency of the misspecified posterior by, respectively, shifting the center by an estimate of the bias, and rescaling the log likelihood using a properly chosen temperature parameter. Since the Bernstein-von Mises theorem involves convergence in total variation norm, and this norm is invariant with respect to location shifts, the resulting corrected posterior for volatility still admits a Bernstein-von Mises theorem but with a correct center and efficient variance equal to the Cram\'{e}r-Rao lower bound.

In a model with infinite jump activity, we can similarly ignore the jump part and consider a misspecified model, but it is impossible to conclude an unconditional Bernstein-Von Mises theorem from the analogous result for the conditional posterior given a fixed path of the jump process $J$. The main reason is that for a jump process with infinite activity, the realized quadratic variation $[J]_n = \sum_{i=1}^{n} \Delta_i J^2 = \sum_{i=1}^{n} (J_{t_{i}}- J_{t_{i-1}})^2$ does not converge to the quadratic variation $[J]= \sum_{0\le t<T} (J_t - J_{t-})^2$ for almost every path of $J$ (i.e., a.s. convergence does not hold but merely convergence in probability). The almost sure consistency is necessary to prove the properties of the posterior, which is required when proving the local asymptotic normality (LAN) of the likelihood and an optimal convergence rate of the posterior mean. The satisfaction of these two conditions facilitates the establishment of a Bernstein-von Mises theorem under misspecification (see BVM). In Martin's model, because the jump part $J$ is assumed to have finitely many jumps in a finite time interval, $J$ can be expressed as a finite summation of the discontinuities. Thus, there exists $n_0\in\mathbb{N}$, such that for $n>n_0$, the quadratic variation $[J] $ is exactly equal to $ [J]_n$. However, with infinite jump activity, the convergence of $[J]_n$ to $[J]$ does not hold for almost every path of $J$. This complication leads to the failure of the usual conditions used to prove a Bernstein-von Mises theorem.

On the other hand, for a general semimartingale, it is well-known that $[J]_n$ does converge to $[J]$ in probability JacodAndProtterDiscret. Furthermore, for L\'evy processes, a rather good rate of convergence of $O_{p}(n^{-1/2})$ can be obtained (see Lemma (ref) below). We find that this weaker convergence (i.e. in probability rather than almost surely) is enough to demonstrate the desired property of the posterior by applying an unconditional version of the Bernstein-Von Mises theorem and skipping the intermediate results under the conditional probability measure given the jump part.

Besides the infinite jump activity complicating the nonparametric part of the model, the parametric part is also affected by the presence of the noise $\varepsilon$. Good news is that the variance of the noise, $\sigma_\varepsilon^2$, is an additional nuisance parameter. The adjusted posterior for volatility, and the associated Bernstein-von Mises theorem, must include corrections for deliberately ignoring the presence of microstructure noise.

A semiparametric version of the misspecified BvM Theorem

As explained in the previous section, the misspecified Bernstein-von-Mises Theorem of BVM plays a crucial role in proving the asymptotic properties of the purposely misspecified posterior. To accommodate the more complicated settings of our model, the result needs to be generalized to a semiparametric version, which is stated as follows.

thmConsider the space $\Omega^{(n)}:=\Omega_{1}^{(n)}\times \Omega_{2}:=\mathbb{R}^{n}\times D([0,\infty))$ (D represents the Skorokhod space of all c\`{a}dl\`{a}g $\mathbb{R}$-valued functions) and a collection of semiparametric models on $\Omega^{(n)}$, \[ \left\{P^{(n)}_{\left( (\theta, \eta), \nu \right)} : (\theta, \eta) \in \Theta, \nu\in U \right\}, \] where $\Theta$ is an open subset of $\mathbb{R} \times \mathbb{R}^d$ and $U$ is an open subset of an infinite dimensional topological space $\mathbb{F}$. Let $P_0^{(n)}:={P}^{(n)}_{\left( (\theta^*, \eta^*), \nu^{*} \right)}$ and let $Z^{(n)}=(Z_{1},\dots,Z_{n})$ and $\{Y_{t}\}_{t\geq{}0}$ be the canonical processes on $\Omega^{(n)}$ defined for $\omega=(\omega_{1},\omega_{2})\in \Omega_{1}^{(n)}\times \Omega_{2}$ as $Z_{i}(\omega)=\omega_{1i}$ and $Y_{t}(\omega)=\omega_{2}(t)$, respectively. Define $\Phi = g( \theta, \eta, Y_{\cdot})$ and $\Phi^\dag = g( \theta^*, \eta^*, Y_{\cdot})$, where $Y_{\cdot}$ denotes the sample path of $J$ and $g:\Theta\times D([0,\infty))\to\Theta'\subset \mathbb{R}$ is a known deterministic function. Our data consists of $X^{(n)}:=(X_{1},\dots,X_{n}):= T(Z^{(n)},Y_{\cdot})$, where the function $T:\mathbb{R}^{n}\times D([0,\infty)]\to\mathbb{R}^{n}$ is known. Suppose there are purposely misspecified models for $X^{(n)}$ denoted as $\tilde{P}_{{\vartheta}}(\cdot):=\tilde{P}(\cdot |{\vartheta})$, $\vartheta\in \Theta'$, which are distributions on $\mathbb{R}^n$ parameterized by $\vartheta$ with densities $ \tilde{p}_{{\vartheta}}$. Let $\Pi$ be a prior distribution with a density $\pi$ that is continuous and positive on $\Theta'$ . Define the misspecified posterior distribution based on $\Pi$ and $\tilde{P}_{{\vartheta}}(\cdot)$ as $$ \Pi^n(\varphi \in B|X^{(n)}) = \frac{\int_B \tilde{p}_\varphi(X^{(n)}) \pi(\varphi) \, d\varphi}{\int \tilde{p}_\xi(X^{(n)}) \pi(\xi)\, d\xi}, \quad B\in \mathcal{B}(\Theta').$$ Assume $\{ \tilde{P}_\vartheta , \vartheta \in \Theta' \}$ satisfy a stochastic local asymptotic normality (LAN) condition relative to a given sequence $\delta_n \rightarrow0$ as norming rate, i.e. there exist some random quantities $\Delta_{n,}$ and $V_{n}$ such that for every compact set $K\in\mathbb{R}$ and $\epsilon>0$, \begin{align} {P}^{(n)}_{ 0} \left( \sup_{h\in K} \left| \log \frac{\tilde{p}_{\Phi^\dag+\delta_n h} }{\tilde{p}_{\Phi^\dag} } (X^{(n)})-V_{\Phi^\dag} \Delta_{n,\Phi^\dag} h -\frac{1}{2}V_{\Phi^\dag}h^2 \right| > \epsilon \right) \rightarrow 0, \quad as n \rightarrow \infty. \end{align} Also, for any sequence of constants $M_n \rightarrow \infty$, the posterior $\Pi^n$ is assumed to satisfy \begin{equation} {\Pi}^n \left( {| \varphi-\Phi^\dag|} >\delta_n M_n | X^{(n)} \right) \stackrel{{P}^{(n)}_{ 0}}{\rightarrow} 0, \quad n\rightarrow \infty. \end{equation} Then, $\Pi^n$ converges to a sequence of normal distributions in total variation: \begin{equation*} {P}^{(n)}_{ 0} \left( \sup_B\left| {\Pi}^n \left( (\varphi - \Phi^\dag )/\delta_n \in B | X^{(n)} \right) - N_{\Delta_{n,\Phi^\dag}, V_{\Phi^\dag}^{-1}}(B) \right| >\epsilon\right) \rightarrow 0, \quad n\rightarrow \infty. \end{equation*}

The proof of the above result follows the original proof in BVM. The main modifications are changing the almost sure convergence to convergence in probability, and adding a nuisance parameter which does not affect the proof.

remarkCondition ((ref)) above is equivalent to $$ {{P}^{(n)}_{ 0}} \left[ {P}^{(n)}_{ 0}\left( \sup_{h\in K} \left| \log \frac{\tilde{p}_{\Phi^\dag+\delta_n h} }{\tilde{p}_{\Phi^\dag} } -V_{\Phi^\dag} \Delta_{n,\Phi^\dag} h -\frac{1}{2} V_{\Phi^\dag}h^2 \right| > \zeta\ \bigg|\ Y_{\cdot} \right) > \epsilon\right] \rightarrow 0, $$ for all $\zeta, \epsilon>0$. This means that as $n\rightarrow \infty$, the set of those ${Y_{\cdot}} $ which satisfy the condition will cover its probability space with probability 1. This is weaker than the misspecified Bernstein-von-Mises Theorem in BVM when applying their theorem with ${P}^{(n)}_{ 0} ( \cdot |\ {Y_{\cdot}})$, which implies for almost all paths ${Y_{\cdot}} $, $$ {P}^{(n)}_{ 0} \left( \sup_{h\in K} \left| \log \frac{\tilde{p}_{\Phi^\dag+\delta_n h} }{\tilde{p}_{\Phi^\dag} } - V_{\Phi^\dag}\Delta_{n,\Phi^\dag} h -\frac{1}{2} V_{\Phi^\dag}h^2 \right| > \eta\ \bigg|\ {Y_{\cdot}} \right)\rightarrow 0, \text{ for all $\epsilon>0$}. $$ The second condition ((ref)) and conclusion can be compared with their counterparts in BVM in the same way.

The misspecified model

Our methodology starts with a misspecified model ignoring the drift and the jump component. Namely, $Y_t$ is assumed to follow

equation[equation omitted — 116 chars of source]

This means that we first misinterpret the increments of the underlying process $X$ as independent Gaussian variables, with mean zero, and variance $\theta$.

Under the misspecified model (ref), our target of estimation is still $\theta$, but what it represents changes because of the misspecification. In the absence of jumps, $\theta$ measures the total variation of the underlying process $X$ per unit time and, hence, it can efficiently be estimated by the scaled realized quadratic variation,

equation[equation omitted — 82 chars of source]

which coincides with the maximum likelihood estimator (MLE) of the parameter $\theta$ in the underlying misspecified model $X_{t}=\theta^{1/2}W_{t}$. However, under the model $X_{t}=\theta^{1/2}W_{t}+J_{t}$, $\theta$ merely controls the variation of the continuous component and, in the infill limit, the realized quadratic variation ((ref)) will aggregate both the true volatility, $\theta^*$, and the scaled variation introduced by the jump process $J$, $T^{-1}[J]$, where $[J]:=\sum_{s\leq{}T}(\Delta Y_{s})^{2}$. Throughout, this total variation is denoted as

equation[equation omitted — 75 chars of source]

which takes values on the random parameter domain $$ \Theta' := \left\{ \theta + T^{-1}[J]; \theta\in\Theta\right\}. $$ For any sample path of $J$, $\Theta'$ is an open set in $(0,+\infty)$, and $0\notin \bar{\Theta}'$. Furthermore, there exists some deterministic constant $\delta_0 >0$ such that $\Theta\subset (\delta_0,+\infty)$ and, hence, $\Theta'\subset (\delta_0,+\infty)$.

In \S (ref), we explicitly write the misspecified likelihood function and the corresponding MLE for $\theta$ under the model (ref). Bayesian inference under this misspecified model is proposed in \S (ref). We will show that, given that the data $Y$ is misinterpreted by the model (ref), the posterior of $\theta$ can be approximated by a normal distribution. Further extensions are subsequently considered.

Misspecified likelihood function and MLE

Let us first note that, because of the presence of the noise $\varepsilon$, the increments $\Delta_j Y = Y_{t_j}-Y_{t_{j-1}}$, $j = 1,2,\ldots,n$, are not independent. To deal with the dependency and write an explicit likelihood function, we follow {gloter_jacod_2001_est, lan} and transform the observed data $\{\Delta_j Y\}_{j}$ into independent random variables $\{R_j\}_{j}$ via ${\textbf{R}} = (P_n) (\Delta {\textbf{Y}})$, where ${\textbf{R}} = (R_1, \ldots, R_n)'$, $\Delta{\textbf{Y}} = (\Delta_1Y,\Delta_2Y,\ldots,\Delta_{n-1}Y,\Delta_n Y)'$, and $P_n$ is a symmetric orthogonal matrix with entries $$ p_{ij}^n := \sqrt{\frac{2}{n+1}}\sin \frac{ij \pi}{n+1}, \quad i,j = 1,2,\ldots,n.$$ gloter_jacod_2001_est, lan showed that, under the misspecified model (ref), ${R}_j$ is Gaussian distributed, with mean zero, and variance equal to $$\lambda_j^n(\theta) := \theta \Delta_{n} + 2{\sigma_{\varepsilon}}^2 \left(1-\cos \frac{j\pi}{n+1} \right),\quad j=1,2,\ldots,n.$$ For future reference, let us also note that under the true model (ref), the conditional distribution of $R_j$ given $J$ is $$R_{j}|J\sim\mathcal{N} \left( \mu \Delta_{n} + \sum_{i=1}^n p_{ij}\Delta_i J, \lambda_j^n(\theta) \right).$$

Based on these Gaussian variables, the likelihood function of the parameters $\theta$ and $\sigma_\varepsilon^2$ given the data $\{\Delta_j Y\}$ can be explicitly written under the misspecified model. However, note that only $\theta$ is the parameter of interest, while $\sigma_\varepsilon^2$ is merely the nuisance parameter. Instead of writing the likelihood function based on $\lambda_j^n(\theta)$ and maximizing it over a two dimensional space, we replace the nuisance parameter, $\sigma_\varepsilon^2$, with its consistent estimator $\hat{\sigma}_\varepsilon^2 =\frac{1}{2n} \sum_{j=1}^n \Delta_j Y^2$, and then, obtain a pseudo-likelihood function for $\theta$. The properties of $\hat{\sigma}_\varepsilon^2$ and the rationale of the replacement {are} further demonstrated in Lemmas (ref) and (ref). Then, it is natural to {consider the following} misspecified log likelihood function $\tilde{l}_n$ of $\theta$ given the data $\Delta Y_1, \Delta Y_2,\ldots,\Delta Y_n$:

equation[equation omitted — 353 chars of source]

The corresponding MLE $\tilde{\theta}_n$ is the root of the score function

equation[equation omitted — 326 chars of source]

We further assume that the MLE $\tilde{\theta}_n$ is unique.

remarkThe misspecified likelihood function (ref) can be simplified and directly applied to a model without the microstructure noise (i.e., $Y=X$ in ((ref))) by taking $\sigma_\varepsilon^2 = 0$ and $\hat{\sigma}_\varepsilon^2 = 0$. Then, \begin{equation} \tilde{l}_n(\theta) = -\frac{1}{2} \sum_{j=1}^{n} \left\{\log \theta \Delta_{n} + \frac{R_j^2}{\theta \Delta_{n} } \right\}=-\frac{1}{2} \sum_{j=1}^{n} \left\{\log \theta \Delta_{n} + \frac{\Delta_j Y^2}{\theta \Delta_{n} } \right\}. \end{equation} In this case, the MLE can be obtained in closed form as \begin{equation} \tilde{\theta}_n = \frac{1}{T}\sum_{i=1}^n R_i^2 = \frac{1}{T} \sum_{i=1}^n (\Delta_i Y)^2 = \frac{1}{T} \sum_{i=1}^n (\Delta_i X)^2. \end{equation} Thus, the misspecified model is consistent with the one in Martin and, hence, the model with finite jump activity can be viewed as a particular case of our results.

Bernstein-von Mises Theorems

We assume that the prior distribution $\Pi$ of $\theta$ possesses a continuous and positive density $\pi$ on $(\delta_0,+\infty)$. Denote ${P_*}$ as the distribution of the process $\{ Y_t \}_{t\ge 0}$ under the true model (ref), and ${E_*}$ as the corresponding expectation. Based on the prior $\Pi$ and the likelihood function (ref), we introduce the Gibbs posterior $\Pi^n$ Zhang2006gibbs,jiang2008gibbs with temperature parameters $\kappa_{n}$ as

align[align omitted — 196 chars of source]

where $A$ is a Borel set of $\mathbb{R}^+$. The Gibbs posterior increases the flexibility of the Bayesian procedure, which would allow us to further correct for the misspecification. Specifically, the misspecification causes the posterior for volatility to contract too quickly, making the Bayes estimator (e.g. the posterior mean) superefficient. Rescaling the likelihood flattens out the likelihood and also the posterior, slowing down the contraction of the posterior. Choosing the temperature parameter optimally will make the posterior contract at the efficient rate established by frequentist asymptotic analysis. We assume that $\kappa_{n}$ converges in probability to a random variable $\kappa^\dagger$ as $n\rightarrow \infty$ under the true measure ${P}_*$. Note that $\kappa_n$ may be data-dependent, and therefore it is possible that the random variable $\kappa^\dag$ also depends on the data under ${P}_*$.

Our main result states that, as the sample size $n$ increases, the misspecified posterior based on $\Pi$ and the misinterpreted data $\{\Delta Y_i\}$ will be approximately normal and centered at the maximum likelihood estimator $\tilde{\theta}_n$ obtained from the misspecified likelihood (ref) under the true measure $P_*$. The asymptotic variance is equal to the temperature parameter $\kappa^\dag$ times the inverse of the Fisher information of the misspecified likelihood. We give two versions. The first result covers situations where the microstructure noise $\varepsilon$ can be ignored. This is the case when, for instance, we use medium range frequencies such as 5-minute or daily observations. We achieve the standard $n^{-1/2}$ rate of convergence. The second result covers the more realistic case where the microstructure noise is explicitly incorporated in the model. This is needed when working with ultra high frequencies. This comes at the cost of a slower $n^{-1/4}$ rate of convergence.

thmSuppose that the data $Y_{t_{0}},\dots,Y_{t_{n}}$ is generated according to ((ref))-((ref)) with $\varepsilon_{t}\equiv 0$ and Assumption (ref)-1 is satisfied. Then, the misspecified posterior defined in ((ref)) with $\tilde{l}_{n}$ given as in ((ref)) and $\kappa_{n}\stackrel{P_{*}}{\to} \kappa^{\dagger}$, for some positive r.v. $\kappa^{\dagger}$, can be approximated by a normal distribution in the sense that $$ TV\left(\Pi_n, \ \mathcal{N}( \tilde{\theta}_n, 2\kappa^\dag \theta^{\dagger 2}n^{-1} ) \right) \stackrel{{P}_*}{\rightarrow} 0, \quad \text{ as }n\rightarrow \infty,$$ where $TV$ represents the total variation distance, $\tilde{\theta}_{n}$ is the MLE ((ref)), and $\theta^{\dagger}$ is defined in ((ref)).
thmUnder the framework and assumptions (ref), (ref), and (ref) above, the misspecified posterior $\Pi^n$ defined in (ref) with $\tilde{l}_{n}$ given as in ((ref)) and $\kappa_{n}\stackrel{P_{*}}{\to} \kappa^{\dagger}$, for some positive r.v. $\kappa^{\dagger}$, is such that $$ TV\left(\Pi_n, \ \mathcal{N}(\tilde{\theta}_n, 8\kappa^\dag \theta^{\dagger 3/2} {\sigma_{\varepsilon}} n^{-1/2} ) \right) \stackrel{{P_*}}{\rightarrow} 0, \quad \text{ as }n\rightarrow \infty,$$ where $\tilde{\theta}_{n}$ is the corresponding MLE (i.e., the root of the score function ((ref))) and $\theta^{\dagger}$ is defined in ((ref)).

The proofs of the two theorems utilize Theorem (ref), and are contained in the Appendix.

remarkIt is worth nothing that Theorem (ref) holds without any restriction on the Blumenthal-Getoor index $\alpha$. In fact, this result holds for a large class of pure-jump semimartingales $J$ and even quite general stochastic volatility models (see Section (ref)). The restriction of $\alpha<1$ is needed when correcting the posterior as shown below.

Correcting for misspecification

The main conclusions of Theorems (ref) and (ref), namely, as $n\rightarrow \infty$,

equation*[equation* omitted — 357 chars of source]

state that the misspecified posterior $\Pi_n$ is approximately normally distributed, centered at $\tilde{\theta}_n$, which is a biased estimator of $\theta^{*}$ in the presence of jumps. Furthermore, the asymptotic variance may not be the most efficient either since we ignored the drift and the jump components on purpose. To adjust the bias and variance, what we need is a consistent estimator for the true parameter $\theta^*$, which admits a feasible central limit theorem. In what follows, we will first propose a general correction procedure and the corresponding Bernstein-von Mises theorem for any estimator with these two properties. Concrete instances of these estimators for both the no-noise and the general cases are presented thereafter.

Suppose we have an estimator $\hat{\theta}_n$ of $\theta^*$ such that

align[align omitted — 213 chars of source]

where, in accordance with Theorems (ref) and (ref), the rate of convergence $\beta$ is $-\frac{1}{4}$ when $\sigma_\varepsilon \neq 0$, and $-\frac{1}{2}$ when $\sigma_\varepsilon = 0$.

Our goal is to adjust the posterior so that it centers at $\hat{\theta}_{n}$ and matches the asymptotic variance of $\hat{\theta}_{n}$. For the center, we simply shift the posterior by the right amount, while for the asymptotic variance, we adjust the temperature parameter. Concretely, define the estimator

align[align omitted — 89 chars of source]

The notation $\widehat {[J]}_n$ comes from the fact that this is a consistent estimator for the quadratic variation of the jump component $J$, because, as shown in the Appendix (see (ref) and (ref)), $\tilde{\theta}_n$ converges to $\theta^\dag = \theta^* + T^{-1}[J]$ and $\hat{\theta}_{n}$ is a consistent estimator of $\theta^{*}$ by construction. We will then adjust the location of the posterior by subtracting $T^{-1}\widehat {[J]}_n$ (this operation will necessarily center the posterior at $\tilde{\theta}_{n}-T^{-1}\widehat {[J]}_n=\hat{\theta}_{n}$). To adjust the variance, we adopt a sequence of the temperature parameters and its limit of the form:

align[align omitted — 127 chars of source]

where the quantities $\hat{V}_n$ and $\hat{V}_{asy,n}$ are suitable consistent estimators of ${V}$ and ${V_{asy}}$, respectively. The choice of these estimators will be specified below in \S (ref)-\S (ref).

Finally, we can define the adjusted misspecified posterior $\widetilde{\Pi}_n $ as one having the density function

equation[equation omitted — 120 chars of source]

where $\pi_n$ is the misspecified posterior obtained in Theorems (ref) and (ref) with $\kappa_{n}$ and $\kappa^\dagger$ defined in (ref). Asymptotic normality of the adjusted posterior is established by the following result.

thmWith the same conditions as in Theorem (ref) or Theorem (ref) except for the temperature parameter $\kappa_n$ defined as in (ref), the adjusted posterior $\widetilde{\Pi}_n$ defined above can be approximated by a normal distribution in the sense that, \begin{equation} TV\left(\widetilde{\Pi}_n, \ \mathcal{N}\left(\hat{\theta}_n, V n^{2 \beta} \right) \right) \stackrel{P_*}{\rightarrow} 0\quad as n\rightarrow \infty. \end{equation}

A location shift in Theorem (ref) or Theorem (ref) with $\kappa_n$ defined in (ref) gives us the proof of Theorem (ref).

This theorem illustrates that any type of $1-\alpha$ credible interval ($CI_{B,\alpha}$) of $\widetilde{\Pi}_n$ is asymptotically the same as a $1-\alpha$ confidence interval for $\theta^{*}$ based on $\mathcal{N}(\hat{\theta}_n, Vn^{2\beta} )$. The upper and lower bounds of the $CI_{B,\alpha}$ can then be approximated by $\hat{\theta}_n \pm \sqrt{Vn^{2\beta}} z_{\alpha/2}$ as $n\rightarrow\infty$, where $z_{\alpha/2}$ is the $\alpha/2$ quantile of the standard normal distribution. Because $\hat{\theta}_n$ satisfies a central limit theorem with asymptotic variance $V$, we have that $$ P_* ( \theta\in CI_{B,\alpha}) \approx P_* ( \theta\in \hat{\theta}_n \pm \sqrt{Vn^{2\beta}} z_{\alpha/2}) = P_* ( \hat{\theta}_n \in \theta\pm \sqrt{Vn^{2\beta}} z_{\alpha/2}) \approx 1-\alpha.$$ Therefore, the $1-\alpha$ credible interval has approximately the correct repeated sampling coverage under $P_*$, which indicates frequentist validation of the Bayesian inference based on the adjusted posterior.

Correction for a model without microstructure noise

When the variance $\sigma^{2}_{\varepsilon}$ of the noise is $0$, we can use the thresholded realized quadratic variation of mancini2009:

align[align omitted — 118 chars of source]

where $\eta_n$ is a threshold proportional to $n^{-w}$ for some suitable exponent $w$. Consistency of $\hat{\theta}_{n}$ is established in mancini2009 for any $w\in(0,1/2)$ when $J$ consists of the superposition of a general finite-jump activity process and an independent L\'evy process. cont2011 showed that $\hat{\theta}_n$ satisfies a central limit theorem with asymptotic variance $2 \theta^{* 2}n^{-1}$ under Assumption (ref) provided that $w\in\left(\frac{1}{4-2\alpha}, \frac{1}{2}\right)$. The existence of $w$ is guaranteed because $\alpha<1$ and, hence, $\frac{1}{4-2\alpha}< \frac{1}{2}$.

With the estimator $\hat{\theta}_{n}$ described above, we apply Theorem (ref) with $\beta = -1/2$, $V = 2 \theta^{* 2}$, and the temperature parameters taken as

align[align omitted — 184 chars of source]

By Slutsky's Theorem, it is clear that $\kappa_{n} \rightarrow \kappa^\dag$ in $P_*$-probability. We then obtain the following.

corUsing the same conditions as in Theorem (ref) except for the temperature parameter $\kappa_n$ defined as in ((ref)), and assume (ref), the adjusted posterior $\widetilde{\Pi}^n$ with density ((ref)) can be approximated by a normal distribution in the sense that, $$ TV\left(\widetilde{\Pi}_n, \ \mathcal{N}(\hat{\theta}_n , 2 \theta^{* 2}n^{-1} ) \right) \stackrel{P_*}{\rightarrow} 0\quad \text{ as }n\rightarrow \infty.$$
remarkAs we will show in Section (ref) below, the result above also holds for stochastic volatility models and more general pure-jump processes $J$.

Correction for the general model

When the variance of the noise is positive, one possible solution is to adopt the estimator $\hat{\Sigma}_n$ proposed in Jing2014jumpest, {which} combines the thresholding approach of mancini2009 with the pre-averaging method of JACOD2009preaveraging (see also JacodAndProtterDiscret for a detailed exposition of the theory). The pre-averaging method is used to {mitigate} the effect of the noise $\varepsilon$. Utilizing this method, we formulate several overlapping blocks of increments, and calculate proxies of the increments of the uncontaminated process $X$ by taking the weighted average of the increments of $Y$ within {each} block. Then, the estimator is defined as the sum of the squares of those new quasi-increments that are less than some threshold, and is further debiased using {an} estimator of the variance of the noise. This estimator meets our requirements, when we include both infinitely many jumps with bounded variation and normally distributed microstructure noise. For completeness, we describe the key aspects of this estimator below.

The estimator depends on two parameters: the length of the block $k_n$ and the weight function $g$. The latter satisfies the following regularity conditions:

itemize$g$ is continuous on $[0,1]$, piecewise $C^1$ with a piecewise Lipschitz derivative $g'$, and • $g(0) = g(1) = 0$, and $\bar{g} = \int_0^1 g^2(s)\,ds<\infty$.

One simple and common choice is $g(s) = s\wedge (1-s)$. Next, for some constant $c$, let $k_n = \lfloor cn^{1/2}\rfloor$ (the notation $\lfloor a\rfloor$ defines the largest interger that is smaller than $a$), $c_1 = c\bar{g}$, $c_2 = \int_0^1 (g'(s))^2\,ds/c$, and also define

align*[align* omitted — 286 chars of source]

where we recall that $\hat{\sigma}_\varepsilon^2 =\frac{1}{2n} \sum_{j=1}^n \Delta_j Y^2$ and {the} threshold $u_n$ satisfies $$ u_nn^{w_1}\rightarrow 0, \; u_nn^{w_2}\rightarrow \infty, \; \text{ as } n\rightarrow\infty,$$ for some $ 0\le w_1< w_2<{1}/{4} \text{ and } w_1 >1/(8-4\beta).$ The estimator is consistent and admits a central limit theorem. More specifically, by Theorems 1 and 3 in Jing2014jumpest, (ref) holds with $\hat{\theta}_n = \hat{\Sigma}_n$, $\beta = -1/4$, and \[ V = V_{noise}:= \frac{c}{\bar{g}^2}\left[ 4\theta^2 \Phi_{22} + \frac{2\theta\sigma_\varepsilon^2}{c^2}\Phi_{12} + \frac{\sigma_\varepsilon^4}{c^4}\Phi_{11} \right], \] where $\Phi_{ij} = \int_0^1 \phi_i(x)\phi_j(x)\,dx$, $\phi_1 = \int_x^1 g'(y)g'(y-x)\,dy$, and $\phi_2(x) = \int_x^1 g(y)g(y-x)\,dy$.

The temperature parameters in (ref) can be defined as

align[align omitted — 367 chars of source]

The convergence of $\kappa_n$ to $\kappa^\dag$ can be established through the consistency of $\hat{\Sigma}_n$ and $\hat{\sigma}_\varepsilon^2$ for $\theta^*$ and $\sigma_\varepsilon^2$, respectively, as well as the property that when $X_n = O_{P_*}(1)$ and $Y_n \stackrel{P_*}{\rightarrow} 0$, then $X_n Y_n \stackrel{P_*}{\rightarrow} 0$.

Then, we have the following corollary of Theorem (ref).

corWith the same conditions as in Theorem (ref) and with the temperature parameter $\kappa_n$ defined as in (ref), the adjusted posterior $\widetilde{\Pi}_n$ defined above can be approximated by a normal distribution in the sense that, $$ TV\left(\widetilde{\Pi}_n, \ \mathcal{N}\left(\hat{\Sigma}_n, \frac{c}{\bar{g}^2}\left[ 4\theta^2 \Phi_{22} + \frac{2\theta\sigma_\varepsilon^2}{c^2}\Phi_{12} + \frac{\sigma_\varepsilon^4}{c^4}\Phi_{11} \right] n^{-1/2} \right) \right) \stackrel{P_*}{\rightarrow} 0,\ \text{ as }n\rightarrow \infty.$$

Extension to more general semimartingales without noise

Thus far, we have assumed constant parameters for both the drift and diffusion components and a L\'evy process for the jump component $J$. In this section, we show that, in fact, when the microstructure noise can be ignored, the purposely misspecified posterior approach can also be applied to stochastic volatility models and more general jump processes $J$. As mentioned before, it is generally believe that the microstructure noise is relatively negligible when using medium range frequencies such as 5-minute or daily observations.

We consider the model

align[align omitted — 97 chars of source]

where $W$ is a Wiener process, $J$ is a suitable pure-jump semimartingale, and $\beta = \{\beta_t\}_{t\ge 0}$ and $\sigma = \{\sigma_t\}_{t\ge 0}$ are c\`{a}dl\`{a}g adapted processes. The parameter of interest is the scaled integrated variance

equation[equation omitted — 81 chars of source]

We again use the misspecified model (ref) for $X$ with $\varepsilon=0$. The corresponding log likelihood function would then be the same as in Remark (ref) with associated MLE

equation[equation omitted — 87 chars of source]

An analysis of the proof of Theorem (ref) reveals that the key for the result therein is the CLT stated in Lemma (ref). Specifically, what is needed is that the misspecified MLE ((ref)) converges to ((ref)) at the rate $O_{p}(n^{-1/2})$ (see Eq. ((ref)) in the proof). Jacod2008 (see Theorem 2.12 and Remark 2.13 therein) shows an analogous CLT to that of Lemma (ref) (with the same rate of convergence) under the more general setting ((ref)) when $\sigma$ and $J$ are of the form:

align*[align* omitted — 531 chars of source]

where $W'$ is a Wiener process independent of $W$ and $\mu$ is a Poisson random measure on $\mathbb{R}_{+}\times \mathbb{R}$ with predictable compensator $\nu(ds,dx)=dsdx$, independent of $(W,W')$. The coefficients of $\sigma$ and $J$ (including $\delta:\Omega\times \mathbb{R}_{+}\times \mathbb{R}\to\mathbb{R}\backslash\{0\}$ and $\tilde{\delta}:\Omega\times \mathbb{R}_{+}\times \mathbb{R}\to\mathbb{R}\backslash\{0\}$) are random processes satisfying standard conditions for the integrals therein to be well defined.

As explained in Section (ref), the step to correct the center and variance of the misspecified posterior $\Pi_{n}$ requires an estimator $\hat{\theta}_{n}$ of $\theta^{*}$ enjoying a CLT with a rate of $n^{-1/2}$. As it turns out, the thresholded realized quadratic variation of mancini2009, defined in ((ref)), does again the job at least in the case of bounded variation jump process $J$. Specifically, Jacod2008 (see Theorems 2.4 and 2.11 therein) obtains a feasible CLT for ((ref)) under the same framework as above, but with an additional condition on $J$ that amounts to $J$ having bounded variation paths.

When the microstructure noise is taken into account, the extension is not as direct as for the no noise case, because after applying an orthonormal transformation to remove the autocovariance introduced by the noise, similar to that at the beginning of Section (ref), the distribution of the transformed data does not depend anymore only on the target parameter $\theta^* = T^{-1}\int_0^T \sigma_t^2\,dt$. Instead, the variance of each transformed data depends on a weighted sum of the `volatility' of each increments. Then, analyzing the transformed data using the same procedure as before can only provide us an estimator of some value larger than the integrated volatility, but not about the exact parameter $\theta^{*}$.

Simulation

This section discusses the finite sample performance of the adjusted posterior defined in Theorem (ref). We aim to show the plausibility of the limit ((ref)) at large sample size. This is demonstrated through comparing the empirical coverage probability of the credible interval derived from the adjusted posterior and the confidence interval from its corresponding asymptotic normal distribution in the theorem. We also aim to compare the “purposely misspecified" method with the frequentist central limit theorem (CLT) (ref).

Infinite jump activity without noise

In order to incorporate infinite jump activity, the jump component is set be a variance gamma process

equation[equation omitted — 37 chars of source]

where $a = -0.2$, $b= 0.2$, $\{G_t\}_{t\geq{}0} $ is Gamma process such that $G_h\Gamma(\Delta_n/c,c)$, with $c=0.23$, and $\{B_{t}\}_{t\geq{}0}$ is an Wiener process independent of the Wiener process $W$. For the drift and diffusion components, let $ \mu = 0.1$ and $ \theta = 0.3$. All the parameters are taken from mancini2009. For simplicity, we adopt the widely-used threshold $\eta_n = n^{-w}$, where $w \in (0,0.5)$ and $n$ is the sample size. This is a possible and conventional choice in terms of consistency and efficiency. In the following simulation, we use $w = 0.39$.

For the prior of $\theta$, an inverse gamma distribution is applied with shape and scale both equal to one. Since the temperature parameters do not affect the conjugacy, the misspecified posterior and the adjusted posterior both follow inverse gamma distribution.

Single sample path

First of all, 5000 equally spaced observations are simulated based on the parameters defined above (sample size $n=5000$). The adjusted posterior $\widetilde{\Pi}_n$ is generated using Corollary (ref). The results are shown in Figure (ref). The adjusted posterior for one possible sample path is plotted as the dashed line and compared with the corresponding asymptotic normal distribution $\mathcal{N}(\hat{\theta}_n, 2\theta^{*2}n^{-1})$ (the solid line). These two lines can hardly be distinguished from each other. Moreover, they are both roughly centered at the true volatility 0.3. This true volatility also lies between the dashed vertical lines which mark out the 95% highest posterior density (HPD) interval of the adjusted posterior.

SCfigure[][ht] \caption{Comparison of adjusted posterior and asymptotic normal distribution for one sample path. The solid line represents the asymptotic normal distribution in Corollary (ref). The dashed line is the adjusted posterior. 95% HPD interval lies between the two black dashed lines. }

It turns out that the adjusted posterior recovers the asymptotic normal distribution, which proves the validity of Corollary (ref). This suggests that the adjusted posterior will be centered at an efficient estimator with optimal variance when the sample size is large enough.

Point estimators

In the second step, we evaluate the consistency of the point estimators. The biases of the means of two distributions defined in Corollary (ref) are compared: the mean of the adjusted posterior $\widetilde{\Pi}_n$, and the mean of the asymptotic normal distribution $\hat{\theta}_n=\tilde{\theta}_n - T^{-1}\widehat{[J]}_n$, which is also the threshold estimator in mancini2009. We also consider the misspecified posterior adjusted by the latent realized quadratic variation of the jump component $T^{-1}{[J]}_n$ instead of $ T^{-1}\widehat{[J]}_n$. The corresponding asymptotic normal distribution has mean $\hat{\theta}_n^*=\tilde{\theta}_n - T^{-1}{[J]}_n$. The analysis of these four point estimators is based on 1000 simulations. For each simulation, 5000 equally spaced observations are generated and used to calculate the biases.

The distribution of the biases is plotted in Figure (ref). The solid line is formed by the biases of the threshold estimator, while the dashed line is formed by the biases of the mean of the adjusted posterior $\widetilde{\Pi}_n$. The dotted line represents the bias of $\hat{\theta}_n^*$. The bias of the mean of the adjusted posterior using the realized quadratic variation is represented by the dashed-dotted line.

SCfigure[][ht] \caption{Bias of point estimators. The biases of the mean of the asymptotic normal distribution in Corollary (ref) forms the solid lines. While the dashed line is the distribution of the mean of the adjusted-posterior. The dotted and the dashed-dotted lines represent the distributions of the means of asymptotic normal distribution and posterior in Theorem (ref) with location shift equals the realized quadratic variation of the jumps. }

The similarity of the solid and the dashed lines as well as the similarity of the dotted and the dashed-dotted lines suggest that the posterior mean and the mean of the asymptotic normal distribution have similar behavior in terms of their difference with the true volatility. The biases are relatively small since the volatility is 0.3 while most of the biases are within 0.01 range.

remarkWe may increase the accuracy of the adjusted posterior $\widetilde{\Pi}_n$ by using a better estimator of the quadratic variation of the jump $[J]$ to correct the misspecified posterior $\Pi_n$ defined in Theorem (ref). While the dashed and the solid lines have higher probability for the positive values, the dotted and the dashed-dotted lines are more symmetric. This suggests that the right-skewed tendency of our posterior mean might be because of the poor estimates for the jump component. The better the estimation of the $[J]$ we applied, the closer distribution to the symmetric dashed-dotted line we will get. One approach is to optimize the threshold $\eta_{n}$ in the threshold parameter $\hat{\theta}_n$.

Comparison based on confidence interval

In order to evaluate the accuracy of the inference, we compare the empirical coverage probability of the credible interval of the posterior with the confidence intervals of the asymptotic distribution and of the CLT based on the threshold estimator. For simplicity, in this section, we use “CI" to represent both the credible interval and the confidence interval. We increase the sample size $n$ to 105000, which is approximately the number of the stock data obtained within one year with 5-minute interval. The coverage probabilities of the 95% CIs based on 1000 repeats are listed in Table (ref). For each repeat, we simulate a sample path with 105000 observations.

table[table omitted — 659 chars of source]

Therefore, it can be concluded that the HPD interval has the highest coverage probability among all the CIs derived from various distributions defined above. However, all the coverage probabilities are slightly less than $0.95$. We may fill this gap by adopting a more refined frequentist estimator $\hat{\theta}_{n}$ to serve as the center of the posterior or by utilizing a more accurate misspecified model.

Except the conjugate prior, an non-informative prior, the uniform distribution, and an exponential distribution are also for the simulation based on the same model. The results are compatible with the inverse-gamma prior.

L\'evy Model with microstructure noise

In order to illustrate the results in presence of both infinitely many jumps and microstructure noise, we conduct simulation for the following model from Jing2014jumpest: $$ X_t = W_t + J_t, \quad Y_{i/n} = X_{i/n} + \varepsilon_{i/n}, \quad \varepsilon_{i/n} \sim \mathcal{N}(0, 0.01^2), $$ for $i = 1,2,\ldots, n$. The jump part $J_t$ is a trimmed symmetric $\beta$-stable process with $\beta = 0.5$. The trimmed process means that after we simulated the increments of all the jumps, the largest $2\%$ of them (ranked by absolute values) will be discard to match the behaviour for high-frequency tick-by-tick data. To allow a comparison, simulation is conducted based on exactly the same parameters described in the paper except one constant $c$, which determines the length of the preaveraging blocks by $ k_n = \lfloor c\Delta_n^{-1/2} \rfloor.$ The choice of $c$ is not clearly stated in the paper. Thus, we choose the same $c=1/3$ as in the original work JACOD2009preaveraging. The sample size is set as $n=15600$ and $\Delta_n = 1/7800$ taken from Jing2014jumpest.

For the adjusted posterior, the data is divided into two parts. The first half is used to evaluate the estimator $\hat{\Sigma}_n$ in Jing2014jumpest, which is used in the prior, and the rest is used to make inference. The prior is chosen to be a truncated normal distribution with lower boundary $0$, centered at $\hat{\Sigma}_n$ obtained using the first half of the data, and standard deviation $0.06$. We generate 20000 MCMC samples, in which the first 5000 are burned. We generate 1000 times posteriors, each based on 20000 MCMC samples. The results can be summarized as follows.

table[table omitted — 218 chars of source]

For the comparison of the point estimators, we analyse the mean bias and the standard errors of the frequentist estimator $\hat{\Sigma}_n$ and the maximum a posterior point estimator (MAP) defined as $$ \hat{\theta}_{MAP} = \arg\max_{\theta} \widetilde{\Pi}_n(\theta).$$ Both the mean bias and the standard errors are similar and small, suggesting the estimation accuracy of the two approaches. For inference strength, the coverage probability of the $95\%$ credible interval of the adjusted posterior is slightly better than the confidence interval derived from the CLT of the estimator $\hat{\Sigma}_n$.

table[table omitted — 377 chars of source]
figure[figure omitted — 877 chars of source]

Conclusion

In this paper, we consider an infinite activity model with microstructure noise over a fixed time horizon. A “purposely misspecified" posterior is proposed for the volatility, the variation of the diffusion component. We theoretically and empirically prove that the posterior can be approximated by a normal distribution centered at a suitable estimator with the optimal variance. Thus, valuable inference can be developed based on the Bayesian credible intervals, whose empirical coverage probability is shown to be close to the nominal one by simulation.

Compared to Martin, we generalize the feature of finite many jumps to infinite jumps, propose an extension to handle stochastic volatility and general It\^o jump processes of bounded variation, and add the microstructure noise. These results suggest more possibilities of further use of the method.

The purposely misspecified method contributes to the Bayesian framework by ignoring the infinite-dimensional nuisance parameter and directly making inference of the parameter of interest. It provides us a Bayesian posterior without any requirement for assigning a prior and specifying the likelihood of the complicated jump part. Thus, the difficulties of deriving the full posterior and obtaining a marginalized posterior can also be avoided.

The recentering and rescaling procedure is also highly flexible. Any consistent and efficient estimator can be used as a correction. Furthermore, the variance can be adjusted in response to the information. For example, when the variance of the volatility $\theta^*$ is foreknown, the temperature parameter can be set to the correct variance over the optimal one.

Considering the good performance of the misspecified model and the complexity of the nonparametric nuisance part, it may be noteworthy to apply the “purposely misspecified" method to other semiparametric problems.

thebibliography{39} \bibitem{Barndorff-NielsenKernel} Barndorff-Nielsen, O., Hansen, P., Lunde, A. and Shephard, N. (2008). Designing Realised Kernels to Measure the Ex-Post Variation of Equity Prices in the Presence of Noise. Stochastic Processes and their Applications. 76(6) 1481 - 1536. \bibitem{levyprocessbook} Barndorff-Nielsen, O. E., Mikosch, T. and Resnick, S. (2001). L\'{e}vy Processes: Theory and Applications.. {Birkhauser}. \bibitem{infjump} \textsc{Carr, P. and Geman, H. and Madan, D. and Yor, M. } (2002). \textit{The Fine Structure of Asset Returns: An Empirical Investigation}. \textit{The Journal of Business}. \textbf{75}(2) 305-332. \bibitem{Jones99bayesianestimation} \textsc{Christopher S. J.} (1999). \textit{{Bayesian} estimation of continuous-time finance models. {Working} Paper}. \textit{Rochester University}. Working Paper. \bibitem{CHRISTENSEN2010} \textsc{Christensen, K., Oomen, R. and Podolskij, M.} (2010). \textit{Realised quantile-based estimation of the integrated variance}. \textit{Journal of Econometrics}. \textbf{159}(1) 74 - 98. \bibitem{Cont:2003} \textsc{Cont, R. and Tankov, P.} (2004). \textit{{Financial modelling with Jump Processes}}. \textit{{Chapman & Hall/CRC Finance}}. Math. Ser., Chapman & Hall/CRC, Boca Raton, FL. \bibitem{cont2011} \textsc{Cont, R. and Mancini, C.} (2011). \textit{Nonparametric tests for pathwise properties of semimartingales}. \textit{Bernoulli}. \textbf{17}(2) 781-813. \bibitem{laplace} \textsc{de Bruijn, N. G. } (1961). \textit{Asymptotic Methods in Analysis}. {North Holland, Amsterdam}. \bibitem{Eraker2003impact} \textsc{Eraker, B., Johannes, M. S. and Polson, N.} (2003). \textit{The Impact of Jumps in Volatility and Returns}. \textit{Journal of Finance}. \textbf{58} 1269-1300. \bibitem{FIGUEROA2008jumpmoment} \textsc{Figueroa-L\'{o}pez, J. E.} (2008). \textit{Small-time moment asymptotics for L\'{e}vy processes}. \textit{Statistics & Probability Letters}. \textbf{78}(18) 3355 - 3365. \bibitem{FLKiseop2015} \textsc{Figueroa-L\'opez, J. E. and Lee, K.} (2017). \textit{Estimation of a noisy subordinated Brownian Motion via two-scale power variations}. \textit{Journal of Statistical Planning and Inference}. \textbf{189} 16-37. \bibitem{lan} \textsc{Gloter, A. and Jacod, J.} (2001). \textit{Diffusions with measurement errors. I. Local Asymptotic Normality}. \textit{ESAIM: Probability and Statistics}. \textbf{5} 225–242. \bibitem{gloter_jacod_2001_est} \textsc{Gloter, A. and Jacod, J.} (2001). \textit{Diffusions with measurement errors. II. Optimal estimators}. \textit{ESAIM: Probability and Statistics}. \textbf{5} 243–260. \bibitem{Griffin2016} \textsc{Griffin, J. E.} (2016). \textit{Flexibly Modelling Volatility and Jumps Using Realised and Bi-Power Variation}. University of Kent, School of Mathematics, Statistics and Actuarial Science. \bibitem{hansen2006moderatefreq} \textsc{Hansen, P. and Lunde, A.} (2006). \textit{Realized Variance and Market Microstructure Noise}. \textit{Journal of Business & Economic Statistics}. \textbf{24}(2) 127-161. \bibitem{he2014efficient} \textsc{He, C. and Wang, Z.} (2014). \textit{Efficient Estimation of Stochastic Diffusion Models with Leverage Effects and {L\'{e}vy} Jumps}. \textit{Journal of Information & Computational Science}. \textbf{11}(2) 367–382. \bibitem{Jacod2007} \textsc{Jacod J.} (2007). \textit{Asymptotic properties of power variations of L\'evy processes}. \textit{ESAIM: Probability and Statistics}. \textbf{11} 173-196. \bibitem{Jacod2008} \textsc{Jacod J.} (2008). \textit{Asymptotic properties of realized power variations and related functionals of semimartingales}. \textit{Stochastic Processes and Their Applications}. \textbf{118} 517-559. \bibitem{JACOD2009preaveraging} \textsc{Jacod, J., Li, Y, Mykland, P., Podolskij, M. and Vetter M.} (2009). \textit{Microstructure noise in the continuous case: The pre-averaging approach}. \textit{Stochastic Processes and their Applications}. \textbf{119}(7) 2249 - 2276. \bibitem{JacodAndProtterDiscret} \textsc{Jacod, J. and Protter, P.E.} (2012). \textit{Discretization of Processes}. \textit{Stochastic Modelling and Applied Probability}. \textbf{67}. \bibitem{JacodSum} \textsc{Jacod, J. and Todorov, V.} (2014). \textit{Efficient estimation of integrated volatility in presence of infinite variation jumps}. \textit{The Annals of Statistics}. \textbf{42}(3) 1029--1069. \bibitem{jasra2011a} \textsc{Jasra, A., Stephens, D. , Doucet, A, and Tsagaris, T.} (2011). \textit{Inference for {L\'{e}vy}-Driven Stochastic Volatility Models via Adaptive Sequential {Monte Carlo}}. \textit{Scandinavian Journal Of Statistics}. \textbf{38}(1) 1-22. \bibitem{jiang2008gibbs} \textsc{Jiang, W. and Tanner, M.} (2008). \textit{{Gibbs} posterior for variable selection in high-dimensional classification and data mining}. \textit{The Annals of Statistics}. \textbf{36}(10) 2207-2231. \bibitem{Jing2014jumpest} \textsc{Jing,B., Liu, Z. and Kong,X. } (2014). \textit{On the Estimation of Integrated Volatility With Jumps and Microstructure Noise}. \textit{Journal of Business & Economic Statistics}. \textbf{32}(3) 457-467. \bibitem{BVM} \textsc{Kleijn, B.J.K. and van der Vaart, A.W.} (2012). \textit{The {Bernstein-von Mises} theorem under misspecification}. \textit{Electronic Journal of Statistics}. \textbf{6} 354--381. \bibitem{Kou2017} \textsc{Kou, S. , Yu, C. and Zhong, H.} (2017). \textit{Jumps in Equity Index Returns Before and During the Recent Financial Crisis: A {Bayesian} Analysis}. \textit{Management Science}. \textbf{63}(4) 988-1010. \bibitem{Li2008Bayesian} \textsc{Li, H and Wells, M. and Yu, C. } (2008). \textit{A {Bayesian} Analysis of Return Dynamics with {L\'{e}vy} Jumps}. \textit{The Review of Financial Studies}. \textbf{21}(5) 2345-2378. \bibitem{mancini2009} \textsc{Mancini, C.} (2009). \textit{Non-parametric Threshold Estimation for Models with Stochastic Diffusion Coefficient and Jumps}. \textit{Scandinavian Journal of Statistics}. \textbf{36} 270-296. \bibitem{Martin} \textsc{Martin, R., Ouyang, C. and Domagni, F.} (2018). \textit{‘{Purposely misspecified}’ posterior inference on the volatility of a jump diffusion process}. \textit{Statistics & Probability Letters}. \textbf{134} 106-113. \bibitem{PODOLSKIJ2009IA} \textsc{Podolskij, M. and Vetter, M.} (2009). \textit{Bipower-type estimation in a noisy diffusion setting}. \textit{Stochastic Processes and their Applications}. \textbf{119}(9) 2803 - 2831. \bibitem{podolskij2009} \textsc{Podolskij, M. and Vetter, M.} (2009). \textit{Estimation of volatility functionals in the simultaneous presence of microstructure noise and jumps}. \textit{Bernoulli}. \textbf{15}(3) 634--658. \bibitem{Szerszen2009Baye} \textsc{Szerszen, P.} (2009). \textit{{Bayesian} analysis of stochastic volatility models with {L\'{e}vy} jumps: Application to risk analysis}. \textit{Board of Governors of the Federal Reserve System}. 2009-40. \bibitem{peterintro} \textsc{Tankov, P.} (2007). \textit{{L\'{e}vy} Processes in Finance and Risk Management}. \textit{Wilmott Magazine}. Sept-Oct, 89-97. \bibitem{xiu2010} \textsc{Xiu, D.} (2010). \textit{Quasi-maximum likelihood estimation of volatility with high frequency data}. \textit{Journal of Econometrics}. \textbf{159}(1) 235 - 250. \bibitem{Yang2017jump} \textsc{Yang, H. and Kanniainen, J.} (2017). \textit{Jump and Volatility Dynamics for the {S&P} 500: Evidence for Infinite-Activity Jumps with Non-Affine Volatility Dynamics from Stock and Option Markets}. \textit{Review of Finance}. \textbf{21}(2) 811-844. \bibitem{Yu2011MCMC} \textsc{Yu, C. and Li, H. and Wells, M.} (2011). \textit{{MCMC} Estimation of {L\'{e}vy} Jump Models Using Stock and Option Prices}. \textit{Mathematical Finance}. \textbf{21}(3) 383-422. \bibitem{TSRV} \textsc{Zhang, L. and Mykland, P. and A$\ddot{i}$t-Sahalia, Y.} (2005). \textit{A Tale of Two Time Scales: Determining Integrated Volatility with Noisy High-Frequency Data}. \textit{Journal of the American Statistical Association}. \textbf{100}(472) {1394-1411}. \bibitem{Zhang2006gibbs} \textsc{Zhang, T.} (2006). \textit{Information-theoretic upper and lower bounds for statistical estimation}. \textit{IEEE Transactions on Information Theory}. \textbf{52}(4) 1307-1321. \bibitem{zhou1996noise} \textsc{Zhou, B.} (1996). \textit{High-Frequency Data and Volatility in Foreign-Exchange Rates}. \textit{Journal of Business & Economic Statistics}. \textbf{14}(1) 45-52