EconBase
← Back to paper

Bias optimal vol-of-vol estimation: the role of window overlapping

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

90,785 characters · 15 sections · 96 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Bias-Optimal Vol-of-Vol Estimation: The Role of Window Overlapping

abstractThe simplest and most natural vol-of-vol estimator, the pre-estimated spot variance based realized variance (PSRV), is typically plagued by a large finite-sample bias. In this paper, we analytically show that allowing for the overlap of consecutive local windows to pre-estimate the spot variance may correct for this bias. In particular, we provide a feasible rule for the bias-optimal selection of the length of local windows when the volatility is a CKLS process. The effectiveness of this rule for practical applications is supported by numerical and empirical analyses.

{JEL codes:} C13, C14, C58, G10

Introduction

Estimating the volatility of asset volatility (hereinafter vol-of-vol) is relevant in many areas of mathematical finance, such as the calibration of stochastic volatility of volatility models (bnv, scm), the hedging of portfolios against volatility of volatility risk (hsst), the estimation of the leverage effect (kx, aflwy), and the inference of future returns (btz), along with spot volatilities (mz).

The literature offers a number of consistent estimators for the integrated vol-of-vol. The first estimator to appear was the one proposed by bnv, termed pre-estimated spot-variance based realized variance (PSRV), which is, in fact, simply the realized variance of the unobservable spot variance, computed using estimates of the latter. Later, v derived two sophisticated versions of the simple PSRV: one that allows for a central limit theorem with the optimal rate of convergence, but also for negative values, and another that preserves positivity at the expense of a slower rate of convergence. Note that the simple PSRV and its sophisticated versions are consistent when the price and volatility processes are continuous semimartingales, in the absence of microstructure noise contaminations. Further, Fourier-based estimators of the integrated vol-of-vol were introduced by scm and ct. In particular, the estimator by scm is asymptotically unbiased in the presence of market microstructure noise, while the estimator by ct allows for a central limit theorem in the presence of jumps in the price and volatility processes.

The numerical studies in aflwy and scm show that both realized and Fourier-based integrated vol-of-vol estimators may carry a substantial finite-sample bias unless the selection of the tuning parameters involved in their computation is carefully optimized. However, this is a rather unexplored issue, which we aim to explore. To do so, we focus on the simple PSRV, since it represents the most intuitive and easy-to-implement vol-of-vol estimator. Furthermore, asymptotically-optimal estimators do not necessarily guarantee the best finite-sample performance, as pointed out in the extensive study by go on integrated volatility estimators and confirmed for integrated vol-of-vol estimators by the numerical studies in aflwy and scm. Thus, there is no reason to expect a priori that the simple PSRV would show worse finite-sample performance than its sophisticated version with optimal rate of convergence.

As mentioned, the PSRV is the realized volatility of the unobservable spot volatility process, computed from discrete estimates of the latter. In other words, the PSRV is the sum of the squared increments of estimates of the unobservable spot volatility on a discrete grid. These estimates are obtained as local averages of the price realized volatility. Formally, the locally averaged realized variance and the PSRV are defined as follows.

definition{Locally averaged realized variance} Suppose that the log-price process $p$ is observable on an equally-spaced grid of mesh size $\delta_N$, with $\delta_N \to 0$ as $N \to \infty$. Also, let $k_N=O(\delta_N^b)$, $b \in (-1,0),$ be a sequence of positive integers such that $k_N \to \infty$ and define the local window $W_N:=k_N\delta_N$ such that $W_N \to 0$ as $N \to \infty$. The locally averaged realized variance at time $t$ is defined as $$\hat \nu_N(t) := \displaystyle \frac{1}{k_N \delta_N} \sum_{j=1}^{k_N} \Big[ p(\lfloor t/ \delta_N\rfloor \delta_N - k_N \delta_N + j\delta_N) - p(\lfloor t/ \delta_N\rfloor \delta_N - k_N \delta_N + (j-1) \delta_N) \Big]^2,$$ where $\lfloor\,\cdot\,\rfloor$ denotes the floor function.
definition{Pre-estimated spot-variance based realized variance} Suppose that the log-price process $p$ is observable on an equally-spaced grid of mesh size $\delta_N$, with $\delta_N \to 0$ as $N \to \infty$. The pre-estimated spot-variance based realized variance (PSRV) on the interval $[\tau, \tau + h]$ is defined as $$PSRV_{[\tau, \tau + h],N} := \sum_{i=1}^{\lfloor h/\Delta_N \rfloor} \Big[ \hat \nu_N(\tau+i \Delta_N) - \hat \nu_N(\tau+(i-1) \Delta_N) \Big]^2,$$ where: \begin{itemize} • $\hat\nu_N(\cdot) $ is the locally averaged realized variance in Definition (ref), with $k_N=O(\delta_N^b)$, $b \in (-1,0)$; • $\Delta_N= O(\delta_N^c)$, $c \in (0,1)$, is the locally averaged realized variance sampling frequency. \end{itemize}

The following propositions summarize the asymptotic properties of the locally averaged realized variance and the PSRV. For further details, see Chapter 8 inaj.

propositionLet the log-price process $p$ be a continuous semimartingale and let the process $\nu$ denote its instantaneous volatility. Then $\hat \nu_N(t)$ is a consistent local estimator of $\nu(t)$ as $N \to \infty$.
proofSee Chapter 8.1 in aj.
propositionLet the log-price process $p$ and the spot volatility process $\nu$ be continuous semimartingales. Then the PSRV is a consistent estimator of the quadratic variation of the volatility process $\langle \nu, \nu \rangle_{[\tau,\tau +h]}$ if $b \in (-1/2,0)$ and $ c \in (0, -b/2)$.
proofSee Proposition 8 in bnv.
remarkNote that the requirements for rates $b$ and $c$ that guarantee consistency imply that $\frac{W_N}{\Delta_N} \to 0$ as $N \to \infty$. Indeed, as one can easily verify, $-1/2<b<0$ and $0<c<-b/2$ imply $ c < 1+b, $ which, in turn, implies $ \frac{W_N}{\Delta_N} \to 0$ as $N \to \infty$.

In practical applications, when computing PSRV values, one has to select the spot volatility estimation grid. Moreover, since the spot volatility is estimated as an average of the price realized volatility over a local window, the length of the latter must also be selected. More specifically, the figure below details the different quantities involved in the computation of the PSRV: the time horizon $h$; the log-price sampling frequency $\delta_N:=\frac{h}{N}$; the spot volatility sampling frequency $\Delta_N:=\lambda_N\delta_N$, $\lambda_N= min(N, \lceil \lambda \delta_N^{c-1}\rceil),$ $\lambda >0$, $c \in (0,1)$; the size of the local window to estimate the spot volatility $W_N=k_N\delta_N$, $k_N= \lceil \kappa \delta_N^b \rceil$, $\kappa >0$, $b \in (-1,0)$; and the spot volatility estimates $\hat\nu(s)$, $s=\tau+j\Delta_N$, $j=0,1,...,\lfloor h / \Delta_N \rfloor$. Note that $\lceil\,\cdot\,\rceil$ denotes the ceiling function. \\

figure[figure omitted — 1,123 chars of source]

As a consequence, for given values of the asymptotic rates $b$ and $c$, the finite-sample performance of the PSRV (i.e., the performance of the PSRV for a fixed $N$) depends on the selection of two tuning parameters: $\lambda$, which determines the mesh of the spot volatility estimation grid and $\kappa$, which determines the length of the local window used to estimate the spot volatility.

With regard to the selection of $\kappa$, note that the efficient computation of the spot volatility in finite samples may require the selection of fairly long local windows (see, e.g., LeeM, afl and ZB). This in turn suggests that the finite-sample efficient implementation of the PSRV over a given period (e.g., one day) may require the use of price observations from the previous period(s) (e.g., day(s)). At the same time, this might imply that it is optimal to allow consecutive local windows to overlap in finite samples, that is, $W_N>\Delta_N$ for $N$ finite. This aspect is confirmed by the numerical study in scm, which shows that it is optimal to select $\kappa$ such that $W_N>\Delta_N$ in finite samples. The aim of this paper is to gain insight into the bias-reducing effect due to window-overlapping from an analytic perspective. To do so, we follow an approach inspired by the one used in afl to solve the “leverage effect puzzle".

The “leverage effect puzzle" pertains to the absence of correlation between log-price and (estimated) volatility changes at high-frequencies, observed in empirical samples. afl solve this puzzle by showing analytically that a substantial bias masks the presence of correlation unless log-price and volatility estimates changes are computed on a suitably sparse grid. The aim is not to solve the problem of the efficient non-parametric estimation of the leverage at high-frequencies, but rather to obtain insight into the puzzle by solving it in a widely used parametric setting that allows for fully explicit computations. This paper is written in the same spirit. In fact, we do not address the general problem of the efficient non-parametric vol-of-vol estimation from high-frequency prices, but, rather, our aim is to obtain insight from an analytic perspective into why the PSRV, the simplest and most natural vol-of-vol estimator, is plagued by a large bias in finite samples and investigate the role of window-overlapping as a tool for reducing such large bias.

Outline of the paper

To achieve this aim, we proceed as follows. In Section (ref) we perform a preliminary numerical exercise that uncovers the crucial role of the local-window parameter $\kappa$ in determining the finite-sample performance of the PSRV and, at the same time, shows that the latter is basically insensitive to the selection of the grid parameter $\lambda$. In particular, it is evident from simulations that the PSRV finite-sample bias is optimized by selecting $\kappa$ such that consecutive local windows to estimate the spot volatility overlap. Numerical results of Section (ref) confirm those by scm and motivate the analytic study of Section (ref).

In Section (ref), we address the problem of the optimal selection of PSRV tuning parameters in finite samples from an analytic perspective. To do so, in the spirit of afl, we assume a widely-used parametric form for the data-generating process, which allows us to obtain the full explicit PSRV finite-sample bias expression. Specifically, we assume the price to be a continuous semimartingale and the volatility to be a CIR process (see cir). In general, independently of the parametric assumption on the data-generating process, the PSRV finite-sample bias expression differs in case window overlapping is allowed, i.e., when $W_N > \Delta_N$, or not, i.e., $W_N \le \Delta_N$. Consequently, in Section (ref) we study both cases.

In the no-overlapping case we adopt a conventional approach and isolate the dominant term of the bias as $N\to \infty$, thereby showing that a value of $\kappa$ that annihilates the dominant term of the bias does not always exist and, even when it does exist, its computation would be basically unfeasible in practice, as it depends on the drift parameters of the volatility, which cannot be reliably estimated on a fixed time horizon, due to the fact that their consistent estimation is possible only in the classic long-sample asymptotics setting (see, e.g., kankris). In addition, when the optimal value of $\kappa$ exists, for typical orders of magnitude of the CIR parameters it actually satisfies the no-overlapping constraint only at ultra-high frequencies ($< 1$ second).

In the overlapping case, instead, the natural expansion as $N \to \infty$ is precluded, as the consistency of the PSRV requires that consecutive windows do not overlap as the number of price observations grows to infinity (see Remark (ref)). Therefore, in this case we adopt a novel approach and expand sequentially the bias expression as the tuning parameter $\lambda$ and the time horizon $h$ go to zero, based on the fact that, in practical applications, $h$ and $\lambda$ are typically very close to zero. This approach yields a dominant term of the bias which is independent of the tuning parameter $\lambda$ and is annihilated by selecting the asymptotic rate of $k_N$ as $b=-1/2$, the asymptotic rate of $\Delta_N$ as $c<1/2$, and the local-window tuning parameter as $\kappa =2\sqrt{\nu(\tau)}{\gamma}^{-1}$, where $\nu(t) $ and $\gamma$ denote, respectively, the spot variance process at the initial time $\tau$ and the CIR diffusion parameter. This analytic result shows that, when overlapping is allowed, it is possible to select $\kappa$ such that the bias is effectively optimized in practical applications and supports the numerical evidence on the bias-reducing effect due to window overlapping collected in Section (ref) and in scm. However, this rule to select $\kappa$ is unfeasible unless reliable estimates of $\nu(\tau)$ and $\gamma$ are available. Accordingly, in Appendix B we detail a simple procedure to estimate $\nu(\tau)$ and $\gamma$.

In Section (ref) we also address the problem of the bias-optimal implementation of the PSRV in the more realistic situation where the price process is contaminated by an i.i.d. microstructure noise process at high frequencies. Again, we distinguish between the overlapping and no-overlapping case and derive, for each case, the exact parametric expression of the extra bias term due to microstructure noise. However, in both cases it emerges that this extra term depends not only on some moments of the noise process but also on the drift parameters of the volatility process, which cannot be consistently estimated over a fixed time horizon (see, e.g., kankris). This precludes the possibility of efficiently subtracting the bias due to noise in small samples. As a solution, we suggest sampling prices on a suitably sparse grid as in the seminal paper by abdh, so that the presence of noise becomes negligible and the bias optimal rule to select the local-window parameter $\kappa$ can still be applied. The efficiency of this solution is verified numerically in Section (ref) for typical values of the noise-to-signal ratio. For completeness, we also analyze the noise bias expression in the no-overlapping case. In particular, we exploit this expression to derive the asymptotic rate of divergence of the PSRV bias as $N$ tends to infinity.

Additionally, as a byproduct of the PSRV bias analysis, in Section (ref) we quantify the bias reduction following the assumption that the initial value of the volatility process is equal to the long-term volatility parameter, in the case of both the PSRV and the locally averaged realized variance. This is a very common assumption in the literature, typically made in simulation studies where a mean-reverting process drives the spot volatility (see, e.g., among many others, afl, scm, v).

In Section (ref) we use a heuristic approach based on dimensional analysis to generalize the rule for the selection of $\kappa$ to the case of a volatility process belonging to the CKLS class (see ckls). Specifically, we find that it is optimal, in terms of bias reduction, to select $ \kappa=2\frac{\nu(\tau)}{\sqrt{\xi(\tau)}}$, where $\nu(t)$ is the variance process and $\xi(t)$ is the variance-of-variance process, while $\tau$ is the initial time of the estimation horizon. Note that in the absence of price and volatility jumps (a condition required for the PSRV to be consistent), the semi-parametric stochastic volatility model where the price is a semimartingale and the volatility is a CKLS process represents a fairly flexible model. In fact, the CKLS framework encompasses a number of widely-used models for financial applications. Indeed, besides the CIR model, which determines the volatility dynamics in the popular Heston model (heston) and its generalized version with stochastic leverage by veraart2, the CKLS family includes, e.g., the model by brschw and the model by cox, which appear, respectively, in the continuous-time GARCH stochastic volatility model by nelson and 3/2 stochastic volatility model by 32.

In Section (ref) we perform an extensive numerical study where we test the performance of the feasible rule to select $\kappa$ derived in Section (ref) and generalized in Section (ref). The results confirm that this rule is effective in reducing the PSRV bias. We underline that this rule does not require the estimation of the drift parameters of the CIR process, which can not be consistently estimated on a fixed time horizon. Finally, in Section (ref) we illustrate the results of an empirical study, in which we compute PSRV values from high-frequency S&P 500 prices, selecting $\kappa$ based on the bias-optimal rule. Section (ref) summarizes our conclusions. Finally, Appendix A contains the proofs and Appendix B illustrates the feasible procedure that we propose to select $\kappa$ from sample prices.

Preliminary results

The finite-sample accuracy of the PSRV requires the careful selection of the tuning parameters $\kappa$ and $\lambda$. In this section we gain some preliminary insight into this issue by performing a numerical study, whose result motivate the analytic investigation of Section (ref). In particular, we simulate observations from the following data-generating process, where the volatility is a CIR process. Note that this data-generating process is also used for the analytic study in Section (ref).

assumption{Data-generating process} For $t \in [0,T]$, $T>0$, the dynamics of the log-price process $p(t)$ and the spot volatility process $\nu(t)$ read: $$\begin{cases} \displaystyle p(t) = p(0)+ \int_0^t \sqrt{\nu(s)} dW(s)\\ \displaystyle\nu(t) = \nu(0) + \theta \int_0^t \Big( \alpha -\nu(s)\Big) ds + \gamma\int_0^t \sqrt{\nu(s)} dZ(s)\end{cases},$$ where $W$ and $Z$ are two Brownian motions on $(\Omega, \mathcal{F}, (\mathcal{F}_t)_{t \ge 0}, P)$, $p(0) \in \mathbb{R}$ is the initial price, and the strictly positive constants $ \nu(0),\theta, \alpha$ and $\gamma $ denote, respectively, the initial volatility and the speed of mean reversion, long-term mean and vol-of-vol parameters. We also assume that $2\alpha \theta > \gamma^2$ to ensure that $\nu(t)$ is a.s. positive $\forall t \in [0,T]$.

In particular, we simulate one thousand 1-year trajectories of 1-second observations, with a year composed of 252 trading days of 6 hours each. We consider three scenarios determined by the following sets of model parameters: Set 1: $(\alpha, \theta,\gamma, \rho,\nu(0))= (0.2,5,0.5,-0.2,0.2)$; Set 2: $(\alpha, \theta,\gamma, \rho,\nu(0))=(0.02,10,0.25, -0.8, 0.03)$; Set 3: $(\alpha, \theta,\gamma, \rho,\nu(0))= (0.2,5,0.5, -0.2, 0.4)$.

The first set of parameters, Set 1, is taken from scm and v and is used as the baseline scenario. The second, Set 2, represents the opposite scenario. In fact, the volatility generated by Set 2 is lower than the volatility generated by Set 1, since the long term mean, $\alpha$, and the speed of mean reversion, $\theta$, are, respectively, much lower and much higher than in Set 1. The second scenario is also characterized by a lower volatility of the volatility, which is captured by the parameter $\gamma$ and a more pronounced leverage effect, which is captured by the correlation parameter $\rho$. The third set of parameters, Set 3, differs from the first only in that the initial value of the volatility, $\nu(0)$, is twice the long term volatility, $\alpha$. In this regard, note that if the initial volatility $\nu(0)$ is equal to $\alpha$, the spot volatility has a constant unconditional mean over time under Assumption (ref) (see Appendix A in bz). Setting $\nu(0)=\alpha$ is a simplifying assumption typically adopted in numerical studies where a mean-reverting volatility process is used (see, e.g., among many others, afl, scm, v).

We estimate daily values of the PSRV in these three scenarios from simulated prices sampled with a 1-minute frequency. For the estimation, we set $b=-1/2$ and $c=1/4$ \footnote{This choice of $b$ and $c$ satisfies the constraints for asymptotic unbiasedness (see Theorem (ref)). Moreover, note that the selection $b=-1/2$ is also performed in the numerical exercises of aflwy and scm.} and study the sensitivity of the bias to different values of $\kappa$ and $\lambda$. With respect to $\lambda$, we consider values in the set $(0.0002, 0.0004, 0.0006, 0.0010, 0.0019, 0.0029, 0.0057)$, which correspond to $\Delta_N$ equal to $1,2,3,5,10,15,30$ minutes, respectively, thereby preserving the high-frequency nature of the estimator. As for $\kappa$, we consider values in the set $(0.017, 0.033, 0.05, 0.1, 0.2, 0.4, 0.5, 1, 1.5, 2, 2.5, 3)$, which correspond to $W_N$ equal to (approximately) $5, 10, 15, 30, 60, 120, 150, 300, 450, 600, 750, 900 $ minutes, respectively. Overall, these sets of values for $\lambda$ and $\kappa$ allow us to consider both cases when window overlapping occurs, that is when $W_N > \Delta_N$, and cases when it does not occur, that is when $W_N \le \Delta_N$. Figure (ref) summarizes the results of the numerical exercise for values of $\kappa$ that lead to a relative bias smaller than an absolute value of $1$. This happens for $\kappa=1.5, 2, 2.5, 3$. Instead, for $\kappa$ smaller than $1.5$, the relative bias rapidly explodes for all values of $\lambda$ considered, as shown in Figure (ref).

figure[figure omitted — 656 chars of source]
figure[figure omitted — 728 chars of source]

As one can easily verify, the values of $\kappa$ in Figure (ref) imply that local windows for estimating the spot volatility overlap, for all values of $\lambda$ considered. Consequently, Figure (ref) tells us that window overlapping is crucial in order to optimize the relative bias of the PSRV even when {$\Delta_N >> \delta_N$}. This confirms the numerical results in scm. Furthermore, one can also easily check that the combinations of $\lambda$ and $\kappa$ such that overlapping does not occur are all included in Figure (ref), where the relative bias is larger than $1$ and rapidly increases as $\kappa$ becomes smaller, for any $\lambda$ considered, reaching the order magnitude $10^3$ when $W_N$ equals $5$ minutes.

Moreover, focusing on Figure (ref), it is worth noting that the bias-optimal selection of $\kappa$ is strongly dependent on the parameters of the data-generating process. In fact, the same value of $\kappa$ may lead to very different values of the bias in the three scenarios considered: for instance, the selection $\kappa=2$ leads to a relative bias of approximately $-20\%$ in scenario 1, $- 50\%$ in scenario 2 and $-3\%$ in scenario 3. At the same time, Figure (ref) also tells that the bias is not very sensitive to the selection of $\lambda$. Finally, Panel a) of Figure (ref) suggests that, for all values of $\lambda$ considered, the bias-optimal value of $\kappa$ is close to 2 in scenario 1. This indication is in line with the numerical findings by scm, where, based on the same parameter set as in scenario 1, the optimal value of $\kappa$ is found to be approximately equal to 2.

In sum, our preliminary numerical study shows not only that allowing for window overlapping is critical to avoid obtaining highly biased vol-of-vol estimates, but also that the selection of $\kappa$ is crucial {for optimizing} the PSRV finite-sample bias and, in particular, it is critical to uncover the dependence between the bias-optimal value of $\kappa$ and the parameters of the data-generating process. Gaining a more in-depth understanding of these numerical findings is what motivates our analytic study in the next section.

Analytic results

In this section we analyze the PSRV finite-sample bias in a parametric setting, namely assuming that the volatility is a CIR process, so that a fully explicit formula of the latter can be obtained. We treat the overlapping case (i.e., the case when $W_N > \Delta_N$) and the no-overlapping case (i.e., $W_N\le\Delta_N$) separately as, in general, the finite-sample bias expression differs in the two cases, independently of the parametric model used. Lemma (ref) details the explicit expression of the PSRV bias for $N$ fixed.

lemmaLet Assumption (ref) hold, with $Z$ independent of $W$. Further, let $N$ be fixed. If $W_N \le \Delta_N$, the bias of the PSRV in Definition (ref) reads \begin{equation} E\Big[PSRV_{[\tau,\tau+h],N}- \langle\nu,\nu\rangle_{[\tau,\tau+h]} \Big] = \gamma^2 \alpha h (A_N-1) + \gamma^2 \Big( E[\nu(\tau)] - \alpha \Big) \frac{1-e^{-\theta h}}{\theta} (B_N-1) + C_ N. \end{equation} Instead, if $W_N > \Delta_N$, it reads \begin{equation} E\Big[PSRV_{[\tau,\tau+h],N}- \langle\nu,\nu\rangle_{[\tau,\tau+h]} \Big] = \gamma^2 \alpha h (A_N-1) + \gamma^2 \Big( E[\nu(\tau)] - \alpha \Big) \frac{1-e^{-\theta h}}{\theta} (B_N-1) + C_N +O_N. \end{equation} The parametric expressions of $A_N$, $B_N$, $C_N$ and $O_N$ are rather cumbersome and thus are reported in the Appendix (see, respectively, equations ((ref)), ((ref)), ((ref)) and ((ref)) in the proof to Lemma (ref)).
proofSee Appendix A.
remarkThe bias in the case $W_N \le \Delta_N$ differs from that in the case $W_N > \Delta_N $ for the presence of the extra term $O_N$, which appears due to the fact that the parametric expression of $E[RV (\tau+i\Delta_N, k_N\delta_N)RV (\tau+i\Delta_N-\Delta_N, k_N\delta_N)]$ differs in the two cases. See the proof of Lemma (ref) for the definition of the quantity $RV$, which is only used in the Appendix, and further details.
remarkThe explicit bias expression in Lemma (ref) is derived under the simplifying assumption that $Z$ is independent of $W$, which rules out leverage effects. The sensitivity of the PSRV bias to the presence of leverage effects is studied numerically in Section (ref), where simulations suggest that such effects are a negligible source of finite-sample bias, thereby preserving the practical relevance of the results derived in this section. In the literature, numerical and empirical studies of the impact of leverage effects on analytical results derived under the no-leverage assumption are found, e.g., in bnvj, csda and references therein.

In the next subsections we investigate the existence of a rule to select the tuning parameters $\kappa$ and $\lambda$ in both the cases $W_N > \Delta_N$ and $W_N \le \Delta_N$. To do so, we first isolate the leading term of the bias in each case and then verify whether the latter can be canceled by an $ad$ $hoc$ selection of tuning parameters. We address overlapping case first, as it is the one relevant for practical applications, based on the results of the simulation studies in Section (ref) and in scm.

The relevant case for practical applications: $W_N > \Delta_N$

When $W_N > \Delta_N$, the natural expansion of the bias as the number of sampled price observations $N$ tends to infinity is precluded, because the consistency of the PSRV requires that $\frac{W_N}{\Delta_N}\to 0$ as $N \to \infty$. Thus, we determine the leading term of the bias through an alternative asymptotic expansion, which exploits some natural, non-restrictive constraints on the magnitude of the tuning parameter $\lambda$ and the time horizon $h$. Specifically, we first regard the bias in equation ((ref)) as a function of $\lambda$ and we perform its Taylor expansion with base point $\lambda=0$. Then, regarding each term of this expansion as a function of $h$, we perform their Taylor expansions with base point $h=0$. The choice of the base point $\lambda=0$ is supported by the fact that the largest feasible values of $\lambda$ {are} very small, e.g., on the order of $10^{-3}$ when $c<1/2$ and $\delta_N$ is equal to one minute (see Figure (ref) for the case $c=1/4$). Note that a value of $\lambda$ is feasible if it satisfies $\Delta_N:=\lambda\delta_N^c < h$. The choice of base point $h=0$ is instead supported by the fact that in the literature on high-frequency econometrics, the typical time horizon used to estimate the integrated quantities is one trading day, i.e., $h=1/252\approx 4 \cdot 10^{-3}$. The order of this sequential expansion is rather natural: intuitively, we first take the limit $\lambda \to$ 0 to approximate the integral of the vol-of-vol in an infill-asymptotics sense, then take the limit $h \to$ 0 to localize the estimate of the integral near the initial time $\tau$. This approach leads to the following result.

theoremLet Assumption (ref) hold, with $Z$ independent of $W$. Further, let $W_N>\Delta_N$. Then, for $N$ fixed, as $\lambda\to 0, h\to 0$ \begin{equation} E\Big[PSRV_{[\tau,\tau+h],N} - \langle\nu,\nu\rangle_{[\tau,\tau+h]} \Big] =\begin{cases} \Bigg( \frac{ 4E[\nu(\tau)]^2}{ {\kappa}^2\delta_N^{1+2b}} -\gamma^2 E[\nu(\tau)] \Bigg)h + O(h^{1-b}) +O(\lambda) \quad if \quad b\ge-1/2, c<-b \\ -\gamma^2 E[\nu(\tau)] h + O(h^{-2b}) +O(\lambda) \quad if \quad b<-1/2, c<1+b \\ \end{cases}. \end{equation} Moreover, let $(\mathcal{F}^{\nu}_t)_{t\ge0}$ be the natural filtration associated with the process $\nu$. Then, for $N$ fixed, as $\lambda\to 0, h\to 0$ \begin{equation} E\Big[PSRV_{[\tau,\tau+h],N} - \langle\nu,\nu\rangle_{[\tau,\tau+h]} | \mathcal{F}^{\nu}_\tau\Big] = \begin{cases} \Bigg( \frac{4\nu(\tau)^{2 }}{ {\kappa}^2\delta_N^{1+2b}} - \gamma^2\nu(\tau) \Bigg)h + O(h^{1-b}) +O(\lambda) \quad if \quad b\ge-1/2, c<-b \\ -\gamma^2\nu(\tau) h + O(h^{-2b}) +O(\lambda) \quad if \quad b<-1/2, c<1+b \\ \end{cases}. \end{equation}
proofSee Appendix A.
remarkThe expansion in Theorem (ref) is performed under the asymptotic constraints on rates $b$ and $c$ that ensure the asymptotic unbiasedness of the PSRV under Assumption (ref) (see Theorem (ref)).
remarkThe conditional bias expansion in equation ((ref)) allows the dominant term of the bias {to be expressed} in terms of $\nu(\tau)$ and $\gamma$, two quantities {that} can be consistently estimated over a fixed time horizon. This is crucial for the existence of a feasible procedure to select $\kappa$, as detailed below. Instead, the unconditional expression in equation ((ref)) depends on $E[\nu(\tau)]$, whose parametric expression in turn depends on the drift parameters of the volatility and thus cannot be consistently estimated over a fixed time horizon (see, e.g., kankris). In particular, it holds $E[\nu(\tau)]=(\nu(0)-\alpha)e^{-\theta \tau} +\alpha$ (see equation (4) in Section 2.2.1 of bz).

Figure (ref) compares the true finite-sample bias of the daily PSRV in equation (ref) with the dominant term of the expansion in equation (ref) as functions of the tuning parameter $\kappa$. Specifically, the panels refer to the three parameter sets already used in Section (ref), that is Set 1 (panel a)), Set 2 (panel b)), and Set 3 (panel c)). Note that we have set $b=-1/2$, $c=1/4$, $\lambda=0.0006$, $h=1/252$, and $N=360$. The corresponding $\delta_N$ and $\Delta_N$ are equal to 1 minute and (approximately) 3 minutes, as we consider 6-hr trading days. The approximation of the true bias with the dominant term of the expansion is very accurate.

figure[figure omitted — 659 chars of source]

Based on the conditional bias expansion in equation ((ref)), we make the following considerations on the optimal selection of tuning parameters in finite samples. First, we note that the dominant term of the bias can be annihilated simply by suitably selecting $\kappa$ for any feasible value of $\lambda$ when $b\ge -1/2$, $c<-b$. Instead, when $b<-1/2$, $c<1+b$, the dominant term of the bias is independent of $\kappa$ and $\lambda$. Specifically, when $b\ge -1/2$, $c<-b$, the suitable selection is

equation[equation omitted — 102 chars of source]

{ However, since $\kappa$ is a tuning parameter, it is not allowed to depend on $N$. Therefore, the only admissible choice is $b=-1/2$ and $c<1/2$, so that the suitable selection becomes}

equation[equation omitted — 97 chars of source]

Further, if $\nu(0)=\alpha$, then $E[\nu(\tau)]=\alpha$ (see equation (4) in Section 2.2.1 of bz) and thus, based on equation ((ref)), it is immediate to see that the bias-optimal value of $\kappa$ reduces to $ \displaystyle \frac{2 \sqrt{\alpha}}{\gamma}.$ Interestingly, this analytic result supports the optimal selections of $b$ and $\kappa$ determined numerically in the literature. Indeed, for the first parameter set in the numerical exercise in Section (ref), Set 1, which is also used in scm, $\kappa^*$ is equal to $1.79$, a value compatible with the numerical result in scm, where the optimal $\kappa$ is said to be approximately equal to $2$. Note also that the numerical studies in aflwy, scm both select $b=-1/2$. With regard to this selection of $b$, the following remark is in order.

remarkThe selection $b=-1/2$ does not satisfy the consistency constraint in Proposition (ref). However, to achieve consistency, it is sufficient to select $b=-1/2+\epsilon$ and $c=1/4 -\epsilon$, with $\epsilon$ strictly positive but arbitrarily small and, for such selection of $b$, bearing in mind ((ref)), the impact of $\epsilon$ on the optimal selection of $\kappa$ will be negligible in finite-sample exercises. Therefore, in finite-sample exercises we can select $\kappa=\kappa^*$ as in ((ref)).

Furthermore, the following remark regarding the selection $\kappa=\kappa^*$ is in order.

remark{The overlapping condition $W_N>\Delta_N$ implies a constraint on the price grid $\delta_N$. In particular, if $\kappa=\kappa^*$, for $b=-1/2$, $c<1/2$, $W_N>\Delta_N$ is equivalent to $\delta_N > \delta^*:=\displaystyle \Big( \frac{\kappa^*}{ \lambda}\Big)^{\frac{1}{c-1/2}} $. The threshold $\delta^*$ is very small for typical orders of magnitude of $\alpha, \theta $ and $\gamma$, $h$ corresponding to one trading day and any feasible value of $\lambda$. For example, for the values of the parameters in Set 1 (see Section (ref)), $\lambda=0.0006$ and $c=1/4$ (so that, if $\delta_N=1$ minute, then $\Delta_N:=\lambda\delta_N^c \approx 3$ minutes), we have $\delta^*=7.5 \cdot 10^{-8}$ seconds and thus the constraint $\delta_N>\delta^*$ is largely satisfied at the most commonly available price sampling frequencies.}

However, Equation ((ref)) implies that the bias-optimal selection $\kappa:=\kappa^{*}$ is unfeasible unless reliable estimates of $\nu(\tau)$ and $\gamma$ are available. In Appendix B we detail a simple feasible procedure to obtain $\kappa^{*}$. In a nutshell, the procedure is as follows. First, we estimate $\nu(\tau)$ using the Fourier spot volatility estimator by mm. Then we estimate $\gamma$ via a simple indirect inference method.

The case $W_N \le \Delta_N$

The finite-sample bias expression for $W_N \le \Delta_N$ in equation ((ref)) is the starting point to derive the asymptotic constraints on rates $b$ and $c$ that ensure the asymptotic unbiasedness of the PSRV. In this regard, we obtain the following result, which is based on the asymptotic expansion of the bias in the limit $N \to \infty$.

theoremLet Assumption (ref) hold, with $Z$ independent of $W$. Then, if $b\ge -1/2$ and $c < -b$ or $b<-1/2$ and $c<1+b $, $\frac{W_N}{\Delta_N} \to 0 $ as $N \to \infty$ and the PSRV as given in Definition (ref) is asymptotically unbiased, i.e., $$ E\Big[PSRV_{[\tau,\tau+h],N} - \langle\nu,\nu\rangle_{[\tau,\tau+h]} \Big] \to 0 \quad \textit{as} \quad N \to \infty. $$ In particular, as $N \to \infty$, \begin{eqnarray} E\Big[PSRV_{[\tau,\tau+h],N} - \langle\nu,\nu\rangle_{[\tau,\tau+h]} \Big]= a_1 \Delta_N+ a_2 \frac{1}{k_N\Delta_N}+a_3 \frac{k_N \delta_N}{\Delta_N} + o\Big(\Delta_N\Big) + o\Big(\frac{1}{k_N\Delta_N}\Big) + o \Big(\frac{k_N \delta_N}{\Delta_N}\Big), \end{eqnarray} where: \begin{eqnarray*} &&a_1= - \frac{\theta }{2} \gamma^2\alpha h + \frac{\theta}{2} \gamma^2 (E[\nu(\tau)] -\alpha)\frac{1-e^{-\theta h}}{\theta} + \frac{\theta}{2} (1-e^{-2 \theta h})\Big[(E[\nu(\tau)]-\alpha)^2+\frac{\gamma^2}{\theta}\Big(\frac{\alpha}{2}-E[\nu(\tau)]\Big)\Big] , \\ &&a_2= \frac{2}{\theta}\gamma^2\alpha h + \frac{4}{\theta}\gamma^2 (E[\nu(\tau)] -\alpha)\frac{1-e^{-\theta h}}{\theta} +\frac{2}{ \theta }(1-e^{-2 \theta h})\Big[(E[\nu(\tau)]-\alpha)^2+\frac{\gamma^2}{\theta}\Big(\frac{\alpha}{2}-E[\nu(\tau)]\Big)\Big] \\ && \quad\quad + 4 \alpha^2 h +\frac{8 \alpha (E[\nu(\tau)]-\alpha)(1-e^{-\theta h})}{\theta },\\ && a_3=-\gamma^2 (E[\nu(\tau)] -\alpha)\frac{1-e^{-\theta h}}{\theta}. \end{eqnarray*}
proofSee Appendix A.

A bias-optimal rule for the selection of the tuning parameters $\kappa$ and $\lambda$ when $W_N \le \Delta_N$ is given in the following corollary to Theorem (ref). Unfortunately, this bias-optimal rule is of little interest for practical applications, as explained in Remark (ref).

corollaryThe leading term of the PSRV finite-sample bias expansion in Eq. ((ref)) can be canceled in the case $b=-1/2$ and $c=1/4$, provided that there exists a solution ${(\tilde\kappa , \tilde\lambda ) \in \mathbb{R}_{>0}\times \mathbb{R}_{>0}}$ to the following system: $$\begin{cases} \displaystyle a_3 \kappa^2 + a_1 \lambda^2 \kappa + a_2=0 \\ W_N \le \Delta_N \end{cases}.$$ If a solution $(\tilde\kappa , \tilde\lambda )\in \mathbb{R}_{>0}\times \mathbb{R}_{>0}$ exists, the corresponding bias-optimal selection of $W_N$ and $\Delta_N$ reads $$W_N=\tilde\kappa \delta_N^{1/2}, \ \Delta_N=\tilde\lambda \delta_N^{1/4}.$$
proofSee Appendix A.
remarkFor $b =-1/2$ and $c=1/4$, the no-overlapping condition $W_N \le \Delta_N$ is equivalent to $ \delta_N \le (\lambda/ \kappa)^{4}$. Assuming that a positive solution $\tilde\kappa(\lambda)$ to $a_3 \kappa^2 + a_1 \lambda^2 \kappa + a_2=0$ exists for some $\lambda>0$, we define the “no-overlapping" threshold for $\delta_N$ as {$\delta^*(\lambda):= (\lambda / \tilde\kappa(\lambda))^{4}$}. For the three sets of CIR parameters used in the numerical study in Section (ref), Figure (ref) plots the threshold $\delta^*(\lambda)$ as a function of $\lambda\in (0,\lambda^*]$, where $\lambda^*$ is the largest admissible value of $\lambda$ such that $\lambda \delta^*(\lambda)^{1/4}\le h$, i.e., such that $\Delta_N \le h$ when $\delta_N$ is equal to the “no-overlapping" threshold. Specifically, Figure (ref) shows that the sampling frequency corresponding to $\delta^*(\lambda)$ is bounded by a value smaller than, respectively, $0.02$ (see Panel a)), $0.05$ (see Panel b)) and $0.125$ seconds (see Panel c)). This suggests that for typical values of the CIR parameters, the system in Corollary 1 may be solved only for ultra-high frequencies. Also, note that the solution (if it exists) depends on $a_1$, $a_2$ and $a_3$, which in turn depend on the expected initial volatility and all CIR parameters, including the drift parameters, which can not be consistently estimated over a fixed time horizon.
figure[figure omitted — 863 chars of source]

The impact of noise on the bias

In empirical applications one can only observe the noisy price $\tilde p (t)$, that is, the efficient price contaminated by a noise component that originates from market microstructure frictions, such as bid-ask bounce effects and price rounding. Here, we assume that the noise component is an i.i.d. process independent of the efficient price process, as in the seminal paper by roll. For a general discussion of the statistical models of microstructure noise, see jli.

assumption{Data-generating process in the presence of market microstructure noise} The observable price process $\tilde p $ is given by $$\tilde p (t) = p(t) + \eta(t),$$ where $p(t)$ represents the efficient price process and evolves according to Assumption (ref) while $\eta(t)$ is a sequence of i.i.d.\ random variables independent of $p(t)$, such that $E[\eta(t)]=0$, $E[\eta(t)^2]=V_{\eta} < \infty $ and $E[\eta(t)^4]=Q_{\eta} < \infty$ $\forall t$.

The presence of noise clearly changes the PSRV bias expression, introducing an extra term, as illustrated in the following lemma. Note that the parametric form of the extra bias term due to the presence of noise is different in the overlapping and no-overlapping cases.

lemmaLet Assumption (ref) hold, with $Z$ independent of $W$, and let $N$ be fixed. Moreover, let $\widetilde{PSRV}_{[\tau,\tau+h],N}$ denote the PSRV in Definition (ref), computed from noisy price observations. If $W_N \le \Delta_N$, then \begin{equation} E\Big[\widetilde{PSRV}_{[\tau,\tau+h],N}- \langle\nu,\nu\rangle_{[\tau,\tau+h]} \Big] = \gamma^2 \alpha h (A_N-1) + \gamma^2 \Big( E[\nu(\tau)] - \alpha \Big) \frac{1-e^{-\theta h}}{\theta} (B_N-1) + C_ N +D_N; \end{equation} Instead, if $W_N > \Delta_N$, then \begin{equation} E\Big[\widetilde{PSRV}_{[\tau,\tau+h],N}- \langle\nu,\nu\rangle_{[\tau,\tau+h]} \Big] = \gamma^2 \alpha h (A_N-1) + \gamma^2 \Big( E[\nu(\tau)] - \alpha \Big) \frac{1-e^{-\theta h}}{\theta} (B_N-1) + C_N +O_N+ D^{*}_N. \end{equation} The parametric expressions of $A_N$, $B_N$, $C_N$ and $O_N$ are as in Lemma (ref), while that of the extra term due to the presence of noise $D_N$ (resp., $D^*_N$) in the no-overlapping case (resp., overlapping case) is as follows: \begin{equation} D_N =[4(Q_\eta+V_\eta^2)+16\alpha V_\eta\delta_N]h\frac{1}{k_N\delta_N^2\Delta_N}+\frac{8}{\theta}V_\eta (\alpha -E[\nu(\tau)])(1-e^{-\theta h})\frac{(1+e^{-\theta \Delta_N})(1-e^{-\theta k_N\delta_N})}{(1-e^{-\theta \Delta_N})k_N^2\delta_N^2}; \end{equation} \begin{equation}\begin{split} D^*_N&=[ 4\left(Q_\eta+V^2_\eta\right)+ 16\alpha V_\eta \delta_N] h\frac{1}{k_N^2\delta_N^3}+ \frac{8}{\theta}V_\eta(\alpha-E[\nu(\tau)]) (1-e^{-\theta h})\frac{1}{(1-e^{-\theta\Delta_N}) k_{N}^{ 2} \delta_{N}^{ 2} } \Bigg\{ \frac{(2+ k_{N} )}{2k_N\delta_N}\Bigg[ \frac{(e^{\theta k_N \delta_N - \theta \Delta_N}-1)(k_N\delta_N+\Delta_N)}{k_N\delta_N-\Delta_N}\\ &+(e^{-\theta \Delta_N}-e^{\theta k_N\delta_N})\Bigg]+ \frac{k_N}{2\Delta_N}(1+e^{\theta k_N\delta_N})(1-e^{-\theta\Delta_N})\Bigg\}.\end{split} \end{equation}
proofSee Appendix A.
remarkFrom the proof of Theorem (ref) in Appendix A, one can easily see that the expressions of $D_N$ is the same for any continuous mean-reverting volatility model, as their computation only depends on the drift of $\nu$ in Assumption (ref). The same holds also for $D^*_N$ in the overlapping case.

Ideally, in the overlapping case, if one could efficiently estimate the extra bias due to noise $D^*_N$ and subtract it, then the bias-optimal rule to select $\kappa$ could still be applied effectively. Unfortunately, $D^*_N$ can not be consistently estimated over a fixed time horizon, as it depends on the drift parameters of the volatility $\alpha$ and $\theta$, whose consistent estimation can not be achieved on a fixed time horizon\footnote{Note that $Q_\eta$ and $V_\eta$ can, instead, be estimated consistently for $T$ fixed, see for instance ZMA}. As a solution, we suggest to sample prices on a suitably sparse grid, as done for the realized variance in the seminal paper by abdh, so that the extra bias term induced by the presence of noise becomes negligible and the bias optimal rule to select the local-window parameter $\kappa$ can still be applied. The efficiency of this solution is verified numerically in Section (ref).

Finally, for completeness, we also study the asymptotic behavior of the additional bias due to noise in the no-overlapping case, $D_N$. More precisely, in the next theorem we derive its rate of divergence as $N \to \infty$.

theoremLet Assumption (ref) hold, with $Z$ independent of $W$. Moreover, let $\widetilde{PSRV}_{[\tau,\tau+h],N}$ denote the PSRV in Definition (ref), computed from noisy price observations. Then, if either $b \ge -\frac{1}{2}$ and $c < -b $ or $b < -\frac{1}{2} $ and $c < b+1 $, $\frac{W_N}{\Delta_N} \to 0$ as $N \to \infty$ and $\widetilde{PSRV}_{[\tau,\tau+h],N}$ is asymptotically biased, i.e., $$ E\Big[\widetilde{PSRV}_{[\tau,\tau+h],N} -\langle\nu,\nu\rangle \Big]_{[\tau,\tau+h],N} \to \infty, \ \ as \ \ N \to \infty, $$ since the bias term $D_N$ in equation ((ref)) of Lemma (ref) diverges as $N \to \infty.$ In particular, we have \begin{equation*} k_N\delta_N^2\Delta_N D_N =4(Q_\eta+V^2_\eta)h +O(\delta_N) \quad and\quad k_N\delta_N^2\Delta_N \to 0 , \,\,\, \delta_N\to 0 . \end{equation*}
proofSee Appendix A.

The bias-reducing effect of the assumption $\nu(0)=\alpha$

As mentioned in Section (ref), if $\nu(0)=\alpha$, then $E[\nu(\tau)]=\alpha$. Lemmas (ref) and (ref) quantify the bias reduction ensuing from assuming that $ \nu(0) = \alpha$. Indeed, this assumption cuts off the entire source of bias $B_N$ and part of the sources of bias $D_N$ (see equation ((ref))) or $D^*_N$ (see equation ((ref))). The finite-sample bias reduction ensuing from the assumption $\nu(0)=\alpha$ is not peculiar to the PSRV, though. {In fact}, this simplifying assumption is also beneficial for reducing the finite-sample bias of the locally averaged realized variance, as shown in the next theorem.

theoremLet Assumption (ref) hold. Moreover, let $\hat\nu(\tau)$ denote the locally averaged realized variance in Definition (ref) at time $\tau$. Then, if $b \in (-1,0)$, $\hat\nu(\tau)$ is asymptotically unbiased, i.e., \begin{equation*} E[\hat\nu(\tau) - \nu(\tau)] =( \nu(0) -\alpha)e^{-\theta \tau}\frac{e^{\theta k_N \delta_N}-1- \theta k_N \delta_N}{\theta k_N \delta_N}, \end{equation*} and, as $N \to \infty$, we have \begin{equation*} { E[\hat\nu(\tau) - \nu(\tau)]=\frac{\theta}{2} ( \nu(0) -\alpha)e^{-\theta\,\tau}k_N\delta_N+o(k_N \delta_N), \quad k_N\delta_N\to 0}. \end{equation*} Let Assumption (ref) hold. Moreover, let $w(\tau)$ denote the locally averaged realized variance in Definition (ref) at time $\tau$ computed from noisy price observations. Then, $\forall$ $b \in (-1,0)$, $w(\tau)$ is asymptotically biased, i.e., \begin{equation*} E[w(\tau) - \nu(\tau)] =( \nu(0) -\alpha)e^{-\theta \tau}\frac{e^{\theta k_N \delta_N}-1- \theta k_N \delta_N}{\theta k_N \delta_N} + \frac{2V_\eta}{\delta_N}, \end{equation*} and, as $N \to \infty$, we have \begin{equation*} { E[w(\tau) - \nu(\tau)]=\frac{\theta}{2} ( \nu(0) -\alpha)e^{-\theta\,\tau}k_N\delta_N+\frac{2V_\eta}{\delta_N}+o(k_N\delta_N), \quad \quad k_N \delta_N\to 0}. \end{equation*}
proofSee Appendix A.

This theorem has two interesting implications. First, under Assumption (ref), the locally averaged realized variance is unbiased in finite samples if and only if $ \nu(0) =\alpha$. Second, under Assumption (ref), if $\alpha > \nu(0) $, the presence of noise could actually compensate for the negative bias originating from the first term of the bias expression. This also holds for the PSRV finite-sample bias, provided that the term $D_N$ (resp., $D^*_N$) in Lemma (ref) is of opposite sign with respect to the sum of the other terms in the bias expression.

Generalization via dimensional analysis

In this section we propose a heuristic approach, based on dimensional analysis, to generalize the rule for the bias-optimal selection of $\kappa$ in equation ((ref)), derived under the assumption that the volatility is a CIR process, to the more general case where the volatility follows a process in the CKLS class (see ckls). Specifically, the stochastic volatility model we assume as the data-generating process is now as follows.

assumption{Data-generating process} For $t \in [0,T]$, $T>0$, the dynamics of the log-price process $p(t)$ and the spot volatility process $\nu(t)$ follow $$\begin{cases} \displaystyle p(t) =p(0)+ \int_0^t \mu(s)ds + \int_0^t \sqrt{\nu(s)} dW(s) \\ \displaystyle \nu(t) = \nu(0) + \theta \int_0^t \Big( \alpha -\nu(s)\Big) ds + \gamma\int_0^t \nu(s)^\beta dZ(s) \end{cases}, $$ where $W$ and $Z$ are two correlated Brownian motions on $(\Omega, \mathcal{F}, (\mathcal{F}_t)_{t \ge 0}, P)$, $\mu(t)$ is a continuous adapted process, $ \beta \ge 1/2$, $p(0) \in \mathbb{R}, \nu(0), \theta, \alpha, \gamma >0$, and $2\alpha \theta > \gamma^2$ if $\beta=1/2$.

The stochastic volatility model in Assumption (ref) is quite flexible to reproduce empirical prices behaviour in the absence of price and volatility jumps. In fact it incorporates a number of widely-used stochastic volatility models with continuous price and volatility paths as special cases. For example, if $\beta = 1/2$, one obtains the model by heston; if $\beta=1$ one finds the continuous-time Garch model by nelson; if $\beta=3/2$, one gets the 3/2 model by 32. Further, by allowing for a stochastic correlation between $W$ and $Z$, Assumption (ref) includes also the generalized Heston model with stochastic leverage introduced by veraart2. Finally, note that Assumption (ref) also includes a price drift. The numerical study in Section (ref) confirms that the impact of the latter on the PSRV finite-sample bias is negligible.

We now use dimensional analysis to heuristically derive a rule for the bias-optimal selection of $\kappa$ under Assumption (ref). We test the efficacy of this rule in the numerical study of Section (ref), with overwhelming results. Note that dimensional analysis is typically used in physics and engineering to make an educated guess about the solution to a problem without performing a full analytic study (see, e.g., kyle, krishnam).

The basic concept of dimensional analysis is that one can only add quantities with the same units\footnote{{Dimensional analysis is also called a unit-factor method or a factor-label method, since a conversion factor is used to evaluate the units.}}. Accordingly, when applying dimensional analysis, the first step entails identifying the units of the quantities appearing in the equations being studied. In this specific analysis, we start with the units of the quantities appearing in the model given in Assumption (ref). Let $dim[q]$ denote the unit/dimension of the quantity $q$. The log-return $dp(t)$, $t>0$, is a dimensionless quantity (i.e., a pure number) since it is the logarithm of a ratio of prices (the ratio of quantities with the same units is dimensionless). Instead, the quadratic variation of the Wiener processes $W$ and $Z$ has the dimension of $time$ since $W$ and $Z$ are continuous random walks. As a consequence, we have $dim[dW(t)]=dim[dZ(t)]=time^{1/2}$ (see, for example, WILMOTT or the square-root-of-time rule in sqrtTime). Now consider the dynamics of the log-price, bearing in mind that we cannot add or subtract quantities with different measurement units. The dimension of the left-hand side must then be equal to those of the addenda on the right-hand side, thereby implying that $dim[\mu] =1/time$ and $\dim[\nu(t)]=1/time$. Thus, from the dynamics of $\nu(t)$, we have $dim[\alpha]=1/time$, $dim[\theta]=1/time$ and $dim[\gamma\nu(t)^\beta \,dZ(t)]=1/time$. The latter implies $dim[\gamma]dim[\nu(t)^\beta]dim[dZ(t)]=1/time$. Therefore, bearing in mind that $dim[\nu(t)^\beta]dim[dZ(t)]=1/time^{\beta-1/2}$, we obtain $dim[\gamma] =1/time^{-\beta+3/2}$.

Now, without loss of generality, let $ \nu(0) =\alpha$ and consider the dominant term in the expansion of Theorem (ref), i.e., the term $$\Big(\frac{4 \alpha^2}{\kappa^2\delta_N^{1+2b}} -\gamma^2\alpha\Big)h.$$ Since the dominant term of the PSRV bias must clearly have the same dimension as the expected quadratic variation of $\nu$ over any generic interval of length $h$, i.e., $\gamma^2 \alpha h$, we have

$$dim\Big[\Big(\frac{4 \alpha^2}{\kappa^2\delta_N^{1+2b}} -\gamma^2\alpha\Big)h\Big] =dim[ \gamma^2 \alpha h] = 1/time{^2},$$

and, as one can easily verify, this implies $dim[\kappa]=time^{-b}$ (alternatively, one can show that $dim[\kappa]=time^{-b}$ by simply noting that $k_N = \kappa\delta_N^b$ is dimensionless and $dim[\delta_N^{b}]=time^{b}$).

Now observe that the leading term of any expansion of the PSRV finite-sample bias must have dimension equal to $1/time{^2}$. Based on this observation, we conjecture that the leading term of the expansion in Theorem (ref) under Assumption (ref) is

$$\Big(\frac{4 E[\nu(\tau)]^2}{\kappa^2\delta_N^{1+2b}} -\gamma^2E[\nu(\tau)^{2\beta}]\Big)h,$$

whose dimension is $1/time^2$, as one can easily check by recalling that $dim[\kappa]=time^{-b}$, $\dim[\nu(t)]=1/time$ and $dim[\gamma] =1/time^{-\beta+3/2}$. Accordingly, if one conditions the bias to the natural filtration of $\nu(t)$ up to time $t=\tau$, the generalized bias-optimal value of $\kappa$, for $b=-1/2$ and $c<1/2$, reads

equation[equation omitted — 93 chars of source]

Note that equation ((ref)) can be rewritten in non-parametric form as

equation*[equation* omitted — 68 chars of source]

where $\xi(t):=\gamma^2\nu(t)^{2\beta}$ is the vol-of-vol process. This result, while offering insight into the non-parametric solution to the problem of the bias-optimal selection of $\kappa$, is problematic in terms of feasibility as it requires the estimation of the spot vol-of-vol $\xi(t)$ at $t=\tau$, a challenging issue which has not been addressed so far in the literature to the best of our knowledge and goes beyond the scope of this paper.

Our conjecture is based on the origin of the two addenda in the leading term of the bias (see Theorem (ref)) in the CIR framework. In fact, bearing in mind the the leading term is $$ \displaystyle \Bigg( \frac{4E[\nu(\tau)]^2}{k^2\delta_N^{1+2b}} -\gamma^2 E[\nu(\tau)] \Bigg)h \,, $$ we note that the second addendum, i.e., $\gamma^2E[\nu(\tau)]h$, comes from the expected quadratic variation of the volatility process. More specifically, it originates from the leading term of the following expansion:

eqnarray*[eqnarray* omitted — 120 chars of source]

Instead, the first addendum, i.e., $\frac{4E[\nu(\tau)]^2}{k^2\delta_N^{1+2b}}$, is due to the drift of the volatility process.

Thus in the case of the CKLS model, the first addendum remains unchanged since the drift of the process is the same for any $\beta$, while the second addendum changes according {to} the expected quadratic variation of the volatility process, {which}, for small $h$, reads

eqnarray*[eqnarray* omitted — 130 chars of source]

since $E\Big[ \langle\nu,\nu\rangle_{[\tau,\tau+h]} \Big] = \gamma^2 \int_{\tau}^{\tau+h} \,E[\nu(s)^{2\beta}] ds$.

Numerical results

Numerical results in the CIR setting

As detailed in Section (ref), in the absence of microstructure noise and assuming $\nu(\tau)$ to be observable and $\gamma$ to be known, the finite-sample bias of the PSRV is optimized, under Assumption (ref) and for any $\nu(0)$, by selecting $b=-1/2$, $c<1/2$ and $\kappa=\kappa^{*}:= 2\sqrt{\nu(\tau)}\gamma^{-1}$. In this subsection, we give numerical confirmation of the optimality of this rule for the selection of $\kappa$ in three progressively more realistic scenarios, where incremental sources of biases are added.

In the first scenario, we simulate log-price paths under Assumption (ref) and compute daily PSRV values from noise-free price observations assuming that the CIR parameters are known and the initial volatility value $\nu(\tau)$ is observable. In this scenario, we use two price sampling frequencies, that is, $\delta_N=1$ minute and $\delta_N= 5$ minutes. Results show that the bias generated by the price discrete sampling is relatively small, e.g., less than $5\%$ if $\delta_N = 1$ minute when $\kappa = \kappa^{*}$ (see Table (ref)).

In the second scenario, we simulate log-price paths under Assumption (ref) and compute PSRV values from noisy prices while assuming that the CIR parameters are known and the initial volatility value $\nu(\tau)$ is observable. As the PSRV is not robust to the presence of noise contaminations in the price process, here we only consider the sampling frequency $\delta_N= 5$ minutes, as recommended in the seminal paper {by} abdh, where the authors suggest that this sampling frequency reduces the impact of noise on returns while still falling within a high-frequency framework. Indeed, a comparison of the numerical results obtained in these first two scenarios shows that the impact of the price noise on the PSRV estimates is relatively small at the 5-minute sampling frequency, when $\kappa =\kappa^{*}$ is used.\\ In the third scenario, we still simulate the log-price path under Assumption (ref), but the value of the initial volatility, $\nu(\tau)$, is now unobservable and the model parameter $\gamma$ is unknown. Thus, we compute PSRV values from noisy prices by selecting $\kappa=\hat\kappa^{*}:= \frac{2\sqrt{\hat \nu (\tau)}}{\hat\gamma}$. Here, $\hat\nu(\tau)$ and $\hat\gamma$ are obtained through the estimation procedure detailed in Appendix B. A comparison of the results obtained in these different scenarios shows that the PSRV finite-sample bias reduction obtained with the feasible selection $\kappa=\hat\kappa^{*}$ is very similar to the reduction obtained with the unfeasible selection $\kappa=\kappa^{*}$. For the simulation of each scenario, we use the three realistic sets of parameters from Section (ref). For each parameter set, we simulate one thousand 1-year trajectories of 1-second observations.

The noise component $\eta$ in Assumption (ref) is simulated as an i.i.d.\ Gaussian process, with noise-to-signal ratio $\zeta$ ranging from 0.5 to 3.5, as in the numerical exercise proposed in scm. We define the noise-to-signal ratio $\zeta$ as in scm, i.e., $\zeta := \frac{std(\Delta \eta)}{std(r)}$, where $\Delta \eta$ denotes a generic increment of the i.i.d.\ process $\eta$ under Assumption (ref) and $r$ denotes the noise-free log-return at the maximum sampling frequency available, which is equal to 1 second in our numerical exercise. From the simulated prices, we compute daily PSRV values, that is, we set a small time horizon $h$, i.e., $h=1/252$. Recall that the bias-optimal rule for the selection of $\kappa$ is valid when $b=-1/2$ and $c<1/2$. Accordingly, we set $b=-1/2$ and $c=1/4$ in our numerical study. \\ Tables (ref)--(ref) summarize the results of our numerical exercises and, to make the results of the three parameter sets comparable, we report the values of the relative bias. Since we simulate 6-hr days, $N$ is equal to $360$ when $\delta_N=1$ minute and $72$ when $\delta_N= 5 $ minutes. Note that the overlapping condition $W_N>\Delta_N$ is always satisfied for the values of $\Delta_N$ in Table (ref). In particular, the average length of $W_N$ is approximately equal to: 530 minutes for Set 1, 410 minutes for Set 2 and 580 for Set 3, when $\delta_N = 1 $ minute; 1200 minutes for Set 1, 930 minutes for Set 2 and 1310 minutes for Set 3, when $\delta_N =5$ minutes. These averages are computed over all simulated days and are stable across the three scenarios. Recall that the length of $W_N$ varies by day, as it depends on $\kappa^{*}$, which in turn depends on the volatility value at the beginning of each day, i.e., $\nu(\tau)$ (in scenarios 1 and 2), or its estimate, i.e., $\hat\nu(\tau)$ (in scenario 3).

table[table omitted — 1,650 chars of source]
table[table omitted — 4,358 chars of source]

Table (ref) shows that for $\delta_N= 1$ minute and $\Delta_N \le 3$ minutes, the bias is almost negligible (i.e., less than $1\%$) when $\nu(0) =\alpha$, while it is slightly larger but still acceptable (i.e., between $3\%$ and $4\%$) when $\nu(0)=2\alpha$. This is in line with equation (ref) in Lemma (ref), where it is evident that the source of bias $B_N$ is eliminated when $\nu(0) =\alpha$, which implies $E[\nu(\tau)]=\alpha$. With a price sampling frequency of five minutes, the bias is still acceptable, around $6\%$ at worst. Additionally, Table (ref) shows that in the presence of noise, price sampling at five-minute intervals to avoid microstructure frictions represents an acceptable compromise, as the bias is less than $15\%$ even in the presence of very intense microstructure effects. Finally, Table (ref) shows that the statistical error related to the estimation of $\gamma$ and $\nu(\tau)$ could actually partially compensate for the bias due to the presence of noise, especially when the common assumption $\nu(0)=\alpha$ is violated.

Finally, an important remark is in order. The three realistic parameter sets that we have used all imply the presence of leverage effects, as each includes a negative $\rho$. To meet the simplifying no-leverage assumption under which the results in Section (ref) are derived, for each scenario we have also performed additional simulations under the hypothesis of the independence between the Brownian motion driving the price and the volatility, keeping the same values of the variance parameters $\alpha, \theta, \gamma$, and $\nu(0)$ that characterize the scenario. However, the numerical results obtained in the absence of leverage are basically indistinguishable from those illustrated in Tables (ref)--(ref), thereby suggesting that the leverage is a negligible source of bias. Such additional numerical results are not reported here for brevity, but are available from authors.

table[table omitted — 1,294 chars of source]
table[table omitted — 1,268 chars of source]

Numerical results in the more general CKLS setting

We conclude this section by testing the efficacy of the generalized, conjecture-based, criterion for the bias-optimal selection of $\kappa$ under Assumption (ref), i.e., under the assumption that the volatility evolves as a CKLS model. In this case, the feasible version of the bias-optimal rule to select $\kappa$ is given by $\hat\kappa^{**}=\frac{2 \hat\nu(\tau)^{1-\beta}}{\hat\gamma }$, for $b=-1/2$, $c<1/2$.

To test the efficacy of this criterion, we repeat the numerical exercise previously performed in scenario 1 under Assumption (ref), considering three different values of $\beta$: $\beta=1/2$, corresponding to the model by heston, which differs from the model of Assumption (ref) only in the presence of a price drift; $\beta=1$, corresponding to the continuous-time GARCH model by nelson; and $\beta=3/2$, corresponding to the 3/2 model by 32. For all parameter sets, $\mu$ is set equal to 0.05. Tables 4,5 and 6 show that our general criterion for the bias-optimal selection of $\kappa$ under Assumption (ref) is effective, as it gives satisfactory results in terms of relative bias. Note that the case $\beta=1/2$ is of interest only in that it confirms that the criterion for the bias-optimal selection of $\kappa$ derived analytically under Assumption (ref), i.e., $\kappa = \kappa^{*}$, is also effective in the presence of a price drift.

table[table omitted — 1,273 chars of source]

Empirical study

We conclude the paper with an empirical analysis, where we apply the bias-optimal criterion for selecting $\kappa$ in equation ((ref)) to compute daily PSRV estimates. The dataset is composed of two 1-year samples of S&P 500 1-minute prices relative to the years 2016 and 2017, respectively. The two samples are analyzed separately since the volatility of these two time series behaves very differently. In fact, the year 2016 is characterized by volatility spikes (due, e.g., to uncertainty pertaining to the so-called Brexit in the month of June or the U.S. presidential election in the month of November), while the year 2017 is characterized by low volatility, as one can see in Figure (ref). Analyzing the two series separately allows for validation of the feasible rule for the selection of $\kappa$ in two very different scenarios.

figure[figure omitted — 182 chars of source]

We proceed as follows. First, through the method detailed in Appendix B, we obtain non-parametric estimates of the process $\nu$ at the beginning of each day and estimates of $\gamma$ under Assumption (ref), for the three different values of $\beta$ considered in the numerical exercise of Section (ref). The results of the estimation of $\gamma$ are shown in Table (ref)\footnote{The estimates of the process $\nu$ at the beginning of each day are not reported for the sake of brevity. See Chapter 4 in mrs for a detailed study which demonstrates the finite-sample accuracy of the Fourier estimator of the spot volatility suggested in Appendix B.}.

table[table omitted — 610 chars of source]

Then, based on $R^2$ values, we assume the Heston model ($\beta=1/2$) as the data generating process for both samples. Consequently, we select $b=-1/2$, $c=1/4$, $\kappa = 2 \hat \gamma^{-1}\sqrt{\hat\nu(\tau)}$ and compute daily PSRV values from empirical prices sampled at the frequency $\delta_N=5$ minutes. The resulting selection of $W_N$ is approximately equal, on average, to $450$ minutes in 2016 and $275$ minutes in 2018. Note that the selection $\delta_N= 5$ minutes is justified by the fact that we assume the impact of microstructure contaminations to be negligible at that sampling frequency, based on the application of the Hausman test by asx for the presence of noise, which tells that the impact of noise at the 5-minute frequency is negligible in our samples, confirming a well-known stylized fact (see abdh).

Additionally, we have performed the jump-detection test by cpr. Precisely, we have performed this test on 5-minute returns, as it is not robust to the presence of noise. Based on the results of the jump test, we have computed daily PSRV values from 5-minute prices according to the following procedure for the removal of days with jumps\footnote{Note that the analytical results in Section (ref) are derived under the assumption of absence of jumps in the price and volatility. The literature on non-parametric jump tests provides large and robust empirical evidence, mainly based on US markets, that volatility jumps are accompanied by price jumps (see, e.g., JT,BRcoJ,BW). Thus removing days with price jumps from the estimation basically also takes care of jumps in the volatility. In the absence of jumps, a model in the CKLS class for the volatility could provide a reasonable trade-off between accuracy in reproducing empirical features of prices and parsimony in terms of parameters to be estimated, as pointed out, e.g., in Christ2010 and Goard.}. Assume that time is measured in days and that we are interested in computing the daily PSRV on $[\tau, \tau +1]$. Further, without loss of generality, assume that the bias-optimal value of $W_N$ is equal to 2.5 days for $\delta_N$ equal to 5 minutes. If jumps have occurred within the interval $[\tau-1,\tau+1]$, i.e., on the day of interest or the day before, we do not compute the PSRV on $[\tau,\tau+1]$; if jumps have occurred in $[\tau-2,\tau-1)$, but not in $[\tau-1,\tau+1],$ at any instant $\tau+i\Delta_N$, $i=0,1,\ldots,\lfloor \Delta_N^{-1} \rfloor$, we select $W_N=1+i\Delta_N \le2$ {to pre-estimate} the spot volatility; instead, if jumps have occurred in $[\tau-3,\tau-2)$, but not in $[\tau-2,\tau+1]$, at any instant $\tau+i\Delta_N$, we select $W_N=min(2+i\Delta_N,2.5)$; finally, if no jump has occurred within the period $[\tau-3,\tau+1]$, at any instant $\tau+i\Delta_N$, we select $W_N=2.5.$ Overall, based on this procedure, the days for which we do not compute the PSRV amount to $12.25\%$ of the sample in 2016 and $8.30\%$ of the sample in 2017.

Figures (ref) and (ref) show the daily PSRV values obtained for four different values of $\lambda$ corresponding to a spot volatility estimation frequency $\Delta_N$ equal to 5, 10, 15, and 30 minutes, respectively.

figure[figure omitted — 125 chars of source]
figure[figure omitted — 125 chars of source]

Comparing the dynamics of the VIX$^2$ index in Figure (ref) with those of the PSRV, one notices that when the VIX$^2$ spikes, the vol-of-vol also spikes (see, e.g., the behavior of the plots at the end of June 2016) and, viceversa, when the VIX$^2$ is low and stable (e.g., in 2017) the vol-of-vol is also low and stable. This evidence corroborates the goodness of our vol-of-vol estimates. Finally, note that for either of the two samples, the plots for different values of $\Delta_N$ are basically indistinguishable. With respect to the bias-optimal selection of $\lambda$ (i.e., $\Delta_N$), this evidence confirms what emerges from the analytic study in Section (ref): the impact of the selection of $\lambda$ (i.e., $\Delta_N$) on PSRV values is marginal, if not negligible.

Conclusions

The pre-estimated spot-variance based realized variance (PSRV) by bnv, the simplest and most natural consistent estimator of the integrated vol-of-vol, is typically affected by a substantial finite-sample bias. The main contribution of this paper is to show, analytically, that local-window overlapping in finite samples effectively reduces this bias. This result confirms the findings of scm, based on simulations.

The paper is written in the spirit of afl. In afl, a parametric data-generating process, namely the Heston model, is used to obtain a fully explicit bias expression for the price-volatility correlation, the most natural leverage estimator, which is very biased at high frequencies. Based on the full explicit knowledge of the bias, the authors are able to isolate the sources of bias that affect the simple leverage estimator and derive a feasible strategy to correct for them. In this paper we follow a similar approach. Assuming that the volatility is a CIR process, we obtain the full explicit expression of the PSRV finite-sample bias. Crucially, we show that this expression differs in the overlapping case and the no-overlapping case and, most importantly, that a feasible bias-correction strategy for finite samples can be derived only in the overlapping case.

Further, using dimensional analysis, we generalize the feasible bias-correction strategy to hold under the assumption that the volatility process belongs to the more general CKLS class, which encompasses a number of widely-used parametric models. Numerical results corroborate the validity of the generalized rule in that nearly unbiased vol-of-vol estimates are obtained for two other models in the CKLS class, namely, the continuous-time GARCH model and the 3/2 model.

In the paper, the impact of microstructure noise on the PSRV bias is also investigated. First, we derive the exact analytic expression of the extra bias due to noise, which differs in the overlapping and no-overlapping cases. This extra bias can not be consistently estimated over a fixed time horizon and then subtracted, as it depends, other than on some moments of the noise process, on the drift of the volatility. As a solution, we propose to apply the feasible rule for the bias-optimal selection of the local-window parameter on sparsely-sampled prices, following abdh. Numerical evidence of the efficacy of this solution is provided.

Finally, as a byproduct of this analysis, we quantify, for both the PSRV and the locally averaged realized variance, the bias reduction ensuing from the assumption that the initial value of the volatility is equal to its long-term mean, which is very common in simulation studies found in the literature.

Declarations

Funding: Not applicable.

Conflicts of interest: The authors have no conflicts of interests to declare.

\noindentAvailability of data and material: Not applicable.

\noindentCode availability: See supplementary file.