Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,148 characters · 5 sections · 28 citation commands
Data-driven fixed-point tuning for truncated realized variations
The continuous part of the quadratic variation of an It\^o semimartingale, commonly known as the integrated volatility, plays an outsize role in financial econometrics, and its estimation in various settings based on discrete observations has been a major focus in the literature at various points in the past 20+ years. The semimartingale $X$ commonly represents the log-price of a financial asset, and its integrated volatility serves as a measure of the overall uncertainty inherent in the continuous part of $X$ over a given time period.
Among the variety of available methods for integrated volatility estimation, the truncated realized variation (TRV), introduced in mancini:2001, was one of the first and remains among the most popular approaches to-date that is jump-robust, in the sense that it can still provide reliable estimates of integrated volatility when jumps occur in the process $X$. Other well known jump-robust methods for estimating integrated volatility include bipower variations and their extensions barndorff-nielsen:2004,barndorff-nielsen:shephard:winkel:2006,corsi:pirino:reno:2010 or those based on empirical characteristic functions todorov:tauchen:2012,jacod:todorov:2014,jacod:todorov:2018, among others, giving the practitioner a wide array of choices at their disposal for estimation of integrated volatility in modeling contexts where jumps may be present.
To choose an estimator among this array of options, currently one must first decide between two distinct classes: either asymptotically efficient approaches, like TRV, which require selection of tuning parameters, or alternatively “tuning-free” estimators but at the unfortunate expense of asymptotic efficiency. From the perspective of minimizing variance, asymptotically efficient approaches are preferable, but their use in practice necessitates the critical additional step of specifying the tuning parameter values themselves. This consequential step can significantly impact estimation performance, but current asymptotic theory does not offer direct guidelines for choosing parameters explicitly, which can be an extremely delicate matter in practice. For instance, even in idealized asymptotic settings, appropriate choices often depend on a priori unknown properties of $X$ and can determine whether or not a given estimator retains even the basic requirement of consistency. In the absence of theoretically supported approaches for specifying explicit values of these parameters, the practical use of tuning-parameter-based methods remains entirely reliant on heuristics. The purpose of the present work is to address this gap.
In the case of TRV, the tuning parameter of importance is called the threshold, denoted hereafter as $\varepsilon>0$, indicating a level above which increments are discarded from the estimation procedure. Concretely, given a discretely observed semimartingale $X=\{X_t\}_{t\geq 0}$ at times $0=t_0<t_1<\ldots<t_n=T$, the TRV is defined as
where $\Delta_{i}^{n}X:=X_{t_{i}}-X_{t_{i-1}}$ is the $i^{th}$ increment of $X$, often assumed to be observed on a regular sampling grid, so that $t_i-t_{i-1}=:h_n$ for all $i$. Statistical properties of TRV have been extensively studied when $\varepsilon=\varepsilon(h_n)$ is a deterministic function of the time step $h_n$ such that $\varepsilon(h_n)\to 0$ at specified rates as $h_n\to 0$. In mancini:2009, when either the jump component of the process $X$ is of finite activity or is a pure-jump L\'evy process with infinite jump activity, TRV was shown to be consistent whenever
A consistency statement for TRV encompassing a broader class of semimartingales was given in jacod:2008, but for the more restrictive case of power thresholds, namely, thresholds of the form
Under finite jump activity, central limit theorems for TRV were established under the threshold constraint (ref) in mancini:2009; in the infinite-activity case, they were established in jacod:2008 for more general semimartingales based on thresholds satisfying (ref) under the additional assumption the volatility itself is a semimartingale, and also in cont:mancini:2011,mancini:2011 for general c\`adl\`ag volatility processes but for L\'evy-type jump behavior, both under additional constraints on $\varepsilon$ related to the Blumenthal-Getoor index of $X$.
While asymptotic constraints such as (ref) and (ref) may be informative for threshold selection, they do not concretely indicate how one should make an explicit choice for $\varepsilon$ in a given context. Moreover, even if a particular deterministic choice for $\varepsilon$ may lead to good estimation performance under a given model, the same choice of $\varepsilon$ under a perturbed version of the same model can lead to dramatically worse estimation performance. To illustrate this point, the left panel of Figure (ref), below, shows histograms of the relative estimation errors for TRV using a fixed, deterministically chosen threshold value under two different parameter settings of the same model. While TRV performs satisfactorily with this deterministic threshold value under one of the parameter settings, it performs poorly with the same threshold value under alternate parameter settings, even though the expected quadratic variation of $X$ is the same in both cases. In contrast, the right panel of Figure (ref) displays histograms of relative estimation errors for the approach developed in this paper, where satisfactory performance is maintained across both settings.
Though Monte Carlo studies or empirical insights may help in choosing the value of $\varepsilon$ deterministically in a given setting, an arguably more natural approach is to select thresholds through some data-driven procedure, permitting the threshold itself to depend on observed data. Indeed, random, data-driven parameter tuning is often done in numerical studies in the literature -- without theoretical support -- to illustrate finite-sample behavior of estimators and to improve their numerical performance.
\footnotetext{Specifically our approach as described in (6b) in Section (ref), though similar behavior holds in all other cases.} However, by their very nature, data-driven parameter selection procedures introduce considerable statistical dependencies and associated theoretical challenges that are otherwise absent when parameters are chosen deterministically. Consequently, despite the practicality and potential benefits of data-driven parameter selection, the literature on TRV and related methods employing data-driven tuning procedures has remained relatively scarce. For instance, in the case of finite activity jumps, it was stated without proof in a remark in mancini:reno:2011 that consistency holds for time-dependent random thresholds (possibly different for each increment $\Delta_i^n X$) of the form $c_{t_i}\varepsilon$, where $\varepsilon=\varepsilon(h_n)$ satisfies (ref) and $\{c_t\}_{t\geq 0}$ is a stochastic process that is a.s. bounded above and bounded away from $0$. Later, in figueroa-lopez:mancini:2019, consistency was rigorously established under finite jump activity for possibly data-dependent time-varying thresholds of the type $\sqrt{2(1+\eta)M_i h_n \log (1/h_n)}$, where $\eta>0$ and $M_i$ are random variables satisfying $M_i\in[ \inf_{s\in [t_{i-1}, t_i]} \sigma_s^2,\, \sup_{s\in [0,T]} \sigma_s^2 ]$ a.s. To the authors' knowledge, these statements comprise the totality of asymptotic theory for TRV with data-driven thresholds, and there is currently no theoretical support in the literature for data-driven parameter tuning of TRV outside consistency statements in the finite activity setting.
Moreover, in spite of the considerable focus on asymptotic properties of TRV with the threshold constraints (ref) and (ref), recent work figueroa-lopez:mancini:2019,figueroa-lopez:nisen:2013 has demonstrated that certain optimal choices of threshold do not satisfy these asymptotic conditions, leaving a substantive gap in the available asymptotic theory even within the scope of deterministic thresholding. Optimal-type thresholds can lead to substantial gains in finite sample estimation performance, and their explicit expressions can serve as a more direct guideline for threshold selection, making them ideal choices for practitioners. However, their direct use, even to first-order approximation, is complicated by the fact that they depend on the volatility itself. For instance, under an idealized constant volatility assumption and general finite jump activity, the MSE-optimal threshold $\varepsilon_n^\star$ admits the approximation:
where $\sigma>0$ is the volatility. Under L\'evy stable-like infinite jump activity, the MSE-optimal threshold is the same as $\varepsilon^{\star}_n$ up to an additional multiplicative constant depending on the Blumenthal-Getoor index figueroa-lopez:gong:han:2022. Though this expression cannot be used directly in practice due to its dependence on knowledge of the volatility, it lends itself naturally to fixed-point iterative procedures, as suggested in figueroa-lopez:nisen:2013 and figueroa-lopez:mancini:2019, whose asymptotic theory has remained unestablished, until now.
In this work, we consider two classes of iterative procedures for jump-robust estimation of the integrated volatility based on data-driven parameter tuning. Our procedures are designed to turn the otherwise infeasible threshold (ref) into a feasible one, and will be seen to stem from solutions $\xi$ to random fixed-point equations of the type
for an appropriate function $\Phi_n$ and sequence $r_n\to0$. Viewed as random, data-dependent thresholds, our procedures extend the asymptotic theory beyond deterministic thresholding to accommodate automatic, data-driven calibration of TRV, and further extends current asymptotic theory beyond the general rate constraints imposed in (ref) and (ref), allowing for time-dependent thresholding, ultimately leading to substantial gains in finite-sample performance and more principled threshold selection procedures. Part of our analysis is based on relating our proposed iterative estimators to oracle-like sequences of estimators; this general approach may be of use for parameter tuning in other jump-robust methods in the literature.
This paper is organized as follows. Section (ref) introduces the model, estimation framework and some notation used throughout the paper. Section (ref) contains our main results, including instances of uniform and time-varying thresholding, and Section (ref) contains some numerical illustrations concerning finite-sample estimation performance. The proofs of the main results and auxiliary lemmas are given in Appendices (ref) and (ref).
We consider a 1-dimensional It\^o semimartingale $X=(X_{t})_{t\geq 0}$ defined on a complete filtered probability space $(\Omega,\mathscr{F},(\mathscr{F}_{t})_{t\geq 0},\mathbb{P})$ of the form
Above, $W$ is a standard Brownian motion, $b,\gamma,\sigma$ are c\'adl\'ag adapted, $L=\{L_t\}_{t\geq 0}$ is a pure-jump infinite-activity L\'evy process, and $J=\{J_t\}_{t\geq 0}$ is a general pure-jump process with finite jump activity. We refer to Assumption (ref) for complete conditions on all driving processes and coefficients.
We suppose that on a fixed and finite time interval $[0,T]$, $n$ observations, $X_{t_1}, X_{t_2}, \ldots, X_{t_n}$, of the continuous-time process $X$ are available at known times $0=t_0<t_1<\ldots<t_n=T$. We assume sampling times are evenly spaced, and denote the time step between observations as $h_n:=T/n$. Our estimation target is the integrated volatility (or integrated variance) of $X$ defined as $$ C_T :=\int_0^T\sigma_s^2ds. $$
We consider two classes of estimators of $C_T$. The first class of estimators we consider are based on an iterative scheme that proceeds as follows:
The second class of estimators we consider are based on time-varying (or local) thresholding, namely estimators $\widehat C^*_n$ of the type
for appropriate data-driven local thresholds $B^*_n(i)$, $i=1,\ldots, n$, arising from fixed-point equations, whose precise definition is deferred to Section (ref). The central focus of this work is to study the classes of estimators defined by (ref) and (ref). Below we state our main assumptions relating to the model (ref).
Our assumptions, in particular, do not require $\sigma$ to be a semimartingale, which is important in rough volatility modeling. Note that the condition on $\gamma$ in (ref) is satisfied whenever $\{\gamma_t\}_{t\geq{}0}$ is an It\^o semimartingale with locally bounded characteristics.
We use the following standard notation throughout the paper: for two sequences $a_n,b_n>0$,
In this section, we study the asymptotic properties of the estimator $\widehat C_n$ introduced in Section (ref). As stated in the introduction, in figueroa-lopez:mancini:2019 it was shown that the first-order asymptotic behavior of the MSE-optimal threshold $\varepsilon_n^\star$ under the idealized assumption of constant volatility takes the form (ref), which cannot be implemented feasibly in practice as it depends on knowledge of the volatility itself. However, exploiting this relationship is the driving principle behind the iterative algorithm leading to the estimators $\widehat C_n$. The proposed method can be seen as a natural mechanism to make such a threshold feasible by taking the sequence $r_n=r(h_n)=2h_n\log(1/h_n)$ in the iterative procedure (ref). As we will see, our iterative approach will enable us to asymptotically “attain” the infeasible threshold $\varepsilon^\star_n$, in principle rendering near-MSE-optimal behavior possible in practice.
In general, it is a nontrivial task to establish asymptotic properties of the $\widehat C_{n}$ defined in (ref), even drawing upon results from the existing literature, which has almost exclusively focused on deterministic uniform thresholding. For instance, in spite of the fact that $\widehat C_n$ satisfies the random fixed-point equation
such an expression offers little insight into finding closed-form expressions for $\widehat C_n$.
The central idea in our approach rests on relating the sequence of iterates $\widehat C_{n,j}$ to an iterative sequence of “oracle-like” estimators $\widetilde C_{n,j}(y_n)$ that make use of the (unknown) location of jumps of size $y_n>0$ or larger, where $y_n\to 0$ at an appropriate rate. More concretely, for each $y\in(0,1)$, by virtue of the L\'evy-It\^o decomposition of $L$, we may reexpress
where $\mu$ is the jump measure of $L$ with intensity $\nu (dx)dt$, and $\widetilde \mu(dx,dt)= \mu(dx,dt)-\nu(dx)dt$ is the corresponding compensated jump measure. Above, $H_t(y)$ is a compound Poisson process with finite jump activity satisfying $ H_t(y)= \sum_{i=1}^{N_t(y)}\zeta_i(y)$, where $N_t(y)$ is a Poisson process with rate $\lambda(y)= \int_{|x|>y} \nu(dx)$, and $\{\zeta_i(y)\}_{i\geq 1}$ are i.i.d. and supported on $(-\infty,y)\cup(y,\infty)$ with distribution $\frac{{\bf 1}_{\{|x|\geq y\}}\nu(dx)}{\nu(|x|>y)}$. For each $y>0$, we first define the random set
which consists of all indices corresponding to intervals where no “large" jumps have occurred. For a sequence $y=y_n\to 0$, we then define an oracle analog of TRV that eliminates any increments corresponding to time intervals in which “large” jumps of $X$ occur:
To connect $\mathscr C_n(y)$ with $\widehat C_n$, we then construct an iterative sequence $\{\widetilde C_{n,j}(y)\}_{j \geq 1}$, analogous to (ref), by setting $\widetilde C_{n,0}(y):=\widehat{C}_{n,0}$ (so that the oracle sequence has the same initial value as the original sequence $\widehat C_{n,j}$) and recursively define, for $j\geq{}1$,
Though the variables $\widetilde C_{n,j}(y_n)$ and $\mathscr C_{n}(y)$ are not feasible estimators themselves, their asymptotic behavior in fact completely determines that of $\widehat C_n$ provided the auxiliary sequence $y_n$ tends to 0 at an appropriate rate. In our arguments, we demonstrate that $$ \widetilde C_{n,n+1}(y_n) \leq \widehat C_n\leq \mathscr C_{n}(y_n)+ R_n, $$ for an appropriate asymptotically negligible remainder $R_n$. The above relation allows us to analyze $\widehat C_n$ in terms of the array of oracle iterates $\{\widetilde C_{n,j}(y_n)\}_{{ j\geq 1}}$ and the oracle itself $\mathscr C_{n}(y_n)$. We then demonstrate that $\{\widetilde C_{n,j}(y_n)\}_{ j\geq 1}$ are all asymptotically equivalent to the oracle $\mathscr C_n(y_n)$ (Proposition (ref)); effectively reducing the problem to the analysis of $\mathscr C_n(y_n)$, which is considerably simpler.
We now proceed to describe the class of initial estimators we consider in our procedure. Apart from some mild regularity conditions, they are required only to be consistent for $C_T$ when the underlying process is continuous, allowing for a great deal of flexibility in the choice of initialization. More specifically, for a generic process $Y$, let
where $F:\mathbb{R}^d\to [0,\infty)$ satisfies, for some $\delta_0 \in (0,2]$ and for all $\mathbf x,\mathbf y\in \mathbb{R}^d$ with $\|\mathbf y\|\vee\|\mathbf x\|\leq 1,$
for some $K<\infty$. Above, $\|\mathbf x\|_\infty=\max_{1\leq{}i\leq{}d}|x_{i}|$. An initial estimate $\widehat C_{n,0}$ is said to belong to class $\mathcal C$ if $\widehat C_{n,0}=\widehat C_{n,0}(X)$, where $\widehat C_{n,0}(\cdot)$ is given by (ref), and satisfies
where $(\sigma\! \cdot\! W)_{t}:=\int_0^t \sigma_s dW_s$.
We now state our first main result.
Observe that the upper bounds on $r_n$ in (i)--(ii) above depend on the jump activity index $\alpha$ and become more restrictive as $\alpha$ increases. Bearing this in mind, we make the following remarks.
\
For the best possible finite-sample performance, heuristically one should set the threshold $\varepsilon$ as small as possible -- to remove as many jumps as possible -- but allow it to remain large enough so that a sufficient number increments remain to ultimately yield efficient estimates of $C_T$. From this perspective, the asymptotic lower bound on the rate $r_n$ given in the hypotheses of Theorem (ref) may appear unsatisfactory, as it precludes rates as fast as the optimal threshold $\varepsilon^\star_n\sim \sqrt{2\sigma^2 h_n\log(1/h_n)}$ in the constant volatility case. It is natural to suspect that faster rates may be possible for potential improvement in $\widehat C_n$. However, the next result shows this is not true, in general.
Note that the estimator $\widehat C_n$ defined in ((ref)) is a TRV with threshold $\varepsilon_n= \big(c_0\widehat C_n h_n \log(1/h_n)\big)^{1/2}$, which is approximately equal to $\vartheta_n$ if $\widehat C_n$ remains a consistent estimator under this threshold choice. In that case, the above result suggests that $\widehat C_n$ will not be rate-efficient in general with the threshold rate $r_n=c_0 h_n\log(1/h_n)$ and may remove too many increments even if jumps are completely absent from the process $X$. In particular, the proof of Proposition (ref) illustrates that efficiency losses can result from volatility paths that exhibit significant jumps. A natural way to remedy this is to consider localized thresholds that adapt to the volatility level. In this way, thresholds corresponding to periods of high volatility are increased, and conversely, thresholds for periods of low volatility are decreased, so as to prevent efficiency losses that might otherwise occur with uniform thresholding. This is the central motivation behind our second class of estimators, which utilize spot volatility estimates to locally tune the threshold.
To this end, for a given even integer $k_n\leq n$ and $B>0$, we define
where we set $\Delta_i^n X =0 $ if $i\leq 0$ or $i >n$. The above estimator is a type of kernel-based estimator of the spot volatility $\sigma_{t_i}^2$, as defined in fan:wang:2008,kristensen:2010, with kernel function $K(x)=\frac{1}{2}{\bf 1}_{[-1,1]}$ and bandwidth $b_n=k_n h_n$ (see jacod:protter:2011 for the asymptotic theory of the estimator in the case of one-sided uniform kernels $K(x)={\bf 1}_{[0,1]}$ and figueroa-lopez:li:2020,figueroa-lopez:wu:2022 for general kernels). Our second thresholding scheme for the localized thresholding estimator $\widehat C_n^*$ then proceeds as follows:
Let us now introduce the class of initial estimates $\mathcal C^\text{spot}$ for time-varying thresholds, which is essentially a localized analog of the class $\mathcal C$ of initial estimates defined in Section (ref). To this end, for a generic process $Y$, define
where $F:\mathbb{R}^d\to [0,\infty)$, and for convenience we set $F( \Delta_i^n Y,\ldots,\Delta_{i+d-1}^n Y)=0$ if $i\leq 0$ or $i+d-1>n$. We say the initializing threshold constants $\widehat c_{n,0}(i)$ belong to the class $\mathcal C^\text{spot}$ if $\widehat c_{n,0}(i)=\widehat c_{n,0}(i;X)$, $i=1,\ldots n$, where $\widehat c_{n,0}(i;\,\cdot\,\,)$ are of the form (ref) and satisfy
We are now in a position to state our second main result.
Statistical errors for spot volatility estimation are known to be substantially larger by comparison to the $O_P(n^{-1/2})$--sized errors that occur in estimation of integrated volatility $C_T$ (for instance, optimal choices of $k_n$ in spot volatility estimation lead to errors of order $n^{-1/4}$; see, e.g., figueroa-lopez:wu:2022,jacod:protter:2011). Interestingly enough, the estimator $\widehat C^{*}_n$ utilizes the comparatively noisier estimates of spot volatility in an auxiliary manner to lead to potentially improved estimates of $C_T$.
In this section, we compare the finite-sample performance of $\widehat C_n$ and $\widehat C_n^*$ against standard tuning approaches for TRV in the literature based on simulated data from the following stochastic volatility model:
Above, $W$ and $B$ are two correlated standard Brownian motions with covariation $d\langle W, B\rangle_t=\rho dt$, $L$ is a CGMY L\'{e}vy process independent of $W$ and $B$, and $J$ is an inhomogeneous compound Poisson process independent of all other processes with intensity $\{\lambda(t)\}_{t\geq{}0}$ and jump distribution $\varrho(dx)$.
Based on a 6.5 hour trading day and 252 trading days per year, we consider time horizons of $T\in\{\frac{1}{252}, \frac{5}{252},\frac{1}{12}\}$, corresponding to 1 day, 1 week, and 1 month, respectively, at the 5-minute ($h_n=(\frac{1}{252})(\frac{1}{6.5})(\frac{5}{60})$) sampling frequency. For illustration, we examine five separate scenarios we now describe. Unless otherwise stated, for ease of comparison the parameters for $\{\sigma_t^2\}_{t\geq{}0}$ are set as:
With these parameter choices, the annualized expected integrated variance is $(1/T)\mathbb{E} C_T = (0.2)^2$, and, in all settings, parameters are chosen so that the expected annualized realized volatility is approximately $\sqrt{(1/T)\mathbb{E} (\text{RV}_n)}\approx 0.275$ (Models 1,2,4,5, below) or $0.3$ (Model 3), which are realistic for financial data.
We compare 6 types of estimators based on TRV: two instances of standard approaches, and two instances each of the iterative estimator $\widehat C_n$ and of the localized iterative estimator $\widehat C^*_n$ with different types of initializations. Specifically, for $\widehat C_n$ we use the following estimators as initializations: $$ \text{RV}_n = \sum_{i=1}^n(\Delta_i^nX)^2,\qquad \text{BV}_n= \frac{\pi}{2}\sum_{i=2}^n |\Delta_{i-1}^n X| |\Delta_{i}^n X|. $$ For $\widehat C_n^*$, we use their localized counterparts, denoted by $$ \hat \sigma_n^2(\ell) = \frac{1}{h_nk_n}\sum_{i=\ell-k_n/2+1}^{\ell+k_n/2}(\Delta_i^n X)^2,\quad \text{BV}^{\text{spot}}_n(\ell)= \frac{1}{h_nk_n}\frac{\pi}{2}\sum_{i=\ell-k_n/2+1}^{\ell+k_n/2} |\Delta_{i-1}^n X| |\Delta_{i}^n X|. $$ We consider the following estimation procedures:
We simulate $m=5000$ paths for each Model 1-5. Denoting by $\widehat{\mathcal C}$ one of the estimators in (1)-(6), on the $j$--th realization we compute the estimator value, $\widehat{\mathcal C}_j$, the corresponding true integrated volatility, $C_{T,j}$, and report
The results are displayed in Tables (ref)-(ref); the smallest bias and MSE for each time horizon are shown in bold.
In general, we see that when jumps are present (Models 1-4), both the iterative estimator $\widehat C_n$ and localized iterative estimator $\widehat C_n^*$ can outperform the standard tuning choice (2) for $\text{TRV}$ both in terms of relative error and MSE by significant margins, with reductions in bias often by 50% or more by comparison to (2) and reductions in $\sqrt{\text{MSE}}$ as high as 40%. As anticipated, deterministic tuning (1) performs rather poorly by comparison to approaches (2)-(6) on all time horizons, and although the (non-iterative) bipower-tuned TRV in (2) leads to a substantial improvement over (1), it is uniformly outperformed by (4)-(6) on all time horizons considered and also outperformed by (3) except on daily time horizons.
In general, the localized estimators (6a,6b) have the largest relative performance gains compared to standard procedures (1)-(2) over longer time horizons, which is somewhat expected, ranging from 13%-23% reduction in $\sqrt{\text{MSE}}$ at daily horizons to 28%-40% reduction in $\sqrt{\text{MSE}}$ at monthly horizons compared to (2). Also, iterative approaches with jump-robust initializations (4,6a,6b) generally have improved performance compared to those without jump-robust initializations (3,5a,5b). Furthermore, for the localized estimators, the choice $k_n=h_n^{-0.6}$ (5b,6b) tends to lead to improvement relative to the choice $k_n=h_n^{-0.5}$ (5a,6a) over longer time horizons. The best performance in terms of both relative error and MSE is typically achieved by (6b).
Comparing performance across Models 1-4, we see that all iterative approaches (3)-(6) are generally more robust against increased levels of jump activity as well as time-varying jump behavior compared to (1)-(2). In both Models 2 and 3, on longer time horizons, the performance advantage of the localized estimators over uniform approaches is typically larger by comparison to the performance advantage they have over uniform approaches in Model 1. Though all estimators (1)-(6) have better overall performance under finite jump activity (Model 4) relative to settings with infinite activity (Models 1-3), the iterative approaches still retain performance advantages over standard choices (1)-(2) even without an infinite activity component in the model.
Turning to the jump-free case (Model 5), we note that all estimators perform very similarly in terms of both bias and MSE and are typically slightly negatively biased. Over weekly and monthly time horizons, the localized estimators (5-6) incur a very slight increase in bias (appx. 0.15%) compared to uniform thresholding (2), and the deterministic TRV has marginally smaller $\sqrt{\text{MSE}}$ compared to (2)-(6).
We note that although a slight increase in bias occurs in the localized estimators (5)-(6) in the absence of jumps, it is relatively small relative to the potential performance gain one may attain if jumps are present. Since jumps are generically expected in many types of data, for use in practice we recommend the localized estimator with jump-robust initialization and settings of (6b). However, if a simpler implementation is desired, or one wants to avoid the potential marginal additional bias when jumps are absent, method (4) is a reasonable alternative. We remark that in any case, these choices (4,6b) in the presence of jumps can significantly outperform the common choice in the literature (2).
Regarding computational considerations, in Table (ref) we report the empirical distribution of the number of iterations required for stabilization for both the localized thresholding and uniform thresholding approaches (i.e., $j_n$, as in (ref), and $j_n^*$ as in (ref)) across all models with jumps (Models 1-4) on 1-month time horizons. Both $\widehat C_n$ and $\widehat C_n^*$ stabilize rather quickly, with roughly 98% of all estimates stabilizing in 4 or fewer iterations, and the global and local thresholding approaches take roughly the same number of iterations. Though not included in Table (ref), in the jump-free case (Model 5), all observed instances of estimators stabilized in 3 or fewer iterations, with the vast majority taking 1 or 2; also, shorter time horizons typically required fewer iterations to stabilize in all settings.
Unreported simulation studies suggest localized estimators can have further performance gains relative to uniform thresholding approaches when the time horizon is extended or when additional inhomogeneities are incorporated into the model such as volatility jumps. Generally performance improvement of $\widehat C_n$ and $\widehat C_n^*$ relative to standard-type TRV tuning (1) and (2) becomes more dramatic as the overall proportion of jump variation increases relative to the quadratic variation of $X$, or when the activity of either jump component ($L$ or $J$) is increased, and substantive performance gains are typically observed provided at least one of these components is present. We also remark that at daily horizons, with relatively small sample size ($n=78$) there is little difference between uniform thresholding (3)-(4) and the localized thresholding (5)-(6), except for the rates $r_n$ and $r_n^*$; not included in this study is a detailed examination of the optimal choice of $k_n$, which could be of future interest, though $k_n = h_n^{-0.6}$ seems to reasonably well in most scenarios.