EconBase
← Back to paper

An Efficient Multi-scale Leverage Effect Estimator under Dependent Microstructure Noise

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

96,849 characters · 23 sections · 31 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Holistic Multi-Scale Inference of the Leverage Effect: Efficiency under Dependent Microstructure Noise

abstractThis paper addresses the long-standing challenge of estimating the leverage effect from high-frequency data contaminated by dependent, non-Gaussian microstructure noise. We depart from the conventional reliance on pre-averaging or volatility “plug-in” methods by introducing a holistic multi-scale framework that operates directly on the leverage effect. We propose two novel estimators: the Subsampling-and-Averaging Leverage Effect (SALE) and the Multi-Scale Leverage Effect (MSLE). Central to our approach is a shifted window technique that constructs a noise-unbiased base estimator, significantly simplifying the multi-scale architecture. We provide a rigorous theoretical foundation for these estimators, establishing central limit theorems and stable convergence results that remain valid under both noise-free and dependent-noise settings. The primary contribution to estimation efficiency is a specifically designed weighting strategy for the MSLE estimator. By optimizing the weights based on the asymptotic covariance structure across scales and incorporating finite-sample variance corrections, we achieve substantial efficiency gains over existing benchmarks. Extensive simulation studies and an empirical analysis of 30 U.S. assets demonstrate that our framework consistently yields smaller estimation errors and superior performance in realistic, noisy market environments.

Keywords: High-frequency data; Realized volatility; Subsampling; Variance reduction; Robust estimation

Introduction

The leverage effect, or the observed negative correlation between asset returns and their volatility changes, is a prominent stylized fact in financial econometrics black1976StudiesStockPrice, christie1982StochasticBehaviorCommon. It captures the asymmetry in volatility responses to positive and negative shocks in asset prices and is widely attributed to mechanisms such as financial leverage and asymmetric information in markets. Accurate estimation of the leverage effect is not only central to our understanding of asset price dynamics but also has critical implications for the pricing and hedging of derivative securities, especially in the presence of volatility skews.

However, obtaining reliable estimates of the leverage effect in high-frequency data is complicated by market microstructure (MMS) noise. The well-known “volatility signature plot” demonstrates a key challenge: with high-frequency observations, simple realized volatility estimators are largely biased and thus inconsistent due to noise zhou1996HighFrequencyDataVolatility, andersen2000GreatRealizations, aitsahalia2005HowOftenSample, patton2011DatabasedRankingRealised, aitsahalia2019HausmanTestPresence; simple leverage effect estimators suffer from a similar issue. Existing methods for leverage effect estimation have primarily focused on mitigating the impact of such noise by either adopting the pre-averaging technique jacod2009MicrostructureNoiseContinuous, podolskij2009BipowertypeEstimationNoisy, mykland2016DataCleaningInference under the assumption of independent and identically distributed (i.i.d.) Gaussian noise or using extra information. For instance, wang2014EstimationLeverageEffect and aitsahalia2017EstimationContinuousDiscontinuous utilize the pre-averaging methods to estimate leverage effect under i.i.d. Gaussian noise, yuan2020LeverageEffectHighfrequency adopts the parametric setting for microstructure noise proposed by li2016EfficientEstimationIntegrated that incorporates trading information to address noise, chong2024VolatilityVolatilityLeverage utilizes high-frequency short-dated option to recover spot volatility process, thereby estimating leverage effect, while some other works bandi2012TimevaryingLeverageEffects, kalnina2017NonparametricEstimationLeverage, curato2022StochasticLeverageEffect,yang2023EstimationLeverageEffect estimate leverage effect without explicitly addressing microstructure noise. While these methods represent significant advances, their reliance on i.i.d. Gaussian noise-related assumptions or auxiliary data remains restrictive in practice. Specifically, empirical evidence has consistently shown that microstructure noise exhibits serial dependence and higher-order moments jacod2017StatisticalPropertiesMicrostructure, aitsahalia2019HausmanTestPresence, li2020DependentMicrostructureNoise, da2021WhenMovingAverageModels, li2022ReMeDIMicrostructureNoise. These features not only violate standard modeling assumptions, but also exacerbate bias and variance in leverage effect estimation, underscoring the need for more flexible methodologies that can accommodate complex noise structures.

In this paper, we propose a novel multi-scale framework for estimating the leverage effect that explicitly accounts for microstructure noise exhibiting more flexible noise structures, allowing for stationary, dependent noise with nontrivial higher-order moments, which are commonly observed in empirical financial data. Specifically, we introduce two new estimators: the Subsampling-and-Averaging Leverage Effect (SALE) estimator and the Multi-Scale Leverage Effect (MSLE) estimator. In particular, the MSLE estimator aggregates a series of weighted SALE estimators computed across multiple time scales, exploiting their complementary properties to achieve improved convergence rates. Our methodology draws inspiration from the principles underlying the Two-Scale Realized Volatility (TSRV) and Multi-Scale Realized Volatility (MSRV) estimators zhang2005TaleTwoTime, zhang2006EfficientEstimationStochastic, aitsahalia2011UltraHighFrequency, but adapts and extends them for the specific task of estimating the leverage effect. Crucially, we do not merely use their methods as plug-in estimators for spot volatility; we develop a holistic multi-scale approach for the leverage effect itself. A key innovation in our construction is the use of a shifted window for estimating spot volatility, as illustrated schematically in Figure (ref) and (ref). This shift is important as it not only helps decouple noise components to achieve unbiasedness and variance reduction with respect to noise, but also fundamentally simplifies certain aspects of the multi-scale estimation procedure compared with the classical constructions in TSRV and MSRV.

The second primary contribution of this work is the explicit demonstration of multi-scale benefits in efficiency beyond mere noise mitigation. Specifically, for the noise-free setting, we show that our MSLE estimator, through a proper weighting scheme, achieves a notable reduction in asymptotic variance compared with the base estimator by exploiting the covariance structure of SALE estimators of different subsampling scales. Furthermore, in the noisy setting, this efficiency advantage becomes even more pronounced when benchmarked against the pre-averaging estimator, particularly at lower noise levels that are more representative of practical scenarios.

Recognizing a potential gap between standard theoretical assumptions and empirical reality regarding MMS noise, another major contribution of this work lies in the in-depth study of optimal weight assignment under diverse and realistic noise magnitudes. Classical asymptotic analyses often assume that noise variance is of constant order, thus dominating the shrinking latent increments as the sampling frequency increases. However, empirical evidence suggests a more complex picture. For instance, aitsahalia2019HausmanTestPresence finds that improvements in market liquidity allow employing simple volatility estimators at higher frequencies, while kalnina2008EstimatingQuadraticVariation and da2021WhenMovingAverageModels also explore the case of shrinkage MMS noise. A recent work by chong2025WhenFrictionsAre investigates the rough noise model that captures a more subtle interplay between the latent price process and noise. This implies that weighting schemes based solely on the strict noise-dominance assumption might suffer from modeling error when applied to real data. Motivated by this, we conduct a fine-grained analysis of the MSLE's asymptotic variance components across different subsampling scales and noise conditions. Moreover, implementing truly optimal weights would require precise values of asymptotic covariances between SALE estimators at sampling scales, which are infeasible in practice. To overcome this, we develop a computationally efficient approximate weighting strategy that works for a wide spectrum of noise conditions, yielding robust finite-sample performance as demonstrated in Monte Carlo simulations.

Our theoretical analysis establishes central limit theorems and stable convergence results for both the noise-free and noisy settings, demonstrating that the proposed estimators achieve consistency and asymptotic normality. Under the noise-free setting, both estimators attain the optimal convergence rate of order $n^{-1/4}$. In the presence of MMS noise, the MSLE estimator achieves a convergence rate of order $n^{-1/9}$. While this theoretical rate is slightly slower than the $n^{-1/8}$ achieved by the pre-averaging approach, our simulations consistently demonstrate that the MSLE estimator outperforms the pre-averaging method in practical settings across a wide range of noise levels and sample sizes. To support feasible inference, we construct consistent estimators for the asymptotic variances in both noise-free and noisy regimes, enabling the implementation of feasible central limit theorems. Monte Carlo simulations, employing both independent and dependent noise with various distributions, validate both the feasible and infeasible central limit theorems.

To assess the finite-sample performance of the proposed estimators, we conduct another extensive simulation study that encompasses a variety of settings with realistic time horizons and noise conditions. This study verifies the outstanding efficiency of the proposed MSLE estimator with the approximate weighting strategy. The result shows that: (i) under the noise-free setting, our method outperforms the base estimator, whereas (ii) under the noisy setting, our method consistently outperforms the pre-averaging approach, and this advantage is particularly pronounced in empirically relevant scenarios. This superior finite-sample behavior is attributable to a combination of factors, detailed in Section (ref): (i) the more advantageous noise-free asymptotic variance of the SALE estimator, (ii) the relatively smaller impact of noise under realistic settings, and (iii) the enhanced performance of MSLE over the individual SALE estimators.

To examine the performance of our estimators in practical applications, we also conduct an empirical analysis using high-frequency financial data. The high-frequency trading data of 30 assets including ETFs and individual stocks from the U.S. stock market are utilized. The leverage effects are estimated adaptively based on the microstructure noise characteristics in each period, and general negative correlation between the returns and volatility changes is verified. This study demonstrates the practical flexibility and effectiveness of the MSLE estimator in capturing the leverage effect under realistic market conditions, further validating the advantages observed in the simulation experiments.

The remaining paper is arranged as follows. Section (ref) introduces the model settings, notations, and estimators. Section (ref) states the main theoretical results, including limit theorems for the SALE and MSLE estimators under both noise-free and noisy conditions. Section (ref) discusses the issue of variance reduction, including variance approximations under both noise-free and noisy settings, and proposes practical strategies for optimizing the performance of MSLE. Section (ref) provides a detailed simulation study to validate the theoretical properties, examining the asymptotic behavior and finite-sample performance under various settings of microstructure noise. Section (ref) reports the results of an empirical study using real-world high-frequency data to demonstrate the practical utility of the proposed methods. Proofs, feasible central limit theorems, and further elaborations on Section (ref) to (ref) are provided in Supplementary Material.

Methodology

Model Settings

The noise-contaminated log-price $Y_{t_i}$ is observed at $t_i = i\Delta_n = iT/n$ for $i = 0, 1, \dots, n$:

align[align omitted — 71 chars of source]

where $X_{t_i}$ denotes the latent log-price, and $\varepsilon_i$ denotes the noise. For simplicity, for any stochastic process $V$, we denote: (i) $V_i \coloneqq V_{t_i}$, (ii) $\Delta V_i \coloneqq V_{i+1} - V_i$, and (iii) $\Delta_k V_i \coloneqq V_{i+k} - V_i$.

assumption[Underlying processes] Let both log-price process $(X_t)_{t \geq 0}$ and volatility process $(\sigma_t)_{t \geq 0}$ be Itô processes defined on a filtered probability space $(\Omega, \calF, (\calF_t)_{t \geq 0}, \bbP)$: \begin{align} X_t &= X_0 + \int_0^t \mu_s \mathrm{d} s + \int_0^t \sigma_s \mathrm{d} W_s, \\ \sigma_t &= \sigma_0 + \int_0^t a_s \mathrm{d} s + \int_0^t f_s \mathrm{d} W_s + \int_0^t g_s \mathrm{d} B_s, \end{align} where $W_t$ and $B_t$ are independent Brownian motions, and $\mu_t, a_t, f_t, g_t$ are adapted càdlàg locally bounded processes. In addition, $f_t$ and $g_t$ are Itô processes and the volatility path $\sigma_t^2$ is bounded away from zero.
assumption[Noise] Let $\{\varepsilon_i\}_{i=0}^n$ be mean-zero, identically distributed random variables independent of $\mathcal{F}$, with $\nu_k \coloneqq \bbE[\varepsilon_i^k]$ and $\nu_4 < \infty$. In terms of serial dependency, $\{\varepsilon_i\}_{i=0}^n$ are specified as either: \begin{enumerate}[label=(\alph*)] • independent ($\varepsilon_i \perp \varepsilon_j$ for all $i \neq j$); or • $q$-dependent ($\varepsilon_i \perp \varepsilon_j$ for $|i-j| > q$) and stationary up to the fourth moment. \end{enumerate}

Let $\langle X, \sigma^2 \rangle_T = \int_0^T 2\sigma_s^2 f_s \mathrm{d} s$ denote the true leverage effect parameter. For any estimator of interest, we maintain a clear distinction between its infeasible version $\widetilde{\langle X, \sigma^2 \rangle}_T$ based on latent data $\{X_i\}_{i=0}^n$, and its feasible version $\widehat{\langle X, \sigma^2 \rangle}_T$ based on observed data $\{Y_i\}_{i=0}^n$. The statistical properties of the estimator are investigated from two aspects:

enumerate• Asserting its unbiasedness with respect to noise: \begin{align} \underbrace{ \bbE\bigl( \widehat{\langle X, \sigma^2 \rangle}_T \big| \calF \bigr) - \widetilde{\langle X, \sigma^2 \rangle}_T }_{bias due to noise} = 0. \end{align} • Assessing its total variance by decomposing it into the variance due to discretization and the expected variance due to noise: \begin{align} \mathrm{Var}\bigl(\widehat{\langle X, \sigma^2 \rangle}_T \bigr) = \underbrace{ \mathrm{Var}\bigl( \widetilde{\langle X, \sigma^2 \rangle}_T \bigr) }_{variance due to discretization} + \: \bbE\bigl[\, \underbrace{ \mathrm{Var}\bigl( \widehat{\langle X, \sigma^2 \rangle}_T \big| \calF \bigr) }_{variance due to noise} \,\bigr]. \end{align}
remarkWhile some literature wang2014EstimationLeverageEffect, kalnina2017NonparametricEstimationLeverage defines the leverage effect as $\langle X, F(\sigma^2) \rangle_T = \int_0^T 2 F'(\sigma_s^2) \sigma_s^2 f_s \mathrm{d} s$ for a general function $F \in \bbC^2$, we focus on the canonical case where $F(x) = x$. Beyond notational simplicity, this allows us to concentrate on a primary challenge addressed in this work: robust estimation under complex, dependent noise, where a general $F$ function poses significant additional difficulties. However, it is worth noting that the results of our noise-free estimators (Theorems (ref) and (ref)) can be easily extended to the general case.

Estimators and The Robustness to Noise

For simplicity, the following notations are used for spot volatility estimation:

align[align omitted — 553 chars of source]

Here, $H$ represents the subsampling scale, whereas $k$ and $s$ represent the size and shift of the windows for estimating spot volatility. Specifically, the last parameter $s$ can be omitted when $s = 1$, as this work primarily focuses on this case.

figure[figure omitted — 8,590 chars of source]

All-Observation Estimator

To start with, consider an all-observation Leverage Effect (LE) estimator that directly utilizes all noisy observations,

align[align omitted — 173 chars of source]

Here, the window size $k_n$ satisfies that $k_n \to \infty$ and $k_n \Delta_n \to 0$ as $n\to \infty$. The estimator is very similar to the continuous leverage effect estimator\footnote{Since jumps are not included, the truncation term in their original estimator is omitted here.} studied by aitsahalia2017EstimationContinuousDiscontinuous. The only difference is that our estimator shifts the windows for spot volatility estimates outward by $\Delta_{n}$, as illustrated by Figure (ref) and (ref). To see the reason, consider applying their estimator directly to noisy observations, and it follows that there is a divergent bias due to noise of

align[align omitted — 222 chars of source]

when $\nu_3$ is not strictly zero. In contrast, the estimator in Equation (ref) is unbiased due to noise with a smaller variance, as described by the next Proposition (ref).

propositionUnder Assumptions (ref) and (ref)(ref), as $n\to\infty$, we have \begin{align} \bbE \big( \widehat{\langle X, \sigma^2 \rangle}^{\rm (all)}_T \big| \calF \bigr) &= \widetilde{\langle X, \sigma^2 \rangle}^{\rm (all)}_T, \\ \Delta_n^3 k_n^2 \mathrm{Var}\bigl( \widehat{\langle X, \sigma^2 \rangle}^{\rm (all)}_T \big| \calF \bigr) &\xrightarrow{p} (8 \nu_2 \nu_4 + 16 \nu_2^3 + 8 \nu_3^2) T. \end{align}

With the window shift, the bias due to noise is eliminated, and the variance due to noise is reduced, while the asymptotic behavior of the noise-free estimator $\widetilde{\langle X, \sigma^2 \rangle}_T^{\rm (all)}$ remains the same as $\widetilde{\langle X, \sigma^2 \rangle}^{[\text{AFLWY17}]}_{T}$. As a direct result of this unbiasedness, a debias step in TSRV or MSRV is no longer needed. However, the all-observation estimator is not consistent in the presence of noise: the variance due to noise is $O(n^3 k_n^2)$ while the noise due to discretization is $O(k_n^{-1} + n^{-1} k_n)$, resulting in an exploding total variance.

SALE: Subsampling-and-Averaging Estimator

To mitigate the impact of noise, we employ a subsampling procedure: a subsample of observations is denoted by a pair of integers $(H, h)$, where $H \geq 1$ and $1 \leq h \leq H$ are the scale and index of the subsample. The $j$th observation in subsample $(H, h)$ corresponds to the original index $j_{H,h} = jH + h - 1$, where $j = 0, 1, \dotsc, n_{H,h}$ and $n_{H,h} = \lfloor (n - h + 1)/H \rfloor$. Assume that $k_n\to\infty$, $H_n\to\infty$ and $k_nH_n\Delta_n \to 0$ as $n\to\infty$. A subsampling estimator is constructed by applying the all-observation estimator to a subsample of observations:

align[align omitted — 208 chars of source]

As a variant of the all-observation estimator, the subsampling estimator is not consistent. Nor is it statistically sound, as it fails to utilize the full information in the observed data. Therefore, by taking average over the subsampling estimators with the same scale, we obtain an Subsampling-and-Averaging Leverage Effect (SALE) estimator:

align[align omitted — 288 chars of source]

which can be viewed as summation of overlapping sparse increments. The SALE estimator is consistent under Assumption (ref) with proper choices of $k_n$ and $H_n$. This is primarily attributed to the fact that the correlation due to noise between different subsampling estimators in Equation (ref) can be controlled by the dependence level of noise. Specifically, if Assumption (ref)(ref) holds, this correlation becomes zero. Thus, the variance due to noise significantly reduces, as the following proposition describes.

propositionSuppose that Assumptions (ref) and (ref)(ref) hold, and that $H_n > 2q$. Define the generalized autocorrelation functions (ACFs) of noise for any $l \in \bbZ$ as \begin{align} \rho_2(l) &= \mathrm{Corr}(\varepsilon_i, \varepsilon_{i+l}) = \frac{\bbE [ \varepsilon_{i}\varepsilon_{i+l} ]}{\nu_2}, \\ \rho_3(l) &= \mathrm{Corr}(\varepsilon_i, \varepsilon_{i+l}^2) = \frac{\bbE [ \varepsilon_{i}\varepsilon_{i+l}^{2} ]}{\sqrt{\nu_2(\nu_4 - \nu_2^2)}}, \\ \rho_4(l) &= \mathrm{Corr}( \varepsilon_{i}^{2}, \varepsilon_{i+l}^{2} ) = \frac{\bbE[ \varepsilon_{i}^{2} \varepsilon_{i+l}^{2} ] - \nu_2^2}{\nu_4 - \nu_2^2}. \end{align} As $n\to\infty$, we have \begin{gather} \Delta_n^3 H_n^4 k_n^2 \mathrm{Var} \bigl( \widehat{\langle X, \sigma^2 \rangle}^{(H_n)}_T \big| \calF \bigr) \xrightarrow{p} \Phi T, \\ where \Phi = 8 \nu_2 (\nu_4 - \nu_2^2) \sum_{l=-q}^{q} \bigl( \rho_2(l)\rho_4(l) + \rho_3(l) \rho_3(-l) \bigr) + 24 \nu_2^3 \sum_{l=-q}^{q} \rho_2^3(l). \end{gather} Specifically, under Assumptions (ref) and (ref)(ref), Equation (ref) holds with $\Phi = 8 \nu_2 \nu_4 + 16 \nu_2^3 + 8 \nu_3^2$.

The SALE estimator has a variance due to noise of $O(n^3 H_n^{-4} k_n^{-2})$, and a variance due to discretization of $O(k_n^{-1} + n^{-1} H_n k_n)$. Suppose $H_n \propto n^a$ and $k_n \propto (n/H)^b$, the optimal total variance $O(n^{-1/7})$ is achieved at $a=5/7$ and $b=1/2$. Despite being consistent, this is far from the optimal rate $O(n^{-1/4})$ of pre-averaging approach.

MSLE: Multi-Scale Estimator

Consider a set of scales $1 \leq H_1 < \dots < H_{M_n} \leq n^a$ for some $a\in(0,1)$, where $M_n>0$. An Multi-Scale Leverage Effect (MSLE) estimator is defined as a weighted average of SALE estimators at different scales:

align[align omitted — 160 chars of source]

where the weight vector $\bm w=(w_1, \dotsc, w_{M_n})$ satisfies $\bm w^T \bm 1_{M_n} = 1$ with $\|\bm w\|_1$ bounded.

When the noise is i.i.d., the covariance due to noise between a pair of SALE estimators are non-zero only when a scale is double the other, and the corresponding correlation coefficient is small (for example, 0.1 for Gaussian noise). For serial dependent noise, the case is more complicated and analytical results are hard to derive. Instead, numerical calculation is available (see Supplementary Material).

propositionUnder Assumptions (ref) and (ref)(ref), and suppose $k_p = \lfloor \beta \lfloor n/H_p\rfloor^b \rfloor$ holds for all $p\in\{1, \dotsc, M_n\}$ with some constant $\beta>0$ and $b\in(0,1)$. For any $p, q \in \{1, \dotsc, M_n\}$, as $n\to\infty$, we have \begin{gather} \Delta_n^3 H_p^2 H_q^2 k_p k_q \mathrm{Cov} \bigl( \widehat{\langle X, \sigma^2 \rangle}^{(H_p)}_T, \widehat{\langle X, \sigma^2 \rangle}^{(H_q)}_T \big| \calF \bigr) \xrightarrow{p} F_{p,q} T. \\ where F_{p,q} = (8\nu_2\nu_4 + 16\nu_2^3 + 8\nu_3^2) 1_{\{p=q\}} + 2\nu_2(\nu_4-\nu_2^2) (1_{\{H_p/H_q=2\}} + 1_{\{H_q/H_p=2\}}), \end{gather} and thus \begin{align} \frac {\mathrm{Var} \bigl( \widehat{\langle X, \sigma^2 \rangle}^{\rm (MS)}_T \big| \calF \bigr)} { \Delta_n^{-3} \sum_{p=1}^{M_n} \sum_{q=1}^{M_n} \frac{w_p}{H_p^2 k_p} \cdot F_{p,q} \cdot \frac{w_q}{H_q^2 k_q}} \xrightarrow{p} T. \end{align}
remarkThe “double scale” terms in $F_{p,q}$ can be removed by using a further shifted spot volatility window with $s=2$ in Equation (ref), as illustrated in Figure (ref). A similar proposition can be established with $F_{p,q} = (8\nu_2\nu_4 + 8\nu_2^3 + 8\nu_3^2) 1_{\{p=q\}}$, eliminating the cross terms and reducing the variance.

The MSLE estimator effectively reduces the variance due to noise. For example, setting $H_p=p$ for $p=1, \dotsc, M_n$, $M_n = \lfloor n^a \rfloor$, and $w_p \propto p^{4-2b}$, the variance due to noise is $O(n^{3-5a-2b+2ab})$, while the variance due to discretization is $O(n^{-(1-a)(b \land (1-b))})$. Thus, by selecting $a=5/9$ and $b = 1/2$, an optimal convergence rate of $n^{1/9}$ is achieved for MSLE, close to the optimal rate $n^{1/8}$ of pre-averaging approach.

Main Results

Central Limit Theorems for SALE

We start by establishing the following theorem for the noise-free version of SALE. Two scenarios for the scale $H_n$ are considered: either it is fixed, or it goes to infinity as $n \to \infty$. Hereafter, we use $\xrightarrow{\rm st}$ to denote stable convergence in law.

assumptionSuppose that $H_n$ and $k_n$ satisfy one of the following conditions: \begin{enumerate}[label=(\alph*)] • $H_n=H$ is a given positive integer, $k_n = \lfloor \beta \lfloor n/H \rfloor^b \rfloor$ for some $\beta > 0$ and $b \in (0,1)$. • $H_n = \lfloor \alpha n^a \rfloor$, $k_n = \lfloor \beta \lfloor n/H_n \rfloor^b \rfloor$ for some $\alpha, \beta > 0$ and $a, b \in (0,1)$. \end{enumerate}
theorem\begin{enumerate}[label=(\arabic*)] • • Under Assumptions (ref) and (ref)(ref), let $u_n = \sqrt{k_n \land (k_n H \Delta_n)^{-1}}$. There exist a standard Brownian motion $(W_{1,t})_{t\geq 0}$ independent of $\calF$ and a predictable process $(\zeta_{1,t})_{t\geq 0}$ such that, as $n\to\infty$, \begin{gather} u_n \bigl( \widetilde{\langle X, \sigma^2 \rangle}^{(H)}_T - \langle X, \sigma^2 \rangle_T \bigr) \xrightarrow{\rm st} \int_0^T \zeta_{1,t} \mathrm{d} W_{1,t}, \\ \int_0^T \zeta_{1,t}^2 \mathrm{d} t = \frac{u_n^2}{k_n} \left(\frac{8}{3} + \frac{4}{3H^2}\right) \int_0^T \sigma_s^6 \mathrm{d} t + u_n^2 k_n H \Delta_n \frac{2}{3} \int_0^T \sigma_t^2 \mathrm{d} \langle \sigma^2, \sigma^2 \rangle_t. \end{gather} • Under Assumptions (ref) and (ref)(ref), there exist a standard Brownian motion $(W_{1,t})_{t\geq 0}$ independent of $\calF$ and a predictable process $(\zeta_{1,t})_{t\geq 0}$ such that, as $n\to\infty$, \begin{gather} n^{\frac{1}{2}(1-a)(b\land (1-b))} \bigl( \widetilde{\langle X, \sigma^2 \rangle}^{(H_n)}_T - \langle X, \sigma^2 \rangle_T \bigr) \xrightarrow{\rm st} \int_0^T \zeta_{1, t} \mathrm{d} W_{1, t}, \\ \int_{0}^{T} \zeta_{1, t}^{2} \mathrm{d} t = \frac{8 \alpha^b}{3 \beta} \int_0^T \sigma_t^6 \mathrm{d} t \cdot 1_{(0, 1/2]}(b) + \frac{2 \alpha^{1-b} \beta T}{3} \int_0^T \sigma_t^2 \mathrm{d} \langle \sigma^2, \sigma^2 \rangle_t \cdot 1_{[1/2, 1)}(b). \end{gather} \end{enumerate}
remarkTaking $H=1$ in Theorem (ref)(ref) yields the central limit theorem for the all-observation estimator, which has a same asymptotic variance as the continuous leverage effect estimator in aitsahalia2017EstimationContinuousDiscontinuous.
remarkSimilar to existing work on leverage effect estimation wang2014EstimationLeverageEffect, aitsahalia2014HighFrequencyFinancialEconometrics, aitsahalia2017EstimationContinuousDiscontinuous, kalnina2017NonparametricEstimationLeverage, yang2023EstimationLeverageEffect, the asymptotic variance is determined by the spot volatility estimation, which consists of two sources of errors: the price variation error and the volatility variation error aitsahalia2017EstimationContinuousDiscontinuous, corresponding to the first and second terms in Equation (ref) or (ref). Intuitively, increasing $k_n$ leads to a wider spot volatility estimation window and thus a larger sample size for that estimation, which reduces the price variation error. On the other hand, this increases the volatility variation error, because the estimated volatility becomes less “local”. The optimal choice of $k_n$ is thus a trade-off between these two sources of error.

Next, we establish the following theorem for the noisy version of SALE. Notably, Assumption (ref)(ref) is not considered, as a finite $H$ leads to a divergent variance due to noise and thus an inconsistent estimator.

theoremUnder Assumptions (ref), (ref)(ref) and (ref)(ref), suppose that $H_n > 2q$ and $4a + 2b - 2ab > 3$. Let $r = [(1-a)(b\land (1-b))] \land [4a+2b-2ab-3]$ and let $\Phi$ be as defined in Equation (ref). There exist a standard Brownian motion $(W_{2, t})_{t\geq 0}$ independent of $\calF$ and a predictable process $(\zeta_{2, t})_{t\geq 0}$, such that, as $n\to\infty$, \begin{align} & n^{\frac{1}{2}r} \bigl( \widehat{\langle X, \sigma^2 \rangle}^{(H_n)}_T - \langle X, \sigma^2 \rangle_T \bigr) \xrightarrow{\rm st} \int_{0}^{T} \zeta_{2, t} \mathrm{d} W_{2, t}, \\ \int_{0}^{T} \zeta_{2, t}^{2} \mathrm{d} t & = \frac{8 \alpha^{b}}{3 \beta} \int_{0}^{T} \sigma_{t}^{6} \mathrm{d} t \cdot 1_{\{(1-a)b\}}(r) \notag \\ & \qquad + \frac{2 \alpha^{1-b} \beta T}{3} \int_{0}^{T} \sigma_{t}^{2} \mathrm{d} \langle \sigma^{2}, \sigma^{2} \rangle_t \cdot 1_{\{(1-a)(1-b)\}}(r) \notag \\ & \qquad + \frac{1}{\alpha^{4-2b} \beta^2 T^3} \int_{0}^{T} \Phi \mathrm{d} t \cdot 1_{\{4a+2b-2ab-3\}}(r). \end{align}

Central Limit Theorems for MSLE

To establish the limit theorems for MSLE, the following conditions on scales, window sizes and weights are established, and the covariances due to discretization between SALE estimators are given by Proposition (ref) and (ref).

assumptionThe scales $\{H_p\}_{p=1}^{M_n}$ satisfy $1 \leq H_1 < \dotsc < H_{M_n} \leq n^a$ for some $a\in(0,1)$. The window sizes are $k_p = \lfloor \beta \lfloor n/H_p \rfloor^b \rfloor$ for all $p \in \{1, \dotsc, M_n\}$ for some $\beta > 0$ and $b \in (0, 1)$. The weight vector $\bm w=(w_1, \dotsc, w_{M_n})$ satisfies that $\bm w^T \bm 1_{M_n} = 1$ and that $\|\bm w\|_1$ is bounded.
propositionSuppose that Assumptions (ref) and (ref) hold. For any $1 \leq q \leq p \leq M_n$, let $u_n = \sqrt{k_p \land (k_p H_p \Delta_n)^{-1}}$. There exist $v_{p,q}^{(1)}, v_{p,q}^{(2)} \in [0, 1)$ that depend on $H_p, H_q$ and $n$, such that, as $n\to\infty$, \begin{gather} u_n \begin{pmatrix} \widetilde{\langle X, \sigma^2 \rangle}^{(H_p)}_T - \langle X, \sigma^2 \rangle_T \\ \widetilde{\langle X, \sigma^2 \rangle}^{(H_q)}_T - \langle X, \sigma^2 \rangle_T \end{pmatrix} \xrightarrow{\rm st} \calN\left(0, u_n^2 \begin{bmatrix} \Sigma_{p,p}^{\mathrm{(disc)}} & \Sigma_{p,q}^{\mathrm{(disc)}} \\ \Sigma_{p,q}^{\mathrm{(disc)}} & \Sigma_{q,q}^{\mathrm{(disc)}} \end{bmatrix}\right), \\ where \Sigma_{p,q}^{\mathrm{(disc)}} = \frac{1}{k_p} \cdot 4 v_{p,q}^{(1)} \frac{H_q}{H_p} \cdot \int_0^T \sigma_t^6 \mathrm{d} t + k_pH_p\Delta_n \cdot \frac{2}{3} v_{p,q}^{(2)} \left(\frac{k_qH_q}{k_pH_p}\right)^2 \cdot \int_0^T \sigma_t^2 \mathrm{d} \langle \sigma^2, \sigma^2 \rangle_t, \end{gather} and the limiting process is independent of $\calF$.
remarkFactors $v_{p,q}^{(1)}$ and $v_{p,q}^{(2)}$, arising from the grid structures of scales $H_p$ and $H_q$, contribute to price variation error and volatility variation error, respectively. Their definitions are detailed in Supplementary Material.
propositionSuppose that the conditions of Proposition (ref) hold. Consider two sequences of scales $H_p$ and $H_q$ indexed by $n$, satisfying that (i) $H_p \geq H_q$ for all $n$, (ii) as $n\to\infty$, $H_p, H_q \to \infty$ and $H_{q} / H_{p} \to \rho$ for some constant $\rho \in (0, 1]$. Then, as $n\to\infty$, we have \begin{align} v_{p,q}^{(1)} \to 1 - \frac{\rho}{3} \quad and \quad v_{p,q}^{(2)} \to 1. \end{align}

For the asymptotic behavior of MSLE, a specific set of consecutive scales are considered, and the weights are defined with a continuous bounded function.

assumptionSuppose that $\{H_p\}_{p=1}^{M_n}$ and $\bm{w}$ satisfy the following conditions: \begin{enumerate}[label=(\alph*)] • $H_p = m_n+p$ for all $p \in \{1, \dotsc, M_n\}$. Defining $H_n^* = H_{M_n}$, the sequences of positive integers $m_n$ and $M_n$ are selected such that, as $n\to\infty$, $H_n^* / n^a \to \alpha$, $m_n / H_n^* \to c$ for some constants $\alpha > 0$, $a\in(0,1)$, $c\in(0, 1)$. • $w_p = \frac{1-c}{M_n}\phi(c+\frac{p}{H_n^*})$ for all $p \in \{1, \dotsc, M_n\}$, where $\phi: [c, \infty) \to \bbR$ is a continuous bounded function satisfying that $\int_c^1 \phi(x) \mathrm{d} x = 1$. \end{enumerate}
theoremSuppose that Assumptions (ref), (ref) and (ref) hold. There exist a standard Brownian motion $(W_{3, t})_{t\geq 0}$ independent of $\calF$ and a predictable process $(\zeta_{3, t})_{t\geq 0}$, such that, as $n\to\infty$, \begin{align} & \quad n^{\frac{1}{2}(1-a)(b\land (1-b))} \bigl( \widetilde{\langle X, \sigma^2 \rangle}^{\rm (MS)}_T - \langle X, \sigma^2 \rangle_T \bigr) \xrightarrow{\rm st} \int_0^T \zeta_{3, t} \mathrm{d} W_{3, t}, \\ \int_0^T \zeta_{3, t}^{2} \mathrm{d} t &= \frac{8\alpha^b}{\beta} \int_c^1 \int_c^x \phi(x) \phi(y) x^b \left(\frac{y}{x} - \frac{y^2}{3x^2}\right) \mathrm{d} y \mathrm{d} x \cdot \int_0^T \sigma_t^6 \mathrm{d} t \cdot 1_{(0, 1/2]}(b) \notag \\ & \quad + \frac{4\alpha^{1-b}\beta T}{3} \int_c^1 \int_c^x \phi(x) \phi(y) \frac{y^{2(1-b)}}{x^{1-b}} \mathrm{d} y \mathrm{d} x \cdot \int_0^T \sigma_t^2 \mathrm{d} \langle \sigma^2, \sigma^2 \rangle_t \cdot 1_{[1/2, 1)}(b). \end{align}
theoremUnder Assumptions (ref), (ref)(ref), (ref) and (ref), suppose that $5a+2b-2ab > 3$, and let $r = [(1-a)(b\land (1-b))] \land [5a+2b-2ab-3]$, $F_1 = 8\nu_2\nu_4 + 16\nu_2^3 + 8\nu_3^2$ and $F_2 = 2\nu_2(\nu_4-\nu_2^2)$. There exist a standard Brownian motion $(W_{4, t})_{t\geq 0}$ independent of $\calF$ and a predictable process $(\zeta_{4, t})_{t\geq 0}$, such that, as $n\to\infty$, \begin{align} & \qquad \qquad n^{\frac{1}{2}r} \bigl( \widehat{\langle X, \sigma^2 \rangle}^{\rm (MS)}_T - \langle X, \sigma^2 \rangle_T \bigr) \xrightarrow{\rm st} \int_0^T \zeta_{4, t} \mathrm{d} W_{4, t}, \\ \notag \int_0^T \zeta_{4, t}^{2} \mathrm{d} t &= \frac{8\alpha^b}{\beta} \int_c^1 \int_c^x \phi(x) \phi(y) x^b \left(\frac{y}{x} - \frac{y^2}{3x^2}\right) \mathrm{d} y \mathrm{d} x \cdot \int_0^T \sigma_t^6 \mathrm{d} t \cdot 1_{\{(1-a)b\}}(r) \\ & \quad \notag + \frac{4\alpha^{1-b}\beta T}{3} \int_c^1 \int_c^x \phi(x) \phi(y) \frac{y^{2(1-b)}}{x^{1-b}} \mathrm{d} y \mathrm{d} x \cdot \int_0^T \sigma_t^2 \mathrm{d} \langle \sigma^2, \sigma^2 \rangle_t \cdot 1_{\{(1-a)(1-b)\}}(r) \\ & \quad \notag + \frac{1}{\alpha^{5-2b}\beta^2 T^3} \int_c^1 \phi^2(x) x^{-(4-2b)} \mathrm{d} x \cdot \int_0^T F_1 \mathrm{d} t \cdot 1_{\{5a+2b-2ab-3\}}(r) \\ & \quad + \frac{1_{(0, 1/2]}(c)}{2^{2-b}\alpha^{5-2b}\beta^2 T^3} \int_c^{1/2} \phi(x) \phi(2x) x^{-(4-2b)} \mathrm{d} x \cdot \int_0^T F_2 \mathrm{d} t \cdot 1_{\{5a+2b-2ab-3\}}(r). \end{align}

Note that the asymptotic variances in Theorems (ref) to (ref) are unobservable. Their consistent estimators and feasible central limit theorems are detailed in Supplementary Material.

Practical Aspects: Variances and Weights

Asymptotic Variances in Practice

Accurate asymptotic variance is important for parameter tuning. For SALE, it helps pin down the optimal scale; for MSLE, it helps decide the optimal weight distributions. Despite theoretical correctness, the accuracy of derived variances may be affected by two situations in practice: (i) small noise and (ii) violation of conditions in Proposition (ref).

The asymptotic variances due to noise in Proposition (ref), (ref) and (ref) are established based on non-shrinkaging noise assumptions. However, a small noise correction could be necessary in practice, as some terms of small order become more pronounced as noise becomes smaller. For all-observation estimators, we have

align[align omitted — 423 chars of source]

Figure (ref) compares the simulated performance of Equation (ref) and Equation (ref). Similar correction for SALE estimators are provided in Supplementary Material.

figure[figure omitted — 820 chars of source]

As for asymptotic variance due to discretization, the limit expression in Equation (ref) does not work for $H_q / H_p \to 0$, and can be inaccurate when any of $n, H_p, H_q$ is not large enough (Theorem (ref)(ref) is an example with fixed $H_p = H_q$). Consequently, Proposition (ref) will improve the accuracy of asymptotic variances in such cases.

Approximate Weights of MSLE Estimators

Suppose that Assumption (ref) holds. Let $\bm{\Sigma} \in \bbR^{M_n \times M_n}$ denote the total asymptotic covariance matrix between scales, the optimal weight assignment can be obtained by solving

align[align omitted — 179 chars of source]

The solution and the corresponding minimum are given by

align[align omitted — 333 chars of source]

However, direct application of Equation (ref) faces challenges in practice: (i) the covariance matrix cannot be observed and therefore estimated values are needed; (ii) the solution $\bm w^*$ can be numerically unstable, sensitive to estimation errors of $\bm \Sigma$; and (iii) calculating Equation (ref) requires matrix inversion, which has an expensive time complexity of $O(M_n^3)$. To address these challenges, we construct approximate weights for MSLE estimators. For simplicity, hereafter in this section, we suppose that Assumption (ref)(ref) holds, and let

align[align omitted — 257 chars of source]

Detailed derivations for Equations (ref) and (ref) presented in this section can be found in Supplementary Material.

Approximation in the Noise-Free Case

In the absence of noise, the MSLE estimator can be used to enhance the statistical efficiency of the all-observation estimator by employing a more optimal weight assignment, rather than only allocating all weight to the $H=1$ scale. For generality, Definition (ref) provides a closed-form expression for the approximate weights applicable to all scales within $(m_n, m_n+M_n]$. It is obtained by taking the limit $m_n \to \infty$ in Equation (ref), where $\bm\Sigma$ is approximated by using Proposition (ref) and assuming that $s_2 \ll s_1$.

definition[Approximate weights] Suppose that Assumptions (ref) and (ref)(ref) hold. The approximate weights are given by $\bm{\widetilde w} = (\bm 1_{M_n}^\sfT \bm{\widetilde\omega})^{-1} \bm{\widetilde\omega}$, where \begin{align} \bm{\widetilde\omega}_p = \begin{cases} 2(m_n+1)^{-1/2}, & p = 1, \\ (m_n+p)^{-3/2}, & p = 2, \dots, M_n-1, \\ (m_n+M_n)^{-1/2}, & p = M_n. \end{cases} \end{align}

Apart from being computationally and statistically efficient, $\bm{\widetilde{w}}$ is numerically stable: since $\widetilde{w}_p > 0$ holds for any $p = 1, \dots, M_n$, we always have $\sum_{p=1}^{M_n} |\widetilde{w}_p| = \sum_{p=1}^{M_n} \widetilde{w}_p = 1$.

Approximation in the Noisy Case

Let $\varphi(x) = x^{-3/2} \phi(x)$, $\gamma=s_1/(3(s_1+s_2))$, $\lambda=-(s_1+s_2)/s_3$, and

align[align omitted — 178 chars of source]

The optimal $\phi(x)$ in Assumption (ref)(ref) is related to a Fredholm integral equation:

align[align omitted — 106 chars of source]

where $k \in \bbR$ is a constant such that $\int_c^1 \varphi(x) x^{3/2} \mathrm{d} x = 1$. Equation (ref) relies on several simplifications: (i) the conditions of Theorem (ref) hold with $a=5/9$ and $b=1/2$, corresponding to the optimal convergence rate of MSLE estimators; (ii) the sparse off-diagonal terms of covariance due to noise are omitted, as explained in Section (ref); and (iii) finite-sample corrections in Section (ref) are not considered. These simplifications are made to isolate the dominant asymptotic behavior for analytical tractablity.

figure[figure omitted — 1,907 chars of source]

Figure (ref) shows the impact of the noise level in this simplified situation. The asymptotic variances of SALEs and the optimal weights of the MSLE are presented on a fixed set of scales. Note that a smaller $|\lambda|$ represents a larger noise magnitude. For a small $|\lambda|$, the variances due to noise dominate in most scales. Specifically, as $|\lambda| \to 0$, the integral term in Equation (ref) vanishes, and the optimal weights are given by $\phi(x) \to 4x^3$. Despite having a closed-form expression, the result is not useful in practice, as noise always dominates the total variance, so the estimation error is too large. On the other hand, as $|\lambda| \to \infty$, noise has negligible contributions to the total variances, and Equation (ref) becomes ill-posed. However, the proposed weights in Definition (ref) offer good approximations in this case.

Beyond these simplifications, the intuition behind our multi-scale approach and approximate weighting strategy is illustrated by the signature plot in Figure (ref), which compares different estimators across various scales (or pre-averaging window lengths) for a simulated path under realistic noise. While illustrative, these patterns are systematic and confirmed by the extensive simulations in Section (ref).

figure[figure omitted — 263 chars of source]

The plot reveals two key advantages. First, SALE outperforms the pre-averaging estimator, as its minimum asymptotic variance (top panel) is smaller, leading to a more efficient optimal estimator.\footnote{The fundamental reason is that the SALE estimator has a smaller coefficient of price variation error as shown in Theorem (ref), which is $8/3$, compared to a larger coefficient $4$ in the pre-averaging estimator.} Second, MSLE can improve upon SALE by averaging SALE estimates across an appropriate range of scales (bottom panel) using an effective weighting strategy, yielding a more accurate estimate with a tighter confidence interval.

Based on this, we propose the following weighting method: (i) find an optimal scale $\overline{H}_n$ for SALE estimators by minimizing the total variance, using the corrections in Section (ref); and (ii) allocate weights by applying Definition (ref) with $m_n = \overline{H}_n - 1$. This data-driven strategy ensures that MSLE leverages the most informative scales, improving upon the optimal SALE in two ways: the variance due to discretization is reduced through the approximate weights, and the variance due to noise is reduced because SALE estimators at subsequent scales exhibit smaller noise contributions.

Monte Carlo Simulations

Data Generating Processes

The Heston model heston1993ClosedFormSolutionOptions is used to generate the discrete values of the underlying continuous processes. The model is defined as

align[align omitted — 272 chars of source]

where $W_{t}$ and $B_{t}$ are independent Brownian motions, with the leverage effect being $\langle X, \sigma^2 \rangle_T = \gamma \rho \int_{0}^{T} V_{t} \mathrm{d} t$. The parameters are set as follows: $\mu=0.02, \kappa=5, \theta=0.04, \gamma=0.5, \rho=-0.7$. The initial values are set as $X_0 = 0, V_0 = 0.02$.

Let the variance of noise random variables $\{\varepsilon_i\}_{i=0}^n$ be $\varsigma^2$. For independent noises, three distributions are considered: (i) normal: $\varepsilon_i \sim \calN(0, \varsigma^2)$; (ii) uniform: $\varepsilon_i \sim \mathrm{Unif}(-\sqrt{3}\varsigma, \sqrt{3}\varsigma)$; and (iii) skew-normal: $\varepsilon_i$ has a PDF of $f(x) = 2\omega^{-1}\phi_0(\omega^{-1}(x-\xi)) \Phi_0(\alpha \omega^{-1} (x-\xi))$, where $\phi_0$ and $\Phi_0$ are the PDF and CDF of $\calN(0, 1)$, and $\xi=-\omega\delta\sqrt{2/\pi}$, $\omega=\varsigma (1-2\delta^2/\pi)^{-1/2}$, $\delta=\alpha (1+\alpha^2)^{-1/2}$, with the shape parameter $\alpha = 1$. For dependent noises, consider

align[align omitted — 467 chars of source]

Specifically, the AR(1) process is included to evaluate the robustness of the proposed estimators. While AR(1) noise is not $q$-dependent and thus technically violates Assumption (ref), it serves as a benchmark model for persistent, serially correlated noise jacod2017StatisticalPropertiesMicrostructure, li2022ReMeDIMicrostructureNoise. As our subsequent simulations will confirm, the proposed estimators are indeed robust to this moderate violation, preserving their asymptotic normality and superior finite-sample efficiency.

Asymptotic Normality

This section validates the central limit theorems by examining the distribution of standardized estimation errors. For each estimator, the error is standardized using both its infeasible and feasible asymptotic variance. According to the results in Section (ref), these standardized errors should converge to a standard normal distribution.

We simulate 5000 paths for each scenario, covering noise-free, independent noise, and dependent noise. For data generation, we set $T=1/252$, $n=23400$, $\varsigma = 0.005$, $\theta_1=\pm 0.7$, $\theta_2=0.5$, and $\phi=0.7$. For estimators, we set $\beta=1/2$, $b=1/2$, with scales and weights detailed in Table (ref).

table[table omitted — 632 chars of source]

Table (ref) presents the summary statistics for the standardized errors. Across all scenarios, these statistics closely match those of a standard normal distribution, corroborating our theoretical results (Theorems (ref) to (ref) and their feasible versions). Additional Q-Q plots in the Supplementary Material further support these findings.

table[table omitted — 3,762 chars of source]

Finite-Sample Performance: Superior Efficiency

Efficiency in the Noise-Free Case

The finite-sample efficiency of the MSLE estimator, using both optimal and approximate weights, is compared against the all-observation estimator. To evaluate the performance across different sample sizes, four common time horizons are considered for $T$: one day $(T=1/252)$, one week $(T=5/252)$, two weeks $(T=10/252)$ and one month $(T=22/252)$. For each $T$, the sample size is set to $n = 23400 \times 252T$, and 1000 paths are simulated. The MSLE estimators are computed with scales $H_p = 1, 2, \dots, \lfloor 0.5 n^{0.5}\rfloor$, and the optimal weights are given by Equation (ref) and (ref).

figure[figure omitted — 2,377 chars of source]
table[table omitted — 1,714 chars of source]

Figure (ref) and Table (ref) summarize the results. The findings clearly demonstrate the superiority of the MSLE estimator over the all-observation estimator, evidenced by its smaller asymptotic variance, lower finite-sample RMSE, and higher finite-sample efficiency.

A key practical insight is that the MSLE with approximate weights achieves efficiency nearly identical to that with optimal weights, but offers significantly better numerical stability. By construction, the approximate weights are non-negative, thus ensuring $\|\bm{w}\|_1 = 1$. In contrast, the optimal weights can become negative, causing their $L^1$-norm to grow with the sample size. This not only increases numerical instability, but also potentially violates our theoretical requirement in Assumption (ref), making the approximate weighting scheme a more robust choice for practical implementation.

Efficiency in the Noisy Case

The finite-sample efficiency of the MSLE estimator using approximate weights is compared against the pre-averaging LE estimator in aitsahalia2017EstimationContinuousDiscontinuous in a noisy setting. While the pre-averaging estimator has a slightly faster theoretical convergence rate ($n^{-1/8}$) than the MSLE estimator ($n^{-1/9}$), we demonstrate that MSLE achieves superior finite-sample efficiency, especially in realistic scenarios. To this end, we again vary $T$ to assess the performance across different sample sizes.

The simulation uses dependent AR(1) noise with $\phi = 0.7$, and three noise levels: small ($\varsigma=10^{-4}$), medium ($\varsigma=10^{-3.5}$), and large ($\varsigma=10^{-3}$). Notably, empirical evidence suggests that real-world noise levels are closer to the “small” setting christensen2014FactFrictionJumps, which is further supported by our empirical study in Section (ref). For the MSLE estimator, the noise ACF is truncated at $q=3$, the scales are set to $H_p=7, 8, \dots, \lfloor n^{5/9} \rfloor$, and for simplicity, the same weight allocation is used for all paths in each case. For the pre-averaging estimator, since a closed-form optimal tuning parameter is unavailable, we grant it an advantage by ex-post selecting the pre-averaging window that yields the minimum RMSE, from a wide grid of candidates ($5, 10, 30, 60, 90, 120, 180, 240,$ and $300$).

figure[figure omitted — 1,870 chars of source]
table[table omitted — 2,314 chars of source]

Figure (ref) (for $\varsigma=10^{-4}$) and Table (ref) (for all noise levels) present the results. The findings confirm that the MSLE estimator consistently and substantially outperforms the pre-averaging estimator in terms of finite-sample RMSE and efficiency across all sample sizes and noise levels. Crucially, the advantage is most pronounced in the small-noise setting, which is the most empirically relevant scenario. Even as the time horizon increases to one month, MSLE's lead remains significant, demonstrating that the theoretical convergence rate is not the only determinant of the finite-sample performance. This highlights the practical power of the proposed estimators and the approximate weighting strategy. Furthermore, this superior performance is achieved without resorting to the infeasible ex-post parameter tuning that was granted to the pre-averaging estimator, underscoring the robustness and practical utility of our methods.

An additional study for the i.i.d. noise case, along with supplementary information of the simulation details, are provided in Supplementary Material.

Empirical Study

The high-frequency trading data for a selection of assets, covering the regular trading hours from 2014 to 2023 (2,516 trading days), are collected from the TAQ database. The data are cleaned before analysis, retaining only regular trades and removing erroneous entries.\footnote{A practical and detailed guideline on high-frequency data cleaning is offered by barndorffnielsen2009RealizedKernelsPractice. While we follow most of the steps therein, some are omitted. For example, the entries with Sale Conditions `I' (odd lot trade) and `C' (cash trade) are retained because of their significant contribution in our dataset. We also remove the “bounceback” outliers described by aitsahalia2011UltraHighFrequency.} The dataset consists of 15 ETFs and 15 individual stocks, as listed in Table (ref). The ETFs track the performance of the S&P 500, NASDAQ 100, Dow Jones Industrial Average, Russell 2000 indices, as well as the 11 sectors of the S&P 500. The individual stocks are selected to represent a range of liquidity and volatility levels, covering various sectors such as technology, consumer goods, healthcare, and entertainment, thus providing a diverse set of assets for the empirical study. Among the 30 assets, XLC and XLRE were issued partway through the sample period. Therefore, our analysis for them begins at the start of their second year, in 2017 and 2019, respectively. After data cleaning, we resample the data to obtain 1-second and 5-second returns.

We apply the jump test proposed by aitsahalia2012TestingJumpsNoisy to identify and remove trading days with the presence of jumps for each stock. This test is a robustified version of the test introduced by aitsahalia2009TestingJumpsDiscretely, incorporating the pre-averaging method to deal with the MMS noise. After computing the standardized statistics with 5-second intraday data, we apply the universal threshold technique proposed by bajgrowicz2016JumpsHighFrequencyData to eliminate spurious jump detections. This method is more stringent than the FDR procedure and is designed to asymptotically remove all spurious detections, thereby minimizing data loss. As a result, 909 asset-days, comprising 1.2% of the entire dataset of 73,910 asset-days, are identified as containing jumps and excluded from further analysis. The numbers of days with jumps for each asset are listed in Table (ref).

We estimate the leverage effects for both weekly (defined as every five trading days) and monthly (defined as natural months) periods using both 1-second and 5-second data for each stock. The estimation proceeds in several steps, showcasing the flexibility and robustness of our framework in handling real-world data complexities.

enumerate• The ReMeDI estimator proposed by li2022ReMeDIMicrostructureNoise is used to estimate the moments $\nu_2$, $\nu_4$ and the generalized acfs $\rho_2(l)$, $\rho_3(l)$, and $\rho_4(l)$ of the MMS noise in each period. The existence and the dependence level of noise are determined by its autocovariances. The results show that the noise in the dataset is small, while the dependencies are common. For example, the analysis of noise in monthly 5-second frequency data shows that: (i) for the ETFs, 33.9% of asset-months exhibit significant noise, among which 71.5% are dependent and the average noise scale is $\varsigma=9.6\times 10^{-5}$; while (ii) for the stocks, 63.5% of the asset-months exhibit significant noise, among which 60.8% are dependent and the average noise scale is $\varsigma=2.0\times 10^{-4}$. • The scale $\overline{H}_n$ in the approximate weights of the MSLE estimator is determined by minimizing the total asymptotic variances of SALE estimators, where the asymptotic variances are estimated using 1-minute pre-averaging return data. With the existence of noise, additional lower bounds for $\overline{H}_n$ are applied, such that: (i) $\overline{H}_n$ satisfies $\overline{H}_n \geq 2\hat q + 1$, where $\hat q$ is the estimated dependence level; and (ii) the minimum values of $\overline{H}_n$ are 20 for 1-second data (corresponding to 20 seconds) and 12 for 5-second data (corresponding to 60 seconds). The former is a condition for the proposed SALE and MSLE estimators, while the latter is a rather conservative manual intervention, which mitigates numerical instability at the cost of larger asymptotic variance. • The leverage effect and its asymptotic variance are estimated using the MSLE estimator with approximate weights, where the number of scales is set to $M_n=50$ for computational efficiency.
figure[figure omitted — 443 chars of source]

Figure (ref) illustrates the dynamic nature of the leverage effect for AMZN, showcasing our estimator's ability to capture its time-varying behavior. The monthly estimates reveal significant fluctuations, clearly capturing major market stress events such as the February 2018 “Volpocalypse” and the COVID-19 sell-off in early 2020. The weekly estimates, while more volatile, provide a higher-resolution view of these dynamics.

table[table omitted — 6,418 chars of source]
figure[figure omitted — 407 chars of source]

As summarized in Table (ref) and visualized in Figure (ref), a negative leverage effect is predominant across most assets, particularly within established large-cap and defensive stocks, consistent with financial theory. The notable exceptions are the “meme stocks” AMC and GME, where retail-investor-driven speculative trading results in extreme volatility and a weaker or non-negative leverage effect. This demonstrates our method's ability to uncover such asset-specific idiosyncrasies.

Conclusion

We introduce a multi-scale framework for the robust and efficient estimation of the leverage effect from high-frequency data contaminated by complex, serially dependent microstructure noise. We construct two estimators, SALE and MSLE, by combining the shifted window, subsampling, and multi-scale techniques. We develop the asymptotic theory, and design an effective weighting strategy for the MSLE estimator. Beyond noise robustness, a central merit of our framework is its superior efficiency. In the absence of noise, the MSLE estimator improves the efficiency of the base estimator. Under a realistic setting of noise and sample size, the SALE estimator already outperforms the standard pre-averaging estimator, and the MSLE estimator further improves upon this, delivering consistently more accurate and reliable estimates. Extensive simulations and empirical applications have validated the asymptotic theory, finite-sample performance, and the practical robustness and flexibility of the proposed methods.