The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
96,849 characters
Holistic Multi-Scale Inference of the Leverage Effect: Efficiency under Dependent Microstructure Noise
\maketitle
\begin{abstract}
This paper addresses the long-standing challenge of estimating the
leverage effect from high-frequency data contaminated by dependent,
non-Gaussian microstructure noise. We depart from the conventional
reliance on pre-averaging or volatility ``plug-in'' methods by
introducing a holistic multi-scale framework that operates directly
on the leverage effect. We propose two novel estimators: the
Subsampling-and-Averaging Leverage Effect (SALE) and the
Multi-Scale Leverage Effect (MSLE). Central to our approach is a
shifted window technique that constructs a noise-unbiased base
estimator, significantly simplifying the multi-scale architecture.
We provide a rigorous theoretical foundation for these estimators,
establishing central limit theorems and stable convergence results
that remain valid under both noise-free and dependent-noise
settings. The primary contribution to estimation efficiency is a
specifically designed weighting strategy for the MSLE estimator. By
optimizing the weights based on the asymptotic covariance structure
across scales and incorporating finite-sample variance corrections,
we achieve substantial efficiency gains over existing benchmarks.
Extensive simulation studies and an empirical analysis of 30 U.S.
assets demonstrate that our framework consistently yields smaller
estimation errors and superior performance in realistic, noisy
market environments.
\end{abstract}
\noindent
\textbf{Keywords:}
High-frequency data; Realized volatility; Subsampling; Variance
reduction; Robust estimation
\section{Introduction}
The leverage effect, or the observed negative correlation between
asset returns and their volatility changes, is a prominent stylized
fact in financial econometrics \citep{black1976StudiesStockPrice,
christie1982StochasticBehaviorCommon}. It captures the asymmetry in
volatility responses to positive and negative shocks in asset prices
and is widely attributed to mechanisms such as financial leverage and
asymmetric information in markets. Accurate estimation of the
leverage effect is not only central to our understanding of asset
price dynamics but also has critical implications for the pricing and
hedging of derivative securities, especially in the presence of
volatility skews.
However, obtaining reliable estimates of the leverage effect in
high-frequency data is complicated by market microstructure (MMS)
noise. The well-known ``volatility signature plot'' demonstrates a
key challenge: with high-frequency observations, simple realized
volatility estimators are largely biased and thus inconsistent due to
noise \citep{zhou1996HighFrequencyDataVolatility,
andersen2000GreatRealizations, aitsahalia2005HowOftenSample,
patton2011DatabasedRankingRealised,
aitsahalia2019HausmanTestPresence}; simple leverage effect estimators
suffer from a similar issue. Existing methods for leverage effect
estimation have primarily focused on mitigating the impact of such
noise by either adopting the pre-averaging technique
\citep{jacod2009MicrostructureNoiseContinuous,
podolskij2009BipowertypeEstimationNoisy,
mykland2016DataCleaningInference} under the assumption of independent
and identically distributed (i.i.d.) Gaussian noise or using extra
information. For instance, \cite{wang2014EstimationLeverageEffect}
and \cite{aitsahalia2017EstimationContinuousDiscontinuous} utilize
the pre-averaging methods to estimate leverage effect under i.i.d.
Gaussian noise, \cite{yuan2020LeverageEffectHighfrequency} adopts the
parametric setting for microstructure noise proposed by
\cite{li2016EfficientEstimationIntegrated} that incorporates trading
information to address noise,
\cite{chong2024VolatilityVolatilityLeverage} utilizes high-frequency
short-dated option to recover spot volatility process, thereby
estimating leverage effect, while some other works
\citep[\emph{e.g.}][]{bandi2012TimevaryingLeverageEffects,
kalnina2017NonparametricEstimationLeverage,
curato2022StochasticLeverageEffect,yang2023EstimationLeverageEffect}
estimate leverage effect without explicitly addressing microstructure
noise. While these methods represent significant advances, their
reliance on i.i.d. Gaussian noise-related assumptions or auxiliary
data remains restrictive in practice. Specifically, empirical
evidence has consistently shown that microstructure noise exhibits
serial dependence and higher-order moments
\citep{jacod2017StatisticalPropertiesMicrostructure,
aitsahalia2019HausmanTestPresence,
li2020DependentMicrostructureNoise, da2021WhenMovingAverageModels,
li2022ReMeDIMicrostructureNoise}. These features not only violate
standard modeling assumptions, but also exacerbate bias and variance
in leverage effect estimation, underscoring the need for more
flexible methodologies that can accommodate complex noise structures.
In this paper, we propose a novel multi-scale framework for
estimating the leverage effect that explicitly accounts for
microstructure noise exhibiting more flexible noise structures,
allowing for stationary, dependent noise with nontrivial higher-order
moments, which are commonly observed in empirical financial data.
Specifically, we introduce two new estimators: the
Subsampling-and-Averaging Leverage Effect (SALE) estimator and the
Multi-Scale Leverage Effect (MSLE) estimator. In particular, the MSLE
estimator aggregates a series of weighted SALE estimators computed
across multiple time scales, exploiting their complementary
properties to achieve improved convergence rates. Our methodology
draws inspiration from the principles underlying the Two-Scale
Realized Volatility (TSRV) and Multi-Scale Realized Volatility (MSRV)
estimators \citep{zhang2005TaleTwoTime,
zhang2006EfficientEstimationStochastic,
aitsahalia2011UltraHighFrequency}, but adapts and extends them for
the specific task of estimating the leverage effect. Crucially, we do
not merely use their methods as plug-in estimators for spot
volatility; we develop a holistic multi-scale approach for the
leverage effect itself. A key innovation in our construction is the
use of a shifted window for estimating spot volatility, as
illustrated schematically in Figure~\ref{fig:base-estimators-1} and
\ref{fig:base-estimators-2}. This shift is important as it not only
helps decouple noise components to achieve unbiasedness and variance
reduction with respect to noise, but also fundamentally simplifies
certain aspects of the multi-scale estimation procedure compared with
the classical constructions in TSRV and MSRV.
The second primary contribution of this work is the explicit
demonstration of multi-scale benefits in efficiency beyond mere noise
mitigation. Specifically, for the noise-free setting, we show that
our MSLE estimator, through a proper weighting scheme, achieves a
notable reduction in asymptotic variance compared with the base
estimator by exploiting the covariance structure of SALE estimators
of different subsampling scales. Furthermore, in the noisy setting,
this efficiency advantage becomes even more pronounced when
benchmarked against the pre-averaging estimator, particularly at
lower noise levels that are more representative of practical scenarios.
Recognizing a potential gap between standard theoretical assumptions
and empirical reality regarding MMS noise, another major contribution
of this work lies in the in-depth study of optimal weight assignment
under diverse and realistic noise magnitudes. Classical asymptotic
analyses often assume that noise variance is of constant order, thus
dominating the shrinking latent increments as the sampling frequency
increases. However, empirical evidence suggests a more complex
picture. For instance, \cite{aitsahalia2019HausmanTestPresence} finds
that improvements in market liquidity allow employing simple
volatility estimators at higher frequencies, while
\cite{kalnina2008EstimatingQuadraticVariation} and
\cite{da2021WhenMovingAverageModels} also explore the case of
shrinkage MMS noise. A recent work by
\cite{chong2025WhenFrictionsAre} investigates the rough noise model
that captures a more subtle interplay between the latent price
process and noise. This implies that weighting schemes based solely
on the strict noise-dominance assumption might suffer from modeling
error when applied to real data. Motivated by this, we conduct a
fine-grained analysis of the MSLE's asymptotic variance components
across different subsampling scales and noise conditions. Moreover,
implementing truly optimal weights would require precise values of
asymptotic covariances between SALE estimators at sampling scales,
which are infeasible in practice. To overcome this, we develop a
computationally efficient approximate weighting strategy that works
for a wide spectrum of noise conditions, yielding robust
finite-sample performance as demonstrated in Monte Carlo simulations.
Our theoretical analysis establishes central limit theorems and
stable convergence results for both the noise-free and noisy
settings, demonstrating that the proposed estimators achieve
consistency and asymptotic normality. Under the noise-free setting,
both estimators attain the optimal convergence rate of order
$n^{-1/4}$. In the presence of MMS noise, the MSLE estimator achieves
a convergence rate of order $n^{-1/9}$. While this theoretical rate
is slightly slower than the $n^{-1/8}$ achieved by the pre-averaging
approach, our simulations consistently demonstrate that the MSLE
estimator outperforms the pre-averaging method in practical settings
across a wide range of noise levels and sample sizes. To support
feasible inference, we construct consistent estimators for the
asymptotic variances in both noise-free and noisy regimes, enabling
the implementation of feasible central limit theorems. Monte Carlo
simulations, employing both independent and dependent noise with
various distributions, validate both the feasible and infeasible
central limit theorems.
To assess the finite-sample performance of the proposed estimators,
we conduct another extensive simulation study that encompasses a
variety of settings with realistic time horizons and noise
conditions. This study verifies the outstanding efficiency of the
proposed MSLE estimator with the approximate weighting strategy. The
result shows that: (i) under the noise-free setting, our method
outperforms the base estimator, whereas (ii) under the noisy setting,
our method consistently outperforms the pre-averaging approach, and
this advantage is particularly pronounced in empirically relevant
scenarios. This superior finite-sample behavior is attributable to a
combination of factors, detailed in
Section~\ref{sec:practical-weight}: (i) the more advantageous
noise-free asymptotic variance of the SALE estimator, (ii) the
relatively smaller impact of noise under realistic settings, and
(iii) the enhanced performance of MSLE over the individual SALE estimators.
To examine the performance of our estimators in practical
applications, we also conduct an empirical analysis using
high-frequency financial data. The high-frequency trading data of 30
assets including ETFs and individual stocks from the U.S. stock
market are utilized. The leverage effects are estimated adaptively
based on the microstructure noise characteristics in each period, and
general negative correlation between the returns and volatility
changes is verified. This study demonstrates the practical
flexibility and effectiveness of the MSLE estimator in capturing the
leverage effect under realistic market conditions, further validating
the advantages observed in the simulation experiments.
The remaining paper is arranged as follows. Section~\ref{sec:method}
introduces the model settings, notations, and estimators.
Section~\ref{sec:results} states the main theoretical results,
including limit theorems for the SALE and MSLE estimators under both
noise-free and noisy conditions. Section~\ref{sec:practical}
discusses the issue of variance reduction, including variance
approximations under both noise-free and noisy settings, and proposes
practical strategies for optimizing the performance of MSLE.
Section~\ref{sec:simulation} provides a detailed simulation study to
validate the theoretical properties, examining the asymptotic
behavior and finite-sample performance under various settings of
microstructure noise. Section~\ref{sec:empirical} reports the results
of an empirical study using real-world high-frequency data to
demonstrate the practical utility of the proposed methods.
Proofs, feasible central limit theorems, and further elaborations on
Section~\ref{sec:practical} to \ref{sec:empirical} are provided in
Supplementary Material.
\section{Methodology}
\label{sec:method}
\subsection{Model Settings}
The noise-contaminated log-price $Y_{t_i}$ is observed at $t_i = i\Delta_n
= iT/n$ for $i = 0, 1, \dots, n$:
\begin{align}\label{eq:observation}
Y_{t_i} = X_{t_i} + \varepsilon_i,
\end{align}
where $X_{t_i}$ denotes the latent log-price, and $\varepsilon_i$ denotes
the noise. For simplicity, for any stochastic process $V$, we denote:
(i) $V_i \coloneqq V_{t_i}$,
(ii) $\Delta V_i \coloneqq V_{i+1} - V_i$, and
(iii) $\Delta_k V_i \coloneqq V_{i+k} - V_i$.
\begin{assumption}[Underlying processes]\label{ass:process}
Let both log-price process $(X_t)_{t \geq 0}$ and volatility process
$(\sigma_t)_{t \geq 0}$ be Itô processes defined on a filtered
probability space $(\Omega, \calF, (\calF_t)_{t \geq 0}, \bbP)$:
\begin{align}
X_t &= X_0 + \int_0^t \mu_s \mathrm{d} s + \int_0^t \sigma_s \mathrm{d} W_s, \\
\sigma_t &= \sigma_0 + \int_0^t a_s \mathrm{d} s + \int_0^t f_s \mathrm{d} W_s +
\int_0^t g_s \mathrm{d} B_s,
\end{align}
where $W_t$ and $B_t$ are independent Brownian motions, and $\mu_t,
a_t, f_t, g_t$ are adapted càdlàg locally bounded processes. In
addition, $f_t$ and $g_t$ are Itô processes and the volatility path
$\sigma_t^2$ is bounded away from zero.
\end{assumption}
\begin{assumption}[Noise]\label{ass:noise}
Let $\{\varepsilon_i\}_{i=0}^n$ be mean-zero, identically distributed
random variables independent of $\mathcal{F}$, with $\nu_k
\coloneqq \bbE[\varepsilon_i^k]$ and $\nu_4 < \infty$. In terms of serial
dependency, $\{\varepsilon_i\}_{i=0}^n$ are specified as either:
\begin{enumerate}[label=(\alph*)]
\item \label{ass:noise-iid} independent ($\varepsilon_i \perp \varepsilon_j$
for all $i \neq j$); or
\item \label{ass:noise-dep} $q$-dependent ($\varepsilon_i \perp \varepsilon_j$
for $|i-j| > q$) and stationary up to the fourth moment.
\end{enumerate}
\end{assumption}
Let $\langle X, \sigma^2 \rangle_T = \int_0^T 2\sigma_s^2 f_s \mathrm{d} s$ denote the true
leverage effect parameter. For any estimator of interest, we maintain
a clear distinction between its infeasible version $\widetilde{\langle X, \sigma^2 \rangle}_T$ based on
latent data $\{X_i\}_{i=0}^n$, and its feasible version $\widehat{\langle X, \sigma^2 \rangle}_T$
based on observed data $\{Y_i\}_{i=0}^n$. The statistical properties
of the estimator are investigated from two aspects:
\begin{enumerate}
\item Asserting its \emph{unbiasedness with respect to noise}:
\begin{align}
\underbrace{
\bbE\bigl( \widehat{\langle X, \sigma^2 \rangle}_T \big| \calF \bigr)
-
\widetilde{\langle X, \sigma^2 \rangle}_T
}_{\text{bias due to noise}}
= 0.
\end{align}
\item Assessing its total variance by decomposing it into the
\emph{variance due to discretization} and the expected
\emph{variance due to noise}:
\begin{align}\label{eq:variance-decomposition}
\mathrm{Var}\bigl(\widehat{\langle X, \sigma^2 \rangle}_T \bigr)
=
\underbrace{
\mathrm{Var}\bigl( \widetilde{\langle X, \sigma^2 \rangle}_T \bigr)
}_{\text{variance due to discretization}}
+ \:
\bbE\bigl[\,
\underbrace{
\mathrm{Var}\bigl( \widehat{\langle X, \sigma^2 \rangle}_T \big| \calF \bigr)
}_{\text{variance due to noise}}
\,\bigr].
\end{align}
\end{enumerate}
\begin{remark}
While some literature \citep[for
example,][]{wang2014EstimationLeverageEffect,
kalnina2017NonparametricEstimationLeverage} defines the leverage
effect as $\langle X, F(\sigma^2) \rangle_T = \int_0^T 2
F'(\sigma_s^2) \sigma_s^2 f_s \mathrm{d} s$ for a general function $F \in
\bbC^2$, we focus on the canonical case where $F(x) = x$. Beyond
notational simplicity, this allows us to concentrate on a primary
challenge addressed in this work: robust estimation under complex,
dependent noise, where a general $F$ function poses significant
additional difficulties. However, it is worth noting that the
results of our noise-free estimators
(Theorems~\ref{thm:SALE-clt-noisefree} and
\ref{thm:MSLE-clt-noisefree}) can be easily extended to the general case.
\end{remark}
\subsection{Estimators and The Robustness to Noise}
For simplicity, the following notations are used for spot volatility estimation:
\begin{align}
\label{eq:spot-vol-right}
\widehat{\sigma}_+^2 (i, H, k, s)
=
\frac{1}{kH\Delta_n} \sum_{j=s+1}^{k+s} (\Delta_H Y_{i+jH})^2
& \quad \text{for} \quad \sigma_{i+H}^2,
\\
\label{eq:spot-vol-left}
\widehat{\sigma}_-^2 (i, H, k, s)
=
\frac{1}{kH\Delta_n} \sum_{j=-k-s}^{-s-1} (\Delta_H Y_{i+jH})^2
& \quad \text{for} \quad \sigma_i^2,
\\
\label{eq:spot-vol-delta}
\widehat{\delta} (i, H, k, s)
=
\widehat{\sigma}_+^2 (i, H, k, s) - \widehat{\sigma}_-^2 (i, H, k, s)
& \quad \text{for} \quad \Delta_H \sigma_i^2.
\end{align}
Here, $H$ represents the subsampling scale, whereas $k$ and $s$
represent the size and shift of the windows for estimating spot
volatility. Specifically, the last parameter $s$ can be omitted when
$s = 1$, as this work primarily focuses on this case.
\begin{figure}[!ht]
\centering
\begin{subfigure}[c]{\textwidth}
\centering
\resizebox{\textwidth}{!}{
\begin{tikzpicture}
\draw[-latex] (-15, 0) -- (17, 0);
\foreach \x in {-14, -13} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {-10, ..., -6} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {-3, ..., 4} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {7, ..., 11} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {14, ..., 16} {\draw (\x, 0) -- (\x, 0.1);}
\node[above] at (-11.5, -0.15) {$\cdots\cdots$};
\node[above] at (-4.5, -0.15) {$\cdots\cdots$};
\node[above] at (5.5, -0.15) {$\cdots\cdots$};
\node[above] at (12.5, -0.15) {$\cdots\cdots$};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (1,
0.2) -- (0, 0.2);
\node[above] at (0.5, 0.3) {$\Delta Y_i$};
\node[below] at (0, 0) {\scriptsize{$i$}};
\node[below] at (1, 0) {\scriptsize{$i+1$}};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (0,
0.2) -- (-7, 0.2);
\node[above] at (-3.5, 0.3) {$\widehat{\sigma}_-^2 (i, 1, k_n, 0)$};
\node[below] at (-7, 0) {\scriptsize{$i-k_n$}};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (8,
0.2) -- (1, 0.2);
\node[above] at (4.5, 0.3) {$\widehat{\sigma}_+^2 (i, 1, k_n, 0)$};
\node[below] at (8, 0) {\scriptsize{$i+k_n+1$}};
\foreach -5/11 in {-7/0, 1/8} {
\fill [fill=red, opacity=0.2] (-5,0.08) -- (-5,
0.0) -- (11, 0.0) -- (11, 0.08) -- cycle;
}
\foreach -5/11 in {0/1} {
\fill [fill=blue, opacity=0.2] (-5,0.08) -- (-5,
0.0) -- (11, 0.0) -- (11, 0.08) -- cycle;
}
\end{tikzpicture}
}
\caption{The continuous leverage effect estimator in
\cite{aitsahalia2017EstimationContinuousDiscontinuous}}
\label{fig:base-estimators-1}
\end{subfigure}
\begin{subfigure}[c]{\textwidth}
\centering
\resizebox{\textwidth}{!}{
\begin{tikzpicture}
\draw[-latex] (-15, 0) -- (17, 0);
\foreach \x in {-14, -13} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {-10, ..., -6} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {-3, ..., 4} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {7, ..., 11} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {14, ..., 16} {\draw (\x, 0) -- (\x, 0.1);}
\node[above] at (-11.5, -0.15) {$\cdots\cdots$};
\node[above] at (-4.5, -0.15) {$\cdots\cdots$};
\node[above] at (5.5, -0.15) {$\cdots\cdots$};
\node[above] at (12.5, -0.15) {$\cdots\cdots$};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (1,
0.2) -- (0, 0.2);
\node[above] at (0.5, 0.3) {$\Delta Y_i$};
\node[below] at (0, 0) {\scriptsize{$i$}};
\node[below] at (1, 0) {\scriptsize{$i+1$}};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (-1,
0.2) -- (-8, 0.2);
\node[above] at (-4.5, 0.3) {$\widehat{\sigma}_-^2 (i, 1, k_n, 1)$};
\node[below] at (-1, 0) {\scriptsize{$i-1$}};
\node[below] at (-8, 0) {\scriptsize{$i-k_n-1$}};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (9,
0.2) -- (2, 0.2);
\node[above] at (5.5, 0.3) {$\widehat{\sigma}_+^2 (i, 1, k_n, 1)$};
\node[below] at (2, 0) {\scriptsize{$i+2$}};
\node[below] at (9, 0) {\scriptsize{$i+k_n+2$}};
\foreach -5/11 in {-8/-1, 2/9} {
\fill [fill=red, opacity=0.2] (-5,0.08) -- (-5,
0.0) -- (11, 0.0) -- (11, 0.08) -- cycle;
}
\foreach -5/11 in {0/1} {
\fill [fill=blue, opacity=0.2] (-5,0.08) -- (-5,
0.0) -- (11, 0.0) -- (11, 0.08) -- cycle;
}
\end{tikzpicture}
}
\caption{The all-observation leverage effect estimator defined in
Equation~\eqref{eq:all-observation}}
\label{fig:base-estimators-2}
\end{subfigure}
\begin{subfigure}[c]{\textwidth}
\centering
\resizebox{\textwidth}{!}{
\begin{tikzpicture}
\draw[-latex] (-15, 0) -- (17, 0);
\foreach \x in {-14, -13} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {-10, ..., -6} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {-3, ..., 4} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {7, ..., 11} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {14, ..., 16} {\draw (\x, 0) -- (\x, 0.1);}
\node[above] at (-11.5, -0.15) {$\cdots\cdots$};
\node[above] at (-4.5, -0.15) {$\cdots\cdots$};
\node[above] at (5.5, -0.15) {$\cdots\cdots$};
\node[above] at (12.5, -0.15) {$\cdots\cdots$};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (1,
0.2) -- (0, 0.2);
\node[above] at (0.5, 0.3) {$\Delta Y_i$};
\node[below] at (0, 0) {\scriptsize{$i$}};
\node[below] at (1, 0) {\scriptsize{$i+1$}};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (-2,
0.2) -- (-9, 0.2);
\node[above] at (-5.5, 0.3) {$\widehat{\sigma}_-^2 (i, 1, k_n, 2)$};
\node[below] at (-2, 0) {\scriptsize{$i-2$}};
\node[below] at (-9, 0) {\scriptsize{$i-k_n-2$}};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (10,
0.2) -- (3, 0.2);
\node[above] at (6.5, 0.3) {$\widehat{\sigma}_+^2 (i, 1, k_n, 2)$};
\node[below] at (3, 0) {\scriptsize{$i+3$}};
\node[below] at (10, 0) {\scriptsize{$i+k_n+3$}};
\foreach -5/11 in {-9/-2, 3/10} {
\fill [fill=red, opacity=0.2] (-5,0.08) -- (-5,
0.0) -- (11, 0.0) -- (11, 0.08) -- cycle;
}
\foreach -5/11 in {0/1} {
\fill [fill=blue, opacity=0.2] (-5,0.08) -- (-5,
0.0) -- (11, 0.0) -- (11, 0.08) -- cycle;
}
\end{tikzpicture}
}
\caption{The shifted all-observation leverage effect estimator
(see Remark~\ref{rem:shifted-spot-vol})}
\label{fig:base-estimators-3}
\end{subfigure}
\begin{subfigure}[c]{\textwidth}
\centering
\resizebox{\textwidth}{!}{
\begin{tikzpicture}
\draw[-latex] (-15, 0) -- (17, 0);
\foreach \x in {-14, -13} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {-10, ..., -6} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {-3, ..., 4} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {7, ..., 11} {\draw (\x, 0) -- (\x, 0.1);}
\foreach \x in {14, ..., 16} {\draw (\x, 0) -- (\x, 0.1);}
\node[above] at (-11.5, -0.15) {$\cdots\cdots$};
\node[above] at (-4.5, -0.15) {$\cdots\cdots$};
\node[above] at (5.5, -0.15) {$\cdots\cdots$};
\node[above] at (12.5, -0.15) {$\cdots\cdots$};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (2,
0.2) -- (0, 0.2);
\node[above] at (1, 0.3) {$\Delta Y_i$};
\node[below] at (0, 0) {\scriptsize{$i$}};
\node[below] at (2, 0) {\scriptsize{$i+2$}};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (-2,
0.2) -- (-14, 0.2);
\node[above] at (-8, 0.3) {$\widehat{\sigma}_-^2 (i, 2, k_n', 1)$};
\node[below] at (-2, 0) {\scriptsize{$i-2$}};
\node[below] at (-14, 0) {\scriptsize{$i-2k_n'-2$}};
\draw[decorate,decoration={brace,amplitude=5pt,mirror}] (16,
0.2) -- (4, 0.2);
\node[above] at (10, 0.3) {$\widehat{\sigma}_+^2 (i, 2, k_n', 1)$};
\node[below] at (4, 0) {\scriptsize{$i+4$}};
\node[below] at (16, 0) {\scriptsize{$i+2k_n'+4$}};
\foreach -5/11 in {-14/-2, 4/16} {
\fill [fill=red, opacity=0.2] (-5,0.08) -- (-5,
0.0) -- (11, 0.0) -- (11, 0.08) -- cycle;
}
\foreach -5/11 in {0/2} {
\fill [fill=blue, opacity=0.2] (-5,0.08) -- (-5,
0.0) -- (11, 0.0) -- (11, 0.08) -- cycle;
}
\end{tikzpicture}
}
\caption{The subsampling leverage effect estimator defined in
Equation~\eqref{eq:subsample-noisy} with $H=2$}
\label{fig:subsample-estimators}
\end{subfigure}
\caption{
Increments in the base and subsampling estimators. The increments
in the base estimators shown in panels
(\subref{fig:base-estimators-1}) to
(\subref{fig:base-estimators-3}) are used as proxies for
$\int_{t_i}^{t_{i+1}} \mathrm{d} \langle X, \sigma^2 \rangle_t$, whereas the increment in the
subsampling estimator shown in panel
(\subref{fig:subsample-estimators}) is used as a proxy for
$\int_{t_i}^{t_{i+H}} \mathrm{d} \langle X, \sigma^2 \rangle_t$.
}
\label{fig:base-estimators}
\end{figure}
\subsubsection{All-Observation Estimator}
To start with, consider an \emph{all-observation Leverage Effect}
(LE) estimator that directly utilizes all noisy observations,
\begin{align}\label{eq:all-observation}
\widehat{\langle X, \sigma^2 \rangle}^{\rm (all)}_T
=
\sum_{i=k_n+1}^{n-k_n-2} (\Delta Y_{i})
\: \widehat{\delta} (i, 1, k_n).
\end{align}
Here, the window size $k_n$ satisfies that $k_n \to \infty$ and $k_n
\Delta_n \to 0$ as $n\to \infty$. The estimator is very similar to
the continuous leverage effect estimator\footnote{Since jumps are not
included, the truncation term in their original estimator is omitted
here.} studied
by~\citet{aitsahalia2017EstimationContinuousDiscontinuous}. The only
difference is that our estimator shifts the windows for spot
volatility estimates outward by $\Delta_{n}$, as illustrated by
Figure~\ref{fig:base-estimators-1}~and~\ref{fig:base-estimators-2}.
To see the reason, consider applying their estimator directly to
noisy observations, and it follows that there is a divergent bias due
to noise of
\begin{align}
\bbE\bigl(
\widehat{\langle X, \sigma^2 \rangle}^{[\text{AFLWY17}]}_T \big| \calF
\bigr)
-
\widetilde{\langle X, \sigma^2 \rangle}^{[\text{AFLWY17}]}_T
=
2 \nu_3 T^{-1} k_n^{-1} n^2
+ O_p(n),
\end{align}
when $\nu_3$ is not strictly zero. In contrast, the estimator in
Equation~\eqref{eq:all-observation} is unbiased due to noise with a
smaller variance, as described by the next
Proposition~\ref{prop:all-observation-noise}.
\begin{proposition}\label{prop:all-observation-noise}
Under Assumptions~\ref{ass:process} and
\ref{ass:noise}\ref{ass:noise-iid}, as $n\to\infty$, we have
\begin{align}
\bbE \big(
\widehat{\langle X, \sigma^2 \rangle}^{\rm (all)}_T
\big| \calF
\bigr)
&=
\widetilde{\langle X, \sigma^2 \rangle}^{\rm (all)}_T,
\\
\Delta_n^3 k_n^2
\mathrm{Var}\bigl( \widehat{\langle X, \sigma^2 \rangle}^{\rm (all)}_T \big| \calF \bigr)
&\xrightarrow{p}
(8 \nu_2 \nu_4 + 16 \nu_2^3 + 8 \nu_3^2) T.
\label{eq:all-observation-noise-variance}
\end{align}
\end{proposition}
With the window shift, the bias due to noise is eliminated, and the
variance due to noise is reduced, while the asymptotic behavior of
the noise-free estimator $\widetilde{\langle X, \sigma^2 \rangle}_T^{\rm (all)}$ remains the same as
$\widetilde{\langle X, \sigma^2 \rangle}^{[\text{AFLWY17}]}_{T}$. As a direct result of this
unbiasedness, a debias step in TSRV or MSRV is no longer needed.
However, the all-observation estimator is not consistent in the
presence of noise: the variance due to noise is $O(n^3 k_n^2)$ while
the noise due to discretization is $O(k_n^{-1} + n^{-1} k_n)$,
resulting in an exploding total variance.
\subsubsection{SALE: Subsampling-and-Averaging Estimator}
To mitigate the impact of noise, we employ a subsampling procedure: a
subsample of observations is denoted by a pair of integers $(H, h)$,
where $H \geq 1$ and $1 \leq h \leq H$ are the scale and
index of the subsample. The $j$th observation in subsample $(H, h)$
corresponds to the original index $j_{H,h} = jH + h - 1$, where $j =
0, 1, \dotsc, n_{H,h}$ and $n_{H,h} = \lfloor (n - h + 1)/H \rfloor$.
Assume that $k_n\to\infty$, $H_n\to\infty$ and $k_nH_n\Delta_n \to 0$
as $n\to\infty$. A subsampling estimator is constructed by applying
the all-observation estimator to a subsample of observations:
\begin{align}\label{eq:subsample-noisy}
\widehat{\langle X, \sigma^2 \rangle}^{(H_n, h)}_T
=
\sum_{i=k_n+1}^{n_{H_n,h}-k_n-2}
(\Delta_{H_n} Y_{i_{H_n, h}})
\: \widehat{\delta} (i_{H_n, h}, H_n, k_n).
\end{align}
As a variant of the all-observation estimator, the subsampling
estimator is not consistent. Nor is it statistically sound, as it
fails to utilize the full information in the observed data.
Therefore, by taking average over the subsampling estimators with the
same scale, we obtain an \emph{Subsampling-and-Averaging Leverage
Effect} (SALE) estimator:
\begin{align}\label{eq:SALE-noisy}
\widehat{\langle X, \sigma^2 \rangle}^{(H_n)}_T
=
\frac{1}{H_n}
\sum_{h=1}^{H_n} \widehat{\langle X, \sigma^2 \rangle}^{(H_n, h)}_T
=
\frac{1}{H_n}
\sum_{i=(k_n+1)H_n}^{n-(k_n+2)H_n}
(\Delta_{H_n} Y_i)
\: \widehat{\delta} (i, H_n, k_n),
\end{align}
which can be viewed as summation of overlapping sparse increments.
The SALE estimator is consistent under Assumption~\ref{ass:noise}
with proper choices of $k_n$ and $H_n$. This is primarily attributed
to the fact that the correlation due to noise between different
subsampling estimators in Equation~\eqref{eq:SALE-noisy} can be
controlled by the dependence level of noise. Specifically, if
Assumption~\ref{ass:noise}\ref{ass:noise-iid} holds, this correlation
becomes zero. Thus, the variance due to noise significantly reduces,
as the following proposition describes.
\begin{proposition}\label{prop:SALE-noise}
Suppose that Assumptions~\ref{ass:process} and
\ref{ass:noise}\ref{ass:noise-dep} hold, and that $H_n > 2q$.
Define the generalized autocorrelation functions (ACFs) of noise
for any $l \in \bbZ$ as
\begin{align}\label{eq:general-acf}
\rho_2(l) &= \mathrm{Corr}(\varepsilon_i, \varepsilon_{i+l})
= \frac{\bbE [ \varepsilon_{i}\varepsilon_{i+l} ]}{\nu_2}, \\
\rho_3(l) &= \mathrm{Corr}(\varepsilon_i, \varepsilon_{i+l}^2)
= \frac{\bbE [ \varepsilon_{i}\varepsilon_{i+l}^{2} ]}{\sqrt{\nu_2(\nu_4 - \nu_2^2)}}, \\
\rho_4(l) &= \mathrm{Corr}( \varepsilon_{i}^{2}, \varepsilon_{i+l}^{2} )
= \frac{\bbE[ \varepsilon_{i}^{2} \varepsilon_{i+l}^{2} ] - \nu_2^2}{\nu_4 - \nu_2^2}.
\end{align}
As $n\to\infty$, we have
\begin{gather}
\label{eq:SALE-var-noise}
\Delta_n^3 H_n^4 k_n^2
\mathrm{Var} \bigl( \widehat{\langle X, \sigma^2 \rangle}^{(H_n)}_T \big|
\calF \bigr)
\xrightarrow{p}
\Phi T, \\
\label{eq:Phi}
\text{where }
\Phi = 8 \nu_2 (\nu_4 - \nu_2^2)
\sum_{l=-q}^{q} \bigl( \rho_2(l)\rho_4(l) + \rho_3(l) \rho_3(-l) \bigr)
+ 24 \nu_2^3 \sum_{l=-q}^{q} \rho_2^3(l).
\end{gather}
Specifically, under Assumptions~\ref{ass:process} and
\ref{ass:noise}\ref{ass:noise-iid},
Equation~\eqref{eq:SALE-var-noise} holds with
$\Phi = 8 \nu_2 \nu_4 + 16 \nu_2^3 + 8 \nu_3^2$.
\end{proposition}
The SALE estimator has a variance due to noise of $O(n^3 H_n^{-4}
k_n^{-2})$, and a variance due to discretization of $O(k_n^{-1} +
n^{-1} H_n k_n)$. Suppose $H_n \propto n^a$ and $k_n \propto
(n/H)^b$, the optimal total variance $O(n^{-1/7})$ is achieved at
$a=5/7$ and $b=1/2$. Despite being consistent, this is far from the
optimal rate $O(n^{-1/4})$ of pre-averaging approach.
\subsubsection{MSLE: Multi-Scale Estimator}\label{sec:MSLE}
Consider a set of scales $1 \leq H_1 < \dots < H_{M_n} \leq n^a$ for
some $a\in(0,1)$, where $M_n>0$. An \emph{Multi-Scale Leverage
Effect} (MSLE) estimator is defined as a weighted average of SALE
estimators at different scales:
\begin{align}\label{eq:MSLE-noisy}
\widehat{\langle X, \sigma^2 \rangle}^{\rm (MS)}_T = \sum_{p=1}^{M_n} w_{p} \widehat{\langle X, \sigma^2 \rangle}^{(H_p)}_T,
\end{align}
where the weight vector $\bm w=(w_1, \dotsc, w_{M_n})$ satisfies $\bm
w^T \bm 1_{M_n} = 1$ with $\|\bm w\|_1$ bounded.
When the noise is i.i.d., the covariance due to noise between a pair
of SALE estimators are non-zero only when a scale is double the
other, and the corresponding correlation coefficient is small (for
example, 0.1 for Gaussian noise). For serial dependent noise, the
case is more complicated and analytical results are hard to derive.
Instead, numerical calculation is available (see Supplementary Material).
\begin{proposition}\label{prop:MSLE-noise}
Under Assumptions~\ref{ass:process} and
\ref{ass:noise}\ref{ass:noise-iid}, and suppose $k_p = \lfloor
\beta \lfloor n/H_p\rfloor^b \rfloor$ holds for all $p\in\{1,
\dotsc, M_n\}$ with some constant $\beta>0$ and $b\in(0,1)$. For
any $p, q \in \{1, \dotsc, M_n\}$, as $n\to\infty$, we have
\begin{gather}
\Delta_n^3 H_p^2 H_q^2 k_p k_q
\mathrm{Cov} \bigl(
\widehat{\langle X, \sigma^2 \rangle}^{(H_p)}_T, \widehat{\langle X, \sigma^2 \rangle}^{(H_q)}_T
\big| \calF
\bigr)
\xrightarrow{p}
F_{p,q} T. \\
\label{eq:MSLE-noise-Fpq}
\text{where }
F_{p,q} = (8\nu_2\nu_4 + 16\nu_2^3 + 8\nu_3^2) 1_{\{p=q\}} +
2\nu_2(\nu_4-\nu_2^2) (1_{\{H_p/H_q=2\}} + 1_{\{H_q/H_p=2\}}),
\end{gather}
and thus
\begin{align}
\frac
{\mathrm{Var} \bigl( \widehat{\langle X, \sigma^2 \rangle}^{\rm (MS)}_T
\big| \calF \bigr)}
{\displaystyle \Delta_n^{-3} \sum_{p=1}^{M_n} \sum_{q=1}^{M_n}
\frac{w_p}{H_p^2 k_p} \cdot F_{p,q} \cdot \frac{w_q}{H_q^2 k_q}}
\xrightarrow{p}
T.
\end{align}
\end{proposition}
\begin{remark}\label{rem:shifted-spot-vol}
The ``double scale'' terms in $F_{p,q}$ can be removed by using a
further shifted spot volatility window with $s=2$ in
Equation~\eqref{eq:spot-vol-delta}, as illustrated in
Figure~\ref{fig:base-estimators-3}. A similar proposition can be
established with $F_{p,q} = (8\nu_2\nu_4 + 8\nu_2^3 + 8\nu_3^2)
1_{\{p=q\}}$, eliminating the cross terms and reducing the variance.
\end{remark}
The MSLE estimator effectively reduces the variance due to noise. For
example, setting $H_p=p$ for $p=1, \dotsc, M_n$, $M_n = \lfloor n^a
\rfloor$, and $w_p \propto p^{4-2b}$, the variance due to noise is
$O(n^{3-5a-2b+2ab})$, while the variance due to discretization is
$O(n^{-(1-a)(b \land (1-b))})$. Thus, by selecting $a=5/9$ and $b =
1/2$, an optimal convergence rate of $n^{1/9}$ is achieved for MSLE,
close to the optimal rate $n^{1/8}$ of pre-averaging approach.
\section{Main Results}\label{sec:results}
\subsection{Central Limit Theorems for SALE}
We start by establishing the following theorem for the noise-free
version of SALE. Two scenarios for the scale $H_n$ are considered:
either it is fixed, or it goes to infinity as $n \to \infty$.
Hereafter, we use $\xrightarrow{\rm st}$ to denote stable convergence in law.
\begin{assumption}\label{ass:SALE-para}
Suppose that $H_n$ and $k_n$ satisfy one of the following conditions:
\begin{enumerate}[label=(\alph*)]
\item \label{ass:SALE-para-finite} $H_n=H$ is a given positive
integer, $k_n = \lfloor \beta \lfloor n/H \rfloor^b \rfloor$
for some $\beta > 0$ and $b \in (0,1)$.
\item \label{ass:SALE-para-asym} $H_n = \lfloor \alpha n^a \rfloor$,
$k_n = \lfloor \beta
\lfloor n/H_n \rfloor^b \rfloor$ for some $\alpha, \beta > 0$ and
$a, b \in (0,1)$.
\end{enumerate}
\end{assumption}
\begin{theorem}\label{thm:SALE-clt-noisefree}
\begin{enumerate}[label=(\arabic*)]
\item []
\item \label{thm:SALE-clt-noisefree-1} Under
Assumptions~\ref{ass:process} and
\ref{ass:SALE-para}\ref{ass:SALE-para-finite}, let $u_n =
\sqrt{k_n \land (k_n H \Delta_n)^{-1}}$. There exist a standard
Brownian motion $(W_{1,t})_{t\geq 0}$ independent of $\calF$
and a predictable process $(\zeta_{1,t})_{t\geq 0}$ such that,
as $n\to\infty$,
\begin{gather}
u_n \bigl( \widetilde{\langle X, \sigma^2 \rangle}^{(H)}_T - \langle X, \sigma^2 \rangle_T \bigr)
\xrightarrow{\rm st}
\int_0^T \zeta_{1,t} \mathrm{d} W_{1,t},
\\
\label{eq:avar-SALE-clt-noisefree-1}
\int_0^T \zeta_{1,t}^2 \mathrm{d} t =
\frac{u_n^2}{k_n} \left(\frac{8}{3} + \frac{4}{3H^2}\right)
\int_0^T \sigma_s^6 \mathrm{d} t +
u_n^2 k_n H \Delta_n \frac{2}{3} \int_0^T \sigma_t^2 \mathrm{d}
\langle \sigma^2, \sigma^2 \rangle_t.
\end{gather}
\item \label{thm:SALE-clt-noisefree-2} Under
Assumptions~\ref{ass:process} and
\ref{ass:SALE-para}\ref{ass:SALE-para-asym}, there exist a
standard Brownian motion $(W_{1,t})_{t\geq 0}$ independent of
$\calF$ and a predictable process $(\zeta_{1,t})_{t\geq 0}$
such that, as $n\to\infty$,
\begin{gather}
n^{\frac{1}{2}(1-a)(b\land (1-b))}
\bigl( \widetilde{\langle X, \sigma^2 \rangle}^{(H_n)}_T - \langle X, \sigma^2 \rangle_T \bigr)
\xrightarrow{\rm st}
\int_0^T \zeta_{1, t} \mathrm{d} W_{1, t},
\\
\label{eq:avar-SALE-clt-noisefree-2}
\int_{0}^{T} \zeta_{1, t}^{2} \mathrm{d} t
=
\frac{8 \alpha^b}{3 \beta} \int_0^T \sigma_t^6 \mathrm{d} t
\cdot 1_{(0, 1/2]}(b)
+
\frac{2 \alpha^{1-b} \beta T}{3} \int_0^T \sigma_t^2 \mathrm{d}
\langle \sigma^2, \sigma^2 \rangle_t
\cdot 1_{[1/2, 1)}(b).
\end{gather}
\end{enumerate}
\end{theorem}
\begin{remark}
Taking $H=1$ in
Theorem~\ref{thm:SALE-clt-noisefree}\ref{thm:SALE-clt-noisefree-1}
yields the central limit theorem for the all-observation estimator,
which has a same asymptotic variance as the continuous leverage
effect estimator in~\citet{aitsahalia2017EstimationContinuousDiscontinuous}.
\end{remark}
\begin{remark}\label{remark:spot-vol-errors}
Similar to existing work on leverage effect
estimation~\citep{wang2014EstimationLeverageEffect,
aitsahalia2014HighFrequencyFinancialEconometrics,
aitsahalia2017EstimationContinuousDiscontinuous,
kalnina2017NonparametricEstimationLeverage,
yang2023EstimationLeverageEffect}, the asymptotic variance is
determined by the spot volatility estimation, which consists of two
sources of errors: the \emph{price variation error} and the
\emph{volatility variation
error}~\citep{aitsahalia2017EstimationContinuousDiscontinuous},
corresponding to the first and second terms in
Equation~\eqref{eq:avar-SALE-clt-noisefree-1} or
\eqref{eq:avar-SALE-clt-noisefree-2}. Intuitively, increasing $k_n$
leads to a wider spot volatility estimation window and thus a
larger sample size for that estimation, which reduces the price
variation error. On the other hand, this increases the volatility
variation error, because the estimated volatility becomes less
``local''. The optimal choice of $k_n$ is thus a trade-off between
these two sources of error.
\end{remark}
Next, we establish the following theorem for the noisy version of
SALE. Notably,
Assumption~\ref{ass:SALE-para}\ref{ass:SALE-para-finite} is not
considered, as a finite $H$ leads to a divergent variance due to
noise and thus an inconsistent estimator.
\begin{theorem}\label{thm:SALE-clt-noisy}
Under Assumptions~\ref{ass:process},
\ref{ass:noise}\ref{ass:noise-dep} and
\ref{ass:SALE-para}\ref{ass:SALE-para-asym}, suppose that $H_n >
2q$ and $4a + 2b - 2ab > 3$. Let $r = [(1-a)(b\land (1-b))] \land
[4a+2b-2ab-3]$ and let $\Phi$ be as defined in
Equation~\eqref{eq:Phi}. There exist a standard Brownian motion
$(W_{2, t})_{t\geq 0}$ independent of $\calF$ and a predictable
process $(\zeta_{2, t})_{t\geq 0}$, such that, as $n\to\infty$,
\begin{align}
& n^{\frac{1}{2}r}
\bigl(
\widehat{\langle X, \sigma^2 \rangle}^{(H_n)}_T - \langle X, \sigma^2 \rangle_T
\bigr)
\xrightarrow{\rm st}
\int_{0}^{T} \zeta_{2, t} \mathrm{d} W_{2, t},
\\
\int_{0}^{T} \zeta_{2, t}^{2} \mathrm{d} t
& =
\frac{8 \alpha^{b}}{3 \beta} \int_{0}^{T} \sigma_{t}^{6} \mathrm{d} t
\cdot 1_{\{(1-a)b\}}(r)
\notag
\\ & \qquad
+
\frac{2 \alpha^{1-b} \beta T}{3} \int_{0}^{T} \sigma_{t}^{2}
\mathrm{d} \langle \sigma^{2}, \sigma^{2} \rangle_t
\cdot 1_{\{(1-a)(1-b)\}}(r)
\notag
\\ & \qquad
+ \frac{1}{\alpha^{4-2b} \beta^2 T^3}
\int_{0}^{T} \Phi \mathrm{d} t
\cdot 1_{\{4a+2b-2ab-3\}}(r).
\end{align}
\end{theorem}
\subsection{Central Limit Theorems for MSLE}
To establish the limit theorems for MSLE, the following conditions on
scales, window sizes and weights are established, and the covariances
due to discretization between SALE estimators are given by
Proposition~\ref{prop:SALE-acov-noisefree} and \ref{prop:adj-factor-asym}.
\begin{assumption}\label{ass:MSLE-para}
The scales $\{H_p\}_{p=1}^{M_n}$ satisfy $1 \leq H_1 < \dotsc <
H_{M_n} \leq n^a$ for some $a\in(0,1)$. The window sizes are $k_p =
\lfloor \beta \lfloor n/H_p \rfloor^b \rfloor$ for all $p \in \{1,
\dotsc, M_n\}$ for some $\beta > 0$ and $b \in (0, 1)$. The weight
vector $\bm w=(w_1, \dotsc, w_{M_n})$ satisfies that $\bm w^T \bm
1_{M_n} = 1$ and that $\|\bm w\|_1$ is bounded.
\end{assumption}
\begin{proposition}\label{prop:SALE-acov-noisefree}
Suppose that Assumptions~\ref{ass:process} and \ref{ass:MSLE-para}
hold. For any $1 \leq q \leq p \leq M_n$, let $u_n = \sqrt{k_p
\land (k_p H_p \Delta_n)^{-1}}$. There exist $v_{p,q}^{(1)},
v_{p,q}^{(2)} \in [0, 1)$ that depend on $H_p, H_q$ and $n$, such that,
as $n\to\infty$,
\begin{gather}
u_n
\begin{pmatrix}
\widetilde{\langle X, \sigma^2 \rangle}^{(H_p)}_T - \langle X, \sigma^2 \rangle_T \\
\widetilde{\langle X, \sigma^2 \rangle}^{(H_q)}_T - \langle X, \sigma^2 \rangle_T
\end{pmatrix}
\xrightarrow{\rm st}
\calN\left(0,
u_n^2
\begin{bmatrix}
\Sigma_{p,p}^{\mathrm{(disc)}} &
\Sigma_{p,q}^{\mathrm{(disc)}} \\
\Sigma_{p,q}^{\mathrm{(disc)}} &
\Sigma_{q,q}^{\mathrm{(disc)}}
\end{bmatrix}\right),
\\
\label{eq:SALE-acov-disc}
\text{where }
\Sigma_{p,q}^{\mathrm{(disc)}} =
\frac{1}{k_p} \cdot 4 v_{p,q}^{(1)} \frac{H_q}{H_p}
\cdot \int_0^T \sigma_t^6 \mathrm{d} t
+
k_pH_p\Delta_n \cdot \frac{2}{3} v_{p,q}^{(2)}
\left(\frac{k_qH_q}{k_pH_p}\right)^2
\cdot \int_0^T \sigma_t^2 \mathrm{d} \langle \sigma^2, \sigma^2 \rangle_t,
\end{gather}
and the limiting process is independent of $\calF$.
\end{proposition}
\begin{remark}
Factors $v_{p,q}^{(1)}$ and $v_{p,q}^{(2)}$, arising from the grid
structures of scales $H_p$ and $H_q$, contribute to price variation
error and volatility variation error, respectively. Their
definitions are detailed in Supplementary Material.
\end{remark}
\begin{proposition}\label{prop:adj-factor-asym}
Suppose that the conditions of
Proposition~\ref{prop:SALE-acov-noisefree} hold. Consider two
sequences of scales $H_p$ and $H_q$ indexed by $n$, satisfying that
(i) $H_p \geq H_q$ for all $n$, (ii) as $n\to\infty$, $H_p, H_q \to
\infty$ and $H_{q} / H_{p} \to \rho$ for some constant $\rho \in (0, 1]$.
Then, as $n\to\infty$, we have
\begin{align}\label{eq:adj-factor-asym}
v_{p,q}^{(1)} \to 1 - \frac{\rho}{3}
\quad \text{and} \quad
v_{p,q}^{(2)} \to 1.
\end{align}
\end{proposition}
For the asymptotic behavior of MSLE, a specific set of consecutive
scales are considered, and the weights are defined with a continuous
bounded function.
\begin{assumption}\label{ass:MSLE-para-detail}
Suppose that $\{H_p\}_{p=1}^{M_n}$ and $\bm{w}$ satisfy the
following conditions:
\begin{enumerate}[label=(\alph*)]
\item \label{ass:MSLE-para-detail-scale} $H_p = m_n+p$ for all $p
\in \{1, \dotsc, M_n\}$. Defining $H_n^* = H_{M_n}$, the
sequences of positive integers $m_n$ and $M_n$ are selected
such that, as $n\to\infty$, $H_n^* / n^a \to \alpha$, $m_n /
H_n^* \to c$ for some constants $\alpha > 0$, $a\in(0,1)$, $c\in(0, 1)$.
\item \label{ass:MSLE-para-detail-weight} $w_p =
\frac{1-c}{M_n}\phi(c+\frac{p}{H_n^*})$ for all $p \in \{1,
\dotsc, M_n\}$, where $\phi: [c, \infty) \to \bbR$ is a
continuous bounded function satisfying that $\int_c^1 \phi(x)
\mathrm{d} x = 1$.
\end{enumerate}
\end{assumption}
\begin{theorem}\label{thm:MSLE-clt-noisefree}
Suppose that Assumptions~\ref{ass:process}, \ref{ass:MSLE-para} and
\ref{ass:MSLE-para-detail} hold. There exist a standard Brownian
motion $(W_{3, t})_{t\geq 0}$ independent of $\calF$ and a
predictable process $(\zeta_{3, t})_{t\geq 0}$, such that, as $n\to\infty$,
\begin{align}
& \quad
n^{\frac{1}{2}(1-a)(b\land (1-b))}
\bigl( \widetilde{\langle X, \sigma^2 \rangle}^{\rm (MS)}_T - \langle X, \sigma^2 \rangle_T \bigr)
\xrightarrow{\rm st}
\int_0^T \zeta_{3, t} \mathrm{d} W_{3, t},
\\
\int_0^T \zeta_{3, t}^{2} \mathrm{d} t
&=
\frac{8\alpha^b}{\beta} \int_c^1 \int_c^x \phi(x) \phi(y)
x^b \left(\frac{y}{x} - \frac{y^2}{3x^2}\right) \mathrm{d} y \mathrm{d} x
\cdot \int_0^T \sigma_t^6 \mathrm{d} t \cdot 1_{(0, 1/2]}(b)
\notag
\\ & \quad
+ \frac{4\alpha^{1-b}\beta T}{3} \int_c^1 \int_c^x \phi(x)
\phi(y) \frac{y^{2(1-b)}}{x^{1-b}} \mathrm{d} y \mathrm{d} x \cdot
\int_0^T \sigma_t^2 \mathrm{d} \langle \sigma^2, \sigma^2
\rangle_t \cdot 1_{[1/2, 1)}(b).
\label{eq:avar-MSLE-clt-noisefree-2}
\end{align}
\end{theorem}
\begin{theorem}\label{thm:MSLE-clt-noisy}
Under Assumptions~\ref{ass:process},
\ref{ass:noise}\ref{ass:noise-iid}, \ref{ass:MSLE-para} and
\ref{ass:MSLE-para-detail}, suppose that $5a+2b-2ab > 3$, and let
$r = [(1-a)(b\land (1-b))] \land [5a+2b-2ab-3]$, $F_1 = 8\nu_2\nu_4
+ 16\nu_2^3 + 8\nu_3^2$ and $F_2 = 2\nu_2(\nu_4-\nu_2^2)$. There
exist a standard Brownian motion $(W_{4, t})_{t\geq 0}$ independent
of $\calF$ and a predictable process $(\zeta_{4, t})_{t\geq 0}$,
such that, as $n\to\infty$,
\begin{align}
& \qquad \qquad
n^{\frac{1}{2}r}
\bigl( \widehat{\langle X, \sigma^2 \rangle}^{\rm (MS)}_T - \langle X, \sigma^2 \rangle_T \bigr)
\xrightarrow{\rm st} \int_0^T \zeta_{4, t} \mathrm{d} W_{4, t},
\\
\notag
\int_0^T \zeta_{4, t}^{2} \mathrm{d} t
&=
\frac{8\alpha^b}{\beta} \int_c^1 \int_c^x \phi(x) \phi(y) x^b
\left(\frac{y}{x} - \frac{y^2}{3x^2}\right) \mathrm{d} y \mathrm{d} x \cdot
\int_0^T \sigma_t^6 \mathrm{d} t \cdot 1_{\{(1-a)b\}}(r)
\\ & \quad \notag
+ \frac{4\alpha^{1-b}\beta T}{3} \int_c^1 \int_c^x \phi(x)
\phi(y) \frac{y^{2(1-b)}}{x^{1-b}} \mathrm{d} y \mathrm{d} x \cdot \int_0^T
\sigma_t^2 \mathrm{d} \langle \sigma^2, \sigma^2 \rangle_t \cdot
1_{\{(1-a)(1-b)\}}(r)
\\ & \quad \notag
+ \frac{1}{\alpha^{5-2b}\beta^2 T^3} \int_c^1 \phi^2(x)
x^{-(4-2b)} \mathrm{d} x \cdot \int_0^T F_1 \mathrm{d} t \cdot 1_{\{5a+2b-2ab-3\}}(r)
\\ & \quad
+ \frac{1_{(0, 1/2]}(c)}{2^{2-b}\alpha^{5-2b}\beta^2 T^3}
\int_c^{1/2} \phi(x) \phi(2x) x^{-(4-2b)} \mathrm{d} x \cdot \int_0^T
F_2 \mathrm{d} t \cdot 1_{\{5a+2b-2ab-3\}}(r).
\end{align}
\end{theorem}
Note that the asymptotic variances in
Theorems~\ref{thm:SALE-clt-noisefree} to \ref{thm:MSLE-clt-noisy} are
unobservable. Their consistent estimators and feasible central limit
theorems are detailed in Supplementary Material.
\section{Practical Aspects: Variances and Weights}
\label{sec:practical}
\subsection{Asymptotic Variances in Practice}
\label{sec:practical-variance}
Accurate asymptotic variance is important for parameter tuning. For
SALE, it helps pin down the optimal scale; for MSLE, it helps decide
the optimal weight distributions. Despite theoretical correctness,
the accuracy of derived variances may be affected by two situations
in practice: (i) small noise and (ii) violation of conditions in
Proposition~\ref{prop:adj-factor-asym}.
The asymptotic variances due to noise in
Proposition~\ref{prop:all-observation-noise}, \ref{prop:SALE-noise}
and \ref{prop:MSLE-noise} are established based on non-shrinkaging noise
assumptions. However, a small noise correction could be necessary in
practice, as some terms of small order become more pronounced as
noise becomes smaller. For all-observation estimators, we have
\begin{align}
\Delta_n^3 & k_n^2
\mathrm{Var}\bigl( \widehat{\langle X, \sigma^2 \rangle}^{\rm (all)} \big| \calF \bigr)
\xrightarrow{p}
\notag
\\
& \Bigl(
(8\nu_2\nu_4 + 16\nu_2^3 + 8\nu_3^2)
+ \frac{k_n}{n^2} \cdot (8\nu_4 + 16\nu_2^2) \int_0^T \sigma_t^2 \mathrm{d} t
+ \frac{1}{n^2} \cdot 8T \nu_2 \int_0^T \sigma_t^4 \mathrm{d} t
\Bigr) T.
\label{eq:all-observation-noise-variance-corrected}
\end{align}
Figure~\ref{fig:avar-noise} compares the simulated performance of
Equation~\eqref{eq:all-observation-noise-variance-corrected} and
Equation~\eqref{eq:all-observation-noise-variance}. Similar
correction for SALE estimators are provided in Supplementary Material.
\begin{figure}[!htp]
\centering
\centering
\includegraphics[width=0.5\textwidth]{figs/avar_noise_average.pdf}
\caption{
Simulated asymptotic variance due to noise of all-observation
estimators under different noise levels. A fixed Heston path from
Section~\ref{sec:simulation} is used for $X$, whereas i.i.d.
$\calN(0, \varsigma^2)$ random variables are used for $\varepsilon$.
``Empirical'': average of 5000 realizations of $\bigl( \widehat{\langle X, \sigma^2 \rangle}^{\rm
(all)}_T - \widetilde{\langle X, \sigma^2 \rangle}^{\rm (all)}_T \bigr)^2$. ``Dominant'': variance
calculated with
Equation~\eqref{eq:all-observation-noise-variance}.
``Corrected'': variance calculated with
Equation~\eqref{eq:all-observation-noise-variance-corrected}.
}
\label{fig:avar-noise}
\end{figure}
As for asymptotic variance due to discretization, the limit
expression in Equation~\eqref{eq:adj-factor-asym} does not work for
$H_q / H_p \to 0$, and can be inaccurate when any of $n, H_p, H_q$ is
not large enough
(Theorem~\ref{thm:SALE-clt-noisefree}\ref{thm:SALE-clt-noisefree-1}
is an example with fixed $H_p = H_q$). Consequently,
Proposition~\ref{prop:SALE-acov-noisefree} will improve the accuracy
of asymptotic variances in such cases.
\subsection{Approximate Weights of MSLE Estimators}
\label{sec:practical-weight}
Suppose that Assumption~\ref{ass:MSLE-para} holds. Let $\bm{\Sigma}
\in \bbR^{M_n \times M_n}$ denote the total asymptotic covariance
matrix between scales, the optimal weight assignment can be
obtained by solving
\begin{align}
\underset{\bm{w} \in \bbR^{M_n}}{\text{minimize}} & \quad V(\bm{w})
= \bm{w}^\sfT \bm{\Sigma} \bm{w}, \\
\text{subject to} & \quad \bm{w}^\sfT \bm{1}_{M_n} = 1.
\end{align}
The solution and the corresponding minimum are given by
\begin{align}\label{eq:weight-optimization-solution}
\bm{w}^*
=
\frac
{\bm{\Sigma}^{-1} \bm{1}_{M_n}}
{\bm{1}_{M_n}^\sfT \bm{\Sigma}^{-1} \bm{1}_{M_n}},
\quad
V(\bm{w}^*)
=
\frac{1}{\bm{1}_{M_n}^\sfT \bm{\Sigma}^{-1} \bm{1}_{M_n}}
=
\Biggl(\sum_{p=1}^{M_n}\sum_{q=1}^{M_n}
(\bm{\Sigma}^{-1})_{p,q}\Biggr)^{-1}.
\end{align}
However, direct application of
Equation~\eqref{eq:weight-optimization-solution} faces challenges in
practice: (i) the covariance matrix cannot be observed and therefore
estimated values are needed; (ii) the solution $\bm w^*$ can be
numerically unstable, sensitive to estimation errors of $\bm \Sigma$;
and (iii) calculating
Equation~\eqref{eq:weight-optimization-solution} requires matrix
inversion, which has an expensive time complexity of $O(M_n^3)$.
To address these challenges, we construct approximate weights for
MSLE estimators. For simplicity, hereafter in this section, we suppose that
Assumption~\ref{ass:MSLE-para-detail}\ref{ass:MSLE-para-detail-scale}
holds, and let
\begin{align}
s_1 = \frac{4}{\beta} \int_0^T \sigma_t^6 \mathrm{d} t, \quad
s_2 = \frac{2\beta T}{3}
\int_0^T \sigma_t^2 \mathrm{d} \langle \sigma^2, \sigma^2 \rangle_t, \quad
s_3 = \frac{8\nu_2\nu_4 + 16\nu_2^3 + 8\nu_3^2}{\alpha^{9/2} \beta^2 T^2}.
\end{align}
Detailed derivations for Equations~\eqref{eq:approx-weights} and
\eqref{eq:fredholm} presented in this section can be found in
Supplementary Material.
\subsubsection{Approximation in the Noise-Free Case}
\label{sec:practical-weight-noisefree}
In the absence of noise, the MSLE estimator can be used to enhance
the statistical efficiency of the all-observation estimator by
employing a more optimal weight assignment, rather than only
allocating all weight to the $H=1$ scale. For generality,
Definition~\ref{def:approx-weights} provides a closed-form expression
for the approximate weights applicable to all scales within $(m_n,
m_n+M_n]$. It is obtained by taking the limit $m_n \to \infty$ in
Equation~\eqref{eq:weight-optimization-solution}, where $\bm\Sigma$
is approximated by using Proposition~\ref{prop:adj-factor-asym} and
assuming that $s_2 \ll s_1$.
\begin{definition}[Approximate weights]\label{def:approx-weights}
Suppose that Assumptions~\ref{ass:MSLE-para} and
\ref{ass:MSLE-para-detail}\ref{ass:MSLE-para-detail-scale} hold.
The approximate weights are given by $\bm{\widetilde w} = (\bm
1_{M_n}^\sfT \bm{\widetilde\omega})^{-1} \bm{\widetilde\omega}$, where
\begin{align}\label{eq:approx-weights}
\bm{\widetilde\omega}_p =
\begin{cases}
2(m_n+1)^{-1/2}, & p = 1, \\
(m_n+p)^{-3/2}, & p = 2, \dots, M_n-1, \\
(m_n+M_n)^{-1/2}, & p = M_n.
\end{cases}
\end{align}
\end{definition}
Apart from being computationally and statistically efficient,
$\bm{\widetilde{w}}$ is numerically stable: since $\widetilde{w}_p >
0$ holds for any $p = 1, \dots, M_n$, we always have
$\sum_{p=1}^{M_n} |\widetilde{w}_p| = \sum_{p=1}^{M_n} \widetilde{w}_p = 1$.
\subsubsection{Approximation in the Noisy Case}
\label{sec:practical-weight-noisy}
Let $\varphi(x) = x^{-3/2} \phi(x)$, $\gamma=s_1/(3(s_1+s_2))$,
$\lambda=-(s_1+s_2)/s_3$, and
\begin{align}\label{eq:fredholm-kernel}
K(x,y) = x^{3/2} y^{3/2} (x\lor y)^{1/2} \Biggl(
\frac{x\land y}{x\lor y} - \gamma \left(\frac{x\land y}{x\lor y}\right)^2
\Biggr).
\end{align}
The optimal $\phi(x)$ in
Assumption~\ref{ass:MSLE-para-detail}\ref{ass:MSLE-para-detail-weight}
is related to a Fredholm integral equation:
\begin{align}\label{eq:fredholm}
\varphi(x) = kx^{3/2} + \lambda \int_c^1 K(x,y) \varphi(y) \mathrm{d} y,
\end{align}
where $k \in \bbR$ is a constant such that $\int_c^1 \varphi(x) x^{3/2}
\mathrm{d} x = 1$.
Equation~\eqref{eq:fredholm} relies on several simplifications:
(i) the conditions of Theorem~\ref{thm:MSLE-clt-noisy} hold with
$a=5/9$ and $b=1/2$, corresponding to the optimal convergence rate
of MSLE estimators;
(ii) the sparse off-diagonal terms of covariance due to noise are
omitted, as explained in Section~\ref{sec:MSLE}; and
(iii) finite-sample corrections in
Section~\ref{sec:practical-variance} are not considered.
These simplifications are made to isolate the dominant asymptotic
behavior for analytical tractablity.
\begin{figure}[!htp]
\centering
\begin{subfigure}[b]{0.24\linewidth}
\includegraphics[width=\linewidth]{figs/avar_weight_lamb=-1.pdf}
\caption{$\lambda=-1$}
\end{subfigure}
\hfill
\begin{subfigure}[b]{0.24\linewidth}
\includegraphics[width=\linewidth]{figs/avar_weight_lamb=-10.pdf}
\caption{$\lambda=-10$}
\end{subfigure}
\hfill
\begin{subfigure}[b]{0.24\linewidth}
\includegraphics[width=\linewidth]{figs/avar_weight_lamb=-100.pdf}
\caption{$\lambda=-10^2$}
\end{subfigure}
\hfill
\begin{subfigure}[b]{0.24\linewidth}
\includegraphics[width=\linewidth]{figs/avar_weight_lamb=-1000.pdf}
\caption{$\lambda=-10^3$}
\end{subfigure}
\begin{subfigure}[b]{0.24\linewidth}
\includegraphics[width=\linewidth]{figs/avar_weight_lamb=-10000.pdf}
\caption{$\lambda=-10^4$}
\end{subfigure}
\hfill
\begin{subfigure}[b]{0.24\linewidth}
\includegraphics[width=\linewidth]{figs/avar_weight_lamb=-100000.pdf}
\caption{$\lambda=-10^5$}
\end{subfigure}
\hfill
\begin{subfigure}[b]{0.24\linewidth}
\includegraphics[width=\linewidth]{figs/avar_weight_lamb=-1000000.pdf}
\caption{$\lambda=-10^6$}
\end{subfigure}
\hfill
\begin{subfigure}[b]{0.24\linewidth}
\includegraphics[width=\linewidth]{figs/avar_weight_lamb=-10000000.pdf}
\caption{$\lambda=-10^7$}
\end{subfigure}
\caption{
The effect of varying noise levels on the asymptotic variances of
SALE estimators (top panels) and the resulting optimal weights
for the MSLE estimator (bottom panels). Each subfigure
corresponds to a different noise level, parameterized by
$\lambda$, while the range of scales is fixed to isolate the
effect of noise. Asymptotic variances are shown on a logarithmic
scale. Parameters: $n=23400$, $m_n=20$, $M_n=180$, $H_n^*=200$,
$s_1=1/2$, $s_2 = 1/2$, $s_3 = -1/\lambda$.
}
\label{fig:avar-weight}
\end{figure}
Figure~\ref{fig:avar-weight} shows the impact of the noise level in
this simplified situation. The asymptotic variances of SALEs and the
optimal weights of the MSLE are presented on a fixed set of scales.
Note that a smaller $|\lambda|$ represents a larger noise magnitude.
For a small $|\lambda|$, the variances due to noise dominate in most
scales. Specifically, as $|\lambda| \to 0$, the integral term in
Equation~\eqref{eq:fredholm} vanishes, and the optimal weights are
given by $\phi(x) \to 4x^3$. Despite having a closed-form expression,
the result is not useful in practice, as noise always dominates the
total variance, so the estimation error is too large. On the other
hand, as $|\lambda| \to \infty$, noise has negligible contributions
to the total variances, and Equation~\eqref{eq:fredholm} becomes
ill-posed. However, the proposed weights in
Definition~\ref{def:approx-weights} offer good approximations in this case.
Beyond these simplifications, the intuition behind our multi-scale
approach and approximate weighting strategy is illustrated by the
signature plot in Figure~\ref{fig:signature}, which compares
different estimators across various scales (or pre-averaging window
lengths) for a simulated path under realistic noise. While
illustrative, these patterns are systematic and confirmed by the
extensive simulations in Section~\ref{sec:simulation-finite-sample}.
\begin{figure}[!ht]
\centering
\includegraphics[width=0.7\textwidth]{figs/signature_plot_simulation.pdf}
\caption{
Signature plot for a simulated path with time horizon $T=5/252$
and noise scale $\varsigma=3\times10^{-4}$.
}
\label{fig:signature}
\end{figure}
The plot reveals two key advantages. First, SALE outperforms the
pre-averaging estimator, as its minimum asymptotic variance (top
panel) is smaller, leading to a more efficient optimal
estimator.\footnote{The fundamental reason is that the SALE estimator
has a smaller coefficient of price variation error as shown in
Theorem~\ref{thm:SALE-clt-noisefree}, which is $8/3$, compared to a
larger coefficient $4$ in the pre-averaging estimator.} Second, MSLE
can improve upon SALE by averaging SALE estimates across an
appropriate range of scales (bottom panel) using an effective
weighting strategy, yielding a more accurate estimate with a tighter
confidence interval.
Based on this, we propose the following weighting method: (i) find an
optimal scale $\overline{H}_n$ for SALE estimators by minimizing the
total variance, using the corrections in
Section~\ref{sec:practical-variance}; and (ii) allocate weights by
applying Definition~\ref{def:approx-weights} with $m_n =
\overline{H}_n - 1$. This data-driven strategy ensures that MSLE
leverages the most informative scales, improving upon the optimal
SALE in two ways: the variance due to discretization is reduced
through the approximate weights, and the variance due to noise is
reduced because SALE estimators at subsequent scales exhibit smaller
noise contributions.
\section{Monte Carlo Simulations}\label{sec:simulation}
\subsection{Data Generating Processes}
The Heston model \citep{heston1993ClosedFormSolutionOptions} is used
to generate the discrete values of the underlying continuous
processes. The model is defined as
\begin{align} \mathrm{d} X_t &= \left(\mu - \frac{\sigma_t^2}{2}\right) \mathrm{d}
t + \sigma_t \mathrm{d} W_t, \\ \mathrm{d} \sigma_t^2 &= \kappa (\theta -
\sigma_t^2) \mathrm{d} t + \gamma \sigma_t \left(\rho \mathrm{d} W_t +
\sqrt{1-\rho^2} \mathrm{d} B_t \right),
\end{align} where $W_{t}$ and $B_{t}$ are independent Brownian
motions, with the leverage effect being $\langle X, \sigma^2 \rangle_T = \gamma \rho
\int_{0}^{T} V_{t} \mathrm{d} t$. The parameters are set as follows:
$\mu=0.02, \kappa=5, \theta=0.04, \gamma=0.5, \rho=-0.7$. The initial
values are set as $X_0 = 0, V_0 = 0.02$.
Let the variance of noise random variables $\{\varepsilon_i\}_{i=0}^n$ be
$\varsigma^2$. For independent noises, three distributions are considered:
(i) normal: $\varepsilon_i \sim \calN(0, \varsigma^2)$;
(ii) uniform: $\varepsilon_i \sim \mathrm{Unif}(-\sqrt{3}\varsigma, \sqrt{3}\varsigma)$; and
(iii) skew-normal: $\varepsilon_i$ has a PDF of $f(x) =
2\omega^{-1}\phi_0(\omega^{-1}(x-\xi)) \Phi_0(\alpha \omega^{-1}
(x-\xi))$, where $\phi_0$ and $\Phi_0$ are the PDF and CDF of
$\calN(0, 1)$, and $\xi=-\omega\delta\sqrt{2/\pi}$,
$\omega=\varsigma (1-2\delta^2/\pi)^{-1/2}$, $\delta=\alpha
(1+\alpha^2)^{-1/2}$, with the shape parameter $\alpha = 1$.
For dependent noises, consider
\begin{align}
\text{(i) MA(2) process:}
& \quad
\varepsilon_{i} = e_{i} + \theta_{1}e_{i-1} + \theta_{2}e_{i-2}, \quad
e_i \overset{\text{i.i.d.}}{\sim} \calN(0,
\varsigma^2(1+\theta_1^2+\theta_2^2)^{-1});
\text{ and}
\label{eq:ma2}
\\
\text{(ii) AR(1) process:}
& \quad
\varepsilon_i = \phi \varepsilon_{i-1} + e_i, \quad
e_i \overset{\text{i.i.d.}}{\sim} \calN(0, \varsigma^2\sqrt{1-\phi^2})
\text{ with } \phi \in (-1, 1).
\label{eq:ar1}
\end{align}
Specifically, the AR(1) process is included to evaluate the
robustness of the proposed estimators. While AR(1) noise is not
$q$-dependent and thus technically violates
Assumption~\ref{ass:noise}, it serves as a benchmark model for
persistent, serially correlated noise
\citep{jacod2017StatisticalPropertiesMicrostructure,
li2022ReMeDIMicrostructureNoise}. As our subsequent simulations will
confirm, the proposed estimators are indeed robust to this moderate
violation, preserving their asymptotic normality and superior
finite-sample efficiency.
\subsection{Asymptotic Normality}
\label{sec:simulation-asym-normal}
This section validates the central limit theorems by examining the
distribution of standardized estimation errors. For each estimator,
the error is standardized using both its infeasible and feasible
asymptotic variance. According to the results in
Section~\ref{sec:results}, these standardized errors should converge
to a standard normal distribution.
We simulate 5000 paths for each scenario, covering noise-free,
independent noise, and dependent noise. For data generation, we set
$T=1/252$, $n=23400$, $\varsigma = 0.005$, $\theta_1=\pm 0.7$,
$\theta_2=0.5$, and $\phi=0.7$. For estimators, we set $\beta=1/2$,
$b=1/2$, with scales and weights detailed in Table~\ref{tab:scales-weights}.
\begin{table}[!ht]
\centering
\caption{Scales and weights used in SALE and MSLE.}
\label{tab:scales-weights}
\begin{tabular}{llll}
\toprule
\multirow{2.5}{*}{\textbf{Noise}} &
\multicolumn{1}{c}{\textbf{SALE}} & \multicolumn{2}{c}{\textbf{MSLE}} \\
\cmidrule(lr){2-2} \cmidrule(lr){3-4}
& $H$ & $\{H_p\}_{p=1}^{M_n}$ & $\{w_p\}_{p=1}^{M_n}$ \\
\midrule
None & 1, 15 & $\{1, 2, \dotsc, 15\}$ & $w_p \propto H_p^{-3/2}$ \\
Independent & 1, 15 & $\{1, 2, \dotsc, 15\}$ & $w_p \propto H_p^3$ \\
Dependent & 15 & $\{11, 12, \dotsc, 15\}$ & $w_p \propto H_p^3$ \\
\bottomrule
\end{tabular}
\end{table}
Table~\ref{tab:standardized-errors} presents the summary statistics
for the standardized errors. Across all scenarios, these statistics
closely match those of a standard normal distribution, corroborating
our theoretical results (Theorems~\ref{thm:SALE-clt-noisefree} to
\ref{thm:MSLE-clt-noisy} and their feasible versions). Additional Q-Q
plots in the Supplementary Material further support these findings.
\begin{table}[!htb]
\centering
\caption{
Summary statistics for standardized estimation errors under
different settings. The standardized error is computed as
$(\text{Estimate} - \text{True Value}) / \sqrt{\text{Asymptotic
Variance}}$, using both infeasible and feasible versions of the
asymptotic variance. The mean, standard deviation and the 25th,
50th and 75th percentiles are reported. ``LE'' in the
``Estimator'' column represents the all-observation estimator
(SALE with $H=1$), whereas ``SALE'' represents SALE with $H=15$.
}
\label{tab:standardized-errors}
\resizebox{0.95\textwidth}{!}{
\begin{tabular}{lllrrrrrrrrrr}
\toprule
\multicolumn{2}{l}{\multirow{2}{*}{\textbf{Noise Setting}}} &
\multirow{2}{*}{\textbf{Estimator}} &
\multicolumn{5}{c}{\textbf{Infeasible}} &
\multicolumn{5}{c}{\textbf{Feasible}} \\
\cmidrule(lr){4-8} \cmidrule(lr){9-13}
& & & \textbf{Mean} & \textbf{Std} & $\bm{Q_1}$ & $\bm{Q_2}$ &
$\bm{Q_3}$ & \textbf{Mean} & \textbf{Std} & $\bm{Q_1}$ &
$\bm{Q_2}$ & $\bm{Q_3}$ \\
\midrule
\multicolumn{2}{l}{\multirow{3}{*}{\textbf{None}}} & LE &
-0.016 & 1.013 & -0.693 & -0.013 & 0.683 &
-0.014 & 1.013 & -0.693 & -0.013 & 0.684 \\
& & SALE & 0.012 & 0.990 & -0.650 & 0.012 & 0.685 &
0.013 & 0.996 & -0.652 & 0.012 & 0.687 \\
& & MSLE & -0.012 & 1.013 & -0.694 & 0.005 & 0.678 &
-0.010 & 1.014 & -0.696 & 0.005 & 0.676 \\
\cmidrule(lr){1-13}
\multirow{9}{*}{\textbf{Independent}} & \multirow{3}{*}{Normal} & LE
& -0.022 & 1.004 & -0.693 & -0.030 & 0.650 &
-0.022 & 0.998 & -0.691 & -0.030 & 0.643 \\
& & SALE & -0.008 & 1.003 & -0.693 & -0.007 & 0.673 &
-0.007 & 0.991 & -0.684 & -0.007 & 0.666 \\
& & MSLE & -0.005 & 1.047 & -0.718 & 0.000 & 0.702 &
-0.005 & 1.019 & -0.699 & 0.000 & 0.685 \\
\cmidrule(lr){2-13}
& \multirow{3}{*}{Uniform} & LE & 0.008 & 1.009 & -0.653 &
0.019 & 0.680 & 0.008 & 1.002 & -0.651 & 0.019 & 0.676 \\
& & SALE & 0.002 & 0.998 & -0.679 & -0.009 & 0.673 &
0.002 & 0.983 & -0.667 & -0.009 & 0.663 \\
& & MSLE & 0.004 & 1.023 & -0.707 & 0.008 & 0.689 &
0.004 & 0.989 & -0.683 & 0.007 & 0.668 \\
\cmidrule(lr){2-13}
& \multirow{3}{*}{Skew-normal} & LE & 0.010 & 1.004 & -0.662
& 0.018 & 0.685 & 0.010 & 0.999 & -0.655 & 0.018 & 0.678 \\
& & SALE & 0.008 & 1.010 & -0.681 & 0.015 & 0.691 &
0.008 & 0.999 & -0.676 & 0.014 & 0.686 \\
& & MSLE & 0.010 & 1.048 & -0.699 & 0.001 & 0.716 &
0.010 & 1.021 & -0.680 & 0.001 & 0.696 \\
\cmidrule(lr){1-13}
\multirow{6}{*}{\textbf{Dependent}} & \multirow{2}{*}{MA(2)
($\theta_1=0.7$)} & SALE & -0.013 & 1.012 & -0.706 & -0.024 & 0.676 &
-0.012 & 1.004 & -0.698 & -0.024 & 0.671 \\
& & MSLE & -0.008 & 1.032 & -0.727 & -0.007 & 0.696 &
-0.008 & 1.022 & -0.721 & -0.007 & 0.689 \\
\cmidrule(lr){2-13}
& \multirow{2}{*}{MA(2) ($\theta_1=-0.7$)} & SALE & 0.001 &
1.022 & -0.682 & 0.008 & 0.680 & 0.001 & 1.015 & -0.677 &
0.008 & 0.678 \\
& & MSLE & -0.001 & 1.073 & -0.717 & 0.004 & 0.727 &
-0.001 & 1.063 & -0.709 & 0.004 & 0.723 \\
\cmidrule(lr){2-13}
& \multirow{2}{*}{AR(1)} & SALE & -0.011 & 1.006 & -0.693 &
-0.015 & 0.666 & -0.010 & 0.996 & -0.690 & -0.015 & 0.657 \\
& & MSLE & -0.005 & 1.040 & -0.705 & -0.008 & 0.685 &
-0.003 & 1.028 & -0.695 & -0.008 & 0.680 \\
\midrule
\multicolumn{3}{l}{\textbf{Asymptotic Value (Standard Normal)}}
& 0.000 & 1.000 & -0.674 & 0.000 & 0.674 &
0.000 & 1.000 & -0.674 & 0.000 & 0.674 \\
\bottomrule
\end{tabular}
}
\end{table}
\subsection{Finite-Sample Performance: Superior Efficiency}
\label{sec:simulation-finite-sample}
\subsubsection{Efficiency in the Noise-Free Case}
The finite-sample efficiency of the MSLE estimator, using both
optimal and approximate weights, is compared against the
all-observation estimator. To evaluate the performance across
different sample sizes, four common time horizons are considered for
$T$: one day $(T=1/252)$, one week $(T=5/252)$, two weeks
$(T=10/252)$ and one month $(T=22/252)$. For each $T$, the sample
size is set to $n = 23400 \times 252T$, and 1000 paths are simulated.
The MSLE estimators are computed with scales $H_p = 1, 2, \dots,
\lfloor 0.5 n^{0.5}\rfloor$, and the optimal weights are given by
Equation~\eqref{eq:weight-optimization-solution} and
\eqref{eq:SALE-acov-disc}.
\begin{figure}[!ht]
\centering
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_estimate_nday=1.pdf}
\caption{1 day}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_estimate_nday=5.pdf}
\caption{5 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_estimate_nday=10.pdf}
\caption{10 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_estimate_nday=22}
\caption{22 days}
\end{subfigure}
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_error_nday=1.pdf}
\caption{1 day}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_error_nday=5.pdf}
\caption{5 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_error_nday=10.pdf}
\caption{10 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_error_nday=22.pdf}
\caption{22 days}
\end{subfigure}
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_avar_nday=1.pdf}
\caption{1 day}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_avar_nday=5.pdf}
\caption{5 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_avar_nday=10.pdf}
\caption{10 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisefree_avar_nday=22.pdf}
\caption{22 days}
\end{subfigure}
\caption{
The performances of the MSLE and all-observation estimators for
each setting of $T$ in the noise-free setting. The first row
shows the true and estimated values of leverage effect, the
second row shows the estimation errors, and the third row shows
the asymptotic variances.
}
\label{fig:noisefree-estimate}
\end{figure}
\begin{table}[!ht]
\centering
\caption{
Finite-sample performances of the MSLE and all-observation
estimators in the noise-free setting. The finite-sample relative
efficiency is compared with the all-observation estimator.
}
\label{tab:noisefree-rmse}
\resizebox{0.95\textwidth}{!}{
\begin{tabular}{rrrrrrrrrrrr}
\toprule
\multicolumn{1}{c}{\multirow{2.5}{*}{\textbf{Days}}} &
\multicolumn{2}{c}{\textbf{True Value}} &
\multicolumn{3}{c}{\textbf{RMSE}} &
\multicolumn{3}{c}{\textbf{Relative Efficiency}} &
\multicolumn{3}{c}{$\bm{\|w\|_1}$} \\
\cmidrule(lr){2-3} \cmidrule(lr){4-6} \cmidrule(lr){7-9}
\cmidrule{10-12}
& \multicolumn{1}{c}{Mean} & \multicolumn{1}{c}{Std} &
\multicolumn{1}{c}{Approximate} & \multicolumn{1}{c}{Optimal} &
\multicolumn{1}{c}{LE} & \multicolumn{1}{c}{Approximate} &
\multicolumn{1}{c}{Optimal} & \multicolumn{1}{c}{LE} &
\multicolumn{1}{c}{Approximate} & \multicolumn{1}{c}{Optimal} &
\multicolumn{1}{c}{LE} \\
\midrule
1 & \num{-2.80e-05} & \num{2.77e-06} & \num{2.55e-05} &
\num{2.44e-05} & \num{2.93e-05} & 1.32 & 1.44 & 1.00 & 1.00 &
13.77 & 1.00 \\
5 & \num{-1.48e-04} & \num{3.10e-05} & \num{4.57e-05} &
\num{4.50e-05} & \num{5.32e-05} & 1.36 & 1.40 & 1.00 & 1.00 &
24.77 & 1.00 \\
10 & \num{-3.06e-04} & \num{8.98e-05} & \num{6.43e-05} &
\num{6.44e-05} & \num{7.28e-05} & 1.28 & 1.28 & 1.00 & 1.00 &
28.53 & 1.00 \\
22 & \num{-7.44e-04} & \num{2.72e-04} & \num{1.06e-04} &
\num{1.07e-04} & \num{1.14e-04} & 1.16 & 1.13 & 1.00 & 1.00 &
34.83 & 1.00 \\
\bottomrule
\end{tabular}
}
\end{table}
Figure~\ref{fig:noisefree-estimate} and
Table~\ref{tab:noisefree-rmse} summarize the results. The findings
clearly demonstrate the superiority of the MSLE estimator over the
all-observation estimator, evidenced by its smaller asymptotic
variance, lower finite-sample RMSE, and higher finite-sample efficiency.
A key practical insight is that the MSLE with approximate weights
achieves efficiency nearly identical to that with optimal weights,
but offers significantly better numerical stability. By construction,
the approximate weights are non-negative, thus ensuring $\|\bm{w}\|_1
= 1$. In contrast, the optimal weights can become negative, causing
their $L^1$-norm to grow with the sample size. This not only
increases numerical instability, but also potentially violates our
theoretical requirement in Assumption~\ref{ass:MSLE-para}, making the
approximate weighting scheme a more robust choice for practical
implementation.
\subsubsection{Efficiency in the Noisy Case}
The finite-sample efficiency of the MSLE estimator using approximate
weights is compared against the pre-averaging LE estimator
in~\citet{aitsahalia2017EstimationContinuousDiscontinuous} in a noisy
setting. While the pre-averaging estimator has a slightly faster
theoretical convergence rate ($n^{-1/8}$) than the MSLE estimator
($n^{-1/9}$), we demonstrate that MSLE achieves superior
finite-sample efficiency, especially in realistic scenarios. To this
end, we again vary $T$ to assess the performance across different
sample sizes.
The simulation uses dependent AR(1) noise with $\phi = 0.7$, and
three noise levels: small ($\varsigma=10^{-4}$), medium
($\varsigma=10^{-3.5}$), and large ($\varsigma=10^{-3}$). Notably,
empirical evidence suggests that real-world noise levels are closer
to the ``small'' setting~\citep{christensen2014FactFrictionJumps},
which is further supported by our empirical study in
Section~\ref{sec:empirical}. For the MSLE estimator, the noise ACF is
truncated at $q=3$, the scales are set to $H_p=7, 8, \dots, \lfloor
n^{5/9} \rfloor$, and for simplicity, the same weight allocation is
used for all paths in each case. For the pre-averaging estimator,
since a closed-form optimal tuning parameter is unavailable, we grant
it an advantage by \emph{ex-post} selecting the pre-averaging window
that yields the minimum RMSE, from a wide grid of candidates ($5, 10,
30, 60, 90, 120, 180, 240,$ and $300$).
\begin{figure}[!ht]
\centering
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisy_dependent_estimate_nday=1_noise_std=0.0001.pdf}
\caption{1 day}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisy_dependent_estimate_nday=5_noise_std=0.0001.pdf}
\caption{5 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisy_dependent_estimate_nday=10_noise_std=0.0001.pdf}
\caption{10 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisy_dependent_estimate_nday=22_noise_std=0.0001.pdf}
\caption{22 days}
\end{subfigure}
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisy_dependent_error_nday=1_noise_std=0.0001.pdf}
\caption{1 day}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisy_dependent_error_nday=5_noise_std=0.0001.pdf}
\caption{5 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisy_dependent_error_nday=10_noise_std=0.0001.pdf}
\caption{10 days}
\end{subfigure}
\hfill
\begin{subfigure}{0.24\textwidth}
\includegraphics[width=\textwidth]{figs/performance_noisy_dependent_error_nday=22_noise_std=0.0001.pdf}
\caption{22 days}
\end{subfigure}
\caption{The performances of the MSLE and pre-averaging LE
estimators for each setting of $T$ in the dependent noise setting
($\varsigma=10^{-4}$). The first row shows the true and estimated
values of leverage effect, and the second row shows the estimation error.}
\label{fig:dependent-estimate}
\end{figure}
\begin{table}[!ht]
\centering
\caption{
Finite-sample performances of the MSLE and pre-averaging LE
estimators in the dependent noise setting. The finite-sample
relative efficiency is compared with the pre-averaging estimator.
}
\label{tab:dependent-rmse}
\resizebox{0.95\textwidth}{!}{
\begin{tabular}{lrrrrrrr}
\toprule
\multicolumn{1}{c}{\multirow{2.5}{*}{$\bm{\varsigma}$}} &
\multicolumn{1}{c}{\multirow{2.5}{*}{\textbf{Days}}} &
\multicolumn{2}{c}{\textbf{True Value}} &
\multicolumn{2}{c}{\textbf{RMSE}} &
\multicolumn{2}{c}{\textbf{Relative Efficiency}} \\
\cmidrule(lr){3-4} \cmidrule(lr){5-6} \cmidrule(lr){7-8}
& & \multicolumn{1}{c}{Mean} & \multicolumn{1}{c}{Std} &
\multicolumn{1}{c}{MSLE} & \multicolumn{1}{c}{Pre-Averaging LE}
& \multicolumn{1}{c}{MSLE} & \multicolumn{1}{c}{Pre-Averaging LE} \\
\midrule
\multirow{4}{*}{$10^{-4}$} & 1 & \num{-2.80e-05} &
\num{2.77e-06} & \num{4.78e-05} & \num{7.11e-05} & 2.21 & 1.00 \\
& 5 & \num{-1.48e-04} & \num{3.10e-05} & \num{8.67e-05} &
\num{1.21e-04} & 1.95 & 1.00 \\
& 10 & \num{-3.06e-04} & \num{8.98e-05} & \num{1.14e-04} &
\num{1.62e-04} & 2.00 & 1.00 \\
& 22 & \num{-7.44e-04} & \num{2.72e-04} & \num{1.74e-04} &
\num{2.50e-04} & 2.05 & 1.00 \\
\cmidrule(lr){1-8}
\multirow{4}{*}{$10^{-3.5}$} & 1 & \num{-2.80e-05} &
\num{2.77e-06} & \num{6.65e-05} & \num{9.28e-05} & 1.95 & 1.00 \\
& 5 & \num{-1.48e-04} & \num{3.10e-05} & \num{1.18e-04} &
\num{1.62e-04} & 1.86 & 1.00 \\
& 10 & \num{-3.06e-04} & \num{8.98e-05} & \num{1.63e-04} &
\num{2.20e-04} & 1.81 & 1.00 \\
& 22 & \num{-7.44e-04} & \num{2.72e-04} & \num{2.67e-04} &
\num{3.65e-04} & 1.87 & 1.00 \\
\cmidrule(lr){1-8}
\multirow{4}{*}{$10^{-3}$} & 1 & \num{-2.80e-05} &
\num{2.77e-06} & \num{9.63e-05} & \num{1.18e-04} & 1.50 & 1.00 \\
& 5 & \num{-1.48e-04} & \num{3.10e-05} & \num{1.82e-04} &
\num{2.27e-04} & 1.56 & 1.00 \\
& 10 & \num{-3.06e-04} & \num{8.98e-05} & \num{2.53e-04} &
\num{3.07e-04} & 1.48 & 1.00 \\
& 22 & \num{-7.44e-04} & \num{2.72e-04} & \num{3.82e-04} &
\num{4.86e-04} & 1.62 & 1.00 \\
\bottomrule
\end{tabular}
}
\end{table}
Figure~\ref{fig:dependent-estimate} (for $\varsigma=10^{-4}$) and
Table~\ref{tab:dependent-rmse} (for all noise levels) present the
results. The findings confirm that the MSLE estimator consistently
and substantially outperforms the pre-averaging estimator in terms of
finite-sample RMSE and efficiency across all sample sizes and noise
levels. Crucially, the advantage is most pronounced in the
small-noise setting, which is the most empirically relevant scenario.
Even as the time horizon increases to one month, MSLE's lead remains
significant, demonstrating that the theoretical convergence rate is
not the only determinant of the finite-sample performance. This
highlights the practical power of the proposed estimators and the
approximate weighting strategy. Furthermore, this superior
performance is achieved without resorting to the infeasible
\emph{ex-post} parameter tuning that was granted to the pre-averaging
estimator, underscoring the robustness and practical utility of our methods.
An additional study for the i.i.d. noise case, along with
supplementary information of the simulation details, are provided in
Supplementary Material.
\section{Empirical Study}\label{sec:empirical}
The high-frequency trading data for a selection of assets, covering
the regular trading hours from 2014 to 2023 (2,516 trading days), are
collected from the TAQ database. The data are cleaned before analysis,
retaining only regular trades and removing erroneous
entries.\footnote{A practical and detailed guideline on
high-frequency data cleaning is offered by
\cite{barndorffnielsen2009RealizedKernelsPractice}. While we follow
most of the steps therein, some are omitted. For example, the entries
with \emph{Sale Conditions} `I' (odd lot trade) and `C' (cash trade)
are retained because of their significant contribution in our
dataset. We also remove the ``bounceback'' outliers described by
\cite{aitsahalia2011UltraHighFrequency}.} The dataset consists of 15
ETFs and 15 individual stocks, as listed in Table~\ref{tab:emp}. The
ETFs track the performance of the S\&P 500, NASDAQ 100, Dow Jones
Industrial Average, Russell 2000 indices, as well as the 11 sectors
of the S\&P 500. The individual stocks are selected to represent a
range of liquidity and volatility levels, covering various sectors
such as technology, consumer goods, healthcare, and entertainment,
thus providing a diverse set of assets for the empirical study. Among
the 30 assets, XLC and XLRE were issued partway through the sample
period. Therefore, our analysis for them begins at the start of their
second year, in 2017 and 2019, respectively. After data cleaning, we
resample the data to obtain 1-second and 5-second returns.
We apply the jump test proposed by
\cite{aitsahalia2012TestingJumpsNoisy} to identify and remove trading
days with the presence of jumps for each stock. This test is a
robustified version of the test introduced by
\cite{aitsahalia2009TestingJumpsDiscretely}, incorporating the
pre-averaging method to deal with the MMS noise. After computing the
standardized statistics with 5-second intraday data, we apply the
universal threshold technique proposed by
\cite{bajgrowicz2016JumpsHighFrequencyData} to eliminate spurious
jump detections. This method is more stringent than the FDR procedure
and is designed to asymptotically remove all spurious detections,
thereby minimizing data loss. As a result, 909 asset-days, comprising
1.2\% of the entire dataset of 73,910 asset-days, are identified as
containing jumps and excluded from further analysis. The numbers of
days with jumps for each asset are listed in Table~\ref{tab:emp}.
We estimate the leverage effects for both weekly (defined as every
five trading days) and monthly (defined as natural months) periods
using both 1-second and 5-second data for each stock. The estimation
proceeds in several steps, showcasing the flexibility and robustness
of our framework in handling real-world data complexities.
\begin{enumerate}
\item The ReMeDI estimator proposed by
\cite{li2022ReMeDIMicrostructureNoise} is used to estimate the
moments $\nu_2$, $\nu_4$ and the generalized acfs $\rho_2(l)$,
$\rho_3(l)$, and $\rho_4(l)$ of the MMS noise in each period. The
existence and the dependence level of noise are determined by its
autocovariances. The results show that the noise in the dataset
is small, while the dependencies are common. For example, the
analysis of noise in monthly 5-second frequency data shows that:
(i) for the ETFs, 33.9\% of asset-months exhibit significant
noise, among which 71.5\% are dependent and the average noise
scale is $\varsigma=9.6\times 10^{-5}$; while (ii) for the
stocks, 63.5\% of the asset-months exhibit significant noise,
among which 60.8\% are dependent and the average noise scale is
$\varsigma=2.0\times 10^{-4}$.
\item The scale $\overline{H}_n$ in the approximate weights of the
MSLE estimator is determined by minimizing the total asymptotic
variances of SALE estimators, where the asymptotic variances are
estimated using 1-minute pre-averaging return data. With the
existence of noise, additional lower bounds for $\overline{H}_n$
are applied, such that: (i) $\overline{H}_n$ satisfies
$\overline{H}_n \geq 2\hat q + 1$, where $\hat q$ is the
estimated dependence level; and (ii) the minimum values of
$\overline{H}_n$ are 20 for 1-second data (corresponding to 20
seconds) and 12 for 5-second data (corresponding to 60 seconds).
The former is a condition for the proposed SALE and MSLE
estimators, while the latter is a rather conservative manual
intervention, which mitigates numerical instability at the cost
of larger asymptotic variance.
\item The leverage effect and its asymptotic variance are estimated
using the MSLE estimator with approximate weights, where the
number of scales is set to $M_n=50$ for computational
efficiency.
\end{enumerate}
\begin{figure}[!ht]
\centering
\begin{subfigure}{\textwidth}
\includegraphics[width=\textwidth]{figs/AMZN_month_5s.pdf}
\caption{Month data, 5-second frequency}
\end{subfigure}
\begin{subfigure}{\textwidth}
\includegraphics[width=\textwidth]{figs/AMZN_week_5s.pdf}
\caption{Week data, 5-second frequency}
\end{subfigure}
\caption{Leverage effect estimation for AMZN (Amazon.com, Inc.).}
\label{fig:estimation-amzn}
\end{figure}
Figure~\ref{fig:estimation-amzn} illustrates the dynamic nature of
the leverage effect for AMZN, showcasing our estimator's ability to
capture its time-varying behavior. The monthly estimates reveal
significant fluctuations, clearly capturing major market stress
events such as the February 2018 ``Volpocalypse'' and the COVID-19
sell-off in early 2020. The weekly estimates, while more volatile,
provide a higher-resolution view of these dynamics.
\begin{table}[!ht]
\centering
\caption{Data descriptions and results of empirical study.}
\label{tab:emp}
\resizebox{1.00\textwidth}{!}{
\begin{tabular}{lllrrrrrrrrrrrrrr}
\toprule
\multicolumn{1}{c}{\multirow{3.5}{*}{\textbf{Type}}} &
\multicolumn{1}{c}{\multirow{3.5}{*}{\textbf{Ticker}}} &
\multicolumn{1}{c}{\multirow{3.5}{*}{\textbf{Name}}} &
\multicolumn{1}{c}{\multirow{3.5}{*}{\textbf{\shortstack{Average
\\ Daily \\ Observations}}}} &
\multicolumn{1}{c}{\multirow{3.5}{*}{\textbf{\shortstack{Average
\\ Daily \\ Volume}}}} &
\multicolumn{1}{c}{\multirow{3.5}{*}{\textbf{\shortstack{Annualized
\\ Volatility}}}} &
\multicolumn{1}{c}{\multirow{3.5}{*}{\textbf{\shortstack{Trading
\\ Days}}}} &
\multicolumn{1}{c}{\multirow{3.5}{*}{\textbf{\shortstack{Days
\\ with \\ Jumps}}}} & \multicolumn{8}{c}{\textbf{Signs of
Leverage Effects (\%)}} \\
\cmidrule(lr){9-16}
& & & & & & & & \multicolumn{2}{c}{\textbf{M, 1-sec}} &
\multicolumn{2}{c}{\textbf{M, 5-sec}} &
\multicolumn{2}{c}{\textbf{W, 1-sec}} &
\multicolumn{2}{c}{\textbf{W, 5-sec}} \\
\cmidrule(lr){9-10} \cmidrule(lr){11-12} \cmidrule(lr){13-14}
\cmidrule(lr){15-16}
& & & & & & & & \multicolumn{1}{c}{$\bm{-}$} &
\multicolumn{1}{c}{$\bm{+}$} & \multicolumn{1}{c}{$\bm{-}$} &
\multicolumn{1}{c}{$\bm{+}$} & \multicolumn{1}{c}{$\bm{-}$} &
\multicolumn{1}{c}{$\bm{+}$} & \multicolumn{1}{c}{$\bm{-}$} &
\multicolumn{1}{c}{$\bm{+}$} \\
\midrule
\multirow{15}{*}{\textbf{ETF}} & SPY & SPDR S\&P 500 ETF Trust
& 428321 & \num{9.27e+07} & 0.175 & 2516 & 5 & 89.2 & 10.8 &
93.3 & 6.7 & 83.3 & 16.7 & 86.7 & 13.3 \\
& QQQ & Invesco QQQ Trust & 219580 & \num{4.14e+07} & 0.215 &
2516 & 3 & 88.3 & 11.7 & 90.8 & 9.2 & 84.9 & 15.1 & 86.7 & 13.3 \\
& DIA & SPDR Dow Jones Industrial Average ETF Trust & 35501 &
\num{4.53e+06} & 0.174 & 2516 & 11 & 86.7 & 13.3 & 90.0 & 10.0
& 77.4 & 22.6 & 82.1 & 17.9 \\
& IWM & iShares Russell 2000 ETF & 153098 & \num{2.98e+07} &
0.221 & 2516 & 5 & 86.7 & 13.3 & 89.2 & 10.8 & 77.0 & 23.0 &
80.4 & 19.6 \\
& XLC & The Communication Services Select Sector SPDR ETF Fund
& 27626 & \num{4.65e+06} & 0.242 & 1393 & 6 & 88.3 & 11.7 &
86.7 & 13.3 & 75.1 & 24.9 & 76.3 & 23.7 \\
& XLY & The Consumer Discretionary Select Sector SPDR Fund &
49127 & \num{5.63e+06} & 0.208 & 2516 & 12 & 88.3 & 11.7 & 90.0
& 10.0 & 74.2 & 25.8 & 79.8 & 20.2 \\
& XLP & The Consumer Staples Select Sector SPDR Fund & 45530 &
\num{1.19e+07} & 0.146 & 2516 & 21 & 70.8 & 29.2 & 75.8 & 24.2
& 62.5 & 37.5 & 61.9 & 38.1 \\
& XLE & The Energy Select Sector SPDR Fund & 112759 &
\num{2.10e+07} & 0.298 & 2516 & 24 & 71.7 & 28.3 & 75.8 & 24.2
& 62.5 & 37.5 & 64.9 & 35.1 \\
& XLF & The Financial Select Sector SPDR Fund & 72205 &
\num{5.54e+07} & 0.221 & 2516 & 15 & 81.7 & 18.3 & 87.5 & 12.5
& 71.2 & 28.8 & 71.6 & 28.4 \\
& XLV & The Health Care Select Sector SPDR Fund & 63047 &
\num{9.89e+06} & 0.170 & 2516 & 19 & 82.5 & 17.5 & 88.3 & 11.7
& 69.0 & 31.0 & 68.8 & 31.2 \\
& XLI & The Industrial Select Sector SPDR Fund & 65217 &
\num{1.15e+07} & 0.196 & 2516 & 13 & 80.8 & 19.2 & 90.8 & 9.2 &
71.4 & 28.6 & 73.8 & 26.2 \\
& XLB & The Materials Select Sector SPDR Fund & 38370 &
\num{6.33e+06} & 0.206 & 2516 & 20 & 73.3 & 26.7 & 77.5 & 22.5
& 69.0 & 31.0 & 68.8 & 31.2 \\
& XLRE & The Real Estate Select Sector SPDR Fund & 15748 &
\num{4.21e+06} & 0.214 & 2069 & 30 & 73.8 & 26.2 & 75.0 & 25.0
& 59.8 & 40.2 & 63.5 & 36.5 \\
& XLK & The Technology Select Sector SPDR Fund & 60643 &
\num{1.04e+07} & 0.226 & 2516 & 7 & 89.2 & 10.8 & 91.7 & 8.3 &
83.3 & 16.7 & 83.7 & 16.3 \\
& XLU & The Utilities Select Sector SPDR Fund & 66683 &
\num{1.48e+07} & 0.191 & 2516 & 48 & 64.2 & 35.8 & 65.0 & 35.0
& 55.4 & 44.6 & 57.1 & 42.9 \\
\cmidrule(lr){1-16}
\multirow{15}{*}{\textbf{Stock}} & AAPL & Apple Inc. & 368263 &
\num{1.37e+08} & 0.284 & 2516 & 22 & 76.7 & 23.3 & 75.8 & 24.2
& 70.2 & 29.8 & 73.6 & 26.4 \\
& AMC & AMC Entertainment Holdings, Inc. & 104278 &
\num{2.77e+06} & 1.352 & 2516 & 56 & 50.8 & 49.2 & 48.3 & 51.7
& 47.8 & 52.2 & 49.6 & 50.4 \\
& AMZN & Amazon.com, Inc. & 164947 & \num{8.02e+07} & 0.332 &
2516 & 21 & 78.3 & 21.7 & 76.7 & 23.3 & 68.7 & 31.3 & 68.8 & 31.2 \\
& CLX & The Clorox Company & 16352 & \num{1.18e+06} & 0.227 &
2516 & 70 & 52.5 & 47.5 & 55.8 & 44.2 & 49.6 & 50.4 & 51.2 & 48.8 \\
& CPB & The Campbell's Company & 18902 & \num{2.32e+06} & 0.236
& 2516 & 66 & 51.7 & 48.3 & 49.2 & 50.8 & 54.8 & 45.2 & 53.6 & 46.4 \\
& GME & GameStop Corp. & 51874 & \num{1.82e+07} & 1.080 & 2516
& 52 & 43.3 & 56.7 & 49.2 & 50.8 & 47.2 & 52.8 & 48.2 & 51.8 \\
& KO & The Coca-Cola Company & 78632 & \num{1.42e+07} & 0.180 &
2516 & 44 & 65.8 & 34.2 & 67.5 & 32.5 & 55.4 & 44.6 & 59.1 & 40.9 \\
& MRK & Merck \& Co., Inc. & 68200 & \num{1.08e+07} & 0.214 &
2516 & 52 & 64.2 & 35.8 & 56.7 & 43.3 & 56.0 & 44.0 & 50.8 & 49.2 \\
& MSFT & Microsoft Corporation & 231381 & \num{3.02e+07} &
0.271 & 2516 & 16 & 72.5 & 27.5 & 73.3 & 26.7 & 65.5 & 34.5 &
66.3 & 33.7 \\
& NVDA & NVIDIA Corporation & 204440 & \num{4.58e+08} & 0.464 &
2516 & 30 & 71.7 & 28.3 & 76.7 & 23.3 & 68.1 & 31.9 & 68.7 & 31.3 \\
& PEP & PepsiCo, Inc. & 44764 & \num{4.69e+06} & 0.184 & 2516 &
44 & 66.7 & 33.3 & 68.3 & 31.7 & 52.0 & 48.0 & 54.6 & 45.4 \\
& PFE & Pfizer Inc. & 113288 & \num{2.79e+07} & 0.226 & 2516 &
40 & 65.8 & 34.2 & 61.7 & 38.3 & 57.3 & 42.7 & 57.1 & 42.9 \\
& PG & The Procter \& Gamble Company & 61423 & \num{8.27e+06} &
0.182 & 2516 & 61 & 69.2 & 30.8 & 65.8 & 34.2 & 59.1 & 40.9 &
61.5 & 38.5 \\
& TAP & Molson Coors Beverage Company & 16751 & \num{1.83e+06}
& 0.282 & 2516 & 67 & 60.0 & 40.0 & 59.2 & 40.8 & 52.0 & 48.0 &
52.6 & 47.4 \\
& TSLA & Tesla, Inc. & 368486 & \num{1.13e+08} & 0.557 & 2516 &
29 & 75.8 & 24.2 & 74.2 & 25.8 & 63.7 & 36.3 & 63.7 & 36.3 \\
\bottomrule
\end{tabular}
}
\end{table}
\begin{figure}[!ht]
\centering
\includegraphics[width=0.5\textwidth]{figs/stocks.pdf}
\caption{Individual stocks included in the empirical study. The
sizes of the circles represent the average number of daily
observations in our dataset, while the colors represent the
proportion of negative leverage effect detected using monthly data
sampled at 1-second frequency.}
\label{fig:stocks}
\end{figure}
As summarized in Table~\ref{tab:emp} and visualized in
Figure~\ref{fig:stocks}, a negative leverage effect is predominant
across most assets, particularly within established large-cap and
defensive stocks, consistent with financial theory. The notable
exceptions are the ``meme stocks'' AMC and GME, where
retail-investor-driven speculative trading results in extreme
volatility and a weaker or non-negative leverage effect. This
demonstrates our method's ability to uncover such asset-specific idiosyncrasies.
\section{Conclusion}
We introduce a multi-scale framework for the robust and efficient
estimation of the leverage effect from high-frequency data
contaminated by complex, serially dependent microstructure noise. We
construct two estimators, SALE and MSLE, by combining the shifted
window, subsampling, and multi-scale techniques. We develop the
asymptotic theory, and design an effective weighting strategy for the
MSLE estimator. Beyond noise robustness, a central merit of our
framework is its superior efficiency. In the absence of noise, the
MSLE estimator improves the efficiency of the base estimator. Under a
realistic setting of noise and sample size, the SALE estimator
already outperforms the standard pre-averaging estimator, and the
MSLE estimator further improves upon this, delivering consistently
more accurate and reliable estimates. Extensive simulations and
empirical applications have validated the asymptotic theory,
finite-sample performance, and the practical robustness and
flexibility of the proposed methods.
\bibliographystyle{plainnat}
\bibliography{ref.bib}
\pagebreak