Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Nonparametric Estimation of Large Spot Volatility Matrices for High-Frequency Financial Data
\centerline{\bf Abstract}
In this paper, we consider estimating spot/instantaneous volatility matrices
of high-frequency data collected for a large number of assets. We first
combine classic nonparametric kernel-based smoothing with a generalised
shrinkage technique in the matrix estimation for noise-free data under a
uniform sparsity assumption, a natural extension of the approximate sparsity
commonly used in the literature. The uniform consistency property is derived
for the proposed spot volatility matrix estimator with convergence rates
comparable to the optimal minimax one. For the high-frequency data
contaminated by microstructure noise, we introduce a localised
pre-averaging estimation method that reduces the effective magnitude of the noise.
We then use the estimation tool developed in the noise-free scenario, and derive the uniform
convergence rates for the developed spot volatility matrix estimator. We further combine the kernel smoothing with the shrinkage technique to estimate the time-varying volatility matrix of the high-dimensional noise
vector. In addition, we consider large spot volatility matrix estimation in time-varying factor models with observable risk factors and derive the uniform convergence property. We provide numerical studies including simulation and empirical application to examine the performance of the proposed estimation methods in finite samples.
\noindentKeywords: \ Brownian semi-martingale, Factor model, Kernel smoothing,
Microstructure noise, Sparsity, Spot volatility matrix, Uniform consistency.
Introduction
\setcounter{equation}{0}
Modelling high-frequency financial data is one of the most important topics
in financial economics and has received increasing attention in recent
decades. Continuous-time econometric models such as the It\^{o}
semimartingale are often employed in the high-frequency data analysis. One
of the main components in these models is the volatility function or matrix. In the low-dimensional setting
(with a single or a small number of assets), the realised volatility is often used to estimate the integrated volatility over
a fixed time period AB98, BS02, BS04, ABDL03. In practice,
it is not uncommon that the high-frequency financial data are contaminated
by the market microstructure noise, which leads to biased realised
volatility if the noise is ignored. Hence,
various modification techniques such as the two-scale, pre-averaging and
realised kernel have been introduced to account for the microstructure noise
and produce consistent volatility estimation ZMA05, BHLS08, KL08, JLMPV09, PV09,
CKP10, PHL16. S05, ABD10 and AJ14 provide
comprehensive reviews for estimating volatility with high-frequency
financial data under various settings.
In practical applications, financial economists often have to deal with the situation
that there are a large amount of high-frequency financial data collected for
a large number of assets. A key issue is to estimate the large volatility
structure for these assets, which has applications in various areas such as
the optimal portfolio choice and risk management. Partly motivated by
developments in large covariance matrix estimation for low-frequency data in
the statistical literature, WZ10, TWZ13 and KWZ16
estimate the large volatility matrix under an approximate sparsity
assumption BL08; ZL11 and XZ18 study large
volatility matrix estimation using the large-dimensional random matrix
theory BS10; and LF18 propose a nonparametric
eigenvalue-regularised integrated covariance matrix for high-dimensional
asset returns. Given that there often exists co-movement between a large
number of assets and the co-movement is driven by some risk factors which
can be either observable or latent, FFX16, AX17, DLX19
extend the methodologies developed by FLM11, FLM13 to estimate the large volatility matrix by imposing a
continuous-time factor model structure on the high-dimensional and
high-frequency financial data, and AX19 study the principal component
analysis of high-frequency data and derive the asymptotic distribution for
the realised eigenvalues, eigenvectors and principal
components.
The estimation methodologies in the aforementioned literature often rely on the
realised volatility (or covariance) matrices, measuring the integrated
volatility structure over a fixed time interval. In practice, it is often
interesting to further explore the actual spot/instantaneous volatility
structure and its dynamic change over a certain time interval, which is a
particularly important measurement for the financial assets when the market
is in a volatile period (say, the global financial crisis or COVID-19
outbreak). For a single financial asset, FW08 and K10
introduce a kernel-based nonparametric method to estimate the spot
volatility function and establish its asymptotic properties including the
point-wise and global asymptotic distribution theory and uniform
consistency. For the noise-contaminated high-frequency data, ZB14
combine the two-scale realised volatility with the kernel-weighted technique
to estimate the spot volatility, whereas KK16 propose a
kernel-weighted pre-averaging spot volatility estimation method. Other
nonparametric spot volatility estimation methods can be found in FFL07
and FL20. It seems straightforward to extend this local
nonparametric method to estimate the spot volatility matrix for a small
number of assets. However, a further extension to the setting with vast
financial assets is non-trivial. There is virtually no work on estimating
the large spot volatility matrix except K18, which considers estimating
large spot volatility matrices and their integrated versions under the
continuous-time factor model structure for noise-free high-frequency data.
The main methodological and theoretical contributions of this paper are summarised as follows.
itemize• {\em Large spot volatility matrix estimation with noise-free high-frequency data}. We use the nonparametric kernel-based smoothing method to estimate the volatility and co-volatility
functions as in FW08 and K10, and then apply a generalised
shrinkage to off-diagonal estimated entries. With small off-diagonal entries
forced to be zeros, the resulting large spot volatility matrix estimate
would be non-degenerate with stable performance in finite samples. We derive the consistency property for the proposed spot volatility matrix estimator uniformly over the entire time interval under a uniform sparsity
assumption, which is also adopted by CXW13, CL16 and CLL19 in the low-frequency data setting. In particular, the derived
uniform convergence rates are comparable to the optimal minimax rate in
large covariance matrix estimation CZ12. The number of
assets is allowed to be ultra large in the sense that it can grow at an
exponential rate of $1/\Delta$ with $\Delta$ being the sampling frequency.
• {\em Large spot volatility matrix estimation with noise-contaminated
high-frequency data and time-varying noise volatility matrix estimation}. When the high-frequency data are contaminated by the microstructure noise, we extend KK16's localised pre-averaging estimation method to the high-dimensional data setting. Specifically, we first pre-average the log price data via a kernel
filter and then apply the same estimation method to the kernel fitted
high-frequency data (at pseudo-sampling time points) as in the noise-free
scenario. The microstructure noise vector is assumed to be heteroskedastic with the time-varying covariance matrix satisfying the uniform sparsity assumption. We show that the existence of microstructure noises slows down the uniform convergence rates, see Theorem (ref). Furthermore, we combine the kernel
smoothing with generalised shrinkage to estimate the time-varying noise
volatility matrix and derive its uniform convergence property. To the best of our knowledge, there is virtually no work on large time-varying noise volatility matrix estimation for high-frequency data.
• {\em Large spot volatility matrix estimation with risk factors}. Since the uniform sparsity assumption is often too restrictive, we relax this restriction in Section (ref) and consider large spot volatility matrix estimation in the time-varying factor model at high frequency, i.e., a large number of asset prices are driven by a small number of observable common factors. By imposing the sparsity restriction on the spot idiosyncratic volatility matrix, we obtain the so-called “low-rank plus sparse" spot volatility structure. A similar structure (with constant betas) is adopted by FFX16 and DLX19 in estimation of large integrated volatility matrices. We use the kernel smoothing method to estimate the spot volatility and covariance of the observed asset prices and factors as well as the time-varying betas, and apply the shrinkage technique to the estimated spot idiosyncratic volatility matrix. We derive the uniform convergence property of the developed matrix estimates, partly extending the point-wise convergence property in K18. The developed methodology and theory can be further modified to tackle the noise-contaminated high-frequency data.
We argue that all three of the scenarios we consider above may be practically relevant. Microstructure noise is considered important in the very highest frequency of data,
whereas researchers working with five minute data, say, often ignore the noise. For the lower frequency of data there is a lot comovement in returns and the factor model is
designed to capture that comovement, whereas at the ultra high frequency, comovement is less of an issue; indeed, under the so-called Epps effect this comovement shrinks to zero
with sampling frequency.
The rest of the paper is organised as follows. In Section (ref), we
estimate the large spot volatility matrix in the noise-free high-frequency
data setting and give the uniform consistency property. In Section (ref)
, we extend the methodology and theory to the noise-contaminated data
setting and further estimate the time-varying noise volatility matrix. Section (ref) considers the large spot volatility matrix with systematic factors. Section (ref) reports the
simulation studies and Section (ref) provides an empirical application. Section (ref) concludes the paper. Proofs of the main theoretical results are available in Appendix A. The supplementary document contains proofs of some technical lemmas and propositions and discussions on the spot precision matrix estimation and the asynchronicity issue. Throughout the
paper, we let $\Vert\cdot\Vert_2$ be the Euclidean norm of a vector; and for
a $d\times d$ matrix ${\mathbf{A}}=(A_{ij})_{d\times d}$, we let $\Vert {
\mathbf{A}}\Vert$ and $\Vert {\mathbf A}\Vert_F$ be the matrix spectral norm and Frobenius norm, $\vert{\mathbf{A}}
\vert_1=\sum_{i=1}^d\sum_{i=1}^d |A_{ij}|$, $\Vert{\mathbf{A}}
\Vert_1=\max_{1\leq j\leq d}\sum_{i=1}^d |A_{ij}|$, $\Vert{\mathbf{A}}
\Vert_{\infty,q}=\max_{1\leq i\leq d}\sum_{j=1}^d |A_{ij}|^q$ and $\Vert {
\mathbf{A}}\Vert_{\max}=\max_{1\leq i\leq d}\max_{1\leq j\leq d} |A_{ij}|$.
Estimation with noise-free data
\setcounter{equation}{0}
Suppose that ${\mathbf{X}}_t=\left(X_{1,t},\cdots,X_{p,t}\right)^{^\intercal}
$ is a $p$-variate Brownian semi-martingale solving the following stochastic
differential equation:
equation[equation omitted — 110 chars of source]
where ${\mathbf{W}}_t=\left(W_{1,t},\cdots,W_{p,t}\right)^{^\intercal}$ is a
$p$-dimensional standard Brownian motion, ${\boldsymbol{\mu}}
_t=(\mu_{1,t},\cdots,\mu_{p,t})^{^\intercal}$ is a $p$-dimensional drift
vector, and ${\boldsymbol{\sigma}}_t=\left(\sigma_{ij,t}\right)_{p\times p}$
is a $p\times p$ matrix. The spot volatility matrix of ${\mathbf{X}}_t$ is
defined as
equation[equation omitted — 155 chars of source]
Our main interest lies in estimating ${\boldsymbol{\Sigma}}_t$ when the size
$p$ is large. As in CXW13 and CL16, we assume that the true
spot volatility matrix satisfies the following uniform sparsity condition: $
\left\{{\boldsymbol{\Sigma}}_t:\ 0\leq t\leq T\right\}\in \mathcal{S}
(q,\varpi(p), T)$, where
equation[equation omitted — 246 chars of source]
where $0\le q<1$, $\varpi(p)$ is larger than a positive constant, $T$ is a fixed positive number and $\Lambda$ is a positive
random variable satisfying $\mathsf{E}[\Lambda]\leq C_\Lambda<\infty$. This
is a natural extension of the approximate sparsity assumption
BL08. Section (ref) below will relax this assumption and consider estimating large spot volatility matrices with systematic factors. The asset prices are assumed to be collected
over a fixed time interval $[0,T]$ at $0,\Delta,2\Delta,\cdots,n\Delta$,
where $\Delta$ is the sampling frequency and $n=\lfloor T/\Delta\rfloor$
with $\lfloor\cdot\rfloor$ denoting the floor function. In the main text, we
focus on the case of equidistant time points in the high-frequency data
collection. The asynchronicity issue will be discussed in Appendix C.2 of the supplement.
For each $1\leq i,j\leq p$, we estimate the spot co-volatility $\Sigma_{ij,t}
$ by
equation[equation omitted — 109 chars of source]
with
\[
K_{h}^\ast(t_k-t)=K_h\left(t_k-t\right)/\left[\Delta\sum_{l=1}^nK_h\left(t_l-t\right)\right],
\]
where $t_k=k\Delta$, $K_h(u)=h^{-1}K(u/h)$, $K(\cdot)$ is a kernel function,
$h$ is a bandwidth shrinking to zero and $\Delta
X_{i,k}=X_{i,t_k}-X_{i,t_{k-1}}$. The use of $K_h^\ast(t_k-t)$ rather than $K_h(t_k-t)$ in the estimation ((ref)) is to correct a constant bias when $t$ is close to the boundary points $0$ and $T$. A naive method of estimating the spot
volatility matrix ${\boldsymbol{\Sigma}}_{t}$ is to directly use $\widehat\Sigma_{ij,t}$ to form an estimated matrix. However, this
estimate often performs poorly in practice when the number of assets is very
large (say, $p>n$). To address this issue, a commonly-used technique is to
apply a shrinkage function to $\widehat\Sigma_{ij,t}$ when $i\neq j$,
forcing very small estimated off-diagonal entries to be zeros. Let $
s_\rho(\cdot)$ denote a shrinkage function satisfying the following three
conditions: (i) $\vert s_\rho(u)\vert \leq \vert u\vert$ for $u\in\mathscr{R}
$; (ii) $s_\rho(u)=0$ if $\vert u\vert \leq \rho$; and (iii) $\vert
s_\rho(u)-u\vert\leq \rho$, where $\rho$ is a user-specified tuning
parameter. With the shrinkage function, we construct the following
nonparametric estimator of ${\boldsymbol{\Sigma}}_{t}$:
equation[equation omitted — 235 chars of source]
where $\rho_1(t)$ is a tuning parameter which is allowed to change over $t$
and $I(\cdot)$ denotes the indicator function. Section (ref) discusses the choice of $\rho_1(t)$, ensuring that $\widehat{\boldsymbol\Sigma}_t$ is positive definite in finite samples. Our estimation method of the spot volatility matrix can be seen as a natural extension of the kernel-based large sparse covariance matrix estimation CXW13, CL16,CLL19 from the low-frequency data setting to the high-frequency one. We next give some technical assumptions which are
needed to derive the uniform convergence property of $\widehat{\boldsymbol{
\Sigma}}_{t}$.
assumption(i)\ $\{\mu_{i,t}\}$ and $
\{\sigma_{ij,t}\}$ are adapted locally bounded processes with continuous
sample path.
(ii)\ With probability one,
\begin{equation*}
\min_{1\leq i\leq p}\inf_{0\leq s\leq T}\Sigma_{ii, s}>0,\ \ \min_{1\leq
i\neq j\leq p}\inf_{0\leq s\leq T} \Sigma_{ij,s}^\ast>0,
\end{equation*}
where $\Sigma_{ij, s}^\ast=\Sigma_{ii, s}+\Sigma_{jj, s}+2\Sigma_{ij, s}$.
For the spot covariance process $\{\Sigma_{ij,t}\}$, there exist $\gamma\in(0,1)$ and $B(t,\epsilon)$, a positive random function slowly
varying at $\epsilon=0$ and continuous with respect to $t$, such that
\begin{equation}
\max_{1\leq i,j\leq p}\left\vert
\Sigma_{ij,t+\epsilon}-\Sigma_{ij,t}\right\vert\leq
B(t,\epsilon)|\epsilon|^\gamma+o(|\epsilon|^\gamma),\ \ \epsilon\rightarrow0.
\end{equation}
assumption(i)\ The kernel $K(\cdot)$ is a bounded
and Lipschitz continuous function with a compact support $[-1,1]$. In
addition, $\int_{-1}^1 K(u)du=1$.
(ii)\ The bandwidth $h$ satisfies that $h\rightarrow0$ and $\frac{h}{
\Delta\log (p\vee \Delta^{-1})}\rightarrow\infty$.
(iii) Let the time-varying tuning parameter $\rho_1(t)$ in the
generalised shrinkage be chosen as
\begin{equation*}
\rho_1(t)=M(t)\zeta_{\Delta,p},\ \ \zeta_{\Delta,p}=h^{\gamma}+\left[\frac{
\Delta\log (p\vee \Delta^{-1})}{h}\right]^{1/2},
\end{equation*}
where $\gamma$ is defined in ((ref)) and $M(t)$ is a positive function satisfying that
\begin{equation*}
0<C_M\le \inf_{0\leq t\leq T}M(t)\leq\sup_{ 0\leq t\leq
T}M(t)\leq\overline{C}_M<\infty.
\end{equation*}
remarkAssumption (ref) imposes some mild restrictions on
the drift and volatility processes. By a typical localisation procedure as
in Section 4.4.1 of JP12, the local boundedness condition in
Assumption (ref)(i) can be strengthened to the bounded condition over the entire
time interval, i.e., with probability one,
\begin{equation*}
\max_{1\leq i\leq p}\sup_{0\leq s\leq T} |\mu_{i,s}|\leq C_\mu<\infty,\ \
\max_{1\leq i\leq p}\sup_{0\leq s\leq T} \Sigma_{ii,s} \leq C_\Sigma<\infty,
\end{equation*}
which are the same as Assumption A2 in TWZ13 and Assumptions (A.ii)
and (A.iii) in CHLZ20. It may be possible to relax the uniform boundedness restriction (when $T$ is allowed to diverge) at the cost of more lengthy proofs KK16. Assumption (ref)(ii) gives the smoothness condition on the spot covariance process, crucial to derive the uniform
asymptotic order for the kernel estimation bias. When the spot covariance
is driven by continuous semimartingales, ((ref)) holds with $
\gamma<1/2$ RY99. Assumption (ref)(i)
contains some commonly-used conditions for the kernel function. Assumption (ref)(ii)(iii) imposes some mild conditions on the bandwidth and time-varying shrinkage parameter. In particular, when $p$
diverges at a polynomial rate of $1/\Delta$, Assumption (ref)(ii) reduces to the
conventional bandwidth restriction. Assumption (ref)(iii) is comparable to that
assumed by CL16 and CLL19. It is worthwhile to point out that the developed methodology and theory still hold when the time-varying tuning parameter in Assumption (ref)(iii) is allowed to vary over entries in the spot volatility matrix estimation, which is expected to perform well in finite samples. For example, we set $\rho_{ij}(t)=\rho(t)(\widehat{\Sigma}_{ii,t}\widehat{\Sigma}_{jj,t})^{1/2}$ in the numerical studies and shrink the $(i,j)$-entry to zero if $\widehat{\Sigma}_{ij,t}\leq \rho(t)(\widehat{\Sigma}_{ii,t}\widehat{\Sigma}_{jj,t})^{1/2}$.
The following theorem gives the uniform convergence property (in the matrix spectral norm) for the
spot volatility matrix estimator $\widehat{\boldsymbol{\Sigma}}_{t}$ under the uniform sparsity assumption.
theoremSuppose that Assumptions (ref) and (ref) are
satisfied, and $\left\{{\boldsymbol{\Sigma}}_t:\ 0\leq t\leq T\right\}\in
\mathcal{S}(q,\varpi(p), T)$. Then we have
\begin{equation}
\sup_{0\leq t\leq T}\left\Vert\widehat{\boldsymbol{\Sigma}}_{t}-{
\boldsymbol{\Sigma}}_t\right\Vert=O_P\left(\varpi(p)
\zeta_{\Delta,p}^{1-q}\right),
\end{equation}
where $\varpi(p)$ is defined in ((ref)) and $\zeta_{\Delta,p}$ is
defined in Assumption (ref)(iii).
remark(i) The first term of $\zeta_{\Delta,p}$ is $h^{\gamma}$, which is the bias rate due to application of the local
smoothing technique. It is slower than the conventional $h^2$-rate since we do not assume existence of smooth derivatives of $\Sigma_{ij,t}$ (with respect to $t$). The second
term of $\zeta_{\Delta,p}$ is square root of $\Delta h^{-1}\log (p\vee
\Delta^{-1})$, a typical uniform asymptotic rate for the kernel estimation
variance component. The uniform convergence rate in ((ref)) is
also similar to those obtained by CL16 and CLL19 in the
low-frequency data setting (disregarding the bias order). Note that the dimension $p$ affects the uniform convergence rate via $\varpi(p)$ and $\log (p\vee \Delta^{-1})$ and the
estimation consistency may be achieved in the ultra-high dimensional setting
when $p$ diverges at an exponential rate of $n=\lfloor T/\Delta\rfloor$. Treating $(nh)$ as
the “effective" sample size in the local estimation procedure and disregarding the bias rate $h^{\gamma}$, the rate in ((ref)) is comparable to the optimal minimax rate in large covariance
matrix estimation CZ12.
(ii) If we further assume that $\Sigma_{ij,t}$ is deterministic with continuous second-order derivative with respect to $t$, and $K(\cdot)$ is symmetric, we may improve the kernel estimation bias order. In fact, following the proof of Theorem (ref), we may show that
\begin{equation}
\sup_{h\leq t\leq T-h}\left\Vert\widehat{\boldsymbol{\Sigma}}_{t}-{
\boldsymbol{\Sigma}}_t\right\Vert=O_P\left(\varpi(p)
\zeta_{\Delta,p,\star}^{1-q}\right),
\end{equation}
where $\zeta_{\Delta,p,\star}=h^{2}+\left[\frac{\Delta\log (p\vee
\Delta^{-1})}{h}\right]^{1/2}$. The above uniform consistency property only holds over the trimmed time interval $[h, T-h]$ due to the kernel boundary effect. In practice, however, it is often important to investigate the spot volatility structure near the boundary points. For example, when we consider one trading day as a
time interval, it is particularly interesting to estimate the spot
volatility matrix near the opening and closing times which are peak times in
stock market trading. To address this issue, we may replace $K_h^\ast(t_k-t)$ in ((ref)) by a boundary kernel weight defined by
\[
K_{h,t}^\star(t_k-t)=K_t\left(\frac{t_k-t}{h}\right)/\left[\Delta\sum_{l=1}^nK_t\left(\frac{t_l-t}{h}\right)\right],
\]
where $K_t(\cdot)$ is a boundary kernel satisfying $\int_{-t/h}^{(T-t)/h}uK_t(u)du=0$ (a key condition to improve the bias order near the boundary points). Examples of boundary kernels can be found in FG96 and LR07. With this adjustment in the kernel estimation, we can extend the uniform consistency result ((ref)) to the entire interval $[0,T]$.
Estimation with contaminated high-frequency data
\setcounter{equation}{0}
In practice, it is not uncommon that high-frequency financial data are
contaminated by the market microstructure noise. The kernel estimation
method proposed in Section (ref) would be biased if the noise is
ignored in the estimation procedure. Consider the
following additive noise structure:
equation[equation omitted — 162 chars of source]
where $t_k=k\Delta$, $k=1,\cdots,n$, ${\mathbf{Z}}_{t}=\left(Z_{1,t},
\cdots,Z_{p,t}\right)^{^\intercal}$ is a vector of observed asset prices at
time $t$, and ${\boldsymbol{\xi}}_k=(\xi_{1,k},\cdots,\xi_{p,k})^{^\intercal}
$ is a $p$-dimensional vector of noises with nonlinear heteroskedasticity, ${
\boldsymbol{\omega}}(\cdot)=\left[\omega_{ij}(\cdot)\right]_{p\times p}$ is
a $p\times p$ matrix of deterministic functions, and ${\boldsymbol{\xi}}
_k^\ast=\left(\xi_{1,k}^\ast,\cdots,\xi_{p,k}^\ast\right)^{^\intercal}$
independently follows a $p$-variate identical distribution. The noise
structure defined in ((ref)) is similar to the setting considered in
KL08 which also contains a nonlinear mean function and allows the
existence of endogeneity for a single asset. Throughout this section, we
assume that $\{{\boldsymbol{\xi}}_k^\ast\}$ is independent of the Brownian
semimartingale $\{{\mathbf{X}}_t\}$.
Estimation of the spot volatility matrix
To account for the microstructure noise and produce consistent volatility
matrix estimation, we apply the pre-averaging
technique as the realised kernel estimate BHLS08 can be seen as a
member of the pre-averaging estimation class whereas the two-scale estimate
ZMA05 can be re-written as the realised kernel estimate with the
Bartlett-type kernel (up to the first-order approximation). The
pre-averaging method has been studied by JLMPV09, PV09 and
CKP10 in estimating the integrated volatility for a single asset and
is further extended by KWZ16 and DLX19 to the large
high-frequency data setting. KK16 use a
localised pre-averaging technique to estimate the spot volatility function
for a single asset and derive the uniform convergence rate for the developed
estimate. A similar technique is also used by XL02 to improve
convergence of the nonparametric spectral density estimator for time series
with general autocorrelation for low-frequency data.
We first pre-average the observed high-frequency data via a kernel filter,
i.e.,
equation[equation omitted — 123 chars of source]
with $L_{b}^\dagger(t_k-\tau)=L_b\left(t_k-\tau\right)/\int_0^TL_b(s-\tau)ds$, where $L_b(u)=b^{-1}L(u/b)$, $L(\cdot)$ is a kernel function and $b$ is a bandwidth. Let $\Delta
\widetilde{X}_{i,l}=\widetilde{X}_{i,\tau_l}-\widetilde{X}_{i,\tau_{l-1}}$,
where $\widetilde{X}_{i,\tau_l}$ is the $i$-th component of $\widetilde{
\mathbf{X}}_{\tau_l}$ and $\tau_0,\tau_1,\cdots,\tau_N$ are the
pseudo-sampling time points in the fixed interval $[0,T]$ with equal
distance $\Delta_\ast=T/N$. Replacing $\Delta X_{i,k}$ by $\Delta\widetilde{X
}_{i,l}$ in ((ref)), we estimate the spot
co-volatility $\Sigma_{ij,t}$ by
equation[equation omitted — 143 chars of source]
where
\[
K_{h}^\dagger(\tau_l-t)=K_h\left(\tau_l-t\right)/\left[\Delta_\ast\sum_{k=1}^NK_h\left(\tau_k-t\right)\right].
\]
Furthermore, to obtain a stable spot volatility matrix estimate in
finite samples when the dimension $p$ is large, as in ((ref)), we
apply shrinkage to $\widetilde{\Sigma}_{ij,t}$, $1\leq i\neq j\leq p$, and
subsequently construct
equation[equation omitted — 247 chars of source]
where $\rho_2(t)$ is another time-varying shrinkage parameter. We next give
some conditions needed to derive the uniform consistency property of $
\widetilde{\boldsymbol{\Sigma}}_t$.
assumption(i)\ Let $\{{\boldsymbol{\xi}}_k^\ast\}$
be an independent and identically distributed (i.i.d.) sequence of
p-dimensional random vectors. Assume that $\mathsf{E}(\xi_{i,k}^\ast)=0$ and
\begin{equation*}
\mathsf{E}\left[\exp\left(s|{\mathbf{u}}^{^\intercal}{\boldsymbol{\xi}}
_{k}^\ast|\right)\right]\leq C_\xi<\infty,\ \ 0<s\leq s_0,
\end{equation*}
for any $p$-dimensional vector ${\mathbf{u}}$ satisfying $\Vert{\mathbf{u}}
\Vert_2=1$.
(ii)\ The deterministic functions $\omega_{ij}(\cdot)$ are bounded
uniformly over $i,j\in\{1,\cdots,p\}$, and satisfy that
\begin{equation*}
\max_{1\leq i\leq p}\sup_{0\leq t\leq T}\sum_{j=1}^p\omega_{ij}^2(t)\leq
C_\omega<\infty.
\end{equation*}
assumption(i)\ The kernel function $L(\cdot)$ is
Lipschitz continuous and has a compact support $[-1,1]$. In addition, $
\int_{-1}^1 L(u)du=1$.
(ii)\ The bandwidth $b$ and the dimension $p$ satisfy that
\begin{equation*}
b\rightarrow0,\ \ \frac{\Delta^{2\iota-1}b}{\log (p\vee \Delta^{-1})}
\rightarrow\infty,\ \ p\Delta\exp\{-s\Delta^{-\iota}\}\rightarrow0,
\end{equation*}
where $0<\iota<1/2$ and $0<s\leq s_0$.
(iii) Let $\nu_{\Delta,p,N}=\sqrt{N\log(p\vee \Delta^{-1})}\left[b^{1/2}+(\Delta^{-1}b)^{-1/2}\right]\rightarrow0$ and the time-varying
tuning parameter $\rho_2(t)$ be chosen as $\rho_2(t)=M(t)\left(\zeta_{N,p}^
\ast+\nu_{\Delta,p,N}\right)$, where $M(t)$ is defined as in Assumption (ref)(iii) and $\zeta_{N,p}^\ast$ is defined as $\zeta_{\Delta,p}$ with $N$
replacing $\Delta^{-1}$.
remarkWe allow nonlinear heteroskedasticity on the
microstructure noise. The i.i.d. restriction on ${\boldsymbol{\xi}}_i^\ast$
may be weakened to some weak dependence conditions KWZ16,
DLX19 at the cost of more lengthy proofs. The moment condition in
Assumption (ref)(i) is weaker than the sub-Gaussian condition
BL08, TWZ13 which is commonly used in large covariance
matrix estimation when the dimension $p$ is ultra large. The boundedness
condition on $\omega_{ij}(\cdot)$ in Assumption (ref)(ii) is similar to the
local boundedness restriction in Assumption (ref)(i). Assumption (ref)(ii) imposes
some mild restrictions on $b$ and $p$, which imply that there is a
trade-off between them. When $\iota$ is larger, $p$ diverges at a faster
exponential rate of $1/\Delta$ but the bandwidth condition becomes more
restrictive. If $p$ is divergent at a polynomial rate of $1/\Delta$, we may
let $\iota$ be sufficiently close to zero, and then the bandwidth condition
reduces to the conventional one as in Assumption (ref)(ii). The condition $
\nu_{\Delta,p,N}\rightarrow0$ in Assumption (ref)(iii) is crucial to show that
the error of the kernel filter $\widetilde{\boldsymbol{X}}_\tau$ tends to
zero asymptotically, whereas the form of the time-varying shrinkage
parameter $\rho_2(t)$ is relevant to the uniform convergence rate of $
\widetilde{\Sigma}_{ij,t}$ (see Proposition (ref)).
theoremSuppose that Assumptions (ref)(i)(ii), (ref)(i),
(ref) and (ref) are satisfied, and Assumption (ref)(ii) holds with
$\Delta^{-1}$ replaced by $N$. When $\left\{{\boldsymbol{\Sigma}}_t:\ 0\leq
t\leq T\right\}\in \mathcal{S}(q,\varpi(p), T)$, we have
\begin{equation}
\sup_{0\leq t\leq T}\left\Vert\widetilde{\boldsymbol{\Sigma}}_{t}-{
\boldsymbol{\Sigma}}_t\right\Vert=O_P\left(\varpi(p) \left[
\zeta_{N,p}^\ast+\nu_{\Delta,p,N}\right]^{1-q}\right),
\end{equation}
where $\zeta_{N,p}^\ast$ and $\nu_{\Delta,p,N}$ are defined in Assumption (ref)(iii).
remarkThe uniform convergence rate in ((ref))
relies on $\varpi(p)$, $\zeta_{N,p}^\ast$ and $\nu_{\Delta,p,N}$. With the
high-frequency data collected at pseudo time points with sampling frequency $
\Delta_\ast=T/N$, the rate $\zeta_{N,p}^\ast$ is comparable to $
\zeta_{\Delta,p}$ for the noise-free kernel estimator in Section (ref).
The rate $\nu_{\Delta,p,N}$ is due to the error of the kernel filter $
\widetilde{\boldsymbol{X}}_\tau$ in the first step of the local
pre-averaging estimation procedure. In particular, when $q=0$, $\varpi(p)$
is bounded, $b=\Delta^{1/4}$ and $h=N^{-\frac{1}{2\gamma+1}}$ with $
N=\Delta^{-\frac{2\gamma+1}{2\left(4\gamma+1\right)}}$, the uniform
convergence rate in ((ref)) becomes $\Delta^{\frac{\gamma}{
2(4\gamma+1)}}\sqrt{\log(p\vee \Delta^{-1})}$. Furthermore, if $\gamma=1/2$, the rate is simplified to $\Delta^{1/12}\sqrt{\log(p\vee
\Delta^{-1})}$, comparable to those derived by ZB14 and KK16
in the univariate high-frequency data setting.
Estimation of the time-varying noise volatility matrix
It is often interesting to further explore the volatility
structure of microstructure noise. CHLT21 estimate
the constant covariance matrix for high-dimensional noise and derive the
optimal convergence rates for the developed estimate. In the present paper,
we consider the time-varying noise covariance matrix defined by
equation[equation omitted — 177 chars of source]
It is sensible to assume that $\left\{{\boldsymbol{\Omega}}(t):\ 0\leq t\leq
T\right\}$ satisfies the uniform sparsity condition as in ((ref)). For
each $1\leq i,j\leq p$, we estimate $\Omega_{ij}(t)$ by the kernel smoothing
method:
equation[equation omitted — 146 chars of source]
where $h_1$ is a bandwidth, $\Delta Z_{i,t_{k}}=Z_{i,t_k}-Z_{i,t_{k-1}}$ and $K_{h_1}^\ast(t_k-t)$ is defined similarly to $K_{h}^\ast(t_k-t)$ in ((ref)) but with $h_1$ replacing $h$.
As in ((ref)) and ((ref)), we again apply shrinkage to $
\widehat{\Omega}_{ij}(t)$, $1\leq i\neq j\leq p$, and construct
equation[equation omitted — 241 chars of source]
where $\rho_3(t)$ is a time-varying shrinkage parameter. To derive the
uniform consistency property of $\widehat{\boldsymbol{\Omega}}(t)$, we need
to impose stronger moment condition on ${\boldsymbol{\xi}}_k^\ast$ and
smoothness restriction on $\Omega_{ij}(\cdot)$.
assumption(i)\ For any $p$-dimensional vector ${
\mathbf{u}}$ satisfying $\Vert{\mathbf{u}}\Vert_2=1$,
$\mathsf{E}\left[\exp\left(s({\mathbf{u}}^{^\intercal}{\boldsymbol{\xi}}
_{k}^\ast)^2\right)\right]\leq C_\xi^\star<\infty$, $0<s\leq s_0$.
(ii)\ The time-varying function $\Omega_{ij}(t)$ satisfies that
\begin{equation*}
\max_{1\leq i,j\leq p}\left\vert
\Omega_{ij}(t)-\Omega_{ij}(s)\right\vert\leq
C_\Omega|t-s|^{\gamma_1},
\end{equation*}
where $C_\Omega$ is a positive constant and $0<\gamma_1<1$.
(iii)\ The bandwidth $h_1$ and the dimension $p$ satisfy that
\begin{equation*}
h_1\rightarrow0,\ \ \frac{\Delta^{2\iota_\star-1}h_1}{\log (p\vee
\Delta^{-1})}\rightarrow\infty,\ \
p\Delta^{-1}\exp\{-s\Delta^{-\iota_\star}/C_\omega\}\rightarrow0,
\end{equation*}
where $0<\iota_\star<1/2$, $0<s\leq s_0$ and $C_\omega$ is defined in
Assumption (ref)(ii).
remarkAssumption (ref)(i) strengthens the moment condition in Assumption (ref)(i) and is equivalent to the sub-Gaussian condition, see Assumption A1 in TWZ13. The smoothness condition in
Assumption (ref)(ii) is similar to ((ref)), crucial to derive the
asymptotic order of the kernel estimation bias. The restrictions on $h_1$
and $p$ in Assumption (ref)(iii) are similar to those in Assumption (ref)(ii),
allowing $p$ to be divergent to infinity at an exponential
rate of $1/\Delta$.
In the following theorem, we state the uniform consistency result for $
\widehat{\boldsymbol{\Omega}}(t)$ with convergence rate comparable to that
in Theorem (ref).
theoremSuppose that Assumptions (ref), (ref)(i), (ref)
and (ref) are satisfied, and Assumption (ref)(ii)(iii) holds when $
\rho_1(t)$, $\zeta_{\Delta,p}$ and $h$ are replaced by $\rho_3(t)$, $
\delta_{\Delta,p}$ and $h_1$, respectively, where $\delta_{\Delta,p}=h_1^{
\gamma_1}+\left[\frac{\Delta\log (p\vee \Delta^{-1})}{h_1}\right]^{1/2}$. If
$\left\{{\boldsymbol{\Omega}}(t):\ 0\leq t\leq T\right\}\in \mathcal{S}
(q,\varpi(p), T)$, we have
\begin{equation}
\sup_{0\leq t\leq T}\left\Vert\widehat{\boldsymbol{\Omega}}(t)-{
\boldsymbol{\Omega}}(t)\right\Vert=O_P\left(\varpi(p)\delta_{\Delta,p}^{1-q}
\right).
\end{equation}
remarkIf the bandwidth parameter $h_1$ in ((ref)) is the same as $h$ in ((ref)), we may find that the uniform convergence rate $O_P\left(\varpi(p)\delta_{\Delta,p}^{1-q}\right)$ would be
the same as that in Theorem 1. Treating $(nh_1)$ as the “effective" sample
size and disregarding the bias order, we may show that the uniform
convergence rate in ((ref)) is comparable to the optimal minimax rate
derived by CHLT21 for the constant noise covariance matrix
estimation. Meanwhile, the kernel estimation bias order $h_1^{\gamma_1}$ may be improved by strengthening the smoothness condition on $\Omega_{ij}(\cdot)$ and adopting the boundary kernel weight as suggested in Remark (ref)(ii).
Estimation with observed factors
\setcounter{equation}{0}
The large spot volatility matrix estimation with the shrinkage technique developed in Sections (ref) and (ref) heavily relies on the uniform sparsity assumption ((ref)). However, the latter may be too restrictive in practice since the price processes of a large number of assets are often driven by some common factors such as the market factors, resulting in strong correlation among assets and failure of the sparsity condition. To address this problem, we next consider the nonparametric time-varying regression at high frequency:
equation[equation omitted — 97 chars of source]
where ${\boldsymbol\beta}(t)=\left[\beta_{1}(t),\cdots,\beta_p(t)\right]^{^\intercal}$ is a $p\times k$ matrix of time-varying betas (or factor loadings), ${\mathbf F}_t$ and ${\mathbf X}_t$ are $k$-variate and $p$-variate continuous semi-martingales defined by
equation[equation omitted — 216 chars of source]
respectively, ${\boldsymbol{\mu}}_t^F$ and ${\boldsymbol{\mu}}_t^X$ are drift
vectors, ${\boldsymbol{\sigma}}_t^F=\left(\sigma_{ij,t}^F\right)_{k\times k}$, ${\boldsymbol{\sigma}}_t^X=\left(\sigma_{ij,t}^X\right)_{p\times p}$, ${\mathbf{W}}_t^F$ and ${\mathbf{W}}_t^X$ are
$k$-dimensional and $p$-dimensional standard Brownian motions. For the time being, we assume that ${\mathbf Y}_t$ and ${\mathbf F}_t$ are observable and noise free but ${\mathbf X}_t$ is latent. Extension of the methodology and theory to the noise-contaminated high-frequency data will be considered later in this section.
Estimation of the constant betas via the ratio of realised covariance to realised variance is proposed by BS04, and extension to time-varying beta estimation has been studied by MZ06, RTT15 and AKX20, some of which allow jumps in the semi-martingale processes. The main interest of this section lies in estimating the large spot volatility structure ${\boldsymbol\Sigma}_t^Y$ of ${\mathbf Y}_t$. Letting ${\boldsymbol\Sigma}_t^F={\boldsymbol\sigma}_t^F\left({\boldsymbol\sigma}_t^F\right)^{^\intercal}$ and ${\boldsymbol\Sigma}_t^X={\boldsymbol\sigma}_t^X\left({\boldsymbol\sigma}_t^X\right)^{^\intercal}$, and assuming orthogonality between ${\mathbf X}_t$ and ${\mathbf F}_t$, see Assumption (ref)(iii) below, it follows from ((ref)) that
equation[equation omitted — 156 chars of source]
As in FLM11, FLM13, we impose the uniform sparsity restriction on ${\boldsymbol\Sigma}_t^X$ instead of ${\boldsymbol\Sigma}_t^Y$, i.e., $
\left\{{\boldsymbol{\Sigma}}_t^X:\ 0\leq t\leq T\right\}\in \mathcal{S}(q,\varpi(p), T)$. This is a reasonable assumption in practical applications as the asset prices, after removing the influence of systematic factors, are expected to be weakly correlated. FFX16 and DLX19 use a similar framework with constant betas to estimate large integrated volatility matrices.
Suppose that we observe ${\mathbf Y}_t$ and ${\mathbf F}_t$ at regular points: $t_k=k\Delta$, $k=1,\cdots,n$, as in Sections (ref) and (ref). Let ${\boldsymbol\Sigma}_t^{YF}$ be the spot covariance between ${\mathbf Y}_t$ and ${\mathbf F}_t$. We may use the kernel smoothing method as in ((ref)) to estimate ${\boldsymbol\Sigma}_t^{Y}$, ${\boldsymbol\Sigma}_t^{F}$ and ${\boldsymbol\Sigma}_t^{YF}$, i.e.,
eqnarray[eqnarray omitted — 409 chars of source]
where $\Delta{\mathbf Y}_k={\mathbf Y}_{t_k}-{\mathbf Y}_{t_{k-1}}$, $\Delta{\mathbf F}_k={\mathbf F}_{t_k}-{\mathbf F}_{t_{k-1}}$, and $K_h^\ast(t_k-t)$ is defined as in ((ref)). Consequently, the time-varying betas ${\boldsymbol\beta}(t)$ and the spot idiosyncratic volatility matrix ${\boldsymbol\Sigma}_t^X$ are estimated by
equation[equation omitted — 218 chars of source]
and
equation[equation omitted — 288 chars of source]
With the uniform sparsity condition, it is sensible to further apply shrinkage to $\widehat\Sigma_{ij,t}^X$, i.e.,
equation[equation omitted — 253 chars of source]
where $\rho_4(t)$ is a time-varying shrinkage parameter. We finally estimate ${\boldsymbol\Sigma}_t^Y$ as
equation[equation omitted — 388 chars of source]
We need the following assumption to derive the uniform convergence property for $\widehat{\boldsymbol\Sigma}_t^{X,s}$ and $\widehat{\boldsymbol\Sigma}_t^{Y,s}$.
assumption{\em (i) Assumption (ref) is satisfied for $\{X_t\}$ defined in ((ref)) (with minor notational changes).}
{\em (ii) Let $\{{\boldsymbol{\mu}}_t^F\}$, $\{{\boldsymbol{\sigma}}_t^F\}$ and $\{{\boldsymbol{\Sigma}}_t^F\}$ satisfy the boundedness and smoothing conditions as in Assumption (ref).}
{\em (iii) For any $1\leq i\leq p$ and $1\leq j\leq k$, $\left[X_{it}, F_{jt}\right]=0$ for any $t\in[0,T]$, where $X_{i,t}$ is the $i$-th element of ${\mathbf X}_t$, $F_{j,t}$ is the $j$-th element of ${\mathbf F}_t$, and $[\cdot,\cdot]$ denotes the quadratic covariation. }
{\em (iv) The time-varying beta function $\beta_i(\cdot)$ satisfies that
\[
\max_{1\leq i\leq p}\sup_{0\leq t\leq T}\left\Vert \beta_i(t)\right\Vert_2\leq C_\beta<\infty,\ \ \max_{1\leq i\leq p}\left\Vert \beta_i(t)-\beta_i(s)\right\Vert_2\leq C_\beta|t-s|^{\gamma},
\]
where $\gamma$ is the same as that in Assumption (ref)(ii). In addition, there exists a positive definite matrix ${\boldsymbol\Sigma}_\beta(t)$ (with uniformly bounded eigenvalues) such that}
\begin{equation}
\sup_{0\leq t\leq T}\left\Vert \frac{1}{p}{\boldsymbol\beta}(t)^{^\intercal}{\boldsymbol\beta}(t)-{\boldsymbol\Sigma}_\beta(t)\right\Vert=o(1).
\end{equation}
remarkThe uniform boundedness and smoothness conditions imposed on the drift and spot volatility functions of ${\mathbf X}_t$ and ${\mathbf F}_t$ in Assumption (ref)(i)(ii) are the same as those in Assumption (ref). This is crucial to ensure that the uniform convergence rates of $\widehat{\boldsymbol\Sigma}_t^Y$, $\widehat{\boldsymbol\Sigma}_t^F$ and $\widehat{\boldsymbol\Sigma}_t^{YF}$ (in the max norm) derived in Proposition (ref) are the same as that in Proposition (ref). The orthogonality condition in Assumption (ref)(iii) is commonly used to consistently estimate the time-varying factor model FFX16, DLX19. Assumption (ref)(iv) is a rather mild restriction on time-varying betas and may be strengthened to improve the estimation bias order, see the discussion in Remark (ref)(ii). The condition ((ref)) indicates that all the factors are pervasive.
We next present the convergence property of $\widehat{\boldsymbol\Sigma}_t^{X,s}$ and $\widehat{\boldsymbol\Sigma}_t^{Y,s}$ defined in ((ref)) and ((ref)), respectively. Due to the nonparametric factor regression model structure ((ref)), the largest $k$ eigenvalues of ${\boldsymbol\Sigma}_t^Y$ are spiked, diverging at a rate of $p$. Hence, ${\boldsymbol\Sigma}_t^Y$ cannot be consistently estimated in the absolute term. To address this problem, as in FLM11, FLM13, we measure the spiked volatility matrix estimate in the following relative error:
\[\left\Vert \widehat{\boldsymbol\Sigma}_t^{Y,s}-{\boldsymbol\Sigma}_t^Y\right\Vert_{{\boldsymbol\Sigma}_t^Y}=\frac{1}{\sqrt{p}}\left\Vert \left({\boldsymbol\Sigma}_t^Y\right)^{-1/2}\left(\widehat{\boldsymbol\Sigma}_t^{Y,s}-{\boldsymbol\Sigma}_t^Y\right) \left({\boldsymbol\Sigma}_t^Y\right)^{-1/2} \right\Vert_F,\]
where the normalisation factor $p^{-1/2}$ is used to guarantee that $\left\Vert{\boldsymbol\Sigma}_t^Y\right\Vert_{{\boldsymbol\Sigma}_t^Y}=1$.
theoremSuppose that Assumptions (ref)(i)(ii) and (ref) are satisfied, and Assumption (ref)(iii) holds with $\rho_1(t)$ replaced by $\rho_4(t)$. When $\left\{{\boldsymbol{\Sigma}}_t^X:\ 0\leq t\leq T\right\}\in \mathcal{S}(q,\varpi(p), T)$, we have
\begin{equation}
\sup_{0\leq t\leq T}\left\Vert\widehat{\boldsymbol{\Sigma}}_{t}^{X,s}-{\boldsymbol{\Sigma}}_t^X\right\Vert=O_P\left(\varpi(p)\zeta_{\Delta,p}^{1-q}\right),
\end{equation}
where $\varpi(p)$ is defined in ((ref)) and $\zeta_{\Delta,p}$ is defined in Assumption (ref)(iii); and
\begin{equation}
\sup_{0\leq t\leq T}\left\Vert \widehat{\boldsymbol\Sigma}_t^{Y,s}-{\boldsymbol\Sigma}_t^Y\right\Vert_{{\boldsymbol\Sigma}_t^Y}=O_P\left(p^{1/2}\zeta_{\Delta,p}^2+\varpi(p)\zeta_{\Delta,p}^{1-q}\right).
\end{equation}
remarkAlthough ${\mathbf X}_t$ is latent in model ((ref)), the uniform convergence rate for $\widehat{\boldsymbol{\Sigma}}_{t}^{X,s}$ in ((ref)) is the same as that in Theorem (ref) when ${\mathbf X}_t$ is observable. Treating $(nh)$ as the effective sample size in kernel estimation and disregarding the bias order in $\zeta_{\Delta,p}$, the uniform convergence rate for $\widehat{\boldsymbol{\Sigma}}_{t}^{Y,s}$ in ((ref)) is comparable to the convergence rates derived by FLM11 in low frequency and FFX16 in high frequency. To guarantee uniform consistency in the relative matrix estimation error, we have to further assume that $p\zeta_{\Delta,p}^4=o(1)$, limiting the divergence rate of the asset number, i.e., $p$ can only diverge at a polynomial rate of $n=\lfloor T/\Delta\rfloor$.
We next modify the above methodology and theory to accommodate microstructure noise in the asset prices and factors. Assume that
equation[equation omitted — 234 chars of source]
where ${\boldsymbol{\omega}}_Y(\cdot)$ and ${\boldsymbol{\omega}}_F(\cdot)$ are matrices of deterministic functions similar to ${\boldsymbol{\omega}}(\cdot)$, and $\{{\boldsymbol{\xi}}_{Y,k}^\ast\}$ and $\{{\boldsymbol{\xi}}_{F,k}^\ast\}$ are i.i.d. sequences of random vectors similar to $\{{\boldsymbol{\xi}}_{k}^\ast\}$. Since both ${\mathbf Y}_t$ and ${\mathbf F}_t$ are latent, we need to first adopt the kernel pre-averaging technique proposed in Section (ref) to obtain the approximation of ${\mathbf Y}_t$ and ${\mathbf F}_t$, and then apply the kernel smoothing and generalised shrinkage as in ((ref))--((ref)). This results in a three-stage estimation procedure which we describe as follows.
enumerate• As in ((ref)), we pre-average the noise-contaminated ${\mathbf{Z}}_{Y,t_k}$ and ${\mathbf{Z}}_{F,t_k}$ via the kernel filter:
\begin{equation}
\widetilde{\mathbf{Y}}_\tau=\frac{T}{n}\sum_{k=1}^n L_b^\dagger(t_k-\tau){\mathbf{Z}}_{Y,t_k},\quad
\widetilde{\mathbf{F}}_\tau=\frac{T}{n}\sum_{k=1}^n L_b^\dagger(t_k-\tau){\mathbf{Z}}_{F,t_k},
\end{equation}
where $L_b^\dagger(t_k-\tau)$ is defined as in ((ref)) and we consider $\tau$ as the pseudo-sampling time points: $\tau_l=l\Delta_\ast$, $l=0,1,\cdots, N=\lfloor T/\Delta_\ast\rfloor$.
• With $\widetilde{\mathbf Y}_{\tau_l}$ and $\widetilde{\mathbf F}_{\tau_l}$, $l=1,\cdots,N$, we estimate ${\boldsymbol\Sigma}_t^Y, {\boldsymbol\Sigma}_t^F$ and ${\boldsymbol\Sigma}_t^{YF}$ by the kernel smoothing as in ((ref))--((ref)):
\begin{eqnarray}
&&\widetilde{\boldsymbol\Sigma}_t^Y=\sum_{l=1}^N K_h^\dagger(\tau_l-t)\Delta \widetilde{\mathbf Y}_l\Delta \widetilde{\mathbf Y}_l^{^\intercal},\notag\\
&&\widetilde{\boldsymbol\Sigma}_t^F=\sum_{l=1}^N K_h^\dagger(\tau_l-t)\Delta \widetilde{\mathbf F}_l\Delta \widetilde{\mathbf F}_l^{^\intercal},\notag\\
&&\widetilde{\boldsymbol\Sigma}_t^{YF}=\sum_{l=1}^N K_h^\dagger(\tau_l-t)\Delta \widetilde{\mathbf Y}_l\Delta \widetilde{\mathbf F}_l^{^\intercal},\notag
\end{eqnarray}
where $K_h^\dagger(\tau_l-t)$ is defined as in ((ref)), $\Delta \widetilde{\mathbf Y}_l=\widetilde{\mathbf Y}_{\tau_l}-\widetilde{\mathbf Y}_{\tau_{l-1}}$ and $\Delta \widetilde{\mathbf F}_l=\widetilde{\mathbf F}_{\tau_l}-\widetilde{\mathbf F}_{\tau_{l-1}}$. Furthermore, estimate ${\boldsymbol\beta}(t)$ and ${\boldsymbol\Sigma}_t^X$ by
\[
\widetilde{\boldsymbol\beta}(t)=\widetilde{\boldsymbol\Sigma}_t^{YF}\left(\widetilde{\boldsymbol\Sigma}_t^{F}\right)^{-1},\quad \widetilde{\boldsymbol\Sigma}_t^X=\left(\widetilde\Sigma_{ij,t}^X\right)_{p\times p}=\widetilde{\boldsymbol\Sigma}_t^Y-\widetilde{\boldsymbol\Sigma}_t^{YF}\left(\widetilde{\boldsymbol\Sigma}_t^{F}\right)^{-1}\left(\widetilde{\boldsymbol\Sigma}_t^{YF}\right)^{^\intercal}.
\]
• Apply the generalised shrinkage to $\widetilde\Sigma_{ij,t}^X$, i.e.,
\[
\widetilde{\boldsymbol{\Sigma}}_{t}^{X,s}=\left(\widetilde\Sigma_{ij,t}^{X,s}\right)_{p
\times p}\ \ \mathrm{with}\ \
\widetilde\Sigma_{ij,t}^{X,s}=s_{\rho_5(t)}(\widetilde\Sigma_{ij,t}^X)I(i\neq
j)+\widetilde\Sigma_{ii,t}^X I(i=j),
\]
where $\rho_5(t)$ is the shrinkage parameter, and then estimate ${\boldsymbol\Sigma}_t^Y$ by
\[
\widetilde{\boldsymbol\Sigma}_t^{Y, s}=\widetilde{\boldsymbol\beta}(t)\widetilde{\boldsymbol\Sigma}_t^F\widetilde{\boldsymbol\beta}(t)^{^\intercal}+\widetilde{\boldsymbol\Sigma}_t^{X,s}=\widetilde{\boldsymbol\Sigma}_t^{YF}\left(\widetilde{\boldsymbol\Sigma}_t^{F}\right)^{-1}\left(\widetilde{\boldsymbol\Sigma}_t^{YF}\right)^{^\intercal}+\widetilde{\boldsymbol\Sigma}_t^{X,s}.
\]
As shown in Theorem (ref), the existence of microstructure noises slows down the uniform convergence rates. Following the proof of Lemma B.1 in Appendix B, we may show that
\[
\max_{0\leq l\leq N}\left\vert \widetilde{\mathbf{Y}}_{\tau_l}-{\mathbf{Y}}_{\tau_l}\right\vert_{\max}+\max_{0\leq l\leq N}\left\vert \widetilde{\mathbf{F}}_{\tau_l}-{\mathbf{F}}_{\tau_l}\right\vert_{\max}=O_P\left(\nu_{\Delta,p,N}\right),
\]
where $\vert\cdot\vert_{\max}$ denotes the $L_\infty$-norm of a vector, and $\nu_{\Delta,p,N}$ is defined in Assumption (ref)(iii). Modifying Proposition (ref) and the proof of Theorem (ref) in Appendix A, we can prove that ((ref)) and ((ref)) hold but with $\zeta_{\Delta,p}$ replaced by $\zeta_{N,p}^\ast+\nu_{\Delta,p,N}$ defined in Assumption (ref)(iii), i.e.,
eqnarray[eqnarray omitted — 453 chars of source]
Monte-Carlo Study
\setcounter{equation}{0}
In this section, we report the Monte-Carlo simulation studies to assess the numerical performance of the proposed large spot volatility matrix and time-varying noise volatility matrix estimation methods under the sparsity condition and the factor-based spot volatility matrix estimation. Here we only consider the synchronous high-frequency data. Additional simulation results for asynchronous high-frequency data are provided in the supplement.
Simulation for sparse volatility matrix estimation
{\em 5.1.1.\ \ Simulation setup}
We generate the noise-contaminated high-frequency data according to model ((ref)), where ${\boldsymbol{\omega}}(t)$ is taken as the Cholesky decomposition of the noise covariance matrix ${\boldsymbol\Omega}(t)=\left[\Omega _{ij}(t)\right] _{p\times p}$, ${\boldsymbol{\xi }}_{k}^{\ast }=\left(\xi _{1,k}^{\ast },\cdots ,\xi _{p,k}^{\ast }\right)^{^{\intercal }}$ is an independent $p$-dimensional random vector of cross-sectionally independent standard normal random variables, the latent return process $\mathbf{X}_{t}$ of $p$ assets is generated from the following drift-free model:
equation[equation omitted — 111 chars of source]
$\mathbf{W}_{t}^{X}=\left( W_{1,t}^{X},\cdots ,W_{p,t}^{X}\right)^{^{\intercal }}$ is a standard $p$-dimensional Brownian motion, and ${\boldsymbol{\sigma }}_{t}$ is chosen as the Cholesky decomposition of the spot covariance matrix ${\boldsymbol{\Sigma }}_{t}=\left(\Sigma_{ij,t}\right) _{p\times p}$. In the simulation, we consider the volatility matrix estimation over the time interval of a full trading day, and set the sampling interval to be $15$ seconds, i.e., $\Delta =1/(252\times 6.5\times 60\times 4)$, to generate synchronous data. We consider three structures in ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$: “banding", “block-diagonal", and “exponentially decaying". Following WZ10, we generate the diagonal elements of ${\boldsymbol\Sigma}_{t}$ from the following geometric Ornstein-Uhlenbeck model BS02:
equation*[equation* omitted — 178 chars of source]
where $\mathbf{W}_{t}^\ast=\left(W_{1,t}^\ast,\cdots,W_{p,t}^\ast\right)^{^\intercal}$ is a standard $p$-dimensional Brownian motion independent of $\mathbf{W}_{t}^{X}$, and $\iota_{i}$ is a random number generated uniformly between $-0.62$ and $-0.30$, reflecting the leverage effects. The diagonal elements of ${\boldsymbol\Omega}(t)$ are defined as daily cyclical deterministic functions of time:
\[
\Omega _{ii}\left( t\right) =c_{i}\left\{ \frac{1}{2}\left[ \cos \left(2\pi t/T\right) +1\right] \times \left( \overline\omega -\underline\omega\right) +\underline\omega\right\},
\]
where $\overline\omega=1$ and $\underline\omega=0.1$ reflect the observation by KL08 that the noise level is high at both the opening and the closing times of a trading day and is low in the middle of the day, and the scalar $c_{i}$ controls the noise ratio for each asset which is chosen to match the highest noise ratio considered by WZ10. As in BS02, BS04, we define a continuous-time stochastic process $\kappa^{\Sigma}_t$ by
eqnarray[eqnarray omitted — 260 chars of source]
where $W_{t}^\diamond$ is a standard univariate Brownian motion independent of $\mathbf{W}_{t}^{X}$ and $\mathbf{W}_{t}^\ast$. Let
\[
\kappa^\Omega_t =\frac{\overline\kappa-\underline\kappa}{2}\left[ \cos \left( 2\pi t/T\right) +1\right] +\underline\kappa,
\]
where $\overline\kappa=0.5$ and $\underline\kappa=-0.5$. We will use $\kappa^{\Sigma}_t$ and $\kappa^{\Omega}_t$ to define the off-diagonal elements in ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$, respectively, which are specified as follows.
itemize• Banding structure for ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$: The off-diagonal elements are defined by
\[
\Sigma _{ij,t} =\left(\kappa^\Sigma_{t}\right)^{\vert i-j\vert }\sqrt{\Sigma_{ii,t}\Sigma_{jj,t}}\cdot I\left( \left\vert i-j\right\vert \leq 2\right),\]
and
\[\Omega_{ij}(t) =\left(\kappa^\Omega_t\right) ^{\vert i-j\vert }\sqrt{\Omega _{ii}( t) \Omega _{jj}(t) }\cdot I\left( \left\vert i-j\right\vert \leq 2\right),
\]
for $1\leq i\neq j\leq p$.
• Block-diagonal structure for ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$: The off-diagonal elements are defined by
\[
\Sigma _{ij,t} =\left(\kappa^\Sigma_{t}\right)^{\vert i-j\vert }\sqrt{\Sigma_{ii,t}\Sigma_{jj,t}}\cdot I\left( (i,j)\in {\cal B}\right),
\]
\[
\Omega _{ij}(t)=\left(\kappa^\Omega_t\right) ^{\vert i-j\vert }\sqrt{\Omega _{ii}( t) \Omega _{jj}(t) }\cdot I\left((i,j)\in {\cal B}\right) ,
\]
for $1\leq i\neq j\leq p$, where ${\cal B}$ is a collection of row and column indices $(i,j)$ located within our randomly generated diagonal blocks \footnote{As in DLX19, to generate blocks with random sizes, we fix the largest block size at $20$ when $p=200$ and randomly generate the sizes of the remaining blocks from a random integer uniformly picked between $5$ and $20$. When $p=500$, the largest size is $40$, and the random integer is uniformly picked between $10$ and $40$. Block sizes are randomly generated but fixed across all Monte Carlo repetitions.}.
• Exponentially decaying structure for ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$: The off-diagonal elements are defined by
\begin{equation}
\Sigma _{ij,t} =\left(\kappa^\Sigma_{t}\right)^{\vert i-j\vert }\sqrt{\Sigma_{ii,t}\Sigma_{jj,t}},\ \ \Omega_{ij}(t) =\left(\kappa^\Omega_t\right) ^{\vert i-j\vert }\sqrt{\Omega _{ii}( t) \Omega _{jj}(t) },\ \ 1\leq i\neq j\leq p.
\end{equation}
It is clear that the sparsity condition is not satisfied when the off-diagonal elements of ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$ are exponentially decaying as in ((ref)). The number of assets $p$ is set as $p=200$ and $500$ and the replication number is $R=200$
{\em 5.1.2.\ \ Volatility matrix estimation}
In the simulation studies, we consider the following volatility matrix estimates.
itemize• Noise-free spot volatility matrix estimate $\widehat{\boldsymbol\Sigma }_{t}$. This infeasible estimate serves as a benchmark in comparing the numerical performance of various estimation methods. As in Section (ref), we apply the kernel smoothing method to estimate $\Sigma_{ij,t}$ by directly using the latent return process ${\mathbf X}_t$, where the bandwidth is determined by the leave-one-out cross validation. We apply four shrinkage methods to $\widehat{\Sigma}_{ij,t}$ for $i\neq j$: hard thresholding (Hard), soft thresholding (Soft), adaptive LASSO (AL) and smoothly clipped absolute deviation (SCAD). For comparison, we also compute the naive estimate without applying any regularisation technique.
• Noise-contaminated spot volatility matrix estimate $\widetilde{\boldsymbol\Sigma}_{t}$. We combine the kernel smoothing with pre-averaging in Section (ref) to estimate $\Sigma_{ij,t}$ by using the noise-contaminated process ${\mathbf Z}_t$. As in the noise-free estimation, we apply four shrinkage methods to $\widetilde{\Sigma}_{ij,t}$ for $i\neq j$ and also compute the naive estimate without applying the shrinkage.
• Time-varying noise volatility matrix estimate $\widehat{\boldsymbol\Omega}(t)$. We combine the kernel smoothing with four shrinkage techniques in the estimation as in Section (ref) and also the naive estimate without shrinkage.
The choice of tuning parameter in shrinkage is similar to that in DLX19. For example, in the noise-free spot volatility estimate, we set the tuning parameter as $\rho_{ij}(t)=\rho(t)(\widehat{\Sigma}_{ii,t}\widehat{\Sigma}_{jj,t})^{1/2}$ where $\rho(t)$ is chosen as the minimum value among the grid of values on $[0,1]$ such that the shrinkage estimate of the spot volatility matrix is positive definite. To evaluate the estimation performance of $\widehat{\boldsymbol\Sigma }_{t}$, we consider $21$ equidistant time points on $[0,T]$ and compute the following Mean Frobenius Loss (MFL) and Mean Spectral Loss (MSL) over $200$ repetitions:
eqnarray*[eqnarray* omitted — 393 chars of source]
where $t_{j}$, $j=1,2,\cdots,21$ are the equidistant time points on the interval $[0,T]$, and $\boldsymbol{\widehat{\Sigma}}_{t_{j}}^{(m)}$ and $\boldsymbol{\Sigma }_{t_{j}}^{(m)}$ are respectively the estimated and true spot volatility matrices at $t_{j}$ for the $m$-th repetition. The “MFL" and “MSL" can be similarly defined for $\widetilde{\boldsymbol\Sigma}_{t}$ and $\widehat{\boldsymbol\Omega}(t)$.
{\em 5.1.3.\ \ Simulation results}
Table 1 reports the simulation results when the dimension is $p=200$. The three panels in the table (from top to bottom) report the results where the true volatility matrix structures are banding, block-diagonal, and exponentially decaying, respectively. In each panel, the MFL results are reported on the left, whereas the MSL results are on the right. The first two rows of each panel contain the MFL and MSL results for the spot volatility matrix estimation whereas the third row contains the results for the time-varying noise volatility matrix estimation.
For the noise-free estimate $\widehat{\boldsymbol\Sigma }_{t}$, when the volatility matrix structure is banding, the performance of the four shrinkage estimators are substantially better than that of the naive estimate (without any shrinkage). In particular, the results of the soft thresholding, adaptive LASSO and SCAD are very similar and their MFL and MSL values are approximately one third of those of the naive estimator. Meanwhile, the performance of the hard thresholding is less accurate (despite the much stronger level of shrinking used), but is still much better than the naive estimate. These results show that the shrinkage technique is an effective tool in estimating the sparse volatility matrix. Similar results are obtained for the noise-contaminated estimate $\widetilde{\boldsymbol\Sigma}_{t}$. Unsurprisingly, due to the microstructure noise, the MFL and MSL values of the local pre-averaging estimates are noticeably higher than the corresponding values of the noise-free estimates. We next turn the attention to the time-varying noise volatility matrix estimate $\widehat{\boldsymbol\Omega}(t)$. As in the spot volatility matrix estimation, the naive method again produces the highest MFL and MSL values. The performance of the four shrinkage estimators are similar with the adaptive LASSO and SCAD being slightly better than the hard and soft thresholding. The simulation results for the block-diagonal and exponentially decaying covariance matrix settings, reported in the middle and bottom panels of Table 1, are fairly close to those for the banding setting. Overall, the results in Table 1 show that the shrinkage methods perform well not only in the sparse covariance matrix settings but also in the non-sparse one (i.e., the exponentially decaying setting).
center[center omitted — 4,165 chars of source]
{\scriptsize The selected bandwidths are $h^\ast=90$ for $\widehat{\boldsymbol\Sigma}_{t}$, $h^\ast=90$ and $b^\ast=4$ for $\widetilde{\boldsymbol\Sigma}_{t}$, and $h_1^{\ast}=90$ for $\widehat{\boldsymbol\Omega}(t)$, where $h^{\ast}=h/\Delta$, $b^{\ast}=b/\Delta$, and $h_{1}^{\ast}=h_1/\Delta$.}
The simulation results when the dimension is $p=500$ are reported in Table 2. Overall the results are very similar to those in Table 1, so we omit the detailed discussion and comparison to save the space.
center[center omitted — 3,953 chars of source]
{\scriptsize The selected bandwidths are $h^{\ast }=240$ for $\widehat{\boldsymbol\Sigma}_{t}$, $h^{\ast }=240$, $b^{\ast }=4$ for $\widetilde{\boldsymbol\Sigma}_{t}$ and $h_1^{\ast }=240$ for $\widehat{\boldsymbol\Omega}(t)$, where $h^{\ast}=h/\Delta$, $b^{\ast}=b/\Delta$, and $h_{1}^{\ast}=h_1/\Delta$.}
Simulation for factor-based spot volatility matrix estimation
{\em 5.2.1.\ \ Simulation setup}
We generate ${\mathbf Y}_t$ via ((ref)), where the $p$-dimensional idiosyncratic returns follow the dynamics of $d\mathbf{X}_{t}$ defined in ((ref)). In this simulation, we only consider $p=500$. As in AKX20, we adopt a three-factor model, where the factors $\mathbf{F}_{t}=\left(F_{1,t},F_{2,t},F_{3,t}\right) ^{^\intercal}$ are generated by
equation*[equation* omitted — 524 chars of source]
The factor volatilities are driven by
equation*[equation* omitted — 183 chars of source]
where ${\sf E}[ dW_{k,t}^{F}d\widetilde{W}_{k,t}] =\rho_{k}dt$, allowing for potential leverage effects in the factor dynamics. Both $W_{k,t}^{F}$ and $\widetilde{W}_{k,t}$ are standard univariate Brownian motions. In the simulation, we set $\left( \tilde{\kappa}_{1},\tilde{\kappa}_{2},\tilde{
\kappa}_{3}\right) =\left( 3,4,5\right) $, $\left( \tilde{\alpha}_{1},\tilde{
\alpha}_{2},\tilde{\alpha}_{3}\right) =\left( 0.09,0.04,0.06\right) $, $
\left( \tilde{\nu}_{1},\tilde{\nu}_{2},\tilde{\nu}_{3}\right) =\left(
0.3,0.4,0.3\right) $, $\left( \mu _{1}^{F},\mu _{2}^{F},\mu _{3}^{F}\right)
=\left( 0.05,0.03,0.02\right) $, $\left( \rho _{1},\rho _{2},\rho
_{3}\right) =\left( -0.6,-0.4,-0.25\right) $ and $\left( \rho _{12},\rho
_{13},\rho _{23}\right) =\left( 0.05,0.10,0.15\right) $.
We consider the following three cases for generating the time-varying beta processes: $\beta_{i}(t)=\left[\beta_{i,1}(t), \beta_{i,2}(t), \beta_{i,3}(t)\right]^{^\intercal}$, $i=1,\cdots,p$.
itemize• Constant betas. The factor loadings are constants over time, i.e., $\beta _{i,l}(t)=\beta _{i,l}$, $i=1,\cdots,p$ and $l=1,2,3$. For each $i$, we set $\beta _{i,1}\sim {\sf U}\left(
0.25,2.25\right)$ and $\beta _{i,2},\beta _{i,3}\sim {\sf U}\left(-0.5,0.5\right)$.
• Deterministic time-varying betas. Consider the following deterministic function:
\begin{equation*}
\beta _{i,l}\left( t\right) =\frac{1}{2}\left[ \cos \left( \pi (
t-\omega _{i,l})/T \right) +1\right] \times \left( \overline{\beta} _{i,l}-\beta_{i,l}\right) +\beta _{i,l},\ \ i=1,\cdots,p,\ \ l=1,2,3,
\end{equation*}
where $\omega _{i,1},\omega_{i,2},\omega _{i,3}\sim {\sf U}(0, 2T)$, $(\underline{\beta} _{i,1},\overline{\beta}_{i,1})$ is a pair of two random numbers from ${\sf U}\left( 0.25,2.25\right)$ whereas
$(\underline{\beta}_{i,2},\overline{\beta}_{i,2})$ and $(\underline{\beta}_{i,3},\overline{\beta}_{i,3})$ are pairs of random numbers from ${\sf U}\left( -0.5,0.5\right)$.
• Stochastic time-varying betas. As in AKX20, we consider the following diffusion process:
\[
d\beta _{i,l}(t)=\kappa _{i,l}^{\beta }\left( \alpha _{i,l}^{\beta }-\beta
_{i,l}(t)\right) dt+\upsilon _{i,l}^{\beta }dW_{i,l,t}^{\beta},\ \ i=1,\cdots,p,\ \ l=1,2,3,
\]
where $W_{i,l,t}^{\beta}$ are standard Brownian motions independently over $i$ and $l$, $\kappa _{i,1}^{\beta},\kappa_{i,2}^{\beta},\kappa _{i,3}^{\beta}\sim {\sf U}(1,3)$, $\alpha _{i,1}^{\beta }\sim{\sf U}( 0.25,2.25)$, $\alpha _{i,2}^{\beta},\alpha_{i,3}^{\beta }\sim {\sf U}(-0.5, 0.5)$ and $\upsilon _{i,1}^{\beta },\upsilon _{i,2}^{\beta},\upsilon _{i,3}^{\beta } \sim {\sf U}(2,4)$.
{\em 5.2.2.\ \ Simulation Results}
The spot idiosyncratic volatility matrix is estimated via ((ref)). For ease of comparison, we use exactly the same bandwidth as in our first experiment. The results for the noise-free and noise-contaminated spot idiosyncratic volatility matrix estimates $\widehat{\boldsymbol{\Sigma }}_{t}$ and $\widetilde{\boldsymbol{\Sigma }}_{t}$ measured by MFL and MSL are reported in Table 3, which reveal some desirable observations. Firstly, we note that our estimation results in terms of MFL and MSL are almost identical across different types of dynamics of factor loadings, indicating that the developed estimation procedure is robust in finite samples to different assumption of the factor loading dynamics as long as they satisfy our smooth restriction, see Assumption (ref)(iv). Secondly, the MFL and MSL values are similar to those reported in Table 2 which were obtained based on data generating model without common factors. This means that the proposed nonparametric time-varying high frequency regression can effectively remove common factors, resulting in accurate estimation of the spot idiosyncratic volatility matrix.
The factor-based spot volatility matrix of ${\mathbf Y}_t$ is estimated via ((ref)). As discussed in Section (ref), we measure the accuracy of the spiked volatility matrix estimate by the relative error defined above Theorem (ref), i.e., consider the following Mean Relative Loss (MRL):
\[
\text{MRL}=\frac{1}{200}\sum_{m=1}^{200}\left( \frac{1}{21}
\sum_{j=1}^{21}\left\Vert \widehat{\boldsymbol{\Sigma }}_{t_{j}}^{Y,(m)}-
\boldsymbol{\Sigma }_{t_{j}}^{Y,(m)}\right\Vert _{\boldsymbol{\Sigma }
_{t_{j}}^{Y,(m)}}\right) .
\]
The relevant results are reported in Table 4, where $\widehat{\boldsymbol\Sigma }_{t}^{Y}$ and $\widetilde{\boldsymbol\Sigma }_{t}^{Y}$ denote the noise-free and noise-contaminated factor-based spot volatility matrix estimates, respectively. We can see that the performance of the shrinkage estimates is substantially better than that of the naive estimate. Unsurprisingly, due to the presence of microstructure noise, the MRL results of $\widetilde{\boldsymbol\Sigma }_{t}^{Y}$ are much higher than those of $\widehat{\boldsymbol\Sigma }_{t}^{Y}$. As in Table 3, our proposed estimation is robust to different factor loading dynamics.
landscape\begin{center}
Table 3: Estimation results for the spot idiosyncratic volatility matrices
{
\begin{tabular}{llllllllclllll}
\hline\hline
& & & \multicolumn{11}{c}{\textquotedblleft Banding"} \\ \cline{4-14}
$\beta $ Dynamics & & & \multicolumn{5}{c}{Frobenius Norm} & &
\multicolumn{5}{c}{Spectral Norm} \\ \cline{4-8}\cline{10-14}
& & & \ Naive & \ Hard & \ Soft & \ AL & \ SCAD & \multicolumn{1}{l} & \
Naive & \ Hard & \ Soft & \ AL & \ SCAD \\ \cline{4-8}\cline{10-14}
Constant & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9037} &
\multicolumn{1}{r}{4.2461} & \multicolumn{1}{r}{5.2485} & \multicolumn{1}{r}{
4.9880} & \multicolumn{1}{r}{3.9910} & \multicolumn{1}{l}{MSL} &
\multicolumn{1}{r}{3.8887} & \multicolumn{1}{r}{0.6359} & \multicolumn{1}{r}{
0.7291} & \multicolumn{1}{r}{0.7154} & \multicolumn{1}{r}{0.5720} \\
& $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.6646
} & \multicolumn{1}{r}{19.3752} & \multicolumn{1}{r}{18.3036} &
\multicolumn{1}{r}{17.7552} & \multicolumn{1}{r}{18.1388} &
\multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{11.1910} & \multicolumn{1}{r}{
2.3576} & \multicolumn{1}{r}{2.2653} & \multicolumn{1}{r}{2.2160} &
\multicolumn{1}{r}{2.2554} \\
Deterministic & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9127}
& \multicolumn{1}{r}{4.1916} & \multicolumn{1}{r}{5.2503} &
\multicolumn{1}{r}{4.9898} & \multicolumn{1}{r}{3.9842} & \multicolumn{1}{l}{
MSL} & \multicolumn{1}{r}{3.8901} & \multicolumn{1}{r}{0.6313} &
\multicolumn{1}{r}{0.7288} & \multicolumn{1}{r}{0.7144} & \multicolumn{1}{r}{
0.5712} \\
& $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.5672
} & \multicolumn{1}{r}{19.3633} & \multicolumn{1}{r}{18.2947} &
\multicolumn{1}{r}{17.7267} & \multicolumn{1}{r}{18.1284} &
\multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.9662} & \multicolumn{1}{r}{
2.3571} & \multicolumn{1}{r}{2.2636} & \multicolumn{1}{r}{2.2128} &
\multicolumn{1}{r}{2.2536} \\
Stochastic & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9099} &
\multicolumn{1}{r}{4.2123} & \multicolumn{1}{r}{5.2498} & \multicolumn{1}{r}{
4.9893} & \multicolumn{1}{r}{3.9872} & \multicolumn{1}{l}{MSL} &
\multicolumn{1}{r}{3.8896} & \multicolumn{1}{r}{0.6331} & \multicolumn{1}{r}{
0.7289} & \multicolumn{1}{r}{0.7149} & \multicolumn{1}{r}{0.5717} \\
& $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.7323
} & \multicolumn{1}{r}{19.3896} & \multicolumn{1}{r}{18.3164} &
\multicolumn{1}{r}{17.7839} & \multicolumn{1}{r}{18.1538} &
\multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{11.3262} & \multicolumn{1}{r}{
2.3603} & \multicolumn{1}{r}{2.2708} & \multicolumn{1}{r}{2.2203} &
\multicolumn{1}{r}{2.2602} \\ \hline
& & & \multicolumn{11}{c}{\textquotedblleft Block-diagonal"} \\
\cline{4-14}
$\beta $ Dynamics & & & \multicolumn{5}{c}{Frobenius Norm} & &
\multicolumn{5}{c}{Spectral Norm} \\ \cline{4-8}\cline{10-14}
& & & \ Naive & \ Hard & \ Soft & \ AL & \ SCAD & \multicolumn{1}{l} & \
Naive & \ Hard & \ Soft & \ AL & \ SCAD \\ \cline{4-8}\cline{10-14}
Constant & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9047} &
\multicolumn{1}{r}{5.6802} & \multicolumn{1}{r}{6.4718} & \multicolumn{1}{r}{
5.9421} & \multicolumn{1}{r}{5.4710} & \multicolumn{1}{l}{MSL} &
\multicolumn{1}{r}{3.9741} & \multicolumn{1}{r}{0.8722} & \multicolumn{1}{r}{
1.1481} & \multicolumn{1}{r}{0.9106} & \multicolumn{1}{r}{0.9014} \\
& $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.7266
} & \multicolumn{1}{r}{19.8195} & \multicolumn{1}{r}{18.8114} &
\multicolumn{1}{r}{18.3162} & \multicolumn{1}{r}{18.6638} &
\multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.9751} & \multicolumn{1}{r}{
2.8701} & \multicolumn{1}{r}{2.7656} & \multicolumn{1}{r}{2.7097} &
\multicolumn{1}{r}{2.7551} \\
Deterministic & $\widehat{\boldsymbol{\Sigma }}_{t}$ & \ MFL &
\multicolumn{1}{r}{21.9137} & \multicolumn{1}{r}{5.6821} &
\multicolumn{1}{r}{6.4738} & \multicolumn{1}{r}{5.9436} & \multicolumn{1}{r}{
5.4729} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{3.9754} &
\multicolumn{1}{r}{0.8718} & \multicolumn{1}{r}{1.1479} & \multicolumn{1}{r}{
0.9103} & \multicolumn{1}{r}{0.9012} \\
& $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.6284
} & \multicolumn{1}{r}{19.8161} & \multicolumn{1}{r}{18.8043} &
\multicolumn{1}{r}{18.2953} & \multicolumn{1}{r}{18.6547} &
\multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.7452} & \multicolumn{1}{r}{
2.8706} & \multicolumn{1}{r}{2.7663} & \multicolumn{1}{r}{2.7092} &
\multicolumn{1}{r}{2.7559} \\
Stochastic & $\widehat{\boldsymbol{\Sigma }}_{t}$ & \ MFL &
\multicolumn{1}{r}{21.9108} & \multicolumn{1}{r}{5.6811} &
\multicolumn{1}{r}{6.4732} & \multicolumn{1}{r}{5.9433} & \multicolumn{1}{r}{
5.4722} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{3.9751} &
\multicolumn{1}{r}{0.8721} & \multicolumn{1}{r}{1.1480} & \multicolumn{1}{r}{
0.9104} & \multicolumn{1}{r}{0.9013} \\
& $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.7955
} & \multicolumn{1}{r}{19.8314} & \multicolumn{1}{r}{18.8237} &
\multicolumn{1}{r}{18.3434} & \multicolumn{1}{r}{18.6767} &
\multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{11.1149} & \multicolumn{1}{r}{
2.8719} & \multicolumn{1}{r}{2.7691} & \multicolumn{1}{r}{2.7142} &
\multicolumn{1}{r}{2.7584} \\ \hline
& & & \multicolumn{11}{c}{\textquotedblleft Exponentially decaying"} \\
\cline{4-14}
$\beta $ Dynamics & & & \multicolumn{5}{c}{Frobenius Norm} &
\multicolumn{1}{l} & \multicolumn{5}{c}{Spectral Norm} \\
\cline{4-8}\cline{10-14}
& & & \ Naive & \ Hard & \ Soft & \ AL & \ SCAD & \multicolumn{1}{l} & \
Naive & \ Hard & \ Soft & \ AL & \ SCAD \\ \cline{4-8}\cline{10-14}
Constant & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9057} &
\multicolumn{1}{r}{6.0626} & \multicolumn{1}{r}{6.7715} & \multicolumn{1}{r}{
6.1573} & \multicolumn{1}{r}{5.7617} & \multicolumn{1}{l}{MSL} &
\multicolumn{1}{r}{4.0142} & \multicolumn{1}{r}{0.9106} & \multicolumn{1}{r}{
1.1898} & \multicolumn{1}{r}{0.9453} & \multicolumn{1}{r}{0.9388} \\
& $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.8728
} & \multicolumn{1}{r}{20.3802} & \multicolumn{1}{r}{19.2709} &
\multicolumn{1}{r}{18.7715} & \multicolumn{1}{r}{19.1262} &
\multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.8858} & \multicolumn{1}{r}{
2.9381} & \multicolumn{1}{r}{2.8260} & \multicolumn{1}{r}{2.7707} &
\multicolumn{1}{r}{2.8154} \\
Deterministic & $\widehat{\boldsymbol{\Sigma }}_{t}$ & \ MFL &
\multicolumn{1}{r}{21.9147} & \multicolumn{1}{r}{6.0709} &
\multicolumn{1}{r}{6.7737} & \multicolumn{1}{r}{6.1589} & \multicolumn{1}{r}{
5.7637} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{4.0156} &
\multicolumn{1}{r}{0.9112} & \multicolumn{1}{r}{1.1896} & \multicolumn{1}{r}{
0.9450} & \multicolumn{1}{r}{0.9387} \\
& $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.7746
} & \multicolumn{1}{r}{20.3564} & \multicolumn{1}{r}{19.2632} &
\multicolumn{1}{r}{18.7460} & \multicolumn{1}{r}{19.1173} &
\multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.6538} & \multicolumn{1}{r}{
2.9354} & \multicolumn{1}{r}{2.8247} & \multicolumn{1}{r}{2.7673} &
\multicolumn{1}{r}{2.8140} \\
Stochastic & $\widehat{\boldsymbol{\Sigma }}_{t}$ & \ MFL &
\multicolumn{1}{r}{21.9118} & \multicolumn{1}{r}{6.0636} &
\multicolumn{1}{r}{6.7730} & \multicolumn{1}{r}{6.1585} & \multicolumn{1}{r}{
5.7630} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{4.0151} &
\multicolumn{1}{r}{0.9106} & \multicolumn{1}{r}{1.1897} & \multicolumn{1}{r}{
0.9451} & \multicolumn{1}{r}{0.9388} \\
& $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.9430
} & \multicolumn{1}{r}{20.3820} & \multicolumn{1}{r}{19.2839} &
\multicolumn{1}{r}{18.8017} & \multicolumn{1}{r}{19.1405} &
\multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{11.0295} & \multicolumn{1}{r}{
2.9381} & \multicolumn{1}{r}{2.8291} & \multicolumn{1}{r}{2.7745} &
\multicolumn{1}{r}{2.8173} \\ \hline\hline
\end{tabular}
}
\end{center}
center[center omitted — 4,425 chars of source]
Empirical Study
\setcounter{equation}{0}
We apply the proposed methods to the intraday returns of the S&P 500 component stocks to demonstrate the effectiveness of our nonparametric spot volatility matrix estimation in revealing time-varying patterns. We consider the 5-minute returns of the S&P 500 stocks collected in September 2008. On September 15 Lehman Brothers filed for bankruptcy, causing shockwaves throughout the global financial system. Hence, it is interesting to examine how the spot volatility structure of the returns evolved during this one-month period. In addition, to demonstrate the effectiveness of our model with observed risk factors in explaining the systemic component of the dependence structure, we also collect the 5-minute returns of twelve factors. The first three factors are constructed in AKX20 as our proxy for the market (MKT), small-minus-big market capitalisation (SMB), and high-minus-low price-earning ratio (HML). The other nine factors are the widely available sector SDPR ETFs, which are intended to tract the following nine largest S&P sectors: Energy (XLE), Materials (XLB), Industrials (XLI), Consumer Discretionary (XLY), Consumer Staples (XLP), Health Care (XLV), Financial (XLF), Information Technology (XLK), Utilities (XLU). We sort our stocks according to their GICS (Global Industry Classification Standard) codes, so that they are grouped by sectors in the above order. Consequently, the correlation (sub)matrix for stocks within each sector corresponds to a block on the diagonal of the full correlation matrix FFX16.
We only use stocks that are included in the S&P 500 index and whose GICS codes are unchanged in September 2008. We also exclude stocks that do not belong to any of the above nine sectors. This leaves us with a total of $p=482$ stocks. All the returns are synchronised via the previous-tick subsampling technique Zh11, and overnight returns are removed because of potential dividends and stock splits. Consequently, we have $1638$ time series
observations for each of the $482$ stocks. For the 5-minute returns, we may assume that the potential impact of microstructure noises are negligible. The smoothing parameter in our kernel estimation is chosen as $h=2/252$ (equivalent to 2 trading days)\footnote{We experimented three bandwidth choices, namely, 1 day, 2 days and 3 days.
We found that $h=1/252$ ($h=3/252$) produced clearly undersmoothed
(oversmoothed) time series of estimated deciles of the cross sectional
distribution of the variances and pairwise correlations of our returns,
whereas $h=2/252$ seems to be most reasonable. Our qualitative
conclusion is unaffected by the choices of $h$ within the range of 1 to 3
days.}.
We start with estimating the spot volatility matrices of the total returns (i.e., the observed returns) without incorporating the observed factors or applying any shrinkage. To visualise the potential time variation of the estimated spot matrices, as in BHMR19, we plot the time series of deciles of the distribution of the estimated spot variances and the pairwise correlations\footnote{The nine decile levels we use in this study are the 10th, 20th,..., and 90th percentiles.}. The patterns of the spot variances and correlations in Figure (ref)(a) and (b) reveal some clear evidence of time variation in our sampling period. We note that the distributions of the variances are relatively narrow and stay low on the first few days of the month. However, close to Lehman Brothers' announcement on the 15th, they start to rise and get wider quite rapidly and reach the peak around the 17th and the 18th. The spot variances at the peak are much higher than those on the earlier days of the month. The distributions return to the earlier level in the following week. In contrast, the distributions of pairwise spot correlations also start to shift up around the same time, but quickly reach the peak on the 16th (only one day after the bankruptcy news), and then dip to a relatively low point around the 19th before returning to the earlier level. Such time-varying features in the dynamics covariance structure are quite interesting and sensible, reflecting the impact of market news. Hence, our proposed spot volatility matrix estimation methodology provides a useful tool for revealing such dynamics.
figure[figure omitted — 179 chars of source]
To examine whether it is appropriate to directly apply shrinkage techniques to the spot volatility matrices of the total returns, following FFX16 we plot in Figure (ref)(a) and (b) their sparsity patterns on the 16th and the 19th of September \footnote{Recall that as in DLX19 our tuning parameter used for each pairwise spot covariance in our shrinkage method is proportional to the product of the spot standard deviations of the returns of that pair of assets. Therefore, the sparsity pattern is effectively determined by the spot correlation matrix.}. The deep blue dots correspond to the locations of pairwise correlations that are at least $0.15$, whereas the white dots correspond to those smaller than $0.15$. Note that the covariance structure of the total returns is very dense on these two days. Therefore, it is not appropriate to directly apply the shrinkage technique as in Sections (ref) and (ref). Meanwhile, although both are quite dense, we can still clearly see their differences. Consistent with our observation from the decile plots of the correlations, we can see that the plot for the 16th is almost completely covered by blue dots, but the plot for the 19th in contrast has significantly more areas covered in white.
figure[figure omitted — 215 chars of source]
We next incorporate the twelve observed factors in the large spot volatility matrix estimation as suggested in Section (ref). In particular, we are interested on estimating the spot idiosyncratic volatility matrix, which is expected to satisfy the sparsity restriction. To save space, we choose to only report results using the SCAD shrinkage due to its satisfactory performance in our simulations. In Figure (ref)(c) and (d), we plot the deciles of the estimated spot idiosyncratic variances and correlations over trading days. In Figure (ref)(c), we observe a significant upward shift of the distribution of the spot variances of the idiosyncratic returns around the time of Lehman Brothers' bankruptcy, indicating that the observed factors may not fully capture the time variation of the spot variances. In contrast, the deciles of the spot correlations in Figure (ref)(d) seem to be quite flat throughout the entire month, suggesting that the systematic factors may explain the time variation in the distribution of the pairwise correlations better than that of the variances.
We finally plot the sparsity patterns of the two estimated spot idiosyncratic volatility matrices on 16 and 19 September in Figure (ref)(c) and (d), respectively. Unlike Figure (ref)(a) and (b), we note that the estimated spot idiosyncratic volatility matrices are highly sparse on the two days. This is consistent with our observation from Figure (ref)(d), confirming that the observed factors can effectively account for the time variation in the spot covariance structure of the returns. Meanwhile, we also note that the two idiosyncratic volatility matrices are clearly not diagonal and still carry some visible time variation. Lastly, it is worth mentioning that the estimated spot idiosyncratic volatility matrices do not exhibit significant correlations within the blocks along the diagonal lines, except for some very limited actions in the lower right corner of the two matrices (the lower right corner corresponds to the XLU sector according to our sorting).
Conclusion
We developed nonparametric estimation methods for large spot volatility matrices under the uniform sparsity assumption. We allowed for microstructure noise and observed common risk factors and employed kernel smoothing and generalised shrinkage. In each scenario we obtained the uniform convergence rates for the large estimated covariance matrices and these reflect the smoothness and sparsity assumptions we made. The simulation results show that the proposed estimation methods work well in finite samples for both the noise-free and noise-contaminated data. The empirical study demonstrated the effectiveness on S&P 500 stocks five minute data. Several issues can be further explored. For example, it is worthwhile to further study the spot precision matrix estimation which is briefly discussed in Appendix C.1 of the supplement and explore its application to optimal portfolio choice.
Acknowledgements
The authors would like to thank a Co-Editor and two reviewers for the constructive comments, which helped to improve the article. The first author's research was partly supported by the BA Talent Development Award (No. TDA21$\backslash$210027). The second author’s research was partly supported by the BA/Leverhulme Small Research Grant funded by the Leverhulme Trust (No. SRG1920/ 100603).
thebibliography{99}
{ \harvarditem{A\"{\i}t-Sahalia \harvardand\ Jacod}{2014}{AJ14}
A\"{\i}t-Sahalia, Y. & J. Jacod (2014) High-Frequency
Financial Econometrics. Princeton University Press. }
{ \harvarditem{A\"{\i}t-Sahalia, Kalnina \harvardand\ Xiu}{2020}{AKX20}
A\"{\i}t-Sahalia, Y., I. Kalnina, & D. Xiu (2020) High-frequency factor models and regressions. Journal of Econometrics 216, 86--105. }
{ \harvarditem{A\"{\i}t-Sahalia \harvardand\ Xiu}{2017}{AX17}
A\"{\i}t-Sahalia, Y. & D. Xiu (2017) Using principal component
analysis to estimate a high dimensional factor model with high-frequency
data. Journal of Econometrics 201, 384--399. }
{ \harvarditem{A\"{\i}t-Sahalia \harvardand\ Xiu}{2019}{AX19}
\textsc{A\"{\i}t-Sahalia, Y. & D. Xiu} (2019) Principal component
analysis of high-frequency data. \emph{Journal of the American Statistical
Association} 114, 287--303. }
{ \harvarditem{Andersen and Bollerslev}{1998}{AB98} \textsc{
Andersen, T. G. & T. Bollerslev} (1998) Answering the skeptics: yes,
standard volatility models do provide accurate forecasts. \emph{
International Economic Review} 39, 885--905. }
{ \harvarditem{Andersen, Bollerslev \harvardand\
Diebold}{2010}{ABD10} \textsc{Andersen, T. G., T. Bollerslev, & F. X. Diebold} (2010) Parametric and nonparametric volatility measurement. In \emph{
Handbook of Financial Econometrics: Tools and Techniques} (Y. A\"{\i}
t-Sahalia and L. P. Hansen, eds.), 67--137. }
{ \harvarditem{Andersen {\em et al.}}{2003}{ABDL03} \textsc{
Andersen, T. G., T. Bollerslev, F. X. Diebold, & P. Labys} (2003)
Modeling and forecasting realized volatility. \emph{Econometrica} 71,
579--625. }
{ \harvarditem{Bai \harvardand\ Silverstein}{2010}{BS10}
\textsc{Bai, Z. & J. W. Silverstein} (2010) \emph{Spectral Analysis of
Large Dimensional Random Matrices}. Springer Series in Statistics, Springer.
}
{ \harvarditem{Barndorff-Nielsen \harvardand\
Shephard}{2002}{BS02} \textsc{Barndorff-Nielsen, O. E. & N. Shephard}
(2002) Econometric analysis of realized volatility and its use in
estimating stochastic volatility models. \emph{Journal of the Royal
Statistical Society Series B} 64, 253--280. }
{ \harvarditem{Barndorff-Nielsen \harvardand\
Shephard}{2004}{BS04} \textsc{Barndorff-Nielsen, O. E. & N. Shephard}
(2004) Econometric analysis of realized covariation: High frequency based
covariance, regression and correlation in financial economics. \emph{
Econometrica} 72, 885--925. }
{ \harvarditem{Barndorff-Nielsen {\em et al.}}{2008}{BHLS08}
\textsc{Barndorff-Nielsen, O. E., P. R. Hansen, A. Lunde, & N. Shephard}
(2008) Designing realised kernels to measure the ex-post variation of
equity prices in the presence of noise. \emph{Econometrica} 76, 1481--1536. }
{ \harvarditem{Bibinger {\em et al.}}{2019}{BHMR19}\textsc{Bibinger, M., N. Hautsch, P. Malec, & M. Reiss} (2019) Estimating the spot covariation of asset prices - statistical theory and empirical evidence. \emph{Journal of Business and Economic Statistics} 37(3), 419--435. }
{ \harvarditem{Bickel \harvardand\ Levina}{2008}{BL08} \textsc{
Bickel, P. & E. Levina} (2008) Covariance regularization by
thresholding. \emph{Annals of Statistics} 36, 2577--2604. }
{ \harvarditem{Cai {\em et al}}{2020}{CHLZ20} \textsc{Cai, T.
T., J. Hu, Y. Li, & X. Zheng} (2020) High-dimensional minimum
variance portfolio estimation based on high-frequency data. \emph{Journal of
Econometrics} 214, 482--494. }
{ \harvarditem{Cai \harvardand\ Zhou}{2012}{CZ12} \textsc{Cai,
T. T. & H. H. Zhou} (2012) Optimal rates of convergence for
sparse covariance matrix estimation. \emph{Annals of Statistics} 40,
2389--2420. }
{ \harvarditem{Chang {\em et al}.}{2021}{CHLT21} \textsc{Chang,
J., Q. Hu, C. Liu, & C. Tang} (2021) Optimal covariance matrix
estimation for high-dimensional noise in high-frequency data. Working paper
available at \url{https://arxiv.org/abs/1812.08217.} }
{ \harvarditem{Chen, Li \harvardand\ Linton}{2019}{CLL19}
\textsc{Chen, J., D. Li, & O. Linton} (2019) A new
semiparametric estimation approach of large dynamic covariance matrices with
multiple conditioning variables. \emph{Journal of Econometrics} 212,
155--176. }
{ \harvarditem{Chen, Mykland \harvardand\ Zhang}{2020}{CMZ20} \textsc{Chen, D., P. A. Mykland, & L. Zhang} (2020) The five trolls under the bridge: principal component analysis with asynchronous and noisy high frequency data. {\em Journal of the American Statistical Association} 115, 1960--1977. }
{ \harvarditem{Chen, Xu \harvardand\ Wu}{2013}{CXW13} \textsc{
Chen, X., M. Xu, & W. Wu} (2013) Covariance and precision
matrix estimation for high-dimensional time series. \emph{Annals of
Statistics} 41, 2994--3021. }
{ \harvarditem{Chen \harvardand\ Leng}{2016}{CL16} \textsc{
Chen, Z.& C. Leng} (2016) Dynamic covariance models. \emph{
Journal of the American Statistical Association} 111, 1196--1207. }
{ \harvarditem{Christensen, Kinnebrock \harvardand\
Podolskij}{2010}{CKP10} \textsc{Christensen, K., S. Kinnebrock, & M.
Podolskij} (2010) Pre-averaging estimators of the ex-post covariance
matrix in noisy diffusion models with non-synchronous data. \emph{Journal of
Econometrics} 159, 116--133. }
{ \harvarditem{Dai, Lu \harvardand\ Xiu}{2019}{DLX19} \textsc{
Dai, C., K. Lu, & D. Xiu} (2019) Knowing factors or factor loadings, or
neither? Evaluating estimators for large covariance matrices with noisy and
asynchronous data. \emph{Journal of Econometrics} 208, 43--79. }
{ \harvarditem{Epps}{1979}{Ep79} \textsc{Epps, T. W.} (1979)
Comovements in stock prices in the very short run. \emph{Journal of the
American Statistical Association} 74, 291--298. }
{ \harvarditem{Fan, Fan \harvardand\ Lv}{2007}{FFL07} \textsc{
Fan, J., Y. Fan, & J. Lv} (2007) Aggregation of nonparametric estimators
for volatility matrix. \emph{Journal of Financial Econometrics} 5, 321--357.
}
{ \harvarditem{Fan, Furger \harvardand\ Xiu}{2016}{FFX16}
\textsc{Fan, J., A. Furger, & D. Xiu} (2016) Incorporating global
industrial classification standard into portfolio allocation: A simple
factor-based large covariance matrix estimator with high frequency data.
\emph{Journal of Business and Economic Statistics} 34, 489--503. }
{ \harvarditem{Fan \harvardand\ Gijbels}{1996}{FG96}\textsc{Fan, J. & I. Gijbels} (1996) {\em Local Polynomial Modelling and Its Applications}. Chapman and Hall, London. }
{ \harvarditem{Fan, Liao \harvardand\ Mincheva}{2011}{FLM11}
\textsc{Fan, J., Y. Liao, & M. Mincheva} (2011)
High-dimensional covariance matrix estimation in approximate factor models.
\emph{Annals of Statistics} 39, 3320--3356. }
{ \harvarditem{Fan, Liao \harvardand\ Mincheva}{2013}{FLM13}
\textsc{Fan, J., Y. Liao, & M. Mincheva} (2013) Large
covariance estimation by thresholding principal orthogonal complements (with
discussion). \emph{Journal of the Royal Statistical Society, Series B} 75,
603--680. }
{ \harvarditem{Fan \harvardand\ Wang}{2008}{FW08} \textsc{Fan,
J. & Y. Wang} (2008) Spot volatility estimation for high-frequency data.
\emph{Statistics and Its Interface} 1, 279--288. }
{ \harvarditem{Figueroa-L\'{o}pez \harvardand\ Li}{2020}{FL20}
\textsc{Figueroa-L\'{o}pez, J. E. & C. Li} (2020) Optimal
kernel estimation of spot volatility of stochastic differential equations.
\emph{Stochastic Processes and Their Applications} 130, 4693--4720. }
{ \harvarditem{Hayashi \harvardand\ Yoshida}{2005}{HY05}
\textsc{Hayashi, T. & N. Yoshida} (2005) On covariance
estimation of non-synchronously observed diffusion processes. \emph{Bernoulli
} 11, 359--379. }
{ \harvarditem{Jacod {\em et al.}}{2009}{JLMPV09} \textsc{
Jacod, J., Y. Li, P. A. Mykland, M. Podolskij, & M. Vetter} (2009)
Microstructure noise in the continuous case: the pre-averaging approach.
\emph{Stochastic Processes and Their Applications} 119, 2249--2276. }
{ \harvarditem{Jacod \harvardand\ Protter}{2012}{JP12} \textsc{
Jacod, J. & P. Protter} (2012) \emph{Discretization of Processes}.
Springer. }
{ \harvarditem{Kalnina \harvardand\ Linton}{2008}{KL08} \textsc{
Kalnina, I. & O. Linton} (2008) Estimating quadratic variation
consistently in the presence of endogenous and diurnal measurement error.
\emph{Journal of Econometrics} 147, 47--59. }
{ \harvarditem{Kanaya \harvardand\ Kristensen}{2016}{KK16}
\textsc{Kanaya, S. & D. Kristensen} (2016) Estimation of stochastic
volatility models by nonparametric filtering. \emph{Econometric Theory} 32,
861--916. }
{ \harvarditem{Kim, Wang \harvardand\ Zou}{2016}{KWZ16} \textsc{
Kim, D., Y. Wang, & J. Zou} (2016) Asymptotic theory for large
volatility matrix estimation based on high-frequency financial data. \emph{
Stochastic Processes and Their Applications} 126, 3527--3577. }
{ \harvarditem{Kong}{2018}{K18} \textsc{Kong, X.} (2018) On
the systematic and idiosyncratic volatility with large panel high-frequency
data. \emph{Annals of Statistics} 46, 1077--1108. }
{ \harvarditem{Kristensen}{2010}{K10} \textsc{Kristensen, D.}
(2010) Nonparametric filtering of the realized spot volatility: a
kernel-based approach. \emph{Econometric Theory} 26, 60--93. }
{ \harvarditem{Lam \harvardand\ Feng}{2018}{LF18} \textsc{Lam,
C. & P. Feng} (2018) A nonparametric eigenvalue-regularized integrated
covariance matrix estimator for asset return data. \emph{Journal of
Econometrics} 206, 226--257. }
{ \harvarditem{Li \harvardand\ Racine}{2007}{LR07} \textsc{Li,
Q. & J. Racine} (2007) \emph{Nonparametric Econometrics}.
Princeton University Press, Princeton. }
{ \harvarditem{Mykland \harvardand\ Zhang}{2006}{MZ06} \textsc{Mykland, P. A. & L. Zhang} (2006) ANOVA for diffusions and It\^o processes. {\em Annals of Statistics} 34, 1931--1963.}
{ \harvarditem{Park, Hong \harvardand\ Linton}{2016}{PHL16}
\textsc{Park, S., S. Y. Hong, & O. Linton} (2016) Estimating the
quadratic covariation matrix for asynchronously observed high frequency
stock returns corrupted by additive measurement error. \emph{Journal of
Econometrics} 191, 325-347. }
{ \harvarditem{Podolskij \harvardand\ Vetter}{2009}{PV09}
\textsc{Podolskij, M. & M. Vetter} (2009) Estimation of volatility
functionals in the simultaneous presence of microstructure noise and jumps.
\emph{Bernoulli} 15, 634--658. }
{ \harvarditem{Rei{\ss}, Todorov \harvardand\ Tauchen}{2015}{RTT15} \textsc{Rei{\rm {\ss}}, M., V. Todorov, & G. Tauchen} (2015) Nonparametric test for a constant beta between It\^o semi-martingales based on high-frequency data. {\em Stochastic Processes and Their Applications} 125, 2955--2988.}
{ \harvarditem{Revuz \harvardand\ Yor}{1999}{RY99} \textsc{
Revuz, D. & M. Yor} (1999) \emph{Continuous Martingales and Brownian
Motion}. Grundlehren der mathematischen Wissenschaften 293, Springer. }
{ \harvarditem{Shephard}{2005}{S05} \textsc{Shephard, N.}
(2005) \emph{Stochastic Volatility: Selected Readings}. Oxford University
Press. }
{ \harvarditem{Tao, Wang \harvardand\ Zhou}{2013}{TWZ13}
\textsc{Tao, M., Y. Wang, & H. H. Zhou} (2013) Optimal sparse
volatility matrix estimation for high-dimensional It\^{o} processes with
measurement errors. \emph{Annals of Statistics} 41, 1816--1864. }
{ \harvarditem{Wang \harvardand\ Zou}{2010}{WZ10} \textsc{Wang,
Y. & J. Zou} (2010) Vast volatility matrix estimation for
high-frequency financial data. \emph{Annals of Statistics} 38, 943--978. }
{ \harvarditem{Xia \harvardand\ Zheng}{2018}{XZ18} \textsc{Xia,
N. & X. Zheng} (2018) On the inference about the spectral distribution
of high-dimensional covariance matrix based on high-frequency noisy
observations. \emph{Annals of Statistics} 46, 500--525. }
{ \harvarditem{Xiao \harvardand\ Linton}{2002}{XL02} \textsc{
Xiao, Z. & O. Linton} (2002) A nonparametric prewhitened covariance
estimator. \emph{Journal of Time Series Analysis} 23, 215--250. }
{ \harvarditem{Zhang}{2011}{Zh11} \textsc{Zhang, L.} (2011)
Estimating covariation: Epps effect, microstructure noise. \emph{Journal of
Econometrics} 160, 33--47. }
{ \harvarditem{Zhang, Mykland \harvardand\
A\"{\i}t-Sahalia}{2005}{ZMA05} \textsc{Zhang, L., P. A. Mykland, & Y. A\"{\i}
t-Sahalia} (2005) A tale of two time scales: Determining integrated
volatility with noisy high-frequency data. \emph{Journal of the American
Statistical Association} 100, 1394--1411. }
{ \harvarditem{Zheng \harvardand\ Li}{2011}{ZL11} \textsc{
Zheng, X. & Y. Li} (2011) On the estimation of integrated covariance
matrices of high dimensional diffusion processes. \emph{Annals of Statistics}
39, 3121--3151. }
{ \harvarditem{Zu \harvardand\ Boswijk}{2014}{ZB14} \textsc{Zu,
Y. & H. P. Boswijk} (2014) Estimating spot volatility with
high-frequency financial data. \emph{Journal of Econometrics} 181, 117--135.
}