EconBase
← Back to paper

Nonparametric Estimation of Large Spot Volatility Matrices for High-Frequency Financial Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

118,246 characters · 12 sections · 94 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric Estimation of Large Spot Volatility Matrices for High-Frequency Financial Data

\centerline{\bf Abstract}

In this paper, we consider estimating spot/instantaneous volatility matrices of high-frequency data collected for a large number of assets. We first combine classic nonparametric kernel-based smoothing with a generalised shrinkage technique in the matrix estimation for noise-free data under a uniform sparsity assumption, a natural extension of the approximate sparsity commonly used in the literature. The uniform consistency property is derived for the proposed spot volatility matrix estimator with convergence rates comparable to the optimal minimax one. For the high-frequency data contaminated by microstructure noise, we introduce a localised pre-averaging estimation method that reduces the effective magnitude of the noise. We then use the estimation tool developed in the noise-free scenario, and derive the uniform convergence rates for the developed spot volatility matrix estimator. We further combine the kernel smoothing with the shrinkage technique to estimate the time-varying volatility matrix of the high-dimensional noise vector. In addition, we consider large spot volatility matrix estimation in time-varying factor models with observable risk factors and derive the uniform convergence property. We provide numerical studies including simulation and empirical application to examine the performance of the proposed estimation methods in finite samples.

\noindentKeywords: \ Brownian semi-martingale, Factor model, Kernel smoothing, Microstructure noise, Sparsity, Spot volatility matrix, Uniform consistency.

Introduction

\setcounter{equation}{0}

Modelling high-frequency financial data is one of the most important topics in financial economics and has received increasing attention in recent decades. Continuous-time econometric models such as the It\^{o} semimartingale are often employed in the high-frequency data analysis. One of the main components in these models is the volatility function or matrix. In the low-dimensional setting (with a single or a small number of assets), the realised volatility is often used to estimate the integrated volatility over a fixed time period AB98, BS02, BS04, ABDL03. In practice, it is not uncommon that the high-frequency financial data are contaminated by the market microstructure noise, which leads to biased realised volatility if the noise is ignored. Hence, various modification techniques such as the two-scale, pre-averaging and realised kernel have been introduced to account for the microstructure noise and produce consistent volatility estimation ZMA05, BHLS08, KL08, JLMPV09, PV09, CKP10, PHL16. S05, ABD10 and AJ14 provide comprehensive reviews for estimating volatility with high-frequency financial data under various settings.

In practical applications, financial economists often have to deal with the situation that there are a large amount of high-frequency financial data collected for a large number of assets. A key issue is to estimate the large volatility structure for these assets, which has applications in various areas such as the optimal portfolio choice and risk management. Partly motivated by developments in large covariance matrix estimation for low-frequency data in the statistical literature, WZ10, TWZ13 and KWZ16 estimate the large volatility matrix under an approximate sparsity assumption BL08; ZL11 and XZ18 study large volatility matrix estimation using the large-dimensional random matrix theory BS10; and LF18 propose a nonparametric eigenvalue-regularised integrated covariance matrix for high-dimensional asset returns. Given that there often exists co-movement between a large number of assets and the co-movement is driven by some risk factors which can be either observable or latent, FFX16, AX17, DLX19 extend the methodologies developed by FLM11, FLM13 to estimate the large volatility matrix by imposing a continuous-time factor model structure on the high-dimensional and high-frequency financial data, and AX19 study the principal component analysis of high-frequency data and derive the asymptotic distribution for the realised eigenvalues, eigenvectors and principal components.

The estimation methodologies in the aforementioned literature often rely on the realised volatility (or covariance) matrices, measuring the integrated volatility structure over a fixed time interval. In practice, it is often interesting to further explore the actual spot/instantaneous volatility structure and its dynamic change over a certain time interval, which is a particularly important measurement for the financial assets when the market is in a volatile period (say, the global financial crisis or COVID-19 outbreak). For a single financial asset, FW08 and K10 introduce a kernel-based nonparametric method to estimate the spot volatility function and establish its asymptotic properties including the point-wise and global asymptotic distribution theory and uniform consistency. For the noise-contaminated high-frequency data, ZB14 combine the two-scale realised volatility with the kernel-weighted technique to estimate the spot volatility, whereas KK16 propose a kernel-weighted pre-averaging spot volatility estimation method. Other nonparametric spot volatility estimation methods can be found in FFL07 and FL20. It seems straightforward to extend this local nonparametric method to estimate the spot volatility matrix for a small number of assets. However, a further extension to the setting with vast financial assets is non-trivial. There is virtually no work on estimating the large spot volatility matrix except K18, which considers estimating large spot volatility matrices and their integrated versions under the continuous-time factor model structure for noise-free high-frequency data.

The main methodological and theoretical contributions of this paper are summarised as follows.

itemize• {\em Large spot volatility matrix estimation with noise-free high-frequency data}. We use the nonparametric kernel-based smoothing method to estimate the volatility and co-volatility functions as in FW08 and K10, and then apply a generalised shrinkage to off-diagonal estimated entries. With small off-diagonal entries forced to be zeros, the resulting large spot volatility matrix estimate would be non-degenerate with stable performance in finite samples. We derive the consistency property for the proposed spot volatility matrix estimator uniformly over the entire time interval under a uniform sparsity assumption, which is also adopted by CXW13, CL16 and CLL19 in the low-frequency data setting. In particular, the derived uniform convergence rates are comparable to the optimal minimax rate in large covariance matrix estimation CZ12. The number of assets is allowed to be ultra large in the sense that it can grow at an exponential rate of $1/\Delta$ with $\Delta$ being the sampling frequency. • {\em Large spot volatility matrix estimation with noise-contaminated high-frequency data and time-varying noise volatility matrix estimation}. When the high-frequency data are contaminated by the microstructure noise, we extend KK16's localised pre-averaging estimation method to the high-dimensional data setting. Specifically, we first pre-average the log price data via a kernel filter and then apply the same estimation method to the kernel fitted high-frequency data (at pseudo-sampling time points) as in the noise-free scenario. The microstructure noise vector is assumed to be heteroskedastic with the time-varying covariance matrix satisfying the uniform sparsity assumption. We show that the existence of microstructure noises slows down the uniform convergence rates, see Theorem (ref). Furthermore, we combine the kernel smoothing with generalised shrinkage to estimate the time-varying noise volatility matrix and derive its uniform convergence property. To the best of our knowledge, there is virtually no work on large time-varying noise volatility matrix estimation for high-frequency data. • {\em Large spot volatility matrix estimation with risk factors}. Since the uniform sparsity assumption is often too restrictive, we relax this restriction in Section (ref) and consider large spot volatility matrix estimation in the time-varying factor model at high frequency, i.e., a large number of asset prices are driven by a small number of observable common factors. By imposing the sparsity restriction on the spot idiosyncratic volatility matrix, we obtain the so-called “low-rank plus sparse" spot volatility structure. A similar structure (with constant betas) is adopted by FFX16 and DLX19 in estimation of large integrated volatility matrices. We use the kernel smoothing method to estimate the spot volatility and covariance of the observed asset prices and factors as well as the time-varying betas, and apply the shrinkage technique to the estimated spot idiosyncratic volatility matrix. We derive the uniform convergence property of the developed matrix estimates, partly extending the point-wise convergence property in K18. The developed methodology and theory can be further modified to tackle the noise-contaminated high-frequency data.

We argue that all three of the scenarios we consider above may be practically relevant. Microstructure noise is considered important in the very highest frequency of data, whereas researchers working with five minute data, say, often ignore the noise. For the lower frequency of data there is a lot comovement in returns and the factor model is designed to capture that comovement, whereas at the ultra high frequency, comovement is less of an issue; indeed, under the so-called Epps effect this comovement shrinks to zero with sampling frequency.

The rest of the paper is organised as follows. In Section (ref), we estimate the large spot volatility matrix in the noise-free high-frequency data setting and give the uniform consistency property. In Section (ref) , we extend the methodology and theory to the noise-contaminated data setting and further estimate the time-varying noise volatility matrix. Section (ref) considers the large spot volatility matrix with systematic factors. Section (ref) reports the simulation studies and Section (ref) provides an empirical application. Section (ref) concludes the paper. Proofs of the main theoretical results are available in Appendix A. The supplementary document contains proofs of some technical lemmas and propositions and discussions on the spot precision matrix estimation and the asynchronicity issue. Throughout the paper, we let $\Vert\cdot\Vert_2$ be the Euclidean norm of a vector; and for a $d\times d$ matrix ${\mathbf{A}}=(A_{ij})_{d\times d}$, we let $\Vert { \mathbf{A}}\Vert$ and $\Vert {\mathbf A}\Vert_F$ be the matrix spectral norm and Frobenius norm, $\vert{\mathbf{A}} \vert_1=\sum_{i=1}^d\sum_{i=1}^d |A_{ij}|$, $\Vert{\mathbf{A}} \Vert_1=\max_{1\leq j\leq d}\sum_{i=1}^d |A_{ij}|$, $\Vert{\mathbf{A}} \Vert_{\infty,q}=\max_{1\leq i\leq d}\sum_{j=1}^d |A_{ij}|^q$ and $\Vert { \mathbf{A}}\Vert_{\max}=\max_{1\leq i\leq d}\max_{1\leq j\leq d} |A_{ij}|$.

Estimation with noise-free data

\setcounter{equation}{0}

Suppose that ${\mathbf{X}}_t=\left(X_{1,t},\cdots,X_{p,t}\right)^{^\intercal} $ is a $p$-variate Brownian semi-martingale solving the following stochastic differential equation:

equation[equation omitted — 110 chars of source]

where ${\mathbf{W}}_t=\left(W_{1,t},\cdots,W_{p,t}\right)^{^\intercal}$ is a $p$-dimensional standard Brownian motion, ${\boldsymbol{\mu}} _t=(\mu_{1,t},\cdots,\mu_{p,t})^{^\intercal}$ is a $p$-dimensional drift vector, and ${\boldsymbol{\sigma}}_t=\left(\sigma_{ij,t}\right)_{p\times p}$ is a $p\times p$ matrix. The spot volatility matrix of ${\mathbf{X}}_t$ is defined as

equation[equation omitted — 155 chars of source]

Our main interest lies in estimating ${\boldsymbol{\Sigma}}_t$ when the size $p$ is large. As in CXW13 and CL16, we assume that the true spot volatility matrix satisfies the following uniform sparsity condition: $ \left\{{\boldsymbol{\Sigma}}_t:\ 0\leq t\leq T\right\}\in \mathcal{S} (q,\varpi(p), T)$, where

equation[equation omitted — 246 chars of source]

where $0\le q<1$, $\varpi(p)$ is larger than a positive constant, $T$ is a fixed positive number and $\Lambda$ is a positive random variable satisfying $\mathsf{E}[\Lambda]\leq C_\Lambda<\infty$. This is a natural extension of the approximate sparsity assumption BL08. Section (ref) below will relax this assumption and consider estimating large spot volatility matrices with systematic factors. The asset prices are assumed to be collected over a fixed time interval $[0,T]$ at $0,\Delta,2\Delta,\cdots,n\Delta$, where $\Delta$ is the sampling frequency and $n=\lfloor T/\Delta\rfloor$ with $\lfloor\cdot\rfloor$ denoting the floor function. In the main text, we focus on the case of equidistant time points in the high-frequency data collection. The asynchronicity issue will be discussed in Appendix C.2 of the supplement.

For each $1\leq i,j\leq p$, we estimate the spot co-volatility $\Sigma_{ij,t} $ by

equation[equation omitted — 109 chars of source]

with \[ K_{h}^\ast(t_k-t)=K_h\left(t_k-t\right)/\left[\Delta\sum_{l=1}^nK_h\left(t_l-t\right)\right], \] where $t_k=k\Delta$, $K_h(u)=h^{-1}K(u/h)$, $K(\cdot)$ is a kernel function, $h$ is a bandwidth shrinking to zero and $\Delta X_{i,k}=X_{i,t_k}-X_{i,t_{k-1}}$. The use of $K_h^\ast(t_k-t)$ rather than $K_h(t_k-t)$ in the estimation ((ref)) is to correct a constant bias when $t$ is close to the boundary points $0$ and $T$. A naive method of estimating the spot volatility matrix ${\boldsymbol{\Sigma}}_{t}$ is to directly use $\widehat\Sigma_{ij,t}$ to form an estimated matrix. However, this estimate often performs poorly in practice when the number of assets is very large (say, $p>n$). To address this issue, a commonly-used technique is to apply a shrinkage function to $\widehat\Sigma_{ij,t}$ when $i\neq j$, forcing very small estimated off-diagonal entries to be zeros. Let $ s_\rho(\cdot)$ denote a shrinkage function satisfying the following three conditions: (i) $\vert s_\rho(u)\vert \leq \vert u\vert$ for $u\in\mathscr{R} $; (ii) $s_\rho(u)=0$ if $\vert u\vert \leq \rho$; and (iii) $\vert s_\rho(u)-u\vert\leq \rho$, where $\rho$ is a user-specified tuning parameter. With the shrinkage function, we construct the following nonparametric estimator of ${\boldsymbol{\Sigma}}_{t}$:

equation[equation omitted — 235 chars of source]

where $\rho_1(t)$ is a tuning parameter which is allowed to change over $t$ and $I(\cdot)$ denotes the indicator function. Section (ref) discusses the choice of $\rho_1(t)$, ensuring that $\widehat{\boldsymbol\Sigma}_t$ is positive definite in finite samples. Our estimation method of the spot volatility matrix can be seen as a natural extension of the kernel-based large sparse covariance matrix estimation CXW13, CL16,CLL19 from the low-frequency data setting to the high-frequency one. We next give some technical assumptions which are needed to derive the uniform convergence property of $\widehat{\boldsymbol{ \Sigma}}_{t}$.

assumption(i)\ $\{\mu_{i,t}\}$ and $ \{\sigma_{ij,t}\}$ are adapted locally bounded processes with continuous sample path. (ii)\ With probability one, \begin{equation*} \min_{1\leq i\leq p}\inf_{0\leq s\leq T}\Sigma_{ii, s}>0,\ \ \min_{1\leq i\neq j\leq p}\inf_{0\leq s\leq T} \Sigma_{ij,s}^\ast>0, \end{equation*} where $\Sigma_{ij, s}^\ast=\Sigma_{ii, s}+\Sigma_{jj, s}+2\Sigma_{ij, s}$. For the spot covariance process $\{\Sigma_{ij,t}\}$, there exist $\gamma\in(0,1)$ and $B(t,\epsilon)$, a positive random function slowly varying at $\epsilon=0$ and continuous with respect to $t$, such that \begin{equation} \max_{1\leq i,j\leq p}\left\vert \Sigma_{ij,t+\epsilon}-\Sigma_{ij,t}\right\vert\leq B(t,\epsilon)|\epsilon|^\gamma+o(|\epsilon|^\gamma),\ \ \epsilon\rightarrow0. \end{equation}
assumption(i)\ The kernel $K(\cdot)$ is a bounded and Lipschitz continuous function with a compact support $[-1,1]$. In addition, $\int_{-1}^1 K(u)du=1$. (ii)\ The bandwidth $h$ satisfies that $h\rightarrow0$ and $\frac{h}{ \Delta\log (p\vee \Delta^{-1})}\rightarrow\infty$. (iii) Let the time-varying tuning parameter $\rho_1(t)$ in the generalised shrinkage be chosen as \begin{equation*} \rho_1(t)=M(t)\zeta_{\Delta,p},\ \ \zeta_{\Delta,p}=h^{\gamma}+\left[\frac{ \Delta\log (p\vee \Delta^{-1})}{h}\right]^{1/2}, \end{equation*} where $\gamma$ is defined in ((ref)) and $M(t)$ is a positive function satisfying that \begin{equation*} 0<C_M\le \inf_{0\leq t\leq T}M(t)\leq\sup_{ 0\leq t\leq T}M(t)\leq\overline{C}_M<\infty. \end{equation*}
remarkAssumption (ref) imposes some mild restrictions on the drift and volatility processes. By a typical localisation procedure as in Section 4.4.1 of JP12, the local boundedness condition in Assumption (ref)(i) can be strengthened to the bounded condition over the entire time interval, i.e., with probability one, \begin{equation*} \max_{1\leq i\leq p}\sup_{0\leq s\leq T} |\mu_{i,s}|\leq C_\mu<\infty,\ \ \max_{1\leq i\leq p}\sup_{0\leq s\leq T} \Sigma_{ii,s} \leq C_\Sigma<\infty, \end{equation*} which are the same as Assumption A2 in TWZ13 and Assumptions (A.ii) and (A.iii) in CHLZ20. It may be possible to relax the uniform boundedness restriction (when $T$ is allowed to diverge) at the cost of more lengthy proofs KK16. Assumption (ref)(ii) gives the smoothness condition on the spot covariance process, crucial to derive the uniform asymptotic order for the kernel estimation bias. When the spot covariance is driven by continuous semimartingales, ((ref)) holds with $ \gamma<1/2$ RY99. Assumption (ref)(i) contains some commonly-used conditions for the kernel function. Assumption (ref)(ii)(iii) imposes some mild conditions on the bandwidth and time-varying shrinkage parameter. In particular, when $p$ diverges at a polynomial rate of $1/\Delta$, Assumption (ref)(ii) reduces to the conventional bandwidth restriction. Assumption (ref)(iii) is comparable to that assumed by CL16 and CLL19. It is worthwhile to point out that the developed methodology and theory still hold when the time-varying tuning parameter in Assumption (ref)(iii) is allowed to vary over entries in the spot volatility matrix estimation, which is expected to perform well in finite samples. For example, we set $\rho_{ij}(t)=\rho(t)(\widehat{\Sigma}_{ii,t}\widehat{\Sigma}_{jj,t})^{1/2}$ in the numerical studies and shrink the $(i,j)$-entry to zero if $\widehat{\Sigma}_{ij,t}\leq \rho(t)(\widehat{\Sigma}_{ii,t}\widehat{\Sigma}_{jj,t})^{1/2}$.

The following theorem gives the uniform convergence property (in the matrix spectral norm) for the spot volatility matrix estimator $\widehat{\boldsymbol{\Sigma}}_{t}$ under the uniform sparsity assumption.

theoremSuppose that Assumptions (ref) and (ref) are satisfied, and $\left\{{\boldsymbol{\Sigma}}_t:\ 0\leq t\leq T\right\}\in \mathcal{S}(q,\varpi(p), T)$. Then we have \begin{equation} \sup_{0\leq t\leq T}\left\Vert\widehat{\boldsymbol{\Sigma}}_{t}-{ \boldsymbol{\Sigma}}_t\right\Vert=O_P\left(\varpi(p) \zeta_{\Delta,p}^{1-q}\right), \end{equation} where $\varpi(p)$ is defined in ((ref)) and $\zeta_{\Delta,p}$ is defined in Assumption (ref)(iii).
remark(i) The first term of $\zeta_{\Delta,p}$ is $h^{\gamma}$, which is the bias rate due to application of the local smoothing technique. It is slower than the conventional $h^2$-rate since we do not assume existence of smooth derivatives of $\Sigma_{ij,t}$ (with respect to $t$). The second term of $\zeta_{\Delta,p}$ is square root of $\Delta h^{-1}\log (p\vee \Delta^{-1})$, a typical uniform asymptotic rate for the kernel estimation variance component. The uniform convergence rate in ((ref)) is also similar to those obtained by CL16 and CLL19 in the low-frequency data setting (disregarding the bias order). Note that the dimension $p$ affects the uniform convergence rate via $\varpi(p)$ and $\log (p\vee \Delta^{-1})$ and the estimation consistency may be achieved in the ultra-high dimensional setting when $p$ diverges at an exponential rate of $n=\lfloor T/\Delta\rfloor$. Treating $(nh)$ as the “effective" sample size in the local estimation procedure and disregarding the bias rate $h^{\gamma}$, the rate in ((ref)) is comparable to the optimal minimax rate in large covariance matrix estimation CZ12. (ii) If we further assume that $\Sigma_{ij,t}$ is deterministic with continuous second-order derivative with respect to $t$, and $K(\cdot)$ is symmetric, we may improve the kernel estimation bias order. In fact, following the proof of Theorem (ref), we may show that \begin{equation} \sup_{h\leq t\leq T-h}\left\Vert\widehat{\boldsymbol{\Sigma}}_{t}-{ \boldsymbol{\Sigma}}_t\right\Vert=O_P\left(\varpi(p) \zeta_{\Delta,p,\star}^{1-q}\right), \end{equation} where $\zeta_{\Delta,p,\star}=h^{2}+\left[\frac{\Delta\log (p\vee \Delta^{-1})}{h}\right]^{1/2}$. The above uniform consistency property only holds over the trimmed time interval $[h, T-h]$ due to the kernel boundary effect. In practice, however, it is often important to investigate the spot volatility structure near the boundary points. For example, when we consider one trading day as a time interval, it is particularly interesting to estimate the spot volatility matrix near the opening and closing times which are peak times in stock market trading. To address this issue, we may replace $K_h^\ast(t_k-t)$ in ((ref)) by a boundary kernel weight defined by \[ K_{h,t}^\star(t_k-t)=K_t\left(\frac{t_k-t}{h}\right)/\left[\Delta\sum_{l=1}^nK_t\left(\frac{t_l-t}{h}\right)\right], \] where $K_t(\cdot)$ is a boundary kernel satisfying $\int_{-t/h}^{(T-t)/h}uK_t(u)du=0$ (a key condition to improve the bias order near the boundary points). Examples of boundary kernels can be found in FG96 and LR07. With this adjustment in the kernel estimation, we can extend the uniform consistency result ((ref)) to the entire interval $[0,T]$.

Estimation with contaminated high-frequency data

\setcounter{equation}{0}

In practice, it is not uncommon that high-frequency financial data are contaminated by the market microstructure noise. The kernel estimation method proposed in Section (ref) would be biased if the noise is ignored in the estimation procedure. Consider the following additive noise structure:

equation[equation omitted — 162 chars of source]

where $t_k=k\Delta$, $k=1,\cdots,n$, ${\mathbf{Z}}_{t}=\left(Z_{1,t}, \cdots,Z_{p,t}\right)^{^\intercal}$ is a vector of observed asset prices at time $t$, and ${\boldsymbol{\xi}}_k=(\xi_{1,k},\cdots,\xi_{p,k})^{^\intercal} $ is a $p$-dimensional vector of noises with nonlinear heteroskedasticity, ${ \boldsymbol{\omega}}(\cdot)=\left[\omega_{ij}(\cdot)\right]_{p\times p}$ is a $p\times p$ matrix of deterministic functions, and ${\boldsymbol{\xi}} _k^\ast=\left(\xi_{1,k}^\ast,\cdots,\xi_{p,k}^\ast\right)^{^\intercal}$ independently follows a $p$-variate identical distribution. The noise structure defined in ((ref)) is similar to the setting considered in KL08 which also contains a nonlinear mean function and allows the existence of endogeneity for a single asset. Throughout this section, we assume that $\{{\boldsymbol{\xi}}_k^\ast\}$ is independent of the Brownian semimartingale $\{{\mathbf{X}}_t\}$.

Estimation of the spot volatility matrix

To account for the microstructure noise and produce consistent volatility matrix estimation, we apply the pre-averaging technique as the realised kernel estimate BHLS08 can be seen as a member of the pre-averaging estimation class whereas the two-scale estimate ZMA05 can be re-written as the realised kernel estimate with the Bartlett-type kernel (up to the first-order approximation). The pre-averaging method has been studied by JLMPV09, PV09 and CKP10 in estimating the integrated volatility for a single asset and is further extended by KWZ16 and DLX19 to the large high-frequency data setting. KK16 use a localised pre-averaging technique to estimate the spot volatility function for a single asset and derive the uniform convergence rate for the developed estimate. A similar technique is also used by XL02 to improve convergence of the nonparametric spectral density estimator for time series with general autocorrelation for low-frequency data.

We first pre-average the observed high-frequency data via a kernel filter, i.e.,

equation[equation omitted — 123 chars of source]

with $L_{b}^\dagger(t_k-\tau)=L_b\left(t_k-\tau\right)/\int_0^TL_b(s-\tau)ds$, where $L_b(u)=b^{-1}L(u/b)$, $L(\cdot)$ is a kernel function and $b$ is a bandwidth. Let $\Delta \widetilde{X}_{i,l}=\widetilde{X}_{i,\tau_l}-\widetilde{X}_{i,\tau_{l-1}}$, where $\widetilde{X}_{i,\tau_l}$ is the $i$-th component of $\widetilde{ \mathbf{X}}_{\tau_l}$ and $\tau_0,\tau_1,\cdots,\tau_N$ are the pseudo-sampling time points in the fixed interval $[0,T]$ with equal distance $\Delta_\ast=T/N$. Replacing $\Delta X_{i,k}$ by $\Delta\widetilde{X }_{i,l}$ in ((ref)), we estimate the spot co-volatility $\Sigma_{ij,t}$ by

equation[equation omitted — 143 chars of source]

where \[ K_{h}^\dagger(\tau_l-t)=K_h\left(\tau_l-t\right)/\left[\Delta_\ast\sum_{k=1}^NK_h\left(\tau_k-t\right)\right]. \] Furthermore, to obtain a stable spot volatility matrix estimate in finite samples when the dimension $p$ is large, as in ((ref)), we apply shrinkage to $\widetilde{\Sigma}_{ij,t}$, $1\leq i\neq j\leq p$, and subsequently construct

equation[equation omitted — 247 chars of source]

where $\rho_2(t)$ is another time-varying shrinkage parameter. We next give some conditions needed to derive the uniform consistency property of $ \widetilde{\boldsymbol{\Sigma}}_t$.

assumption(i)\ Let $\{{\boldsymbol{\xi}}_k^\ast\}$ be an independent and identically distributed (i.i.d.) sequence of p-dimensional random vectors. Assume that $\mathsf{E}(\xi_{i,k}^\ast)=0$ and \begin{equation*} \mathsf{E}\left[\exp\left(s|{\mathbf{u}}^{^\intercal}{\boldsymbol{\xi}} _{k}^\ast|\right)\right]\leq C_\xi<\infty,\ \ 0<s\leq s_0, \end{equation*} for any $p$-dimensional vector ${\mathbf{u}}$ satisfying $\Vert{\mathbf{u}} \Vert_2=1$. (ii)\ The deterministic functions $\omega_{ij}(\cdot)$ are bounded uniformly over $i,j\in\{1,\cdots,p\}$, and satisfy that \begin{equation*} \max_{1\leq i\leq p}\sup_{0\leq t\leq T}\sum_{j=1}^p\omega_{ij}^2(t)\leq C_\omega<\infty. \end{equation*}
assumption(i)\ The kernel function $L(\cdot)$ is Lipschitz continuous and has a compact support $[-1,1]$. In addition, $ \int_{-1}^1 L(u)du=1$. (ii)\ The bandwidth $b$ and the dimension $p$ satisfy that \begin{equation*} b\rightarrow0,\ \ \frac{\Delta^{2\iota-1}b}{\log (p\vee \Delta^{-1})} \rightarrow\infty,\ \ p\Delta\exp\{-s\Delta^{-\iota}\}\rightarrow0, \end{equation*} where $0<\iota<1/2$ and $0<s\leq s_0$. (iii) Let $\nu_{\Delta,p,N}=\sqrt{N\log(p\vee \Delta^{-1})}\left[b^{1/2}+(\Delta^{-1}b)^{-1/2}\right]\rightarrow0$ and the time-varying tuning parameter $\rho_2(t)$ be chosen as $\rho_2(t)=M(t)\left(\zeta_{N,p}^ \ast+\nu_{\Delta,p,N}\right)$, where $M(t)$ is defined as in Assumption (ref)(iii) and $\zeta_{N,p}^\ast$ is defined as $\zeta_{\Delta,p}$ with $N$ replacing $\Delta^{-1}$.
remarkWe allow nonlinear heteroskedasticity on the microstructure noise. The i.i.d. restriction on ${\boldsymbol{\xi}}_i^\ast$ may be weakened to some weak dependence conditions KWZ16, DLX19 at the cost of more lengthy proofs. The moment condition in Assumption (ref)(i) is weaker than the sub-Gaussian condition BL08, TWZ13 which is commonly used in large covariance matrix estimation when the dimension $p$ is ultra large. The boundedness condition on $\omega_{ij}(\cdot)$ in Assumption (ref)(ii) is similar to the local boundedness restriction in Assumption (ref)(i). Assumption (ref)(ii) imposes some mild restrictions on $b$ and $p$, which imply that there is a trade-off between them. When $\iota$ is larger, $p$ diverges at a faster exponential rate of $1/\Delta$ but the bandwidth condition becomes more restrictive. If $p$ is divergent at a polynomial rate of $1/\Delta$, we may let $\iota$ be sufficiently close to zero, and then the bandwidth condition reduces to the conventional one as in Assumption (ref)(ii). The condition $ \nu_{\Delta,p,N}\rightarrow0$ in Assumption (ref)(iii) is crucial to show that the error of the kernel filter $\widetilde{\boldsymbol{X}}_\tau$ tends to zero asymptotically, whereas the form of the time-varying shrinkage parameter $\rho_2(t)$ is relevant to the uniform convergence rate of $ \widetilde{\Sigma}_{ij,t}$ (see Proposition (ref)).
theoremSuppose that Assumptions (ref)(i)(ii), (ref)(i), (ref) and (ref) are satisfied, and Assumption (ref)(ii) holds with $\Delta^{-1}$ replaced by $N$. When $\left\{{\boldsymbol{\Sigma}}_t:\ 0\leq t\leq T\right\}\in \mathcal{S}(q,\varpi(p), T)$, we have \begin{equation} \sup_{0\leq t\leq T}\left\Vert\widetilde{\boldsymbol{\Sigma}}_{t}-{ \boldsymbol{\Sigma}}_t\right\Vert=O_P\left(\varpi(p) \left[ \zeta_{N,p}^\ast+\nu_{\Delta,p,N}\right]^{1-q}\right), \end{equation} where $\zeta_{N,p}^\ast$ and $\nu_{\Delta,p,N}$ are defined in Assumption (ref)(iii).
remarkThe uniform convergence rate in ((ref)) relies on $\varpi(p)$, $\zeta_{N,p}^\ast$ and $\nu_{\Delta,p,N}$. With the high-frequency data collected at pseudo time points with sampling frequency $ \Delta_\ast=T/N$, the rate $\zeta_{N,p}^\ast$ is comparable to $ \zeta_{\Delta,p}$ for the noise-free kernel estimator in Section (ref). The rate $\nu_{\Delta,p,N}$ is due to the error of the kernel filter $ \widetilde{\boldsymbol{X}}_\tau$ in the first step of the local pre-averaging estimation procedure. In particular, when $q=0$, $\varpi(p)$ is bounded, $b=\Delta^{1/4}$ and $h=N^{-\frac{1}{2\gamma+1}}$ with $ N=\Delta^{-\frac{2\gamma+1}{2\left(4\gamma+1\right)}}$, the uniform convergence rate in ((ref)) becomes $\Delta^{\frac{\gamma}{ 2(4\gamma+1)}}\sqrt{\log(p\vee \Delta^{-1})}$. Furthermore, if $\gamma=1/2$, the rate is simplified to $\Delta^{1/12}\sqrt{\log(p\vee \Delta^{-1})}$, comparable to those derived by ZB14 and KK16 in the univariate high-frequency data setting.

Estimation of the time-varying noise volatility matrix

It is often interesting to further explore the volatility structure of microstructure noise. CHLT21 estimate the constant covariance matrix for high-dimensional noise and derive the optimal convergence rates for the developed estimate. In the present paper, we consider the time-varying noise covariance matrix defined by

equation[equation omitted — 177 chars of source]

It is sensible to assume that $\left\{{\boldsymbol{\Omega}}(t):\ 0\leq t\leq T\right\}$ satisfies the uniform sparsity condition as in ((ref)). For each $1\leq i,j\leq p$, we estimate $\Omega_{ij}(t)$ by the kernel smoothing method:

equation[equation omitted — 146 chars of source]

where $h_1$ is a bandwidth, $\Delta Z_{i,t_{k}}=Z_{i,t_k}-Z_{i,t_{k-1}}$ and $K_{h_1}^\ast(t_k-t)$ is defined similarly to $K_{h}^\ast(t_k-t)$ in ((ref)) but with $h_1$ replacing $h$. As in ((ref)) and ((ref)), we again apply shrinkage to $ \widehat{\Omega}_{ij}(t)$, $1\leq i\neq j\leq p$, and construct

equation[equation omitted — 241 chars of source]

where $\rho_3(t)$ is a time-varying shrinkage parameter. To derive the uniform consistency property of $\widehat{\boldsymbol{\Omega}}(t)$, we need to impose stronger moment condition on ${\boldsymbol{\xi}}_k^\ast$ and smoothness restriction on $\Omega_{ij}(\cdot)$.

assumption(i)\ For any $p$-dimensional vector ${ \mathbf{u}}$ satisfying $\Vert{\mathbf{u}}\Vert_2=1$, $\mathsf{E}\left[\exp\left(s({\mathbf{u}}^{^\intercal}{\boldsymbol{\xi}} _{k}^\ast)^2\right)\right]\leq C_\xi^\star<\infty$, $0<s\leq s_0$. (ii)\ The time-varying function $\Omega_{ij}(t)$ satisfies that \begin{equation*} \max_{1\leq i,j\leq p}\left\vert \Omega_{ij}(t)-\Omega_{ij}(s)\right\vert\leq C_\Omega|t-s|^{\gamma_1}, \end{equation*} where $C_\Omega$ is a positive constant and $0<\gamma_1<1$. (iii)\ The bandwidth $h_1$ and the dimension $p$ satisfy that \begin{equation*} h_1\rightarrow0,\ \ \frac{\Delta^{2\iota_\star-1}h_1}{\log (p\vee \Delta^{-1})}\rightarrow\infty,\ \ p\Delta^{-1}\exp\{-s\Delta^{-\iota_\star}/C_\omega\}\rightarrow0, \end{equation*} where $0<\iota_\star<1/2$, $0<s\leq s_0$ and $C_\omega$ is defined in Assumption (ref)(ii).
remarkAssumption (ref)(i) strengthens the moment condition in Assumption (ref)(i) and is equivalent to the sub-Gaussian condition, see Assumption A1 in TWZ13. The smoothness condition in Assumption (ref)(ii) is similar to ((ref)), crucial to derive the asymptotic order of the kernel estimation bias. The restrictions on $h_1$ and $p$ in Assumption (ref)(iii) are similar to those in Assumption (ref)(ii), allowing $p$ to be divergent to infinity at an exponential rate of $1/\Delta$.

In the following theorem, we state the uniform consistency result for $ \widehat{\boldsymbol{\Omega}}(t)$ with convergence rate comparable to that in Theorem (ref).

theoremSuppose that Assumptions (ref), (ref)(i), (ref) and (ref) are satisfied, and Assumption (ref)(ii)(iii) holds when $ \rho_1(t)$, $\zeta_{\Delta,p}$ and $h$ are replaced by $\rho_3(t)$, $ \delta_{\Delta,p}$ and $h_1$, respectively, where $\delta_{\Delta,p}=h_1^{ \gamma_1}+\left[\frac{\Delta\log (p\vee \Delta^{-1})}{h_1}\right]^{1/2}$. If $\left\{{\boldsymbol{\Omega}}(t):\ 0\leq t\leq T\right\}\in \mathcal{S} (q,\varpi(p), T)$, we have \begin{equation} \sup_{0\leq t\leq T}\left\Vert\widehat{\boldsymbol{\Omega}}(t)-{ \boldsymbol{\Omega}}(t)\right\Vert=O_P\left(\varpi(p)\delta_{\Delta,p}^{1-q} \right). \end{equation}
remarkIf the bandwidth parameter $h_1$ in ((ref)) is the same as $h$ in ((ref)), we may find that the uniform convergence rate $O_P\left(\varpi(p)\delta_{\Delta,p}^{1-q}\right)$ would be the same as that in Theorem 1. Treating $(nh_1)$ as the “effective" sample size and disregarding the bias order, we may show that the uniform convergence rate in ((ref)) is comparable to the optimal minimax rate derived by CHLT21 for the constant noise covariance matrix estimation. Meanwhile, the kernel estimation bias order $h_1^{\gamma_1}$ may be improved by strengthening the smoothness condition on $\Omega_{ij}(\cdot)$ and adopting the boundary kernel weight as suggested in Remark (ref)(ii).

Estimation with observed factors

\setcounter{equation}{0}

The large spot volatility matrix estimation with the shrinkage technique developed in Sections (ref) and (ref) heavily relies on the uniform sparsity assumption ((ref)). However, the latter may be too restrictive in practice since the price processes of a large number of assets are often driven by some common factors such as the market factors, resulting in strong correlation among assets and failure of the sparsity condition. To address this problem, we next consider the nonparametric time-varying regression at high frequency:

equation[equation omitted — 97 chars of source]

where ${\boldsymbol\beta}(t)=\left[\beta_{1}(t),\cdots,\beta_p(t)\right]^{^\intercal}$ is a $p\times k$ matrix of time-varying betas (or factor loadings), ${\mathbf F}_t$ and ${\mathbf X}_t$ are $k$-variate and $p$-variate continuous semi-martingales defined by

equation[equation omitted — 216 chars of source]

respectively, ${\boldsymbol{\mu}}_t^F$ and ${\boldsymbol{\mu}}_t^X$ are drift vectors, ${\boldsymbol{\sigma}}_t^F=\left(\sigma_{ij,t}^F\right)_{k\times k}$, ${\boldsymbol{\sigma}}_t^X=\left(\sigma_{ij,t}^X\right)_{p\times p}$, ${\mathbf{W}}_t^F$ and ${\mathbf{W}}_t^X$ are $k$-dimensional and $p$-dimensional standard Brownian motions. For the time being, we assume that ${\mathbf Y}_t$ and ${\mathbf F}_t$ are observable and noise free but ${\mathbf X}_t$ is latent. Extension of the methodology and theory to the noise-contaminated high-frequency data will be considered later in this section.

Estimation of the constant betas via the ratio of realised covariance to realised variance is proposed by BS04, and extension to time-varying beta estimation has been studied by MZ06, RTT15 and AKX20, some of which allow jumps in the semi-martingale processes. The main interest of this section lies in estimating the large spot volatility structure ${\boldsymbol\Sigma}_t^Y$ of ${\mathbf Y}_t$. Letting ${\boldsymbol\Sigma}_t^F={\boldsymbol\sigma}_t^F\left({\boldsymbol\sigma}_t^F\right)^{^\intercal}$ and ${\boldsymbol\Sigma}_t^X={\boldsymbol\sigma}_t^X\left({\boldsymbol\sigma}_t^X\right)^{^\intercal}$, and assuming orthogonality between ${\mathbf X}_t$ and ${\mathbf F}_t$, see Assumption (ref)(iii) below, it follows from ((ref)) that

equation[equation omitted — 156 chars of source]

As in FLM11, FLM13, we impose the uniform sparsity restriction on ${\boldsymbol\Sigma}_t^X$ instead of ${\boldsymbol\Sigma}_t^Y$, i.e., $ \left\{{\boldsymbol{\Sigma}}_t^X:\ 0\leq t\leq T\right\}\in \mathcal{S}(q,\varpi(p), T)$. This is a reasonable assumption in practical applications as the asset prices, after removing the influence of systematic factors, are expected to be weakly correlated. FFX16 and DLX19 use a similar framework with constant betas to estimate large integrated volatility matrices.

Suppose that we observe ${\mathbf Y}_t$ and ${\mathbf F}_t$ at regular points: $t_k=k\Delta$, $k=1,\cdots,n$, as in Sections (ref) and (ref). Let ${\boldsymbol\Sigma}_t^{YF}$ be the spot covariance between ${\mathbf Y}_t$ and ${\mathbf F}_t$. We may use the kernel smoothing method as in ((ref)) to estimate ${\boldsymbol\Sigma}_t^{Y}$, ${\boldsymbol\Sigma}_t^{F}$ and ${\boldsymbol\Sigma}_t^{YF}$, i.e.,

eqnarray[eqnarray omitted — 409 chars of source]

where $\Delta{\mathbf Y}_k={\mathbf Y}_{t_k}-{\mathbf Y}_{t_{k-1}}$, $\Delta{\mathbf F}_k={\mathbf F}_{t_k}-{\mathbf F}_{t_{k-1}}$, and $K_h^\ast(t_k-t)$ is defined as in ((ref)). Consequently, the time-varying betas ${\boldsymbol\beta}(t)$ and the spot idiosyncratic volatility matrix ${\boldsymbol\Sigma}_t^X$ are estimated by

equation[equation omitted — 218 chars of source]

and

equation[equation omitted — 288 chars of source]

With the uniform sparsity condition, it is sensible to further apply shrinkage to $\widehat\Sigma_{ij,t}^X$, i.e.,

equation[equation omitted — 253 chars of source]

where $\rho_4(t)$ is a time-varying shrinkage parameter. We finally estimate ${\boldsymbol\Sigma}_t^Y$ as

equation[equation omitted — 388 chars of source]

We need the following assumption to derive the uniform convergence property for $\widehat{\boldsymbol\Sigma}_t^{X,s}$ and $\widehat{\boldsymbol\Sigma}_t^{Y,s}$.

assumption{\em (i) Assumption (ref) is satisfied for $\{X_t\}$ defined in ((ref)) (with minor notational changes).} {\em (ii) Let $\{{\boldsymbol{\mu}}_t^F\}$, $\{{\boldsymbol{\sigma}}_t^F\}$ and $\{{\boldsymbol{\Sigma}}_t^F\}$ satisfy the boundedness and smoothing conditions as in Assumption (ref).} {\em (iii) For any $1\leq i\leq p$ and $1\leq j\leq k$, $\left[X_{it}, F_{jt}\right]=0$ for any $t\in[0,T]$, where $X_{i,t}$ is the $i$-th element of ${\mathbf X}_t$, $F_{j,t}$ is the $j$-th element of ${\mathbf F}_t$, and $[\cdot,\cdot]$ denotes the quadratic covariation. } {\em (iv) The time-varying beta function $\beta_i(\cdot)$ satisfies that \[ \max_{1\leq i\leq p}\sup_{0\leq t\leq T}\left\Vert \beta_i(t)\right\Vert_2\leq C_\beta<\infty,\ \ \max_{1\leq i\leq p}\left\Vert \beta_i(t)-\beta_i(s)\right\Vert_2\leq C_\beta|t-s|^{\gamma}, \] where $\gamma$ is the same as that in Assumption (ref)(ii). In addition, there exists a positive definite matrix ${\boldsymbol\Sigma}_\beta(t)$ (with uniformly bounded eigenvalues) such that} \begin{equation} \sup_{0\leq t\leq T}\left\Vert \frac{1}{p}{\boldsymbol\beta}(t)^{^\intercal}{\boldsymbol\beta}(t)-{\boldsymbol\Sigma}_\beta(t)\right\Vert=o(1). \end{equation}
remarkThe uniform boundedness and smoothness conditions imposed on the drift and spot volatility functions of ${\mathbf X}_t$ and ${\mathbf F}_t$ in Assumption (ref)(i)(ii) are the same as those in Assumption (ref). This is crucial to ensure that the uniform convergence rates of $\widehat{\boldsymbol\Sigma}_t^Y$, $\widehat{\boldsymbol\Sigma}_t^F$ and $\widehat{\boldsymbol\Sigma}_t^{YF}$ (in the max norm) derived in Proposition (ref) are the same as that in Proposition (ref). The orthogonality condition in Assumption (ref)(iii) is commonly used to consistently estimate the time-varying factor model FFX16, DLX19. Assumption (ref)(iv) is a rather mild restriction on time-varying betas and may be strengthened to improve the estimation bias order, see the discussion in Remark (ref)(ii). The condition ((ref)) indicates that all the factors are pervasive.

We next present the convergence property of $\widehat{\boldsymbol\Sigma}_t^{X,s}$ and $\widehat{\boldsymbol\Sigma}_t^{Y,s}$ defined in ((ref)) and ((ref)), respectively. Due to the nonparametric factor regression model structure ((ref)), the largest $k$ eigenvalues of ${\boldsymbol\Sigma}_t^Y$ are spiked, diverging at a rate of $p$. Hence, ${\boldsymbol\Sigma}_t^Y$ cannot be consistently estimated in the absolute term. To address this problem, as in FLM11, FLM13, we measure the spiked volatility matrix estimate in the following relative error: \[\left\Vert \widehat{\boldsymbol\Sigma}_t^{Y,s}-{\boldsymbol\Sigma}_t^Y\right\Vert_{{\boldsymbol\Sigma}_t^Y}=\frac{1}{\sqrt{p}}\left\Vert \left({\boldsymbol\Sigma}_t^Y\right)^{-1/2}\left(\widehat{\boldsymbol\Sigma}_t^{Y,s}-{\boldsymbol\Sigma}_t^Y\right) \left({\boldsymbol\Sigma}_t^Y\right)^{-1/2} \right\Vert_F,\] where the normalisation factor $p^{-1/2}$ is used to guarantee that $\left\Vert{\boldsymbol\Sigma}_t^Y\right\Vert_{{\boldsymbol\Sigma}_t^Y}=1$.

theoremSuppose that Assumptions (ref)(i)(ii) and (ref) are satisfied, and Assumption (ref)(iii) holds with $\rho_1(t)$ replaced by $\rho_4(t)$. When $\left\{{\boldsymbol{\Sigma}}_t^X:\ 0\leq t\leq T\right\}\in \mathcal{S}(q,\varpi(p), T)$, we have \begin{equation} \sup_{0\leq t\leq T}\left\Vert\widehat{\boldsymbol{\Sigma}}_{t}^{X,s}-{\boldsymbol{\Sigma}}_t^X\right\Vert=O_P\left(\varpi(p)\zeta_{\Delta,p}^{1-q}\right), \end{equation} where $\varpi(p)$ is defined in ((ref)) and $\zeta_{\Delta,p}$ is defined in Assumption (ref)(iii); and \begin{equation} \sup_{0\leq t\leq T}\left\Vert \widehat{\boldsymbol\Sigma}_t^{Y,s}-{\boldsymbol\Sigma}_t^Y\right\Vert_{{\boldsymbol\Sigma}_t^Y}=O_P\left(p^{1/2}\zeta_{\Delta,p}^2+\varpi(p)\zeta_{\Delta,p}^{1-q}\right). \end{equation}
remarkAlthough ${\mathbf X}_t$ is latent in model ((ref)), the uniform convergence rate for $\widehat{\boldsymbol{\Sigma}}_{t}^{X,s}$ in ((ref)) is the same as that in Theorem (ref) when ${\mathbf X}_t$ is observable. Treating $(nh)$ as the effective sample size in kernel estimation and disregarding the bias order in $\zeta_{\Delta,p}$, the uniform convergence rate for $\widehat{\boldsymbol{\Sigma}}_{t}^{Y,s}$ in ((ref)) is comparable to the convergence rates derived by FLM11 in low frequency and FFX16 in high frequency. To guarantee uniform consistency in the relative matrix estimation error, we have to further assume that $p\zeta_{\Delta,p}^4=o(1)$, limiting the divergence rate of the asset number, i.e., $p$ can only diverge at a polynomial rate of $n=\lfloor T/\Delta\rfloor$.

We next modify the above methodology and theory to accommodate microstructure noise in the asset prices and factors. Assume that

equation[equation omitted — 234 chars of source]

where ${\boldsymbol{\omega}}_Y(\cdot)$ and ${\boldsymbol{\omega}}_F(\cdot)$ are matrices of deterministic functions similar to ${\boldsymbol{\omega}}(\cdot)$, and $\{{\boldsymbol{\xi}}_{Y,k}^\ast\}$ and $\{{\boldsymbol{\xi}}_{F,k}^\ast\}$ are i.i.d. sequences of random vectors similar to $\{{\boldsymbol{\xi}}_{k}^\ast\}$. Since both ${\mathbf Y}_t$ and ${\mathbf F}_t$ are latent, we need to first adopt the kernel pre-averaging technique proposed in Section (ref) to obtain the approximation of ${\mathbf Y}_t$ and ${\mathbf F}_t$, and then apply the kernel smoothing and generalised shrinkage as in ((ref))--((ref)). This results in a three-stage estimation procedure which we describe as follows.

enumerate• As in ((ref)), we pre-average the noise-contaminated ${\mathbf{Z}}_{Y,t_k}$ and ${\mathbf{Z}}_{F,t_k}$ via the kernel filter: \begin{equation} \widetilde{\mathbf{Y}}_\tau=\frac{T}{n}\sum_{k=1}^n L_b^\dagger(t_k-\tau){\mathbf{Z}}_{Y,t_k},\quad \widetilde{\mathbf{F}}_\tau=\frac{T}{n}\sum_{k=1}^n L_b^\dagger(t_k-\tau){\mathbf{Z}}_{F,t_k}, \end{equation} where $L_b^\dagger(t_k-\tau)$ is defined as in ((ref)) and we consider $\tau$ as the pseudo-sampling time points: $\tau_l=l\Delta_\ast$, $l=0,1,\cdots, N=\lfloor T/\Delta_\ast\rfloor$. • With $\widetilde{\mathbf Y}_{\tau_l}$ and $\widetilde{\mathbf F}_{\tau_l}$, $l=1,\cdots,N$, we estimate ${\boldsymbol\Sigma}_t^Y, {\boldsymbol\Sigma}_t^F$ and ${\boldsymbol\Sigma}_t^{YF}$ by the kernel smoothing as in ((ref))--((ref)): \begin{eqnarray} &&\widetilde{\boldsymbol\Sigma}_t^Y=\sum_{l=1}^N K_h^\dagger(\tau_l-t)\Delta \widetilde{\mathbf Y}_l\Delta \widetilde{\mathbf Y}_l^{^\intercal},\notag\\ &&\widetilde{\boldsymbol\Sigma}_t^F=\sum_{l=1}^N K_h^\dagger(\tau_l-t)\Delta \widetilde{\mathbf F}_l\Delta \widetilde{\mathbf F}_l^{^\intercal},\notag\\ &&\widetilde{\boldsymbol\Sigma}_t^{YF}=\sum_{l=1}^N K_h^\dagger(\tau_l-t)\Delta \widetilde{\mathbf Y}_l\Delta \widetilde{\mathbf F}_l^{^\intercal},\notag \end{eqnarray} where $K_h^\dagger(\tau_l-t)$ is defined as in ((ref)), $\Delta \widetilde{\mathbf Y}_l=\widetilde{\mathbf Y}_{\tau_l}-\widetilde{\mathbf Y}_{\tau_{l-1}}$ and $\Delta \widetilde{\mathbf F}_l=\widetilde{\mathbf F}_{\tau_l}-\widetilde{\mathbf F}_{\tau_{l-1}}$. Furthermore, estimate ${\boldsymbol\beta}(t)$ and ${\boldsymbol\Sigma}_t^X$ by \[ \widetilde{\boldsymbol\beta}(t)=\widetilde{\boldsymbol\Sigma}_t^{YF}\left(\widetilde{\boldsymbol\Sigma}_t^{F}\right)^{-1},\quad \widetilde{\boldsymbol\Sigma}_t^X=\left(\widetilde\Sigma_{ij,t}^X\right)_{p\times p}=\widetilde{\boldsymbol\Sigma}_t^Y-\widetilde{\boldsymbol\Sigma}_t^{YF}\left(\widetilde{\boldsymbol\Sigma}_t^{F}\right)^{-1}\left(\widetilde{\boldsymbol\Sigma}_t^{YF}\right)^{^\intercal}. \] • Apply the generalised shrinkage to $\widetilde\Sigma_{ij,t}^X$, i.e., \[ \widetilde{\boldsymbol{\Sigma}}_{t}^{X,s}=\left(\widetilde\Sigma_{ij,t}^{X,s}\right)_{p \times p}\ \ \mathrm{with}\ \ \widetilde\Sigma_{ij,t}^{X,s}=s_{\rho_5(t)}(\widetilde\Sigma_{ij,t}^X)I(i\neq j)+\widetilde\Sigma_{ii,t}^X I(i=j), \] where $\rho_5(t)$ is the shrinkage parameter, and then estimate ${\boldsymbol\Sigma}_t^Y$ by \[ \widetilde{\boldsymbol\Sigma}_t^{Y, s}=\widetilde{\boldsymbol\beta}(t)\widetilde{\boldsymbol\Sigma}_t^F\widetilde{\boldsymbol\beta}(t)^{^\intercal}+\widetilde{\boldsymbol\Sigma}_t^{X,s}=\widetilde{\boldsymbol\Sigma}_t^{YF}\left(\widetilde{\boldsymbol\Sigma}_t^{F}\right)^{-1}\left(\widetilde{\boldsymbol\Sigma}_t^{YF}\right)^{^\intercal}+\widetilde{\boldsymbol\Sigma}_t^{X,s}. \]

As shown in Theorem (ref), the existence of microstructure noises slows down the uniform convergence rates. Following the proof of Lemma B.1 in Appendix B, we may show that \[ \max_{0\leq l\leq N}\left\vert \widetilde{\mathbf{Y}}_{\tau_l}-{\mathbf{Y}}_{\tau_l}\right\vert_{\max}+\max_{0\leq l\leq N}\left\vert \widetilde{\mathbf{F}}_{\tau_l}-{\mathbf{F}}_{\tau_l}\right\vert_{\max}=O_P\left(\nu_{\Delta,p,N}\right), \] where $\vert\cdot\vert_{\max}$ denotes the $L_\infty$-norm of a vector, and $\nu_{\Delta,p,N}$ is defined in Assumption (ref)(iii). Modifying Proposition (ref) and the proof of Theorem (ref) in Appendix A, we can prove that ((ref)) and ((ref)) hold but with $\zeta_{\Delta,p}$ replaced by $\zeta_{N,p}^\ast+\nu_{\Delta,p,N}$ defined in Assumption (ref)(iii), i.e.,

eqnarray[eqnarray omitted — 453 chars of source]

Monte-Carlo Study

\setcounter{equation}{0}

In this section, we report the Monte-Carlo simulation studies to assess the numerical performance of the proposed large spot volatility matrix and time-varying noise volatility matrix estimation methods under the sparsity condition and the factor-based spot volatility matrix estimation. Here we only consider the synchronous high-frequency data. Additional simulation results for asynchronous high-frequency data are provided in the supplement.

Simulation for sparse volatility matrix estimation

{\em 5.1.1.\ \ Simulation setup}

We generate the noise-contaminated high-frequency data according to model ((ref)), where ${\boldsymbol{\omega}}(t)$ is taken as the Cholesky decomposition of the noise covariance matrix ${\boldsymbol\Omega}(t)=\left[\Omega _{ij}(t)\right] _{p\times p}$, ${\boldsymbol{\xi }}_{k}^{\ast }=\left(\xi _{1,k}^{\ast },\cdots ,\xi _{p,k}^{\ast }\right)^{^{\intercal }}$ is an independent $p$-dimensional random vector of cross-sectionally independent standard normal random variables, the latent return process $\mathbf{X}_{t}$ of $p$ assets is generated from the following drift-free model:

equation[equation omitted — 111 chars of source]

$\mathbf{W}_{t}^{X}=\left( W_{1,t}^{X},\cdots ,W_{p,t}^{X}\right)^{^{\intercal }}$ is a standard $p$-dimensional Brownian motion, and ${\boldsymbol{\sigma }}_{t}$ is chosen as the Cholesky decomposition of the spot covariance matrix ${\boldsymbol{\Sigma }}_{t}=\left(\Sigma_{ij,t}\right) _{p\times p}$. In the simulation, we consider the volatility matrix estimation over the time interval of a full trading day, and set the sampling interval to be $15$ seconds, i.e., $\Delta =1/(252\times 6.5\times 60\times 4)$, to generate synchronous data. We consider three structures in ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$: “banding", “block-diagonal", and “exponentially decaying". Following WZ10, we generate the diagonal elements of ${\boldsymbol\Sigma}_{t}$ from the following geometric Ornstein-Uhlenbeck model BS02:

equation*[equation* omitted — 178 chars of source]

where $\mathbf{W}_{t}^\ast=\left(W_{1,t}^\ast,\cdots,W_{p,t}^\ast\right)^{^\intercal}$ is a standard $p$-dimensional Brownian motion independent of $\mathbf{W}_{t}^{X}$, and $\iota_{i}$ is a random number generated uniformly between $-0.62$ and $-0.30$, reflecting the leverage effects. The diagonal elements of ${\boldsymbol\Omega}(t)$ are defined as daily cyclical deterministic functions of time: \[ \Omega _{ii}\left( t\right) =c_{i}\left\{ \frac{1}{2}\left[ \cos \left(2\pi t/T\right) +1\right] \times \left( \overline\omega -\underline\omega\right) +\underline\omega\right\}, \] where $\overline\omega=1$ and $\underline\omega=0.1$ reflect the observation by KL08 that the noise level is high at both the opening and the closing times of a trading day and is low in the middle of the day, and the scalar $c_{i}$ controls the noise ratio for each asset which is chosen to match the highest noise ratio considered by WZ10. As in BS02, BS04, we define a continuous-time stochastic process $\kappa^{\Sigma}_t$ by

eqnarray[eqnarray omitted — 260 chars of source]

where $W_{t}^\diamond$ is a standard univariate Brownian motion independent of $\mathbf{W}_{t}^{X}$ and $\mathbf{W}_{t}^\ast$. Let \[ \kappa^\Omega_t =\frac{\overline\kappa-\underline\kappa}{2}\left[ \cos \left( 2\pi t/T\right) +1\right] +\underline\kappa, \] where $\overline\kappa=0.5$ and $\underline\kappa=-0.5$. We will use $\kappa^{\Sigma}_t$ and $\kappa^{\Omega}_t$ to define the off-diagonal elements in ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$, respectively, which are specified as follows.

itemize• Banding structure for ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$: The off-diagonal elements are defined by \[ \Sigma _{ij,t} =\left(\kappa^\Sigma_{t}\right)^{\vert i-j\vert }\sqrt{\Sigma_{ii,t}\Sigma_{jj,t}}\cdot I\left( \left\vert i-j\right\vert \leq 2\right),\] and \[\Omega_{ij}(t) =\left(\kappa^\Omega_t\right) ^{\vert i-j\vert }\sqrt{\Omega _{ii}( t) \Omega _{jj}(t) }\cdot I\left( \left\vert i-j\right\vert \leq 2\right), \] for $1\leq i\neq j\leq p$. • Block-diagonal structure for ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$: The off-diagonal elements are defined by \[ \Sigma _{ij,t} =\left(\kappa^\Sigma_{t}\right)^{\vert i-j\vert }\sqrt{\Sigma_{ii,t}\Sigma_{jj,t}}\cdot I\left( (i,j)\in {\cal B}\right), \] \[ \Omega _{ij}(t)=\left(\kappa^\Omega_t\right) ^{\vert i-j\vert }\sqrt{\Omega _{ii}( t) \Omega _{jj}(t) }\cdot I\left((i,j)\in {\cal B}\right) , \] for $1\leq i\neq j\leq p$, where ${\cal B}$ is a collection of row and column indices $(i,j)$ located within our randomly generated diagonal blocks \footnote{As in DLX19, to generate blocks with random sizes, we fix the largest block size at $20$ when $p=200$ and randomly generate the sizes of the remaining blocks from a random integer uniformly picked between $5$ and $20$. When $p=500$, the largest size is $40$, and the random integer is uniformly picked between $10$ and $40$. Block sizes are randomly generated but fixed across all Monte Carlo repetitions.}. • Exponentially decaying structure for ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$: The off-diagonal elements are defined by \begin{equation} \Sigma _{ij,t} =\left(\kappa^\Sigma_{t}\right)^{\vert i-j\vert }\sqrt{\Sigma_{ii,t}\Sigma_{jj,t}},\ \ \Omega_{ij}(t) =\left(\kappa^\Omega_t\right) ^{\vert i-j\vert }\sqrt{\Omega _{ii}( t) \Omega _{jj}(t) },\ \ 1\leq i\neq j\leq p. \end{equation}

It is clear that the sparsity condition is not satisfied when the off-diagonal elements of ${\boldsymbol{\Sigma}}_{t}$ and ${\boldsymbol\Omega}(t)$ are exponentially decaying as in ((ref)). The number of assets $p$ is set as $p=200$ and $500$ and the replication number is $R=200$

{\em 5.1.2.\ \ Volatility matrix estimation}

In the simulation studies, we consider the following volatility matrix estimates.

itemize• Noise-free spot volatility matrix estimate $\widehat{\boldsymbol\Sigma }_{t}$. This infeasible estimate serves as a benchmark in comparing the numerical performance of various estimation methods. As in Section (ref), we apply the kernel smoothing method to estimate $\Sigma_{ij,t}$ by directly using the latent return process ${\mathbf X}_t$, where the bandwidth is determined by the leave-one-out cross validation. We apply four shrinkage methods to $\widehat{\Sigma}_{ij,t}$ for $i\neq j$: hard thresholding (Hard), soft thresholding (Soft), adaptive LASSO (AL) and smoothly clipped absolute deviation (SCAD). For comparison, we also compute the naive estimate without applying any regularisation technique. • Noise-contaminated spot volatility matrix estimate $\widetilde{\boldsymbol\Sigma}_{t}$. We combine the kernel smoothing with pre-averaging in Section (ref) to estimate $\Sigma_{ij,t}$ by using the noise-contaminated process ${\mathbf Z}_t$. As in the noise-free estimation, we apply four shrinkage methods to $\widetilde{\Sigma}_{ij,t}$ for $i\neq j$ and also compute the naive estimate without applying the shrinkage. • Time-varying noise volatility matrix estimate $\widehat{\boldsymbol\Omega}(t)$. We combine the kernel smoothing with four shrinkage techniques in the estimation as in Section (ref) and also the naive estimate without shrinkage.

The choice of tuning parameter in shrinkage is similar to that in DLX19. For example, in the noise-free spot volatility estimate, we set the tuning parameter as $\rho_{ij}(t)=\rho(t)(\widehat{\Sigma}_{ii,t}\widehat{\Sigma}_{jj,t})^{1/2}$ where $\rho(t)$ is chosen as the minimum value among the grid of values on $[0,1]$ such that the shrinkage estimate of the spot volatility matrix is positive definite. To evaluate the estimation performance of $\widehat{\boldsymbol\Sigma }_{t}$, we consider $21$ equidistant time points on $[0,T]$ and compute the following Mean Frobenius Loss (MFL) and Mean Spectral Loss (MSL) over $200$ repetitions:

eqnarray*[eqnarray* omitted — 393 chars of source]

where $t_{j}$, $j=1,2,\cdots,21$ are the equidistant time points on the interval $[0,T]$, and $\boldsymbol{\widehat{\Sigma}}_{t_{j}}^{(m)}$ and $\boldsymbol{\Sigma }_{t_{j}}^{(m)}$ are respectively the estimated and true spot volatility matrices at $t_{j}$ for the $m$-th repetition. The “MFL" and “MSL" can be similarly defined for $\widetilde{\boldsymbol\Sigma}_{t}$ and $\widehat{\boldsymbol\Omega}(t)$.

{\em 5.1.3.\ \ Simulation results}

Table 1 reports the simulation results when the dimension is $p=200$. The three panels in the table (from top to bottom) report the results where the true volatility matrix structures are banding, block-diagonal, and exponentially decaying, respectively. In each panel, the MFL results are reported on the left, whereas the MSL results are on the right. The first two rows of each panel contain the MFL and MSL results for the spot volatility matrix estimation whereas the third row contains the results for the time-varying noise volatility matrix estimation.

For the noise-free estimate $\widehat{\boldsymbol\Sigma }_{t}$, when the volatility matrix structure is banding, the performance of the four shrinkage estimators are substantially better than that of the naive estimate (without any shrinkage). In particular, the results of the soft thresholding, adaptive LASSO and SCAD are very similar and their MFL and MSL values are approximately one third of those of the naive estimator. Meanwhile, the performance of the hard thresholding is less accurate (despite the much stronger level of shrinking used), but is still much better than the naive estimate. These results show that the shrinkage technique is an effective tool in estimating the sparse volatility matrix. Similar results are obtained for the noise-contaminated estimate $\widetilde{\boldsymbol\Sigma}_{t}$. Unsurprisingly, due to the microstructure noise, the MFL and MSL values of the local pre-averaging estimates are noticeably higher than the corresponding values of the noise-free estimates. We next turn the attention to the time-varying noise volatility matrix estimate $\widehat{\boldsymbol\Omega}(t)$. As in the spot volatility matrix estimation, the naive method again produces the highest MFL and MSL values. The performance of the four shrinkage estimators are similar with the adaptive LASSO and SCAD being slightly better than the hard and soft thresholding. The simulation results for the block-diagonal and exponentially decaying covariance matrix settings, reported in the middle and bottom panels of Table 1, are fairly close to those for the banding setting. Overall, the results in Table 1 show that the shrinkage methods perform well not only in the sparse covariance matrix settings but also in the non-sparse one (i.e., the exponentially decaying setting).

center[center omitted — 4,165 chars of source]

{\scriptsize The selected bandwidths are $h^\ast=90$ for $\widehat{\boldsymbol\Sigma}_{t}$, $h^\ast=90$ and $b^\ast=4$ for $\widetilde{\boldsymbol\Sigma}_{t}$, and $h_1^{\ast}=90$ for $\widehat{\boldsymbol\Omega}(t)$, where $h^{\ast}=h/\Delta$, $b^{\ast}=b/\Delta$, and $h_{1}^{\ast}=h_1/\Delta$.}

The simulation results when the dimension is $p=500$ are reported in Table 2. Overall the results are very similar to those in Table 1, so we omit the detailed discussion and comparison to save the space.

center[center omitted — 3,953 chars of source]

{\scriptsize The selected bandwidths are $h^{\ast }=240$ for $\widehat{\boldsymbol\Sigma}_{t}$, $h^{\ast }=240$, $b^{\ast }=4$ for $\widetilde{\boldsymbol\Sigma}_{t}$ and $h_1^{\ast }=240$ for $\widehat{\boldsymbol\Omega}(t)$, where $h^{\ast}=h/\Delta$, $b^{\ast}=b/\Delta$, and $h_{1}^{\ast}=h_1/\Delta$.}

Simulation for factor-based spot volatility matrix estimation

{\em 5.2.1.\ \ Simulation setup}

We generate ${\mathbf Y}_t$ via ((ref)), where the $p$-dimensional idiosyncratic returns follow the dynamics of $d\mathbf{X}_{t}$ defined in ((ref)). In this simulation, we only consider $p=500$. As in AKX20, we adopt a three-factor model, where the factors $\mathbf{F}_{t}=\left(F_{1,t},F_{2,t},F_{3,t}\right) ^{^\intercal}$ are generated by

equation*[equation* omitted — 524 chars of source]

The factor volatilities are driven by

equation*[equation* omitted — 183 chars of source]

where ${\sf E}[ dW_{k,t}^{F}d\widetilde{W}_{k,t}] =\rho_{k}dt$, allowing for potential leverage effects in the factor dynamics. Both $W_{k,t}^{F}$ and $\widetilde{W}_{k,t}$ are standard univariate Brownian motions. In the simulation, we set $\left( \tilde{\kappa}_{1},\tilde{\kappa}_{2},\tilde{ \kappa}_{3}\right) =\left( 3,4,5\right) $, $\left( \tilde{\alpha}_{1},\tilde{ \alpha}_{2},\tilde{\alpha}_{3}\right) =\left( 0.09,0.04,0.06\right) $, $ \left( \tilde{\nu}_{1},\tilde{\nu}_{2},\tilde{\nu}_{3}\right) =\left( 0.3,0.4,0.3\right) $, $\left( \mu _{1}^{F},\mu _{2}^{F},\mu _{3}^{F}\right) =\left( 0.05,0.03,0.02\right) $, $\left( \rho _{1},\rho _{2},\rho _{3}\right) =\left( -0.6,-0.4,-0.25\right) $ and $\left( \rho _{12},\rho _{13},\rho _{23}\right) =\left( 0.05,0.10,0.15\right) $.

We consider the following three cases for generating the time-varying beta processes: $\beta_{i}(t)=\left[\beta_{i,1}(t), \beta_{i,2}(t), \beta_{i,3}(t)\right]^{^\intercal}$, $i=1,\cdots,p$.

itemize• Constant betas. The factor loadings are constants over time, i.e., $\beta _{i,l}(t)=\beta _{i,l}$, $i=1,\cdots,p$ and $l=1,2,3$. For each $i$, we set $\beta _{i,1}\sim {\sf U}\left( 0.25,2.25\right)$ and $\beta _{i,2},\beta _{i,3}\sim {\sf U}\left(-0.5,0.5\right)$. • Deterministic time-varying betas. Consider the following deterministic function: \begin{equation*} \beta _{i,l}\left( t\right) =\frac{1}{2}\left[ \cos \left( \pi ( t-\omega _{i,l})/T \right) +1\right] \times \left( \overline{\beta} _{i,l}-\beta_{i,l}\right) +\beta _{i,l},\ \ i=1,\cdots,p,\ \ l=1,2,3, \end{equation*} where $\omega _{i,1},\omega_{i,2},\omega _{i,3}\sim {\sf U}(0, 2T)$, $(\underline{\beta} _{i,1},\overline{\beta}_{i,1})$ is a pair of two random numbers from ${\sf U}\left( 0.25,2.25\right)$ whereas $(\underline{\beta}_{i,2},\overline{\beta}_{i,2})$ and $(\underline{\beta}_{i,3},\overline{\beta}_{i,3})$ are pairs of random numbers from ${\sf U}\left( -0.5,0.5\right)$. • Stochastic time-varying betas. As in AKX20, we consider the following diffusion process: \[ d\beta _{i,l}(t)=\kappa _{i,l}^{\beta }\left( \alpha _{i,l}^{\beta }-\beta _{i,l}(t)\right) dt+\upsilon _{i,l}^{\beta }dW_{i,l,t}^{\beta},\ \ i=1,\cdots,p,\ \ l=1,2,3, \] where $W_{i,l,t}^{\beta}$ are standard Brownian motions independently over $i$ and $l$, $\kappa _{i,1}^{\beta},\kappa_{i,2}^{\beta},\kappa _{i,3}^{\beta}\sim {\sf U}(1,3)$, $\alpha _{i,1}^{\beta }\sim{\sf U}( 0.25,2.25)$, $\alpha _{i,2}^{\beta},\alpha_{i,3}^{\beta }\sim {\sf U}(-0.5, 0.5)$ and $\upsilon _{i,1}^{\beta },\upsilon _{i,2}^{\beta},\upsilon _{i,3}^{\beta } \sim {\sf U}(2,4)$.

{\em 5.2.2.\ \ Simulation Results}

The spot idiosyncratic volatility matrix is estimated via ((ref)). For ease of comparison, we use exactly the same bandwidth as in our first experiment. The results for the noise-free and noise-contaminated spot idiosyncratic volatility matrix estimates $\widehat{\boldsymbol{\Sigma }}_{t}$ and $\widetilde{\boldsymbol{\Sigma }}_{t}$ measured by MFL and MSL are reported in Table 3, which reveal some desirable observations. Firstly, we note that our estimation results in terms of MFL and MSL are almost identical across different types of dynamics of factor loadings, indicating that the developed estimation procedure is robust in finite samples to different assumption of the factor loading dynamics as long as they satisfy our smooth restriction, see Assumption (ref)(iv). Secondly, the MFL and MSL values are similar to those reported in Table 2 which were obtained based on data generating model without common factors. This means that the proposed nonparametric time-varying high frequency regression can effectively remove common factors, resulting in accurate estimation of the spot idiosyncratic volatility matrix.

The factor-based spot volatility matrix of ${\mathbf Y}_t$ is estimated via ((ref)). As discussed in Section (ref), we measure the accuracy of the spiked volatility matrix estimate by the relative error defined above Theorem (ref), i.e., consider the following Mean Relative Loss (MRL): \[ \text{MRL}=\frac{1}{200}\sum_{m=1}^{200}\left( \frac{1}{21} \sum_{j=1}^{21}\left\Vert \widehat{\boldsymbol{\Sigma }}_{t_{j}}^{Y,(m)}- \boldsymbol{\Sigma }_{t_{j}}^{Y,(m)}\right\Vert _{\boldsymbol{\Sigma } _{t_{j}}^{Y,(m)}}\right) . \] The relevant results are reported in Table 4, where $\widehat{\boldsymbol\Sigma }_{t}^{Y}$ and $\widetilde{\boldsymbol\Sigma }_{t}^{Y}$ denote the noise-free and noise-contaminated factor-based spot volatility matrix estimates, respectively. We can see that the performance of the shrinkage estimates is substantially better than that of the naive estimate. Unsurprisingly, due to the presence of microstructure noise, the MRL results of $\widetilde{\boldsymbol\Sigma }_{t}^{Y}$ are much higher than those of $\widehat{\boldsymbol\Sigma }_{t}^{Y}$. As in Table 3, our proposed estimation is robust to different factor loading dynamics.

landscape\begin{center} Table 3: Estimation results for the spot idiosyncratic volatility matrices { \begin{tabular}{llllllllclllll} \hline\hline & & & \multicolumn{11}{c}{\textquotedblleft Banding"} \\ \cline{4-14} $\beta $ Dynamics & & & \multicolumn{5}{c}{Frobenius Norm} & & \multicolumn{5}{c}{Spectral Norm} \\ \cline{4-8}\cline{10-14} & & & \ Naive & \ Hard & \ Soft & \ AL & \ SCAD & \multicolumn{1}{l} & \ Naive & \ Hard & \ Soft & \ AL & \ SCAD \\ \cline{4-8}\cline{10-14} Constant & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9037} & \multicolumn{1}{r}{4.2461} & \multicolumn{1}{r}{5.2485} & \multicolumn{1}{r}{ 4.9880} & \multicolumn{1}{r}{3.9910} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{3.8887} & \multicolumn{1}{r}{0.6359} & \multicolumn{1}{r}{ 0.7291} & \multicolumn{1}{r}{0.7154} & \multicolumn{1}{r}{0.5720} \\ & $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.6646 } & \multicolumn{1}{r}{19.3752} & \multicolumn{1}{r}{18.3036} & \multicolumn{1}{r}{17.7552} & \multicolumn{1}{r}{18.1388} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{11.1910} & \multicolumn{1}{r}{ 2.3576} & \multicolumn{1}{r}{2.2653} & \multicolumn{1}{r}{2.2160} & \multicolumn{1}{r}{2.2554} \\ Deterministic & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9127} & \multicolumn{1}{r}{4.1916} & \multicolumn{1}{r}{5.2503} & \multicolumn{1}{r}{4.9898} & \multicolumn{1}{r}{3.9842} & \multicolumn{1}{l}{ MSL} & \multicolumn{1}{r}{3.8901} & \multicolumn{1}{r}{0.6313} & \multicolumn{1}{r}{0.7288} & \multicolumn{1}{r}{0.7144} & \multicolumn{1}{r}{ 0.5712} \\ & $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.5672 } & \multicolumn{1}{r}{19.3633} & \multicolumn{1}{r}{18.2947} & \multicolumn{1}{r}{17.7267} & \multicolumn{1}{r}{18.1284} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.9662} & \multicolumn{1}{r}{ 2.3571} & \multicolumn{1}{r}{2.2636} & \multicolumn{1}{r}{2.2128} & \multicolumn{1}{r}{2.2536} \\ Stochastic & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9099} & \multicolumn{1}{r}{4.2123} & \multicolumn{1}{r}{5.2498} & \multicolumn{1}{r}{ 4.9893} & \multicolumn{1}{r}{3.9872} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{3.8896} & \multicolumn{1}{r}{0.6331} & \multicolumn{1}{r}{ 0.7289} & \multicolumn{1}{r}{0.7149} & \multicolumn{1}{r}{0.5717} \\ & $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.7323 } & \multicolumn{1}{r}{19.3896} & \multicolumn{1}{r}{18.3164} & \multicolumn{1}{r}{17.7839} & \multicolumn{1}{r}{18.1538} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{11.3262} & \multicolumn{1}{r}{ 2.3603} & \multicolumn{1}{r}{2.2708} & \multicolumn{1}{r}{2.2203} & \multicolumn{1}{r}{2.2602} \\ \hline & & & \multicolumn{11}{c}{\textquotedblleft Block-diagonal"} \\ \cline{4-14} $\beta $ Dynamics & & & \multicolumn{5}{c}{Frobenius Norm} & & \multicolumn{5}{c}{Spectral Norm} \\ \cline{4-8}\cline{10-14} & & & \ Naive & \ Hard & \ Soft & \ AL & \ SCAD & \multicolumn{1}{l} & \ Naive & \ Hard & \ Soft & \ AL & \ SCAD \\ \cline{4-8}\cline{10-14} Constant & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9047} & \multicolumn{1}{r}{5.6802} & \multicolumn{1}{r}{6.4718} & \multicolumn{1}{r}{ 5.9421} & \multicolumn{1}{r}{5.4710} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{3.9741} & \multicolumn{1}{r}{0.8722} & \multicolumn{1}{r}{ 1.1481} & \multicolumn{1}{r}{0.9106} & \multicolumn{1}{r}{0.9014} \\ & $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.7266 } & \multicolumn{1}{r}{19.8195} & \multicolumn{1}{r}{18.8114} & \multicolumn{1}{r}{18.3162} & \multicolumn{1}{r}{18.6638} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.9751} & \multicolumn{1}{r}{ 2.8701} & \multicolumn{1}{r}{2.7656} & \multicolumn{1}{r}{2.7097} & \multicolumn{1}{r}{2.7551} \\ Deterministic & $\widehat{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{21.9137} & \multicolumn{1}{r}{5.6821} & \multicolumn{1}{r}{6.4738} & \multicolumn{1}{r}{5.9436} & \multicolumn{1}{r}{ 5.4729} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{3.9754} & \multicolumn{1}{r}{0.8718} & \multicolumn{1}{r}{1.1479} & \multicolumn{1}{r}{ 0.9103} & \multicolumn{1}{r}{0.9012} \\ & $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.6284 } & \multicolumn{1}{r}{19.8161} & \multicolumn{1}{r}{18.8043} & \multicolumn{1}{r}{18.2953} & \multicolumn{1}{r}{18.6547} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.7452} & \multicolumn{1}{r}{ 2.8706} & \multicolumn{1}{r}{2.7663} & \multicolumn{1}{r}{2.7092} & \multicolumn{1}{r}{2.7559} \\ Stochastic & $\widehat{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{21.9108} & \multicolumn{1}{r}{5.6811} & \multicolumn{1}{r}{6.4732} & \multicolumn{1}{r}{5.9433} & \multicolumn{1}{r}{ 5.4722} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{3.9751} & \multicolumn{1}{r}{0.8721} & \multicolumn{1}{r}{1.1480} & \multicolumn{1}{r}{ 0.9104} & \multicolumn{1}{r}{0.9013} \\ & $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.7955 } & \multicolumn{1}{r}{19.8314} & \multicolumn{1}{r}{18.8237} & \multicolumn{1}{r}{18.3434} & \multicolumn{1}{r}{18.6767} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{11.1149} & \multicolumn{1}{r}{ 2.8719} & \multicolumn{1}{r}{2.7691} & \multicolumn{1}{r}{2.7142} & \multicolumn{1}{r}{2.7584} \\ \hline & & & \multicolumn{11}{c}{\textquotedblleft Exponentially decaying"} \\ \cline{4-14} $\beta $ Dynamics & & & \multicolumn{5}{c}{Frobenius Norm} & \multicolumn{1}{l} & \multicolumn{5}{c}{Spectral Norm} \\ \cline{4-8}\cline{10-14} & & & \ Naive & \ Hard & \ Soft & \ AL & \ SCAD & \multicolumn{1}{l} & \ Naive & \ Hard & \ Soft & \ AL & \ SCAD \\ \cline{4-8}\cline{10-14} Constant & $\widehat{\Sigma }_{t}$ & \ MFL & \multicolumn{1}{r}{21.9057} & \multicolumn{1}{r}{6.0626} & \multicolumn{1}{r}{6.7715} & \multicolumn{1}{r}{ 6.1573} & \multicolumn{1}{r}{5.7617} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{4.0142} & \multicolumn{1}{r}{0.9106} & \multicolumn{1}{r}{ 1.1898} & \multicolumn{1}{r}{0.9453} & \multicolumn{1}{r}{0.9388} \\ & $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.8728 } & \multicolumn{1}{r}{20.3802} & \multicolumn{1}{r}{19.2709} & \multicolumn{1}{r}{18.7715} & \multicolumn{1}{r}{19.1262} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.8858} & \multicolumn{1}{r}{ 2.9381} & \multicolumn{1}{r}{2.8260} & \multicolumn{1}{r}{2.7707} & \multicolumn{1}{r}{2.8154} \\ Deterministic & $\widehat{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{21.9147} & \multicolumn{1}{r}{6.0709} & \multicolumn{1}{r}{6.7737} & \multicolumn{1}{r}{6.1589} & \multicolumn{1}{r}{ 5.7637} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{4.0156} & \multicolumn{1}{r}{0.9112} & \multicolumn{1}{r}{1.1896} & \multicolumn{1}{r}{ 0.9450} & \multicolumn{1}{r}{0.9387} \\ & $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.7746 } & \multicolumn{1}{r}{20.3564} & \multicolumn{1}{r}{19.2632} & \multicolumn{1}{r}{18.7460} & \multicolumn{1}{r}{19.1173} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{10.6538} & \multicolumn{1}{r}{ 2.9354} & \multicolumn{1}{r}{2.8247} & \multicolumn{1}{r}{2.7673} & \multicolumn{1}{r}{2.8140} \\ Stochastic & $\widehat{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{21.9118} & \multicolumn{1}{r}{6.0636} & \multicolumn{1}{r}{6.7730} & \multicolumn{1}{r}{6.1585} & \multicolumn{1}{r}{ 5.7630} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{4.0151} & \multicolumn{1}{r}{0.9106} & \multicolumn{1}{r}{1.1897} & \multicolumn{1}{r}{ 0.9451} & \multicolumn{1}{r}{0.9388} \\ & $\widetilde{\boldsymbol{\Sigma }}_{t}$ & \ MFL & \multicolumn{1}{r}{30.9430 } & \multicolumn{1}{r}{20.3820} & \multicolumn{1}{r}{19.2839} & \multicolumn{1}{r}{18.8017} & \multicolumn{1}{r}{19.1405} & \multicolumn{1}{l}{MSL} & \multicolumn{1}{r}{11.0295} & \multicolumn{1}{r}{ 2.9381} & \multicolumn{1}{r}{2.8291} & \multicolumn{1}{r}{2.7745} & \multicolumn{1}{r}{2.8173} \\ \hline\hline \end{tabular} } \end{center}
center[center omitted — 4,425 chars of source]

Empirical Study

\setcounter{equation}{0}

We apply the proposed methods to the intraday returns of the S&P 500 component stocks to demonstrate the effectiveness of our nonparametric spot volatility matrix estimation in revealing time-varying patterns. We consider the 5-minute returns of the S&P 500 stocks collected in September 2008. On September 15 Lehman Brothers filed for bankruptcy, causing shockwaves throughout the global financial system. Hence, it is interesting to examine how the spot volatility structure of the returns evolved during this one-month period. In addition, to demonstrate the effectiveness of our model with observed risk factors in explaining the systemic component of the dependence structure, we also collect the 5-minute returns of twelve factors. The first three factors are constructed in AKX20 as our proxy for the market (MKT), small-minus-big market capitalisation (SMB), and high-minus-low price-earning ratio (HML). The other nine factors are the widely available sector SDPR ETFs, which are intended to tract the following nine largest S&P sectors: Energy (XLE), Materials (XLB), Industrials (XLI), Consumer Discretionary (XLY), Consumer Staples (XLP), Health Care (XLV), Financial (XLF), Information Technology (XLK), Utilities (XLU). We sort our stocks according to their GICS (Global Industry Classification Standard) codes, so that they are grouped by sectors in the above order. Consequently, the correlation (sub)matrix for stocks within each sector corresponds to a block on the diagonal of the full correlation matrix FFX16.

We only use stocks that are included in the S&P 500 index and whose GICS codes are unchanged in September 2008. We also exclude stocks that do not belong to any of the above nine sectors. This leaves us with a total of $p=482$ stocks. All the returns are synchronised via the previous-tick subsampling technique Zh11, and overnight returns are removed because of potential dividends and stock splits. Consequently, we have $1638$ time series observations for each of the $482$ stocks. For the 5-minute returns, we may assume that the potential impact of microstructure noises are negligible. The smoothing parameter in our kernel estimation is chosen as $h=2/252$ (equivalent to 2 trading days)\footnote{We experimented three bandwidth choices, namely, 1 day, 2 days and 3 days. We found that $h=1/252$ ($h=3/252$) produced clearly undersmoothed (oversmoothed) time series of estimated deciles of the cross sectional distribution of the variances and pairwise correlations of our returns, whereas $h=2/252$ seems to be most reasonable. Our qualitative conclusion is unaffected by the choices of $h$ within the range of 1 to 3 days.}.

We start with estimating the spot volatility matrices of the total returns (i.e., the observed returns) without incorporating the observed factors or applying any shrinkage. To visualise the potential time variation of the estimated spot matrices, as in BHMR19, we plot the time series of deciles of the distribution of the estimated spot variances and the pairwise correlations\footnote{The nine decile levels we use in this study are the 10th, 20th,..., and 90th percentiles.}. The patterns of the spot variances and correlations in Figure (ref)(a) and (b) reveal some clear evidence of time variation in our sampling period. We note that the distributions of the variances are relatively narrow and stay low on the first few days of the month. However, close to Lehman Brothers' announcement on the 15th, they start to rise and get wider quite rapidly and reach the peak around the 17th and the 18th. The spot variances at the peak are much higher than those on the earlier days of the month. The distributions return to the earlier level in the following week. In contrast, the distributions of pairwise spot correlations also start to shift up around the same time, but quickly reach the peak on the 16th (only one day after the bankruptcy news), and then dip to a relatively low point around the 19th before returning to the earlier level. Such time-varying features in the dynamics covariance structure are quite interesting and sensible, reflecting the impact of market news. Hence, our proposed spot volatility matrix estimation methodology provides a useful tool for revealing such dynamics.

figure[figure omitted — 179 chars of source]

To examine whether it is appropriate to directly apply shrinkage techniques to the spot volatility matrices of the total returns, following FFX16 we plot in Figure (ref)(a) and (b) their sparsity patterns on the 16th and the 19th of September \footnote{Recall that as in DLX19 our tuning parameter used for each pairwise spot covariance in our shrinkage method is proportional to the product of the spot standard deviations of the returns of that pair of assets. Therefore, the sparsity pattern is effectively determined by the spot correlation matrix.}. The deep blue dots correspond to the locations of pairwise correlations that are at least $0.15$, whereas the white dots correspond to those smaller than $0.15$. Note that the covariance structure of the total returns is very dense on these two days. Therefore, it is not appropriate to directly apply the shrinkage technique as in Sections (ref) and (ref). Meanwhile, although both are quite dense, we can still clearly see their differences. Consistent with our observation from the decile plots of the correlations, we can see that the plot for the 16th is almost completely covered by blue dots, but the plot for the 19th in contrast has significantly more areas covered in white.

figure[figure omitted — 215 chars of source]

We next incorporate the twelve observed factors in the large spot volatility matrix estimation as suggested in Section (ref). In particular, we are interested on estimating the spot idiosyncratic volatility matrix, which is expected to satisfy the sparsity restriction. To save space, we choose to only report results using the SCAD shrinkage due to its satisfactory performance in our simulations. In Figure (ref)(c) and (d), we plot the deciles of the estimated spot idiosyncratic variances and correlations over trading days. In Figure (ref)(c), we observe a significant upward shift of the distribution of the spot variances of the idiosyncratic returns around the time of Lehman Brothers' bankruptcy, indicating that the observed factors may not fully capture the time variation of the spot variances. In contrast, the deciles of the spot correlations in Figure (ref)(d) seem to be quite flat throughout the entire month, suggesting that the systematic factors may explain the time variation in the distribution of the pairwise correlations better than that of the variances.

We finally plot the sparsity patterns of the two estimated spot idiosyncratic volatility matrices on 16 and 19 September in Figure (ref)(c) and (d), respectively. Unlike Figure (ref)(a) and (b), we note that the estimated spot idiosyncratic volatility matrices are highly sparse on the two days. This is consistent with our observation from Figure (ref)(d), confirming that the observed factors can effectively account for the time variation in the spot covariance structure of the returns. Meanwhile, we also note that the two idiosyncratic volatility matrices are clearly not diagonal and still carry some visible time variation. Lastly, it is worth mentioning that the estimated spot idiosyncratic volatility matrices do not exhibit significant correlations within the blocks along the diagonal lines, except for some very limited actions in the lower right corner of the two matrices (the lower right corner corresponds to the XLU sector according to our sorting).

Conclusion

We developed nonparametric estimation methods for large spot volatility matrices under the uniform sparsity assumption. We allowed for microstructure noise and observed common risk factors and employed kernel smoothing and generalised shrinkage. In each scenario we obtained the uniform convergence rates for the large estimated covariance matrices and these reflect the smoothness and sparsity assumptions we made. The simulation results show that the proposed estimation methods work well in finite samples for both the noise-free and noise-contaminated data. The empirical study demonstrated the effectiveness on S&P 500 stocks five minute data. Several issues can be further explored. For example, it is worthwhile to further study the spot precision matrix estimation which is briefly discussed in Appendix C.1 of the supplement and explore its application to optimal portfolio choice.

Acknowledgements

The authors would like to thank a Co-Editor and two reviewers for the constructive comments, which helped to improve the article. The first author's research was partly supported by the BA Talent Development Award (No. TDA21$\backslash$210027). The second author’s research was partly supported by the BA/Leverhulme Small Research Grant funded by the Leverhulme Trust (No. SRG1920/ 100603).

thebibliography{99} { \harvarditem{A\"{\i}t-Sahalia \harvardand\ Jacod}{2014}{AJ14} A\"{\i}t-Sahalia, Y. & J. Jacod (2014) High-Frequency Financial Econometrics. Princeton University Press. } { \harvarditem{A\"{\i}t-Sahalia, Kalnina \harvardand\ Xiu}{2020}{AKX20} A\"{\i}t-Sahalia, Y., I. Kalnina, & D. Xiu (2020) High-frequency factor models and regressions. Journal of Econometrics 216, 86--105. } { \harvarditem{A\"{\i}t-Sahalia \harvardand\ Xiu}{2017}{AX17} A\"{\i}t-Sahalia, Y. & D. Xiu (2017) Using principal component analysis to estimate a high dimensional factor model with high-frequency data. Journal of Econometrics 201, 384--399. } { \harvarditem{A\"{\i}t-Sahalia \harvardand\ Xiu}{2019}{AX19} \textsc{A\"{\i}t-Sahalia, Y. & D. Xiu} (2019) Principal component analysis of high-frequency data. \emph{Journal of the American Statistical Association} 114, 287--303. } { \harvarditem{Andersen and Bollerslev}{1998}{AB98} \textsc{ Andersen, T. G. & T. Bollerslev} (1998) Answering the skeptics: yes, standard volatility models do provide accurate forecasts. \emph{ International Economic Review} 39, 885--905. } { \harvarditem{Andersen, Bollerslev \harvardand\ Diebold}{2010}{ABD10} \textsc{Andersen, T. G., T. Bollerslev, & F. X. Diebold} (2010) Parametric and nonparametric volatility measurement. In \emph{ Handbook of Financial Econometrics: Tools and Techniques} (Y. A\"{\i} t-Sahalia and L. P. Hansen, eds.), 67--137. } { \harvarditem{Andersen {\em et al.}}{2003}{ABDL03} \textsc{ Andersen, T. G., T. Bollerslev, F. X. Diebold, & P. Labys} (2003) Modeling and forecasting realized volatility. \emph{Econometrica} 71, 579--625. } { \harvarditem{Bai \harvardand\ Silverstein}{2010}{BS10} \textsc{Bai, Z. & J. W. Silverstein} (2010) \emph{Spectral Analysis of Large Dimensional Random Matrices}. Springer Series in Statistics, Springer. } { \harvarditem{Barndorff-Nielsen \harvardand\ Shephard}{2002}{BS02} \textsc{Barndorff-Nielsen, O. E. & N. Shephard} (2002) Econometric analysis of realized volatility and its use in estimating stochastic volatility models. \emph{Journal of the Royal Statistical Society Series B} 64, 253--280. } { \harvarditem{Barndorff-Nielsen \harvardand\ Shephard}{2004}{BS04} \textsc{Barndorff-Nielsen, O. E. & N. Shephard} (2004) Econometric analysis of realized covariation: High frequency based covariance, regression and correlation in financial economics. \emph{ Econometrica} 72, 885--925. } { \harvarditem{Barndorff-Nielsen {\em et al.}}{2008}{BHLS08} \textsc{Barndorff-Nielsen, O. E., P. R. Hansen, A. Lunde, & N. Shephard} (2008) Designing realised kernels to measure the ex-post variation of equity prices in the presence of noise. \emph{Econometrica} 76, 1481--1536. } { \harvarditem{Bibinger {\em et al.}}{2019}{BHMR19}\textsc{Bibinger, M., N. Hautsch, P. Malec, & M. Reiss} (2019) Estimating the spot covariation of asset prices - statistical theory and empirical evidence. \emph{Journal of Business and Economic Statistics} 37(3), 419--435. } { \harvarditem{Bickel \harvardand\ Levina}{2008}{BL08} \textsc{ Bickel, P. & E. Levina} (2008) Covariance regularization by thresholding. \emph{Annals of Statistics} 36, 2577--2604. } { \harvarditem{Cai {\em et al}}{2020}{CHLZ20} \textsc{Cai, T. T., J. Hu, Y. Li, & X. Zheng} (2020) High-dimensional minimum variance portfolio estimation based on high-frequency data. \emph{Journal of Econometrics} 214, 482--494. } { \harvarditem{Cai \harvardand\ Zhou}{2012}{CZ12} \textsc{Cai, T. T. & H. H. Zhou} (2012) Optimal rates of convergence for sparse covariance matrix estimation. \emph{Annals of Statistics} 40, 2389--2420. } { \harvarditem{Chang {\em et al}.}{2021}{CHLT21} \textsc{Chang, J., Q. Hu, C. Liu, & C. Tang} (2021) Optimal covariance matrix estimation for high-dimensional noise in high-frequency data. Working paper available at \url{https://arxiv.org/abs/1812.08217.} } { \harvarditem{Chen, Li \harvardand\ Linton}{2019}{CLL19} \textsc{Chen, J., D. Li, & O. Linton} (2019) A new semiparametric estimation approach of large dynamic covariance matrices with multiple conditioning variables. \emph{Journal of Econometrics} 212, 155--176. } { \harvarditem{Chen, Mykland \harvardand\ Zhang}{2020}{CMZ20} \textsc{Chen, D., P. A. Mykland, & L. Zhang} (2020) The five trolls under the bridge: principal component analysis with asynchronous and noisy high frequency data. {\em Journal of the American Statistical Association} 115, 1960--1977. } { \harvarditem{Chen, Xu \harvardand\ Wu}{2013}{CXW13} \textsc{ Chen, X., M. Xu, & W. Wu} (2013) Covariance and precision matrix estimation for high-dimensional time series. \emph{Annals of Statistics} 41, 2994--3021. } { \harvarditem{Chen \harvardand\ Leng}{2016}{CL16} \textsc{ Chen, Z.& C. Leng} (2016) Dynamic covariance models. \emph{ Journal of the American Statistical Association} 111, 1196--1207. } { \harvarditem{Christensen, Kinnebrock \harvardand\ Podolskij}{2010}{CKP10} \textsc{Christensen, K., S. Kinnebrock, & M. Podolskij} (2010) Pre-averaging estimators of the ex-post covariance matrix in noisy diffusion models with non-synchronous data. \emph{Journal of Econometrics} 159, 116--133. } { \harvarditem{Dai, Lu \harvardand\ Xiu}{2019}{DLX19} \textsc{ Dai, C., K. Lu, & D. Xiu} (2019) Knowing factors or factor loadings, or neither? Evaluating estimators for large covariance matrices with noisy and asynchronous data. \emph{Journal of Econometrics} 208, 43--79. } { \harvarditem{Epps}{1979}{Ep79} \textsc{Epps, T. W.} (1979) Comovements in stock prices in the very short run. \emph{Journal of the American Statistical Association} 74, 291--298. } { \harvarditem{Fan, Fan \harvardand\ Lv}{2007}{FFL07} \textsc{ Fan, J., Y. Fan, & J. Lv} (2007) Aggregation of nonparametric estimators for volatility matrix. \emph{Journal of Financial Econometrics} 5, 321--357. } { \harvarditem{Fan, Furger \harvardand\ Xiu}{2016}{FFX16} \textsc{Fan, J., A. Furger, & D. Xiu} (2016) Incorporating global industrial classification standard into portfolio allocation: A simple factor-based large covariance matrix estimator with high frequency data. \emph{Journal of Business and Economic Statistics} 34, 489--503. } { \harvarditem{Fan \harvardand\ Gijbels}{1996}{FG96}\textsc{Fan, J. & I. Gijbels} (1996) {\em Local Polynomial Modelling and Its Applications}. Chapman and Hall, London. } { \harvarditem{Fan, Liao \harvardand\ Mincheva}{2011}{FLM11} \textsc{Fan, J., Y. Liao, & M. Mincheva} (2011) High-dimensional covariance matrix estimation in approximate factor models. \emph{Annals of Statistics} 39, 3320--3356. } { \harvarditem{Fan, Liao \harvardand\ Mincheva}{2013}{FLM13} \textsc{Fan, J., Y. Liao, & M. Mincheva} (2013) Large covariance estimation by thresholding principal orthogonal complements (with discussion). \emph{Journal of the Royal Statistical Society, Series B} 75, 603--680. } { \harvarditem{Fan \harvardand\ Wang}{2008}{FW08} \textsc{Fan, J. & Y. Wang} (2008) Spot volatility estimation for high-frequency data. \emph{Statistics and Its Interface} 1, 279--288. } { \harvarditem{Figueroa-L\'{o}pez \harvardand\ Li}{2020}{FL20} \textsc{Figueroa-L\'{o}pez, J. E. & C. Li} (2020) Optimal kernel estimation of spot volatility of stochastic differential equations. \emph{Stochastic Processes and Their Applications} 130, 4693--4720. } { \harvarditem{Hayashi \harvardand\ Yoshida}{2005}{HY05} \textsc{Hayashi, T. & N. Yoshida} (2005) On covariance estimation of non-synchronously observed diffusion processes. \emph{Bernoulli } 11, 359--379. } { \harvarditem{Jacod {\em et al.}}{2009}{JLMPV09} \textsc{ Jacod, J., Y. Li, P. A. Mykland, M. Podolskij, & M. Vetter} (2009) Microstructure noise in the continuous case: the pre-averaging approach. \emph{Stochastic Processes and Their Applications} 119, 2249--2276. } { \harvarditem{Jacod \harvardand\ Protter}{2012}{JP12} \textsc{ Jacod, J. & P. Protter} (2012) \emph{Discretization of Processes}. Springer. } { \harvarditem{Kalnina \harvardand\ Linton}{2008}{KL08} \textsc{ Kalnina, I. & O. Linton} (2008) Estimating quadratic variation consistently in the presence of endogenous and diurnal measurement error. \emph{Journal of Econometrics} 147, 47--59. } { \harvarditem{Kanaya \harvardand\ Kristensen}{2016}{KK16} \textsc{Kanaya, S. & D. Kristensen} (2016) Estimation of stochastic volatility models by nonparametric filtering. \emph{Econometric Theory} 32, 861--916. } { \harvarditem{Kim, Wang \harvardand\ Zou}{2016}{KWZ16} \textsc{ Kim, D., Y. Wang, & J. Zou} (2016) Asymptotic theory for large volatility matrix estimation based on high-frequency financial data. \emph{ Stochastic Processes and Their Applications} 126, 3527--3577. } { \harvarditem{Kong}{2018}{K18} \textsc{Kong, X.} (2018) On the systematic and idiosyncratic volatility with large panel high-frequency data. \emph{Annals of Statistics} 46, 1077--1108. } { \harvarditem{Kristensen}{2010}{K10} \textsc{Kristensen, D.} (2010) Nonparametric filtering of the realized spot volatility: a kernel-based approach. \emph{Econometric Theory} 26, 60--93. } { \harvarditem{Lam \harvardand\ Feng}{2018}{LF18} \textsc{Lam, C. & P. Feng} (2018) A nonparametric eigenvalue-regularized integrated covariance matrix estimator for asset return data. \emph{Journal of Econometrics} 206, 226--257. } { \harvarditem{Li \harvardand\ Racine}{2007}{LR07} \textsc{Li, Q. & J. Racine} (2007) \emph{Nonparametric Econometrics}. Princeton University Press, Princeton. } { \harvarditem{Mykland \harvardand\ Zhang}{2006}{MZ06} \textsc{Mykland, P. A. & L. Zhang} (2006) ANOVA for diffusions and It\^o processes. {\em Annals of Statistics} 34, 1931--1963.} { \harvarditem{Park, Hong \harvardand\ Linton}{2016}{PHL16} \textsc{Park, S., S. Y. Hong, & O. Linton} (2016) Estimating the quadratic covariation matrix for asynchronously observed high frequency stock returns corrupted by additive measurement error. \emph{Journal of Econometrics} 191, 325-347. } { \harvarditem{Podolskij \harvardand\ Vetter}{2009}{PV09} \textsc{Podolskij, M. & M. Vetter} (2009) Estimation of volatility functionals in the simultaneous presence of microstructure noise and jumps. \emph{Bernoulli} 15, 634--658. } { \harvarditem{Rei{\ss}, Todorov \harvardand\ Tauchen}{2015}{RTT15} \textsc{Rei{\rm {\ss}}, M., V. Todorov, & G. Tauchen} (2015) Nonparametric test for a constant beta between It\^o semi-martingales based on high-frequency data. {\em Stochastic Processes and Their Applications} 125, 2955--2988.} { \harvarditem{Revuz \harvardand\ Yor}{1999}{RY99} \textsc{ Revuz, D. & M. Yor} (1999) \emph{Continuous Martingales and Brownian Motion}. Grundlehren der mathematischen Wissenschaften 293, Springer. } { \harvarditem{Shephard}{2005}{S05} \textsc{Shephard, N.} (2005) \emph{Stochastic Volatility: Selected Readings}. Oxford University Press. } { \harvarditem{Tao, Wang \harvardand\ Zhou}{2013}{TWZ13} \textsc{Tao, M., Y. Wang, & H. H. Zhou} (2013) Optimal sparse volatility matrix estimation for high-dimensional It\^{o} processes with measurement errors. \emph{Annals of Statistics} 41, 1816--1864. } { \harvarditem{Wang \harvardand\ Zou}{2010}{WZ10} \textsc{Wang, Y. & J. Zou} (2010) Vast volatility matrix estimation for high-frequency financial data. \emph{Annals of Statistics} 38, 943--978. } { \harvarditem{Xia \harvardand\ Zheng}{2018}{XZ18} \textsc{Xia, N. & X. Zheng} (2018) On the inference about the spectral distribution of high-dimensional covariance matrix based on high-frequency noisy observations. \emph{Annals of Statistics} 46, 500--525. } { \harvarditem{Xiao \harvardand\ Linton}{2002}{XL02} \textsc{ Xiao, Z. & O. Linton} (2002) A nonparametric prewhitened covariance estimator. \emph{Journal of Time Series Analysis} 23, 215--250. } { \harvarditem{Zhang}{2011}{Zh11} \textsc{Zhang, L.} (2011) Estimating covariation: Epps effect, microstructure noise. \emph{Journal of Econometrics} 160, 33--47. } { \harvarditem{Zhang, Mykland \harvardand\ A\"{\i}t-Sahalia}{2005}{ZMA05} \textsc{Zhang, L., P. A. Mykland, & Y. A\"{\i} t-Sahalia} (2005) A tale of two time scales: Determining integrated volatility with noisy high-frequency data. \emph{Journal of the American Statistical Association} 100, 1394--1411. } { \harvarditem{Zheng \harvardand\ Li}{2011}{ZL11} \textsc{ Zheng, X. & Y. Li} (2011) On the estimation of integrated covariance matrices of high dimensional diffusion processes. \emph{Annals of Statistics} 39, 3121--3151. } { \harvarditem{Zu \harvardand\ Boswijk}{2014}{ZB14} \textsc{Zu, Y. & H. P. Boswijk} (2014) Estimating spot volatility with high-frequency financial data. \emph{Journal of Econometrics} 181, 117--135. }