The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
67,911 characters
Low-Rank Structured Nonparametric Prediction of Instantaneous Volatility
\bibliographystyle{ecta}
\title{Low-Rank Structured Nonparametric Prediction of Instantaneous Volatility}
\date{July 2025}
\author{
Sung Hoon Choi\thanks{
Department of Economics, University of Connecticut, Storrs, CT 06269, USA.
E-mail: \texttt{sung\[email removed]}.} \\
\and Donggyu Kim\thanks{Department of Economics, University of California, Riverside, CA 92521, USA.
Email: \texttt{[email removed]}.}
}
\maketitle
\pagenumbering{arabic}
\begin{abstract}
\onehalfspacing
Based on Itô semimartingale models, several studies have proposed methods for forecasting intraday volatility using high-frequency financial data.
These approaches typically rely on restrictive parametric assumptions and are often vulnerable to model misspecification.
To address this issue, we introduce a novel nonparametric prediction method for the future intraday instantaneous volatility process during trading hours, which leverages both previous days' data and the current day's observed intraday data.
Our approach imposes an interday-by-intraday matrix representation of the instantaneous volatility, which is decomposed into a low-rank conditional expectation component and a noise matrix.
To predict the future conditional expected volatility vector, we exploit this low-rank structure and propose the Structural Intraday-volatility Prediction (SIP) procedure.
We establish the asymptotic properties of the SIP estimator and demonstrate its effectiveness through an out-of-sample prediction study using real high-frequency trading data.
\\
\noindent \textbf{Key words:}
Diffusion process, high-frequency financial data, low-rank matrix, matrix completion.
\end{abstract}
\vfill\eject
\doublespacing
\section{Introduction}
The analysis of volatility is essential in financial econometrics and statistics, and it has broad applications in hedging, option pricing, risk management, and portfolio allocation.
In recent years, the widespread availability of high-frequency financial data has led to the development of several nonparametric integrated volatility estimation methods \citep{ait2010high, barndorff2008designing, barndorff2011multivariate, bibinger2014estimating, christensen2010pre, fan2018robust, fan2007multi, jacod2009microstructure, xiu2010quasi, zhang2005tale, zhang2006efficient, zhang2011estimating}.
These high-frequency-based estimators have significantly improved our understanding of low-frequency (interday) market dynamics and have motivated the development of various conditional volatility models built on realized measures \citep{andersen2003modeling, corsi2009simple, hansen2012realized, kim2016unified, shephard2010realising}.
Parallel to these developments, nonparametric procedures for estimating instantaneous (or spot) volatility have been proposed to capture intraday volatility dynamics \citep{fan2008spot,figueroa2022kernel, foster1996continuous, kristensen2010nonparametric, mancini2015spot, todorov2019nonparametric, todorov2023bias, zu2014estimating}.
Leveraging these well-performing estimators, several studies have illustrated U-shaped or reverse J-shaped intraday volatility patterns across various markets \citep{admati1988theory, andersen1997intraday, andersen2019time, hong2000trading, li2023robust}.
In contrast to the extensive literature on predicting daily volatility processes, relatively less attention has been given to forecasting intraday volatility.
For example, \cite{engle2012forecasting} proposed a modified GARCH model in which the conditional variance is expressed as the product of daily, diurnal, and stochastic intraday components to capture intraday volatility dynamics, while \cite{zhang2024volatility} explored several nonparametric machine learning methods for forecasting intraday realized volatility, leveraging the commonality observed in intraday volatility patterns.
Recently, \cite{TIP-PCA} considered the interday-by-intraday instantaneous volatility matrix process and proposed a semiparametric prediction procedure for the one-day-ahead instantaneous volatility process by leveraging both interday and intraday volatility dynamics.
However, existing methods for forecasting intraday volatility suffer from notable limitations.
For example, many methods rely on strong distributional assumptions inherent in parametric models and often fail to incorporate interday dynamics effectively.
In practice, practitioners are often interested in forecasting the remaining future intraday volatility process conditional on the observed intraday data up to the present time.
Therefore, it is both practically relevant and methodologically important to develop a prediction procedure that adapts to current intraday information and is robust to potential model misspecification.
This paper introduces a novel nonparametric prediction approach based on a low-rank structure for forecasting the remaining future instantaneous volatility process during intraday trading.
Specifically, we represent the entire instantaneous volatility process as a matrix, where each row corresponds to a day and each column represents an intraday time point.
We impose a low-rank plus noise structure on this matrix.
The parameter of interest is the instantaneous volatility process after the current intraday time.
That is, on the current day (denoted as the $D$th day), only the initial portion of the row is observed, while the remaining part is the target of prediction.
To forecast this remaining portion, we exploit the low-rank matrix structure.
For example, the future instantaneous volatility vector can be represented by the singular components of the observed part of the low-rank matrix.
We use this structure to construct a nonparametric prediction of the future volatility vector.
This method is called the Structural Intraday-volatility Prediction (SIP) procedure.
It is worth noting that as long as the interday-by-intraday instantaneous volatility matrix exhibits a low-rank structure, the SIP method can effectively predict the remaining future volatility vector without relying on any specific parametric dynamics such as autoregressive models.
Therefore, the SIP procedure is robust to model misspecification arising from incorrect parametric assumptions.
Furthermore, we derive the convergence rate of the predicted instantaneous volatility under the SIP framework and assess its empirical performance via an out-of-sample prediction study using high-frequency financial data.
The SIP method consistently outperforms alternative approaches in terms of both out-of-sample prediction accuracy and Value-at-Risk (VaR) forecasting performance.
The remainder of the paper is organized as follows.
In Section \ref{setup}, we introduce the model and propose the SIP prediction procedure.
Section \ref{asymp} presents the asymptotic properties of the SIP estimator.
In Section \ref{simulation}, we evaluate the finite sample performance of the proposed method through simulation.
Section \ref{empiric} applies the proposed method to real high-frequency financial data to forecast the remaining future instantaneous volatility process during trading hours.
Finally, the conclusion is provided in Section \ref{conclusion}.
All technical proofs are provided in the Appendix \ref{proofs}.
\section{Model Setup and Estimation Procedure} \label{setup}
Throughout this paper, we denote by $\|\bfm A\|_{F}$, $\|\bfm A\|_{2}$ (or $\|\bfm A\|$ for short), $\|\bfm A\|_{1}$, $\|\bfm A\|_{\infty}$, and $\|\bfm A\|_{\max}$ the Frobenius norm, operator norm, $l_{1}$-norm, $l_{\infty}$-norm, and element-wise norm, which are defined, respectively, as $\|\bfm A\|_{F} = \mathrm{tr}^{1/2}(\bfm A^{\top}\bfm A)$, $\|\bfm A\|_{2} = \lambda_{\max}^{1/2}(\bfm A^{\top}\bfm A)$, $\|\bfm A\|_{1} = \max_{j}\sum_{i}|a_{ij}|$, $\|\bfm A\|_{\infty} = \max_{i}\sum_{j}|a_{ij}|$, and $\|\bfm A\|_{\max} = \max_{i,j}|a_{ij}|$.
When $\bfm A$ is a vector, the maximum norm is denoted as $\|\bfm A\|_{\infty}=\max_{i}|a_{i}|$, and both $\|\bfm A\|$ and $\|\bfm A\|_{F}$ are equal to the Euclidean norm.
\subsection{A Model Setup} \label{model}
We consider the following jump diffusion process: for the $i$-th day and intraday time $t \in [0,1],$
\begin{equation}\label{diffusion-def}
dX_{i,t}= \mu_{i,t} dt + \sigma_{i,t} dB_{i,t} + J_{i,t} d P_{i,t},
\end{equation}
where $X_{i,t}$ is the log price of an asset, $\mu_{i,t}$ is a drift process, $B_{i,t}$ is a one-dimensional standard Brownian motion, $J_{i,t}$ is the jump size, and $P_{i,t}$ is the Poisson process with the intensity $\mu_{J}$.
For a given intraday time sequence, for each $i =1,\dots, D$ and $j=1,\dots, n$, we denote the instantaneous volatility process as $c_{i,j} := \sigma_{i,t_{j}}^2$, where $0< t_{1} < \cdots < t_{n} =1$.
Then, we assume that the discrete-time instantaneous volatility process satisfies the following low-rank structure:
\begin{equation}\label{Sigma}
\bfsym \Sigma_{D, n} = (c_{i,j}) _{D \times n} = \bfm U \ensuremath{\boldsymbol{\Lambda}} \bfm V^{\top} + \bfm E \equiv \bfm A + \bfm E,
\end{equation}
where $\bfm U= (u_{i,k} ) _{i=1,\ldots, D, k=1,\ldots,r}$ is the left singular vector matrix, $\bfm V = (v_{j,k}) _{j=1, \ldots, n, k=1,\ldots, r}$ is the right singular vector matrix, $\ensuremath{\boldsymbol{\Lambda}}= \mathrm{Diag} (\lambda_1, \ldots, \lambda_r)$ is the singular value matrix, and $\bfm E = (e_{i,j})_{D\times n}$ is the random noise matrix.
We only impose the above low-rank structure in \eqref{Sigma} to predict future remaining intraday volatility.
It allows for considerable modeling flexibility and can implicitly incorporate various time-series features.
For example, the left singular vector can represent interday time-series dynamics, while the right singular vector can account for intraday patterns, such as U-shaped or reverse J-shaped \citep{TIP-PCA}.
It is worth noting that, unlike \citet{TIP-PCA}, we do not assume any explicit dynamic or parametric form for the volatility evolution.
We note that $\bfsym \Sigma_{D, n}$ is a $D \times n$ matrix, which is distinct from a covariance matrix and contains only positive elements.
In this paper, our target is to predict the remaining future instantaneous volatility vector for the $D$th day.
Specifically, we assume that trading data is available up to a fraction $\omega = \frac{n_1}{n}$ of the $D$th day.
The objective is to predict the remaining $n_2 \times 1$ instantaneous vector, $(c_{D,n_1+1},\dots, c_{D,n})$, where $n = n_1 + n_2$.
Given the current information $\mathcal{F}_{D,n_1}$, we can predict the instantaneous volatility as follows: for $j = n_1+1, \ldots, n$,
\begin{equation} \label{conditional exp}
E\left[\sigma_{D,j}^2 | \mathcal{F}_{D,n_1}\right] = \sum_{k=1}^r \lambda_k u_{D,k} v_{j,k} \,\, \text{ a.s.}
\end{equation}
We can write the model \eqref{Sigma} in a partitioned matrix form as follows:
$$
\bfsym \Sigma_{D,n} =
\begin{array}{c@{}c}
\footnotesize \begin{array}{cc} n_1 & n_2 \end{array} \\
\left[
\begin{array}{cc}
\Sigma_{11} & \Sigma_{12} \\
\Sigma_{21} & \Sigma_{22}
\end{array}
\right]& \footnotesize \begin{array}{c} D-1 \\ 1 \end{array}
\end{array},
\bfm A =
\begin{array}{c@{}c}
\footnotesize \begin{array}{cc} n_1 & n_2 \end{array} \\
\left[
\begin{array}{cc}
A_{11} & A_{12} \\
A_{21} & A_{22}
\end{array}
\right]& \footnotesize \begin{array}{c} D-1 \\ 1 \end{array}
\end{array},
\bfm E =
\begin{array}{c@{}c}
\footnotesize \begin{array}{cc} n_1 & n_2 \end{array} \\
\left[
\begin{array}{cc}
E_{11} & E_{12} \\
E_{21} & E_{22}
\end{array}
\right]& \footnotesize \begin{array}{c} D-1 \\ 1 \end{array}.
\end{array}
$$
We note that $\Sigma_{11}, \Sigma_{12}$, and $\Sigma_{21}$ are estimable using high-frequency log-price observations.
Since $\bfm A$ is the low rank matrix, $A_{22}$ is given by
\begin{equation}\label{target}
A_{22} = A_{21}(A_{11})^\dagger A_{12} = A_{21}(V_{11}\Lambda_{11}^{-1}U_{11}^{\top})A_{12},
\end{equation}
where $U_{11}$ is the $(D-1)\times r$ left singular vector matrix, $V_{11}$ is the $n_1 \times r$ right singular vector matrix, and $\Lambda_{11}$ is the $r\times r$ singular value matrix of $A_{11}$.
Since $A_{11}$, $A_{21}$, and $A_{12}$ are quantities related to the observed period, using this relationship \eqref{target}, we can predict $A_{22}$ without a specific parametric modeling assumption.
Specifically, as long as we can estimate $A_{11}$, $A_{21}$, and $A_{12}$ well using the available data, $A_{22}$ can be estimated using the plug-in scheme.
We note that if we treat the remaining intraday volatility vector $A_{22}$ as missing components, this problem becomes similar to the matrix completion \citep{cai2010singular, candes2012exact, candes2010matrix}.
Especially in this problem, the missing has a determined structure as in \citet{bai2021matrix, cai2016structured,fan2019structured, yan2024entrywise}.
From this point of view, this paper extends the matrix completion concept to predict time series patterns.
Due to the imperfections in the trading process, practitioners cannot directly observe the true underlying log-stock price $X_{i,t}$ in \eqref{diffusion-def} \citep{ait2009high}.
To account for this, we assume that the high-frequency intraday observations $X_{i,t_s}, s=1, \dots, m,$ are contaminated by microstructure noise as follows:
\begin{equation} \label{def-obervation}
Y _{i,t_s} = X_{i,t_s} + e_{i,t_s}, \quad i=1,\ldots,D, s=1,\ldots,m,
\end{equation}
where $e_{i,t_s}$ is the microstructure noise process with a mean of zero and a variance of $\eta_{ii}$.
\subsection{Structural Intraday-volatility Prediction} \label{estimation procedure}
To predict the remaining future instantaneous volatility vector $A_{22}$ nonparametrically, we use the low-rank structural equation \eqref{target}.
Since $A_{11}$, $A_{12}$, and $A_{21}$ are latent low-rank matrices, we need to estimate them using a well-performing instantaneous volatility estimator based on the observed high-frequency data.
For example, we need a large dimension to estimate the latent low-rank matrix and use the blessing of dimensionality \citep{li2018embracing}.
To do this, we choose the large block matrices such as $(\Sigma_{11}, \Sigma_{12})$ and $(\Sigma_{11} ^{\top}, \Sigma_{21} ^{\top})$.
On the other hand, we employ a well-performing nonparametric instantaneous volatility estimator to estimate the instantaneous volatility. Details can be found in Remark \ref{remark_spot} below.
The specific estimation procedure is as follows:
\begin{enumerate}
\item Using high-frequency log-price observations, we compute the instantaneous volatility estimators $\widehat{c}_{i,j}$ up to the $D$th day $n_1$ intraday time point (see Remark \ref{remark_spot}).
Then, we define the following submatrices: $\widehat{\Sigma}_{11} = (\widehat{c}_{i,j})_{(D-1)\times n_1}, \widehat{\Sigma}_{12} = (\widehat{c}_{i,j})_{(D-1)\times n_2}$, and $\widehat{\Sigma}_{21} = (\widehat{c}_{i,j})_{1\times n_1}$, which are the initial estimated matrices for $\Sigma_{11}$,$\Sigma_{12}$, and $\Sigma_{21}$, respectively.
\item We conduct the singular value decomposition to the large submatrices $\widehat{\Sigma}_{1\bullet} \equiv [\widehat{\Sigma}_{11} \: \widehat{\Sigma}_{12}]$ and $\widehat{\Sigma}_{\bullet 1} \equiv [\widehat{\Sigma}_{11}^{\top} \: \widehat{\Sigma}_{21}^{\top}]^{\top}$.
The columns of $\widehat{U}_{11}$ are defined as the $r$ leading left singular vector of $\widehat{\Sigma}_{1\bullet}$, and the columns of $\widehat{V}_{11}$ are defined as the $r$ leading right singular vector of $\widehat{\Sigma}_{\bullet 1}$.
Note that $\widehat{U}_{11}$ and $\widehat{V}_{11}$ are estimators for the orthonormal basis of column vectors of $U_{11}$ and $V_{11}$ defined in \eqref{target}, respectively.
\item Finally, we estimate the conditional expectation of the remaining instantaneous volatility vector on the $D$th day, $E[(c_{D,n_1 + 1},\dots,c_{D,n}) | \mathcal{F}_{D,n_{1}}]$, by
$$
\widetilde{\Sigma}_{22} := (\widetilde{c}_{D,n_{1}+ 1},\dots, \widetilde{c}_{D,n}) = \widehat{\Sigma}_{21}\widehat{V}_{11}(\widehat{U}_{11}^{\top}\widehat{\Sigma}_{11}\widehat{V}_{11})^{-1}\widehat{U}_{11}^{\top}\widehat{\Sigma}_{12}.
$$
\end{enumerate}
\begin{remark} \label{remark_spot}
For the SIP procedure, we first need to estimate the instantaneous volatilities using the observed data.
In the high-frequency literature, several nonparametric methods for instantaneous volatility estimation have been proposed \citep{fan2008spot, figueroa2022kernel, foster1996continuous, kristensen2010nonparametric, mancini2015spot, todorov2019nonparametric, todorov2023bias, zu2014estimating}.
Any well-performing estimator of the instantaneous volatility, denoted by $\widehat{c}_{i,j}$, that satisfies Assumption \ref{assum_spotvol}(iii) can be used within our framework.
In the numerical study, we adopt the jump-robust pre-averaging method proposed by \cite{figueroa2022kernel}, which is described explicitly in \eqref{pre-spot} in Section \ref{simulation}.
\end{remark}
\begin{remark}
To implement the SIP procedure, we need to select the number of factors.
The number of latent factors, $r$, can be determined using data-driven methods \citep{ahn2013eigenvalue, bai2002determining, onatski2010determining}.
For instance, $r$ can be chosen by maximizing the singular value ratio or the largest singular value gap, given by $\max_{k\leq r_{\max}}\frac{\widehat{\lambda}_{k}}{\widehat{\lambda}_{k+1}}$ or $\max_{k\leq r_{\max}}(\widehat{\lambda}_{k}-\widehat{\lambda}_{k+1})$, where $\widehat{\lambda}_{1}\geq \widehat{\lambda}_{2}\geq \cdots \geq \widehat{\lambda}_{r_{\max}}$ are the leading singular values of $\widehat{\bfsym \Sigma}_{1\bullet} \equiv (\widehat{c}_{i,j})_{(D-1)\times n}$ for a predetermined maximum number of factors, $r_{\max}$.
In this paper, we choose $r=1$ in the numerical study (Sections \ref{simulation} and \ref{empiric}), following the eigenvalue ratio method proposed by \cite{ahn2013eigenvalue}.
\end{remark}
In summary, given the estimated instantaneous volatility process, we construct a structured matrix that includes the remaining future intraday volatility vector.
To forecast the future intraday volatility vector, we utilize the low-rank structure described above.
We refer to this procedure as the Structural Intraday-volatility Prediction (SIP) method.
The SIP method nonparametrically predicts the future instantaneous volatility vector using historical trading data along with current intraday market observations.
Crucially, by incorporating current intraday information, the prediction is formulated under a low-rank matrix assumption without relying on any specific parametric model, such as an autoregressive structure.
This feature makes the SIP procedure robust to model misspecification arising from incorrect parametric assumptions.
The numerical results in Sections \ref{simulation} and \ref{empiric} demonstrate that the SIP method effectively predicts the remaining future instantaneous volatility vector.
\section{Asymptotic Properties} \label{asymp}
In this section, we develop the asymptotic properties of the SIP estimator.
To achieve this, the following technical conditions are needed.
\begin{assum} \label{assum_spotvol} ~
\begin{itemize}
\item[(i)] For $k \leq r$, the eigengap satisfies $|\lambda_{k+1} - \lambda_{k}| = O_{P}(\sqrt{nD})$ and $\lambda_{r+1} = 0$.
\item[(ii)] For some bounded $\mu_{1}$ and $\mu_{2}$, the left and right singular vectors satisfy
$$
\max_{i\leq D, k\leq r}|u_{i,k}|^2 \leq \frac{\mu_1}{D}, \text{ and } \max_{j\leq n, k\leq r}|v_{j,k}|^2 \leq \frac{\mu_2}{n}.
$$
In addition, for each $k \leq r$, $u_{k} = (u_{1,k},\dots,u_{D,k})^{\top}$ and $v_{k} = (v_{1,k},\dots,v_{n,k})^{\top}$ are assumed to be constant vectors.
\item[(iii)] For each $i\leq D$ and $j\leq n$, the estimated instantaneous volatility $\widehat{c}_{i,j}$ satisfies
$$
\widehat{c}_{i,j}-c_{i,j} = \upsilon_{i,j} + \varsigma_{i,j},
$$
where $\upsilon_{i,j}$ follows the martingale difference sequence and $\varsigma_{i,j}$ is the estimation bias term such that $E(\upsilon_{i,j} | {\cal F}_{i,t_j})=0$ a.s.
For any $s \leq D$ and $j'\leq n$, we assume $E( c_{i,j'}\upsilon_{i,j} | {\cal F}_{i,t_j}) = 0$ and $E( c_{s,j} c_{s,j'}\upsilon_{i,j} | {\cal F}_{i,t_j}) = 0$ a.s.
In addition, for all $i\neq s$ and any $j,j'\leq n$, we assume $E(\upsilon_{i,j}\upsilon_{s,j'})=0$.
We further assume that $c_{i,j}$, $\upsilon_{i,j}$, and $\varsigma_{i,j}$ are sub-Gaussian random variables; that is, there exists a constant $K>0$ such that for all $i\leq D$, $j\leq n$, and all $\tau \in \mathbb{R}$,
$$
E[\exp(\tau \alpha_{i,j})] \leq \exp\left(\frac{K^{2}\tau^{2}}{2}\right) \quad \text{for} \quad \alpha_{i,j} \in \{c_{i,j}, m^{1/8} \upsilon_{i,j}, \rho_{m}^{-1} \varsigma_{i,j}\},
$$
where $\rho_m$ is the convergence rate of the bias term.
\item[(iv)] Let $\ensuremath{\boldsymbol{\Omega}}_{e_j,D\times D}$ denote the $D\times D$ covariance matrix of $\bfm e_{\cdot j} = (e_{1,j},\ldots, e_{D,j})^{\top}$ for each $j = 1,\dots, n$. Define $\ensuremath{\boldsymbol{\Omega}}_{j,D\times D} = \frac{1}{n_1}U_{\bullet 1}\Lambda_{\bullet 1}^{2}U_{\bullet 1}^{\top} + \ensuremath{\boldsymbol{\Omega}}_{e_j,D\times D}$. The instantaneous volatility matrix, $\bfsym \Sigma_{\bullet 1} \equiv [\Sigma_{11}^{\top} \Sigma_{21}^{\top}]^{\top}$, satisfies
$$
\|\bfsym \Sigma_{\bullet 1}\bfsym \Sigma_{\bullet 1}^{\top} - \sum_{j=1}^{n_{1}}\ensuremath{\boldsymbol{\Omega}}_{j,D\times D}\|_{\max} = O_{P}(\sqrt{n_{1}\log D}).
$$
Similarly, let $\ensuremath{\boldsymbol{\Omega}}_{e_i,n\times n}$ denote the $n\times n$ covariance matrix of $\bfm e_{i \cdot} = (e_{i,1},\ldots, e_{i,n})^{\top}$ for each $i = 1,\dots, D$. Define $\ensuremath{\boldsymbol{\Omega}}_{i, n\times n} = \frac{1}{D}\bfm V\ensuremath{\boldsymbol{\Lambda}}^{2}\bfm V^{\top} + \ensuremath{\boldsymbol{\Omega}}_{e_{i},n\times n}$. The instantaneous volatility matrix, $\bfsym \Sigma$, satisfies
$$
\|\bfsym \Sigma^{\top}\bfsym \Sigma - \sum_{i=1}^{D}\ensuremath{\boldsymbol{\Omega}}_{i,n\times n}\|_{\max} = O_{P}(\sqrt{D\log n}).
$$
\item[(v)] The covariance matrices $\ensuremath{\boldsymbol{\Omega}}_{e_{j},D\times D} =\mathrm{cov}(\bfm e_{\cdot j})$ for each $j = 1,\ldots,n$ and $\ensuremath{\boldsymbol{\Omega}}_{e_{i},n\times n} =\mathrm{cov}(\bfm e_{i \cdot})$ for each $i = 1,\ldots, D$, where $\bfm e_{\cdot j} = (e_{1,j},\ldots, e_{D,j})^{\top}$ and $\bfm e_{i \cdot} = (e_{i,1},\ldots, e_{i,n})^{\top}$, satisfy
\begin{align*}
\max_{j\leq n}\|\ensuremath{\boldsymbol{\Omega}}_{e_{j},D\times D}\|_{1} = O(\varphi_D) \quad \text{and} \quad
\max_{i \leq D}\|\ensuremath{\boldsymbol{\Omega}}_{e_{i},n\times n}\|_{1} = O(\varphi_n),
\end{align*}
where $\varphi_{D} = o(D)$ and $\varphi_{n} = o(n)$, respectively.
In addition, $\|\bfm E\| = O_{P}(\max(\sqrt{D},\sqrt{n}))$.
\item[(vi)] There exist positive constants $C_{1}$ and $C_{2}$ such that, for each $k\leq r$, $j, l\leq n$, and $i,q \leq D$,
\begin{align*} &|\mathrm{cov}(\sum_{i=1}^{D}u_{i,k}e_{i,j},\sum_{i=1}^{D}u_{i,k}e_{i,l})|\leq C_{1}\varphi_{D}\rho_{1}(|j-l|),\quad \text{where} \quad \sum_{h=1}^{\infty}\rho_{1}(h) <\infty, \\
&|\mathrm{cov}(\sum_{j=1}^{n}v_{j,k}e_{i,j},\sum_{j=1}^{n}v_{j,k}e_{q,j})|\leq C_{2}\varphi_{n}\rho_{2}(|i-q|),\quad \text{where} \quad \sum_{h=1}^{\infty}\rho_{2}(h) <\infty.
\end{align*}
In addition, for each $k \leq r$, $j \leq n$, and $i \leq D$, we have $E[(\sum_{i=1}^{D}u_{i,k}e_{i,j})^4] = O(\varphi_{D}^2)$ and $E[(\sum_{j=1}^{n}v_{j,k}e_{i,j})^4] = O(\varphi_{n}^2)$.
\end{itemize}
\end{assum}
\begin{remark}
Assumption \ref{assum_spotvol}(i) ensures a pervasive assumption, which is essential for the analysis of low-rank matrix structures (see, e.g., \citealp{candes2010matrix, cho2017asymptotic, fan2018large}).
Given that our instantaneous volatility matrix is of size $D \times n$, this eigengap condition suggests that the leading eigenvalues are associated with the low-rank structure scale on the order of $\sqrt{nD}$.
Assumption \ref{assum_spotvol}(ii) imposes the conventional incoherence condition, which ensures effective entrywise control in low-rank matrix estimation.
Assumption \ref{assum_spotvol}(iii) is the sub-Gaussian conditions that can be justified under some mild assumptions on the process $X$, the microstructure noise, and the kernel function \citep{figueroa2022kernel,kim2016asymptotic, kim2016sparse}.
Moreover, this condition also holds for heavy-tailed data given the bounded fourth moments condition using a truncated estimation scheme \citep{fan2018robust, shin2023adaptive}.
In practice, estimation of instantaneous volatility at time $t$ typically uses data up to and including time $t$, leading to the martingale difference property $E(\upsilon_{i,j} | {\cal F}_{i,t_j})=0$ almost surely.
The resulting instantaneous volatility estimator is asymptotically unbiased, and the bias vanishes at a rate faster than $m^{-\frac{1}{8}}$, which is the optimal convergence rate of the instantaneous volatility estimator in the presence of microstructure noise.
To justify the effectiveness of the smoothing technique, we require an uncorrelatedness condition: $E( c_{i,j'}\upsilon_{i,j} | {\cal F}_{i,t_j})= E( c_{s,j} c_{s,j'}\upsilon_{i,j} | {\cal F}_{i,t_j}) = 0$ almost surely.
This condition holds when the spot volatility process is adapted to ${\cal F}_{i,t_j}$ and the noise is independent.
Such conditions are mild and are commonly satisfied in standard time series settings like ARMA models.
Assumption \ref{assum_spotvol}(iv) is the element-wise convergence condition for analyzing large matrix inference.
Under the sub-Gaussian condition and mixing time dependency, we can easily obtain this condition (see \citealp{fan2018large, fan2018eigenvector, vershynin2010introduction, wang2017asymptotics}).
Assumption \ref{assum_spotvol}(v) imposes the conventional condition on the idiosyncratic covariance matrix, such as sparsity, which is commonly assumed in empirical applications \citep{boivin2006more, fan2016incorporating}.
We note that our framework allows for heterogeneity across both the intraday and interday dimensions.
Assumption \ref{assum_spotvol}(vi) requires weak temporal dependence in both the linear and quadratic forms of the projected error terms, as well as bounded fourth moments.
\end{remark}
We obtain the following elementwise convergence rate of the predicted instantaneous volatility using the SIP method.
\begin{thm} \label{main_thm}
Suppose that $r$ is fixed, $\log D = o(n_{1})$, $\log n = o(D)$, $\varphi_{D} \sqrt{n_{1}} \leq D$, $\varphi_{n}\sqrt{D} \leq n_{1}$ and Assumption \ref{assum_spotvol} hold.
As $D,n_1,n,m \rightarrow \infty$, we have
\begin{align*}
&\max_{j\leq n_2}\left|\widetilde{c}_{D,n_1+j} - E[c_{D,n_1+j} | \mathcal{F}_{D,n_1}]\right| = O_{P}\Bigg( \rho_{m} + \frac{1}{m^{\frac{1}{4}}} + \frac{\varphi_{D}}{D} + \sqrt{\frac{\log D}{n_1}}+ \frac{\varphi_{n}}{n_{1}}+ \sqrt{\frac{\log n}{D}}\Bigg).
\end{align*}
\end{thm}
Theorem \ref{main_thm} shows that the proposed SIP estimator consistently predicts the future instantaneous volatility process.
The convergence rate of the SIP estimator is given by $m^{-1/4} + \frac{1}{\sqrt{D}} + \frac{1}{\sqrt{n_1}}$ up to the log order and the sparsity level, when $\rho_m < m^{-1/4}$.
Each term in this rate corresponds to a distinct component of the estimation challenge:
$D^{-1/2}$ reflects the statistical cost of learning the interday (daily) time-series dynamics;
$n_1^{-1/2} $ captures the cost of estimating intraday volatility patterns;
$ m^{-1/4}$ arises from estimating the unobserved instantaneous volatility using high-frequency returns.
Notably, while the optimal nonparametric convergence rate for estimating instantaneous volatility is $m^{-1/8}$, the SIP method achieves a faster rate of $m^{-1/4}$.
This improvement stems from the low-rank structure, which allows the SIP estimator to benefit from a law-of-large-numbers-type averaging across the volatility surface.
In addition, the term $ \rho_m$ arises from the bias component in the spot volatility estimator.
Specifically, the bias originates from the drift term, which typically decays at the rate $m^{-1}$; however, due to the subsampling approach used to mitigate microstructure noise, the effective convergence rate $\rho_m$ is generally of order $m^{-1/2}$.
\section{Simulation Study} \label{simulation}
In this section, we examined the finite sample performance of the proposed SIP method through simulations.
We generated high-frequency observations that incorporate both interday HAR model feature and the intraday periodic pattern \citep{TIP-PCA} as follows: for $i= 1,\dots, D$, $s=0,\dots, m$, and $t_{s} = s/m$,
\begin{align*}
&Y_{i,t_{s}} = X_{i,t_{s}} + e_{i,t_{s}},\\
&dX_{i,t} = (\mu - \sigma_{i,t}^2/2)dt + \sigma_{i,t} dB_{i,t} + J_{i,t}dP_{i,t},\\
&\sigma_{i,t_{s}}^2 = \widetilde{\sigma}_{i}^2 h(t_{s}) + \varepsilon_{i,t_{s}},
\end{align*}
where we set microstructure noise as $e_{i,t_{s}} \sim \mathcal{N}(0,0.0005^2)$ and the initial value as $X_{1,0} = 1$; $B_{i,t}$ is a standard Brownian motion; for the jump part, we set $J_{i,t} \sim \mathcal{N}(-0.01, 0.02^2)$ and $P_{i,t+\Delta} - P_{i,t} \sim \text{Poisson}(36\Delta/252)$; $\widetilde{\sigma}_{i} = b_{0} + b_{1}\widetilde{\sigma}_{i-1} + b_{2}\frac{1}{5}\sum_{s=1}^{5}\widetilde{\sigma}_{i-s} + b_{3}\frac{1}{22}\sum_{s=1}^{22}\widetilde{\sigma}_{i-s} + \zeta_{i}$, $\zeta_{i} \sim \mathcal{N}(0,1)$, $h(t_{s}) = \gamma_{0} + \gamma_{1}(t_{s}-0.6)^2$, and $\varepsilon_{i,t_{s}} = q(t_{s})\xi_{i,t_{s}}$, where $q(t_{s})^{2} = 0.1+0.5(2t_{s}-1)^{2}$ and $\xi_{i,t_{s}} \sim \mathcal{N}(0,0.01^2)$.
The model parameters were set to be
\begin{align*}
\mu = 0.05/252, \, \gamma_{0} = 0.04/252, \, \gamma_{1} = 0.5/252, \\
b_{0} = 0.5, \, b_{1} =0.372, \, b_{2} = 0.343, \, b_{3} = 0.224.
\end{align*}
The normalized parameter values above imply the daily time unit, and we adapted the estimated coefficients studied in \cite{corsi2009simple} to generate $\widetilde{\sigma}_{i}$.
We note that in each simulation, the instantaneous volatility process was generated until all instantaneous volatility values were positive, based on the data-generating process described above.
We set $m =$ 23,400, which indicates that the data are observed every second over a period of 6.5 trading hours per day.
For each simulation, we used the jump robust pre-averaging method \citep{figueroa2022kernel} to estimate the instantaneous volatility, $c_{i,\tau}=\sigma_{i,\frac{\tau}{n}}^2$, at a frequency of every 5 minutes (i.e., $n=78$) for each $i$-th day as follows: for $\tau = 1, \dots, n$,
\begin{equation}\label{pre-spot}
\widehat{c}_{i,\tau} = \frac{1}{\phi_{k_{n}}(g)}\sum_{s=1}^{m-k_{n}+1}K_{b_{m}}(t_{s-1}-\frac{\tau}{n})\left(\bar{Y}_{i,s}^2 - \frac{1}{2}\widehat{Y}_{i,s}\right)\mathbf{1}_{\{|\bar{Y}_{i,s}|\leq \nu_{m}\}},
\end{equation}
where $K_{b}(x) = K(x/b)/b$, the bandwidth size $b_{m} = 1/n$, the weight function $g(x) = 2x \wedge (1-x)$,
\begin{align*}
&\bar{Y}_{i,s} = \sum_{l=1}^{k_{m}-1}g\left(\frac{l}{k_{m}}\right)(Y_{i,t_{s+l}}-Y_{i,t_{s+l-1}}), \qquad \phi_{k_{m}}(g) = \sum_{i=1}^{k_{m}}g\left(\frac{i}{k_{m}}\right)^2,\\
&\widehat{Y}_{i,s} = \sum_{l=1}^{k_{m}}\biggl\{g\left(\frac{l}{k_{m}}\right)-g\left(\frac{l-1}{k_{m}}\right)\biggl\}^2 (Y_{i,t_{s+l}}-Y_{i,t_{s+l-1}})^2,
\end{align*}
$\mathbf{1}_{\{\cdot\}}$ is an indicator function, and $\nu_{m} = 1.8\sqrt{\text{BPV}}(k_{m}/m)^{0.47}$, where the bipower variation $\text{BPV} = \frac{\pi}{2}\sum_{s=2}^{m}|Y_{i,t_{s-1}} - Y_{i,t_{s-2}}| |Y_{i,t_{s}} - Y_{i,t_{s-1}}|$.
We used the uniform kernel function and the data-driven approach to obtain the preaveraging window size, $k_m$, as suggested in Section 3.1 of \cite{figueroa2022kernel}.
With the instantaneous volatility estimates available up to the $n_1$-th intraday data point on the $D$th day, we examined the out-of-sample performance of estimating the remaining $(1\times n_2)$ instantaneous volatility vector, where $n_1 = \omega n$ and $n_2 = n- n_1$.
For comparison, we considered the SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA methods in predicting $c_{D,\tau}$ for $\tau = n_1+1, \ldots, n$.
AVE represents estimates obtained from the column mean of $\widehat{\bfsym \Sigma}_{D-1,n}$, while AR generates predictions using an autoregressive model of order 1, applied to each column of $\widehat{\bfsym \Sigma}_{D-1,n}$.
SARIMA provides $n_2$-step-ahead forecasts using a seasonal ARIMA(1,1,1) model \citep{sheppard2010financial}, where the seasonal cycle length is set to $n$.
HAR-D extends the HAR model by incorporating diurnal effects and prior intraday components in addition to daily, weekly, and monthly realized volatilities \citep{zhang2024volatility}.
XGBoost \citep{chen2016xgboost}, a decision-tree-based ensemble learning method, accounts for nonlinear dependencies, using the same hyperparameters as in \cite{zhang2024volatility}.
Since SARIMA, HAR-D, and XGBoost require a time series format, we first vectorized the estimates as $(\widehat{c}_{1,1},\dots, \widehat{c}_{D,n_1})$ before applying these methods.
PC corresponds to the last row of the estimated low-rank matrix obtained from the best rank-$r$ approximation of $\widehat{\bfsym \Sigma}_{D-1,n}$.
TIP-PCA uses ex-post daily, weekly, and monthly realized volatilities and the intraday time sequence as observable covariates that explain the left and right singular vectors, as proposed by \cite{TIP-PCA}, given $\widehat{\bfsym \Sigma}_{D-1,n}$.
For SIP, TIP-PCA, and PC, we set rank 1 as determined by the eigenvalue ratio method \citep{ahn2013eigenvalue}.
We generated high-frequency data with $m = 23,400$ over 200 consecutive days, using subsampled log prices from the last $D =$ 50, 100, 150, and 200 days.
To assess prediction accuracy, we computed the mean squared prediction error (MSPE) as
$$
\frac{1}{n}\sum_{\tau=n_1+1}^{n}(\widetilde{c}_{D,\tau}-c_{D,\tau})^{2},
$$
where $\widetilde{c}_{D,\tau}$ denotes the instantaneous volatility estimator of each method above.
Finally, we calculated the sample averages of MSPEs across 500 simulations.
\begin{figure}
\includegraphics[width=\linewidth]{Varying_D.eps}
\centering
\caption{MSPE$\times 10^9$ for the SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA against $D$ with fixed $\omega = \{0.1, 0.5\}$.} \label{varying_D}
\end{figure}
Figure \ref{varying_D} presents the average MSPEs of the estimators for the remaining future intraday instantaneous volatility process across different values of dimensionality, $D = {50, 100, 150, 200}$, with fixed $n$ and $\omega$.
Figure \ref{varying_rho}, on the other hand, plots the average MSPEs against varying values of $\omega = {0.05, 0.1, 0.2, \dots, 0.9}$, while keeping $D$ and $n$ fixed.
Note that the target future volatility remains consistent across different $D$ values, as the same subsampled data are used in each simulation setting.
In contrast, in Figure \ref{varying_rho}, the MSPEs exhibit a U-shaped pattern as $\omega$ varies.
This behavior arises because the target future volatility differs with each value of $\omega$, and the underlying data-generating process follows a U-shaped intraday volatility pattern.
As a result, direct comparisons of MSPE values across different $\omega$ values are not appropriate.
Instead, MSPEs should be interpreted relatively across competing methods for each fixed value of $\omega$.
\begin{figure}
\includegraphics[width=\linewidth]{Varying_rho.eps}
\centering
\caption{MSPE$\times 10^9$ for the SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA against $\omega$ with fixed $D=\{50, 100\}$.} \label{varying_rho}
\end{figure}
Figures \ref{varying_D} and \ref{varying_rho} demonstrate that the SIP method consistently achieves the best performance among the competing approaches.
This may be because by incorporating information from previous intraday volatility patterns on the $D$th day, SIP effectively predicts the remaining future instantaneous volatility process through a nonparametric framework.
Moreover, the MSPEs of SIP tend to decrease as the number of daily observations $D$ increases, which supports the theoretical results established in Section \ref{asymp}.
In contrast, the TIP-PCA, PC, AVE, and AR methods perform poorly, as they are unable to leverage current intraday observations on the $D$th day within their estimation procedures.
In Figure \ref{varying_rho}, we observe that the performance of SARIMA improves as $\omega$ increases.
This suggests that, given sufficient real-time intraday information, SARIMA can better incorporate both periodic intraday patterns and daily autoregressive dynamics.
Nevertheless, SIP outperforms SARIMA and all other competing methods across the entire range of $\omega$, particularly when $\omega$ is small.
\section{Empirical Study} \label{empiric}
In this section, we applied the proposed SIP method to predict intraday instantaneous volatility vector using real high-frequency trading data.
Specifically, we obtained intraday data of the S\&P 500 index ETF (SPY) and the 11 Global Industry Classification Standard (GICS) sector ETFs (XLC, XLY, XLP, XLE, XLF, XLV, XLI, XLB, XLRE, XLK, and XLU) from July 2021 to June 2022, sourced from the Trade and Quote (TAQ) database in the Wharton Research Data Services (WRDS) system.
We used high-frequency data subsampled at 1-second intervals, excluding trading days with early market closures.
This subsampling mitigates the effects of irregular observation times \citep{li2023robust}.
Using the log prices, we implemented the jump robust pre-averaging estimation procedure, as defined in Section \ref{simulation}, to estimate the instantaneous variance at 5-minute intervals.
Using the in-sample period data, we applied SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA, as described in Section \ref{simulation}, to predict the remaining instantaneous volatilities given each $\omega = \{0.1, 0.5, 0.9\}$.
We employed the rolling window scheme with a 63-day (i.e., one quarter) in-sample period, where the last $n_2 = (1-\omega)n$ instantaneous volatilities on the final in-sample day are the target volatility.
The out-of-sample period was from October 2021 to June 2022 (i.e., $q=189$ days).
To evaluate the performance of the predicted instantaneous volatility, we employed the mean squared prediction errors (MSPE) and QLIKE loss function \citep{patton2011volatility}, defined as follows: for each $\omega$ and $n_2 = (1-\omega)n$,
\begin{align*}
&\text{MSPE} = \frac{1}{qn_2}\sum_{i=1}^{q}\sum_{j = 1}^{n_2}(\widetilde{c}_{i,j} - \widehat{c}_{i,j})^{2},\\
&\text{QLIKE} = \frac{1}{qn_2}\sum_{i=1}^{q}\sum_{j = 1}^{n_2}(\log{\widetilde{c}_{i,j}} + \frac{\widehat{c}_{i,j}}{\widetilde{c}_{i,j}}),
\end{align*}
where $\widetilde{c}_{i,j}$ represents the predicted instantaneous volatility on the $j$th intraday time point of the $i$th out-of-sample day, obtained from one of the SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA methods.
We predicted the remaining future conditional expected instantaneous volatilities using in-sample period data.
Since the true conditional expected instantaneous volatility is unknown, we assessed the significance of differences in prediction performances using the Diebold and Mariano (DM) test \citep{diebold2002comparing} based on both MSPE and QLIKE.
We compared the proposed SIP method against alternative approaches.
Given that multiple hypothesis tests were conducted repeatedly throughout this section, it is crucial to control the False Discovery Rate (FDR).
To account for this, we applied the Benjamini–Hochberg (BH) procedure \citep{benjamini1995controlling} at a significance level of $\alpha = 0.05$ to adjust the $p$-values and mitigate the risk of false positives across all hypothesis tests performed in this section.
\begin{table}[htbp]
\centering
\caption{MSPEs ($\times 10^9$) for SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA across 12 ETFs under varying values of $\omega \in \{0.1, 0.5, 0.9\}$.} \label{MSPE}
\scalebox{0.91}{
\begin{tabular}{lllllllll}
\hline
& SIP & AVE & AR & SARIMA & HAR-D & XGBoost & PC & TIP-PCA \\
\midrule
\multicolumn{9}{c}{$\omega = 0.1$} \\
\cmidrule(lr){2-9}
SPY & \textbf{2.259} & 3.555\textsuperscript{***} & 3.647\textsuperscript{***} & 3.099\textsuperscript{***} & 3.590\textsuperscript{***} & 2.858\textsuperscript{***} & 2.839\textsuperscript{***} & 2.693\textsuperscript{***} \\
XLC & \textbf{2.981} & 4.669\textsuperscript{***} & 5.073\textsuperscript{***} & 4.757\textsuperscript{***} & 5.299\textsuperscript{***} & 4.478\textsuperscript{***} & 3.873\textsuperscript{***} & 3.681\textsuperscript{***} \\
XLY & \textbf{6.959} & 12.196\textsuperscript{***} & 11.554\textsuperscript{***} & 11.769\textsuperscript{***} & 13.272\textsuperscript{***} & 11.328\textsuperscript{***} & 9.838\textsuperscript{***} & 9.472\textsuperscript{***} \\
XLP & \textbf{0.569} & 0.833\textsuperscript{***} & 0.816\textsuperscript{***} & 0.793\textsuperscript{***} & 0.784\textsuperscript{***} & 0.818\textsuperscript{***} & 0.641\textsuperscript{***} & 0.634\textsuperscript{***} \\
XLE & \textbf{7.089} & 10.871\textsuperscript{***} & 10.362\textsuperscript{***} & 12.145\textsuperscript{***} & 10.327\textsuperscript{***} & 8.715\textsuperscript{***} & 8.418\textsuperscript{***} & 8.282\textsuperscript{***} \\
XLF & \textbf{1.772} & 3.534\textsuperscript{***} & 3.079\textsuperscript{***} & 2.661\textsuperscript{***} & 3.101\textsuperscript{***} & 2.678\textsuperscript{***} & 2.595\textsuperscript{***} & 2.500\textsuperscript{***} \\
XLV & \textbf{0.884} & 1.316\textsuperscript{***} & 1.197\textsuperscript{***} & 1.071\textsuperscript{***} & 1.251\textsuperscript{***} & 1.226\textsuperscript{***} & 0.994\textsuperscript{***} & 0.995\textsuperscript{***} \\
XLI & \textbf{1.642} & 2.626\textsuperscript{***} & 2.626\textsuperscript{***} & 2.204\textsuperscript{***} & 2.511\textsuperscript{***} & 2.358\textsuperscript{***} & 1.900\textsuperscript{***} & 1.911\textsuperscript{***} \\
XLB & \textbf{2.024} & 3.114\textsuperscript{***} & 3.115\textsuperscript{***} & 2.836\textsuperscript{***} & 3.005\textsuperscript{***} & 2.778\textsuperscript{***} & 2.434\textsuperscript{***} & 2.393\textsuperscript{***} \\
XLRE & \textbf{1.558} & 2.422\textsuperscript{***} & 2.343\textsuperscript{***} & 2.087\textsuperscript{***} & 2.325\textsuperscript{***} & 2.188\textsuperscript{***} & 1.945\textsuperscript{***} & 1.918\textsuperscript{***} \\
XLK & \textbf{6.256} & 10.667\textsuperscript{***} & 10.154\textsuperscript{***} & 8.183\textsuperscript{***} & 10.816\textsuperscript{***} & 10.222\textsuperscript{***} & 8.273\textsuperscript{***} & 8.072\textsuperscript{***} \\
XLU & \textbf{0.814} & 1.175\textsuperscript{***} & 1.139\textsuperscript{***} & 1.108\textsuperscript{***} & 1.160\textsuperscript{***} & 1.223\textsuperscript{***} & 0.988\textsuperscript{***} & 0.981\textsuperscript{***} \\
\midrule
\multicolumn{9}{c}{$\omega = 0.5$} \\
\cmidrule(lr){2-9}
SPY & \textbf{3.012} & 4.114\textsuperscript{***} & 4.488\textsuperscript{***} & 3.125\textsuperscript{***} & 4.212\textsuperscript{***} & 3.446\textsuperscript{***} & 3.345\textsuperscript{***} & 3.234\textsuperscript{***} \\
XLC & \textbf{3.675} & 4.857\textsuperscript{***} & 5.831\textsuperscript{***} & 3.907\textsuperscript{***} & 5.868\textsuperscript{***} & 4.668\textsuperscript{***} & 4.223\textsuperscript{***} & 4.010\textsuperscript{***} \\
XLY & \textbf{8.327} & 12.452\textsuperscript{***} & 12.470\textsuperscript{***} & 9.307\textsuperscript{***} & 13.671\textsuperscript{***} & 10.188\textsuperscript{***} & 10.316\textsuperscript{***} & 10.136\textsuperscript{***} \\
XLP & \textbf{0.703} & 1.005\textsuperscript{***} & 1.041\textsuperscript{***} & 0.742\textsuperscript{***} & 0.945\textsuperscript{***} & 0.972\textsuperscript{***} & 0.771\textsuperscript{***} & 0.768\textsuperscript{***} \\
XLE & \textbf{6.196} & 8.504\textsuperscript{***} & 8.301\textsuperscript{***} & 6.747\textsuperscript{***} & 8.793\textsuperscript{***} & 6.685\textsuperscript{***} & 7.097\textsuperscript{***} & 7.035\textsuperscript{***} \\
XLF & \textbf{2.214} & 3.892\textsuperscript{***} & 3.587\textsuperscript{***} & 2.627\textsuperscript{***} & 3.545\textsuperscript{***} & 2.861\textsuperscript{***} & 2.902\textsuperscript{***} & 2.834\textsuperscript{***} \\
XLV & \textbf{1.022} & 1.534\textsuperscript{***} & 1.439\textsuperscript{***} & 1.044 & 1.459\textsuperscript{***} & 1.757\textsuperscript{***} & 1.179\textsuperscript{***} & 1.177\textsuperscript{***} \\
XLI & \textbf{1.911} & 2.904\textsuperscript{***} & 3.171\textsuperscript{***} & 2.013\textsuperscript{***} & 2.750\textsuperscript{***} & 2.457\textsuperscript{***} & 2.106\textsuperscript{***} & 2.129\textsuperscript{***} \\
XLB & \textbf{2.442} & 3.421\textsuperscript{***} & 3.611\textsuperscript{***} & 2.601\textsuperscript{***} & 3.245\textsuperscript{***} & 3.029\textsuperscript{***} & 2.667\textsuperscript{***} & 2.656\textsuperscript{***} \\
XLRE & \textbf{1.850} & 2.765\textsuperscript{***} & 2.777\textsuperscript{***} & 2.070\textsuperscript{***} & 2.573\textsuperscript{***} & 2.536\textsuperscript{***} & 2.218\textsuperscript{***} & 2.190\textsuperscript{***} \\
XLK & \textbf{7.727} & 11.448\textsuperscript{***} & 11.655\textsuperscript{***} & 8.401\textsuperscript{***} & 12.085\textsuperscript{***} & 9.823\textsuperscript{***} & 9.191\textsuperscript{***} & 9.067\textsuperscript{***} \\
XLU & \textbf{0.937} & 1.344\textsuperscript{***} & 1.334\textsuperscript{***} & 0.968\textsuperscript{**} & 1.335\textsuperscript{***} & 1.250\textsuperscript{***} & 1.113\textsuperscript{***} & 1.126\textsuperscript{***} \\
\midrule
\multicolumn{9}{c}{$\omega = 0.9$} \\
\cmidrule(lr){2-9}
SPY & \textbf{2.302} & 4.383\textsuperscript{***} & 3.623\textsuperscript{***} & 2.715 & 4.240\textsuperscript{***} & 2.555 & 3.295\textsuperscript{***} & 3.387\textsuperscript{***} \\
XLC & \textbf{2.703} & 4.956\textsuperscript{***} & 4.519\textsuperscript{***} & 3.286\textsuperscript{*} & 5.361\textsuperscript{***} & 2.704 & 3.892\textsuperscript{***} & 3.955\textsuperscript{***} \\
XLY & \textbf{5.483} & 11.780\textsuperscript{***} & 9.936\textsuperscript{***} & 7.261\textsuperscript{**} & 13.361\textsuperscript{***} & 5.768 & 9.009\textsuperscript{***} & 9.712\textsuperscript{***} \\
XLP & \textbf{0.592} & 1.228\textsuperscript{***} & 1.016\textsuperscript{***} & 0.763\textsuperscript{**} & 1.116\textsuperscript{***} & 0.749\textsuperscript{**} & 0.852\textsuperscript{***} & 0.910\textsuperscript{***} \\
XLE & 4.317 & 6.357\textsuperscript{***} & 5.428\textsuperscript{***} & \textbf{3.936} & 6.476\textsuperscript{***} & 4.050 & 5.601\textsuperscript{***} & 6.117\textsuperscript{***} \\
XLF & \textbf{1.548} & 4.240\textsuperscript{***} & 3.015\textsuperscript{***} & 1.908\textsuperscript{**} & 3.312\textsuperscript{***} & 1.690 & 2.822\textsuperscript{***} & 2.892\textsuperscript{***} \\
XLV & \textbf{0.943} & 1.924\textsuperscript{***} & 1.555\textsuperscript{***} & 1.102\textsuperscript{*} & 1.722\textsuperscript{***} & 1.271\textsuperscript{***} & 1.356\textsuperscript{***} & 1.445\textsuperscript{***} \\
XLI & \textbf{1.491} & 3.456\textsuperscript{***} & 2.417\textsuperscript{***} & 1.679\textsuperscript{*} & 2.889\textsuperscript{***} & 1.984\textsuperscript{***} & 2.258\textsuperscript{***} & 2.478\textsuperscript{***} \\
XLB & \textbf{2.011} & 4.099\textsuperscript{***} & 3.362\textsuperscript{***} & 2.405\textsuperscript{**} & 3.710\textsuperscript{***} & 2.172 & 2.931\textsuperscript{***} & 3.188\textsuperscript{***} \\
XLRE & \textbf{1.722} & 3.692\textsuperscript{***} & 3.138\textsuperscript{***} & 1.959 & 3.373\textsuperscript{***} & 2.652\textsuperscript{***} & 2.860\textsuperscript{***} & 2.925\textsuperscript{***} \\
XLK & \textbf{5.594} & 10.868\textsuperscript{***} & 9.580\textsuperscript{***} & 7.009\textsuperscript{*} & 11.661\textsuperscript{***} & 7.044 & 8.139\textsuperscript{***} & 8.830\textsuperscript{***} \\
XLU & \textbf{0.916} & 1.803\textsuperscript{***} & 1.481\textsuperscript{***} & 1.211\textsuperscript{**} & 1.748\textsuperscript{***} & 1.182\textsuperscript{**} & 1.398\textsuperscript{***} & 1.523\textsuperscript{***} \\
\bottomrule
\end{tabular}}
\centering
\begin{tablenotes}
\item Note: Bold numbers indicate the lowest MSPE for each ETF, while $^{***}$, $^{**}$, and $^{*}$ indicate rejection of the null hypothesis against SIP at the 1\%, 5\%, and 10\% significance levels, respectively, based on the Diebold-Mariano (DM) test.
\end{tablenotes}
\end{table}
\begin{table}[htbp]
\centering
\caption{QLIKEs for SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA across 12 ETFs under varying values of $\omega \in \{0.1, 0.5, 0.9\}$.} \label{QLIKE}
\scalebox{0.91}{
\begin{tabular}{lllllllll}
\hline
& SIP & AVE & AR & SARIMA & HAR-D & XGBoost & PC & TIP-PCA \\
\midrule
\multicolumn{9}{c}{$\omega = 0.1$} \\
\cmidrule(lr){2-9}
SPY & \textbf{-9.287} & -8.970\textsuperscript{***} & -9.114\textsuperscript{***} & -5.377\textsuperscript{***} & -7.974\textsuperscript{***} & -9.208\textsuperscript{***} & -9.213\textsuperscript{***} & -9.071\textsuperscript{***} \\
XLC & \textbf{-9.116} & -8.863\textsuperscript{***} & -8.942\textsuperscript{***} & 0.942\textsuperscript{***} & -8.255\textsuperscript{***} & -8.990\textsuperscript{***} & -9.046\textsuperscript{***} & -8.955\textsuperscript{***} \\
XLY & \textbf{-8.521} & -8.243\textsuperscript{***} & -8.347\textsuperscript{***} & 10.962\textsuperscript{***} & -6.636\textsuperscript{***} & -8.285\textsuperscript{***} & -8.452\textsuperscript{***} & -8.435\textsuperscript{***} \\
XLP & \textbf{-9.576} & -9.436\textsuperscript{***} & -9.489\textsuperscript{***} & -5.412\textsuperscript{***} & -9.166\textsuperscript{***} & -9.531\textsuperscript{***} & -9.547\textsuperscript{***} & -9.543\textsuperscript{***} \\
XLE & \textbf{-8.210} & -8.095\textsuperscript{***} & -8.131\textsuperscript{***} & -4.371\textsuperscript{***} & -7.621\textsuperscript{***} & -8.175\textsuperscript{***} & -8.179\textsuperscript{***} & -8.179\textsuperscript{***} \\
XLF & \textbf{-8.490} & -8.385\textsuperscript{***} & -8.434\textsuperscript{***} & 1.489\textsuperscript{***} & -8.037\textsuperscript{***} & -8.465\textsuperscript{***} & -8.454\textsuperscript{***} & -8.455\textsuperscript{***} \\
XLV & \textbf{-9.563} & -9.375\textsuperscript{***} & -9.452\textsuperscript{***} & -5.433\textsuperscript{***} & -8.594\textsuperscript{***} & -9.451\textsuperscript{***} & -9.539\textsuperscript{***} & -9.495\textsuperscript{***} \\
XLI & \textbf{-9.286} & -9.057\textsuperscript{***} & -9.167\textsuperscript{***} & -6.441\textsuperscript{***} & -8.464\textsuperscript{***} & -9.201\textsuperscript{***} & -9.269\textsuperscript{***} & -8.926\textsuperscript{***} \\
XLB & \textbf{-9.239} & -9.016\textsuperscript{***} & -9.088\textsuperscript{***} & -7.276\textsuperscript{***} & -7.454\textsuperscript{***} & -9.135\textsuperscript{***} & -9.202\textsuperscript{***} & -9.184\textsuperscript{***} \\
XLRE & \textbf{-9.197} & -9.036\textsuperscript{***} & -9.077\textsuperscript{***} & -5.944\textsuperscript{***} & -7.678\textsuperscript{***} & -9.082\textsuperscript{***} & -9.137\textsuperscript{***} & -9.140\textsuperscript{***} \\
XLK & \textbf{-8.661} & -8.374\textsuperscript{***} & -8.501\textsuperscript{***} & -6.775\textsuperscript{***} & -7.012\textsuperscript{***} & -8.569\textsuperscript{***} & -8.587\textsuperscript{***} & -8.563\textsuperscript{***} \\
XLU & \textbf{-9.222} & -9.119\textsuperscript{***} & -9.144\textsuperscript{***} & -7.066\textsuperscript{***} & -8.418\textsuperscript{***} & -9.160\textsuperscript{***} & -9.178\textsuperscript{***} & -9.185\textsuperscript{***} \\
\midrule
\multicolumn{9}{c}{$\omega = 0.5$} \\
\cmidrule(lr){2-9}
SPY & \textbf{-9.337} & -9.006\textsuperscript{***} & -9.161\textsuperscript{***} & -9.123\textsuperscript{***} & -8.207\textsuperscript{***} & -9.225\textsuperscript{***} & -9.271\textsuperscript{***} & -9.083\textsuperscript{***} \\
XLC & \textbf{-9.181} & -8.918\textsuperscript{***} & -9.001\textsuperscript{***} & -8.955\textsuperscript{***} & -8.168\textsuperscript{***} & -8.986\textsuperscript{***} & -9.111\textsuperscript{***} & -9.001\textsuperscript{***} \\
XLY & \textbf{-8.640} & -8.340\textsuperscript{***} & -8.453\textsuperscript{***} & -8.033\textsuperscript{**} & -5.726\textsuperscript{***} & -8.543\textsuperscript{***} & -8.572\textsuperscript{***} & -8.541\textsuperscript{***} \\
XLP & \textbf{-9.623} & -9.456\textsuperscript{***} & -9.520\textsuperscript{***} & -9.577\textsuperscript{***} & -9.082\textsuperscript{***} & -9.559\textsuperscript{***} & -9.591\textsuperscript{***} & -9.582\textsuperscript{***} \\
XLE & \textbf{-8.388} & -8.265\textsuperscript{***} & -8.306\textsuperscript{***} & -8.257\textsuperscript{***} & -7.639\textsuperscript{***} & -8.355\textsuperscript{***} & -8.348\textsuperscript{***} & -8.346\textsuperscript{***} \\
XLF & \textbf{-8.559} & -8.439\textsuperscript{***} & -8.493\textsuperscript{***} & -8.532\textsuperscript{***} & -8.046\textsuperscript{**} & -8.529\textsuperscript{***} & -8.517\textsuperscript{***} & -8.517\textsuperscript{***} \\
XLV & \textbf{-9.657} & -9.417\textsuperscript{***} & -9.512\textsuperscript{***} & -9.430\textsuperscript{**} & -8.067\textsuperscript{***} & -9.524\textsuperscript{***} & -9.608\textsuperscript{***} & -9.547\textsuperscript{***} \\
XLI & \textbf{-9.369} & -9.100\textsuperscript{***} & -9.223\textsuperscript{***} & -9.276\textsuperscript{***} & -8.140\textsuperscript{***} & -9.287\textsuperscript{***} & -9.341\textsuperscript{***} & -9.061\textsuperscript{***} \\
XLB & \textbf{-9.316} & -9.064\textsuperscript{***} & -9.147\textsuperscript{***} & -9.177\textsuperscript{***} & -6.553\textsuperscript{***} & -9.212\textsuperscript{***} & -9.279\textsuperscript{***} & -9.257\textsuperscript{***} \\
XLRE & \textbf{-9.232} & -9.048\textsuperscript{***} & -9.091\textsuperscript{***} & -8.969\textsuperscript{***} & -6.833\textsuperscript{***} & -9.111\textsuperscript{***} & -9.158\textsuperscript{***} & -9.161\textsuperscript{***} \\
XLK & \textbf{-8.764} & -8.446\textsuperscript{***} & -8.590\textsuperscript{***} & -8.561\textsuperscript{***} & -6.418\textsuperscript{***} & -8.429\textsuperscript{*} & -8.688\textsuperscript{***} & -8.652\textsuperscript{***} \\
XLU & \textbf{-9.294} & -9.170\textsuperscript{***} & -9.203\textsuperscript{***} & -9.231\textsuperscript{***} & -7.461\textsuperscript{***} & -9.237\textsuperscript{***} & -9.241\textsuperscript{***} & -9.245\textsuperscript{***} \\
\midrule
\multicolumn{9}{c}{$\omega = 0.9$} \\
\cmidrule(lr){2-9}
SPY & \textbf{-9.065} & -8.754\textsuperscript{***} & -8.892\textsuperscript{***} & -9.063 & -8.485\textsuperscript{***} & -9.013\textsuperscript{**} & -8.947\textsuperscript{***} & -8.795\textsuperscript{***} \\
XLC & \textbf{-8.858} & -8.606\textsuperscript{***} & -8.702\textsuperscript{***} & -8.819\textsuperscript{**} & -8.057 & -8.803\textsuperscript{***} & -8.760\textsuperscript{***} & -8.662\textsuperscript{***} \\
XLY & -8.425 & -8.138\textsuperscript{***} & -8.261\textsuperscript{***} & -8.413 & -5.270\textsuperscript{***} & \textbf{-8.435} & -8.333\textsuperscript{***} & -8.304\textsuperscript{***} \\
XLP & -9.250 & -9.089\textsuperscript{***} & -9.155\textsuperscript{***} & -9.253 & \textbf{-10.141} & -9.224\textsuperscript{***} & -9.199\textsuperscript{***} & -9.187\textsuperscript{***} \\
XLE & -8.289 & -8.187\textsuperscript{***} & -8.224\textsuperscript{***} & -8.281 & -8.118\textsuperscript{***} & \textbf{-8.300}\textsuperscript{**} & -8.229\textsuperscript{***} & -8.229\textsuperscript{***} \\
XLF & \textbf{-8.404} & -8.268\textsuperscript{***} & -8.340\textsuperscript{***} & -8.402 & -8.303 & -8.400 & -8.349\textsuperscript{***} & -8.347\textsuperscript{***} \\
XLV & \textbf{-9.265} & -9.049\textsuperscript{***} & -9.124\textsuperscript{***} & -9.245 & -6.789\textsuperscript{*} & -9.236\textsuperscript{***} & -9.196\textsuperscript{***} & -9.146\textsuperscript{***} \\
XLI & \textbf{-8.993} & -8.738\textsuperscript{***} & -8.869\textsuperscript{***} & -8.962 & -8.556\textsuperscript{***} & -8.972\textsuperscript{***} & -8.930\textsuperscript{***} & -8.623\textsuperscript{***} \\
XLB & \textbf{-8.955} & -8.706\textsuperscript{***} & -8.807\textsuperscript{***} & -8.924 & -8.773 & -8.861 & -8.893\textsuperscript{***} & -8.858\textsuperscript{***} \\
XLRE & \textbf{-8.804} & -8.621\textsuperscript{***} & -8.645\textsuperscript{***} & -8.680 & -1.301 & -8.709\textsuperscript{***} & -8.722\textsuperscript{***} & -8.702\textsuperscript{***} \\
XLK & \textbf{-8.554} & -8.263\textsuperscript{***} & -8.391\textsuperscript{***} & -8.481\textsuperscript{*} & -8.245\textsuperscript{***} & -8.544 & -8.454\textsuperscript{***} & -8.418\textsuperscript{***} \\
XLU & \textbf{-8.919} & -8.794\textsuperscript{***} & -8.824\textsuperscript{***} & -8.914 & -8.757\textsuperscript{***} & -8.893\textsuperscript{***} & -8.853\textsuperscript{***} & -8.847\textsuperscript{***} \\
\bottomrule
\end{tabular}}
\centering
\begin{tablenotes}
\item Note: Bold numbers indicate the lowest QLIKE for each ETF, while $^{***}$, $^{**}$, and $^{*}$ indicate rejection of the null hypothesis against SIP at the 1\%, 5\%, and 10\% significance levels, respectively, based on the Diebold-Mariano (DM) test.
\end{tablenotes}
\end{table}
Tables \ref{MSPE} and \ref{QLIKE} report the MSPE and QLIKE results, respectively, under varying values of $\omega \in \{0.1, 0.5, 0.9\}$.
As discussed in Section \ref{simulation}, the target future volatility changes with $\omega$.
Consequently, direct comparisons of MSPE and QLIKE values across different $\omega$ values are inappropriate.
Instead, performance should be evaluated relative to other methods for each given $\omega$.
Tables \ref{MSPE}--\ref{QLIKE} show that the SIP method consistently outperforms other methods across different values of $\omega$.
This implies that incorporating the early intraday trading data enhances the accuracy of instantaneous volatility prediction.
We note that the performance of SARIMA and XGBoost improves when $\omega = 0.9$.
This may be attributed to the availability of a larger amount of information on the last trading day, which improves their prediction performance for a smaller set of target volatilities.
\begin{table}[htbp]
\centering
\caption{Number of cases where the adjusted $p$-value exceeds 0.05 for SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA across 12 ETFs at each $q_0 = \{0.01, 0.02, 0.05, 0.1, 0.2\}$, based on the LRuc, LRcc, and DQ tests under varying values of $\omega \in \{0.1, 0.5, 0.9\}$.}
\label{VaR}
\scalebox{0.91}{
\begin{tabular}{lrrrrr|rrrrr|rrrrr}
\toprule
& \multicolumn{5}{c}{LRuc} & \multicolumn{5}{c}{LRcc} & \multicolumn{5}{c}{DQ} \\
\cmidrule(lr){2-6} \cmidrule(lr){7-11} \cmidrule(lr){12-16}
$q_{0}$ & 0.01 & 0.02 & 0.05 & 0.1 & 0.2 & 0.01 & 0.02 & 0.05 & 0.1 & 0.2 & 0.01 & 0.02 & 0.05 & 0.1 & 0.2 \\
\midrule
& \multicolumn{15}{c}{$\omega = 0.1$} \\
\cmidrule(lr){2-16}
SIP & 12 & 12 & 12 & 12 & 12 & 11 & 12 & 12 & 12 & 12 & 7 & 6 & 7 & 7 & 9 \\
AVE & 2 & 1 & 2 & 11 & 12 & 2 & 1 & 5 & 7 & 12 & 1 & 1 & 1 & 2 & 5 \\
AR & 4 & 3 & 6 & 12 & 12 & 6 & 4 & 5 & 12 & 12 & 1 & 2 & 3 & 3 & 6 \\
SARIMA & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 1 & 2 \\
HAR-D & 0 & 0 & 0 & 4 & 12 & 0 & 0 & 0 & 3 & 12 & 1 & 1 & 1 & 2 & 4 \\
XGBoost & 7 & 5 & 6 & 9 & 12 & 6 & 3 & 4 & 8 & 12 & 2 & 2 & 2 & 4 & 5 \\
PC & 7 & 7 & 12 & 12 & 12 & 9 & 11 & 10 & 12 & 12 & 4 & 3 & 3 & 4 & 7 \\
TIP-PCA & 12 & 10 & 12 & 12 & 12 & 9 & 10 & 11 & 12 & 12 & 4 & 3 & 4 & 5 & 7 \\
\midrule
& \multicolumn{15}{c}{$\omega = 0.5$} \\
\cmidrule(lr){2-16}
SIP & 12 & 12 & 12 & 12 & 12 & 11 & 12 & 12 & 12 & 12 & 6 & 7 & 7 & 9 & 10 \\
AVE & 3 & 3 & 5 & 12 & 12 & 5 & 3 & 8 & 11 & 12 & 2 & 2 & 3 & 4 & 6 \\
AR & 7 & 6 & 11 & 12 & 12 & 7 & 8 & 11 & 12 & 12 & 3 & 3 & 4 & 5 & 8 \\
SARIMA & 5 & 6 & 11 & 12 & 12 & 5 & 7 & 12 & 12 & 12 & 2 & 3 & 4 & 7 & 10 \\
HAR-D & 0 & 0 & 1 & 4 & 12 & 0 & 1 & 1 & 3 & 12 & 1 & 1 & 2 & 3 & 5 \\
XGBoost & 5 & 5 & 4 & 8 & 12 & 6 & 5 & 5 & 11 & 12 & 2 & 2 & 3 & 5 & 7 \\
PC & 10 & 12 & 12 & 12 & 12 & 10 & 12 & 12 & 12 & 12 & 5 & 5 & 5 & 6 & 8 \\
TIP-PCA & 10 & 12 & 12 & 12 & 12 & 10 & 10 & 11 & 12 & 12 & 5 & 4 & 5 & 6 & 8 \\
\midrule
& \multicolumn{15}{c}{$\omega = 0.9$} \\
\cmidrule(lr){2-16}
SIP & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 8 & 9 & 10 & 11 & 12 \\
AVE & 12 & 12 & 12 & 12 & 12 & 11 & 11 & 12 & 12 & 12 & 6 & 6 & 7 & 9 & 10 \\
AR & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 7 & 8 & 8 & 10 & 11 \\
SARIMA & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 7 & 7 & 9 & 11 & 11 \\
HAR-D & 6 & 8 & 12 & 12 & 12 & 6 & 8 & 12 & 12 & 12 & 4 & 4 & 6 & 9 & 11 \\
XGBoost & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 7 & 8 & 8 & 10 & 11 \\
PC & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 12 & 6 & 7 & 8 & 10 & 11 \\
TIP-PCA & 12 & 12 & 12 & 12 & 12 & 11 & 12 & 12 & 12 & 12 & 7 & 7 & 9 & 10 & 11 \\
\bottomrule
\end{tabular}}
\end{table}
To further evaluate the performance of the proposed method, we conducted a 5-minute frequency Value at Risk (VaR) estimation for the remainder of the trading day, given each $\omega$.
Specifically, we first predicted the remaining conditional expected instantaneous volatilities on the $D$th day using the SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA, based on the same in-sample period data as in the previous analysis.
We then computed quantiles using historical standardized 5-minute returns.
Specifically, we first standardized in-sample 5-minute returns with the estimated conditional instantaneous volatilities; then, we derived sample quantiles at levels 0.01, 0.02, 0.05, 0.1, and 0.2.
Using these sample quantiles and predicted instantaneous volatilities, we obtained 5-minute frequency VaR values for each prediction method.
We used a fixed in-sample period as one quarter and employed a rolling window scheme with the out-of-sample period spanning 189 days, which is consistent with the previous analysis.
Based on the estimated VaR values, we conducted the likelihood ratio unconditional coverage (LRuc) test \citep{kupiec1995techniques}, the likelihood ratio conditional coverage (LRcc) test \citep{christoffersen1998evaluating}, and the dynamic quantile (DQ) test with lag 4 \citep{engle2004caviar}.
Table \ref{VaR} reports the number of cases where the adjusted $p$-value using the BH procedure exceeds 0.05 for the 12 ETFs at each $q_0 \in \{0.01, 0.02, 0.05, 0.1,0.2\}$, based on the LRuc, LRcc, and DQ tests under varying values of $\omega \in \{0.1,0.5,0.9\}$.
From Table \ref{VaR}, we find that the SIP method consistently outperforms other methods across all hypothesis tests.
We note that when $\omega$ is small, the TIP-PCA method also performs well relative to other methods.
This may be because when $\omega$ is small, the amount of current trading day's information is small.
Thus, TIP-PCA, which uses only the information up to the previous day's information, can account for the interday and intraday dynamics by incorporating the interday HAR structure and the U-shaped intraday volatility feature.
Additionally, Table \ref{VaR} shows that the performance of SARIMA gets better as $\omega$ grows, aligning with the results in Table \ref{MSPE}.
However, SARIMA exhibits poor performance when $\omega$ is small in terms of VaR estimation.
Table \ref{VaR} indicates that SIP outperforms TIP-PCA and other methods, particularly for lower quantiles, which are generally more difficult to predict.
This result confirms that the proposed nonparametric prediction method, SIP, by leveraging current intraday volatility information, significantly enhances both the prediction accuracy of future instantaneous volatilities and the effectiveness of risk management.
\section{Conclusion} \label{conclusion}
This paper proposes a novel nonparametric prediction procedure for forecasting the remaining future instantaneous volatility process during the current intraday trading period.
The proposed method, Structural Intraday-volatility Prediction (SIP), leverages both previous days' data and partially observed current-day data by considering the missing component structure of the approximately low-rank matrix representation of the instantaneous volatility process.
That is, this paper extends the low-rank matrix completion to predict time series patterns, which makes it possible to forecast future intraday volatility without relying on parametric modeling assumptions.
As a result, SIP is robust to model misspecification.
We establish the asymptotic properties of the SIP estimator and prove its consistency.
In the empirical analysis, the SIP method outperforms existing alternatives in terms of out-of-sample forecasting accuracy and Value at Risk (VaR) estimation for the remaining intraday volatility.
Notably, SIP demonstrates strong predictive performance even when only a small fraction of intraday observations is available.
This finding highlights that early intraday information plays a crucial role in forecasting the remaining volatility, and that the low-rank structure effectively captures intraday volatility dynamics without requiring a specific model form such as an autoregressive process.
\bibliography{myReferences}
\vfill\eject