Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,920 characters · 8 sections · 39 citation commands
Low-Rank Structured Nonparametric Prediction of Instantaneous Volatility
\pagenumbering{arabic}
\eject
\doublespacing
The analysis of volatility is essential in financial econometrics and statistics, and it has broad applications in hedging, option pricing, risk management, and portfolio allocation. In recent years, the widespread availability of high-frequency financial data has led to the development of several nonparametric integrated volatility estimation methods ait2010high, barndorff2008designing, barndorff2011multivariate, bibinger2014estimating, christensen2010pre, fan2018robust, fan2007multi, jacod2009microstructure, xiu2010quasi, zhang2005tale, zhang2006efficient, zhang2011estimating. These high-frequency-based estimators have significantly improved our understanding of low-frequency (interday) market dynamics and have motivated the development of various conditional volatility models built on realized measures andersen2003modeling, corsi2009simple, hansen2012realized, kim2016unified, shephard2010realising. Parallel to these developments, nonparametric procedures for estimating instantaneous (or spot) volatility have been proposed to capture intraday volatility dynamics fan2008spot,figueroa2022kernel, foster1996continuous, kristensen2010nonparametric, mancini2015spot, todorov2019nonparametric, todorov2023bias, zu2014estimating. Leveraging these well-performing estimators, several studies have illustrated U-shaped or reverse J-shaped intraday volatility patterns across various markets admati1988theory, andersen1997intraday, andersen2019time, hong2000trading, li2023robust.
In contrast to the extensive literature on predicting daily volatility processes, relatively less attention has been given to forecasting intraday volatility. For example, engle2012forecasting proposed a modified GARCH model in which the conditional variance is expressed as the product of daily, diurnal, and stochastic intraday components to capture intraday volatility dynamics, while zhang2024volatility explored several nonparametric machine learning methods for forecasting intraday realized volatility, leveraging the commonality observed in intraday volatility patterns. Recently, TIP-PCA considered the interday-by-intraday instantaneous volatility matrix process and proposed a semiparametric prediction procedure for the one-day-ahead instantaneous volatility process by leveraging both interday and intraday volatility dynamics. However, existing methods for forecasting intraday volatility suffer from notable limitations. For example, many methods rely on strong distributional assumptions inherent in parametric models and often fail to incorporate interday dynamics effectively. In practice, practitioners are often interested in forecasting the remaining future intraday volatility process conditional on the observed intraday data up to the present time. Therefore, it is both practically relevant and methodologically important to develop a prediction procedure that adapts to current intraday information and is robust to potential model misspecification.
This paper introduces a novel nonparametric prediction approach based on a low-rank structure for forecasting the remaining future instantaneous volatility process during intraday trading. Specifically, we represent the entire instantaneous volatility process as a matrix, where each row corresponds to a day and each column represents an intraday time point. We impose a low-rank plus noise structure on this matrix. The parameter of interest is the instantaneous volatility process after the current intraday time. That is, on the current day (denoted as the $D$th day), only the initial portion of the row is observed, while the remaining part is the target of prediction. To forecast this remaining portion, we exploit the low-rank matrix structure. For example, the future instantaneous volatility vector can be represented by the singular components of the observed part of the low-rank matrix. We use this structure to construct a nonparametric prediction of the future volatility vector. This method is called the Structural Intraday-volatility Prediction (SIP) procedure. It is worth noting that as long as the interday-by-intraday instantaneous volatility matrix exhibits a low-rank structure, the SIP method can effectively predict the remaining future volatility vector without relying on any specific parametric dynamics such as autoregressive models. Therefore, the SIP procedure is robust to model misspecification arising from incorrect parametric assumptions. Furthermore, we derive the convergence rate of the predicted instantaneous volatility under the SIP framework and assess its empirical performance via an out-of-sample prediction study using high-frequency financial data. The SIP method consistently outperforms alternative approaches in terms of both out-of-sample prediction accuracy and Value-at-Risk (VaR) forecasting performance.
The remainder of the paper is organized as follows. In Section (ref), we introduce the model and propose the SIP prediction procedure. Section (ref) presents the asymptotic properties of the SIP estimator. In Section (ref), we evaluate the finite sample performance of the proposed method through simulation. Section (ref) applies the proposed method to real high-frequency financial data to forecast the remaining future instantaneous volatility process during trading hours. Finally, the conclusion is provided in Section (ref). All technical proofs are provided in the Appendix (ref).
Throughout this paper, we denote by $\|\bfm A\|_{F}$, $\|\bfm A\|_{2}$ (or $\|\bfm A\|$ for short), $\|\bfm A\|_{1}$, $\|\bfm A\|_{\infty}$, and $\|\bfm A\|_{\max}$ the Frobenius norm, operator norm, $l_{1}$-norm, $l_{\infty}$-norm, and element-wise norm, which are defined, respectively, as $\|\bfm A\|_{F} = \mathrm{tr}^{1/2}(\bfm A^{\top}\bfm A)$, $\|\bfm A\|_{2} = \lambda_{\max}^{1/2}(\bfm A^{\top}\bfm A)$, $\|\bfm A\|_{1} = \max_{j}\sum_{i}|a_{ij}|$, $\|\bfm A\|_{\infty} = \max_{i}\sum_{j}|a_{ij}|$, and $\|\bfm A\|_{\max} = \max_{i,j}|a_{ij}|$. When $\bfm A$ is a vector, the maximum norm is denoted as $\|\bfm A\|_{\infty}=\max_{i}|a_{i}|$, and both $\|\bfm A\|$ and $\|\bfm A\|_{F}$ are equal to the Euclidean norm.
We consider the following jump diffusion process: for the $i$-th day and intraday time $t \in [0,1],$
where $X_{i,t}$ is the log price of an asset, $\mu_{i,t}$ is a drift process, $B_{i,t}$ is a one-dimensional standard Brownian motion, $J_{i,t}$ is the jump size, and $P_{i,t}$ is the Poisson process with the intensity $\mu_{J}$. For a given intraday time sequence, for each $i =1,\dots, D$ and $j=1,\dots, n$, we denote the instantaneous volatility process as $c_{i,j} := \sigma_{i,t_{j}}^2$, where $0< t_{1} < \cdots < t_{n} =1$. Then, we assume that the discrete-time instantaneous volatility process satisfies the following low-rank structure:
where $\bfm U= (u_{i,k} ) _{i=1,\ldots, D, k=1,\ldots,r}$ is the left singular vector matrix, $\bfm V = (v_{j,k}) _{j=1, \ldots, n, k=1,\ldots, r}$ is the right singular vector matrix, $\ensuremath{\boldsymbol{\Lambda}}= \mathrm{Diag} (\lambda_1, \ldots, \lambda_r)$ is the singular value matrix, and $\bfm E = (e_{i,j})_{D\times n}$ is the random noise matrix. We only impose the above low-rank structure in (ref) to predict future remaining intraday volatility. It allows for considerable modeling flexibility and can implicitly incorporate various time-series features. For example, the left singular vector can represent interday time-series dynamics, while the right singular vector can account for intraday patterns, such as U-shaped or reverse J-shaped TIP-PCA. It is worth noting that, unlike TIP-PCA, we do not assume any explicit dynamic or parametric form for the volatility evolution. We note that $\bfsym \Sigma_{D, n}$ is a $D \times n$ matrix, which is distinct from a covariance matrix and contains only positive elements.
In this paper, our target is to predict the remaining future instantaneous volatility vector for the $D$th day. Specifically, we assume that trading data is available up to a fraction $\omega = \frac{n_1}{n}$ of the $D$th day. The objective is to predict the remaining $n_2 \times 1$ instantaneous vector, $(c_{D,n_1+1},\dots, c_{D,n})$, where $n = n_1 + n_2$. Given the current information $\mathcal{F}_{D,n_1}$, we can predict the instantaneous volatility as follows: for $j = n_1+1, \ldots, n$,
We can write the model (ref) in a partitioned matrix form as follows: $$ \bfsym \Sigma_{D,n} =
, \bfm A =
, \bfm E =
$$ We note that $\Sigma_{11}, \Sigma_{12}$, and $\Sigma_{21}$ are estimable using high-frequency log-price observations. Since $\bfm A$ is the low rank matrix, $A_{22}$ is given by
where $U_{11}$ is the $(D-1)\times r$ left singular vector matrix, $V_{11}$ is the $n_1 \times r$ right singular vector matrix, and $\Lambda_{11}$ is the $r\times r$ singular value matrix of $A_{11}$. Since $A_{11}$, $A_{21}$, and $A_{12}$ are quantities related to the observed period, using this relationship (ref), we can predict $A_{22}$ without a specific parametric modeling assumption. Specifically, as long as we can estimate $A_{11}$, $A_{21}$, and $A_{12}$ well using the available data, $A_{22}$ can be estimated using the plug-in scheme. We note that if we treat the remaining intraday volatility vector $A_{22}$ as missing components, this problem becomes similar to the matrix completion cai2010singular, candes2012exact, candes2010matrix. Especially in this problem, the missing has a determined structure as in bai2021matrix, cai2016structured,fan2019structured, yan2024entrywise. From this point of view, this paper extends the matrix completion concept to predict time series patterns.
Due to the imperfections in the trading process, practitioners cannot directly observe the true underlying log-stock price $X_{i,t}$ in (ref) ait2009high. To account for this, we assume that the high-frequency intraday observations $X_{i,t_s}, s=1, \dots, m,$ are contaminated by microstructure noise as follows:
where $e_{i,t_s}$ is the microstructure noise process with a mean of zero and a variance of $\eta_{ii}$.
To predict the remaining future instantaneous volatility vector $A_{22}$ nonparametrically, we use the low-rank structural equation (ref). Since $A_{11}$, $A_{12}$, and $A_{21}$ are latent low-rank matrices, we need to estimate them using a well-performing instantaneous volatility estimator based on the observed high-frequency data. For example, we need a large dimension to estimate the latent low-rank matrix and use the blessing of dimensionality li2018embracing. To do this, we choose the large block matrices such as $(\Sigma_{11}, \Sigma_{12})$ and $(\Sigma_{11} ^{\top}, \Sigma_{21} ^{\top})$. On the other hand, we employ a well-performing nonparametric instantaneous volatility estimator to estimate the instantaneous volatility. Details can be found in Remark (ref) below. The specific estimation procedure is as follows:
In summary, given the estimated instantaneous volatility process, we construct a structured matrix that includes the remaining future intraday volatility vector. To forecast the future intraday volatility vector, we utilize the low-rank structure described above. We refer to this procedure as the Structural Intraday-volatility Prediction (SIP) method. The SIP method nonparametrically predicts the future instantaneous volatility vector using historical trading data along with current intraday market observations. Crucially, by incorporating current intraday information, the prediction is formulated under a low-rank matrix assumption without relying on any specific parametric model, such as an autoregressive structure. This feature makes the SIP procedure robust to model misspecification arising from incorrect parametric assumptions. The numerical results in Sections (ref) and (ref) demonstrate that the SIP method effectively predicts the remaining future instantaneous volatility vector.
In this section, we develop the asymptotic properties of the SIP estimator. To achieve this, the following technical conditions are needed.
We obtain the following elementwise convergence rate of the predicted instantaneous volatility using the SIP method.
Theorem (ref) shows that the proposed SIP estimator consistently predicts the future instantaneous volatility process. The convergence rate of the SIP estimator is given by $m^{-1/4} + \frac{1}{\sqrt{D}} + \frac{1}{\sqrt{n_1}}$ up to the log order and the sparsity level, when $\rho_m < m^{-1/4}$. Each term in this rate corresponds to a distinct component of the estimation challenge: $D^{-1/2}$ reflects the statistical cost of learning the interday (daily) time-series dynamics; $n_1^{-1/2} $ captures the cost of estimating intraday volatility patterns; $ m^{-1/4}$ arises from estimating the unobserved instantaneous volatility using high-frequency returns. Notably, while the optimal nonparametric convergence rate for estimating instantaneous volatility is $m^{-1/8}$, the SIP method achieves a faster rate of $m^{-1/4}$. This improvement stems from the low-rank structure, which allows the SIP estimator to benefit from a law-of-large-numbers-type averaging across the volatility surface. In addition, the term $ \rho_m$ arises from the bias component in the spot volatility estimator. Specifically, the bias originates from the drift term, which typically decays at the rate $m^{-1}$; however, due to the subsampling approach used to mitigate microstructure noise, the effective convergence rate $\rho_m$ is generally of order $m^{-1/2}$.
In this section, we examined the finite sample performance of the proposed SIP method through simulations. We generated high-frequency observations that incorporate both interday HAR model feature and the intraday periodic pattern TIP-PCA as follows: for $i= 1,\dots, D$, $s=0,\dots, m$, and $t_{s} = s/m$,
where we set microstructure noise as $e_{i,t_{s}} \sim \mathcal{N}(0,0.0005^2)$ and the initial value as $X_{1,0} = 1$; $B_{i,t}$ is a standard Brownian motion; for the jump part, we set $J_{i,t} \sim \mathcal{N}(-0.01, 0.02^2)$ and $P_{i,t+\Delta} - P_{i,t} \sim \text{Poisson}(36\Delta/252)$; $\widetilde{\sigma}_{i} = b_{0} + b_{1}\widetilde{\sigma}_{i-1} + b_{2}\frac{1}{5}\sum_{s=1}^{5}\widetilde{\sigma}_{i-s} + b_{3}\frac{1}{22}\sum_{s=1}^{22}\widetilde{\sigma}_{i-s} + \zeta_{i}$, $\zeta_{i} \sim \mathcal{N}(0,1)$, $h(t_{s}) = \gamma_{0} + \gamma_{1}(t_{s}-0.6)^2$, and $\varepsilon_{i,t_{s}} = q(t_{s})\xi_{i,t_{s}}$, where $q(t_{s})^{2} = 0.1+0.5(2t_{s}-1)^{2}$ and $\xi_{i,t_{s}} \sim \mathcal{N}(0,0.01^2)$. The model parameters were set to be
The normalized parameter values above imply the daily time unit, and we adapted the estimated coefficients studied in corsi2009simple to generate $\widetilde{\sigma}_{i}$. We note that in each simulation, the instantaneous volatility process was generated until all instantaneous volatility values were positive, based on the data-generating process described above. We set $m =$ 23,400, which indicates that the data are observed every second over a period of 6.5 trading hours per day.
For each simulation, we used the jump robust pre-averaging method figueroa2022kernel to estimate the instantaneous volatility, $c_{i,\tau}=\sigma_{i,\frac{\tau}{n}}^2$, at a frequency of every 5 minutes (i.e., $n=78$) for each $i$-th day as follows: for $\tau = 1, \dots, n$,
where $K_{b}(x) = K(x/b)/b$, the bandwidth size $b_{m} = 1/n$, the weight function $g(x) = 2x \wedge (1-x)$,
$\mathbf{1}_{\{\cdot\}}$ is an indicator function, and $\nu_{m} = 1.8\sqrt{\text{BPV}}(k_{m}/m)^{0.47}$, where the bipower variation $\text{BPV} = \frac{\pi}{2}\sum_{s=2}^{m}|Y_{i,t_{s-1}} - Y_{i,t_{s-2}}| |Y_{i,t_{s}} - Y_{i,t_{s-1}}|$. We used the uniform kernel function and the data-driven approach to obtain the preaveraging window size, $k_m$, as suggested in Section 3.1 of figueroa2022kernel.
With the instantaneous volatility estimates available up to the $n_1$-th intraday data point on the $D$th day, we examined the out-of-sample performance of estimating the remaining $(1\times n_2)$ instantaneous volatility vector, where $n_1 = \omega n$ and $n_2 = n- n_1$. For comparison, we considered the SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA methods in predicting $c_{D,\tau}$ for $\tau = n_1+1, \ldots, n$. AVE represents estimates obtained from the column mean of $\widehat{\bfsym \Sigma}_{D-1,n}$, while AR generates predictions using an autoregressive model of order 1, applied to each column of $\widehat{\bfsym \Sigma}_{D-1,n}$. SARIMA provides $n_2$-step-ahead forecasts using a seasonal ARIMA(1,1,1) model sheppard2010financial, where the seasonal cycle length is set to $n$. HAR-D extends the HAR model by incorporating diurnal effects and prior intraday components in addition to daily, weekly, and monthly realized volatilities zhang2024volatility. XGBoost chen2016xgboost, a decision-tree-based ensemble learning method, accounts for nonlinear dependencies, using the same hyperparameters as in zhang2024volatility. Since SARIMA, HAR-D, and XGBoost require a time series format, we first vectorized the estimates as $(\widehat{c}_{1,1},\dots, \widehat{c}_{D,n_1})$ before applying these methods. PC corresponds to the last row of the estimated low-rank matrix obtained from the best rank-$r$ approximation of $\widehat{\bfsym \Sigma}_{D-1,n}$. TIP-PCA uses ex-post daily, weekly, and monthly realized volatilities and the intraday time sequence as observable covariates that explain the left and right singular vectors, as proposed by TIP-PCA, given $\widehat{\bfsym \Sigma}_{D-1,n}$. For SIP, TIP-PCA, and PC, we set rank 1 as determined by the eigenvalue ratio method ahn2013eigenvalue. We generated high-frequency data with $m = 23,400$ over 200 consecutive days, using subsampled log prices from the last $D =$ 50, 100, 150, and 200 days. To assess prediction accuracy, we computed the mean squared prediction error (MSPE) as $$ \frac{1}{n}\sum_{\tau=n_1+1}^{n}(\widetilde{c}_{D,\tau}-c_{D,\tau})^{2}, $$ where $\widetilde{c}_{D,\tau}$ denotes the instantaneous volatility estimator of each method above. Finally, we calculated the sample averages of MSPEs across 500 simulations.
Figure (ref) presents the average MSPEs of the estimators for the remaining future intraday instantaneous volatility process across different values of dimensionality, $D = {50, 100, 150, 200}$, with fixed $n$ and $\omega$. Figure (ref), on the other hand, plots the average MSPEs against varying values of $\omega = {0.05, 0.1, 0.2, \dots, 0.9}$, while keeping $D$ and $n$ fixed. Note that the target future volatility remains consistent across different $D$ values, as the same subsampled data are used in each simulation setting. In contrast, in Figure (ref), the MSPEs exhibit a U-shaped pattern as $\omega$ varies. This behavior arises because the target future volatility differs with each value of $\omega$, and the underlying data-generating process follows a U-shaped intraday volatility pattern. As a result, direct comparisons of MSPE values across different $\omega$ values are not appropriate. Instead, MSPEs should be interpreted relatively across competing methods for each fixed value of $\omega$.
Figures (ref) and (ref) demonstrate that the SIP method consistently achieves the best performance among the competing approaches. This may be because by incorporating information from previous intraday volatility patterns on the $D$th day, SIP effectively predicts the remaining future instantaneous volatility process through a nonparametric framework. Moreover, the MSPEs of SIP tend to decrease as the number of daily observations $D$ increases, which supports the theoretical results established in Section (ref). In contrast, the TIP-PCA, PC, AVE, and AR methods perform poorly, as they are unable to leverage current intraday observations on the $D$th day within their estimation procedures. In Figure (ref), we observe that the performance of SARIMA improves as $\omega$ increases. This suggests that, given sufficient real-time intraday information, SARIMA can better incorporate both periodic intraday patterns and daily autoregressive dynamics. Nevertheless, SIP outperforms SARIMA and all other competing methods across the entire range of $\omega$, particularly when $\omega$ is small.
In this section, we applied the proposed SIP method to predict intraday instantaneous volatility vector using real high-frequency trading data. Specifically, we obtained intraday data of the S&P 500 index ETF (SPY) and the 11 Global Industry Classification Standard (GICS) sector ETFs (XLC, XLY, XLP, XLE, XLF, XLV, XLI, XLB, XLRE, XLK, and XLU) from July 2021 to June 2022, sourced from the Trade and Quote (TAQ) database in the Wharton Research Data Services (WRDS) system. We used high-frequency data subsampled at 1-second intervals, excluding trading days with early market closures. This subsampling mitigates the effects of irregular observation times li2023robust. Using the log prices, we implemented the jump robust pre-averaging estimation procedure, as defined in Section (ref), to estimate the instantaneous variance at 5-minute intervals. Using the in-sample period data, we applied SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA, as described in Section (ref), to predict the remaining instantaneous volatilities given each $\omega = \{0.1, 0.5, 0.9\}$. We employed the rolling window scheme with a 63-day (i.e., one quarter) in-sample period, where the last $n_2 = (1-\omega)n$ instantaneous volatilities on the final in-sample day are the target volatility. The out-of-sample period was from October 2021 to June 2022 (i.e., $q=189$ days).
To evaluate the performance of the predicted instantaneous volatility, we employed the mean squared prediction errors (MSPE) and QLIKE loss function patton2011volatility, defined as follows: for each $\omega$ and $n_2 = (1-\omega)n$,
where $\widetilde{c}_{i,j}$ represents the predicted instantaneous volatility on the $j$th intraday time point of the $i$th out-of-sample day, obtained from one of the SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA methods. We predicted the remaining future conditional expected instantaneous volatilities using in-sample period data. Since the true conditional expected instantaneous volatility is unknown, we assessed the significance of differences in prediction performances using the Diebold and Mariano (DM) test diebold2002comparing based on both MSPE and QLIKE. We compared the proposed SIP method against alternative approaches. Given that multiple hypothesis tests were conducted repeatedly throughout this section, it is crucial to control the False Discovery Rate (FDR). To account for this, we applied the Benjamini–Hochberg (BH) procedure benjamini1995controlling at a significance level of $\alpha = 0.05$ to adjust the $p$-values and mitigate the risk of false positives across all hypothesis tests performed in this section.
Tables (ref) and (ref) report the MSPE and QLIKE results, respectively, under varying values of $\omega \in \{0.1, 0.5, 0.9\}$. As discussed in Section (ref), the target future volatility changes with $\omega$. Consequently, direct comparisons of MSPE and QLIKE values across different $\omega$ values are inappropriate. Instead, performance should be evaluated relative to other methods for each given $\omega$. Tables (ref)--(ref) show that the SIP method consistently outperforms other methods across different values of $\omega$. This implies that incorporating the early intraday trading data enhances the accuracy of instantaneous volatility prediction. We note that the performance of SARIMA and XGBoost improves when $\omega = 0.9$. This may be attributed to the availability of a larger amount of information on the last trading day, which improves their prediction performance for a smaller set of target volatilities.
To further evaluate the performance of the proposed method, we conducted a 5-minute frequency Value at Risk (VaR) estimation for the remainder of the trading day, given each $\omega$. Specifically, we first predicted the remaining conditional expected instantaneous volatilities on the $D$th day using the SIP, AVE, AR, SARIMA, HAR-D, XGBoost, PC, and TIP-PCA, based on the same in-sample period data as in the previous analysis. We then computed quantiles using historical standardized 5-minute returns. Specifically, we first standardized in-sample 5-minute returns with the estimated conditional instantaneous volatilities; then, we derived sample quantiles at levels 0.01, 0.02, 0.05, 0.1, and 0.2. Using these sample quantiles and predicted instantaneous volatilities, we obtained 5-minute frequency VaR values for each prediction method. We used a fixed in-sample period as one quarter and employed a rolling window scheme with the out-of-sample period spanning 189 days, which is consistent with the previous analysis.
Based on the estimated VaR values, we conducted the likelihood ratio unconditional coverage (LRuc) test kupiec1995techniques, the likelihood ratio conditional coverage (LRcc) test christoffersen1998evaluating, and the dynamic quantile (DQ) test with lag 4 engle2004caviar. Table (ref) reports the number of cases where the adjusted $p$-value using the BH procedure exceeds 0.05 for the 12 ETFs at each $q_0 \in \{0.01, 0.02, 0.05, 0.1,0.2\}$, based on the LRuc, LRcc, and DQ tests under varying values of $\omega \in \{0.1,0.5,0.9\}$. From Table (ref), we find that the SIP method consistently outperforms other methods across all hypothesis tests. We note that when $\omega$ is small, the TIP-PCA method also performs well relative to other methods. This may be because when $\omega$ is small, the amount of current trading day's information is small. Thus, TIP-PCA, which uses only the information up to the previous day's information, can account for the interday and intraday dynamics by incorporating the interday HAR structure and the U-shaped intraday volatility feature. Additionally, Table (ref) shows that the performance of SARIMA gets better as $\omega$ grows, aligning with the results in Table (ref). However, SARIMA exhibits poor performance when $\omega$ is small in terms of VaR estimation. Table (ref) indicates that SIP outperforms TIP-PCA and other methods, particularly for lower quantiles, which are generally more difficult to predict. This result confirms that the proposed nonparametric prediction method, SIP, by leveraging current intraday volatility information, significantly enhances both the prediction accuracy of future instantaneous volatilities and the effectiveness of risk management.
This paper proposes a novel nonparametric prediction procedure for forecasting the remaining future instantaneous volatility process during the current intraday trading period. The proposed method, Structural Intraday-volatility Prediction (SIP), leverages both previous days' data and partially observed current-day data by considering the missing component structure of the approximately low-rank matrix representation of the instantaneous volatility process. That is, this paper extends the low-rank matrix completion to predict time series patterns, which makes it possible to forecast future intraday volatility without relying on parametric modeling assumptions. As a result, SIP is robust to model misspecification. We establish the asymptotic properties of the SIP estimator and prove its consistency.
In the empirical analysis, the SIP method outperforms existing alternatives in terms of out-of-sample forecasting accuracy and Value at Risk (VaR) estimation for the remaining intraday volatility. Notably, SIP demonstrates strong predictive performance even when only a small fraction of intraday observations is available. This finding highlights that early intraday information plays a crucial role in forecasting the remaining volatility, and that the low-rank structure effectively captures intraday volatility dynamics without requiring a specific model form such as an autoregressive process.
\eject