EconBase
← Back to paper

Heteroscedasticity test of high-frequency data with jumps and microstructure noise

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

57,949 characters · 11 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Heteroscedasticity test of high-frequency data with jumps and microstructure noise

frontmatter\address[NUS]{Department of Mathematics, National University of Singapore, Singapore} \address[UM]{Department of Mathematics, University of Macau, Macau SAR, China} \address[UMZRI]{UMacau Zhuhai Research Institute, Zhuhai, China} \address[ZN]{School of Finance, Zhongnan University of Economics and Law, Wuhan 430073, China} \begin{abstract} In this paper, we are interested in testing if the volatility process is constant or not during a given time span by using high-frequency data with the presence of jumps and microstructure noise. Based on estimators of integrated volatility and spot volatility, we propose a nonparametric way to depict the discrepancy between local variation and global variation. We show that our proposed test estimator converges to a standard normal distribution if the volatility is constant, otherwise it diverges to infinity. Simulation studies verify the theoretical results and show a good finite sample performance of the test procedure. We also apply our test procedure to do the heteroscedasticity test for some real high-frequency financial data. We observe that in almost half of the days tested, the assumption of constant volatility within a day is violated. And this is due to that the stock prices during opening and closing periods are highly volatile and account for a relative large proportion of intraday variation. \\ \\ JEL Classification: C12, C14, G10 \end{abstract} \begin{keyword} High-frequency data \sep Jumps \sep Market microstructure noise \sep Heteroscedasticity \sep Nonparametric test \end{keyword}

Introduction

It is well known that the logarithmic price of an asset is necessarily to be modeled as a semi-martingale process under the assumption of arbitrage-free and frictionless market. The coefficient process driving the standard Brownian motion part, which is called the volatility process, serves as a measurement of risk in finance. Due to the wide applications of volatility in pricing of asset and derivative, portfolio selection, hedging and risk management, there are tons of research works on estimating the volatility. Many quantities targeting to measure the magnitude of the volatility, such as integrated volatility, spot volatility, realized Laplace transform of volatility, are proposed (see AJ2014 for their concrete definitions and a comprehensive introduction). Besides, it is also important to investigate the dynamic structure of the volatility process. Up until now, numerous models which have been proposed and widely applied are the ones in but not limited to BS1973, V1977, CIR1985, C1992, DH1993. Or in another way, specific functional forms of the volatility process may be postulated before one can do goodness-of-fit tests to verify the correctness, related references are AS1996, CW1999, DLW2003, DPV2006, DP2008, VD2012, christensen2018diurnal, and references therein. Among them, one of the most basic questions have been tried to be answered is that whether the volatility process is constant or not over a period of time, say a day\footnote{Regarding the daily pattern of the volatility process, it reaches an agreement in TT1997,christensen2018diurnal,andersen2019time that there are two distinct sources of variation for many financial asset return series. One of them is a deterministic diurnal component representing the fixed daily pattern. The other one is a stochastic part fluctuating around the fixed one, which brings in randomness and captures volatility clustering. Recently, christensen2018diurnal concluded that the re-scaled log-returns are often close to homoscedastic within a trading day and the fixed diurnal pattern accounts for a rather significant fraction of intraday variation in the volatility. But they also found that important sources of heteroscedasticity remain present in the data after annihilating the diurnal effect. Thus, the diurnal pattern is not sufficient to explain daily variation of the volatility.}. Putting forward a procedure to answer such a question is also the purpose of this paper. In the most of the previous literatures, the test procedures are constructed based on a continuous diffusion assumption, while we consider the underlying data generating process of the return as a general It$\hat{\text{o}}$ semi-martingale where the jumps are involved. Besides, the presence of market microstructure noise is also taken into account in our paper. From a theoretical perspective, we contribute to propose a new nonparametric heteroscedasticity test procedure by using high-frequency data and further extend it to different settings incorporating the jumps and the market microstructure noise.

Our goodness-of-fit test procedure is based on the estimation of integrated volatility and spot volatility, which are well documented in existing literatures and many methods are valid under different settings. The integrated volatility quantifies the fluctuation of the asset price over a fixed time period, while the spot volatility measures the variation instantaneously. We take a special case of diffusion process for an example to explain the mechanism implicated in our test. The stochastic process is discretely observed at evenly distributed points on the fixed time interval $[0,1]$. The asymptotic setup is of infilled type, namely, the mesh between the observation grids shrinks to zero. Under such a setting, we know that the estimators of the integrated volatility, for example, realized volatility and realized power variations (see e.g. ABDL2003; BN2004; BGJPS2006; J2008) are constructed based on all observations, while corresponding bounded kernel versions of the estimation of the spot volatility, as used in FW2008 and K2010, only use the local data near a fixed time point. When the volatility remains constant over $[0,1]$, then the integrated volatility and the spot volatility at any given time second coincide, and the estimators of the former one have a faster convergence rate than the ones for the latter quantity. When the volatility varies over the time interval, the estimation of the spot volatility enables us to recover the time-varying volatility process, while the estimators of the integrated volatility give us a random variable. We construct a test statistic by integrating the squared differences between a sequence of spot volatility estimators over blocks with shrinking time length and an integrated volatility estimator. If the volatility is constant, then the scaled differences asymptotically distribute as a standard normal distribution, and the partial sum of the centered squared differences behaves asymptotically like a discrete martingale. Our statistic is shown to be asymptotic normal under the null hypothesis of constant volatility, while it diverges to infinity at an appropriate rate if the volatility process is time-varying. The test statistic is easy to compute and our test procedure can be naturally extended to other scenarios after taking the jumps and the market microstructure noise into consideration. Similar idea is also adopted in T2017 to test time-varying jump activity index for a pure jump semi-martingale defined on a fixed time interval.

We start our discussion as described above with continuous diffusion model, which is the most commonly-used one for the return process. But it has been shown, in BNS2006, AJ2009, AJ2010, AJ2009deg, JKLM2012 and references therein, that it is not adequate to describe the various fluctuation patterns of financial asset price because of the presence of jumps, which may be due to the news shocks from the markets. The mixing jumps bring in extra bias compared with the estimation of volatility under the continuous framework. If the number of jumps is finite, two well-behaved estimators are realised multi-power variation estimator and realised threshold quadratic variation estimator, respectively. The former one was given in BN2004, BSW2006 and J2008, while the latter one was proposed by M2009 and MR2011. The cases regarding more active jump intensity, like infinite activity or even infinite variation, are considered in JLK2014, JT2014, LLL2018 and among many others.

Apart from jumps, the existence of market microstructure noise in the observation procedure, which may caused by the presence of a bid-ask spread and the corresponding bounces, the differences in trading sizes and in representation of the prices, the different informational content of price changes, the discreteness of price changes, and data errors, also brings in bias. To eliminate the bias, one of the most effective and easy to implement way is by averaging the raw data before we apply the aforementioned estimation procedure, which is called pre-averaging approach proposed in PV2009b and further extended in JLMPV2009. The other approaches are two time-scaled and multi time-scaled estimators proposed in ZMA2005 and Zhang2006; the realised kernel method proposed in BNHLS2008a; the quasi-maximum likelihood method in Xiu2010, the local moment method proposed by bibinger2014estimating, and etc. Based on existing methods of volatility estimation in the presence of jumps and market microstructure noise, we extend our heteroscedasticity test procedure and verify our theoretical results.

The rest of this paper is organized as follows. In Section (ref), we give out our model setup and the asymptotic theoretical properties. We firstly illustrate our test under continuous semi-martingale assumption, and then extend our theories to the framework with jumps and market microstructure noise by using thresholding and pre-averaging techniques separately. In Section (ref), we verify our theoretical results and test the finite sample performance of the proposed tests via Monte Carlo studies. Our tests are applied to some real high-frequency financial data sets for empirical analysis in Section (ref). In subsequent Section (ref), we conclude our paper. Technical proofs are postponed to Appendix.

Theoretical results

In this section, we firstly construct our heteroscedasticity test procedure by modeling the logarithmic price process as a continuous It$\hat{\text{o}}$ semi-martingale. If a jump part of finite activity is further involved in the underlying data-generating process, the test procedure can be naturally extended after thresholding the raw observed data. Finally, we incorporate the presence of market microstructure noise into the observation procedure, namely the observed data at a given time equals to the value of the underlying process at that time plus another stochastic error term. We apply the pre-averaging technique before implementing the heteroscedasticity test to eliminate the bias due to the noise. Detailed descriptions and assumptions regarding the jumps and the market microstructure noise will be given later. By combining the techniques used for eliminating the effects of the jumps and the market microstructure noise, we can also extend the test procedure to the situation with simultaneous presence of the jumps and the market microstructure noise. Since the extension can be evidently seen from our previous results, we omit its detailed proof and discussion in this paper.

Throughout the paper, all the processes are defined on the time interval $[0,1]$. We denote $V_{i/n}$ to be the value of the process $V$ at the time point $i/n$ and define $ \Delta_i^n V = V_{i/n} - V_{(i-1)/n} $ for $i=1,\dots, n$. The whole test procedure is based on an infill asymptotic setting, namely $n\rightarrow \infty$, which gives us the high-frequency data. We use the notations $\rightarrow^{p}, \rightarrow^{d}, \rightarrow^{ds}$ to denote convergence in probability, convergence in distribution and stable convergence, respectively. In general, we say $\mathcal{F}$-stable convergence of a sequence $X_n$ to $X$ defined on an extension of $(\Omega, \mathcal{F}, \mathcal{F}_t, P)$, if for any bounded Lipschitz function $g$ and any bounded $\mathcal{F}$-measurable $\mathcal{Q}$, as $n\rightarrow \infty $, it holds that

align*[align* omitted — 85 chars of source]

where $\mathbf{E}'$ stands for the expectation on an extension space. The detailed definition and more properties of stable convergence can be found in JS2003.

Continuous semi-martingale

At first, we present our methodology in a benchmark setup, which excludes jumps and market microstructure noise when modeling the high-frequency data. We denote $X$ to be the logarithmic price process of an asset, and $X$ is set to be an one-dimensional continuous It$\hat{\text{o}}$ semi-martingale defined on the filtered probability space $(\Omega,\mathcal{F},(\mathcal{F}_{t})_{t\geq 0}, \mathcal{P})$ with the following form:

equation[equation omitted — 79 chars of source]

where $b$ and $\sigma$ are progressively measurable processes, and $B$ is a standard Wiener process. We also assume that the volatility process $\sigma$ to be a continuous It$\hat{\text{o}}$ semi-martingale on the same filtered probability space $(\Omega,\mathcal{F},(\mathcal{F}_{t})_{t\geq 0}, \mathcal{P})$, and it can be represented as

equation[equation omitted — 154 chars of source]

where $b^{\sigma}$, $D^{\sigma}$ and $D^{'\sigma}$ are adapted, c$\grave{\text{a}}$dl$\grave{\text{a}}$g stochastic processes, $b^{\sigma}$ is further predictable and locally bounded, and $B^{'}$ is another standard Wiener process independent of $B$. It is required that $\sigma$ is bounded away from 0, that is, $\sigma_t > 0$ for $0 \leq t \leq 1$ almost surely. We note that the common driving standard Wiener process $B$ in $X$ and $\sigma$ accommodates the leverage effect in finance, which depicts the dependence structure between these two stochastic processes. Such continuous semi-martingale models for the log-price processes and the volatility process are widely used in vast existing high-frequency literature for volatility estimation, e.g., BNHLS2008a, MZ2009, JLK2014, and etc.

In this paper, we are interested in investigating the pattern of the volatility process. Specifically, we want to test if the volatility process is constant or not during a given time period. To this end, we partition the sample space $\Omega$ into two complementary subsets

align*[align* omitted — 138 chars of source]

The null hypothesis can then be written as $\mathcal{H}_0: \omega \in \Omega^{c}$, while the alternative $\mathcal{H}_a: \omega \in \Omega^{v}$. Our target then turns to proposing a test with a pre-set asymptotic significance level and with power going to one to test the null hypothesis, as $n\rightarrow \infty$.

We start demonstrating our theories with the estimation of integrated volatility $IV := \int_0^1 \sigma_s^2 ds$. It is well known that the most frequently used estimator of the integrated volatility is the so-called realized volatility, which is defined as

equation[equation omitted — 67 chars of source]

It is shown in BNS2007 that, under our setting,

equation[equation omitted — 138 chars of source]

where $N(0,1)$ denotes standard normal distribution with mean 0 and variance 1. And the central limit theorem result can be turned feasible when we replace the integrated quarticity $\int_0^t\sigma_s^4ds$ by its consistent estimators. Based on the estimation of the integrated volatility, the estimation of spot volatility $\sigma_{\tau}^2$ at any given time $\tau$ can be correspondingly proposed by applying the kernel method described in FW2008, K2010, YFLZZ2014 and LLL2018. For example, using the specific one-side uniform kernel function $K(u) = 1_{ \{ 0 \leq u\leq 1 \} }$, we can obtain an estimator of $\sigma_{\tau}^2$ as

equation[equation omitted — 183 chars of source]

where $k_n$ is the number of intervals after the time point $\tau$ and lie closest to $\tau$. Following the theoretical results in aforementioned references, we conclude that under our setting, and if further $k_n\rightarrow \infty$ and $k_n^2/n \rightarrow 0$ hold, as $n \rightarrow \infty$, we have

equation[equation omitted — 194 chars of source]

The consistency result $(\widehat{\sigma^2}_\tau^{n}(k_n))^2 \rightarrow^{p} \sigma_{\tau}^4$ further implies that a feasible central limit theorem can be obtained if we replace $\sigma_{\tau}^4$ by $(\widehat{\sigma^2}_\tau^{n}(k_n))^2$.

Now, we state the first test procedure, which is based on the estimators of the integrated volatility and the spot volatility discussed above.

thm$X$ follows the process in (ref). \begin{enumerate} • For $ \omega \in \Omega$, if as $n \rightarrow \infty$, $k_n \rightarrow \infty$ and $k_n/n \rightarrow 0$, then it holds that, as $n \rightarrow \infty$, \begin{equation} \frac{k_n}{n} \sum_{j=0}^{\lfloor n/k_n \rfloor-1} (\widehat{\sigma^2}_{jk_n/n}^{n}(k_n)-\widehat{IV}^{n} )^2 \rightarrow^{p} \int_{0}^{1} (\sigma_s^2 - IV)^2ds. \end{equation} • For $\omega \in \Omega^{c}$, if as $n \rightarrow \infty$, $k_n \rightarrow \infty$ and $k_n^2/n \rightarrow 0$, then it holds that, as $n \rightarrow \infty$, \begin{equation} \mathcal{T}^{n}(k_n):= \sqrt{\frac{k_n}{2n}} \sum_{j=0}^{\lfloor n/k_n \rfloor-1 } \Big\{ \big(\frac{\sqrt{k_n}(\widehat{\sigma^2}_{jk_n/n}^{n}(k_n)-\widehat{IV}^{n})}{\sqrt{2}\widehat{IV}^{n}}\big)^2 - 1\Big\} \rightarrow^{d} N(0,1). \end{equation} • Denote $ z_{\alpha} $ as the $\alpha$-quantile of standard normal distribution, if as $n \rightarrow \infty$, $k_n \rightarrow \infty$ and $k_n^2/n \rightarrow 0$, then it holds that, as $n \rightarrow \infty$, \begin{align} \begin{cases} &\mathcal{P}( \mathcal{T}^{n}(k_n) > z_{1-\alpha}|\Omega^{c}) \rightarrow \alpha, \ if \ \mathcal{P}(\Omega^{c}) > 0, \\ &\mathcal{P}( \mathcal{T}^{n}(k_n) > z_{1-\alpha}|\Omega^{v}) \rightarrow 1. \end{cases} \end{align} \end{enumerate}

The intuition is as follows. If the volatility process is constant over $[0,1]$, then the magnitudes of the integrated volatility and the spot volatility at any given time are equal. So that from the conclusion $(\ref{cltspo})$, it holds that $ \sqrt{k_n}(\widehat{\sigma^2}_{jk_n/n}^{n}(k_n) - IV) $ with $j=0,...,\lfloor n/k_n \rfloor-1$ are asymptotically uncorrelated and normally distributed with mean zero and variance $2\sigma^4_{0}$. Then, by martingale central limit theorem in HH1980, we can obtain

equation*[equation* omitted — 271 chars of source]

After replacing $IV$ by its estimator $\widehat{IV}^{n}$ above, we note that the independence structure between the terms in the summation is broken up. But it does not affect the asymptotic conclusion, because $\widehat{IV}^{n}$ converges to $IV$ at a faster rate, compared with the convergence rate of $(\widehat{\sigma^2}_{jk_n/n}^{n}(k_n) - \sigma^2_{jk_n/n})$ to zero. After substituting $\sigma^2_{0}$ with its consistent estimator $\widehat{IV}^{n}$, we obtain the result ((ref)). Whether the volatility process is constant or not, we always have ((ref)), where the left hand side term is approximately the Riemann sums of the term on the right hand side. Plugging ((ref)) into ((ref)), we see that if the constant volatility assumption is violated, then the quantity $\mathcal{T}^{n}(k_n)$ in ((ref)) will tend to infinity at a fast rate of $\sqrt{nk_n}$. The different asymptotic properties of $\mathcal{T}^{n}(k_n)$ for constant volatility and time-varying volatility lead to conclusion ((ref)) and enable us to do the hypothesis testing in our way.

Finite activity jump

Now, we consider the setting where the underlying logarithmic price process is modeled as the combination of the continuous process $X$ and another pure jump process $J$, which is restricted to be of finite activity. That is, we have

equation[equation omitted — 46 chars of source]

and the process $Y$, instead of $X$ in the last part, is observed at the time points $\frac{i}{n}$, for $i=0,1,...,n$. We write $J_t = \sum_{j=1}^{N_t} \gamma_{\tau_j}$, where $N_t$ is a non-explosive counting process with possibly time varying intensity, $\gamma_{\tau_j}$ is the size of the jump at time $\tau_j$. These jump sizes are not necessarily i.i.d random variables, nor independent of $N$.

To eliminate the influence of jumps on estimating the integrated volatility, M2009 proposed a thresholding technique to discriminate the time intervals with jumps from those without jumps. Such a filtering procedure can be done by using a deterministic threshold function $r(x)$ satisfying the following conditions:

asuThe function $r(x): \mathbf{R} \rightarrow \mathbf{R}$ satisfies, $\lim_{x\rightarrow 0}\displaystyle\frac{x\log(\displaystyle\frac{1}{x})}{r(x)} = 0$, $\lim_{x\rightarrow 0} r(x) = 0$.

It is shown that for the sample paths, with probability one, there are jumps between $[(i-1)/n, i/n]$ if $(\Delta_i^nY)^2 > r(1/n)$. Since these intervals with jumps are finite, excluding observed data in these intervals has no influence on the asymptotic properties of the estimator of the integrated volatility. Consequently, the thresholding versions of the estimators of the integrated volatility (called truncated realised volatility) and the spot volatility are formalized as

equation[equation omitted — 367 chars of source]

The same conclusions in ((ref)) and ((ref)) also hold if we replace $\widehat{IV}^{n}$ and $\widehat{\sigma^2_\tau}^{n}(k_n)$ with $\widehat{IV}^{n,Thr}$ and $\widehat{\sigma^2_\tau}^{n,Thr}(k_n)$, respectively. As a by-product, their detailed proofs are also given as we prove the following main theorem in Appendix.

thm$X$ follows the process in (ref), and Assumption (ref) hold. \begin{enumerate} • For $ \omega \in \Omega$, if $k_n \rightarrow \infty$ and $k_n/n \rightarrow 0$ hold, as $n \rightarrow \infty$, then we have, as $n \rightarrow \infty$, \begin{equation} \frac{k_n}{n} \sum_{j=0}^{\lfloor n/k_n \rfloor-1} (\widehat{\sigma^2}_{jk_n/n}^{n,Thr}(k_n)-\widehat{IV}^{n,Thr} )^2 \rightarrow^{p} \int_{0}^{1} (\sigma_s^2 - IV)^2ds. \end{equation} • For $\omega \in \Omega^{c}$, if $k_n \rightarrow \infty$, $k_n^2/n \rightarrow 0$ and $\sqrt{k_n}\log{n}/\sqrt{n} \rightarrow 0$ hold, as $n \rightarrow \infty$, then we have, as $n \rightarrow \infty$, \begin{equation} \mathcal{T}^{n,Thr}(k_n):= \sqrt{\frac{k_n}{2n}} \sum_{j=0}^{\lfloor n/k_n \rfloor -1 } \Big( \big(\frac{\sqrt{k_n}(\widehat{\sigma^2}_{jk_n/n}^{n,Thr}(k_n)-\widehat{IV}^{n,Thr})}{\sqrt{2}\widehat{IV}^{n,Thr}}\big)^2 - 1\Big) \rightarrow^{d} N(0,1). \end{equation} • Denote $ z_{\alpha} $ as the $\alpha$-quantile of standard normal distribution, if $k_n \rightarrow \infty$, $k_n^2/n \rightarrow 0$ and $\sqrt{k_n}\log{n}/\sqrt{n} \rightarrow 0$ hold, as $n \rightarrow \infty$, then we have, as $n \rightarrow \infty$, \begin{align} \begin{cases} &\mathcal{P}( \mathcal{T}^{n,Thr}(k_n) > z_{1-\alpha}|\Omega^{c}) \rightarrow \alpha, \ if \ \mathcal{P}(\Omega^{c}) > 0, \\ &\mathcal{P}( \mathcal{T}^{n,Thr}(k_n) > z_{1-\alpha}|\Omega^{v}) \rightarrow 1. \end{cases} \end{align} \end{enumerate}

To reduce the effect of jumps (finite activity or infinite activity) in estimating the integrated volatility, another alternative method is the so-called realized multi-power variation estimator (see BN2004, BSW2006 and J2008), which diminishes the effect of jumps by using the products of the consecutive absolute increments $|\Delta_{i}^n Y|$. Theoretically, both of these two estimators are rate-efficient, but the truncated realised volatility is more efficient than the realised multi-power variation estimator in the sense of having a smaller variance. Indeed, the realised multi-power variation estimator is mainly biased by large jumps but is less affected by small jumps, while on the contrary, the truncated realised volatility is problematic in removing small jumps but eliminates large jumps effectively. In Veraart2011, the properties of these two estimators are analyzed and compared comprehensively, their finite sample performances are verified by numerous Monte Carlo studies under different models. Furthermore, a combination of these two estimators breeds a new estimator called truncated realized multi-power variation estimator therein, which achieves the best effect of finite sample performance, since such a combination compensates the weaknesses of these two estimators. We note that our test procedure can be constructed accordingly by using these estimators mentioned, but we only consider the truncated realised volatility version here from the perspective of both simplicity and efficiency.

rmkThe restriction of finite activity on the jump process $J$ can be relaxed to some extent, for example, the case of L\'{e}vy jumps of infinite activity with finite variation. It can be shown that the same conclusions in above theorem also hold for this relatively relax condition, but we only consider finite jumps for simplicity of the proof procedure. More on related properties and analyses can be found in MR2011 and JLK2014.
rmkOne possible choice for $r(x)$ is the power function $cx^{\omega}$, with $c$ being a constant and $\omega \in (0,1)$. A time varying version of $r(x)$ (may be stochastic) is considered in MR2011. Furthermore, AJ2009 point out that the value of $c$ should be proportional to the “average" value of $\sigma_t$, which could be consistently estimated by the multi-power variation estimator mentioned above. The specific setting of the parameters $c$ and $\omega$ are also discussed in Veraart2011, supported by a great deal of simulation studies.

Market microstructure noise

In this part, the data generating process of log-price is still modeled as the continuous semi-martingale $X$, but the observation procedure is conducted with disturbance. Mathematically, the observed data $Z_{i/n}$ at $\frac{i}{n}$ for $i=0,1,...,n$ are the underlying process $X_{i/n}$ contaminated by another market microstructure noise term $\epsilon_{i/n}$, that is

equation[equation omitted — 65 chars of source]

For the convenience of description, we define $\epsilon_t$ over the whole time span for $t \in [0,1]$. About the process $\epsilon$, we assume that there exists a transition probability $Q_t(\omega,dx) $ from $(\Omega,\mathcal{F}_t) $ into $R$. We endow the space $\Omega' = R^{[0,\infty)}$ with the product Borel $\sigma$-field $\mathcal{F}'$ and with the probability $\mathcal{Q}(\omega,d\omega')$ which is the product $\otimes_{t \geq 0} Q_t(\omega,\cdot)$. The process $Z$ is called the “canonical process" on $(\Omega', \mathcal{F}')$, with the filtration $\mathcal{F}' = \sigma(Z_s: s\leq t)$. We then work in the filtered probability space $(\Omega'', \mathcal{F}^{''}, \mathcal{F}_{t \geq 0}^{''}, \mathcal{P})$ with

equation*[equation* omitted — 261 chars of source]

And the following assumption is satisfied:

asuWe have \begin{equation*} \int x Q_t(\omega,dx) = X_t(\omega), \end{equation*} and the process \begin{equation*} \alpha_t(\omega) = \int x^2 Q_t(\omega,dx) - X_t(\omega)^2 = \mathbf{E}[(Z_t)^2|\mathcal{F}](\omega) - X_t(\omega)^2 \end{equation*} is c$\grave{a}$dl$\grave{a}$g(necessarily ($\mathcal{F}_t$)- adapted), and the process \begin{equation*} \beta_t(\omega) = \int x^8Q_t(\omega,dx) \end{equation*} is locally bounded.

Before giving our estimators of the integrated volatility and the spot volatility, we firstly need to pre-average the raw increments with a function $g$ supported on the interval $[0,1]$ satisfying

asuThe function $g$ is continuous and piecewise differentiable with a piecewise Lipschitz derivative $g'$, \begin{equation*} g(0) = g(1) = 0, \qquad 0 < \int_{0}^{1} g(s)^2ds < \infty. \end{equation*}

Denote the shorthand $g_i^n = g(i/p_n)$, then the pre-averaged increments for any process $V$ is defined as

equation*[equation* omitted — 148 chars of source]

Now, these treated increments are used to construct our estimator of the integrated volatility, which is given by

equation[equation omitted — 140 chars of source]

with $\varphi_n = \displaystyle\frac{1}{p_n}\sum_{i=1}^{p_n}(g_i^n)^2$. Similarly in an aforementioned way of kernel smoothing, an estimator of the spot volatility can be obtained as

equation[equation omitted — 224 chars of source]

where $l_n$ is the widow width of the kernel estimation.

In view of ((ref)), we have $\overline{Z}_{jp_n}^n = \overline{X}_{jp_n}^n + \overline{\epsilon}_{jp_n}^n$. Some simple variance calculations show that $\overline{X}_{jp_n}^n = O_p(\sqrt{\displaystyle\frac{p_n}{n}})$ and $\overline{\epsilon}_{jp_n}^n = O_p(\sqrt{\displaystyle\frac{1}{p_n}})$. If $p_n\rightarrow \infty$ and $p_n^2/n \rightarrow \infty$ hold, as $n\rightarrow \infty$, then $\overline{Z}_{jp_n}^n$ is dominated by $\overline{X}_{jp_n}^n$, and the effect of the market microstructure noise can then be neglected. As a consequence, we have $\widehat{IV}^{n,Pre}(p_n) \rightarrow^{p} \int_{0}^{1}\sigma_s^2ds$ and $\widehat{\sigma^2}_{kp_nl_n/n}^{n,Pre}(p_n,l_n) \rightarrow^{p} \sigma_{kp_nl_n/n}^2$. If further that the conditions $p_n^5/n^3$ and $\sqrt{nl_n}/p_n \rightarrow 0$ are satisfied, we have the following central limit theorems:

align[align omitted — 311 chars of source]

We also give a sketch of their proofs in Appendix (Lemma (ref)) as a by-product of this paper. Based on these results, we can establish our test procedure as

thm$X$ follows the process in (ref), and Assumptions (ref)-(ref) hold. \begin{enumerate} • For $ \omega \in \Omega$, if as $n \rightarrow \infty$, $p_n \rightarrow \infty$, $l_n \rightarrow \infty$, and $p_n^{2}/n \rightarrow \infty$ hold, then we have, as $n \rightarrow \infty$, \begin{equation} \frac{p_nl_n}{n} \sum_{k=0}^{\lfloor n/(p_nl_n) \rfloor-1 } (\widehat{\sigma^2}_{kp_nl_n/n}^{n,Pre}(p_n,l_n)-\widehat{IV}^{n,Pre}(p_n) )^2 \rightarrow^{p} \int_{0}^{1} (\sigma_s^2 - IV)^2ds. \end{equation} • For $\omega \in \Omega^{c}$, if as $n \rightarrow \infty$, $p_n \rightarrow \infty$, $l_n \rightarrow \infty$, $\sqrt{nl_n}/p_n\rightarrow 0$ and $p_n^{5}/n^3 \rightarrow 0$ hold, then we have, as $n \rightarrow \infty$, \begin{equation} \mathcal{T}^{n,Pre}(p_n,l_n):= \sqrt{\frac{p_nl_n}{2n}} \sum_{k=0}^{\lfloor n/(p_nl_n) \rfloor-1 } \Big( \big(\frac{\sqrt{l_n}(\widehat{\sigma^2}_{kp_nl_n/n}^{n,Pre}(p_n,l_n)-\widehat{IV}^{n,Pre}(p_n) )}{\sqrt{2}\widehat{IV}^{n,Pre}(p_n)}\big)^2 - 1\Big) \rightarrow^{d} N(0,1). \end{equation} • Denote $ z_{\alpha} $ as the $\alpha$-quantile of standard normal distribution, if as $n \rightarrow \infty$, $p_n \rightarrow \infty$, $l_n \rightarrow \infty$, $\sqrt{nl_n}/p_n\rightarrow 0$ and $p_n^{3/2}/n \rightarrow 0$ hold, then we have, as $n \rightarrow \infty$, \begin{align} \begin{cases} &\mathcal{P}( \mathcal{T}^{n,Pre}(p_n,l_n) > z_{1-\alpha}|\Omega^{c}) \rightarrow \alpha, \ if \ \mathcal{P}(\Omega^{c}) > 0, \\ &\mathcal{P}( \mathcal{T}^{n,Pre}(p_n,l_n) > z_{1-\alpha}|\Omega^{v}) \rightarrow 1. \end{cases} \end{align} \end{enumerate}

On one hand, we consider constructing our estimators of the integrated volatility and the spot volatility by using non-overlapping pre-averaged data for simplicity, instead of the overlapping case considered in JLMPV2009. It has no harm to our theoretical results, but at a cost of reducing the number of pre-averaged data. On the other hand, as mentioned above, we diminish the effect of the noise by choosing $p_n^2/n \rightarrow \infty$. Alternatively, we can also take $p_n = O_p(n^{1/2})$, then $\overline{X}_{jp_n}^n$ and $\overline{\epsilon}_{jp_n}^n $ are of the same order. In this case, the effect of the noise should be removed by subtracting an estimator of the variance of the noise. Furthermore, the presence of the noise also deforms the variances of the asymptotic distributions in ((ref)) and ((ref)), thus new estimators of these variances are necessarily to be reconstructed. It is viable to extend our test procedure to the setting with $p_n = O_p(n^{1/2})$ and overlapping pre-averaged data, but such a consideration can complicate our test procedure to an undesirable degree. We mention that such a setting may be considered as a sole work for our future research. Readers who are interested in this setting can refer to JLMPV2009 for the detailed discussion when it comes to the estimation of the integrated volatility.

rmkIf the presence of jump process and market microstructure noise are both considered simultaneously, we can obtain similar results by combining the thresholding technique and the pre-averaging method. The extension can be obviously seen from our previous derivation, thus we omit the detailed discussion here. Related papers can be referred to are JLK2014 and references therein.

Monte Carlo study

We now conduct some Monte Carlo simulation studies to examine our test procedure and investigate the finite sample performance of our test estimator in the cases of constant volatility and stochastic volatility. As discussed in the last theoretical section, we consider three different scenarios where continuous semi-martingale, involvement of finite jumps and contamination from market microstructure noise are used for modeling the log-price process. For the notations, we follow their definitions given in previous sections for the old ones, and shall specify later where new ones are used.

Simulation design

The latent log-price process $X=(X_{t})_{0 \leq t\leq 1}$ is generated from the following two stochastic differential equations, one of them considers constant volatility while the other one considers stochastic volatility.

$ \bullet $ Model 1--The constant volatility model

gather[gather omitted — 33 chars of source]

with $X_0=1$ and $\sigma=1$.

$ \bullet $ Model 2--The Heston model with stochastic volatility

align[align omitted — 161 chars of source]

with the parameters $\kappa=5, \alpha=0.04,\gamma=5$, $\rho=-\sqrt{0.5}$, $X_0=1$ and $\sigma_0=1$. We follow the parameter setting in WM2014 to calibrate the model to real financial data.

Regarding the jump component $J_t = \sum_{j=1}^{N_t} \gamma_{\tau_j}$ in Section (ref), we consider the jump size $\gamma_{\tau_j} \sim N(0,\sigma_{\kappa}^{2})$, the number of jumps up to time point $t$, $N_{t}\sim Poisson(\lambda t)$, which is a Poisson distribution with parameter $\lambda$. We firstly generate the process $N_{t}$ within $t\in[0,1]$, and in subsequence generate $\gamma_{\tau_j}$ independently. We fix $\sigma_{\kappa}=0.5$ and choose different jump intensity by setting the parameter $\lambda=10, 20, 50, 100$. For estimating the spot volatility, the window-width is set as $k_n=\lfloor\frac{\theta}{\sqrt{\Delta_n}}\rfloor$ with $\theta=1.2$ for satisfying the theoretical conditions in Theorem (ref). We apply the thresholding technique to filter the jumps by setting the truncation level $\nu_n=4\sqrt{BV_n} \cdot \Delta_n^{\varpi}$ with $\varpi=0.499$ and

align*[align* omitted — 85 chars of source]

The quantity $BV_n$ is the realized bipower variation estimator introduced in BN2004 and serves as a consistent estimator of the integrated volatility which is robust to jumps.

For the market microstructure noise term $\epsilon_t$ in Section (ref), it is mixed in the observed prices at $t= 0, 1/n, ..., n/n$, with $\epsilon_{t} \sim N(0,\eta^2)$. The noise terms are independently and identically distributed with different strengths $\eta=0.001, 0.01, 0.05$. Recall that $p_n$ is the number of increments based on the raw data used for pre-averaging, $l_n$ is the number of non-overlapping pre-averaged blocks used for the kernel estimation of the spot volatility. For Theorem (ref), we take $p_n = \lfloor c \Delta_n^{-(1/2+\chi)}\rfloor$ with $c=1/3$ and $\chi=0.05$, and $l_n = \lfloor a \Delta_n^{-b} \rfloor$ with $a = 2$ and $b = 0.17$, which satisfy our theoretical requirement.

For each experiment, we simulate 5000 runs of daily sample paths by using the Euler discretization method. We consider different sampling frequencies with $ n= 23400, 11700, 7800$, $4680, 2340, 1170, 780$, corresponding to sampling at every 1, 2, 3, 5, 10, 20, 30 seconds respectively, over a 6.5-hour trading day in the U.S. stock market.

Simulation results

Table (ref) records the empirical size (based on constant volatility) and power (based on time-varying volatility) of the heteroscedasticity test when finite activity jumps are present in the logarithmic price process. The setting $\lambda = 0$, which corresponds to the continuous semi-martingale model without jumps, is also documented for comparison. We observe desirable size performances, meaning that the probability of type I error is acceptable, for $\lambda=0, 10, 20$ and all $n$ considered. For fixed $\lambda$, the magnitude of size approaches to corresponding nominal confidence level as the sampling frequency increases, this is even more evident for relatively larger $\lambda$. As for the influence of jumps, we see that more intensive jumps always worsen the performance of size, and the extent is more obvious when the sample size $n$ is relatively smaller. We find that almost all the values of power are 1 for all $n$, $\lambda$ and the three nominal levels, which shows our test is quite powerful in detecting the time variation in volatility process. This is inline with our theoretical analysis in Section (ref) that our test estimator diverges at a fast rate if the volatility process is not constant.

table[table omitted — 2,093 chars of source]

Table (ref) documents the size and power of the heteroscedasticity test in the presence of market microstructure noise. The finite sample performances of both size and power are satisfying for $\eta = 0.001, 0.01$. For fixed $\eta$, as the sampling frequency increases, the size and power perform better in the sense of getting closer to corresponding nominal confidence levels and 1, respectively. This phenomenon is even more distinct for relatively larger $\eta$. Regarding the effect of the market microstructure noise, a larger $\eta$ deviates the values of Power away from 1 and yields a larger type II error. Moreover, such a deterioration is even worse for relatively smaller sample size $n$. We also find that all the values of size are close to corresponding nominal confidence levels for all the different parameters $n$ and $\eta$ considered. This justifies that the pre-averaging technique works well in demolishing the disturbance from the market microstructure noise for the estimation of volatility (the integrated volatility or/and the spot volatility). The results of power are more sensitive to the presence of market microstructure noise because our statistics $\mathcal{T}^{n,Pre}(p_n,l_n)$ in Theorem (ref) diverges to infinity in a relatively slow rate which depends on the parameters $p_n,l_n$, when the constant volatility assumption is violated.

table[table omitted — 1,559 chars of source]

To verify the accuracy of the normal approximations of our test statistics, namely (ref) in Theorem (ref), (ref) in Theorem (ref) and (ref) in Theorem (ref), we demonstrate Q-Q plots and histograms for the finite estimates of $\mathcal{T}^{n}(k_n)$, $\mathcal{T}^{n,Thr}(k_n)$ and $\mathcal{T}^{n,Pre}(p_n,l_n)$ under the constant volatility model in Figures (ref)--(ref). It is shown that all the histograms approximate standard normal distribution closely and the Q-Q plots are almost linear, which proves the asymptotic normality of these three quantities.

figure[figure omitted — 565 chars of source]
figure[figure omitted — 579 chars of source]
figure[figure omitted — 580 chars of source]

Real data analysis

In this section, we apply our proposed heteroscedasticity test statistics to high-frequency data from the NYSE TAQ database. We use the transaction price data of the International Business Machines (IBM) in the whole year of 2011, with a total of 252 trading days. For various reasons, raw trading data contains numerous errors. Therefore, the data is not immediately suitable for analysis and data-cleaning is an essential step when dealing with tick-by-tick data. Following the pre-filtering routine of BNHLS2009, we collect all transactions from 9:30 to 16:00, delete entries with zero prices, merge multiple transactions with the same time stamp by taking the weighted average of all prices and sample every 5 seconds in calendar time. We consider 5-minute data for $\mathcal{T}^{n}(k_n)$ in Theorem (ref) and $\mathcal{T}^{n,Thr}(k_n)$ in Theorem (ref) to avoid the influence of market microstructure noise, and 5-second data for $\mathcal{T}^{n,Pre}(p_n,l_n)$ in Theorem (ref).

Recall that we set $k_n=\lfloor\frac{\theta}{\sqrt{\Delta_n}}\rfloor$ and $p_n = \lfloor c \Delta_n^{-(1/2+\chi)}\rfloor$ for the number of raw data used for kernel smoothing in the noise-free setting and number of raw data used for pre-averaging in the noisy setting, respectively. Figure (ref) depicts the proportion of the day with time-varying volatility tested at significance levels of 10%, 5% and 1% as a function of $\theta$ in the frictionless cases and a function of $c$ in the noisy setting. Regarding other related parameters not mentioned, they remain the same as the ones in the simulation section. When we implement the test by using 5-minute high-frequency sampling to diminish the influence from the market microstructure noise, the proportions of heteroscedasticity volatility are insensitive to the choice of $\theta$. For the three different significance levels, similar patterns are observed for each scenario with small deviation in the magnitude of heteroscedasticity proportion for all range of $\theta$. If we remove the jumps by the truncation method, the proportions reduce by around 10% for all the three significance levels compared to the case without removing the jumps. This is inline with the intuition that the presence of jumps makes the price process more volatile. For the scenario of considering removing the market microstructure noise by using pre-averaged 5-second data, the heteroscedasticity proportion decreases as $c$ increases. This verifies that the pre-averaging methodology mitigates the impact of the market microstructure noise better for relative larger $c$, which corresponds to the case that more raw data are used for pre-averaging.

figure[figure omitted — 1,023 chars of source]

In Figure (ref), we demonstrate the cross-sectional average of intraday volatility curves estimated by the spot volatility estimators $\widehat{\sigma^2}_{\tau}^{n}(k_n)$ in Theorem (ref) for the continuous setting, $\widehat{\sigma^2_\tau}^{n,Thr}(k_n)$ in Theorem (ref) for the setting with jumps, and $\widehat{\sigma^2}_{kp_nl_n/n}^{n,Pre}(p_n,l_n)$ in Theorem (ref) for the noisy setting respectively. It is shown that the volatility estimates near the opening time or the closing time are relatively larger than other time points in the middle time span. Moreover, the estimated volatility curves are roughly with sharp decreases or increases around pre-scheduled macroeconomic announcements (e.g., at 10:00 or 14:00). This makes the whole volatility curve like a reverted “J"-shape, which is also found in christensen2018diurnal. In fact, this happens for most of the days in a year. There are also empirical literatures explaining the phenomenon. For example, the period covering 9:30 and 10:00 is associated with market-wide news such as FOMC meetings and macroeconomic reports, which make the stock prices to be more volatile, as discussed in LM2008, Lee2012, and etc.

figure[figure omitted — 1,020 chars of source]
table[table omitted — 809 chars of source]

To quantify how does the variation of stock price during the opening and closing time affect our heteroscedasticity test procedure, we record the heteroscedasticity proportion results within three different time spans, namely 09:30-16:00, 10:00-15:30 and 10:30-15:00, in Table (ref). For all three significance levels of 10%, 5% and 1%, we see that the proportions decrease as we gradually discard data obtained in the opening and closing periods for our heteroscedasticity test. The observation implies that the variation during the opening and closing periods leads to a test result of time-varying intraday volatility for most of the days tested.

Conclusion

In this paper, we propose a new nonparametric way to do the heteroscedasticity test for high-frequency data. The test procedure is based on the estimations of integrated volatility and spot volatility, for which a great deal of existing literatures can be found. Our test procedure is easy to conduct and can be naturally extended to different settings, such as the cases in the presence of jumps and market microstructure noise. Our Monte Carlo simulation studies show the good finite sample performance of the asymptotic theory. Finally, we also apply our test procedure to do the heteroscedasticity test for some real high-frequency financial data. The empirical studies indicate that the volatility is not constant in most of days, and the opening and closing periods account for a relatively large proportion of intraday heteroscedasticity. This paper also enlighten us on testing whether the covariance structure between different assets is constant or not during a given time interval, as a future work.

Acknowledgement

Qiang Liu's work is supported by MOE-AcRF Grant of Singapore (No. R-146-000-258-114), Zhi Liu gratefully acknowledges financial support from FDCT of Macau (No. 202/2017/A3) and NSFC (No. 11971507), Chuanhai Zhang's research is supported in part by Humanity and Social Science Youth Foundation of Chinese Ministry of Education (No. 18YJC790210) and in part by the Fundamental Research Funds for the Central Universities, Zhongnan University of Economics and Law (2722019PY038).