Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
94,058 characters · 21 sections · 0 citation commands
Pre-averaging estimators of the ex-post covariance matrix in noisy diffusion models with non-synchronous data
\thispagestyle{empty}
\setcounter{page}{1}
The theory of financial economics is often cast in multivariate settings, where the covariance structure of assets plays a key role to the solution of fundamental economic problems, such as optimal asset allocation and risk management. In recent years, a broader access to financial high-frequency data has improved our ability to accurately estimate and draw inference about financial covariation. The underlying idea is to use quadratic covariation, which we can estimate using realised covariance, as an ex-post measure, whose increments can be studied to learn about the properties of the true asset return covariation \citep*[e.g.][]{andersen-bollerslev-diebold-labys:03a,barndorff-nielsen-shephard:04a}.
In practice, implementing realised covariance is hampered by two empirical phenomena, namely the presence of market microstructure noise (e.g., price discreteness or bid-ask spread bounce) and non-synchronous trading. The impact of microstructure noise has received much attention in the univariate setting, where its effect on the realised variance has been well-documented. This builds on previous work in the noiseless case, including \citet*{andersen-bollerslev-diebold-labys:01a}, \citet*{barndorff-nielsen-shephard:02a}, or \citet*{mykland-zhang:06a,mykland-zhang:09a}. A key to understanding the nature of the noise and a possible tool of how to deal with it is that microstructure noise induces autocorrelation in high-frequency returns and this leads to a bias problem \citep*[see, e.g.,][]{zhou:96a,ait-sahalia-mykland-zhang:05a,hansen-lunde:06a}. Currently, there are three main univariate approaches, where the damage caused by the noise is explicitly fixed: the two-scale subsampler proposed by \citet*{zhang-mykland-ait-sahalia:05a} or its multi-scale version of \citet*{zhang:06a}, the realised kernel introduced in \citet*{barndorff-nielsen-hansen-lunde-shephard:08a}, which relies on autocovariance-based corrections, and finally the pre-averaging estimator of \citet*{podolskij-vetter:09a} and \citet*{jacod-li-mykland-podolskij-vetter:09a}.
The multivariate version of this problem is, however, more complicated in that not only does the estimator need to be robust against various types of noise, it also has to cope with non-synchronous trading \citep*[see, e.g.,][]{fisher:66a}. Asynchronicity causes high-frequency covariance estimates to be biased towards zero as the sampling frequency increases. This feature of the data, the so-called Epps effect, was highlighted by \citet*{epps:79a}. Intuitively, as the sampling frequency is increased, there are more and more zero-returns in the presence of non-synchronous trading, and this will dominate realised covariance and related statistics (e.g. realised correlation). \citet*{hayashi-yoshida:05a} introduced an estimator, which is capable of dealing with non-synchronous data, but not with market microstructure noise. More recently, \citet*{zhang:08a} extended the two-scale RV to integrated covariance estimation in the simultaneous presence of noise and non-synchronicity, while in concurrent and independent work \citet*{barndorff-nielsen-hansen-lunde-shephard:08c} proposed a multivariate realised kernel. Additional work in this growing line of research includes \citet*{malliavin-mancino:02a}, \citet*{martens:03a}, \citet*{reno:03a}, \citet*{bandi-russell:05a}, \citet*{griffin-oomen:06a}, \citet*{large:07a}, \citet*{voev-lunde:07a}, and \citet*{boudt-croux-laurent:08a}, among others.
In this paper, we propose to use a "modulated" realised covariance (MRC) to estimate the ex-post integrated covariance. The econometric technique employed here for dealing with microstructure noise relies on rather simple pre-averaging of the high-frequency data, which makes the estimator both intuitive to understand and trivial to implement. It relates to previous work in the univariate case, where pre-averaging has been suggested in \citet*{podolskij-vetter:09a} and \citet*{jacod-li-mykland-podolskij-vetter:09a}. The current article draws ideas from these papers, but the multivariate extension is challenging, as it faces the additional complexity of non-synchronous trading and requires that the resulting estimator be positive semi-definite.
The pre-averaging approach depends on a bandwidth parameter, or window length, that grows with the sample and dictates the amount of averaging to be carried out. In turn, the choice of this tuning parameter controls the influence of microstructure noise on the MRC and, hence, also its asymptotic properties. In the optimal case, called balanced pre-averaging, this leads to an efficient $n^{- 1 / 4}$ rate of convergence, which is known to be the fastest attainable \citep*[see, e.g.,][]{gloter-jacod:01a,gloter-jacod:01b}. This baseline MRC estimator, however, needs a bias-correction to be consistent for the integrated covariance. As a result, it is not guaranteed to be positive semi-definite in finite samples, though our empirical work indicates this shortcoming is not too much of a concern for more recent data. Nonetheless, as we show in the paper, it is straightforward to design a positive semi-definite estimator by increasing the pre-averaging window length slightly, which can also serve to make the MRC robust against more general noise processes.
The MRC is, in all its essence, a realised covariance computed on the back of pre-averaged high-frequency returns. As such, it depends on receiving synchronous observations as input, which clashes with the irregular spacing of real high-frequency data. We propose two distinct ways in which pre-averaging can be applied in the context of non-synchronous trading. First, we use traditional imputation schemes to map asynchronous data onto a common time grid, for example using previous-tick or refresh time, where the latter approach has been used in \citet*{barndorff-nielsen-hansen-lunde-shephard:08c}. An MRC computed from such returns will be asymptotically robust to non-synchronous trading. Second, we extend the \citet*{hayashi-yoshida:05a} estimator to the case of microstructure noise by using pre-averaging and show that it is consistent. This second estimator has the property that it can be implemented directly on the irregular, non-synchronous and noisy observations without any form of imputation. It therefore omits throwing away information in the sample and further avoids potential biases arising from artificially imputed returns.
An appealing feature of pre-averaging is that it is a general statistical tool that can be applied to many estimation problems. This proves useful in our setting, because as usual the mixed Gaussian central limit theorems feature an unknown conditional covariance matrix. In practice, this must be robustly estimated from sample data in the presence of noise and non-synchronous trading to make the distributional results feasible, such that confidence bands for elements of the integrated covariance matrix can be constructed. We outline how this can be done based on pre-averaged high-frequency data.
The paper progresses as follows. In section 2, we formulate the theoretical setup and define the MRC estimator. In section 3, we first show consistency of the MRC based on balanced pre-averaging and then derive its asymptotic distribution. As discussed above, this estimator needs a bias-correction, so we carry on to study a modified MRC estimator, in which the degree of pre-averaging is increased. We also discuss the application of the MRC to non-synchronous data, show its relation to the multivariate realised kernel, and finally we derive a pre-averaged version of the Hayashi-Yoshida estimator. In section 4, we propose an estimator of the conditional covariance matrix that appears in the central limit theorem of MRC, which can be used to transform infeasible limit results into feasible ones. In section 5, the focus is shifted towards regression and correlation analysis. A simulation study is undertaken in section 6 to uncover the finite sample properties of our estimators, while an empirical illustration is conducted in section 7. Section 8 draws conclusions and presents some ideas for future work. The appendix contains the derivations of all our theoretical results.
We consider a vector of log-prices $X$ defined on a probability space $(\Omega^0, \mathcal F^0, P^0)$ and equipped with an information filtration $(\mathcal F_t^0)_{t \geq 0}$. $X$ has dimension $d$ - the number of assets under consideration.
A standard no-arbitrage condition suggests security prices must be semimartingales \citep*[see, e.g.,][]{back:91a,delbaen-schachermayer:94a}. These processes obey the fundamental theorem of asset pricing and, as a result, are used extensively to model the evolution of asset prices through time. In accordance with this, we model $X$ as a semimartingale that follows the equation
where $a = \left( a_{t} \right)_{t \geq 0}$ is a $d$-dimensional predictable locally bounded drift vector, $\sigma = \left( \sigma_{t} \right)_{t \geq 0}$ an adapted c\`{a}dl\`{a}g $d \times d$ covolatility matrix and $W = \left( W_{t} \right)_{t \geq 0}$ is $d$-dimensional Brownian motion.
This model is a Brownian semimartingale, or stochastic volatility model with drift, which permeates financial economics \citep*[cf.,][for a review]{ghysels-harvey-renault:96a}. We think of this construct as governing an underlying efficient price process - the price that would prevail in the absence of market frictions, which we then subject to microstructure noise.
Of importance to our analysis is the quadratic covariation process of $X$, which is defined as
for any sequence of deterministic partitions $0 = t_0 < t_1 < ... < t_n = t$ with $\sup_{i} \left\{ t_{i} - t_{i - 1} \right\} \to 0$ for $n \to \infty$. In our setting, the quadratic covariation of $X$ is given by
where $\Sigma = \sigma \sigma'$. The quadratic covariation is pivotal in financial economics \citep*[see, e.g., the reviews by][]{barndorff-nielsen-shephard:07a,andersen-bollerslev-diebold:09a}, and we thus take Eq. (ref) as defining the target that we are interested in estimating. We note that for obvious reasons the matrix in Eq. (ref) is also called the integrated covariance and both terms are used interchangeably.
Throughout the remainder of the paper, and without loss of generality, we restrict the clock $t$ to evolve in the unit interval $[0,1]$, which we think of as representing the passing of an economic event, for example a trading day.
In practice, market microstructure noise leads to a departure from the pure semimartingale model. Microstructure noise has many forms, including price discreteness and bid-ask spread bounce, which creates spurious variation in asset prices. As a result, we do not observe $X$ from Eq. (ref) in the market but a process $Y$, which is the efficient price distorted by noise. More precisely, we consider the process $Y$, observed at time points $i / n$, $i = 0, 1, \ldots, n$, which is given as
where $(\epsilon_{t})$ is an i.i.d. process with $X \perp\!\!\!\perp \epsilon$ (the symbol $\perp\!\!\!\perp$ is used to denote stochastic independence).
The noise process can be constructed as follows. We define a second probability space $(\Omega^1, \mathcal F^{1}, ( \mathcal{F}_{t}^{1})_{t \geq 0}, P^1)$, where $\Omega^{1}$ denotes $\mathbb{R}^{[0,1]}$ and $\mathcal{F}^{1}$ the product Borel-$\sigma$-field on $\Omega^{1}$. Next, let $Q$ be a probability measure on $\mathbb{R}$ ($Q$ is the marginal distribution of $\epsilon$). For any $t\in [0,1]$, $P_t^{1}=Q$ and $P^1$ denotes the product $\otimes_{t \in [0,1]} P_t^{1}$. The filtered probability space $(\Omega, \mathcal F, (\mathcal{F}_{t})_{t \geq 0}, P)$, on which we define the process $Y$, is given as
The multivariate noise process $\epsilon$ is assumed to satisfy:
where $\Psi$ is a positive definite $d \times d$-matrix.
Remark \ The empirical results found by \citet*{hansen-lunde:06a} show that both the i.i.d. assumption on $(\epsilon_t)$ and the independence $X \perp\!\!\!\perp \epsilon$ can be called into question when sampling the data at very high frequencies, e.g., below the 1-minute mark \citep*[see also][]{diebold-strasser:08a}. \citet*{jacod-li-mykland-podolskij-vetter:09a} consider more general types of (1-dimensional) noise processes. Roughly speaking, they assume that the errors $\epsilon_t$'s are, conditionally on $X$, centered and independent. The asymptotic theory developed in this paper still holds true for the multivariate version of such noise processes, but we restrict attention to models of the form in Eq. (ref) to ease the exposition.
It is intuitive that under mean zero i.i.d. microstructure noise some form of smoothing of the observed log-price $Y$ should tend to diminish the impact of the noise. Effectively, we are going to approximate $X_{t}$, $X$ being a continuous function of $t$, by an average of observations of $Y$ in a neighborhood of $t$, the noise being averaged away.
Here, we describe in more detail how to conduct the pre-averaging. In particular, we consider a sequence of integers, $k_{n}$, and a number $\theta \in (0, \infty)$ such that
An example of this would be $k_{n} = \left\lfloor \theta \sqrt{n} \right \rfloor$.
We also choose a function $g$ on $[0,1]$, which is continuous, piecewise continuously differentiable with a piecewise Lipschitz derivative $g'$ with $g(0) = g(1) = 0$ and which satisfies $\int_0^1 {g^2 \left( s \right) \text{d}s > 0}$. Furthermore, we introduce the following functions and numbers that are associated with $g$:
The functions $\phi_1$ and $\phi_2$ are assumed to be $0$ outside the interval $[0,1]$.
Next, with any process $V = (V_t)_{t \ge 0}$ we associate the following random variables
Applying this notation to $Y$, it can be seen that $\Delta_{i}^{n} Y$ represents the noisy high-frequency returns, while $\bar{Y}_{i}^{n}$ is the pre-averaged return data, using the weight function $g$. It follows that the stochastic order of $\bar{Y}_{i}^{n} = \bar{X}_{i}^{n} + \bar{ \epsilon}_{i}^{n}$ is controlled by the sequence $k_{n}$, since
Thus, taking $k_{n} = O( \sqrt{n})$ implies that the orders of the two terms in Eq. (ref) are equal, so that $\bar{Y}_{i}^{n} = O_{p} \left( n^{-1/4} \right)$. This is called balanced pre-averaging and delivers the best rate of convergence. As shown below, it is also useful to look at cases in which a higher order of $k_{n}$ is chosen. This results in a suboptimal rate of convergence, but it has some potentially valuable side-effects on the robustness and finite sample properties of our estimator.
The pre-averaging window length, $k_{n}$, depends on the tuning parameter $\theta$, which needs to be chosen by the user. We will later discuss how to sensibly make this choice.
The core statistic of this paper is the multivariate extension of the estimator, which was introduced in \citet*{jacod-li-mykland-podolskij-vetter:09a}. We call it the modulated realised covariance (MRC) and define it as
The factor $n / (n - k_{n} + 2)$ is a finite sample correction for the true number of summands in $MRC \left[ Y \right]_{n}$ relative to the sample size $n$. It is sometimes left out in the presentation below, but it is always included in implementations on data.
Remark \ The sum of outer products in Eq. (ref) is a realised covariance based on pre-averaged data. To build some intuition for our approach, we explain the usage of $\bar Y_i^{n}$ in more detail. Suppose $k_{n}$ is an even number and write
which is a simple average of $Y$ over $k_{n} / 2$ terms. Because of this pre-averaging, $\hat Y_i^{n}$ will be closer to the efficient price $X_{ \frac{i}{n}}$. Next, we compute the realised covariation estimator based on these filtered increments by setting
(However, as we shall see this induces a bias, which is a function of $\Psi$). This method was originally proposed by \citet*{podolskij-vetter:09a} and using the above definition of $\bar Y_i^{n}$ corresponds to choosing the weight function
which is the most intuitive example. Here we explicitly give the numerical values of the asymptotic constants for this choice of function $g$, as it is the one used for all our simulations and empirical work:
Remark \ As noted in \citet*{jacod-li-mykland-podolskij-vetter:09a}, to avoid biases in small samples we should replace the asymptotic constants and functions $\psi_{1}$, $\psi_{2}$, $\phi_{1}$, $\phi_{2}$, $\Phi_{11}$, $\Phi _{12}$, and $\Phi _{22}$ by their Riemann approximations:
These are the actual terms that appear in the computations of the MRC. Note that for all appropriate indices of $i$ and $j$, $\psi_{i}^{k_{n}} \to \psi_{i}$, $\phi_{i}^{k_{n}} \to \phi_{i}$, $\Phi_{ij}^{k_{n}} \to \Phi_{ij}$ as $n \to \infty$ at smaller order than $n^{-1 / 4}$, so the finite sample versions can be replaced with the asymptotic constants in all the limit theorems given below, including the central limit theorems.
Our first result inspects the probability limit of $MRC \left[ Y \right]_{n}$.
A couple of points are worth highlighting. First, as Theorem (ref) shows $MRC \left[ Y \right]_{n}$ is consistent for $\int_{0}^{1} \Sigma_{s} \text{d}s$ up to a bias-correction. The bias term depends on the unknown $\Psi$, which must be estimated from the data.
We set
which is linked to the univariate estimator proposed by \citet*{ait-sahalia-mykland-zhang:05a} and \citet*{bandi-russell:06a,bandi-russell:08a}. Then, we obtain the convergence $\hat{ \Psi}_{n} \overset{p}{ \to} \Psi$, such that
Hence, for the remainder of the paper, we shall incorporate the bias-correction term into the definition of $MRC \left[ Y \right]_{n}$.\footnote{Since $\sum_{i = 1}^{n} \Delta_{i}^{n} Y \left( \Delta_{i}^{n} Y \right)' = 2 n \Psi + \int_{0}^{1} \Sigma_{s} \text{d}s + o_{p}(n^{-1})$, where the error of this approximation has expectation zero, the bias-corrected $MRC \left[ Y \right]_{n}$ actually estimates $\left(1 - \frac{ \psi_{1}^{k_{n}}}{ \theta^{2} \psi_{2}^{k_{n}}} \frac{1}{2n} \right) \int_{0}^{1} \Sigma_{s} \text{d}s$ and thus needs to be rescaled by $1 / \left(1 - \frac{ \psi_{1}^{k_{n}}}{ \theta^{2} \psi_{2}^{k_{n}}} \frac{1}{2n} \right)$. We include this rescaling in our simulations and empirical work but omit it throughout the remainder of the text to simplify notation.} In doing so, we are no longer ensured that $MRC \left[ Y \right]_{n}$ is positive semi-definite in finite samples. To deal with this problem, we propose an alternative formulation of the MRC below, which uses a longer pre-averaging window $k_{n}$ to avoid this step. Second, since $\hat{ \Psi}_{n}$ is a $\sqrt{n}$-estimator of $\Psi$, it will not affect the CLT of $MRC \left[ Y \right]_{n}$, as the latter converges at a slower rate. Third, our initial MRC estimator is based on synchronous data, assuming all components of the log-price vector $Y$ are observed contemporaneously. We will later extend the MRC to the non-synchronous setting.
We proceed with the central limit theorem for $MRC \left[ Y \right]_{n}$. As in \citet*{jacod-li-mykland-podolskij-vetter:09a}, we only require a moment condition on $\epsilon$ to prove this result.
The notion of stable convergence is used, which we describe here. A sequence of random variables $V^{n}$ is said to converge stably in law towards $V$, where $V$ is defined on an appropriate extension $(\Omega', \mathcal{F}', P')$ of the probability space $(\Omega, \mathcal{F},P)$, if and only if for any $\mathcal{F}$-measurable, bounded random variable $W$ and any bounded, continuous function $f$, the convergence
holds.
We write this as $V^{n} \overset{d_{s}}{ \to} V$ and note that stable convergence is a slightly stronger mode of convergence than weak convergence, or convergence in law, which is the special case obtained by taking $W = 1$ \citep*[see, e.g.,][for further details on stable convergence]{renyi:63a,aldous-eagleson:78a}. \citet*{jacod-shiryaev:03a} discuss the extension of this concept to stable convergence of processes. The key reason we require the convergence in law stably is that the conditional covariance matrix in the CLT of $MRC[Y]_{n}$, $\text{avar}_{\text{MRC}_{n}}$, is a function of $\sigma$ and therefore random, and the usual convergence in law is insufficient to ensure joint convergence of the bivariate vector $(MRC[Y]_{n}, \text{avar}_{\text{MRC}_{n}})$, which we need to apply the delta method to the joint asymptotic distribution and to construct confidence intervals.
Because $B \perp\!\!\!\perp \mathcal{F}$, we can write the convergence statement in Theorem (ref) as follows:
where
is the conditional covariance matrix. This means that the asymptotic distribution of $MRC \left[ Y \right]_{n}$ is mixed normal. To make use of this result to construct confidence intervals for elements of $\int_{0}^{1} \Sigma_{u} \text{d}u$ in practice, we need to estimate $\text{avar}_{\text{MRC}}$, which is addressed in section 4.\footnote{The assumption that data be equidistant is not required for the consistency to hold true. It would also apply under identical observation times (i.e., synchronous but non-equidistant data). Here, $MRC \left[ Y \right]_{n}$ is consistent with $\Delta_{i}^{n} V$ redefined as $\Delta_{i}^{n} V = V_{t_{i}} - V_{t_{i-1}}$ for any process $V$ and $k_{n} = \theta \sqrt{n} + o(n^{1 / 4})$. This is not surprising: the realised covariance also remains consistent for irregular observations (by definition). The CLT also holds, but here the variable of integration $\text{d}s$ needs to be replaced by $\text{d}H_{s}$, where $ H_{s}$ is the so-called "quadratic variation of time", see, e.g., \citet*{barndorff-nielsen-hansen-lunde-shephard:08c}.}
The $\text{avar}_{\text{MRC}}$ matrix in Theorem (ref) depends on the parameter $\theta$, or in other words the window size $k_{n}$. If the purpose is to estimate some one-dimensional parameters (such as realised covariance, regression or correlation) by real-valued functions of the MRC, it is in principle possible to minimize $\text{avar}_{\text{MRC}}$ by choosing the "best" $\theta$ for a fixed function $g$.\footnote{The optimal choice of $\theta$ will depend on the original real-valued functions of the MRC. In this sense, there is no universal optimal $\theta$.}
To illustrate this point, we focus on the estimation problem in the univariate case, $d = 1$. In this situation, the expressions reduce to
where $IV = \int_{0}^{1} \sigma_{s}^{2} \text{d}s$ and $IQ = \int_{0}^{1} \sigma_{s}^{4} \text{d}s$ are called the integrated variance and integrated quarticity, respectively. Minimizing this term with respect to $\theta$ results in solving a quadratic equation. Thus, for the optimal choice of $\theta$, say $\theta^{*}$, we get $\theta^{*} = f \left(IV, IQ, \Psi \right)$. Then, we can use an iterative procedure to find an approximation of $\theta^{*}$: (i) Choose a "reasonable" value for $\theta$ and compute $\hat{IV}$, $\hat{IQ}$ and $\hat{ \Psi}$, (ii) from these compute $\hat{ \theta}^{*}$ by setting $\hat{ \theta}^{*} = f ( \hat{IV}, \hat{IQ}, \hat{ \Psi} )$ and (iii) then, assuming the values of $\hat{ \theta}^{*}$ converge, repeat this process until $\hat{ \theta}^{*}$ has stabilized.
In practice, a simple guide that informs us about how to select $\theta$ for small values of $n$ would certainly be valuable. Unfortunately, $\theta$ comes from asymptotic statistics and therefore it does not give any precise instructions on this issue. Nonetheless, some plausible range of values of $\theta$ can be inferred from previous work in this area. For example, \citet*{jacod-li-mykland-podolskij-vetter:09a} report that the univariate pre-averaged realised variance measure is fairly robust to the choice of $k_{n}$, and they suggest to take $\theta = 1/3$. \citet*{christensen-oomen-podolskij:10a} use simulations to gauge the influence of sample size and noise on the optimal choice of $\theta$. Interestingly, they show that the MSE curve of their pre-averaged quantile-based realised variance is highly asymmetric in $k_n$ and they generally prefer to use a slightly higher value of $k_{n}$ than what would be optimal. In their empirical analysis, they also show that conservative choices of $k_{n}$ helps to heavily reduce the detrimental effects of price discreteness. In the simulation section and empirical illustration below, we therefore pick conservative values of $k_n$ by choosing $\theta = 1$, which seems to work well.
In the previous section, we used an optimal pre-averaging window length to construct the MRC, which balances the impact of the noise with the estimation of the integrated covariance matrix. This choice leads to an optimal rate of convergence - $n^{- 1 / 4}$ - but requires that we subtract an estimate of $\Psi$ to eliminate the bias induced by noise, and the final estimator is then not positive semi-definite in general. Here, we demonstrate how positive semi-definite estimates of $\int_{0}^{1} \Sigma_{s} \text{d}s$ can be formed by increasing the bandwidth parameter $k_n$ as to kill the influence of the noise, rather than balancing it. This comes at the cost of slowing down the speed at which MRC converges to the true integrated covariation.
Now, we take:
for some $0<\delta < 1/2$, and set
The following result shows that $MRC \left[ Y \right]_{n}^{ \delta}$ is consistent without a bias-correction.
In Theorem (ref), the properties of the noise process do not show up in the stochastic limit of Eq. (ref), because the influence of the noise is negligible by the choice of order for $k_{n}$ made in Eq. (ref) (refer back to Eq. (ref)). This has some appealing advantages. First, $MRC \left[ Y \right]_{n}^{ \delta}$ is positive semi-definite by construction. Second, although we state and prove this result in the i.i.d. noise case, Theorem (ref) does in fact allow for more general noise dynamics than in Theorem (ref). In particular, so long as $\mathbb{E} \left( \epsilon_{t} \mid X \right) = 0$ and $\bar{ \epsilon}_{i}^{n}$ admits asymptotic normality at the usual rate $k_{n}^{- 1 / 2}$, the theorem will hold (so we do not require any assumptions on the dependence between $X$ and $\epsilon$). Of course, the rate $k_{n}^{- 1 / 2}$ is achieved in the i.i.d. case, but there are more general cases where it also holds (e.g., for $q$-dependent and mixing processes).
To show the CLT, we require a restriction on $\delta$. This is because the bias caused by the noise, which is negligible in Theorem (ref), becomes more substantial when multiplying with the rate of convergence.
Theorem (ref) amounts to a classical bias-variance trade-off. It shows the expected result that using a longer pre-averaging window $k_{n} = O(n^{1/2 + \delta})$ averages enough to make $MRC \left[ Y \right]_{n}^{ \delta}$ consistent without a bias-correction, but it also slows down its rate of convergence. The larger is $\delta$, the harder is this effect. Note that the asymptotic variance depends solely on the volatility process $\sigma$, while the two noise-dependent terms (a cross-term and a pure noise term) appearing in Eq. (ref) are wiped out.
The optimal choice $\delta = 0.1$ results in a $n^{- 1 / 5}$ rate of convergence, and a large sample bias of the order $n^{-1 / 5} \frac{ \psi_{1}}{ \theta^{2} \psi_{2}} \Psi$. If desired, the bias can be estimated using $\hat{ \Psi}_{n}$, but it should be small for more recent data and could be ignored.\footnote{If we do bias-correct $MRC \left[ Y \right]_{n}^{ \delta}$, the resulting estimator is then again not ensured to be positive semi-definite.}
This result bears some resemblance to that of the multivariate realised kernel of \citet*{barndorff-nielsen-hansen-lunde-shephard:08c}. A flat-top kernel estimator can be designed to converge at rate $n^{-1/4}$, but the resulting estimator may go negative. To guarantee positive semi-definiteness, the authors propose kernel corrections that are not entirely flat-top and they show this results in $n^{-1 /5}$ rate of convergence and a non-zero asymptotic mean as produced here as well.
We can indeed make a stronger link between the multi-scale RV, the kernel approach and our pre-averaging method, when estimating quadratic variation. Here, we provide some more details about this relationship. To fix ideas, we take $d = 1$ (i.e., the univariate setting) and compare only to the kernel approach. \citet*{barndorff-nielsen-hansen-lunde-shephard:08a} then show how the multi-scale RV fits into the realised kernel setting.
When we explicitly include the bias-correction and ignore finite sample issues, the $MRC \left[ Y \right]_{n}$ estimator is given by the formula
Consider a flat-top kernel-based estimator with kernel weights
where $\phi_{2}(s)$ and $\psi_{2}$ are defined as in Section (ref). We call the function $k$ a flat-top kernel, if (i) $k(0) = 1$, (ii) $k(1) = 0$ and (iii) $k'(0) = k'(1) = 0$. We note that $k(s) = \phi_2(s) / \psi_2$ satisfies all three conditions: (i) and (ii) are trivial, while the third condition follows from the identity:
and integration by parts (recall that $g(0) = g(1) = 0$).
In the next step, we map the MRC statistic into a kernel-like one as follows
with
and
for $1\leq h\leq k_n-2$.
An example and some remarks are now in order.
Example \ Take $g(x) = \min(x, 1 - x)$ for $x\in [0,1]$, which is our canonical weight function. Then, the corresponding kernel $k(s)= \phi_2(s) / \psi_2$ is the Parzen kernel.
Remark \ Apart from border terms, i.e. terms close to 0 and 1, the pre-averaging estimator coincides with the kernel-based estimator using the flat-top kernel function $k(s) = \phi_2(s)/\psi_2$. Both estimators possess the optimal rate of convergence $n^{-1/4}$, but the border terms affect their asymptotic distribution. In fact, to arrive at their central limit theorem for the realised kernel, \citet*{barndorff-nielsen-hansen-lunde-shephard:08c} need to apply some averaging (or "jittering") to edge terms, while the pre-averaging estimator is asymptotically mixed normal "by construction".
Remark \ Using the definition $k(s)= \phi_2(s)/\psi_2$, we learn that for every weight function $g$ there exists a flat-top kernel $k$. However, the converse statement is not true in general. Thus, the class of flat-top kernels is larger than our class of pre-averaging functions, but this does not appear to be a noticeable disadvantage for practical applications \citep*[][for example, advocate using the Parzen kernel in practice, which corresponds to our $g$ function by the example]{barndorff-nielsen-hansen-lunde-shephard:08c}.
It is worthwhile to underscore that pre-averaging is a general approach, which can be used to estimate various characteristics of semimartingales, when they are cloaked with noise. \citet*{jacod-li-mykland-podolskij-vetter:09a}, for example, pre-average realised variance, but it has also been used in jump-robust inference in conjunction with the bipower variation statistic \citep*[e.g.,][]{podolskij-vetter:09a} or the quantile-based realised variance \citep*[e.g.,][]{christensen-oomen-podolskij:10a}. This is also the reason, why we can estimate the asymptotic conditional variance in the CLT using the same type of estimator to deliver a feasible result. Other estimation methods are typically designed to estimate quadratic variation and cannot, in general, be used to solve other estimation problems.
Non-synchronous trading has long been a recognized feature of multivariate high-frequency data \citep*[see, e.g.,][]{fisher:66a,lo-mackinlay:90a}. This causes spurious cross-autocorrelation amongst asset price returns sampled at regular intervals in calendar time, as new information gets built into prices at varying intensities. It is well-known that high-frequency realised covariance estimates, using for example the previous-tick method to align prices, are biased towards zero in this setting \citep*[e.g.,][]{epps:79a}.
Motivated by these shortcomings of realised covariance, a number of alternative procedures have been proposed in the literature. \citet*{scholes-williams:77a} and \citet*{dimson:79a} suggested to include leads and lags of sample autocovariances of high-frequency returns into the realised covariance to correct for stale prices. This results in a bias reduction but also increases the variance of the estimator \citep*[e.g.,][]{griffin-oomen:06a}. Importantly, the lead-lag realised covariance is generally inconsistent in the presence of noise. More recently, \citet*{zhang:08a} extends the two-scale RV to the multivariate setting and shows that it is consistent under non-synchronous prices and noise. An analytic derivation of the impact of Epps effect on realised covariance under previous-tick sampling is also given. In independent and concurrent work to ours, \citet*{barndorff-nielsen-hansen-lunde-shephard:08c} propose a multivariate realised kernel. They rely on the concept of refresh time to match prices in calendar time and show that stale pricing errors under refresh time sampling do not change the asymptotic distribution of the multivariate realised kernel.
A related approach can be taken to make the MRC robust to non-synchronous observation times. In particular, under suitable regularity conditions on the sampling times, non-synchronous trading does not alter the asymptotic distribution of the MRC, when constructing synchronous time series of prices. A multivariate time series of high-frequency returns obtained in this fashion can therefore be plugged into the asymptotic theory developed above without concern. The proof of these results are highly technical, but their validity should be clear in the light of the comparison with the multivariate realised kernel given above. Therefore, we omit the detailed proofs of these results. Instead, we conduct some simulations in section (ref) to verify the correctness of these conjectures.
\citet*{hayashi-yoshida:05a,hayashi-yoshida:08a} develop an alternative procedure for covariance measurement in the noise-free case, which is based on the original non-synchronous data \citep*[see, e.g.,][for related work]{dejong-nijman:97a,martens:03a,palandri:06a,corsi-audrino:07a}. This estimator has the profound advantage that it does not throw away information that is typically lost using a synchronization procedure. Here, we show how this estimator can also be made robust to noise by using pre-averaging.
Given the vector of log-prices $Y = (Y^{1}, \ldots, Y^{d})'$, which is defined by the noisy diffusion model in Eq. (ref), we now assume that the component processes $(Y^{k})$ are observed at non-random time points $t_{i}^{(k)}$, for $i = 1, \ldots, n_{k}$, with $(t_{i}^{(k)})$ being a partition of the interval $[0,1]$ and $k = 1, \ldots, d$. In addition, we need some regularity conditions on the sampling such that all time schemes are comparable in the following way: $\max |t_{i}^{(k)} - t_{i - 1}^{(k)}| \rightarrow 0$ as $n_{k} \to \infty$, for $k = 1, \ldots, d$ and
for $1 \leq k, l \leq d$ and some $K > 0$, where $K$ is independent of $n_{k}$, for $k = 1, \ldots, d$. The latter condition says that data from one process do not cluster in any single interval of the others. Finally, the following condition is needed
where $c>0$ is a constant independent of $n_k$ for all $1\leq k\leq d$. This condition implies that $n_k \max|t_i^{(k)} - t_{i-1}^{(k)}|\leq c$, which could be assumed instead.
A consistent estimator of the integrated covariance $\int_{0}^{1} \Sigma_{s}^{kl} \text{d}s$ between assets $Y^{k}$ and $Y^{l}$ can then be constructed as follows. We set $n = \sum_{k = 1}^{d} n_{k}$ and define the statistic
where $\psi_{HY} = \int_{0}^{1} g(x) \text{d}x$ and $\mathbb{I}_{ \left\{ \bullet \right\} }$ is the indicator function discarding pre-averaged returns that do not overlap in time. Hence, $HY[Y]_n^{(k,l)}$ is a pre-averaged version of the \citet*{hayashi-yoshida:05a} estimator. Note that under the previous conditions, $n$, $n_{k}$ and $n_{l}$ are of the same order and that $n$ controls the universal pre-averaging window $k_{n}$.\footnote{Our choice here implies that $k_{n}$ is identical across all pairs of asset combinations. In general, the best way of choosing $k_{n}$ depends on what is being estimated (e.g., covariance, correlation or beta). Thus, if one only cares about efficiency in pairwise covariance estimation, a better $HY[Y]_n^{(k,l)}$ estimator might be based on a local $k_n$, for example $k_{n}^{k,l} = \theta_{k,l} \sqrt{n} + o(n^{1/2})$. This would also serve to make the estimated covariances universe independent (i.e., they will not change by adding or removing assets). To make a more qualified, statistical argument on this issue, however, we need to compare an expression for the asymptotic variance of $HY[Y]_{n}^{(k,l)}$ under different ways of choosing $k_n$ (e.g., from a CLT). This is beyond the scope of this paper and will be left for future research, where we hope to shed more light on the subject.}
Interestingly, Theorem (ref) shows the somewhat surprising result that there is no asymptotic noise-induced bias in $HY[Y]_{n}^{(k,l)}$, not even when the spacings $(t_{i}^{(k)})$ and $(t_{j}^{(l)})$ are identical. This can be seen as follows. First, under the i.i.d. structure on the noise only products of the form $\epsilon_{t_{i}^{(k)}}^{k} \epsilon_{t_{j}^{(l)}}^{l}$ with $t_{i}^{(k)} = t_{j}^{(l)} = t$ contribute to potential bias. We consider that set of points and assume that $k_{n} \leq i \leq n_{1} - k_{n}$ and $k_{n} \leq j \leq n_{2} - k_{n}$ (this is innocent, for the summands which do not fulfill this are negligible). Then, an inspection of $HY[Y]_{n}^{(k,l)}$ shows that all products $\epsilon_{t}^{k} \epsilon_{t}^{l}$ appear with the factor
in front. But $\sum_{j = 0}^{k_{n} - 1} g( \frac{j + 1}{k_{n}} ) - g( \frac{j}{k_{n}}) = 0$, because $g(0) = g(1) = 0$, so these terms drop out of the summation.
While $HY[Y]_{n}^{(k,l)}$ has the advantage that it is free of prior alignment of log-prices and hence does not throw away information in the sample, it does suffer from being a pairwise estimator, which means that once we assemble all the single variance/covariance estimates into a full covariance matrix, it is not guaranteed to be positive semi-definite. Still, there are some problems in financial economics, in which it is only the accuracy of the estimator that matters and positive semi-definiteness is less important, for example in asset allocation and risk management under gross exposure constraints \citep*[e.g.][]{fan-zhang-yu:09a}. Moreover, our empirical results show that $(HY[Y]_{n}^{(k,l)})_{1 \leq k,l \leq d}$ does not fail to be positive semi-definite on a single day for the $d = 5$-dimensional vector of asset prices considered there.
Proposition (ref) shows that the rate of convergence associated with $HY[Y]_{n}^{(k,l)}$ is $n^{-1/4}$. A proof of the much stronger result, the CLT for $HY[Y]_{n}^{(k,l)}$, will be given in a companion paper to this one \citep*[see][]{christensen-podolskij-vetter:10a}. In the empirical section, we gauge the properties of this estimator on actual data and find that the pre-averaged version performs very well.
In order to make Theorem (ref) and (ref) feasible, we need to estimate the asymptotic covariance matrix $\text{avar}_{\text{MRC}}$, as it appears in Eq. (ref). Here, we give an explicit estimator of $\text{avar}_{\text{MRC}}$. More precisely, we present an estimator of the asymptotic covariance matrix of the vectorized statistic in Eq. (ref).
First, we set
where $\text{vec}(\cdot)$ is the vectorisation operator that stacks columns of a matrix below one another. Next, we define the statistic
We should note that $V_n(g)$ depends on both the bandwidth parameter $\theta$ and the pre-averaging function $g$, and that it is positive semi-definite by construction. Moreover, for any $1\leq k,k',l,l'\leq d$, we get the following convergence
where $a_B(g,\theta )=\theta^2 \psi_{2}^2$, $a_M(g,\theta )= \psi_{1} \psi_{2}$ and $a_N(g,\theta ) = \frac{ \psi_{1}^2}{ \theta^2}$ \citep*[the proof of this result is achieved by using arguments alike the ones presented in the proof of Theorem (ref); see][for further details]{kinnebrock-podolskij:08b}.
What needs to be estimated is
with all constants referring to a given function $g_0$. Suppose that $g_{0}(x) = \min(x,1 - x)$ as above. Now, we take three different functions $g_1$, $g_2$ and $g_3$ that satisfy all the conditions of $g_0$, and which are chosen such that the matrix
is invertible. We can then construct a weight vector
Finally, we form the statistic
where $V_n(g_k)$ are the above estimators associated with the functions $g_k, k = 1,2,3$. Then, it holds that
for any $1\leq k,k',l,l'\leq d$.\footnote{The weights $\Big(\frac{2 \Phi _{22} \theta}{\psi _2^2 }, \frac{2 \Phi _{12} }{\psi _2^2 \theta}, \frac{2 \Phi _{11} }{\psi _2^2 \theta^3} \Big)$ in Eq. (ref) are the coefficients in front of $(\int_0^1 \Lambda_u \text{d}u, \int_0^1 \Theta _u \text{d}u, \Upsilon)$ in the expression for the asymptotic variance of MRC. This choice results in a consistent estimator of $\text{avar}_{\text{MRC}}$. More generally, one can estimate any linear combination of $(\int_0^1 \Lambda_u \text{d}u, \int_0^1 \Theta _u \text{d}u, \Upsilon)$ by choosing an appropriate weight vector in Eq. (ref). For example, the multivariate version of the integrated quarticity, $\int_0^1 \Lambda_u \text{d}u$, can be estimated by plugging the linear combination $(1,0,0) A^{-1} (g_1,g_2,g_3)$ into Eq. (ref).}
There are various classes of functions from which the $g_k$ can be selected, for example $g(x) = x^{a} \left( 1 - x \right)^{b}$ with $a,b \geq 1$ or $g(x) = \sin(c \pi x)$ for integer $c$. If one manages to find a set of functions $g_k$ for which the coefficients $C^{(k)}(g_1,g_2,g_3)$ are all positive, $k = 1, 2$ and $3$, then the statistic in Eq. (ref) is guaranteed to be positive semi-definite, because it is a linear combination of positive semi-definite statistics $V_n(g_k)$ using positive weights $C^{(k)}(g_1,g_2,g_3)$. It appears to be quite hard to find such a combination in practice, and so far we did not succeed at this. Nonetheless, $\widehat{\text{avar}}_{\text{MRC},n}$ remains a consistent estimator.
The results in section 3 and 4 can be applied in order to compute confidence intervals for some functionals of $\int_{0}^{1} \Sigma_{u} \text{d}{u}$ that are important in practice, such as covariance, regression and correlation. For the $i$th and $j$th asset, these quantities are given by
Theorem (ref), (ref) or (ref) can be invoked to provide consistent estimates of $\int_0^1 \Sigma_u^{ij} \text{d}u$, $\beta ^{(ji)}$ and $\rho^{(ji)}$, e.g. for $MRC \left[Y \right]_{n}$
for any $1\leq i,j\leq d$. The estimators in (ref) are called modulated realised covariance, regression and correlation, respectively. In the next theorem we present the associated feasible central limit theorems, which follow from Theorem (ref) and the delta method for stable convergence.
All the required terms are easy to compute, so it is rather simple to implement the estimators.
We now demonstrate the finite sample accuracy of some of the asymptotic results developed above using simulations. The design of our Monte Carlo study, which we briefly describe here, is identical to the analysis conducted in \citet*{barndorff-nielsen-hansen-lunde-shephard:08c}.
Specifically, to simulate log-prices we consider the following bivariate stochastic volatility model
where $B^{(i)}$ and $W$ are independent Brownian motions. In this model, the term $\rho^{(i)} \sigma_{t}^{(i)} \text{d}B_{t}^{(i)}$ is an idiosyncratic component, while $\sqrt{1 - [\rho^{(i)}]^2} \sigma_{t}^{(i)} \text{d}W_{t}$ is a common factor.
The spot volatility is modeled as $\sigma_{t}^{(i)} = \exp (\beta_0^{(i)} + \beta_1^{(i)} \varrho_{t}^{(i)})$ with an Ornstein-Uhlenbeck specification for $\varrho_{t}^{(i)}$: $\text{d} \varrho_{t}^{(i)} = \alpha^{(i)} \varrho_{t}^{(i)} \text{d}t + \text{d}B_{t}^{(i)}$. This implies that there is perfect correlation between the innovations of $\rho^{(i)} \sigma_{t}^{(i)} \text{d}B_{t}^{(i)}$ and $\sigma_{t}^{(i)}$, while it is $\rho^{(i)}$ between the increments of $X_{t}^{(i)}$ and $\varrho_{t}^{(i)}$. Finally, the magnitude of correlation between the two underlying price processes $X_{t}^{(1)}$ and $X_{t}^{(2)}$ is $\sqrt{1 - [\rho^{(1)}]^2} \sqrt{1 - [\rho^{(2)}]^2}$. The reported results are based on the following configuration of parameters for both processes: $(a^{(i)}, \beta_0^{(i)}, \beta_1^{(i)}, \alpha^{(i)}, \rho^{(i)}) = (0.03, -5/16, 1(8,-1/40,-0.3)$, so that $\beta_0^{(i)} = [\beta_1^{(i)}]^{2} / [2 \alpha^{(i)}]$. We note that this particular choice of parameters also means that the volatility process has been normalized, in the sense that $\mathbb{E} ( \int_0^1 [\sigma_s^{(i)}]^2 \text{d}s) = 1$.
We simulate 1,000 paths of this model over the interval $[0,1]$. Motivated by our empirical data, we let $[0,1]$ represent 6.5 hours worth of trading, which is then further decomposed into $N = 23,400$ subintervals of equal length $1 / N$ ($N$ denoting the number of seconds in 6.5 hours). In constructing noisy prices $Y^{(i)}$, we first generate a complete high-frequency record of $N$ equidistant observations of the efficient price $X^{(i)}$ using a standard Euler scheme. The initial values for the $\varrho_{t}^{(i)}$ processes at each simulation run is drawn randomly from the stationary distribution of $\varrho_{t}^{(i)}$, which is $\varrho_{t}^{(i)} \sim N(0,[-2\alpha^{(i)}]^{-1})$.\footnote{Note that the Ornstein-Uhlenbeck process admits an exact discretization \citep*[see, e.g.,][]{glassermann:04a}. We use that result here to avoid discretization errors in approximating the continuous time distribution of $\text{d} \varrho^{(i)}$ over discrete time steps of size $\Delta t = 1 / N$.}
We add simulated microstructure noise $Y^{(i)} = X^{(i)} + \epsilon^{(i)}$ by taking
This choice again follows \citet*{barndorff-nielsen-hansen-lunde-shephard:08c} and means that the variance of the noise process increases with the level of volatility of $X^{(i)}$, as documented by \citet*{bandi-russell:06a}. $\gamma^{2}$ takes the values 0, 0.001, 0.01, which covers scenarios with no noise through low-to-high levels of noise.
Finally, we extract irregular, non-synchronous data from the complete high-frequency record using Poisson process sampling to generate actual observation times, $\{ t_{j}^{(i)} \}$. In particular, we consider two independent Poisson processes with intensity parameter $\lambda = \left( \lambda_1, \lambda_2 \right)$. Here $\lambda_{i}$ denotes the average waiting time (in seconds) for new data from process $Y^{(i)}$, so that an average day will have $N / \lambda_{i}$ observations of $Y^{(i)}, i = 1,2$. We vary $\lambda_{1}$ through $(3, 5, 10, 30, 60)$ to capture the influence of liquidity on the performance of our estimators, and we set $\lambda_{2} = 2 \lambda_{1}$ such that on average $Y^{(2)}$ refreshes at half the pace of $Y^{(1)}$.
In Table (ref), we report the results of the simulations. As the results for estimating the variance components of the $2 \times 2$ covariance matrix are as expected compared to prior work, we give focus here towards estimating the integrated covariance, correlation and beta. Also, because the refresh time sampled MRC estimators perform somewhat better than the previous-tick based MRC estimators, we only report the results for the MRC estimators based on refresh time sampling ($n = RT$).\footnote{All unreported results are available upon request.} In the three panels of the table, we therefore provide the bias and root mean squared error (rmse) of the various estimators in terms of estimating the integrated covariance, correlation and beta. We compare our results to a standard realised covariance sampled at either a 1-minute or 15-minute frequency.
Looking at the table, we see that the MRC estimators are very efficient across all scenarios of noise and non-synchronous trading considered here, matching or outperforming the standard realised covariance. The $HY[Y]^{(k,l)}$ estimator is less efficient, owing largely to a larger finite sample bias in this estimator. We will study the possibilities of making finite sample adjustments to $HY[Y]^{(k,l)}$, elsewhere.
Above, we conjectured without a proof that the MRC was robust to stale prices and retained its rate of convergence under non-synchronous trading. If this is to be true, and ignoring finite sample biases, we should then expect the rmse of the two MRC estimators to decrease at rate $n^{-1/4}$ and $n^{-1/5}$. As can be seen from the table, this is exactly what we find. For example, when estimating the integrated covariance and going from $\lambda = (60,120)$ to $\lambda = (3,6)$, i.e. an approximate twenty-fold increase in sample size, the rmse of $MRC[Y]_n$ decreases roughly by half and the rmse of $MRC[Y]_n^{\delta}$ by slightly less than a half, which is consistent with the rates given above.
All in all, the simulation results show that the estimators proposed in this paper are very good at estimating the integrated covariance, correlation and beta across a wide range of noise and liquidity scenarios, and that the asymptotic predictions given above are also reasonable guides to their finite sample behavior.
To illustrate some empirical features of the pre-averaging theory developed above, we retrieved high-frequency data for a five-dimensional vector of assets from Wharton Research Data Services (WRDS). We picked four equities at random from the S&P 500 constituents list as of July 1, 2009. We then added a 5th element, namely the S&P 500 Depository Receipt (ticker symbol SPY), which is an exchange-traded fund that tracks the large-cap segment of the U.S. stock market. As such, it can be viewed as generating market-wide index returns. The four remaining stocks are the following (with ticker symbol and industry classification in parenthesis): Bristol-Myers Squibb (BMY, health care), Lockheed Martin (LMT, industrials), Oracle (ORCL, information technology) and Sara Lee (SLE, consumer staples), thus representing a broad category of industries. We use both trades and quotes data for the sample period that covers the whole of 2006, which results in 251 trading days.
Table (ref) reports some descriptive statistics for our universe of stocks and sample period. As can be seen, these equities display varying degrees of liquidity with ORCL and SPY being the most liquid, while LMT and SLE are the least liquid. Also reported in the table is the univariate noise ratio statistic, $\gamma$, which is a noise-to-signal measure that describes the level of microstructure noise to integrated variance \citep*[see, e.g.,][for further details on the noise ratio]{oomen:06a}. Generally speaking, there is a tendency for more frequently traded companies to contain less microstructure noise, the notable exception being the transaction data for ORCL.
As a preliminary step, we subdued the sample data to some cleaning procedures. Pre-cleaning high-frequency data is necessary, because the raw data has many invalid observations (e.g., data with misplaced decimal points, or trades that are reported out-of-sequence). Our filter is roughly identical to that used by \citet*{barndorff-nielsen-hansen-lunde-shephard:08c} with some minor differences. Here, we briefly describe the filtering rules we employ.
Trades and quotes: The following rules are applied to both trades and quotes data. a) We keep data from a single exchange: Pacific for SPY and primary exchange for the 4 remaining equities, see Table (ref), b) we delete data with time stamps outside the regular exchange opening hours from 9:30am to 4:00pm, c) we delete rows with a transaction price, bid or ask quote of zero, and d) we aggregate data with identical time stamp using volume-weighted average prices (using total transaction volume or quoted bid and ask volume, respectively).
Trades only: We delete entries with a correction indicator $\neq$ 0 or with abnormal sales condition.
Quotes only: We delete quotes with negative spreads and rows where the quoted spread exceeds 10 times the median spread for that day.
Table (ref) also reports how many observations that are lost by passing these filters through the data. It should be noted that the "Trades Only" and "Quotes Only" filters generally tend to reduce the sample by only a very small fraction.
Here, we inspect the outcome of applying the estimators introduced above, after which we look at transforms of the covariance matrix. As a comparison, we also compute the standard realised covariance from 15-min, 5-min, and 15-sec previous-tick data.
We implement both the $MRC \left[ Y \right]_{n}$ and $MRC \left[ Y \right]_{n}^{ \delta}$ estimators of section 2 and 3.4, respectively. Recall that $MRC \left[ Y \right]_{n}$ converges at rate $n^{-1 / 4}$, it needs to be corrected for bias, and as a result is not guaranteed to be positive semi-definite. We base $MRC \left[ Y \right]_{n}^{\delta}$ on $\delta = 0.1$, among many plausible choices, which results in a $n^{-1 / 5}$ rate of convergence, a small finite sample bias that we omit correcting and, hence, positive semi-definiteness by construction. Two sampling schemes are used, calendar time and refresh time sampling, which yields a total of four combinations. We use pre-averaging windows found as $k_{n}= \lfloor \theta n^{ \delta'} \rfloor$, where $\theta = 1$ and $\delta' = \left(0.5, 0.6 \right)$ for our two choices. The selection of $\theta$ follows the conservative rule discussed above. We set $n = 390$ for the calendar time-based estimators, which is 1-minute sampling, while $n$ is determined automatically by the data for the refresh time sampling scheme.
In Table (ref), we report the sample covariance matrix estimates averaged across the 251 days. This table is constructed in the usual way, displaying the results based on transaction prices in the upper diagonal (including the main diagonal), while the strict lower diagonal elements are the corresponding results based on quotation data. Consistent with prior literature, we see from the table that the standard realised covariance is suffering from Epps effect, when sampling runs quickly. All estimated covariance terms lie in the positive region, but for $Cov^{15s}$ they are heavily compressed towards zero. This is less of a concern for $Cov^{avg(5m,15m)}$, which should tend to capture the average level of the covariance structure well, while not being seriously influenced by microstructure frictions and Epps effect. Turning next to the estimators proposed in this paper, we note that the time series average of both MRC versions and the noise-robust HY estimator are in line with that produced by $Cov^{avg(5m,15m)}$, showing that they appear free of any systematic bias. The \citet*{hayashi-yoshida:05a} estimator produces a strong downwards bias in the covariance estimates, when it is applied directly to noisy and irregular high-frequency data. This reaffirms previous empirical work \citep*[see, e.g.,][]{barndorff-nielsen-hansen-lunde-shephard:08c,griffin-oomen:06a,voev-lunde:07a}, so these results are not surprising or novel. The pre-averaged version $HY[Y]_{n}^{(k,l)}$, however, does a much better job and tends to agree with the average level of other noise-robust estimators.
Finally, we turn to the issue of positive semi-definiteness. As noted above, $MRC \left[ Y \right]_{n}$ and $HY[Y]_{n}^{(k,l)}$ are not guaranteed to possess this property. Nonetheless, they do not fail to be positive semi-definite on a single instance across our sample period. This is true for both the transaction and quotation data, and both combinations of the bias-corrected MRC estimator. Thus, while theoretically a concern, this problem does not appear to occur frequently in practice, although the conclusion might change for other data sets.
We now focus on estimating $\beta^{(ji)}$ by $\hat{ \beta}_{n}^{(ji)}= MRC \left[ Y \right]_{n}^{i, j} / MRC \left[ Y \right]_{n}^{i, i}$, where we take $i = \text{SPY}$ and form regressions by using $j = \text{BMY, LMT, ORCL, SLE}$. This type of regression, where individual equity covariances with the market are regressed onto a market-wide realised variance measure, is important in financial economics, for example within the conditional CAPM \citep*[see, e.g.,][]{jagannathan-wang:96a,lettau-ludvigson:01a}, since only systematic risk should be rewarded with expected excess returns.
In Figure (ref), we plot the MRC-based betas from transaction prices and refresh time sampling. The corresponding plots from the other estimators proposed in this paper are qualitatively similar. As in \citet*{barndorff-nielsen-hansen-lunde-shephard:08c}, we smooth the daily beta estimates by passing them through an ARMA(1,1) filter. The figure shows that beta is time-varying and predictable, and that it tends to fluctuate around its mean level. Evidently, the estimated processes exhibit substantial memory with autoregressive roots at 0.90 or higher \citep*[see also][who study the persistence of quarterly realised beta estimated from daily asset returns]{andersen-bollerslev-diebold-wu:06a}.
In this paper, we present a simple solution, based on applying pre-averaging to financial high-frequency data, to the problem of how to estimate the multivariate ex-post integrated covariance matrix, possibly in the simultaneous presence of market microstructure noise and non-synchronous trading.
A modulated realised covariance (MRC) estimator is introduced. The MRC bears close resemblance to a standard realised covariance, being a sum of outer products of high-frequency returns, but it relies instead on pre-averaging to reduce the harmful impact of microstructure noise. We study the properties of this new estimator by showing its consistency and asymptotic mixed normality under mild conditions on the dynamics of the price process. As shown in the paper, the MRC can be configured to possess an optimal rate of convergence or to guarantee positive semi-definite covariance matrix estimates. In the presence of non-synchronous trading, we also outline how to modify the MRC by using an imputation scheme (for example the previous-tick rule or refresh time sampling) to match high-frequency prices in time. An MRC constructed on the back of such artificial returns will again be consistent for the integrated covariance.
Another novelty developed in this paper is a pre-averaged version of the \citet*{hayashi-yoshida:05a} estimator that can be implemented directly on the raw noisy {\em and\/} non-synchronous observations, without any prior alignment of prices. We also show the consistency of this estimator and derive a rate for its variance, but otherwise we defer further theoretical analysis of finite sample improvements and asymptotic properties of the pre-averaged Hayashi-Yoshida estimator to future research \citep*[see][]{christensen-podolskij-vetter:10a}.
We demonstrate with a set of simulations that these newly proposed estimators can bring substantial efficiency gains with them compared to a standard realised covariance in the realistic setting with both microstructure noise and non-synchronous trading. Furthermore, an empirical illustration highlights their applicability to real high-frequency data. We therefore look forward to future applications of these estimators, including an investigation of their informational content about future volatility. Being able to produce good forecasts of future volatility is paramount in financial economics. Thus, as an example, it could be interesting to apply our estimators in the context of portfolio choice to calculate their economic value, akin to \citet*{fleming-kirby-ostdiek:01a,fleming-kirby-ostdiek:02a}, \citet*{bandi-russell:06a} and others. However, generally they should also be useful in many other areas, including asset- and option pricing or risk management.