EconBase
← Back to paper

Cointegration in high frequency data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

91,496 characters · 22 sections · 90 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Cointegration in high frequency data

abstractIn this paper, we consider a framework adapting the notion of cointegration when two asset prices are generated by a driftless It\^{o}-semimartingale featuring jumps with infinite activity, observed regularly and synchronously at high frequency. We develop a regression based estimation of the cointegrated relations method and show the related consistency and central limit theory when there is cointegration within that framework. We also provide a Dickey-Fuller type residual based test for the null of no cointegration against the alternative of cointegration, along with its limit theory. Under no cointegration, the asymptotic limit is the same as that of the original Dickey-Fuller residual based test, so that critical values can be easily tabulated in the same way. Finite sample indicates adequate size and good power properties in a variety of realistic configurations, outperforming original Dickey-Fuller and Phillips-Perron type residual based tests, whose sizes are distorted by non ergodic time-varying variance and power is altered by price jumps. Two empirical examples consolidate the Monte-Carlo evidence that the adapted tests can be rejected while the original tests are not, and vice versa.

Keywords: cointegration; deflation; high frequency data; It\^{o}-semimartingale; residual based test; truncation; unit root test \\

Introduction

It is often the case that time series cruise componentwise, but that a linear combination of the components does not drift apart. Since the seminal papers of granger1981some and engle1987co, cointegration has spread across and way beyond the field of econometrics. The authors bring forward a residual based two-step strategy to test for the presence of cointegration. The first step corresponds to estimation of cointegrating relations via regression. The second step is closely related to test of unit roots. More specifically, residual based tests are designed to test the null of no cointegration by relying on a unit root test on the residuals. If the null of unit root test is not rejected, then the null of no cointegration is also not rejected. Among the class of unit root tests, standard Dickey-Fuller (DF) tests originally from dickey1981likelihood and augmented procedure from said1984testing, and Phillips-Perron tests from phillips1988testing, are among the most popular. phillips1987time shows that this type of unit root tests is typically robust to many weakly dependent and heterogeneously distributed time series. As for cointegration tests, a large sample limit theory for the residual based tests is investigated in phillips1990asymptotic. In particular, the critical values tabulated for the pure unit root tests are altered due to the error made in the first step. In the cited paper, and to the best of our knowledge in the rest of the literature on cointegration, the asymptotics is low frequency in the sense that the time gap between two observations $\Delta $ is fixed while the horizon time $T \rightarrow \infty$.

There is a solid body of empirical work employing cointegration in financial economics. In that field, the most common specification of cointegration definition is that all the components of the observed vector $x_t$ are unit roots (e.g. random walks) and that there exists a vector $\alpha$ so that $\alpha' x_t$ is stationary. Examples of application include and are not limited to nominal dollar spot exchange rates, e.g. baillie1989common and diebold1994cointegration, price discovery, e.g. hasbrouck1995one, and pairs trading, e.g. caldeira2013selection and the review of krauss2017statistical. Although the major part of earlier empirical studies was confined to data observed on a daily basis, nowadays they frequently incorporate data observed on the intraday basis, see e.g. elyasiani2001interdependence, hasbrouck2003intraday, pati2011intraday or yang2012intraday among others. This increasing use of high frequency data is perfectly natural, and should even be encouraged given the basic statistical principle that one shall not throw away data. As a matter of fact, it echoes similar changes in other areas of the literature, the best example being the diversification and sophistication of efficient variance estimation measures over the past two decades.

Our concern in this paper is to back up the empirical use of high frequency data by theoretically validated high frequency robust estimation of cointegrating relations and tests. More specifically, the difference between the classical time series framework and the high frequency framework we consider is that in the latter $\Delta \rightarrow 0$ while keeping $T \rightarrow \infty$, and the observed time series $x_t$ will be generated by a driftless It\^{o}-semimartingale. This environment, quite standard in high frequency financial econometrics, typically accommodates for a variety of stylized facts, such as time-varying variance featuring jumps, leverage effect and price jumps, all of which are salient in high frequency data. Price jumps can happen at deterministic times (e.g. macroeconomic news announcements) or not, are usually of random size, and quite frequent, at least ten a year on average according to Table 8 (p. 484) in huang2005relative. For the purpose of simplicity, we restrict ourselves to the two dimensional case.

The question of testing for no cointegration with high frequency data is of practical relevance since some of its stylized facts can lead to significant distortions of the power of the usual tests. In an extensive Monte Carlo experiment, krauss2017power implement ten leading cointegration tests in a variety of set-up tailored with high-frequency features. They use an AR(1) with normal innovations as benchmark, but also look at non-normality effects employing t-distributed innovations, GARCH effects, nonlinearities and price jumps. They typically find that among a cohort of high frequency stylized facts, price jumps deteriorate the power the most.

In fact, the current theoretical framework to test for no cointegration is quite restricted and only accommodates for ergodic time-varying variance. Spurious cointegration can occur in the presence of a mere jump in long term equilibrium variance, as documented by noh2003behaviour. In other words, this means that non ergodic time-varying variance can distort the size of the tests. Obviously, time-varying variance is not solely a high frequency stylized fact, as sensier2004testing report that most of the real and price variables in the dataset from stock1999comparison reject the null hypothesis of constant variance against one regime shift alternative. As far as we know, the only attempt to accommodate for time-varying variance in a more flexible cointegration model is kim2010cointegrating, but unfortunately in that paper the authors test the reverse null of cointegration.

Our aim is to adapt the residual based DF procedure developed in phillips1990asymptotic to test for the null hypothesis of no cointegration in high frequency data. More specifically, the residual based tests of the cited paper are not theoretically robust to the two aforementioned high frequency stylized facts pointed by the literature as distorting size and power -time-varying variance and price jumps-. Accordingly, we develop two separate adaptations, one for each feature, that we eventually combine with each other, as the two problems are not of the same nature and cannot simply be tackled with a unique straightforward adaptation. As far as we know, no cointegration test has been tuned to those two high frequency stylized facts, yet there are accordingly two very related fields from the literature.

The first active field is about testing for the presence of a unit root incorporating time-varying variance. cavaliere2008time apply a time transformation to the data prior to use the standard unit root tests and retrieve the same asymptotics. cavaliere2007testing employ the time deformation to simulate valid critical values for standard unit root tests. cavaliere2008bootstrap and cavaliere2009heteroskedastic consider a related method involving the wild bootstrap. Wild bootstrap tests formed from a feasible generalized least squares are developed in boswijk2018adaptive. Finally, beare2018unit preestimate volatility and exploit it to deflate the returns prior the use of the standard tests. Unfortunately, we were not able to find track of the use of those nice methods in the context of residual based cointegration test. In addition, the theoretical framework of all those papers is restricted to deterministic volatility and no price jumps, except for cavaliere2009heteroskedastic who allow for random variance.

Another very dynamic field is about residual based cointegration tests with regime changes, which is typically related to price jumps in the It\^{o}-semimartingale setup. In a univariate context, unit root tests were designed in perron1989great, perron1990testing, banerjee1992recursive, perron1992nonstationarity and zivot1992further. gregory1996residual extend the tests in a cointegration framework. In all those papers, the number of shifts is at most equal to unity. Moreover, hatemi2008tests allows for two possible breaks. Finally, maki2012tests proposes tests which permits an arbitrary number of breaks in principle, although the method is seldom used with more than five breaks in practice. The strong limitation of the regime change approach is that jumps are required to happen at deterministic times, with deterministic size, known number of jumps, and not very often, all of which are not consistent with the aforementioned high-frequency stylized facts.

To obtain the robustness of the residual based test to time-varying variance, the most natural candidate consists in a residual based test based on the aforementioned time-varying variance robust unit root tests. Unfortunately, a transformation of data solely at the unit root test step entails hybrid-behaved tests to the extent that the limit distribution is unidentifiable, at least to us. To circumvent this difficulty, we deflate the returns prior to the two steps. As for that to price jumps, we simply consider a truncation method as in mancini2009non, which is commonly used in the literature on high frequency econometrics when dealing with jumps with random size and times.

Our theoretical contribution is divided into two parts. First, we provide a regression procedure based on deflated and truncated asset price returns in order to estimate the cointegrated system and establish the related consistency and central limit theory when there is cointegration. This is related to Theorem (ref) in Section (ref). Second, we develop a DF type residual-based cointegration test, also based on deflated and truncated observations, along with its central limit theory in the absence of cointegration and under local alternatives. We also provide with an equivalent of the expression in the presence of cointegration (see Theorem (ref) and Proposition (ref)). In particular we show that under the null hypothesis of no cointegration the limit distribution of the statistic is that of a classical DF test, and that no local power is lost due to the truncation and the deflation.

The finite sample from this paper corroborates that from cited papers and our theoretical findings. On the one hand, a typical environment of cointegrated assets observed at high frequency distorts badly size and power of the DF or Phillips-Perron type residual based unit root tests, and as expected we can isolate the effect of time-varying variance as impacting the former, and that of price jumps as altering the latter. On the other hand, the proposed test size is appropriate and its power is good in the simultaneous presence of both stylized facts. Two empirical examples illustrate the fact that the proposed test can be in disagreement with the original tests, indicating the practical relevance of asset prices truncation and deflation.

One big limitation of our technique is that we have no drift. Such a simple assumption is quite restrictive since the concept of cointegration is mainly used to analyze the long run comovements of multiple time series and the role of stochastic drift can be very important in the long run relationship. Correspondingly, we examine the case of linear drifting It\^{o}-semimartingales and show that, even though theoretical derivations are possible, the limit distribution of the statistic is quite complex, indicating that combining drift and deflation is a difficult task. We are well aware that linear drift is rather a strong assumption, but overall there is a lack of literature on drift. Some exceptions include perron2016residuals, which deals with optimal methods for unit-root and cointegration tests when a trend is present (but no time-varying volatility), and beare2018unit (Section 4.3), which gives a general detrending method for time-linear trends in the presence of time-varying variance for the unit-root test. Theorem 4.3 from the latter work gives the related theoretical asymptotic results. In particular, the author already noticed that the modified limit distribution under the null strongly depends on the shape of volatility and so has to be computed numerically every time the method is used. Our own method detailed in Section 4.2 uses a similar detrending device (adapted to our own high-frequency scheme). Our Theorem 4.5 essentially generalizes the result of beare2018unit to the case of cointegration and presents the same feature, i.e that the detrended data yields limit distribution under the null which explicitly depends on the volatility curve. To our knowledge, there is no method (even in unit root literature) which deals with time linear trends in the presence of time-varying volatility and which yields a limit distribution independent of the volatility process. In addition, finite sample sheds light on the fact that drift does not seem to affect both size and power of all the considered residual based tests, at least in a range of realistic configurations.

Another obvious limitation of our approach is that the framework considered rules out market microstructure noise. This is mitigated by the fact that the vast majority of empirical studies related to cointegration operating with financial time series does not sample with a frequency higher than five minutes so that if the asset is liquid enough the data are reasonably free of market microstructure noise (see, e.g., ait2019hausman). To temper as much as possible the effect of microstructure noise, we do not sample faster than every ten minutes in both our numerical and empirical studies.

The remaining of this paper is structured as follows. The framework which is a natural adaptation to high frequency data is given in Section 2. Estimation of cointegrated relations, based on regression on the truncated and deflated price, and its related limit theory under cointegration is provided in Section 3. A DF-type test of the null hypothesis that there is no cointegration against the alternative that there is cointegration, together with its limit theory, is developed in Section 4. In particular, this test's robustness to time-varying variance and price jumps, which is based respectively on deflation and truncation, is established. Section 5 is devoted to a Monte Carlo experiment, to shed light on the size and power twist of the original DF and Phillips-Perron based cointegration tests and the good behavior of the adapted DF test in a variety of realistic configurations. In Section 6, a brief empirical study, which corroborates the fact that the proposed test is not always in accordance with the original tests, is conducted. We conclude in Section 7. Proofs and part of the numerical results have been relegated to the Appendix for the sake of clarity.

The framework: a natural adaptation to high frequency data

In this section, we introduce our general framework, which in particular accommodates the definition of no cointegration (engle1987co) when two driftless It\^{o}-processes including jumps with infinite activity are observed synchronously and regularly. Our introduced framework also specifies the notion of cointegration and weak cointegration.

We introduce a few key concepts from the aforementioned paper and the existing literature on cointegration for time series that will help motivate the framework of our own study. For the sake of clarity, we restrict ourselves to the case of a pair of processes, but all the definitions can be extended to the multivariate case with no major difficulty. Let us consider two unit root time series $x_t$ and $y_t$ (i.e. with explosive variance and with stationary increment processes $\Delta x_t$ and $\Delta y_t$). As pointed out in engle1987co, in general, any linear combination $y_t - \alpha x_t$ will be again a unit root process drifting away from zero. In that case, $x_t$ and $y_t$ are said not to be cointegrated. However, it may also happen that for some $\alpha$, the time series $y_t- \alpha x_t$ does not wander far from zero, and not only its increments but also the series itself is stationary. The couple $(x_t,y_t)$ is then said to be cointegrated with cointegration vector $(1,-\alpha)^T$. When conducting tests, and for the sake of tractability, both notions are often naturally embedded (see e.g. model (4.7) of engle1987co) in an AR(1) model specification as follows. We assume that $x_t$ is a unit root process, and that \bea y_t = \alpha x_t + \epsilon_t, with \epsilon_t = \rho \epsilon_{t-1} + u_t, \eea where $u_t$ is a non-trivial stationary process possibly correlated with $\Delta x_t$. Then, $0 < \rho < 1$ yields a cointegrated system, whereas $\rho = 1 $ implies that $\epsilon_t$ is a unit root process, hence no cointegration. Moreover, the regime $0 < \rho < 1 $ with $\rho \to 1$, yields a process $\epsilon_t$ which is a nearly integrated process (as first introduced in nabeya1994local), and accordingly the system ((ref)) can be said weakly cointegrated. Residual based tests for the null of no cointegration usually pre-estimate $\alpha$ by, for instance, an ordinary least squares (OLS) regression, and then run a unit root test (i.e $\rho=1$ versus $\rho<1$) on the estimated residual $\widehat{\epsilon}_t = y_t - \widehat{\alpha} x_t$. As detailed in phillips1990asymptotic, it is of importance to mention that such a two-step procedure affects the limit distribution of the test statistic so that testing for cointegration does not amount to directly testing for a unit root in $\widehat{\epsilon}_t$. In particular and of practical relevance, the critical values are altered. This deviation from the unit root case stems from the inconsistency of $\widehat{\alpha}$ (hence $\widehat{\epsilon}_t$) under the null of no cointegration.

In view of the above discussion, our first goal consists in introducing a model similar to ((ref)) where now $\Delta x_t$ and $u_t$ are replaced by increments of driftless It\^{o}-semimartingales. We return to the case of drifting processes in a detailed discussion at the end of Section (ref). We assume that we observe regularly (i.e. at times $t_0 := 0, t_1 := \Delta, \cdots, t_n := n \Delta$ with $\Delta := T/n$) between 0 and the horizon time $T$ (depending on $n$) two \cadlag (right continuous with left limits) processes $X$ and $Y$.\footnote{A full specification of the model actually involves the stochastic basis $\calb=(\Omega,\proba,\calf,\F)$, where $\F=(\calf_t)_{t\geq0}$ is a filtration satisfying the usual conditions and $\calf = \vee_{t \geq0} \calf_t$. We assume that all the processes are $\F$-adapted. Also, when referring to It\^{o}-semimartingale, we automatically mean that the statement is relative to $\F$.} For any process $A$, we use the following conventions for $i\in\{0,\cdots,n\}$: \bea A_i & := & A_{t_i},\\ \Delta A_i & := & A_i - A_{i-1} for i \neq 0,\\ \Delta A_0 & := & 0. \eea In what follows, we assume that $X$ and $Y$ may be further decomposed as \bea X_t = X_t^c + J_t^X and Y_t = Y_t^c + J_t^Y, t\in[0,T] \eea where $X^c$ and $Y^c$ are the continuous parts of $X$ and $Y$, and $J^X$ and $J^Y$ are pure jump processes such that, for $U \in \{X,Y\}$, $t \in[0,T]$, $$ J_t^U = \int_{[0,t] \times E} \delta^U (s,z) \mu^U(ds, dz),$$ where $\mu^U$ is a Poisson random measure on $\reels_+ \times E$ for $E$ some auxiliary Polish space, $\nu^U$ is the compensator of $\mu^U$ of the form $\nu^U(ds,dz)=ds \otimes \lambda^U(dz) $ where $\lambda^U$ is a $\sigma$-finite measure, and where $\delta^U$ is a predictable function on $\Omega \times \reels_+ \times E$. Moreover, we assume that there exists $r \in [0,1)$ such that \bea \sup_{t \in \reels_+} \esp\int_E (|\delta^U(t,z)|^r \vee |\delta^U(t,z)|^8) \lambda^U(dz) < +\infty. \eea In particular, although the jump processes may feature an infinite number of jumps on bounded time intervals, ((ref)) ensures that the jumps are summable on $[0,T]$ and of order at most $T$, since it implies that \bea \esp \sum_{0 <s \leq T} |\Delta J_s^U| \leq \esp \sum_{0 <s \leq T} |\Delta J_s^U|^r \vee |\Delta J_s^U|^8 \leq KT \eea for some constant $K \geq 0$. The summability property of the jumps is used extensively in our proofs (see e.g Lemma (ref) in the Appendix) in order to control most deviations involving jumps. Unfortunately, results such as those in Lemma (ref) may not hold when the jump processes are not of finite variation. As far as we know, it is not clear whether relaxing ((ref)) to the case of non-summable jumps (for instance assuming $r \in [0,2)$ only) is possible in the present framework, which is why we leave it for future research. In particular, note that it is well-known that when the jumps are not summable, slower rates of convergence for even simple quantities such as threshold realized volatility should be expected (jacod2014remark), indicating that this case is non-standard.

We now assume that $X$ and $Y$ satisfy a relation of the same nature as ((ref)). Assuming first that $J^X = J^Y = 0$, we naturally adapt ((ref)) as follows. We assume that there exist $c_0, \alpha_0 \in \reels$ such that for any $i \in \{1,\cdots,n\}$, we have

\bea Y_i^c = c_0 + \alpha_0 X_i^c + \epsilon_i, with \epsilon_i = \rho \epsilon_{i-1} + \Delta Z_i, \epsilon_0 = Z_0 = 0 \eea where $\rho \in [0,1]$, and may depend on $n$, and where $X^c$ and $Z$ are two continuous It\^{o}-martingales of the form

eqnarray[eqnarray omitted — 175 chars of source]

where $\sigma^X$ and $\sigma^Z$ are \cadlag adapted processes, and $W^X$ and $W^Z$ are Brownian motions featuring possibly non-trivial high frequency correlation $d\langle W^X, W^Z \rangle_t = r_t dt$. Therefore, at time $t \in [0,T]$, and up to the multiplicative term $(\sigma_{t/T}^M)^2$, the two dimensional process $(X,Z)$ features a squared volatility equal to $$ \Sigma_t = \left(

matrix[matrix omitted — 101 chars of source]

\right).$$ Finally, $\sigma^M$, which is the common deterministic volatility component, is a \caglad (left-continuous with right limits) function from $[0,1]$ to $\reels_+ - \{0\}$. The component $\sigma^M$ can be interpreted as the market volatility (i.e. common to all the stocks), whereas $\sigma^X$ and $\sigma^Z$ correspond to the idiosyncratic component of volatility. For example, $\sigma^M$ can be a linear trend, while $\sigma^X$ and $\sigma^Z$ can both be a product of daily U-shape and random stochastic component such as Heston model. Further examples are considered in our finite sample analysis.

remark*The components $\sigma^X$ and $\sigma^Z$ may differ, will be assumed ergodic and typically account for the long time regularities (e.g. seasonality) of $X$ and $Z$. Having ergodic returns is in line with most of the literature on cointegration, see for instance phillips1990asymptotic, or the more recent work of perron2016residuals. On the other hand, $\sigma^M$ encompasses possibly non ergodic trends in volatility and is assumed to be a common factor in $X$ and $Y$. As far as we know, and as detailed in Section (ref), if the non ergodic components differ in $X$ and $Y$, then constructing a test statistic which is numerically reliable and whose distribution is identifiable under the null of no cointegration remains an open and difficult question that we set aside in this paper. Note that adding such a non ergodic component scaled in time from 0 to $T$ is common practice in the literature on tests for unit root processes, as in cavaliere2005unit, cavaliere2007testing and more recently beare2018unit among others. In our case, however, $\sigma^M$ is taken \cadlag whereas earlier works on unit root processes assumed the function to satisfy a Lipschitz condition except for, at most, a finite number of points of discontinuity. Note also that, as mentioned in the aforementioned papers, the fact that $\sigma^M$ is assumed deterministic can be easily relaxed to random and independent of the main filtration $\F$. Then all the convergences can be taken conditionally to $\sigma^M$. Finally, to our knowledge, there is no existing literature on a test of no cointegration which is robust to the presence of a common non ergodic volatility component. \\

Just as $\Delta X_i^c$ is the continuous time counterpart of $\Delta x_t$, $\Delta Z_i$ now plays the role of $u_t$ in ((ref)). Note also that the presence of an intercept $c_0$ in the regression is just a convenient way to center the residual process $\epsilon$ without loss of generality. Moreover, as $\rho$ controls how close the residual is from a unit root process in ((ref)), $\rho$ now controls how close $\epsilon$ is from an It\^{o}-martingale in ((ref)), with the two extreme cases being $\rho = 1$ where $\epsilon_i = Z_i$, and $\rho = 0$ where $\epsilon_i = \Delta Z_i$ for $i \in \{1,\ldots,n\}$. When $0 < \rho < 1$, $\epsilon_i$ lies somewhere between an It\^{o}-martingale and an increment of an It\^{o}-martingale.

remark*(limitation in terms of economic/statistical modeling) Note that as the model was designed with the (necessary and non straightforward) development of asymptotic theories in mind, there is (at least) an important side-effect feature in terms of economic/statistical modeling. If $0 \leq \rho < 1$, $\epsilon_i$ and $Y^c_i$ are $\Delta$ dependent. Highly related literature working under the cointegration error, $\epsilon_i$, being independent of $\Delta$ includes bandi2014nonparametric, kanaya2011nonparametric and kim2018unit.

In general, when $J^X \neq 0$ or $J^Y \neq 0$, it would be natural to simply replace $X^c$ and $Y^c$ by $X$ and $Y$ in ((ref)), but it turns out that imposing such a constraint on the jump processes would imply that $\Delta J^Y = \alpha_0 \Delta J^X$. If the jumps are seen as structural breaks in the processes $X$ and $Y$, it is a very strong assumption which would require substantial support from empirical data. More importantly, letting $J^X$ and $J^Y$ free of constraint does not affect our strategy (and the related limit theory) for analyzing $X$ and $Y$, which will consist in first getting rid of the jump components using the truncation approach of mancini2009non and then directly work with the estimated continuous components. Accordingly, for the sake of generality, we keep ((ref)) even in the presence of jumps. The cointegration relation between $X$ and $Y$ yields for $i\in \{1,\ldots,n\}$ that \bea Y_i = c_0 + J_i^Y - \alpha_0J^X_i + \alpha_0 X_i + \epsilon_i, \eea which can be seen as cointegration (if $\rho <1$) with multiple level shifts. Cointegration with breaks has been studied in gregory1996residual (see Model 2) in the case of a single shift, and extended to the case of an arbitrary large (but known) number of deterministic breaks in maki2012tests. In Equation ((ref)), the shifts are $\Delta J_s^Y - \alpha_0 \Delta J_s^X$ for $s \in [0,T]$, which may be in infinite number, are of random sizes, and can feature endogeneity.

We now adapt the notion of no cointegration introduced in engle1987co and discussed above to the case of driftless It\^{o}-semimartingales.

definition*(no cointegration) Two \cadlag processes $X$ and $Y$ are said not cointegrated if any linear combination $Y - \alpha X$, $\alpha \in \reels$, is a driftless It\^{o}-semimartingale whose volatility component $\sigma$ is such that $\proba-\liminf_{T \to +\infty} T^{-1}\int_0^T \sigma_s^2ds > 0$.

Definition (ref) is thus a straightforward adaptation where time series have been replaced by \cadlag processes and unit root processes have been replaced by driftless It\^{o}-semimartingales with non-trivial volatility components so that they are indeed explosive as $T \to +\infty$. Let us now get back to the model ((ref)) and turn our attention to the description of different settings of $\rho$ and their impact on the relationship between $X$ and $Y$. Following the time series framework, $Y$ and $X$ are not cointegrated if $\rho = 1$ and, of course, if for any $x \in \reels^2 - \{0\}$, $\proba-\liminf_{T \to +\infty} T^{-1}x^T\int_0^T \Sigma_s^2dsx > 0$, since in that case, ((ref)) reads for any $i \in \{1,\ldots,n\}$ \beas Y_i^c = c_0 + \alpha_0 X_i^c + Z_i. \eeas We now turn our attention to the notion of cointegration. We follow again the time series case and say that $X$ and $Y$ satisfying ((ref)) are cointegrated if $\rho \in [0,1)$ and does not depend on $n$. In that case, note that this implies the existence of $c_0, \alpha_0 \in \reels$ such that

\beas Y_i^c = c_0 + \alpha_0 X_i^c + \epsilon_i, \eeas where $\epsilon_i$ is not the value of an It\^{o}-martingale (and is of one order of magnitude smaller). Finally, we will also consider the intermediary situation where $X$ and $Y$ are weakly cointegrated, which corresponds to the case where $\rho = 1 - \beta/n$ for some $\beta > 0$. Henceforth we will accordingly always assume that $X$ and $Y$ are generated according to one of the following setting:

enumerate[{(}i{)}] • Cointegration $0 \leq \rho < 1 $ (independent of $n$). • Weak cointegration $\rho = 1- \beta/n$, with $\beta > 0$. • No cointegration $\rho = 1$.

We end this section with a brief remark on the connection between the above definition of cointegration and continuous-time mean reverting residuals. It sounds indeed reasonable to expect that if $X^c$ and $Y^c$ are such that for any $t \in [0,T]$ $$ Y_t^c = c_0 + \alpha_0 X_t^c + \cale_t $$ where $\cale$ is mean-reverting around $0$, i.e if for instance $$ d\cale_t = - \theta \cale_t dt + \sigma^\cale(t)dW_t$$ for some $\theta >0$ and where $W$ is a standard Brownian motion, then $\cale$ remains of order $O_\proba(1)$ for any $t \in [0,T]$ and $Y^c$ and $X^c$ should be cointegrated to a certain degree. The next remark shows that this mean-reverting residual framework precisely corresponds asymptotically to the case of weak cointegration, where $X^c$ and $Y^c$ satisfy ((ref)) for $\rho = 1-\beta/n$ with $\beta = \theta T$, and where the observation frequency $n \to +\infty$.

remark*Assume that $X$ and $Y$ satisfy ((ref)) with cointegration parameter $\rho = 1 - \theta T/n$ with $\theta >0$. Then, the interpolating residual \beas \epsilon_t = \epsilon_{i} for t \in [t_i, t_{i+1}) \eeas is such that when $n \to +\infty$ $$ \epsilon \to^{u.c.p} \cale = \left(\int_0^{t} e^{-\theta (t-s)} dZ_s\right)_{t \in [0,T]} $$ where $\to^{u.c.p}$ stands for the uniform convergence in probability on any compact. Therefore, $\cale$ enjoys the mean-reverting dynamics $$ d\cale_t = -\theta \cale_tdt + \sigma^Z(t) \sigma^M(t/T) dW_t^Z, t \in [0,T].$$ In particular, when $\sigma^Z(t) = \sigma^Z$ and $\sigma^M=1$, $\cale$ is an Ornstein-Uhlenbeck process.
proofWe have the representation $\epsilon_t = \sum_{j=1}^n \rho^{\Delta^{-1}(t_i-t_j)} (Z_{t_i \wedge t_j} - Z_{t_i \wedge t_{j-1}} )$ for $t \in [t_i, t_{i+1})$. Therefore, the convergence towards $\cale$ is a direct consequence of the convergence $\rho^{ \Delta^{-1} s} = (1-\theta T/n)^{ns/T} \to e^{-\theta s}$ for any $s \in [0,T]$ along with Theorem I.4.31(iii) from jacod2013limit (p. 47).

Estimation of cointegrated systems

Construction of the estimator

We now focus on estimating the couple $(\alpha_0, c_0)$ based on the discrete observations of $X$ and $Y$. Naturally, one can expect $(\alpha_0, c_0)$ to be identifiable only when $X$ and $Y$ are cointegrated. Indeed, we will see that when $\rho = 1$ (i.e. non cointegration case), our proposed estimator is inconsistent. Accordingly, any estimation of the cointegration parameters is untrustworthy without performing a test of no cointegration, question that we set aside in this section, and that is treated in Section (ref). In any case, constructing the estimator does not require any knowledge whatsoever about the cointegration level $\rho \in [0,1]$.

We first consider the case where $J^Y = J^X = 0$ and $\sigma^M = 1$. We adapt the classical OLS estimator proposed in engle1987co, and resulting from ((ref)) seen as a linear regression where $\epsilon$ is the noise process. Recall that $X$ and $Z$ may be correlated, so that the regression model induced by ((ref)) features endogeneity. In particular, this rules out the alternative regression based on the high-frequency returns of $X$ and $Y$ \bea \Delta Y_i^c = c_0 + \alpha_0 \Delta X_i^c + \Delta \epsilon_i, with \Delta \epsilon_i = (\rho-1) \epsilon_{i-1} + \Delta Z_i, \eea because $\Delta Y_i^c$, $\Delta X_i^c$ and $\Delta \epsilon_i$ being of the same order $\sqrt{\Delta}$, the OLS estimator based on ((ref)) would be inconsistent due to non-zero correlation between $\Delta \epsilon_i$ and $\Delta X_i^c$. Conversely, the regression based on ((ref)) is robust to endogeneity because $X_i$ and $Y_i$ are of order $\sqrt{T}$ whereas when $\rho < 1$, $\epsilon_i$ remains of order $\sqrt{\Delta}$.

In general, $X$ and $Y$ contain jumps and a non-constant common volatility component. Accordingly, we now estimate $(\alpha_0, c_0)$ in a three-step procedure consisting in first getting rid of those two features and then applying the aforementioned OLS estimation. At this point, in view of the representation ((ref)), it seems natural to adapt the methodology of gregory1996residual and derive a modified OLS estimator which estimates and cancels the effect of the jumps seen as level shifts in ((ref)). However, the time required to run the break-robust OLS estimation greatly increases with the number of breaks. maki2012tests proposed an alternative and less time-consuming method, but it also presents several drawbacks. First, the number of breaks (or at least an upper bound $k$) must be known. Second, the limit distribution of the test statistic depends on $k$, meaning that it may be necessary to calculate critical values if $k$ is larger than five, which is the highest value for which they have been reported. Finally, the method is supported by a numerical study only (again, for models with five breaks at most). Since we allow for a potentially high number of jumps in ((ref)), both approaches are inadapted, which is why we henceforth adopt the truncation method of mancini2009non. As explained in the asymptotic theory section, it is independent of the number of shifts and robust to all the aforementioned features of the jumps, both for estimation and test. Accordingly, we remove the increments of $X$ and $Y$ such that at least one of them is greater than a given threshold in absolute value. More precisely, for $U \in \{X,Y\}$, we compute the truncated process $\calt(U) = (\calt(U)_0,\ldots, \calt(U)_n)$ such that $\calt(U)_0 = U_0$, and for $i \in \{1,\ldots,n\}$ \beas \Delta \calt(U)_i = \Delta U_i \mathbb{1}_{\{ |\Delta X_i| \leq a \Delta^{\overline{\omega}}\} \cap \{ |\Delta Y_i| \leq a \Delta^{\overline{\omega}}\}}, \eeas that is \beas \calt(U)_i = U_0 + \sum_{j=1}^i \Delta U_j \mathbb{1}_{\{ |\Delta X_j| \leq a \Delta^{\overline{\omega}}\} \cap \{ |\Delta Y_j| \leq a \Delta^{\overline{\omega}}\}}, \eeas for some constant $a > 0$ and some exponent $\overline{\omega} \in (0,1/2)$ satisfying additional constraints stated in the next section.

remark*The practitioner should keep in mind that the choice of the threshold tuning parameters $(a,\overline{\omega})$ is very important in practice. On the one hand, if the threshold is too loose, then some small jumps will not be properly removed. On the other hand, when the truncation is tight, some returns will be mistakenly truncated away. From a theoretical perspective, both constraints are crucial, and the power of the test procedure will be badly affected while the cointegration estimator may become inconsistent if they are not satisfied. From a numerical perspective, however, it seems that not being able to remove jumps distorts the test procedure far more than removing too many returns. Details about the constraints that the truncation exponent $\overline{\omega}$ must satisfy can be found in Assumption $\textnormal{\textbf{[C]}}$.

Second, we deflate the returns of both truncated processes $\calt(X)$ and $\calt(Y)$ by a consistent estimator $\sqrt{C_i}$ of $\sigma_{t_i-}^M$ up to some multiplicative constant. The procedure is similar to what was proposed in beare2018unit for the case of a unit root process. Hereafter, we take $C_{i}$ as the standard local realized volatility on the truncated returns of $X$ (Using the returns of $Y$ would yield an estimator of $\sigma_{t_i-}^M$ up to a coefficient which depends on $\rho$). First, for two indices $0 \leq l < k < i$, $i\in \{1,\ldots,n\}$, we define \bea RV_{i,k,l} = \sum_{j=(i-k) \vee 1}^{(i-l-1) \vee 1} \Delta X_j^2 \mathbb{1}_{\{ |\Delta X_j| \leq a \Delta^{\overline{\omega}} \}}, \eea where $a$ and $\overline{\omega}$ were defined before. Next, we take $k = [T^\gamma \Delta^{-1}]$, $l = [T^{\gamma'} \Delta^{-1}]$ (which both implicitly depend on $n$), where $[x]$ is the floor of $x$, and $0 < \gamma' < \gamma < 1$. The local window considered for ((ref)) is thus such that the number of observations $k-l \to+\infty$ and at the same time the length of the window $(k-l)\Delta = o(T)$. Moreover, realized volatility is calculated over the interval $[t_{i-k},t_{i-l-1}]$ and not $[t_{i-k},t_{i-1}]$ in order to preserve the martingale structure of some transformations of $\epsilon$ when $\rho < 1$ and thus circumvent some technical difficulties that arise in the proofs. We then define for $i \in \{1,\ldots,n\}$ \beas C_{i} &=& T^{-\gamma}RV_{i,k,l} if RV_{i,k,l} > 0 and i > 2k , \\ C_{i} &=&+\infty otherwise. \eeas

Given that $l$ is negligible with respect to $k$, it does not affect the limit theory of $C_i$. We then compute the deflated version of $\calt(U) \in \{\calt(X),\calt(Y)\}$, $\calt( U)^{def}$ such that $\calt (U)_0^{def} = U_0$, and for $i \in \{1,\ldots,n\}$, $$ \Delta \calt (U)_i^{def} = \frac{\Delta \calt (U)_i}{\sqrt{C_{i}}},$$ that is \beas \calt (U)_i^{def} = U_0 + \sum_{j=1}^i \frac{\Delta U_j}{\sqrt{C_{j}}} \mathbb{1}_{\{ |\Delta X_j| \leq a \Delta^{\overline{\omega}}\} \cap \{ |\Delta Y_j| \leq a \Delta^{\overline{\omega}}\}}. \eeas Key to our analysis is that both operations $\calt$ and '$def$' naturally preserve the cointegration relationship, in the sense that for any $i \in \{1,\ldots,n\}$ \bea \calt(Y)_i^{def} = c_0 + \alpha_0 \calt(X)_i^{def} + \calt(\epsilon)_i^{def}, \eea with the new residual \beas \calt(\epsilon)_i^{def} = \sum_{j=1}^i \frac{\Delta \epsilon_{j}}{\sqrt{C_j}} \mathbb{1}_{\{ |\Delta X_j| \leq a \Delta^{\overline{\omega}}\} \cap \{ |\Delta Y_j| \leq a \Delta^{\overline{\omega}}\}}. \eeas Finally, for two processes $A,B$, their associated OLS estimator is defined as $$ \textnormal{OLS}[A,B] = \left( \frac{\sum_{i=1}^n (B_i - \overline{B})(A_i - \overline{A})}{\sum_{i=1}^n (A_i - \overline{A})^2},\overline{B} -\frac{\sum_{i=1}^n (B_i - \overline{B})(A_i - \overline{A})}{\sum_{i=1}^n (A_i - \overline{A})^2}\overline{A}\right)$$ with $\overline{A} = n^{-1}\sum_{i=1}^n A_i$ and $\overline{B} = n^{-1}\sum_{i=1}^n B_i$. The general cointegration estimator is defined as

eqnarray[eqnarray omitted — 129 chars of source]
remark*When $J^X = J^Y = 0$ and in the absence of truncation, $(\widehat{\alpha},\widehat{c})$ is the OLS estimator of a the linear transformation of $X$ and $Y$ where their respective returns have been multiplied by the weights $w_i = C_i^{-1/2}$. This is similar (yet not equal) to the GLS of kim2010cointegrating, where the authors pre-estimate the time-varying variance of the noise $\epsilon$ by a standard OLS, and then construct the associated GLS cointegration estimator consisting in putting similar weights directly in front of $X_i$ and $Y_i$.

Assumptions and high-frequency framework

We now proceed to give an asymptotic framework along with reasonable conditions under which the OLS estimator introduced in ((ref)) is consistent (assuming, of course, $\rho <1$, i.e. cointegration). We also give a stronger setting on the jump processes which ensures a central limit theory for $(\widehat{\alpha},\widehat{c})$. We will use the same framework when testing for no cointegration in the next section. Our first assumption specifies the high frequency asymptotics that is considered in this paper.\\

\noindentAssumption [A]: $n \to + \infty$, $T \to +\infty$, $\Delta = T/n \to 0$. \\

Such a double asymptotic is consistent with the high-frequency context ($\Delta \to 0$) and the fact that cointegration is a long-run phenomenon ($T \to +\infty$). Next, we assume that the volatility matrix $\Sigma$ has bounded moments up to some order $p_0$ and is ergodic.\\

\noindentAssumption [B]: There exists $p_0 \geq 8$ such that $\sup_{t \in \reels_+} \esp \|\Sigma_t \|^{p_0} < +\infty$, where for a matrix $M$, $\|M\| := \sum_{i,j} |m_{i,j}|$. Moreover, there exists a positive definite matrix $\Omega = (\omega_{ij})_{1\leq i,j \leq 2}$ such that \bea \epsilon(T) := \sup_{u\in [0,1]} \esp\left| \frac{1}{T} \int_{uT}^{T} \Sigma_tdt - (1-u)\Omega \right|^2 \to 0, T \to +\infty. \eea Moreover, $\sigma^X$ is asymptotically bounded from below with probability $1$, i.e. there exists $\underline{\sigma}^X >0$ such that $\proba - \liminf_{t \to + \infty} \sigma_t^X \geq \underline{\sigma}^X$.

remark*The definition of ergodicity stated in $\textnormal{\textbf{[B]}}$ is quite flexible. For instance, it encompasses most combinations of stationary ergodic processes and periodic processes. The asymptotic boundedness of $\sigma^X$ away from $0$ is assumed to avoid degenerate behaviors of the statistics due to the deflation operation. Similar long-run high-frequency asymptotics and ergodic settings can be found in the recent literature, see e.g. christensen2018diurnal and andersen2019time where the volatility process is the product of a stationary mixing component and a periodic component. Finally, note that $\Sigma_t$ may be correlated with $(W^X,W^Z)$ so that the Brownian integrals feature leverage effect.

Now we turn to our third assumption, which states conditions on the truncation parameters, and an additional condition on the relationship between $T$ and $n$.\\ \noindentAssumption [C]: We have $\frac{1}{4-r} < \overline{\omega} < \frac{1}{2} - \frac{3}{2p_0}$. Moreover, let $e_1 = \frac{1}{2\overline{\omega}(1-r)} \vee \frac{1}{\overline{\omega}(4-r)-1}$ if $r > 0$ and $e_1 > \frac{1}{4\overline{\omega}-1}$ if $r = 0$, and $e_2 = \frac{4}{p_0(1/2 -\overline{\omega})} \vee \frac{1}{p_0(1/2-\overline{\omega}) + 2\overline{\omega} - 3/2}$. We have $T^{1+e_1\vee e_2}n^{-1} \to 0$.

remark*In particular, the second condition in $\textnormal{\textbf{[C]}}$ implies that $T$ must tend to infinity slowly enough compared to $n$. In the case where $\Sigma_t$ admits moments of any order ($p_0 = +\infty$), we note that Condition $\textnormal{\textbf{[C]}}$ can be simplified as $1/(4-r) < \overline{\omega} < 1/2$, and taking $\overline{\omega}$ arbitrarily close to $1/2$, the condition on $T$ becomes $T^{1+ \eta + \frac{1}{1-r}}n^{-1} \to 0$ for $\eta >0$ arbitrarily small. Finally, note that for jumps of finite activity $r=0$, this can be further simplified as $T^{2+ \eta} n^{-1} \to 0$ which is stronger than the condition $\Delta \to 0$ stated in Assumption $\textnormal{\textbf{[A]}}$. Of course, if $J^X = J^Y = 0$ and the truncation step is ignored in the estimators, all the stated results hold with $e_1 = 0$ in $\textnormal{\textbf{[C]}}$. If moreover $p_0 = +\infty$ then $e_2 = 0$ too and $\textnormal{\textbf{[C]}}$ reduces to $\Delta \to 0$.

Finally, we state an additional more restrictive assumption on the jumps, that we will use only to derive a central limit theorem under cointegration for $(\widehat{\alpha}, \widehat{c})$ (Theorem (ref)) and an equivalent of the Dickey-Fuller statistics under the alternative of cointegration (Proposition (ref)). It plays no role in the derivation of the consistency of the OLS and of the Dickey-Fuller test under any of the hypotheses. \\ \noindentAssumption [D]: $J^X$ and $J^Y$ are two sequences of jump processes such that for $U \in \{X,Y\}$, $\sup_{j \in \{1,\ldots,n\}, u \in [0,1]} |J_{t_j \wedge uT_n}^U-J_{t_{j-1} \wedge uT_n}^U| = O_\proba(\Delta^{1/2})$, and $T^{-1} \sum_{0<s \leq T} |\Delta J_s^Y - \alpha_0 \Delta J_s^X| = o_\proba(T^{-1/2}n^{-1})$. \\ The above assumption states that the jump processes are asymptotically negligible in the regression, and satisfy an additional cointegration condition. Hereafter, we will always implicitly assume that [A]-[C] hold. As for [D], we will explicitly state whether it is assumed or not.

Asymptotic theory of the OLS estimator under cointegration

We now give the asymptotic properties of $C_i$, and of the OLS estimator under the cointegration regime $\rho < 1$. We start with $C_i$. Since $0 < \gamma < 1$, note that the local time window $T^{\gamma} \to +\infty$ so that the ergodic theory for $\Sigma$ kicks in, and at the same time the scaled time window $T^{\gamma-1} \to 0$ so that for $t \in [t_{i-k_n}, t_{i-l_n}]$, $\sigma_{t/T}^M \approx \sigma_{(t_i/T)-}^M$ by left continuity of $\sigma^M$.

proposition*For any $i \in \{1,\ldots,n\}$, when $n \to +\infty$, $$\esp |C_i - (\sigma_{(t_{i}/T)-}^M)^2 \omega_{11}|^2 \to 0.$$

As proved in the Appendix (see Lemma (ref)), it turns out that the above $\mathbb{L}^2$ convergence is even uniform outside a set of indices whose cardinality is negligible with respect to $n$. A full uniformity as in Lemma 4.1 of beare2018unit is impossible here because $\sigma^M$ may have jumps (whereas in the aforementioned paper $\sigma^M$ is assumed differentiable). We now focus on the asymptotic properties of $(\widehat{\alpha},\widehat{c})$ when $X$ and $Y$ are cointegrated. In what follows, we let $W = (W^1,W^2)$ be a standard Brownian motion on $[0,1]$. Moreover, defining \bea L = \l(

matrix[matrix omitted — 122 chars of source]

\r) = \l(

matrix[matrix omitted — 107 chars of source]

\r), \eea with $r_\infty = \omega_{12}/\sqrt{\omega_{11} \omega_{22}}$, we also let $B = (B^1,B^2) = L W$ be a two dimensional Brownian motion on $[0,1]$, with covariance matrix $L L^T = \Omega$. Below, for a process $V$ on $[0,1]$, we let $\overline{V} = \int_0^1 V_udu$.

theorem*Assume that $X$ and $Y$ are cointegrated, that is $\rho < 1$ and is fixed. Assume further that $1/2 \leq \gamma < 1$ and $0< \gamma^{'} < \gamma$. Then we have $$ \widehat{\alpha} - \alpha_0\to^\proba0 \textnormal{ and } T^{-1/2}(\widehat{c} -c_0) \to^\proba 0.$$ Moreover, under the additional assumption $\textnormal{\textbf{[D]}}$, we have the convergence in distribution \beas \l(\begin{matrix} n(\widehat{\alpha} - \alpha_0) \\ nT^{-1/2}(\widehat{c} - c_0) \end{matrix}\r) \to^d \frac{1}{1-\rho}\l(\begin{matrix} \frac{\omega_{12} +\int_0^1(B_s^1-\overline{B}^1) dB_s^2}{\int_0^1 (B_s^1 - \overline{B}^1)^2ds} \\ \omega_{11}^{-1/2}B_1^2 - \frac{\omega_{12} +\int_0^1(B_s^1-\overline{B}^1) dB_s^2}{\int_0^1 (B_s^1 - \overline{B}^1)^2ds}\omega_{11}^{-1/2}\overline{B}^1 \end{matrix}\r). \eeas

The fast rate $n$ for the estimation of $\alpha_0$ is in line with the literature on cointegration estimation, see for instance Proposition 1 of engle1987co and Theorem 7 in kim2010cointegrating. Similarly, the fact that $X_T^c$ and $Y_T^c$ are of order $T^{1/2}$ implies that one can consistently estimate only $T^{-1/2}c_0$, with the same rate $n$ as for $\alpha_0$. Note also that in the exogenous residual case $\omega_{12} = 0$, the limit distribution for $\widehat{\alpha}$ corresponds to the one of a classical OLS on an homoskedastic cointegration regression as in Lemma 2.1 of phillips1988asymptotic up to the mean terms $\overline{W}^1$, due to the presence of the constant $c_0$ in our regression. In other words, the jumps and the heteroskedasticity coming from $\sigma^M$ do not impact the limit theory of $\widehat{\alpha}$ and $\widehat{c}$, and no efficiency is lost due to the truncation and the deflation. Finally, note that, as explained in the following remark, the above limit distribution is actually mixed normal.

remark*Rewriting $B$ as $LW$ and conditioning on $W^1$, we can specify the above limit as the following mixed normal distribution. Defining $I[W^1] = \int_0^1 (W_s^1 - \overline{W}^1)dW_s^1 = (W_1^1)^2/2-1/2-(\overline{W}^1)W_1^1$, $J[W^1] = \int_0^1 (W_s^1-\overline{W}^1)^2ds$, $K[W^1] = \int_0^1 (W_s^1)^2ds$, $r_\infty=\omega_{12}/\sqrt{\omega_{11}\omega_{22}}$, and $v_\epsilon := \frac{\omega_{22}}{\omega_{11}(1-\rho)^2}$, we have \beas \l(\begin{matrix} n(\widehat{\alpha} - \alpha_0) \\ nT^{-1/2}(\widehat{c} - c_0) \end{matrix}\r) \to^d \sqrt{v_\epsilon} \calm\caln \l( r_\infty \mathcal{B}[W^1], \frac{ (1 - r_\infty^2)}{ J[W^1]} \mathcal{V}[W^1]\r) \eeas where $$ \mathcal{B}[W^1] = \l(\begin{matrix} \frac{1+I[W^1]}{ J[W^1]} \\ W_1^1 - \frac{1+ I[W^1]}{J[W^1]} \overline{W}^1 \end{matrix}\r) \textnormal{ and } \mathcal{V}[W^1] = \l(\begin{matrix} 1 & \overline{W}^1 \\ \overline{W}^1 & K[W^1] \end{matrix}\r).$$ Moreover, when $B^1$ and $B^2$ are independent, $r_\infty=0$, and the above limit becomes \beas \l(\begin{matrix} n(\widehat{\alpha} - \alpha_0) \\ nT^{-1/2}(\widehat{c} - c_0) \end{matrix}\r)\to^d \sqrt{\frac{v_\epsilon}{ J[W^1]}}\calm\caln \l(0, \l(\begin{matrix} 1 & \overline{W}^1 \\ \overline{W}^1 & K[W^1] \end{matrix}\r)\r). \eeas

Remark (ref) suggests that it is possible to construct a studentized version of Theorem (ref), essential for the computation of confidence intervals and significance tests. To do so, we need to estimate the different quantities appearing in the bias and the variance of the mixed normal distribution. We construct the estimated residuals \bea \widehat{\epsilon}_i = \calt (Y)_i^{def} - \widehat{c} - \widehat{\alpha} \calt (X)_i^{def}, \eea and then estimate $v_\epsilon$, $\rho$, and $r_\infty$ as follows. \bea \widehat{v}_\epsilon = T^{-1} \sum_{i=1}^n \widehat{\epsilon}_i^2, \eea

\bea \widehat{\rho} = \frac{\sum_{i=2}^{n} \Delta \widehat{\epsilon}_i \widehat{\epsilon}_{i-2}}{\sum_{i=1}^{n} \Delta \widehat{\epsilon}_i \widehat{\epsilon}_{i-1}}, \eea

\bea \widehat{r}_\infty = \frac{\sum_{i=2}^n \Delta \calt(X)_i^{def} (\widehat{\epsilon}_i - \widehat{\rho} \widehat{\epsilon}_{i-1})}{\sqrt{\sum_{i=2}^n (\widehat{\epsilon}_i - \widehat{\rho} \widehat{\epsilon}_{i-1})^2}}. \eea Note that for the estimation of $\rho$, we have preferred the formula of ((ref)) over the more classical estimator $\tilde{\rho} = \sum_{i=1}^n\widehat{\epsilon}_i \widehat{\epsilon}_{i-1}/\sum_{i=1}^n \widehat{\epsilon}_i^2$, because the former is robust to truncation under $\textbf{[A]-[C]}$ whereas it is not theoretically clear whether the latter is consistent without Assumption [D]. We also substitute $\calt(X)^{def}$ to all the quantities involving $W^1$ appearing in the central limit theorem. That is, letting \beas \overline{\calt(X)}^{def} &=& T^{-1/2}n^{-1} \sum_{i=1}^n \calt(X)_i^{def}\\ I[\calt(X)^{def}]&=& \frac{(\calt(X)_n^{def})^2 - T}{2T} - T^{-1/2}\overline{\calt(X)}^{def} \calt(X)_n^{def}\\ J[\calt(X)^{def}]&=& n^{-1} \sum_{i=1}^n (T^{-1/2}\calt(X)_i^{def} - \overline{\calt(X)}^{def})^2\\ K[\calt(X)^{def}] &=& J[\calt(X)^{def}] + (\overline{\calt(X)}^{def} )^2, \eeas we introduce the estimators for the asymptotic biases and variances of $\widehat{\alpha}$ and $\widehat{c}$: \beas B_{\widehat{\alpha}} = n^{-1} \sqrt{\widehat{v}_\epsilon} \widehat{r}_\infty\frac{1+I[\calt(X)^{def}]}{J[\calt(X)^{def}]} and B_{\widehat{c}} = n^{-1}T^{1/2}\sqrt{\widehat{v}_\epsilon} \widehat{r}_\infty \l(T^{-1/2}\calt(X)_n^{def} - \frac{1+I[\calt(X)^{def}]}{J[\calt(X)^{def}]} \overline{\calt(X)}^{def}\r), \eeas and \beas V_{\widehat{\alpha}} = \widehat{v}_\epsilon\frac{1-\widehat{r}_\infty^2}{J[\calt(X)^{def}]} and V_{\widehat{c}} = \widehat{v}_\epsilon\frac{(1-\widehat{r}_\infty^2) K[\calt(X)^{def}]}{J[\calt(X)^{def}]}. \eeas we have the following studentized version of the central limit theory for $\widehat{\alpha}$ and $\widehat{c}$.

proposition*Assume that $X$ and $Y$ are cointegrated, that is $\rho < 1$ and is fixed. Assume that $1/2 \leq \gamma < 1$ and $0< \gamma^{'} < \gamma$. We have $$ \widehat{\rho} \to^\proba \rho.$$ Moreover, if we assume $\textnormal{\textbf{[D]}}$, then we also have $$ (\widehat{v}_\epsilon, \ \widehat{r}_\infty) \to^\proba (v_\epsilon, r_\infty),$$ \beas \frac{n(\widehat{\alpha} - \alpha_0 - B_{\widehat{\alpha}})}{\sqrt{V_{\widehat{\alpha}}}} \to^d \caln(0,1), \eeas and \beas \frac{nT^{-1/2}(\widehat{c} - c_0 - B_{\widehat{c}})}{\sqrt{V_{\widehat{c}}}} \to^d \caln(0,1). \eeas

Residual based test for cointegration

Construction of the test and limit theory

It is perhaps of even more crucial importance to infer from the data whether $X$ and $Y$ are cointegrated or not in the first place prior to analyzing the estimated cointegration coefficients. Consequently, we now give a test for the null hypothesis that there is no cointegration against the alternative of cointegration. We let

eqnarray*[eqnarray* omitted — 120 chars of source]

Recall that both hypotheses induce respectively the following models on the continuous parts of $X$ and $Y$:

eqnarray*[eqnarray* omitted — 236 chars of source]

and that the parameter $\rho$ controls how far $\calh_1$ is from $\calh_0$. As it is standard in the literature on tests for unit roots and cointegration (see e.g. pesavento2004analytical, beare2018unit) and useful to derive the local power of our test, we embed $\calh_0$ in the family of local alternatives $\widetilde{\calh}_1^{n,\beta}$ defined as \bea \widetilde{\calh}_1^{n, \beta} & : & \rho = 1 - \frac{\beta}{n}, with \beta \geq 0, \eea which implies the following model on the continuous parts of $X$ and $Y$, \beas \widetilde{\calh}_1^{n, \beta} & : & Y_i^c = c_0+\alpha_0 X_i^c + \epsilon_i with \epsilon_i = \rho \epsilon_{i-1} + \Delta Z_i, \rho = 1 - \frac{\beta}{n}, and \beta \geq 0, \eeas that corresponds to the notion of weak cointegration introduced at the end of Section (ref) when $\beta >0$, and is simply $\calh_0$ when $\beta = 0$. The canonical test in the unit root literature is the so-called Dickey-Fuller test on residuals of dickey1981likelihood. It has been extended to many directions, such as, among others, the augmented Dickey-Fuller (ADF) test, robust to a residual following an AR(p) specification, and the $Z_t$ and $Z_\alpha$ tests of phillips1987time, robust to autocorrelated returns under the null hypothesis of a unit root process. These tests have been later adapted to cointegration, and their asymptotic properties derived in phillips1990asymptotic. In the present work, we choose to focus on the DF approach, performed on the estimated residuals resulting from the OLS estimation of ((ref)). Before we state the main result of this section, we briefly recall the construction of the test statistic. Recall that the estimated residuals are defined for $i \in \{1,\ldots,n\}$ as \beas \widehat{\epsilon}_i = \calt (Y)_i^{def} - \widehat{c}- \widehat{\alpha} \calt (X)_i^{def}. \eeas

Then, the associated DF statistic $\Psi$ is the $t$-statistic of the coefficient $\phi$ in the linear regression $$ \Delta \widehat{\epsilon}_i = \phi \widehat{\epsilon}_{i-1} + \eta_i, \textnormal{ } i \in \{1,\ldots,n\},$$ that is, first estimating $\phi$ with $$ \widehat{\phi} = \frac{\sum_{i=1}^n\Delta \widehat{\epsilon}_i \widehat{\epsilon}_{i-1}}{\sum_{i=1}^n \widehat{\epsilon}_i^2},$$ and estimating the standard deviation of $\widehat{\phi}$ with $$s_{\widehat{\phi}} = \sqrt{\frac{n^{-1} \sum_{i=1}^n (\Delta \widehat{\epsilon}_i - \widehat{\phi} \widehat{\epsilon}_{i-1})^2}{\sum_{i=1}^n \widehat{\epsilon}_{i-1}^2}},$$ $\Psi$ is defined as \bea \Psi = \frac{\widehat{\phi}}{s_{\widehat{\phi}}}. \eea We now proceed to derive the asymptotic distribution of $\Psi$ under $\widetilde{\calh}_1^{n,\beta}$, for any $\beta \geq 0$. In particular, recall that $\calh_0$ is covered by Theorem (ref) below, since $\calh_0 = \widetilde{\calh}_1^{n,0}$. We need to define a few quantities before we state the main result. As in the previous section, we consider $W = (W^1,W^2)$ a standard Brownian motion on $[0,1]$, and we define the two dimensional process on $[0,1]$ and for $\beta \geq 0$ $$ J(\beta)_u = \int_0^u e^{-\beta(u-s)}\sigma_s^M dW_s, \textnormal{ }u\in [0,1].$$ Next, letting $\lambda = (r_\infty/\sqrt{1-r_\infty^2},1)^T$, we consider $$ \xi(\beta)_u = W^2 -\beta\int_0^u (\sigma_s^M)^{-1} \lambda^TJ(\beta)_sds, \textnormal{ }u\in [0,1]$$ and finally $$ H(\beta) = (W^1 - \overline{W}^1, \xi(\beta) - \overline{\xi(\beta)}),$$ where we recall that for a process $(V_u)_{u \in [0,1]}$, $\overline{V} = \int_0^1 V_udu$. Note that under $\calh_0$, $\beta=0$ and $H(0)$ is simply $W-\overline{W}$. We finally introduce $$ \kappa(\beta) := \left(\frac{\int_0^1 H(\beta)_u^1 H(\beta)_u^2du}{\int_0^1 (H(\beta)_u^1)^2du},-1\right)^T.$$ The next proposition shows that the OLS estimator (and therefore the associated estimated residual process) is inconsistent under $\widetilde{\calh}_{1}^{n,\beta}$.

proposition*Let $ \beta \geq 0$. Let $L_{\alpha_0} = \sqrt{\frac{ \omega_{22}}{\omega_{11}}}\l( \sqrt{1-r_\infty^2}, r_\infty\r)^T$, and $L_{c_0} = -\sqrt{\frac{\omega_{22}}{\omega_{11}}}\l(\overline{W}^1, \overline{\xi(\beta)}\r)^T$. Then, under $\widetilde{\calh}_1^{n,\beta}$, we have the joint convergences $$ \widehat{\alpha} - \alpha_0 \to^d L_{\alpha_0}^T \kappa(\beta),$$ and $$ T^{-1/2}(\widehat{c} - c_0) \to^d L_{c_0}^T \kappa(\beta).$$

We are now ready to state the limit distribution of $\Psi$ under any local alternative $\widetilde{\calh}_{1}^{n,\beta}$. Define $$ Q(\beta) = \kappa(\beta)^T H(\beta) = \frac{\int_0^1 H(\beta)_u^1 H(\beta)_u^2du}{\int_0^1 (H(\beta)_u^1)^2du} H(\beta)^1 - H(\beta)^2.$$

theorem*Let $ \beta \geq 0$. Under $\widetilde{\calh}_1^{n,\beta}$, $$ \Psi \to^d \frac{\int_0^1 Q(\beta)_s dQ(\beta)_s}{\sqrt{\kappa(\beta)^T \kappa(\beta)\int_0^1 Q(\beta)_s^2ds} }.$$ In particular, under $\calh_0$, we have $$ \Psi \to^d \frac{\int_0^1 Q_s dQ_s}{\sqrt{\kappa^T \kappa\int_0^1 Q_s^2ds} }$$ where $$ \kappa:= \kappa(0)= \left(\frac{\int_0^1 (W_u^1 - \overline{W}^1) (W_u^2 - \overline{W}^2)du}{\int_0^1 (W_u^1 - \overline{W}^1)^2du},-1\right)^T$$ and $$ Q := Q(0) = \kappa^T(W-\overline{W}).$$

The next proposition gives the behavior of $\Psi$ (and proves the consistency of the test) under $\calh_1$, and provides an equivalent of the statistics under the stronger assumption $\textnormal{\textbf{[D]}}$.

proposition*Under $\calh_1$, we have $$ \Psi \to^\proba - \infty.$$ Moreover, under $\textnormal{\textbf{[D]}}$, we have $$\Psi \sim^\proba - n^{1/2} \sqrt{\frac{1-\rho}{1+\rho}}.$$

When $\beta = 0$, the limit distribution of $\Psi$ in Theorem (ref) is the same as the one of the ADF statistic in Theorem 4.2 of phillips1990asymptotic, up to the mean component $\overline{W}$ coming from the fact that an intercept is present in the regression. Therefore, under $\calh_0$, as in the previous section, the truncation and the deflation completely cancel the impact of jumps and that of the non ergodic volatility $\sigma^M$ that may affect $\Psi$. Moreover, in Proposition (ref), the divergence rate of $\Psi$ also corresponds to the standard one (see Theorem 5.1 in phillips1990asymptotic). However, under the local alternative $\widetilde{\calh}_1^{n,\beta}$ with $\beta >0$, the limit of $\Psi$ depends on the shape of $\sigma^M$, so that the local power of the test may be affected by a non ergodic volatility component. This feature was already present in the time-varying variance robust unit root tests of beare2018unit. More importantly, without jumps and if $\sigma^M = 1$, a careful examination of Theorem 1 in pesavento2004analytical shows that the limit distribution of the standard DF and our own modified test coincide for any $\beta \geq 0$: no local power is lost when applying the truncation and the deflation even in the absence of those features. Finally, as a direct corollary of Theorem (ref) and Proposition (ref), we conclude this section with the consistency of the modified DF test.

corollary*Let $\delta \in (0,1)$ and $q_\delta$ be the $\delta$-quantile of $\int_0^1 Q_sdQ_s/\sqrt{\kappa^T\kappa \int_0^1 Q_s^2ds}$. Then the test statistic $\Psi$ satisfies $$ \proba(\Psi < q_\delta | \calh_0) \to \delta \textnormal{ and } \proba(\Psi <q_\delta | \calh_1 ) \to 1.$$

Testing for cointegration with drifting It\^{o}-semimartingales

We now examine how the testing procedure can be adapted if the processes $X^c$ and $Z$ feature drift terms. We only partially address the problem, and restrict ourselves to the simple case of linear trends. Dealing simultaneously with general drifts, even ergodic ones, and a non ergodic volatility component is a difficult matter (at least to us) that we set aside in this work. As a matter of fact, we show in this section that even with linear drifts a natural adaptation of our DF statistic already yields a complex limit distribution that depends on the curve $u \to \sigma_u^M$ even under the null hypothesis (so that critical values must be estimated everytime the test is run). The new model for $X$ and $Y$ is

eqnarray*[eqnarray* omitted — 177 chars of source]

where $b^X,b^Z \in \reels$.

The testing procedure can be modified as follows to be drift robust. We first truncate the returns, then detrend the processes by subracting to each return the quantity $\cald(U) = n^{-1} (\calt(U)_n - \calt(U)_0)$, and finally deflate by $\sqrt{C_i}$ each truncated and detrented return. This yields the new process for $U \in \{X,Y\}$ \beas \breve{\calt}(U)_i^{def} = U_0 + \sum_{j=1}^i C_j^{-1/2}(\Delta U_j \mathbb{1}_{\{ |\Delta X_j| \leq a \Delta^{\overline{\omega}}\} \cap \{ |\Delta Y_j| \leq a \Delta^{\overline{\omega}}\}} - \cald(U)). \eeas Next, we apply the testing procedure of the previous section to $\breve{\calt}(X)_i^{def}$ and $\breve{\calt}(Y)_i^{def}$ in lieu of $\calt(X)_i^{def}$ and $\calt(Y)_i^{def}$. We denote by $\breve{\Psi}$ the associated statistic. In the following theorem, for $(V_u)_{u \in[0,1]}$ a process, we define $(\breve{V})_{u \in [0,1]}$ such that for $u\in[0,1]$,

$$\breve{V}_u = \int_0^1 V_sds + \l(\int_0^1 (\sigma_s^M)^{-1} sds - \int_{u}^1 (\sigma_s^M)^{-1} ds\r) \int_0^1 \sigma_s^M dV_s$$ whenever the integrals make sense.

theorem*Let $ \beta \geq 0$. Under $\widetilde{\calh}_1^{n,\beta}$, $\breve{\Psi}$ converges to the same limit as $\Psi$ in Theorem (ref) except that $\overline{W}$ and $\overline{\xi(\beta)}$ are respectively replaced by $\breve{W}$ and $\breve{\xi}(\beta)$ in the expression of $H(\beta)$, $\kappa(\beta)$ and $Q(\beta)$.

In particular, note that even in the case of a constant drift, the limit distribution of $\breve{\Psi}$ now depends on $\sigma^M$ even under $\calh_0$. Therefore, one needs to estimate the curve of $\sigma^M$ and then compute the related critical values by, for instance, Monte-Carlo simulations. This sheds light on the lack of applicability of the above procedure, and also indicates that dealing with a drift and time-varying volatility at the same time is a complex procedure.

Since the drift seems to be having a negligible impact in our numerical studies, at least in a realistic model and when it is calibrated to values usually encountered in empirical data, it seems more reasonable to use the simpler statistic $\Psi$ whose critical values are known and independent of the model at hand. Correspondingly, we will focus entirely on that statistic in the following finite sample experiment. In particular, we will not implement Monte-Carlo method to preestimate the curve of $\sigma^M$.

Finite sample

In this section, we conduct a Monte Carlo experiment in two steps. First, we investigate that the deflated and truncated based OLS method to estimate the cointegrated relations performs reasonably well, and that it outperforms the classical OLS procedure in a general model incorporating all the features of high frequency data in case of cointegration. Very related to estimation of relation methods is that of autocorrelation of the residuals' level, which we also look at. Second, we examine the size and power properties of the modified Dickey-Fuller residual based test for the null of no cointegration against the alternative of cointegration. In addition, we explore how the modified test performs relative to four standard residual based tests from the literature on cointegration in a variety of models, and more specifically in the presence of which feature the new test outperforms the standard procedures.

Setup

Overall, eight different models are generated. An overview is reported on Table (ref). One model (i.e. Model 8) is general and includes all the aforementioned features of high frequency data, whereas each remaining model (i.e. Model 1-7) includes one specific feature. In what follows, and for the sake of brevity, we sharpen our focus on Model 3, Model 7 and Model 8. Additional tables and comments related to the other models can be found in the Appendix. We simulate $M=1,000$ Monte Carlo paths of high-frequency returns, where each path consists of $T=2$ years of generated returns. A year is divided into 252 working days, each of them being set to 6.5 hours of trading activity, i.e. 23,400 seconds. Each path is simulated via an Euler scheme with related step set to 10 seconds.

Sampling gap

In accordance with our empirical examples, we consider the gap between two observations $\Delta$ ranging from 10 minutes, i.e. 600 seconds which sets the number of observations to $n=19,656$ across the two simulated years, to 2 days, yielding $n=252$ observations. With one observation every 10 minutes, we are enough in the high frequency regime so that the limit theory related to the truncation method kicks in, but not too much into it so that we prevent as much as possible from market microstructure effects. When the gap is two days, this is purely low frequency setting.

Simulation mechanism

We simulate $X_t^c$ and $Z_t$ as:

eqnarray*[eqnarray* omitted — 198 chars of source]

where the correlation between $W_t^X$ and $W_t^Z$ is set to $\overline{\rho} = 0.2$, i.e. $d\langle W^X, W^Z \rangle_t = \overline{\rho} dt$. Here, contrary to the theoretical setting in ((ref)), the two processes can incorporate non-zero drifts which are set to $b_t^X = 0.03 (1 + W_t^{X,b})$, and $b_t^Z = 0.02 (1+ W_t^{Z,b})$, where $W_t^{X,b}$ and $W_t^{Z,b}$ are two independent Brownian motions. Depending on the model at hands, the market volatility can be constant, linear, or including 1 jump, and may respectively take the following forms:

eqnarray[eqnarray omitted — 188 chars of source]

where we fix $\tilde{\sigma} = \sqrt{0.1}$. For $V \in \{ X,Z\}$, the idiosyncratic component of the volatility is split into a U-shape intraday seasonality component and Heston model with jumps specified as

eqnarray*[eqnarray* omitted — 64 chars of source]

where

eqnarray*[eqnarray* omitted — 203 chars of source]

with $C=0.75$, $A=0.25$, $D=0.89$, $a=10$, $c=10$, the volatility jump process is defined as $dJ_t^{\sigma,V} = M_t^{\sigma,V} S_t^{\sigma,V} dN_t^{\sigma,V}$, where the volatility jump magnitude $M_t^{\sigma,V}$ is distributed as $\mathcal{N}(0.5,0.1)$, the signs of the jumps $S_t^{\sigma,V} = \pm 1$ are i.i.d symmetric, $N_t^{\sigma,V}$ is a homogeneous Poisson process with parameter $\bar{\lambda} = 10T/252$ (with that setting volatility jumps occur randomly on average ten times a year), $\alpha = 5$, $\bar{\sigma}^2 = 1$, $\delta = 0.4$, $\bar{W}_t^V$ is a standard Brownian motion correlated to $W^V$ with $d\langle W^V,\bar{W}^V \rangle_t = \overline{\phi} dt$, $\overline{\phi} = -0.75$, $(\sigma_{0,SV}^V)^2$ is sampled from a Gamma distribution of parameters $(2\alpha\bar{\sigma}^2/\delta^2,\delta^2/2\alpha)$, which corresponds to the stationary distribution of the CIR process. To obtain more information about the model one can consult clinet2018efficient. The model is inspired directly from andersen2012jump and ait2019hausman.

In addition, for $V\in \{X,Y\}$ the price jumps are generated via $dJ_t^{V} = M_t^{V} S_t^{V} dN_t^{V}$, where the price jump magnitude $M_t^{V}$ is distributed as $\mathcal{N}(\tilde{\sigma}/\sqrt{10} ,\tilde{\sigma}/10^{3/2})$, the signs of the jumps $S_t^{V} = \pm 1$ are i.i.d symmetric, $N_t^{\sigma,V}$ is a homogeneous Poisson process with parameter $\bar{\lambda} = 10T/252$ (with that setting jumps occur on average 10 times a year and the contribution of jumps to the total quadratic variation of the price process is around 50%, both of which are roughly in line with empirical findings in huang2005relative).

Finally, the parameter related to the autocorrelation of residuals' level introduced in ((ref)) is obviously set to $\rho=1$ in case of no cointegration (i.e. null hypothesis) and chosen equal to $\rho=0.8,0.9$ when there is cointegration (i.e. in the alternative).

Concurrent methods

We implement four concurrent leading methods, all of which have already been mentioned: the DF test and the ADF test, and the Phillips-Perron tests $Z_\alpha$ and $Z_\tau$, which are tuned to cointegration in phillips1990asymptotic.

Remaining tuning parameters

We choose $\overline{\omega} = 0.48$, $a = a_0 \widehat{\sigma}_{MLE}$, $a_0 = 4$, where $\widehat{\sigma}_{MLE}$ is the daily volatility MLE, consistently with the parameter values of the numerical study (Section 5, p. 301) in clinet2019testing and ait2019hausman (except for $a_0 = 4$, because the original value ($a_0=3$) was yielding too many jumps detection in our case). Parameters related to the deflation are set to $\gamma= 1/2$ and $\gamma'=0.01$.

table[table omitted — 574 chars of source]
table[table omitted — 1,795 chars of source]
table[table omitted — 892 chars of source]
figure[figure omitted — 196 chars of source]
table[table omitted — 1,133 chars of source]
table[table omitted — 1,134 chars of source]
table[table omitted — 1,134 chars of source]

Results

Estimation of cointegrated relations via modified OLS

Table (ref) reports the bias and standard deviation of the two parameters of cointegrated relations in the case of the modified OLS and standard OLS in a general model when there is cointegration. It is clear that the standard OLS is defectively biased for any level of subsampling. On the contrary, the modified OLS works well when the frequency of subsampling is high enough, but is equally biased when the frequency decreases. This is due to the truncation method which performs more poorly when the frequency decreases. Thus, a limitation of our method when there is a price jump component is that it requires to sample at reasonable high frequencies, i.e. up to one hour.

Table (ref) reports the estimated autocorrelation of the residuals' level $\rho$. The corresponding signature plots can be found on Figure (ref). The standard estimator is off, notably in the presence of cointegration (i.e. $\rho<1$). The modified estimator is quite reliable when subsampling up to one hour, but insufficient with lower frequencies. The standard and adapted DF can be seen as testing respectively $\widetilde{\rho}=1$ and $\widehat{\rho}=1$.

Validity of modified DF tests

We turn now to the behavior of the size and power of the tests. Table (ref)-(ref) report the size and power of the tests for a variety of models.

Table (ref) report the size and power of modified DF and that of the alternative methods when there is one break in market volatility. It is clear that sizes of the concurrent methods are distorted when market volatility is non constant. Reversely, sizes of the modified DF are satisfactory at any level of sampling and for both configurations. The powers of all the methods are not affected. This indicates that the deflation provides a real advantage in practice when market volatility is non constant.

Table (ref) reports the statistical properties in case of breaks in price process. We can see that the power of the concurrent methods is distorted when price features jumps. In case of the modified DF, the power is adequate when the sampling frequency is high enough, but not suitable when the frequency decreases. This is what to be expected using the truncation method, and definitely a limitation of our method. Nonetheless, we can see that the truncation is beneficial for whoever implements standard residual based tests for no cointegration with high frequency data.

Finally, Table (ref) is concerned with a general model featuring all the aforementioned high frequency features. Mostly, the idiosyncratic effects add to each other, although the sizes of the concurrent methods are somehow not as badly impacted as in the pure non constant market volatility case.

Empirical examples

We illustrate our methodology by studying two empirical examples where in particular the modified tests results deviate from that of standard tests. The two pairs of stocks considered are Action Construction Equipment Limited (ACE) - Alexion Pharmaceuticals (ALXN) and CMS Energy Corporation (CMS) - Eversource Energy (ES), all of which traded on the S&P500.\footnote{The data were obtained through Reuters and provided by the Chair of Quantitative Finance of Ecole Centrale Paris.} In line with our numerical study, we consider a two-year-long period, i.e. 2012-2013, and subsample with frequency ranging from ten minutes to two days to conduct the tests.

ACE-ALXN case

Table (ref) reports the tests results. The corresponding signature plot of estimated autocorrelation of the residuals' level can be found on Figure (ref). The modified DF rejects the null of no cointegration at the highest frequencies, with estimated autocorrelation level around 0.90. On the contrary, the concurrent tests do not reject the null hypothesis. This is an echo of the results available on Table (ref) in the case $\rho=0.9$. It seems that there is cointegration, and that due to price jumps, the alternative tests do not reject the null hypothesis. As in the numerical study, the tests results related to the modified DF are unstable when the subsample frequency is higher or equal to two hours. The signature plot in Figure (ref) is also a replica of that in Figure (ref) related to the case $\rho=0.9$, and corroborates the aforementioned analysis.

table[table omitted — 591 chars of source]
figure[figure omitted — 232 chars of source]

CMS-ES case

Table (ref) reports the tests results. The related signature plot of estimated autocorrelation of the residuals' level is available on Figure (ref). This case is reverse from the previous case. The modified DF does not reject the null of no cointegration at the highest frequencies, while the concurrent tests do reject the null hypothesis. For this particular pair of stocks, results are to be compared with size results in Table (ref). It seems that we should trust modified DF, which indicates no cointegration, whereas the concurrent tests are altered due to time-varying market volatility. Here again the test results related to modified DF are unstable when subsampling with lower frequencies.

table[table omitted — 589 chars of source]
figure[figure omitted — 231 chars of source]

Final remarks

We have explored the challenges posed by the use of cointegration methods along with high frequency data. In terms of theoretical contribution, we have adapted the problem to the in-fill asymptotics case. We have provided a modified OLS to estimate cointegration relations when there is cointegration, together with its related central limit theory. We have also developed a (non ergodic) time-varying volatility and price-jump robust DF estimator, along with its limit theory.

In terms of applied contribution, we have seen in finite sample that some of the residual based concurrent methods to test for no cointegration are not sufficient when the model accommodates high frequency features, whereas our modified DF showed adequate size and reasonable power. Two empirical examples corroborated the fact that modified DF and standard tests can disagree in practice.