EconBase
← Back to paper

The drift burst hypothesis

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

106,425 characters · 13 sections · 19 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The drift burst hypothesis-0.50cm

abstractThe drift burst hypothesis postulates the existence of short-lived locally explosive trends in the price paths of financial assets. The recent U.S. equity and treasury flash crashes can be viewed as two high-profile manifestations of such dynamics, but we argue that drift bursts of varying magnitude are an expected and regular occurrence in financial markets that can arise through established mechanisms of liquidity provision. We show how to build drift bursts into the continuous-time It\^{o} semimartingale model, discuss the conditions required for the process to remain arbitrage-free, and propose a nonparametric test statistic that identifies drift bursts from noisy high-frequency data. We apply the test and demonstrate that drift bursts are a stylized fact of the price dynamics across equities, fixed income, currencies and commodities. Drift bursts occur once a week on average, and the majority of them are accompanied by subsequent price reversion and can thus be regarded as “flash crashes.” The reversal is found to be stronger for negative drift bursts with large trading volume, which is consistent with endogenous demand for immediacy during market crashes. JEL Classification: G10; C58. Keywords: flash crashes; gradual jumps; volatility bursts; liquidity; nonparametric statistics; microstructure noise

\thispagestyle{empty}

Introduction

\setcounter{page}{1}

The orderly functioning of financial markets is viewed by most regulators as their first and foremost objective. It is therefore unsurprising that the recent flash crashes in the U.S. equity and treasury markets are subject to intense debate and scrutiny, not least because they raise concerns around the stability of the market and the integrity of its design \citep*[see e.g.][and Figure (ref) for an illustration]{cftc-sec:10a,cftc-sec:11a,cftc-sec:15a}. Moreover, there is growing consensus that flash crashes of varying magnitude are becoming more frequent across financial markets.\footnote{Nanex Research has reported hundreds of flash crashes across all major financial markets, see \url{http://www.nanex.net/NxResearch/}. Related work includes \citet*{golub-keane-poon:12a}.} The distinct price evolution over such events -- with highly directional and sustained price moves -- poses three direct challenges to the academic community. First, how can one formally model such dynamics? The literature on continuous-time finance has focused extensively on the volatility and jump components of the price process, but as we show, these are not sufficient to explain the observed dynamics. Secondly, how can one identify or test for the presence of such features in the data? And third, can such events be reconciled within the theory of price formation in the presence of market frictions? This paper addresses all these challenges.

Suppose the log-price of a traded asset, $X = (X_{t})_{t \geq 0}$, has the dynamic:

equation[equation omitted — 123 chars of source]

where $\mu_{t}$ is the drift, $\sigma_{t}$ is the volatility, $W_{t}$ a Brownian motion, and $J_{t}$ is a jump process. In a conventional setup with locally bounded coefficients, over a vanishing time interval $\Delta \rightarrow 0$, the drift is $O_{p}( \Delta)$ (as is the jump term) and swamped by a diffusive component of larger order $O_{p}( \sqrt{ \Delta})$. Hence, much of the infill asymptotics is unaffected by the presence of a drift and the theory therefore invariably neglects it. Also, in empirical applications, particularly those relying on intraday data over short horizons, the drift term is generally small and estimates of it are subject to considerable measurement error \citep*[e.g.][]{merton:80a}. Thus, the common recommendation is to ignore it. Yet, to explain such events as those in Figure (ref), it is hard to see how the drift component can be dismissed. Our starting point is therefore -- what we refer to as -- the drift burst hypothesis, which postulates the existence of short-lived locally explosive trends in the price paths of financial assets. The objective of this paper is to build theoretical and empirical support for the hypothesis, thereby contributing towards a better understanding of financial market dynamics. We show how drift bursts can be embedded in Eq. (ref). Next, we develop a feasible nonparametric identification strategy that enables the online detection of drift bursts from high-frequency data. What is tested here is a drift explosion against the null hypothesis of a jump-diffusion model.\footnote{We assume that volatility is strictly positive, so pure-jump It\^{o} semimartingale processes are ruled out.} The empirical application demonstrates that drift bursts are a stylized fact of the price process.

figure[figure omitted — 826 chars of source]

An exploding drift is unconventional in continuous-time finance, but there are a number of theoretical models of price formation that back the idea. grossman-miller:88a examine a collection of risk-averse market makers that provide immediacy in exchange for a positive expected (excess) return $\mu = E(P_1/P_0-1)$ of the form:

equation[equation omitted — 90 chars of source]

where $P_{0}$ is the initial price, $P_{1}$ is the price at which the market maker trades, $\sigma$ is the standard deviation of the price move, $s$ is the size of the order that requires execution, $\gamma$ is the risk aversion of market makers, and $M$ is the number of market makers competing for the order. Eq. (ref) illustrates that the drift can dominate the volatility when liquidity demand ($s$) is unusually high or the willingness or capacity of the collective market makers to absorb order flow is impaired (i.e. increased risk aversion $\gamma$ or fewer active market makers $M$). This prediction fits the 2010 equity flash crash in that the price drop appeared to be accompanied by increased risk aversion and a rapid decline in the number of participating market makers. \citet*{cftc-sec:10a} write “some market makers and other liquidity providers widened their quote spreads, others reduced offered liquidity, and a significant number withdrew completely from the markets.” The subsequent price reversal observed in Figure (ref) is also predicted by this model as the excess return is only temporary and the long-run price level returns to $P_{0}$. In related work, campbell-grossman-wang:93a show that as liquidity demand increases (as measured by trading volume), the price reaction and subsequent reversal grow in size. Alternative mechanisms that can generate these price dynamics include trading frictions as in huang-wang:09a, predatory trading and forced liquidation as in brunnermeier-pedersen:05a,brunnermeier-pedersen:09a, or agents with tournament-type preferences and an aversion to missing out on trends as in Johnson:16a.\footnote{There are also a number of practical mechanisms that cause or amplify price drops and surges, such as margin calls on leveraged positions (i.e. forced liquidation), dynamic hedging of short-gamma positions, stop-loss orders, or technical momentum trading strategies.} While this literature provides valuable insights and hypotheses regarding price dynamics, the testable implications often relate to confounding measures such as the unconditional serial correlation of price returns. Moreover, because the theory is typically cast as a two-period model, it does not readily translate into an econometric identification strategy of the impacted sample paths in continuous-time. In addition to the above, there is an extensive econometrics literature on bubble detection \citep*[see e.g.][]{phillips-yu:11a}, which is potentially related to our paper, but it operates on a much longer time scale than the one analyzed here. We deliver a foundation that increases the depth of the empirical work that can be conducted in these areas.

With the drift burst hypothesis and the corresponding It\^{o} semimartingale price process in place, we develop a nonparametric identification strategy for the online detection of drift burst sample paths from intraday noisy high-frequency data. Our approach aims to establish whether the observed price movement is generated by the drift rather than be the result of volatility. Unsurprisingly, it requires estimation of the local drift and volatility coefficients which is nontrivial for several reasons. First, from \citet*{merton:80a} we know that even when the drift term is a constant, it cannot be estimated consistently over a bounded time interval. Secondly, while infill asymptotics do provide consistent volatility estimates, in practice microstructure effects complicate inference. Building on the \citet*{bandi:02a,kristensen:10a} for coefficient estimation and \citet*{newey-west:87a,andrews:91a,barndorff-nielsen-hansen-lunde-shephard:08a,jacod-li-mykland-podolskij-vetter:09a} for the robustification to microstructure noise, we formulate a nonparametric kernel-based filtering approach that delivers estimates of the local drift and volatility on the basis of which we construct the test statistic. Under the null hypothesis of no drift burst, the test is asymptotically standard normal, but it diverges -- and therefore has power under the alternative -- when the drift explodes sufficiently fast. When calculated sequentially and using potentially overlapping data, the critical values of the test are determined on the basis of extreme value theory, as in lee-mykland:08a. A simulation study confirms that the test is capable of identifying drift bursts. Interestingly, applying the test to high-frequency data for the days of the U.S. equity and treasury market flash crashes, displayed in Figure (ref), we find that they constitute highly significant drift bursts.

The introduction of drift bursts via an exploding drift coefficient provides an essential ingredient, which helps to reconcile a number of phenomena observed in financial markets. The first is the occurrence of flash crashes, where highly directional and sustained price movements are reversed shortly after. While there is a substantial body of research that looks at the 2010 equity market flash crash \citep*[a partial list includes][]{easley-prado-ohara:11a, madhavan:12a, andersen-bondarenko-kyle-obizhaeva:15a, kirilenko-kyle-samadi-tuzun:17a, menkveld-yueshen:18a}, there has been no attempt thus far to analyze these events in a more systematic fashion. Our test procedure lays down a framework that makes this possible. The second is that of “gradual jumps” -- in \citet*{barndorff-nielsen-hansen-lunde-shephard:09a} terminology -- where the price converges in a rapid but continuous fashion to a new level. This relates to a puzzle put forward by christensen-oomen-podolskij:14a, who find that the total return variation that can be attributed to the jump component is an order of magnitude smaller than previously reported by extensive empirical literature \citep*[see also][]{bajgrowicz-scaillet-treccani:16a}. In particular, they show that jumps identified using data sampled at a five-minute frequency often vanish when viewed at the highest available tick frequency and instead appear as sharp but continuous price movements. They show that spurious detection of jumps at low frequency can be explained by an erratic volatility process. However, because volatility merely leads to wider price dispersion, it fails to reconcile the often steady and directional price evolution over such episodes. We argue that the drift burst hypothesis constitutes a more intuitive and appealing mechanism that can explain the reported over-estimation of the total jump variation.

We undertake an empirical analysis to determine the prevalence of drift bursts in practice and to characterize their basic features. To that end, we employ a comprehensive set of high-quality tick data covering some of the most liquid futures contracts across the equity, fixed income, currency, and commodity markets. We calculate the drift burst test statistic at five-second intervals over a multi-year sample. Our findings demonstrate that drift bursts are an integral part of the price process across all asset classes. Over the full sample we identify more than one thousand highly significant events, or roughly one per week. For most of the drift bursts we detect, the magnitude of the price drop or surge typically ranges between 25 and 200 basis points, with only a handful of extreme moves between 3% and 8%. We find that roughly two thirds of the drift bursts are followed by price reversion, which means that many of the identified events resemble (mini) flash crashes that are symptomatic of liquidity shocks. Consistent with the literature on price formation, particularly that in huang-wang:09a of endogenous trading imbalances generated by costly market presence, we find that trading volume during a drift burst is highly correlated with the subsequent price reversal. The post-drift burst return can therefore -- as predicted by the theory -- be interpreted as a compensation for supplying immediacy during times of substantial market stress.

The remainder of the paper is organized as follows. Section (ref) introduces the drift burst hypothesis and describes the theoretical framework. Section (ref) develops the identification strategy on the basis of noisy high-frequency data. Section (ref) includes an extensive simulation study that demonstrates the properties of the test. The empirical application is found in Section (ref), while Section (ref) concludes.

The hypothesis

The log-price process $X = (X_{t})_{t \geq 0}$, introduced in Eq. (ref), is an It\^{o} semimartingale defined on a filtered probability space $( \Omega, \mathcal{F}, (\mathcal{F}_{t})_{t \geq 0}, \mathcal{P})$. We assume the following about $X$.

assumption$X$ is described by the dynamics in Eq. (ref), where $X_{0}$ is $\mathcal{F}_{0}$-measurable, $\mu = ( \mu_{t})_{t \geq 0}$ is a locally bounded and predictable drift, $\sigma = ( \sigma_{t})_{t \geq 0}$ is an adapted, c\`{a}dl\`ag, locally bounded and almost surely (a.s.) strictly positive volatility, $W = (W_{t})_{t \geq 0}$ is a standard Brownian motion and $J = (J_{t})_{t \geq 0}$ is a pure-jump process.

The above is a standard formulation for continuous-time arbitrage-free price processes. It represents our frictionless null, where the price is in a “normal” state with non-explosive (locally bounded) coefficients $\mu_{t}$ and $\sigma_{t}$. We do not restrict the model in any essential way, other than by imposing mild regularity conditions on the driving terms, which are listed in Assumption (ref) in Appendix (ref). As such, it encompasses a wide range of specifications and is compatible with time-varying expected returns, stochastic volatility, leverage effects, and jumps (both of finite and infinite activity) in the log-price and volatility. Below, we further add pre-announced jumps, explosive volatility, and additive microstructure noise. Note, however, that pure-jump processes without a Brownian term are not analyzed, because we require $\sigma_{t}$ to be strictly positive in Assumption (ref).

To introduce the drift burst hypothesis, we momentarily enforce that $X$ has continuous sample paths, i.e. $\text{d}J_{t} = 0$. The jump process is fully reactivated and driving functions are generalized in our theoretical results.

As $\mu$ and $\sigma$ are locally bounded under Assumption (ref), it follows that for a fixed time point $\tau_{\textrm{db}}$:

equation[equation omitted — 313 chars of source]

as $\Delta \rightarrow 0$. Thus, the drift is much smaller than the volatility, because $\Delta \ll \sqrt{ \Delta}$. This is consistent with the notion that over short time intervals the main contributor to the log-return is volatility. It is this feature that has led the econometrics literature to largely neglect the drift.

However, the drift can prevail in an alternative model where, in a neighborhood of $\tau_{\textrm{db}}$, $\mu$ is allowed to diverge in such a way that:

equation[equation omitted — 179 chars of source]

with $0 < \gamma_{ \mu} < 1/2$. We refer to an exploding drift coefficient as a drift burst and to $\tau_{\textrm{db}}$ as a drift burst time. The condition $\gamma_{ \mu} > 0$ ensures continuity of the sample path.

An example of an exploding drift leading to a drift burst is:

equation[equation omitted — 247 chars of source]

with $1/2 < \alpha < 1$ and $a_1$, $a_2$ constants. Setting $\gamma_{ \mu} = 1 - \alpha$, this formulation is consistent with Eq. (ref). This specification of the drift can capture flash crashes when $a_1$ and $a_2$ are of the opposite sign (see e.g. Panel A in Figure (ref)). It can also accommodate gradual jumps without reversion, e.g. when $a_2=0$.\footnote{The drift burst specification in Eq. (ref) also allows for gradual jumps that start off strong and then decelerate (when $a_1=0$ and $a_2\neq0$), akin to price behavior observed around, for instance, scheduled news announcements, where the first order price impact tends to be realized immediately, but it may then be followed by a gradual continuation as the market interprets and fully incorporates the shock.}

The log-price with explosive coefficients, replacing the drift in Eq. (ref) with a drift as in Eq. (ref), is denoted by $\widetilde{X}$. As a process it is still a semimartingale, which is necessary -- but not sufficient -- to exclude arbitrage from the model \citep*[e.g.][]{delbaen-schachermayer:94a}.\footnote{While explosive drift does not impede the semimartingale structure, it can on the other hand negatively affect nonparametric estimation of volatility from high-frequency data, as unveiled by Example 3.4.2 in \citet*{jacod-protter:12a}, because the volatility is completely swamped by the drift. This is consistent with the findings of \citet*{li-todorov-tauchen:15a}, who note that standard OLS estimation of their proposed jump regression is seriously affected by the inclusion of two outliers in the sample. Incidentally, these are the equity flash crash of May 6, 2010 (in Figure (ref)) and the hoax tweet of April 23, 2014 (in Figure (ref)).} To prevent arbitrage, a further condition (imposed by Girsanov's Theorem) is necessary for the existence of an equivalent martingale measure:

equation[equation omitted — 191 chars of source]

which is known as a “structural condition”.\footnote{The structural condition is sufficient together with the exponential moment condition $E\bigg[ \exp \Big\{- \int_{ \tau_{\textrm{db}}- \Delta}^{ \tau_{ \textrm{db}} + \Delta} \frac{\mu_{s}}{ \sigma_{s}} \text{d}W_{s} - \frac{1}{2} \int_{ \tau_{ \textrm{db}}- \Delta}^{ \tau_{\textrm{db}} + \Delta} \left( \frac{ \mu_{s}}{ \sigma_{s}} \right)^{2} \text{d}s \Big\} \bigg] = 1$, see Theorem 4.2 in \citet*[][]{karatzas-shreve:98a}.} This cannot hold if the drift explodes in the neighborhood of $\tau_{\textrm{db}}$, but the volatility remains bounded, as it allows for a so-called “free lunch with vanishing risk,” see Definition 10.6 in \citet*{bjork:03a}. Thus, explosive volatility is a necessary condition for drift bursts in a market free of arbitrage.

We say there is a volatility burst, if

equation[equation omitted — 188 chars of source]

with $0<\gamma_{\sigma}<1/2$. As above, an example of a bursting volatility is:

equation[equation omitted — 111 chars of source]

with $0 < \beta < 1/2$ and $b>0$. We here restrict $\beta$ to ensure that $\int_{\tau_{\textrm{db}}- \Delta}^{ \tau_{ \textrm{db}} + \Delta} \left( \sigma_{s}^{ \text{vb}} \right)^{2} \text{d}s < \infty$, so that the stochastic integral in Eq. (ref) can be defined. In this case,

equation[equation omitted — 159 chars of source]

so that $\gamma_{ \sigma} = 1/2- \beta$.

We can thus introduce a “canonical” alternative model:

equation[equation omitted — 139 chars of source]

for which $\mu_{t}^{\text{db}} / \sigma_{t}^{\text{vb}} \rightarrow \infty$ as $t \rightarrow \tau_{ \text{db}}$ if $\alpha > \beta$. The structural condition in Eq. (ref) is readily fulfilled when $\alpha - \beta < 1/2$. Thus, this example shows that the drift coefficient can explode locally (even after normalizing by an exploding volatility) either preserving absence of arbitrage (when $0< \alpha - \beta < 1/2$), or allowing local arbitrage opportunities (when $\alpha - \beta> 1/2$). However, as our econometric analysis in Theorem (ref) reveals, we are only able to detect an explosive drift with $\alpha - \beta > 1/2$, that is when a short-lived absence of arbitrage is found. The intuition for this result is that under the condition in Eq. (ref) we can always switch to an equivalent probability measure with no drift and unaltered volatility. Hence, there is no way to detect an explosive drift in the arbitrage-free setting.

In Panel B of Figure (ref), we show simulated sample paths of the associated log-price in this framework. A flash crash is generated from Eq. (ref). We set the drift burst rate at $\alpha = 0.65$ and $\alpha = 0.75$ with a volatility burst parameter $\beta = 0.2$. While the former parametrization preserves absence of arbitrage (since $\alpha- \beta<1/2$), the latter does not. Visually, however, the price dynamics in both scenarios is pronounced and qualitatively identical.

The notion that volatility can burst during market turbulence or dislocation is uncontroversial. \citet*{kirilenko-kyle-samadi-tuzun:17a} and \citet*{andersen-bondarenko-kyle-obizhaeva:15a} report elevated levels of volatility during the equity flash crash \citep*[see also][]{bates:18a}. However, as illustrated by Figure (ref) and proved formally in Theorem (ref), a volatility burst in itself is not sufficient to capture the gradual jump or flash crash dynamics regularly observed in practice. The introduction of a separate drift burst component as we propose in this paper is a convenient and effective tool to reconcile continuous-time semimartingale theory with such empirical observations.

figure[figure omitted — 849 chars of source]

Identification

We now develop a nonparametric approach to detect drift bursts in real data. We propose a test statistic that exploits the message of Eq. (ref), namely if there is a drift burst in the price process at time $\tau_{\text{db}}$, the drift can prevail over volatility and locally dominate log-returns in the vicinity of $\tau_{\text{db}}$. The test statistic thus compares a suitably rescaled estimate of $\mu_{t} / \sigma_{t}$ based on high-frequency data in a neighborhood of $t$. Later, we prove that our “signal-to-noise” measure uncovers drift bursts if they are sufficiently strong.

We extend existing work on nonparametric kernel-based estimation of the coefficients of diffusion processes to estimate $\mu_{t}$ and $\sigma_{t}$ bandi:02a,kristensen:10a. We assume that $X$ is recorded at times $0 = t_{0} < t_{1} < \ldots < t_{n} = T$, where $\Delta_{i,n} = t_{i} - t_{i-1}$ is the time gap between observations and $T$ is fixed. The sampling times are potentially irregular, as formalized in Assumption (ref) in Appendix (ref). The discretely sampled log-return over $[t_{i-1},t_{i}]$ is $\Delta_{i}^{n} X = X_{t_{i}} - X_{t_{i-1}}$, and we define:

equation[equation omitted — 193 chars of source]

where $h_{n}$ is the bandwidth of the mean estimator and $K$ is a kernel. We also set:

equation[equation omitted — 243 chars of source]

where $h_{n}'$ is the bandwidth of the volatility estimator.

The bandwidths $h_{n}$, $h_{n}'$ and the kernel $K$ are assumed to fulfill some weak regularity conditions that are succinctly listed in Assumption (ref) in Appendix (ref).

In absence of a drift burst, the proof of Theorem (ref) stated below shows that, as $n \rightarrow \infty$:

equation[equation omitted — 180 chars of source]

where $\mu_{t}^{*} = \mu_{t} + \int_{ \mathbb{R}} \delta(t,x) I_{ \left\{| \delta(t,x)| > 1 \right\}} \lambda( \text{d}x)$, $K_{2}$ is a kernel-dependent constant, and the above convergence is stable in law.

As shown by Eq. (ref), $\hat{ \mu}_{t}^{n}$ is asymptotically unbiased for the (jump compensated) drift term. It is inconsistent, because the variance explodes as $h_{n} \rightarrow 0$. This appears to rule out drift burst detection via $\hat{ \mu}_{t}^{n}$. On the other hand, if we rescale the left-hand side of Eq. (ref) with $\hat{ \sigma}_{t}^{n}\sqrt{K_{2}}$, it appears the right-hand side has a standard normal distribution.\footnote{Lemma (ref) in Appendix (ref) shows that $\hat{ \sigma}_{t}^{n}$ is a consistent estimator of $\sigma_{t-}$.} It is this insight that facilitates the construction of a test statistic that can identify drift bursts, as we prove in Theorem (ref), which describes the behavior of the $t$-statistic under the null of locally bounded coefficients, whereas Theorem (ref) does it under the alternative of explosive drift and volatility.

The test statistic is defined as:

equation[equation omitted — 135 chars of source]

$T_{t}^{n}$ has an intuitive interpretation with the indicator kernel. Here, it is the ratio of the drift part to the volatility part of the log-return over the interval $[t-h_{n},t]$ (as $h_{n}, h_{n}' \rightarrow 0$, this holds for any valid kernel).

theoremAssume that $X$ is a semimartingale as defined by Eq. (ref), and that Assumption (ref) and (ref) -- (ref) are fulfilled. As $n \rightarrow \infty$, it holds that: \begin{equation} T_{t}^{n} \overset{d}{ \rightarrow} N(0,1). \end{equation}
proofSee Appendix (ref). $\blacksquare$

Theorem (ref) shows that, in absence of a drift burst, the $t$-statistic in Eq. (ref) has a limiting standard normal distribution.\footnote{While this statement appears to follow trivially from Eq. (ref) -- i.e., via application of Slutsky's Theorem -- this is not true. In general, we can only use Eq. (ref) to deduce Eq. (ref), if $\sigma_{t-}$ is a constant. In our paper, where $\sigma_{t-}$ is a random variable, the definition of convergence in distribution does not support such a conclusion. We therefore prove in Appendix (ref) that the convergence in Eq. (ref) is in law stably, which is a stronger form of convergence that helps to recover this feature \citep*[the concept is explained in e.g.][]{jacod-protter:12a}. Moreover, we allow for leverage effects and jumps. In both these directions, Theorem (ref) extends kristensen:10a.} We note that under the null, $T_{t}^{n}$ does not depend on $\mu_{t-}^{*}$ and $\sigma_{t-}$ for large $n$. Thus, although it is not possible to consistently estimate $\mu_{t-}^{*}$, we can exploit its asymptotic distribution to form a test of the drift burst hypothesis, since a large $t$-statistic signals that the realized log-return is mostly induced by drift.\footnote{\citet*{todorov-tauchen:14a} build a related type of self-normalized statistic to test whether the price process is jump-diffusion or of pure-jump type. Their procedure has power against the pure-jump alternative. It is therefore important in future work to establish the properties of our procedure in a pure-jump setting to ensure it has size control in such models.}

The drift burst alternative is formalized by an exploding $\mu_{t}$ term. We note that the alternative is broad enough to allow $\sigma_{t}$ to co-explode with the drift.

theoremAssume that $\widetilde{X}$ is of the form: \begin{equation} {d} \widetilde{X}_{t} = {d} X_{t} + \frac{c_{1,t}}{ (\tau_{{db}} - t)^{ \alpha}} {d}t + \frac{c_{2,t}}{ (\tau_{{db}} - t)^{ \beta}} {d}W_{t}, \end{equation} where $\tau_{\text{\upshape{db}}}>0$, $\text{\upshape{d}}X_{t}$ is the model with locally bounded coefficients in Eq. (ref) for which the conditions of Theorem (ref) hold, $c_{1,t}$ and $c_{2,t}$ are adapted stochastic processes adhering to identical conditions as for $\mu_{t}$ and $\sigma_{t}$ listed in Assumption (ref). Moreover, $\alpha$ and $\beta$ are constants such that $0 \leq \beta < 1/2$ and $0 < \alpha < 1$. Then, as $n \rightarrow \infty$, it holds that: \begin{equation} T_{ \tau_{\text{{db}}}}^{n}\left\{ \begin{array}{lc} \overset{a.s.}{ \rightarrow} \pm \infty, & \text{if} \quad \alpha - \beta > 1/2, \\[0.10cm] \overset{d}{\rightarrow} c_{K, \beta} N(0,1) + d_{K, \beta, c_{1}, c_{2}}, & \text{if} \quad \alpha - \beta = 1/2, \\[0.10cm] \overset{d}{\rightarrow} c_{K, \beta} N(0,1), & \text{if} \quad \alpha - \beta < 1/2, \end{array} \right. \end{equation} where \begin{equation} c_{K, \beta} = \sqrt{ \frac{ \int_{ \mathbb{R}}K^{2}(x)|x|^{-2\beta} \text{{d}}x}{K_{2} \int_{\mathbb{R}}K(x)|x|^{-2 \beta} \text{{d}}x}} \quad \text{and} \quad d_{K, \beta,c_{1},c_{2}} = \frac{c_{1, \tau_{\text{{db}}}}}{c_{2, \tau_{\text{{db}}}}} \frac{ \int_{ \mathbb{R}}K^2(x)|x|^{- \beta-1/2} \text{{d}}x}{ \sqrt{ K_{2} \int_{ \mathbb{R}}K(x)|x|^{-2 \beta} \text{{d}}x}}. \end{equation}
proofSee Appendix (ref). $\blacksquare$

Theorem (ref) implies the explosion of the $t$-statistic under the alternative, when the drift term explodes fast enough relative to the volatility (i.e. $\alpha- \beta> 1/2$). The condition $\alpha- \beta> 1/2$ is equivalent to require that, in a neighborhood of $\tau_{ \text{\upshape{db}}}$, the log-return is dominated by drift. In a frictionless economy this allows for a short-lived arbitrage around $\tau_{ \text{\upshape{db}}}$. In practice, to make the test statistic large, it is of course enough that the mean log-return is significantly nonzero over a small time interval. Our simulations confirm that the condition $\alpha- \beta>1/2$ is not needed to achieve power in small samples.\footnote{As detailed in Appendix (ref), we can estimate $\alpha$ and $\beta$ using a parametric maximum likelihood approach based on Eq. (ref). The sample averages across detected events in the empirical high-frequency data analyzed in Section (ref) are $\bar{ \hat{ \alpha}}_{ \text{ML}} = 0.6250$ and $\bar{ \hat{ \beta}}_{ \text{ML}} = 0.1401$. Hence, assuming that bursts are of the form in Eq. (ref), the process is found to be right at the margin of being arbitrage-free.}

An implication of Theorem (ref) is that $T_{ \tau_{\text{\upshape{db}}}}^{n}$ does not explode because of a volatility burst, even if it occurs without a drift burst. With $\alpha - \beta < 1/2$ the $t$-statistic is normally distributed in the limit with a controllable asymptotic variance that depends on the kernel shape and $\beta$. With a left-sided exponential kernel $K(x) = \exp(-|x|)$, for $x \leq 0$, adopted below, $c_{K, \beta} =2^{ \beta} \leq \sqrt{2}$. At the border, $\alpha - \beta=1/2$, the $t$-statistic is bounded but can take arbitrarily large values as $\alpha \rightarrow 1$ and $\beta \rightarrow 1/2$, because $\lim_{ \beta \rightarrow 1/2} d_{K, \beta, c_{1}, c_{2}} = \infty$.\footnote{\citet*{hengartner-linton:96a} study nonparametric regression estimation at design poles and zeros. A nearly identical set of “bias” and “variance” kernel constants to $d_{K, \beta, c_{1}, c_{2}}$ and $c_{K, \beta}$ also appear in the asymptotic distribution of their estimator.}

Thus, the large values of the test statistic frequently observed in the empirical application are neither explained by volatility bursts nor by jumps. We conclude that -- in our setting -- a significant $t$-statistic can only be induced by a drift explosion.\footnote{A likelihood ratio test based on the model in Eq. (ref), also reported in Appendix (ref), strongly rejects $\mathcal{H}_{0}^{ \prime} : \beta = 0$ against $\mathcal{H}_{1}^{ \prime} : \beta > 0$ during a drift burst, yielding empirical support for the volatility co-exploding with the drift.} These theoretical statements are also corroborated by our simulation analysis.

remarkIn this paper, we are testing against drift explosions. However, other mechanisms can possibly cause the $t$-statistic to deviate from the asymptotic normal distribution. First, pure-jump processes without a Brownian component can be a potential alternative \citep*[e.g.][]{todorov-tauchen:14a}. Second, it is possible to work with the persistence of jumps in a Hawkes-type process to construct a “jump burst.” Third, one can develop a market microstructure model with endogenous noise that depends on the fundamental value of the asset \citep*[e.g.][]{li-linton:20a}. These ideas to describe the departures from the jump-diffusion It\^{o} semimartingale framework that our paper reveals, via the drift burst hypothesis, constitute promising avenues for future research. In the end, it is the economic application that dictates how we label such deviations.

Inference via the maximum statistic

The asymptotic theory asserts that $T_{t}^{n}$ is standard normal in absence of a drift burst, whereas it grows arbitrarily large under the alternative, as we approach a drift burst time. It suggests that a viable detection strategy is to compute the $t$-statistic progressively over time and reject the null when $|T_{t}^{n}|$ gets significantly large. This leads to a multiple testing problem, which can cause size distortions, if the quantile function of the standard normal distribution is used to determine a critical value of $T_{t}^{n}$.

To control the family-wise error rate, we evaluate a standardized version of the maximum of the absolute value of our $t$-statistic using extreme value theory.\footnote{\citet*{bajgrowicz-scaillet-treccani:16a} and \citet*{lee-mykland:08a} also exploit these ideas in the high-frequency framework to devise an unbiased jump-detection test, while in a related context \citet*{andersen-bollerslev-dobrev:07a} propose a Bonferroni correction. The latter was another viable tool to avoid systematic overrejection of the null hypothesis.} We compute $\left(T_{t_{i}^{*}}^{n} \right)_{i=1}^{m}$ at $m$ equispaced time points $t_{i}^{*} \in (0,T]$, where $T$ is fixed. We set:

equation[equation omitted — 115 chars of source]

The crucial point is that, in addition to $T_{t_{i}^{*}}^{n} \overset{d}{ \rightarrow} N(0,1)$ under the null, the $T_{t_{i}^{*}}^{n}$'s are also independent -- up to error terms that are asymptotically negligible -- if $m$ does not grow too fast. It follows that a normalized version of $T_{m}^{*}$ has a limiting Gumbel distribution, as $m \rightarrow \infty$ at a suitable rate.

theoremAssume that the conditions of Theorem (ref) hold. Then, if $n \rightarrow \infty$, $m \rightarrow \infty$ such that $mh_{n} \rightarrow 0$ and $\sqrt{ \log(m)} \bigg( \frac{m}{ \sqrt{nh_{n}}}+(mh_{n})^{-B} \sqrt{ \log(m)} + m^{- \Gamma/2} \sqrt{ \log(m)} \bigg) \rightarrow 0$, it further holds that: \begin{equation} (T_{m}^{*} - b_{m})a_{m} \overset{d}{ \rightarrow} \xi, \end{equation} where \begin{equation} a_{m} = \sqrt{2 \log(m)}, \qquad b_{m} = a_{m} - \frac{1}{2} \frac{ \log(\pi \log(m))}{a_{m}}, \end{equation} and the CDF of $\xi$ is the Gumbel, i.e. $P( \xi \leq x) = \exp( -\exp(-x))$.
proofSee Appendix (ref). $\blacksquare$

The alternative to Theorem (ref) is that there is at least one drift burst (as formalized in Theorem (ref)) in the time interval $[0,T]$, whence the maximum statistic $T_{m}^{*}$ diverges rendering the test consistent.

In practice, the Gumbel distribution is conservative, as the convergence in (ref) is slow and because of residual dependence in the $t$-statistics due to small sample effects, microstructure noise and pre-averaging (introduced in Section (ref)). In Appendix (ref), we propose a simulation-based procedure to determine data-driven critical values.

Robustness to a pre-announced jump

We here study an extended model that -- on top of the drift, volatility and jump component in Eq. (ref) -- has a “pre-announced” jump \citep*[see e.g.][]{jacod-li-zheng:17a,dubinsky-johannes-kaeck-seeger:18a}, where the jump time is fixed across sample paths. We show these types of jumps do not compromise drift burst detection.

theoremAssume that ${ \mathpalette\double@widetilde{X} }$ is of the form: \begin{equation} {d} { \mathpalette\double@widetilde{X} }_{t} = {d}X_{t} + {d}J_{t}', \end{equation} where $\text{\upshape{d}}X_{t}$ is the model in Eq. (ref) such that the conditions of Theorem (ref) hold, while $J_{t}' = J \cdot I_{\{0< \tau_{J} \leq t\}}$, $\tau_{J}$ is a stopping time, and $J$ is $\mathcal{F}_{ \tau_{J}}$-measurable. Then, as $n \rightarrow \infty$, it holds that: \begin{equation} T_{ \tau_{J}}^{n} \overset{p}{ \rightarrow} \sqrt{ \frac{K(0)}{K_{2}}} \cdot \operatorname{sign}(J). \end{equation}
proofSee Appendix (ref). $\blacksquare$

We can readily select a kernel that can tell apart the occurrence of a fixed jump from a drift explosion. In particular, the left-sided exponential kernel advocated above has $\displaystyle \sqrt{ K(0)/K_{2}} = \sqrt{2}$, so that $|T_{ \tau_{J}}^{n}| \overset{p}{ \rightarrow} \sqrt{2}$. Thus, our proposed $t$-statistic is -- asymptotically -- small under the null (standard normal distributed) and pre-announced jump alternative ($\sqrt{2}$ in absolute value), while it is large (diverging) under the drift burst alternative.

Robustness to microstructure noise

In practice, we do not measure the true, efficient log-price $X_{t_{i}}$, because transaction and quotation data are disrupted by multiple layers of “noise” or “friction” \citep*[e.g.][]{black:86a,stoll:00a}. In this section, we show how to modify our test for drift bursts, so it is resistant to such features of the market microstructure at the tick level. To incorporate noise, we suppose that:

equation[equation omitted — 97 chars of source]

where $\epsilon_{t_{i}}$ is an additive error term.

assumption$( \epsilon_{t_{i}})_{i=0}^{n}$ is adapted and independent of $X$. Moreover, $E[ \epsilon_{t_{i}}] = 0$, $E \big[ \epsilon_{t_{i}}^{4} \big] < \infty$, and denoting the autocovariance function by $\gamma_{k} = E[\epsilon_{t_{i}}\epsilon_{t_{i+k}}]$ for any integer $k\geq 0$, we further assume $\gamma_{k}$ is finite, independent of $i$ and $n$, and such that $\gamma_{k} = 0$ for $k > Q$, where $Q \geq 0$ is an integer (i.e., $Q$-dependent noise).

This is a standard noise model in financial econometrics. It allows for autocorrelation in the noise process, but it rules out dependence between the noise and fundamental price, as in \citet*{li-linton:20a}. The difficulty brought by noise is then that in order to do inference about drift bursts in $X$, we are forced to work with the contaminated high-frequency record of $Y$.

A direct application of the $t$-statistic in Eq. (ref) to the noise-contaminated returns $\Delta_{i}^{n} Y$ is powerless, because the noise asymptotically dominates the other shocks and overwhelms the signal of an exploding drift. A solution to this problem is to slightly slow down the accumulation of noise by pre-averaging $Y_{t_{i}}$, as in \citet*{jacod-li-mykland-podolskij-vetter:09a}. We define a pre-averaged increment for any stochastic process $V$:

equation[equation omitted — 215 chars of source]

where $k_{n}$ is the pre-averaging window, $g_{j}^{n} = g(j/k_{n})$ and $H_{j}^{n} = g_{j+1}^{n}-g_{j}^{n}$ with $g: [0,1] \mapsto \mathbb{R}$ continuous and piecewise continuously differentiable with a piecewise Lipschitz derivative $g^{\prime}$, such that $g(0)=g(1)=0$ and $\int_{0}^{1}g^{2}(s) \text{d}s< \infty$. Absent a drift burst, it follows that:

equation[equation omitted — 317 chars of source]

As Eq. (ref) shows, the noise is reduced by a factor $\sqrt{k_{n}}$. The drift and volatility of $X$ are enhanced by $\sqrt{k_{n}}$, while leaving their relative order unchanged. Intuitively, it therefore suffices with minimal pre-averaging to bring down the noise enough and make a fast drift burst with $\alpha$ close to one dominate the divergence of the asymptotic variance in the drift estimator.

The drift burst $t$-statistic in Eq. (ref) is then refined as a noise-robust version, where the drift estimator and its long-run variance are computed from the pre-averaged return series:

equation[equation omitted — 311 chars of source]

with

equation[equation omitted — 281 chars of source]

and

equation[equation omitted — 643 chars of source]

where $w: \mathbb{R}_{+} \rightarrow \mathbb{R}$ is a kernel with $w(0)=1$ and $w(x) \rightarrow 0$ as $x \rightarrow \infty$, and $L_{n}$ is the lag length that determines the number of autocovariances in (ref). $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \sigma\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$ is a heteroscedasticity and autocorrelation consistent (HAC)-type statistic \citep*[e.g.][]{newey-west:87a, andrews:91a}. The extra complexity is required to account for any noise dependence and the serial correlation induced by pre-averaging to consistently estimate the asymptotic variance of $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \mu\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$.

theoremSet $Y_{t_{i}} = X_{t_{i}} + \epsilon_{t_{i}}$, where $X$ is defined by Eq. (ref) and $\epsilon$ is defined by Assumption (ref). Suppose that Assumption (ref) -- (ref) are fulfilled. For every fixed $t \in (0,T]$, as $n \rightarrow \infty$, $k_{n} \rightarrow \infty$, $L_{n} \rightarrow \infty$ such that $k_{n}h_{n} \rightarrow 0$, $k_{n} h_{n}' \rightarrow 0$, $\frac{k_{n}}{nh_{n}} \rightarrow 0$, $\frac{k_{n}}{nh_{n}'} \rightarrow 0$, and $\frac{L_{n}}{nh_{n}'} \rightarrow 0$, it holds that: \begin{equation*} \mkern 1.5mu\overline{\mkern-1.5muT\mkern-1.5mu}\mkern 1.5mu_{t}^{n} \overset{d}{ \rightarrow} N(0,1). \end{equation*}
proofSee Appendix (ref). $\blacksquare$

The conditions $k_{n}h_{n} \rightarrow 0$ and $\frac{k_{n}}{nh_{n}} \rightarrow 0$ (and the associated ones with $h_{n}'$) call for moderate pre-averaging, so that the number of pre-averaged terms is not too large. The condition $\frac{L_{n}}{nh_{n}'} \rightarrow 0$ means the lag length also cannot grow too fast when estimating the long-run variance of the drift estimator.

theoremSet $Y_{t_{i}} = \widetilde{X}_{t_{i}} + \epsilon_{t_{i}}$, where $\widetilde{X}$ is defined as in Theorem (ref) and everything else is maintained as in Theorem (ref). As $n \rightarrow \infty$, $k_{n} \rightarrow \infty$, $L_n\rightarrow\infty$ such that $k_{n} h_{n} \rightarrow 0$, $k_{n} h_{n}' \rightarrow 0$, $\frac{k_{n}}{nh_{n}} \rightarrow 0$, $\frac{k_{n}}{nh_{n}'} \rightarrow 0$, and $\frac{L_{n}}{nh_{n}'} \rightarrow 0$, and $k_{n}h_{n}^{2(1- \alpha)} \rightarrow \infty$, it holds that $| \mkern 1.5mu\overline{\mkern-1.5muT\mkern-1.5mu}\mkern 1.5mu_{ \tau_{ \text{\upshape{db}}}}^{n}| \overset{p}{ \rightarrow} \infty$ for $\alpha- \beta > 1/2$.
proofSee Appendix (ref). $\blacksquare$

This confirms that in presence of noise the pre-averaged test statistic converges in law to a standard normal under the null of no drift burst, while it again diverges at a drift burst time, if $\alpha - \beta > 1/2$. Under the alternative, pre-averaging cannot be too moderate ($k_{n}h_{n}^{2(1- \alpha)} \rightarrow \infty$), otherwise the signal of an exploding drift is not enhanced enough relative to the noise. The condition also shows that with $\alpha$ closer to $1$, we need less pre-averaging.

Simulation study

In this section, we adopt a Monte Carlo approach to further explore the $t$-statistic proposed in Eq. (ref) as a tool to uncover drift bursts in $X$. The overall goal is to investigate the size and power properties of our test and figure out how “small” drift bursts we are able to detect under the alternative, amid also an exploding volatility, the presence of infinity-activity price jump processes, and microstructure noise.

We simulate a driftless heston:93a-type stochastic volatility (SV) model:

align[align omitted — 256 chars of source]

where $W$ and $B$ are standard Brownian motions with $E( \text{d}W_{t} \text{d}B_{t}) = \rho \text{d}t$. Thus, the drift-to-volatility ratio of the efficient log-price is $\mu_{t}/\sigma_{t} = 0$.

We configure the variance process to match key features of real financial high-frequency data. As consistent with prior work \citep*[e.g.][]{ait-sahalia-kimmel:07a}, we assume the annualized parameters of the model are $( \kappa, \theta, \xi, \rho) = (5, 0.0225, 0.4, -\sqrt{0.5})$. We note $\theta$ implies an unconditional standard deviation of log-returns of roughly 15% p.a., which aligns with what we observe across assets in our empirical study. A total of $\text{1,000}$ repetitions is generated via an Euler discretization. In each simulation, $\sigma_{t}^{2}$ is initiated at random from its stationary law $\sigma_{t}^{2} \sim \text{Gamma}(2 \kappa \theta \xi^{-2}, 2 \kappa \xi^{-2})$. The sample size is $n = \text{23,400}$, which is representative of the liquidity in the futures contracts analyzed in Section (ref) (see Table (ref)). It corresponds to second-by-second sampling in a 6.5 hours trading session.

We create drift and volatility bursts with the parametric model:

equation[equation omitted — 283 chars of source]

with $\tau_{\text{db}} = 0.5$ fixed. Here, the price experiences a short-lived flash crash at $\tau_{\text{db}}$, as consistent with our empirical finding that most of the identified drift bursts are followed by partial or full recovery.\footnote{To ensure $X$ reverts during a pure volatility burst, we recenter the log-return series associated with $\sigma_{t}^{\text{vb}}$, so that $\int_{0}^{T} \sigma_{t}^{\text{vb}} \text{d}W_{t} = 0$. This has almost no impact on the outcome of the $t$-statistic, but it makes the price processes comparable across settings.} The window $[0.475, 0.525]$ can be interpreted as making the duration of the bursts last about 20 minutes. We set $\alpha = (0.55, 0.65, 0.75)$ and $\beta = (0.1, 0.2, 0.3, 0.4)$ to gauge their impact on our $t$-statistic.\footnote{Note that as $\alpha - \beta > 1/2$ for some of these combinations, the model is not always devoid of arbitrage.} In particular, fixing $a = 3$ we induce a cumulative return $\int_{0}^{ \tau_{\text{db}}} \mu_{t}^{ \text{db}} \text{d}t$ of about $-0.5\%$ (with opposite sign after the crash) for $\alpha = 0.55$ to slightly less than $-1.5\%$ for $\alpha = 0.75$, as comparable to what we observe in the real data. Also, with $b = 0.15$ our choices of $\beta$ produce a 25% ($\beta = 0.1$) to more than 100% ($\beta = 0.4$) increase in the standard deviation of log-returns in the drift burst window relative to its unconditional level across simulations. A drift burst is therefore accompanied by highly elevated volatility, making it challenging to detect the signal.

To explicate the robustness of our $t$-statistic, in each simulation we superposition on top of the continuous sample path model in Eq. (ref) a L\'{e}vy process with jump density given by:

equation[equation omitted — 90 chars of source]

where $\upsilon > 0$ and $\psi > 0$. This defines a pure-jump tempered stable process with activity index $\upsilon$. We assume $\upsilon = 0.5$, such that an infinite-activity finite-variation process with both small and large jumps appearing randomly in $X$ is produced. We further set $\lambda = 3$ and calibrate $\psi$ so that on average 20% of the quadratic variation is induced by the jump component. This is broadly in line with previous studies, e.g. \citet*{ait-sahalia-jacod-li:12a, ait-sahalia-xiu:16a}. The process is simulated as the difference between two positive tempered stable processes, as outlined in \citet*{todorov-tauchen-grynkiv:14a}. These are generated with the acceptance-rejection algorithm of \citet*{baeumer-meerschaert:10a}. Note that the discretization is exact for the selected value of $\upsilon$. Also, in agreement with our theoretical exposition the jump process is active both under the null and alternative hypothesis.

The noisy log-price is:

equation[equation omitted — 78 chars of source]

where $\epsilon_{i/n} \sim N \left(0, \omega_{i/n}^{2} \right)$ and $\displaystyle \omega_{i/n} = \gamma \frac{ \sigma_{i/n}}{ \sqrt{n}}$, so the noise is both conditionally heteroscedastic, serially dependent (via $\sigma$), and positively related to the riskiness of the efficient log-price \citep*[e.g.][]{bandi-russell:06a, oomen:06a, kalnina-linton:08a}. We should note that this is a more general microstructure noise than allowed by Assumption (ref). $\gamma$ is the noise-to-volatility ratio. We set $\gamma = 0.5$, which amounts to a medium contamination level \citep*[e.g.][]{christensen-oomen-podolskij:14a}. To reduce the noise, we pre-average $Y_{i/n}$ locally within a block of length $k_{n} = 3$ and based on the weight function $g(x)=\min(x,1-x)$.\footnote{In the Online Appendix, we present a comprehensive analysis with $\gamma = 0.5, 2$ and $5$ and pre-averaging horizon $k_{n} = 1, \ldots, 10$. The results do not differ materially from those reported here. A modest loss of power is noted, however, if $h_{n}$ is small and $k_{n}$ is large.}\textsuperscript{,}\footnote{With equidistant data, it follows that if $k_{n}$ is even and $g(x) = \min(x,1-x)$, the pre-averaged return in Eq. (ref) can be rewritten as $\Delta_{i}^{n} \bar{Y} = \frac{1}{k_{n}} \sum_{j=1}^{k_{n}/2} Y_{ \frac{i+k_{n}/2+j}{n}} - \frac{1}{k_{n}} \sum_{j=1}^{k_{n}/2} Y_{ \frac{i+j}{n}}$. Thus, the sequence $(2 \Delta_{i}^{n} \bar{Y})_{i=1}^{n-k_{n}+2}$ can be interpreted as constituting a new set of increments from a price process that is constructed by averaging of the rescaled noisy log-price series, $(Y_{i/n})_{i=0}^{n}$, in a neighbourhood of $i/n$, thus making the use of the term pre-averaging and the associated notation transparent.} $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \mu\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$ and $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \sigma\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$ are constructed from $( \Delta_{i}^{n} \mkern 1.5mu\overline{\mkern-1.5muY\mkern-1.5mu}\mkern 1.5mu)_{i=0}^{n-2k_{n}+1}$ based on Eq. (ref) and (ref) with a left-sided exponential kernel $K(x) = \exp(-|x|)$, for $x \leq 0$.

A Parzen kernel is selected for $w$:

equation[equation omitted — 225 chars of source]

This choice has some profound advantages in our framework. First, the Parzen kernel ensures that $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \sigma\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$ is positive, so we can always compute the $t$-statistic, which is not true for a general weight function. Secondly, the efficiency of the Parzen kernel is near-optimal, e.g. \citet*{andrews:91a, barndorff-nielsen-hansen-lunde-shephard:09a}. The slight loss of efficiency brings the distinct merit that $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \sigma\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$ can be computed on the back of the first $L_{n}$ lags of the autocovariance function, whereas more efficient weight functions typically require all $n$ lags. In the high-frequency framework $n$ is large, so the latter can be prohibitively slow to compute. In contrast, $L_{n}$ is typically small compared to $n$ in practice, rendering our choice of kernel much less time-consuming.

We set $L_{n} = Q^{*} + 2(k_{n}-1)$ to estimate $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \sigma\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$. Here, $2(k_{n}-1)$ is due to pre-averaging and we compute $Q^{*}$ from $( \Delta_{i}^{n} Y)_{i=1}^{n}$ as a data-driven measure of noise dependence based on automatic lag selection \citep*[see, e.g.][]{newey-west:94a, barndorff-nielsen-hansen-lunde-shephard:09a}.\footnote{In our simulations, the average value of $Q^{*}$ is 11.9, while its interquartile range is 8 -- 17.} The bandwidth for $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \mu\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$ is varied in $h_{n} = (120,300,600)$ seconds. We use a larger bandwidth of $h_{n}' = 5h_{n}$ for $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \sigma\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$ to better capture persistence in volatility and estimate the microstructure-induced return variation.

A new value of $\mkern 1.5mu\overline{\mkern-1.5muT\mkern-1.5mu}\mkern 1.5mu_{t}^{n}$ is recorded at every 60th transaction update.\footnote{We set a burn-in period of a full volatility bandwidth before testing to allow for a sufficient number of observations to construct $\mkern 1.5mu\overline{\mkern-1.5muT\mkern-1.5mu}\mkern 1.5mu_{t}^{n}$.} We extract $T_{m}^{*} = \max_{i = 1, \ldots, m} | \mkern 1.5mu\overline{\mkern-1.5muT\mkern-1.5mu}\mkern 1.5mu_{t_{i}^*}^{n}|$ based on the resulting $m = 341$ tests in each sample. The simulation-based approach explained in Appendix (ref) is adopted to find a critical value of $T_{m}^{*}$.

figure[figure omitted — 930 chars of source]

Figure (ref) reports Q-Q plots of the distribution of $\mkern 1.5mu\overline{\mkern-1.5muT\mkern-1.5mu}\mkern 1.5mu_{t}^{n}$ in absence of a drift burst. Panel A is based on the \citet*{heston:93a}-type SV+jump model. As readily seen, the Gaussian curve is an accurate description of the sampling variation of $\mkern 1.5mu\overline{\mkern-1.5muT\mkern-1.5mu}\mkern 1.5mu_{t}^{n}$, although the $t$-statistic is slightly thin-tailed with $h_{n} = 120$. This is caused by a modest correlation between $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \mu\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$ and $\hat{ \mkern 1.5mu\overline{\mkern-1.5mu \sigma\mkern-1.5mu}\mkern 1.5mu}_{t}^{n}$, which are computed from partly overlapping data; an effect that is more pronounced for small bandwidths. In Panel B, the process featuring a large volatility burst ($\beta = 0.4$) is plotted.\footnote{The results for other values of $\beta$ fall in between those of Panel A and B and are therefore not reported.} While the volatility burst puts more mass into the tails of the distribution of $\mkern 1.5mu\overline{\mkern-1.5muT\mkern-1.5mu}\mkern 1.5mu_{t}^{n}$, the normal continues to be a good approximation.

table[table omitted — 3,682 chars of source]

This is corroborated by Table (ref), where we compute how often $T_{m}^{*}$ leads to rejection of the null hypothesis of no drift burst for three significance levels $c = 5\%, 1\%, 0.5\%$. There are several interesting findings. Look at the columns with $\mu_{t}^{ \text{db}} \equiv 0$, which report the results in absence of a drift burst (i.e., size). Without a volatility burst ($\beta = 0.0$), the test is conservative compared to the nominal level if $h_{n}$ is small, as also reflected in Figure (ref). As $\beta$ increases, $T_{m}^{*}$ is mildly inflated yielding a tiny size distortion, but the effect is only present for large $\beta$ and $h_{n}$. Otherwise, the test is roughly unbiased. This suggests our $t$-statistic is adaptive and highly robust to even substantial shifts in spot variance, so that we do not falsely pick up an explosion in volatility as a significant drift burst.

Turn next to the alternative with a drift burst (i.e., power). As expected, the power is increasing in $\alpha$, holding $\beta$ fixed, while it is decreasing in $\beta$, holding $\alpha$ fixed. In general, the test has decent power and is capable of identifying a true explosion in the drift coefficient, except those causing a minuscule cumulative log-return and that are coupled with a large volatility burst. While the test has excellent ability to discover the largest drift bursts, which from a practical point of view are arguably also the most important, it is intriguing that we can uncover many of the smaller ones as well. At last, higher values of $h_{n}$ improve the rejection rate under the alternative, but the marginal gain of going from $h_{n} = 300$ to $h_{n} = 600$ is negligible. This suggests -- on the one hand -- that $h_{n}$ should not be too narrow, as it erodes the power, while -- on the other -- it should neither be too wide, as this creates a small size distortion.

Drift bursts in financial markets

Data

table[table omitted — 1,820 chars of source]

We apply the drift burst test statistic developed above to a comprehensive set of intraday tick data, covering a broad range of financial assets. We use trades and quotes with milli-second precision timestamps for futures contracts traded on the Chicago Mercantile Exchange (CME). We select the most actively traded futures contract for each of the main assets classes, namely the Euro FX for currencies (6E), Crude oil for energy (CL), the E-mini S&P500 for equities (ES), Gold for precious metals (GC), Corn for agricultural commodities (ZC), and the 10-Year Treasury Note for rates (ZN). These futures contracts are amongst the most liquid financial instruments in the world. To illustrate, the average daily notional volume traded in just a single ES contract on the CME is comparable to the trading volume of the entire U.S. cash equity market covering over 5,000 stocks traded across more than ten different exchanges.\footnote{See \url{https://batstrading.com/market_summary/} for daily U.S. equity market volume statistics.} The sample period is January 2012 -- December 2017, except we backdate ES to January 2010 in order to capture the May 2010 flash crash. While the CME is open nearly all day, we restrict attention to the more liquid European and U.S. trading sessions: from 01:00 -- 15:15 Chicago time or 07:00 -- 21:15 London time. The only exception is Corn, where we use data from 08:30 -- 13:20 Chicago time. Outside of these hours, trading is minimal in this contract. Table (ref) provides informative summary statistics of the data.

Implementation of test

We construct for each series a mid-quote as the average of the best bid and offer available at any point in time. A quote update is retained if the mid-quote changes. The remaining data are pre-averaged with $k_{n} = 3$.\footnote{The Online Appendix contains the empirical analysis with $k_{n} = 1$ (i.e., no pre-averaging), 5 and 10. The results are broadly speaking in line with those reported here.} We then calculate the drift burst test statistic on a regular five-second grid and only include values that are preceded by a mid-quote revision. As our primary interest is to identify short-lived drift bursts, we continue with a five-minute bandwidth for the drift. We base the spot volatility on a 25-minute bandwidth with the Parzen kernel and $L_{n} = 2(k_{n}-1) + 10$ lags for the HAC-correction. As in the simulations, a left-sided exponential kernel is adopted: $K(x) = \exp(-|x|)$, for $x \leq 0$.\footnote{In practice, a backward-looking kernel is preferred to allow for real-time updating. Moreover, it enchances the power for testing the drift burst hypothesis. The intuition is that after a drift burst, the pre-dominant tendency of the price to reverse combined with a high and persistent level of volatility, delivers a lower value of the test statistic if a two-sided kernel is employed.}

table[table omitted — 1,653 chars of source]

Table (ref) reports selected descriptive measures of the calculated drift burst $t$-statistics over the full sample. Judging by the standard deviation and kurtosis of the test statistic, the distribution is close to standard normal, as consistent with the asymptotic theory under the null of no drift burst. This is remarkable, because the test is applied to relatively short intraday intervals across a wide range of asset classes and is therefore exposed to substantial changes in liquidity conditions, diurnal effects, or to futures contracts where the minimum price increment -- and hence the microstructure noise -- is relatively large (e.g. ZN). Concentrating on the tails, we identify a large number of drift bursts.\footnote{To account for the rolling calculation of the test statistic and avoid double counting of events, we allow for at most one drift burst to be established over any five-minute window at which the test statistic attains a local extremum and exceeds a set critical value.} At a critical value of 4.5, for instance, there are 202 drift bursts in the E-mini S&P500 futures, or about one every two weeks. They are more prevalent in Euro FX, Gold, and Oil contracts but less frequent in the Treasury and Corn futures. The number of expected false positives (computed as described in Appendix (ref)) is practically zero.

figure[figure omitted — 787 chars of source]

Panel A of Figure (ref) indicates the location of drift bursts for the various securities, while Panel B reports the pooled monthly counts. The results lend support to the perception that flash crashes are increasing over time. A time trend is positive for each asset but lacks statistical significance for some due to the limited number of observations. The pooled results, however, indicate a statistically significant time trend of 5--10% per annum in the number of identified drift bursts with a regression $t$-statistic of 2.5 for both scenarios drawn in Panel B.

figure[figure omitted — 987 chars of source]

Figure (ref) shows some examples of single-asset drift bursts identified by the test. A multi-asset drift burst is presented in Figure (ref), which plots the evolution of the six asset prices during the Twitter hoax flash crash of April 23, 2013. The figures show that drift bursts are a stylized feature of the price process and, in some instances, systemic to the market. It is evident that neither jumps nor volatility are driving the price dynamics observed in these examples. The drift burst hypothesis is a plausible alternative to model the data.

figure[figure omitted — 1,584 chars of source]

Reversion after the drift burst

We now investigate in more detail whether the mean reversion experienced during the flash crashes in Figure (ref) and Figure (ref) is more broadly associated with the price dynamics of a drift burst. Let $\{ t_{j} \}$ denote the set of time points, where drift bursts are identified. The start of a drift burst is set to five minutes before $t_{j}$.\footnote{The 5-minute frequency is arbitrary, but often used in practice and consistent with our bandwidth. As a robustness check, we also applied an endogenous event window, where the duration of the drift burst is defined relative to the latest point in time prior to $t_{j}$, where the absolute value of the $t$-statistic is below one. The results are in line with those we report here and are available at request.} We sample the mid-quote process at this frequency pre- and post-drift burst and calculate:

equation[equation omitted — 161 chars of source]

where $R_{t_{j}}^{-}$ is the five-minute log-return during the $j$th drift burst, while $R_{t_{j}}^{+}$ is the corresponding post-drift burst log-return. In Figure (ref), we plot $R_{t_{j}}^{+}$ against $R_{t_{j}}^{-}$ pooled across the various asset markets. We observe that drift bursts can be associated with both positive and negative returns, but most of them are reversals and the percentage of “gradual jumps” is roughly one third.

To gauge the magnitude of the reversion and evaluate whether drift bursts are subject to short-term return predictability, we run the following regression:

equation[equation omitted — 89 chars of source]

A value of $b$ different from zero indicates predictability conditional on a drift burst, and $b<0$ means a drift burst tends to be followed by a retracement of the price.

figure[figure omitted — 610 chars of source]

In Table (ref), we show the regression results for each futures contract separately and with all assets pooled together. We find a strong mean reversion with estimates of $b$ that are negative and highly significant (the intercept estimate is insignificant and omitted). This is consistent across asset classes, also for those with a relatively small number of observations. The regression $R^{2}$ tends to lie in the 25% - 40% range, with a few exceptions, and is about 30% on average. This suggests a substantial predictive power of the pre-drift burst return to forecast the post-drift burst return. The fraction of reversals, as measured by counting the relative occurrence where $R_{t_{j}}^{+}$ and $R_{t_{j}}^{-}$ are of opposite sign, rarely drops below 65%. As a robustness check, we remove the top decile of the most significant drift bursts and rerun the regression. The outcome is in the right-hand side of the table. As expected, the regression $R^{2}$ drops but the finding remains overly evident in the data with frequent reversals and significant mean reversion. The critical value is set to 4.0 and 4.5 here, but the results do not change much for larger and more conservative values.

table[table omitted — 3,036 chars of source]

Asymmetric reversal, trading volume and liquidity

\citet*{grossman-miller:88a} suggest a large drift-to-volatility can be caused by exogenous demand for immediacy. In this framework, the return reversal represents a premium paid to market makers supplying liquidity against one-sided order flow \citep*[e.g.][also shows that returns on short-term reversal strategies can be thought of as a liquidity signature]{nagel:12a}. The empirical results in Section (ref) align with this interpretation. Their model, however, is symmetric in the order imbalance, as expressed by Eq. (ref).

huang-wang:09a show how trading imbalances arise endogenously due to costly market presence. Such a mechanism always leads to selling pressure, which tends to attenuate rallies and exacerbate sell-offs, resulting in market crashes absent news on fundamentals (Result 1 -- 3 in their paper). As in grossman-miller:88a, the model induces negative serial correlation in observed returns (Result 4a), but it is asymmetric. Negative returns exhibit stronger correlation than positive returns (Result 4b); volatility is higher during a negative return and high volume event (Result 5); abnormal volume implies larger future returns and stronger reversion (Result 6); and negative returns accompanied by high volume exhibit stronger reversals (Result 7). We here conduct an empirical assessment to check whether our sample of pre- and post-drift burst returns are consistent with these testable implications. We restrict attention to the E-mini S&P500 futures contract ES \citep*[][also base their analysis on a stock paying a risky dividend]{huang-wang:09a}. The results for other assets are available in the Online Appendix.

We construct a $5$-minute grid as above. We then compute for each interval $j$:

itemize$R_{t_{j}}^{-}$ and $V_{t_{j}}^{-}$: the 5-minute log-return and traded volume,\footnote{We construct a normalized gross trading volume measure, which is defined by taking the gross trading volume in notional value and normalizing it by the average trading rate at that time of the day. This allows to depurate our volume series for the pronounced intraday swings in trading intensity, which makes the analysis more robust to time-of-the-day effects. Still, the results based on raw gross trading volume are broadly consistent with what we report here.} • $R_{t_{j}}^{+}$: the subsequent 5-minute log-return, • $\rho$: correlation between $R_{t_{j}}^{-}$ and $R_{t_{j}}^{+}$, • $\sigma_{p} = \sqrt{(R_{t_{j}}^{-})^2 + (R_{t_{j}}^{+})^2}$: a volatility measure.
table[table omitted — 2,561 chars of source]

huang-wang:09a report $R_{t_{j}}^{-}$, $R_{t_{j}}^{+}$, $V_{t_{j}}^{-}$, $\rho$ and $\sigma_p$ based on simulated data.\footnote{They employ the notation $R_{1/2}$ for $R_{t_{j}}^{-}$, $R_{1}$ for $R_{t_{j}}^{+}$, and $V_{1/2}$ for $V_{t_{j}}^{-}$.} In contrast, Table (ref) -- shaped as Table 1 in that paper for comparison -- is computed from actual market returns. The full sample results in Panel A are calculated from all available 5-minute intervals. We observe a tiny negative return autocorrelation, which is slightly stronger if the preceding return is less than average or volume is above average. Volatility increases marginally with volume. The predicted effects are thus present, but extremely weak.

We now condition on $t_{j}$ being a drift burst time (with critical value set to 4.0), allowing at most for one significant event per day. Grouping by the sign of $R_{t_{j}}^{-}$, we are left with 253 negative and 208 positive drift bursts. In Panel B -- C of Table (ref), we tabulate the corresponding averages. The results are striking. A drift burst of either sign is associated with high volatility, large volume, and substantial negative serial correlation. The typical negative drift burst entails a price move of $-37.39$bps compared to $+36.41$bps for a positive one.\footnote{These price changes are far out in the tails of the return distribution. A $-37.39$bps drop in five minutes translates to a loss of $-31.41\%$ in seven hours. In our sample it represents a tenfold increase in magnitude compared to a typical 5-minute negative stock market return and corresponds to a four standard deviations draw, or about 1 in 25,000.} The subsequent $5$-minute log-return is $+8.01$bps and $-5.44$bps, on average. The serial correlation is $-0.7958$ ($-0.2573$) for negative (positive) drift bursts. Looking further at price drops, we split the sample into negative drift bursts with a smaller and larger return or volume than that of an average event. The negative serial correlation is most pronounced in drops that are below average (more negative) or accompanied by high trading volume. These effects are not present in positive drift bursts. At last, volatility is markedly higher for tail drift bursts (in either direction) and those with large volume.

table[table omitted — 1,338 chars of source]

Table (ref) reports conditional averages of the post-drift burst return $R_{t_{j}}^{+}$, which is double-sorted by pre-drift burst return $R_{t_{j}}^{-}$ and volume $V_{t_{j}}^{-}$. As before, the table is shaped as Table 2 in huang-wang:09a to facilitate comparison. It reinforces that the reversion is stronger if the price change or volume is large. The results are again more pronounced for negative drift bursts.

As an alternative way to describe the double-conditioning of Table (ref), we estimate the forecasting equation of campbell-grossman-wang:93a:

equation[equation omitted — 123 chars of source]

where an interaction term $R_{t_{j}}^{-}V_{t_{j}}^{-}$ is added to Eq. (ref). This approach is suggested by \citet*{huang-wang:09a} in order to assess if the return predictability is driven by trading volume. The parameter estimates of Eq. (ref) are displayed in Table (ref) along with $t$-statistics (in parenthesis) based on newey-west:87a-robust standard errors. Once we control for volume, the serial correlation observed during a negative drift burst is largely subsumed by the interaction term and the coefficient on $R_{t_{j}}^{-}$ is heavily reduced, albeit it remains significant for drops below average. This does not occur for positive drift bursts. We again split the sample in half according to the size of the pre-drift burst return and -- as predicted by \citet*{huang-wang:09a} -- we note that the volume-return variable is more important for large negative drops. On the other hand, for positive drift bursts it is largely irrelevant in explaining the post-drift burst return.

To summarize, Table (ref) -- (ref) confirm the theoretical predictions made by \citet*{grossman-miller:88a} and \citet*{huang-wang:09a}. The reversion in drift bursts is consistent with market makers absorbing large orders and charging a fee for this service. The asymmetry between negative and positive drift bursts is a symptom of endogenous demand for liquidity due to costly market participation. In the latter case, we confirm several testable implications on trading volume (regarded as a proxy of order imbalance), in that compensation increases with volume, which is larger during negative drift bursts. Overall, our findings suggest the regular occurrence of “flash crashes” documented here is consistent with existing theories of liquidity provision.

table[table omitted — 2,431 chars of source]

Conclusion

The drift burst hypothesis is proposed as a theoretical framework for the modeling of distinct and sustained trends in the price paths of financial assets. We show how drift bursts -- defined as a short-lived locally explosive drift coefficient -- can be embedded into standard continuous-time models and demonstrate that the arbitrage-free property is preserved if the volatility co-explodes during a drift burst, something we provide strong empirical support for. Applying a novel methodology for drift burst identification to a comprehensive set of tick data covering six major asset classes, we deliver unprecedented insights into potentially disruptive but poorly understood events, such as flash crashes. In contrast to the existing literature, which has mostly regarded these events as market glitches, we show that they are instead a regular, stylized fact in the markets, whose dynamic features match theoretical predictions from the market microstructure literature of liquidity provision.

The paper contributes towards a better understanding of the microstructure dynamics of financial markets and can help to inform the regulatory policy agenda going forward by shedding light on a number of important questions: Who triggers a flash crash? Who supplies liquidity during the event? What is the role played by high-frequency traders? And what, if anything, can be done to prevent flash crashes in the future?