EconBase
← Back to paper

Real-Time Detection of Local No-Arbitrage Violations

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

69,147 characters · 18 sections · 68 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Real-Time Detection of Local No-Arbitrage Violations

\thispagestyle{empty} \abstract{ This paper focuses on the task of detecting local episodes involving violation of the standard It\^o semimartingale assumption for financial asset prices in real time that might induce arbitrage opportunities. Our proposed detectors, defined as stopping rules, are applied sequentially to continually incoming high-frequency data. We show that they are asymptotically exponentially distributed in the absence of It\^o semimartingale violations. On the other hand, when a violation occurs, we can achieve immediate detection under infill asymptotics. A Monte Carlo study demonstrates that the asymptotic results provide a good approximation to the finite-sample behavior of the sequential detectors. An empirical application to S&P 500 index futures data corroborates the effectiveness of our detectors in swiftly identifying the emergence of an extreme return persistence episode in real time.\\}

JEL classification: C12, C53, G10, G17

Keywords: asset price, high-frequency data, It\^o semimartingale violation, real-time detection, stopping rule

\pagenumbering{arabic} {6.5mm}

Introduction

The no-arbitrage principle is central to modern asset pricing theory (ross1976arbitrage). delbaen1994general demonstrate, within a frictionless setting, that arbitrage opportunities are precluded if and only if the price process constitutes a semimartingale. The standard class of no-arbitrage price processes in financial economics is the It\^o semimartingale, which is a semimartingale with characteristics that are absolutely continuous in time. However, recent work document episodic violations of the It\^o semimartingale assumption that, absent transaction costs, might induce arbitrage opportunities. A prominent example is the gradual jump identified by barndorff2009realized and further studied by christensen2014fact. It occurs when an apparent return jump, identified from lower frequency data, instead reflects a strongly drifting, yet (near) continuous price path, when observed at higher frequencies. A related phenomenon is the so-called flash crash, where a sudden collapse in price is reversed rapidly; see, e.g., the work on the May 2010 events in the S&P 500 e-mini futures market by kirilenko2017flash and menkveld2019flash.\footnote{In the presence of trading costs and uncertainty surrounding the data generating process, such episodes are not necessarily true arbitrage opportunities, see, e.g., the discussion in andersen2021volatility.}

The “explosive” price paths characterizing such events are unlikely to be generated by an It\^o semimartingale. To accommodate these occurrences, alternative models have been developed for violation episodes, including the drift burst model proposed by christensen2020drift and the persistent noise by andersen2021volatility, as a stochastic generalization of the former. These models contain a parameter $\tau$, indicating a random point in time located within a neighborhood in which an It\^o semimartingale violation occurs. Such episodes typically involve turbulent market conditions with extreme realized volatility, raising concerns of evaporating liquidity and general market malfunction. Moreover, our standard measures for monitoring return volatility are potentially subject to large biases, when the semimartingale assumption is violated. Consequently, identification of the onset, indicated by $\tau$, as well as duration of the extreme return drift episode is of great interest for regulators, industry practitioners and academics. This is the goal of this paper.

The detection problem can be addressed from two distinct perspectives. The more common is the offline approach where the researcher observes the full dataset, and then conducts a “one-shot” procedure to identify whether and when violations have occurred. The vast literature on ex-post detection of structural breaks in macroeconomic and financial time series data, including andrews1993tests, bai1996testing, bai1998estimating, and elliott2014pre, falls within this category. More recently, bucher2017nonparametric develop inference procedures for a change point in the jump intensity parameter for a L\'evy process with high-frequency data following this approach.

A more challenging and, for practitioners, investors and regulators, arguably more relevant perspective is the real-time setting, where data arrive continuously, and one wishes to detect the violation in a timely manner. This objective has some resemblance to that of the recent now-casting literature in macroeconomics, where the aim is to update the assessment of the current and future state of the economy as new data are received in real time. For early initial work on a formal statistical framework in this setting, see, e.g., evans2005we and giannone2008nowcasting, while corresponding work utilizing financial data can be found in, e.g., andreou2013should and banbura2013now. Another related macro-econometric literature is initiated by diba1988explosive regarding the detection of macroeconomic bubbles and crises. This methodology is extended by phillips2011explosive using a recursive procedure based on right-tailed unit root tests to detect and locate the origin and terminal dates for bubbles. A series of subsequent studies provide further theoretical modifications, e.g., phillips2011dating, phillips2014specification, and phillips2015testingb, while empirical applications are provided by phillips2018financial and phillips2015testinga.

Our objective is closely aligned with the latter real-time or online detection procedures. From a technical perspective, our approach is rooted in the statistic literature on the sequential detection problem for a change point. This literature typically deals with the mean of i.i.d samples or with the drift within a continuous-time model featuring both a drift and a scaled standard Brownian motion component. The literature goes back to, at least, Abraham Wald's sequential analysis. His work inspired the widely-known CUSUM rule by page1954continuous and subsequently the Bayesian rule developed by shiryaev1963optimum and roberts1966comparison.\footnote{For more studies and optimality results, see, e.g., moustakides1986optimal, ritov1990decision, shiryaev1996minimax, and moustakides2004optimality. See also srivastava1993comparison for a comparison and figueroa2019change for the case of a L\'evy process. On the financial econometrics side, andreou2006monitoring accommodates the CUSUM test for strongly dependent financial time series data, and andreou2008quality applies real-time detection procedures to the structural parameters of credit risk models by monitoring the associated Radon-Nikodym derivative process.} We follow the former rule, but also deviate substantially, because that statistic requires knowledge of the alternative measure after the change, which is not a natural assumption in our setting. Specifically, our detector is based on the generalized likelihood ratio (GLR) statistic, which profiles out the unknown alternative parameter. For studies on this GLR-CUSUM procedure under an i.i.d setting, see siegmund1995using, pollak1975approximations, lai1979nonlinear, and lai1995sequential, among others.

Our paper differs from the aforementioned work in two key aspects. First, we rely on infill asymptotics and exploit the feature of accessing an asymptotically increasing number of observations of the process locally. This helps us formalize the notion of rapid detection of local It\^o semimartingale violations. These deviations from the semimartingale dynamics are only local in nature, unlike the earlier sequential detection literature which deals with detection of a permanent change. Second, we abandon the use of size versus power to characterize the properties of our detection procedure because we, by necessity, must apply our tests sequentially. If a test with fixed critical value is applied sequentially on a constant flow of newly arriving data, the null will inevitably be rejected repeatedly, and the traditional type I error literally explodes (with probability approaching one). To address this issue, we follow the sequential detection literature to design and evaluate the performance of our procedure using alternative metrics: the average run length (ARL) and false detection rate (FDR) or, more comprehensively, an asymptotic probability bound on the false detection (BFD) rate versus a corresponding bound on the detection delay (BDD) (see lorden1971procedures and lai1995sequential). Specifically, for the null probability measure, when there is no It\^o semimartingale violation, ARL measures the expected sample size until the first false detection, while FDR measures the probability of a false detection within a given period (and we refer to this period as BFD). For the alternative, BDD provides a probability bound on the number of observations before we achieve successful detection following a violation.\footnote{The sequential detection literature typically measures the behavior under the alternative via the {\it Expected Detection Delay}, or EDD. Characterizing the behavior of the latter is more complicated within our high-frequency setting. Thus, we focus on the probability estimate BDD instead.} Consequently, one strives to develop a detector that, conditional on a reasonably large ARL/small FDR, achieves the smallest possible BDD. An alternative would be to develop a time-varying boundary function -- in contrast to a constant threshold -- to control test size uniformly, see, e.g., chu1996monitoring. Unfortunately, in a real-time sequential setting, such procedures struggle with the detection of structural changes that arrive late within the period. There is work seeking to alleviate this drawback, e.g., leisch2000monitoring, horvath2004monitoring, aue2004delay, aue2006change, and horvath2007sequential. However, these procedures still tend to generate a significant detection delay, and we do not pursue this direction here.

As noted, our objective is to design statistical devices that detect local It\^o semimartingale violations swiftly and reliably after their occurrence, using continually incoming high-frequency data. Towards this end, we propose GLR-CUSUM type detectors as stopping times based on an estimated Brownian motion component of the latent asset price, which we recover from high-frequency return data along with short-dated options. We first establish the accuracy of this Brownian motion estimator, exploiting the option-based spot volatility estimate of todorov2019nonparametric to standardize the high-frequency returns. This approach avoids complications stemming from the inconsistency of standard volatility measurements from returns, caused by local It\^o semimartingale violations, and it exploits the efficiency offered by option data for recovering spot volatility. Next, following the asymptotic setting in the sequential detection literature, we show that our detector is approximately exponentially distributed under the null (Theorem (ref)) and develop a bound for the BDD under the alternative (Theorem (ref)). The former result implies that the ARL is the mean and FDR a percentile of the distribution. In turn, the latter result implies that our detectors achieve immediate detection of It\^o semimartingale violations under infill asymptotics. A Monte Carlo experiment finds that our theoretical results provide good guidance for the finite-sample properties of the detectors under a realistically calibrated simulation setting.

Finally, we apply our detection methods empirically to one-minute S&P 500 equity (SPX) index futures data to investigate the pervasiveness of the It\^o semimartingale violation phenomenon and to assess whether our detector provides timely alerts regarding the failures. We further extend the procedure to obtain an identification rule for the duration of the violation. It is defined as the union of the time intervals on which our GLR-CUSUM statistic exceeds the chosen threshold. With a threshold specification for our detector leading to about 6% daily FDR, we find these violations to be common in the SPX data --- with a little more than 1,000 such violations across 3,500 days. That said, more than half of the violations last for less than $10$ minutes, and only a small proportion exceeds $20$ minutes. Visual inspection suggests that, in most instances, our procedure detects such incidents within just a few minutes of their occurrence.

The remainder of the paper is organized as follows. In Section (ref) we describe the setting and formulate the problem. In Section (ref) we introduce our real-time detectors and study their asymptotic properties. Next, we carry out a Monte Carlo study in Section (ref) to illustrate the finite-sample performances of the detectors, and we then use them in an empirical application in Section (ref). Section (ref) concludes.

Setup

The observed log-price process, $Y_t$, is defined on a filtered probability space $\left(\Omega, \mathcal{F}, (\mathcal{F}_t)_{t \geq 0}, \mathbb{P}\right)$. We assume that it may be decomposed as,

align[align omitted — 70 chars of source]

where $X$ represents an underlying efficient and arbitrage-free log-price. In addition, the observed price is contaminated by the noise component $H$, which is the source of brief episodes characterized by extreme return persistence, such as the so-called gradual jumps and flash crashes. Sections (ref) and (ref) introduce our fairly standard specification for the observation scheme and $X$, while Section (ref) provides a detailed account of the dynamics for the noise component $H$.

The Observation Scheme

We assume the price is observed over a given time interval $[0, T]$ at equidistant times, $t_i = i\Delta_n$, for all $i = 0, 1, \dots, nT$, where $\Delta_n = 1/n$ is the increment length. Let the $i$-th high-frequency log-return be denoted by,

align[align omitted — 70 chars of source]

Without loss of generality, we normalize the trading day to unity. Hence, $T$ refers to the number of trading days, and $n = 1/\Delta_n$ (assumed integer) is the number of observations per (trading) day. The time span of the data, $T$, is assumed fixed throughout.

Following the standard infill asymptotic framework, we let the sampling frequency, $\Delta_n$, shrink to $0$ (or, equivalently, $n \to \infty$). This equidistant sampling scheme can readily be relaxed to a non-equidistant one, requiring $\max_{\{i\in{1,\dots,n T\}}}(t_i - t_{i-1}) \to 0$.

The efficient log-price $X$

The efficient log-price process $X$ is an It\^o semimartingale,

align[align omitted — 173 chars of source]

where the initial value $X_0$ is $\mathcal{F}_0$-measurable, the drift $b_t$ takes value in $\mathbb{R}$, $W = (W_t)_{t \geq 0}$ is a one-dimensional standard Wiener process, $\delta: \, \mathbb{R}_{+}\times\mathbb{R}\mapsto\mathbb{R}$ \, is a predictable mapping, and $\mu$ is a Poisson random measure on $\mathbb{R}_{+}\times\mathbb{R}$ with predictable compensator (or intensity measure) $\nu(dt,dx)=dt\otimesF(dx)$.

We impose the following assumption on the efficient price process.

assumptionThere exists arbitrary small $\mathbb{T}>0$ such that for the process $X$ in equation ((ref)), we have, for $t\in[0,\mathbb{T}]$ and $\forall p\geq 1$, \begin{equation} \mathbb{E}_0|b_t| \,+\, \mathbb{E}_0\left(\int_{\mathbb{R}}|\delta(t,x)|F(dx)\right) \,+\, \mathbb{E}_0|\sigma_t|^p < C_0(p), \end{equation} where $C_0(p)$ is $\mathcal{F}_0$-adapted random variable that depends on $p$, and further \begin{equation} \mathbb{E}_0|\sigma_t-\sigma_s|^2 \leq C_0 \, |t-s|, for $s,t\in[0,\mathbb{T}]$, \end{equation} where $C_0$ is $\mathcal{F}_0$-adapted random variable.

A few comments about this assumption are warranted. First, since our testing procedure concerns the asset price dynamics during an It\^o semimartingale violation initiated close to time zero, the assumption focuses on the behavior of $X$ in the vicinity of $0$. Second, the first part of Assumption (ref) involves the existence of conditional moments, but since $\mathbb{T}>0$ can be arbitrary small, these moment conditions are relatively weak. Third, the jump part of $X$ is modeled as an integral against a Poisson measure and, therefore, we restrict attention only to finite variation jumps. Finally, the “smoothness in expectation” condition for $\sigma$ will be satisfied whenever $\sigma$ is an It\^o semimartingale, which is the standard way of modeling stochastic volatility. We note, however, that this condition will not hold for the class of rough stochastic volatility models.\footnote{For this class of models, the right-hand side of the second equation in Assumption (ref) will be replaced by $C_0 \, |t-s|^{2H}$, for $0<H<1/2$ capturing the degree of roughness of the volatility path.}

remarkWe refer to the model ((ref)) for the efficient asset price $X$, along with Assumption (ref), as the standard It\^o semimartingale model since it embeds most specifications used in current existing work. We note, however, that there is a gap between a semimartingale, required to preclude the existence of arbitrage opportunities, and the standard It\^o semimartingale considered above. The gap includes models in which the semimartingale characteristics are not absolutely continuous in time as well as models with infinite variation jumps. We do not consider those in our analysis.
remarkWe can readily accommodate an empirically relevant extension of the standard It\^o semimartingale model to the case, where the efficient price path features discontinuities triggered by economic announcements at pre-specified points in time. Since the announcement times are known a priori, one may remove the associated price increments from the analysis and proceed as if $X$ follows a standard It\^o semimartingale. To minimize notational complexity, we do not formally introduce this extension.

The persistent noise $H$

Our main objective is to develop inference tools to detect episodic violations of the It\^o semimartingale assumption, which is almost universally imposed in standard asset pricing theory. From a practical perspective, the occurrences of gradual jumps or, in more extreme manifestations, flash crashes, are of particular interest because the associated drift burst in the returns can induce severe biases in the high-frequency measurement of the (efficient) return variation. We include such short-lived episodes of extreme return persistence in our setup through the Persistent Noise (PN) model of andersen2021volatility, which is designed explicitly to accommodate these types of empirically observed phenomena. The PN model is a stochastically extended version of the Drift Burst (DB) model by christensen2020drift, and the real-time detection procedures developed below applies for either specification. We initially consider a simplified setting with only a single episode involving semimartingale violations.

The PN model initiates a violation episode at a random point, $\tau_n$, i.e., $H_t = 0$ for $t < \tau_n$ and in our analysis $\tau_n\downarrow 0$. At $\tau_n$, the efficient price may display a discrete jump, $\Delta X_{\tau_n} \ne 0$. If an efficient price jump occurs, it likely reflects the arrival of new information or unusual ongoing trading activity. To the extent this information is not common knowledge or not readily interpretable, a gap will emerge between the efficient price, reflecting rational processing of all relevant information, and the market price which is determined by the interaction of risk averse and merely partially informed agents.\footnote{For an in-depth discussion of this feature across a large cross-section of stocks, see AACH.} In other words, the efficient price may deviate from the market price, implying that the noise component is active, $H_{\tau_n} \neq 0$. In fact, if the information is not immediately observed or processed by market participants, we have $H_{\tau_n} = \Delta H_{\tau_n} = - \Delta X_{\tau_n}$, implying $\Delta Y_{\tau_n} = \Delta X_{\tau_n} - \Delta X_{\tau_n} = 0$. That is, the price reaction is delayed. An intermediate case is where the news generates a partial response. For example, if $H_{\tau_n} = - \Delta X_{\tau_n}/2$, the price jump underreacts to the new information. Such events may trigger a period of intense information acquisition and price discovery, inducing rapid price adjustment. After initiation at $\tau_n$, the future path of the noise component is modeled through a specific functional form, $g(t), t \in [\tau_n,\overline{\tau}]$. Finally, we ensure that the PN event terminates at some random point, $\overline{\tau} > \tau_{n\,}$, so that we have $H(t) = 0$ for $t > \overline{\tau}$.

We are now in position to formally introduce the PN model.

{\bf The Persistent Noise (PN) model}: For some non-negative sequence $\tau_n\downarrow 0$, we let,

align[align omitted — 134 chars of source]

where $\eta_{\tau_n}$ is an $\mathcal{F}_{\tau_n}$-adapted random variable, $f$ is a continuous and bounded function, and $g$ is given as,

align[align omitted — 235 chars of source]

for some $\mathcal{F}_{\tau_n}$-adapted random $\overline{\tau}>\tau_n$.

The strength of the local return drift in the vicinity of the PN initiation point, $\tau_n$, in equation ((ref)) increases as $\vartheta_{pn}$ takes on lower values, while the duration of the event is controlled by the realization of $\overline{\tau}$. We also note that $g(s)$ is zero, indicating no It\^o semimartingale violation, until $\tau_n$. Finally, $\eta_{\tau_n}$ allows for a random response to the events triggering the permanent noise component, which may or may not be associated with a jump in the efficient price.

To illustrate the dynamics of the PN model, Figure (ref) plots a simulated sample path for the efficient price $X$ (blue) across the 6.5 hour trading day along with the associated path for the observed price $Y$ (black). The efficient price path features a positive jump at $\tau_n = 0.25$ (at $97$ minute), but (informational) frictions prevent this from being recognized by market participants, so there is no instantaneous jump in the observed price. In terms of the model, we register a corresponding negative latent noise jump at $\tau_n$, namely $f(\Delta X_{\tau_n}, \eta_{\tau_n}) = - \Delta X_{\tau_n}$. This is a purely mechanical effect. If there is no observed price change, but the efficient price jumps, then the deviation between the two -- the noise term $H$ -- must offset the efficient price jump.

As noted, Figure (ref) is designed to replicate the gradual jump phenomenon. It is evident that this PN episode also may reflect the recovery from a so-called flash crash that hits its nadir at $\tau_n$. Thus, one may replicate the typical flash crash pattern by combining this “recovery phase” with an inverted version of the PN path prior to $\tau_n$, see andersen2021volatility for a more detailed discussion.

figure[figure omitted — 295 chars of source]

From a technical perspective, it is convenient to introduce a version of the DB model by christensen2020drift.\footnote{christensen2020drift allow for simultaneous volatility and drift bursts. We do not include this feature, as we find no systematic evidence for significant elevation in option-based volatility measures at such times.}

{\bf The Drift Burst (DB) model}: For some non-negative sequence $\tau_n\downarrow 0$, we have,

align[align omitted — 159 chars of source]

for a constant $c$ and $\tau_n \leq \overline{\tau}$. $\tau_n$ and $\overline{\tau}$ have the same interpretation as in the PN model.

The DB model may be viewed as a differential version of the PN model by observing that $\int_{t_1}^{t_2}(s-\tau)^{-\vartheta}ds = (1-\vartheta)^{-1}(t_2-\tau)^{1-\vartheta} - (1-\vartheta)^{-1}(t_1-\tau)^{1-\vartheta}$ for $0\leq \tau \leq t_1<t_2$. Therefore, the asymptotic analysis will be identical for the PN model and the DB model with $\vartheta = 1-\vartheta_{pn}$. We will exploit this equivalence in our theoretical arguments.

Henceforth, we denote the probability and expectation by $\mathbb{P}_{\tau_n}$ and $\mathbb{E}_{\tau_n}$ for the case where the observed log-price is generated as above, and by $\mathbb{P}_{\infty}$ and $\mathbb{E}_{\infty}$ if there are no It\^o semimartingale violations (that is, when $\tau_n = \infty$).

Real-Time Detection

In this section, we propose GLR-CUSUM type detectors as stopping rules for the local It\^o semimartingale violations illustrated above. Next, we develop their theoretical properties. This involves characterizing the accuracy for our estimate of the Brownian motion driving the price, followed by approximation results for the distribution of the stopping rules, leading to analytic formulas for ARL and FDR, and finally a probability result for BDD which, in turn, implies a probability bound on the speed of detection.

The GLR-CUSUM stopping rule

Our local It\^o semimartingale violation detector is built upon a local estimator of the Brownian motion driving the price, defined as,

align[align omitted — 210 chars of source]

for some $l_1,l_2\in\mathbb{N}$ such that $0 \leq l_1 < l_2 \leq nT$, and $\hat{\sigma}_{i\Delta_n}$ is an option-based nonparametric spot volatility estimator defined later. As detection of the semimartingale violation inevitably is subject to some delay, the option-based $\hat{\sigma}$ enables us to avoid the bias in volatility estimation induced by the bursting drift following $\tau_n$ (see andersen2021volatility).

Our CUSUM-type detector, employing the generalized likelihood ratio (GLR) statistic, is now defined by the stopping rule,

align[align omitted — 195 chars of source]

where $w_n$ controls the maximum window length for $\widehat{W}_{k,l}$, and $\xi$ is a chosen threshold.\footnote{The terminology of generalized likelihood ratio statistic here refers to (the square root of $2$ times) the likelihood of a Brownian motion with drift $\mu$ versus drift of $0$, $\mu \, W_{k,l} - \mu^2(l-k)\Delta_n/2$, with the unknown alternative $\mu$ being replaced by its maximum likelihood estimate $\widehat\mu = W_{k,l}/\sqrt{(l-k)\Delta_n}$.} In words, at time point $l$, we calculate the absolute value of the Brownian increment estimate $\widehat{W}_{k,l}$ over the interval $[k,l]$ (normalized by $1/\sqrt{(l-k)\Delta_n\,}$) for each point $k$ before $l$, and then take the maximum. If the maximum exceeds the threshold $\xi$, the detector reports an alarm, signaling detection of an It\^o semimartingale violation.

Recovering the Driving Brownian Motion for the Price

Our theoretical results concerning detection of local semimartingale violations hinge on the accuracy of the estimator $\widehat{\sigma}_{t}$. The following assumption is sufficient.

assumptionFor some arbitrary small $\mathbb{T}>0$, we have $\inf_{t\in[0,\mathbb{T}]}\sigma_t>0$ and \begin{equation} \mathbb{E}_{0\,} |\widehat{\sigma}_t - \sigma_t|^2 \leq C_0 \, \delta_n^2, t\in[0,\mathbb{T}], \end{equation} for some $\mathcal{F}_0$-adapted positive random variable $C_0$, and a deterministic sequence $\delta_n\rightarrow 0$. In addition, we have $\widehat{\sigma}_t>C_0/l_n$, for $t\in[0,\mathbb{T}]$, some $\mathcal{F}_0$-adapted positive random variable $C_0$, and $l_n = \log(1/\delta_n)$.

Several features of the above assumption are noteworthy. First, we impose conditions on the volatility estimator only in the vicinity of zero. Second, we require that $\sigma_t$ is non-vanishing in a neighborhood of zero. This is important for the behavior of the volatility estimator, but also for the overall validity of our detection procedure, as the diffusive component of $X$ drives our asymptotic results. Third, the deterministic sequence $\delta_n$ captures the rate of convergence of the option-based volatility estimator. This rate is determined by the mesh of the strike grid of the available options as well as their tenor, see todorov2019nonparametric.

The above assumption is natural in the absence of a local semimartingale violation. Its validity is less clear if there is a violation of the type discussed in the previous section in the vicinity of $t=0$. However, note that a drift burst or persistent noise incident will not affect the norm of the conditional characteristic function of the price increment which forms the basis for the estimator of todorov2019nonparametric. Hence, the consistency of our spot volatility estimator should not be affected by the local semimartingale violation. That said, the quality of the option-based estimator might worsen during such periods due to poor option data quality. To guard against this possibility, we impose some filters on the option data in our application.

To facilitate the statement of our next result, we introduce the following notation,

align*[align* omitted — 296 chars of source]

for some $l_1,l_2\in\mathbb{N}$ such that $0 \leq l_1 < l_2 \leq nT$. The issue in this section is how closely $\widehat{W}_{l_1,\,l_2}$ approximates $ \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{W}}]{\kern-.3pt\bigwedge\kern-.3pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.4ex}} \stackon[1pt]{W}{\scalebox{-1}{\tmpbox}} _{l_1,\,l_2}$. The key result is provided by the following proposition.

propositionSuppose Assumptions (ref) and (ref) hold and $H\equiv 0$. Let $k_n\to\infty$ and $k_n\Delta_n\to 0$ as $\Delta_n \to 0$. Let $0\leq l_1^n<l_2^n$ be such that $l_1^n\Delta_n\rightarrow 0$ as $\Delta_n\rightarrow 0$ and, as before, $l_n = \log(1/\delta_n)$. Then, for a positive $\xi>0$, we have, \begin{equation} \begin{split} &\mathbb{P}_{0}\left(\max_{|l_2^n-l_1^n|\leq k_n} \frac{ |\widehat{W}_{l_1^n,l_2^n} - \savestack{\tmpbox}{\stretchto{ \scaleto{ \scalerel*[\widthof{\ensuremath{W}}]{\kern-.3pt\bigwedge\kern-.3pt} {\rule[-\textheight/2]{1ex}{\textheight}} }{\textheight} }{0.4ex}} \stackon[1pt]{W}{\scalebox{-1}{\tmpbox}} _{l_1^n,l_2^n}| }{\sqrt{(l_2^n-l_1^n) \, \Delta_n}} \, > \, \xi\right) \\ & \leq C_0 \, k_n \, \left(\frac{ \sqrt{k_n} \, l_n \, \delta_n + \sqrt{k_n \, \Delta_n} }{\xi} \,+\, \frac{k_n \, \Delta_n}{ 1\wedge \xi^2 } \,+\, \Delta_n^{1-\varpi}\right) \, , \end{split} \end{equation} for some $\mathcal{F}_0$-adapted positive random variable $C_0$ that does not depend on $k_n$, $\Delta_n$ or $\xi$.

Average Run Length (ARL)

We first analyze the behavior of the stopping time $\mathbf{N}^{\rm w}(\xi)$ under the null measure, $\mathbb{P}_{\infty}$, where there is no It\^o semimartingale violation, i.e., $\tau_n = \infty$. More specifically, our interest is in evaluating the ARL, formally defined as $\mathbb{E}_{\infty}[\, \mathbf{N}^{\rm w}(\xi) \, ]$. We show in Theorem (ref) below that $\mathbf{N}^{\rm w}(\xi)$, under $\mathbb{P}_{\infty}$, is approximately exponentially distributed, leading to an approximation result for ARL. Theorem (ref) follows by adapting the theoretical framework in siegmund1995using, based on the random fields analysis in siegmund1988approximate to our setting (see also yao1993boundary). We use the symbol “$x_n \sim y_n$” to denote “$x_n$ is asymptotic equivalent to $y_n$”, that is, $\lim_{n\to\infty} x_n/y_n = 1$.

theoremDenote the standard normal CDF and PDF by $\Phi$ and $\phi$. Let $w_n$ be such that $w_n \sim a_n \, \xi^2$, with $a_n$ being a deterministic sequence converging to $a\in (0,\infty]$. Suppose the conditions in Proposition (ref) hold. Then, if $(l_n \, \delta_n + \sqrt{\Delta_{n\,}})/\, [\xi \, \phi(\xi)] \rightarrow 0$, as $\xi\to\infty$, $\mathbf{N}^{\rm w}(\xi)$ will be asymptotically exponentially distributed with expectation, \begin{align} \mathbb{E}_{\infty}[\, \mathbf{N}^{\rm w}(\xi) \, ] \sim \, \frac{1}{ D_a \, \xi \, \phi(\xi) } \,\, . \end{align} Here $D_a$ is a positive constant depending on $a$ such that $D_a\toD$, as $a\to\infty$, where $D = \int_0^\infty x \, \nu^2(x) \, dx$ \, with \, $\nu(x) = 2x^{-2} \exp\left[-2\sum_{n=1}^{\infty}n^{-1} \, \Phi(-x\sqrt{n}/2)\right]$ \, for $x>0$.

The proof of Theorem (ref) is provided in Appendix (ref). Following siegmund1995using, it is useful for numerical evaluation to apply the approximation $\nu(x) = \exp(-\rho x) + o(x^2)$ for $x \to 0$, where the value of the constant $\rho$ is about $0.583$. Based on this, we can numerically determine $D$ to take the value $0.735$.

We desire for the ARL, under the null, to be as large as possible or, equivalently, $\xi$ to be as large as possible. The optimal such $\xi$, up to a log term and for a given $w_n$, will be $\xi$ such that $\phi(\xi) \sim (\delta_n \, l_n \vee \sqrt{\Delta_n})$. It is reasonable to assume that the size of the error in estimating spot volatility from options is much smaller than the high-frequency approximation error, that is, $\delta_n \, l_n/\sqrt{\Delta_n} \rightarrow 0$. Under this condition and the indicated choice of $\xi$, we obtain from Theorem (ref) that,

equation[equation omitted — 162 chars of source]

We note that this ARL value can be obtained with $w_n$ (the window size) taking values within a wide range --- from an order of $\xi^2$ through using (almost) all available observations on a given day (retaining $w_n \, \Delta_n \rightarrow 0$). For the case of $w_n/ \, \xi^2 \rightarrow \infty$, we may denote our detector $\mathbf{N}(\xi)$, and replace $D_a$ by $D$, as $D_a \to D$ with $a \to \infty$.\footnote{For the standard i.i.d.\ case, more specific guidelines in selecting $w_n$ can be found in lai1995sequential.}

False Detection Rate (FDR)

Given the nature of our sequential detection procedure, we cannot rely on the regular notion of test size to control the rejection rate under the null hypothesis. The preceding subsection instead focuses on the concept of ARL, establishing this quantity (the ARL) as the mean value of the asymptotic distribution for stopping time $\mathbf{N}^{\rm w}(\xi)$ under $\mathbb{P}_{\infty}$. However, this distributional result has wider implications. For example, it allows us to control the false detection rate (FDR) as,

align[align omitted — 119 chars of source]

where $\ell_n$ indicates an asymptotically increasing sequence of observations, while $\alpha$ is the bound for the FDR -- the probability of triggering (at least) one false alarm after observing no more than $\ell_n$ price increments. If $\alpha$ is asymptotically shrinking, then we refer to $\ell_n$ as the bound (in probability) on false detection or BFD.

The FDR is often viewed as a more robust measure than ARL. Typically, the objective is to ensure an expected long duration (ARL) before a false alarm is triggered. The problem is that a stopping rule may feature a large ARL and still retain a high probability of false alarms within short periods. Although this does not apply in our case, we still proceed with a result concerning the limiting behavior of FDR under the null, since it provides a guideline for selecting a threshold in empirical applications. Moreover, it renders the Monte Carlo studies more computationally tractable, as we avoid simulating an excessively long sample to assess the ARL, especially when we experiment with large thresholds. Below, we use the symbol “$x_n \lesssim y_n$” to denote “$x_n$ is asymptotic equivalent to or less than $y_n$”, that is, $\lim_{n\to\infty} x_n/y_n \leq 1$.

corollarySuppose the conditions in Theorem (ref) hold and let $\ell_n$ be such that $\ell_n \, \xi \, \phi(\xi) \to 0$, as $\xi\to\infty$. Then, we have, \begin{align} \mathbb{P}_{\infty}[\, \mathbf{N}^{\rm w}(\xi) < \ell_{n\,}] \lesssim \ell_n \, D_a \, \xi \, \phi(\xi) \, . \end{align}

This corollary follows directly from Theorem (ref). Specifically, the approximate exponential distribution of $\mathbf{N}^{\rm w}(\xi)$ indicates $\mathbb{P}_{\infty}[\mathbf{N}^{\rm w}(\xi) < \ell_n] \sim 1 - \exp(-\ell_n \, D_a \, \xi \, \phi(\xi))$, and the result then follows from the inequality $1 - \exp(-x) \leq x$ for $x$ close to $0$. A sequence $\ell_n$ satisfying the conditions of Corollary (ref) constitutes a BFD.

Bound on Detection Delay (BDD)

When an It\^o semimartingale violation occurs at some time $\tau_n\downarrow 0$, we wish to detect it as quickly as possible, using discrete high-frequency observations of $Y$ starting at time 0. We refer to a deterministic sequence $T_n\rightarrow\infty$ such that $\mathbb{P}\left( \, \mathbf{N}^{\rm w}(\xi) \, > \, \tau_n/\Delta_n+T_n \, \right) ~\rightarrow~ 0$, when $H$ is nontrivial, as a bound (in probability) on detection delay or BDD. We want $T_n$ as small as possible. Theorem (ref) below provides a lower bound on BDD.

theoremSuppose $b_s = b_0$, $\sigma_s = \sigma_0$ and $\Delta X_s = 0$, almost surely, for $s$ in a neighborhood of zero and that Assumption (ref) holds. Let $H$ be given by the drift burst model in equation ((ref)) for some $c\neq 0$. Let $w_n$ be such that $w_n \sim a_n \, \xi^2$, with $a_n$ being a deterministic sequence converging to $a\in (2/c^2,\infty]$. Finally, let $\varpi\in (0,1/2)$ and $\vartheta\in(1/2,1)$ as well as $\xi\Delta_n^{\iota}\rightarrow 0$ for any $\iota>0$. We then have, \begin{equation} \mathbb{P}\left( \, \mathbf{N}^{\rm w}(\xi) \, > \, \tau_n/\Delta_n+T_n \, \right) \rightarrow 0 \, , \end{equation} for any sequence $T_n\rightarrow\infty$ such that $T_n \, \Delta_n\rightarrow 0$,\, $w_n/T_n\rightarrow 0$, \, $T_n/\Delta_n^{(1/2-\vartheta)/\vartheta}\rightarrow\infty$ \, and $\sqrt{w_n} ~ l_n \, \delta_n \, \Delta_n^{\varpi-1/2} / \, (\xi \, T_n) \rightarrow 0$.

We invoke a number of simplifying assumptions in the derivation of Theorem (ref). Primarily, we assume that $X$ is continuous and features constant drift and constant diffusive volatility in a (shrinking) neighborhood of zero. This style of assumption is common in the high-frequency financial econometrics literature and --- as is typically the case --- it should be possible to relax at the cost of more complicated derivations.

Theorem (ref) is proved by constructing an alternative measure, say $\widetilde\mathbb{P}$, under which the semimartingale violations are milder. Specifically, the expected value of the increments under $\widetilde\mathbb{P}$ are smaller in asymptotic order of magnitude than under $\mathbb{P}$. Consequently, it takes longer to detect the change, and thus the BDD under $\widetilde\mathbb{P}$ is an upper bound for that under $\mathbb{P}$.

The object characterized by the BDD condition ((ref)) is a diverging sequence of observations, $(T_n)$, sufficiently large to ensure almost sure detection, in our infill asymptotic setting, of an It\^o semimartingale violation. It is natural to require that our detector identifies such violations before triggering a false alarm with probability approaching one or, equivalently, BDD $<$ BFD. From Theorem (ref), we need $T_n/\Delta_n^{(1/2-\vartheta)/\vartheta}\rightarrow\infty$, while BFD, from Corollary (ref) and the discussion thereafter, is of asymptotic order $1/\, (\xi \, \phi(\xi))$. Thus, we must have $\xi \, \phi(\xi) \, \Delta_n^{(1/2-\vartheta)/\vartheta} \to \, 0$ to ensure BDD $<$ BFD. This condition may be satisfied for any $\vartheta<1$, provided $\xi$ is chosen as large as possible, while still guaranteeing that Theorem (ref) applies (recall the discussion after Theorem (ref)).

The BDD embeds two sources of delay (as shown in our proof): (i) the truncation of price increments to guard against jumps; (ii) the requisite cumulation of non-truncated returns in our detection statistic. Asymptotically, the first delay term dominates the second. Nonetheless, in finite samples, the BDD will be determined by both effects.

remarkWe can also introduce a double-window-limited GLR rule, defined as, \begin{align} \mathbf{N}^{\rm ww}(\xi) := \inf\left\{l > 0: \max_{l - w_n - r_n \leq k < l - r_n}\frac{|\widehat{W}_{k,l}|}{\sqrt{(l-k)\Delta_n}} > \xi \right\}, \end{align} where $r_n$ imposes a restriction on the minimum span for the statistic $\widehat{W}_{k,l}$. This helps reduce the false detection error under the null induced by extreme values for the statistic $\widehat{W}_{k,l}$ due to only a few ($\, < r_n$) increments. Consequently, this will generate a longer ARL/smaller FDR under the null, and a slightly larger BDD under the alternative. We do not analyze $\mathbf{N}^{\rm ww}(\xi)$ theoretically, but evaluate its performance numerically via simulation. The Monte Carlo study in Section (ref) shows that setting $r_n$ to, say, $5$, we obtain effective protection against outliers without affecting the detection delay by much. We note that the literature on testing for bubbles in macroeconomic settings often imposes a similar minimum duration restriction on the bubble period (see, e.g., phillips2011explosive, phillips2011dating, and phillips2015testingb).

Simulation Study

In this section we carry out a simulation study to assess the finite-sample performance of our procedure for swift detection of local It\^o semimartingale violations.

Simulation setting

To generate sample paths for the efficient log-price $X$, we simulate the following Heston type stochastic volatility model with jumps,

align[align omitted — 181 chars of source]

where $W_{1}$ and $W_{2}$ are standard Brownian motions with correlation $\mathbb{E}(dW_{1,t}, dW_{2,t}) = \rho \, d t$, and $J_t$ denotes the jump term for which we employ a Poisson process with intensity $p_X$ and jump size distribution $\mathcal{N}(0, \lambda_X^2)$\,.

In terms of parameter specification, we set the initial efficient log-price $X_0$ to $\log(1200)$, the drift term $b_t$ to zero, and the unit of time to one trading day. The volatility process $\sigma_t^2$ is initiated at its unconditional mean of $\gamma$ on day one while, for other days, it is initiated at the ending value of the previous day. The annualized parameter vector for the Heston model is set to $(\kappa, \gamma, \varsigma, \rho) = (5, 0.0225, 0.4, -\sqrt{0.5})$. For the jump components, we let $p_X = 3/5$, corresponding to 3 jump per week on average, and $\lambda_X = 0.5\%$.

For the It\^o semimartingale violation term $H$, we employ the DB and PN models ((ref)) and ((ref))--((ref)), respectively, to generate gradual-jump type patterns. Specifically, for the DB model, we set $c = 3$, $\tau_n = 0.25$ and $\overline\tau = 0.4$, and explore $\vartheta \in \{0.55, 0.65, 0.75\}$. For the PN model, we add a jump in $X$ at $\tau_n = 0.25$ each day. We let $f(\DeltaX_\tau, \eta_\tau) = \eta_\tau \, \DeltaX_\tau$ with $\eta_\tau = -1$, $\overline{\tau} = 0.4$, and explore $\vartheta_{pn} \in \{0.45, 0.35, 0.25\}$, corresponding to $\DeltaX_\tau = \{1.4\%, 2.0\%, 3.0\%\}$. Finally, $\widehat\sigma_t = \sigma_t\times(1 + 0.02 \times Z)$, where $Z$ is a standard normal random variable, serves as proxy for the (noisy) option-based volatility estimate.

Simulation results

We first present results for the null measure $\mathbb{P}_\infty$, i.e., no It\^o semimartingale violations.

table[table omitted — 989 chars of source]

In Table (ref), we report the ARL values for our proposed detectors $\mathbf{N}(\xi)$ (with $w_n = \infty$), $\mathbf{N}^{\rm w}(\xi)$ ($w_n = 30$ minutes), and $\mathbf{N}^{\rm ww}(\xi)$ ($(w_n, r_n) = (30, 5)$ minutes) based on our simulation setting. We also provide the theoretical ARL values for the former two detectors based on Theorem (ref). Comparing the first two and the next two columns we find that, for a wide range of threshold values ranging from $3.4$ to $4.0$, the Monte Carlo values for $\mathbf{N}(\xi)$ and $\mathbf{N}^{\rm w}(\xi)$ are close to their corresponding theoretical values, indicating that the asymptotic approximation ((ref)) captures the finite-sample performance well. When comparing the ARLs of $\mathbf{N}(\xi)$ to those of $\mathbf{N}^{\rm w}(\xi)$, we find that the latter are slightly larger than the former. This indicates that the window limit $w_n$ does not induce any major deterioration under the null, even with the relatively small value (here, $30$ minutes). On the contrary, in the last column, we find the ARL values for $\mathbf{N}^{\rm ww}(\xi)$ to be significantly larger than the previous ones within the same settings. Evidently, the window limit $r_n$ has a significant impact under the null measure by ignoring false detections induced by just a couple of increments. Thus, there is room to experiment along this dimension in order to improve the FDR, although it is critical to also monitor the associated increase in the detection delay under the alternative.

table[table omitted — 1,304 chars of source]

The above findings are corroborated by the FDR results in Table (ref), where we use a wider range of the threshold values, taking advantage of the fact that FDR can be assessed through simulations over deliberately chosen shorter time intervals. In particular, for each replication, we simulate one-minute prices for a day, and obtain the FDR as the fraction of replications with (at least) one false detection across our 5,000 replica. In general, the Monte Carlo FDRs are close to their theoretical values for $\mathbf{N}(\xi)$ and $\mathbf{N}^{\rm w}(\xi)$ by Corollary (ref), implying accurate analytic approximations. The window limit $w_n$ does not impact the FDRs by much, while a moderate choice for $r_n$ may significantly reduce the false detection rate under the null. Finally, we note that this table can be used as a guide for choosing $\xi$, as well as $w_n$ and/or $r_n$, in empirical applications, perhaps following further simulations for alternative asset price dynamics.

figure[figure omitted — 556 chars of source]

We round off our simulation study for the null measure $\mathbb{P}_{\infty}$ with Figure (ref). It displays histograms for the duration until (false) alarm associated with our detectors $\mathbf{N}(\xi)$ ($w_n = \infty$) and $\mathbf{N}^{\rm w}(\xi)$ ($w_n = 30$ minutes) based on 1,000 replications. The plots visually corroborate the approximate exponential distribution results in Theorem (ref).

We now turn to $\mathbb{P}_{\tau_n}$ --- the probability measure with local It\^o semimartingale violation at time $\tau_n$ under the setting of Section (ref). Table (ref) reports the (true) detection rates (denoted $\mathbf{r}$), i.e., what fraction of the 1,000 replications contain (at least) one detection, and the estimated expected detection delay or EDD defined as,

align[align omitted — 131 chars of source]

The EDD is regarded as the analogue of the ARL for the alternative measure $\mathbb{P}_{\tau_n}$. In fact, ARL versus EDD is the canonical criteria in the sequential detection literature (see, e.g., lorden1971procedures), as size versus power for hypothesis testing. Unfortunately, an explicit approximation for the EDD is not available in our case, and we can only derive BDD, which is a bound in probability on $\mathbf{N}^{\rm w}(\xi)$. Reporting EDD in the simulation, however, makes it easier to assess the size of $\mathbf{N}^{\rm w}(\xi)$ under the alternative.

table[table omitted — 1,491 chars of source]

We consider both DB and PN gradual jump patterns as specified above. In both cases, we explore our detectors $\mathbf{N}^{\rm w}(\xi)$ with $w_n = 30$ minutes and $\mathbf{N}^{\rm ww}(\xi)$ with $(w_n, r_n) = (30, 5)$ minutes under three sampling frequencies: $10$-seconds, $30$-seconds and $60$-seconds.

Both detectors provide satisfactory performance, delivering high detection rates and short detection delays. Specifically, comparing vertically within each panel, we see that $\mathbf{r}$ increases and EDD shrinks, as the violations grow more severe. Likewise, comparing horizontally within each panel, we find the detectors improving with higher sampling frequency --- generating larger $\mathbf{r}$'s and shorter detection delays (in seconds, obtained by multiplying the entries accordingly with $10$, $30$ and $60$). Experiments featuring even higher sampling frequencies, in turn, confirm our (asymptotic) immediate detection result. However, implementation at very high frequencies within an actual market setting is impacted by microstructure noise, so it is arguably less practically applicable unless one employs a scheme to explicitly account for ultra high-frequency frictions.

Empirical Application

Data

We apply our detectors to one-minute S&P 500 equity (SPX) index futures data covering January 2007 - December 2020. We only use data for the regular trading hours and eliminate days with reduced trading hours, producing a sample of $3,524$ days. Following our theoretical framework, we use the nonparametric option-based volatility estimator, SV, proposed by todorov2019nonparametric. We rescale SV using the previous day's truncated volatility (TV) so that they are at the same level.

Detection of semimartingale violations for S&P 500

We apply $\mathbf{N}^{\rm ww}(\xi)$ with $(w_n,r_n) = (30,5)$ minutes and $\xi = 4$ for our one-minute data each day, implying we initiate the procedure $30$ minutes after the market open. Given this design, we lose power in terms of detecting violation periods shorter than $5$ minutes. From Table (ref), we see that this detector has a probability of about $6\%$ to trigger a false alarm at the daily level.

We further equip our detection with an ad hoc identification rule to pin down It\^o semimartingale violation regions. Specifically, we define it as the union of all (probably overlapping) intervals, $[\,l_a, l_b\,]$, such that

align[align omitted — 158 chars of source]

That is, in words, we define the region as the union of the span of GLR-CUSUM statistics that exceed the chosen threshold $\xi$.

Figures (ref) and (ref) provide illustrations for two trading days, August 7, 2007 and August 30, 2019, with detected It\^o semimartingale violations. The price paths are red, $\mathbf{N}^{\rm ww}(\xi)$ rejections are indicated by dark grey vertical lines, and the violation regions by light grey areas. In the bottom panel, we also provide $30$-minutes rolling window TV in blue and the option-based SV (rescaled by the same day's TV average) in orange.

On August 7, 2007, the price displays a gradual upward trend, until an abrupt 20 point crash over 2:15--2:30 pm, followed by a rapid recovery, taking the price beyond the original level. Our detector swiftly signals a violation as the flash-crash pattern emerges, triggering an alarm within a few minutes. As shown in the bottom plot, this dramatic price pattern induces an explosion in the local realized volatility measure, which is not consistent with our option-based SV measure, confirming the potential for dramatic biases in return-based volatility measurement. Similar findings hold for our second example on August 30, 2019, where a gradual jump initiates before 11:00 am. Again, our detector captures it expeditiously.

figure[figure omitted — 625 chars of source]

Table (ref) reports the daily count of It\^o semimartingale violation regions detected by $\mathbf{N}^{\rm ww}(\xi)$ with window length $(w_n,r_n) = (30,5)$ minutes under various values for the threshold $\xi$. We also split the trading day (of $390$ minutes) into three periods --- the “Morning” for the first $130$ minutes, the “Noon” for the middle $130$ minutes, and the “Afternoon” for the rest --- and report the associated count of violation regions initiated within that period. Finally, Figure (ref) plots the corresponding histograms for the duration of these violation regions for the scenario with $\xi = 4$.

figure[figure omitted — 628 chars of source]

We find that “Morning” generates most violations, with the “Noon” containing only slightly fewer, while “Afternoon” produces the least, even if the difference is not striking. In sum, the violations are frequent and they are not particularly prone to occur in specific regions of the trading day. Finally, from the histograms in Figure (ref), we note that the duration of a typical violation is short, with the clear majority lasting less than 20 minutes. Moreover, the violation regions starting in the morning tend to be somewhat shorter lived than those initiated during noon or afternoon.

table[table omitted — 951 chars of source]
figure[figure omitted — 560 chars of source]

A manifestation of the It\^o semimartingale violation during the detected PN episodes is the divergence between return- and option-based spot volatility estimators. The bottom panels of Figures (ref) and (ref) show that the divergences are large in these two specific cases. Table (ref) reports the sample mean of the return- and option-based volatility estimates $TV$ and $SV$ over the hour before and after the initiation of a PN episode. The results reveal that the two volatility proxies are close before the PN episode but feature a substantial gap over the following hour, with the average $TV$ about $35\%$ higher than the average $SV$. This gap only grows larger, if we exclude days with FOMC announcements from the calculation, so the discrepancy is not driven by this type of macroeconomic announcements. These findings demonstrate how local deviations from the semimartingale assumption can distort the measurement of volatility.

table[table omitted — 729 chars of source]

Finally, Table (ref) reports the number of PN episodes for each year across the sample given different threshold levels for detection. The PN events appear near uniformly distributed over the sample, with no particular clustering apparent during periods characterized by generally volatile or tranquil market conditions. Overall, we conclude that PN episodes are largely idiosyncratic events with no particular tendency to occur during specific times within the trading day or during specific market conditions.

table[table omitted — 953 chars of source]

Conclusion

In this paper, we focus on real-time detection of local It\^o semimartingale violations that have drawn increasing attention in the recent literature. We propose CUSUM-type detectors as stopping rules exploiting high-frequency data. We show that they possess desirable theoretical properties under infill asymptotics. Specifically, for a suitably chosen average run length (the average sample length until a false alarm), our detectors enable “quick” detection of a violation, once it occurs. Our formal interpretation of rapid detection is that the bound on the detection delay (BDD) shrinks to zero under infill asymptotics, as the sampling interval $\Delta_n$ goes to zero. These properties are corroborated through simulations calibrated to reflect key features of market data. Finally, we apply our detector to S&P 500 equity (SPX) index futures data. We obtain timely detection of a nontrivial number of short-lived episodes involving likely semimartingale violations. Such turbulent market events are critical for real-time decision making. For example, they may indicate temporary market malfunction motivating exchange or regulatory action, they may induce the termination of trading strategies and algorithms among active market participants triggering a period of fleeting liquidity, and they signal problems in extracting reliable high-frequency based market volatility and risk measures.