Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,288 characters · 11 sections · 63 citation commands
Microstructural Foundations of Rough Noise
\and Anders Midtgaard Norlyk. } }
{\bf Keywords}: Ambit stochastics, market microstructure, microstructure noise, rough noise, Gaussian moving averages.\\
\doublespacing
It is a well-known stylized fact in financial econometrics that prices observed at high frequencies are contaminated by microstructure noise, stemming from bid--ask bounces, transaction costs, rounding, and related frictions Hasbrouck2007,JacodEtAl2017. The observed logarithmic price is typically decomposed as $Y_t = X_t + Z_t$, where the efficient price $X$ is a continuous It\^o semimartingale and $Z$ is additive noise AitSahaliaEtAl2005. When prices are observed on a grid $\{i\Delta_n : i=1,\ldots,[T/\Delta_n]\}$, the noise is conventionally modeled as a discrete-time sequence $Z_{i\Delta_n}=\varepsilon_i$, specified, for example, as white noise (ChristensenEtAl2010,preAvg09), as an AR- or MA-type series (PeterAsgerRV06,SahaliaMyklandZhang11), or as arising from rounding (DelattreJacod1997,LiMykland2010). The noise also has well-documented consequences for inference. In particular, realized variance, which consistently estimates integrated volatility in the absence of noise andersen_etal01a, diverges as the sampling frequency increases AndersenEtAl2000.
Recently, ChongEtAl2021 argued that discrete noise sequences are incompatible with two empirical features of high-frequency data. First, realized variance diverges at a rate slower than the $\Delta_n^{-1}$ implied by stationary discrete noise. Second, observed price increments shrink as $\Delta_n\to0$, which discrete noise cannot produce. In short, the noise is non-shrinking while its increments shrink. They therefore propose modeling $Z$ as a continuous-time process with sample paths rougher than those of Brownian motion, which accommodates both features. The roughness of the noise matters for inference, as classical noise-robust techniques such as preaveraging preAvg09 are no longer consistent estimators of integrated volatility under rough noise ChongEtAl2021.
While the rough-noise model has appealing properties on the macroscopic scale, an important question is whether it is consistent with a realistic microstructural description of prices. That is, does there exist a tick-by-tick model, replicating the salient features of ultrahigh-frequency prices, whose scaling limit is of the form $Y_t=X_t+Z_t$ with $Z$ a continuous-time rough process? Providing such a microstructural foundation is the aim of the present paper.
Our first contribution is a microstructural foundation for rough noise. Building on the model of fleeting price moves of Shephard2017, we model tick-by-tick prices on an integer grid, decomposed into a permanent and a transitory component. The permanent component is an integer-valued L\'evy process, the continuous-time analogue of a random walk on the tick grid, while the transitory component is a trawl process capturing short-lived price distortions and their subsequent correction. In an asymptotic framework inspired by Rosenbaum2018, we show that the rescaled price process converges to a Bachelier-type semimartingale plus a Gaussian moving average whose kernel is determined explicitly by the trawl function. By choosing the trawl function appropriately, the limiting noise attains any roughness index $\alpha\in(-1/2,0]$, matching the local behavior, and hence the roughness, of the noise specification of ChongEtAl2021. Moreover, the mechanism is transparent, as roughness arises when distortions of very short duration dominate, being created and corrected almost instantaneously.
Our second contribution is an inference methodology applicable directly on data at the highest frequency. We construct a GMM estimator of the model parameters from second moments of price increments across multiple lags, and establish its consistency and asymptotic distribution. Since the null hypothesis of non-rough noise places the roughness index on the boundary of the parameter space, standard asymptotics fail. Using results of Andrews2002, we show that the estimator remains $\sqrt{n}$-consistent under the null, with a one-sided normal limit with a point mass at zero, and we derive a feasible test of the null of no roughness against a rough alternative. A simulation study shows that the estimator is approximately unbiased and that the test has good size and power in samples corresponding to a single trading day.
Our third contribution is empirical. We apply the methodology to tick-by-tick transaction data on the Dow Jones Industrial Average constituents in 2024. Because the estimator operates at the tick level, we obtain a separate roughness estimate for each ticker-day, in contrast to existing evidence based on multi-day rolling windows. The daily estimates reveal that rough noise is present but episodic. About one fifth of ticker-days with detectable noise reject the null of no roughness, the estimated roughness is typically mild, and detections concentrate on days dominated by short-run price reversals rather than being explained by broad illiquidity. The episodic nature of rough noise also reconciles our estimates with the stronger, stable roughness reported by ChongEtAl2021. Their estimator is computed over multi-day rolling windows and is dominated by the roughest day in each window, so occasional rough days can generate stable rough estimates even when the typical day is not rough.
Our paper relates to a literature that builds tick-level models consistent with semimartingale limits. AitSahaliaJacod2020 construct a model of transaction prices that remain positive, move by at most one tick at a time, and exhibit rapid runs of successive up- or down-ticks, and show that, after rescaling, such a model can converge to an It\^o semimartingale with stochastic volatility and jumps. BacryEtAl2013b,Bacry2013 use Hawkes processes to reproduce the divergence of realized variance and the Epps effect Epps1979. None of these models generate rough noise in the limit. Our setup is instead analogous to the microstructural foundations of rough volatility, in which Rosenbaum2018 obtain rough volatility as the scaling limit of nearly unstable Hawkes models. In our model, roughness enters through a different channel, namely the behavior of the trawl function near zero, and it appears in the noise rather than in the volatility.
The remainder of the paper is organized as follows. Section (ref) motivates the tick-by-tick specification, combining microstructure arguments with empirical evidence from transaction data. Section (ref) introduces the model and establishes its macroscopic limit. Section (ref) develops the GMM estimator, its asymptotic theory, and the test for rough noise. Section (ref) examines finite-sample performance, Section (ref) contains the empirical application, and Section (ref) concludes.
In this section, we motivate the tick-by-tick model that serves as the microscopic starting point of the paper. The observed price is decomposed into an efficient price component and a microstructure noise component, $Y_t = X_t + Z_t$ for $t\ge 0$. Our goal is to specify these two components in a way that is consistent with the tick structure of transaction prices while retaining a clear distinction between permanent price changes and transitory distortions.
In the high-frequency literature, the efficient price is typically modeled at the macroscopic scale as an It\^o semimartingale. At the tick level, however, transaction prices live on a discrete grid, and a microscopic model should respect this discreteness directly. Prices also evolve in continuous time, and both components should reflect this. For the efficient price the continuous-time formulation is standard, while for the noise it is motivated by ChongEtAl2021, as discussed in the introduction. A third requirement concerns the economic role of the efficient price. Abstracting from time variation in risk premia, an efficient-market view implies that its increments behave approximately as martingale increments, whose classical discrete-time benchmark is the random walk Campbell1998. These three requirements---discreteness, continuous time, and martingale-like increments---are jointly met by an integer-valued L\'evy process, the continuous-time analogue of a random walk on the tick grid, which we therefore adopt for the permanent component.
The transitory component \(Z_t\) should capture temporary deviations of transaction prices from the efficient price. The canonical mechanism is bid--ask bounces. Transaction prices alternate between bid and ask quotes, which mechanically generates reversals in observed price changes, and hence negative autocorrelation in high-frequency returns, even when the efficient price itself has martingale increments Roll1984. Related mechanisms whose price impact is transitory rather than informational, such as temporary liquidity pressure and inventory control, have similar effects (Stoll1978,Hasbrouck2007,Ait2009), and negative serial dependence in high-frequency returns is a well-documented feature of transaction prices (SahaliaMyklandZhang11,JacodEtAl2017).
ChongEtAl2021 emphasize a related empirical feature, namely that the presence of microstructure noise depends on whether prices are constructed from transactions on a single exchange or from all exchanges.\footnote{This distinction is closely related to Rule P.3 in the cleaning procedure of Barndorff2009, which recommends constructing prices from a single exchange rather than consolidating transactions across exchanges.} Bid--ask bounces generates negative autocorrelation already at the single-venue level, but it cannot explain why the dependence strengthens once transactions are consolidated across exchanges. A natural candidate for the additional dependence is cross-venue latency. If different venues incorporate information with different delays, switching between them in the consolidated tape can generate rapid reversals, even when each venue's own price has uncorrelated increments. In a simple Hasbrouck1995 model with prices on two exchanges, both following a random walk with one lagging the other, consolidated returns exhibit negative first-order autocorrelation. The calculations are detailed in online Appendix D.
We next examine whether this mechanism is visible in our data. Using tick data throughout 2024 for the constituents of the Dow Jones Industrial Average, we test, for each ticker and day, whether one-second returns have zero first-order autocorrelation, using a Wald test with HAC standard errors, and we compare prices constructed from a single exchange with prices constructed from all exchanges. Since the efficient price is common across venues, a systematic change in autocorrelation between the two constructions is naturally attributed to cross-venue microstructure effects. The results are shown in (ref). For single-exchange prices, the distribution of test statistics is relatively close to the standard normal benchmark and the null is rejected on \(12\%\) of ticker-days, whereas for all-exchange prices the distribution shifts sharply to the left and the null is rejected on \(55\%\). Consolidating transactions across venues thus introduces or amplifies short-lived reversals in observed prices. The same all-exchange construction and autocorrelation test reappear in the empirical application in (ref), where the latter serves as a pre-test for the presence of microstructure noise.
The negative dependence in (ref) indicates mean reversion in observed prices that is detectable already at the one-second horizon. To capture such short-lived distortions, we follow the continuous-time model of fleeting discrete price moves of Shephard2017 and model the noise as a trawl process, a stationary integer-valued process in which each distortion appears, persists for a random duration, and is then corrected. The distribution of these durations is governed by a deterministic function, the trawl function, which thereby controls the dependence structure of the noise.
(ref) formalizes this discussion. Both components of the observed price are generated by a single integer-valued L\'evy basis evaluated on two different sets in space--time. Evaluating the basis on a growing rectangle yields the permanent component, an integer-valued L\'evy process, while evaluating it on a moving set shaped by the trawl function yields the transitory component. The behavior of the trawl function near zero then determines the properties of the noise in the macroscopic limit. When it is unbounded at zero, distortions of arbitrarily short duration dominate, and the accumulation of these microscopic reversals produces rough noise at the macroscopic scale.
In this section, we introduce the model of Shephard2017 in detail. In view of the preceding discussion, this model will serve as our microscopic description of price formation at ultra-high frequency. Our aim is to show that, under a suitable scaling, this microscopic model gives rise to a Bachelier-type price model with additive rough noise, belonging to the class of mixed semimartingale models considered by ChongEtAl2021.
The macroscopic target of this section is the mixed semimartingale model of ChongEtAl2021. There, the observed price is $Y_t = X_t + Z_t$, where $X$ is an It\^o semimartingale and the noise component satisfies $Z_t = Z_0+\int_0^t g(t-s)\rho_s\,dW_s$. Here $(\rho_t)_{t\geq 0}$ is adapted and locally bounded, while the kernel $g$ is of the form
where $g_0$ is smooth with $g_0(0)=0$, and $K_H$ is a constant depending on $H$. The noise is therefore a Gaussian moving average whose sample paths are rougher than those of Brownian motio, in the sense made precise in (ref), its roughness index is $\alpha=H-1/2<0$.
This model does not respect the tick structure of markets and is therefore not suitable at the highest frequencies. However, microstructure noise is still modelled explicitly, so the setting is not noise-free. With this in mind, we view it as a model for prices at high frequency, but not at ultra-high frequency: price discreteness is no longer a dominant feature, but microstructure noise remains present.
In the remainder of this section, we first introduce the microscopic price model, then establish its macroscopic limit, and finally construct a trawl function whose limit matches the local behavior, and hence the roughness, of the target kernel (ref).
Let \(L\) be an integer-valued, time-homogeneous L\'evy basis on \(\mathcal S\times\mathbb R\), where \(\mathcal S\subseteq\mathbb R_+\). That is, \(L\) assigns an integer-valued random variable \(L(E)\) to each Borel set \(E\subseteq\mathcal S\times\mathbb R\) with finite Lebesgue measure, is additive over disjoint sets, and the resulting random variables are independent. We write \(L_1\) for the L\'evy seed of \(L\), which may be identified in distribution with \(L(E_0)\) for any Borel set \(E_0\subseteq\mathcal S\times\mathbb R\) with \(\operatorname{Leb}(E_0)=1\), and we denote its L\'evy measure by \(\nu\).
Fix \(b\in(0,1)\), and let \(d:[0,\infty)\to\mathcal S\) be a non-increasing function satisfying that $\int_0^\infty d(s) - b ds <\infty$. Define $A = \{(x,s):s\leq 0,\ b\leq x<d(-s)\}$, and let $A_t=A+(0,t) = \{(x,s):s\leq t,\ b\leq x<d(t-s)\}$, with $t\geq 0$. Finally, define $B_t=[0,b)\times(0,t]$, for $t\geq 0$. We then set $C_t=A_t\cup B_t$ and define the microscopic price process by \[ P_t = V_0+L(C_t) = V_0+L(A_t)+L(B_t), \qquad t\geq 0. \] The term $L(B_t)$ is a Lévy process and captures permanent price changes: once an event enters $B_t$, it never leaves. We therefore interpret $L(B_t)$ as the efficient price component. By contrast, $L(A_t)$ is a trawl process. The set $A_t$ moves through the Lévy basis over time, so points enter and later leave $A_t$. Hence $L(A_t)$ captures temporary price changes and is interpreted as microstructure noise. The function $d$ controls the shape and overlap of the sets $A_t$ over time and therefore determines the dependence structure of this transitory component. We refer to it as the trawl function. Since $A_t$ and $B_t$ are disjoint for each $t\geq0$, and since $L$ is independently scattered, the two components are independent. Finally, the initial value $V_0$ captures permanent price changes prior to time zero.
As argued in Shephard2017, this specification captures several key features of tick-by-tick prices: prices move on a discrete grid, evolve in continuous time, and may exhibit a large fraction of fleeting price changes. The parameter \(b\) controls the relative size of the permanent component: larger values of \(b\) imply that a larger share of price changes is permanent rather than fleeting.
To bridge the microscopic and macroscopic scales, we introduce an asymptotic framework inspired by Rosenbaum2018.\footnote{Our setup is similar to that of Rosenbaum2018 in the sense that we work with a sequence of processes that become increasingly active as \(T\) increases. The difference is that roughness arises there from a nearly unstable Hawkes kernel, whereas in our model it arises from the behavior of the trawl function near zero.} For each \(T\in\mathbb N\), we work on a probability space $(\Omega^T,\mathcal F^T,\mathbb P^T)$, on which \(L^T\) is an integer-valued, time-homogeneous L\'evy basis on \([0,T]\times\mathbb R\). Let \(L_1^T\) denote the associated L\'evy seed, and let \(\nu^T\) denote its L\'evy measure. For each \(T\), the trawl component is generated by a \(T\)-dependent trawl function \(d^T\), with associated trawl sets $A_t^T = \{(x,s):s\leq t,\ b\leq x<d^T(t-s)\}$ for $t\geq 0$. We equip the probability space with the completed right-continuous filtration \((\mathcal F_t^T)_{t\geq 0}\) generated by \(\{L^T(A_s^T):0\leq s\leq t\}\), \(\{L^T(B_s):0\leq s\leq t\}\), and \(V_0^T\).
To obtain a non-degenerate macroscopic limit, we let the activity of the L\'evy basis increase linearly with \(T\).
The trawl functions are also \(T\)-dependent. The next assumption requires \(d^T\) to converge monotonically to a limiting trawl function \(d\), while ensuring that each finite-\(T\) trawl set is bounded and has finite Lebesgue measure.
For each fixed \(T\in\mathbb N\), the microscopic price process is
We are going to study the asymptotic behavior of \(P^T=(P_t^T)_{t\geq 0}\) as \(T\to\infty\).
To derive the convergence of the price process, we will require a number of assumptions on the moments of the Lévy seed. As argued in Barndorff2012, we may represent the Lévy process associated to the Lévy seed in terms of the difference of two, independent, discrete subordinators, $\tilde{L}_t^T$ and $\bar{L}_t^T$, whose Lévy measures $\tilde{\nu}^T$, $\bar{\nu}^T$ are the restrictions of $\nu^T$ to respectively the positive and negative half axes. We will state the assumptions on the Lévy seed in terms of these.
Finally, we will make the very mild assumptions that our efficient price process initially is almost surely finite.
In this subsection, we state the main convergence result for the microscopic price process. The result shows that, after rescaling, the permanent component converges to Brownian motion, while the trawl component converges to a Gaussian moving average. The proof relies on auxiliary convergence results for trawl processes, which are stated in online Appendix B.
We first impose a condition linking the limiting trawl function to the kernel of the Gaussian moving-average process that appears in the limit. These processes take the form $Y_t=\int_{-\infty}^t g(t-s)\,\mathrm d W_s$ for $t\in\mathbb R$, where \(W\) is a Brownian motion on $\mathbb{R}$ and \(g\) is a square-integrable kernel.
To describe the local regularity of such processes, we recall the notion of roughness used throughout the paper. Let \(U=(U_t)_{t\geq 0}\) be a continuous process with stationary increments. Following Bennedsen2021, we say that \(U\) has roughness index \(\alpha\) if its second-order variogram satisfies $\mathbb E[(U_h-U_0)^2]=h^{2\alpha+1}L(h)$ for $h>0$, where $L$ is continuously differentiable, slowly varying at zero,\footnote{By a function that is slowly varying at zero, we mean a function with the property that for any $a>0$, $\lim_{h\to0}\frac{L(ah)}{L(h)}=1$.} and bounded away from zero in a neighbourhood of zero. The index $\alpha$ describes the local regularity of the sample paths, with smaller values of $\alpha$ corresponding to rougher local behaviour. Brownian motion has roughness index $0$, and processes with $\alpha<0$ are rougher than Brownian motion. For a Gaussian moving average, the roughness is determined by the local behavior of the kernel near zero. In particular, if $g(x)$ behaves like $x^{\alpha}$ as $x\downarrow 0$, then $Y$ has roughness index $\alpha$ BLP2021. The following assumption specifies the relation between the limiting trawl function \(d\) and the kernel \(g\).
For the functional convergence statement, we also record the following sufficient condition.
We can now state the main limit theorem for the full price process.
The first part of (ref) is the main convergence result used in the paper. It shows that the permanent component converges to Brownian motion, while the fleeting trawl component converges to a Gaussian moving average. The second part gives a sufficient condition under which the convergence can be strengthened from finite-dimensional convergence to functional convergence. As we will see in the following subsection, this sufficient condition is satisfied in the non-rough case but fails for the rough trawl functions considered below. Thus, in the rough-noise case, the link between the microscopic and macroscopic models is established through finite-dimensional convergence.
From this section onward, we work with the $T=1$ price process, i.e.\ the one generated by a Lévy basis on $[0,1]\times\mathbb{R}$. We refer to this as the standard-time scale, corresponding to the tick-by-tick observation frequency in our data. Since $T$ is fixed henceforth, we suppress the $T$-superscripts to ease notation.
We derive a generalized method of moments (GMM) estimator for the parameters of the model in (ref), characterize its asymptotic behavior, and develop a formal test for the presence of rough noise. All proofs are deferred to online Appendices A and C. The parameters are $\theta = (\alpha,\lambda,b,\kappa_2(L_1))^\top$, where $\alpha$ and $\lambda$ parameterize the trawl function $d$ and $\kappa_2(L_1)$ is the variance of the Lévy seed. The roughness index $\alpha$ is the central object of interest, since it governs whether the noise is rough.
Our estimator is semiparametric. The model accommodates a wide range of Lévy seeds, but our interest is in the roughness of the noise, which is controlled by the trawl function rather than by the Lévy seed itself. We therefore neither assume a parametric form for the Lévy seed nor estimate its parameters beyond the second cumulant $\kappa_2(L_1)$. The first cumulant, corresponding to the mean, is set to zero by (ref).
We propose a GMM estimator similar to that of Shephard2017. For $h>0$, let $\Delta P_t^h := P_{t+h} - P_t$ denote the price increment over a time step of length $h$ at time $t$. For each $m \in \mathbb{N}$, let $D_m \subset \mathbb{N}$ be a finite set of $m$ positive integers, written $D_m = \{l_1,\dots,l_m\}$, which index the lags used in estimation. Define the vector of increments $Y_t^{(m)} = \left( \Delta P_{t\delta}^{l_1 \delta},\Delta P_{t\delta}^{l_2 \delta},\dots,\Delta P_{t\delta}^{l_m \delta}\right)$ for $t = 1,\dots,n-l_m$, where $\delta = 1/n$ for some $n \in \mathbb{N}$. Let $\Theta$ denote the parameter space and set $D(k,\theta) := \mathbb{E}\!\left[\left(\Delta P_1^{l_k\delta}\right)^2\right]$, $k = 1,\dots,m$. The sample moment function is $g_{n, m}(\theta) = \frac{1}{n-l_m} \sum_{t=1}^{n-l_m} h\!\left(Y_t^{(m)}, \theta\right)$, where $h:\mathbb{R}^m \times \Theta \to \mathbb{R}^m$ stacks the centered second moments, $h\!\left(Y_t^{(m)}, \theta\right) = \left((\Delta P_{t}^{l_1 \delta})^2 - D(1, \theta),\dots,(\Delta P_{t}^{l_m \delta})^2 - D(m, \theta)\right)^{\top}$. The GMM estimator of $\theta_0$ is
where $A_{n, m}$ is the diagonal weighting matrix with $i$th diagonal entry $1/(l_i\delta)^2$. The weights $1/(l_i\delta)^2$ are (approximate) inverse-variance weights under the null hypothesis $\alpha = 0$, where $\mathbb{E}[(\Delta P_t^h)^2]$ scales linearly in $h$. This is not efficient in general, but consistency of $\widehat\theta_{0,\mathrm{GMM}}^{n,m}$ does not depend on $A_{n,m}$. Consistency and the asymptotic distribution of $\widehat\theta_{0,\mathrm{GMM}}^{n,m}$ are established in the next subsection.
Since we work with integer-valued Lévy bases, the corresponding Lévy seed has characteristic triplet $(\gamma, 0, \eta)$, where $\gamma = \int_{\mathbb{R}} \xi \mathbb{I}_{[-1,1]}(\xi)\,\eta(d\xi) = \sum_{\xi=-1}^{1} \xi\,\eta(\xi)$. We impose the following assumptions.
We first establish asymptotic identification of the parameter vector.
Asymptotic identification, in the sense of Definition 3.3 of Stelzer2015 and Curato2019, ensures that the parameters are nearly identified for sufficiently large $m$. Since exact identification is needed for consistency, we strengthen (ref) to an assumption; the simulation study in (ref) provides further evidence that the parameters are identified in practice.
In the remainder of this subsection we assume that $m$ is sufficiently large.
We next derive the asymptotic distribution, distinguishing the case where $\theta_0$ lies in the interior of $\Theta$ from the boundary case $\alpha_0=0$ that is relevant for testing non-rough noise. In the interior a standard $\sqrt{n}$ central limit theorem applies; on the boundary we invoke Andrews2002, so the rate remains $\sqrt{n}$ but the limit is a one-sided normal with a point mass at zero.
To assess the finite-sample performance of the GMM estimator, we conduct a Monte Carlo study, using the trawl function in (ref) so that the simulated process resembles real tick-by-tick data.
We take $\mathcal{T}\in\{5850,11700,23400\}$ time periods, interpreted as seconds. This corresponds to a quarter trading day, half a trading day, and a full trading day, respectively. The grid spacing is $\delta=0.1$, giving one observation every tenth of a second. The underlying Lévy basis is Skellam with intensity $\nu\bigl(\mathbb{Z}\backslash\{0\}\bigr)=6.28$ per second, so each tick is $\pm 1$, which means we expect just over $6$ price changes per second. This corresponds to our empirical average. The GMM moment conditions use $m=100$ lags $D_m=\{l_1,\dots,l_{100}\}$ ranging from $l_1=1$ to $l_{100}=100$ (in units of $\delta$), equidistantly-spaced. The weighting matrix is the diagonal $A_{n,m}$ from (ref), and the long-run covariance estimator $\widehat{\Sigma}_a^k(\widehat\theta^n_{\mathrm{full}})$ truncates lag covariances at $k=1000$. We simulate $R=1000$ replications and report the results in (ref).
We see that the GMM estimator performs well across roughness levels and sample sizes: bias is negligible and the Monte Carlo standard deviations are small in every setting.
Next, we investigate the size and power of the test. We perform $1000$ replications at each of $10$ values of $\alpha$ between $-0.15$ and $0$, testing $H_0:\alpha=0$ at the $5\%$ significance level. The rejection rates are shown in Panel A of (ref).
The power curve has the expected shape: rejection rates approach $1$ quickly as $\alpha$ moves away from the null. The rejection rate at $\alpha_0=0$ exceeds the nominal $5\%$ level but improves substantially as the sample size grows, suggesting that the size distortion is a finite-sample effect (our largest $\mathcal{T}=23400$ corresponds to a single trading day).
We next examine the distribution of the test statistic under the null $\alpha_0=0$. By (ref), $T_n\overset{d}{\to} Z\mathbbm{1}_{\{Z\le 0\}}$ with $Z\sim\mathrm{N}(0,1)$, which places mass $1/2$ at zero and is half-normal on the negative half-line. Panel B of (ref) plots the empirical CDF of $T_n$ for $\mathcal{T}\in\{5850,11700,23400,46800\}$ against the limit CDF $F(x)=\Phi(x)\mathbbm{1}_{\{x<0\}}+\mathbbm{1}_{\{x\ge 0\}}$.
For smaller samples, the empirical distribution is somewhat heavier-tailed than the limit, consistent with the inflated size in Panel A. However, the fit improves substantially as $\mathcal{T}$ grows, suggesting once again that this is a finite-sample phenomenon. We also note that the median of the test statistics, which according to theory should be exactly zero, is zero up to the 4th decimal point for the two larger sample sizes. Overall, the simulations seem to support our theoretical results.
We apply the estimation strategy from (ref) to tick-by-tick transaction data for the constituents of the Dow Jones Industrial Average (DJIA) during 2024, producing a roughness estimate for each ticker-day. The DJIA composition is taken as of the beginning of the year. Motivated by the cross-venue evidence in (ref), we use trades from all available exchanges. The raw transaction data is otherwise processed following Barndorff2009.\footnote{We follow the trade-based cleaning rules of Barndorff2009, with two deviations. First, we keep trades from all available exchanges, dropping their single-exchange restriction (rule P3), since the cross-venue evidence in (ref) motivates the all-exchange construction. Second, we omit the quote-based filters, since we do not have access to quote data.}
As discussed in (ref), in a tick-by-tick model it is natural to interpret microstructure noise as a source of short-run reversals in observed transaction prices.
This motivates two empirical questions. First, how prevalent is rough noise among DJIA constituents in 2024? Second, are rough-noise detections associated with observable trading characteristics that reflect the high-frequency frictions underlying microstructure noise? We address the first by reporting the share of ticker-days for which the test of (ref) rejects the null of non-rough noise, together with descriptive statistics of the parameter estimates. We address the second by following Ait2009 in regressing the rough-noise indicator on standard liquidity, activity, and price-level measures.
We use the same estimation configuration as in the simulation study of (ref): prices sampled at $\delta=0.1$ seconds, $m=100$ equidistantly-spaced lags, the diagonal weighting matrix $A_{n,m}$, and the long-run covariance estimator $\widehat\Sigma_a^k$ truncated at $k=1000$. Since the model parameters are not identifiable in the absence of microstructure noise, we restrict the sample using the HAC Wald pre-test from (ref) for the null of non-negative first-order autocorrelation, $H_0:\rho(1)\geq 0$, at the $5\%$ level, retaining only ticker-days for which the null is rejected. Across the full sample, $64\%$ of ticker-day observations exhibit evidence of microstructure noise. For these noisy ticker-days, we estimate the model parameters and apply the test of (ref), which rejects non-rough noise at the $5\%$ level when the statistic falls below $-1.645$, the corresponding quantile of the half-normal-with-point-mass limit distribution. We find that $20.4\%$ of noisy ticker-days exhibit rough noise, a non-negligible share.
(ref) reports the empirical distribution of the parameter estimates across noisy ticker-days. The distribution of $\alpha$ is concentrated near zero, with a median of essentially zero and only the most rough $5\%$ of estimates falling below $-0.18$. Conditioning on the subsample where the null of non-rough noise is rejected, the median estimate is $-0.08$.
At first glance, our findings seem to contrast sharply with ChongEtAl2021, who report a stable roughness level over the period 2013--2022 corresponding in our notation to $\alpha \approx -0.20$. The apparent disagreement, however, is consistent with the difference in estimation methodology between the two studies. They estimate based on power variation over a 5-day rolling window, which will asymptotically be dominated by the roughest day in the window. A single day with $\alpha < 0$ pulls $\widehat\alpha$ toward its value even when the surrounding four days are non-rough. Their $\alpha \approx -0.20$ is therefore consistent with rough noise occurring episodically rather than systematically, with rolling estimates reflecting the most rough day in each window rather than a typical day. Our day-by-day estimator unbundles this signal and reveals substantial day-to-day variation in the underlying roughness, with most days exhibiting no roughness at all.\footnote{The estimator occasionally hits the boundary at $\alpha=-1/2$ or $\lambda=0$. These cases coincide with days where $b$ is estimated close to its upper boundary of $1$, on which the trawl component shrinks toward zero and the noise parameters $\alpha$ and $\lambda$ become hard to identify.}
We next study which empirical characteristics explain the presence of rough noise. Ait2009 show that the magnitude of market microstructure noise is closely related to stock liquidity, with estimated noise and noise-to-signal ratios lower for more liquid stocks. We follow the spirit of their exercise but replace the estimated noise magnitude with a binary indicator for rough noise.
Specifically, we estimate linear probability models $ROUGH_{i,d} = \beta_0 + \beta' X_{i,d} + \varepsilon_{i,d}$, where $ROUGH_{i,d} = \mathbf{1}\{T_{i,d} \le -1.645\}$ is the rough-noise indicator from the test of (ref). The critical value corresponds to a 5% test under the distribution of our test-statistic under the null. A positive coefficient in $\beta$ indicates that the corresponding variable is associated with a higher probability of detecting rough noise.
The vector $X_{i,d}$ is chosen to mirror the liquidity measures in Ait2009 as closely as possible given our data: \[
\] Here, $\sigma_{i,d}$ is the intraday realized volatility based on $5$-minute returns, $LOGTRADESIZE_{i,d}$ and $LOGNTRADE_{i,d}$ are the logarithms of the average trade size and of the number of intraday trades, $MONTHVOL_{i,d}$ is a lagged 21-trading-day realized-volatility measure, $LOGP_{i,d}$ is the logarithm of the daily closing price, and $ILLIQ_{i,d}$ is the daily Amihud illiquidity measure Amihud2002, computed as absolute daily return divided by dollar volume.
Since we do not have quoted bid--ask spreads, we follow Roll1984 and proxy the average spread by the Roll spread, $\text{RollSpread}_{i,d} = 2\sqrt{\max\{-\gamma_{1,i,d},\,0\}}$, where $\gamma_{1,i,d}$ is the first-order autocovariance of intraday price changes. The Roll spread is a direct function of the negative autocovariance and therefore captures short-run price reversals. As established in (ref), this is the mechanism through which microstructure noise enters our framework. Rough noise corresponds to this mechanism in its extreme form, with reversals that are frequent, short-lived, and large relative to the price level.
The results are reported in (ref). The dominant empirical pattern is the central role of short-run price reversals as a predictor of rough noise. The standard liquidity proxies studied by Ait2009, which they find to be closely related to the magnitude of microstructure noise, play a more nuanced role in our setting, a contrast we return to at the end of the section.
The strongest evidence comes from Roll spread. In the individual regression, it is positive and highly significant ($t = 7.60$) and explains $4.59\%$ of the variation. In the joint specification it remains positive and significant ($t = 3.12$), adding about $2.5$ percentage points of explained variation on top of the other regressors. Since Roll spread is a direct function of the negative first-order autocovariance of intraday price changes, this result points to short-run price reversals as the main empirical driver of rough noise, matching the mechanism emphasized by the model.
A closely related pattern emerges for $LOGTRADESIZE$. In the individual regression, average trade size is in fact the single strongest predictor of rough noise, with an $R^2$ of $6.43\%$, higher than that of Roll spread itself. Once Roll spread is included in the joint specification, however, the coefficient on $LOGTRADESIZE$ collapses from $0.23$ to $0.04$ and becomes insignificant. The interpretation is a mediation pattern. Larger trades consume more liquidity and generate larger temporary price impact Kyle1985, Hasbrouck1991, which mechanically appears as negative autocovariance in observed returns. Conditioning on Roll spread, the direct measure of these reversals, the explanatory content of trade size is exhausted. The fact that no independent trade-size effect remains is informative: it suggests that the relevant trade-size channel for rough noise operates through temporary impact (reversals) rather than through permanent impact (information), since the latter would not be captured by Roll spread.
Intraday volatility and lagged monthly volatility enter the joint specification with opposite signs: $\sigma$ is strongly negative ($t = -6.37$) while $MONTHVOL$ is positive and significant ($t = 3.23$). Rough noise is therefore most likely on days with low current intraday volatility within stocks that have been historically volatile.
The coefficient on $LOGP$ is negative and significant ($t = -2.89$ in the joint with Roll spread). The negative coefficient is consistent with rough noise being driven by reversals that are large relative to the price level, not by reversals in absolute terms.
The standard liquidity proxies $ILLIQ$ and $LOGNTRADE$ have no significant individual effect on rough-noise detection ($t = -0.69$ and $t = 0.92$). Neither variable directly captures the magnitude or frequency of short-run reversals; trading frequency in particular is necessary but not sufficient for reversals.
This pattern of results helps explain why the relation between liquidity and rough noise is less direct than the relation between liquidity and noise magnitude documented by Ait2009. Illiquidity may widen spreads, increasing the size of each reversal, but it may also reduce trading frequency and therefore reduce the number of reversals. The two effects can offset each other, leaving little net association between standard illiquidity measures and rough-noise detection. The empirical evidence here suggests that rough noise is best understood as a reversal-driven phenomenon, most prevalent on days when trading generates many large reversals relative to fundamental price movements, regardless of whether the underlying stock is liquid or illiquid in the broader sense.
In this paper, we argue that microstructure noise can be understood as mechanisms that induce price reversals and negative autocorrelation in high-frequency returns. With this interpretation, a trawl process is a natural choice for a parametric model of microstructure noise. Based on this insight, we work with a microstructural model for tick-by-tick price changes that explicitly separates permanent price changes from fleeting price changes due to noise. We show that this model converges to a standard semimartingale model for the permanent price process plus a rough noise term originating from fleeting price changes at the macroscopic scale. In this sense, we provide a microstructural foundation for the rough noise model suggested by ChongEtAl2021, based on a realistic tick-by-tick description. We further derive an explicit trawl function for the microstructural model, which gives rise to a Gaussian moving average with a gamma kernel as the limiting model for the noise.
We derive an estimation method for the tick-by-tick model based on the generalized method of moments and prove that the estimator is consistent. We further derive two limit theorems for the estimator: one where all parameters are in the interior of the parameter space, and one where the roughness parameter is allowed to lie on the boundary, corresponding to non-rough noise. Using these limit theorems, we construct a feasible test statistic for hypotheses on the roughness parameter, which in particular allows us to test whether the noise is rough.
Through simulation, we show that the model parameters can be estimated accurately in finite samples. A second simulation study shows that the proposed test has power against rough alternatives, and that its distribution under the null approaches the theoretical limiting distribution as the sample size increases.
Finally, we apply the methodology to tick-level transaction data from the TAQ database for the DJIA constituents in 2024. Motivated by the findings in (ref), we use trades from all available exchanges, since the presence of detectable microstructure noise depends strongly on whether prices are constructed from all exchanges or from a single exchange. We first identify stock-day pairs with evidence of microstructure noise, since the model parameters are not identifiable in the absence of noise. About two-thirds of the sample exhibits evidence of noise, while only a small fraction of the noisy stock-day pairs are classified as rough according to our test. The estimated roughness parameters on rough days are typically closer to zero than those found previously, a finding made possible by the use of daily data.
We then study which market characteristics are associated with roughness detection. Following the spirit of Ait2009, we regress a roughness detection dummy on liquidity, volatility, trading-activity, and transaction-cost proxies. In contrast to their finding that less liquid stocks have larger microstructure noise, we do not find that roughness is explained by broad illiquidity alone. Instead, roughness appears to be associated with short-run price reversals: it is more likely when reversals are large relative to overall return variation. Hence, rough noise is better understood as a reversal-driven phenomenon than as a simple consequence of illiquidity.
Overall, the empirical results support the central mechanism of the paper: fleeting microstructure effects at the tick level can generate rough noise at the macroscopic scale. They also suggest that the appearance of rough noise depends on the balance between trading frequency, reversal size, and efficient-price variation. Understanding the market-design and economic forces behind this balance is a promising direction for future research.