Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
183,377 characters · 20 sections · 54 citation commands
From orders to prices: A stochastic description of the limit order book to forecast intraday returns
\singlespace
{\let\relax}
\thispagestyle{empty}
\pagestyle{plain} \pagenumbering{Roman} \doublespacing \pagenumbering{arabic} \setcounter{page}{1}
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringIntroduction} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
Ever since Glosten94 raised the question whether 'the electronic open limit order book [was] inevitable', limit order books (LOBs) have become the most important way of trading, resulting in more than 75% of exchanges around the world (including the recently sprouting cryptocurrency exchanges) using order driven systems and relying (at least partially) on limit order books Jain03. Nevertheless, 'no comprehensive and realistic models (either statistical or economic) exist' Hasbrouck07 which describe the deep-rooted mechanisms of limit order markets in their entirety. While Hasbrouck07's statement is more than 10 years old and the literature has made significant progress, a dynamic model which comprehensively describes the interaction of individual orders and is able to incorporate the (strategic) behavior of market participants has, to the best of our knowledge, not yet been developed. GouldPWMFH13 provide an overview about research on the dynamics of the LOB and identify (roughly speaking) two branches of research. One of them originates in the field of physics and focuses mainly on idealized models to describe statistical features of the LOB system, focusing on dynamic order flows. The second branch is rooted in the economics literature which tends to treat order flows as static. According to GouldPWMFH13, economists primarily focus on the (strategic) behavior of traders, but neglect the dynamical structure. Of course, this reduction is too simplistic and there are multiple attempts to combine strategic behavior and the order book dynamics. For example, Parlour98 develops a stylized, dynamic model for the LOB and strategic order placement. Nevertheless, only few models approach the subject by heuristically incorporating statistical regularities observed in market microstructure and incorporate trader interaction based on these statistical observations. A notable exception is HautschH12 who use high-frequency cointegrated vector autoregressive models to shift the spotlight on the order flow of incoming orders and therewith, based on empirical analysis, draw attention to the intersection of LOB mechanics and strategic behavior of market participants. They show that the revelation of trading intentions through limit order placements affects the LOB states. Large07a shows how trades of different size (marketable incoming orders) affect future states of the LOB. Paulsen14 derives macroscopic limiting models for a microscopic LOB system in which bid, ask, and transaction prices drive the dynamics of the order flow. These limiting models serve as first order approximations of the stochastic processes that describe the system.
This paper aims to link statistical models of dynamic order flow and theoretical models which consider strategic interaction. Adapting ideas from Paulsen14 and inspired by Baez2018, we relate the LOB dynamics to the modeling of reaction networks BaezP17 and present a simple but comprehensive description of the microscopic order book mechanisms. Based on operator algebra, we construct the LOB dynamics bottom up from the elementary events of the book, namely the entry and exit of orders which enables us to model the LOB system by a Markov process. Furthermore, the Hamiltonian, i.e., the operator which governs the time dynamics of the system, can be constructed from these elementary events.\footnote{In economics and for that matter as well in optimal control theory, the Hamiltonian is know as a function that describes optimality conditions with respect to some control variable for some dynamical system. A popular example is the Ramsey-Cass-Koopmanns model in economics. In our model, the Hamiltonian directly describes the LOB dynamics and is based on the mechanical algebra of the LOB. In principle, it is therefore similar to the Hamiltonian used in classical mechanics.} Similar to Paulsen14, our approach allows to incorporate both the dynamics of the fundamental LOB mechanisms as well as the (strategic) behavior of market participants. Using the event log of the first quarter of 2004 from XETRA, we develop, simulate and empirically evaluate the implications of our operator formulation. By describing the interaction of individual orders in a purely statistical fashion, the state space of the order book system is worked out. It depends only on the rates of arriving and canceled orders as well as on the current state of the LOB. Based on a limited set of key variables, we show that the return dynamics can be approximated rather accurately by a linear model. This variable set also allows to forecast returns, arrival rates, and other measures such as order book imbalance and liquidity. Compared to other prediction models in the literature such as ZhouPHTZ18, we are able to reduce the root mean squared prediction error ($RMSPE$) decisively by a factor of 1/10.
Similar to our approach, ContST10 explore the idea of modeling the order book as a Markov process depending on the rates of arrivals and cancellations. They work out closed form solutions for the probability of an increase in the midprice, execution of an order at the best bid price (before a change of the best ask price), and execution of both a buy and a sell order at the best quotes before the price moves. ContD13 show that such a model can also be used to calculate the distribution of the duration between transactions. Unfortunately, as ContD12 state, these models are based on several assumptions which empirical research has shown to be incorrect BouchaudMP02,Hasbrouck07. In particular, the time intervals between order arrivals and cancellations are neither independent nor exponentially distributed and orders are not equally sized. Another model which is based on the Markov properties of the order book has been developed by DanielsFIS02 and extended by SmithFGK03. However, the model imposes broad restrictions on the functional structure, the parameters and the assumptions about the stochastic processes which govern the LOB dynamics. Again, some of their assumptions like equal order size or balanced order flow, order placement with uniform probability, among others are 'too simple to be literally true' SmithFGK03, but the resulting insights provide a useful foundation for LOB modeling.
In our model, we do not assume a specific functional structure for the stochastic processes of the order book. Instead, our model realistically describes the dynamics for any order size, with an arbitrary order placement probability and regardless of whether the order flow is balanced or not. We also allow the time interval between LOB events to directly depend on the state of the order book. While we maintain the assumption of an exponential distribution of the intervals, we only require them to be conditionally exponentially distributed.
The paper proceeds as follows. In (ref), we introduce the description of a typical LOB by an algebra of operators. Simultaneously, we present selected empirical characteristics of the order book which guide our model development. A key insight is that the LOB dynamics depends on the rates of order arrivals and cancellations at different bid and ask prices. (ref) presents the XETRA order book data and an analysis of observed arrival and cancellation rate distributions. In (ref), we report the results of a simulation study where we identify key drivers of order flow and the statistical distribution of events in the order book on short time scales. Finally, (ref) holds the empirical analysis of the XETRA LOB and (ref) concludes.
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringThe Model} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
The limit order book is the place where traders' orders meet. These orders carry information about a trader's willingness to accept a certain price, the limit price, in exchange for the chosen number of instruments, or vice-versa. The price level at which two orders are matched is called the reference price. Within the LOB, we distinguish buy (ask) orders and sell (bid) orders. Perceiving these two order types as species which populate price and order size levels inspired the use of the mathematical tools presented in Baez2018. At any point in time, the exchange keeps track of all orders within the LOB. Aside from the market side, the location within the LOB is defined by three key components: the limit price and the number of securities, both determined by the trader, and the time when the order arrives in the LOB.
If traders require immediacy, they rely on market orders which can be thought to have an infinite (bid order) or zero (bid order) limit price depending on the market side they were issued from. These orders are matched immediately and, normally, do not reside in the order book for an extended amount of time. They enter the LOB at the best price level of all limit orders currently residing in the LOB on the opposite market side.
The majority of orders, however, are designed to remain in the book for some time at their specified limit price level. These are called limit orders which (as we will show in (ref)) make up for roughly 90% of all orders in the XETRA LOB while only 3% are market orders. Limit orders have a well defined location in the price dimension. If their designated price location is behind the best price level of all limit orders which currently reside in the LOB on the other side of the market, they are matched (partially) before they can reach their designated limit price level. The smallest populated price level on the sell side of the market is the best ask and the highest price level on the bid side is the best bid. We will refer to one or the other as the best quote.
There are generally further order types that are only submitted to the market if certain conditions are met. For example, stop orders are inserted in the LOB once the reference price hits a certain threshold price. As soon as such an order enters the LOB, however, it is effectively equivalent to a market or a limit order. Hence, conditional orders can be perceived as more sophisticated versions of limit or market orders and can in principle be incorporated in the framework presented below. For the purpose of this paper, we restrict our considerations, thus, to plain market and limit orders.
In the following, we describe order creation and cancellation in the LOB as determined by the rules of a typical order book. For this purpose, we borrow the so-called Dirac or Bra-Ket Notation from physics where the state of a system is denoted by a ket $\ket{\psi}$. In (ref), we will discuss the underlying notion of a state in detail. For now, we can refer to any possible configuration of the order book with $\ket{\psi}$. Furthermore, we can assign weights (probabilities) to each possible configuration and refer to such a weighted bundle of pure states by $\ket{\psi}$. To begin with, however, we consider the empty order book (or vacuum) $\ket{0}$. From this vacuum state, more complicated order book states are created by successively acting on it with creation and annihilation operators. As we will see below, the rules of the LOB induce certain commutation relations in the algebra of these operators. It is convenient to also introduce the notation
which represents an empty ledger.
\let\theLOBsafehaven\theLOBsafehaven
For example, if $n$ ask orders are residing in the book, each with its associated limit price $k_i$ and size $q_i$, $i \in {1,\ldots,n}$, the order book state is given by
Analogously, for $m$ bid orders residing in the book with specified limit prices $k_j$ and sizes $q_j$ with $j \in {1,\ldots,m}$, the string of operators
describes the current state. Note that, so far, the rules only describe the successive submission of orders. In particular, we do not yet have a rule that would allow us to reorder the queue of creation operators. Put differently, creation operators generally do not commute: $a^+_{k,q} a^+_{s,p} \neq a^+_{s,p} a^+_{k,q}$. As a result, the strings of creation operators of ask and bid type are time-ordered.
Clearly, the probability of a cancellation must be zero if there is no order in the book. This means that when an annihilation operator acts on the empty order book, it generates a state with probability mass zero:
There is a standard argument that there is always one more possibility to create and then delete an object than deleting and then creating one. In terms of the operators this argument is represented by the commutation relation\footnote{ In physics, these commutation relations are known as canonical commutation relations.}
where $[A,B]:= AB - BA$ denotes the commutator of two operators. In fact this relation directly follows from (ref) and (ref)
Furthermore, since the cancellation of an order $a^+_{k,q}$ does not influence other orders $a^+_{s,p}$, we also have
whenever $s\neq k$ and $p\neq q$. We can summarize these algebraic relations as follows:
where $\delta_{ij}$ is the Kronecker-Delta, defined on an index set $\mathcal{I}\ni i,j $ by
In fact, these commutation relations are usually viewed as the defining properties of creation and annihilation operators.
In comparison to the commutation relation of ask orders, the order of annihilation and creation operators is reversed since bid orders act from the left.
When there are several identical limit orders, i.e., orders with the same price level and quantity, we can distinguish their position in the order queue by means of their time stamp. In contrast, for cancellation orders, an observer cannot predict which of the identical limit orders is supposed to be canceled. The algebraic formalism captures this uncertainty: up to normalization, the commutation relations lead to a stochastically mixed state that contains each possible cancellation, for example
\setcounter{LOB}{2}
The order book state is the result of successive order submissions and the corresponding string of operators is strictly ordered by time. Hence, we get a price-time ordering by rearranging the operators into groups with identical price level whilst maintaining the time ordering within each group. This can be achieved by letting ask and bid orders commute whenever they have different price levels $k\neq s$:
Using these relations, the order book state can always be written in the price-time ordered form
where $k_{i}\leq k_{i+1}$. Whenever $k_i = k_{i+1}$, the order nearer to $|0|$ was submitted first.
Given a LOB state in price-time ordered form, the priority of an order is encoded by its distance to $|0|$, where orders closer to $|0|$ have higher priority.
Recall for the case $q=p$ that orders of size $0$ are equivalent to the identity operator.
Since incoming bid (ask) orders always act on a state from the left (right), we need to commute them through older orders until they reach their designated position in the price-time ordered queue. This means orders automatically 'walk the book'\footnote{ In the LOB literature, 'walking the book' usually refers to an arriving, marketable order that is executed against several orders on the opposite market side. We borrow this notion of the walking order and extend it. In our case, every order 'walks through the book', however, only marketable orders encounter orders on their way to their destined limit price level. } until they reach their destined price level $k$. Along its walk, an order may encounter orders of the other market side and will then be executed as described by Rule (ref). It follows that market orders are described by creation operators $a^+_{k=0,q}$ and $b^+_{k=\infty,q}$, which will walk all the way through the book until they are completely executed.
Let us stress that the price level $k$ of an order $a^+_{k,q}$ is not necessarily the transaction price at which the order will be executed. Instead, the transaction price is usually determined by the price level of the 'settled order' that is encountered by the 'walking order'. The 'settled order', however, may depend on the trading mode (see (ref)). Also note that we did not yet specify when orders are executed and when a transaction will take place. The reason is that such rules do not add further algebraic relations. The question when orders are matched is not relevant for the description of the current state of the book. However, it is relevant for the time dynamics, i.e., if we examine time series of order book states. The question is whether matching occurs after each event, as continuous trading dictates, or whether matching is only conducted hypothetically after each event to produce indicative prices like in the pre-auction phase or at the end of an auction. These different modes matter for the evolution of the book. A detailed discussion of transactions and related issues follows in (ref).
As we have seen in (ref), at any given time the state of the order book is given by a price-time ordered string of creation operators
where $k_i \leq k_{i+1}$ for all $i\in S=\{1,\ldots, n\}$. Each operator in this string is specified by its market side $m\in \{a,b\} = \mathcal{M}$, the price level $k\in \mathcal{K}\subset \mathbb{R}$, and the order size $q\in \mathcal{Q}\subset \mathbb{R}$. Here $\mathcal{K}$ and $\mathcal{Q}$ are the price and quantity levels at which orders can be submitted to the LOB. These typically are discrete subsets of $\mathbb{R}$. In principle, price and quantity levels can become arbitrarily large, but in practice one can introduce a cutoff for both prices and quantities at a large enough value. As a result, $\mathcal{K}$ and $\mathcal{Q}$, can be thought of as finite sets.
It is convenient to introduce a partial ordering $\leq$ on creation operators via the following set of relations:
Then any price-time ordered string of creation operators is equivalent to a monotonically increasing map
from some finite set $S\subset \mathbb{N}$ to the set of creation operators. We denote the associated states by
The collection $\{\ket{z}\}$ fully describes the possible configurations of the order book at any given moment. We may refer to these states as pure states. Note that $S$, $\mathcal{M}$, $\mathcal{K}$, and $\mathcal{Q}$ are countable sets, so the set $\{\ket{z}\}$ is countable as well.
Clearly, any prediction of the future state of the order book must be probabilistic. So, while pure states are potentially observable, mixed states are not. Mixed states are composed of several pure states in which each pure state is weighted with some probability mass. For this reason, mixed and pure states are elements of the vector space spanned by the pure states $\ket{z}$:
In reality, the order book must be in a stochastically normalized state, i.e., a state ${\ket{\psi}=\sum_{\ket{z}} p(z) \ket{z}}$ where $0\leq p(z)\leq 1$ for all $\ket{z}$ and $\sum_{\ket{z}\in \mathcal{H}} p(z) = 1$. In this case, (the coefficient) $p(z)$ is the probability that the state $\ket{z}$ will be realized.
In (ref), we describe the time evolution of an initial state $\ket{z_0}$ at time $t_0$. We will see that (by construction) time evolution produces a state
In the literature, this object is usually referred to as generalized probability generating function Weber2017. For brevity, we often write $\ket{\psi(t)}$ and drop the reference to the conditional nature of the generalized probability generating function.
It is customary to denote linear functionals by 'bra' vectors $\bra{\psi}\in\mathcal{H}^\ast$ and introduce the dual basis $\bra{z}$, which satisfies
In this notation, the conditional probability to find a state $\ket{z}$ at time $t$ in $\ket{\psi(t); z_0,t_0}$ is given by
In this section, we introduce dynamics to the order book, i.e., we explain how an order book state evolves over time. Throughout this section, we closely follow Baez2018, where the general theory of stochastic time evolution is laid out in great detail.
The future state of the order book arises from acting on an initial state with the order operators introduced in (ref). This means that we are automatically in the situation of a Markov process.
The only issue is that the rate (probability) of incoming orders can depend on the history of the order book. It is, however, not sensible to assume that the entire history of the order book affects the properties of arrival and cancellation rates as old configurations of the LOB are usually not relevant for the decision process of market participants. They usually seek to maximize the probability of order execution based on the current state of the order book and possibly a very narrow history of preceding order book configurations. This only implies that arrival rates may be dependent on several preceding states of the LOB, which is not in contradiction to the Markov property per se. It would only mean that the process governing the LOB dynamics might be of Markov order higher than 1. However, theoretically, by appropriately extending the state space, every Markov process of finite order can be expressed as a Markov process of order one. Thus, we assume that the dynamics of the order book follow a Markov process of order 1.
As a continuous Markov process, the order book satisfies the Master equation vanKampen1992,Weber2017. In our notation, the Master equation is given by
where the so-called Hamiltonian operator ${H}$ encodes all information on the transition probabilities between order book states.
A solution of the Master equation is provided by a stochastic time evolution operator $U(t,t_0)$ via
If the Hamiltonian is time-independent, the time evolution operator is remarkably easy:
If the Hamiltonian is time dependent, the time evolution operator can similarly be written as
but the evaluation of this expression is typically more involved.
We assume that other variables, like news from outside the order book, may impact the rates of incoming orders. However, these variables are pre-determined outside the mechanism of the order book. This may lead to time dependent arrival and cancellation rates. In the system description, this would mean that the Hamiltonian is time dependent. Nevertheless, a time-independent approximation of such a system may still serve as a good approximation if the time intervals are small enough. In (ref), we will take an indeterministic approach towards those other variables and regard them as predetermined outside the book. Again, this is not in contradiction to the Markov property of the LOB system which is the key assumption for the Master equation (ref) to hold. At this point, it is also interesting to note that beyond the model presented in this paper, the LOB may be an open Markov process, which can be described by relying on a compositional model framework -- in the sense of BaezP17 -- and would allow to incorporate trader behavior.
Given the above considerations, the dynamics of the order book are fully described by the choice of a Hamiltonian $H$. Baez2018 show how $H$ can be constructed from infinitesimal stochastic operators which describe the elementary transitions that can take place in a system.
In the LOB, there are four possible transitions for each price level $k$ and each quantity $q$:
where $a^\pm_{k,q}$, $b^\pm_{k,q}$ are the creation and annihilation operators of (ref). The number operators $N_{k,q}^A = a^+_{k,q}a^-_{k,q}$ and $N_{k,q}^B = b^+_{k,q}b^-_{k,q}$ return the number of active bid and ask orders on price level $k$ and of quantity $q$ when they act on a pure LOB state (see (ref)).
As mentioned above, the Hamiltonian of the LOB is a combination of elementary transitions
where each transition is weighted by its arrival rate $\alpha$ or cancellation rate $\omega$, respectively. Also note that the arrival rates need to be scaled such that the time evolution operator ${U}(t,t_0)$ is indeed stochastic and maps one stochastically normalized state to another.
Generally, the arrival and cancellation rates in a LOB are observed to be time dependent. Intraday patterns of order flow have been documented for example by BiaisHS95. Even for the very recent development of international Bitcoin markets, in which trading is possible 24/7, ErossMU19 document activity patterns related to the opening and closing of major markets. The observed clustering of transactions in time can be conceived as the result of time dependent arrival and cancellation rates. These are usually modeled using Autoregressive Conditional Duration (ACD) models EngleR98,FernandesG06. We may treat the arrival and cancellation rates, especially on small and intermediate time scales, as mainly determined by the state of the order book, in the sense that the distributions of the rates across $k$ and $q$ depend on the current state of the order book -- for example via current best bid/ask or the spread. With this dependence on the current state, we allow for a quite general feedback mechanism between the current state of the order book and arrival and cancellation rates. If the state of the LOB changes by an event, the arrival and cancellation rates may change subsequently, as well. We investigate both static and dynamic specifications of arrival rates. In (ref), we will investigate static distributions using empirical unconditional frequencies and a uniform as well as a theoretical discrete Gaussian exponential (DGX) distribution for arrival and cancellation rates across price levels. The latter can be justified heuristically by the characteristics found in our data as described in (ref), in particular (ref). We will, in one simulation scenario, also allow the parameters of the assumed DGX distribution, for arrival and cancellation rates across relative price levels, to depend on the spread. In (ref), we measure the arrival rates during fixed non-overlapping time intervals and therewith allow them to vary over time. We also incorporate in our empirical analysis in (ref) the idea of conditional autoregressive arrival and cancellation rates and include lagged terms of arrival rates, moments of the spread and the distance to the opposite best quote. Sampling the LOB data on different time intervals, i.e., taking snapshots of the current state at different frequencies (for example 1, 2, and 5 minutes), allows to relate the moments of the relative integer distance $d_l$ (as defined in (ref) in (ref)) and the quantity of incoming and canceled orders $q$ to price changes and other observables of the system. Empirical tests of these implications can be found in (ref).
For now, we focus on the conceptual implications of these empirical findings and on how they affect the set up of the time evolution of the LOB system. Thus, we denote the arrival rate of an order at price $k$ and quantity $q$ as $\alpha_M(k,q)$, $M\in \mathcal{M}$. Since the distance to the opposite market side $d$ and the prevalent spread $\Delta$ depend on the current state of the order book, the arrival rates must be considered to be operators. When $\alpha_M(k,q)$ acts on a pure state $\ket{z}$, it returns an arrival rate which depends on the values of $d$ and $\Delta$ that are realized in the state $\ket{z}$:
A similar operator yields the distribution of cancellation rates corresponding to the current state of the order book:
In the Hamiltonian given in (ref) the operators $\alpha_M(k,q)$ and $\omega_M(k,q)$ act on the state first, thus determining the rate of the corresponding transition $E^M_{k,q}$ that acts on the state subsequently.
A specific configuration $\ket{\psi}$ of the order book contains an enormous amount of information. Usually, the focus lies on selected descriptive quantities which can be extracted from the order book at any state. We will call these quantities observables and describe them by the action of an operator $O$ on pure order book states $\ket{z}$. The value of $O$ for a given state $\ket{z}$ can be calculated as
More generally, given a state $\ket{\psi}$, the $\nu$th conditional moment of the observable $O$ is given by\footnote{In quantum mechanics, a similar relation holds, known as the Born rule $\bra{\Psi} \hat{O}^\nu \ket{\Psi}$. Since we work with stochastic probabilities (and not with quantum mechanical amplitudes), $\bra{\Psi}$ needs to be replaced by the sum over all dual basis vectors $\bra{z}$.}
Similarly, we can calculate the expected value of sums and products of distinct operators. This gives rise to covariance and correlation measures, e.g.,
Combined with the time evolution of an initial state $\ket{\psi_0}$, we obtain the moments of an observable's probability distribution at time $t$ ($t>t_0$) as
Note that the expected value in (ref) is a conditional expectation. It is conditional on the state $\ket{\psi_0}$ at time $t_0$. Later on, in (ref), to make this conditioning clear, we will denote conditional moments with ${\operatorname*{\mathbb{E}}}_{t_0}[O^\nu]$. The following example illustrates the rationale behind the formula in (ref). Consider the operator $\beta_B$ which extracts the value of the best bid order in a state $\ket{z}$ (cf.\ (ref)). Furthermore, assume that at $t_0 = 0$, the initial state is given in price-time ordered form by
Clearly, since $k_1 < k_2 <k_3$, the best bid is currently at price level $k_2$. According to (ref), the expected value of $\beta_B$ at time $t$ is given by an infinite sum. To begin with, we consider terms for which the state does not change during the time period $\Delta t = t-t_0$. These include of course the identity operator $1$ in Equation (ref). But $H$ is build from terms of the form $\alpha(p,q) (b^+_{p,q} -1)$, so there are additional contributions at any order in $H^k$. Together, they contribute the following term to the expected value of the best bid price only for price level $k_2$:
The expression in parenthesis is of course nothing but the probability that the state will not change within $\Delta t$.
For other price levels, these contributions are different. Therefore, we next investigate the linear terms in $H$ which describe the entry of a single order. There are only three cases in which the best bid changes. First, we may observe an entry of a limit bid order with a price level in between best ask and best bid. It's contribution to the expected value is $\sum_{k_2 < k < k_3, q} k \alpha_B(k,q) \Delta t$. Second, the order residing on $k_2$, the current best bid price level, may be canceled. In this case the contribution to the expected value is $k_1 \omega(k_2,q_2) \Delta t$. Third, an arrival of an ask order above the best bid which exhausts the best bid order's quantity may arrive. In this case we get $\sum_{k\geq k_2,q\geq q_2} k_3 \alpha_A(k,q) \Delta t$. In all other cases of incoming limit orders the value of $\beta_B$ remains at $k_2$.
A similar analysis is possible at order $H^2$. This would entail interaction terms of two orders entering the book: $a^+_{k,q} \alpha_A(k,q) b^+_{\ell,r} \alpha_B(\ell,r)$. There are again several different cases which depend on the type, price level and quantity of the incoming orders, each contributing differently to the expected value. More generally, at order $H^n$ one encounters the probabilities that $n$ orders enter the book and influence the best bid during time period $\Delta t$. For short time periods, such higher order contributions are negligible compared to the linear contributions because they depend on products of arrival rates, which are typically very small. However, for long time periods the powers ${\Delta t}^n$ will eventually dominate.
The example above illustrates an important property of the model. According to (ref), the moments of observables depend solely on the (current) distributions $\alpha_M(k,q)$ and $\omega_M(k,q)$. These distributions, if they vary sufficiently slowly, may be measured or modeled from the event stream of the book. Therefore, a testable hypothesis implied by our model is whether the distributional moments of $k$ (or equivalently $d_l$ for that matter) and $q$ (calculated by perceiving $\alpha_M(k,q)$ and $\omega_M(k,q)$ as their underlying probability functions) can be used to predict the expected value of observables, including price changes and inter-transaction duration.
In the following, we present a selection of observables, which are important for our analysis.
A basic observable is the number of active orders on price level $k$ with size $q$. It can be described for the bid and ask side by the number operators
These operators can be utilized to extract several other observables. In particular, the total number of active orders on price level $k$ of ask or bid type $M\in \mathcal{M} =\{A,B\}$
the quantity of active orders on price level $k$ and the total quantity on each market side $M \in\mathcal{M}$
or the volume of active orders at price level $k$ and the total volume on each market side
There are also operators that describe a global aspect of the configuration of an order book state $\ket{z}$, e.g. the best bid and best ask prices $\beta_M$, $M\in\mathcal{M}$. In the following, let the state of the LOB be
Then the best bid and best ask operators act on $\ket{z}$ as follows:
Note that $k$ on the right hand side is not an operator but the price level associated with the best quote. Combining the two, one obtains the spread operator $\Delta$ and mid price operator $\beta_\text{mid}$ as
Harris2003 defines liquidity as "the ability to trade large size quickly, at low cost, when you want to trade." According to the same source, the notion of liquidity incorporates four dimensions: immediacy of trade execution for a given size, depth, width, and resilience of the market. Therefore, the spread itself is used frequently as a liquidity measure in the literature.
There are multiple approaches to measure liquidity and we rely on the exchange liquidity measure (${XLM}$) which is based on the concept of implementation shortfall, introduced by Gomber2002. It covers three dimensions of liquidity: depth, width, and immediacy. The ${XLM}$ (also known as XETRA Liquidity Measure) is composed of liquidity measures for the ask side (${XLM}_A$) and the bid side of the market (${XLM}_B$),
where
The ${XLM}$ depends on the volume weighted price which can be realized immediately on each side of the market for a round trip order with a certain volume $\bar{V}$, i.e., simultaneously submitting marketable ask and bid orders with a total volume of $\bar{V}$. In other words, the ${XLM}$ measures the cost of a round trip order (in basis points).
Up to now, we have deferred the discussion of transactions since, strictly speaking, they are not necessary to set up the order book states. In this section we first discuss the trading modes of the XETRA order book and explain how one can augment the LOB states to also record information about transactions. This will allow us to study the transaction price and transaction rates, which were so far not available in the order book state.
The XETRA order book is organized as continuous trading augmented by opening-, intraday-, and closing-auctions. Before stating the rules for these modes, we make a small change in notation: Instead of the symbol $|0|$ for the empty book, we record via $|T_{k,q;t}|$ the last price $k$, quantity $q$, and time $t$ at which a transaction occurred. \let\theLOBsafehaven\theLOBsafehaven
\setcounter{LOB}{5}
The rules above are illustrated in (ref). We can now introduce the transaction price, transaction quantity, and transaction volume operators, which extract the corresponding numbers from the last recorded transaction. Let the state of the book be
The operators are then defined as follows
Furthermore, we can extract the time at which the last transaction occurred via $T_t \ket{z}= t \ket{z}$. This is the basis for an important observable, the inter-trade duration $T_{\Delta t} = t_2 -t_1$, i.e., the time between two transaction. In our current setup, $T_{\Delta t}$ cannot be expressed as an operator which only acts on the current state, while in practice we can calculate the time intervals from remembering earlier transactions' time stamps.
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringData} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
The data used in the present paper were provided by Deutsche B\"orse in 2004 and have been previously used by GrammigHR04. They consist of all recorded order book events of the XETRA system\footnote{XETRA Release 7.0} for trading of the 30 stocks that constituted the German stock market index, DAX between January 2 and March 26, 2004. Additionally,\ Deutsche B\"orse provided the open order positions in their books as of January 1, 2004, 12 pm. The data allow for the full recovery of the order book. Over the three months period, 228,275,832 events were recorded. Additionally, 2,282 initial positions are available at the beginning of the period. The data cover order arrivals, (partial) matches, changes and cancellations. The XETRA trading system allows for limit and market orders. It is also possible to mix the two standard order types (market and limit orders) with a market-to-limit order (MTL). An MTL order is filled on the best limit price in the book, either fully or partially. If an MTL order is matched only partially on the best ask or bid price, it enters the LOB on the best limit price with the remaining order size. Additionally, iceberg orders (ICE) are allowed for which only a fraction of the total volume chosen by the issuer is displayed to market participants.
Market and limit orders can also have a specified stop price. They are then called stop orders. Different from the limit price, i.e., the price upon which the trader is willing to trade, the stop price specifies a price level from which onwards the trader is willing to submit an order. Hence, if the reference price exceeds (in case of bid order) or undercuts (ask order) the stop price, the order is inserted in the book.
Furthermore, the XETRA system also allowed so called XETRA BEST execution orders during the sample period at hand. The BEST execution orders are matched against incoming market or crossing limit orders at a price level just before the currently prevailing best ask and bid prices. In that way they introduce an extra hidden layer of occupied price levels in front of the prevailing best ask and bid prices in the LOB. XETRA BEST orders can be market or limit orders. Market-to-limit orders, iceberg orders and stop-orders cannot be submitted as XETRA BEST orders. Also, if XETRA BEST orders are not executed they do not enter the LOB event log at all. This happens if the associated price level is better than the current best quote. If they are matched in a trade, they are recorded as the counterparty of a transaction. If they are submitted as limit orders and the associated price level is worse than the best quote, the enter as regular limit orders. In the latter case, we cannot distinguish regular limit orders from XETRA BEST orders in our data.
For all orders validity constraints can be set. Users may specify a termination date up to which the order is valid. Without such a restriction, orders are valid for 90 days. Iceberg orders, however, are only good for the day. Traders may also specify whether orders are valid only for specific trading phases such as auctions, or during which auction the order shall be valid. Orders with such a restricted validity reside in the book and become active during the trading phase for which they are valid.
With regard to execution restrictions, the XETRA system allows for the following two specifications for market, limit or MTL orders. First, the Fill-or-Kill order (FOK) is either filled entirely or canceled. FOK orders are only recorded as entries to the book if successfully filled. If no immediate filling is possible, FOK orders are canceled without notification within the LOB event log. Second the Immediate-or-Cancel order (IOC) is filled as far as possible upon entry, or canceled. Similar to the FOK order, a record of the order is only entered in the LOB event log in case of a successful (partial) filling.
All these different order types and restrictions can be incorporated in the time evolution set out in (ref) by introducing different types of arrival and cancellation rates for the related events as elements of the Hamiltonian. These order types, however, do not affect the generality of the algebra set up in (ref).
(ref) in (ref) provides an overview of the distribution of the events in the LOB log. It presents the total number of submitted limit, market, iceberg, and MTL orders along with their relative occurrences on the bid and ask side.
In empirical investigations of LOB data, it is frequently observed that the distributions of arrival and cancellation rates show a relatively stable connection with the current distance $d = |\beta_M - k|$ of the respective price level $k$ to the best active price level on the opposite market side $\beta_M$, $M\in\mathcal{M} = \{A,B\}$ BouchaudMP02. We find a similar phenomenon in the XETRA data. We set the distance $d$ of all crossing, arriving orders to 0 and define the logarithmic relative integer distance as
Taking logarithms exposes the heavy tails of the distribution across price levels and its similarities with the DGX-distribution (see (ref) for a short description and BiFCK01 for further details). For our simulation study in (ref), we use the DGX distribution to model order arrivals across $d_l$. (ref) displays the empirical logarithmic frequencies of order arrivals and cancellations at the logarithmic relative integer distance $d_l$ of their limit order price to the best ask or best bid price, respectively. The red line indicates a DGX distribution (truncated at 1). The logarithmic frequencies of different types of marketable orders (limit, stop, iceberg and market orders) are displayed separately at $d_{l0}=0$ in (ref).
In our data, we also observe a small correlation between $d_l$ and the prevailing spread $\Delta = v_A - v_B$ in logarithms. (ref) presents scatter plots of $d_l$ against the logarithmic integer spread $\log(100\Delta)$. For large spreads there is a stronger correlation between events that are mainly concerned with price levels around the best limit price of the same market side. When the spread is small, it seems that events occur more evenly spread out up to as much away as EUR 4 from the best limit price of the opposite market side. We also note that events on the bid side are less dispersed across price levels than events on the ask side as the natural limit price for a bid order is a price level of 0. For ask orders, theoretically, no such limit exists.
With respect to order size, we observe no clear pattern between order size and the logarithmic relative integer distance in the data, only a very small negative correlation can be noted. Budget restrictions of market participants would suggest that order sizes further away from the opposite best quote on the ask side bring order sizes down with growing price levels. For the bid side, the inverse argument should hold, nevertheless, a small negative correlation can also be observed for the bid side. However, as the correlations are low, the price level and quantity, at least in logs, may justify the approximating assumption of stochastic independence used in several scenarios of the simulation study in Section (ref). (ref) lists the correlations between the order volume $q$ associated with a certain event and the logarithmic integer distance to the best opposite quote for the MEO stock. As can be seen, the approximating assumption that the two variables $q$ and $d_l$ are independent does not mirror reality exactly. Nevertheless, we will make the assumption several times in this paper as it keeps the estimation and simulation manageable.
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringSimulation Study} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
In order to simulate the LOB, several assumptions have to be imposed on the functional structure of the arrival and cancellation rates included in the Hamiltonian $H$ of the model developed in (ref). The following section gives a short overview on the calibration of the stochastic simulation algorithm (SSA) that we use, which was developed by Gillespie1977. With the algorithm, we simulate an artificial history of the order book. (ref) presents the theory and the successive steps of the algorithm in more detail. The possibility to exactly simulate the system with the SSA is a direct consequence of the model. Note that the assumptions made about the immanent functional structure of the arrival and cancellation rates are the crucial ingredients of the model. We therefore explore possible calibrations of our model in the subsequent simulation study. Our goal is not to fit the simulation results to an observed LOB history as closely as possible. Nonetheless, our choice of parameters will often be guided by empirical observations. The simulation, however, has the purpose to offer insights into the sensitivity of the order book dynamics, especially with respect to transaction price dynamics, when the structure of arrival and cancellation rates are changed.
We consider three theoretical probability mass functions for the arrival and cancellation rates across price levels (denoted by $p_{K,M}(\cdot)$): The uniform distribution (uni), a discrete log-normal distribution with fixed parameters (fix), and a discrete log-normal distribution which depends on the prevailing spread (dyn). For the probability distribution across volume levels (denoted by $p_{Q,M}(\cdot)$), we only consider a power law distribution (pow). The power-law distribution is empirically motivated to capture the heavy tails of the volume distribution. The distribution of order size and the heavy tails can be seen in (ref) which depicts the frequencies of order arrivals and cancellations for an exemplary stock in our dataset, the retail company Metro (MEO).
(ref) lists all the functional specifications as well as a description on how market orders are incorporated in the distributional setup. Iceberg orders, stop orders, or fill-or-kill restrictions are neglected in the simulation study, as the events marked by these order types only make up for less than 1% of all events in our data set.
Additionally, we also investigate cases in which $p_{K,M}(\cdot)$ and $p_{Q,M}(\cdot)$ are described by the empirical univariate frequency distributions in our sample across $k$ and $q$, respectively (emp). We also utilize the joint frequency distribution of the observed pairs $(k,q)$ in one scenario (emp,emp). Note that although we use the empirical frequencies, the rates are fixed over the entire simulation run. Thus, in the scenario 'emp', no dynamic feedback between the state of the book and the arrival and cancellation rates is introduced.
For all combinations of these distributional specifications (in total 8 scenarios\footnote{The scenarios are (uni,pow), (uni,emp), (fix,pow), (fix,emp), (dyn,pow), (dyn,emp), (emp,pow) and (emp,emp) where the list of pairs utilizes the introduced abbreviations and states the distribution across $k$ in the first coordinate and the one across $q$ in the second coordinate.}) for each stock, we simulate 200 realizations of LOB evolutions over half a trading day (4 hours). The state at the end of our sample, i.e., after the opening auction on March 31, 2004 at 9h00 CET, serves as a starting point for the simulation. We also ran the simulations for a different starting point, namely the positions after the opening auction at 9h00 CET of January 2, 2004. The results are not qualitatively different for the two simulation setups.
We first turn to the results from the uni scenario which are depicted in (ref). In this scenario, a lot of simulation runs ended with an empty order book, indicated by the sudden stop of the individual (red) price paths. Furthermore, the variance of transaction prices, which is induced by the uniform distribution, is rather high, especially, when using the empirical volume distribution. This can be seen in Table (ref) which shows the average mean and standard deviation of the simulated logarithmic transaction price changes across all simulation runs for all the different scenarios. The 'emp/pow' as well as the two 'uni' scenarios result in a high standard deviation for the simulated returns. Compared to the observed empirical values, the simulated values in these scenarios are implausibly high. The average of the time series means and standard deviations of the observed logarithmic transaction price changes across all thirty stocks for the same time period are displayed in the last row in (ref). For the time after the opening auction (9h00 - 12h00) on January 2, 2020 the values are $\bar{\mu}_{\text{emp}} = 0.02\cdot 10^3$ and $\bar{\sigma}_{\text{emp}} = 0.67\cdot 10^3$.
Note that for these simulations, the average event rates on each market side (which is denoted $\bar{r}_{0,M,i,j,\cdot}$ in (ref) in (ref)) are the same as in the case of the fixed and dynamic arrival and cancellation rates. We may associate the uniform distribution across price levels with somewhat uninformed traders who, regardless of the price, randomly submit orders in the vicinity of the current best quote. With the uniform distribution, the mean and variance of the price level is rather high compared to the DGX specifications in other simulation scenarios as well as the empirically observed equivalents. Throughout all simulation scenarios, we see that a higher mean and variance in the distribution across price levels are related to a higher variance in transaction prices. This result is interesting since, as noted before, arrival rates are linked to trader's behavior. So, if traders are uniformed on how the asset should be valued and constantly shift their valuation with no clear tendency and/or if traders are indifferent between immediate execution and delayed execution, transaction prices become highly volatile.
Second, the results for the fixed DGX distribution across price levels are presented in (ref). We can see that the power law very rarely induces large jumps in transaction prices due to the extremely large order sizes that are possible under this distributional scheme. In general, however, differences between the volume distributions are not obvious, neither in the 'uni' scenario nor in the 'fix' scenario. The average of the time series means and standard deviations for the simulations with the fixed DGX distributions as provided in (ref) are close to the empirical ones. One very interesting result for the 'fix' scenario concerns the parameters of the DGX distribution. The distributional parameters $\mu$ and $\sigma$ are almost identically defined: The parameter $\mu$ of the DGX distribution for incoming and canceled orders on the bid side has been specified slightly higher (at $\mu_{B,a} = 1{.}766$ and $\mu_{B,c}= 1{.}674$) than the one for the ask side ($\mu_{A,a}= 1{.}726$ and $\mu_{A,c}= 1{.}620$) to match estimated parameters from empirically observed frequencies. However, this small difference, does not seem to have any effect. In order to analyse the effect, we ran the simulation of the 'fix' scenario only for the stock MEO again with adjusted parameters: Leaving the distribution of the cancellation rates untouched, only the parameters of DGX distribution across arrival rates are altered to $\mu_{B,a} = 0{.}1$ and $\mu_{A,a}= 2$ as well as $\sigma_{B,a} = 0{.}2$ and $\sigma_{A,a}= 1$. Therewith, the distribution of bid order arrivals is much more dense around the best ask price. The result can be seen in (ref).
The behavior of the transaction prices also depends crucially on the cancellation rates as well as the volume distribution. In the case where order size is distributed according to a power law distribution, order sizes on both market sides are identically distributed. Furthermore, the power law distribution generates a lot of small orders while large orders are very rare. At the same time, the initial position on the bid side contains several large orders close to the best bid price (which are more likely to be canceled). The ask side consists of several medium sized orders close to the best ask, while the large orders rest deep in the book. So, the frequent small orders inserted at or close to the best ask price are not able to move the market upwards permanently due to the medium sized orders sitting in the book on the ask side. At the same time large orders at the front of the bid side (from the initial position) are canceled frequently. One rare large ask order generated by the power law, thus, is able to move the bid price quite a lot. The longer the simulation is running, the more likely it is for a large ask order to occur and the more likely it is that the large orders at the top of the bid side are already canceled. This makes it easier for the ask side to move the best ask down and therefore transaction prices deteriorate.
In the empirical distribution, the volume distribution on the bid side dominates the volume distribution of the ask side. Thus, bid orders inserted into the book are larger in size than inserted ask orders. Since bid orders are inserted close to or at the best ask price level, the bid side moves the market soon after the simulation starts upwards. This leads to a hefty upward drift of transaction prices in each simulated transaction path and empty order books on the ask side. Hence, we can conclude that limit order distributions with a high probability mass in the vicinity of the best quote in one market side push transaction prices in the direction of the opposite market side if the order size distribution allows for frequent medium sized orders.
Third, the simulated transaction prices resulting from the scenario with dynamical shifting and scaling DGX distributions which depend on the prevailing spread, are shown in (ref). In this scenario the moments of the DGX distribution depend on the prevailing spread. The dynamical adjustment of the arrival and cancellation rates across price levels is balanced, So, the mean and standard deviation of the DGX distribution across price levels is the same on both market sides. Furthermore, the large jumps induced by the power law distribution are still rare, but rather pronounced. These jumps are, however, not sufficient to cause an increase in volatility. In fact, the scenarios with a power law distribution exhibit on average a slightly smaller volatility which might be due to the fact that the order size is rather small. For the power law distributed order size, the time series mean and standard deviation of the dynamically adjusting simulation scenario are close to the time series mean and standard deviation of real observed logarithmic transaction changes as presented in (ref). The scenario with the empirical order size distribution is too volatile.
Fourth, using the unconditional empirical frequency distributions as well as the empirical rates $\bar{r}_{0,M,i,j,\cdot}$, the results presented in (ref) deviate from empirical stylized facts in that the resulting paths are much more volatile. The imbalance between ask and bid order arrivals and cancellations exhibited by some stocks, together with the observation that on average ask orders arrive closer to the best bid than bid orders to the best ask, drags transaction prices on average slightly down (see (ref)). The smaller distance to the best quote, on average, of ask orders can be seen in (ref) as well as in (ref). It seems that in our sample, sellers tend to seek quicker order execution by placing their orders close to the buy side. Buyers on the other hand, test their fortunes and patiently wait for a good deal to occur deeper in the book. This can be clearly seen in (ref): For the same average spread, arriving bid orders are placed on average further away from the opposite market side than arriving ask orders.
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringEmpirical Analysis} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
In our empirical analysis, we focus on three variables that we deem most important to market participants. The first is the logarithmic return which could be achieved based on a buy-and-sell strategy in subsequent intervals. The second is the return of a sell-and-buy strategy. The third one is the exchange liquidity measure (XLM) as introduced in (ref), calculated with a round-trip of EUR 100{.}000. To implement them, we sample the data in intervals of fixed length $\Delta t$ (1, 2, 5, 10, 15, 30, 45, 50, 60, 120, 240 minutes). Let $t$ denote the last point in time of some arbitrary interval and $t-1$ the last point in the previous interval. Then the logarithmic returns of a buy-and-sell ($\Delta p_{t,b}$) and a sell-and-buy strategy ($\Delta p_{t,s}$) are given as
For the $XLM$ the last observation in the respective interval is taken.
We have shown in (ref) that the moments of some observable $O$ of the LOB system can be expressed as
We can perceive the right hand side of this equation as an intricate function of the arrival and cancellation rates. These rates depend on the event rates $\bar{r}_{0,M,i,j,e}$ of arrivals and cancellations of the various order types, the relative logarithmic integer price level $d_l$, the order size $q$, the spread $\Delta$, and possibly other variables. Thus, we can formulate a linear approximation for the expectation of any observable in the moments of exactly these variables.
By repeated Taylor series approximations of the terms in (ref) in the variables $d_l$, $q$, and $\Delta$ around their respective mean, collecting terms, taking expectations and neglecting terms with an order higher than four (see (ref)), we get for the expected value of some observable a linear approximation of the form
where $\mu_{x,t} ={\operatorname*{\mathbb{E}}}_{t_0,M,i,j,e}[x_t]$. Note that the indices in the subscript of the expected values indicate their conditioning set as outlined in (ref). Hence, the tuple of subscripts $(t_0,M,i,j,e)$ indicates that the conditional expectation is formed with information available at time $t_0$ for stock $i$ given the market side $M$, order type $j$, and event type $e$. As we are not interested in $\gamma_{0,i}$, $\rho_{0,M,i,j,e}$, $\kappa_{M,i,j,e,v}$, $\xi_{M,i,j,e,v}$, and $\delta_{i,v}$, we can collect all possible products in (ref) in the parameters $\gamma_{0,i},\gamma_{1,i}, \ldots \gamma_{1,R}$. In total, this yields 1{.}971 parameters. Hence, the estimation of the parameters is sensibly feasible using ordinary least-squares up to non-overlapping intervals with a length of 15 minutes. For non-overlapping 15 minute intervals, we can get 2{.}164 observations from the 64 trading days in our sample. Increasing the interval length beyond 15 minutes makes the use of overlapping intervals and rolling variable calculation necessary. While this is in principle feasible, we restrict the analysis for the specification in (ref) to non-overlapping intervals and sampling frequencies below 15 minutes.
To increase the length of the intervals, we use three alternative specifications which entail less parameters. In the first alternative, the moments of the spread are only included additively:
This specification entails the estimation of 399 parameters. This enables us to estimate the model on interval lengths of up to one hour.
In the second alternative, the moments of the spread and the moments of the order size are included additively:
This reduces the number of parameters to 95 and makes interval lengths of up to 4 hours possible.
In the third alternative, we employ a completely additive structure:
This reduces the number of parameters which have to be estimated further to 43.
For each of the four specifications in Equations (ref) to (ref), we use a formulation in which the moments on both sides of the equations are estimated contemporaneously, i.e., at the same time $t$. Naturally, this is only possible in-sample. Additionally, we also investigate specifications of Equations (ref) to (ref) in which the moments on the right hand side are estimated at $t-1$ to describe the expectation of the observable on the left hand side. This formulation allows for out-of-sample evaluation of the model in terms of predictive power assuming that the moments in one interval anchor the moments of the subsequent interval.
We evaluate our model across several measures in- and out-of-sample.
For the in-sample evaluation, we consider the adjusted and unadjusted $R^2$ as well as the root mean squared error ($RMSE$). Furthermore, following ZhouPHTZ18, we also use the direction prediction accuracy ($DPA$) defined as
The in-sample results are depicted in Figures (ref) to (ref). In each figure, the contemporaneous models are depicted in subfigures (a) and (b) for returns and (e) for the $XLM$ measure while the results for the specifications which use only past information to model the current state are presented in subfigures (c) and (d) for returns and (f) for the $XLM$ measure. Every line in the graph represents the results for one stock. The highlighted thicker line is the average of the respective measure across all stocks. There are a few noteworthy results. First, the DPA as well as the $R^2$ measures (adjusted and unadjusted) allow us to reject the hypothesis that the contemporaneous as well as the lagged models have no significance in explaining the data. Our results suggest that the extensive model in (ref) overfits the data with growing sampling frequency and, hence, less observations. This can be seen by the drastically increasing $R^2$ and $DPA$ values when the sampling frequency is increased. The adjusted $R^2$ should account for this effect, and indeed remains rather stable. However, when the degrees of freedom of the model become sparse, the adjusted $R^2$ is not able to correct the full extent of the overfitting. Nonetheless, the high values at the highest frequencies indicate that the contemporaneous model describes high-frequency returns very well. In all measures, the contemporaneous specification in (ref) (blue) turns out to be superior to the specifications in (ref) (red), (ref) (green) and (ref) (orange) , i.e., it has a higher direction prediction accuracy, better fit in terms of higher $R^2$ values, and results in a lower root mean squared error.
Subfigures (c), (d), and (f) in Figures (ref) to (ref) present the results using the lagged specifications of Equations (ref) to (ref), i.e., they compare their out-of-sample forecast performance. Recall that these specifications rely heavily on the assumption that the arrival rates stay constant for some (very) short time horizons. Therefore, the results regarding the performance of the different models turns out to be different compared to the in-sample evaluation above. Now, the specifications in Equations (ref) and (ref), which use far less interaction terms and have a rather small number of parameters, capture the dynamics of returns almost as well as the other two specifications in Equations (ref) and (ref) which becomes apparent, for example, when looking at the $DPA$ in Figure (ref). Only considering the adjusted $R^2$ (Figure (ref)), the specification in Euqation (ref) has a better performance. It also slightly decreases the $RMSE$.
For the $XLM$ measure, the performance of the highly parameterized models is better. In our opinion, this is due to the following two reasons. First, the $XLM$ measure changes with each event, so that on a high frequency the volatility of the $XLM$ measure is higher than the volatility of returns. The second reason is rooted in the definition of the $XLM$ measure as a fraction of observables which in general requires a higher polynomial degree to arrive at a sensible approximation. The more complex models in Equations (ref) and (ref) are, thus, better suited to provide such an approximation. Note also that we take the last observation for the $XLM$ within an interval which is highly volatile. From a trading perspective, the average $XLM$ over the interval may be better suited to describe liquidity of the market during the interval and might be a less volatile measure. Nevertheless, we chose the more volatile measure for our analysis. Our results should therefore pose a lower bound with respect to the modeling accuracy.
Nevertheless, for intraday returns, the relatively parsimonious models in Equations (ref) and (ref) perform remarkably well. Compared to other studies like ZhouPHTZ18 who use high-dimensional neural networks (without contemporaneous information), we find a decisively smaller $RMSE$. For example, the out-of-sample forecast error reported by ZhouPHTZ18 for their best model (GAN for minimizing forecast error loss and direction prediction loss with training sets with a length of 20 days and test sets with a length of 5 days) is 0.0079 with a DPA of 69% on a 1 minute interval. Even though the in-sample results in our case are not exactly comparable, we can note that for the models with lagged information on 1 minute intervals an average $RMSE$ of 0.0010 can be achieved together with a DPA slightly above 85% on average. We also see that, on high frequencies, the models with lagged information all result in similar $RMSE$ and $DPA$ values. Again the results for $R^2$ show that overfitting is a problem for lower sampling frequencies. These only differ in $R^2$ (adjusted and unadjusted). Nevertheless, the size of the adjusted $R^2$ of around 15% on 1 minute intervals is remarkable.
All models perform better at intervals of 45 minutes than on other frequencies. This is rooted in an above-average precision of the prediction of the return in the last interval of the day. When using 45 minutes intervals, the last interval of the day is only 30 minutes long. Hence, estimated moments based on the observations within the previous 45min interval are used to predict the last, shorter, 30min-long interval. If these shorter end-of-day intervals are not excluded from the data, the model performs better compared to other frequencies. This suggests another avenue for further research: Can the model performance be improved by using longer intervals to estimate sample moments to then make predictions on smaller time horizons. Put differently, is it in general beneficial to calculate moments on the last hour of data to forecast the next ten minutes? The result for the 45min interval suggest there may be some benefits. However, answering the question would entail a rolling calculation of means on overlapping intervals which is beyond the scope of the current paper.
We also evaluate our models based on out-of-sample predictions using rolling windows. Depending on the size of the non-overlapping intervals, we vary the number of intervals included in the rolling window.\footnote{ Recall that when we talk about using non-overlapping intervals, we mean the procedure to use the observations associated to the last interval of the rolling window to predict the observation of the next interval. Alternatively, one could use overlapping-intervals, i.e., update the observation at each new event to predict the next interval. However, this procedure is computationally very demanding and, therefore, out of the scope of the current paper. } (ref) presents the lengths of these windows for the frequencies considered.
In addition to the root mean squared prediction error ($RMSPE$) and the out-of-sample $DPA$, we also use the $R^2$ of a Mincer-Zarnowitz regression MincerZ69 to evaluate the model.
Results are reported in Table (ref) for selected 1 and 5 minute intervals. As can be seen, for the return series, the precision of the out-of-sample prediction is remarkably high. On a 1 minute frequency, we are able to predict the direction of the next price change with an average accuracy at or over 80%, irrespective of the model. But even on the 5 minute frequency, the accuracy only falls slightly below 75%. For both buy and sell returns, we deem the $R^2$ rather high and the $RMSPE$ rather low given that we intend to predict financial returns on a ultra-high frequency. \footnote{ For the stock ADS, we also observe a 100% accuracy when predicting the direction of the next price change. Also the $R^2$ of the Mincer-Zarnowitz regression is around 50%. However, it needs to be mentioned that for this stock, only 19 out-of-sample predictions are made in total. Since, based on 1 minute intervals, some moments that enter the right hand side of Equations (ref) to (ref) cannot be calculated, we drop the observations for these intervals from our sample. In effect ADS has 10{.}019 valid observations on a 1 minute frequency. Due to such missing values also out-of-sample 1 minute results for DB1, FME and HEN3 are not reported since less than 10{.}000 observations are valid for these stocks. This is why we do not include the 1-min out-of-sample results for ADS in the figures. }
For the $XLM$, results are somewhat different. The precision of the forecast is poor which is in line with the in-sample results. Again, this is due to the high variability of the measure and its structure. The $XLM$ changes with each event and is a highly nonlinear function in the arrival rates (see Equations (ref) to (ref)). Therefore, the linear approximation may be poor and the approximation for longer time horizons may be especially poor. In this line of argument, it is worth mentioning that the model with the highest complexity in our considerations (specified in (ref)) performs best in all evaluation measures.
The results for all stocks and all frequencies are presented in Figures (ref) to (ref). As we can see the smaller the interval, the better the forecasting ability of all our linear models. The extensive linear approximation in (ref) predicts the direction of the returns very well on small intervals. The $R^2_{MZ}$ of the Mincer-Zarnowitz regression of above 2% for the sell strategy is above what we had expected for returns on ultra-high frequencies. On intervals longer than 10 minutes, the predictive ability of all three linear approximations is, however, poor. It can also be noted, that the sparse model formulations in Equations (ref) and (ref) perform just as well, or even better in some situations, than the heavily parameterized formulation in (ref) in the case of the return series. This is not true for the $XLM$. For the $XLM$, the more complex formulations in Equation (ref) and (ref) perform better in all measures. Especially, the constant $RMSE$ and the increasing $DPA$ and $R^2_{MZ}$ up to 5-min intervals are remarkable. The variance of the $DPA$ in (ref)c shows how noisy the $XLM$ and the associated forecasts are and by how much the more complex model is able to reduce this variability.
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringConclusion} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
In the paper, we propose a model for the limit order book which describes the dynamics in the book by a continuous Markov process and allows to forecast returns. The mathematical formulation is based on the operator algebra which we borrow from physics. Our model closely describes the reality of the order book and identifies the arrival and cancellation rates as the key ingredients of the book's dynamics. By means of a simulation study, we show that the distribution of order arrival rates across price levels determines the shape of the book and, as a consequence, the transaction price evolution. Varying the type and shape of arrival and cancellation rates across prices and volume, we find that the moments of price levels and quantity levels of incoming and canceled orders are important determinants for the evolution of the book.
In an empirical study which is based on a linearized version of our model, we estimate three different specifications on non-overlapping intervals of various frequencies. As we have the entire record of the XETRA order book for 3 months at our disposition, we can include a large number of parameters in the estimation such that an evaluation of the model becomes feasible. In-sample, all considered models exhibit a good fit in terms of $R^2$, $RMSE$ and direction prediction accuracy (DPA). Our fully parameterized model seems to overfit the data on lower frequencies. Nevertheless, when using only past information, the values for the adjusted $R^2$ range for the minute-by-minute intervals around or over 10% whereas the direction is correctly predicted in around 70% of all cases.
To evaluate the robustness of our results we also conduct an out-of-sample test of the model. We use one-step-ahead forecasts on various frequencies and evaluate the accuracy with the $R^2_{\text{MZ}}$ of a Mincer-Zarnowitz regression, the $DPA$ as well as the $RMSPE$. We find an extraordinarily good fit for ultra-high frequency returns as the $R^2_{\text{MZ}}$ is generally above $2\%$. In addition, on low frequencies, the $DPA$ is around 80% which suggests that we can predict the directional change of the next return very precisely. The time varying estimates of the parameters as well as the short forecasting horizon make the return prediction astonishingly accurate. We also try to predict liquidity at the end of each interval with the $XLM$ measure. The measure cannot be forecasted well for longer time intervals with adequate accuracy. Even on very short time horizons, the best fitting model is barely able to predict the direction of the next change of the $XLM$ in more than 50% of the cases. This result is rooted in the definition and the very volatile nature of the $XLM$ measure.
On the basis of the event log of XETRA for the first quarter of 2004, we have, nevertheless, shown that our model describes the LOB data well, both in- and out-of-sample. The data requirements are rather high as knowledge about price and quantity levels of incoming and canceled orders are required. This sort of data is usually not available. Even though returns may be predicted, market impact of actual trading strategies as well as order costs may hamper profitability of a trading strategy based on our model. Nevertheless, we are convinced that our empirical analysis provides new lower limits of forecast accuracy, as we have made several approximating decisions in the course of this paper. In addition, for time horizons beyond 1 minute other variables may possibly help to predict returns or any other measure in the order book.
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef\Appendix\sAppendix *{Acknowledgments} The authors acknowledge support by the state of Baden-W\"urttemberg through bwHPC.
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef\Appendix\sAppendix *{Data availability statement} The data that support the findings of this study are available from the corresponding author upon reasonable request.
\onehalfspace \addcontentsline{toc}{section}{References}
\addtocontents{toc}{\null Appendices} \renewcommand \thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef\Appendix\sAppendix } \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
\setcounter{section}{0} \setcounter{figure}{0} \setcounter{equation}{0} \renewcommand\Alph{section}{\Alph{section}} \renewcommand\Alph{section}\arabic{equation}{\Alph{section}\arabic{equation}} \renewcommand\Alph{section}\arabic{figure}{\Alph{section}\arabic{figure}}
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringDistribution of events} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
{\singlespacing
}
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringSimulation Specification} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
This appendix presents the theoretical and algorithmical details of the stochastic simulation algorithm (SSA) used to simulate an artificial history of the LOB.
The starting point of the SSA is based on the probability that within the next interval $\delta\tau$ no event occurs which we can denote in our notation as
where the diagonal elements in $H$ are obtained by
This is the negative sum of the rates of all possible events conditional on the book being in state $\ket{z}$.
Gillespie1977 shows how to formulate this probability for some event $\mu$ to happen during the interval $\tau$ without the operator algebra. First, the probability that an order arrives during the interval $d\tau$ is $r_\mu d\tau$, where $r_\mu$ is the rate corresponding to the event. In our case, $r_\mu$ may be some rate from the set of arrival or cancellation rates, $\alpha_M(k,q)\delta\tau$ or $\omega_M(k,q)\delta\tau$. In fact, we may label all possible events with integer numbers and let $\mu$ be a specific integer denoting a specific event. Setting $\tau = \delta \tau + d\tau$, the probability that given the state $\ket{\Psi(t_0)}$ at time $t_0$ the next reaction $\mu$ will happen during the next interval of $\tau$, denoted $P(\tau,\mu)$, can be written as the product of the probability that nothing will happen during $\delta \tau$ and the probability that $\mu$ will happen during $d\tau$:
From (ref), Gillespie1977 deduces that the probability that nothing happens during $\tau$, can be formulated as
Noting that $\tau = \delta \tau + d\tau$ by definition, bringing all terms involving $P_0$ to the left hand side, dividing both sides by $d\tau$ and taking limits for $\delta \tau \to 0$, yields a differential equation that is solved by setting
Substituting (ref) into (ref), the probability that $\mu$ will happen during the next time interval $\tau$ is given by
where in our case $r_0 = \sum_{k,q,M} \alpha_M(k,q;z) + \sum_{k,q,M} \omega_M(k,q;z) $.
From (ref), we may randomly generate the pair $(\tau,\mu)$, i.e., the time when an event occurs $\tau$ and which event will happen $\mu$. As we have set up the rates as price and size specific, by generating the event $\mu$ we also specify the price location and the size which are affected by the event. By noting that (ref) determines an exponential distribution with scale parameter $r_0$, we can first sample $\tau$ by drawing $u_1$ from a uniform distribution $\mathcal{U}(0,1)$ and calculating
Having determined when an event occurs, we may now ask the question what will happen. By numerically specifying the rates for all possible events $r_\nu$ and drawing a second realization $u_2$ from a uniform distribution $\mathcal{U}(0,1)$, we may find the integer $\mu$ by solving
for $\mu$. In other words, by drawing $u_1$ and $u_2$, we can simulate an answer to the question when something will happen with $u_1$ and, with $u_2$, what as well as where it will take place. In fact, we also draw a third realization from a uniform distribution $u_3$, to answer the question what size is affected (see (ref) for details). Having drawn an event and it's characteristics, the current state of the system can be updated. This may change the rates $r_\nu$ and their sum $r_0$. Note that by sampling the events in this fashion, the events are conditionally independent. They may not be independent as the rates are conditional on the current state (and under the assumption of a higher Markov order also on finitely many previous states) of the LOB.
In order to simulate the LOB dynamics, we have to specify the rates of all possible events and how they depend on the current state. In our case, all possible events comprise the order arrivals and order cancellations. Thus, we have to find a functional form for the respective rates $\alpha_M$ and $\omega_M$. In our specifications presented in (ref), we let $\alpha_M$ and $\omega_M$ be functions of the quantity $q$ and the price level $k$ (or more precisely of the integer distance to the opposite best quote $d_l$). For simplicity, we will assume that all rates $\alpha_M(k,q)$ and $\omega_M(k.q)$ are separable in $k$ and $q$ such that
As the rates are proportional to the probability distribution of arriving (or canceled) orders across price and size, this means that the size of arriving (or canceled orders) is stochastically independent of the price level they concern. In (ref), we see that for lower distances to the opposite best quote the size of arriving and canceled order is equally spread out across possible size levels. A clear relationship between the price level and the size level is not visible. In the absence of such a clear relationship, we find the approximating assumption that the size and price level are stochastically independent justifiable.
We also decompose the arrival rates further by setting the general intensity of events for each market side $\bar{r}_{0,M,i}$ to the average event rate over the entire sample of stock $i$, where $\bar{r}_{0,M,i}$ is defined as
which is calculated as the number of events on one market side divided by the total number of events. Note that since we have several order types, the arrival rates may be split into market orders as well as marketable and non-marketable limit orders. The empirical frequencies for $\bar{r}_{0,M,i}$ are reported in the last column in (ref).
Hence, the arrival and cancellation rates for limit orders can be described by the partitioning of the average event rate $\bar{r}_{0,M,i,j,a}$ across price levels $k$ and order sizes $q$:
where $\bar{r}_{0,M,i,j,a}$ is the rate for an order of type $j$ (market or limit order) for stock $i$ to arrive and $\bar{r}_{0,M,i,L,c}$ is the rate for a limit order (i.e. $j=L$) to be canceled. $p_{K,M}(k;\boldsymbol{\theta}_{M,a})$ denotes the discrete probability mass function of order arrivals across the integer price levels $k$ given some parameter set $\boldsymbol{\theta}_{M,a}$ and similarly $p_{Q,M}(q;\boldsymbol{\phi}_{M,a})$ is the discrete probability mass function of order arrivals or cancellations across order sizes. The index $a$ indicates the parameters for order arrivals, $c$ the parameters for cancellations. The index $M$ denotes the market side.
In order to simulate the order book, several probabilities and other conventions have to be specified. Therefore, we go through the terms in (ref) and present how we have chosen to specify $\alpha_{M}(k,q)$ and $\omega_{M}(k,q)$. For convenience, recall (ref) as
(ref) illustrates the components of (ref).
Recall that we chose three theoretical scenarios for the distributions across price levels $p_{K,M}(\cdot)$: First, the uniform distribution (uni), second, a discrete log-normal distribution with fixed parameters (fix), and third, a discrete log-normal distribution with dynamic parameters where the parameters depend on the prevailing spread (dyn). For the distribution across order sizes, we only consider one theoretical specification: a power law distribution. Additionally, we also consider the unconditional empirical frequencies of incoming and canceled orders as observed in the first quarter of 2004, both across price and size levels.
The first element of (ref) is $\bar{r}_{0,M,i,j,e}$, the rate for an arrival ($e=a$) or a cancellation ($e=c$) of order type $j$ on market side $M$ for stock $i$. We first need to specify the order types that we include in the simulation. In (ref), we have depicted limit orders and market orders across relative integer distances to the best quote to show that there is a somewhat stable distribution across price levels when the best quote is used as a fix point. At the zero level, we have plotted the marketable orders split up into different types.
(ref) shows the percentages of the different types of marketable orders in detail. In general, approximately 10% of all incoming orders (cancellations excluded) are marketable. In fact, about half of those marketable orders are arriving on the best quote i.e. with $d=0$. Around a quarter is due to market orders with no limit price $d<-\infty$ and another quarter are marketable limit orders i.e. with $d<0$. Marketable iceberg and stop order are tiny in comparison. The inverse of the last column are the limit orders that are submitted before the best quote. As depicted in (ref), in our simulation scenarios, we only treat market and limit orders separately. We do not distinguish iceberg and stop orders, since they are market and limit orders with some additional features. Thus, when we use the unconditional empirical frequencies for $\bar{r}_{0,M,i,j,e}$, we calculate
where $n_{M,i,j,e}$ is the number of arrivals ($e=a$) or cancellations ($e=c$) of order type $j$ on market side $M$ for stock $i$ observed during the entire first quarter of 2004. $\Delta T$ refers to the total trading time during this period. In our case, $\Delta T$ is specified to be 64 trading days. As we restrict the simulation to continuous trading, we only include events during the 8h28m of continuous trading to calculate the frequencies. In our sample, option settlement is conducted in three dates. On these three days, further 3 minutes have to be subtracted from the continuous trading phase. In total, we have $(64\cdot(8+28/60)\cdot60-3\cdot3)\cdot60 = 1{,}950{,}180s$ of continuous trading time in our sample. Note that we can also decompose the unconditional empirical rates according to
where $n$ refers to a number of events and the indices specify which characteristic is relevant for counting. $n_{\cdot,i,\cdot,\cdot}$ means that only the index $i$ (referring to the event concerning stock $i$) is relevant to determine the number of events. Categories marked with a $\cdot$ in the index are summed over. In other words, $n_{\cdot,i,\cdot,\cdot}$ denotes the number of events concerning stock $i$. In the empirical scenarios, all elements of (ref) can be observed. In theory, we can craft theoretical scenarios to investigate, ceteris paribus, the sensitivity of the LOB dynamics to changes in just one conditional frequency in (ref). In this paper, we chose to focus on the sensitivity of the order book dynamics to changes in the distribution across price and quantity levels.
In the scenarios that entail a theoretical distribution, we do not use the empirical values observed in our sample. We also chose to focus on the distribution of arrival rates across price and size levels. Thus, we set the values summarized in (ref). The rates are specified in the unit [orders/second]. They approximately mirror the observed values in reality but we fix them to parity, so that the two sides of the market are symmetric and balanced.
One peculiarity in the theoretical scenarios 'fix' and 'dyn' is that we treat marketable limit orders below or above the best quote as market orders. Marketable limit orders on the best-quote, i.e., with $d=0$, are modeled together with the rest of the limit orders as they approximately seem to fit into the discrete logarithmic distributions across price levels (cp.\ (ref)). In the scenario 'uni', we separate the market orders and the marketable limit orders (strictly) below or above the best quote up to $d=-10$.
{\singlespacing
}
For the probability distribution of order arrivals across price levels specified in the factor $p_{K,M}(\cdot)$, we distinguish three theoretical scenarios and one scenario using unconditional empirical frequencies.
The easiest approach to define the arrival rates across price levels is a uniform distribution. In this scenario, we assume that the arrivals of orders are concentrated on the first 90 integer price levels before the best quote of the opposite market. Additionally, marketable limit orders are also allowed to cross the best quote up to 10 price levels. In essence, this means that the arrivals of bid and ask orders are concentrated on 100 price levels around the best quote of the opposite market where the arrival rate on each price level is 0{.}0012 orders per second.
For the cancellations, we distribute the probability for an order cancellation uniformly among the occupied price levels.
Empirical frequencies of (non-marketable) limit orders across relative price levels exhibit pronounced probability mass at the tails of the distribution. For the distribution across price levels, in the scenario 'fix', we use a discrete Gaussian exponential distribution (DGX) as presented by BiFCK01. The distribution is especially useful in cases where the random variable to be modeled is discrete and has pronounced probability mass at the tails. It is particularly nice that the DGX reduces to the generalized Zipf distribution when $\mu \to -\infty$. Thus, it is flexible enough to incorporate situation where the probability distribution is a straight line in log-log-plots and cases in which it exhibits some curvature. A short summary of the DGX distribution is given in (ref).
In the simulation scenario with a fixed probability distribution, we chose to set the values as outlined in Table (ref).
The values are the empirical mean and standard deviation across incoming orders of a random sample over several stocks. Note that the mean of arrivals is slightly higher on the bid side of the market, i.e., orders are more likely to arrive deeper in the book. Also the variance of order arrivals is higher. The same holds for cancellations. So while there are more arrivals deeper in the book, slightly more orders deep in the book are also canceled.
Similar to the case in which we use a DGX distribution with fixed parameters $\mu$ and $\sigma$, in the simulation scenario with a dynamical distribution across price levels, we also use the DGX distribution as the fundamental distribution. However, in this case we specify the parameters of the distribution to depend on the prevailing spread. The functional relationship we use is the following:
The functional relation is inspired by a scatter plot of $\hat{\mu}_i$ and $\hat{\sigma}_i$ estimated on the unconditional frequencies of order arrivals (and cancellations) across price levels against the average spread $\Delta_i$ for each stock $i$.\footnote{We are aware of the fact that the expectation of the functional relationship between DGX parameters and spread is not the same as a function for the log-likelihood in dependence of the expectation of the spread i.e. ${\operatorname*{\mathbb{E}}}_{t_0}[\mu(\Delta)] \neq \mu({\operatorname*{\mathbb{E}}}_{t_0}[\Delta])$. } This scatter plot is depicted in (ref). Note also, that we have switched the scale of standard deviation and expectation as observed in the data on purpose. In that way, we hope to get an impression on how an increase in variance and a decrease of the mean may affect the characteristics of the order book evolution. The 'dyn' scenario is theoretically also motivated by the quest to study the sensitivity of the LOB system to feedback reactions between the state of the book and traders' order submission behavior.
We also simulate one scenario where we take the empirical frequencies observed across price levels into account. The empirical log-frequencies are depicted in (ref).
For the distribution across volume, we employ two different specifications. In one specification, we use a power law distribution. Even though, in our data at hand, we find that a power law does not match the volume distribution. This can be seen by sheer eyeballing of (ref). Nevertheless, the good fit of the power law distribution to describe order sizes has been shown in various articles BouchaudMP02, GopikrishnanPGS00, MaslovM01.
The probability mass function of a discrete power law distribution where the smallest value of the support is 1, is theoretically defined as ClausetSCN09
where is the Hurwitz zeta function $\zeta(\lambda) = \sum_{n=0}^\infty (n+1)^{-\lambda}$. We fix the parameter in all simulations at $\lambda = 1.6$ which is close to empirically observed values.
According to ClausetSCN09 (ClausetSCN09, Appendix D), given a random number $u \in [0,1]$, we can generate an integer realization $\tilde{x}$ from the power law distribution by calculating
where $\lfloor\cdot\rfloor$ signify the floor operator which cuts off the decimal places of the argument.
Note that since we assume independence of $p_{Q,M}(\cdot)$ and $p_{K,M}(\cdot)$, and each arriving order surely has to be assigned a size, we simply generate a third realization from a uniform distribution $\mathcal{U}(0,1)$ to determine the volume. In other words, before randomly generating what size is affected with $u_3$, we answer the question when something will happen with $u_1$ and, with $u_2$, what as well as where (i.e.\ at which limit price level) it will happen, as described in (ref).
We also simulate one scenario where we use the empirical frequencies observed across quantity levels to simulate the LOB evolution. The empirical log-frequencies are depicted in (ref).
For the case where both the distribution across price levels and the distribution across price levels are sampled from the empirically observed frequencies, we take the joint frequencies (not the product of the marginal frequencies) to sample both size and price of an incoming or canceled order.
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringDiscrete Gaussian Exponential Distribution (DGX)} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip}
In the simulation, as described in (ref), we use the DGX for the simulation of order arrivals and cancellations across price levels.
The probability mass function of the distribution can be defined according to BiFCK01 as
where the normalizing constant $A(\mu,\sigma)$ is defined as
In (ref), we use a slightly modified version of the DGX by truncating the distribution at 1 to show how the DGX can be fit to the data. The truncated DGX can be derived from the truncated (log-)normal distribution for continuous values and has the following probability distribution function:
The normalization factor $A_T(\mu,\sigma)$ is similarly defined by
The parameters $\mu$ and $\sigma$ can be estimated using a maximum likelihood specification as described in BiFCK01.
\thispagestyle{plain} \suppressfloats[t] \@afterindentfalse \secdef \refstepcounter{section} \addcontentsline{toc}{section} {\numberline {\Alph{section}}?} { \appendixname\, \Alph{section} \appendixname\ \Alph{section} \centeringTaylor Series Expansion of Linear Models} \sectionmark{?} \@afterheading \addvspace{\baselineskip}
\@afterheading \addvspace{\baselineskip} In this appendix the Taylor Series expansion that justifies the model specifications in Equations (ref) to (ref) is derived.
Starting point of the derivation is the decomposition of the event rate in (ref). First, for the sake of brevity, we introduce the intensity of events at the price level $k$ with affected order size $q$ given a prevailing spread $\Delta$.
where the index $M$ denotes the market side (either bid or ask), $i$ indicates the instrument, $j$ denotes the order type (limit or market) and $e$ denotes the event type (arrival or cancellation). The right hand side follows when writing arrival $\alpha$ and cancellation rates $\omega$ separately.
As we have seen in (ref), the Hamiltonian $H$ is directly constructed from the event rates. As a direct consequence, the conditional probability to find the system in state $\ket{z}$ given that it has been in $\ket{z_0}$ at time $t_0$ can be expressed as described in (ref) or in more detail as
As one can see, by construction the conditional distribution is polynomial in the (time integral over) arrival and cancellation rates, and, depending on the succession of orders in $\ket{z}$ (see e.g. (ref)), further combinatorical factors have to be introduced (which include the factorial $w!$). Only regarding the terms up to order one the conditional probability can be written as
This first order approximation may fit the conditional probability for short time horizons well, however, for longer time horizons the interactions between order arrivals and cancellations may become the more important factor.
Nevertheless, with this approximation, we may view the conditional expected value of some observable $O$ given the state of the order book at time $t_0$ of stock $i$ at some future time $t >t_0$ as
where $E$ and $C$ denote the order entry and cancellation operators laid out in (ref) and $O_{i,t_0}$ is the realization of the observable at time $t_0$.
We may generalize this notion for the conditional expected value to some function that depends on the intensity of events, and, thus, again on order size $q$, the price level $k$ and additionally further variables that determine arrival and cancellation rates, e.g. the spread.
Making the very crude assumption that the expected value of some observable is linear in the intensity of events, we could formulate the approximation as
Decomposing $r_{M,i,j,e}(k,q \mid \Delta)$ as done in (ref) and additionally assuming that the the average intensity $\bar{r}_{0,M,i,j,a}(\Delta)$ is some function of the prevailing spread yields
Now, expanding each term by a Taylor series expansion around the respective expected value we have the following expansions for $p_{Q,M}(q;\boldsymbol{\phi}_{M,e})$
Taking expectations with respect to $p_{Q,M}(\operatorname*{\mathbb{E}}[q];\boldsymbol{\phi}_{M,e})$ yields an approximation of the expected value in moments of order 2 and higher
Using then, again, the first order Taylor approximation of $p_{Q,M}(\operatorname*{\mathbb{E}}[q];\boldsymbol{\phi}_{M,e})$ around $0$ reintroduces the first moment
Thus, we may write the expected value of $p_{Q,M}(q;\boldsymbol{\phi}_{M,e})$ as a linear function of the moments
Changing variables in $p_{K,M}(k;\boldsymbol{\theta}_{M,e})$ by considering the relative distance $d_l$ to the opposing best quote instead of the absolute price level $k$, the same can be done for $\operatorname*{\mathbb{E}}[p_{K,M}(k;\boldsymbol{\theta}_{M,e})]$
Last but not least, we may model the expected value of the event specific intensity by its expected value shifted by a event specific factor $\rho_{0,M,i,j,e}$ and the expected value of an event unspecific function $f(\Delta)$ solely dependent on the spread
Approximating $\operatorname*{\mathbb{E}}[f(\Delta)]$ by an expansion in moments as above yields
Reinserting the Taylor expansions in Equations (ref), (ref) and (ref) up to order four together with (ref) in (ref) yields (ref)