Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
41,211 characters · 11 sections · 38 citation commands
Granger causality on horizontal sum of Boolean algebras
In the past there were several attempts to model causality. Yet, there exists no mathematical definition of this notion. Loosely understood is such a relationship between a cause $\xi$ and its effect $\eta$. The usual dependence relationship between $\xi$ and $\eta$ is symmetric (often expressed by a type of correlation). However, there are situations where we have non-symmetric dependence between variables, e.g., the past of a time-series may influence its future, but not vice-versa. Or in hydrology, high water level in a river may cause high water level dawn the river flow, but usually not vice-versa. Thus, we need also a non-symmetric relationship between a cause and its effect.
The aim of this paper is to provide a mathematically rigorous model of non-symmetric causality. It will be based on Granger's approach to causality G. From the algebraic point of view, our model will be based on horizontal sums of Boolean algebras. A preliminary version of our model was published in BKN1. This is an extended version of the paper BKN1.
Our paper is organized as follows: in section (ref) we provide a historical overview of various models of causality. In section (ref) we recall the Granger's definition of causality for stationary time-series. Section (ref) contains basic definitions and some known facts on orthomodular lattices. In section (ref) we provide a theoretical background for our model of causality. It has two subsections:\\ -- (ref) contains a theory of bivariate states on orthomodular lattices,\\ -- (ref) contains an approach that enables to add non-compatible observables.
Finally, section (ref) contains our model of causality on horizontal sums of Boolean algebras. It has again two subsections:\\ -- (ref) provides a comparison of random vectors on Boolean algebras and of vectors of observables on horizontal sums of Boolean algebras,\\ -- (ref) is devoted to the Granger's model of causality modified for horizontal sums of Boolean algebras.
In the 20th century causal inference was frequently associated with multiple correlation and regression. As it is well known, the regression of Y on X produces coefficient estimates that are not the algebraic inverses of those produced from the regression of X on Y. Although regressions may have a natural causal direction, there is nothing in the data on their own that reveals which direction is the correct one – each of them is an equally appropriate rescaling of a symmetrical and non-causal correlation. This is a problem of observational equivalence. For example, we can mention the problem of econometric identification: how to distinguish a supply curve from a demand curve. A standard solution to this identification problem is to look for additional causal determinants that discriminate between otherwise simultaneous relationships. Possible solution of this problem gives us the language of exogenous and endogenous variables. Exogenous variables can also be regarded as the causes of the endogenous ones Hoo08.
In the 1930s Jan Tinbergen Tin introduced structural models in modern econometrics. These models express causality in a diagram that uses arrows to indicate causal connections among time-dated variables. Another approach is known as process analysis. Process analysis emphasizes the asymmetry of causality, typically grounded in Hume's criterion of temporal precedence Mor. Wold's process analysis belongs to the time-series tradition that ultimately produces Granger causality and vector autoregression. The Wold's approach relates causality to the invariance properties of the structural econometric model. This approach emphasizes the distinction between endogenous and exogenous variables and the identification and estimation of structural parameters. Herbert Simon Sim has shown that causality could be defined in a structural econometric model, not only between exogenous and endogenous variables, but also among the endogenous variables themselves. And he has shown that the conditions for a well-defined causal order are equivalent to the well-known conditions for identification Hoo08.
Hans Reichenbach GE12,Rei56, taken the idea that simultaneous correlated events must have prior common causes, tried to use them to infer the existence of unobserved and unobservable events and to infer causal relations from statistical relations. Reichenbach's common cause principle is a time-asymmetric principle that can be formulated as follows: simultaneous correlated events have a prior common cause that screens off the correlation. It means, if simultaneous values of quantities $A$ and $B$ are correlated, then there are common causes $C_1,C_2,\ldots,C_n$ such that conditioned upon any combination of values of these quantities at an earlier time, the values of $A$ and $B$ are probabilistically independent, see Af10,Uff99. Reichenbach's common cause principle was adopted by Penrose and Percival pp62 into the law of conditional independence and by Spirtes et al. SGS93 into causal Markov condition.
Some open problems concerning Reichenbachian common cause systems are formulated and solved by Hofer-Szabó and Rédei in many papers. Hofer-Szabó and Rédei HR06 have shown that given any non-strict correlation in $(\Omega ,\mathcal S, P)$ and given any finite natural number $n>2$, the probability space $(\Omega ,\mathcal S, P)$ can be embedded into a larger probability space in such manner that the larger space contains a Reichenbachian common cause system of size $n$ for the correlation.
Another approach, given by Clive W. J. Granger G, introduces the data-based concept without direct reference to background economic theory. This concept has become a fundamental notion for studying dynamic relationships among time series. Granger's causality is an example of the modern probabilistic approach to causality, and it is a natural successor to Hume (see, e.g., Sup70). Where Hume requires constant conjunction of cause and effect, probabilistic approaches are content to identify cause with a factor that raises the probability of the effect: $A$ causes $B$ if $P(B|A) > P(B)$, where the vertical ‘$|$’ indicates ‘conditioned on’. The asymmetry of causality is secured by requiring the cause $(A)$ to occur before the effect $(B)$ (see Hoo08). But the probability criterion is not enough on its own to produce asymmetry since $P(B|A) > P(B)$ implies $P(A|B) > P(A)$.
Granger's causality helps us to understand and measure the relative roles of different causal systems, e.g., between commodity prices and exchange rates. Granger causality has important implications in financial decision making, especially for market participants with short horizons. From a macroeconomic perspective, this can also be useful for interpreting exchange rate movements, financial market monitoring and monetary policy. Basic economic reasoning on currency demand suggests that the currencies of countries whose exports depend heavily on a particular commodity should be strongly influenced by its price, so commodity price movements should lead (Granger-cause) exchange rate movements (macroeconomic/trade mechanism).
In statistics, the notion of causality is usually identified with a kind of stochastic dependence. Of course, this dependence (e.g. between two random variables) is a symmetric notion. In G, Granger defined a causality between two stationary time-series ${\mathbf X}=\{X_t\}_{t\in{\mathbb Z}}$ and ${\mathbf Y}=\{Y_t\}_{t\in{\mathbb Z}}$ in a non-symmetric way. There are two basic principles upon which this notion of causality (and a relationship between a cause and its effect) is based.
The precise definition of Granger's causality is the following:
In fact, this notion of causality is based on the Kolmogorovian conditional probability theory. Granger's theory is used especially in econometrics and finance to model one-sided dependencies, as we have mentioned in the historical overview.
We have already mentioned in Introduction that our model of causality works on horizontal sums of Boolean algebras. Since horizontal sums of Boolean algebras are special cases of orthomodular lattices, we recall some basic facts also on general orthomodular lattices. For more information on orthomodular lattices and their properties one can consult, e.g., D-P,Kalm-83,PtakPulm,Var.
In this paper, for the sake of brevity, we will write briefly an orthomodular lattice $L$, skipping the operations whenever it will not cause any confusion. In general, an orthomodular lattice is not distributive. For arbitrary $a,b\in L$ just the following property is guaranteed \[ (a\wedge b)\vee (a\wedge b' )\leq a. \] On the other hand, if $L$ is distributive then it is a Boolean algebra.
An orthomodular sub-lattice $L_1$ of $L$ is an orthomodular lattice such that $L_1\subset L$, with operations inherited from $L$ and possessing the same greatest and least elements ${\mathbf 1}_L$ and ${\mathbf 0}_L$, respectively. A distributive orthomodular sub-lattice $\cal B$ is called a Boolean sub-algebra of $L$.
Every orthomodular lattice $L$ is a collection of blocks R2. A block is the maximal set of pairwise compatible elements of $L$, i.e. $L=\bigcup_j B_j$, where blocks $B_j$ have operations inherited from $L$. Each block in $L$ is a Boolean algebra.
In this paper we will deal only with $\sigma$-complete orthomodular lattices $L$. Such $\sigma$-complete orthomodular lattices are called orthomodular $\sigma$-lattices ($\sigma$-OML, for brevity).
As it was proven by Greechie Gree, there exist orthomodular lattices with no state.
Directly from the properties of $\sigma $-homomorphism it follows that $R(x)$ is a Boolean sub-$\sigma$-algebra of $L$ (e.g., PtakPulm,Var).
If $x\in\mathcal O$ and $m$ is a $\sigma$-additive state on $L$, then $m_x(B)=m(x(B))$, $B\in\mathcal{B({\mathbb R})}$ is a probability distribution of $x$.
Let $(\Omega ,\mathcal S,P)$ be a probability space. Then $\mathcal S$ is a Boolean $\sigma$-algebra and $P$ is a $\sigma$-additive state. Hence $\mathcal S$ is a $\sigma$-OML. Furthermore, if $\xi $ is a random variable on $(\Omega ,\mathcal S, P)$, then $\xi^{-1}$ is an observable. It means that, if we have an observable $x$ on a $\sigma$-OML $L$, we are in the same situation as in the classical probability space. We use only another language for the standard situation. Problems occur if we have more then just one observable, and their ranges are not compatible.
In this section we show how it is possible to introduce causality on orthomodular lattices between observables. As a first step we need conditional states and joint distributions (s-maps).
Conditional states and s-maps were introduced in Ns,Nc resp., and their properties were studied for example in NP08. For a given $\sigma$-OML $L$ with a $\sigma$-additive state, $L_{0}$ will denote the set of all elements $a\in L$ for which there exists a $\sigma$-additive state $m_a$ such that $m_a(a)=1$. In this paper we will assume that
In fact, $f(\cdot|{\mathbf 1}_L)$ plays the role of a prior state (prior probability distribution) in the classical definition of independence. This means that Definition (ref) is just re-written from the Kolmogorovian probability theory. But unlike the Kolmogorovian theory, the independence of elements of a $\sigma$-OML is not necessarily symmetric. In Nc a conditional state $f$ was constructed in such a way that there are elements $a,b\in L$ for which $f(b|{\mathbf 1}_L)=f(b|a)$ and $f(a|{\mathbf 1}_L)\neq f(a|b)$ (see also Example (ref) later in this paper). This fact implies that the well-known Bayes Theorem may be violated on a $\sigma$-OML.
Let $\cal P$ denote the system of all s-maps on $L$ (for fixed $L$), which are $\sigma$-additive in both variables. The relationship between s-maps $p\in{\cal P}$ and conditional states is given by the following proposition
Let $L$ be a $\sigma$-OML and $p\in {\cal P}$ be an s-map on $L$. Denote $\mu_p(a)=p(a,a)$ for all $a\in L$. Then the following statements hold:
In what follows, for a given s-map $p\in{\cal P}$ we use the notation $\mu_p(a)=p(a,a)$ for all $a\in L$.
We will use the following notation
We say that an s-map $p:L\times L\to[0,1]$ is strongly causal if it is causal and there exists a pair of elements $a,b\in L$ such that $a$ is dependent on $b$ but $b$ is independent of $a$.
The s-map $p_1$ from Example (ref) is strongly causal. The s-map $p_2$ from that example is causal, but not strongly causal.
An important notion for our considerations is also that of a conditional expectation.
Since for arbitrary observable $y$ $R(y)$ is a Boolean sub-$\sigma$-algebra of $L$ we will write simply $E_p(x|y)=E_p(x|R(y))$.
For compatible observables $x,y$ on a $\sigma$-OML $L$ due to Theorem (ref) there exist a $\sigma $-homomorphism $h$ and real functions $f,g$ such that $x=h\circ f$ and $y=h\circ g$. This means that $x+y$ is defined by $x+y=(f+g)\circ h$. If $x,y$ are non-compatible then we cannot apply this procedure and $x+y$ does not exist in this sense.
In N-K14 a sum of non-compatible observables was defined.
The following basic properties of $\oplus_p$ are proven in N-K14.
Before turning our attention to the Granger causality, we should say something on random vectors and stochastic processes as a generalization of random vectors.
We will deal with a measurable space $(\Omega,{\cal S})$ where ${\cal S}$ is a $\sigma$-algebra of measurable events. Denote ${\cal B}$ the $\sigma$-algebra of Borel subsets of ${\mathbb R}$. A random variable $\xi:\Omega\to{\mathbb R}$ is an ${\cal S}$-measurable function, i.e., for every $B\in{\cal B}$ $\xi^{-1}(B)\in{\cal S}$.
Further, let ${\cal B}^2$ and ${\cal S}^2$ denote the direct products ${\cal B}\times{\cal B}$ and ${\cal S}\times{\cal S}$, respectively. By $\sigma({\cal B}^2)$ and $\sigma({\cal S}^2)$ we will denote the least set $\sigma$-algebra containing the corresponding direct products.
Let $\xi$ and $\eta$ be ${\cal S}$-measurable functions. In the Kolmogorovian probability theory the random vector $(\xi,\eta)$ is modelled as a bivariate function such that for every $B\in\sigma({\cal B}^2)$ $(\xi,\eta)^{-1}(B)\in\sigma({\cal S}^2)$. This model works perfectly if $\xi$ and $\eta$ are measurable simultaneously (e.g., two parameters measured on the same objects). But also in this case we are usually interested in knowing probabilities for $P(\xi^{-1}(A),\eta^{-1}(B))$ where $A,B\in{\cal B}$. This means that instead of constructing $\sigma({\cal B}^2)$ and $\sigma({\cal S}^2)$ it is enough (might be up to some exceptions) to work with the corresponding direct products ${\cal B}^2$ and ${\cal S}^2$. Thus the model becomes slightly different from the Kolmogorovian one, especially when we extend this consideration to stochastic processes.
A different situation occurs if we consider a random vector $(\xi,\eta)$, but $\xi$ and $\eta$ are not simultaneously measurable. Of course one possibility how to model this situation is to stay within the Kolmogorovian model. In this case we know that $P((\xi,\eta)^{-1}\in A\times B)=P((\eta,\xi)^{-1}\in B\times A)$, where $A,B\in{\cal B}$. Instead of random variables $\xi$ and $\eta$ we can use observables $\xi^{-1}$ and $\eta^{-1}$. The fact that observables $\xi^{-1}$ and $\eta^{-1}$ are not simultaneously measurable, can be interpreted as their non-compatibility. We have seen in Example (ref) that unlike the probability measure, s-maps are not necessarily symmetric. This means, if we denote $a=\xi^{-1}(A)$ and $B=\eta^{-1}(B)$, we might get $p(a,b)\ne p(b,a)$. However, to get non-compatibility, we must leave Boolean algebras and switch to more general structures. We will consider two copies of the $\sigma$-algebra ${\cal S}$ denoted by ${\cal S}_1$ and ${\cal S}_2$. Assume ${\cal S}_1\cap{\cal S}_2=\{\emptyset,\Omega\}$. $\emptyset$ and $\Omega$ are the bottom and top elements, respectively, of these two $\sigma$-algebras. This means that we can make their horizontal sum in the same way as we have made it with blocks $B_1$ and $B_2$ in Example (ref) when we constructing the OML $L_1$. The corresponding horizontal sum of ${\cal S}_1$ and ${\cal S}_2$ will be denoted by $\tilde{\cal S}$. In such a way for arbitrary $A,B\in{\cal B}$ we have $(\xi^{-1}(A),\eta^{-1}(B))\in\tilde{\cal S}\times\tilde{\cal S}$ and $(\eta^{-1}(B),\xi^{-1}(A))\in\tilde{\cal S}\times\tilde{\cal S}$. In this situation we have one s-map $p$ modelling the (possibly non-symmetric) distribution of both vectors of observables, $(\xi^{-1},\eta^{-1})$ and $(\eta^{-1},\xi^{-1})$.
Assume that $\{\mathbb X_t\}_{t\in T}$ is a stochastic process. For every time-stamp $t\in T$, $X_t$ is a ${\cal S}$-measurable random variable where ${\cal S}$ is a Boolean $\sigma$-algebra. If we want to model causality (in the sense of non-symmetric dependence), we have to make the same procedure as above (with random vectors) when we have abandoned Boolean algebras and considered horizontal sums of Boolean algebras, instead.
We will consider $card(T)$ copies of the $\sigma$-algebra ${\cal S}$, i.e., we will have a family $\{{\cal S}_t\}_{t\in T}$ and we make their horizontal sum. By $\hat{\cal S}$ we denote the resulting horizontal sum. For every time-stamp $t\in T$ and every Borel set $A\in{\cal B}$ we will have $X^{-1}_t(A)\in\hat{\cal S}$. Then, for $s\ne t$, $X^{-1}_t$ and $X^{-1}_s$ are non-compatible observables. We know already that there exists a joint distribution of $X^{-1}_t$ and $X^{-1}_s$ (or equivalently, conditional distribution $f(X^{-1}_s|X^{-1}_t)$ which is interesting especially when $s>t$), and by Proposition (ref), having the conditional distribution $f(X^{-1}_s|X^{-1}_t)$, there exists also their sum.
{\bf Granger causality.} Assume that we have two (not necessarily stationary) stochastic processes, $\{\mathbf X_t\}_{t\in T}$ and $\{\mathbf Y_t\}_{t\in T}$, where $T$ is a set of all possible time-stamps. According to Definitions 2 and 5 in G2, $\{\mathbf Y_t\}_{t\in T}$ causes $\{\mathbf X_t\}_{t\in T}$ if $F(X_{t+1}|Y_t)\ne F(X_{t+1})$, where $F(\cdot|\cdot)$ is a conditional distribution function and $F(\cdot)$ is an unconditioned distribution function.
To model causality between stochastic processes $\{\mathbf X_t\}_{t\in T}$ and $\{\mathbf Y_t\}_{t\in T}$, we need to have an equivalent of a measurable space such that for every $t,s\in T$ observables $X^{-1}_t$ and $Y^{-1}_s$ are non-compatible. This means that we need two copies of $\hat{\cal S}$ and make their horizontal sum. We denote this newly constructed lattice by $\hat{\cal S}_2$. In this way we get that $F_{(X^{-1}_t,Y^{-1}_t)}$ and $F_{(Y^{-1}_t,X^{-1}_t)}$ may be different functions.
In experiment we are not able to distinguish the order $(X^{-1}_t,Y^{-1}_t)$ and $(Y^{-1}_t,X^{-1}_t)$ if we measure $X$ and $Y$ at the same time stamp. This means that, as we have already commented in Remark (ref), measuring the non-symmetric causality experimentally has to follow exactly what Granger proposed in G,G2.
In this paper we have shown parallels between the Granger causality G and modelling of causality on horizontal sums of Boolean algebras which is based on s-maps and conditional states Ns,Nc,NP08. The basic property of Granger's causality is its non-symmetry, i.e., the ability to distinguish between a cause and its effect. Causality based on s-maps and conditional states on orthomodular lattices (and on horizontal sums of Boolean algebras as special orthomodular lattices) bears the same property of non-symmetry. This non-symmetry is suitable for modelling of causality (dependencies) in stochastic processes (as we have shown in Section (ref)) where we are able, in a natural way, to distinguish the cause and its effect. As we have pointed out in Remark (ref) such non-symmetry (order-dependence) may occur also when measuring two different parameters, $\xi$ and $\eta$, by designing two different experiments -- first measuring $\xi$ then $\eta$, or vice versa.