EconBase
← Back to paper

On the equivalence between the Kinetic Ising Model and discrete autoregressive processes

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

31,822 characters · 5 sections · 33 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On the equivalence between the Kinetic Ising Model and discrete autoregressive processes

\affil[1]{Scuola Normale Superiore, piazza dei Cavalieri 7, 56126 Pisa, Italy} \affil[2]{UZH Blockchain Center, University of Z\"{u}rich, R\"{a}mistrasse 71, 8006 Z\"{u}rich, Switzerland} \affil[3]{URPP Social Networks, University of Z\"{u}rich, Andreasstrasse 15, 8050 Z\"{u}rich, Switzerland}

\affil[4]{Department of Mathematics, University of Bologna, piazza di Porta San Donato 5, 40126 Bologna, Italy}

abstractBinary random variables are the building blocks used to describe a large variety of systems, from magnetic spins to financial time series and neuron activity. In Statistical Physics the Kinetic Ising Model has been introduced to describe the dynamics of the magnetic moments of a spin lattice, while in time series analysis discrete autoregressive processes have been designed to capture the multivariate dependence structure across binary time series. In this article we provide a rigorous proof of the equivalence between the two models in the range of a unique and invertible map unambiguously linking one model parameters set to the other. Our result finds further justification acknowledging that both models provide maximum entropy distributions of binary time series with given means, auto-correlations, and lagged cross-correlations of order one. We further show that the equivalence between the two models permits to exploit the inference methods originally developed for one model in the inference of the other.

Introduction

The dynamics of a large variety of systems, from Physics to Economics and Finance, can be represented as time series of binary variables. The most telling examples are the {\it spin systems} in Statistical Physics, where magnetic moments of the particles in a lattice are described as two-state variables, or {\it binary time series} in Quantitative Finance, capturing for instance the occurrence of extreme events of prices calcagnile2018collective,hong2009granger, or buy and sell orders in the order book of financial markets bouchaud2002statistical. Different models have been introduced to capture the multivariate interaction structure of such binary systems, in particular the {\it Kinetic Ising Model} (KIM) fredrickson1984kinetic in Physics and the {\it discrete autoregressive processes} jacobs1978discrete,jacobs1983stationary in time series analysis, together with the (Markovian) multivariate generalization recently introduced by mazzarisi2020tail, namely the{\it Vector Discrete AutoRegressive Process} VDAR(1). In this paper we prove analytically that, under some condition, the KIM is equivalent to the VDAR(1) model. Furthermore it is well known that the Ising model, in both static schneidman2006weak and kinetic marre2009prediction version, is a maximum entropy model, given mean magnetizations and pairwise correlations (at lag one in the kinetic case), see also jaynes1982rationale,presse2013principles,marcaccioli2020correspondence, and, among other aspects, maximum entropy arguments can be used to define without ambiguity the temperature in such nonequilibrium spin systems sastre2003nominal. Here, by exploiting the equivalence between the two models, we prove also that the Markov chain associated with the vector discrete autoregressive process VDAR(1) can be interpreted as the maximum entropy distribution of binary random variables with given {\it means}, {\it auto-correlations}, and {\it lagged cross-correlations} (of order one). Thus, the KIM and the VDAR(1) should be preferable to other models in absence of prior information on other metrics, following the principle of maximum entropy.

The Kinetic Ising Model fredrickson1984kinetic, derrida1987exactly, crisanti1988dynamics, sides1998kinetic, was originally proposed as the out of equilibrium version of the classical Ising spin glass ising1925beitrag, edwards1975theory, kirkpatrick1978infinite to describe a Markovian dynamics of {\it spins} $\sigma_t^i$, i.e. binary random variables taking values $-1$ and $1$, interacting with each other according to some generic matrix of couplings. The KIM has found countless applications in many contexts such as neuroscience roudi2011mean, capone2015inferring, computational biology imparato2007ising, machine learning coolen1993dynamics, dunn2013learning, campajola2019inference, and economics and finance bouchaud2013, campajola2020unveiling, campajola2020modelling.

In mathematical terms, the KIM is a logistic regression model specified by the following transition probability for a set of $N$ spins $\boldsymbol{\sigma}_t\equiv\{\sigma_t^i\}_{i=1,...,N}$,

dmathp_{KIM}(\boldsymbol{\sigma}_{t} \vert \boldsymbol{\sigma}_{t-1}, \bm{J}, \bm{h}) = Z^{-1}_{t-1} \exp \left( \sum_{ i, j =1}^N \sigma^i_{t} J_{ij} \sigma^j_{t-1} + \sum_{i=1}^N \sigma^i_{t} h_i \right)

where $\bm{J}\equiv\{J_{ij}\}_{i,j=1,...,N}$ is a matrix of real-valued parameters giving the multivariate auto-regressive structure of the model or, equivalently, representing the couplings between spins, $\bm{h}\equiv\{h_i\}_{i=1,...,N}$ is a set of variable-specific parameters representing the external magnetic fields associated with each spin, and $Z_{t-1}$ is the partition function $Z_{t-1} = \prod_{i=1}^N 2 \cosh (\sum_j J_{ij} \sigma^j_{t-1} + h_i)$ guaranteeing that the probability distribution is properly normalized.

The KIM of Eq. ((ref)) is a {\it maximum entropy} jaynes1957information, barnett2013information model for a set of binary random variables which display on average given {\it means}, and both {\it auto-} and (lagged) {\it cross-correlations}. For the sake of clarity and in preparation to the section below, let us move from the spin variables $\sigma_t^i\in\{-1,1\}$ to the binary variables $X_t^i\in\{0,1\}$.

Given a set of $N$ binary variables $\bm{X}_t\equiv\{X_t^i\}_{i=1,...,N}^{t=1,...,T}$, let us consider the following metrics,

equation[equation omitted — 66 chars of source]
equation[equation omitted — 100 chars of source]

related (under stationarity conditions) to the mean of the binary random variables and the correlation between them, respectively.

The metric ((ref)) for $i=j$ is known as {\it stability} hanneke2010discrete, which is connected with the sample auto-correlation of a binary sequence, i.e. $\sum_{t}X_t^iX_{t-1}^i$. Similarly, when $i \neq j$ the metric is related to lagged cross-correlations. The maximum entropy probability distribution of $\bm{X}_1,\bm{X}_2,...,\bm{X}_T$, i.e. the one maximizing the entropy $-\sum_{\bm{X}_1,...,\bm{X}_T}p(\bm{X}_1,...,\bm{X}_T)\log p(\bm{X}_1,...,\bm{X}_T)$ while preserving on average some given values for the metrics ((ref)) and ((ref)), has transition probability (by assuming a given initial condition $\bm{X}_0$ and exploiting the Markov property)

equation[equation omitted — 273 chars of source]

where $\bm{J}=\{J_{ij}\}_{i,j=1,...,N}$ and $\bm{h}=\{h_i\}_{i=1,...,N}$ are $N^2+N$ Lagrange multipliers solving the maximum entropy problem jaynes1957information, park2004statistical. It is trivial to show that the transition probability of the Kinetic Ising Model ((ref)) can be stated equivalently as the maximum entropy probability ((ref)) in terms of the binary variables $\boldsymbol{X}_t$ $\forall i,t$, through the relation $X_t^i = \frac{1 + \sigma^i_t}{2}$ (with the same parameters $\bm{J}$ and $\bm{h}$).

The VDAR(1) model describes the dependence structure of a set of binary random variables which has Markov property and Bernoulli marginal distribution. It has been proposed originally in its univariate version jacobs1978discrete and followed by several extensions such as the Discrete AutoRegressive Moving Average (DARMA) model jacobs1983stationary and recently proposed in its multivariate formulation mazzarisi2020tail, the VDAR model. Models from this family have seen applications in genetics dehnert2003discrete, queueing theory kim2008mean, temporal networks williams2019effects and recently in financial systems, as methods to forecast order flows taranto2014adaptive or to identify preferential lending between banks mazzarisi2020dynamic.

In terms of the $N$ binary variables $\bm{X}_t\equiv\{X_t^i\}_{i=1,...,N}$ with $X_t^i\in\{0,1\}$ (and initial condition $\bm{X}_0$), the VDAR(1) process describes the evolution of $X_t^i$ as

equation[equation omitted — 72 chars of source]

with $V_t^i\sim\mathcal{B}(\nu_i)$ a Bernoulli random variable with parameter $\nu_i\in[0,1]$, $A_t^i\sim\mathcal{M}(\lambda_{i1},...\lambda_{iN})$ a multinomial random variable taking integer value in \{1,...,N\}, with parameters $\lambda_{i1},...\lambda_{iN}$ such that $\sum_{j=1}^N\lambda_{ij}=1$, and $Z_t^i\sim\mathcal{B}(\chi_i)$ with $\chi_i\in[0,1]$. In other words, the VDAR(1) process captures the (multivariate) mechanism of copying from the past: with probability $\nu_i$, $X_t^i$ is copied from the past and, in this case, $\lambda_{ij}$ is the probability that $X_t^i$ is equal to $X_{t-1}^j$ (including also the past itself with probability $\lambda_{ii}$); otherwise, with probability $1-\nu_i$, $X_t^i$ is not copied and is instead sampled according to a Bernoulli marginal with probability $\chi_i$. Hence, the VDAR(1) model describes $N$ binary random variables with both Markov property and some autoregressive dependency structure, similarly to the KIM. The model is formalized by the transition probability

dmathp_{VDAR}(\boldsymbol{X}_{t} \vert \boldsymbol{X}_{t-1}; \boldsymbol{\pi}) = \prod_{i=1}^N \left[ \nu_i \left( \sum_{j=1}^N \lambda_{ij} \delta_{X_t^i,X_{t-1}^j}\right) +(1-\nu_i)(\chi_i)^{X_t^i}(1-\chi_i)^{1-X_t^i}\right]

where $\delta_{X_t^i,X_{t-1}^j}$ is the Kronecker delta, $\boldsymbol{\pi} = \{\lbrace \nu_i \rbrace, \lbrace \lambda_{ij}\rbrace, \lbrace \chi_i \rbrace\}_{i,j=1,...,N}$.

Notice that the model has $N^2+N$ parameters, exactly as the Kinetic Ising Model. It is thus immediate to ask the question whether a mapping between the two models exists, as well as finding under which conditions the two models can be considered equivalent. In the following, we indicate the KIM model as $\lbrace \lbrace \boldsymbol{X}_t \rbrace, p_{KIM}, \boldsymbol{\theta} \rbrace$ with set of parameters $\boldsymbol{\theta} \equiv (\bm{J},\bm{h})$, while the VDAR(1) model is summarized as $\lbrace \lbrace \boldsymbol{X}_t \rbrace, p_{VDAR}, \boldsymbol{\pi} \rbrace$. Calling $\Theta = \mathbb{R}^{N \times N} \times \mathbb{R}^N$ the space of all possible KIM parameters $\boldsymbol{\theta}$ and $\Pi$ the space of all possible VDAR(1) parameters $\boldsymbol{\pi}$,

definitionThe KIM and the VDAR(1) models are said to be equivalent on $(\hat{\Theta},\hat{\Pi})$ if there exist an unique invertible map $f:\hat{\Pi}\subseteq\Pi \rightarrow \hat{\Theta}\subseteq\Theta$ such that $$p_{KIM}(\boldsymbol{X}_t \vert \boldsymbol{X}_{t-1}; f(\boldsymbol{\pi})) = p_{VDAR}(\boldsymbol{X}_t \vert \boldsymbol{X}_{t-1}; \boldsymbol{\pi})$$ for any $\boldsymbol{X}_t$ and $\boldsymbol{X}_{t-1}$.

Model equivalence

Before stating the main theorem, let us show that this mapping exists in the trivial cases of $N=1$ and $N=2$. For $N=1$, both $J$ and $h$ are scalar parameters, while the DAR(1) model, namely the univariate version of VDAR(1), has two parameters, i.e. $\nu$ and $\chi$ ($\lambda = 1$ by design, since we can copy only the past value of the single variable), thus it is trivial to prove that

subequations\begin{equation} h = \frac{1}{4} \log \left( \frac{\frac{\chi}{1-\chi}+\nu}{\frac{1}{\chi}-(1-\nu)}\right) \end{equation} \begin{equation} J = \frac{1}{4}\log\left(1+\frac{\nu}{(1-\nu)^2\chi(1-\chi)}\right) \end{equation}

One can notice that here $J$ is strictly positive as long as $\nu, \chi > 0$, while it is $J=0$ if and only if $\nu = 0$: this suggests that the VDAR(1) model is indeed a restricted version of the KIM, with the elements of the coupling matrix restricted to positive values. Intuitively, in the KIM $J_{ij}<0$ implies that spin $i$ tends to take the opposite value of the past state of $j$, whereas the VDAR model describes only positive (or zero) correlations.

When $N=2$, both models have $6$ free parameters, three parameters associated with each variable (spin) $i=1,2$, thus one can map the two models by considering three independent configurations of $\boldsymbol{X}_{t-1}$, e.g. $\{X_{t-1}^1,X_{t-1}^2\}=\{1,1\},\{0,1\},\{0,0\}$ and one possible realization of $X_t^i$, e.g. $X_t^i=1$, for both cases $i=1$ and $i=2$. Then, by matching the transition probabilities for the two models, one obtains the following system

dmath\begin{pmatrix} 1 & 1 & -1 \\ 1 & -1 & -1 \\ -1 & -1 & -1 \\ \end{pmatrix} \begin{pmatrix} J_{i1} \\ J_{i2} \\ h_i \\ \end{pmatrix} =\frac{1}{2} \begin{pmatrix} \log \left(\frac{1}{(1-\nu_i)\chi_i}-1\right) \\ \log \left(\frac{1}{\nu_i(1-\lambda_{i})+(1-\nu_i)\chi_i}-1\right) \\ \log \left(\frac{1}{\nu_i+(1-\nu_i)\chi_i}-1\right) \\ \end{pmatrix}

$\forall$ $i = 1,2$.

Hence, there exists a {\it unique} mapping $f:\Pi \rightarrow \Theta$ as long as the linear system of equations ((ref)) admits a solution in the domain of parameters: (i) the solution exists because the matrix in the left hand side of ((ref)) is invertible, then (ii) the mapping $f$ admits the inverse $f^{-1}:\Theta\vert_{J_{ij}\geq 0} \rightarrow \Pi$ in the restricted codomain $J_{ij}\geq 0$ $\forall i,j$ (as can be verified by simple computations).

Given these premises, we can now move to the main result of this paper, by stating

theoremThe Kinetic Ising Model $\lbrace \lbrace \boldsymbol{X}_t \rbrace, p_{KIM}, \boldsymbol{\theta} \rbrace$ is equivalent to the VDAR(1) model $\lbrace \lbrace \boldsymbol{X}_t \rbrace, p_{VDAR}, \boldsymbol{\pi} \rbrace$ if and only if $J_{ij} \geq 0$ $\forall i,j$.

In order to prove the theorem above, let us first prove the existence of a map $f:\Pi \rightarrow \Theta$ from the VDAR(1) model to the KIM, for any set of parameters $\bm{\pi}\in\Pi$. In particular, this map is {\it unique} because of the linearity of the mapping problem, see below. Second, we prove that the mapping of parameters is {\it invertible} on its codomain $f(\Pi)\subset \Theta$, corresponding to the set of positive couplings $ J_{ij}\geq0$, $\forall i,j$. Thus, the two models are equivalent under such condition.

Let us start by constructing the system of equations generating the mapping for the generic case $N>2$. Following the same procedure used to obtain Eq. ((ref)), we find

equation[equation omitted — 797 chars of source]

$\:\:\:\forall i=1,...,N$.

Similarly to the case $N=2$, the above system is obtained by considering $n\equiv N+1$ independent configurations for $\boldsymbol{X}_{t-1}$ and the transition probability associated with $X_t^i=1$. By matching the transition probabilities ((ref)) and ((ref)) associated with the $N+1$ independent configurations, one finds the system of equations ((ref)) for variable $i$. Then, one can repeat the same procedure for all $i$s, thus obtaining $N$ systems of $(N+1)$ linear equations in $(N+1)$ unknowns, namely $J_{i1},J_{i2},...,h_i$, each one characterized by the same matrix $M_n$ in Eq. ((ref)).

Defining $\Lambda(x) \equiv \frac{e^{2x}}{1 + e^{2x}}$, the matching of probabilities associated with the $N+1$ independent configurations read as

dmath\makebox[\textwidth]{ $ \begin{cases} \Lambda(h_i + \sum_{j \geq 1} J_{ij}) = \nu_i + (1-\nu_i)\chi_i \:\: &\mbox{if} \:\: X_{t-1}^1=1, X_{t-1}^2=1, ..., X_{t-1}^N =1 \\ \Lambda(h_i - J_{i1} + \sum_{j > 1} J_{ij}) = \nu_i (\sum_{j=2}^N \lambda_{ij}) + (1-\nu_i) \chi_i \:\: &\mbox{if} \:\: X_{t-1}^1=0, X_{t-1}^2=1, ..., X_{t-1}^N=1; \\ \Lambda(h_i - \sum_{j\leq 2}J_{ij} + \sum_{j > 2} J_{ij}) = \nu_i (\sum_{j=3}^N \lambda_{ij}) + (1-\nu_i) \chi_i \:\: &\mbox{if} \:\: X_{t-1}^1=0, X_{t-1}^2=0, ..., X_{t-1}^N=1; \\ ... & ...\\ \Lambda(h_i - \sum_{j \leq n} J_{ij}) = (1-\nu_i)\chi_i \:\: & \mbox{if} \:\: X^1_{t-1}=0, X_{t-1}^2=0, ..., X_{t-1}^N=0, \end{cases}$}

then, Eq. ((ref)) is obtained by applying $\Lambda^{-1}(y)\equiv\frac{1}{2}\log\left(\frac{1}{y}-1\right)$ to both sides of Eqs. ((ref)).

Given this result, a unique mapping $f:\Pi \rightarrow \Theta$ exists as long as there exists the inverse of the matrix $M_n$ in Eq. ((ref)), i.e. if the determinant of $M_n$ is non-zero. We then start by proving the following

propositionGiven the determinant of the matrix $M_{n-1}$, then the determinant of the matrix $M_{n}$ is \begin{equation} \det(M_{n})=(-1)^n 2\det(M_{n-1}) \end{equation}

\paragraph*{\bf Proof of Proposition (ref).} By means of the minor expansion formula (by using the minors associated with the elements of the first row), the determinant of $M_n$ can be computed as

equation[equation omitted — 824 chars of source]

In the previous formula, one can notice that the first $n-2$ minors of the sum in the right hand side are zero, because the last two columns of each $(n-1)\times (n-1)$ matrix are indeed equal (two $(n-1)\times 1$ vectors of $-1$). Thus, Eq. ((ref)) is simplified as

equation[equation omitted — 242 chars of source]

where we notice that the last two minors of ((ref)) are equal to each other and correspond to the determinant of $M_{n-1}$. Eq. ((ref)) then completes the proof of the proposition.

Thanks to this result, we are now able to prove the existence of the mapping from the VDAR(1) model to the KIM model, expressed by

propositionGiven $\bm{\pi}\in\Pi$, there exists a solution of the problem of Eq. ((ref)) for any $N >0$ and this solution is unique.

\paragraph*{\bf Proof of Proposition (ref).} For $N=1$, the solution can be explicitly computed as showed in Eqs. ((ref)) and ((ref)). For $N=2$ the problem in Eq. ((ref)) is equivalent to Eq. ((ref)) and $\det(M_3)=4$, thus there exists the inverse of the matrix $M_3$ and the solution is uniquely determined by solving the linear system of Eq. ((ref)). Because of Proposition (ref), the determinant of $M_n$ is different from zero, in particular

dmath*\det(M_n)=(-1)^{\sum_{l=4}^n l}(2^{n-3})\det(M_3)=(-1)^{\sum_{l=4}^n l}(2^{n-3})4,

$\forall n>3$ (or, equivalently, $\forall N>2$), thus resulting in the existence of the inverse matrix of $M_n$. Hence, the solution of the problem in Eq. ((ref)) can be uniquely determined. This completes the proof of the proposition.

\paragraph*{\bf Proof of Theorem (ref).} Given Propositions (ref) and (ref), we have proved there exists a unique mapping $f:\Pi \rightarrow \Theta$ from the VDAR(1) model to the KIM. To complete the proof, we are now left with the existence of the {\it inverse} of the map $f$ in its codomain $f(\Pi)\subset\Theta$. By restricting to such subset of parameters, the two models are equivalent. In particular, we now prove the last claim of Theorem (ref) which states that the two models are equivalent if and only if $J_{ij} \geq 0$ $\forall i,j$.

Let us start by proving that, if $\boldsymbol{\pi} \in \Pi$, then $J_{ij} \geq 0$. To this end let us go back to Eq. ((ref)) and notice that, combining the equations by taking the difference between the first and the second, between the second and the third and so on, we obtain the following $N$ relations

equation[equation omitted — 554 chars of source]

By definition $\nu_i \lambda_{ij} \geq 0$ for any $i,j$ (because it represents a probability), then it is

equation*[equation* omitted — 66 chars of source]

Since $\Lambda(x)$ is a monotonically increasing function of $x$, the previous inequality is fulfilled if and only if $J_{ij} \geq 0$ $\forall i,j$. Thus, this condition is necessarily true if $\boldsymbol{\pi}$ is in the domain of $f$.

By following the same steps in the opposite direction it is straightforward to prove the reverse relation, that is $J_{ij} \geq 0$ is a sufficient condition to have $f^{-1}(\boldsymbol{\theta}) \in \Pi$. Indeed for any $J_{ij} \geq 0$, the product $\nu_i \lambda_{ij}$ is $0 \leq \nu_i \lambda_{ij} \leq 1$ $\forall i,j$ given the system ((ref)) and $\Lambda(x) \in [0,1]$ for any $x$. Then, by summing all the equations in system ((ref)), one obtains

equation*[equation* omitted — 86 chars of source]

which is also positive and smaller than $1$ if $J_{ij} \geq 0$ $\forall j$. Then, it follows that all the $\lambda_{ij}$ are $0 \leq \lambda_{ij} \leq 1$ $\forall i,j$. Finally, combining the first and last lines of Eq. ((ref)), one finds that $0 \leq \chi_i \leq 1$. This procedure can be repeated for all variables $i=1,...,N$, thus obtaining the inverse mapping $f^{-1}:\Theta\vert_{J_{ij}\geq0}\rightarrow \Pi$ in the subset of the codomain $\Theta$ defined by the condition $J_{ij}\geq0$ $\forall i,j=1,...,N$. This concludes the proof.

Practical implications in model inference

figure[figure omitted — 651 chars of source]

Having formally demonstrated the equivalence between the KIM and the VDAR opens the door to cross-contamination between the literatures in which they were developed. In this section we show that one can use inference methods of the KIM developed in the statistical literature, namely the mean field method mezard2011exact, for improving standard inference methods for discrete autoregressive processes, namely the Yule-Walker equations.

Specifically, a popular method in time series literature for the inference of autoregressive models is the method of moments, which generates the so-called Yule-Walker equations matching the empirically measured moments with the ones implied by the model parameters tsay2014financial. In the context of the VDAR(1) it can be shown mazzarisi2020tail that the Yule-Walker equations read

eqnarray[eqnarray omitted — 259 chars of source]

where $\tilde{\mathbf{X}}_t = \mathbf{X}_t - \mathbb{E}(\mathbf{X}_t)$, $\boldsymbol \phi$ is a $N$-dimensional vector and $\boldsymbol\Psi$ is a $N \times N$ matrix. Solving these linear systems for $\boldsymbol\phi$ and $\boldsymbol\Psi$ allows to obtain the VDAR parameters as $\nu_i = \sum_j \boldsymbol\Psi_{ij}$, $\lambda_{ij} = \boldsymbol\Psi_{ij}/\nu_i$ and $\chi_i = \boldsymbol\phi_i/(1 - \nu_i)$.

On the other hand, the statistical mechanics literature has developed suitable approximation methods of maximum likelihood estimation of KIM. In particular, we consider the method developed by M\'{e}zard and Sakellariou mezard2011exact, which takes advantage of a mean field approximation to infer the parameters of the Kinetic Ising Model. The method is exact if the $J$ generating the data has all $J_{ij} \neq 0$ and can be adapted to a sparse version through $\ell_1$-regularization or decimation methods decelle2015inference. It has been extensively used in applications of the model to real financial and neural data campajola2020modelling, roudi2015multi.

The equivalence between KIM and VDAR we proved in this paper allows us to use the mean field method for doing inference on a VDAR model. To test this idea, we apply mean field and Yule-Walker method on simulated data and compare them both in speed and accuracy. The speed is measured by the time needed to perform the inference as a function of $N$ on a regular commercial laptop, while the accuracy is measured by the bias and root mean squared error (RMSE) of the estimator of $J$ over 500 simulations. We simulate the KIM with uniformly distributed $J_{ij} \sim \mathcal{U}(1/2N, 1/N)$ to keep the model far from the dynamic ferromagnetic transition crisanti1988dynamics, for $T = 10N$ time steps and $N$ ranging from $10$ to $10,000$. For the sake of simplicity we consider $h_i = 0 \, \forall \, i$ in our simulations. We show the results of the comparison in Figure (ref), where it is clear that the Yule-Walker method is faster on small scale systems but is also less accurate. In particular we see that the method of moments has a positive bias and a relatively large RMSE, whereas the mean field method has close to zero bias and a smaller RMSE. The computational effort required in the two methods is comparable as the biggest contribution comes from the inversion of the covariance matrix, typically achieved in $O(N^3)$ operations, hence they present similar execution time for large $N$. Thus the mean field method is a better choice for the inference of the KIM/VDAR(1), and this result is somewhat expected, since mean field is a likelihood-based method whereas the Yule-Walker equations are not.

This simple analysis shows that the equivalence between the two models can be leveraged to identify inference methods, originally developed for the KIM, that can be used in the inference of the VDAR. As shown above, this can lead to significant improvements in performance when compared to standard VDAR inference methods.

Conclusions

In conclusion, the VDAR(1) model is equivalent to the Kinetic Ising Model thanks to the existence of a unique mapping for both the binary random variables and the parameters as long as the $\bm{J}$ parameters of the Kinetic Ising Model are positive or zero, as a consequence of the fact that the $\bm{\nu}$ and $\bm{\lambda}$ parameters only account for non-negative lagged correlations among random variables. Moreover, since the two models can be interpreted as the maximum entropy distribution of binary random sequences with given means, and both auto- and cross-correlations (only non-negative correlations for the specific case of the VDAR(1) model), both of them represent further the best choice in describing such binary random sequences in absence of prior information on other metrics, according to the principle of maximum entropy.

There are several directions in which future research can go to take advantage of our equivalence theorem. We have already shown that inference methods developed in the statistical physics literature can be used to improve model estimation; another straightforward application of this theorem is that of defining an extension of the VDAR(1) including negative correlations, as the equivalent to the KIM without the restriction on the positive $J$ elements. Finally, it is common in autoregressive models to consider Markov chains of order higher than 1, that is where there are explicit parameters linking the value of $X_t$ to the value of $X_{t-k}$ with $k = 1, \dots, p$, $p > 1$. The properties of a Kinetic Ising Model with higher order interactions have not been studied yet to the best of our knowledge, thus opening another interesting perspective to explore in the context of this cross-contamination between very active research fields.

Acknowledgments

FL acknowledges partial support by the European Integrated Infrastructure for Social Mining and Big Data Analytics (SoBigData++, Grant Agreement \#871042). DT acknowledges GNFM-Indam for financial support.