Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
78,402 characters · 11 sections · 49 citation commands
Gravity Models of Networks: Integrating Maximum-Entropy and Econometric Approaches
\email{[email removed]}
\pacs{89.75.Fb; 02.50.Tt; 89.65.Gh}
The World Trade Web (WTW), formed by the import/export relationships between world countries, has received a lot of attention both in traditional economic studies and in more recent investigations. Indeed, as recent crises have clearly pointed out, `network effects matter' Schweitzer2009: trade linkages are, in fact, the most important channels of interaction between world countries, directly transmitting economic shocks Fagiolo2010b,Saracco2016,Squartini2015a or indirectly transmitting financial ones Schiavo2010,Squartini2013,Kali2007,Kali2010,Starnini2019). It therefore comes with no surprise that many contributions have focused on the analysis of the network properties of the WTW. Indeed, the (ever-increasing) data availability over the last years has motivated researchers (coming from disciplines as different as economics, network science and social sciences) to explore the architecture of the WTW from an empirical point of view, complementing the traditional theoretical knowledge in classical trade economics.
These efforts have produced a large wealth of literature characterizing various stylized facts of the WTW architecture Barigozzi2010,Fronczak2012a,Fronczak2012b,Serrano2003,Garlaschelli2004,Garlaschelli2005,Fagiolo2010,Fagiolo2008a,Fagiolo2008b, Schweitzer2009, Vitali2011,Schiavo2010. A first stream of contributions has focused on the purely binary properties of the WTW, i.e. the structural features that require the knowledge of only the presence of connections (trade relationships) between countries, irrespective of their intensity. These studies have highlighted a large (compared with most of the other real-world networks) density, a right-skewed and heavy-tailed degree distribution, disassortative mixing by degree (indicating that countries having many trade partners are, on average, connected with countries having few partners), a hierarchical organization (indicating that partners of well-connected countries are less interconnected than those of poorly-connected ones), the presence of a core of countries trading with almost everyone else, and finally the presence of (a sort of) bow-tie structure Garlaschelli2004,Garlaschelli2005,Fagiolo2008b,Vitali2011,Jeroen2019. A second stream of contributions has focused on the weighted properties of the WTW, where the weight attached to each link represents the volume of the corresponding trade relationship and the strength (sum of all link weights) of each node represents the overall trade volume of the corresponding country. These studies have highlighted right-skewed and heavy-tailed (generally log-normal) weight and strength distributions (indicating that few intense trade connections coexist with a majority of low-intensity ones), disassortative mixing by strength (indicating that countries whose trade volume is large are, on average, connected with countries whose trade volume is small), and a large weighted clustering coefficient (indicating that the trade volume of the partners of well-connected countries is larger than the trade volume of the partners of poorly connected ones) Fagiolo2008a.\\
From the modelling side, two broad classes of models of the international trade system can be roughly identified: econometric models and network models (the latter mainly rooted into statistical physics). Both can be traced back to the earliest model of international trade, proposed in 1962 by the physics-educated Dutch economist Jan Tinbergen (considered the founding father of econometrics together with Ragnar Frisch). Tinbergen modelled the import/export flows between countries via a relationship that is formally analogous to the law of gravity, where the `masses' of countries are replaced by GDPs and the inter-country distances are replaced by suitably defined geographic distances Tinbergen1962. This is the celebrated Gravity Model (GM) and can be shown to reproduce the positive trade volumes quite accurately Anderson2011, Bergeijk2010,DeBenedictis2011,Almog2019. Although our focus here is on international trade, it is worth mentioning that the GM has been used successfully to describe the positive link weights associated to processes relevant to many other networks as well, including migration flows ravenstein,migrationGM (which actually represent the earliest ravenstein application of the GM), mobility and traffic patterns zipf,balcan,koreaGM, communication streams LambiotteGM, and spreading phenomena Ma2016,Li2019. Despite its success, the most evident limitation of the GM consists in predicting that each country establishes a trading relationship with every other country, a result that is in stark contrast with empirical data, where `missing' trade relationships are actually found to be a significant proportion (up to one half) of the number of all possible pairs of countries. Since systematically overestimating the number of connections is known to lead to a significant miscalculation of network effects Schweitzer2009, one should avoid the use of the pure GM as a reliable network model of the WTW Almog2019.
From the Sixties on, a lot of work has been done in econometrics to overcome the aforementioned limitation. Initially, Eaton and Tamura Eaton1995 and Martin and Pham Martin2008 suggested to employ a tobit-like estimation procedure, by rounding to zero the trade flows below a certain threshold. Later, the serious conceptual problem posed by the arbitrariness of the chosen threshold motivated Helpman, Melitz and Rubinstein Helpman2008 to propose a two-step estimation procedure: first, a probit model is employed to estimate the probability of observing a trade relationship between any two nodes; subsequently, the corresponding trade flow is estimated via an Ordinary-Least-Squares (OLS) regression whose parameters are tuned on the entire set of positive weights. An alternative algorithm is the one proposed by Silva and Tenreyro Silva2006, who estimated the gravity equation in a multiplicative - rather than additive - fashion by employing a Poisson Pseudo-Maximum Likelihood (PPML) method and obtained good estimates of trade flows, robust to heteroskedasticity. In the following years, PPML has been proven to be very sensitive to the `excess zeros' of the dependent variables. Hence, a different class of two-step estimation procedures has been devised, i.e. the so-called `zero-inflated' (ZI) methods Burger2009,Winkelmann2008: these models are defined by a logit estimation, aimed at establishing the presence of a link, followed by either a Poisson (ZIP) or a negative binomial (ZINB) regression, aimed at estimating the corresponding trade volume. Later in 2011, Due\ nas and Fagiolo Duenas2011 proved that the weighted structure of the WTW is well predicted by GMs if and only if its topological structure is specified as well. This result marks an important shift from a perspective whose only focus use to be the prediction of trade flows to a perspective where the topological structure of the network becomes one of the reconstruction targets.\\
Coming to the physics-inspired approaches to international trade, the majority of proposed models are maximum-entropy models of networks Squartini2011a,Squartini2011b,Squartini2011c,Mastrandrea2014,Bargigli2014,DSBook,Almog2019. The maximum-entropy (ME) framework Jaynes1957a,Jaynes1957b,Jaynes1982,Cimini2019,Parisi2020,DSBook,Squartini2011b allows for purely structural network properties to be constrained in order to derive the maximally unbiased probability distribution compatible with those properties. A series of publications over the years Garlaschelli2004,Squartini2011a, Squartini2017,Squartini2015b,Cimini2021,Squartini2011c,Almog2019 has showed that fixing the degree sequence of world countries is enough to successfully reproduce higher-order binary properties such as the average nearest-neighbour degree and the clustering coefficient. By contrast, this result is no longer retrieved when the analysis is reformulated within a purely weighted framework: for instance, it turns out that the strength sequence is not an effective constraint in reproducing the average nearest-neighbour strength and the weighted clustering coefficient Squartini2011b,Fagiolo2012. In fact, accuracy in the reconstruction of weighted properties is recovered only if the degrees and the strengths are constrained jointly Mastrandrea2014b,Garlaschelli2009,Gabrielli2019.\\
Econometric and network approaches have mostly proceeded along parallel tracks, with little interaction so far. The only line of research that has systematically looked for the possibility of an integration between the two approaches is the one reformulating maximum-entropy network ensembles as hidden variable (or fitness Caldarelli2002) models Garlaschelli2004,Garlaschelli2005,Almog2015,Almog2017,Almog2019. In this approach, the maximum-entropy construction of the network ensemble is formally preserved, but the Lagrange multipliers usually viewed as free parameters tuned in order to enforce the chosen constraints are instead identified with empirical macroeconomic factors, most notably the GDP of countries. Building on those results, as well as on the success of physics-inspired models in reconstructing various kinds of socio-economic and financial systems Cimini2021,Squartini2018, here we explore the potential of the maximum-entropy formalism to provide a viable framework for econometrics too.\\
The rest of the paper is organized as follows. Section (ref) reviews the traditional GM and its performance in reproducing the positive WTW weights. Section (ref) is devoted to the description of purely econometric models and their application to the analysis of trade (networks). Section (ref) is devoted to the description of maximum-entropy network models and their application to the WTW. Section (ref) illustrates the main results of the paper, i.e. the integration of purely econometric and physics-inspired models. Section (ref) concludes and presents an outlook on possible future extensions.
From a merely econometric point of view, the simplest exercise is that of reproducing the realized (positive) trade volumes (link weights) of the WTW. To this aim, we have considered two different datasets. The first one is curated by Gleditsch Gleditsch2002 and includes yearly trade volumes, yearly GDP values (both reported in millions of US dollars) and the (time-independent) matrix of geographic distances between capital cities of all countries in the data. The second one is the BACI dataset, a detailed description of which can be found in baci1,baci2. For both datasets we have selected and analyzed eleven years: 1990-2000 for the Gleditsch dataset and 2007-2017 for the BACI one. The year 2000 of the Gleditsch dataset - i.e. the one with the largest number of countries ($176$) - is the snapshot we have selected to graphically illustrate all the results of our analyses. We have always considered the undirected (symmetrized) version of the weighted trade matrix, whose generic entry reads $w_{ij}=\frac{\text{exp}_{ij}+\text{exp}_{ji}}{2}$, i.e. $w_{ij}$ is the bilateral trade volume defined as the arithmetic mean of the export volume from country $i$ to country $j$ and of the export volume from country $j$ to country $i$.
Our econometric exercise is carried out by comparing the empirical, positive weights of the WTW, in the year 2000, with four different specifications of the following econometric function
where $\omega_i=\frac{\text{GDP}_i}{\overline{\text{GDP}}}$ is the GDP of country $i$ divided by the arithmetic mean of GDPs, $d_{ij}$ is the geographic distance between the capitals of countries $i$ and $j$ and $\underline{\theta}=(\rho,\beta,\gamma)$. The first specification of the model above is characterized by the assignment $\rho=\beta=1$ and $\gamma=0$ and its performance in reproducing the positive weights of the WTW is shown in Fig. (ref)\subref{fig1a}. The overall poor performance of this version of the GM signals the presence of both a scaling problem - estimates are positively correlated with observations, only shifted to the right - and of a dimensional problem - in fact, pure numbers appear on the right-hand side while the WTW weights are measured in (multiples of) dollars.
A second specification of the GM solves both problems at once. This specification is characterized by the assignment $\beta=1$ and $\gamma=0$ while $\rho$ is treated as a free parameter, tuned by requiring that the total weight is reproduced, i.e. that $W=\sum_{i<j}w_{ij}=\sum_{i<j}\langle w_{ij}\rangle_\text{GM}=\langle W\rangle_\text{GM}$: the fit is, now, much more accurate, as shown in Fig. (ref)\subref{fig1b}. The model can be further enriched by adding dyadic factors such as the geographic distances between capitals: some more accuracy in the description of the empirical data is indeed gained, as the reduced dispersion of the cloud of points around the identity witness. Quite remarkably, the picture does not change much if, now, we let the entire set of parameters to be tuned as described in Appendix A (see Fig. (ref)\subref{fig1c}).
Although the GM leads to a good prediction of the positive weights, its intrinsic limitation is that it does not allow the topological structure of the WTW to be correctly recovered. In fact, by outputting only positive weights, it induces a trivial, fully-connected structure: upon defining $a_{ij}=\Theta[\langle w_{ij}\rangle_\text{GM}]$, $\forall\:i<j$, it is evident that $a_{ij}=1$, $\forall\:i<j$. In order to overcome such a limitation, one can `refine' the plain gravity model by `dressing' it with a probability distribution capable of accounting for the null entries as well.
In very general terms, we need to define a statistical network model, i.e. a set of mathematical relationships between the random variables that are of interest for our network description. Two broad classes of such models can be identified, i.e. the econometric ones and the ones rooted into statistical physics. In what follows, we will deal with (either econometric or physics-rooted) discrete statistical models.
Let us start with the description of some of the most representative members of the econometric class of models.
The simplest model in this class prescribes to consider $\langle w_{ij}\rangle_\text{GM}=\rho(\omega_i\omega_j)^{\beta}d_{ij}^{\gamma}$ as the expected value of a Poisson probability mass function:
since $\langle w_{ij}\rangle_\text{Pois}=z_{ij}$, the explanatory power of the GM is retained upon requiring that $\langle w_{ij}\rangle_\text{Pois}=\langle w_{ij}\rangle_\text{GM}$, i.e. by posing
The expected topology of the network is determined by the expected adjacency matrix entries implied by the model, which is captured by the expression $\langle a_{ij}\rangle_\text{Pois}=p_{ij}^\text{Pois}=1-q_{ij}^\text{Pois}(0)=1-e^{-z_{ij}}$ (see Appendix B for a detailed description of the procedure to estimate the parameters of the Poisson model).
The main drawback of the Poisson model is that of predicting a variance of the weights that is necessarily equal to their average value. In general, this may be different from what empirical analyses suggest. In order to overcome this problem, econometricians have considered a different probability mass function, namely the negative binomial one with the introduction of an overdispersion\footnote{In the Poisson case, one has that $\sigma^2_\text{Pois}[w_{ij}]=z_{ij}=\langle w_{ij}\rangle_\text{Pois}$; hence, variance cannot be adjusted independently from the mean. In the negative binomial case, instead, $\sigma^2_\text{NB}[w_{ij}]= z_{ij}(1+\alpha z_{ij})=z_{ij}\left(1+\frac{z_{ij}}{m}\right)=\langle w_{ij}\rangle_\text{NB}(1+\alpha \langle w_{ij}\rangle_\text{NB})=\langle w_{ij}\rangle_\text{NB}\left(1+\frac{\langle w_{ij}\rangle_\text{NB}}{m}\right)$ and the variance can be increased to overcome the problem of overestimating the link density.} parameter $\alpha=m^{-1}$ Hilbe:
One finds that $\langle w_{ij}\rangle_\text{NB}=m\alpha z_{ij}=z_{ij}$. The requirement that the expected value of the negative binomial distribution coincides with the prediction coming from the GM, i.e. $\langle w_{ij}\rangle_\text{NB}=\langle w_{ij}\rangle_\text{GM}$, can be, again, realized by posing $z_{ij}=\rho(\omega_i\omega_j)^{\beta}d_{ij}^{\gamma}$. Predictions about the topology are, now, carried out via the expression $\langle a_{ij}\rangle_\text{NB}=p_{ij}^\text{NB}=1-q_{ij}^\text{NB}(0)=1-\left(\frac{1}{1+\alpha z_{ij}}\right)^m$ (see Appendix B for a detailed description of the procedure to estimate the parameters of the negative binomial model).
The main drawback of the econometric models above is that of failing in reproducing the link density of the WTW. For instance, the latter equals $c=\frac{2L}{N(N-1)}\simeq0.63$ in year 2000. It turns out that, while the Poisson model overestimates this quantity, the negative binomial one underestimates it, i.e.
where $\langle c\rangle_\text{Pois}=\sum_{i<j}p_{ij}^\text{Pois}\simeq0.68$ and $\langle c\rangle_\text{NB}\sum_{i<j}p_{ij}^\text{NB}\simeq0.60$. For this reason, econometricians have defined the so-called zero-inflated (ZI) models, i.e. two-step recipes whose general form reads
a relationship indicating that the probability of the (network represented by the) weighted adjacency matrix $\mathbf{W}$ can be obtained as the product of the probability $P(\mathbf{A})$ of observing the purely binary adjacency matrix $\mathbf{A}$ and the conditional probability $Q(\mathbf{W}|\mathbf{A})$ - where, for consistency, $\mathbf{A}=\Theta[\mathbf{W}]$, i.e. $a_{ij}=\Theta[w_{ij}]$, $\forall\:i<j$, the position $a_{ij}=\Theta[w_{ij}]$ meaning that $a_{ij}=1$ whenever $w_{ij}>0$ and $a_{ij}=0$ if and only if $w_{ij}=0$.\\
The simplest ZI model is the Poisson one, defined by the positions
Notice that $1-p_{ij}^\text{ZIP}=\frac{1}{1+G_{ij}}+\frac{G_{ij}}{1+G_{ij}}e^{-z_{ij}}$, i.e. the connection between nodes $i$ and $j$ can be missing either because a link is not there (with probability $\frac{1}{1+G_{ij}}$) or because a link is there but has zero weight (with probability $\frac{G_{ij}}{1+G_{ij}}e^{-z_{ij}}$). Consistently, $i$ and $j$ are connected because the weight is not zero (with probability $1-e^{-z_{ij}}$). In order to `dress' the GM, we need to identify some of the parameters of the Poisson model with the usual econometric function. Since
we can make the identification $z_{ij}=\rho(\omega_i\omega_j)^{\beta}d_{ij}^{\gamma}$. Upon doing so, we are treating $z_{ij}$ as an `effective' conditional weight: in fact, Eq. ((ref)) can be understood as describing an aleatory experiment that combines a logit with a full Poisson step. According to this interpretation, $z_{ij}$ would represent a Poisson-like expected weight, conditional to the success of the logit step, i.e. $z_{ij}=\frac{\langle w_{ij}\rangle_\text{ZIP}}{p_{ij}^\text{logit}}$, with $p_{ij}^\text{logit}=\frac{G_{ij}}{1+G_{ij}}$.
A second econometric identification is, however, needed: we will proceed by imposing
(see Appendix B for a detailed description of the procedure to estimate the parameters of the ZIP model).
The ZI version of the negative binomial model, instead, is defined by
where, as the for the plain negative binomial model, $\alpha=m^{-1}$ and $\tau_{ij}=\left(\frac{1}{1+\alpha z_{ij}}\right)^m$. Moreover, as for the zero-inflated Poisson (ZIP) model, $1-p_{ij}^\text{ZINB}=\frac{1}{1+G_{ij}}+\frac{G_{ij}}{1+G_{ij}}\tau_{ij}$, i.e. the connection between nodes $i$ and $j$ can be missing either because a link is not there (with probability $\frac{1}{1+G_{ij}}$) or because a link is there but has zero weight (with probability $\frac{G_{ij}}{1+G_{ij}}\tau_{ij}$); consistently, $i$ and $j$ are connected because the weight is not zero (with probability $1-\tau_{ij}$). In order to `dress' the GM, we need to identify some of the parameters of the negative binomial model with the usual econometric function. Upon considering that
we can make the identification $z_{ij}=\rho(\omega_i\omega_j)^{\beta}d_{ij}^{\gamma}$ and $G_{ij}=\delta\omega_i\omega_j$.
As for the ZIP case, we are treating $z_{ij}$ as a negative binomial-like expected weight, conditional to the success of a logit step, i.e. $z_{ij}=\frac{\langle w_{ij}\rangle_\text{ZINB}}{p_{ij}^\text{logit}}$, with $p_{ij}^\text{logit}=\frac{G_{ij}}{1+G_{ij}}$ (see Appendix B for a detailed description of the procedure to estimate the parameters of the ZINB model).\\
Let us notice that, while the ZIP model provides a better estimation of the link density than the Poisson model, the ZINB and the negative binomial ones basically perform in the same way. In fact,
since $c=\frac{2L}{N(N-1)}\simeq0.63$, $\langle c\rangle_\text{ZIP}=\sum_{i<j}p_{ij}^\text{ZIP}\simeq0.63$ and $\langle c\rangle_\text{ZINB}=\sum_{i<j}p_{ij}^\text{ZINB}\simeq0.60$, a result suggesting that both variants of the negative binomial model will perform poorly in reproducing the binary properties of the WTW.
The members of the second class of network models are the ones defined within the framework of traditional statistical mechanics. All of them can be derived by performing a constrained maximization of Shannon entropy Squartini2015b where the constraints represent the available information about the system at hand.\\
The simplest, yet non trivial, ME model that can be considered comes from the maximization of the binary Shannon functional
constrained to reproduce the entire degree sequence, $\{k_i(\mathbf{A})\}_{i=1}^N$, of the network. This model is known under the name of Undirected Binary Configuration Model (UBCM) and has been shown to accurately reproduce many (binary) properties of a wide spectrum of real-world systems Squartini2011a.
The UBCM is described by the probability mass function
which is factorized into the product of Bernoulli probability mass functions (one for each pair of nodes) with
(where $x_i$ is the Lagrange multiplier controlling for the degree of node $i$). Importantly, the logit model admitting the presence of a single global constant can be derived from entropy maximization upon re-parametrizing the Lagrange multipliers of the UBCM and imposing the total number of links as the only constraint Squartini2018). The identification $x_i\equiv\sqrt{\delta}\omega_i$, in fact, leads to
Although the functional form above is not the most general one (for instance, dyadic factors such as geographic distances could be added as well), it is the form we will adopt in what follows.
In the network literature, the logit model (in its formulation above) has been popularized Garlaschelli2004 as one particular case of the so-called fitness model Caldarelli2002 and as the so-called density-corrected Gravity Model (dcGM) Cimini2015 and has been proven to perform remarkably well for the task of reconstructing the topology of networks from partial information Squartini2018.\\
Since we are interested in reproducing the structural properties of weighted networks, we need to complement the purely binary step above with a recipe for reconstructing weights. The entropy-based framework handles such a requirement via the maximization of conditional Shannon functionals allowing the specification of $P(\mathbf{A})$ to be disentangled from that of $Q(\mathbf{W}|\mathbf{A})$ Parisi2020.
When discrete weighted models are considered, a useful quantity is the conditional Shannon entropy
where the first sum runs over all binary configurations within the ensemble $\mathbb{A}$ and the second sum runs over all weighted configurations that are compatible with each specific binary structure represented by the adjacency matrix $\mathbf{A}$, i.e. such that $\mathbb{W}_\mathbf{A}=\{\mathbf{W}:\Theta[\mathbf{W}]=\mathbf{A}\}$.
Conditional maximization proceeds by specifying a set of weighted constraints that, in the discrete case, reads
the first condition ensuring the normalization of the conditional probability mass function and the vector $\{C_\alpha(\mathbf{W})\}$ representing the `proper' set of weighted constraints. Differentiating the corresponding Lagrangean functional with respect to $Q(\mathbf{W}|\mathbf{A})$ and equating the result to zero leads to
where $H(\mathbf{W})=\sum_\alpha\psi_\alpha C_\alpha$ is the so-called Hamiltonian, listing the constrained, weighted quantities, and $Z_\mathbf{A}=\sum_{\mathbb{W}_\mathbf{A}}e^{-H(\mathbf{W})}$ is the partition function for fixed $\mathbf{A}$. The explicit functional form of $Q(\mathbf{W}|\mathbf{A})$ can be obtained only once the functional form of the constraints has been specified as well.\\
To this aim, let us consider the Hamiltonian
where weights are modelled as non-negative integer variables, i.e. $w_{ij}\in\mathbb{N}$, $\forall\:i<j$. This choice induces a conditional probability mass function reading
(with $e^{-\psi_{ij}}=e^{-\phi_{ij}}=y_{ij}$). Let us, now, turn the model above into a proper econometric one. To this aim, let us proceed by analogy. All zero-inflated econometric recipes identify $z_{ij}$ with a conditional expected weight, a prescription that in our case, would translate into its identification with $\langle w_{ij}|a_{ij}\rangle=\frac{1}{1-y_{ij}}$. This choice, however, would lead to an inconsistency, since $z_{ij}>0$ is a positive real number while $\langle w_{ij}|a_{ij}\rangle$ must necessarily exceed 1, as it represents the expected weight conditional to the existence of a connection. An alternative, consistent econometric identification is
which, in turn, induces a conditional probability mass function
Models of the kind are known as hurdle models: quite remarkably, entropy maximization allows us to recover them in a fully principled way, i.e. by eliminating the (otherwise unavoidable) ambiguity that accompanies the choice of the distribution (supposedly) describing the positive values of an economic system.
The hurdle-geometric model derived above, however, suffers from a number of limitations, the most relevant of which is that of failing in reproducing basic network quantities such as the WTW total weight. As an illustrative example, while $W=\sum_{i<j}w_{ij}\simeq 10^9$, in the year 2000, we find that $\langle W\rangle_\text{h-g}\simeq 10^5$. In order to overcome such a limitation, we have considered the conditional probability mass function induced by the Hamiltonian
Identifying $e^{-\phi_0}=y_0$ and $e^{-\phi_{ij}}=y_{ij}=\frac{z_{ij}}{1+z_{ij}}$ (see Eq. (ref)), we arrive at the modified econometric model
We are now ready to fully specify the suite of discrete entropy-models that we will compare with the aforementioned, purely econometric ones. To this aim, we need to fully specify the functional form
the two most obvious choices are represented by the models
and
that combine the weighted, conditional step induced by the Hamiltonian defined in Eq. ((ref)) with the purely binary logit model and with the Undirected Binary Configuration Model, respectively. The acronyms stand for `two-step fitness' model and `two-step' model and recall the names originally used to define them Garlaschelli2005,Almog2017.\\
Less trivial choices are represented by models whose both binary and weighted portions are jointly determined by the constraints. They can all be recovered as specifications of the generic Hamiltonian
in what follows, we will consider two different instances of such a function, defined by the choices $\theta_{ij}=\theta_0$ and $\theta_{ij}=\theta_i+\theta_j$. In other words, while we let the weighted parts of these models coincide and read as in Eq. ((ref)), we allow for the binary part to vary, either constraining the total number of links, $L$, or the entire degree sequence $\{k_i(\mathbf{A})\}_{i=1}^N$. In the first case, our Hamiltonian reads
and instances the model in Eq. ((ref)) with
(having posed $e^{-\theta_0}=x$, $e^{-\phi_0}=y_0$ and $e^{-\phi_{ij}}=y_{ij}=\frac{z_{ij}}{1+z_{ij}}$); in the second case, it reads
and instances the model in Eq. ((ref)) with
(having posed $e^{-\theta_i}=x_i$, $e^{-\phi_0}=y_0$ and $e^{-\phi_{ij}}=y_{ij}=\frac{z_{ij}}{1+z_{ij}}$). Appendix C provides a detailed description of the procedure we have adopted to estimate the parameters entering into the definition of our basket of discrete ME models.\\
So far, we have turned entropy-based models into econometric ones via a suitable econometric transformation of the Lagrange multipliers defining the proper `physical' models. The entropy-based formalism, however, also offers the opportunity to define statistical models in a fully data-driven fashion. To this aim, let us consider the Hamiltonian
that constrains both degrees and strengths. The model induced by the latter ones is called Undirected Enhanced Configuration Model (UECM) and represents the best-performing one for the task of network reconstruction in presence of full information about the constraints Mastrandrea2014,Garlaschelli2009.\\
Remarkably, all models considered in the previous Section can be compactly derived by combining a logit-like probability mass function describing the binary network structure with the conditional expression defined in Eq. ((ref)). To prove this, it is enough to notice that all the Bernoulli-like probability mass functions characterizing our model can be compactly rewritten as
where
Let us now test and compare the performance of our two broad classes of models in reproducing the topological properties of the World Trade Web. To this aim, let us consider both the local properties, such as the degrees and the strengths, and the higher-order ones such as the average nearest neighbors degree (ANND) and the clustering coefficient (BCC), i.e.
we will also consider their weighted counterparts, i.e. the average nearest neighbors strength (ANNS) and the weighted clustering coefficient (WCC), defined as
We will also test the accuracy of the reconstruction provided by the methods in our basket by considering properties like the true positive rate (TPR)
i.e. the percentage of links correctly recovered by a given reconstruction method, the specificity (SPC)
i.e. the percentage of zeros correctly recovered by a given reconstruction method, the positive predictive value (PPV)
i.e. the percentage of links correctly recovered by a given reconstruction method with respect to the total number of predicted links and the accuracy (ACC)
measuring the overall performance of a given reconstruction method in correctly placing both links and zeros.
Fig. (ref) sums up the comparisons carried out between the econometric models and the ME ones. The comparison between the empirical cumulative density function (CDF) of the degrees and the ones output by the econometric models reveals the latter ones to be able to predict an overall similar functional form (see Fig. (ref)\subref{fig2a} and Fig. (ref)\subref{fig2d}); still, the prediction obtained by any of the ME models is much closer to the empirical trend. More quantitatively, one can implement the Kolmogorov-Smirnov (KS) test to check the goodness of any of the models considered in the present work to replicate the empirical degrees: while any of the ME models provides estimates of the degrees that are compatible with the empirical CDF (at the significance level of $5\%$), only the ZIP model predicts degrees that are compatible with the empirical ones: in fact, the p-values of the ME models read $p_{(1)}\simeq 0.06$, $p_{(2)}\simeq 0.99$, $p_\text{TS}\simeq 0.99$, $p_\text{TSF}\simeq 0.32$ while the p-values of the econometric models read $p_\text{Pois}\simeq 0.001$, $p_\text{NB}\simeq 0.0008$, $p_\text{ZIP}\simeq 0.63$, $p_\text{ZINB}\simeq 0.001$.
Coming to the higher-order properties, it is evident that the majority of the econometric models fails to overlap with the empirical cloud of points (see Fig. (ref)\subref{fig2b}): the one providing the best prediction is the ZIP model, whose performance represents quite an improvement with respect to the one provided by the `plain' Poisson model. While this is quite evident for what concerns the prediction of the ANND values, the performances of the ZIP and of the `plain' Poisson model become less different when tested on the BCC values. On the contrary, the performance of the ZINB model closely resembles that of the negative binomial one when tested both on the ANND and on the BCC. As for the local properties, the KS test reveals that the only model outputting predictions compatible with the empirical values (at the significance level of $5\%$) is the ZIP one: in fact, $p_\text{ZIP}^\text{ANND}\simeq 0.38$, $p_\text{ZIP}^\text{BCC}\simeq 0.08$).
For what concerns ME models, the ones performing best are those constraining the degrees, i.e. the model induced by $H_{(2)}$ and its two-step counterpart, whose topological estimation step is carried out by employing $p_{ij}^\text{UBCM}$. The evidence that their performances in reproducing the purely binary structure of a network are very similar lets us suspect that $p_{ij}^{(2)}\simeq p_{ij}^\text{UBCM}$ and conclude that the purely econometric information encoded into $p_{ij}^{(2)}$ does not add much to what is already conveyed by the purely topological one. On the other hand, ME models not constraining the degrees provide predictions differing from the empirical trends to quite a large extent. As the KS test reveals, the only ME model outputting predictions that are not compatible with the empirical values (at the significance level of $5\%$) is the one induced by $H_{(1)}$.
The overall accuracy of our models in reproducing a network topology can be proxied by the index $\Delta_L=|\langle L\rangle-L|/L$ amounting at $\Delta^\text{Pois}_L\gtrsim\Delta^\text{NB}_L=\Delta^\text{ZINB}_L\simeq 6\%$ while $\Delta^\text{ZIP}_L\simeq 0.5\%$ and $\Delta^\text{ME}_L=0$ for each ME model. This is confirmed by our analysis of single link statistics: in fact, $\langle ACC\rangle_{(2)}\simeq 0.83$ attains the largest value, followed by $\langle ACC\rangle_\text{TS}\simeq 0.81$ and $\langle ACC\rangle_\text{ZIP}\simeq 0.77$. Remarkably, $\langle PPV\rangle_{(2)}\simeq 0.86$ attains the largest value, indicating that the ME model induced by $H_{(2)}$ is the one placing links best among all the models in our basket.\\
Let us now consider the weighted properties (see Fig. (ref)). Overall, the distribution of the strengths is reproduced quite well by all models considered here, although no one explicitly constrains them. This seems to indicate that the purely econometric information `feeded' into our models indeed plays a role - which, however, is limited to ensure that the intensive margins (and the related properties, as we will see) are accurately predicted. The larger explanatory power of econometric models becomes now evident: all of them output predictions that are compatible with the empirical values. Although the same result holds true for ME models, the latter ones are outperformed by purely econometric models - the best performing ones in predicting the strengths being the Poisson-like ones.
Coming to the higher-order properties, let us notice that the best performing econometric models in reproducing the ANNS values are the ZIP and the `plain' Poisson ones whose performances differ less than in the ANND case - although the KS test lets the ZIP model win. On the other hand, the ZINB and the negative binomial models (whose performances are, again, very similar) completely fail in capturing the empirical values. All predictions from ME models overlap with the empirical ANNS values: as the KS test reveals, the only ME model outputting predictions that are not compatible with the empirical values (at the significance level of $1\%$) is the one induced by $H_{(1)}$. For what concerns the values of the WCC, both the econometric and the ME models perform quite satisfactorily in capturing its rising trend. However, the KS test reveals that only the econometric models and the TS model output predictions compatible with the WCC empirical values (at the significance level of $1\%$).
To proxy the accuracy of our models in reproducing the weighted network structure we have considered the index $\Delta_W=|\langle W\rangle-W|/W$ amounting at $\Delta^\text{ZINB}_W\simeq 95\%$, $\Delta^\text{NB}_W=60\%$, $\Delta^\text{ZIP}_W\simeq 0.3\%$ and $\Delta^\text{Pois}_W\simeq 0$ while $\Delta^\text{ME}_W=0$ for the ME models constraining the total weight and $\Delta^\text{TS}_W\gtrsim\Delta^\text{TSF}_W\simeq 0.2\%$ for the ME two-step ones.
In order to understand if the conclusions above can be generalized, let us calculate the accuracy of all models in our basket, for all years constituting our two datasets. The results, summed up in Tab. (ref), confirm that model $H_{(2)}$ systematically outperforms all competing models. As an additional test, we have calculated the percentage of times the empirical values of our network statistics are compatible with their ensemble distributions, via KS tests at the significance level of $5\%$: the results, shown in Fig. (ref), point out that ME models are the ones for which compatibility is largest.\\
Let us now ask ourselves if a criterion exist to carry out a principled comparison of the performance of the models considered in the present work. The answer is positive and lays in the adoption of the popular Akaike Information Criterion and Bayesian Information Criterion, respectively defined as
and
where $\mathcal{L}$ is the log-likelihood of the tested model evaluated in its maximum, $K$ is the number of parameters characterizing the model itself and $n$ is the cardinality of the set of observations - estimated as $\frac{N(N-1)}{2}$ for undirected network data. Model selection based on these criteria prescribes to rank models according to (either) their AIC or BIC value and choose the one minimizing it. Tab. (ref) shows both the AIC and the BIC values for all the models considered here.
Quite surprisingly, the negative binomial model is the favoured one among the econometric models, followed by its zero-inflated version; however, its bad performance in reproducing the empirical trends makes the choice of including it among the most suitable models for modelling trade data highly questionable. On the other hand, the ZIP model performs much better in reconstructing the trends of both local and higher-order properties although being much less parsimonious than both versions of the negative binomial model. Apparently, then, the question about which model to prefer - i.e. the favoured one by information criteria or the best performing one in reproducing trends? - cannot be properly answered by just considering purely econometric models. On the other hand, such a question can be unambiguously answered as soon as one switches to the class of maximum-entropy models: now, both the AIC- and the BIC-based rankings favour the model described by the Hamiltonian $H_{(2)}$ - the one encoding the information about the degree sequence and the total weight, plus admitting a tunable function of the weights - i.e. precisely the most accurate in replicating many (if not all) empirical trends.
For the sake of comparison, we have included into the basket of maximum-entropy models the Undirected Enhanced Configuration Model (UECM), i.e. the model performing best in presence of complete information about the constraints - degrees and strengths, in the specific case - of a given networked system: as evident from the table, it is disfavoured with respect to the model described by the Hamiltonian $H_{(2)}$, an evidence signalling that while the information encoded into the degrees is essential (i.e. the latter ones must be explicitly constrained), the one carried by the strengths appears to be `less fundamental' since providing a good approximation of them is enough to obtain an overall good reconstruction.
In order to understand if the conclusions above can be generalized, let us calculate the Akaike weights for the models in our basket $\mathcal{M}$. The Akaike weight for the $i$-th model is defined as
with $\Delta_i=\text{AIC}_i-\min\{\text{AIC}_m\}_{m\in\mathcal{M}}$. Results on the dataset curated by Gleditsch show that the negative binomial and the $H_{(2)}$ models `compete', in the sense that $H_{(2)}$ performs best (i.e. $w_{H_{(2)}}\simeq 1$) in the (bunches of) years 1990-1993 and 1997-2000 while the negative binomial model performs best (i.e. $w_{NB}\simeq 1$) in the (bunch of) years 1994-1996. For what concerns the BACI dataset, instead, the competing models are three: in fact, while $H_{(2)}$ performs best in the (bunches of) years 2007, 2009 and 2015-2017, the negative binomial outperforms the others in the (bunches of) years 2008 and 2010-2014; however, the ZINB model has a positive, non-negligible Akaike weight in the years 2008, 2012 and 2014, hence performing as well as the negative binomial one.\\
Let us now consider a couple of additional exercises, carried out on both datasets considered in the present work.
The first one concerns link prediction and was carried out by following the reference Ashraf2019. Specifically, we have approached link prediction from a temporal perspective, inspecting the accuracy achieved by our reconstruction models at time $t+1$ given the knowledge about the network topology (for the maximum-entropy models) and of the other exogenous variables (for the purely econometric models) at time $t$. In other words, we opted for a one-lagged link prediction, calculating the log-likelihood
for each statistical model in our basket, the coefficients $\left\{p_{ij}^{(t)}\right\}_{i,j=1}^N$ being the probabilities output by any model at time $t$ and the coefficients $\left\{a_{ij}^{(t+1)}\right\}_{i,j=1}^N$ being the entries of the adjacency matrix at time $t+1$. When carried out on the pairs of years 1993-1994, 1994-1995, 1995-1996, 1996-1997, 1997-1998, 1998-1999 and 1999-2000 of the dataset curated by Gleditsch and on the pairs of years 2008-2009, 2009-2010, 2010-2011, 2013-2014, 2014-2015,2015-2016 and 2016-2017 of the BACI dataset, the exercise above shows $H_{(2)}$ to outperform not only the entire class of econometric models but also the purely binary, maximum-entropy ones - as confirmed by the Akaike weights induced by the log-likelihood above.
To provide a more refined picture of the performance of the models in our basket in providing one-lagged predictions, we have also calculated their one-lagged accuracy, defined as
with $\langle TP\rangle_{1l}=\sum_{i<j}a_{ij}^{(t+1)}p_{ij}^{(t)}$ and $\langle TN\rangle_{1l}=\sum_{i<j}\left(1-a_{ij}^{(t+1)}\right)\left(1-p_{ij}^{(t)}\right)$. The results are reported in Tab. (ref) and confirm what has been previously said: $H_{(2)}$ is the one performing best.
As a second exercise, we have tested the accuracy of our models in estimating link-specific weights - hence, carrying out what we have called a `weight prediction' exercise. To this aim, we have considered each expected weight and calculated the confidence interval enclosing the $95\%$ of total probability around it. On the practical side, we have sampled 1000 configurations from the ensemble induced by each model in our basket and calculated the (ensemble-induced) 2.5 and 97.5 percentiles for each specific weight; then, we have calculated the percentage of empirical weights `falling' within the corresponding CIs - now treated as error bars `accompanying' the point-estimate of each weight. The results of this exercise are reported in in Tab. (ref): as it can be appreciated, maximum-entropy models compete with both the negative binomial and the ZINB ones - although the latter (slightly) outperform the former.
In the present work, we have compared the performance of two broad classes of statistical models, i.e. the ones rooted into economic theory and the ones rooted into statistical physics (in particular, the ones derived from the maximum-entropy principle) in reconstructing both the binary and the weighted network properties of an economic system such as the WTW.
Although the case study is the same, the two classes of models `reflect' the languages of two disciplines that are still deeply different: while econometricians have traditionally focused on bilateral trade volumes between countries - emphasizing the role played by common borders, language, religion, the presence of regional trade agreements, etc. on trade relationships - network scientists have, instead, paid more attention to the structural and dynamical aspects of network formation, emphasizing the role played by purely structural information in determining the topology itself. The 2008 global financial crisis has dramatically clarified that bilateral trade relationships can explain only a small fraction of the impact that an economic shock, originating in a given country, can have on another country which is not a direct trade partner, urging researchers in economics to adopt a network-aware perspective. This, in turn, has motivated us to carry out a methodological comparison on real-world cases, with the aim of clarifying pros and cons of both approaches.
Researchers in economics have dealt with the issue of reconstructing network topology by approaching the simpler problem of reproducing the number of missing connections - or, equivalently, the link density. For instance, as we see from the year 2000 snapshot of the dataset curated by Gleditsch Gleditsch2002, although the error of the Poisson model in reproducing $L$ is overall small (amounting at $\Delta^\text{Pois}_L\simeq 7\%$), it can be further reduced by adopting the zero-inflated version of it. On the contrary, inflating zeros does not improve the performance of the negative binomial model in reproducing the link density since it already underestimates $L$. The ability of a model in reproducing a global quantity such as the link density proxies its ability in providing a good estimation of local as well as higher-order topological properties (i.e. the degrees, the ANND and the BCC): from this point of view, the zero-inflated Poisson model is the one performing best among the econometric models. However, it is largely disfavoured by information criteria such as AIC and BIC, a result suggesting that it may be not parsimonious enough.
Some of the problems of purely econometric models are solved by looking at a different class of statistical models, i.e. the physics-inspired ones. In particular, the model described by the Hamiltonian $H_{(2)}=\sum_i\theta_ik_i+\psi_0 W+\sum_{i<j}\psi_{ij}w_{ij}$ provides a very accurate reconstruction while being favoured by information criteria. Remarkably, although it is defined by $N+1$ purely topological constraints, both AIC and BIC reveal that the latter are `irreducible', i.e. necessary to provide a satisfactory explanation of the network generating process. For the sake of comparison, Fig. (ref) explicitly shows the performance of the models favoured by the adopted information criteria (i.e. the negative binomial model and its zero-inflated version) with that of the ME model described by $H_{(2)}$: it is evident that the ME model outperforms the purely econometric ones, still achieving a good ranking.
Looking at the class of ME models in more detail, our analysis indicates that the information carried by the strengths is not as `fundamental' as the one carried by the degrees: this is evident upon considering that 1) the UECM is always disfavoured with respect to the models just constraining the degrees, 2) the second best performing ME model is (always) the two step one, defined by a first purely topological step, accounting for the degrees, followed by an econometric-wise estimation of the weights. On top of that, we explicitly notice that structural topological information (e.g. the one provided by the link density or the degree sequence) usually correlates with node-specific economic covariates; hence, excluding such information from the model may lead to the so-called `omitted variable' bias. As proved by AIC and BIC, ME models represent the best compromise between goodness-of-fit and parsimony: in fact, they allow for structural information to be included, keeping the aforementioned type of bias low while leading to a better description of economic systems than that provided by traditional econometric models.
Our findings may indicate a route towards reconciling econometric and maximum-entropy network approaches, suggesting how to build a model that combines the pros of both: the importance of purely structural information (highlighted by physics-inspired models) can be accounted for by a model with a first step that is purely topological in nature (notice, in fact, that the TSF model is disfavoured with respect to the TS one) and a second step that takes care of estimating the weighted structure. Such an estimation can rest upon econometric considerations, driving the re-parametrization of otherwise purely structural models.
D.G. acknowledges support from the Dutch Econophysics Foundation (Stichting Econophysics, Leiden, the Netherlands). T.S. and D.G. also acknowledge support from the European Union Horizon 2020 Program under the scheme `INFRAIA-01-2018-2019 - Integrating Activities for Advanced Communities', Grant Agreement n. 871042, `SoBigData++: European Integrated Infrastructure for Social Mining and Big Data Analytics'.
\setcounter{equation}{0}