Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
52,956 characters · 8 sections · 44 citation commands
Equivalence of inequality indices: Three dimensions of impact revisited
[orcid=0000-0003-4508-3172] \ead{[email removed]} \credit{Conceptualisation of this study, Methodology, Writing}
[orcid=0000-0003-0637-6028] \ead{[email removed]} \ead[url]{https://www.gagolewski.com} \cormark[1] \cortext[cor1]{Corresponding author} \credit{Conceptualisation of this study, Methodology, Data Curation, Investigation, Software, Writing}
[orcid=0000-0002-9391-6477] \ead{[email removed]} \ead[URL]{http://if.pw.edu.pl/ siudem} \credit{Conceptualisation of this study, Methodology, Writing}
[orcid=0000-0002-2869-7300] \ead{[email removed]} \credit{Data Curation, Investigation, Visualisation, Software, Writing}
\shortauthors{Bertoli-Barsotti, Gagolewski, Siudem, Żogała-Siudem}
\affil[a]{University of Bergamo, Department of Economics, Italy}
\affil[b]{Deakin University, Data to Intelligence Research Centre, School of IT, Geelong, VIC 3220, Australia}
\affil[c]{Warsaw University of Technology, Faculty of Physics, ul. Koszykowa 75, 00-662 Warsaw, Poland}
\affil[d]{Systems Research Institute, Polish Academy of Sciences, ul. Newelska 6, 01-447 Warsaw, Poland}
\let\WriteBookmarks\relax
\shorttitle{Equivalence of inequality indices: Three dimensions of impact revisited}
\newwrite\tempfile \immediate\openout\tempfile=\jobname.abs \immediate\write\tempfile{ Inequality is an inherent part of our lives: we see it in the distribution of incomes, talents, resources, and citations, amongst many others. Its intensity varies across different environments: from relatively evenly distributed ones, to where a small group of stakeholders controls the majority of the available resources. We would like to understand why inequality naturally arises as a consequence of the natural evolution of any system. Studying simple mathematical models governed by intuitive assumptions can bring many insights into this problem. In particular, we recently observed (Siudem et al., PNAS 117:13896--13900, 2020) that impact distribution might be modelled accurately by a time-dependent agent-based model involving a mixture of the rich-get-richer and sheer chance components. Here we point out its relationship to an iterative process that generates rank distributions of any length and a predefined level of inequality, as measured by the Gini index.
Many indices quantifying the degree of inequality have been proposed. Which of them is the most informative? We show that, under our model, indices such as the Bonferroni, De Vergottini, and Hoover ones are equivalent. Given one of them, we can recreate the value of any other measure using the derived functional relationships. Also, thanks to the obtained formulae, we can understand how they depend on the sample size. An empirical analysis of a large sample of citation records in economics (RePEc) as well as countrywise family income data, confirms our theoretical observations. Therefore, we can safely and effectively remain faithful to the simplest measure: the Gini index. } \immediate\closeout\tempfile
Given a series of measurements, indicators, scores, counts, or any other numeric values, it is natural to order them from the highest to the lowest. This way, we get better insight into the aspects of reality they are trying to capture. In particular, various rankings (e.g., of universities, movies, or restaurants; see iniguez2022dynamics) aim to make our lives easier by claiming they can separate seeds from the chaff. Analysis or prediction of the size of the top or otherwise extreme values (e.g., voitalov2019scale,pickands1975statistical,marshall2007life) is crucial in risk analysis or disaster prevention. The search for patterns and universalities in sorted data holme2022universality,Newman2005 remains a fundamental, multidisciplinary research topic. This includes the study of ranking dynamics iniguez2022dynamics and the distribution of their static snapshots, from the most straightforward Zipf power-law models Newman2005 to more complex ones Petersen2011,PNAS2020,pricepareto,singh2022quantifying.
Different systems or environments naturally have different sizes and levels of inequality cowell2000measurement,silber2012handbook, i.e., what percentage of the top scorers is in possession of the majority of the resources. As the degree of evenness vs monopoly can be measured by different indicators (e.g., the Gini, Bonferroni, or Hover index), a question arises which one is the most informative.
Furthermore, we would like to understand why inequality naturally arises as a consequence of the natural evolution of any system. Studying simple mathematical models governed by intuitive assumptions can bring many insights into this problem.
In this work, we revisit the recently-proposed 3DI model (three dimensions of impact; PNAS2020,pricepareto), which can be considered a rank-size approach to the problem of describing the mechanisms governing the growth of bibliographic and other networks studied originally by Price deSollaPrice1965.
Namely, consider a process where, in every time step, a system (e.g., a citation network, a cluster of internet portals) grows by one entity (e.g., a new paper, a website). In each iteration, we distribute $m$ impact/wealth units (e.g., citations, links) amongst the already-existing entities:
with $\rho$ representing the extent to which the rich-get-richer rule dominates over pure luck. Inspired by the observation in bertoli2022equivalent, in Section (ref), we will show that the $\rho$ parameter naturally corresponds to the value of the Gini index of the resulting ordered sample of impact measures. We will thus indicate the relationship between the degree of randomness and inequality. Furthermore, in Section (ref), we will re-express some other popular inequality indices in terms of monotone, 1-to-1 functions of the Gini index, hence showing their equivalence in our model. In Section (ref), an analysis of a large database of citations to research papers in economics (RePEc) will confirm our theoretical derivations. We will reach similar conclusions by analysing countrywise income data in Section (ref).
We utilise the following notation. The gamma function is given by $\Gamma(z)=\int_0^\infty t^{z-1} e^{-t}\,dt$, $\Gamma(z+1)=z\Gamma(z)$, the polygamma functions are defined by $\psi^{(m)}(z)=\frac{d^{m+1}}{dz^{m+1}}\log \left(\Gamma(z)\right)$, in particular, the digamma function is $\psi(z)=\psi^{(0)}(z)=\Gamma'(z)/\Gamma(z)$, the Lambert $\mathcal{W}$ function (we always use the principal branch) is the solution of the equation $\mathcal{W}(z) \exp\left[\mathcal{W}(z)\right]=z$, the harmonic number $H_k=\psi(k+1)-\gamma$, where $\gamma \approx 0.577$ is the Euler constant.
For any fixed parameter $G\in[0,1)$, consider an iterative process discussed in OurGiniLorenz whose update formula features a simple affine function of the consecutive elements:
for $N\geq 2$, under the assumption that $p_N^{(N-1,G)} = 0$ and $p_1^{(1,G)}=1$ and:
Then, for any $N\ge 2$, $\boldsymbol{p}^{(N,G)}=\left(p_1^{(N,G)},\dots,p_N^{(N,G)}\right)$ is an ordered probability $N$-vector, i.e., $\sum_{k=1}^N p_k^{(N,G)}=1$ and $1\ge p_1^{(N,G)}\ge p_2^{(N,G)}\ge\dots\ge p_N^{(N,G)}\ge 0$.
\paragraph{Exact formula.} There exists an explicit formula for the individual components of the probability vectors. Namely, we can show that:
where $\Gamma$ is the gamma function and $H_N=\sum_{i=1}^N \frac{1}{i}$ is the $N$-th harmonic number.
\paragraph{Gini index is exactly $G$.} Figure (ref) depicts a few example probability vectors $\boldsymbol{p}^{(7,G)}$ for different $G$s. We see that the $G$ parameter controls the level of inequality of the data distribution: $G=0$ yields all elements equal to $1/N$, and, as $G\to 1$, all probability mass is transferred to the first element.
It turns out that our process generates ordered probability vectors whose Gini's index, $\mathcal{G}\left(\boldsymbol{p}^{(N,G)}\right)$ (Gini1912:index; see also, e.g., \citealp*{ginimeth}), is equal to exactly $G$, i.e., for $N\ge 2$ and $G\neq \frac{1}{2}$ we have:
and for $G=\frac{1}{2}$ it holds:
Thus, quite remarkably, $N$ (size) and $G$ (inequality) are two independent parameters in our model.
Therefore, from now on, we refer to the above as the Gini-stable process; compare OurGiniLorenz.
\paragraph{It is the only affine process of this kind.} Our process transforms an ordered probability vector to a new, longer, ordered probability vector of the same Gini index. What is more, the aforementioned coefficients $a_N$ and $b_N$ make up the only affine transformation that achieves this property.
\paragraph{Relation to the 3DI model.} To strengthen the underlying fundaments further, let us return to the 3DI (three dimensions of impact) model PNAS2020 mentioned in the introduction. Let $X_k(t)$ denote the impact of the $k$-th richest entity at time step $t$ (e.g., the number of citations to the $k$-th most cited paper). We assume ${X}_k(k-1)=0$ for every $k$, i.e., the $k$-th object enters the system with no impact units. The update formula in our model is a mixture\footnote{In PNAS2020 and pricepareto, we only studied the case of $\rho\in(0,1)$. However, let us note that $\rho<0$ is not only possible but also has a nice interpretation gagolewski2022ockham,bertoli2022equivalent. In such a scenario, we initially distribute more than the assumed $m$ citations at random, but then we take away from those who are already rich (rich get less).} of the accidental and rich-get-richer components governed by the $\rho$ parameter:
where $a=(1-\rho)m$ and $p=\rho m$.
First, if we assume $\rho=0$, then we are only left with the accidental component, and our model reduces to the harmonic one (compare cena2022validating):
Note that “purely accidental” does not mean that every entity ends up with the same amount of wealth, as older agents have had more opportunities to become impactful (“the old get richer”).
For $\rho<1$ and $\rho\neq 0$, the solution is:
We, therefore, note that:
with: $$\rho(G) = 2-\frac{1}{G}=\frac{2G-1}{G},$$ or, equivalently: $$G(\rho)=\frac{1}{2-\rho},$$ and $p_k^{(t,G)}$ given by Eq. ((ref)).
We have thus established a beautiful connection between the 3DI model PNAS2020 and the Gini-stable process OurGiniLorenz, and hence the degree of randomness in the impact distribution and the Gini index. Let us also note that in OurGiniLorenz, where we have studied the asymptotic behaviour of the Lorenz curves generated by this model, we have also shown its relationship to the Type II Pareto, exponential, and scaled beta distributions (compare arnold2015:pareto,pickands1975statistical,marshall2007life), extending the results from pricepareto.
\paragraph{Majorisation order.}
Given two $N$-vectors $\boldsymbol{p}$ and $\boldsymbol{p}'$ with $\sum_{i=1}^N p_i=\sum_{i=1}^N p_i'$, we define the majorisation order $\preceq$ marshall2011majorisation in such a way that $\boldsymbol{p} \preceq \boldsymbol{p}'$ if and only if for all $k=1,\dots,N$ it holds $F_k \le F_k'$, where $F_k = \sum_{i=1}^k p_{[i]}$ and $F_k' = \sum_{i=1}^k p_{[i]}'$ are the sums of the $k$ greatest elements in $\boldsymbol{p}$ and $\boldsymbol{p}'$, respectively. In other words, it is the extension of the standard componentwise relation, $\le$, applied over the consecutive cumulative sums of ordered items.
Let us note that such cumulative sums in our model can be expressed using the following simple formula:
\paragraph{Monotonicity w.r.t. the majorisation order.} We can show that for any $N$, $\boldsymbol{p}^{(N,G)}$ as a function of $G$ is monotone with respect to the majorisation order.
See OurGiniLorenz for the proof, where we discuss the Gini-stable process in the context of the more general Lorenz ordering.
\paragraph{Inequality indices.} Let us now consider any function $\phi$ that maps the ordered probability $N$-vectors to the set of real numbers. We say that $\phi$ is Schur-convex, if for any $\boldsymbol{p} \preceq \boldsymbol{p}'$ it holds ${\phi}(\boldsymbol{p})\le{\phi}(\boldsymbol{p}')$. This is equivalent to $\phi$ being increasing in each $F_1,\dots,F_k$ when re-expressed in terms of cumulative sums of ordered elements; see BeliakovETAL2016:penaltyinequality.
Any Schur-convex function normalised such that $\phi(1,0,\dots,0)=1$ and $\phi(1/N,\dots,1/N)=0$ is called an inequality index; see shorrocks1987transfer and marshall2011majorisation,lambert2001distribution,chantreuil2011inequality,zheng2007unit,bosmans2016consistent.
The Gini index is an example function fulfilling these properties. Below we recall some other noteworthy inequality indices (compare, e.g., ciommi,mehran,imedioolmedo,parasite,BeliakovETAL2016:penaltyinequality). Then, we derive the formulae for different inequality indices as functions of $N$ and $G$ in our model. This will enable us to express every index as a function of any another one.
\paragraph{Bonferroni's index.}
For any probability $N$-vector, the Bonferroni index Bonferroni1930:index is given by:
Substituting $F_k=F_k^{(N,G)}$ from Eq. ((ref)) for $\frac{1}{2}\neq G \in (0,1)$, we get:
For $G=\frac{1}{2}$, we get:
\paragraph{De Vergottini's index.}
The De Vergottini index verg1,verg2 is defined as:
Computing the value of the following sum in our model (taking $p_k=p_k^{(N,G)}$ from Eq. ((ref)) with $\frac{1}{2}\neq G \in (0,1)$):
leads to:
Furthermore, for $G=\frac{1}{2}$, we get:
\paragraph{Hoover's index.}
The Hoover index (hoover; also known as the Robin Hood index) is defined by:
It can be thought of as the normalised Manhattan distance to the perfectly equal vector.
We can simplify the above by determining:
and then writing:
Moreover, for $\frac{1}{2}\neq G\in(0,1)$, the above results in:
\paragraph{$\mathcal{P}_q$ indices.}
The $\mathcal{P}_q$ index for $q\in(0,1)$ is defined as:
This index is a normalised version of the percentage of accumulated probability mass in $q 100\%$ of the top elements (compare the famous Pareto 80/20 rule).
In our model, for $\frac{1}{2} \neq G \in(0,1) $, the $\mathcal{P}_q$ index is equal to
Furthermore, if $G=\frac{1}{2}$, then it holds:
\paragraph{Indices as functions of one another.}
To sum up, we have shown that in our model, the formulae for the Bonferroni, De Vergottini, Hoover, and $\mathcal{P}_q$ indices depend only on $N$ and $G$:
respectively (for readability, we only included the case of $G\neq 1/2$). Figure (ref) depicts them for different $N$s.
We note that they are all strictly increasing, continuous functions of $G$. Thus, based on the derived formulas, all the indices can be expressed as one-to-one functions of one another. For instance, given some value of the Bonferroni index $B$, we can obtain the underlying $G^* = \tilde{B}^{-1}(N, B)$ and then compute, say, $\tilde{V}(N, G^*)$. Even though the analytic formulae for the inverses do not exist, this still can easily be solved numerically.
In this sense, we can say that -- in our model -- all the aforementioned inequality indices are equivalent.
Similar derivations can be performed for many other inequality indices, although they will not necessarily enjoy analytic solutions.
Theoretical models are merely approximations of the real-world phenomena under scrutiny. Thus, empirical data might deviate from the assumed idealisations. Also, looking beyond our simple iterative process, we know that uncountably many probability vectors yield a specific Gini, Bonferroni, or any other index. After all, these measures were introduced to respond to the different needs of the practitioners; see, e.g., imedioolmedo,parasite. Some of them are, for example, more responsive to increasing (via the principle of progressive transfers) the amount of probability mass in the tail of the distribution than others. In particular, as reported in ciommi, the Gini, Bonferroni, and De Vergottini indices belong to the class of linear measures introduced in mehran. They note that for the Bonferroni and De Vergottini indices, the effect of a transfer also depends on the position of individuals, making the Bonferroni index more sensitive to transfers that occur at the lower end of the income distribution and the De Vergottini index more sensitive to variations among the richest.
Therefore, we should be interested in verifying how well our formulae describe the underlying structure of real datasets.
Let us thus consider citation data from the RePEc database (Research Papers in Economics; see https://citec.repec.org/), which features $66{,}347$ authors and $1{,}843{,}967$ papers. In the data cleansing step, we have omitted the authors who published less than $5$ cited papers and whose $h$-index was less than $3$.
This resulted in $n=36{,}425$ citation records of the form $\left(x_1^{(i)},\dots,x^{(i)}_{N^{(i)}}\right)$, where $N^{(i)}$ gives the total number of items published by the $i$-th author and $x_k^{(i)}$ gives the number of citations to their $k$-th most cited works.
Figure (ref) shows three example (quite representative of the whole database) vectors from the RePEc database: observed (points) and predicted (lines) values for $p_k$ (left) and their non-normalised versions, $x_k$ (right). We see a good fit over most parts of the data domain. This comes at no surprise, as the proposed model is equivalent to the 3DI model PNAS2020 and we have already seen its usefulness in the case of modelling citations to computer science papers.
For each author, we compute their actual (observed) Gini index $\hat{G}^{(i)} = \mathcal{G}\left(p_1^{(i)},\dots,p^{(i)}_{N^{(i)}}\right)$ with $p_k^{(i)} = X_k^{(i)}/\sum_{j=1}^{N^{(i)}} X_j^{(i)}$.
Then, based on the derived formulae, we compute the predicted $\tilde{B}(N, \hat{G}^{(i)})$, $\tilde{V}(N, \hat{G}^{(i)})$, \dots, and observed $\hat{B}^{(i)} = \mathcal{B}\left(p_1^{(i)},\dots,p^{(i)}_{N^{(i)}}\right)$, $\hat{V}^{(i)} = \mathcal{V}\left(p_1^{(i)},\dots,p^{(i)}_{N^{(i)}}\right)$, \dots, Bonferroni, De Vergottini, and other indices.
Recall that Figure (ref) describes the theoretical relationships between the Gini index and other metrics. If, overall, our model describes the real data well, we should expect to see these dependencies in the case of the RePEc vectors too.
Figure (ref) presents a scatter plot of the values of different indices as functions of $\hat{G}$ for all vectors with $N>200$ (the De Vergottini index was not included as it approaches $0$ for large $N$s; see Remark (ref)). Ideally, they should lie close to the theoretical curves (depicted as well). And this is approximately the case.
Furthermore, Figure (ref) presents similar results, but for vectors of lengths $N=6$ (green), $N=10$ (pink), and $N=50\pm 1$ (blue; the $\pm 1$ part is to increase the number of data points). Additionally, we coloured the areas representing all of the possible values which could be obtained for vectors generated using our model. In other words, for any vector, we expect the pairs $(\hat{G}, \hat{B})$, $(\hat{G}, \hat{V})$, etc. to lie in the grey zone. This is true for the vast majority of the real data points.
The prediction errors for different vector lengths are summarised in Table (ref). The error rates are usually less than 2--3%.
Overall, we observe a very good fit, confirming our theoretical derivations. The choice of the inequality measure is secondary, as the indices can be considered functions of one another. Therefore, it is best to be faithful to the simplest indicator: the Gini index.
Additionally, let us study the Luxembourg Income Study dataset giving the family incomes in nine countries. The original data, in the form of the Lorenz (e.g., Bertoli2019) curves across the deciles ($N=10$), are given in Bishop.
Table (ref) gives the prediction error for the indices studied. Overall, the errors are very small, except for the case of the $\mathcal{P}_q$ index with higher $q$ in Sweden, Germany, and the United Kingdom, where there is some discrepancy between the predicted and the observed Lorenz curve in their tails.
The discussed process yields $\boldsymbol{p}^{(N,G)}$ that are totally ordered by the majorisation relation $\preceq$. Following shorrocks1987transfer (but see also marshall2011majorisation,lambert2001distribution,chantreuil2011inequality,zheng2007unit,bosmans2016consistent), a function $I$ is an index of inequality if and only if it is symmetric and strictly Schur convex.
Hence, for all ordered probability $N$-vectors $\boldsymbol{p}\neq\boldsymbol{q}$, if $\boldsymbol{p}\preceq \boldsymbol{q}$, then $I(\boldsymbol{p})<I(\boldsymbol{q})$, where the direction of the inequality is uniquely determined by the Pigou–Dalton condition (e.g., Patty2019). In particular, for every vector $\boldsymbol{p}^{(N,G)}$, the parameter $G$ can be interpreted as the (normalised) Gini index $\mathcal{G}(\boldsymbol{p}^{(N,G)})=G$. This implies that $\mathcal{G}$ is an order preserving function marshall2011majorisation on $\{\boldsymbol{p}^{(N,G)}\}$ with respect to $G$.
In this study, we also extended this characterisation result to some other notable indices of inequality, that is, the Bonferroni, De Vergottini, Hoover and $\mathcal{P}_q$ indices, deriving their explicit expressions as one-to-one functions of $\mathcal{G}$. For data following closely our model, the indices can be considered equivalent. An analysis of two empirical datasets (citation vectors in economics and countrywise family incomes) confirms our results.
We are indebted to Jose Manuel Barrueco for providing us with a large snapshot of RePEc (Research Papers in Economics) data. All data are freely available at http://citec.repec.org/api.html.
This research was supported by the Australian Research Council Discovery Project ARC DP210100227 (MG).
The authors certify that they have no affiliations with or involvement in any organisation or entity with any financial interest or non-financial interest in the subject matter or materials discussed in this manuscript.
\printcredits