The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
52,914 characters
Equivalence of inequality indices: Three~dimensions~of~impact~revisited
\author[a]{Lucio Bertoli-Barsotti}[orcid=0000-0003-4508-3172]
\ead{[email removed]}
\credit{Conceptualisation of this study, Methodology,
Writing}
\author[b]{Marek Gagolewski}[orcid=0000-0003-0637-6028]
\ead{[email removed]}
\ead[url]{https://www.gagolewski.com}
\cormark[1]
\cortext[cor1]{Corresponding author}
\credit{Conceptualisation of this study, Methodology,
Data Curation, Investigation, Software,
Writing}
\author[c]{Grzegorz Siudem}[orcid=0000-0002-9391-6477]
\ead{[email removed]}
\ead[URL]{http://if.pw.edu.pl/~siudem}
\credit{Conceptualisation of this study, Methodology,
Writing}
\author[d]{Barbara Żogała-Siudem}[orcid=0000-0002-2869-7300]
\ead{[email removed]}
\credit{Data Curation, Investigation, Visualisation, Software, Writing}
\shortauthors{Bertoli-Barsotti, Gagolewski, Siudem, Żogała-Siudem}
\affil[a]{University of Bergamo, Department of Economics, Italy}
\affil[b]{Deakin University, Data to Intelligence Research Centre, School of IT, Geelong, VIC 3220, Australia}
\affil[c]{Warsaw~University of Technology, Faculty of Physics,
ul. Koszykowa 75, 00-662 Warsaw, Poland}
\affil[d]{Systems Research Institute, Polish Academy of Sciences,
ul. Newelska 6, 01-447 Warsaw, Poland}
\let\WriteBookmarks\relax
\shorttitle{Equivalence of inequality indices: Three~dimensions~of~impact~revisited}
\title[mode = title]{Equivalence of inequality indices: Three~dimensions~of~impact~revisited}
\newwrite\tempfile
\immediate\openout\tempfile=\jobname.abs
\immediate\write\tempfile{
Inequality is an inherent part of our lives: we see it in the distribution of incomes, talents, resources, and citations, amongst many others. Its intensity varies across different environments: from relatively evenly distributed ones, to where a small group of stakeholders controls the majority of the available resources. We would like to understand why inequality naturally arises as a consequence of the natural evolution of any system. Studying simple mathematical models governed by intuitive assumptions can bring many insights into this problem. In particular, we recently observed (Siudem et al., PNAS 117:13896--13900, 2020) that impact distribution might be modelled accurately by a time-dependent agent-based model involving a mixture of the rich-get-richer and sheer chance components. Here we point out its relationship to an iterative process that generates rank distributions of any length and a predefined level of inequality, as measured by the Gini index.
Many indices quantifying the degree of inequality have been proposed. Which of them is the most informative? We show that, under our model, indices such as the Bonferroni, De Vergottini, and Hoover ones are equivalent. Given one of them, we can recreate the value of any other measure using the derived functional relationships. Also, thanks to the obtained formulae, we can understand how they depend on the sample size. An empirical analysis of a large sample of citation records in economics (RePEc) as well as countrywise family income data, confirms our theoretical observations. Therefore, we can safely and effectively remain faithful to the simplest measure: the Gini index.
}
\immediate\closeout\tempfile
\begin{keywords}
Gini index $|$
Bonferroni index $|$
power law $|$
inequality $|$
scientometrics $|$
economics
\end{keywords}
\maketitle
\section{Introduction}
Given a series of measurements, indicators, scores, counts, or any other numeric values, it is natural to order them from the highest to the lowest. This way, we get better insight into the aspects of reality they are trying to capture. In particular, various rankings (e.g., of universities, movies, or restaurants; see \citealp{iniguez2022dynamics}) aim to make our lives easier by claiming they can separate seeds from the chaff. Analysis or prediction of the size of the top or otherwise extreme values (e.g., \citealp{voitalov2019scale,pickands1975statistical,marshall2007life}) is crucial in risk analysis or disaster prevention. The search for patterns and universalities in sorted data \citep{holme2022universality,Newman2005} remains a fundamental, multidisciplinary research topic. This includes the study of ranking dynamics \citep{iniguez2022dynamics} and the distribution of their static snapshots, from the most straightforward Zipf power-law models \citep{Newman2005} to more complex ones \citep{Petersen2011,PNAS2020,pricepareto,singh2022quantifying}.
Different systems or environments naturally have different sizes and levels of inequality \citep{cowell2000measurement,silber2012handbook}, i.e., what percentage of the top scorers is in possession of the majority of the resources. As the degree of evenness vs monopoly can be measured by different indicators (e.g., the Gini, Bonferroni, or Hover index), a question arises which one is the most informative.
Furthermore, we would like to understand why inequality naturally arises as a consequence of the natural evolution of any system. Studying simple mathematical models governed by intuitive assumptions can bring many insights into this problem.
In this work, we revisit the recently-proposed 3DI model (three dimensions of impact; \citealp{PNAS2020,pricepareto}), which can be considered a rank-size approach to the problem of describing the mechanisms governing the growth of bibliographic and other networks studied originally by Price \citeyearpar{deSollaPrice1965}.
Namely, consider a process where, in every time step, a system (e.g., a citation network, a cluster of internet portals) grows by one entity (e.g., a new paper, a website). In each iteration, we distribute $m$ impact/wealth units (e.g., citations, links) amongst the already-existing entities:
\begin{itemize}[nosep]
\item
$a=(1-\rho)m$ units totally at random,
\item
$p=\rho m$ units according to the preferential attachment rule,
\end{itemize}
\noindent
with $\rho$ representing
the extent to which the rich-get-richer rule dominates over pure luck. Inspired by the observation in \citep{bertoli2022equivalent}, in Section~\ref{sec:giniprocess}, we will show that the $\rho$ parameter naturally corresponds to the value of the Gini index of the resulting ordered sample of impact measures. We will thus indicate the relationship between the degree of randomness and inequality. Furthermore, in Section~\ref{sec:inequality}, we will re-express some other popular inequality indices in terms of monotone, 1-to-1 functions of the Gini index, hence showing their equivalence in our model. In Section~\ref{sec:repec}, an analysis of a large database of citations to research papers in economics (RePEc) will confirm our theoretical derivations.
We will reach similar conclusions by analysing countrywise income data in Section~\ref{sec:incomedata}.
\section{The Gini-stable process, rich-get-richer, and random distribution of wealth}\label{sec:giniprocess}
We utilise the following notation.
The gamma function is given by $\Gamma(z)=\int_0^\infty t^{z-1} e^{-t}\,dt$, $\Gamma(z+1)=z\Gamma(z)$,
the polygamma functions are defined by
$\psi^{(m)}(z)=\frac{d^{m+1}}{dz^{m+1}}\log \left(\Gamma(z)\right)$,
in particular, the digamma function is $\psi(z)=\psi^{(0)}(z)=\Gamma'(z)/\Gamma(z)$,
the Lambert $\mathcal{W}$ function (we always use the principal branch) is the solution of the equation $\mathcal{W}(z) \exp\left[\mathcal{W}(z)\right]=z$,
the harmonic number $H_k=\psi(k+1)-\gamma$, where $\gamma \approx 0.577$ is the Euler constant.
For any fixed parameter $G\in[0,1)$, consider an iterative process discussed
in \citep{OurGiniLorenz} whose update formula features a simple affine function
of the consecutive elements:
\begin{equation}\label{eq:process}
p_k^{(N,G)} = a_N(G) + b_N(G)\, p_k^{(N-1,G)}, \qquad k=1,\dots,N,
\end{equation}
\noindent
for $N\geq 2$, under the assumption that $p_N^{(N-1,G)} = 0$ and $p_1^{(1,G)}=1$
and:
\begin{eqnarray}
a_N(G) &=& \frac{1-G}{G(N-2)+1} \frac{1}{N}, \label{eq:at}\\
b_N(G) &=& \frac{G(N-1)}{G(N-2)+1}. \label{eq:bt}
\end{eqnarray}
\noindent
Then, for any $N\ge 2$,
$\boldsymbol{p}^{(N,G)}=\left(p_1^{(N,G)},\dots,p_N^{(N,G)}\right)$
is an \emph{ordered probability $N$-vector}, i.e.,
$\sum_{k=1}^N p_k^{(N,G)}=1$
and
$1\ge p_1^{(N,G)}\ge p_2^{(N,G)}\ge\dots\ge p_N^{(N,G)}\ge 0$.
\begin{example}
Here are two example outcomes for $G=0.25$ and $G=0.75$
up to $N=4$:
\begin{center}
\begin{tabular}{lll}
& $\boldsymbol{p}^{(N,0.25)}$ & $\boldsymbol{p}^{(N,0.75)}$ \\
$N=2$ & $( 0.625, 0.375 )$ & $( 0.875, 0.125 )$ \\
$N=3$ & $( 0.450, 0.350, 0.200 )$ & $( 0.798, 0.155, 0.048 )$ \\
$N=4$ & $( 0.350, 0.300, 0.225, 0.125 )$ & $( 0.743, 0.164, 0.068, 0.025 )$ \\
\end{tabular}
\end{center}
\end{example}
\paragraph{Exact formula.}
There exists an explicit formula for the individual components
of the probability vectors. Namely, we can show that:
\begin{equation}\label{eq:pk}
p_k^{(N,G)}=
\begin{cases}
\displaystyle
\frac{1}{N} \frac{1-G}{2G-1} \left(
\frac{\Gamma(N+1)}{\Gamma(N+1+1/G-2)} \frac{\Gamma(k+1/G-2)}{\Gamma(k)}
-1
\right),
&
\text{if } G\neq \frac{1}{2},
\\
\displaystyle
\frac{1}{N} \left(
H_N-H_{k-1}
\right)=\frac{1}{N}\sum_{i=k}^N\frac{1}{i},
&
\text{if } G = \frac{1}{2},
\\
\end{cases}
\end{equation}
where $\Gamma$ is the gamma function
and
$H_N=\sum_{i=1}^N \frac{1}{i}$ is the $N$-th harmonic number.
\paragraph{Gini index is exactly $G$.}
Figure~\ref{fig:model-vectors} depicts a few example probability vectors
$\boldsymbol{p}^{(7,G)}$ for different $G$s.
We see that the $G$ parameter controls the level of inequality
of the data distribution: $G=0$ yields all elements equal to $1/N$,
and, as $G\to 1$, all probability mass is transferred
to the first element.
\begin{figure}[h!]
\centering
\includegraphics[width=0.7\linewidth]{model-vectors.pdf}
\caption{\label{fig:model-vectors} Example ordered probability $N$-vectors
generated using the affine update formula given by Eq.~(\ref{eq:process}).
Parameter $G$ controls the level of inequality, as measured
by the Gini index.}
\end{figure}
It turns out that our process generates ordered probability
vectors whose Gini's index, $\mathcal{G}\left(\boldsymbol{p}^{(N,G)}\right)$ (\citealp{Gini1912:index}; see also, e.g., \citealp*{ginimeth}),
is equal to exactly $G$, i.e., for
$N\ge 2$ and $G\neq \frac{1}{2}$
we have:
\begin{eqnarray*}
\mathcal{G}\left(p_1^{(N,G)},\dots,p_N^{(N,G)}\right)
&=&
\frac{1}{N-1} \sum_{k=1}^N (N-2k+1) p_{k}^{(N,G)}
\\
&=&
\displaystyle
\frac{1-G}{(N-1)(2G-1)}
\frac{G (2 G-1) (N-1)}{1-G}
= G,
\end{eqnarray*}
\noindent
and for $G=\frac{1}{2}$ it holds:
\begin{eqnarray*}
\mathcal{G}\left(p_1^{(N,1/2)},\dots,p_N^{(N,1/2)}\right)
= \frac{(N-1)N}{2(N-1)N}=\frac{1}{2}.
\end{eqnarray*}
\noindent
Thus, quite remarkably, $N$ (size) and $G$ (inequality)
are two \textit{independent} parameters in our model.
Therefore, from now on, we refer to the above as the \emph{Gini-stable process};
compare \citep{OurGiniLorenz}.
\paragraph{It is the only affine process of this kind.}
Our process transforms an ordered probability vector
to a new, longer, ordered probability vector of the same Gini index.
What is more, the aforementioned coefficients $a_N$ and $b_N$ make up
the only affine transformation that achieves this property.
\begin{proposition}
For any $N\ge 3$ and $G\in[0,1)$, let
$\boldsymbol{q}^{(N-1)}$
be an ordered probability $(N-1)$-vector
with the Gini index of $\mathcal{G}\left(q_1^{(N-1)},\dots,q_{N-1}^{(N-1)}\right) = G$.
Then, generating the probability $N$-vector
$\boldsymbol{q}^{(N)}$ using the update formula
given by Eq.~(\ref{eq:process})
yields $\mathcal{G}\left(q_1^{(N)},\dots,q_N^{(N)}\right) = G$
if and only if $a_N$ and $b_N$ are given exactly by
Eq.~(\ref{eq:at}) and Eq.~(\ref{eq:bt}), respectively.
\end{proposition}
\begin{proof}
We have
$\mathcal{G}\left(q_1^{(N-1)},\dots,q_{N-1}^{(N-1)}\right)=
\frac{1}{N-2} \sum_{k=1}^{N-1} (N-2k) q_{k}^{(N-1)}
=G
$.
Moreover,
\begin{eqnarray*}
\mathcal{G}\left(q_1^{(N)},\dots,q_{N}^{(N)}\right)
&=&
\frac{1}{N-1} \sum_{k=1}^N (N-2k+1) q_{k}^{(N)}
\\
&=&
\frac{1}{N-1} \sum_{k=1}^N (N-2k+1) (a_N+b_N q_{k}^{(N-1)})
\\
&=&
\frac{a_N}{N-1} \sum_{k=1}^N (N-2k+1)
+
\frac{b_N}{N-1} \sum_{k=1}^{N-1} (N-2k+1) q_{k}^{(N-1)}
\\
&=&
0
+
\frac{b_N}{N-1} \sum_{k=1}^{N-1} (N-2k) q_{k}^{(N-1)} + \frac{b_N}{N-1}
\\
&=&
b_N \left( \frac{N-2}{N-1} G + \frac{1}{N-1} \right)
\end{eqnarray*}
is equal to $G$
if and only if
$b_N = \frac{G (N-1)}{G (N-2)+ 1 }$,
which is exactly Eq.~(\ref{eq:bt}).
We also need to have $\sum_{k=1}^N p_{k}^{(N)} = 1$.
But after some simple transformations,
we can show that this holds if and only if $a_N = (1 - b_N)/N$,
which corresponds to Eq.~(\ref{eq:at}), QED.
\end{proof}
\paragraph{Relation to the 3DI model.}
To strengthen the underlying fundaments further,
let us return to the 3DI (three dimensions of impact) model \citep{PNAS2020}
mentioned in the introduction.
Let $X_k(t)$ denote the impact of the $k$-th richest entity at time step $t$
(e.g., the number of citations to the $k$-th most cited paper).
We assume ${X}_k(k-1)=0$ for every $k$, i.e., the $k$-th object enters the
system with no impact units.
The update formula in our model is a mixture\footnote{In \citep{PNAS2020} and \citep{pricepareto}, we only studied the case of $\rho\in(0,1)$.
However, let us note that $\rho<0$ is not only possible but also has a nice interpretation
\citep{gagolewski2022ockham,bertoli2022equivalent}. In such a scenario,
we initially distribute more than the assumed $m$ citations at random,
but then we take away from those who are already rich (rich get less).} of the accidental and rich-get-richer components governed by the $\rho$ parameter:
\begin{equation}
{X}_k{(t)} =
\underbrace{{X}_k{(t-1)}}_{\mathrm{previous\; value}}+
\underbrace{\frac{a}{t}}_{\mathrm{accidental\;income}}+
\underbrace{
p\,\frac{{X}_k{(t-1)}+\frac{a}{t}}{ (t-1)m + a}}_{\mathrm{preferential\;gain\;or\;loss}
},\label{eq:master}
\end{equation}
\noindent
where $a=(1-\rho)m$ and $p=\rho m$.
First, if we assume $\rho=0$, then we are only left with the accidental component,
and our model reduces to the harmonic one (compare \citealp{cena2022validating}):
\begin{equation*}
{X}_k(t) = m \sum\limits_{i=k}^t\frac{1}{i} = m \left(H_t-H_{k-1}\right).
\end{equation*}
\noindent
Note that ``purely accidental'' does not mean that every entity ends up with the same amount of wealth, as older agents have had more opportunities to become impactful (``the old get richer'').
For $\rho<1$ and $\rho\neq 0$, the solution is:
\begin{eqnarray}
X_k(t)
&=&
m\frac{1-\rho}{\rho}
\displaystyle
\left(
\frac{\Gamma(t+1)}{\Gamma(t+1-\rho)}\frac{\Gamma(k-\rho)}{\Gamma(k)} - 1
\right)
\nonumber
\\
&=&
m\frac{1-\frac{1}{2-\rho}}{\frac{2}{2-\rho}-1}
\displaystyle
\left(
\frac{\Gamma(t+1)}{\Gamma\left[t-1+(2-\rho)\right]}\frac{\Gamma\left[k-2+(2-\rho)\right]}{\Gamma(k)} - 1
\right).
\label{eq:ikska}
\end{eqnarray}
\noindent
We, therefore, note that:
\begin{equation}
X_k(t) = m t p_{k}^{(t,G)}
\end{equation}
with:
$$\rho(G) = 2-\frac{1}{G}=\frac{2G-1}{G},$$
or, equivalently:
$$G(\rho)=\frac{1}{2-\rho},$$ and $p_k^{(t,G)}$ given by Eq.~(\ref{eq:process}).
\smallskip
We have thus established a beautiful connection between the 3DI model
\citep{PNAS2020} and the
Gini-stable process \citep{OurGiniLorenz},
and hence the degree of randomness in the impact distribution and the Gini index.
Let us also note that in
\citep{OurGiniLorenz}, where we have studied the asymptotic
behaviour of the Lorenz curves generated by this model,
we have also shown its relationship to the Type~II Pareto,
exponential, and scaled beta distributions (compare \citealp{arnold2015:pareto,pickands1975statistical,marshall2007life}),
extending the results from \citep{pricepareto}.
\section{Measuring Inequality}\label{sec:inequality}
\paragraph{Majorisation order.}
Given two $N$-vectors $\boldsymbol{p}$ and $\boldsymbol{p}'$
with $\sum_{i=1}^N p_i=\sum_{i=1}^N p_i'$,
we define the \emph{majorisation order} $\preceq$
\citep{marshall2011majorisation}
in such a way that $\boldsymbol{p} \preceq \boldsymbol{p}'$ if and only if
for all $k=1,\dots,N$ it holds $F_k \le F_k'$,
where $F_k = \sum_{i=1}^k p_{[i]}$ and $F_k' = \sum_{i=1}^k p_{[i]}'$
are the sums of the $k$ greatest elements in
$\boldsymbol{p}$ and $\boldsymbol{p}'$, respectively.
In other words, it is the extension of the standard componentwise relation, $\le$,
applied over the consecutive cumulative sums of ordered items.
Let us note that such cumulative sums in our model
can be expressed using the following simple formula:
\begin{eqnarray}\label{eq:cdf}
&&F_k^{(N,G)}=\sum_{i=1}^k p_i^{(N,G)}
=
\begin{cases}
\displaystyle
\frac{1-G}{2G-1} \left(
\frac{G}{1-G}\frac{\Gamma(N)}{\Gamma(N-1+1/G)} \frac{\Gamma(k-1+1/G)}{\Gamma(k)}
-\frac{k}{N}
\right),
&
\text{if }
G\neq\frac{1}{2},\\
\displaystyle\frac{k}{N} \left(
1+H_N-H_k
\right),
&
\text{if }
G=\frac{1}{2}.
\\
\end{cases}
\end{eqnarray}
\paragraph{Monotonicity w.r.t.~the majorisation order.}
We can show that for any $N$, $\boldsymbol{p}^{(N,G)}$ as a function of $G$
is monotone with respect to the majorisation order.
\begin{theorem}\label{Thm:majorisation}
For any $G \le G'$ and $N$, it holds $\boldsymbol{p}^{(N,G)} \preceq \boldsymbol{p}^{(N,G')}$.
\end{theorem}
\noindent
See \citep{OurGiniLorenz} for the proof, where we discuss the Gini-stable
process in the context of the more general Lorenz ordering.
\paragraph{Inequality indices.}
Let us now consider any function $\phi$ that maps the ordered probability
$N$-vectors to the set of real numbers.
We say that $\phi$ is \textit{Schur-convex}, if for any
$\boldsymbol{p} \preceq \boldsymbol{p}'$ it holds
${\phi}(\boldsymbol{p})\le{\phi}(\boldsymbol{p}')$.
This is equivalent to $\phi$ being increasing in each $F_1,\dots,F_k$
when re-expressed in terms of cumulative sums of ordered elements;
see \citep{BeliakovETAL2016:penaltyinequality}.
Any Schur-convex function normalised such that $\phi(1,0,\dots,0)=1$
and $\phi(1/N,\dots,1/N)=0$ is called an \textit{inequality index};
see \citep{shorrocks1987transfer} and \citep{marshall2011majorisation,lambert2001distribution,chantreuil2011inequality,zheng2007unit,bosmans2016consistent}.
The Gini index is an example function fulfilling these properties.
Below we recall some other noteworthy inequality indices (compare, e.g., \citealp{ciommi,mehran,imedioolmedo,parasite,BeliakovETAL2016:penaltyinequality}).
Then, we derive the formulae for different inequality indices as functions of $N$ and $G$ in our model. This will enable us to express every index as a function of any another one.
\paragraph{Bonferroni's index.}
\label{subsec:bonferroni-index}
For any probability $N$-vector, the Bonferroni index \citep{Bonferroni1930:index} is given by:
\begin{align}\nonumber
\mathcal{B}(p_1,\dots,p_N) =&\nonumber
\frac{
N \sum_{i=1}^{N} \left( 1-\sum_{j=1}^i \frac{1}{N-j+1} \right) p_{i}
}{
N-1
}
=
\frac{
N \sum_{i=1}^{N-1} \frac{1}{N-i} \sum_{j=1}^i p_j
}{
N-1
}
- N
+ \frac{N}{N-1} \sum_{i=1}^N \left(1-\frac{1}{N-i+1}\right)
\\
=&\frac{
N \sum_{i=1}^{N} \frac{1}{i} \sum_{j=1}^{N-i} p_j
}{
N-1
}
+ \frac{N\left(1-\sum_{j=1}^N\frac{1}{j}\right)}{N-1}
=
\frac{
N}{
N-1}\left(
\sum_{k=1}^{N-1} \frac{F_k}{N-k}
+ 1-H_N\right).\nonumber
\end{align}
\noindent
Substituting $F_k=F_k^{(N,G)}$ from Eq.~(\ref{eq:cdf}) for $\frac{1}{2}\neq G \in (0,1)$, we get:
\begin{align}\nonumber
\mathcal{B}\left(\boldsymbol{p}^{(N,G)}\right)
=&\nonumber
\frac{
N}{
(N-1)(2G-1)}\left(
G \sum_{k=1}^{N-1}
\frac{\Gamma(N)\Gamma(k-1+1/G)}{(N-k)\Gamma(N-1+1/G)\Gamma(k)}
-\frac{1-G}{N}\sum_{k=1}^{N-1} \frac{k}{N-k}
+ 1-H_N\right)\\ \nonumber
=&\frac{
N}{
N-1}\left[
\frac{G}{2G-1}\left( H_{N+1/G-2}-H_{1/G-1} \right)
-\frac{1-G}{(2G-1)N}\left(
N H_{N-1}-N+1
\right)
+ 1-H_N
\right]\\\nonumber
=&\frac{
N}{
N-1}\left[
\frac{G}{2G-1}\left( H_{N+1/G-2}-H_{1/G-1} \right)
-\frac{1-G}{2G-1}\left(
H_{N}-1
\right)
+ 1-H_N
\right]\\
=&\frac{
N}{
N-1}\frac{G}{2G-1}\left(
H_{N+1/G-2}-H_{1/G-1}
-H_{N}+1
\right)\nonumber
\\
=&
\frac{
N}{
N-1}\sum_{k=2}^N\frac{1}{k(k+1/G-2)}.\nonumber
\end{align}
\noindent
For $G=\frac{1}{2}$, we get:
\begin{equation}
\mathcal{B}\left(\boldsymbol{p}^{(N,1/2)}\right) =\frac{
N}{
N-1}\sum_{k=2}^N\frac{1}{k^2}.\nonumber
\end{equation}
\begin{remark}
Unlike in the Gini index's case,
the Bonferroni index depends on the vector length $N$.
However, in the limit as $N\to\infty$, we have:
\begin{equation*}
\mathcal{B}\left(\boldsymbol{p}^{(N,G)}\right)
\stackrel{N\to\infty}{\longrightarrow}
\begin{cases}
\displaystyle\frac{G}{2G-1}\left(1-H_{1/G-1}\right), & \mathrm{for\; }G\neq\frac{1}{2}, \\
\quad\\
\displaystyle\frac{\pi^2-6}{6}, & \mathrm{for\; }G=\frac{1}{2}. \\
\end{cases}
\end{equation*}
\end{remark}
\paragraph{De Vergottini's index.}
\label{subsec:vergottini-index}
The De~Vergottini index \citep{verg1,verg2} is defined as:
\begin{align}\nonumber
\mathcal{V}(p_1,\dots,p_N) = &
\frac{1}{\sum_{i=2}^N \frac{1}{i}} \left(
\sum_{k=1}^N \frac{F_k}{k} - 1
\right)
=
\frac{1}{\sum_{i=2}^N \frac{1}{i}} \left(
\sum_{i=1}^N \sum_{j=i}^{N} \frac{p_i}{j} - 1
\right)
\\
=&
\frac{1}{H_N-1}
\sum_{j=1}^N p_j \left(H_N-H_{j-1}-1\right)
=
1- \frac{\sum_{k=1}^N p_k H_{k-1}}{H_N-1}
.\nonumber
\end{align}
\noindent
Computing the value of the following sum in our model
(taking $p_k=p_k^{(N,G)}$ from Eq.~(\ref{eq:pk})
with $\frac{1}{2}\neq G \in (0,1)$):
\begin{align*}
\sum_{k=1}^N p_k^{(N,G)} H_{k-1} =&
\displaystyle
H_N-1 +\frac{G\left(\frac{G-1}{N}+1-2G\right)}{(G-1)(1-2G)}-\frac{\Gamma(1/G-2)\Gamma(N)}{\Gamma(N-1+1/G)},
\end{align*}
\noindent
leads to:
\begin{align*}
\mathcal{V}\left(\boldsymbol{p}^{(N,G)}\right)=
\displaystyle
\frac{1}{H_N-1}\left(\frac{\Gamma(1/G-2)\Gamma(N)}{\Gamma(N-1+1/G)}+\frac{G}{(2G-1)N}-\frac{G}{G-1}\right).
\end{align*}
\noindent
Furthermore, for $G=\frac{1}{2}$, we get:
\begin{equation*}
\mathcal{V}\left(\boldsymbol{p}^{(N,1/2)}\right) = \frac{N-H_N}{N
\left(H_N-1\right)}.
\end{equation*}
\begin{remark}\label{remark:vinfty}
For $G<1$, the De Vergottini index in our model converges to zero as $N\to\infty$.
However, the convergence rate is extremely slow.
\end{remark}
\paragraph{Hoover's index.}
\label{subsec:hoover-index}
The Hoover index (\citep{hoover}; also known as the Robin Hood index)
is defined by:
\begin{equation*}
\mathcal{H}(p_1,\,\dots,\,p_N)=
\frac{N}{2(N-1)} \sum_{k=1}^N\left|p_{k}-\frac{1}{N}\right|.
\end{equation*}
\noindent
It can be thought of as the normalised Manhattan distance
to the perfectly equal vector.
We can simplify the above by determining:
\begin{equation}\label{eq:nu}
\nu = \max\left\{j: p_j \ge \frac{1}{N} \right\},
\end{equation}
\noindent
and then writing:
\begin{align*}
\mathcal{H}(p_1,\,\dots,\,p_N)
\myoptionalpart{\\}{}
=&
\displaystyle
\frac{N}{2(N-1)}\left(
\displaystyle\sum _{k=1}^{\nu}\left(p_k-\frac{1}{N}\right)-\displaystyle\sum _{k=\nu+1}^N\left(p_k-\frac{1}{N}\right)
\right)
\\
=&
\displaystyle
\frac{N}{2(N-1)}\left(
\displaystyle\sum _{k=1}^{\nu}\left(p_k-\displaystyle\frac{1}{N}\right)
-
\left(
\underbrace{\displaystyle\sum _{k=1}^N\left(p_k-\displaystyle\frac{1}{N}\right)}_{=0}
-
\displaystyle\sum _{k=1}^\nu\left(p_k-\displaystyle\frac{1}{N}\right)
\right)
\right)
=
\frac{1}{N-1} \left(N F_{ \nu }- \nu \right).
\end{align*}
\noindent
Moreover, for $\frac{1}{2}\neq G\in(0,1)$, the above results in:
\begin{align*}
\mathcal{H}\left(\boldsymbol{p}^{(N,G)}\right)
=&\displaystyle\frac{G}{2G-1}\left(
\frac{N}{N-1}\frac{\Gamma(N)}{\Gamma(N-1+1/G)} \frac{\Gamma( \nu -1+1/G)}{\Gamma( \nu )}
-\frac{ \nu }{N-1}\right).\label{eq:H_solution}
\end{align*}
\begin{remark}\label{rem:PP}
We proved in \citep{pricepareto} (using a slightly different notation)
that for large $N$ and $G>\frac{1}{2}$,
we can express the solution of the continuous approximation
to Eq.~(\ref{eq:nu}), i.e.,
$p_\nu^{(N,G)}=1/N$, as:
\begin{equation*}
\nu\approx N \left(\frac{1-G}{G}\right)^{G/(2G-1)}.
\end{equation*}
\noindent
This way, we obtain a compact asymptotic formula for the
Hoover index:
\begin{align*}
\mathcal{H}\left(\boldsymbol{p}^{(N,G)}\right)
\stackrel{N\to\infty}{\longrightarrow}&
\displaystyle
\frac{G}{2G-1}\left[
\left(\frac{1-G}{G}\right)^{(1-G)/(2G-1)}
-\left(\frac{1-G}{G}\right)^{G/(2G-1)}\right]
=
\left(\frac{1-G}{G}\right)^{(1-G)/(2G-1)}.
\end{align*}
\noindent
Also note that for $G=\frac{1}{2}$, we can obtain $\nu$ using
the Euler--Maclauren formula, which yields:
\begin{equation*}
p_k^{(N,1/2)}\approx \frac{1}{N}\left(\frac{1}{2k}+\frac{1}{2N}-\log(k/N) \right),
\end{equation*}
\noindent
This allows us to solve equation $p^{(\nu,1/2)}=1/N$ through
the Lambert $\mathcal{W}$ function, which could be further expanded as:
\begin{align*}
\nu
\approx &
\frac{1}{2}\left[ \mathcal{W}\left(\displaystyle\frac{\exp\left(1-\frac{1}{2 N}\right)}{2 N}\right)\right]^{-1}
=
\displaystyle
\frac{1+e}{2 e}+ \frac{N}{e}+\frac{1-e^2}{8 e N}+O\left(N^{-2}\right)
=
\frac{(2N+1)e^{-1}+1}{2 }+O\left(N^{-1}\right),
\end{align*}
\noindent
leading us to:
\begin{align*}
\mathcal{H}\left(\boldsymbol{p}^{(N,1/2)}\right)
\stackrel{N\gg 1}{\approx} &
\displaystyle
\frac{N}{N-1} F^{(N,1/2)}_{
\left\lfloor \frac{2Ne^{-1}+e^{-1}+1}{2} \right\rfloor
}
-
\frac{1}{N-1}
\left\lfloor \frac{2Ne^{-1}+e^{-1}+1}{2} \right\rfloor
\myoptionalpart{\\}{}
\stackrel{N\to\infty}{\longrightarrow}
e^{-1}.
\end{align*}
\end{remark}
\paragraph{$\mathcal{P}_q$ indices.}
\label{subsec:pq-index}
The $\mathcal{P}_q$ index for $q\in(0,1)$ is defined as:
\begin{equation}\label{Eq:pq}
\mathcal{P}_q(p_1,\,\dots,\,p_N)=\frac{\sum_{j=1}^{\lfloor q
N\rfloor}p_j-q}{1-q}=\frac{F_{\lfloor qN \rfloor}-q}{1-q}.
\end{equation}
\noindent
This index is a normalised version of the percentage of accumulated
probability mass in $q 100\%$ of the top elements (compare the famous Pareto 80/20 rule).
In our model, for $\frac{1}{2} \neq G \in(0,1) $, the $\mathcal{P}_q$ index is equal to
\begin{align*}
\mathcal{P}_q\left(\boldsymbol{p}^{(N,G)}\right)
=&
\displaystyle
\frac{G}{(2G-1)(1-q)}\left(\frac{\Gamma(N)}{\Gamma(N-1+1/G)} \frac{\Gamma(\lfloor qN \rfloor-1+1/G)}{\Gamma(\lfloor qN \rfloor)}
-q\right).
\end{align*}
\noindent
Furthermore, if $G=\frac{1}{2}$, then it holds:
\begin{equation}
\mathcal{P}_q\left(\boldsymbol{p}^{(N,1/2)}\right)
= \frac{q}{1-q}(H_N-H_{\lfloor qN \rfloor}).
\end{equation}
\noindent
\begin{remark}
In the limit as $N\to\infty$, we have:
\begin{equation*}
\mathcal{P}_q\left(\boldsymbol{p}^{(N,G)}\right)
\stackrel{N\to\infty}{\longrightarrow}
\begin{cases}
\displaystyle\frac{G}{(2G-1)(1-q)} \left(
q^{1/G-1}
-q
\right), & \text{for }G\neq\frac{1}{2}, \\
\quad\\
\displaystyle\frac{-q\log(q)}{1-q}, & \text{for }G=\frac{1}{2}. \\
\end{cases}
\end{equation*}
\end{remark}
\paragraph{Indices as functions of one another.}\label{subsec:equivalence}
To sum up, we have shown that in our model,
the formulae for the Bonferroni, De Vergottini, Hoover, and $\mathcal{P}_q$
indices depend only on $N$ and $G$:
\begin{align*}
\tilde{B}(N,G) =&
\displaystyle
\frac{
N}{
N-1}\sum_{k=2}^N\frac{1}{k(k+1/G-2)},
\\
\tilde{V}(N, G)=&
\displaystyle
\frac{1}{H_N-1}\left(\frac{\Gamma(1/G-2)\Gamma(N)}{\Gamma(N-1+1/G)}+\frac{G}{(2G-1)N}-\frac{G}{G-1}\right),
\\
\tilde{H}(N, G) =&
\displaystyle
\frac{N}{N-1} \left(\frac{G}{2G-1}
\frac{\Gamma(N) }{\Gamma(N-1+1/G) }
\frac{\Gamma(\nu-1+1/G)}{\Gamma(\nu)}
- \frac{\nu}{N} \right),
\qquad(\text{with }\nu = \max\{j: p_j \ge {1}/{N} \})
\\
\tilde{P}_q(N, G) =&
\displaystyle
\frac{G}{(2G-1)(1-q)}\left(\frac{\Gamma(N)}{\Gamma(N-1+1/G)} \frac{\Gamma(qN-1+1/G)}{\Gamma(qN)}
-q\right),\\
\end{align*}
\noindent
respectively (for readability, we only included the case of $G\neq 1/2$).
Figure~\ref{fig:GvsOthers} depicts them for different $N$s.
\begin{figure}[bht!]
\begin{center}
\includegraphics[width=0.4\textwidth]{cropped-function_BG}
\includegraphics[width=0.4\textwidth]{cropped-function_VG}
\includegraphics[width=0.4\textwidth]{cropped-function_HG}
\includegraphics[width=0.4\textwidth]{cropped-function_pqG}
\end{center}
\caption{
Functional dependence between the Gini index and other inequality measures
for different sample sizes $N$ in our model.
The indexes are one-to-one functions of one another.
Therefore, similar plots could be drawn for $V$ as a function of $B$,
$H$ as a function of $P_q$, etc.
In this context, there is no need to multiply entities without necessity.
The Gini index can be used as the only measure of inequality;
it is the only one that is independent of $N$.
}
\label{fig:GvsOthers}
\end{figure}
We note that they are all strictly increasing, continuous functions of $G$.
Thus, based on the derived formulas, all the indices can be expressed
as one-to-one functions of one another.
For instance, given some value of the Bonferroni index $B$,
we can obtain the underlying $G^* = \tilde{B}^{-1}(N, B)$
and then compute, say, $\tilde{V}(N, G^*)$.
Even though the analytic formulae for the inverses do not exist,
this still can easily be solved numerically.
In this sense, we can say that -- \textit{in our model} --
all the aforementioned inequality indices are equivalent.
Similar derivations can be performed for many other inequality indices,
although they will not necessarily enjoy analytic solutions.
\section{Analysis of citation data from RePEc}\label{sec:repec}
Theoretical models are merely approximations of the
real-world phenomena under scrutiny. Thus,
empirical data might deviate from the assumed idealisations.
Also, looking beyond our simple iterative process,
we know that uncountably many probability vectors yield
a specific Gini, Bonferroni, or any other index.
After all, these measures were introduced to respond to the different needs
of the practitioners; see, e.g., \citep{imedioolmedo,parasite}.
Some of them are, for example, more responsive to
increasing (via the principle of progressive transfers)
the amount of probability mass in the tail of the distribution
than others.
In particular, as reported in \citep{ciommi}, the Gini, Bonferroni,
and De Vergottini indices belong to the class of linear measures introduced in
\citep{mehran}. They note that \textit{for the Bonferroni and De Vergottini
indices, the effect of a transfer also depends on the position of individuals,
making the Bonferroni index more sensitive to transfers that occur at the lower
end of the income distribution and the De Vergottini index more sensitive to
variations among the richest}.
Therefore, we should be interested in verifying how well our formulae
describe the underlying structure of real datasets.
Let us thus consider citation data from the RePEc database
(Research Papers in Economics; see \texttt{https://citec.repec.org/}), which
features $66{,}347$ authors and $1{,}843{,}967$ papers.
In the data cleansing step, we have omitted the authors who published less than
$5$ cited papers and whose $h$-index was less than $3$.
This resulted in $n=36{,}425$ citation records of the form
$\left(x_1^{(i)},\dots,x^{(i)}_{N^{(i)}}\right)$, where
$N^{(i)}$ gives the total number of items published by the $i$-th author
and $x_k^{(i)}$ gives the number of citations to their $k$-th most cited works.
Figure~\ref{fig:repec_3vectors_loglog} shows
three example (quite representative of the whole database)
vectors from the RePEc database:
observed (points) and predicted (lines) values for $p_k$ (left)
and their non-normalised versions, $x_k$ (right).
We see a good fit over most parts of the data domain.
This comes at no surprise, as the proposed model is equivalent
to the 3DI model~\citep{PNAS2020} and we have already
seen its usefulness in the case of modelling citations to computer science papers.
\begin{figure}[ht!]
\begin{center}
\includegraphics[width=0.45\textwidth]{repec_3vectors_loglog}
\includegraphics[width=0.45\textwidth]{repec_3vectors_loglog_2}
\end{center}
\caption{
Three example citation vectors (left: normalised, right: original) and
the corresponding fitted models (Eq.~(\ref{eq:process}));
note the log-log scale. We note a very good fit in each case.
}
\label{fig:repec_3vectors_loglog}
\end{figure}
\bigskip
For each author, we compute their actual (observed) Gini index
$\hat{G}^{(i)} = \mathcal{G}\left(p_1^{(i)},\dots,p^{(i)}_{N^{(i)}}\right)$
with $p_k^{(i)} = X_k^{(i)}/\sum_{j=1}^{N^{(i)}} X_j^{(i)}$.
Then, based on the derived formulae,
we compute the predicted
$\tilde{B}(N, \hat{G}^{(i)})$, $\tilde{V}(N, \hat{G}^{(i)})$, \dots,
and observed
$\hat{B}^{(i)} = \mathcal{B}\left(p_1^{(i)},\dots,p^{(i)}_{N^{(i)}}\right)$,
$\hat{V}^{(i)} = \mathcal{V}\left(p_1^{(i)},\dots,p^{(i)}_{N^{(i)}}\right)$, \dots,
Bonferroni, De Vergottini, and other indices.
Recall that Figure~\ref{fig:GvsOthers} describes the
theoretical relationships between the Gini index and other metrics.
If, overall, our model describes the real data well,
we should expect to see these dependencies in the case of the RePEc vectors too.
Figure~\ref{fig:repec_indG_largeN} presents a scatter plot
of the values of different indices
as functions of $\hat{G}$ for all vectors with $N>200$
(the De Vergottini index was not included as it approaches $0$ for large $N$s;
see Remark \ref{remark:vinfty}).
Ideally, they should lie close to the theoretical curves
(depicted as well). And this is approximately the case.
Furthermore, Figure~\ref{fig:repec} presents similar results, but for vectors of
lengths $N=6$ (green), $N=10$ (pink), and $N=50\pm 1$ (blue;
the $\pm 1$ part is to increase the number of data points).
Additionally, we coloured the areas representing all of the possible values
which could be obtained for vectors generated using our model. In other words,
for any vector, we expect the pairs $(\hat{G}, \hat{B})$,
$(\hat{G}, \hat{V})$, etc.~to lie in the grey zone.
This is true for the vast majority of the real data points.
The prediction errors for different vector lengths are
summarised in Table~\ref{tab:errors}.
The error rates are usually less than 2--3\%.
Overall, we observe a very good fit, confirming our theoretical derivations.
The choice of the inequality measure is secondary, as the indices
can be considered functions of one another. Therefore, it is best
to be faithful to the simplest indicator: the Gini index.
\begin{table}[b!]
\centering
\caption{Left: Mean absolute prediction errors, i.e., the average of
$|\hat{I}^{(i)} - \tilde{I}(N^{(i)}, \hat{G}^{(i)})|$
for all vectors $\boldsymbol{p}^{(i)}$ with $N^{(i)}=N$,
for $N=5,10, 20, 50$ and $N\geq 200$ and different indices $I$.
Right: Mean prediction errors (bias), i.e., the average of
$(\hat{I}^{(i)} - \tilde{I}(N^{(i)}, \hat{G}^{(i)}))$.
Overall, the errors are small and the longer the vector, the better
the precision with which we can predict the values of the remaining
indices.
Note that the errors for the De Vergottini index for large $N$ are not given due
to its limiting behaviour (as per Remark~\ref{remark:vinfty}).
}\label{tab:errors}
\centering
\begin{minipage}{0.4\linewidth}\centering
Mean absolute prediction errors
\begin{tabular}[t]{r|r|r|r|r|r}
& $6$&$10$&$20$&$50$&$\geq 200$ \\\hline
$B$&\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.017}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.019}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.019}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.011}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.006}}\\\hline
$V$&\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.019}}&
\cellcolor[HTML]{FEE18D}{\textcolor{black}{0.022}}&
\cellcolor[HTML]{FDD582}{\textcolor{black}{0.026}}&
\cellcolor[HTML]{FDD582}{\textcolor{black}{0.026}}&
---\\\hline
$H$&\cellcolor[HTML]{FDD380}{\textcolor{black}{0.027}}&
\cellcolor[HTML]{FDD582}{\textcolor{black}{0.026}}&
\cellcolor[HTML]{FEE18D}{\textcolor{black}{0.022}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.017}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.011}}\\\hline
$P_{0.5}$&\cellcolor[HTML]{FBA35B}{\textcolor{black}{0.042}}&
\cellcolor[HTML]{FDBA6B}{\textcolor{black}{0.035}}&
\cellcolor[HTML]{FDCC7A}{\textcolor{black}{0.029}}&
\cellcolor[HTML]{FEE28F}{\textcolor{black}{0.021}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.016}}\\\hline
\end{tabular}
\end{minipage}
$\qquad$
\begin{minipage}{0.4\linewidth}\centering
Bias
\begin{tabular}[t]{r|r|r|r|r|r}
& $6$&$10$&$20$&$50$&$\geq 200$ \\\hline
$B$&\cellcolor[HTML]{FEFEBE}{\textcolor{black}{-0.004}}&
\cellcolor[HTML]{F4FAAE}{\textcolor{black}{-0.013}}&
\cellcolor[HTML]{F2FAAB}{\textcolor{black}{-0.015}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{-0.008}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{-0.003}}\\\hline
$V$&\cellcolor[HTML]{FEFEBE}{\textcolor{black}{-0.001}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.002}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{-0.001}}&
\cellcolor[HTML]{F3FAAC}{\textcolor{black}{-0.014}}&
---\\\hline
$H$&\cellcolor[HTML]{FEE99A}{\textcolor{black}{0.021}}&
\cellcolor[HTML]{FEE99B}{\textcolor{black}{0.02}}&
\cellcolor[HTML]{FEECA0}{\textcolor{black}{0.018}}&
\cellcolor[HTML]{FEF0A7}{\textcolor{black}{0.014}}&
\cellcolor[HTML]{FEF3AB}{\textcolor{black}{0.011}}\\\hline
$P_{0.5}$&\cellcolor[HTML]{FEF1A8}{\textcolor{black}{0.013}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.005}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.002}}&
\cellcolor[HTML]{FEFEBE}{\textcolor{black}{0.007}}&
\cellcolor[HTML]{FEF3AB}{\textcolor{black}{0.011}}\\\hline
\end{tabular}
\end{minipage}
\end{table}
\section{Analysis of income data}\label{sec:incomedata}
Additionally, let us study the Luxembourg Income Study dataset
giving the family incomes in nine countries.
The original data, in the form of the Lorenz (e.g., \citealp{Bertoli2019})
curves across the deciles ($N=10$),
are given in \citep[Table 2]{Bishop}.
Table~\ref{tab:bishop} gives the prediction error for the indices studied.
Overall, the errors are very small, except
for the case of the $\mathcal{P}_q$ index with higher $q$ in Sweden, Germany, and the United Kingdom,
where there is some discrepancy between the predicted and the observed Lorenz curve
in their tails.
\begin{table}[b!]
\caption{\label{tab:bishop} Predicted vs observed indices for family income data \citep[Table 2]{Bishop}; e.g., $\tilde{B}+(\hat{B}-\tilde{B})=0.50-0.01$ denotes that
the predicted $\tilde{B}=0.50$ and the observed $\hat{B}=0.50-0.01$, hence the prediction
error is $0.01$
}
\centering
\begin{tabular}{lrrrrrr}
\toprule
country & $\hat{G}=\tilde{G}$ & $\tilde{B}+(\hat{B}-\tilde{B})$ & $\tilde{V}+(\hat{V}-\tilde{V})$ & $\tilde{H}+(\hat{H}-\tilde{H})$ & $\tilde{P}_{0.5}+(\hat{P}_{0.5}-\tilde{P}_{0.5})$ & $\tilde{P}_{0.9}+(\hat{P}_{0.9}-\tilde{P}_{0.9})$ \\
\midrule
AU & $0.38$ & $0.50-0.01$ & $0.25+0.02$ & $0.29-0.01$ & $0.52-0.02$ & $0.85-0.04$ \\
CA & $0.38$ & $0.50-0.01$ & $0.25+0.02$ & $0.28-0.01$ & $0.51-0.02$ & $0.84-0.02$ \\
NL & $0.34$ & $0.46-0.02$ & $0.22+0.03$ & $0.26-0.01$ & $0.47-0.03$ & $0.82-0.03$ \\
NO & $0.34$ & $0.46-0.01$ & $0.22+0.02$ & $0.26-0.01$ & $0.46-0.02$ & $0.82-0.04$ \\
SE & $0.32$ & $0.44-0.03$ & $0.20+0.02$ & $0.24-0.01$ & $0.44-0.02$ & $0.81-0.10$ \\
CH & $0.42$ & $0.54-0.02$ & $0.29+0.03$ & $0.31-0.02$ & $0.56-0.05$ & $0.87-0.01$ \\
DE & $0.34$ & $0.46-0.03$ & $0.22+0.02$ & $0.26-0.01$ & $0.46-0.03$ & $0.82-0.08$ \\
UK & $0.38$ & $0.50-0.02$ & $0.25+0.02$ & $0.28-0.01$ & $0.51-0.02$ & $0.84-0.07$ \\
US & $0.40$ & $0.52-0.01$ & $0.27+0.01$ & $0.30-0.01$ & $0.54-0.02$ & $0.86+0.00$ \\
\bottomrule
\end{tabular}
\end{table}
\begin{figure}[p!]
\begin{center}
\includegraphics[width=0.45\linewidth]{cropped-repec_indG_largeN.pdf}
\end{center}
\caption{
Inequality indices as functions of the Gini index
for all RePEc citation vectors with $N>200$ (points).
They match the derived theoretical curves (lines) very well,
confirming the validity of our model,
and the main hypothesis of our paper.
}
\label{fig:repec_indG_largeN}
\end{figure}
\begin{figure}[p!]
\begin{center}
\includegraphics[width=0.4\textwidth]{cropped-repec_BG_2.pdf}
\includegraphics[width=0.4\textwidth]{cropped-repec_VG_2.pdf}
\includegraphics[width=0.4\textwidth]{cropped-repec_HG_2.pdf}
\includegraphics[width=0.4\textwidth]{cropped-repec_pqG_2.pdf}
\end{center}
\caption{
Predicted ($\tilde{B}(N, \hat{G})$ vs $\hat{G}$,
$\tilde{V}(N, \hat{G})$ vs $\hat{G}$, etc.; thick curves)
and
observed ($\hat{B}$ vs $\hat{G}$ etc.; points)
values of different inequality measures
for the RePEc citation records of different lengths $N$.
Our model recreates the indices with decent level of precision.
}
\label{fig:repec}
\end{figure}
\section{Conclusions}\label{sec:conclusion}
The discussed process yields $\boldsymbol{p}^{(N,G)}$ that are totally ordered by the majorisation relation $\preceq$. Following \citep{shorrocks1987transfer} (but see also \citealp{marshall2011majorisation,lambert2001distribution,chantreuil2011inequality,zheng2007unit,bosmans2016consistent}), a function $I$ is an index of inequality if and only if it is symmetric and strictly Schur convex.
Hence, for all ordered probability $N$-vectors $\boldsymbol{p}\neq\boldsymbol{q}$, if $\boldsymbol{p}\preceq \boldsymbol{q}$, then $I(\boldsymbol{p})<I(\boldsymbol{q})$, where the direction of the inequality is uniquely determined by the Pigou–Dalton condition (e.g., \citealp{Patty2019}).
In particular, for every vector $\boldsymbol{p}^{(N,G)}$, the parameter $G$ can be interpreted as the
(normalised) Gini index $\mathcal{G}(\boldsymbol{p}^{(N,G)})=G$. This implies that $\mathcal{G}$ is an order preserving function \citep{marshall2011majorisation} on $\{\boldsymbol{p}^{(N,G)}\}$ with respect to $G$.
In this study, we also extended this characterisation result to some other notable indices of inequality, that is, the Bonferroni, De Vergottini, Hoover and $\mathcal{P}_q$ indices, deriving their explicit expressions as one-to-one functions of $\mathcal{G}$.
For data following closely our model, the indices can be considered equivalent.
An analysis of two empirical datasets (citation vectors in economics and countrywise family incomes)
confirms our results.
\subsection*{Acknowledgements}
We are indebted to Jose Manuel Barrueco for providing us with a large snapshot
of RePEc (Research Papers in Economics) data.
All data are freely available at \texttt{http://citec.repec.org/api.html}.
This research was supported by the Australian Research Council Discovery
Project ARC DP210100227 (MG).
\subsection*{Conflict of interest}
The authors certify that they have no affiliations with or involvement in any
organisation or entity with any financial interest or non-financial interest
in the subject matter or materials discussed in this manuscript.
\printcredits
\begin{thebibliography}{38}
\expandafter\ifx\csname natexlab\endcsname\relax\def\natexlab#1{#1}\fi
\ifx\xfnm\relax \def\xfnm[#1]{\unskip,\space#1}\fi
\bibitem[{Arnold(2015)}]{arnold2015:pareto}
Arnold, B.C., 2015.
\newblock Pareto Distributions.
\newblock Chapman and Hall/CRC, New
York, NY, USA.
\newblock doi:10.1201/b18141.
\bibitem[{Beliakov et~al.(2016)Beliakov, Gagolewski and
James}]{BeliakovETAL2016:penaltyinequality}
Beliakov, G., Gagolewski, M.,
James, S., 2016.
\newblock Penalty-based and other representations of economic
inequality.
\newblock International Journal of Uncertainty, Fuzziness and
Knowledge-Based Systems 24(Suppl.1),
1--23.
\newblock doi:10.1142/S0218488516400018.
\bibitem[{Bertoli-Barsotti(2023)}]{bertoli2022equivalent}
Bertoli-Barsotti, L., 2023.
\newblock Equivalent {G}ini coefficient, not shape parameter!
\newblock Scientometrics 128,
867--870.
\newblock doi:10.1007/s11192-022-04571-8.
\bibitem[{Bertoli-Barsotti et~al.(2023)Bertoli-Barsotti, Gagolewski, Siudem and
Żogała Siudem}]{OurGiniLorenz}
Bertoli-Barsotti, L., Gagolewski, M.,
Siudem, G., Żogała Siudem, B.,
2023.
\newblock Gini-stable {L}orenz curves and their relation to the
generalised {P}areto distribution.
\newblock Submitted for publication; see
\texttt{https://arxiv.org/a/gagolewski_m_1.html}.
\bibitem[{Bertoli-Barsotti and Lando(2019)}]{Bertoli2019}
Bertoli-Barsotti, L., Lando, T.,
2019.
\newblock How mean rank and mean size may determine the
generalised {L}orenz curve: {W}ith application to citation analysis.
\newblock Journal of Informetrics 13,
387--396.
\newblock doi:10.1016/j.joi.2019.02.003.
\bibitem[{Bishop et~al.(1991)Bishop, Formby and Smith}]{Bishop}
Bishop, J.A., Formby, J.P.,
Smith, W.J., 1991.
\newblock International comparisons of income inequality:
{T}ests for {L}orenz dominance across nine countries.
\newblock Economica 58,
461--477.
\bibitem[{Bonferroni(1930)}]{Bonferroni1930:index}
Bonferroni, C., 1930.
\newblock Elementi di statistica generale.
\newblock Libreria Seber, Firenze.
\bibitem[{Bosmans(2016)}]{bosmans2016consistent}
Bosmans, K., 2016.
\newblock Consistent comparisons of attainment and shortfall
inequality: {A} critical examination.
\newblock Health Economics 25,
1425--1432.
\bibitem[{Cena et~al.(2022)Cena, Gagolewski, Siudem and Żogała
Siudem}]{cena2022validating}
Cena, A., Gagolewski, M.,
Siudem, G., Żogała Siudem, B.,
2022.
\newblock Validating citation models by proxy indices.
\newblock Journal of Informetrics 16,
101267.
\newblock doi:10.1016/j.joi.2022.101267.
\bibitem[{Chantreuil and Trannoy(2011)}]{chantreuil2011inequality}
Chantreuil, F., Trannoy, A.,
2011.
\newblock Inequality decomposition values.
\newblock Annals of Economics and Statistics/Annales
d'{\'E}conomie et de Statistique 101/102,
13--36.
\bibitem[{Ciommi et~al.(2022)Ciommi, Gigliarano and Giorgi}]{ciommi}
Ciommi, M., Gigliarano, C.,
Giorgi, G., 2022.
\newblock {B}onferroni and de {V}ergottini are back: {N}ew
subgroup decompositions and bipolarization measures.
\newblock Fuzzy Sets and Systems 433,
22--53.
\bibitem[{Cowell(2000)}]{cowell2000measurement}
Cowell, F.A., 2000.
\newblock Measurement of inequality, in:
Atkinson, A., Bourguignon, F. (Eds.),
Handbook of Income Distribution.
Elsevier. volume~1, pp.
87--166.
\bibitem[{Gagolewski et~al.(2022)Gagolewski, Żogała Siudem, Siudem and
Cena}]{gagolewski2022ockham}
Gagolewski, M., Żogała Siudem, B.,
Siudem, G., Cena, A.,
2022.
\newblock {O}ckham's index of citation impact.
\newblock Scientometrics 127,
2829--2845.
\newblock doi:10.1007/s11192-022-04345-2.
\bibitem[{Gini(1912)}]{Gini1912:index}
Gini, C., 1912.
\newblock Variabilità e mutabilità.
\newblock C. Cuppini, Bologna.
\bibitem[{Holme(2022)}]{holme2022universality}
Holme, P., 2022.
\newblock Universality out of order.
\newblock Nature Communications 13,
2355.
\bibitem[{Hoover(1941)}]{hoover}
Hoover, E., 1941.
\newblock Interstate redistribution of population,
1850–1940.
\newblock The Journal of Economic History
1, 199--205.
\bibitem[{Imedio-Olmedo et~al.(2012)Imedio-Olmedo, Parrado-Gallardo and
Bárcena-Martín}]{imedioolmedo}
Imedio-Olmedo, L., Parrado-Gallardo, E.,
Bárcena-Martín, E., 2012.
\newblock Income inequality indices interpreted as measures of
relative deprivation/satisfaction.
\newblock Social Indicators Research 109,
471--491.
\bibitem[{I{\~n}iguez et~al.(2022)I{\~n}iguez, Pineda, Gershenson and
Barab{\'a}si}]{iniguez2022dynamics}
I{\~n}iguez, G., Pineda, C.,
Gershenson, C., Barab{\'a}si, A.L.,
2022.
\newblock Dynamics of ranking.
\newblock Nature communications 13,
1646.
\bibitem[{Lambert(2001)}]{lambert2001distribution}
Lambert, P.J., 2001.
\newblock The distribution and redistribution of income (3rd
Edition).
\newblock Manchester University Press.
\bibitem[{Marshall and Olkin(2007)}]{marshall2007life}
Marshall, A.W., Olkin, I.,
2007.
\newblock Life distributions: {S}tructure of Nonparametric,
Semiparametric, and Parametric Families.
\newblock Springer.
\bibitem[{Marshall et~al.(2011)Marshall, Olkin and
Arnold}]{marshall2011majorisation}
Marshall, A.W., Olkin, I.,
Arnold, B.C., 2011.
\newblock Inequalities: Theory of majorization and its
applications (2nd Edition).
\newblock Springer Science Business Media.
\bibitem[{McVinish and Lester(2020)}]{parasite}
McVinish, R., Lester, R.,
2020.
\newblock Measuring aggregation in parasite populations.
\newblock Journal of the Royal Society Interface
17, 20190886.
\newblock doi:10.1098/rsif.2019.0886.
\bibitem[{Mehran(1976)}]{mehran}
Mehran, F., 1976.
\newblock Linear measures of income inequality.
\newblock Econometrica 44,
805--809.
\bibitem[{Newman(2005)}]{Newman2005}
Newman, M., 2005.
\newblock Power laws, {P}areto distributions and {Z}ipf's law.
\newblock Contemporary Physics 46,
323--351.
\newblock doi:10.1080/00107510500052444.
\bibitem[{Patty and Penn(2019)}]{Patty2019}
Patty, J.W., Penn, E.M.,
2019.
\newblock Measuring fairness, inequality, and big data:
{S}ocial choice since {A}rrow.
\newblock Annual Review of Political Science
22, 435--460.
\newblock doi:10.1146/annurev-polisci-022018-024704.
\bibitem[{Petersen et~al.(2011)Petersen, Stanley and Succi}]{Petersen2011}
Petersen, A.M., Stanley, H.E.,
Succi, S., 2011.
\newblock Statistical regularities in the rank-citation profile
of scientists.
\newblock Scientific Reports 1,
181.
\newblock doi:10.1038/srep00181.
\bibitem[{Pickands~III(1975)}]{pickands1975statistical}
Pickands~III, J., 1975.
\newblock Statistical inference using extreme order
statistics.
\newblock The Annals of Statistics ,
119--131.
\bibitem[{Price(1965)}]{deSollaPrice1965}
Price, D., 1965.
\newblock Networks of scientific papers.
\newblock Science 149,
510--515.
\newblock doi:10.1126/science.149.3683.510.
\bibitem[{Shorrocks and Foster(1987)}]{shorrocks1987transfer}
Shorrocks, A.F., Foster, J.E.,
1987.
\newblock Transfer sensitive inequality measures.
\newblock The Review of Economic Studies
54, 485--497.
\bibitem[{Silber(2012)}]{silber2012handbook}
Silber, J., 2012.
\newblock Handbook of income inequality measurement.
volume~71.
\newblock Springer Science \& Business Media.
\bibitem[{Singh et~al.(2022)Singh, Barme, Ward, Tupikina and
Santolini}]{singh2022quantifying}
Singh, C.K., Barme, E.,
Ward, R., Tupikina, L.,
Santolini, M., 2022.
\newblock Quantifying the rise and fall of scientific fields.
\newblock PLoS ONE 17,
e0270131.
\bibitem[{Siudem et~al.(2022)Siudem, Nowak and Gagolewski}]{pricepareto}
Siudem, G., Nowak, P.,
Gagolewski, M., 2022.
\newblock Power laws, the {P}rice model, and the {P}areto
type-2 distribution.
\newblock Physica A: Statistical Mechanics and its
Applications 606, 128059.
\newblock doi:10.1016/j.physa.2022.128059.
\bibitem[{Siudem et~al.(2020)Siudem, {\.Z}oga{\l}a-Siudem, Cena and
Gagolewski}]{PNAS2020}
Siudem, G., {\.Z}oga{\l}a-Siudem, B.,
Cena, A., Gagolewski, M.,
2020.
\newblock Three dimensions of scientific impact.
\newblock Proceedings of the National Academy of Sciences
117, 13896--13900.
\newblock doi:10.1073/pnas.2001064117.
\bibitem[{Vergottini(1940)}]{verg1}
Vergottini, M.D., 1940.
\newblock Sul significato di alcuni indici di concentrazione.
\newblock Giornale degli Economisti e Annali di Economia
11, 317--347.
\bibitem[{Vergottini(1950)}]{verg2}
Vergottini, M.D., 1950.
\newblock Sugli indici di concentrazione.
\newblock Statistica 10,
445--454.
\bibitem[{Voitalov et~al.(2019)Voitalov, van~der Hoorn, van~der Hofstad and
Krioukov}]{voitalov2019scale}
Voitalov, I., van~der Hoorn, P.,
van~der Hofstad, R., Krioukov, D.,
2019.
\newblock Scale-free networks well done.
\newblock Physical Review Research 1,
033034.
\bibitem[{Yitzhaki and Schechtman(2013)}]{ginimeth}
Yitzhaki, S., Schechtman, E.,
2013.
\newblock The {G}ini methodology: {A} primer on a statistical
methodology.
\newblock Springer.
\bibitem[{Zheng(2007)}]{zheng2007unit}
Zheng, B., 2007.
\newblock Unit-consistent decomposable inequality measures.
\newblock Economica 74,
97--111.
\end{thebibliography}