Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
38,292 characters · 10 sections · 21 citation commands
Generalized Beta Prime Distribution: Stochastic Model of Economic Exchange and Properties of Inequality Indices
Voluminous literature exists chotikapanich2008modelling on application of Generalized Beta Prime (also known as GB2) to wealth/income distributions. Analytical results are known (see Chapters 3 and 8 in mcdonald2008modelling) for Theil and Gini measures of inequality. Numerical and analytical results using GB2 are routinely used in various economic analyses (for a recent example see chen2016influences, chotikapanich2018using). Yet, there seems to be little discussion as to why GB2 is so uniquely suited for modeling wealth/income distributions. Additionally, there seems to be a glaring lack of maximum likelihood estimation (MLE) and Bayesian fitting, as well as of use of more accurate statistical measures of fits, such as Kolmogorov-Smirnov statistics.
Consequently, the main motivation of this work is to connect GB2 with a plausible stochastic model hertzler2003classical of economic exchange and to better understand parametric dependences and symmetries of the distribution and quantities derived thereof, such as Gini, Hoover (also known as Pietra or Schultz) and Theil indices. We also want to conduct a systematic statistical fitting and quantify goodness of fits. Finally, we propose an alternative measure of inequality, which -- similarly to Hoover and Theil L -- is less prone to exaggerating high and low income/wealth when described by power-law distribution dependencies, such as GB2.
This paper is organized as follows. In Section (ref) we summarize the properties of distributions used for fitting in Section (ref). In Section (ref) we discuss the analytic properties of GB2 and inequality indices. Initially we discuss its particular case of Beta Prime, for which closed-form results are simple and easy to understand. In Section (ref) we discuss the stochastic model of economic exchange, whose steady-state distribution is GB2. In Section (ref) we use housing sales data as a proxy to wealth/income distribution for fitting with distributions of Section (ref). In (ref) we introduce a new measure of inequality, DMMS. In (ref) we expand the number of fitting functions. Finally, in (ref)-(ref) we compare market values distribution with that of sale prices and investigate the effect of the most expensive properties on tail parameters.
Tables (ref) - (ref) contain probability density functions (PDF) and cumulative density function (CDF) of the distributions used for fitting in this article and inequality indices derived for them.\footnote{BP is a particular case of GB2. Expanded tables, that contain IGa and Ga -- particular cases of GIGa and GGa -- are given in (ref).} Here $\Gamma (x)$ is the gamma function, $\mathrm{B}(x,y)$ is the beta function, $\psi(x)$ is the digamma function, $erf(x)$ is the error function and $\prescript{}{m}{F}{}_{n}$ is the hypergeometric function. The fat-tailed Generalized Inverse Gamma (GIGa) appear in network models of economy bouchaud2000wealth,ma2013distribution and is related to Generalized Gamma (GGa) by changing the variable of the PDF to its inverse. These distributions and their particular cases also appear in the models of stochastic volatility and stock returns dragulescu2002probability, ma2014model; for the application of BP in the the latter context see dashti2018combined. GIGa and GGa can be also viewed as the limiting cases of GB2 for $p \rightarrow \infty$ and $q \rightarrow \infty$ respectively mcdonald1995generalization. Lognormal (LN) distribution is also widely used in economics and finance limpert2001lognormal.
A succinct summary of inequality indices for many distributions can be found in mcdonald2008modelling. We derived all inequality indices for GIGa, as well as Theil L and Hover for GB2, for this paper. However, we subsequently found that a very nice, closed-form formula for the Hoover index (Prieta there) had been previously obtained for an arbitrary distribution in terms of its CDF sarabia2014explicit and that Theil L had also been previously derived chotikapanich2018using. We point out the difference structure of all the indices which reflect the symmetry associated with the transformation of variable of the distribution to its inverse, which converts low end of GB2 to high end and vice versa (see below) and GIGa to GGa and vice versa.
Gini, Hoover and Theil indices are calculated using the following formulae, where $p(x)$ is a PDF:
From ((ref)) - ((ref)), it is clear that Gini and Theil T exaggerate power-law dependences of a distribution -- low-end for GGa, high-end fat tails for GIGa and both for GB2 (and BP) -- in that they underweigh low end and overweigh fat tails. Consequently, we believe that Hoover and Theil L are more appropriate measures for distributions with power laws, especially with fat tails. In (ref) we introduce yet another measure of inequality, DMMS, which tries to address the same problem. As will be seen below, Theil L is smaller than Theil T both for actual data and for fat-tailed distribution fits, while Hoover and MDDS are smaller than Gini, with MDDS being smaller than Hoover for all fitted distributions.
We begin our analysis of GB2
with the simpler case of BP, $\alpha=1$ in ((ref)).
with the simpler case of BP, $\alpha=1$ in ((ref)). The limiting behaviors of BP at small and large arguments are
and
Here we assume that $p>1$, that is that BP has a maximum (a bell shape), and $q>2$, that is that the variance exists (we will discuss the example to the contrary for market values in the Appendix). The limiting behaviors ((ref)) and ((ref)) underscore the flexibility of BP: large-$p$ behavior mimics exponential decay of IGa (see (ref)) at small values and corresponds to suppression of low-end values, while large-$q$ behavior mimics exponential decay of Ga and corresponds to suppression of high-end values (see also mcdonald1995generalization).
An important property of BP is that under the change of variable to its inverse, $x \to x^{-1}$, which converts low end to high end and vice versa, it transforms as \footnote{Notice that under such transformation $LN \to LN$, $Ga \to IGa$ and $IGa \to Ga$}
From definitions of Gini and Theil it can be then derived -- and immediately verified from the closed form answers in Tables (ref) - (ref) -- that the transformation property ((ref)) leads to $p+1 \leftrightarrow q$ transformation property for the inequality indices:
and
Notice that $p+1 \leftrightarrow q$ implies that $p>1$ and $q>2$ transform properly as well, which supports these conditions for BP. It is said that Theil L is more sensitive to inequality at low end -- due to the $ln(x)$, for $x$ in units of mean, which diverges as $x \to 0$ -- while Theil T is more sensitive to inequality at high end -- due to the $xln(x)$, which diverges as $x \to \infty$. We believe that ((ref)) is a more precise formulation of their roles.
We now turn to some limiting behaviors of Gini and Theil. As per Table (ref),
(Notice that this expression for $G^{BP}$ mcdonald2008modelling is equivalent to the one in Table (ref); the latter is written in a way to underscore $p+1 \leftrightarrow q$ symmetry.) Given the imposed constraints $p>1$ and $q>2$, the maximum value of Gini in BP is
Also notably
and, in particular,
When both parameters become large, Gini tends to zero as
The dependence of BP Gini on $p$ and $q$ is summarized in Fig. (ref). The slight asymmetry in dependences underscores the aforementioned $p+1 \leftrightarrow q$ symmetry and disappears for large $p$ and $q$, as per ((ref)) -- ((ref)). Finally, while ((ref)) is simple enough and can be easily tabulated, a much simpler expression
is within a fraction of a percentage point of exact value for relatively small $p$ and $q$ that are usually of interest, such as in Section (ref); ((ref)) also respects $p+1 \leftrightarrow q$ symmetry. The ratio of expression in ((ref)) to exact BP Gini in ((ref)) is shown in Fig. (ref).
We now turn to Theil. As per Table (ref),
Both Theil T and Theil L are shown in Fig. (ref). The mirror reflection between the two are as per ((ref)) and the slight asymmetry between $p$ and $q$ is per $p+1 \leftrightarrow q$, as was the case for Gini. Additionally,
and for large $p$ and $q$ we just note that
where $\gamma = - \psi(1) \approx 0.577$ is Euler's gamma. For $p,q \gg 1$, the results for $T^{BP}_L (p,q)$ are obtained via $p \leftrightarrow q$. Notice that ((ref)) indicates that Theil's dependence on $p$ and $q$ decouples, as shown in Fig. (ref).
We now turn to the properties of GB2 distribution, whose PDF is given by ((ref)). Its limiting behaviors are given by
and
Here we again assume that $\alpha p>1$, that is that BP has a maximum (a bell shape), and $\alpha q>2$, that is that the variance exists. Just as for BP, the important property of GB2 is that under the change of variable to its inverse, $x \to x^{-1}$, which converts low end to high end and vice versa, it transforms as
Sequentially, similarly to BP, we recognize the symmetry $p+\frac{1}{\alpha} \leftrightarrow q$, which leads to the following relationships for the Gini coefficient
and for Theil coefficients
where we again decoupled dependence on $p$ and $q$. While the $p+\frac{1}{\alpha} \leftrightarrow q$ symmetry for Hoover and between $T_T$ and $T_L$ is easily verified analytically from Table (ref) and ((ref)), we were able to verify one for Gini only numerically.
We now discuss the stochastic model of economic exchange that may be an underlying cause for the GB2 wealth/income distribution. \footnote{Recently, it was proposed that GB2 may also describe market volatility dashti2019implied, yan2019general, which makes this model even more relevant.} As before, we begin with the BP discussion, which is very transparent an easy to understand.
The mean-reverting stochastic differential equation (SDE), whose steady-state distribution is given by BP can be written as hertzler2003classical, dashti2018combined
where $\mathrm{d}W_t^{(2)}$ is the normally distributed Wiener process, $\mathrm{d}W_t^{(2)} \sim \mathrm{N(}0,\, \mathrm{d}t \mathrm{)}$. Stock market researchers will immediately recognize that for $\kappa_2=0$, ((ref)) reduces to the Heston model of stochastic volatility heston1993closed, dragulescu2002probability and for $\kappa_1=0$ to the multiplicative model of stochastic volatility nelson1990arch,fuentes2009universal,ma2014model. The $\kappa_1=0$ model has also been used in the Bouchaud-M\'ezard network model of economic exchange bouchaud2000wealth, ma2013distribution. We introduced the combined model ((ref)) dashti2018combined in order to "marry" the Heston behavior, which seem to work better for long accumulations of stock returns, and the multiplicative behavior, which seems to work better for daily or a few-days returns.
The steady-state distribution of ((ref)) is a BP
with the scale parameter
and the shape parameters
In the context of wealth/income distribution, the first term in the r.h.s. of ((ref)) postulates the convergence to the mean value $\theta$ over the time scale $\propto \gamma^{-1}$ due to a simple give-and-take economic exchange. The divergence from the mean $\theta$ towards low end and high end are exclusively due to stochastic term, which is $\propto x$ for large $x$ and are $\propto \sqrt{x}$ for small $x$. This term can be interpreted as fortunate and unfortunate events leading to gains and losses, such as inheritance, medical expenditures, bull and bear stock markets, etc. But they may also represent deviations from the mean due to work ethic, talent, etc. We underscore that eq. ((ref)) represents a time series for values $x$, which may not necessarily for a particular individual, but rather an abstract economic entity. The steady-state distributions represents, after the relaxation time, the distribution of values $x$ in this time series.
Consider now the following SDE
Its steady-state distribution is hertzler2003classical, dashti2019implied
with the scale parameter
and shape parameters
and
The steady-state distribution of ((ref)) is GIGa for $\kappa_{\alpha}=0$ and GGa for $\kappa_2=0$. For $\alpha=1$ we have mean-reverting models which yield a BP steady-state distribution in general and IGa and Ga for $\kappa_1=0$ and $\kappa_2=0$ respectively.
To understand the role of parameter $\alpha$, we analyze limiting behaviors ((ref)) and ((ref)) vis-a-vis ((ref)) using mean-reverting concept, even for $\alpha \neq 1$ and no fixed mean. Towards this end, we notice that from ((ref)) the tail exponent $q \alpha$ does not depend on $\alpha$ so the high end is not affected, according to ((ref)). On the other hand, front exponent $p \alpha$ grows with alpha, so on transition from $\alpha < 1$ to $\alpha > 1$ the lower end gets suppressed and mid-range gets more populated, according to ((ref)). This can be understood as follows: for $x \ll \beta$ in going from $\alpha < 1$ to $\alpha > 1$ we see the “reversion" to the value smaller than the “average" to larger than “average" per $ \theta x^{1-\alpha}$ in ((ref)), while volatility, including potential "losses," $\propto x^{2-\alpha}$, becomes smaller.\\
We now undertake a numerical analysis of wealth/income distribution based on the sale prices of homes in Hamilton County, Ohio, which includes Cincinnati metropolitan area of about 2.2 million people. The data is for 1970 - 2010 sales and altogether we had 124,203 data points. We adjusted for inflation with the Bureau of Labor Statistics (BLS) Consumer Price Index (CPI) data for Hamilton County, expressing prices alternatively in 1990 and 2010 constant dollars to find only very minor differences -- see below.
We use Maximum Likelihood Estimation (MLE) for fitting and the results of fitting are presented in Fig. (ref), based on 1990 constant dollars, and Tables (ref) and (ref), based on 1990 and 2010 constant dollars respectively. Additionally, we fit the tails of the Cumulative Density Function (CDF) -- assuming power-law tails -- and present the results in Fig. (ref), based on 1990 constant dollars, and Tables (ref) and (ref). For the actual fat-tailed distributions -- GB2, BP and GIGa -- the results of tail fits are compared with the tail parameters calculated from full distribution fits per Tables (ref) and (ref). In our simulations we always computed using both 1990 and 2010 constant dollars to verify the consistency of our fitting, but Tables (ref) and (ref) amply demonstrate that, with the obvious exception of quantities related to the scale parameter, there are only very minor differences in quantities that depend on shape parameters. For this reason, in Figs. (ref) and (ref) and in what follows we present only1990 adjusted data. \\
We argue that the unique property of Generalized Beta Prime (and Beta Prime), that makes it suitable for describing wealth/income and other distributions in natural and life sciences, is that it can mimic various behaviors, such as exponential, both for small and large variable. It also has an important property that the distribution of the inverse variable is also Generalized Beta Prime. In other words, small and large values of variables can be interchanged. Under such transformation, Generalized Beta Prime cleanly enforces the requirement on parameters to preserve bell shape and existence of the variance. Most importantly, however, is that it is a steady-state distribution of a simple stochastic model which can describe many phenomena, including modeling economic exchange and market volatility.
We investigated measures of income inequality, Gini, Hoover and Theil coefficients, in the Generalized Beta Prime framework and derived the relationships ((ref)) and ((ref)) (as well as ((ref)) and ((ref)) and connected them to the aforementioned property ((ref)) (and ((ref))) of transformation to the inverse variable. We also derived parameters of Generalized Beta Prime from those of the stochastic model and thoroughly analyzed how those parameters affect the distribution and the inequality measures. Additionally, we derived a simple but very precise approximation ((ref)) to the exact Beta Prime Gini. We also derived several new Gini, Hoover and Theil coefficients for distributions in Tables (ref) and (ref).
We argued that Gini and Theil T, by definition, exaggerate the influence power law dependencies of the distributions both for low end (underweigh) and especially for high end fat tails (overweigh). For this reason, we believe that Hoover (Pietra, Schultz) and Theil L are more appropriate measures in such cases, especially as far as fat tails are concerned. We also introduced a new scale-independent measure of inequality, DMMS, which estimates, via a distribution-agnostic procedure, the fraction of the low and high end wealth/income. We used Beta Prime to illustrate this measure. In numerical simulations, Hoover was consistently lower than Gini and DMMS smaller than Hoover (for all fitted distributions in the latter case). Likewise, Theil L was consistently lower than Theil T both for data and for all fat-tailed distribution fits.
We used Hamilton County, Ohio home sale prices as a proxy to wealth/income distribution. Given the very large data set, we used Maximum Likelihood Estimation to fit with a number of distributions. While Generalized Beta Prime provided the best fit, the absolute accuracy was not particularly high. Additionally, we fit the tails directly to query the correspondence between those fits and power-law exponents of full-distribution fits. We examined the effect of a tiny fraction of the highest sale prices and showed that cutting those from the distribution noticeably reduce the inequality indices and "fatness" of tails. Finally, we fitted the distribution of market values (asking prices) and found that their tails are considerably "fatter" than those of sale prices -- so much so that the theoretical variance does not exist. We argued that this is because market values are not described by a model of economic exchange and as such are "unphysical."