EconBase
← Back to paper

Rethinking Generalized Beta Family of Distributions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

37,483 characters · 6 sections · 48 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Rethinking Generalized Beta Family of Distributions

frontmatter\fntext[myfootnote]{[email removed]} \address[mymainaddress]{Department of Physics, University of Cincinnati, Cincinnati, Ohio 45221-0011} \begin{abstract} We approach the Generalized Beta (GB) family of distributions using a mean-reverting stochastic differential equation (SDE) for a power of the variable, whose steady-state (stationary) probability density function (PDF) is a modified GB (mGB) distribution. The SDE approach allows for a lucid explanation of Generalized Beta Prime (GB2) and Generalized Beta (GB1) limits of GB distribution and, further down, of Generalized Inverse Gamma (GIGa) and Generalized Gamma (GGa) limits, as well as describe the transition between the latter two. We provide an alternative form to the "traditional" GB PDF to underscore that a great deal of usefulness of GB distribution lies in its allowing a long-range power-law behavior to be ultimately terminated at a finite value. We derive the cumulative distribution function (CDF) of the "traditional" GB, which belongs to the family generated by the regularized beta function and is crucial for analysis of the tails of the distribution. We analyze fifty years of historical data on realized market volatility, specifically for S&P500, as a case study of the use of GB/mGB distributions and show that its behavior is consistent with that of negative Dragon Kings. \end{abstract} \begin{keyword} Generalized Beta Distribution \sep Stochastic Differential Equation \sep Steady-State Distribution \sep Power-Law Tails \sep Negative Dragon Kings \end{keyword}

Introduction

Generalized Beta family of distributions has been widely used in econometrics and econophysics, in particular for understanding of distributions of income and wealth mcdonald1984some, mcdonald1995generalization, chotikapanich2008modelling, chakrabarti2013econophysics, biewen2018econometrics\footnote{There exist numerous extensions of GB as well -- see e.g. gomez2018family}. Family members with scale-free, power-law tails may be of particular interest in this regard chotikapanich2018using, which prompted development of stochastic models of economic exchange that produced power-law-tailed IGa bouchaud2000wealth, GIGa ma2013distribution, and GB2 dashti2020stochastic as their steady-state distributions. These distributions have other numerous applications in various fields, such as physics sepplinen2012scaling, thiery2014loggamma, grange2017loggamma, mathematical finance praetz1972distribution, cox1985theory, nelson1990arch, heston1993closed, dragulescu2002probability, fuentes2009universal, ma2014model, and cognitive psychology dashti2020modeling, to name a few. Of particular importance to this work is that a very general yet simple model of stochastic volatility yields a GB2 steady-state distribution dashti2021combined.

The first main objective of this work was to generalize GB distribution itself, relative to the one introduced in mcdonald1995generalization, by removing the correlations between the two scale parameters. The main purpose of such generalization is to obtain a distribution which has long power-law tails above the first scale, only to be eventually terminated, via a drop-off to zero, at the second scale. The need for such distribution was motivated by our looking into a possibility of Dragon Kings (DK) sornette2009,sornette2012dragon in the realized volatility (RV) dashti2021realized. In particular, it is of great interest to glean into whether the most calamitous events in the stock market -- Savings and Loan Crisis, Tech Bubble, Great Recession and Covid Pandemic -- follow the Black Swan (BS) behavior and stay on the power-law tails or they deviate strongly higher and thus are DK. What we found instead liu2022dragon was behavior consistent with negative Dragon Kings (nDK) pisarenko2012robust, which are a strong drop-offs of the power-law tails, and are well described by GB-type distribution functions, mGB.

While the most general description of stochastic dynamic models for Generalized Beta family of distributions can be found in hertzler2003classical, the model for the top GB -- or for that matter the top Beta (B) distribution -- is notably absent. Therefore, deriving such model was the second main objective of this work. As it turns out, the "traditional" GB distribution -- as defined in mcdonald1995generalization and generalized here -- does not appear to be a steady-state solution of an SDE. However, we did identify an SDE whose steady-state distribution -- modified Beta (mB) -- is very similar in properties to the Beta (B) distribution. We then constructed an mGB from mB using a change of variable to the power of the variable. We also derived an SDE from the mB SDE via the same change of variable, which yielded yet another mGB as its steady-state distribution. Both mGB distributions have key features of the traditional GB, and its generalization here, but seem to describe RV distribution better. There is very little numerical difference between the two mGB distributions, except that the former has a far simpler analytical form than the latter.

This paper is organized as follows. In Sec. (ref) we discuss the generalization of the GB distribution that is cast in terms of two scale and three shape parameters, with one of the shape parameters representing a change of variable to the power of the variable, the power being the said shape parameter; this change of variable also signifies transition from B to GB distribution. We also derive CDF of GB distribution and show that it belongs to a class of incomplete-Beta-function-generated distributions eugene2002betanormal, jones2004families, cordeiro2009new, alexander2012generalized, alzaatrech2013newmethod, lemonte2013extended. In Sec.(ref) we derive two novel SDE that produce mB and two mGB distributions and explore their properties relative to the "traditional" GB distribution. We also develop an SDE framework of producing GB family hierarchy and explain the nature of transition between power-law-tailed distributions and those with exponential-like tails. In Sec. (ref) we investigate distributions of S&P realized volatility vis-a-vis fits by GB and mGB distributions and show that mGB provides a better fit and that termination of the distribution at a finite values is consistent with nDK behavior. We summarize our results in Sec. (ref).

Generalized Beta Distribution

Employing a "dimensionless" variable, we use the following form of the GB PDF:

equation[equation omitted — 339 chars of source]

where $\beta _1$ and $\beta _2$ are scale parameters and $\alpha$, $p$ and $q$ are shape parameters, all positive, $B(p,q)$ is the beta function and $x \leq \beta _1$. This transforms into the form used by McDonald mcdonald1995generalization with the substitution $\beta _1=\beta/(1-c)^{1/\alpha}$ and $\beta _2=\beta/c^{1/\alpha}$, $0\leq c \leq1$, however we do not find it necessary to impose such correlations between $\beta _1$ and $\beta _2$. We derive the following expressions for GB CDF and 1-CDF, the latter being important in investigating power-law dependencies:

equation[equation omitted — 226 chars of source]
equation[equation omitted — 192 chars of source]

where $I(y;p,q)=B(y;p,q)/B(p,q)$ and $B(y;p,q)$ are, respectively, the regularized and incomplete beta functions nist2022digital.

Eqs. ((ref))-((ref)) reproduce well-known GB hierarchies mcdonald1984some, mcdonald1995generalization, hertzler2003classical, chotikapanich2008modelling: GB1 and GB2 for $\beta_2\to +\infty$ and $\beta_1\to +\infty$ respectively, and further B1 and B2 (Beta Prime) for $\alpha=1$. We arrive at the latter two also by taking the limits in reverse order: first $\alpha=1$, which yields Beta (B) distribution and then $\beta_2\to +\infty$ or $\beta_1\to +\infty$. The further down limits of GGa and GIGa are best understood from analysis of an SDE that produces all the above distributions and will be discussed in Sec. (ref). We also observe the special nature of the shape parameter $\alpha$ in that GB, and all its family members, are easily generated form the B family ($\alpha=1$) by a simple change of variable

equation[equation omitted — 115 chars of source]

where $y=x^{\alpha}$ and

equation[equation omitted — 249 chars of source]

with the appropriate rescaling of $\beta_1$ and $\beta_2$. The same holds true also for the corresponding SDE in Sec. (ref), where the mean reversion SDE for variable $x$ produces B family of steady-state distributions, while the same mean-reversion SDE for variable $x^\alpha$ produces GB family. More precisely, the steady-state distribution of the corresponding SDE for the variable $x$ is a modified Beta distribution, mB, which is very similar to B, while the steady--state distribution for variable $x^\alpha$ is a modified GB, mGB. One can also obtain a second modified GB distribution via the change of variable ((ref)) in mB. Again, both mGB are very similar to GB. It should be also emphasized that below B/GB level al members of the respective families can be written in a traditional form - see footnote (ref) in Sec. (ref) below.

It should be noted that since the variable of the regularized beta changes between $0$ and $1$, GB CDF belongs to a class of distributions generated by a seed CDF eugene2002betanormal, jones2004families, cordeiro2009new, alexander2012generalized, alzaatrech2013newmethod, lemonte2013extended, see also srinivasa2010distance, which in this case is given by

equation[equation omitted — 198 chars of source]

In the limiting cases of GB1, $\beta_2 \rightarrow \infty$, and GB2, $\beta_1 \rightarrow \infty$, it reproduces the results of sepanski2007family. The corresponding PDF $f(x;\alpha,\beta _1,\beta _2)=F'(x;\alpha,\beta _1,\beta _2)$, GB PDF can be rewritten as eugene2002betanormal

equation[equation omitted — 109 chars of source]

which produces an alternative form of GB PDF,

equation[equation omitted — 322 chars of source]

equivalent to ((ref)). In particular, it should be mentioned that for $\alpha=1$ the seed distribution ((ref)) contains only scale parameters $\beta_1$ and $\beta_2$, so the effect of generating B distribution via ((ref)) and ((ref)) is the attainment of the shape parameters $p$ and $q$.

In what follows, we will be specifically interested in the $\beta_2\ll\beta_1$ circumstance since for $\beta_2 \ll x \ll \beta_1$ GB exhibits a power-law dependence,

equation[equation omitted — 170 chars of source]

which is also the power-law tail of GB2, but GB PDF eventually terminates at $\beta_1$, which may be the case in some of the possible negative Dragon Kings (nDK) phenomena pisarenko2012robust, such as realized market volatility, which we discuss in Sec. (ref).

Generalized Beta Family as Steady-State Distributions Stochastic Differential Equation

Invoking now an SDE approach to the GB family of distributions, we point out that -- with the exception of GB (and B) themselves -- all members of GB family of distributions where obtained as steady-state distributions of SDE by Hertzler hertzler2003classical. We, however, provide our own, considerably simplified, version of those results and, in addition, obtain the SDE yielding mGB.

First, to underscore the significance of the change of variable $y=x^\alpha$ in generalizing from B to GB family of distributions that was mentioned in Sec. (ref), we first consider the following SDE, which combines the multiplicative nelson1990arch, praetz1972distribution,fuentes2009universal,ma2014model and Heston (Cox-Ingersoll-Ross)cox1985theory,heston1993closed models of stochastic volatility dashti2021combined:

equation[equation omitted — 123 chars of source]

where $\mathrm{d}W_t$ is the Wiener process, $\mathrm{d}W_t \sim \mathrm{N(}0,\, \mathrm{d}t \mathrm{)}$. Its steady-state distribution is a modified B2 dashti2021combined, which can be easily derived using the standard Fokker-Planck formalism risken1996fokker,jacobs2010stochastic:

equation[equation omitted — 223 chars of source]

where

equation[equation omitted — 68 chars of source]
equation[equation omitted — 64 chars of source]

and

equation[equation omitted — 59 chars of source]

This PDF is normalizable for $q > 0$ and $p>0$.\footnote{ Notice, that in dashti2021combined we used the standard B2 PDF,

equation[equation omitted — 131 chars of source]

whence $q=1+\frac{2 \gamma}{\kappa_2^2}$. The reason behind this ambiguity in the definition of $q$ stems from the fact that $p$ and $q$ are independent at the B2/GB2 level in the steady-state PDF of the respective SDE, which is not the case for B/GB as will be shown below. Similar ambiguity with respect to the value of $q$ exists for B1/GB1 for the same reason as for B2/GB2 -- that $p$ and $q$ are independent at that level and, consequently, B1/GB1 PDF can be written either in standard or modified version.}

If we now consider the same mean-reverting process ((ref)) for $y=x^\alpha$, and define $\gamma^\prime=\gamma/\alpha$, $k^\prime=k/\alpha$ and $k_2^\prime=k_2/\alpha$, we obtain a stochastic process given by

equation[equation omitted — 175 chars of source]

whose steady-state distribution is a mGB2, given by (compare with GB2 in dashti2021combined, dashti2021realized, dashti2020stochastic)

equation[equation omitted — 177 chars of source]

where

equation[equation omitted — 96 chars of source]
equation[equation omitted — 114 chars of source]

and

equation[equation omitted — 104 chars of source]

Notice the obvious renormalization of $p$ and $q$ in going from mB2 to mGB2 (or B2 to GB2 dashti2021combined), unlike undergoing the change of variable in accordance with ((ref)). Nonetheless, this analysis confirms that, for simplicity, it is sufficient to obtain an SDE for mB, given that the SDE for mGB will then automatically follow via aforementioned change of variable $x \to x^{\alpha}$. Analytical renormalization of $p$ and $q$ is irrelevant for numerical analysis, since they are obtained from fitting.

The SDE formalism also makes analysis of limiting cases physically appealing. For instance, at the current GB2 level of hierarchy, substituting $\kappa=0$ directly into ((ref)) yields a GIGa (IGa for $\alpha=1$ bouchaud2000wealth) steady-state distribution ma2013distribution,ma2014model,dashti2021combined, while substituting $\kappa_2=0$ yields GGa (Ga for $\alpha=1$ dragulescu2002probability). ((ref)) and ((ref)) then also explain that those two cases correspond to $p \to \infty$ and $q \to \infty$ mcdonald1984some, mcdonald1995generalization, hertzler2003classical, chotikapanich2008modelling respectively, when starting from ((ref)). However, the explicit form of the steady-state distributions indicates that this just corresponds to the exponential-like decay for small values of the variable in the former case and for large values of the variable in the latter.

A generalization to an SDE for the modified B (mB) can be written as

equation[equation omitted — 200 chars of source]

or using

equation[equation omitted — 78 chars of source]

and rescaling, an alternative, simplified form of the SDE producing mB can be written as

equation[equation omitted — 170 chars of source]

But, as before, the SDE given by ((ref)) allows us to easily trace B-hierarchy: $\kappa_1=0$ yields the B2 family discussed above, with further $\kappa=0$ and $\kappa_2=0$ yielding IGa and Ga respectively; conversely, $\kappa_2=0$ yields the B1 family, with further $\kappa_1=0$ yielding Ga. As discussed above, this can be immediately upgraded to the GB hierarchy via the $x \to x^{\alpha}$ change of variable, with the result summarized as follows:

equation[equation omitted — 252 chars of source]

As was mentioned before, GB2/GB1 distributions can be written either in its standard or modified form, depending on the definition of $q$.

It is obvious form ((ref)) that (G)Ga is a tying link between (G)B2 and and (G)B1. This can be easily seen by considering the following SDE:

equation[equation omitted — 151 chars of source]

where $0 \le c \le 1$. For $c=0$ we have $\tilde{\kappa}=\kappa_1$ and B1 above; for $c=1/2$ we have G; for $c=1$ we have $\tilde{\kappa}=\kappa_2$ and B2 above. This simply means that as $c$ changes we observe a transition from a PDF defined on a finite interval, (G)B, to a PDF with a power-law tail, (G)B2, via a distribution with an exponential-like tail.

The PDF and the CDF of the steady-steady distribution obtained from ((ref)) are given respectively by

equation[equation omitted — 319 chars of source]

and

equation[equation omitted — 411 chars of source]

where

equation[equation omitted — 46 chars of source]

and

equation[equation omitted — 98 chars of source]

One version of the mGB PDF and the CDF related to GB can be obtained via the change of variable ((ref)) applied to ((ref)) and ((ref)) and are given respectively by

equation[equation omitted — 421 chars of source]

and

equation[equation omitted — 619 chars of source]

and also

equation[equation omitted — 593 chars of source]

The second term in ((ref)) is the difference between $F_{mGB}$ and $F_{GB}$ in ((ref)) and is zero at zero and at $\beta_1$. Due to this difference, the asymptotic behaviors of $F_{GB}$ and $F_{mGB}$ on approach to $\beta_1$, $x\rightarrow \beta_1$, develop a considerable contrast in the limit $\beta_2 \ll \beta_1$ of interest here:

equation[equation omitted — 312 chars of source]
equation[equation omitted — 432 chars of source]

that is $1-F_{mGB}$ drops off to zero ($F_{mGB}$ saturates to unity) faster than $1-F_{GB}$ due to the factor $\left(\frac{\beta _2}{\beta _1}\right)^\alpha$.

Another mGB distribution is obtained from the SDE for B with $x$ replaced by $x^\alpha$:

equation[equation omitted — 234 chars of source]

PDF and CDF of the steady-state solution are given respectively by

equation[equation omitted — 437 chars of source]

and

equation[equation omitted — 686 chars of source]

where

equation[equation omitted — 56 chars of source]
equation[equation omitted — 116 chars of source]

and $ \, _2F_1$ and $F_1$ are hypergeometric and Appell hypergeometric functions respectively nist2022digital.

Market Volatility as Possible Example of Application of Generalized Beta Distribution

Realized volatility $RV$ is the square root of realized variance, which is defined as follows

equation[equation omitted — 74 chars of source]

where

equation[equation omitted — 55 chars of source]

are daily returns and $S_{i}$ is the reference (closing) price on day $i$. This is an annualized value, where $252$ represents the number of trading days in a year. In particular, $n=1$ are the daily returns and $n=21$ are monthly returns (typical number of trading days in a month). Here we present results for $n=1, 2, 3, 5, 7, 9, 13, 17, 21$.

We fitted distributions of RV for S&P index from 1970 through May of 2021, which covers four major financial upheavals: Savings and Loan Crisis, Tech Bubble, Great Recession, and Covid Pandemic. In a companion manuscript liu2022dragon, we undertake a far more detailed numerical analysis, which includes actual time-series of RV identifying the far-end-tail points of the distribution, linear tail fitting of RV distribution on a log-log scale meant to test conformity to power-law behavior, and p-value test for DK pisarenko2012robust. Here, we limit our discussion to fitting RV distribution using ((ref))-((ref)) and ((ref))-((ref)), including the confidence intervals (CI) of these fits janczura2012black.

Fitting was conducted using Bayesian for PDF and Gradient Descendant for CDF and we find very small differences in the parameters of the distribution between the two techniques -- at most 8% for GB ((ref))-((ref)) and at most 3% for mGB ((ref))-((ref)). Additionally, the two-sample KS statistic between ((ref)) and ((ref)) is $0.0015 \ll 0.0169$, the latter being the standard value for the number of points in our sample, as per knuth1998art with alpha value being $0.05$. Given such negligible difference, relative simplicity of ((ref))-((ref)) vis-a-vis the actual solution ((ref))-((ref)) of the SDE ((ref)), qualifies the former as a near exact solution of the SDE and explains its choice for fitting.

Tables (ref) and (ref) list the parameters of GB and mGB distributions obtained from fitting the RV distribution, the corresponding Kolmogorov-Smirnov (KS) statistic, and the goodness-of-fit KS values for our RV sample size, as per table in massey1985kolmogorov with alpha value being $0.05$. Figs. (ref) and (ref) show plots of KS statistic and scale parameters $\beta_1$ and $\beta_2$ from the tables as a function of $n$. Figs. (ref)$-$(ref) show GB and mGB fits of RV data -- 1-CDF, or complimentary CDF (ccdf) -- with the corresponding 95% CI, on a log-log scale.

One noteworthy feature in those plots is that a relatively small subset of data points moves up from the straight line (on a log-log scale) portion of the GB and mGB tails, ((ref)), and outside their CI, which can be viewed as "potential DK." This feature is further analyzed in liu2022dragon using the p-value test pisarenko2012robust. However, the data invariably falls back in and indicates termination at final values, as intended to be explained by GB and mGB.

table[table omitted — 1,008 chars of source]
table[table omitted — 1,012 chars of source]
figure[figure omitted — 249 chars of source]
figure[figure omitted — 243 chars of source]
landscape\begin{figure}[htbp!] \caption{GB and mGB fits of RV, with respective CI, for $n=1$.} \end{figure} \begin{figure}[tb] \caption{GB and mGB fits of RV, with respective CI, for $n=2$.} \end{figure} \begin{figure}[tb] \caption{GB and mGB fits of RV, with respective CI, for $n=3$.} \end{figure} \begin{figure}[tb] \caption{GB and mGB fits of RV, with respective CI, for $n=5$.} \end{figure} \begin{figure}[tb] \caption{GB and mGB fits of RV, with respective CI, for $n=7$.} \end{figure} \begin{figure}[tb] \caption{GB and mGB fits of RV, with respective CI, for $n=9$.} \end{figure} \begin{figure}[tb] \caption{GB and mGB fits of RV, with respective CI, for $n=13$.} \end{figure} \begin{figure}[tb] \caption{GB and mGB fits of RV, with respective CI, for $n=17$.} \end{figure} \begin{figure}[tb] \caption{GB and mGB fits of RV, with respective CI, for $n=21$.} \end{figure}

Summary

Main results of this article can be summarized as follows:

itemize• We introduced an alternative form of Generalized Beta PDF ((ref)) and evaluated its CDF ((ref)), the latter belonging to a class of Incomplete Beta Function generated distributions. • We identified an SDE ((ref)), which produces a modified Generalized Beta distribution as its steady-state solution, whose PDF and CDF are given respectively by ((ref)) and ((ref)). • A similar yet simpler modified Generalized Beta distribution ((ref))-((ref)) was constructed via a change of variable ((ref)) from a modified Beta distribution ((ref))-((ref)), which is a steady-state distribution of the $\alpha=1$ SDE, ((ref)). • We showed that an expanded version of ((ref)), ((ref)) allows for a physically appealing explanation of the Generalized Beta hierarchy, ((ref)), as well as of natural link between (Generalized) Beta 1 and (Generalized) Beta 2 from ((ref)). • We argued that Generalized Beta distribution is particularly useful in the circumstance when the distribution is characterized by a long power-law tail followed by a drop-off at a finite value of the variable. This has been applied to the distribution of realized volatility of the S&P500 index, which exhibits behavior characteristic of “Negative Dragon Kings." Specifically, we found that power-law tails, leading potentially to "Black Swan" effects, are abruptly terminated, capping realized volatility at very large but finite values. One interesting nuance to the above observation is that a small subset of data in the tails showed upward deviation from the linear part of the tails (on the log-log scale) and outside the confidence intervals of GB and mGB fits (as well as that of linear fit, all confirmed using p-value test liu2022dragon), which can be viewed as potential to develop "Dragon Kings." However, it is always followed by a rapid dropdown, as seen in Figs. (ref) - (ref). The largest values of realized volatility of S&P500 were, as expected, due to the most virulent market calamities: Savings and Loan crisis, Tech Bubble, Financial Crisis and Covid Pandemic. However, at least for S&P 500 -- perhaps due to the broad spectrum of the index and the size and strength of companies it represents -- market mechanism seem to limit its price from below and thus to disallow emergence of “Dragon Kings" or even development of "Black Swans" beyond some limiting value (which is not to state that the upper limit could not move upward in the future). To reiterate, Generalized Beta seems to be well suited for such situations and in the particular case of realized volatility provides a very good overall fit of its distribution.

In the future, we will further examine applicability of Generalized Beta to phenomena bounded from above yet exhibiting a well-established power-law tail over the large portion of the distribution.

Acknowledgments

Portions of analytical calculations were conducted using Wolfram Mathematica. We wish to thank Jeffrey Mills for helpful discussions.