Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
150,959 characters · 12 sections · 177 citation commands
Robust Inference on Income Inequality: t-Statistic Based Approaches
Empirical analyses on income and wealth inequality and those in other fields in economics and finance often face the difficulty that the data is heterogeneous, heavy-tailed or correlated in some unknown fashion (see, among others, the discussion and reviews in
Importantly, many studies going back to V. Pareto indicate that income and wealth distributions are heavy-tailed and follow power laws
with the tail index $\zeta>0$ (see, among others, the discussion and reviews in PS, PS, Atkinson, Atkinson, Gabaix1, Gabaix1, Milanovic1, Milanovic1, Milanovic, AP, AP, APS, APS, Toda, Toda, IIW, IIW, GabaixIneq, GabaixIneq, Ibr, Ibr, TodaWang, TodaWang, and references therein).\footnote{As is well-known, heavy-tailedness and power law distributons are also exhibited by many other key variables in economics and finance, including financial returns, foreign exchange rates, insurance risks and losses from natural disasters, to name a few (see, among others, the reviews in EKM, Gabaix1, IIW, MFE, and references therein.}
Typical empirical results are $\zeta\in (1.5, 3)$ for income, and $\zeta\approx 1.5$ for wealth. Thus, the variance is infinite for wealth and may be infinite for income.\footnote{Importantly, the value of the tail index $\zeta$ in power law income or wealth distributions ((ref)) may be regarded as a measure of upper tail inequality (that is, among the rich), with smaller values of the tail index corresponding to larger inequality in the upper tails. This may be motivated by the fact that, in the case of Pareto distributions with $\zeta>1$ for income or wealth, where ((ref)) holds exactly for all values $x$ greater than a certain threshold $x_m,$ the Gini coefficient of inequality over the whole income/wealth distribution is equal to $1/ (2\zeta -1)$ and is thus decreasing in $\zeta$ (see also the discussion in Atkinson, Atkinson, GabaixIneq, GabaixIneq, that focuses on the analysis and estimation of the top income inequality measure $\eta=1/\zeta,$ Clara, Clara, and Ibr, Ibr). }
More generally, the tail index parameter $\zeta$ of power law distributions ((ref)) characterize the heaviness (the rate of decay) of its tails, with smaller values of $\zeta$ corresponding to more pronounced heavy-tailedness in the distributions, and vice versa. The tail index $\zeta$ governs the likelihood of observing outliers and extreme values of $X,$ e.g., very high income/wealth levels in the case of income and wealth distributions. It is further important as it governs existence of moments of the r.v. $X>0,$ with the moment $EX^p$ of order $p>0$ of $X$ being finite if and only if $\zeta>p.$ In particular, the second moment $EX^2$ of the r.v. $X$ is finite and its variance $Var(X)$ is defined if and only if $\zeta>2,$ and the first moment - the mean $EX$- of $X$ is finite if and only if $\zeta>1.$
Applicability of commonly used approaches to inference on inequality measures based on asymptotic normality becomes problematic under heavy-tailedness, heterogeneity and correlation in the data. For instance, sample inequality measures - estimators of measures of inequality like sample Gini coefficient - converge to non-Gaussian limits given by stable random variables (r.v.'s) under sufficiently pronounced heavy-tailedness with infinite second moments and variances (see Taleb and the discussion in Appendix B).\footnote{As is well-known, finiteness of variances for variables dealt with, such as economic and financial indicators like financial returns and exchange rates, is crucial for applicability of standard statistical and econometric approaches, including regression and least squares methods. Similarly, the problem of potentially infinite fourth moments of (economic and financial) variables and time series dealt with needs to be taken into account in applications of autocorrelation-based methods and related inference procedures in their analysis (see, among others, the discussion in Granger, Granger, EKM, EKM, cont2001empirical, cont2001empirical, Ch. 1 in IIW, IIW, and references therein).}
Provided normal convergence for sample inequality measures holds, asymptotic methods based on it often have poor finite sample properties under the problems of extreme values, outliers and heavy-tailedness in data.\footnote{More generally, poor finite sample properties are often observed for asymptotic methods based on normal convergence of estimators and consistent estimation of their limiting variances under heterogeneity and dependence in observations (e.g., inference approaches based on heteroskedasticity and autocorrelation consistent - HAC - and clustered standard errors, especially with data with pronounced autocorrelation, dependence and heterogeneity, see, among others, Andrews, Andrews, DL, DL, the discussion in Ph, Ph, IM2, IM2, IM1, canay, canay, esarey, esarey, and references therein).} Similar problems are also observed for bootstrap methods (see CF, CF, DF, DF). Bootstrap methods are also known to fail in heavy-tailed infinite variance settings (see the discussion in Section 5 in DF, DF, and references therein).
The problems with inference on inequality measures are discussed in detail in Dufour, Dufour1. These works also emphasize that reliable methods remain scarce for both the one-sample problem of inference on a single inequality index and the two-sample problem of testing for equality of and inference on the difference between two inequality indices. As discussed in Dufour, Dufour1, the latter problem is much more challenging than the former (see also IM1, IM1, for the discussion and the results on robust inference on equality of and the difference between two general parameters of interest under heterogeneity and dependence). Dufour propose permutation tests for the hypothesis of equality of two inequality measures from independent samples which outperforms other asymptotic and bootstrap methods available in the literature (see also canay, canay, for permutation tests of equality of two general parameters of interest under heterogeneity and clustered dependence). As discussed in Dufour1, the latter tests for the two-sample problem for inequality indices are limited to testing the equality of two inequality measures. In particular, they do not provide a way of making inference on a possibly non-zero difference between the two measures considered nor building a confidence interval for the difference.
Dufour1 propose Fieller-type methods for inference on the generalized entropy (GE) class of inequality measures. Among other results, the authors develop approaches to testing and construction of confidence intervals for any possibly non-zero difference between the inequality measures that can be used under independent samples of i.i.d. observations with possibly unequal sizes and equal-sized samples of i.i.d. observation with arbitrary dependence between the samples.
This paper focuses on applications of recently developed $t-$statistic approaches (see IM2, IM2, IM1, and also Ch. 3 in IIW, IIW) in robust inference on income and wealth inequality measures under the problems of heterogeneity, heavy-tailedness and possible dependence in observations. Following the approaches, in particular, a robust large sample test on equality or a non-zero difference of two parameters of interest (e.g., a test of equality of inequality measures in two regions or countries considered) is conducted as follows: The data in the two samples dealt with is partitioned into fixed numbers $q_1, q_2\ge 2$ (e.g., $q_1=q_2=2, 4, 8$) of groups, the parameters (inequality measures dealt with) are estimated for each group, and inference is based on a standard two-sample $t-$test with the resulting $q_1, q_2$ group estimators (see the next section). As follows from the results in IM2, IM1, robust $t-$statistic approaches result in valid inference under general conditions that group estimators of parameters of interest (e.g., inequality measures) considered weakly converge, at an arbitrary rate, to independent normal or scale mixtures of normal r.v.'s. These conditions are typically satisfied in empirical applications even under pronounced heavy-tailedness, heterogeneity and possible dependence in observations.\footnote{In particular, as discussed in IM2 and Ch. 3 in IIW, the asymptotic Gaussianity of group estimators of the parameters of interest typically follows from the same reasoning and holds under the same conditions as the asymptotic Gaussianity of their full-sample estimators.} The approaches proposed in the paper complement and compare favorably with other inference methods available in the literature, including computationally expensive bootstrap procedures and permutation-based inference methods. Importantly, the approaches proposed in the paper can be used in testing and construction of confidence intervals for any possibly non-zero difference between inequality measures under the problems of heterogeneity, heavy-tailedness and possible dependence in the data.
One should also emphasize wider range of applicability of $t-$statistic approaches to inference on inequality measures proposed in the paper as compared to other inference methods available in the literature, including those considered in Dufour, Dufour1. The inference approaches can be used in the case where observations (e.g., on income or wealth levels) in each of the samples considered are dependent among themselves - for instance, due to spatial or clustered dependence (see Conley and Bhat for a review of settings and methods of inference under spatial and clustered dependence, including complex stratified and clustered household surveys), common shocks affecting them (see AndrewsCommon and Hwang for a review of and inference using data with common shock dependence), or, in the case of time series or panel data on income or wealth levels, due to autocorrelation and dependence in observations over time. Further, in the case of testing for equality of inequality measures or inference on their difference in two populations using two samples of possibly dependent observations, as above, the $t-$statistic inference approaches may be used under an arbitrary dependence between the samples as well as under possibly unequal sample sizes.
Application of the robust inference approaches is illustrated by an empirical analysis of income inequality measures and their comparisons across different regions in Russia.
The paper is organized as follows. Section (ref) describes the robust $t-$statistic approaches to inference on inequality measures analyzed in the paper and discusses the conditions for their validity. Section (ref) provides numerical results on finite sample performance of the robust inference approaches dealt with and their comparisons with other inference methods in the literature, with a particular focus on testing equality of two inequality measures and inference on the difference between two inequality indices in Section (ref). Section (ref) presents empirical applications of the robust $t-$statistic approaches in the analysis of income inequality in Russia and comparisons of inequality measures across Russian regions. Section (ref) makes some concluding remarks and discusses some suggestions for future research. Appendix A provides tables on the numerical and empirical results in the paper. Appendix B provides a review of the definitions and asymptotic properties of the Gini, Theil and Generalized Entropy measures referred to in the paper and a discussion of applicability of the robust approaches dealt with in inference on the measures.
We focus on inference on inequality measures using the t-statistic approaches to robust inference under heterogeneity, heavy-tailedness and dependence of largely unknown form recently developed in IM2, IM1. IM2 provide an approach to robust inference on an arbitrary single parameter of interest. IM1 provide approaches to robust testing of equality of two arbitrary parameters of interest and to robust inference on the difference of the parameters.
We refer to, among others, Section 13.F in CF, DF, MO, Ibr, Dufour and Dufour1 for definitions of the most widely inequality measures, including Gini, Generalized Entropy and Theil indices, and their values for different income distributions, including empirically relevant heavy-tailed Pareto, double Pareto and Singh-Maddala distributions (see the next section).
In the context of a one-sample inference on a single (income or wealth) inequality (e.g., a Theil, Generalized entropy - GE - or Gini index) the robust $t-$statistic approaches are implemented as follows.
Throughout the paper, we denote by $T_k$ a r.v. that has a Student-$t$ distribution with $k\ge 1$ degrees of freedom. Further, for $q\ge 2$ and $0<\alpha<1,$ by $cv_{q, \alpha}$ we denote the $(1-\alpha/2)-$quantile of the Student-$t$ distribution with $q-1$ degrees of freedom: $P(\left|T_{q-1}\right|>t_{\alpha }$)= $\alpha$.
Consider the one-sample problem of testing a hypothesis on or constructing a confidence interval for an inequality measure $\mathcal{L}.$ Following the $t-$statistic robust inference approaches in IM2, a (large) sample I${}_{1}$, I${}_{2}$,..., I${}_{N}$ of observations on income or wealth levels $I,$ is partitioned into a fixed number $q\ge 2$ (e.g., $q=2, 4$ or 8) of groups, and the income inequality measure $\mathcal{L}$ is estimated using the data for each group thus resulting in $q$ group empirical income inequality measures ${\widehat{\mathcal{L}}}_j$, $j = 1,...,q.$ The robust test of the null hypothesis $H_0: \mathcal{L}={\mathcal{L}}_0$ against the two-sided alternative $H_a: \mathcal{L}\neq {\mathcal{L}}_0$ is based on the usual $t-$statistic $t^{I}_{\mathcal{L}}$ in the $q$ group empirical inequality measures ${\widehat{\mathcal{L}}}_j$, $j = 1,...,q:$
with $\overline{\widehat{\mathcal{L}}}=\frac{\sum^q_{j=1}{{\widehat{\mathcal{L}}}_j}}{q}$ and $s^2_{\widehat{\mathcal{L}}}=\frac{\sum^q_{j=1}{{\left({\widehat{\mathcal{L}}}_j-\overline{\widehat{\mathcal{L}}}\right)}^2}}{q-1}$ ${\widehat{\mathcal{L}}}_j.$ The above null hypothesis $H_0$ is rejected in favor of the alternative $H_a$ at level $\alpha\le 0.83$ (e.g., at the usual significance level $\alpha=0.05$) if the absolute value $|t_{\mathcal{L}}|$ of the $t-$statistic in group estimates ${\widehat{\mathcal{L}}}_j$ exceeds the $(1-\alpha/2)-$quantile of the standard Student-$t$ distribution with $q-1$ degrees of freedom: $|t_{\mathcal{L}}|>cv_{q, \alpha}.$ The test of $H_0$ against $H_a$ of level $\alpha\le 0.1$ is conducted in the same way if $2\le q\le 14.$ Using the results in bakirov2006student and IM2, one can further calculate the $p-$values of the above $t-$statistic robust tests in the case of an arbitrary number $q$ of groups thus enabling conducting robust tests on the inequality measure $\mathcal{L}$ of an arbitrary level.\footnote{One-sided tests are conducted in a similar way; one may note that quantiles of Student-$t$ distributions with $q-1$ degrees of freedom can also be used in one-sided tests of level $\alpha\le 0.1$ if $q\in \{2, 3\}.$}
By implication, for all $\alpha\le 0.83$ (and all $\alpha\le 0.1$ for $2\le q\le 14$) a confidence interval for the inequality measure $\mathcal{L}$ with asymptotic coverage of at least $1-\alpha$ may be constructed as ${\widehat{\mathcal{L}}}_j\pm cv_{q, \alpha} s_{\widehat{\mathcal{L}}}.$ For instance, the 95% confidence interval for $\mathcal{L}$ is given by $(\overline{\widehat{\mathcal{L}}}-cv_{q, 0.05}s_{\widehat{\mathcal{L}}}\mathrm{,\ \ }\overline{\widehat{\mathcal{L}}}+cv_{q, 0.05}s_{\widehat{\mathcal{L}}}$), where cv${}_{q, 0.05}$ is the 0.975-quantile of the Student-t distribution with q$-$1 degrees of freedom: $P(\left|T_{q-1}\right|>cv_{q, 0.05}$)=0.05.
As follows from IM2, the above approach results in asymptotically valid inference under the assumption that the group empirical income inequality measures ${\widehat{\mathcal{L}}}_j$, $j=1,..., q,$ are asymptotically independent, unbiased and Gaussian of possibly different variances.
The asymptotic validity of the t-statistic based inference approach continues to hold even when the group estimators ${\widehat{\mathcal{L}}}_j$ of $\mathcal{L}$ converge (at an arbitrary rate) to independent but potentially heterogeneous scale mixtures of normal r.v.'s, such as heavy-tailed stable symmetric r.v.'s. It also holds under convergence of the group estimators to conditionally normal r.v.'s which are unconditionally dependent through their second moments or have a common shock-type dependence (see AndrewsCommon, AndrewsCommon, for inference methods under common shock dependence structures, and Hwang, Hwang, for applications of $t-$statistic robust inference approaches in such settings).\footnote{Justification of asymptotic validity of the robust $t-$statistic inference approaches in IM2 is based on a small sample result in bakirov2006student that implies validity of the standard $t-$test on the mean under independent heterogeneous normal observations. Justification of asymptotic validity of the approaches in inference on equality of two parameters in IM1 is based on the analogues of the above small sample result for two-sample $t-$tests and Behrens-Fisher problem obtained therein.} This implies that the t-statistic based robust inference on $\mathcal{L}$ can thus be applied under extremes and outliers in observations generated by heavy-tailedness with infinite variances and, among others, dependence structures that include models with multiplicative common shocks (see Ibr1, Ibr1, Ibr2). The $t-$statistic based approaches do not require at all estimation of limiting variances of estimators of interest, in contrast to inference methods based on consistent, e.g., HAC or clustered, standard errors (see Section (ref)). The numerical analysis in IM0, IM1 and Section 3 in IIW indicates favorable finite sample performance of the $t-$statistic based robust inference approaches in inference on models with time series, panel, clustered and spatially correlated data. See also esarey for a detailed numerical analysis of finite sample performance of different inference procedures, including $t-$statistic and related approaches, under small number of clusters of dependent data and their software (STATA and R) implementation.
The above conditions for asymptotic validity of $t-$statistic approaches to robust inference are typically satisfied in applications, under the appropriate choice of the groups implying asymptotic unbiasedness and independence of group estimators of parameters of interest - inequality measures considered (see below). Namely, the asymptotic Gaussianity (or other weak convergence results, e.g., convergence to heavy-tailed scale mixtures of Gaussian distributions) of group estimators - group empirical inequality measures - ${\widehat{\mathcal{L}}}_j$ typically follows from the same reasoning and holds under the same conditions as the asymptotic Gaussianity (or other relevant asymptotics) of the full-sample estimator - full-sample empirical inequality measure - ${\widehat{\mathcal{L}}}.$
Concerning the choice of the groups, the condition that group estimators of parameters of interest - inequality measures dealt with - should be asymptotically unbiased (and independent) places natural - again typically satisfied in applications - restrictions on formation of groups in applications of $t-$statistic approaches in the context of inference on inequality indices and their comparisons (see also discussion of general $t-$statistic inference approaches in IM2, IM2). For instance, in the problem of inference on a single inequality measure in the whole country, e.g., Russia, using household income surveys with random samples of households in the country and its regions and thus i.i.d data on income levels, groups cannot be chose to be the country regions. This is because each of the group estimators - group inequality measures - will estimate the inequality index considered in the corresponding region but not in the whole country and unbiasedness of the group estimators with the mean asymptotically equal to the country's inequality index of interest will not hold. The “between-region” component of inequality in the whole country would be missed out by the group estimators.
On the other hand, in the problem of testing equality of or inference on the difference between inequality indices in two regions of a country, e.g., Russia as in the empirical application in Section (ref) in this paper, using household income surveys with i.i.d. data on household income levels in the regions considered, the groups in applications of two-sample $t-$statistic approaches can be formed just by taking subsequent observations on incomes in the two samples of i.i.d. income data in the regions (similar to applications of the approaches with time series data, see IM2, IM2). Namely, in the case of inference on equality of inequality indices of interest in two regions using the random samples $I_1, I_2, ..., I_{N_1},$ $Y_1, Y_2, ..., Y_{N_2}$ of (i.i.d.) income levels in them, the $q_1, q_2$ groups in applications of two-sample $t-$statistic approaches based on $\tilde{t}_{\mathcal{L}}$ in ((ref)) can be taken to be the groups $\{I_k, (i-1)N_1/q_1<k\le iN_1/q_1\},$ $\{Y_l, (j-1)N_2/q_2<l\le jN_2/q_2\},$ $i=1, ..., q_1,$ $j=1, ..., q_2,$ of subsequent observations on household incomes in the samples considered. The groups of subsequent observations in the two samples in applications of $t-$statistic approaches based on $\tilde{\tilde{t}}_{\mathcal{L}}$ in ((ref)) with $q_1=q_2=q$ are formed in a similar way. With the above simple choice of groups, asymptotic unbiasedness and independence of group estimators - group inequality measures - holds due to i.i.d.ness of data in the random samples considered.
Let us now turn to testing that the values $L_1$ and $L_2$ of an inequality measure $\mathcal{L}$ are equal in two populations (e.g., for income distributions in two regions of a country) and to inference on the difference $d=L_1-L_2$ between the two inequality indices using the (large) samples $I_1, I_2, ..., I_{N_1}$ and $Y_1, Y_2, ..., Y_{N_2}$ on income or wealth levels in the populations. We first assume that the two samples are independent. Following the $t-$statistic approaches to robust inference on two parameters of interest in IM1, each of the two samples $I_1, I_2, ..., I_{N_1}$ and $Y_1, Y_2, ..., Y_{N_2}$ is partitioned into fixed numbers $q_1, q_2\ge 2$ (e.g., $q_1, q_2=2, 4$ or 8) groups, respectively, and the income inequality measure $\mathcal{L}$ is estimated using the data for each of the groups in the two samples. This thus results in $q_1+q_2$ group empirical income inequality measures ${\widehat{\mathcal{L}}}^{I}_1,$ ..., ${\widehat{\mathcal{L}}}^{I}_{q_1},$ and ${\widehat{\mathcal{L}}}^{Y}_1$, ..., ${\widehat{\mathcal{L}}}^{Y}_{q_2}.$ The robust test of the null hypothesis $H_0: L_1-L_2=d_0$ (e.g., with $d_0=0,$ the test of the hypothesis $H_0: L_1=L_2$ of equality of the values $L_1, L_2$ of the inequality index $\mathcal{L}$ in the two populations considered) against the two-sided alternative $H_a: L_1-L_2\neq d_0$ (resp., with $d_0=0,$ against the two-sample alternative $H_a: L_1\neq L_2$) is based on the usual two-sample $t-$statistic $\tilde{t}_{\mathcal{L}}$ in the $q_1+q_2$ group empirical inequality measures ${\widehat{\mathcal{L}}}^{I}_j$, ${\widehat{\mathcal{L}}}^{Y}_k$, $j = 1,...,q_1,$ $k = 1,...,q_2:$
with $$\overline{\widehat{\mathcal{L}}^I}=\frac{1}{q_1}\sum^{q_1}_{j=1}{{\widehat{\mathcal{L}}}^I_j}, \overline{\widehat{\mathcal{L}}^Y}=\frac{1}{q_2}\sum^{q_2}_{k=1}{{\widehat{\mathcal{L}}}^Y_k},$$ $$ s^2_{\widehat{\mathcal{L}}^I}=\frac{1}{q_1-1}\sum^{q_1}_{j=1}{\left({\widehat{\mathcal{L}}}^I_j-\overline{\widehat{\mathcal{L}}^I}\right)^2}, s^2_{\widehat{\mathcal{L}}^Y}=\frac{1}{q_2-1}\sum^{q_2}_{k=1}{{\left({\widehat{\mathcal{L}}}^Y_k-\overline{\widehat{\mathcal{L}}^Y}\right)}^2}.$$
For the number of groups $q_1, q_2\le 14$ the above null hypothesis $H_0: L_1-L_2=d_0$ is rejected in favor of the alternative $H_a: L_1-L_2\neq d_0$ (resp., with $d_0=0,$ the null hypothesis $H_0: L_1=L_2$ of equality of the inequality measures values in the populations is rejected in favor of the alternative $H_a: L_1\neq L_2$) at level $\alpha \in \{0.001, 0.002, ... , 0.099, 0.10\}$ (e.g., at the usual significance levels $\alpha=0.01, 0.05$ and 0.1) if the absolute value $|\tilde{t}_{\mathcal{L}}|$ of the two-sample $t-$statistic in group empirical inequality measures ${\widehat{\mathcal{L}}}^{I}_j$, ${\widehat{\mathcal{L}}}^{Y}_k$, $j = 1,...,q_1,$ $k = 1,...,q_2,$ exceeds the $(1-\alpha/2)-$quantile of the standard Student-$t$ distribution with $q-1$ degrees of freedom, where $q=\min(q_1, q_2):$ $|\tilde{t}_{\mathcal{L}}|>cv_{q, \alpha}=cv_{\min(q_1, q_2), \alpha}.$ \footnote{As in the one-sample case, one-sided tests are conducted in a similar way.}\footnote{As follows from the analysis in IM1, the described tests may also be used for all $q_1, q_2\le 50$ if $\alpha \in \{0.001, 0.002, ... , 0.083\},$ e.g., for the usual critical values $\alpha=0.01, 0.05.$}
One further obtains that, for $\alpha=0.01, 0.05, 0.1$ and the number of groups $q_1, q_2\le 14,$ $\min(q_1, q_2)=q,$ a confidence interval for the difference $d_0=L_1-L_2$ between the values of the inequality measure $\mathcal{L}$ in two populations with asymptotic coverage of at least $1-\alpha$ may be constructed as \\ ${\widehat{\mathcal{L}}}^I-{\widehat{\mathcal{L}}}^Y\pm cv_{q, \alpha}\sqrt{s^2_{\widehat{\mathcal{L}}^I}/q_1+s^2_{\widehat{\mathcal{L}}^Y}/q_2}.$ For instance, the 95% confidence interval for $\mathcal{L}$ is given by ${\widehat{\mathcal{L}}}^I-{\widehat{\mathcal{L}}}^Y\pm cv_{q, 0.05}\sqrt{s^2_{\widehat{\mathcal{L}}^I}/q_1+s^2_{\widehat{\mathcal{L}}^Y}/q_2},$ where $cv_{q, 0.05}$ is the 0.975-quantile of the Student-t distribution with $\min(q_1, q_2)-1$ degrees of freedom: $P(\left|T_{\min(q_1, q_2)-1}\right|>cv_{q, 0.05}$)=0.05.
As follows from IM1, the two-sample $t-$statistic approach is asymptotically valid under the assumption - as above, typically satisfied in applications - that the group empirical income inequality measures ${\widehat{\mathcal{L}}}^I_j$, $j=1,..., q_1,$ ${\widehat{\mathcal{L}}}^Y_k$, $k=1,..., q_2,$ are asymptotically independent, unbiased and Gaussian of possibly different variances (or converge at an arbitrary rate to independent but potentially heterogeneous scale mixtures of normal r.v.'s, such as heavy-tailed stable symmetric r.v.'s).
Let us now consider the problem of testing for equality of the values $L_1$ and $L_2$ of an inequality measure $\mathcal{L}$ and to inference on the difference $d=L_1-L_2$ between the inequality indices in two populations using income or wealth level samples $I_1, I_2, ..., I_{N_1}$ and $Y_1, Y_2, ..., Y_{N_2}$ of possibly unequal sizes $N_1, N_2$ that may exhibit an arbitrary dependence between them. Suppose that the samples are divided into an equal number $q_1=q_2=q\ge 2$ (e.g., $q_1, q_2=2, 4$ or 8) of groups, and the sample inequality measures - estimates of the inequality index $\mathcal{L}$ of interest - are calculated using the data for each of the $2q$ groups in the two samples. One thus has the group empirical income inequality measures ${\widehat{\mathcal{L}}}^{I}_1,$ ..., ${\widehat{\mathcal{L}}}^{I}_{q},$ and ${\widehat{\mathcal{L}}}^{Y}_1$, ..., ${\widehat{\mathcal{L}}}^{Y}_{q}.$ The robust test of the null hypothesis $H_0: L_1-L_2=d_0$ (e.g., with $d_0=0,$ the test of the hypothesis $H_0: L_1=L_2$ of equality of the values $L_1, L_2$ of the inequality index $\mathcal{L}$ in the two populations) against the two-sided alternative $H_a: L_1-L_2\neq d_0$ (resp., with $d_0=0,$ against the two-sample alternative $H_a: L_1\neq L_2$) may be based on the one-sample $t-$statistic $\tilde{\tilde{t}}_{\mathcal{L}}$ in the $q$ differences ${\widehat{\mathcal{L}}}^{I}_j-{\widehat{\mathcal{L}}}^{Y}_j,$ $j = 1,...,q,$ of the calculated group empirical inequality measures:
with $$\overline{\widehat{\mathcal{L}}^I}=\frac{1}{q}\sum^{q}_{j=1}{{\widehat{\mathcal{L}}}^I_j}, \overline{\widehat{\mathcal{L}}^Y}=\frac{1}{q}\sum^{q}_{j=1}{{\widehat{\mathcal{L}}}^Y_j},$$ $$ s^2_{\widehat{\mathcal{L}}^{I-Y}}=\frac{1}{q-1}\sum^{q}_{j=1}{\left(({\widehat{\mathcal{L}}}^I_j-{\widehat{\mathcal{L}}}^Y_j)-(\overline{\widehat{\mathcal{L}}^I}-\overline{\widehat{\mathcal{L}}^Y})\right)^2}.$$
As in the case of $t-$statistic inference on one parameter, for any $\alpha\le 0.83$ (any $\alpha\le 0.1$ for $2\le q\le 14$), the null hypothesis $H_0: L_1-L_2=d_0$ (for $d_0=0,$ the null hypothesis $H_0: L_1=L_2$ of equality of the values $L_1, L_2$ of the inequality index $\mathcal{L}$ in two populations) is rejected in favor of the two-sided alternative $H_a: L_1-L_2\neq d_0$ (resp., for $d_0=0,$ in favor of the alternative $H_a: L_1\neq L_2$) at level $\alpha$ if the absolute value $|\tilde{\tilde{t}}_{\mathcal{L}}|$ of the $t-$statistic in the differences ${\widehat{\mathcal{L}}}^I_j-{\widehat{\mathcal{L}}}^Y_j,$ $j=1, ..., q$ of group sample inequality measures exceeds the $(1-\alpha/2)-$quantile of the standard Student-$t$ distribution with $q-1$ degrees of freedom: $|\tilde{\tilde{t}}_{\mathcal{L}}|>cv_{q, \alpha}.$ Further, as in the case of $t-$statistic inference on a single inequality measure, the $p-$values of the above tests can be calculated in the case of an arbitrary number $q=q_1=q_2$ of groups thus enabling conducting robust tests of an arbitrary level.
For all $\alpha\le 0.83$ (and all $\alpha\le 0.1$ for $2\le q\le 14$) a confidence interval for the difference $d_0=L_1-L_2$ between the values of the inequality measure $\mathcal{L}$ in two populations with asymptotic coverage of at least $1-\alpha$ may be constructed as ${\widehat{\mathcal{L}}}^I_j-{\widehat{\mathcal{L}}}^Y_j\pm cv_{q, \alpha} s_{\widehat{\mathcal{L}}^{I-Y}}.$ For instance, the 95% confidence interval for $\mathcal{L}$ is given by $({\widehat{\mathcal{L}}}^I_j-{\widehat{\mathcal{L}}}^Y_j- cv_{q, 0.05} s_{\widehat{\mathcal{L}}^{I-Y}}, {\widehat{\mathcal{L}}}^I_j+{\widehat{\mathcal{L}}}^Y_j- cv_{q, 0.05} s_{\widehat{\mathcal{L}}^{I-Y}}),$ where cv${}_{q, 0.05}$ is the 0.975-quantile of the Student-t distribution with q$-$1 degrees of freedom: $P(\left|T_{q-1}\right|>cv_{q, 0.05}$)=0.05.
As above, the $t-$statistic approaches to robust inference based on ((ref)) are asymptotically valid under the assumption that the group empirical income inequality measures ${\widehat{\mathcal{L}}}^I_j$, ${\widehat{\mathcal{L}}}^Y_j$, $j=1,..., q,$ are asymptotically independent across $j$, unbiased and Gaussian of possibly different variances (or have limiting scale mixtures of Gaussian distributions).
In this section, we provide numerical results on finite sample properties of the asymptotic, robust $t-$statistic, bootstrap and permutation approaches to inference and tests on inequality measures. The results are provided for inference on Theil and Gini inequality, similar to the numerical analysis in CF and Dufour.
We first present the results for the one-sample problem of inference on a single inequality measure in Section (ref).
Then, in Section (ref), the numerical results are provided for the two-sample problem of tests on equality of two inequality measures and inference on the difference between two inequality indices.
As in Dufour, the numerical analysis of finite-sample performance of different approaches to inference on inequality measure(s) is based on simulations from Singh-Maddala distributions that were reported to provide a good fit to real-world income distributions in various countries (see the discussion in CF, CF, DF, DF, Dufour, Dufour, and references therein). The cdf of the Singh-Maddala distribution with the scale parameter $b>0$ and the shape parameters $a, c>0$ is given by
(see the above references). Similar to Dufour, the Singh-Madalla distribution with parameters $a, b, c>0$ is denoted by $SM(a, b, c)$ in what follows. It is easy to see that the cdf $F(x)$ of the Singkh-Maddala distribution $SM(a, b, c)$ satisfies $F(x)\sim c\left(\frac{x} {b}\right)^a$ as $x\rightarrow 0,$ and $1-F(x)\sim \left(\frac{x}{b}\right)^{-ac}$ as $x\rightarrow \infty.$ \footnote{As usual, we write $f(x)\sim g(x)$ as $x\rightarrow x_0$ or $x\rightarrow \infty$ for two positive functions $f(x)$ and $g(x)$ if $f(x)/g(x)\rightarrow 1$ as $x\rightarrow x_0$ or $x\rightarrow \infty$.} Therefore, the Singh-Maddala distribution has the (double) power law or (double) Pareto behavior in the lower tails - for small income levels - and the upper tails - for high incomes (see Toda, Toda, for the analysis of double Pareto and related distributions for income).
In particular, the Singh-Maddala distributions belong to the class of distributions with heavy power law tails, so that for for large $x>0$ and r.v.'s (income or wealth levels) $X>0$ with the Singh-Maddala distribution $SM(a, b, c)$ follows power law ((ref)) with the tail index $\zeta=ac.$
Following Dufour, in the numerical experiments, we use the parameter values $a_0=2.8,$ $b_0=100^{-1/2.8},$ $c_0=1.7$ for the Singh-Maddala distribution, with the corresponding tail index $\zeta=a_0c_0=4.76,$ as a benchmark. The Theil index for this distribition equals to 0.1401151, and the Gini index equals to 0.2887138 (see Dufour, Dufour). The Singh-Maddala distribution with the above values for the parameters was also used in CF and DF to demonstrate poor finite-sample performance of asymptotic and bootstrap inference approaches.
Further, as in Dufour, in the numerical experiments, we consider several other Singh-Maddala distributions $SM(a, b_0, c)$ with the above scale parameter $b_0=100^{-1/2.8}$ for which the Theil inequality index and the Gini index are the same as in the case of the distribution $SM(a_0, b_0, c_0),$ and equal to 0.1401151 (the Theil index) and 0.2887138 (the Gini index).
Following Dufour, in simulations involving the Theil index, we take the parameters $(a, c)$ of the Singh-Maddala distributions $SM(a, b_0, c)$ equal to $(2.5, 2.502199),$ $ (2.6, 2.149747),$ $(2.7, 1.894309),$ $(2.8, 1.7),$ $(3.0, 1.4223847),$ $(3.2, 1.2320215),$ $(3.4, 1.0922125),$ $(3.8, 0.8984488),$ \\ $(4.8, 0.6366578)$ and $(5.8, 0.4996163).$ The corresponding tail indices $\zeta$ of these distributions equal to $\zeta=6.26, 5.59, 5.11, 4.76, 4.27, 3.94, 3.71, 3.41, 3.06, 2.9.$
In simulations involving the Gini index, we take, as in Dufour, the parameters $(a, c)$ of the Singh-Maddala distributions $SM(a, b_0, c)$ equal to (2.5,2.640350), (2.6,2.218091), (2.7,1.920967), (2.8,1.7), (3.0,1.3921126), (3.2,1.1866026), (3.4,1.0388049), (3.8,0.8387663), (4.8,0.5784599) and \\ (5.8,0.4473111). The corresponding tail indices $\zeta$ of these distributions equal to $\zeta=$6.6, 5.77, 5.19, 4.76, 4.18, 3.80, 3.53, 3.19, 2.78, 2.59.
We note that the tail index $\zeta=2.78, 2.59, 2.9$ for the considered Singh-Maddala distributions lie in the interval (1.5, 3) as is typically the case for real-world income distributions, as discussed above. Additionally, we also consider more heavy-tailed distributions $SM(a, b_0, c)$ with $(a, c)=(2,1.1),$ $(2,0.7)$ and $b_0=100^{-1/2.8}$. The corresponding tail indices $\zeta$ in power laws ((ref)) for these distributions equal to $\zeta=2.2, 1.4.$
We follow the notation in the previous sections that is largely similar to CF, DF and Dufour. In the numerical analysis in this section, $\mathcal{L}_0=\mathcal{L}(F)$ denotes the true value of the inequality measure $\mathcal{L}$ of interest (e.g., the Theil or Gini inequality index) for a population with the cdf $F$ considered. As before, $\hat{\mathcal{L}}=\hat{\mathcal{L}}(I_1, ..., I_N)$ denotes the full-sample estimator of $\mathcal{L}$ (the full-sample empirical inequality measure) calculated using a sample of observations $I_1, ..., I_N$ from the population. Further, as in Section (ref), $\hat{\mathcal{L}}_j,$ $j=1, ..., q,$ denote the group estimators of $\mathcal{L}$ (group empirical inequality measures) in applications of $t-$statistic inference approaches.
Asymptotic approaches to inference on an inequality measure $\mathcal{L}$ are based on normal approximations to sample distributions of full-sample estimators $\hat{\mathcal{L}}$ of the measures (full-sample empirical inequality measures), more precisely, on standard normal approximations to sample distributions of (full-sample) $t-$statistics $S_{\hat{\mathcal{L}}}=(\hat{\mathcal{L}}-\mathcal{L}_0)/s.e._{\hat{\mathcal{L}}}$ calculated using these estimators, where $s.e._{\mathcal{L}}$ denotes the usual consistent standard error of $\hat{\mathcal{L}}$ (see the formulas for the empirical inequality measures considered and their standard errors in CF, CF, DF, DF, and Dufour, Dufour). As discussed in Section (ref), validity of $t-$statistic robust inference approaches requires weak convergence of group estimators $\hat{\mathcal{L}}_j,$ $j=1, ..., q,$ of the inequality measures $\mathcal{L}$ (without any Studentization/normalization of the group estimators by their standard errors in contrast to the $t-$statistics $S_{\hat{\mathcal{L}}}$ calculated using the full-sample estimators) to possibly heterogeneous Gaussian distributions (or scale mixtures of Gaussian distributions). Further (see the discussion in the introduction and the Section (ref)), asymptotic normality of group estimators $\hat{\mathcal{L}}_j$ holds under the same conditions as in the case of the full-sample estimators $\hat{\mathcal{L}}.$
We, therefore, begin the analysis with an assessment of finite-sample distributions of (full-sample) empirical inequality measures $\hat{\mathcal{L}}$ and the (full-sample) $t-$statistics $S_{\hat{\mathcal{L}}}$ calculated using them. We, in particular, focus on the assessment of closeness of the above finite-sample distributions to Gaussian ones.
We focus on comparisons of finite-sample distributions of the (full-sample) $t-$statistics $S_{\hat{\mathcal{L}}}$ with those of the centered empirical inequality measures normalized by their true standard deviations, that is, of the statistics $Z_{\hat{\mathcal{L}}}=(\hat{\mathcal{L}}-\mathcal{L}_0)/\sigma_{\hat{\mathcal{L}}},$ where $\sigma_{\hat{\mathcal{L}}}^2=Var(\hat{\mathcal{L}}).$ The true values of the standard deviations $\sigma_{\hat{\mathcal{L}}}$ for the populations and sample sizes considered are obtained using direct simulations.
Figures (ref)-(ref) provide kernel estimates of densities of the finite-sample distributions of the statistics $Z_{\hat{\mathcal{L}}}$ and $S_{\hat{\mathcal{L}}}$ for different population distributions and sample sizes.\footnote{The number of replications in all simulation experiments is equal to 100,000.}
Figures (ref) and (ref) provide kernel density functions of the statistics $Z_{\hat{\mathcal{L}}}$ (sample sizes $N=50, 100, 1000$) and $S_{\hat{\mathcal{L}}}$ (sample size $N=100$)\footnote{Qualitatively similar results for other sample sizes $N$ are omitted for brevity and available on request.} for, respectively, the empirical Theil and Gini inequality measures in the case of samples from the Singh-Maddala distributions $SM(a_0, b_0, c_0)$ with the parameters $a_0=2.8,$ $b_0=100^{-1/2.8},$ $c_0=1.7$ and the corresponding tail index $\zeta=4.76,$ discussed before.
In the case of the Theil measure in Figure (ref), we observe some non-Gaussianity in the distribution of the statistics $Z_{\hat{\mathcal{L}}}$ and $S_{\hat{\mathcal{L}}}$ in small and moderate samples. In addition, the density of the $t-$statistic $S_{\hat{\mathcal{L}}}$ for Theil index is considerably (left) skewed in comparison to the densities of the statistic $Z_{\hat{\mathcal{L}}}$.
In the case of the Gini measure in Figure (ref), we can see that the distribution of the statistic $Z_{\hat{\mathcal{L}}}$ is very close to the standard normal even in small samples. In contrast, the distribution of the $t-$statistic $S$ is again skewed towards left.
For Singh-Maddala distributions with heavier tails as in the case of the parameters $(a, c)=(5.8,0.4473111)$ and the corresponding tail index $\zeta=2.59$ in Figure (ref), the finite sample distributions of the statistics $Z_{\hat{\mathcal{L}}}$ and $S_{\hat{\mathcal{L}}}$ for the Gini measure become more skewed (the same is observed for the Theil measure; the results are omitted for brevity and available on request). Skewness is especially pronounced in the case of small samples and the $t-$statistic $S.$
Overall, according to Figures (ref)-(ref), normal approximation appears to perform better for finite-sample distributions of the statistic $Z_{\hat{\mathcal{L}}}$ as compared to those of the full-sample $t-$statistic $S_{\hat{\mathcal{L}}}$ used in asymptotic tests and inference. We further note that the group estimators $\hat{\mathcal L}_j-\mathcal{L}_0$ used in $t-$statistic robust inference approaches are just scaled versions of the statistics $Z_{\hat{\mathcal{L}}}$ calculated using observations in the groups considered. Therefore, the above comparisons are expected to translate into better finite-sample performance of $t-$statistic inference approaches as compared to the asymptotic ones, provided that the number of observations in each of the groups in $t-$statistic approaches in sufficiently large, e.g., greater than 100 (this is usually the case in empirical applications with the number of groups $q=2, 4, 8$). For better size control, the number of groups, $q$, should be chosen to be smaller if the total sample size $N$ is not very large.
Table (ref) provides the results on the empirical size of asymptotic and the $t-$statistic robust tests on the Theil and Gini measures for sample sizes $N=200, 500, 1000$ and Singh-Maddala distributions $SM(a, b_0, c)$ with the parameters $(a, c)=(2.5,2.502199), (3.2,1.2320215), (5.8,0.4996163)$ corresponding to the tail indices $\zeta=6.6, 3.94, 2.9$ in the case of the Theil index and the parameters $(a, c)=(2.5,2.640350), (3.2,1.1866026)$, and $(5.8,0.4473111)$ corresponding to $\zeta=6.26, 3.8, 2.9$ in the case of the Gini index.
In accordance with the above discussion of finite-sample distributions of statistics $Z$ and $S$ like those in Figures (ref)-(ref), the results in the table show that the finite-sample size of both the asymptotic and $t-$statistic robust tests becomes more distorted if the tail index decreases and thus the degree of heavy-tailedness in observations becomes more pronounced. Importantly, however, size distortions for the Gini measure are not so large as for the Theil measure. In the case of the number of groups $q=4$ or $q=8,$ the finite-sample size properties of robust tests based on $t-$statistics in group estimates are usually better than those of the tests based on asymptotic normality of the (full-sample) $t-$statistics for the measures. Further, the finite sample size properties of the robust $t-$statistic-based tests (with $q=4$) appear to be better than those of the asymptotic tests in the cases where each of the groups contains more than 100 observations, in accordance with the discussion of Figures (ref)-(ref).
In this section, we focus on comparisons of the finite-sample performance of the two-sample $t-$statistic robust inference approaches discussed in Section (ref) with permutation and bootstrap tests proposed by Dufour.
In the numerical analysis in this section, $\mathcal{L}^{I}=\mathcal{L}(F_1)$ and $\mathcal{L}^Y=\mathcal{L}(F_2)$ denote the true values of the inequality measure $\mathcal{L}$ of interest (e.g., the Theil or Gini inequality index, as in the previous section) in two populations with cdf's $F_1$ and $F_2$ considered. By $\hat{\mathcal{L}}^I=\hat{\mathcal{L}}^I(I_1, ..., I_{N_1})$ and $\hat{\mathcal{L}}^Y=\hat{\mathcal{L}}^Y(Y_1, ..., Y_{N_2})$ we denote the full-sample estimators of the measure $\mathcal{L}$ (the full-sample empirical inequality measures) calculated using samples of observations $I_1, ..., I_{N_1}$ and $Y_1, ..., Y_{N_2}$ from the two populations. Further, as in Section (ref), by $\hat{\mathcal{L}}_1^I, ..., \hat{\mathcal{L}}_{q_1}^I$ and $\hat{\mathcal{L}}_1^Y, ..., \hat{\mathcal{L}}_{q_2}^Y$ we denote the group estimators of $\mathcal{L}$ (group empirical inequality measures) in the two samples dealt with in applications of $t-$statistic inference approaches.
Asymptotic approaches to testing the hypothesis $H_0: L_1-L_2=d_0$ (e.g., with $d_0=0,$ testing the hypothesis $H_0: L_1=L_2$ of equality of the values $L_1, L_2$ of the measure $\mathcal{L}$ in the two populations considered) against the two-sided alternative $H_a: L_1-L_2\neq d_0$ (resp., with $d_0=0,$ against the two-sample alternative $H_a: L_1\neq L_2$) are based on the normal approximation to the sample distribution of the difference $\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y$ between the full-sample estimators $\hat{\mathcal{L}}^I$ and $\hat{\mathcal{L}}^Y$ (full-sample empirical inequality measures). More precisely, the asymptotic approaches are based on the standard normal approximation to the sample distribution of the two-sample $t-$statistic $S_{\hat{\mathcal{L}}_1-\hat{\mathcal{L}}_2}=(\hat{\mathcal{L}}_1-\hat{\mathcal{L}}_2-d_0)/s.e._{\hat{\mathcal{L}}_1-\hat{\mathcal{L}}_2}$ calculated using the estimators $\hat{\mathcal{L}}_1$ and $\hat{\mathcal{L}}_2,$ where $s.e._{\hat{\mathcal{L}}_1-\hat{\mathcal{L}}_2}$ denotes the usual consistent standard error of the difference $\hat{\mathcal{L}}_1-\hat{\mathcal{L}}_2$ (see the formulas for the empirical inequality measures considered and the standard errors in CF, CF, DF, DF, Dufour, Dufour).
Similar to the previous section, validity of two-sample $t-$statistic robust inference approaches based on ((ref)) requires weak convergence of group estimators $\hat{\mathcal{L}}_j^I,$ $j=1, ..., q_1,$ $\hat{\mathcal{L}}_k^Y,$ $k=1, ..., q_2,$ of the inequality measure $\mathcal{L}$ in the two samples considered to possibly heterogeneous Gaussian distributions (or scale mixtures of Gaussian distributions). As discussed in the introduction and the previous section, asymptotic normality of group estimators $\hat{\mathcal{L}}_j^I,$ $\hat{\mathcal{L}}_k^Y,$ hold under the same conditions as in the case of the full-sample estimators $\hat{\mathcal{L}}^I$ and $\hat{\mathcal{L}}^Y.$ We refer to the previous section for the assessment of finite-sample distributions of the full-sample empirical inequality measures and their closeness to normality.
On the other hand, with $q_1=q_2=q,$ validity of the (two-sample) $t-$statistic robust inference approaches based on ((ref)) - that is, the one-sample $t-$statistic $\tilde{\tilde{t}}_{\mathcal{L}}$ in the $q$ differences ${\widehat{\mathcal{L}}}^{I}_j-{\widehat{\mathcal{L}}}^{Y}_j,$ $j = 1,...,q,$ of the group empirical inequality measures $\hat{\mathcal{L}}_j^I,$ $\hat{\mathcal{L}}_j^Y$ (without any Studentization/normalization of the differences between the group estimators by their standard errors in contrast to the $t-$statistics $S_{\hat{\mathcal{L}}_1-\hat{\mathcal{L}}_2}$ calculated using the full-sample estimators) requires weak convergence of the differences ${\widehat{\mathcal{L}}}^{I}_j-{\widehat{\mathcal{L}}}^{Y}_j,$ $j = 1,...,q,$ of the group estimators to possibly heterogeneous Gaussian distributions (or scale mixtures of Gaussian distributions). Further, asymptotic normality of the differences $\hat{\mathcal{L}}_j^I-\hat{\mathcal{L}}_j^Y$ between the group estimators holds under the same conditions as in the case of the difference $\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y$ between the full-sample estimators.
We, therefore, begin the analysis with an assessment of finite-sample distributions of the difference $\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y$ between the (full-sample) empirical inequality measures $\hat{\mathcal{L}}^I,$ $\hat{\mathcal{L}}^Y$ and of the (full-sample) $t-$statistics $S_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}$ calculated using them. We, in particular, focus on the assessment of closeness of the above finite-sample distributions to Gaussian ones.
We focus on comparisons of finite-sample distributions of the $t-$statistic $S_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}$ with those of the difference $\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y$ normalized by its true standard deviation, that is, of the statistic $Z_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}=(\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y)/\sigma_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y},$ where $\sigma^2_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}=Var(\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y).$
In Figures (ref)-(ref), we present kernel estimates of the finite-sample densities of the statistics $Z_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}$ and $S_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}$ for the difference between the empirical inequality measures in two samples from populations with the same Singh-Maddala distribution.
Figures (ref)-(ref) provide kernel density functions, for sample sizes $N_1=N_2=N,$ of the statistics $Z_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}$ ($N=50, 100, 1000$) and $S_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}$ (sample sizes $N=100$)\footnote{Qualitatively similar results for other sample sizes $N$ are omitted for brevity and available on request.} for the difference between, respectively, the Theil and Gini empirical inequality measures in two samples from the Singh-Maddala distribution $SM(a_0, b_0, c_0)$ with the parameters $a_0=2.8,$ $b_0=100^{-1/2.8},$ $c_0=1.7$ and the corresponding tail index $\zeta=4.76.$ Figure (ref) provides the above kernel density functions for the Gini index in the case of two samples from a more heavy-tailed Singh-Maddala distribiion $SM(a, b_0, c)$ with $(a, c)=(5.8,0.4473111)$ and the tail index $\zeta=2.59.$
According to Figures (ref)-(ref), the finite-sample distributions of the statistic $Z_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}$ and thus of the difference $\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y$ between the estimators of the Theil and Gini measures is approximately symmetric even in rather small samples and also under pronounced heavy-tailedness, with good performance of Gaussian approximations, e.g., as compared to finite-sample distribution of the (full-sample) $t-$statistic $S_{\hat{\mathcal{L}}^I-\hat{\mathcal{L}}^Y}.$ This also holds in the case when the sample sizes are not very different. \footnote{If the sample sizes of two groups are very different, then different partition, $q_1,q_2$ should be used in applications of the $t$-statistic inference approaches.}
Tables (ref)-(ref) provide the results on the finite-sample size properties of the asymptotic, permutation, bootstrap and $t-$statistic robust tests on equality of Theil and Gini measures. As before, we consider two samples $I_1, ..., I_{N_1}$ and $Y_1, ..., Y_{N_2}$ from, respectively, Singh-Maddala distributions $SM(a_I, b_0, c_I)$ and $SM(a_Y, b_0, c_Y),$ with $b_0=100^{-1/2.8}$ and the tail indices $\zeta_I=a_Ic_I,$ $\zeta_Y=a_Yc_Y.$ In simulations, we consider the following settings with identical/different sample sizes $N_1,$ $N_2;$ distributions $S(a_I, b_0, c_I)$ and $S(a_Y, b_0, c_Y)$ in the samples and the number $q_1, q_2$ of groups used in $t-$statistic robust tests.
Tables (ref)-(ref) provide the results on finite-sample size properties of asymptotic, permutation, bootstrap and $t-$statistic robust tests based on ((ref)) and ((ref)) with the equal number of groups $q_1=q_2=q.$
The results in Table (ref) indicate that the size of all the tests, except the asymptotic ones, never exceeds the nominal 5% level. In addition, in a number of cases, it is quite close to the nominal level for the permutation, bootstrap and the robust $t-$statistic tests.
According to Table (ref), the empirical sizes of the robust two-sample $t-$statistic tests based on ((ref)) with $q=4, 8$ in the case of more heavy-tailed distributions and on ((ref)) with $q=4, 8, 12, 16$ in the case of less heavy-tailed distributions are comparable and in some cases are better than those of the permutation and bootstrap tests. Comparing the robust tests based on the two-sample $t-$statistic ((ref)) in group estimators and those based on the one-sample $t-$statistic ((ref)) in the differences of the group estimators, overall, the former tests with the same number of groups $q_1=q_2=q$ appear to have less over-rejections as compared to the latter ones.
The finite-sample size of all the tests, except the asymptotic ones, appears to be good in all parameter settings: e.g., essentially no over-rejections are observed for $t-$statistic inference approaches, including the settings with more pronounced heavy-tailedness and infinite variances in Table (ref). Further, the finite sample properties of the robust $t-$statistic approaches are comparable or similar and in some cases are better than those of the bootstrap and permutation approaches. Importantly, asymptotic normality of sample Theil and Gini measures is lost under infinite variances, as is the case for tail indices $\zeta_I=\zeta_Y=1.4$ in Table (ref) (see Taleb and the discussion in Appendix B). However, according to the results in the table, the $t-$statistic approaches have good finite sample size properties even in such heavy-tailed settings. This is due to robustness of $t-$statistic approaches to heavy-tailedness as they may be used under convergence of group estimators of parameters in consideration to scale mixtures of normals.
In Table (ref), one can observe better size properties for two-sample $t-$statistic inference approaches based on $\tilde{t}_{\mathcal{L}}$ in ((ref)) (with $q_1=q_2=4,8,12$ for all sample sizes and also $q=16$ for large sample sizes) and those based on $\tilde{\tilde{t}}_{\mathcal{L}}$ in ((ref)) (with $q_1=q_2=4,8$ for all sample sizes) in comparison to permutation and bootstrap tests.
Table (ref) is an analogue of Tables (ref) and (ref) with different numbers $q_1, q_2$ of groups used in two-sample $t-$statistic robust inference approaches based on $\tilde{t}_{\mathcal{L}}$ in ((ref)). It provides the results on the empirical size of the tests based on these approaches with asymptotic, bootstrap and permutation tests in the following settings.
According to the results in Table (ref), in the case of two-sample $t-$statistic inference on equality of/the difference between Theil indices, only the choice of $q_1=q_2=4$ leads to size control for all sample sizes considered. Size distortion of the $t-$statistic approaches in the case of Theil indices is apparently due to skewness in finite-sample distributions of (group) empirical Theil inequality measures implying poor quality of normal approximations to them (see Section (ref)). The solution may be to use different number of groups $q_1, q_2$ for different sample size pairs $N_1, N_2$. E.g., according to Table (ref), in the case of inference on equality/the difference between the Theil indices, good size properties of two-sample $t-$statistic approaches are observed with $(q_1, q_2)=(8, 8)$ for $N_1=200, N_2=400$, $(q_1, q_2)=(6, 8)$ for $N_1=200, N_2=600$ and $(q_1, q_2)=(6, 12)$ for $N_1=200, N_2=800$. The finite-sample distributions of (group) empirical Gini measures are not so skewed and better approximated by normal ones as compared to the Theil measures (see Section (ref)). In Table (ref), one observes good size control for different combinations of $q_1$ and $q_2$ in $t-$statistic robust tests of equality of/the difference between the Gini indices except only the cases $q_1=q_2=12$, $q_1=q_2=16$ and $q_1=12, q_2=16$ with rather small number of observations in each of the group. To avoid very conservative size properties, the best choices for the number of groups in applications of $t-$statistic robust tests in the case of Gini indices appear to be $(q_1, q_2)=(8, 8)$, $(q_1, q_2)=(9, 12)$ and $(q_1, q_2)=(8, 16)$ for all sample sizes $N_1, N_2$ considered.
Finally, Table (ref) provides the results for the case of samples with dependent observations, i.e., those with spatially dependent data relevant for studies of income distributions and inequality. Each of the two samples consists of 192 observations with the standard (parameters $\mu=0$ and $\sigma=1$) lognormal distribution located on a rectangular array of unit squares with 16 rows and 12 columns. The observations are generated such that the correlation between the logarithms of two observations is given by $\exp(-\phi d)$ for some $\phi>0,$ where $d$ is the Euclidean distance between the two observations (see Section 3.4 in IM2 for the use of a similar spatially correlated setting in the analysis of finite sample size properties of one-sample $t-$statistic approaches in inference on the mean of Gaussian observations with spatial dependence). The case $\phi=\infty$ to samples of i.i.d. observations.
More precisely, the observations in the samples are given by $I_{ij}=\exp(u_{ij}),$ $Y_{ij}=\exp(v_{ij}),$ $i=1, ..., 16,$ $j=1, ..., 12,$ where $u_{ij}$ and $v_{ij}$ are multivariate mean zero unit variance Gaussian with correlation between $u_{ij}$ and $u_{lk}$ and between $v_{ij}$ and $v_{lk}$ equals $\exp(-\phi\sqrt{(i-l)^2+(j-k)^2}).$
According to the results in Table (ref), the empirical size properties of $t-$statistic tests of equality of Theil and Gini indices in the two samples with spatial dependence are comparable (especially, for the tests based on the two-sample $t-$statistic $\tilde{\tilde t}$ with $q_1=q_2=q=4$ groups and the one-sample $t-$statistic $\tilde t$ in differences with $q=8$) to those of permutation and bootstrap procedures. Furthermore, the finite sample size properties of essentially all robust $t-$statistic tests are better than those of bootstrap and permutation tests under pronounced spatial dependence with $\phi=1.$
Next, we investigate finite-sample power properties of the tests considered. We report finite-sample size adjusted power for two-sample $t-$statistic and permutation tests.\footnote{Size adjustment is not performed for bootstrap tests as they are strongly dominated in terms of power by permutation test in all settings considered, see also Dufour.}. Under size adjustment, the resulting empirical size of a given $t$-statistic-based robust test and its permutation counterpart coincide under the null hypothesis, thereby enabling meaningful power comparisons.
We consider the following simulation designs.
Table (ref) presents the size-adjusted power when the two samples come from different Singh-Maddala distributions $SM(a_0, b_0, c).$ The sample sizes are $N_1=N_2=200,$ and the number of groups is the same for $t-$statistic tests: $q_1=q_2=q.$ The first sample has a fixed Singh-Maddala distribution $SM(a_0, b_0, c_0)$ with $a_0=2.8,c_0=1.7$ and the corresponding tail index $\zeta_I=4.76$ and the distribution of the second sample varies, with $a_0=2.8$, $c=0.7,1.1,1.7,2.7,31.7$ and the corresponding tail indices $\zeta_Y=1.96, 3.08, 4.76, 7.56, 88.76.$ The permutation test appears to be the most powerful although the two-sample $t$-statistic tests (based on $\tilde{t}_{\mathcal{L}}$ in ((ref))) have only slightly lower power (for $q=12,16$). The two-sample tests based on the $t-$statistic $\tilde{t}_{\mathcal{L}}$ in ((ref)) with $q_1=q_2=q$ are always more powerful than those based on the one-sample $t-$statistic $\tilde{\tilde{t}}_{\mathcal{L}}$ in ((ref)) in the differences of the group estimators with the same number of groups. In inference on both Theil and Gini indices, the power of $t-$statistic approaches based on ((ref)) is very similar across $q=8, 12, 16$ if the second distribution is more light-tailed than the first one, so that $c>c_0$ and $\zeta_Y>\zeta_I$ and also very similar for $q=12, 16$ if the second distribution is more heavy-tailed than the first one, with $c<c_0$ and $\zeta_Y<\zeta_I.$ In the former case of more lighted second distribution ($c>c_0$ and $\zeta_Y>\zeta_I$), the best power is exhibited by $t-$statistic tests based on ((ref)) with $q_1=q_2=q=8, 12.$ In the latter case of more heavy-tailed second distribution ($c<c_0$ and $\zeta_Y<\zeta_I$), the most powerful $t-$statistic test for inference on Theil indices is the one based on ((ref)) with $q_1=q_2=16,$ and the second best test is the $t$-statistic test based on ((ref)) with $q_1=q_2=12.$ Also, in the above case where the second distribution is more heavy-tailed than the first one, with $c<c_0$ and $\zeta_Y<\zeta_I,$ the most powerful $t-$statistic test for inference on Gini indices is the test based on ((ref)) with $q_1=q_2=12.$
Table (ref) provides the results on finite sample power properties of different inference approaches in the case of more heavy-tailed distributions. In the numerical analysis in the table, the fist sample has Singh-Maddala distribution $SM(a, b_0, c)$, with $a=2, c=1.1$ and $\zeta_I=2.2,$ and the second sample is from the Singh-Maddala distribution $SM(a, b_0, c)$ with $a=2,$ $c=0.7,0.9,1.1,1.5,3.7$ and the corresponding tail indices $\zeta_Y=1.4, 1.8, 2.2, 3, 7.4.$ The sample sizes are $N_1=N_2=200.$ According to the results in Table (ref), in the case of inference on Theil or Gini indices, the $t-$statistic tests based on $\tilde{t}_{\mathcal{L}}$ in ((ref)) with $q_1=q_2=8, 12, 16$ are typically the most powerful (this is the case for not very lighted second distribution); in particular, they are typically more powerful than permutation tests. In the case of inference on Theil indices, the best power properties are exhibited by the $t-$statistic tests based on with $q=16,$ and the second best test is the $t$-statistic test based on ((ref)) with $q_1=q_2=12.$ The choice of $q=12, 16$ also provides the best power properties for $t-$statistic tests based on $\tilde{t}_{\mathcal{L}}$ in ((ref)) in inference on Gini indices.
Tables (ref) and (ref) provide the results on finite-sample power properties of different inference approaches in the case of heavy-tailed distributions, including those considered in Table (ref) ($a=2,$ $c=0.7,0.9,1.1,1.5$ and the corresponding tail indices $\zeta_Y=1.4, 1.8, 2.2, 3,$ and also $c=3.7,$ $\zeta_Y=7.4$ in the case of Theil indices and $c=2.2$ and $\zeta_Y=4.4$ in the case of Gini indices), and different sample sizes, with $N_1=200,$ $N_2=400$ in the former table and $N_1=400,$ $N_2=200$ in the later one.
One can see that two-sample $t$-statistic tests are typically much more powerful than permutation tests if the more heavy-tailed distribution has larger sample size. Again, two-sample $t$-statistic tests based on $\tilde{t}_{\mathcal{L}}$ in ((ref)) with the number of groups $q_1=q_2=q$ are always more powerful than those based on the one-sample $t-$statistic $\tilde{\tilde{t}}_{\mathcal{L}}$ in ((ref)) in the differences of the group estimators with the same number of groups.
Table (ref) gives the results on finite-sample size adjusted power of different inference approaches in the same same distributional settings as in Table (ref) and sample sizes $N_1=200$ and $N_2=800.$ Similarly, (ref) provides the results on finite-sample size adjusted power properties of the approaches in the same settings as in Table (ref) and sample sizes $N_1=800$ and $N_2=200$. We also consider different combinations of (not necessarily equal) numbers $q_1$ and $q_2$ of groups for $t-$statistic inference approaches. According to the results in Tables (ref) and (ref), if smaller sample is more heavy-tailed then the power of all two-sample $t$-statistic tests is dominated by that of permutation tests. Otherwise, if the larger sample is more heavy-tailed then the power properties of two-sample $t$-statistics tests (except the tests with very small $q_1$ and $q_2$) are typically considerably better than those of permutation test. One can further see that for inference on Theil indices, the best (compromise) choice of the number of groups in $t-$statistic testing approaches will be $q_1=12$, $q_2=6$ or vice versa because this choice leads to correct size and good power in comparison to other size-controlled two-sample $t$-statistic tests. For Gini indices, the finite-sample power properties are not very sensitive to choice $q_1$ and $q_2$. Interestingly, even if the samples differ 4 times as in the tables, the choice $q_1=q_2=8, 12, 16$ leads to a very good size adjusted power and seems to be one of the best across all combinations of $q_1$ and $q_2$. The choice $q_1=12$ and $q_2=9$ also a good choice and leads to power properties of $t-$statistic inference approaches that are comparable or slightly better than in the case $q_1=q_2=8, 12, 16$. The choice of the different number of groups $q_1$ and $q_2$ may be useful if the sizes of two samples differ very much.
Summarizing the results, the two-sample $t$-statistic robust approaches to testing equality of two inequality measures or inference on their difference appear to be useful complements to other inference methods, including computationally expensive bootstrap and permutation-based inference methods. Finite-sample properties of the $t-$statistic inference approaches appear to be better in the case of testing equality and comparisons of Gini measures as compared to the case of the Theil measures as the former measures are more robust to heavy tails.
In applications of two-sample $t$-statistic inference approaches, the appropriate choice of the numbers $q_1$ and $q_2$ of groups is needed. The most simple way to choose the numbers of groups in the case of distributions that are not very different from each other is to have $q_1/q_2$ (approximately) equal to $N_1/N_2$ so that the sizes of all the groups considered are the same. If two distributions have similar tail indices, then in the case of inference on Gini measures, $q_1$ and $q_2$ may be taken to be equal. In general, the size of the groups in the sample from a more heavy-tailed distribution should be larger than the size of the groups from a less heavy-tailed distribution. E.g., in the case of equally sized samples, one should take the number of groups in the more heavy-tailed sample to be less than the number of groups in the less heavy-tailed sample.
This section presents empirical results on comparisons of Gini coefficients in Moscow and Russian regions using the asymptotic, permutation, bootstrap and the robust $t-$statistic inference approaches considered in this paper.
The empirical analysis is based on a large database on the results of household income surveys conducted by the Federal State Statistics Service of Russia (Rosstat) in 2017 (available at $https://www.gks.ru/free\_doc/new\_site/vndn-2017/index.html$;\\ $https://www.gks.ru/free\_doc/new\_site/vndn-2017/OHousehold.html$). The database covers 160,000 households in Russian regions, and provides the data on, among many other variables, households' total income. The analysis of income inequality indices and their comparisons in this section is based on the above income levels of Russian household normalized, following Rosstat's methodology, by the total number of households' members.
Table A.1 in the appendix provides the $p-$values for the above tests of the null hypothesis $H_0: G_M=G_R$ against the alternative $H_a: G_M\neq G_R,$ where $G_M$ is the Gini coefficient in Moscow and $G_R$ is the Gini coefficient in Russian region $R.$ The entries in the table in bold are the $p-$values not greater than 0.05.
The table also provides the values of the Gini coefficients and the (bias-corrected) log-log rank-size regression estimates (with 5% tail truncation) of tail indices $\zeta$ of the income distribution among $N_2$ households surveyed in the regions (see GI, GI). It also provides the values of the ratio $N_1/N_2,$ where $N_1$ is the number of households surveyed in Moscow. It should be noted that if $q_1$ or $q_2> 14$, we can use only the significance level less than 0.083.
The Gini coefficients in Moscow and Russian regions range from 0.236 (Tambov Region) to 0.354 (the Republic of Ingushetia) indicating low to moderate inequality; Tyva Republic has the Gini coefficient of 0.423 (Tyva Republic) indicating high inequality. The value of the Gini coefficient for Moscow is 0.264 indicating rather low inequality.
The point log-log rank-size regression estimates $\zeta$ of tail indices of income distribution in most of Russian regions lie in the interval $(3, 6),$ with the exception of Karachay-Cherkess ($\zeta=2.08$) and Mari El ($\zeta=2.29$) Republics and Krasnodar ($\zeta=2.6$), Kursk ($\zeta=2.75$) and Tyumen ($\zeta=2.71$) regions. The corresponding confidence intervals for tail indices of income distribution in most of Russian regions lie on the right of 2 implying finite second moments and finite variances. The 95% confidence intervals for tail indices of income distribution in Krasnodar, Krasnoyarsk, Stavropol, Khabarovsk, Arkhangelsk, Astrakhan, Belgorod, Vladimir, Volgograd, Vologda, Voronezh, Ivanovo, Tver, Kemerovo, Kurgan, Kursk, Lipetsk, Magadan, Murmansk, Novosibirsk, Omsk, Oryol, Penza, Pskov, Ryazan, Sakhalin, Sverdlovsk, Smolensk, Tambov, Tomsk, Tyumen, Ulyanovsk and Yaroslav regions; Altai, Buryatia, Ingushetia, Kabardino-Balkar, Kalmykia, Karachay-Cherkess, Karelia, Komi, Mari El, Mordovia, North Osetia, Tyva and Sakha Republics; Chukotka, Khanty-Mansi and Nenets Autonomous Districts and Kamchatka Kray intersect with the interval $(1.5, 3)$ where tail indices of income distribution in developed countries typically lie. The 95% confidence intervals for tail indices of income distribution in Amur, Bryansk, Chelyabinsk, Irkutsk, Kaliningrad, Kaluga, Kirov, Kostroma, Leningrad, Moscow (the tail index estimate is 3,96 with the 95% confidence interval $(3.44, 4.48)$), Nizhny Novgorod, Novgorod, Orenburg, Perm, Rostov, Samara, Saratov, Sevastopol and Tula regions; Adygeya, Bashkortostan, Chuvash, Crimea, Dagestan, Khakassia and Tatarstan Republics and Kamchatka, Primorsky and Zabaykalsky Krays lie on the right of 3 thus implying finite third moments and variances.
According to the table, on the base of all the tests considered, including the $t-$statistic tests with most of the values $q_1, q_2,$ the null hypothesis $H_0: G_M=G_R$ is rejected in favor of the alternative $H_a: G_M>G_R$ (at the level 2.5%) for the Republic of Tatarstan, Sevastopol City and Bryansk, Kostroma, Tambov and Tula regions. For Penza, Smolensk and Ulyanovsk regions and Udmurtia, $H_0: G_M=G_R$ is rejected in favor of $H_a: G_M>G_R$ on the base of the asymptotic, bootstrap, permutation and the $t-$statistic tests with some of the values $q_1, q_2$ in the table.
Further, according to all the tests considered, including the $t-$statistic tests for most of the values $q_1, q_2,$ the null hypothesis $H_0: G_M=G_R$ is rejected in favor of the alternative $H_a: G_M<G_R$ (at the level 2.5%) for Amur, Chelyabinsk, Irkutsk, Khabarovsk, Krasnodar, Krasnoyarsk, Kurgan, Moscow, Sakhalin and Jewish and Yamalo-Nenets Autonomous regions as well as for the Republics of Bashkortostan, Buryatia, Dagestan, Ingushetia, Kalmykia, Khakassia and Sakha (Yakutia); Altai, Chechen, Kabardino-Balkar, Karachay-Cherkess, Komi and Tyva Republics; Kamchatka, Primorskiy, Zabaykalsky Krays; Khanty-Mansi and Nenets Autonomous Okrugs and Chukotka Autonomous District. For Astrakhan, Kaliningrad, Kemerovo, Novosibirsk, Omsk, Penza, Smolensk, Sverdlovsk, Tomsk and Tyumen Regions, $H_0: G_M=G_R$ is rejected in favor of $H_a: G_M<G_R$ on the base of the asymptotic, bootstrap, permutation and the $t-$statistic tests for some of the values $q_1, q_2$ in the table.
Two conclusions are interesting to note.
First, income inequality appears to be higher in most of the Russian Regions as compared to Moscow.
Second, the conclusions of all the approaches to testing equality of the Gini coefficients $G_M$ and $G_R$ considered - the asymptotic, bootstrap, permutation and the robust $t-$statistic tests - for the above regions agree among themselves. Two exceptions are Belgorod and Novgorod Regions, where $H_0: G_M=G_R$ is not rejected in favor of $H_a: G_M<G_R$ on the base of the asymptotic, bootstrap, permutation, but is rejected on the base of robust $t-$statistic tests for some values of $q_1, q_2.$
Empirical analyses on inequality measurement and those in other fields in economics and finance often face the difficulty that the data is correlated, heterogeneous or heavy-tailed in some unknown fashion. In particular, as has been documented in numerous studies, observations on many variables of interest, including income, wealth and financial returns, typically exhibit heterogeneity, dependence and heavy tails in the form of commonly observed Pareto or power laws.
The paper focuses on applications of the recently developed t-statistic based robust inference approaches in the analysis of inequality measures and their comparisons under the above problems. Following the approaches, in particular, a robust large sample test on equality of two parameters of interest (e.g., a test of equality of inequality measures in two regions or countries considered) is conducted as follows: The data in the two samples dealt with is partitioned into fixed numbers $q_1, q_2\ge 2$ (e.g., $q_1=q_2=2, 4, 8$) of groups, the parameters (inequality measures dealt with) are estimated for each group, and inference is based on a standard two-sample $t-$test with the resulting $q_1, q_2$ group estimators. Robust $t-$statistic approaches result in valid inference under general conditions that group estimators of parameters (e.g., inequality measures) considered are asymptotically independent, unbiased and Gaussian of possibly different variances, or weakly converge, at an arbitrary rate, to independent scale mixtures of normal random variables. These conditions are typically satisfied in empirical applications even under pronounced heavy-tailedness and heterogeneity and possible dependence in observations.
The methods dealt with in the paper complement and compare favorably with other inference approaches available in the literature. We illustrate application of the proposed robust inference approaches by an empirical analysis of income inequality measures and their comparisons across different regions in Russia.
The $t-$statistic robust inference approaches, including the two-sample approaches for inference on equality of and the difference between parameters of interest considered in this paper are simple to use and have a wide range of applicability in econometric and statistical analysis under the problems of heterogeneity, dependence and heavy-tailedness in observations. The approaches do not require at all estimation of limiting variances of estimators of interest, in contrast to inference methods based on consistent, e.g., HAC or clustered, standard errors that often have pure finite sample properties, especially under pronounced heterogeneity and dependence in observations. In addition, the inference approaches can be used under extremes and outliers in observations generated by heavy-tailedness with infinite variances and also in settings where observations (e.g., on income or wealth levels) in each of the samples considered are dependent among themselves - for instance, due to spatial or clustered dependence, common shocks affecting them, or, in the case of time series or panel data on income or wealth levels, due to autocorrelation and dependence in observations over time. Further, in the case of testing for equality of inequality measures or inference on their difference in two populations using two samples of possibly dependent observations, as above, the $t-$statistic inference approaches may be used under an arbitrary dependence between the samples as well as under possibly unequal sample sizes.
In addition to inference on inequality and wealth indices dealt with in this work, the approaches may also be applied in inference on and comparisons of poverty and concentration indices where, as is well-known, the presence of extreme values, outliers, heavy-tailedness and heterogeneity makes problematic their applicability and the use of asymptotic methods in inference on the indices similar to the case of inequality measures (see, among others, Appendix B.1 in Section E7 in Mand, Mand, DF, DF, and Section 3.3.2 in IIW, IIW) as well as in inference on tail indices in power laws ((ref)) for income and wealth distributions and corresponding measures of top inequality (see the discussion in Section (ref) and references therein). These and other applications of the $t-$statistic robust inference approaches are currently under way by the authors and their co-authors.
\setcounter{section}{0} \setcounter{subsection}{0} \setcounter{figure}{0} \setcounter{table}{0} \setcounter{equation}{0} \renewcommand\Alph{section}{\Alph{section}} \renewcommand\Alph{section}.\arabic{figure}{\Alph{section}.\arabic{figure}} \renewcommandА.\arabic{table}{А.\arabic{table}} \renewcommandB.\arabic{equation}{\arabic{equation}}
{\centerline{{Appendix A: Tables}}}
\setcounter{section}{1} \setcounter{section}{0} \setcounter{subsection}{0} \setcounter{figure}{0} \setcounter{table}{0} \setcounter{equation}{0} \renewcommand\Alph{section}{\Alph{section}} \renewcommand\Alph{section}.\arabic{figure}{\Alph{section}.\arabic{figure}} \renewcommandА.\arabic{table}{А.\arabic{table}} \renewcommandB.\arabic{equation}{\arabic{equation}}
{\centerline{{Appendix B: Inequality measures and their sample analogues }}} In this section, we review the definitions of the widely used Gini and Theil inequality measures, sample analogues of the measures and their asymptotic properties (see, among others, CF, DF, Section 13.F, 17.C in MO, and references therein).
Let $I$ be an (absolutely continuous) nonnegative r.v. (e.g., income or wealth level) with the finite first moment $\mu_I=E[I]<\infty$ and the cdf $F_I(x)$ representing income or wealth distribution in a population, and let $I_1, I_2, ..., I_N$ denote a sample of observations on the r.v. $I.$
As usual, we denote by $\overline{I}_N=N^{-1}\sum_{i=1}^N I_i$ and $s_N^2=(N-1)^{-1}\sum_{i=1}^N (I_i-\overline{I})^2$ the sample mean and sample variance of the observations $I_i.$
Below, we provide the definitions of Theil and Gini inequality measures (denoted by $\mathcal{L}_{Theil}^I$ and $\mathcal{L}_{Gini}^I$ for the population considered) and discuss the standard results on their asymptotic normality.
Theil index The population Theil index is defined by $$\mathcal{L}_{Theil}^I=\frac{E[I \log I]}{\mu_I}-\log(\mu_I).$$
The Theil index is the limiting case of the Generalized Entropy measures. Its sample analogue - sample Theil index - is given by
$$\hat{\mathcal{L}}_{Theil, N}^I=\frac{\frac{1}{N}\sum_{i=1}^N I_i \log(I_i)}{\overline{I}_N}- \log(\overline{I}_N).$$
Under i.i.d. observations $I_1, I_2, ..., I_N,$ the Theil index is asymptotically normal if $E[I^2]<\infty,$ $E[I^2\log I]<\infty$ and $E[I^2 \log^2(I)]<\infty.$ It is easy to see that these conditions are satisfied in the case of r.v.'s with power law distributions ((ref)) (e.g., Singh-Maddala distributions $SM(a, b, c)$ in ((ref)) with $\zeta=ac$) if the tail index $\zeta$ is greater than 2: $\zeta>2.$
Under the above conditions, one has
$$\sqrt{N}(\hat{\mathcal{L}}_{Theil, N}^I-\mathcal{L}_{Theil}^I)\rightarrow_w N(0, v_{Theil, I}^2),$$
where $$v_{Theil, I}^2=\frac{E[I^2 \log^2 I]}{\mu_I^2}+\frac{E[I^2]} {\mu_I^2} \Big(\frac{E[I \log I]}{\mu_I}+1\Big)^2-\frac{2E[I^2 \log I]}{\mu_I^2}\Big(\frac{E[I \log I]}{\mu_I}+1\Big)-1$$
(see, among others, MZ, Cow1, Cow2, CF and Afr for the review of the results on asymptotic normality and the formulas for the liming and sampling variance of different estimators of inequality measures).
Gini coefficient The population Gini coefficient is defined by $$\mathcal{L}_{Gini}^I=0.5 \frac{E|I'-I''|}{\mu_I},$$ where $I'$ and $I''$ are independent copies of the r.v. $I.$
The most commonly used (nonparametric) estimator of the Gini coefficient $\mathcal{L}_{Gini}^I$ is given by its sample analogue (the sample Gini coefficient)
where $U_N$ is the $U-$statistic $U_N=\frac{2}{N(N-1)} \sum_{1\le i<j\le N} |I_i-I_j|$ (we refer to, among others, Hoeffding, Ch. 5 in Serfling and Ch. 4 in KB for the asymptotic theory for general $U-$statistics).
From the results in the above references, it follows that asymptotic normality for the $U-$statistic $U_N$ and the sample Gini coefficient holds if $I_1, I_2, ..., I_N$ are i.i.d. observations with finite second moment $E[I^2]<\infty.$ This holds under power-law distributions ((ref)) (e.g., for Singh-Maddala distributions $SM(a, b, c)$ in ((ref)) with $\zeta=ac$) if the tail index $\zeta$ is greater than 2: $\zeta>2.$ More precisely, under the above conditions (see Hoeffding) $$\sqrt{N}(\hat{\mathcal{L}}_{Gini, N}^I-\mathcal{L}_{Gini}^I)\rightarrow_w N(0, v_{Gini, I}^2),$$ where $v_{Gini, I}^2=(\mathcal{L}_{Gini}^I)^2 \sigma_I^2-2 \mathcal{L}_{Gini}^I E\{I'|I'-I''|\}/\mu_I^2+E(E_{I'}\{|I'-I''|\})/\mu_I^2,$ and $E_{I'}(\cdot)=E_{I'}(\cdot)=E\{\cdot|I'\}$ denotes the expectation conditional on $I'.$
Naturally, the asymptotic normality of the sample Theil and Gini coefficients is lost under infinite second moments and variances: $E[I^2]=\infty.$ For instance, from the results in Taleb it follows that under i.i.d. observations $I_1, I_2, ..., I_N$ that follow a power-law distribution ((ref)) with the tail index $\zeta\in (1, 2)$ (e.g., the Singh-Maddala distribution $SM(a, b, c)$ in ((ref)) with $1<\zeta=ac<2$) and have finite first and infinite second moments, the sample Gini coefficient $\hat{\mathcal{L}}_{Gini, N}$ has an asymptotic right-skewed stable distribution with the index of stability $\zeta.$ Using the standard generalized CLT and the delta-method, it is also not difficult to show that in the case of distributions exhibiting (double) power law behavior in both the lower and the upper (with the tail index $\zeta$), similar to Singh-Maddala distributions $SM(a, b, c)$ with $\zeta=ac,$ the sample Theil index $\hat{\mathcal{L}}_{Theil, N}$ weakly converges to a function of stable r.v.'s with indices of stability that depend on $\zeta.$ The rate of convergence in the above asymptotic results is slower than $\sqrt{N}$ and depends on $\zeta.$ The fact that the tail index $\zeta$ is unknown in practice makes the results useless for (direct) asymptotic inference.\footnote{The situation is somewhat similar to the properties of autocorrelation functions of GARCH-type processes and their squares, where asymptotic normality is lost under tail indices smaller than 4 and infinite fourth moments, as is typically the case for financial returns and foreign exchange rates in real-world markets (see DM, MS and also IPS for asymptotically valid robust $t-$statistic approaches to inference on measures of market (non-)efficiency and volatility clustering based on powers of absolute values of GARCH-type processes, e.g., financial returns).}