Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
51,842 characters · 8 sections · 71 citation commands
Mind the Income Gap: Bias Correction of Inequality Estimators in Small-Sized Samples
\selectlanguage{english}
JEL Codes --- C15, D31
Keywords --- Complex Surveys, Finite Populations, Income Inequality, Small Area Estimation
The interest in reliable local estimates of economic inequality is growing due to the observed increment in the income gap and social exclusion among regions. Specifically, inequality estimates for specific sub-populations - such as areas at a fine level of geographical disaggregation or rather specific socio-demographic groups - are increasingly in demand marquez2019role. Policymakers and stakeholders need these to formulate and implement policies, distribute resources and measure the effect of policy actions at local levels. In addition, their contribution to regional studies is valuable in the process of decomposing spatial spillovers and identifying local areas that drive inequality at national levels cavanaugh2018locating.
When dealing with inequality estimation in specific groups or local scales, a problem of observations scarcity typically arises. Disposable income is generally adopted as the variable of interest and the primary source of data collection is through household surveys. However, since such surveys are not planned for the estimation of target quantities in specific domains, they result in small sample sizes. In this context, small area estimation techniques are applied, integrating survey with auxiliary data to "borrow strength" across areas and, in this way, improve the reliability of estimates.
The small area models can be specified at the unit (individual or household) level; previous proposals dealing with inequality estimation at the unit level are provided by tzavidis2016robust and marchetti2021robust. However, such models require a large amount of data as, generally, the auxiliary variables have to be known for each unit of the population and linked to survey data. This may be hard to get as administrative archives are not publicly accessible at individual level, cross-linked and associated with survey data harmeningframework. On the other hand, small area models defined at the area level are less demanding in terms of data requirements, needing only survey (direct) estimates endowed with related measures of uncertainty and areal covariates rao2015. An application of such models to inequality estimation can be found in benedetti2023.
Area-level models with common specifications, such as the famous Fay-Herriot model fay1979estimates, have the strict assumption of (approximate) unbiasedness of the survey estimators given as input rao2015. In this paper, we focus on tha bias affecting inequality estimators in small samples, often underestimating inequality Deltas2003, Breunig2008. Their bias may depend on the characteristic of the distribution of the variable of interest, i.e. the income variable Breunig2001, and on the uncertainty induced by the sample selection scheme.
Unfortunately, the bias issue is typically neglected when measuring inequality with area-level models, leading to model misspecification and thus to a possible misleading inference. Note that such an aspect deserves attention given that estimates of inequality measures are often used for comparisons across time and locations. Neglecting such bias may bring out discrepancies that, rather than being true inequality gaps, may be due to disparate sample sizes or to different underlying distributions of the variable of interest Breunig2008. In this vein, we propose a bias correction strategy for a large set of inequality measures and we adopt it in an illustrative small area estimation exercise.
Concerning the Gini index, a large body of literature faces the small sample bias issue, such as jasso1979gini, lerman1989improving, Deltas2003, Davidson2009, Ourti2011 in $iid$ samples. The context of application is varied, spanning from economic inequality to crime or concentration of scholarly citations mohler2019reducing, kim2020influence. Fabrizi2016 tackle such an issue in the complex survey case and their correction is indeed considered within a small area estimation framework. However, concerning alternative measures such as Atkinson Indexes and the Generalized Entropy (GE) measures, the literature on bias is very scarce, even in the $iid$ case: some contributions are provided by Giles2005, Schluter2009 and Breunig2008 by adopting different methodological approaches of correction.
Note that income data are collected through household surveys with complex sampling designs that adopts stratification and/or selection of sampling units in more than one stage. Thus, the sample selection process, together with ex-post treatment procedures such as calibration and imputation, invariably introduces a complex correlation structure in the data that has to be taken into account. This makes the development of a theoretically valid bias correction challenging, in contrast to classical $iid$ settings. Furthermore, the bias issue is even exacerbated in income data applications, traditionally affected by extreme values Kerm2007, since inequality measures are known to be highly unrobust to them cowell1996robustness. This aspect depends clearly on the type of measure we are dealing with and it becomes even more cumbersome to handle in the case of small samples.
We investigate the nature of the bias and propose a methodological framework for bias correction. Our proposal constitutes a generalization of the framework of Breunig2008, developed for $iid$ observations, to the finite population and design-based setting. At the same time, we extend the proposal to a wider set of measures from the Gini index to two parametric families of measures: the Atkinson and the Generalized Entropy family. We considered a wide variety of measures as the concurrent estimation of alternative indicators - as opposed to the more commonly used Gini Index - may bring to light a wider picture of the inequality phenomenon. To the best of our knowledge, this is the first proposal of bias correction for the Atkinson and Generalized Entropy indexes in the complex survey case, whereas it provides an extension for the Gini index case with respect to existing proposals as it is made clear in Section (ref).
To our purpose, we take advantage of a methodology based on Taylor's expansions, even if the same analytical results can be obtained through other types of linearization, such as the one proposed by graf2011use. The extension for complex designs has been developed considering Horvitz-Thompson type estimators, and the ultimate clusters technique for design variances and covariances estimation. An advantage of our proposal is that any parametric assumption on income distribution is not required, providing a very flexible framework. Our bias correction proposal is evaluated via simulations showing a noticeable bias reduction for all the measures and leading, in some cases, to approximately unbiased estimators. Lastly, we provide a small area estimation exercise that shows the risk of ignoring ex-ante bias correction.
The paper is organized as follows. The considered inequality measures are defined in Section (ref), while the bias correction strategy is set out in Section (ref) and the bias-correction estimation steps are detailed in Section (ref). A design-based simulation study involving the European Statistics on Income and Living Condition (EU-SILC) income data is provided in Section (ref) to evaluate the magnitude of the bias and the efficacy of our proposal. Lastly, a small area estimation exercise is carried out in Section (ref), to highlight the utility of our proposal in practice. Conclusions are drawn in Section (ref).
The most famous inequality measure is, indeed, the Gini concentration index, employed in social sciences for measuring concentration in the distribution of a positive random variable. There are several equivalent definitions of the Gini index ceriani2015individual, we use the formulation of sen1997economic. Suppose we have a finite population $\mathcal{U}$ of $N (< \infty)$ elements labelled as $\lbrace 1, \dots, N \rbrace$. Let $y_i$ be a characteristic of interest, in our case income, for the $i$-th unit of the finite population, where $y_i \in \mathbb{R}^+$, $i = 1, \dots, N$, and a sample $s_{iid}$ of size $n_{iid}$ is picked through simple random sampling. The Gini index estimator is defined as
with $n_i$ denoting the rank of $i$-th unit and $\hat{\mu}$ the sample mean.
However, the estimation of alternative measures, in addition to the Gini index, may enable a more meaningful assessment of different aspects of economic inequality. The Gini index is decomposable within and between groups only in very specific cases mookherjee1982decomposition, moreover, it is positional (weakly) transfer sensitive, namely because index variations induced by income transfers depend on the ranks of transfer donor and recipient. Lastly, it constitutes a Lorenz dominance-based measure, allowing only for a partial ranking of probability distributions. For example, two very different distributions - one having more inequality amongst poor, the other more inequality against the rich - can have the same index value.
When the distributional dominance fails, welfare-based measures, such as Atkinson Indexes, may provide for a complete ranking among alternative distributions at the expense of more stringent assumptions as to how to represent social welfare Bellu2006. Atkinson index has support [0,1] and is defined as
The parameter $\varepsilon$ expresses the level of inequality aversion: as $\varepsilon$ increases, the index becomes more sensitive to changes at the lower end of the income distribution.
Besides, an additive decomposable family of inequality measures is the Generalized Entropy class. As opposed to the measures seen before, this class has the advantage of being strongly transfer-sensitive, meaning that it reacts to transfers depending on donor and recipient income levels. It is based on the concept of entropy which, when applied to income distributions, has meaning of deviations from perfect equality:
The parameter $\alpha$ sets the sensitivity of the index: a large $\alpha$ induces the index to be more sensitive to the upper tail, and vice versa a small $\alpha$ to the lower tail. $\theta_{GE}(0)$ is the Mean Log Deviation, while $\theta_{GE}(1)$ is the well known Theil index. Atkinson and Generalized Entropy are two interrelated parametric families of measures, as a transformation of the Atkinson Index is a member of the GE class:
In this paper, we consider the estimation of both classes separately, since common parameter values used in one family do not correspond deterministically to parameter values commonly used for the other family. Lastly, we consider the coefficient of variation (CV) as an inequality measure, being linked with a member of the GE family namely $\theta_{GE}(2)=$CV$^2/2$. Its square has been used in some income distribution analyses, including organisation2011divided, even though it seems to be very sensitive to top outliers Atkinson2015.
The bias of inequality estimators in small samples can be due to the structure of inequality measures as a non-linear function of estimators. The bias can be either positive or negative, depending on the characteristics of the reference variable distribution, except for the Mean Log Deviation which has a structurally negative bias as shown further on in this section. Among the measures with non-predictable bias direction, Breunig2001 shows that the bias of CV and GE$(2)$ is negatively related to the skewness of income distribution. This aspect could be analyzed in-depth by imposing a distributional assumption on the income variable, this is beyond the scope of this paper. For GE and Atkinson measures, the limiting behaviour of their bias is described in the following proposition.
We are interested in a variety of non-linear functions of income values as inequality measures are. Let denote with $s$ a sample of size $n$, drawn using a complex sampling design, with $p(s)$ the probability of selecting the particular sample $s \subset \mathcal{U}$ out of the set of all possible samples $\mathcal{Q}$, thus $p(s)\geq 0$ and $\sum_{s\in \mathcal{Q}} p(s) = 1$. The inclusion probability of unit $k$ is denoted with $\pi_k$, being $\pi_{k}=\sum_{s \in \mathcal{Q}_k} p(s)$ with $\mathcal{Q}_k$ the set of all possible samples including unit $k$.
We consider the generic inequality measure written as a function of the mean $\mu$ and $\gamma=\operatorname{\mathbb{E}}[g(y)]$, with $g(\cdot)$ a generic monotone transformation of the income variable. The population value for the generic inequality measure is
with $f(\cdot)$ a twice-differentiable function. The related estimator in our complex survey framework is $\hat{\theta}=f(\hat{\mu}, \hat{\gamma})$ in which Horvitz-Thompson estimators of $\mu$ and $\gamma$ are plugged in, i.e.
where $\boldsymbol w=(w_1, \dots, w_n)=(1/\pi_1, \dots, 1/\pi_n)$ or a treated and calibrated version of it and $N$ is the population size. Note that the results of this section also hold for Hájek type estimators, i.e. with denominator $\hat{N}=\sum_{i=1}^n w_i$, since it is approximately unbiased sarndal2003model. kakwani1990large uses a similar approach to express inequality indices to derive their asymptotic standard error. By simply applying a second-order Taylor's series expansion of the sample estimator around the population values and evaluating its expected value, the bias can be expressed as
notice that $\hat{\mu}$ is unbiased.
In Table (ref), we detail the survey estimators for each inequality measure and their bias formulation based on Equation (ref) along with all relevant quantities. The complex survey estimators of Atkinson and Generalized Entropy measures come from Biewen2006, while as for the Gini index, we employ the alternative formulation defined by sen1997economic and the complex survey estimator proposed by Langel2013. Let denote with $\sqrt{n/(n-1)}$ the standard bias-correction adjustment for the weighted variance; $F(\cdot)$ denotes the cumulative distribution function of the variable of interest and lastly consider $\hat{N}_i=\sum_{k \in s} w_k \mathds{1}(n_k \leq n_i)$. The notation $\mathds{1}(A)$ defines an indicator function, assuming value 1 if $A$ is observed and 0 otherwise.
Note that the bias formulas of Table (ref) can also be reached differently, namely by applying the linearization proposed by graf2011use and extended by vallee2019linearisation, as made explicit in the Appendix. The Graf's methodology requires a separate derivation for each measure. In contrast, Equation (ref) defines a general formulation of the bias which applies to the entire set of considered measures, isolating its components and easing a general interpretation.
Let us denote the Gini index estimator with $\hat{\theta}_G$, its approximate bias in small samples is
with $\gamma$ and $\hat{\gamma}$ as defined in Table (ref) and $\theta_G$ denoting the true value. The derivation of the approximate bias related to the weighted estimator $\hat{\gamma}$ is not trivial. As explained by Langel2013, its numerator is not composed of two simple sums. Indeed the quantity $\hat{N}_k$, an estimator of the rank of unit $k$, is random since its value depends on the selected sample. One solution is to consider the approximate bias of the corresponding $iid$ estimator, i.e. $\operatorname{\mathbb{E}}[\hat{\gamma}-\gamma] =-1/n ( \gamma- \mu/2)$ as derived by Davidson2009, so that:
This correction is in line with Davidson2009 and Fabrizi2016 proposals. However these are based on a first-order Taylor's expansions and thus limited to the first term of the right-hand side equation ((ref)), ours extends it to a second-order expansion. This translates into the fact that, while jasso1979gini, Deltas2003 and Davidson2009 proposals identify the adjusted Gini in $iid$ context as $n(n-1)^{-1}\hat{\theta}_G$, our correction reconsiders the shape of the adjusted estimator with a further order of approximation as
with $a$ equals the sum of the second and third terms of ((ref)).
As clear from Table (ref), the bias correction of GE(2) does not include the coefficient of skewness of the income distribution, as shown by Breunig2001. A reliable estimation of that quantity, while being straightforward in the $iid$ case, appears cumbersome in the case of weighted data being defined on a discrete grid of values. This leads to the non-applicability of Breunig2001 result in our case.
In this section, we detail the estimation of the approximate bias defined in Table (ref) for each measure. Such estimation is not trivial considering that the mentioned expressions depend on design variances and covariances $\mathbb{V}[\hat{\mu}]$, $\mathbb{V}[\hat{\gamma}]$ and $Cov[\hat{\mu}, \hat{\gamma}]$. We consider a complex survey design involving stratification and multi-stage selection, with both Self-Representing (SR), i.e. included at the first stage with probability one, and Non-Self-Representing (NSR) strata. This design is consistent with the majority of income survey designs and, in general, with official statistics household surveys.
We define an unbiased estimator for the variance of Horvitz-Thompson estimators, such as $\hat{\mu}$, when $w_i=1/\pi_i$ as
with $\pi_{i k}$, $\forall i, k \in \mathcal{U}, i\neq k$ denoting the second-order inclusion probabilities i.e. the probability that the sample includes both $i$-th and $k$-th units arnad2017. However generally (a) $w_i \neq 1/\pi_i$ and (b) $\pi_{ik}$, $\forall i, k \in \mathcal{U}, i \neq k$ are difficult to calculate under complex sampling designs.
Therefore, the variance estimator to be considered constitutes an approximation that relies on simplified assumptions. Firstly, we assume that Primary Sampling Units (PSU) are sampled with replacement, and secondly, we reduce multi-stage sampling into a single-stage process by relying on the Ultimate Clusters technique kalton1979ultimate. Moreover, we take into account the hybrid nature of the probability scheme, blending a variance estimator for stratified design associated with the SR strata, including a finite population correction factor, and a typical Ultimate Cluster variance estimator for multi-stage schemes associated with the NSR strata. The latter one is widely used in official statistics, see Osier2013 for Eurostat procedures. Without loss of generality, let us consider a two-stage scheme, where $\hat{\mu}=\sum_{h}\sum_{d}\sum_{i} w_{hdi} y_{hdi}/N$ is a linear estimator of $\mu$, with $h$ the stratum indicator, $d$ the Primary Sampling Unit (PSU) indicator and $i$ the secondary sampling unit (household) indicator. Its variance estimator is as follows:
with $H_{SR}$ self-representative and $H_{NSR}$ non self-representative strata, $M_h$ the number of resident households in strata $h$, $m_h$ the number of sample households in strata $h$, $f_h=m_h/M_h$ a finite population correction factor, $n_h$ the number of PSUs in strata $h$. Consider, moreover, that $\bar{y}_{h} = \sum_{i=1}^{m_h} y_{hi}/ m_h$, $\hat{\mu}_{hd}= \sum_{i=1}^{m_d} w_{hdi}y_{hdi}/N$ with $i$ denoting the household label and $m_d$ the number of sample households in PSU $d$, lastly $\bar{\mu}_{h}=\sum_{d=1}^{n_h}\hat{\mu}_{hd}/n_h$, with $n_h$ being the number of PSU in stratum $h$. Obviously, if $n_h = 1$ for some strata, the estimator ((ref)) cannot be used. A solution is to collapse strata to create “pseudo-strata” so that each pseudo-stratum has at least two PSUs. Common practice is to collapse a stratum with another one that is similar with respect to some survey target variables rust1987strategies.
An estimator of $\mathbb{V}[\hat{\gamma}]$ can be obtained by adopting the same strategy used for $\mathbb{V}[\hat{\mu}]$ in ((ref)). Whereas, regarding the estimation of the design covariance, consider that
Thus, a possible estimator $\hat{Cov}[\hat{\gamma}, \hat{\mu}]$ would be simply obtained by plugging in the variance estimators previously mentioned, while $\mathbb{V}[\hat{\gamma} +\hat{\mu}]$ is estimated by considering $\hat{\gamma} +\hat{\mu}=\sum_{i \in s} w_i (g(y_i)+y_i) /N$. The estimation procedure is completed by replacing $\mu$ and $\gamma$ with $\hat{\mu}$ and $\hat{\gamma}$.
The Gini index estimator differs from the other indexes since $\hat{\gamma}$ is a non-linear statistic. Thus, a linearization of $\hat{\gamma}$ is needed to make it tractable and carry on variance estimation with the procedure described above. We consider again the linearization proposed by graf2011use with the practical adaptation of Graf2014a for inequality estimators. In such adaptation, the linearized variable is merely a function of the partial derivatives with respect to the weights, that in the case of $\hat{\gamma}$ defined for Gini index in Table (ref) is
for a generic unit $k$ where $\mathcal{S}_k=\lbrace i \in s, n_i > n_k \rbrace$. In this way, the estimator can be re-expressed through a linear approximation, namely $\hat{\gamma} \approx \sum_{i \in s} w_i v_i$, and it becomes possible to perform variance estimation of linear statistics.
A design-based simulation study has been conducted to evaluate our bias correction proposal. In this simulation, the cross-section Italian EU-SILC sample (2017 wave) has been assumed as pseudo-population and the 21 NUTS-2 regions have been considered as target domains. The study is based on real income data, in order to check whether this specific framework works with close-to-reality data, affected by peculiar problems (e.g. extreme values, skewness).
For comparison purposes, two simulation scenarios have been carried out. In the first one, the original income data are employed as pseudo-population. In the second one, an extreme values treatment is performed concerning both upper and lower tails, to circumvent non-robustness problems. The resulting dataset is specified as an alternative pseudo-population. We compare the results obtained after the treatment with the ones before treatment to isolate the effect of extreme values when evaluating bias-correction performances (Table (ref)).
The issue of robust estimation of economic indicators through an extreme values treatment in the upper tail of income distribution is well-established in the literature. See Brzezinski2016 for a review and Alfons2013 for a suitable specification for survey data. On the contrary, the issue of treatment of extreme values in the lower tail of income distribution appears less established Kerm2007, masseran2019power. Concerning the upper tail, we operated a semi-parametric Pareto-tail modelling procedure using the Probability Integral Transform Statistic Estimator (PITSE) proposed by Finkelstein2006, which blends very good performances in small samples and fast computational implementation, as suggested by Brzezinski2016. As regards the lower tail, we used an inverse Pareto modification of the PITSE estimator, suggested by masseran2019power. In our simulations, the treatment has been done at a regional level to the original EU-SILC sample and the detection of extreme values has been carried out following MohdSafari2018 by using the Generalized Boxplot procedure. We expect that, when outlying observations are representative, such treatment would highly bias the outcome and thus we do not recommend it.
From both pseudo-populations, we repeatedly select 1,000 two-stage stratified samples, mimicking the sampling strategy adopted in the survey itself: in the first stage, SR strata are always included in the sample, while a stratified sample of PSU in NSR strata is selected; in the second stage, a systematic sample of households is drawn from each PSU included at the first stage. We repeated the drawing for both scenarios involving different sampling rates, 1.5% and 3% respectively. The Relative Bias (RB) and the Absolute Relative Error (ARE) in percentage have been calculated for each region $r$ using the 1,000 iterations as:
where $\theta_r$ is the population value for region $r$ and $\hat{\theta}_{p,r}$ its estimate at iteration $p$. In our simulation setting, the regional sample sizes range from 6 to 96 individuals (from 6 to 32 households) on average over the simulated samples for the 1.5% sampling rate, and from 11 to 196 individuals (10 to 74 households) for the 3% sampling rate.
Concerning the treated pseudo-population scenario, Figure (ref) illustrates the relative bias for each domain of non-corrected measures (grey line) and of corrected measures (blue line) in 3% samples versus the (average) sample size. The negative relation between sample size and average relative bias is clear for both the survey estimator $\hat{\theta}$ and the bias-corrected estimator $\hat{\theta}_{corr}$. This confirms the nature of the bias as a small sample bias and shows the effectiveness of the correction, even if based on a large-n approximation as the Taylor's expansion. The bias reduction is noticeable for all measures, leading to slightly biased estimates depending on the measure. Notice that the bias correction works well for measures that are not particularly sensitive to extreme observations such as the Gini index, GE$(0)$, Atk$(0.5)$ and Atk$(1)$. In the case of CV and GE$(2)$, the correction provides good results, but it seems, however, not to capture all the bias components. This may confirm the results of Breunig2001, suggesting that the coefficient of variation squared and GE(2) bias depends on the coefficient of skewness of the income distribution, not considered in our bias correction.
Bias and error averaged across all areas for each scenario, sampling rate and estimator are shown in Table (ref). By still focusing on treated population results, the correction induces a reduction of the RB spanning from 5% (CV, 3% rate) to 14% (Gini, 1.5% rate) approximately by considering both sampling rates. When the sample size is greater than 20 individuals ($n \geq 20$), the bias-corrected estimators seem to be approximately unbiased. Furthermore, it is important to note that the bias correction induces a slight but negligible error (ARE) increase for every measure, except for the Gini index which presents a relevant increase. This exception may be explained by the shape of the unbiased estimators, as described by ((ref)), where a sum of estimators is multiplied by a factor $n/(n-2)$, which inherently inflates the variance by its square.
Let us focus on comparing the treated population scenario with the non-treated one. In the latter case, bias and error increase dramatically both for $\hat{\theta}$ and for $\hat{\theta}_{corr}$. In particular, the bias is great for some measures estimated on the non-treated scenario due to their non-robustness properties to extreme values. It is the case of Atk ($\varepsilon=2$), extremely sensitive to low-income values (under $100$ euro per year) which is -48% biased on average for the scenario with the smallest sample sizes. Also, GE with $\alpha$ equal to 1 and 2 are highly sensitive to high-income values being -18% and -23% biased. However, the bias correction leads to a bias reduction comparable in magnitude to the one discussed for the treated pseudo-population; it seems not to change in magnitude with respect to the sample size and the presence of extreme values.
To summarize, our results highlight that in the case of populations that are not affected by income extreme values, the bias correction may provide approximately unbiased estimates for a large class of measures at the expense of, in most cases, only a slight error increase. Vice versa, it might be necessary to restrict the attention to the most robust measures such as GE with $\alpha=0$, Atkinson index with $\varepsilon=1$ and Gini Index to obtain estimates affected by a negligible bias. Another important aspect to point out is that, in certain countries, the EU-SILC is based on registers that better capture top incomes, thus, a cross-country comparison of income inequality by effects on a tail-sensitive measure must be another reason for caution Atkinson2015. Such results may constitute a reference when measuring inequality in small samples, however, since the simulation scenario uses specific data, reflections that have been drawn cannot be general or conclusive.
In the previous sections, we propose a method to correct the small sample bias of inequality estimators in complex surveys. Even if bias-corrected, such estimators are still unreliable due to the high variability induced by the small sample size: this means that estimates cannot be released or used for further inference. As a consequence, when measuring inequality at a fine-grained level, it becomes necessary to rely on Small Area Estimation (SAE) techniques. Such estimation techniques take advantage of available auxiliary information to produce estimates with acceptable uncertainty. More specifically, the model-based SAE techniques employ hierarchical models which can be defined both at area-level, linking area-defined survey estimates with areal covariates, or at unit (individual) level, linking individual income data with individual covariates. See tzavidis2018start for an up-to-date review.
In this context, area-level models appear to be less demanding in terms of data requirements and enable the incorporation of design-based properties. Such models constitute a typical framework of application of our bias-correction proposal, as they assume the unbiasedness of survey estimators used as input. As a consequence, their applicability to the estimation of inequality measures is inevitably tied to a preliminary bias correction, in contrast with unit-level models that do not involve survey estimators.
In this section, we perform an SAE exercise by using the sample related to the first iteration of the simulation detailed in Section (ref) for the 3% case. The purpose is not to propose a small area estimation strategy nor to provide a real application of inequality mapping, but rather to illustrate the framework of application of our bias-correction proposal and, especially, to underline the risk of avoiding bias-correction when estimating inequality in small domains. Such exercise is carried out by applying the Fay-Herriot model fay1979estimates, a landmark model in the small area literature, implemented through the package sae molina2015sae to both uncorrected and corrected survey estimators. The objective is to check whether the inclusion of biased or bias-corrected survey estimates in the model may lead to different results. From the whole set of inequality estimators considered in Sections (ref), (ref) and (ref), we perform the exercise on the most popular ones: the Theil index (Generalized Entropy with $\alpha=1$), the Atkinson index with $\varepsilon=1$ and the Gini index.
Specifically, let us consider $\hat{\theta}_1, \dots, \hat{\theta}_M$ as the set of survey estimators referring to a generic inequality measure in $M$ small areas, with corresponding population values $\theta_1, \dots, \theta_M$, and $\boldsymbol x_m$ the set of $p$ areal covariates for area $m$, $m=1, \dots, M$. The classical area-level model is the Fay-Herriot one, specified as follows:
where $D_m$ denotes the sampling variance of the survey estimator, usually assumed to be known to allow for identifiability, $\boldsymbol \beta$ the set of regression coefficients and $\sigma^2$ the model variance. This clearly implies $\mathbb{E}(\hat{\theta}_m)=\theta_m$ $\forall m$, i.e. the unbiasedness of survey estimators. As a consequence, neglecting the bias correction of survey estimators effectively leads to model misspecification.
As mentioned above, the sampling variance is separately estimated from the data and given as input to the small area model. Since our exercise is merely illustrative, we adopt the Monte Carlo variances of the design-based simulation in Section (ref) as sampling variances and simulated covariates for both estimators. However, in real applications, variance estimation is the crux of an SAE procedure. In the case of uncorrected inequality estimators, it may be easily carried out via linearization. Linearized variables for each measure could be derived consistently with Langel2013 for the Gini index and Biewen2006 for the Generalized Entropy and the Atkinson indexes. On the other hand, the variance of bias-corrected estimators adds a new level of complexity since the estimator formula is no longer the classical one. Indeed, it comprises a bias correction component that appears cumbersome to estimate via linearization since it is inherently a result of several linearizations. Therefore, in real applications, we recommend relying on resampling methods. An example of such a method is the design-aware bootstrap procedure developed by Fabrizi2011, fabrizi2020functional. A comprehensive review of the use of bootstrap methods for survey data can be found in lahiri2003impact.
The comparison between uncorrected and corrected survey estimates for all three measures is displayed in Figure (ref). Uncorrected estimates show lower values of inequality in comparison with the corrected ones for all areas and all the measures considered. This is in accordance with the underestimation highlighted by simulation results of Section (ref). The sampling coefficient of variations of both estimators are high, ranging from 0.24 to 3.48 for the Theil index, from 0.18 to 4.45 for the Atkinson index and, lastly from 0.11 to 0.92 for the Gini index depending on the area, with slightly higher values in case of corrected estimators as the bias correction induces mild variance inflation. Such values point out the need for SAE techniques.
The model-based (or EBLUP, Empirical Best Linear Unbiased Predictor) estimates in both cases are compared in Figure (ref). The inequality levels estimated by the misspecified model are lower, resulting in a misleading inference. In particular, this is quite evident in the Gini index case, where the divergence seems to increase at increasing levels of inequality. When the same SAE procedure is applied for all iterations referred to the simulation of Section (ref), a bias evaluation shows that the EBLUP estimated on uncorrected and corrected measures have an average RB of -17.9% and -13.4% respectively for the Generalized Entropy ($\alpha=1$); of -14.6% and -9.4% for the Atkinson index ($\varepsilon=1$) and of -7.7% and -0.4% for the Gini index. This confirms the risk of high underestimation of inequality when neglecting such an issue.
By focusing only on EBLUP results based on corrected estimates, the decrease in terms of error induced by the model is depicted in Figure (ref). The reduction is relevant and testifies that the variance reduction procedure, put in place by the SAE model, is effective. As a consequence, such model-based estimates result to be reliable and ready to be used for further analysis or mapping.
A strategy based on Taylor's expansion has been proposed to correct the small sample bias of inequality estimators. The inequality measures considered are several, as the comparison of diverse measures may enable us to enlighten the specific point of view that each measure provides, like single tiles in a mosaic. Indeed, the well-known Gini and Theil indexes are widely applied in several fields for inequality and concentration estimation.
A sensitivity analysis with respect to outliers and a simulation study have been conducted to study the estimator behaviour to extreme values and the performance of the proposed correction. Results show that survey-based estimators may be biased in small samples, inducing an underestimation that is even greater in the case of populations affected by extreme values. Moreover, simulation results validate the correction proposal as effective, consistently reducing the bias and leading in some cases to approximately unbiased estimators.
An underlined heterogeneity of sensitivities and bias is recorded across measures. As a consequence, our results may help in choosing the most suitable inequality measure depending on the context. The measures which are structurally more sensitive to extreme values appear to be more biased, in particular, GE with $\alpha=2$ and Atkinson with $\epsilon=2$. Therefore, in the case of samples without extreme income values, the bias correction may provide approximately unbiased estimates. On the other hand, if extreme values are observed, it becomes necessary to focus on the most robust measures such as Mean Log Deviation, Atkinson index with $\varepsilon=1$ and Gini Index to be corrected.
An illustrative small area application has been carried out to evaluate the effect of disregarding bias in a typical small-sized sample context. The results obtained show that neglecting it translates into a misleading inference and an inequality underestimation. This is particularly evident in the case of the popular Gini index. In such an application, we use a basic area-level model, the Gaussian one. Indeed, the possibly not-Gaussian sampling distributions of inequality estimators and the unit-interval support of Gini and Atkinson estimators might urge a more refined model, which may lead to model-based estimators with increased performances: this suggests an interesting direction for future research. Further directions also include the extension of this framework to other widely used inequality measures, such as those based on quintiles and the development of a multivariate SAE framework.
The work of Silvia De Nicolò was partially funded by the ALMA IDEA 2022 grant (title: "Social exclusion and territorial disparities: poverty and inequality mapping through advanced methods of small area estimation", project J45F21002000001), part of the European Union - NextGenerationEU funding.