EconBase
← Back to paper

Measuring wage inequality under right censoring

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

83,387 characters · 14 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Measuring wage inequality under right censoring

abstractIn this paper we investigate potential changes which may have occurred over the last two decades in the probability mass of the right tail of the wage distribution, through the analysis of the corresponding tail index. In specific, a conditional tail index estimator is introduced which explicitly allows for right tail censoring (top-coding), which is a feature of the widely used current population survey (CPS), as well as of other surveys. Ignoring the top-coding may lead to inconsistent estimates of the tail index and to under or over statements of inequality and of its evolution over time. Thus, having a tail index estimator that explicitly accounts for this sample characteristic is of importance to better understand and compute the tail index dynamics in the censored right tail of the wage distribution. The contribution of this paper is threefold: i) we introduce a conditional tail index estimator that explicitly handles the top-coding problem, and evaluate its finite sample performance and compare it with competing methods; ii) we highlight that the factor values used to adjust the top-coded wage have changed over time and depend on the characteristics of individuals, occupations and industries, and propose suitable values; and iii) we provide an in-depth empirical analysis of the dynamics of the US wage distribution's right tail using the public-use CPS database from 1992 to 2017. Keywords: Wage inequality, tail index, top-coding, current population survey, wage distribution, Pareto, occupations JEL classification: C18, C24, E24, J11, J31

\thispagestyle{empty}

\onehalfspace

Introduction

The sharp rise in overall wage inequality in the second half of the 20th century has become a stylized fact (Autor(2019) and Goosetal(2014)). Wage inequality growth in the 1980s was followed by a slowdown in the 1990s as a result of divergent trends in the bottom and top of the wage distribution. Both the 90/50 and 50/10 indexes grew rapidly in the early 1980s, and although lower tail inequality virtually stopped growing after 1987 upper-tail inequality kept rising. The deceleration in inequality growth observed in the 1990s resulted mainly from polarization, i.e., from an abrupt stop or reversal of inequality growth in the lower-tail coupled with a sustained secular rise of the upper-tail inequality. According to Autoretal(2008) between 1963 and 2005 the 90th percentile wage rose by more than 55% relatively to the 10th percentile for both men and women.

Existing empirical evidence suggests that the rise in wage inequality is largely explained by shifts in the supply and demand for skills (GoosManning(2007)), and by the erosion of labour market institutions (e.g. unions and minimum wage) (Kalleberg2011). It is documented that the increase in inequality in the 1980s was the result of a secular rise in the demand for skill which faced an abrupt slowdown in the relative supply of high-skilled workers (college or equivalent) in the form of lower attainment and of smaller labor-entering cohorts which originated expanding wage differentials (Autoretal(2008); KatzMurphy(1992); CardDiNardo(2002); and AcemogluAutor(2011)).

The monotonic increase of inequality until the late 1980s followed by the divergent evolution in the top and lower half of the distribution is robust to different measures and samples.\footnote{The result holds for male and female samples separately, considering weekly wages of full-time workers as well as for the March CPS samples (Autoretal(2006)).} Steady growth in the upper-tail inequality can also be seen from the rising share of wages paid to the top 10% and 1% earners (PikettySaez(2003)). However, literature based on public-use CPS data has produced a less than perfect picture of the right tail of the wage distribution because of the top-coding (Armouretal(2016)).

The CPS wage data has historically been censored at the top (top-coded) and ignoring this fact or not adequately handling it may result in inconsistent tail index estimates, lead to understatements of inequality and affect the estimates of its dynamics (Fengetal(2006)).\footnote{Parker(1999) developed a model of wages in which wages follow a Generalized Beta Distribution of the second kind (GB2). Bordleyetal(1996) show that GB2 exhibits better fit to US wage data than alternative distributions. Because the authors are modelling the distribution of total wage and not its components, they only need to know whether total wages are censored or not, and therefore do not need to be concerned with consistency problems in categories as in Burkhauseretal(2004). } In addition, top-coding has changed over time. For instance, the top-coded wage was set at $\$$1923 in 1997 and changed to $\$$2884 from 1998 onward. But even during periods of constant nominal top-coding the data may hide changes in inequality (LevyMurnane(1992)).

While some authors have tried to address the top-coding issues by restricting the sample under analysis, the method presented in this article makes use of the complete set of information available from the public use CPS data, for every year, in a time-consistent fashion, arguably providing better estimates on the level of wage inequality than other available measures. In specific, a conditional tail index regression specifically designed to account for right censoring is used. The tail index (sometimes also referred to as Pareto coefficient) is an important indicator, as it can be interpreted as an inverse measure of concentration of top wages. The lower the value of the index the more concentrated the distribution is.

Several estimation approaches have recently been proposed which consider either non-random or random covariates; see e.g. MaJiangHuang2019 (and references therein). Our contribution falls into the latter class and provides a tail index estimator which takes the top-coding explicitly into consideration, providing in this way more efficient and consistent estimates than methods currently available in the literature. The superior performance of the new approach is illustrated, and it is shown using the public-use CPS database from 1992 to 2017 that the factor values used for the adjustment of the top-coded wages changed over time and across the characteristics of individuals, occupations and industries; moreover it is also shown using the new estimator that the tail index has been decreasing since 1992 suggesting increased concentration in the right tail.

The contribution of this paper is threefold: i) we introduce a conditional tail index estimator that explicitly handles the top-coding problem, evaluate its finite sample performance and compare it to competing methods; ii) we show that the factor values used for the adjustment of top-coded wages change over time and across the characteristics of individuals, occupations and industries, and suitable values are proposed; and iii) we provide an in-depth analysis of the dynamics of the US wage distribution's right tail using the public-use CPS database from 1992 to 2017.

The remainder of the paper is organized as follows. Section 2 introduces the methodology of analysis, the new tail index estimator and a detailed description of the computation of the partial effects; Section 3 presents the results of an in-depth Monte Carlo analysis on the finite sample properties of the new approach and a comparison to existing procedures; Section 4 describes and discusses the results of an empirical analysis of the right tail characteristics of the wage distribution and wage inequality in the US using the CPS database from 1992 to 2017; and finally, in Section 5 presents the main conclusions of the paper. A technical Appendix collects proofs of the results put forward throughout the paper.

Methodology

To reduce the top-coding bias, researchers interested in measuring long-term trends in wages typically impute top-code values to create a consistent series. Until recently one of four approaches has in general been adopted in the literature: (1) the top-coding problem is ignored i.e., top-coded observations are dropped (see e.g. JensenShore(2015)); (2) an ad hoc adjustment of the top-coded wages is made (e.g. Lemieux(2006) multiplied top-coded hourly wages by 1.4, and Autoretal(2008) multiplied top-coded weekly wages by 1.5); (3) a Pareto distribution is used to estimate wages at the top of the distribution (e.g. BernsteinMishel(1997), PikettySaez(2003)); and (4) cell means or rank-proximity swapped data based on the still-censored internal CPS data is used (e.g. Larrimoreetal(2008) and Burkhauser et al., 2008); for a discussion and shortcomings of these approaches see, inter alia, Burkhauser et al. (2010) and Armouretal(2016).

In a recent contribution Armouretal(2016) proposed an alternative approach which consists in estimating the tail index of a censored Pareto distribution. To briefly illustrate the procedure consider first the survival function, $\overline{F}$, of a Pareto distribution\footnote{This distribution was used, for instance, by Harrisonl(1981) to analyse earnings by size in the UK. }

equation[equation omitted — 154 chars of source]

and the corresponding density function, $f_Y(y)=(\alpha y_0^{\alpha})/(y_i^{\alpha+1}). $ A large number of tail index estimation procedures is available in the literature. One widely used approach is the conditional maximum likelihood estimator (MLE) proposed by Hill(1975),

equation[equation omitted — 160 chars of source]

where $m$ is the number of largest order statistics used in the estimation of $\alpha$, $y_{(j)}, \text{ } j=1,...,m$ are the largest $m$ order statistics and $y_0$ is the tail cut off point.

However, recognizing the limitations of this approach when the data is top-coded, Armouretal(2016) proposed an alternative method, which consists of an adaptation of the Hill estimator taking into consideration the censoring. This approach provides an unbiased estimate of the censored Pareto parameter, $\alpha$, while using all available information. In specific, in the case of a censored sample the outcome variable is,

equation[equation omitted — 159 chars of source]

where $y_0$ is the tail cut off point and $y_c$ the top-coded value. Hence, the density function of the censored Pareto distribution is,

equation[equation omitted — 203 chars of source]

and the respective log-likelihood function,

eqnarray[eqnarray omitted — 282 chars of source]

Consequently, the conditional MLE estimator proposed by Armouretal(2016) computed from ((ref)) is,

equation[equation omitted — 139 chars of source]

where $n_0$ is the number of individuals with wages between $y_0$ and $y_c$, $n_c$ is the number of individuals with wages at or above $y_c$, and $n_0$ + $n_c$ = $m$.

The conditional tail index estimator and properties

In this paper, a new approach, also designed to overcome the top-coding bias, is proposed. In specific, a conditional tail index estimator which explicitly takes the right censoring of the data into account and uses covariates in the estimation process is introduced. Correctly estimating this tail index is of importance as it is used, for instance, for the imputation of wages above the top-code. Furthermore, the procedure has the additional advantage of allowing for an in-depth analysis of the determinants that impact the tail index the strongest according to the characteristics of individuals, occupations and industries and whether these impacts have changed over time. The use of different scaling factors to impute wages depending on the different categorizations of individuals has been used previously in the literature, see e.g., MacphersonHirsch(1995) who allow the scaling factors to vary according to gender and over the years.

To introduce the conditional tail index estimation approach we consider observations ($\mathbf{X}_{i},$ $Y_{i}$), where $ Y_{i}\in \mathbb{R}^{1}$ is the response of interest, and $\mathbf{X} _{i}:=(x_{1i},...,x_{pi})^{\prime }\in \mathbb{R}^{p}$ is an associated p-dimensional vector of predictors with $1\leq i\leq n$. In addition, let $F(y|\mathbf{x}; \boldsymbol{\theta}):=P[Y_{i}\leq y|\mathbf{X}_{i}=\mathbf{x}]$ be the cumulative distribution function of $Y_{i}$ conditional on $\mathbf{X}_{i}$, and assume that the corresponding survival function (under no censoring) is,

equation[equation omitted — 171 chars of source]

where $\alpha (\mathbf{x}):=\exp (\mathbf{x}^{\prime }\boldsymbol{\theta}),$ $ \boldsymbol{\theta}\in \mathbb{R}^{p}$ is the unknown vector of coefficients and $\mathcal{L}(y;\mathbf{x})$ is some predictor-dependent slowly varying function, such that $\mathcal{L}(yk;\mathbf{x})/\mathcal{L}(y;\mathbf{x} )\rightarrow 1$ for any $k>0$ as $y\rightarrow \infty .$ Specifically, following Hall(1982) we characterize the slowly varying function as,

equation[equation omitted — 144 chars of source]

where $c_{0}(\mathbf{x})$ and $c_{1}(\mathbf{x})$ are functions in $\mathbf{x}$ with $c_{0}(\mathbf{x})>0,$ $\beta(\mathbf{x})>\alpha(\mathbf{x})$ a positive function and $o(y^{-\beta (\mathbf{x})})$ is the higher-order remainder term. As a result, as $y \rightarrow \infty$, $\mathcal{L}(y;\mathbf{x}) \rightarrow c_{0}(\mathbf{x})$ and $\mathcal{\dot{L}}(y;\mathbf{x}) \rightarrow 0$, where $\mathcal{\dot{L}}(y;\mathbf{x}) = \partial \mathcal{L}(y;\mathbf{x})/\partial y.$

From ((ref)), it follows that the probability density function of $Y_{i}$ conditional on $\mathbf{X}_{i}$ is,

equation[equation omitted — 179 chars of source]

Considering ((ref)) and assuming that $y$ is sufficiently large, it follows that the density in ((ref)) can be approximated as,

equation[equation omitted — 125 chars of source]

see also WangTsai(2009). Thus, the conditional probability function of $Y_i$ given $\mathbf{X}_i$ and $Y_i > y_0$ can be approximated as,

equation[equation omitted — 115 chars of source]

where $y_0$ is the threshold that controls the sample fraction used for estimation. Note that ((ref)) is the approximate conditional Pareto density function of an unrestricted random variable and thus, its use when some form of censoring (such as right censoring\footnote{ An observation is said to be right censored at $y_c$ if the exact value of the observation is not known except that it is greater than or equal to $y_c$.} in the CPS database) is imposed on the data will originate inconsistent tail index parameter estimates.

In the censored case, rather than observing the outcome $y_{i}$, as in the previous section, we effectively observe $w_{i}$ as defined in ((ref)). In this context, the adequately adjusted conditional Pareto density function is,

equation[equation omitted — 261 chars of source]

where $I(.)$ is the indicator function and $f\left( \left. . \right\vert \mathbf{x}\right) $ and $F\left( \left. . \right\vert \mathbf{x}\right) $ correspond to the conditional Pareto density function and the conditional cumulative Pareto distribution function, respectively.

Hence, the negative log-transformed likelihood function for the top-coded data is,

equation[equation omitted — 172 chars of source]

where $g\left( \left. w_{i}\right\vert \mathbf{x}_{i}; y_c;\boldsymbol{\theta }\right)$ is as defined in $(\ref{log_g})$, $w_i$ is given in ((ref)) and $y_c$ is the censoring threshold.

Moreover, since

eqnarray[eqnarray omitted — 605 chars of source]

if we replace $\alpha \left( \mathbf{x}_{i}\right) =\exp \left( \mathbf{x} _{i}^{\prime }\boldsymbol{\theta }\right) $ the approximate negative log-likelihood function in ((ref)) (omitting for simplicity of notation the terms not related to $\boldsymbol{\theta} $) becomes,

eqnarray[eqnarray omitted — 469 chars of source]

Hence, we see from this approximate log-likelihood function that censuring the data imposes a penalty term, which the unrestricted estimator does not take into consideration.

To derive the limit distribution of the parameter estimators and corresponding test statistics we consider, as in WangTsai(2009), the following assumptions:

Assumption A:

enumerate$n_{0}^{-1}\sum \limits_{i=1}^{n}\mathbf{Z}_{ni}\mathbf{Z} _{ni}^{\prime }I\left( w_{i}\geq y_{0}\right) =\mathbf{\Sigma}_{y_0}^{-1/2}\widehat{\mathbf{\Sigma}}_{y_{0}}\mathbf{\Sigma}_{y_0}^{-1/2}\overset{p}{\rightarrow }\mathbf{I}_{p}$, where $\mathbf{Z}_{ni}:=\mathbf{\Sigma}_{y_0}^{-1/2}\mathbf{x}_i$, $\mathbf{I}_p$ is a $p \times p$ identity matrix and $\mathbf{\widehat{\Sigma}}_{y_{0}}:=n_0^{-1}\sum (\mathbf{x}_i\mathbf{x}'_i)I(w_i \geq y_0).$ • (Slowly varying function) We assume that the remainder term $ o(y^{-\beta (\mathbf{x})})$ satisfies $\sup_{\mathbf{x}}y^{\beta (\mathbf{x} )}o(y^{-\beta (\mathbf{x})})\rightarrow 0$ as $y\rightarrow \infty .$

The following theorem characterizes the limit distribution of the MLE estimates of $\boldsymbol{\theta}$.

theoremUnder Assumptions (A1) - (A2) it follows that \begin{equation*} n^{-1/2}\Sigma _{y_{0}}^{-1/2}\Lambda^{-1/2} (\widehat{\boldsymbol{\theta}}-\boldsymbol{\theta} _{0}) \overset{d}{\rightarrow }N(\mathbf{0,I}_{p}), \end{equation*} where $\Lambda:=E(e_i^2|\mathbf{x}_i)$ and $ e_{i}=\left\{ \begin{array}{ll} \exp \left( \mathbf{x}_{i}^{\prime }\boldsymbol{\theta} \right) \log \left( \frac{ w_{i}}{y_{0}}\right) -1 ,\text{ for } & I_{\left\{ y_{0\leq }w_{i}<y_c\right\} } \\ -\exp \left( \mathbf{x}_{i}^{\prime }\boldsymbol{\theta} \right) \log \left( \frac{y_{0}}{ y_c}\right),\text{ for } & I_{\left\{ w_{i}=y_c\right\} } \end{array}. \right. $
corollaryUnder the same conditions of Theorem 2.1 as $n \rightarrow \infty$ it follows that, \begin{equation} T_{j}=n^{-1/2}\left(\Sigma _{y_{0},jj}^{-1/2}\right)\Lambda^{-1/2} \widehat{\theta}_{j}\overset{d}{\rightarrow }N(0,1). \end{equation} where $\Sigma _{y_{0},jj}^{-1/2}$ corresponds to the $(j,j)^{th}$ element of the $\boldsymbol{\Sigma} _{y_{0}}^{-1/2}$ matrix.

Computation of partial effects

A further important and not immediately obvious aspect of the methodology just described relates to the computation of the partial effects of the covariates used in the conditional tail index regression. In specific, for ease of presentation consider

equation*[equation* omitted — 148 chars of source]

where for the sake of simplicity but with no loss of generality we consider $x$ to be a scalar and continuous. Thus, to measure the impact of $x$ on $\alpha \left(x\right) $ and subsequently on $\overline{F}(y|x; \boldsymbol{\theta})$, consider

equation[equation omitted — 263 chars of source]

where $y$ is an extreme value, say the ($1-u$) quantile, with $u \in (0,1)$, such that,

equation[equation omitted — 78 chars of source]

Hence, $\delta$ in ((ref)) measures the probability's percentage variation of an extreme value due to a variation of $x$, $\Delta x$. For example, considering $ u=0.15$ and $\Delta x=1,$ if $\delta =20\%$ then P$\left( Y>y_0\right) ,$ where $y_0$ is the 0.85 quantile, increases by 20% as a result of $\Delta x=1$. Therefore, the variation of $x$ increases the likelihood of observing extreme values by 20%.

For computational purposes, assuming that $\alpha \left( x\right) :=\exp \left( \phi \left( x\right)\right)$, and $\phi(x)$ is some function of $x$, we show in the appendix that

equation[equation omitted — 140 chars of source]

For instance, in the multivariate case, $\phi \left( \mathbf{x}\right) :=\mathbf{x}^{\prime }\mathbf{\beta}$, where $\mathbf{x}$ is a $p\times1$ vector of covariates, $\alpha \left( \mathbf{x}\right) =\exp \left( \mathbf{x}^{\prime }\mathbf{\beta }\right)$, and the impact of $x_{j}$ on $\alpha(\mathbf{x})$ is,

equation[equation omitted — 124 chars of source]

Thus, a negative (positive) coefficient increases (decreases) the likelihood of having more extreme values, i.e. $\beta _{j}<0$ ($\beta _{j}>0$) implies $\delta _{j}>0$ ($\delta _{j}<0$) (this is also obvious from the impact on $\alpha(\mathbf{x})$ since if $\alpha(\mathbf{x})$ decreases (increases), the right tail becomes (less) heavier).

remarkIf we use only a portion of the sample to estimate the model, say for example, all observations larger than $y_0,$ then $y$ is the quantile of order $ 1-u$ of the conditional distribution $P\left( \left. Y<y\right\vert Y>y_0\right) ,$ i.e. $P\left( \left. Y<y\right\vert Y>y_0\right) =1-u.$ To determine the quantile order of the unconditional distribution $P\left( Y<y\right) $ we use the relation $P\left( \left. Y<y\right\vert Y>y_0\right) =1-u\Rightarrow P\left( Y<y\right) =P\left( Y<y_0\right) +\left( 1-u\right) P\left( Y>y_0\right) .$ In the empirical analysis below we use all observations larger than the empirical quantile of 0.80 (see also Misheletal(2013)), so that $P\left( Y>y_0\right) =0.20$ in the above formula. Hence, when $u=0.15$ and $u=0.20,$ we are actually analysing the 96th and 97th quantile, respectively, of the unconditional distribution.
remarkWhen the covariate considered is discrete (e.g. a dummy variable) a simple adaptation of ((ref)) leads to the following formula, which measures the impact of group $d=1$ over $d=0$, $ \delta \left( u\right) :=\left[ \left( 1-u\right) ^{\frac{\alpha \left( \mathbf{x};d=1\right) }{\alpha \left( \mathbf{x};d=0\right) } -1}-1\right] \times 100.$ Given that $\alpha \left( \mathbf{x};d=1\right) $ and $\alpha \left( \mathbf{\ x};d=0\right) $ depend on $\mathbf{x},$ we also need to provide values for $\mathbf{x}$. One possible solution is to replace $\mathbf{x}$ by its respective averages.

Monte Carlo simulation

In this section we evaluate the finite sample properties of the procedures and their performance in imputing mean wages above the top-code.

Finite sample performance of tail index estimators

To evaluate the performance of the conditional tail index estimator introduced in the previous section, we conduct an in-depth Monte Carlo analysis using several data generation processes (DGPs). In specific, data is simulated from the general framework,

eqnarray[eqnarray omitted — 230 chars of source]

where $\beta _{1}=\beta _{2}=1$ and the $k100$%, $k\in(0,1)$, largest observations of the empirical distribution $D(.)$ closely follow a Pareto distribution. We consider the case of right censoring given by the censoring threshold $y_c$ so that the sequence $\left\{ y_{i}\right\} $ is not completely observed. Instead, we observe $ w_{i}=\min \left( y_{i},y_c\right)$.

To be more precise about the framework used to generate the data, we consider that $D(.)$ in ((ref)) is either a Pareto or a Burr distribution\footnote{The cumulative Burr distribution function considered in the simulations is $F(x):=1-(1+x^{-\alpha \rho})^{\frac{1}{\rho}}$ and the corresponding probability density function $f(x):=x^{-1-\alpha \rho}(1+x^{-\alpha \rho})^{-1\frac{1}{\rho}}\alpha$.} and generate samples of size $n\in\{2500, 5000, 10000, 50000\}$. Moreover, we censor the sample considering $y_c=\{\hat{q}_{0.95}^y, \hat{q}_{0.99}^y\}$ which corresponds to the 95th and 99th empirical quantile of $y$. For estimation of the tail index we use the $\lfloor kn\rfloor$ largest observations, with $k=0.2$ when the Pareto distribution is considered and $k=\{0.05, 0.10, 0.20\}$ for the Burr. In the case of samples generated from a Pareto distribution we could have set $k=1$, however, a lower value was considered in order to mimic the conditions typically found in empirical analysis.

Based on the specifications described above we generated 10,000 sequences of $\left\{ y_{i}\right\} $ and $\left\{w_{i}\right\} $ of size $n$ and use in each iteration three estimation methods:

itemize• The tail index regression of WangTsai(2009) applied to the sequence of $\left\{ y_{i}\right\}$. We define the resulting estimator as $\hat{\alpha}$. This method should provide the best results since it is applied to the original uncensored data. • The censored tail index regression introduced in this paper applied to the sequence $\left\{w_{i}\right\} $. The resulting estimator is denoted as $\hat{\alpha}^c$. • The tail index regression of Wang and Tsai (2009) applied to the censored data $\left\{w_{i}\right\}$. The resulting tail index is defined as $\tilde{\alpha}$. This approach will be useful in providing information on the impact of neglecting the censoring on the tail index estimates.

Table 1 provides the bias and RMSEs associated with the estimates of $\beta_1$ and $\beta_2$ in ((ref)) computed based on the three approaches described in i), ii) and iii). The first observation we can make is that, in general, the largest bias and RMSEs (regardless of considering $\beta_1$ or $\beta_2$) result from the use of the approach described in iii), i.e., when the censoring is ignored. On the other hand, it is interesting to observe that the difference in the bias and RMSEs obtained from the approaches described in i) and ii) are relatively small, which suggest that the estimation approach which accounts for the censoring produces results close to those obtained when the sample without censoring is used for estimation as is the case in i).

Moreover, this Table also shows that the bias remains relatively stable and does not decrease as $n$ increases. There are however different patterns according to the values of $k$ and $y_{c}.$ For instance, in Cases 3 to 6, which use the Burr distribution as DGP, a small value of $k$ tends to improve the estimation results given that the tail of the Burr distribution gets closer to the tail of a Pareto distribution.

To further evaluate the estimation performance of the three estimation approaches in i), ii) and iii), Figure 1 plots the ratios of the RMSEs of the $\alpha$ estimates obtained under the these estimations approaches. In specific, the ratios are, \[ Ratio\text{ }1=\frac{RMSE\left( \hat{\alpha}^{c}\left( \mathbf{x}_{i}\right)\right) }{RMSE\left( \hat{ \alpha}\left( \mathbf{x}_{i}\right)\right) },\qquad Ratio\text{ }2=\frac{RMSE\left( \tilde{\alpha}\left( \mathbf{x}_{i}\right) \right) }{RMSE\left( \hat{\alpha}\left( \mathbf{x}_{i}\right)\right) }. \] Since $\hat{\alpha}$, obtained as described in i), is the best estimator, $Ratio$ $1$ and $Ratio$ $2$ are larger than 1, across the different values of $n.$ However, Ratio 1 is just slightly above 1, which means that the censored estimator performs very well and mimics closely the behavior of the best estimator, $\hat{\alpha}$, although the former is based on the censored data. On the contrary, Ratio 2 is substantially higher than 1, which means that, ignoring the censoring when estimating the tail index produces an inconsistent estimator; see Figure 1.

The censoring threshold $y_{c}$ also impacts the estimation results, i.e., the lower its value, the greater is the impact of censoring on estimation, and Ratio 2 tends to be larger (see, for example, the results for Case 4 in Table 1 and Figure 1).

landscape\begin{table}[htbp] \caption{Bias and RMSE of estimators} \begin{center} \scalebox{0.7}{ \begin{tabular}{lSSSSSSSSSSSS} \toprule \multicolumn{13}{c}{Case 1: DGP Pareto (k=0.20 and $\mathbf{y_c = Q_y(0.95)}$)} \\ & \multicolumn{1}{c}{$\hat{\alpha}\left( \mathbf{x}_{i}\right)$} &\multicolumn{1}{c}{$\hat{\alpha}^c\left( \mathbf{x}_{i}\right)$} & \multicolumn{1}{c}{$\tilde{\alpha}\left( \mathbf{x}_{i}\right)$} & \multicolumn{1}{c}{$\hat{\alpha}\left( \mathbf{x}_{i}\right)$} &\multicolumn{1}{c}{$\hat{\alpha}^c\left( \mathbf{x}_{i}\right)$} & \multicolumn{1}{c}{$\tilde{\alpha}\left( \mathbf{x}_{i}\right)$}& \multicolumn{1}{c}{$\hat{\alpha}\left( \mathbf{x}_{i}\right)$} &\multicolumn{1}{c}{$\hat{\alpha}^c\left( \mathbf{x}_{i}\right)$} & \multicolumn{1}{c}{$\tilde{\alpha}\left( \mathbf{x}_{i}\right)$}& \multicolumn{1}{c}{$\hat{\alpha}\left( \mathbf{x}_{i}\right)$} &\multicolumn{1}{c}{$\hat{\alpha}^c\left( \mathbf{x}_{i}\right)$} & \multicolumn{1}{c}{$\tilde{\alpha}\left( \mathbf{x}_{i}\right)$} \\ \midrule n & \multicolumn{3}{c}{2500} & \multicolumn{3}{c}{5000} & \multicolumn{3}{c}{10000}&\multicolumn{3}{c}{50000} \\ bias($\beta_1$) & 0.0025 & 0.0004 & 0.4553 & 0.0011 & -0.0002 & 0.4539 & 0.0004 & -0.0005 & 0.4528 & 0.0001 & -0.0002 & 0.4526 \\ bias($\beta_2$) & 0.0029 & 0.0024 & -0.4241 & 0.0027 & 0.0034 & -0.4237 & 0.0014 & 0.0020 & -0.4246 & 0.0003 & 0.0007 & -0.4257 \\ RMSE($\beta_1$)& 0.0775 & 0.0937 & 0.4611 & 0.0552 & 0.0660 & 0.4567 & 0.0386 & 0.0465 & 0.4542 & 0.0174 & 0.0208 & 0.4529 \\ RMSE($\beta_2$)& 0.1703 & 0.1942 & 0.4422 & 0.1208 & 0.1369 & 0.4329 & 0.0852 & 0.0963 & 0.4292 & 0.0378 & 0.0430 & 0.4267 \\ \multicolumn{13}{c}{Case 2: DGP Pareto (k=0.20 and $\mathbf{y_c = Q_y(0.99)}$)} \\ \midrule n & \multicolumn{3}{c}{2500} & \multicolumn{3}{c}{5000} & \multicolumn{3}{c}{10000}&\multicolumn{3}{c}{50000} \\ bias($\beta_1$) & 0.0025 & -0.0012 & 0.1000 & 0.0011 & -0.0010 & 0.0984 & 0.0004 & -0.0005 & 0.0980 & 0.0001 & -0.0001 & 0.0977 \\ bias($\beta_2$)& 0.0029 & 0.0070 & -0.1202 & 0.0027 & 0.0050 & -0.1202 & 0.0014 & 0.0023 & -0.1219 & 0.0003 & 0.0005 & -0.1229 \\ RMSE($\beta_1$) & 0.0775 & 0.0807 & 0.1258 & 0.0552 & 0.0575 & 0.1126 & 0.0386 & 0.0401 & 0.1051 & 0.0174 & 0.0182 & 0.0992 \\ RMSE($\beta_2$)& 0.1703 & 0.1745 & 0.2006 & 0.1208 & 0.1240 & 0.1659 & 0.0852 & 0.0871 & 0.1460 & 0.0378 & 0.0389 & 0.1281 \\ \multicolumn{13}{c}{Case 3: DGP Burr $\rho$=-2 (k=0.05 and $\mathbf{y_c = Q_y(0.99)}$)} \\ \midrule n & \multicolumn{3}{c}{2500} & \multicolumn{3}{c}{5000} & \multicolumn{3}{c}{10000}&\multicolumn{3}{c}{50000} \\ bias($\beta_1$) & 0.0015 & -0.0118 & 0.3277 & -0.0006 & -0.0084 & 0.3241 & -0.0025 & -0.0065 & 0.3229 & -0.0042 & -0.0061 & 0.3210 \\ bias($\beta_2$)& 0.0388 & 0.0488 & -0.3189 & 0.0186 & 0.0242 & -0.3356 & 0.0134 & 0.0161 & -0.3409 & 0.0073 & 0.0097 & -0.3452 \\ RMSE($\beta_1$)& 0.1450 & 0.1683 & 0.3562 & 0.1021 & 0.1180 & 0.3387 & 0.0719 & 0.0821 & 0.3301 & 0.0323 & 0.0372 & 0.3224 \\ RMSE($\beta_2$) & 0.4088 & 0.4518 & 0.4549 & 0.2833 & 0.3121 & 0.4043 & 0.1991 & 0.2176 & 0.3754 & 0.0891 & 0.0976 & 0.3523 \\ \multicolumn{13}{c}{Case 4: DGP Burr $\rho$=-2 (k=0.10 and $\mathbf{y_c = Q_y(0.95)}$)} \\ \midrule n & \multicolumn{3}{c}{2500} & \multicolumn{3}{c}{5000} & \multicolumn{3}{c}{10000}&\multicolumn{3}{c}{50000} \\ bias($\beta_1$) & -0.0099 & -0.0239 & 0.9137 & -0.0120 & -0.0258 & 0.9098 & -0.0129 & -0.0259 & 0.9079 & -0.0137 & -0.0259 & 0.9071 \\ bias($\beta_2$) & 0.0297 & 0.0380 & -0.6485 & 0.0250 & 0.0402 & -0.6489 & 0.0216 & 0.0372 & -0.6512 & 0.0200 & 0.0371 & -0.6519 \\ RMSE($\beta_1$)& 0.1055 & 0.1596 & 0.9196 & 0.0754 & 0.1129 & 0.9127 & 0.0537 & 0.0820 & 0.9094 & 0.0271 & 0.0434 & 0.9074 \\ RMSE($\beta_2$)& 0.2609 & 0.3563 & 0.6637 & 0.1835 & 0.2492 & 0.6561 & 0.1305 & 0.1773 & 0.6548 & 0.0602 & 0.0857 & 0.6527 \\ \multicolumn{13}{c}{Case 5: DGP Burr $\rho$=-2 (k=0.20 and $\mathbf{y_c = Q_y(0.99)}$)} \\ \midrule n & \multicolumn{3}{c}{2500} & \multicolumn{3}{c}{5000} & \multicolumn{3}{c}{10000}&\multicolumn{3}{c}{50000} \\ bias($\beta_1$) & -0.0364 & -0.0435 & 0.0602 & -0.0378 & -0.0433 & 0.0586 & -0.0385 & -0.0428 & 0.0582 & -0.0386 & -0.0421 & 0.0581 \\ bias($\beta_2$)& 0.0487 & 0.0578 & -0.0733 & 0.0486 & 0.0558 & -0.0732 & 0.0474 & 0.0531 & -0.0748 & 0.0460 & 0.0510 & -0.0762 \\ RMSE($\beta_1$)& 0.0847 & 0.0909 & 0.0966 & 0.0665 & 0.0716 & 0.0798 & 0.0542 & 0.0584 & 0.0693 & 0.0422 & 0.0458 & 0.0606 \\ RMSE($\beta_2$)& 0.1737 & 0.1807 & 0.1739 & 0.1283 & 0.1342 & 0.1344 & 0.0961 & 0.1009 & 0.1089 & 0.0592 & 0.0638 & 0.0840 \\ \multicolumn{13}{c}{Case 6: DGP Burr $\rho$=-2 (k=0.20 and $\mathbf{y_c = Q_y(0.95)}$)} \\ \midrule n & \multicolumn{3}{c}{2500} & \multicolumn{3}{c}{5000} & \multicolumn{3}{c}{10000}&\multicolumn{3}{c}{50000} \\ bias($\beta_1$) & -0.0364 & -0.0551 & 0.4134 & -0.0378 & -0.0557 & 0.4119 & -0.0385 & -0.0559 & 0.4108 & -0.0386 & -0.0553 & 0.4109 \\ bias($\beta_2$) & 0.0487 & 0.0702 & -0.3813 & 0.0486 & 0.0713 & -0.3808 & 0.0474 & 0.0699 & -0.3817 & 0.0460 & 0.0682 & -0.3831 \\ RMSE($\beta_1$)& 0.0847 & 0.1084 & 0.4196 & 0.0665 & 0.0864 & 0.4150 & 0.0542 & 0.0727 & 0.4124 & 0.0422 & 0.0591 & 0.4112 \\ RMSE($\beta_2$)& 0.1737 & 0.2040 & 0.4010 & 0.1283 & 0.1532 & 0.3908 & 0.0961 & 0.1182 & 0.3866 & 0.0592 & 0.0804 & 0.3841 \\ \bottomrule \end{tabular}} \end{center} {\textbf{Note}: $\hat{\alpha}\left( \mathbf{x}_{i}\right)$ is the tail index regression estimate considering the complete sample of data (with no censoring); $\hat{\alpha}^c\left( \mathbf{x}_{i}\right)$ is the censored tail index regression estimate; and $\tilde{\alpha}\left( \mathbf{x}_{i}\right)$ is the uncensored tail index regression estimate computed from censored data. $n$ corresponds to the total sample size, $k$ to the % of observations used for the tail index estimation and $\lfloor kn \rfloor$ is the effective number of observations used in the estimation of the tail index. $y_c$ is the censoring value used and $Q_y(\tau)$ corresponds to the $\tau^{th}$ quantile of $y$.} \end{table}

\FloatBarrier

figure[figure omitted — 1,511 chars of source]

\FloatBarrier

Imputing mean wages

To provide further insights on the usefulness of the procedure introduced in this paper we provide next an analysis of the performance of the different methods for imputing mean wages. In specific, we compare the following three methods:

itemize• the Pareto-imputed mean wage above the top-code $y_c$, \begin{equation} \hat{\tau}_{1}\left( y_{c}\right) =\frac{\hat{\alpha}_{1}}{\hat{\alpha}_{1}-1 }y_{c} \end{equation} where $\hat{\alpha}_{1}$ is the tail index estimate considering an uncensored Pareto distribution as in Section 2 (see e.g. Hill(1975) and NicolauRodrigues(2019) for tail index estimators) and $y_{c}$ is the top-code threshold; • the imputed mean wage based on the approach suggested by Armouretal(2016) \begin{equation} \hat{\tau}_{2}\left( y_{c}\right) =\frac{\hat{\alpha}_{Hill}^c}{\hat{\alpha}_{Hill}^c-1 }y_{c} \end{equation} where $\hat{\alpha}_{Hill}^c$ is a consistent estimate of the tail index parameter computed as in ((ref)). Note that $\tau_{2}\left( y_{c}\right) :=E\left(y_i|y_i>y_c\right)$ ; • the imputed mean wage based on the method introduced in this paper \begin{equation} \hat{\tau}_{3}\left( y_{c}\right) =\frac{\hat{\alpha}\left( \mathbf{x} _{i}\right) }{\hat{\alpha}\left( \mathbf{x}_{i}\right) -1}y_{c} \end{equation} where $\hat{\alpha}\left( \mathbf{x}_{i}\right) =\exp \left( \mathbf{x} _{i}^{\prime }\boldsymbol{\hat{\theta}}\right) .$ Note that $\tau _{3}\left( y_{c}\right) :=E\left( \left. y_{i}\right\vert \mathbf{x} _{i}, y_i>y_c\right) .$

The main difference between $ \hat{\tau}_{3}\left( y_{c}\right)$ and the other two approaches ($ \hat{\tau}_{1}\left( y_{c}\right)$ and $\hat{\tau}_{2}\left( y_{c}\right)$) is that in the former the particular characteristics of the individuals whose wage is above the threshold are taken into account through $\mathbf{x}_{i}$. Interestingly, the estimator $\hat{\tau}_{3}\left( y_{c}\right)$ corresponds to the optimal mean square predictor because $\tau_{3}\left( y_{c}\right)$ is the conditional expectation of $y_{i}$ given $\mathbf{x}_{i}$ and $y_i>y_c$. This follows from the well known result $E \left( y_{i}-E\left( \left. y_{i}\right\vert \mathbf{x}_{i},y_i>y_c\right) \right) ^{2} \leq E \left( y_{i}-g\left( . \right) \right) ^{2}$ where $g\left( . \right) $ is any other predictor of $y_{i}$ given $y_i>y_c$. It turns out that $\tau_{2}\left( y_{c}\right)$ is optimal only if $y_{i}$ is mean-independent of $\mathbf{x}_{i},$ in which case both $\tau_{2}\left( y_{c}\right)$ and $\tau_{3}\left( y_{c}\right)$ coincide.

Thus, having established the superiority of $\tau_{3}\left( y_{c}\right)$, it remains to be shown how much improvement is provided by $\tau_{3}\left( y_{c}\right)$ compared to $\tau_{1}\left( y_{c}\right)$ and $\tau_{2}\left( y_{c}\right)$ when computing the mean wage above $y_c$. The following Monte Carlo study tries to answer this question. Our experiment is based on the following steps:

enumerate• Select a sample size from \[ n \in \left\{ 250,500,1000,2000,5000,20000\right\}; \] • Simulate $y_{i},$ $i=1,2,...,n$ according to a conditional Pareto distribution $ P\left( \alpha \left( \mathbf{x}_{i}\right) \right) $ where $\alpha \left( \mathbf{x}_{i}\right) =\exp \left( 1+2x_{i}\right) $ and $x_{i}\sim U\left( 0,1\right) .$ For each $i=1,2,...,n$ simulate $\alpha \left( \mathbf{x}_{i}\right) $ and then the corresponding $y_{i}$; • All observations above quantile $0.95$, which corresponds to $y_{c},$ are censored, but their original values are saved for comparison purposes (these values are used to assess the estimators' predictive precision). The data used for estimation are $\left\{ w_{i},i=1,2,...,n\right\} $ where $w_{i}=\min \left( y_{i},y_{c}\right) ,$ from which we estimate $\hat{\tau}_{1}\left( y_{c}\right) ,$ $\hat{\tau} _{2}\left( y_{c}\right) $ and $\hat{\tau}_{3}\left( y_{c}\right) .$ • The estimators $\hat{\tau}_{1}\left( y_{c}\right) ,$ $\hat{\tau} _{2}\left( y_{c}\right) $ and $\hat{\tau}_{3}\left( y_{c}\right) $ are used to predict the imputed mean value above the threshold $y_{c}.$ • Steps 2 to 4 are repeated 2000 times and the mean square errors (MSE) of $\hat{\tau }_{1}\left( y_{c}\right) ,$ $\hat{\tau}_{2}\left( y_{c}\right) $ and $\hat{ \tau}_{3}\left( y_{c}\right) $ are computed.

Other combinations of $\alpha \left( \mathbf{x}_{i}\right) $ produce essentially the same results as long as $\alpha \left( \mathbf{x}_{i}\right) \geq 1$, and for this reason we present results only for the case $\alpha \left( \mathbf{x }_{i}\right) =\exp \left( 1+2x_{i}\right) $ (note that $E\left( \alpha \left( \mathbf{x}_{i}\right) \right) =3.7).$ However, $\alpha \left( \mathbf{x}_{i}\right) $ should be set so that $\alpha\left( \mathbf{x}_{i}\right) >1,$ otherwise the conditional and marginal expected values do not exist, and consequently none of the above estimators will be well defined. The case $0<\alpha \left( \mathbf{x}_{i}\right) <1$ should be dealt with using other estimators, such as, for example, the median, $\sqrt[{\hat{\alpha}\left(\mathbf{x}_{i}\right)}]{2}y_{c}$).

Figure 2 illustrates our results. The lines represent two MSE ratios, $Ratio\text{ } 1:=MSE\left( \hat{\tau}_{1}\right) /MSE\left( \hat{\tau}_{3}\right)$ and $Ratio\text{ }2:=MSE\left( \hat{\tau}_{2}\right) /MSE\left( \hat{\tau}_{3}\right)$, computed over different sample sizes (n = (250, 500, 1000, 2000, 5000, 10000)).

figure[figure omitted — 192 chars of source]

\FloatBarrier

As can be observed from Figure (ref), the $\hat{\tau}_{3}(y_c)$ estimator introduced in this paper produces the best results as all MSE ratios are larger than one. The gains are modest (between 1% and 2%) when the sample size is small (n=250), but they increase steadily as the sample size increases. Another conclusion, is that $\hat{\tau}_{2}$ (Armouretal(2016)) is better than the (naive) $\hat{\tau}_{1}$ estimator that does not accommodate the censoring.

Empirical analysis

In this section, we use censored publicly available CPS data to evaluate how the right tail index of the US wage distribution has changed over time and how these changes may differ across the characteristics of individuals, occupations and industries. In specific, we show that the new tail index estimator introduced provides very rich and detailed insights about the right tail distribution of wages. We also assess the sensitivity of the adjustment of the top coded wage to changes over time and across the characteristics of individuals.

Data

For the empirical analysis the March CPS files from IPUMS for the period between 1992 and 2017 are used. The wage measure is top-coded at $\$$1923 between 1989 and 1997, and at $\$$2884 between 1998 and 2017. The sample is restricted to workers between 16 to 64 year-old on full-time full year basis employed during the CPS sample survey reference week (35+ hours per week, 40+ weeks per year). Following \citet{AutorDorn(2013)} the real weekly wage data are weighted by the appropriate CPS weight to provide a measure of the full distribution of weekly wages paid.\footnote{Wages are converted to 2012 values using the GDP personal consumption expenditure deflator.}$^,$\footnote{Workers in our sample come from outgoing rotation groups 4 and 8 and according to Unicon: When the Outgoing Rotation files are produced, two rotations are extracted from each of the twelve months and gathered into a single annual file. The weights on the file must be modified by the user before giving reliable counts. Since the final weight is gathered from 12 months but only 2/8 rotations, the weight on the outgoing file should be divided by 3 (12/4) before it is applied. The earner weight is gathered from 12 months from the 2 rotations. Since those two rotations were originally weighted to give a full sample, the earner weight must be divided by 12, not 3.}

Occupations are defined as job task requirements of the US Department of Labor Dictionary of Occupational Titles (DOT, 1977) and Census occupation classifications for routine, abstract and manual task classifications (AutorDorn(2013)).

In Figure (ref) we present the distribution of the weekly wages for 1992, 1997, 1998, 2007, 2010 and 2017. From 1992 to 2017 the concentration of wages has become more skewed to the right. Between 1992 and 1997 the mass points around the top-code increased and with the relaxing of the top-code in 1998 real values beyond that top-code are potentially observable. The same phenomenon also occurs in the most recent period. In 2017 the mass point around the current top-code used in the CPS data is much larger.

figure[figure omitted — 686 chars of source]

\FloatBarrier

The means and proportions of workers according to their characteristics, occupations and industry, for observations above the 80th percentile, are presented in Table (ref). In contrast to 1992, in 2017 the population in the right tail is older (41.03 years on average in 1992 and 43.83 years in 2017), the percentage of women is larger (25% in 1992 increased to 33% in 2017), and is about one year more educated (15.20 years of education in 1992 and 16.06 years in 2017).

table[table omitted — 2,786 chars of source]

\FloatBarrier

The results in Table (ref) show that the 7% decrease of white individuals in the right tail in 2017 when compared to 1992 seems to be compensated by a 7% increase of individuals from other races (non-white nor non-black). There is a 3% change in the composition of the sample, where the reduction of married individuals is compensated by an increase in individuals that are single. However, married still represent the marital status of the majority of individuals in the right tail (83% in 1992 and 81% in 2017).

A further observation that can be made from the results in Table (ref) is job polarization. A significant growth in employment in the right tail for non-routine cognitive tasks is observed in detriment of routine occupations (see also Autor(2019) and GoosManning(2007)). The decrease in the share of employment is even more significant for non-routine manual tasks (individuals in occupations associated with transportation, construction and mechanics (Transportation) decreased their share of employment in the right tail, from 11% in 1992 to 8% in 2017). The percentage of individuals in occupation Managers, which is the occupation with the largest proportion of individuals, increased from 72% in 1992 to 81% in 2017.

In 1992, the proportion of individuals, working in manufacturing (transports) was 22% (12%) and this proportion decreased to 14% (8%). This decrease in the share of employment was compensated by an increase in the repair (+6% between 1992 and 2017) and finance, personal and public industries (+6% between 1992 and 2017). Note that the industries with the largest number of individuals in the right tail in (1992, 2017) are Personal (28%, 32%), Manufacturing (22%, 14%), Finance (9%, 10%), Repair (4%, 10%) and Public (8%, 9%). However, Manufacturing (22%, 14%), Transport (11%, 8%) and Trade (11%, 9%) see their weight decrease in 2017.

To illustrate the evolution of the proportions of individuals in the different percentiles of the overall wage distribution, Figure 4 plots the proportions in percentiles 0.05 to 0.95 considering different attributes, occupations and industries (in the appendix we present additional plots for all other cases analysed in Table (ref)). From this Figure we distinguish two patterns from 1992 to 2017: Other Race, Female, Single and Personal increase in proportion from 1992 to 2017 across all percentiles, whereas Administration and Trade decrease across all percentiles. The number of Black individuals seems to decrease up to around percentile 80 and increases thereafter.

Moreover, this Figure also shows that individuals that are Black, Female or Single as well as individuals working in Administration and Trade display a downward trend over the percentiles, whereas the number of individuals of Other Races and individuals working in the Personal industry display a different pattern. The former is relatively constant across all percentiles in 1992 but increases for percentiles larger than the median in 2017, and the latter is relatively constant across all percentiles in 1992 and 2017. Interestingly, Finance shows a different pattern when compared to all other covariates. In specific, the largest proportions are observed at the higher percentiles (i.e. from the median onward). This pattern is very similar across all years analysed.

figure[figure omitted — 1,281 chars of source]

\FloatBarrier

Conditional tail index estimation results

Table 3 presents the censored and uncensored tail index regression results for 1992 and 2017. A negative (positive) regression coefficient corresponds to a decrease (increase) in the tail index (ceteris paribus) and hence a larger (smaller) number of extreme values may result as a consequence of changes in the specific variable associated to such a coefficient. Before analyzing the partial effects, as described in Section 2.2, the behavior of a tail index estimate computed as an average of the conditional tail indexes is examined, i.e.,

equation[equation omitted — 142 chars of source]

where $\hat{\alpha}\left( \mathbf{x}_{t,i}\right) =\exp \left( \mathbf{x} _{t,i}^{\prime }\mathbf{\hat{\theta}}_{t}\right) $ (we only report results for 1992 and 2017, however the tail index regression estimation results from 1993 to 2016 can be obtained upon request). The tail index estimate in ((ref)), $\hat{\alpha}_{t}$, provides an estimate of the unconditional tail index after considering the characteristics of all individuals in the sample for each year.

The results in Table 3 show that in general, the estimates based on the method that ignores censored data underestimate the true effects of the variables. This is especially clear in the estimates for 2017 (e.g. female and finance). It is a consequence of the potential inconsistency of the uncensored estimates as highlighted in the Monte Carlo section above. However, the direction of the impact of the covariates suggested by the uncensored estimation is consistent with the results obtained from the censored tail index regression.

Comparing 1992 and 2017 (censored estimation), we generally observed a decrease in the estimates for 2017 (e.g. Transportation and Craft and Precision). In some cases, although in general not statistically significant (see e.g. other races, married no spouse, widowed, Transports and Trade), positive estimates in 1992 become negative in 2017 and vice versa. Considering only the statistically significant covariates it can be observed that Female, Black, Divorced, Single, Low Skill, Craft, Operators, Transportation and Public have a positive impact leading to a reduction in the probability of individuals with these characteristics being in the right tail, whereas Age, Education, Construction, Finance and Repair have a negative impact, originating an increase of the probability of individuals with these characteristics being in the right tail.

table[table omitted — 5,063 chars of source]

\FloatBarrier

figure[figure omitted — 430 chars of source]

\FloatBarrier

Figure (ref) plots the tail index estimate computed as suggested in ((ref)) for each year from 1992 to 2017, based on uncensored and censored tail index regression estimates. As discussed above, the uncensored approach misrepresents the true unconditional tail index, as it ignores the top-coded wages, leading to the overestimation of the true values of $ \alpha _{t}$ (where the bias/inconsistency is due to the fact that the top extreme values are simply not included in the estimation). On the contrary, the censored estimates take the information of the individuals of the top-coded wages, although the true wages are not known, into account. Figure (ref) shows that the censored estimates of $\hat{\alpha }_{t}$ have declined over the last 20 years. In other words, this Figure shows that the probability of observing an extreme value today is higher compared to the 90s or even in the more recent past. This finding supports the idea that upper-tail inequality has increased since the 90s and has become more pronounced over the last 20 years. Autoretal(2008) observed that the 90th percentile wage rose by more than 55$\%$ relative to the 10th percentile between 1963 and 2005, which represents a significant increase. However, our approach suggests that the increase in inequality found by Autoretal(2008) may be a lower bound of the true increase in inequality.

Not adequately handling the top-coding may also lead to overestimation of the tail index in a variety of other applications (this has been recognized in e.g. the analysis of returns to education by Hubbard(2011), and differences in gender and race by BurkhauserLarrimore(2009a) who already use approaches to accommodate for the fact that the wage data is top-coded). Our procedure provides additional flexibility by allowing researchers to evaluate which determinants impact the probability of being in the right tail of the wage distribution and which determinants do not. Thus, permits a more detailed analysis on how inequality has spread across industries, occupation, gender and other population characteristics.

In what follows, we focus on the (conditional) tail index regression, and especially on the partial effects computed as discussed in section 2.2. Figure (ref) groups the partial effect estimates and rearranges them from the lowest effect to the highest. A negative (positive) regression coefficient translates into a positive (negative) partial effect, which is associated with an increase (decrease) in the likelihood of having more extreme values (the right tail becomes (less) heavier).

In the discussion that follows on the partial effects of the covariates, whenever we refer to an extreme event, we are referring to observations which are larger than the 96th quantile ($u=0.15$); see Remark 2 for details.

It is clear that industries such as Finance, Construction and Repair are the ones with more extreme wages. The probability of observing an extreme value increased in 2017 by 4.94% for an individual working in the Finance industry. The probability of an individual working in the public industry having a wage higher than the 96th percentile decreased in 2017 by 2.14%. The biggest increase from 1992 to 2017 was observed for individuals that are black, married without a spouse, or widowed. Older and more educated workers continued to have a significant probability of observing an extreme wage but there was no relevant change between 1992 and 2017. Women observed a positive increase between 1992 and 2017, but the impact is still towards an increase in alpha (although smaller in absolute value than in 1992), i.e., a decrease in the probability of being in the right tail. In specific, women in 2017 have a partial effect on the tail index of -1.26% which corresponds to a decrease in the probability of an extreme value.

The impact in terms of occupation is very interesting. In comparison to an individual working in a non-routine cognitive occupation (managers) all other occupations display a positive contribution to observe an extreme value (although smaller in 2017). However, the picture is very different across occupations. Individuals working in routine occupations such as administrative workers reduced their presence in the right tail from 10% to 8% (see Table (ref)) but the probability to observe an extreme value for these occupations increased in 2017 (-2.14% in 1992 and changed to 2.00% in 2017). The contribution to the probability of observing an extreme value for an individual working as an operator was significantly higher in 1992 than in 2017 (the partial effect was -8.04% in 1992 and it reduced to -1.41% in 2017).

figure[figure omitted — 360 chars of source]

\FloatBarrier

Imputed mean wages

Imputing wages above the top-code

A further important contribution of the approach introduced is its use for the imputation of mean wages above the top-code as described in Section 3.2. Some authors use values to adjust the top-coded wage which are time-varying and differ by group. For instance, MacphersonHirsch(1995) provide separate Pareto estimates according to gender and by year from 1973 to 2014 using public CPS-Merged Outgoing Rotation Groups (CPS-MORG). These authors indicate that these values increase over time and are higher for men than for women (e.g. for 2014 the adjustment coefficient is 2.06 for men and 1.81 for women). In contrast, our analysis is based on the March CPS (outgoing rotation groups 4 and 8), and on the weekly wages instead of annual earnings, however, we also find evidence in favour of changing adjustment parameters.

This renders support to the observation that imputed wages above the top-code, based on a fixed value may lead to misstatement of results, given that this approach considers wages above the top-code to be independent of time, age, gender, race and other personal characteristics; as well as industry and occupation.

To impute wages above the top-code, we consider the estimate of $ E\left( \left. y_{i}\right\vert y_{i}>y_{c};\mathbf{x}_{i}\right) $ given by $\hat{\tau}_{3}\left( \mathbf{x}_{i},y_{c}\right) $ in ((ref)). However, for some individuals in the CPS database, especially those in the highest wage groups, $\hat{\alpha} \left( \mathbf{x}_{i}\right) $ can be in the neighborhood of 1, or even lower than 1, which implies that $E\left( \left. y_{i}\right\vert y_{i}>y_{c};\mathbf{x}_{i}\right) $ does not exist and, therefore, the estimate $\hat{\tau}_{3}\left( \mathbf{x}_{i},y_{c}\right) $ is inadequate. For these cases, we use the conditional median of $y_{i}$ given $y_{i}>y_{c}$ and $\mathbf{x}_{i},$ which is $2^{1/\hat{\alpha}\left( \mathbf{x} _{i}\right) }y_{c}.$ In specific, to accommodate all situations (i.e. low and high values of $\hat{\alpha}\left( \mathbf{x}_{i}\right) )$ we propose the following estimator,

equation[equation omitted — 381 chars of source]

In the empirical application of this statistic we set $c = 1.5$, since using $c=1$ may lead to explosive estimates as $\frac{\hat{\alpha}\left( \mathbf{x}_{i}\right) }{\hat{\alpha}\left( \mathbf{x }_{i}\right) -1}\rightarrow \infty $ as $\hat{\alpha}\left( \mathbf{x} _{i}\right) \rightarrow 1^{+}$. Other values of $c$ in the neighborhood of $ c=1.5$ yield basically the same results. In the application to the CPS data we observed that for the overwhelming majority of estimates $\hat{\alpha}\left( \mathbf{x}_{i}\right) >1.5$. Thus, for most cases, the estimate $\hat{\tau}_{4}\left( \mathbf{x}_{i},y_{c}\right) $ coincides with the second branch of ((ref)), which is the $\hat{\tau}_{3}\left( \mathbf{x}_{i},y_{c}\right) $ estimator in (3.3). Hence, we use $\hat{\tau}_{4}\left( \mathbf{x}_{i},y_{c}\right) $ to impute wages above the top-code over all individuals of the sample across time (see Figure (ref)).

Figure (ref) illustrates the estimates of the imputed wages above the top-code computed from the different approaches discussed in ((ref)), ((ref)) and ((ref)). The $\hat{\tau}_{2}$ and $\hat{\tau}_{4}$ estimates are similar. This result is expected given that for the overall analysis $\hat{\tau}_{4}$ is based on the values of $ \hat{\alpha}_{t}$ computed as in ((ref)), which provides an estimate of the unconditional tail index after considering the characteristics of all individuals in the sample. However, in the case where estimates for a particular group, occupation or industry are considered, the $\hat{\tau}_{3}$ estimates will certainly be different from the $\hat{\tau}_{2}$ estimates (see next section).

figure[figure omitted — 144 chars of source]

\FloatBarrier

When the proportion of wages above the top-code is relatively small (as for example, from 1992 to 2005), the difference between $\hat{\tau}_{1},$ $ \hat{\tau}_{2}$ and $\hat{\tau}_{4}$ is relatively small; however, as more wages are located in the top-coded category (as for example, in the years following 2007), the effect of censored data becomes stronger and the bias (underestimation) produced by the Hill estimator more pronounced ($\hat{\tau}_{1}\left( y_{c}\right) $).

Figure (ref) illustrates the time varying nature of the factor necessary to compute the imputed mean wages. Recall that to overcome the top-coding bias, in the literature, a constant value is frequently used to adjust the top-coded wages. For instance, AutorDorn(2013) consider a sample between 1980-2005; AcemogluAutor(2011), between 1973-2009; Autoretal(2008) from 1963 to 2005; KatzMurphy(1992) between 1963 and 1987; AutorDorn(2013) between 1950 and 2005; Lemieux(2006) from 1973 to 2003; and Beaudry, Green and Sand (2013) from 1979 to 2011. Figure (ref) shows that using a fixed value may have been adequate for pre-1992 data, but that the adjustment factor has increased over time reaching an overall value around 1.85 in 2017.

figure[figure omitted — 172 chars of source]

\FloatBarrier

Imputed wages above the top-code by gender and industry

Figure (ref) illustrates the difference of the imputed mean values for individuals in the Finance, Repair, Personal and Public industries, as well as for women and women working in those industries. The purpose of these graphs is to further highlight the importance of allowing for different scaling factors depending on individuals characteristics and industry, but other graphs considering other characteristics can be plotted using our approach.

The first noticeable result is that the imputed wage of individuals decreases when we compare the wages for individuals in the Finance, Repair, Personal and Public industries. Finance displays the largest and Public the lowest imputed wages of the four industries. With the exception of the Public industry, women's imputed wages are lower for the other three industries and this observation also holds when we condition women's imputed wages on the industry they are in. A further interesting result is that the imputed mean wages display an increasing trend over time in all industries, for women and for women in those industries, which is an indication that the adjustment factors used to compute the imputed wages also changes over time.

figure[figure omitted — 567 chars of source]

\FloatBarrier

This is further highlighted in Figure 10, where the graphs show the top-coded wage adjustment factors between 1992 and 2017, for different combinations of women working in different industries. Women’s top-coded adjustment factor is always smaller than the topcoded adjustment factor that would be applied to males in any industry with the exception of public. This implies that the public industry is the less heavy tailed. On top of that women working in the public industry earn less in the right tail than women in the right tail working in other industries.

While women working in personal and repair are not earning much more nor much less than in other industries we find that women working in finance would need a much higher adjustment factor. This means that this is the industry in which they have been earning more and this result is reinforced with a clear positive trend between 1992 and 2017.

figure[figure omitted — 595 chars of source]

\FloatBarrier

Conclusions

This paper provides three important contributions to the literature. The first corresponds to the introduction of a conditional tail index estimator which explicitly handles the top-coding problem and an indepth evaluation of its finite sample performance and comparison with competing methods. The Monte Carlo simulation exercise shows that the method proposed to estimate the tail index performs well in terms of estimation of the tail index and when used in the imputation of wages above the top-code when the sample is censored, which is an intrinsic feature of the public-use CPS database.

Second, evidence is provided which shows that the factor values used to adjust the top-coded wages have changed over time and across the characteristics of individuals, occupations and industries and an indication of suitable values is proposed. Interestingly, the empirical results show that the upper-tail inequality has increased since the 90s and has become more pronounced over the last 20 years.

Third, an in depth empirical analysis of the dynamics of the US wage distribution's right tail using the public-use CPS database from 1992 to 2017 is provided. The application of the procedure to the CPS data reveals that individuals working in industries such as finance, construction and repair are the ones with the more extreme wages. Moreover, it is also observed that the biggest increases in the probability of observing an extreme wage between 1992 and 2017 was for individuals that are black, married without a spouse, or widowed. Older and more educated workers continued to have a significant probability of observing an extreme wage but there was no relevant change between 1992 and 2017. Women observed a positive increase between 1992 and 2017, but the impact is still towards an increase of alpha (although smaller in absolute values than in 1992) i.e. a decrease in the probability of an extreme wage. Furthermore, it is also noted that in comparison to an individual working in a non-routine cognitive occupation (managers) all occupations observed a positive contribution to observe an extreme value (although smaller in 2017). However, conclusions are different across occupations.

Our analysis also showed that women working in finance (public) would need a higher (lower) adjustment factor to impute the top-coded wages. Furthermore, we also observe that the adjustment factor used to impute top-coded wages should be adapted over time and across characteristics of the individuals especially when using censored data.

\setcounter{equation}{0}