EconBase
← Back to paper

Non-Existent Moments of Earnings Growth

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

72,640 characters · 12 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Non-Existent Moments of Earnings Growth

abstract{4.9mm} The literature often employs moment-based earnings risk measures like variance, skewness, and kurtosis. However, under heavy-tailed distributions, these moments may not exist in the population. Our empirical analysis reveals that population kurtosis, skewness, and variance often do not exist for the conditional distribution of earnings growth. This challenges moment-based analyses. We propose robust conditional Pareto exponents as novel earnings risk measures, developing estimation and inference methods. Using the UK New Earnings Survey Panel Dataset (NESPD) and US Panel Study of Income Dynamics (PSID), we find: 1) Moments often fail to exist; 2) Earnings risk increases over the life cycle; 3) Job stayers face higher earnings risk; 4) These patterns persist during the 2007--2008 recession and the 2015--2016 positive growth period. Keywords: earnings & income risk, heavy tail, conditional Pareto exponent.

Introduction

One of the main goals of the literature on income dynamics is to quantify income risk. Recent studies show that the distribution of income changes has heavier tails than the normal distribution. This evidence has relevant implications for how income risk affects economic decisions.\footnote{See, for instance, \citet*{golosov2016redistribution} and references therein.} Since heavy tails incur enormous risk premium costs in economies with risk-averse agents, quantifying the tail heaviness, along with features of preferences of economic agents, is necessary to measure the social costs of the uncertainties.

However, the existing literature relies primarily on moment-based measures of income risk. For instance, the recent influential paper by \citet*[][]{guvenen2021data} uses skewness and kurtosis to characterize the distributions of individual earnings dynamics in the U.S. and provide their policy implications. They target this set of moments in their proposed procedure to estimate the parameters of a flexible stochastic process of earnings. The estimated process further feeds back into economic research and policy analyses, for instance in the analysis of life-cycle incomplete market models. While sample skewness and sample kurtosis are generally finite, their true population counterparts may not exist when the distribution has a heavy tail. If the population moments do not exist, then their sample counterparts can be arbitrarily large and hence misleading.

In this paper, we address this issue by first proposing a method of estimation and inference about the conditional Pareto exponent, which characterizes the tail heaviness. Second and more importantly, applying our method to the administrative data set for the UK, the New Earnings Survey Panel Dataset (NESPD), and the US Panel Study of Income Dynamics (PSID), we quantify the tail heaviness in the distribution of earnings changes and discuss how it varies with heterogeneous attributes of individuals. We find that the commonly used measures of income risk, such as sample kurtosis, skewness, and even standard deviation, may be misleading since their population counterparts may not exist (i.e., may be infinite) for the conditional distribution of earnings growth given age, gender, and past earnings. Finally, by interpreting the conditional Pareto exponent as a robust measure of the conditional earnings risk, we show that tail earnings risk increases over the life cycle, is higher for job-stayers, and these patterns appear in both the period 2007--2008 of the great recession and the period 2015--2016 of a positive growth despite some differences.

There is extensive literature on earnings and income dynamics models, dating back to the 1970s -- see the comprehensive survey by \citet*{moffitt2018income}. To explain the heavy tail feature, the existing literature exploits the following strategies: 1) fitting data to mixed normal distributions using mixture approaches \citep*[][]{geweke2000empirical,bonhomme2009assessing}; 2) conducting nonparametric density estimates and graphically exhibiting heavier tails than normal distributions with deconvolution and spectral decomposition approaches \citep*[][]{horowitz1996semiparametric,bonhomme2010generalized,arellano2017earnings,botosaru2018nonparametric,hu2019semiparametric}; 3) reporting the sample moments such as skewness and kurtosis and comparing them with counterparts from normal distribution \citep*[][among others]{guvenen2021data,arellano2021}; 4) reporting some quantile-based measures such as Kelly's measure of skewness and Crow-Siddiqui measure of kurtosis \citep*[][among others]{guvenen2021data}; and 5) reporting the mean absolute deviation \citep*[][]{arellano2021}.

The first approach, based on mixed normal distribution, essentially assumes an exponentially decaying tail, which cannot characterize tail heaviness. The second one is based on nonparametric density estimation, which tends to behave poorly in the tails due to few extreme observations. The third approach based on sample moments may yield misleading information if the population moment does not exist due to heavy tails. In particular, the population kurtosis may not exist even if a researcher can always get a finite sample kurtosis. We indeed find that the population kurtosis, skewness, and even variance may not exist. The fourth one is robust to heavy tails: these non-extremal quantile-based measures are always well-defined even if the population moments do not exist. However, they fail to capture extreme income changes and tail events. Among other things, information on tail events is central to quantifying earnings risk under heavy-tailed distributions. Like the fourth one, the fifth approach based on the first moment alone is not informative about extreme income changes and tail events.

Given the above concerns about existing approaches, we aim for an alternative measure of tail heaviness to characterize the earnings risk conditional on individuals' attributes. The conditional Pareto exponent is proposed as our robust measure of tail heaviness. We have the elegant characterization that the conditional kurtosis/skewness/variance is infinite if and only if the conditional Pareto exponent is smaller than the value of four/three/two. We propose methods of estimation and inference about the Pareto exponent in the conditional distribution of earnings risk given observed attributes, such as gender, age, and base-year income level. Unlike the aforementioned approaches, this new method can be used to characterize the tail heaviness robustly even if the rate of tail decay is slower than exponential ones and even if the population kurtosis, skewness, or variance is infinite. We emphasize that we do not make the assumption that data follow a Pareto distribution. Instead, we only assume the regularly varying (RV) tail condition, which nonparametrically encompasses a broad class of distributions, including Pareto, Student-t, Cauchy, and F distributions among many others. Numerous recent papers document that the distribution of income/earnings, both in levels and growth, roughly follows the power law, which necessarily entails ‘extreme’ observations and endorses our assumption -- see for instance \citet*{guvenen2021data} for earnings growth and \citet*{guvenen2022global} for additional empirical evidence that supports our assumption.\footnote{Based on the evidence provided by all papers submitted within the Global Income Dynamics Projects (GRID), \citet*{guvenen2022global} report that the distribution of income growth has thick Pareto tails in all 13 countries currently under the GRID.} The RV tail condition is in line with these observations from the literature.

To illustrate the empirical significance, we apply the proposed method to two data sets: the UK NESPD and the US PSID, with our focus on the former. NESPD is an employer-based survey data set on individual earnings in the UK and has been investigated very recently \citep*[cf.][]{de2021wage,bell2021time}. We contribute to this new literature with the following three findings. First, we find that the distribution of the earnings changes (measured by the difference between the log earnings across two subsequent years) conditional on age, gender, and past earnings exhibits substantially heavy tails. The conditional Pareto exponent is significantly less than four for the majority of the subpopulations, raising the concern that the kurtosis may be infinite. Also, the conditional Pareto exponent is significantly less than three or even two for some subpopulations, and we thus reject the hypothesis of finite conditional skewness and even finite conditional variance. These results are robust across years, including the period of recession in 2007-2008 and the period of positive growth in 2015-2016 among others.

Second, we find remarkable patterns that the tail earnings risk is increasing over the life cycle. This is documented as that 40- and 50-year-old workers face higher earnings risk than 30-year-old workers. Third, we quantify the tail heaviness of the distribution of earnings changes for job-stayers and find that, at low earnings levels, conditional kurtosis does not exist, contrary to what we observe for the whole population. Excluding middle-aged men at top quantiles of earnings, we reject the hypothesis of finite conditional kurtosis for job-stayers at all earnings levels. In Appendix (ref), we also apply our method to the US PSID and obtain similar empirical findings on the tail heaviness of the conditional distribution of income risk.

The above findings closely compare to the findings in \citet*{guvenen2021data} and the income literature inspired by their methodology. In particular, \citet*{guvenen2021data} find in the US data set that the conditional variance of earnings risk is higher for younger workers at the bottom of the income distribution, and the kurtosis increases with age and with lagged income up to the top 5% of the income distribution where it sharply declines. Since the data exhibit substantially heavier tails than those of the normal distribution, these sample moments could be less informative if their population counterparts are indeed infinite. Our results confirm theirs, but now with our new measure of earnings risk being robust to infinite moments. On the other hand, our findings differ from those in \citet*{arellano2021}, who examine the Spanish administrative data, and find that the income risk measured by the conditional mean absolute deviations is inversely related to income and age, and the income risk inequality increases markedly in the recession. The differences can be explained by a number of factors, such as different countries and different sample selection criteria as well as different measures of income risk for different research objects.

{\bf Organization:} The rest of the paper is organized as follows. Section (ref) describes the data to be examined in this article and provides some descriptive analysis of the heavy-tail phenomena in the earnings growth risk. Section (ref) formally introduces our econometric method, and Section (ref) applies it to the UK data set. Section (ref) concludes with some remarks. The Appendix collects asymptotic theories, computational details, Monte Carlo simulations, additional empirical results with the UK and US data sets, and mathematical proofs.

Data and Preview

New Earnings Survey Panel Data

We start with introducing the New Earnings Survey Panel Data (NESPD), an administrative data set at the individual level on UK earnings from the UK Social Security. It is an annual panel running from 1975 to 2016, and it surveys around 1% of the UK workforce. All employees whose National Insurance Number (NIN) ends in a given pair of digits are included in the survey. The NIN number is randomly issued to all UK residents at age 16 and kept constant throughout the lifetime of an individual. The NESPD is a survey directed to all employers whose employees qualify for the sample: the employers complete the questionnaire based on payroll records for their employees. As a result of being directed to the employer, NESPD has a low non-classical measurement error.

The survey reports the employees' gender, age, and detailed work-related information: annual, weekly, and hourly earnings, hours of work, occupation, industry, working area, firms' number of employers, and unionization. This information relates to a specified week in April of each year: the data sample is taken on the first day of April of each calendar year and concerns complete employee records only. The NESPD contains complete information on the employees' working life from the first year they started working until retirement age over the years 1975-2015, as long as the employer answered the questionnaire and the individual was working with the last recorded employer in April.

Given that NESPD contains detailed information on earnings and a long panel component, it has been used for a broad range of topics in the literature: among the others, goos2007lousy have used NESPD to document job polarization in the UK; nickell2003nominal and elsby2016wage have employed this study for analysis of wage rigidities; adam201935 provide an analysis of valid response rates in NESDP. However, the study of income and earnings dynamics in the UK has not received lots of attention, with a couple of notable exceptions in recent years: de2021wage and bell2021time are two main references for earnings dynamics in the UK using NESPD for the analysis.

The information in NESPD is of exceptionally high quality. However, it is worth mentioning that not all workers are covered and there are non-negligible non-response issues. In particular, data in NESPD are collected in a specified week of April. As a result, NESPD might under-sample part-time workers if their weekly earnings fall below the threshold for paying National Insurance and those who moved jobs recently. Furthermore, the data set is unbalanced with possibly non-random missing observations. This issue might be problematic for the representativeness of those at the bottom of the distribution and with unstable spells. We refer to bell2021time and de2021wage for a deeper investigation of the data limitations of NESPD.

Following the earnings literature, we further extract a subset of NESPD, which corresponds to employees with a strong attachment to the labor force. Specifically, we follow de2021wage to use the following selection criteria: we drop the observations below 5% of the median earnings (roughly \pounds1,300 a year), individuals whose total working hours exceed 80 hours per week, and individuals that display negative values in earnings or hourly wages; we do not consider individuals whose hours worked or weekly pay are missing. Earnings are deflated using CPI (2015=100). The earnings measure is the residual obtained by regressing the logarithm of earnings on year and age dummies. We extract a portion of data for a pair of years across the period of the Great Recession, 2007-2008. The above-described sample selection leaves 78,531 individuals.

For the purpose of comparisons, we also use the portion of data for the pair, 2015 and 2016, of the most recent survey years. Note that this period is associated with positive economic growth in the UK. The sample selection procedure described above leaves 95,906 individuals for this period. The results for other years are similar and hence postponed to Appendix (ref) for readability.

Apart from age, gender, and income, we could further consider industry and occupation as heterogeneous attributes in the analysis. NESPD contains information on both industry and occupation with the Standard Industry Classification (SIC) and Standard Occupation Classification (SOC) codes.

We define a measure $Y$ as the difference between the log earnings in 2007 and the log earnings in 2008 (i.e., one-year earnings growth rate). Similarly, we also construct this variable for the period between 2015 and 2016. Figures (ref) and (ref) show the kernel density estimates of the measure $Y$ for the period 2007--2008 and the period 2015--2016, respectively.\footnote{The kernel estimates in these figures are obtained with the Gaussian smoothing kernel. Given the concerns raised by Degiannakis et al. (2023b) and Lahr (2014) that, under the existence of extreme values, the visual output of the kernel estimates may vary significantly under different weighting functions, we check various kernel functions (e.g., Epanechnikov, rectangular, and triangular kernels) in Appendix (ref). Our analysis reveals no substantial differences among them.} In each figure, the left and right panels show the densities for men and women, respectively. The figures illustrate clear departures from normality: each kernel density exhibits a large spike in the middle of the distribution sticking upward out of the reference normal density. Moreover, each kernel density has heavier tails compared to the reference normal density. These features of the estimated densities suggest that the actual distributions of $Y$ indeed have heavier tails than normal distributions, as documented in the previous literature.

We also inspect the sample moments of the distribution of one-year earnings changes conditional on previous earnings. Figure (ref) shows the standard deviation, the Kelly's measure of skewness, and the Crow-Siddiqui measure of kurtosis, respectively, of earnings changes by previous earnings, for men (left three figures) and women (right three figures). The Kelly's measure of skewness is defined as $(Q_{.9}Y-Q_{.5}Y)-(Q_{.5}Y-Q_{.1}Y)/(Q_{.9}Y-Q_{.1}Y)$ and the Crow-Siddiqui measure of kurtosis as $(Q_{.975}Y-Q_{.025}Y)/(Q_{.75}Y-Q_{.25}Y)$, where $Q_{\tau}Y$ denotes the $\tau$-quantile of the distribution of $Y$. As in de2021wage, we find that the Crow-Siddiqui measure of kurtosis is higher than the reference value of 2.91 for the normal distribution, and the Kelly's measure of skewness deviates from zero, ranging from positive values for low earnings groups to negative values for higher values of previous earnings. However, as mentioned in the introduction, these quantile-based measures do not take into account the very tail observations that are above, say $Q_{.975}Y$. These observations are indeed informative about the tail feature of the earnings risk and inequality. In comparison, our proposed method takes all samples into consideration and is explicitly designed to characterize the tail heaviness.

In the rest of the paper, we analyze the heaviness of the tails of the conditional distributions of $Y$ given the earnings level and age in the base year (i.e., the base year is 2007 for the period 2007--2008 and it is 2015 for the period 2015--2016). Specifically, the heaviness of the tails is quantified by the conditional Pareto exponent, and we propose an econometric method of estimation and inference about this measure of tail heaviness. Its value, in particular, informs whether the $r$-th conditional moments of the income growth exist for $r=2,3,4$, and so on. If the conditional Pareto exponent, $\alpha(x_0)$, given $X=x_0$ is less than $r$, then the conditional distribution of $Y$ given $X=x_0$ does not have a finite $r$-th moment. Section (ref) introduces a method of estimation and inference, and Section (ref) presents the empirical results. It turns out that we reject the hypotheses of finite conditional kurtosis, finite conditional skewness, and even finite conditional standard deviation for some subpopulations.

figure[figure omitted — 563 chars of source]
figure[figure omitted — 563 chars of source]
figure[figure omitted — 671 chars of source]

Measurement of Conditional Tail Risk

We now introduce our proposed measure of conditional tail risk, and present the method of estimation and inference. Formal theoretical justifications are relegated to the Appendix. Let $Y$ denote the variable of interest and let $X$ denote a vector of individual characteristics, such as the income level and age in the base year. We first assume that the sample is i.i.d. so that the joint distribution $F_{Y,X}\left(\cdot ,\cdot \right)$ of $(Y,X)$ is unique.

Assumption 1: The sample $(Y,X)$ is i.i.d.

In our application to earnings changes, the cross-sectional i.i.d. assumption is typically imposed explicitly or implicitly. While the literature has emphasized the relevance of time dependence with ARCH effects or stochastic volatility, for several questions ranging from consumption insurance to income mobility, \citep*[see][]{meghir2004income}, it is plausible that the i.i.d. assumption is satisfied when it comes to cross-sectional data of earnings changes. Also, note that this is a sufficient condition that can be relaxed to allow for some form of weak dependence. See Appendix (ref) for related econometric methods that allow for time series data.

Consider a rectangular array $\{(Y_{ij},X_{ij}): i \in \{1,\cdots,I\}, j \in \{1, \cdots, J\}\}$. Such a data structure can be constructed by randomly splitting a cross-sectional data set of size $N$ into an $I \times J$ array such that $I \cdot J \approx N$, which is what we implement in Section (ref). See Appendix (ref) for how to choose $I$ and $J$ in practice under such a construction of an array.

Our second main assumption is that the conditional distribution $F_{Y|X=x_{0}}\left( \cdot \right) $ is regularly varying at infinity, which is essentially equivalent to

equation*[equation* omitted — 113 chars of source]

where $\alpha \left( x_{0}\right) >0$ denotes the $x_0$-conditional Pareto exponent that characterizes the tail heaviness of the conditional distribution of $Y$ given $X=x_0$. A formal definition of regular variation is as follows.

Assumption 2: The conditional distribution $F_{Y|X=x_{0}}\left( \cdot \right) $ is regularly varying at infinity, that is, for all $y>0$

equation*[equation* omitted — 153 chars of source]

where $\alpha \left( x_{0}\right) >0$ denotes the $x_0$-conditional Pareto exponent that characterizes the tail heaviness of the conditional distribution of $Y$ given $X=x_0$. }

We again emphasize that we do not impose the assumption that data follow a Pareto distribution. Instead, this assumption only requires the regularly varying (RV) tail condition. This RV tail condition is mild and satisfied by many commonly used distributions, such as Student-t, F, and gamma distributions among numerous others, as well as nonparametric distributions. Moreover, it has been widely documented that the earnings data sets exhibit a Pareto-type (i.e., RV) tail and hence support this condition \citep*[cf.][]{Gabaix2009, gabaix2016power, guvenen2021data}. See a survey by \citet*{gabaix2016power} and \citet*{guvenen2022global} for additional empirical evidence that supports our assumption. From this existing literature, the assumption of the existence of extreme events is legitimate for the distribution of income and earnings growth. In line with these findings, our data provide suggestive evidence that extremes can occur and, thus, further validate our assumption, as shown in the densities in Figures (ref)--(ref), and in the Pareto plots in Figures (ref)--(ref). See Appendix (ref) for more discussions and primitive conditions.

With the parameter $\alpha\left(x_{0}\right)$, we can characterize the existence of the conditional moments as follows. For any $r\in \mathbb{R}^{+}$,

align*[align* omitted — 203 chars of source]

Thus, the test of finite conditional $r$-th moment can be represented by the competing hypotheses

equation[equation omitted — 94 chars of source]

In particular, we set $r=2$, $3$, and $4$ for tests of the existence of the standard deviation, skewness, and kurtosis, respectively.

To construct a feasible test for (ref), we need to obtain a random sample from the conditional distribution $F_{Y|X=x_{0}}\left( \cdot \right)$. This would be straightforward if $X$ is discrete. However, random sampling from such a conditional distribution is infeasible when $X$ includes at least one continuous random variable, such as the income level, which we use in our data analysis. To overcome this issue, we propose the following procedure to extract the local sample. Let $\left\vert \left\vert \cdot \right\vert \right\vert $ denote the Euclidean norm.

enumerate• For each $i$, find the nearest neighbor (NN) of $\{X_{ij}\}_{j=1}^{J}$ to $x_{0}$ and the induced $Y_{ij}$ associated with the NN, that is $ Y_{ij}=Y_{ij^{\ast }}$ where $\left\vert \left\vert X_{ij^{\ast }}-x_{0}\right\vert \right\vert =\min_{j\in \{1,...,J\}}\left\vert \left\vert X_{ij}-x_{0}\right\vert \right\vert$. Denote the induced $Y$ by $ \{Y_{i,[x_{0}]}\}_{i=1}^{I}$. • Sort $\{Y_{i,[x_{0}]}\}_{i=1}^{I}$ descendingly as $ \{Y_{(1),[x_{0}]}\geq Y_{(2),[x_{0}]}\geq ,...,\geq Y_{(I),[x_{0}]}\}$. Collect the largest $K+1$ of them as \begin{equation*} \mathbf{Y}(x_{0})=\{Y_{(1),[x_{0}]},Y_{(2),[x_{0}]},...,Y_{(K+1),[x_{0}]}\}. \end{equation*} • Use $\mathbf{Y(}x_{0}\mathbf{)}$ to estimate the Pareto exponent $\alpha\left(x_{0}\right)$ by the formla in (ref) below.

We postpone until Appendix (ref) formal discussions of conditions under which this procedure works and why it works in theory. Here, we instead discuss its intuition. First, for each $i$, we select the NN among $\{X_{ij}\}_{j=1}^{J}$ to $x_{0}$. When $J$ (as well as $I$) is large enough, such a NN gets close enough to $x_{0}$ and therefore its induced value $Y_{i,[x_{0}]}$ behaves as if it were generated from $F_{Y|X=x_{0}}$. Second, given the i.i.d. assumption, the subsample $\{Y_{i,[x_{0}]}\}_{i=1}^{I}$ collected from all individuals serves as a random sample approximately generated from $F_{Y|X=x_{0}}$. We can thus use the extreme order statistics of $\{Y_{i,[x_{0}]}\}_{i=1}^{I}$ to estimate the Pareto exponent $\alpha\left(x_{0}\right)$ by using one of the many existing methods. In particular, if we apply the estimator of Hill1975 to this NN sample $\{Y_{i,[x_{0}]}\}_{i=1}^{I}$, then we get an estimator

equation[equation omitted — 162 chars of source]

of the conditional Pareto exponent $\alpha(x_0)$. See Appendix (ref) for how to choose $K$ in practice.

We show in Appendix (ref) that this estimator asymptotically follows the normal distribution as

equation[equation omitted — 165 chars of source]

under suitable conditions. Moreover, for any $x_{0}\neq x_{1}$, we also have the joint asymptotic normality:

equation[equation omitted — 272 chars of source]

We provide a formal statement of these results as Proposition (ref) in Appendix (ref).

Given (ref) and (ref), we can construct the standard t and F tests for our hypothesis testing problem ((ref)). Specifically, we reject the (one-sided) null hypothesis at the 5% nominal level if

equation[equation omitted — 114 chars of source]

where $\Phi ^{-1}\left( \cdot \right) $ denotes the quantile function of the standard normal distribution and $r=2$, $3$, and $4$. Moreover, we reject the null hypothesis that $\alpha (x_{0})=\alpha (x_{1})$ (against the two-sided alternative) at the 5% nominal level if

equation[equation omitted — 189 chars of source]

Appendix (ref) contains a simulation study, which supports our theoretical results.

We end this section with discussions about the related econometrics and statistics literature. To estimate the unconditional Pareto exponent of some underlying distribution, researchers have developed numerous methods, including the popular and widely used Hill1975's (Hill1975) estimator,\footnote{This method fits large values in the sample to the Pareto distribution and implements the maximum likelihood estimation.} and those proposed by Smith1987MLE and Gabaix2011rank. See Fedotenkov2020 for a recent review of other methods. Primitive conditions and asymptotic properties of these estimators are typically studied by using extreme value theory and the Pareto tail approximation Hall1982,HallWelsh1985. We refer to deHaan2007book for a comprehensive review. In recent work, beare2017determination show that the cumulative sum of Markov multiplicative processes generates an approximate Pareto tail. This result supports our theoretical conditions and fits our empirical setup. In general, a pitfall of these methods in our framework is that the unconditional Pareto exponent cannot characterize the effect of important demographic characteristics on growth risk. This issue motivates us to study the conditional Pareto exponent.

Unlike the unconditional Pareto exponent of, say $Y$, estimating its conditional counterpart given some other variable $X$ has received less attention. This is technically more challenging when $X$ contains continuous random variables as one could not obtain a random sample from the conditional distribution of interest, say $F_{Y|X=x_0}$ for some pre-specified value $x_0$. Wang2009 and WangLi2013 impose some parametric assumptions such that the Pareto exponent is a single index function of $X$. Gardes2008 and Gardes2012 develop fully nonparametric methods based on local smoothing. Estimation and inference based on these local smoothing methods sensitively rely on the bandwidth parameter unlike our proposed method based on nearest neighbors. Furthermore, these existing methods underperform compared to ours in terms of bias and standard deviation, as demonstrated through simulation studies presented in Appendix (ref). For these reasons, we suggest using our novel method proposed above.

Finally, in two related papers, trapani2016testing and degiannakis2022 propose randomized testing procedures to test the finiteness of moments of a random variable or the residual of the linear regression. Unlike their proposal, we test the existence of moments of the conditional distribution of a random variable given some other variables. To see the difference, consider, for example, the model that $Y = \mu(X) + \sigma(X)U$ where $\mu(X)$ and $\sigma(X)$ are some functions of $X$ that characterize the location and scale of $Y$ respectively. One can estimate these two functions and further estimate the tail heaviness of $Y$ by that of the residual $\hat{U}$. But such a model by construction assumes that the tail heaviness of $Y$ remains unchanged across $X$ \citep*[cf.][]{WangLi2013}. This feature is neither sufficient for our purpose of measuring the tail heaviness of earnings growth as a function of demographic characteristics nor coherent with the empirical findings.\footnote{The tests developed by trapani2016testing and degiannakis2022 can be generalized to conditional distributions. When $X$ is discrete, we can condition on specific values of $X$ by taking the subsample. Otherwise, we can select the subsample within a local neighborhood of a certain query value $X=x$. We leave this for future research.} Also related is the paper by sasaki2022fixed, who propose inference for the tail index of conditional distributions. Although allowing the tail index to depend on $X$, their inference method builds on the fixed-$k$ asymptotics and cannot be easily extended for estimation. In our context, estimation is valuable and thereby, we can propose a novel measure of earnings risk based on the estimate of the tail index unlike any of these existing papers. As such, our proposal is suitable for studying the tail features of the conditional distribution of income growth and relates these to the results in the income literature.

remarkWe propose to use the estimate of the Pareto exponent of the conditional distribution as a novel measure of tail earnings risk given the conditioning variables, $X$. Our risk measure does `not' depend on $K$ (or any specific quantiles), but depends only on the underlying distribution (and the conditional value $x_0$), like the existing measures in the literature. In particular, the Pareto exponent is uniquely determined by the underlying distribution. See, for example, Gabaix2016 for some commonly used distributions and their corresponding Pareto exponents. In contrast, $K$ is a tuning parameter, which is close in spirit to the bandwidth in kernel regression.

Results

Benchmark Sample

Applying the method introduced in Section (ref) (and also in more detail in Appendix (ref)) to NESPD described in Section (ref), we analyze the conditional tail risk of earnings of adult individuals in the United Kingdom. We define $Y$ by the absolute difference between the log earnings in 2007 and the log earnings in 2008 for our baseline analysis. For the conditioning variables $X$, we include the quantile of earnings level and the age of the individual in the base year (2007), following guvenen2021data. With this setting, we study the conditional Pareto exponent $\alpha(x_0)$ for each point $x_0$ of earnings levels from $\{0.05,0.10,\cdots,0.90,0.95\}$ (in quantile) and ages from $\{30,40,50\}$ for each of men and women.

Figure (ref) illustrates the estimates of the conditional Pareto exponents $\alpha(x_0)$ (in black lines) along with the upper bounds of their one-sided 95% confidence intervals (in gray lines) for 30-, 40- and 50-year-old individuals. The left (respectively, right) panel shows the results for men (respectively, women) in each figure, which suggests different patterns of heterogeneous earnings risk across age, base-year earnings, and gender groups.

figure[figure omitted — 1,001 chars of source]

We describe the empirical findings in the order of age and gender. For 30-year-old men (the top left panel in Figure (ref)), the conditional Pareto exponents (in point estimates) range from 1.4 to 7.0, and the upper bounds of the one-sided 95% confidence intervals range from 2.3 to 11.4. Given any quantile of earnings received in 2007, except the very top and bottom quantiles, the conditional Pareto exponent is significantly less than four, implying that the conditional kurtosis of earnings growth does not exist for these middle-earnings groups of young men. Overall, the kurtosis barely exists for most of the base-year earnings levels even if we fail to reject the hypothesis of finite kurtosis. Furthermore, given that earnings received in 2007 were between the 45th and the 75th percentile, the Pareto exponent is significantly less than three, implying that even the conditional skewness does not exist. For 30-year-old women (the top right panel in Figure (ref)), the conditional Pareto exponents (in point estimates) range from 1.5 to 6.1, and the upper bounds of the one-sided 95% confidence intervals range from 2.5 to 10.0. For this subpopulation, the conditional Pareto exponent is significantly less than four for high- and middle-high-earnings groups of women, those above the 60th percentile of the earnings distribution in 2007, implying that the conditional kurtosis does not exist. Furthermore, given that earnings received in 2007 were at or above the 70th percentile, the Pareto exponent is significantly less than three, implying that even the conditional skewness does not exist. Comparing the results between men and women at age 30, we observe that women were more vulnerable to earnings risk than men at the top quantiles of the base-year income level, above the 75th percentile.

For 40-year-old men (the middle left panel in Figure (ref)), the conditional Pareto exponents (in point estimates) range from 1.2 to 9.1, and the upper bounds of the one-sided 95% confidence intervals range from 1.9 to 14.9. Remarkably, the earnings risks of 40-year-old men are higher than those of 30-year-old men. For this age group of men, we reject the hypothesis of finite kurtosis at any level of base-year earnings, apart from the bottom quantiles, below the 15th percentile. Furthermore, given that earnings received in 2007 were between the 20th and the 35th percentile and above or at the median excluding the very top quantile (between the 75th and 80th percentile), the Pareto exponent is significantly less than three (two), implying that even the conditional skewness (respectively, standard deviation) does not exist. For 40-year-old women (the middle right panel in Figure (ref)), the conditional Pareto exponents (in point estimates) range from 1.4 to 6.4, and the upper bounds of the one-sided 95% confidence intervals range from 2.4 to 11.0. For this age group of women, we reject the hypothesis of finite kurtosis at any earnings level above the 40th percentile. Furthermore, given that earnings received in 2007 were at or above the 65th percentile, the Pareto exponent is significantly less than three, implying that even the conditional skewness does not exist. Comparing the results between men and women at age 40, we observe that men are almost as vulnerable to earnings risk as women for this middle age group, with some difference at the bottom quantiles.

For 50-year-old men (the bottom left panel in Figure (ref)), the conditional Pareto exponents (in point estimates) range from 1.1 to 4.8, and the upper bounds of the one-sided 95% confidence intervals range from 1.8 to 8.0. For this age group of men, we reject the hypothesis of finite kurtosis and finite skewness at any level of base-year earnings, apart from the bottom quantiles, below the 20th percentile. Moreover, we reject the hypothesis of finite standard deviation at the 45th percentile and between the 65th and the 75th percentile. For 50-year-old women (the bottom right panel in Figure (ref)), the estimated conditional Pareto exponents range from 1.1 to 15.0, and the upper bounds of the one-sided 95% confidence intervals range from 1.9 to 25.7. For this age group of women, we reject the hypothesis of finite kurtosis at any level of base-year earnings, apart from the bottom quantiles, below the 30th percentile. Moreover, we reject the hypothesis of finite skewness (standard deviation) at any level of base-year earnings above the 40th percentile (between the 55th and the 70th percentile).

The tuning parameter $K$ is to be selected by a data-driven method (cf. Appendix (ref)) and should not be manipulated by a researcher. Yet, it may be of interest to see the sensitivity of the results to variations in $K$. Table (ref) reports the estimates of the conditional Pareto exponent, $\alpha \left(x_0\right)$, for several values of $K$ for 50-year old men in 2007--2008, conditional on the quantiles of base year income. Our findings are robust across different values of $K$: estimates tend to be higher at lower quantiles of the base year earnings distribution, compared to those at the higher end.\footnote{We obtain similar findings in terms of robustness to $K$ of the Pareto estimates obtained in the following sections for other years or values of the conditioning variables. Results are available from the authors upon request.}

table[table omitted — 1,332 chars of source]

Period of Positive Growth

While our main focus has been on the period of the great recession between 2007 and 2008, we next look at the period of positive growth, between 2015 and 2016. Figure (ref) illustrates the results for earnings growth between 2015 and 2016 for 30-, 40- and 50-year-old individuals. These results share similar qualitative patterns to those reported in Figure (ref). First, 30-year-old men at high quantiles of base-year earnings enjoy less earnings risk. Second, 40-year-old men have an overall higher earnings risk than 30-year-old men. Lastly, and most remarkably, for men the earnings risk is not necessarily lower in the period 2015--2016 than in the period 2007--2008, even though the former period enjoyed positive GDP growth (2.2% in 2015 and 1.8% in 2016) and the latter period suffered from negative GDP growth (0.7% in 2007 and $-$4.4% in 2008). Women, instead, are less vulnerable to earnings risk in the period 2015--2016 than in the period 2007--2008. Overall, there are more similarities than differences despite the contrast between a recession and a positive growth in the UK economy.

figure[figure omitted — 997 chars of source]

In summary, we find that: 1) the kurtosis, skewness, and even standard deviation may not exist for the conditional distribution of earnings growth given certain attributes (age, gender, and earnings); 2) 40- and 50-year-old workers have overall higher earnings risk than 30-year-old workers and we thus document that tail earnings risk is increasing over the life cycle; and 3) these patterns appear both in the period 2007--2008 of great recession and the period 2015--2016 of positive growth for men, while there are some differences for women. Finally, we remark that, although we use the two periods, 2007--2008 and 2015--2016, to draw the above conclusion, we also present additional empirical results for other periods between 2005 and 2016 in Appendix (ref) to demonstrate robustness. In Appendix (ref), we also apply our proposed method to the US PSID and obtain similar empirical findings on the tail heaviness of the conditional distribution of income risk.

Job Stayers

We also experiment with a different set of selection criteria: we restrict the sample to workers who did not change their jobs in the last twelve months. We refer to them as job stayers. Figures (ref) and (ref) illustrate estimates of the conditional Pareto exponents $\alpha(x_0)$ (in black lines) along with the upper bounds of their one-sided 95% confidence intervals (in gray lines) for the two pairs of years: 2007--2008 and 2015--2016, with the restricted sample of job stayers. Figures (ref) and (ref) show that, for almost all attributes, conditional kurtosis does not exist at any level of baseline earnings.

This finding is in line with the US evidence documented by guvenen2021data that job stayers experience different earnings dynamics and, in particular, face earnings innovations that are more leptokurtic, especially at the bottom of the earnings distribution.

More specifically, we find that, at low earnings levels, conditional kurtosis does not exist, in contrast to what we observe for the whole population. Excluding middle-aged men at top quantiles of earnings, we reject the hypothesis of finite conditional kurtosis for job-stayers at all earnings levels.

figure[figure omitted — 1,120 chars of source]
figure[figure omitted — 1,120 chars of source]

Pareto Tail Fit

In this section, we would like to comment more on Assumption 2 by providing further information about the upper tail of the conditional distribution of earnings changes. To this goal, we present the plots of quantiles in Figure (ref). These plots illustrate the goodness of the Pareto fit to the data at the tail. In particular, we plot the ratio $(Y_{(1)}/Y_{(K)}, ...,Y_{(K-1)}/Y_{(K)})$, with $K = 0.3N$ for the empirical quantiles (solid blue lines) against the Pareto-fitted ones (dashed black lines), obtained as $t^{-1/{\hat{\alpha}(x_0)}}$ for $t \in [0,1]$. Here $Y_{(i)}$ denotes the $i$-th largest observation among all $Y_i$'s in the subsample and $\hat{\alpha}(x_0)$ denotes the classic Hill's estimator. The figure shows a very good Pareto fit at the tail, providing suggestive evidence that the tail follows a power law behavior. The evidence in the data also suggests that the largest observations are far away from each other and do not converge to any finite constant. These features imply that a truncation or a finite upper bound is not coherent with our data set. There are sufficient reasons to care about extreme events in untruncated distributions in this application.

figure[figure omitted — 713 chars of source]

Unconditional Tail Risk

Thus far, our analyses have focused on conditional distributions of earnings changes following the literature. In this section, we conduct the analysis for the unconditional distribution of earnings changes. Table (ref) reports the estimates of the Pareto exponents $\alpha$ for men and women, respectively, for the years 2007--2008 and 2015--2016. Table (ref) further shows the sensitivity of the results to different choices of the order statistics $K$ used to build the Hill's estimator.

table[table omitted — 1,473 chars of source]

Furthermore, we provide more information about the upper tail of the unconditional distribution of earnings changes in our data. We report the Pareto plots in Figure (ref) to illustrate the goodness of the Pareto fit to the data at the tail. Specifically, we plot the ratio $(Y_{(1)}/Y_{(K)}, ...,Y_{(K-1)}/Y_{(K)}$, with $K = 0.3N$ for the empirical quantiles (solid blue line) against the Pareto fitted ones (dashed black line), obtained as $t^{-1/{\hat{\alpha}}}$ for $t \in [0,1]$. Here $Y_{(i)}$ denotes the $i$-th largest observations among all $Y_i$'s and $\hat{\alpha}$ again denotes the Hill estimator using these $Y_i$'s. Figure (ref) shows the Pareto tail fit for the sample of men (left panel) and for two other restricted samples of men: job stayers as defined in Section (ref) and those with non-missing observations from $t-3$ to $t+1$ men, with $t$ being 2007.

This evidence is in line with the log-log density of one-year and five-year earnings changes provided in the supplemental material of bell2021time. Specifically, they estimate the slope for the left and right tails respectively for the region ($-$4,$-$1) and (1,4), which can be interpreted as the estimates of the Pareto exponent.

figure[figure omitted — 649 chars of source]

Discussion

Our findings confirm that moment-based measures, such as variance, skewness, or kurtosis, may not be well-defined for some sub-populations. This result points to the importance of using alternatives to the moment-based measures of earnings risk typically used in the literature. We propose a measure of conditional tail earnings risk that overcomes this limitation and offers new insights into the amount of tail risk. Our measure quantifies the amount of risk related to, possibly extreme, shocks individuals experience, given their lagged income and age profile.

Our first finding is that, compared to younger individuals, older individuals are more likely to be hit by potentially large shocks, e.g., health shocks or occupation changes. For the US, guvenen2021data document that earnings growth of males aged 45-55 and earning \$100,000 has a kurtosis of 18, compared to 5 for younger workers earning \$10,000 (and 3 for a Gaussian distribution). These findings highlight the importance of life cycle effects in tail earnings risk, but these estimates are unreliable since kurtosis might not exist in the first place. Our methodology allows measuring this conditional tail risk and indeed reveals that older workers are much more vulnerable to extreme shocks (estimates of our measure are close to 1 along a large part of the distribution of lagged earnings) compared to younger individuals (for which our estimates are close to 2 or larger). Similarly, in our second finding, we show that job stayers face earnings shocks that are more leptokurtic, especially at the bottom of the earnings distribution (at the bottom of the earnings distribution, estimates of our measure go from 5 for the whole sample to around 1 for job stayers). In addition, the frequency and occurrence of leptokurtic shocks do not seem to differ in recessions. A further investigation of the nature of these shocks might shed more light in this direction. To sum up, our results for the whole sample demonstrate that tail earnings risk increases over the life cycle and is higher for job stayers than for movers. Moreover, at least for the male subsample, there are no relevant differences between recession and expansion periods in our proposed measure of tail earnings risk.

It is worth discussing how our methodology and results differ from those of other papers in the literature on earnings risks. guvenen2021data, among others, use moment-based measures, which we have shown to be not well-defined, and complement the analysis with quantile-based measures. Although well-defined when moments may fail to exist, quantile-based measures do not capture extreme events since they do not include observations larger or smaller than some given percentiles. Given our interest in conditional tail earnings risk, the excluded tail observations contain valuable information for the analysis. Among other things, information on tail events is central to quantifying earnings risk under heavy-tailed distributions. Thus, the quantile-based approach may not be informative about extreme income changes and tail events.

Another measure proposed in the literature is the coefficient of variation (CV) of arellano2021, computed as the ratio between the mean absolute deviation of income divided by the mean income, conditional on a set of variables. Estimating the numerator and denominator of the CV are two prediction tasks performed with Poisson regressions with both micro and macro predictors. The main idea is to capture the agent's prediction problem by focusing on features of the predictive distribution of income.\footnote{We refer to arellano2021 for a detailed description of their methodology and the set of predictors used in the analysis. Note that the CV remains well-defined with zero income.} One of their main findings is that younger individuals face higher levels of income risk, and the dispersion of the CV markedly increases in recessions. These results contrast what we obtain with our measure of conditional tail risk. In order to understand what drives the differences between our results and their findings, we build the CV measure of earnings risk with NESPD for the UK to separate the country and institutional effect from the differences due to the different methodology and object of interest. Indeed, the divergences could reflect differences in terms of institutional context, sample selection, and the types of measures used to quantify income risk. There are several constraints in replicating their approach due to some data limitations. For instance, in our case, earnings are the only source of income as we do not observe unemployment benefits. We thus capture a less broad measure of risk compared to theirs.\footnote{When running the regression with both micro and macro variables, we cannot include the highest level of education achieved, unemployment benefits, or the number of days worked in previous years since these variables are not available in our data set.} Acknowledging these data limitations, we computed the CV for UK workers.\footnote{Results for the replication of \citet*{arellano2021} for the UK are available upon request.} We get qualitatively similar findings to those in arellano2021, although the magnitude of the coefficient of variation is different. In particular, we get a smaller dispersion of the CV compared to what they report for Spain. As in their case, however, we obtain that younger individuals face higher risk: for this subgroup, the coefficient of variation is higher, with the 90th percentile of their CV being much larger than the corresponding CV percentile for other age groups. Thus, institutional settings, different sample sizes, and data sets matter for the magnitude of the results, not for the qualitative differences. We can conclude that the main differences between our findings and theirs come from the different concepts of risk we are measuring rather than differences in context (Spain vs. the UK in this particular case). As with quantile-based measures in guvenen2021data, also with the CV measure proposed by arellano2021, one might fail to capture extreme income changes and tail events. Our measure captures a different aspect of earnings risk: it quantifies tail risk, and thus, it should be seen as complementary to their proposal.

We further conduct the following robustness check, focusing on earnings changes with the benchmark sample of men in 2007--2008. Following the approach used in the GRID project, we condition on the average of the three previous earnings, going from $t-1$ backward, rather than on base year earnings, i.e., earnings at time $t$. Figure (ref), in Appendix (ref), plots the estimates, along with the one-sided 95% confidence intervals, of the Pareto exponents $\alpha(x_0)$ of the conditional tail risk for men (left) and women (right) based on the NESPD in the period 2007--2008. This figure illustrates the robustness of our main findings: the increase of tail earning risk over the life cycle.

Finally, in our baseline analysis, we define $Y$ as the absolute value of log earnings changes. We repeat the analysis separately for the left and the right tail of the distribution of earnings changes.\footnote{Results for the right and left tails are available upon request.} For the left and right tails separately considered, the estimates of the Pareto exponent are always not significantly smaller than 4, and point estimates can get much higher values (up to 50 for the right tail, for some subpopulations). Overall, the patterns of the left and right tails resemble those of the absolute changes. However, despite these similarities, there are some remarkable differences, especially at the bottom and top of the earnings distribution. More specifically, for the left tail, estimates tend to be lower (taking a maximum value of around 11) and relatively constant or decreasing over the earnings quantiles. Instead, for the right tail, estimates are much more variable and lower at the bottom of the earnings distribution. This finding is consistent with the patterns of asymmetry of conditional Kelly's skewness documented in Figure (ref): higher earners are more likely to face large negative earnings shocks, because of limited opportunities for large gains and higher risk of significant declines, while the vice versa holds for workers at the bottom of the earnings distribution.

{\bf Implications:} We close this section with a discussion of some alternative methods that are robust to the non-existence of moments. First, if our risk measure implies that the $p$-th moment may not exist, then researchers should select only those calibration/estimation methods that exploit only up to the $(p-1)$-th moments. For instance, many estimation/calibration methods proposed in the literature, including \citet*{daly2022improving} as well as others like \citet*{meghir2004income}, require only up to second moments of earnings/income changes to exist.

Second, there are other estimation strategies that are not based on the method of moments. For instance, the estimation strategy proposed by \citet*{arellano2017earnings} is theoretically robust to the non-existence of moments. Although in their current implementation, they implicitly impose that moments exist, this is not required and, thus, a careful choice of some elements of their implementation, e.g., changing the Laplace tail assumption of the transitory shock, would suffice to make their proposal robust to the existence of moments. \citet*{de2020nonlinear} require that the conditional probability distributions over the bins they define exist, and \citet*{janssens2022finite} only impose some assumptions of continuity and the use of a sufficient number of grid points. Improvements in the discretization procedure might be obtained by careful tailoring of the approximation strategy to the tails (for instance, adapting \citet*{tauchen1986finite} to non-normal errors for the tail behavior) or by applying methods that rely on moments matching (as in \citet*{farmer2017discretizing} and \citet*{lkhagvasuren2023finite}) but restricting them to the existing moments.\footnote{The method proposed by Rouwenhorst (1995) relies on matching the first two moments only and is applicable to Gaussian AR(1) processes. Given the heavy tail feature, it is unappealing for earnings processes.}

Concluding Remarks

The literature often relies on moment-based measures of earnings risk, such as the variance, skewness, and kurtosis. However, such moments may not exist in the population under heavy-tailed distributions. In this paper, we show that the population kurtosis, skewness, and even variance indeed often fail to exist for the conditional distribution of earnings growths given age, gender, and past earnings. Hence, moment-based analyses in the literature may not make sense under heavy-tailed distributions.

Heavy tails of income and earnings risk distributions are costly in economies with risk-averse agents. Despite this common knowledge, the tail heaviness has been arguably less investigated in the literature on earnings and income dynamics. In light of the limited capacity of moment-based measures in quantifying the tail-heaviness, we consider the conditional Pareto exponent as a novel robust measure of conditional earnings risk given observed attributes of individuals and propose a method of estimation and inference about this measure.

Applying the proposed method to the UK NESPD and, in the Appendix, to the US PSID, we obtain the following findings. First, as emphasized above, the population kurtosis, skewness, and even standard deviation might be infinite for the conditional distribution of income growth given certain attributes, such as age, gender, and lagged income, and hence, their sample counterparts might be less informative. Second, 40- and 50-year-old workers have overall higher earnings risk than 30-year-old workers, thus earnings risk increases over the life cycle. Third, conditional earnings risk is higher for job stayers, in particular, at the bottom of the earnings distribution. Fourth, these patterns appear both in the periods of great recession and positive growth, while there are differences as well, especially for women.

To the best of our knowledge, this is the first work to shed light on tail features of heavy-tailed distributions of earnings and income growth with a formal econometric method. While we focused on two specific data sets for empirical analysis, we hope that our proposed method will spur further empirical research with other data sets and for the study of other economies.