Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,106 characters · 12 sections · 70 citation commands
Capital and Labor Income Pareto Exponents across Time and Space
The purpose of this paper is to estimate and document the Pareto exponents for capital and labor income separately for as many countries and years as possible. We say that a positive random variable $X$ obeys a power law with Pareto exponent $\alpha>0$ if the tail probability decays like a power function: $\operatorname{P}(X>x)\sim x^{-\alpha}$ for large $x$.\footnote{More precisely, we say that a random variable $X$ has a Pareto upper tail with exponent $\alpha>0$ if $\operatorname{P}(X>x)=x^{-\alpha}\ell(x)$ for some slowly varying function $\ell$. A function $\ell:(0,\infty)\to \R$ is said to be slowly varying (at infinity) if it is nonzero for sufficiently large $x$ and $\lim_{x\to\infty}\ell(tx)/\ell(x)=1$ for each $t>0$. See BinghamGoldieTeugels1987 for a comprehensive treatment of the theory of regular variation.} In the context of the income distribution, the Pareto exponent characterizes the tail heaviness of high incomes and hence top tail inequality. We remain agnostic about the shape of the income distribution away from the tail. Our study is motivated by the following two observations. First, we are not aware of a comprehensive study that documents the capital and labor income Pareto exponents separately for many countries and years, despite their importance. Second, the Pareto exponent has desirable properties relative to other popular inequality measures such as the Gini coefficient or top income shares.
Consider the first point. Conceptually, capital and labor income are different entities. While the former is the return for providing capital (wealth), the latter is the return for providing labor services, and there is no particular reason to expect a relation between the two. Although these two forms of income are conceptually distinct, it is often put together as just “income” and discussed in the context of inequality and related policies. If capital and labor income are quantitatively different, a policy design based on total income may be misleading. To give one example, consider the theory of optimal taxation Saez2001, where the income Pareto exponent plays an important role. SaezStancheva2018 carefully distinguish capital and labor income and apply the theory of optimal taxation in the United States. They find that with an income elasticity of $e=0.5$, the optimal top marginal tax rate is about 50% for labor and 60% for capital (see their Figure 5). This difference directly comes from the fact that capital and labor income Pareto exponents are distinct. Thus, distinguishing capital and labor income inequality is potentially important for policy designs.
Consider the second point about the desirable properties of Pareto exponents. In the applied literature such as Piketty2003 and piketty-saez2003, top income shares (such as the top 1% income share) are more commonly reported than the income Pareto exponent, perhaps because top shares are summary statistics that can be computed without specifying functional forms or can be understood by non-experts without special knowledge of statistics. However, atkinson2005comparing documents methodological problems regarding the cross-country comparison of top income shares, citing the differences in tax units (\eg, individual or household) and legislation (\eg, whether social security benefits are taxable). One of the reasons such issues arise is because it is not always clear how to define the population and measure small units.\footnote{Imagine how to formally distinguish cities, towns, villages, and settlements; continents, islands, and islets; and inland seas, lakes, and ponds. How to define units and how to measure small units matter for the size distribution of population, land mass, and water surface area.} Because common inequality measures such as the Gini coefficient and top income shares require the knowledge of the entire distribution, these quantities are greatly affected by the definition and measurement of small units. Using the Pareto exponents significantly alleviates these definition and measurement issues because the Pareto distribution is scale invariant (see JessenMikosch2006 for a summary) and its exponent depends only on the tail behavior, not the entire distribution. For example, doubling the income of all households in the top 1% of the income distribution makes the top 1% income share (roughly) twice as large, but the Pareto exponent is unaffected. A similar comment applies to any inequality measure that depends on the entire distribution, such as the Gini coefficient. While we do not claim that the Pareto exponent is the only interesting inequality measure, it is certainly a robust (detail-independent) measure for top tail inequality. See gabaix2009,Gabaix2016JEP for more discussion on the robustness of the Pareto distribution.
In this paper, we use the harmonized Luxembourg Income Study database (hereafter \citetalias{LIS2021}) to document the capital and labor income Pareto exponents across all available 475\xspace country-year observations that span 52\xspace countries over half a century. We document two empirical findings. First, we find that the capital income Pareto exponent is roughly in the range 1--3 (with median 1.46\xspace) and is smaller than the labor income Pareto exponent, which ranges between 2--5 (with median 3.35\xspace). This implies that capital income is more unequally distributed than labor income. This fact is unsurprising and well known for a specific country or year (see, for example, the Lorenz curve in Figure 1 of SaezStancheva2018). However, we are not aware of a comprehensive study that systematically analyzes datasets from many countries and years, and therefore our finding suggests that capital income is generally more unequal than labor income. More specifically, using a statistical test recently developed by hoga2018detecting, we formally test the equality of capital and labor income Pareto exponents and the null is rejected in 86%\xspace of samples. In every single case of rejection, the capital exponent is smaller than the labor exponent. Second, we find that the capital income Pareto exponent is nearly unrelated to the labor exponent. In particular, the correlation between the two exponents across countries is close to zero.
To explain our empirical findings, we build a simple incomplete market model with job ladders and capital income risk. In the model, agents get randomly promoted to the next job ladder. Because individual income follows a random multiplicative process, we obtain a Pareto-tailed labor income distribution. The agents also save assets and face idiosyncratic investment risk, which generates a Pareto-tailed wealth (hence capital income) distribution. Because the capital income Pareto exponent is mainly determined by the asset return distribution, while the labor income Pareto exponent is mainly determined by the income growth distribution, the relation between the two is weak. Furthermore, we analytically characterize the capital and labor income Pareto exponents and show that the former tends to be smaller than the latter for common parametrization. Our results suggest the importance of distinguishing income and wealth inequality.
\paragraph{Related literature}
The power law behavior of income was first recognized by Pareto1895LaLegge,Pareto1896LaCourbe,Pareto1897Cours, who used tabulation data of tax returns in many European countries. More recent research that employs micro data include reed2001 for U.S., reed2003 for U.S., Canada, Sri Lanka, and Bohemia, nirei-souma2007 for Japan, Toda2011PRE,Toda2012JEBO for U.S., Jenkins2017 for U.K., and IbragimovIbragimov2018 for Russia. These papers all concern specific countries and years. BandourianMcDonaldTurley2003 estimate eleven parametric distributions (some of which exhibit Pareto tails) using 82 household labor income datasets from Luxembourg Income Study \citepalias{LIS2021} as we do, though they neither focus on the Pareto exponent nor consider capital income. gabaix2009 mentions “[the] tail exponent of income seems to vary between 1.5 and 3”, citing AtkinsonPiketty2007, though without providing specific details. AtkinsonPiketty2010 document income Pareto exponents across many countries and years estimated from top income share data based on tax returns. However, these estimates are computed from total income, and since (as we document in Section (ref)) the capital income Pareto exponents tend to be smaller than labor exponents, their estimates are best understood as capital income (hence wealth) Pareto exponents. BenhabibBisinLuo2017 make the point that wealth is more skewed than income, citing a few Pareto exponent estimates from BadelDalyHuggettNybom2018 for income and Vermeulen2018 for wealth. Using the 2013-2014 individual income tax data from Romania, Oancea_2018 document that the capital income Pareto exponent (1.44) is smaller than the labor exponent (2.53). Using time series regressions for 21 countries, Bengtsson_2018 document a positive correlation between the gross capital share in national accounts and the top 1% income shares. Their finding can be explained if top income earners tend to have higher capital income share, which is the case if the capital income Pareto exponent is smaller than the labor exponent as we document in this paper. As mentioned in the introduction, there seems to be no comprehensive studies that document the capital and labor income Pareto exponents separately for many countries and years.
In this section we describe the dataset that we use and discuss its limitations.
We use the data from the Luxembourg Income Study \citepalias{LIS2021}, which is a large, harmonized database of micro-level income data that covers over 50 countries worldwide and many years since the late 1960s. In many countries, the data derive from government surveys (for example, the U.S.\ data is based on the Current Population Survey). The \citetalias{LIS2021} data are available at both individual and household level. We focus on the household labor and capital income because
The \citetalias{LIS2021} defines labor income as “cash payments and value of goods and services received from dependent employment, as well as profits/losses and value of goods from self-employment, including own consumption”. Capital income is defined as “cash payments from property and capital (including financial and non-financial assets), including interest and dividends, rental income and royalties, and other capital income from investment in self-employment activity”. Together these two categories make up total factor income. See the \citetalias{LIS2021} 2019 USER GUIDE\footnote{https://www.lisdatacenter.org/wp-content/uploads/files/data-lis-guide.pdf} for a detailed summary on how these data are retrieved and calculated.
Our analysis draws upon datasets from many different countries that are harmonized into a common framework by the \citetalias{LIS2021}. However, many details about the collection of data in the different countries are omitted. For example, we find evidence of top-coding in some countries and years, as the largest income order statistic is equal to the second largest.\footnote{Among all 475\xspace country-year observations, the first and second order statistics are equal in 6 cases for labor income and 11 cases for capital income. Therefore we conjecture that the top-coding issue is not severe.} Top-coding induces an upward bias in the estimation of the Pareto exponent. This issue is not necessarily resolved if, instead, one relies on administrative tax income data, for similar biases arise such as rich households trying to understate their taxable income atkinson2011top. burkhauser2012recent detail a method that can be used to overcome the bias due to top-coding, however at the end of their paper they show that the results are robust even if estimates are based on the top-coded series. For these reasons we treat the datasets as not being top-coded in our analysis.
Another limitation of the \citetalias{LIS2021} database is that it is based on government surveys and the measurement error may be larger compared to administrative data based on tax returns. The fact that the income distribution in administrative data is often reported as tabulations, not micro data, causes no problem for estimating Pareto exponents, as TodaWang2021JAE provide an efficient estimation method for such data. In fact, AtkinsonPiketty2010 document income Pareto exponents across countries and years estimated from top income share data. However, their table is based on total income, and since (as we document in Section (ref)) the capital income Pareto exponents tend to be smaller than labor exponents, the estimates in AtkinsonPiketty2010 are best understood as capital income (hence wealth) Pareto exponents. Since we are not aware of a comprehensive income database that distinguishes capital and labor income, we decided to use the \citetalias{LIS2021} database.
In this section we estimate the capital and labor income Pareto exponents for all countries and years that are available in the \citetalias{LIS2021} database, which spans 52\xspace countries over half a century (475\xspace country-year observations in total). We then formally test for the equality of the Pareto exponents of capital and labor income.
For each country and year, we suppose that the (capital or labor) income observations $\set{X_n}_{n=1}^N$ are independent and identically distributed (\iid) with cumulative distribution function (CDF) $F(x)=\operatorname{P}(X_n \le x)$. The assumption that the upper tail of income obeys a power law with Pareto exponent $\alpha>0$ translates into the regular variation condition
for some slowly varying function $\ell$ (see Footnote (ref)). Note that the assumption on $\ell$ involves only the limit as $x\to\infty$; we are assuming a power law behavior in the upper tail without taking a stance on the shape of the entire distribution.
We are interested in estimating the Pareto exponent $\alpha$ for each country and year. For this purpose, we employ the hill1975simple (maximum likelihood) estimator
Here $X_{(n)}$ denotes the $n$-th largest order statistic from the sample $\set{X_n}_{n=1}^N$ and $k\in \set{1,\dots,N}$ denotes the number of tail observations used to estimate the Pareto exponent. Hall1982 shows that the standard error of $\widehat{\alpha}(k)$ is $\widehat{\alpha}(k)/\sqrt{k}$ under flexible assumptions on the CDF (see (ref) below). We use the Hill estimator because we are interested in formally testing the equality of the capital and labor income Pareto exponents using the method of hoga2018detecting, for which the Hill estimator is required.\footnote{If we are only interested in estimating the Pareto exponents, then there are many alternative methods available. GomesGuillou2015 review 13 commonly used estimators. Fedotenkov2020 reviews more than 100.}
When the population distribution is known to be exactly Pareto (so $\ell$ in (ref) is zero below the minimum size $x_{\min}$ and constant above this threshold), it is well known that the Hill estimator for the full sample ($k=N$) is consistent, asymptotically normal, and asymptotically efficient because it is a maximum likelihood estimator. In practice, the CDF is not exactly Pareto and the researcher needs to select an appropriate value of $k$. For instance, if $F(x)$ satisfies
with some $\beta>0$, then Hall1982 shows that choosing $k = o(N^{2\beta/(2\beta + \alpha)})$ together with $k \to \infty$ as $N \to \infty$ is sufficient for consistency and asymptotic normality (see also embrechts2013modelling).\footnote{Recall that we write $f(x)=o(g(x))$ if for all $\epsilon>0$, we have $\abs{f(x)}\le \epsilon\abs{g(x)}$ for large enough $x$.} Notice that this choice puts a bound on the growth rate of $k$. The expansion (ref) covers a wide range of distributions of interest, such as the $t$-distribution and the type II extreme value distribution danielsson1997tail.
Despite these asymptotic results, it is notoriously difficult to pick $k$ optimally in finite samples Hall1990,ResnickStarica1997,Danielsson_2001. In practice, researchers often plot the Hill estimator (ref) over a range of $k$ to find a flat region or plot the log rank $\log 1,\dots,\log N$ against the log size $\log X_{(1)},\dots,\log X_{(N)}$ to find a region that exhibits a straight line pattern and choose a size threshold to run the log-rank regression.\footnote{gabaix-ibragimov2011 study the asymptotic behavior of log rank regression and show that the standard error is larger by a factor of $\sqrt{2}$ than the Hill estimator. However, they do not discuss how to select the threshold. In their empirical application, they consider the size distribution of population in U.S.\ metropolitan statistical areas, which are already far into the tail and hence the threshold selection is less of an issue. IbragimovIbragimov2018 apply the same methodology to Russian household income data and consider the top 5% and 10% thresholds.} Unfortunately, this graphical approach is not feasible in our setting because \citetalias{LIS2021} does not allow researchers to download the micro data for confidentiality concerns (researchers are required to submit their execution files to conduct statistical analyses) and there is little scope for exploratory graphical data analysis. In this paper we simply use the largest 5% observations, so
which is standard in the literature.\footnote{An alternative approach is to estimate a parametric distribution $F$ that admits a Pareto upper tail by maximum likelihood using the entire sample. The double Pareto-lognormal distribution proposed by reed2003 and reed-jorgensen2004 often performs best. See Toda2012JEBO for a horse race across several parametric distributions in the context of U.S.\ labor income.} Unreported simulations show that our results are robust to using other thresholds such as the largest 1% and 10% observations, or using a data-driven procedure using the Kolmogorov-Smirnov distance as in danielsson2016tail.
We estimate the capital and labor income Pareto exponents for all countries and years available in the \citetalias{LIS2021} database. The database spans 52\xspace countries across the years 1967--2018, with a total of 475\xspace country-year observations. The point estimates of the capital and labor income Pareto exponents for each country and year as well as their standard errors can be found in Table (ref) in Appendix (ref). To avoid small sample issues, we restrict our analysis to countries with at least 1,000\xspace positive observations for income, resulting in (all) 475\xspace country-year pairs for labor income and 342\xspace for capital income. For visibility, Figure (ref) shows the histogram and scatter plot of the capital and labor income Pareto exponents.
Figure (ref) shows the histogram of the capital and labor income Pareto exponents pooled across all available countries and years. The capital and labor income Pareto exponents are generally in the range 1--3 and 2--5 with medians 1.46\xspace and 3.35\xspace, respectively. This suggests that
Figure (ref) shows the scatter plot of the Pareto exponents together with the 45 degree line. The confidence interval of the correlation coefficient $\rho$ is computed assuming all country-year observations are independent and there is no sampling error in the point estimates of the Pareto exponents. Although it is not obvious how to account for these issues, doing so will only widen the confidence interval. Therefore the fact that the na\"{\i}ve confidence interval almost contains zero suggests that the correlation between the two Pareto exponents is weak. Furthermore, for the vast majority of countries and years, the capital income Pareto exponent is smaller than the labor exponent, again suggesting that capital income is more unequal than labor income.
How do the Pareto exponents evolve over time? Because many countries appear only sporadically in the \citetalias{LIS2021} database, we only consider six countries for which the cross-sectional sample size is large and the time series is long, namely Canada, Germany, Switzerland, Taiwan, United Kingdom, and United States. Figure (ref) plots the capital and labor Pareto exponents and their 95% confidence intervals for these countries. The capital exponent tends to be stable at around 1--2 and is smaller than the labor exponent.\footnote{An exception may be Germany before 1983. However, this may be an artifact of the sudden change in sample size, which was over 40,000 until 1983 and about 5,000 since 1984. Thus the 5% rule for capital income may be including too many observations in the body of the distribution before 1983.} Again, capital income appears to be more unequally distributed than the labor income.
We now formally test whether the capital and labor Pareto exponents are equal. In particular our test is
where $\alpha_\text{cap},\alpha_\text{lab}$ denote the capital and labor income Pareto exponents. Testing the null hypothesis $H_0$ is complicated by the fact that there is dependency between labor income $\set{X_{\text{lab},n}}_{n=1}^N$ and capital income $\set{X_{\text{cap},n}}_{n=1}^N$, because individuals who are rich (receive high labor income) tend to be wealthy and receive high capital income. Thus we cannot use the 95% confidence intervals in Figure (ref) to test the equality of Pareto exponents. Instead, we apply the test recently developed by hoga2018detecting, which allows for dependence in the data but assumes restrictions on the growth rate of the tail dependence (see hoga2018detecting). The test is based upon the inverse of the Hill estimator (ref), which we denote by $\widehat{\gamma} \coloneqq 1/\widehat{\alpha}$. The test statistic is defined by
where $t_0 \in (0,1)$ is a tuning parameter and $\widehat{\gamma}(t)$ is the inverse Hill estimator
Using the Hill estimator based on the subsample with only $\floor{kt}$ observations leads to self-normalization of the test statistic $T_N$ and renders a test that is asymptotically pivotal. The limiting distribution is
where $W(t)$ is a standard Brownian motion. Since the test statistic (ref) can be computed using only the Hill estimator and conducting numerical integration, there is no need to estimate the (potentially difficult) tail covariance. The tuning parameter $t_0$ affects the size of the test in finite samples: high values of $t_0$ make the integral in (ref) based on too few differences of $\widehat{\gamma}$, and low values of $t_0$ yield volatile $\widehat{\gamma}$ in (ref) when $t$ is close to $t_0$. Both of these effects may cause size distortions. Therefore we set $t_0 = 0.2$ following the recommendation of hoga2018detecting, who finds that this choice leads to favorable size properties.\footnote{The choice of $t_0$ in hoga2018detecting comes out using an automated selection procedure to choose $k$, which is different than our 5% rule. An earlier version of our paper employs the same automated procedure, which leads to very similar results.} We reject the null $H_0$ when the test statistic $T_N$ is large. According to Table I of hoga2018detecting, the 95 percentile of (ref) for $t_0=0.2$ is 55.44, which we use as the critical value for testing $H_0$ at 5% significance level.
One issue with the test statistic (ref) is that it requires the same number of tail observations $k$ for both cross-sections of capital and labor income. Hence the 5% rule (ref) discussed in Section (ref) becomes problematic as the number of people with capital income in our data set is rather small. Many households do not hold liquid financial wealth and hence have no capital income. The resulting test is thus not feasible since $k_\text{lab}$ based on our 5% rule could be wildly different from $k_\text{cap}$. To overcome this issue, we only test the equality of Pareto exponents for countries that have more than 1,000\xspace positive capital income observations and set $k = \floor{0.05N_\text{cap}}$, where $N_\text{cap}$ is the number of positive capital income observations (these households always have positive labor income), resulting in 342\xspace country-year observations out of 475\xspace. In practice this means that we estimate the Pareto exponent of labor income further in the tail because the sample size of labor income $N_\text{lab}$ tends to be larger than $N_\text{cap}$. However, this is acceptable since the Pareto approximation tends to fit better for smaller $k$. Figure (ref) presents a scatter plot of the estimated labor income Pareto exponent $\widehat{\alpha}_\text{lab}$ using 5% of the full sample ($N_\text{lab}$) and 5% of the sample with positive capital income ($N_\text{cap}$). The fact that most points are close to the 45 degree supports our claim.
Table (ref) in Appendix (ref) shows the test results of the null hypothesis $H_0: \alpha_\text{lab} = \alpha_\text{cap}$. We reject the null in 294\xspace country-year observations out of 342\xspace (86%\xspace) that meet our sample selection criterion. In every single case of rejection, we have $\widehat{\alpha}_\text{lab}>\widehat{\alpha}_\text{cap}$, and therefore we formally confirm the observation in Section (ref) that capital income is more unequally distributed than labor income.
Our empirical analysis in Section (ref) suggests that
To explain these empirical findings, we present a simple dynamic model of consumption and savings, which builds on one of the authors' prior works MaStachurskiToda2020JET,MaToda2021JET. Our model is more specialized but the characterizations are sharper. The proofs of propositions in this section are deferred to Appendix (ref).
Time is discrete and denoted by $t=0,1,2,\dotsc$. Let $a_t$ be the financial wealth of a typical agent at the beginning of period $t$ including current income. The agent chooses consumption $c_t\ge 0$ and saves the remaining wealth $a_t-c_t$. The period utility function is $u:(0,\infty)\to \R$, the discount factor is $\beta>0$, the gross return on wealth between time $t-1$ and $t$ is $R_t>0$, and non-financial income at time $t$ is $Y_t>0$. Thus the agent solves
where the initial wealth $a_0=a>0$ is given, (ref) is the budget constraint, and (ref) implies that the agent cannot borrow (which is without loss of generality according to the discussion in ChamberlainWilson2000). Throughout the rest of the paper we maintain the following assumptions.
These assumptions are similar to Carroll2020, except that we allow for stochastic returns on savings. Note that the asset return $R_{t+1}$ and income growth $G_{t+1}$ are potentially mutually dependent. Due to the \iid assumption, the state variables of the income fluctuation problem (ref) are financial wealth $a_t>0$ and current income $Y_t>0$. Exploiting homotheticity (Assumption (ref)), we can reduce the number of state variables to just one, namely the wealth-income ratio (normalized wealth) $\tilde{a}_t\coloneqq a_t/Y_t$. To see this, letting $\tilde{c}_t\coloneqq c_t/Y_t$ be the consumption-income ratio (normalized consumption), dividing the borrowing constraint (ref) by $Y_t$, we obtain $0\le \tilde{c}_t\le \tilde{a}_t$. Similarly, dividing the budget constraint (ref) by $Y_{t+1}$, we obtain
where $\tilde{R}_{t+1}\coloneqq R_{t+1}/G_{t+1}$ is the asset return relative to income growth. As for the utility function, since
(here we interpret $\prod_{s=1}^0\bullet=1$), assuming $Y_0=1$ (which is without loss of generality) and $\gamma\neq 1$, it follows from (ref) that
where $\tilde{\beta}_t\coloneqq \beta G_t^{1-\gamma}$. The discussion for $\gamma=1$ is similar. Therefore the problem reduces to an income fluctuation problem with CRRA utility, random discount factors $\set{\tilde{\beta}_t}_{t=1}^\infty$, stochastic returns $\set{\tilde{R}_t}_{t=1}^\infty$ on wealth, and constant income ($\tilde{Y}_t\equiv 1$). The general theory of income fluctuation problems with stochastic discounting, returns, and income in a Markovian setting was developed by MaStachurskiToda2020JET. Therefore we immediately obtain the following result. In what follows, we drop the time subscript when no confusion arises.
We now characterize the tail behavior of income and wealth in the context of the income fluctuation problem in Section (ref).
To make the model stationary, suppose that agents survive to the next period with probability $v\in (0,1)$ (perpetual youth model as in yaari1965). Whenever agents die, they are replaced by newborn agents. For simplicity, assume that the discount factor $\beta$ in (ref) already accounts for survival probability and that there is no market for life insurance (allowing for life insurance only changes $R$ to $R/v$ and is thus mathematically equivalent after reparametrization). Without loss of generality, suppose that newborn agents start with income $Y_0=1$. Then the income of a randomly selected agent is $Y_T$, where $T$ is a geometric random variable with mean $\frac{1}{1-v}$. By the assumption on income growth, the log income of a randomly selected agent
is a geometric sum of \iid random variables, for which we can characterize the tail behavior as follows.
To characterize the tail behavior of wealth, we first note that the normalized consumption function $\tilde{c}$ in Proposition (ref) is concave and asymptotically linear with a specific slope.\footnote{CarrollKimball1996 showed the sufficiency of hyperbolic absolute risk aversion (HARA, which includes CRRA) for the concavity of the consumption function. Toda2021JME proved the necessity.}
Using Proposition (ref) and setting $\rho=\min\set{(\E [\beta R^{1-\gamma}])^{1/\gamma},1}$, for high enough asset level, the detrended budget constraint (ref) becomes approximately
which is a random multiplicative process (kesten1973 process). Under specific assumptions, MaStachurskiToda2020JET prove that the upper tail of the stationary distribution of normalized wealth $\tilde{a}_t$ has a Pareto lower bound. Although a sharp characterization of the tail behavior is generally difficult, in our setting it is possible to obtain an exact characterization due to concavity and the \iid assumption.
Proposition (ref) is significant despite its simplicity. According to the model, we always have $\alpha\le \alpha_Y$, typically with a strict inequality as we see in the numerical example below. This is in sharp contrast to canonical incomplete market general equilibrium models such as aiyagari1994, where agents can save using only a risk-free asset. In such models, the impossibility theorem of StachurskiToda2019JET implies that the tail behavior of income and wealth is the same, implying $\alpha=\alpha_Y$ in our setting.\footnote{The original proof in StachurskiToda2019JET contained an error; it has been corrected in StachurskiToda2020Corrigendum.} Therefore, unlike canonical incomplete market models, our model can explain the empirical fact that capital income is more unequal than labor income. The key assumption leading to this conclusion is the presence of stochastic returns.
We discuss an analytically solvable example to build intuition.
We further examine the tail behavior of income and wealth using a numerical example of the income fluctuation problem (ref). Suppose that asset return is \iid lognormal, so $\log R\sim N((\mu-\sigma^2/2)\Delta,\sigma^2\Delta)$, where $\Delta>0$ is the length of one period, $\mu$ is the expected return, and $\sigma$ is volatility. Suppose every period the agent is “promoted” with some probability, so the income growth rate is
where $p\in (0,1)$ is the promotion probability and $g$ is the log income growth rate conditional on promotion. We parametrize the promotion probability as $p=1-\e^{-\Delta/L}$, where $L$ is the expected length of time until a promotion. Using (ref), the labor income Pareto exponent is determined such that
We set the parameter values as in Table (ref). One unit of time corresponds to a year and one period is a quarter, so $\Delta=1/4$. The preference parameters (discount rate and risk aversion) are standard. The death rate of $\eta=0.025$ implies an average (economically active) age of $1/\eta=40$ years. The expected return and volatility roughly correspond to the stock market. We set the labor income Pareto exponent to $\alpha_Y=3$, which is roughly the median value in Figure (ref). Using the survival probability $v=\e^{-\eta\Delta}$ and (ref), the implied value of income growth upon promotion is $g=0.0403$. The wealth Pareto exponent determined by (ref) is then $\alpha=1.201$.
To numerically solve the income fluctuation problem (ref), we discretize the log asset return $\log R$ using a 7-point Gauss-Hermite quadrature and apply policy function iteration (see Appendix (ref)). After solving the individual problem, we apply the Pareto extrapolation algorithm developed in Gouin-BonenfantTodaParetoExtrapolation to accurately compute the stationary (normalized) wealth distribution. Finally, we also simulate an economy with $10^5$ agents. Figure (ref) shows the results.
Figure (ref) shows the normalized consumption function $\tilde{c}(\tilde{a})$ in the range $\tilde{a}\in [0,100]$. Consistent with Proposition (ref), the consumption function is roughly linear for high asset level. Figure (ref) shows the size distributions of income $Y$ normalized wealth $\tilde{a}=a/Y$ in a log-log plot, both from the theoretical model and the simulation. The fact that the tail probability $\operatorname{P}(X>x)$ exhibits a straight line pattern in a log-log plot suggests that the size distributions have Pareto upper tails, consistent with theory. Furthermore, the slope for income is steeper than that of normalized wealth, so wealth (hence capital income) is more unequally distributed than labor income.
Finally, Figure (ref) shows the income and wealth Pareto exponents when we change the income growth rate $g$ in the range $g\in [0.02,0.1]$, fixing other parameters. Because the income Pareto exponent is inversely proportional to income growth by (ref), the labor income Pareto exponent is highly sensitive to income growth. On the other hand, the wealth (capital income) Pareto exponent does not depend much on income growth by the same intuition as in Example (ref). Thus our model is consistent with our empirical findings in Section (ref) that capital and labor income Pareto exponents are only weakly related.