Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
53,461 characters · 6 sections · 54 citation commands
Determination of Pareto exponents in economic models driven by Markov multiplicative processes
\onehalfspacing
This article presents new tools for studying the shape of the stationary distribution of a dynamic economic system. We have in mind models in which a population of economic units experience random multiplicative shocks to their size over time, occasionally perishing at random and being replaced with a new unit. An economic unit could be, for instance, a household with size measured by wealth, or a firm with size measured by market capitalization. Economic units have a type which may vary over time, such as a worker/entrepreneur type for households, or a productivity type for firms. The type affects the multiplicative growth rate, and perhaps also the survival rate, of an economic unit. Many heterogeneous-agent models of the Bewley-Aiyagari kind fit this general description if agents are subject to random mortality rather than having a certain finite or infinite lifespan.
It has long been known that random multiplicative growth is a generative mechanism for power laws. Random multiplicative growth plays a central role in economics and is known as Gibrat's law, while power laws are often referred to as Pareto tails. Early contributions to economics drawing a connection between random multiplicative growth and power laws include Champernowne1953, WoldWhittle1957, and SimonBonini1958. The topic has attracted a resurgence of interest in economics following the publication of gabaix1999 and reed2001; see gabaix2009,gabaix2016 and benhabibbisin2018 for partial surveys. In mathematics, the central contribution is kesten1973. See Mitzenmacher2004 for a discussion of many other relevant contributions across a range of disciplines.
The primary contribution of this article is an equation whose unique positive solution is the Pareto (power law) exponent for the upper tail of the stationary distribution of sizes in a dynamic economic system of the kind described above. The equation is
where $\rho(\cdot)$ is the spectral radius (maximum modulus of eigenvalues) of a square matrix, $\odot$ is the Hadamard (entry-wise) product of matrices of the same size, $\Pi$ is a matrix of transition probabilities for types, $\Upsilon$ is a matrix of survival rates, and $\Psi(z)$ is a matrix of moment generating functions of random growth rates. If equation (ref) admits a unique positive solution (which is the case under mild conditions), say $z=\alpha$, then our main result, Theorem (ref), establishes that $\alpha$ is the Pareto exponent for the upper tail of the stationary distribution of sizes.
In the simple case where there is only a single type of economic unit (which, in the class of models we consider, implies serial independence of the multiplicative shocks to size and is therefore quite restrictive), equation (ref) simplifies to
where $\upsilon$ is the survival rate and $\psi(z)$ is the moment generating function of the random growth rate, both constant over time. Equation (ref) appears as Equation (10) in ManrubiaZanette1999, and now seems to be fairly widely known. Applications include NireiAoki2016 and MukoyamaOsotimehin2019 in macroeconomics, MonteroVillarroel2013, Yamamoto2014 and MeylahnSabhapanditTouchette2015 in statistical physics, and BeareToda2020PhysD in epidemiology. Equation (ref) is also closely related to the Cram\'{e}r-Lundberg estimate of ruin probabilities used in actuarial science EmbrechtsKluppelbergMikosch1997.
Our equation (ref) extends (ref) to a Markov setting in which economic units switch between types over time. There is a growing recognition in the literature that persistent heterogeneity in types plays a critical role in matching salient empirical regularities. benhabib-bisin-zhu2011 study an overlapping generations model in which agents are subject to random labor (additive) and capital (multiplicative) income shocks. They allow these shocks to be persistent through their dependence on a latent Markov state (i.e., type), writing (p. 127) “the i.i.d.\ condition is very restrictive. Positive autocorrelations in [capital and labor income shocks] capture variations in social mobility in the economy, for example, economies in which returns on wealth and labor earning abilities are transmitted across generations.” CaoLuo2017 study a similar model in which investors switch at random between high and low productivity types, while subject to random mortality. Regarding the importance of persistent heterogeneity in types, they write (p. 302) “the model with homogeneous returns, when calibrated to match the salient aggregate statistics, produces a tail index that is an order of magnitude too high compared to the one in the data.” Based on an analysis of Norwegian administrative tax records, FagerengGuisoMalacrinoPistaferri2020 confirm these observations on the importance of type dependence for explaining the upper tail of the wealth distribution, writing that (pp. 118--119) “persistent traits of individual investors (such as financial sophistication, the ability to process and use financial information, the ability to overcome inertia, and—for entrepreneurs—the talent to manage and organize their businesses), are capable of generating persistent differences in returns to wealth that may be as relevant as those conventionally attributed in household finance to differences in risk exposure or scale”. Taking a somewhat different angle, GabaixLasryLionsMoll2016 emphasize the importance of persistent heterogeneity in productivity types not only for generating an income distribution with a tail sufficiently heavy to match U.S.\ data, but also for generating transition dynamics that are sufficiently fast to match the rate of increase of top income inequality. Additional articles discussing the importance of persistent heterogeneity in types for constructing plausibly calibrated models are discussed and cited therein (p. 2095).
Mathematically, the new results introduced in this article concern a class of Markov chains called Markov multiplicative processes with reset, which we define in Section (ref). Loosely, these are processes driven by positive multiplicative shocks dependent on a Markov state variable, and subject to occasional reset to one at a state-dependent rate. After a heuristic discussion in Section (ref) aimed at building intuition, we state our main results in Section (ref). Under mild conditions, these results establish
In Section (ref) we explore refinements obtaining under a non-lattice condition on growth rates, establishing that
We lay out some directions for future research in Section (ref). A mathematical appendix to Sections (ref)--(ref) (Appendix (ref)) contains the proofs of all numbered results, and a second appendix includes a discussion of comparative statics for the Pareto exponent (Appendix (ref)).
Continuous-time versions of the results in this article have been established in a companion article, BeareSeoToda2021. There we study the tail probabilities of a Markov-modulated L\'{e}vy process stopped at a state-dependent Poisson rate. The tail exponents in continuous-time are determined by an equation similar to (ref), but with a spectral abscissa taking the role of the spectral radius.
While we do not provide a substantive economic application of our results in this article, several recent articles and circulated manuscripts provide applications.\footnote{Articles in which our results are discussed sometimes cite an earlier version of this article titled “Geometrically stopped Markovian random growth processes and Pareto tails”, available on the arXiv e-print repository since December 2017: \url{https://arxiv.org/abs/1712.01431}.} Toda2019 uses equation (ref) to determine the Pareto exponent for the wealth distribution in a Huggett economy with stochastic discounting. MaStachurskiToda2020 do the same in a model with stochastic discounting, returns on wealth, and labor income. GomezGouinbonenfant2020 use our results to determine the Pareto exponent of wealth in an economy populated with entrepreneurs and rentiers. Gouinbonenfant2020 uses our results to calculate the Pareto exponent for the distribution of firm size in a model of the labor market, finding that Zipf's law is approximately satisfied. Gouin-BonenfantTodaParetoExtrapolation show how to improve grid-based calculation of aggregate quantities in dynamic economic models by using equation (ref) to extrapolate beyond the grid. BeareSeoToda2021 use a continuous-time version of equation (ref) to study the tails of wealth in a Huggett economy inhabited by agents with constant absolute risk aversion.
Let $\mathbb{Z}_+$ be the set of nonnegative integers and $\mathcal{N}=\left\{ 1,\dots,N \right\} $, a finite set. The elements of $\mathcal{N}$ index the types of agents in a dynamic economic system. We will study the behavior of a homogeneous\footnote{Homogeneity means that $\mathrm{P}(S_{t+1}\leq s',J_{t+1}=n'\mid S_t=s,J_t=n)$ does not depend on $t$.} Markov chain $(S_t,J_t)_{t\in\mathbb{Z}_+}$ on $(0,\infty)\times\mathcal{N}$. In each time period $t$, $S_t$ indicates the size of an agent (e.g.\ their wealth) and $J_t$ their type. The transition probabilities of our Markov chain are parametrized in terms of four objects.
Conditional on having $S_t=s$ and $J_t=n$, the subsequent pair $(S_{t+1},J_{t+1})$ may be generated in the following way.
Steps (ref)-(ref) determine the Markov kernel for $(S_t,J_t)_{t\in\mathbb{Z}}$. Specifically, given $n,n'\in\mathcal{N}$ and $s,s'>0$, the Markov kernel is
We refer to a homogeneous Markov chain on $(0,\infty)\times\mathcal{N}$ with Markov kernel given by (ref) as a Markov multiplicative process with reset. We call the multiplicative factors $G_{nn'}$ gross growth rates, and their logarithms growth rates.
Under further regularity conditions, a Markov multiplicative process with reset can be made stationary with a suitable choice of its distribution at time $t=0$. This will be made clear in Proposition (ref) below.
To build intuition, in this section we provide a brief heuristic derivation of our equation determining the Pareto exponent of a stationary Markov multiplicative process with reset. Let $(S_t,J_t)_{t\in\mathbb{Z}_+}$ be such a process, and conjecture that for all sufficiently large $s>1$ we have
for some constant $y_n>0$ and Pareto exponent $\alpha>0$. Then
where $G_{nn'}>0$ is the gross growth rate drawn from the CDF $\Phi_{nn'}$ in Step (ref) in Section (ref). Applying the law of iterated expectations, we obtain
For real $z$, set
the moment generating function (MGF) of $\log G_{nn'}$. Let $\Psi(z)$ be the $N\times N$ matrix-valued function with entries $\psi_{nn'}(z)$. Noting that $\psi_{nn'}(\alpha)=\operatorname{E}(G_{nn'}^\alpha)$, and combining (ref), (ref) and (ref), we obtain
Letting $y=(y_1,\dots,y_N)^\top$ and collecting (ref) into a vector, we obtain
Thus $\Pi\odot\Upsilon\odot\Psi(\alpha)$ has a unit eigenvalue, with associated left eigenvector $y$.
The matrix $\Pi\odot\Upsilon\odot\Psi(\alpha)$ is nonnegative; suppose that it is also irreducible. The Perron-Frobenius theorem HornJohnson2013, a fundamental result in the theory of Markov chains, asserts that the spectral radius of a square, nonnegative and irreducible matrix is the maximum real eigenvalue of that matrix, and that there are unique (up to positive scalar multiplication) left and right eigenvectors with positive entries corresponding to that eigenvalue. Since $y$ has positive entries and is a left eigenvector of $\Pi\odot\Upsilon \odot \Psi(\alpha)$ associated with a unit eigenvalue, we are thus led to suspect that $\Pi\odot\Upsilon\odot\Psi(\alpha)$ may have spectral radius equal to one. This suspicion, if valid, suggests the possibility that we may be able to determine the Pareto exponent $\alpha$ by finding a positive number $z$ that solves the equation (ref). The fact that this is indeed often possible is the main result of our article, Theorem (ref).
The left eigenvector $y$ has a natural interpretation. It follows from (ref) that
Therefore, if the left eigenvector $y$ is normalized such that its entries sum to one, then we have $\mathrm{P}(J_t=n \mid S_t>s)=y_n$. We may thus interpret $y$ to be, loosely, the distribution of types in the upper tail of the size distribution. We formalize this claim in Theorem (ref).
The preceding discussion is largely heuristic. The conjectured equality (ref) cannot be expected to hold exactly except in very special cases. We will see that (ref) holds approximately for large $s$ in fairly general settings. The above derivations based on (ref) are therefore also approximate. We have skirted technical issues such as the existence of a unique stationary distribution for $(S_t,J_t)$, the finiteness of $\psi_{nn'}(\alpha)$, and the existence of a unique positive solution to (ref). Derivations similar to those above appear in Gouin-BonenfantTodaParetoExtrapolation. We now turn to a rigorous statement of our main results.
In this section we provide a rigorous statement of our primary result, which is that, under regularity conditions, a Markov multiplicative process with reset has a Pareto upper tail with exponent $\alpha$ given by the unique positive solution to equation (ref). We rely on the following regularity conditions.
Assumption (ref)(ref) means that, for any pair of states $(n,n')$, if $J_t$ is in state $n$ then there is a positive probability of it eventually reaching state $n'$ before reset occurs. Assumption (ref)(ref) means that there is some state $n$ such that, if $J_t$ is in state $n$, then reset occurs next period with positive probability. These two conditions together ensure that $J_t$ visits all states infinitely often and that reset occurs infinitely often. Assumption (ref)(ref), imposed without loss of generality, is merely a normalization since the growth rate when $J_t$ transitions from state $n$ to state $n'$ is unidentified if this transition never occurs.
To clarify the heuristic discussion in Section (ref), the first thing we will do is restrict the domain of the matrix-valued function $\Psi(z)$ such that each of its entries $\psi_{nn'}(z)$, defined in (ref), is finite-valued. Define the set
Since MGFs are always convex and are equal to one at zero, $\mathcal{I}$ is a convex set containing zero. We restrict the domain of $\Psi(z)$ to $\mathcal{I}$ and, for $z\in\mathcal{I}$, define the $N\times N$ matrix-valued function
Using (ref), the equation (ref) may now be rewritten more simply as $\rho(\mathrm{A}(z))=1$.
We claimed in Section (ref) that it is often the case that there is a unique positive value of $z$ that solves the equation $\rho(\mathrm{A}(z))=1$. The following result concerning the shape of the function $\rho(\mathrm{A}(z))$ helps to explain why this is true.
Figure (ref) depicts a typical shape for $\rho(\mathrm{A}(z))$ as a function of $z\in\mathcal{I}$. The graph of $\rho(\mathrm{A}(z))$ is convex, less than one at zero (note that $\mathrm{A}(0)=\Pi\odot\Upsilon$), and diverging to infinity at the left and right endpoints of $\mathcal{I}$ (which may in general be finite or infinite). There is a unique positive number $\alpha$ at which $\rho(\mathrm{A}(z))$ crosses one. We will see that $\alpha$ is the Pareto exponent for the upper tail of the stationary distribution of $S_t$. There is also a unique negative number $-\beta$ at which $\rho(\mathrm{A}(z))$ crosses one. We will see that $\beta$ is the Pareto exponent for the lower tail of the stationary distribution of $S_t$.
The graph of $\rho(\mathrm{A}(z))$ does not always have the typical shape depicted in Figure (ref). Three atypical cases in which no positive $z$ solves $\rho(\mathrm{A}(z))=1$ are depicted in Figure (ref). If all growth rates are nonpositive, so that $\mathrm{P}(G_{nn'}\leq1)=1$ for all $n,n'\in\mathcal{N}$, then $\rho(\mathrm{A}(z))$ is nonincreasing in $z$ and we have the case depicted in Figure (ref). If any growth rate does not have a light upper tail, so that the corresponding MGF $\psi_{nn'}(z)$ is infinite for all positive $z$, then the right endpoint of $\mathcal{I}$ is zero and we have the case depicted in Figure (ref). Problems may also arise if the tails of growth rates are insufficiently light. For instance, suppose for simplicity that we have no Markov modulation ($N=1$) and that $\log G$ has PDF
where $a>0$ and $c$ is a positive constant such that the PDF integrates to one. In this case $\operatorname{E}(G^z)$ is finite for $z\leq a$ and infinite for $z>a$, so that $\mathcal{I}=(-\infty,a]$. The existence of a positive solution to $\rho(\mathrm{A}(z))=1$ thus depends on whether $\rho(\mathrm{A}(a))$ is greater than one. The case where $a=1$ and $\upsilon=0.7$ is depicted in Figure (ref); we see that $\rho(\mathrm{A}(a))<1$, so that there is no positive solution to $\rho(\mathrm{A}(z))=1$.
The following result establishes that, to have a unique positive solution to $\rho(\mathrm{A}(z))=1$, it suffices that all growth rates have exponential moments of all orders, with at least one growth rate $\log G_{nn}$ positive with positive probability. An analogous result applies symmetrically to the existence of a unique negative solution to $\rho(\mathrm{A}(z))=1$.
We mentioned at the end of Section (ref) that, under suitable regularity conditions, a Markov multiplicative process with reset can be made stationary with a suitable choice of its distribution at time $t=0$. The following result indicates that Assumption (ref) provides sufficient regularity. It is proved by observing that reset generates a positive recurrent accessible atom of $(S_t,J_t)_{t\in\mathbb{Z}_+}$, which implies the existence of a unique stationary distribution.
We denote by $p$ the $N\times1$ vector of stationary probabilities for $J_t$, and by $q$ the $N\times1$ vector whose $n$th entry is $q_n=\sum_{n'}\pi_{nn'}(1-\upsilon_{nn'})$, the conditional probability of reset in period $t+1$ given $J_t=n$. The unconditional probability of reset is then $r\coloneqq\sum_np_nq_n$. It will also be useful to introduce notation for the subset of $\mathcal{I}$ on which $\mathrm{A}(z)$ has spectral radius less than one. We therefore set $\mathcal{I}_-=\left\{ z\in\mathcal{I}:\rho(\mathrm{A}(z))<1 \right\} $, which is a convex subset of $\mathcal{I}$ by Proposition (ref), nonempty and containing zero under Assumption (ref). In typical cases such as the one depicted in Figure (ref) we have $\mathcal{I}_-=(-\beta,\alpha)$.
The following result, Proposition (ref), is established in the course of proving our main result, Theorem (ref) below, but is also of independent interest. It reveals the form of the conditional MGF of $\log S_t$ given $J_t$, which has domain $\mathcal{I}_-$. The notation $\mathrm{I}$ refers to an $N\times N$ identity matrix.
An immediate consequence of Proposition (ref) is that the MGF of $\log S_t$ is given by $\operatorname{E}(S_t^z)=r\varpi^\top(\mathrm{I}-\mathrm{A}(z))^{-1}1_N$ for $z\in\mathcal{I}_-$, where $1_N$ is an $N\times1$ vector of ones. Proposition (ref) thus completely characterizes the stationary distribution of sizes whenever the interior of $\mathcal{I}_-$ contains zero.
We now state our main result.
Equation (ref) indicates that there is an approximately linear relationship in log-log scale between the tail probability $\mathrm{P}(S_t>s)$ and the threshold $s$ for large $s$, with slope $-\alpha$. In this sense, the upper tail of the distribution of $S_t$ is Pareto with exponent $\alpha$. The fact that this approximately linear relationship in log-log scale arises in distributions with a Pareto upper tail is the basis for the popular log-log rank-size regression method of estimating the Pareto exponent; see e.g.\xspace GabaixIbragimov2011. Note that Theorem (ref) does not assert the convergence of $s^\alpha\mathrm{P}(S_t>s)$ to a positive and finite limit, which would be a stronger notion of Pareto tail decay. This stronger property is not satisfied in general, but may be guaranteed by imposing a non-lattice condition on the distribution of growth rates, as discussed in Section (ref). To obtain the convergence (ref), it suffices that $s^\alpha\mathrm{P}(S_t>s)$ has positive and finite limits inferior and superior. To see why, note that the latter condition implies the existence of positive and finite constants $c_-$ and $c_+$ such that $c_-\leq s^\alpha\mathrm{P}(S_t>s)\leq c_+$ for all sufficiently large $s$. Taking logarithms, dividing by $\log s$, and letting $s\to\infty$, we obtain (ref).
We close this section with four remarks on Theorem (ref) and a brief example.
Examples 3.5 and 3.6 in BeareSeoToda2021 concern a continuous-time reformulation of Example (ref) in which the log-size of agents evolves as a two-state Brownian motion with drift. There, the determination of the Pareto exponent boils down to examining the roots of a quartic polynomial, or of a quadratic polynomial in the absence of a diffusive component. The latter case corresponds closely to the economic model in CaoLuo2017, in which the upper Pareto exponent for the stationary distribution of wealth is given by the unique positive root of a quadratic polynomial.
While Theorem (ref) establishes Pareto decay of the upper tail of the stationary distribution of $S_t$ in the sense of there being an asymptotically linear relationship in log-log scale between the tail probability $\mathrm{P}(S_t>s)$ and threshold $s$ with slope $-\alpha$, it is notable that Theorem (ref) does not assert the convergence of $s^\alpha\mathrm{P}(S_t>s)$ to a positive and finite limit. Such convergence does not hold in general. We show in this section that it obtains under a non-lattice condition on growth rates. We also establish a characterization of upper-tail type shares that obtains under the non-lattice condition.
A simple sufficient (but not necessary) condition for Assumption (ref) is that at least one growth rate is not a discrete random variable. The following result strengthens the conclusion of Theorem (ref) when Assumption (ref) is satisfied, and also shows that in this case a left eigenvector of $\mathrm{A}(\alpha)$ characterizes type shares in the upper tail of $S_t$.
We illustrate the statement about limiting type shares in Theorem (ref) by revisiting Example (ref). {
\addtocounter{exmp}{-1} }
The following simple example shows that the conclusions of Theorem (ref) need not be satisfied if Assumption (ref) is dropped.
Example (ref) illustrates the fragility of Theorem (ref). Since any collection of growth rate distributions satisfying Assumption (ref) can be well-approximated by a collection of growth rate distributions not satisfying it, and vice-versa, it is natural to be skeptical of results which depend on this condition for their validity. Non-lattice conditions have been used elsewhere to establish Pareto tail behavior of stationary solutions to stochastic difference equations.\footnote{See, for instance, (1.11) in Theorem A of kesten1973, (2) in Theorems 1 and 2 of desaporta2005, (A7) in Assumption 1.2 of roitershtein2007, and the “spread out” condition on pp.\ 1408--9 of collamore2009; and in an economic application, see Footnote 56 in benhabib-bisin-zhu2011. Note also that results of this kind typically exclude the possibility of reset: see (1.9) in Theorem A of kesten1973, the requirement that zero be excluded from the state space in Theorems 1 and 2 of desaporta2005, (A5) in Assumption 1.2 of roitershtein2007, and the requirement that $\log A_n$ be finite-valued on p.\ 1408 of collamore2009.} An advantage of our approach is that our primary result, Theorem (ref), does not rely on any non-lattice condition for its validity. The form of Pareto tail decay given in (ref) is in this sense robust.
Given the importance of persistent heterogeneity in productivity types for constructing plausibly calibrated economic models, we expect to see modellers adopting this feature more widely in future, and hope that the results we have presented here may prove useful to them. We conclude by outlining three potential generalizations of our results which may broaden the scope of applications.
First, it may be useful to pursue a relaxation of our assumption of irreducibility. The evolution of types may follow a reducible Markov chain if, for instance, the health of agents randomly and irreversibly deteriorates over time, affecting their access to different productivity types. We rely on irreducibility at various points in our technical arguments. In particular, in the proof of Theorem (ref) in Appendix (ref), which is used to prove Theorems (ref) and (ref), irreducibility is critical to establishing that poles in the MGF of log-size determining the Pareto exponents are simple. The reducible case may require a more general treatment of these poles.
Second, it may be useful to pursue a relaxation of our assumption that the number of types is finite. This would allow, for instance, productivity to be modelled as a general autoregressive process. Such a generalization would likely entail adapting our numerous arguments involving matrices such that they apply with linear operators on infinite dimensional spaces.
Third, it may be useful to generalize our results such that they apply when the random growth of agents is only asymptotically, rather than exactly, multiplicative. For instance, while capital income is naturally associated with multiplicative growth in wealth, the effect of labor income on wealth is additive. In an economy in which agents randomly accrue both capital and labor income, the growth in the wealth of the wealthiest agents is predominantly driven by capital income, and thus approximately multiplicative. Since it is the wealthiest agents which determine the upper Pareto exponent, we expect that our results on the upper Pareto exponent may be applied in a setting of this sort. This has been done in Gouin-BonenfantTodaParetoExtrapolation, with heuristic justification. A rigorous demonstration of the validity of our characterization of the upper Pareto exponent under asymptotically multiplicative growth would place such applications on firmer footing.