EconBase
← Back to paper

Nonparametric and Semiparametric Estimation of Upward Rank Mobility Curves

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

50,412 characters · 11 sections · 47 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric and Semiparametric Estimation of Upward Rank Mobility Curves

abstractWe introduce the upward rank mobility curve as a new measure of intergenerational mobility that captures upward movements across the entire parental income distribution. Our approach extends Bhattacharya2011 by conditioning on a single parental income rank, thereby eliminating aggregation bias. We show that the measure can be characterized solely by the copula of parent and child income, and we propose a nonparametric copula-based estimator with better properties than kernel-based alternatives. For a conditional version of the measure without such a representation, we develop a two-step semiparametric estimator based on distribution regression and establish its asymptotic properties. An application to U.S.\ data reveals that whites exhibit significant upward mobility dominance over blacks among lower-middle-income families. {\bf Keywords}: Intergenerational mobility, upward mobility, copula, distribution regression, mobility dominance. {\bf JEL Codes}: C14, C21, D31, J62.

Introduction

Economists have long been interested in the intergenerational transmission of income, and how to document income mobility across generations has been an active area of research; see Solon1999,Black2011,Corak2013 for comprehensive reviews and Deutscher2023 for a recent survey. Among existing measures, perhaps the most widely used is the intergenerational income elasticity Solon1992, defined as the slope from a linear regression of log child income on log parent income. Yet the linearity assumption underlying the IGE may be restrictive, motivating alternative measures such as the rank-rank slope Chetty2014, which captures the correlation between children's and parents' income ranks.

While the IGE and rank-rank slope provide useful summaries of persistence, they only indicate the extent of intergenerational mobility rather than its direction. To study upward mobility, researchers have employed transition probabilities across income classes Formby2004 or the upward rank mobility measure of Bhattacharya2011, defined as the probability that a child's relative income rank exceeds that of the parents within a given class. These directional measures have been used extensively Jappelli2006,Dardanoni2012,Hnatkovska2013,Corak2014,Bratberg2017,Richey2018,Millimet2020,Collins2022,Bavaro2023, but they face two important limitations. First, discretizing the income distribution into classes introduces aggregation bias due to heterogeneity within groups. Second, there is no consensus on how many classes to use, or whether an optimal choice even exists.\footnote{A similar problem arises in estimating the Gini coefficient from grouped data; see, e.g., Lerman1989,Davies1989.}

This paper introduces the upward rank mobility curve, a new measure that addresses these limitations and facilitates comparisons of upward mobility across the entire parental income distribution. Our measure extends Bhattacharya2011 by conditioning on a single parental income rank rather than an interval, thereby eliminating aggregation bias. We show that the measure can be expressed solely in terms of the copula of parent and child income---a novel characterization in the literature---and propose a nonparametric estimator based on the empirical Bernstein copula Sancetta2004. Compared with the Nadaraya-Watson kernel estimator, our method exhibits improved asymptotic properties. We also provide practical guidance on choosing the smoothing parameter to minimize the asymptotic mean squared error.

Nevertheless, when analyzing differences in upward rank mobility across demographic groups or geographic regions, valid comparisons must rely on ranks derived from the overall income distribution rather than from group-specific ones. In this conditional context, the copula-based representation does not apply. We therefore develop a two-step semiparametric estimator based on distribution regression Foresi1995,Chernozhukov2013, establish its asymptotic properties, and examine its finite-sample performance through simulations. Applying these methods to data from the National Longitudinal Survey of Youth, we investigate interracial differences in upward mobility in the United States. The results show significant upward mobility dominance of whites over blacks among lower-middle-income families, corroborating findings from previous studies Bhattacharya2011,Mazumder2014,Fox2016,Chetty2019.

A large body of research emphasizes that intergenerational mobility can be measured in multiple ways, with distinct metrics capturing different aspects of the transmission process. As highlighted by Deutscher2023, an important distinction is between global and local measures. Global measures, such as the IGE and the rank-rank slope, summarize the joint distribution of parent and child income with a single parameter. Local measures, by contrast, focus on specific parts of the distribution and reveal heterogeneous patterns. Transition probabilities, conditional expected ranks, and directional rank mobility are particularly useful for detecting barriers to upward mobility at the bottom (poverty traps) or persistence at the top (entrenched privilege); see, e.g., Corak2020. Such measures have proven central to understanding nonlinearities in mobility and their links to mechanisms such as credit constraints or neighborhood effects Grawe2004,Durlauf2004.

Another key distinction concerns whether mobility is measured in relative or absolute terms. Relative measures, including the IGE, the rank-rank slope, and the proposed unconditional upward rank mobility curve, compare children's income levels or ranks directly with those of their parents. In contrast, absolute measures assess mobility against an external benchmark, which may be defined by fixed standards of living Chetty2017, or by income ranks derived from the overall population for cross-group comparisons. Our proposed conditional upward rank mobility curve falls into this category and is particularly well suited for the latter application.

Because relative and absolute mobility need not coincide, it is important to consider both perspectives. Indeed, Deutscher2023 show that correlations between them are generally weak, implying that no single metric provides a complete picture of intergenerational mobility. Against this backdrop, the contribution of this paper is to introduce both unconditional and conditional mobility curves, thereby extending the set of directional local measures to encompass both relative and absolute dimensions.

Our work is related to Callaway2024, who study the identification and estimation of intergenerational mobility parameters under two-sided measurement error. They show that a broad class of parameters (including transition matrices, upward mobility measures of Bhattacharya2011, and rank-rank slopes) can be characterized by the copula of parent and child income as well, and they propose a multi-step semiparametric estimator based on quantile regression. By contrast, our contribution is twofold. First, we introduce upward rank mobility curves that eliminate measurement error arising from misclassification into incorrect transition cells (see Millimet2020 for a related argument). Second, we exploit the copula representation to develop a nonparametric estimator in the unconditional case. Thus, while their work addresses robustness to measurement error, ours expands the methodological toolkit for analyzing upward rank mobility.

Our work is also related to Chernozhukov2025, who introduce bivariate distribution regression as a flexible framework for modeling conditional joint distributions and decomposing intergenerational mobility. In this conditional setting, however, our focus is on diagnosing mobility dominance between groups (e.g., racial disparities) that cannot be inferred from conditional joint distributions alone. As noted earlier, such comparisons require unconditional income ranks of parents and children, which are typically unobserved and must be estimated. We therefore establish the weak convergence of the conditional mobility curve estimator accounting for these estimation effects, and use it to conduct uniform inference across the entire parental income distribution.

The rest of the paper is organized as follows. (ref) formally introduces the upward rank mobility curve and the nonparametric copula-based estimator. (ref) extends the analysis to the conditional measure and develops a semiparametric distribution regression procedure, along with its asymptotic properties. (ref) and (ref) present the simulation and empirical studies, respectively, and (ref) concludes.

Upward Rank Mobility Curves

Definition

Let $Y_0$ and $Y_1$ denote parental and child income, with marginal cumulative distribution functions (CDFs) $F_0$ and $F_1$, respectively. For families whose (normalized) parental income rank lies within an interval, $F_0(Y_0)\in[s_1,s_2]$, Bhattacharya2011 propose a measure of upward rank mobility defined as the probability that the child's income rank, $F_1(Y_1)$, exceeds that of the parent by at least $\tau\in[0,1-s_2]$ for $0<s_1<s_2<1$:

equation[equation omitted — 122 chars of source]

Although widely used, this measure is susceptible to aggregation bias arising from heterogeneity among families within $[s_1,s_2]$. In particular, it is unclear whether observed upward mobility originates mainly from families at the lower or upper end of the interval. This bias is further compounded by the common practice of dividing income into only a few categories (i.e., using relatively coarse intervals).

In this paper, we introduce the upward rank mobility curve, which conditions on an exact parental income rank rather than on an interval:

equation[equation omitted — 96 chars of source]

where $0\leq\tau<1-s$ and $0<s<1$. Clearly, $u(\tau,s)=\upsilon(\tau,s,s+t)$ as $t\to 0$. This refinement eliminates aggregation bias and facilitates comparisons of intergenerational upward mobility in terms of stochastic dominance Fields2002 and stochastic monotonicity Lee2009 across the parental income distribution. For example, a formal definition of upward mobility dominance is provided in (ref).

As pointed out by Bhattacharya2011, conditioning on $F_0(Y_0)=s$ typically requires averaging over a bandwidth around $s$ when kernel smoothing is applied. However, selecting an optimal bandwidth that balances bias and variance is challenging as it depends on the unknown marginals $F_0$ and $F_1$. Moreover, the Nadaraya-Watson estimator is well known to suffer from boundary effects, which lead to larger bias when estimating $u(\tau,s)$ near $s=0$ or $s=1$. To overcome these limitations, we propose a novel nonparametric method that avoids estimating the marginal distributions. We also derive an asymptotic expression for the mean squared error, valid uniformly over interior and boundary points, which in turn yields the optimal smoothing parameter for our estimator.

Empirical Bernstein Copula-Based Estimators

We first show that the mobility measure defined in (ref) can be expressed solely in terms of the copula of $(Y_0,Y_1)$, complementing the findings of Callaway2024 that connect relative intergenerational mobility to the copula literature. This result also parallels Chetty2017, who link absolute intergenerational mobility to the copula through decomposition.

proLet $\partial_0C\equiv\partial C(u_0,u_1)/\partial u_0$ denote the first partial derivative of the copula $C$ of $(Y_0,Y_1)$. If $F_0$ and $F_1$ are continuous and strictly increasing, we have \[ u(\tau,s)=1-\partial_0C(s,s+\tau) \] for almost all $\tau\in[0,1-s)$ and $s\in(0,1)$ with respect to Lebesgue measure.

From (ref) (see (ref) for the proof), the problem now reduces to estimating the copula derivative. Our nonparametric estimator builds on the work of Janssen2016, who employ Bernstein estimation for the first-order derivative of a copula. Specifically, let $\{(Y_{i0},Y_{i1}):i=1,\dotsc,n\}$ be a random sample of $(Y_0,Y_1)$. Denote $R_{ij}=\sum_{k=1}^n1\{Y_{kj}\leq Y_{ij}\}$ as the rank of $Y_{ij}$ among $Y_{1j},\dotsc,Y_{nj}$ for $i=1,\dotsc,n$ and $j=0,1$, with $1\{\cdot\}$ being the indicator function. The Bernstein estimator of Janssen2016 is essentially the partial derivative of the empirical Bernstein copula, $C_{m,n}$, introduced by Sancetta2004:

align[align omitted — 675 chars of source]

where $m\in\mathbb{N}$ is the order of the Bernstein polynomial satisfying $m\to\infty$ and $m/n\to0$ as $n\to\infty$, $C_n(u_0,u_1)=n^{-1}\sum_{i=1}^n1\mathopen{}\mathclose\bgroup\originalleft\{R_{i0}/n\leq u_0,R_{i1}/n\leq u_1\aftergroup\egroup\originalright\}$ is the rank-based empirical copula, and $P_{m,k}(u)={m \choose k}u^k(1-u)^{m-k}$ is the binomial probability for $k=0,1,\dotsc,m$. Note that the definition of $C_n$ here differs slightly from that in Janssen2016, but the discrepancy between these empirical copula variants is at most $2/n$; see Fermanian2004.

According to (ref), the empirical Bernstein copula-based (EBC) estimator of $u(\tau,s)$ is defined as

equation[equation omitted — 117 chars of source]

The remaining task is to choose the order $m$. Assuming $C$ has third-order partial derivatives that are Lipschitz continuous in $(0,1)^2$, the results of Swanepoel2013 yield the asymptotic mean squared error of $\widehat{u}^{\text{EBC}}_{m,n}(\tau,s)$, uniformly in $\tau$ and $s$:

align*[align* omitted — 140 chars of source]

where

align*[align* omitted — 306 chars of source]

Therefore, the optimal order that minimizes the asymptotic mean squared error is

equation[equation omitted — 264 chars of source]

where $\lceil x\rceil$ denotes the smallest integer not smaller than $x$.\footnote{When $b(s,s+\tau)=0$, we simply set $m^*=2$.} In practice, however, $m^*$ is difficult to implement because it depends on third-order partial derivatives of the copula, which are typically unknown. Moreover, since $m^*$ varies with $(\tau,s)$, it is impractical to use different orders for each point on the curve. In simulations reported in (ref), we find that setting $m=n^{1/2}$ globally performs as well as, or even better than, the pointwise optimal $m^*$ in finite samples.

rmkAs summarized in Janssen2016, the Bernstein estimator in (ref) for $(u_0,u_1)\in(0,1)^2$ exhibits superior asymptotic properties compared with the Nadaraya-Watson estimator. For comparison, we set $h=m^{-1}$ as the bandwidth following Sancetta2004. The asymptotic variance of the Bernstein estimator is of order $O(m^{1/2}/n)$, whereas that of the Nadaraya-Watson estimator is $O(m/n)$. With respect to bias, the Bernstein estimator achieves an order of $O(m^{-1})$ at both interior and boundary points (and is therefore free from boundary effects), while the Nadaraya-Watson estimator attains $O(m^{-2})$ at interior points and only $O(m^{-1})$ at boundary points. Consequently, the optimal mean squared error of the Bernstein estimator is $O(n^{-4/5})$ uniformly at interior and boundary points, whereas for the Nadaraya-Watson estimator it is $O(n^{-4/5})$ at interior points but deteriorates to $O(n^{-2/3})$ at boundary points.
rmkInstead of minimizing the asymptotic mean squared error, one may alternatively minimize the asymptotic bias by setting $m=n$, i.e., the largest possible order, noting that the asymptotic variance of the Bernstein estimator is of smaller magnitude. Interestingly, such undersmoothing renders the estimator in (ref) identical to the partial derivative of the empirical beta copula, $C_n^\beta$, developed by Segers2017: \begin{align*} \widehat{\partial_0C}_{n,n}(u_0,u_1) &=\frac{\partial}{\partial u_0}C^\beta_n(u_0,u_1)\\ &=\frac{\partial}{\partial u_0}\frac{1}{n}\sum_{i=1}^nF_{n,R_{i0}}(u_0)F_{n,R_{i1}}(u_1)\\ &=\frac{1}{n}\sum_{i=1}^nf_{n,R_{i0}}(u_0)F_{n,R_{i1}}(u_1), \end{align*} where $F_{n,r}(u)=\sum_{s=r}^n{n\choose s}u^s(1-u)^{n-s}$ is the CDF of the beta distribution $\mathcal{B}(r,n+1-r)$ for $r=1,\dotsc,n$, and $f_{n,r}$ is the corresponding probability density function. A key advantage of this estimator is that it eliminates the need to select a smoothing parameter. As shown in (ref), the empirical beta copula-based estimator of $u(\tau,s)$, \begin{equation} \widehat{u}^\beta_n(\tau,s)\equiv1-\widehat{\partial_0C}_{n,n}(s,s+\tau), \end{equation} substantially outperforms competing estimators in terms of bias.

Conditional Upward Rank Mobility Curves

While $u(\tau,s)$ effectively measures upward rank mobility across the entire parental income distribution, examining its conditional counterpart given covariates $X$ can provide further insight into mobility differences across demographic groups or geographic regions. This section extends our analysis to such a conditional measure. However, unlike in (ref), no analogous copula representation is available when ranks are defined with respect to the overall population rather than the group under study, rendering the method in (ref) inapplicable. We therefore propose an alternative two-step semiparametric estimator based on distribution regression and establish its uniform asymptotic properties, which are useful for testing dominance in upward rank mobility.

Definition

Our definition of the conditional upward rank mobility curve is a direct modification of Bhattacharya2011:

equation[equation omitted — 104 chars of source]

where $0\leq\tau<1-s$ and $0<s<1$. Importantly, $F_0(Y_0)$ and $F_1(Y_1)$ here still denote the unconditional income ranks of parents and children. As noted by Deutscher2023, such a conditional measure behaves like an absolute measure, since the ranks are defined relative to a fixed external benchmark. This feature enables meaningful between-group comparisons, rather than restricting attention to mobility within the group defined by $X=x$. Accordingly, following Aaberge2014, we define upward mobility dominance between two groups as follows:

defnThe $x_1$-group is said to exhibit upward mobility dominance over the $x_2$-group on $\mathcal{S}\subseteq(0,1)$ if \[ u_c(x_1;\tau,s)\geq u_c(x_2;\tau,s)\quad\text{for all $s\in\mathcal{S}$} \] and the inequality holds strictly for some $s\in\mathcal{S}$.

Unlike the representation in (ref), (ref) does not admit a straightforward copula-based characterization. A more natural conditional analogue is instead given by

align[align omitted — 176 chars of source]

where $F_{0|X}$ and $F_{1|X}$ denote the conditional income distributions of parents and children given $X$, and $C_x$ is the conditional copula satisfying $\operatorname{P}(Y_0\leq y_0,Y_1\leq y_1|X=x)=C_x(F_{0|X}(y_0|x),F_{1|X}(y_1|x))$. The distinction between (ref) and (ref) is clear in applications. If $X$ denotes race, (ref) captures interracial differences in upward mobility relative to the overall population, while (ref) reflects intraracial differences within the income distribution of a particular racial group. Since our interest lies in between-group comparisons, we focus on (ref) in the remainder of the paper.

Distribution Regression-Based Estimators

To estimate (ref), note that if $F_0$ and $F_1$ are continuous and strictly increasing, the conditional measure can be rewritten as

equation[equation omitted — 82 chars of source]

where $F_{1|0,X}$ is the conditional CDF of $Y_1$ given $Y_0$ and $X$, and $Q_j=F_j^{-1}$ is the quantile function of $Y_j$ for $j=0,1$. Estimation therefore proceeds in two steps: (i) construct an estimator of $F_{1|0,X}(y_1|y_0,x)$; and (ii) evaluate $y_1$ and $y_0$ at the sample counterparts of $Q_1(s+\tau)$ and $Q_0(s)$, respectively.

Our first step follows the distribution regression developed by Foresi1995,Chernozhukov2013:

equation[equation omitted — 105 chars of source]

where $\Lambda$ is a link function, $P(Y_0,X)$ is a vector of polynomials of $Y_0$ and $X$, and $\widehat\theta(y_1)$ is the parameter vector indexed by $y_1$ that maximizes the likelihood: \[ \sum_{i=1}^n(1\{Y_{i1}\leq y_1\}\ln\Lambda(P(Y_{i0},X_i)^\top\theta(y_1))+1\{Y_{i1}>y_1\}\ln(1-\Lambda(P(Y_{i0},X_i)^\top\theta(y_1)))). \] This semiparametric approach provides a flexible way to model the conditional distribution of children's income given parental income and observed covariates. Rather than assuming a fully parametric model, it uses a series of binary regressions at different thresholds, which allows the conditional distribution to be approximated arbitrarily well; see Chernozhukov2013 for details. This approach enables us to evaluate upward rank mobility at specific parental ranks while accommodating rich forms of heterogeneity across groups.

Combining (ref), the distribution regression-based (DR) estimator of $u_c(x;\tau,s)$ is defined as

equation[equation omitted — 131 chars of source]

where $\widehat{Q}_j(p)=\inf\{y\in\mathbb{R}:\widehat F_j(y)\geq p\}$ is the empirical quantile function of $Y_j$ for $j=0,1$, with $\widehat F_j(y)=n^{-1}\sum_{i=1}^n1\{Y_{ij}\leq y\}$ the corresponding empirical CDF. For comparison with the nonparametric EBC estimator in (ref), we also define the DR estimator of the unconditional mobility curve $u(\tau,s)$ as

equation[equation omitted — 123 chars of source]

where $\widehat F_{1|0}$ is constructed analogously to (ref).

Asymptotic Properties

We now establish the weak convergence of the DR estimator defined in (ref). Similar results hold for the conditional estimator in (ref), but are omitted for brevity. In both cases, however, it is important to account for the estimation effects introduced by the empirical quantile functions $\widehat Q_0$ and $\widehat Q_1$ when deriving the limiting processes. The regularity conditions below are adapted from Chernozhukov2013.

asmDenote $\mathcal{Y}_0\subseteq\mathbb{R}$ and $\mathcal{Y}_1\subseteq\mathbb{R}$ as the supports of $Y_0$ and $Y_1$. Suppose that \begin{enumerate}[(i)] • The support $\mathcal{Y}_0\times\mathcal{Y}_1$ is a compact subset of $\mathbb{R}^2$. • The distribution function $F_j$ is uniformly continuous on $\mathcal{Y}_j$ for $j=0,1$. • The density function $f_j$ is bounded away from zero uniformly on $\mathcal{Y}_j$ for $j=0,1$. • The conditional density function $f_{1|0}$ is uniformly continuous and bounded on $\mathcal{Y}_1\times\mathcal{Y}_0$. • $\operatorname{E}\|P(Y_0)\|^2<\infty$ and the minimum eigenvalue of \[ H(y_1)\equiv\operatorname{E}\mathopen{}\mathclose\bgroup\originalleft(\frac{\lambda^2(P(Y_0)^\top\theta(y_1))}{\Lambda(P(Y_0)^\top\theta(y_1))(1-\Lambda(P(Y_0)^\top\theta(y_1)))}P(Y_0)P(Y_0)^\top\aftergroup\egroup\originalright) \] is bounded away from zero uniformly on $\mathcal{Y}_1$, where $\lambda$ is the derivative of $\Lambda$. \end{enumerate}
thmSuppose $F_{1|0}(y_1|y_0)=\Lambda(P(y_0)^\top\theta(y_1))$ is correctly specified for all $y_1\in\mathcal{Y}_1$ and $y_0\in\mathcal{Y}_0$, and (ref) is satisfied. Then, \[ \sqrt{n}\mathopen{}\mathclose\bgroup\originalleft(\widehat u^\text{\upshape DR}(\tau,s)-u(\tau,s)\aftergroup\egroup\originalright)\Rightarrow\Psi(\tau,s), \] where $\Rightarrow$ denotes weak convergence and $\Psi(\tau,s)$ is a zero-mean Gaussian process with covariance function generated by the influence function \begin{align*} \psi(\tau,s,Y_1,Y_0)&=-\mathopen\mathclose\bgroup\originalleft\{ \lambda(P(Q_0(s))^\top\theta(Q_1(s+\tau)))P(Q_0(s))^\top\psi_\theta(Q_1(s+\tau),Y_1,Y_0)\aftergroup\egroup\originalright.\\ &\quad+\lambda(P(Q_0(s))^\top\theta(Q_1(s+\tau)))P(Q_0(s))^\top\theta'(Q_1(s+\tau))\psi_1(\tau,s,Y_1)\\ &\mathopen\mathclose\bgroup\originalleft.\quad+\lambda(P(Q_0(s))^\top\theta(Q_1(s+\tau)))P'(Q_0(s))^\top\theta(Q_1(s+\tau))\psi_0(s,Y_0)\aftergroup\egroup\originalright\}, \end{align*} where $\theta'(\cdot)$ and $P'(\cdot)$ are $(p+1)$-dimensional vectors consisting of elementwise derivatives of $\theta(\cdot)$ and $P(\cdot)$, respectively, and \begin{align*} \psi_\theta(y_1,Y_1,Y_0)&=H^{-1}(y_1)\frac{1\{Y_1\leq y_1\}-\Lambda(P(Y_0)^\top\theta(y_1))}{\Lambda(P(Y_0)^\top\theta(y_1))(1-\Lambda(P(Y_0)^\top\theta(y_1)))}\lambda(P(Y_0)^\top\theta(y_1))P(Y_0),\\ \psi_1(\tau,s,Y_1)&=\frac{(s+\tau)-1\{Y_1\leq Q_1(s+\tau)\}}{f_1(Q_1(s+\tau))},\\ \psi_0(s,Y_0)&=\frac{s-1\{Y_0\leq Q_0(s)\}}{f_0(Q_0(s))}. \end{align*}

(ref) shows that the estimator in (ref) converges weakly to a zero-mean Gaussian process at the parametric rate. However, since the influence function in (ref) depends on unknown nuisance functions and is therefore non-pivotal, we employ the empirical bootstrap, following Chernozhukov2013, to conduct uniform inference for the entire curve. For example, the bootstrap $(1-\alpha)$ uniform confidence band for $u(\tau,s)$ with fixed $\tau$ is given by \[ \mathopen{}\mathclose\bgroup\originalleft[\widehat{u}^{\text{DR}}(\tau,s)-c_{1-\alpha}^B\widehat\sigma^B(\tau,s),~\widehat{u}^\text{DR}(\tau,s)+c_{1-\alpha}^B\widehat\sigma^B(\tau,s)\aftergroup\egroup\originalright], \] where $c_{1-\alpha}^B$ is the $(1-\alpha)$ quantile of $\{\sup_{0<s<1}|(\widehat u_b^\text{DR}(\tau,s)-\widehat u^\text{DR}(\tau,s))/\widehat{\sigma}^B(\tau,s)|\}_{b=1}^B$ based on $B$ bootstrap samples. Here, $\widehat u_b^\text{DR}(\tau,s)$ denotes the bootstrap version of $\widehat u^\text{DR}(\tau,s)$ for $b=1,\dotsc,B$, and $\widehat{\sigma}^B(\tau,s)$ is the pointwise bootstrap standard deviation of $\widehat u^\text{DR}(\tau,s)$. An analogous procedure can be applied to test upward mobility dominance in (ref), through the details are omitted for brevity.

Simulation Study

We examine the finite-sample performance of the EBC estimator in (ref) and the DR estimator in (ref). The data $\{(Y_{i0},Y_{i1})\}_{i=1}^n$ are generated from the following four copulas with standard Gaussian marginal distributions:

enumerate• Gaussian copula: $C_\theta(u_0,u_1) = \Phi_\theta(\Phi^{-1}(u_0),\Phi^{-1}(u_1))$ with $\theta \in (-1,1)$, where $\Phi_\theta$ is the bivariate standard Gaussian CDF with correlation coefficient $\theta$, and $\Phi$ is the univariate standard Gaussian CDF. • Clayton copula: $C_\theta(u_0,u_1) = (u_0^{-\theta}+u_1^{-\theta}-1)^{-1/\theta}$ with $\theta \in (0,\infty)$. • Gumbel copula: $C_\theta(u_0,u_1) = \exp(-((-\ln u_0)^\theta+(-\ln u_1)^\theta)^{1/\theta})$ with $\theta \in [1,\infty)$. • Independence copula: $C(u_0,u_1) = u_0 u_1$.

For the Gaussian, Clayton, and Gumbel copulas, the parameter $\theta$ is calibrated so that Kendall's tau satisfies $\tau_K\in\{1/3,1/2,2/3\}$, corresponding to weak, moderate, and strong dependence between $Y_0$ and $Y_1$, respectively. Under the independence copula, we have $\tau_K=0$.

For the EBC estimator, we consider three specifications for the order of the Bernstein polynomial $m$: (i) $\widehat u^\text{EBC}_{m^*,n}$, where $m^*$ is the optimal order given in (ref); (ii) $\widehat u^\text{EBC}_{n^{1/2},n}$, which uses the intermediate order $n^{1/2}$; (iii) $\widehat u^\text{EBC}_{n,n} = \widehat u^\beta$, which sets $m=n$ and yields the empirical beta copula in (ref). The optimal order $m^*=m^*(\tau,s)$ is computed pointwise under the true copula. Although infeasible in practice, it is included here for benchmarking. For the DR estimator, we examine three combinations of the link function $\Lambda$ and the polynomial $P(Y_0)$: (i) $\widehat u^\text{DR}_\text{Probit}$, using the probit link with $P(Y_0)=Y_0$; (ii) $\widehat u^\text{DR}_\text{Logit}$, using the logit link with $P(Y_0)=Y_0$; (iii) $\widehat u^\text{DR}_\text{Logit,2}$, using the logit link with $P(Y_0)=[Y_0~Y_0^2]^\top$. While all these specifications are initially misspecified, we will show that their performance improves substantially as the polynomial order increases.

table[table omitted — 2,765 chars of source]
table[table omitted — 2,779 chars of source]
table[table omitted — 2,770 chars of source]
table[table omitted — 1,347 chars of source]

(ref) report the simulation results for the root integrated squared bias (RISB) and root integrated mean squared error (RIMSE), defined as

align*[align* omitted — 176 chars of source]

where $\widehat u(\tau,s)$ denotes a generic estimator. For numerical integration, we set $\tau=0$ and evaluate $s$ on the grid $\{0.01,0.02,\dotsc,0.99\}$. Each design is replicated 1,000 times for sample sizes $n\in\{100,200,400\}$.

From (ref), it is unsurprising that the empirical beta copula-based estimator $\widehat u^\beta$ achieves the best RISB performance in most scenarios, as it exploits the maximal order $n$. Under the Gaussian copula, however, the DR estimator with a probit link, $\widehat u^\text{DR}_\text{Probit}$, can outperform $\widehat u^\beta$ particularly in large samples. By contrast, $\widehat u^\beta$ performs poorly in terms of RIMSE due to undersmoothing. The intermediate-order estimator $\widehat u^\text{EBC}_{n^{1/2},n}$ instead delivers consistently good performance, aside from the infeasible benchmark $\widehat u^\text{EBC}_{m^*,n}$. The DR estimators perform similarly to $\widehat u^\text{EBC}_{n^{1/2},n}$, with their RIMSE roughly halving as the sample size quadruples, suggesting convergence close to the parametric $\sqrt{n}$ rate despite misspecification of the conditional distribution. From (ref), we further find that the DR estimators dominate both $\widehat u^\beta$ and $\widehat u^\text{EBC}_{n^{1/2},n}$ under the independence copula. Overall, these results indicate that DR estimators perform well when the link function is appropriately specified and the sample size is sufficiently large.

Interestingly, the RIMSE of $\widehat u^\text{EBC}_{n^{1/2},n}$ can be smaller than that of $\widehat u^\text{EBC}_{m^,n}$ under both the Clayton and Gumbel copulas. To investigate this, we plot the estimated curves from simulation runs under the Gaussian, Clayton, and Gumbel copulas with sample size $n=200$ and Kendall's tau $\tau_K=1/2$. Under the Gaussian copula ((ref)), $\widehat u^\text{EBC}_{m^*,n}$ outperforms $\widehat u^\text{EBC}_{n^{1/2},n}$ because the bias function $b(s,s)$ defined in (ref) is nearly zero around $s\approx 0.5$, with $b(0.5,0.5)=0$ exactly. Consequently, the optimal order $m^*(0,s)$ is small near $s\approx 0.5$, yielding very low variance and superior RIMSE of $\widehat u^\text{EBC}_{m^*,n}$. In contrast, $\widehat u^\text{EBC}_{n^{1/2},n}$ outperforms $\widehat u^\text{EBC}_{m^*,n}$ under the Clayton copula ((ref)) and the Gumbel copula ((ref)), as $\widehat u^\text{EBC}_{m^*,n}$ suffers from larger bias near $s\approx 0$ in the former and near $s\approx 1$ in the latter.

Taken together, these findings underscore that the optimal order $m^*$ depends on the shape of the curve and is derived from asymptotic mean squared error, which does not necessarily guarantee finite-sample optimality. Our simulations suggest that setting $m=n^{1/2}$ offers a practical and robust choice for the EBC estimator.

figure[figure omitted — 426 chars of source]
figure[figure omitted — 426 chars of source]
figure[figure omitted — 424 chars of source]

Empirical Study

In this section, we examine intergenerational upward rank mobility between whites and blacks in the United States. Following Bhattacharya2011, we use data from the National Longitudinal Survey of Youth 1979 (NLSY79), which tracks individuals who were 14--22 years old in 1979 through adulthood until 2018. The survey provides information on annual income from the previous year. To avoid complications arising from labor force participation, we restrict the sample to 2,002 white sons and 276 black sons.

The parental permanent income is measured using total family income reported by the sons while living with their parents from 1979 to 1981, corresponding to income years 1978--1980. We subtract any recorded earnings of the sons (wages, salaries, farm or business income) and average the remaining values across all available years. The sons' permanent income is measured as the average of their annual earnings in 1998, 2000, 2002, and 2004, when they were approximately 40 years old. All income figures are adjusted to 1978 dollars using the CPI-U prior to averaging.\footnote{These settings are aligned with those in Bhattacharya2011.}

figure[figure omitted — 121 chars of source]

(ref) presents the estimated upward rank mobility curves (URMCs) obtained from the EBC and DR methods introduced in (ref). The blue solid line corresponds to the EBC estimator with $m=n^{1/2}$, as recommended by the simulation results in (ref), while the red dashed line corresponds to the DR estimator with a logit link and quadratic polynomial specification. For reference, we also plot the sample analogues of upward rank mobility (URM) defined in (ref), constructed with different numbers of income categories (green and yellow step functions). The figure shows that both URMC estimates are reliable, as they closely match the empirical URMs for parental income ranks in $[0,0.5]$. More importantly, the URMCs reveal a monotonic decline in upward mobility as parental income rank increases---a pattern that the discrete URMs are unable to capture.

figure[figure omitted — 125 chars of source]
figure[figure omitted — 174 chars of source]

We next examine differences in upward rank mobility between racial groups. Since the copula-based estimator is not applicable in this context, we employ the conditional URMC estimator from (ref) to construct race-specific mobility curves, as shown in (ref). The figure indicates that whites consistently exhibit higher upward rank mobility than blacks. Further evidence is provided in (ref), which plots the white-black difference along with 95% pointwise and uniform confidence bands over parental income ranks $[0,0.5]$. As the uniform band lies strictly above zero on the interval $[0.2,0.4]$, we conclude that whites statistically dominate blacks in upward rank mobility among lower-middle-income families in our sample.

Conclusion

This paper proposes the upward rank mobility curve as a tool for assessing intergenerational upward mobility across the entire parental income distribution. We show that the measure can be expressed solely in terms of the copula of parent and child income, and we develop nonparametric and semiparametric estimators for the measure and its conditional variant. An empirical application using the data of Bhattacharya2011 demonstrates systematically higher upward rank mobility for whites than for blacks in the United States. We further provide evidence of significant upward mobility dominance of whites over blacks among lower-middle-income families.