EconBase
← Back to paper

Conditional Rank-Rank Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

102,733 characters · 21 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Conditional Rank-Rank Regression$^*$

abstractRank-rank regression is commonly employed in economic research as a way of capturing the relationship between two economic variables. The slope of this regression is the Spearman rank correlation, a classical measure of association. However, in many applications it is common practice to include covariates to account for differences in association levels between groups as defined by the values of these covariates. This is either done by including the covariates or by modeling the residuals obtained after partialing out the impact of the covariates. In each of these instances the resulting rank-rank regression coefficients can be difficult to interpret. We propose the conditional rank-rank regression, which uses conditional ranks instead of unconditional ranks, to measure average within-group persistence. The coefficient of this new regression corresponds to the average Spearman rank correlation conditional on the covariates, a natural summary measure of within-group association. We develop a flexible estimation approach using distribution regression and establish a theoretical framework for large sample inference. An empirical study on intergenerational income mobility in Switzerland demonstrates the advantages of this approach. The study reveals stronger intergenerational persistence between fathers and sons compared to fathers and daughters, with the within-group persistence explaining 62% of the overall income persistence for sons and 52% for daughters. Smaller families and those with highly educated fathers exhibit greater persistence in economic status.

Introduction

The linear regression of the rank of a variable $Y$ on the rank of another variable $W$ is commonly referred to as a rank-rank regression (RRR). It has become increasingly popular in empirical investigations in economics to analyze policy relevant issues such as, for example, mobility and sorting behavior beller2006intergenerational,dahl2008association,chetty2014land,adermon2018intergenerational,murphy2020top.\footnote{chetverikov2025inferencerankrankregressions have recently documented that 40 articles published between January 2013 and February 2024 in American Economic Review, Journal of Political Economy, Quarterly Journal of Economics, and Review of Economic Studies employ RRRs in their empirical analysis.} A desirable feature of RRR is that its slope coefficient corresponds to the Spearman rank correlation coefficient between $Y$ and $W$, a classical measure of association with desirable properties such as invariance to monotone transformations of the variables and robustness to outliers and heavy tails spearman1904proof,kendall1948rank.\footnote{maasoumi2022generalized questioned the economic interpretation of the RRR as a measure of mobility due to the use of linear regression and proposed alternative measures based on nonparametric regression.}

We propose a general method to control for covariates $X$ in RRRs. Covariates are commonly accounted for in RRRs either via their inclusion as additional regressors (RRRX), or by conducting RRR on residuals after partialling out their influence (RRR.res). We show that these two approaches can have undesirable properties. In particular, their resulting coefficients can be difficult to interpret and produce counterintuitive results. For example, chetverikov2025inferencerankrankregressions noted that the slope of RRRX is no longer the Spearman correlation and might lie outside the interval $[-1,1]$. Moreover, the ranks of the residuals in RRR.res can only be related to meaningful ranks of the original variables under restrictive assumptions on the distribution of the variables of interest conditional on the covariates.\footnote{For example, the ranks of the residuals of the linear regression of $Y$ on $X$ correspond to the ranks of $Y$ conditional on $X$ under the location-shift model $Y = X'\beta + \varepsilon$ where $\varepsilon$ is independent of $X$. However this relationship generally does not hold otherwise.} We call our proposal conditional rank-rank regression (CRRR) because it employs conditional ranks. Intuitively, we replace the linear partialing out of the covariates implicit or explicit in the existing approaches by a nonlinear partialing out. We show that our approach always delivers coefficients that are easy to interpret and which do not have the undesirable properties of those from the existing approaches.

In its canonical form, CRRR regresses the rank of $Y$ conditional on $X$ on the rank of $W$ conditional on $X$. It provides an alternative to RRRX and RRR.res and produces estimates that are easy to interpret. Indeed, we show that the CRRR slope is equal to the Spearman correlation of $Y$ and $W$ conditional on $X$, averaged over the distribution of $X$, which is a natural summary measure of within-group association with similar properties to the Spearman rank correlation. In Statistics, this type of measure of association is so-called covariate-adjusted and it has been frequently used in biostatistic applications gijbels2011conditional,liu2018covariate,eden2022nonparametric,wei2023partial. Another attractive feature of CRRR is that it is suitable for subgroup analysis, i.e. to run separate CRRR by groups defined from a categorical variable. If this variable is included in $X$, the CRRR slope for each group has the interpretation of the average conditional Spearman rank correlation for that group.

CRRR is also attractive for economic applications. For example, hagedorn2017identifying used ranks of residuals of wages on worker and establishment characteristics to analyze labor market sorting. However, these ranks based on residuals are statistical constructions that, in general, do not have a natural interpretation in terms of the original wages. Our proposal is to conduct the analysis with conditional wage ranks, which reflect wage ranks within groups of workers with the same characteristics, and are accordingly closely related to the observed wages. CRRR can also be applied to the analysis of sorting in other matching settings such as marriage and worker-firm markets haltiwanger2015cyclical,guiso2022assortative,haner2024marry.

The interpretation of the CRRR slope is different from the RRR slope. Assume, for example, in an investigation of sorting behavior of married couples that $Y$ is husband's wage, $W$ is wife's wage and $X$ is number of children. The RRR slope without covariates is the correlation between husband's and wife's wage ranks where the ranks are relative to the entire wage distribution. The CRRR slope is the rank correlation where the ranks are relative to the wage distribution of those households with the same number of children. The slope of RRRX is the regression slope of the husband's wage rank on wife's wage rank, where the ranks are relative to the entire wage distribution and the wife's rank is centered to have the same mean within households with the same number of children. The slope of this regression might be difficult to interpret as the centered rank is no longer a rank. We believe that in many settings the CRRR slope better reflects the relationship researchers are intending to capture when they control for covariates as it is closer to the ceteris paribus effect that economic models typically predict. Mathematically, the difference between CRRR and RRRX is the order in the application of the rank and covariate partialing out operators. RRRX obtain ranks first and partials out the covariates second, whereas CRRR reverses the order. The final outcome differs across procedures as the two operators do not commute due to the nonlinearity of the rank operator.

CRRR is different from RRR because they have different dependent variables, conditional versus unconditional (marginal) ranks. In some instances the explicit objective is to predict the marginal rank of $Y$ using $W$ and $X$. For example, in intergenerational mobility applications one might wish to predict the child's income rank using father's income and family household characteristics such as father's immigration status abramitzky2021intergenerational. Here, the marginal rank of $Y$ (child's rank in the entire distribution) might be more interesting than the rank of $Y$ conditional on $X$ (child's rank among those who have the same father's immigration status indicator). The commonly employed approach in these instances is to include $X$ as a covariate in the RRR or to perform subgroup analysis for different values of $X$. We refer to both approaches as RRRX as the subgroup analysis can also be implemented using RRR by including suitable interactions. As noted above, chetverikov2023inference observed that the coefficient of the rank of $W$ in RRRX might be difficult to interpret. In Section (ref) we show through an example that both forms of RRRX can deliver counterintuitive economic results in terms of mobility and predict ranks of $Y$ outside $[0,1]$.\footnote{Another way to avoid this problem is to use a fractional response regression to model the ranks of $Y$ conditional on the ranks of $W$ and $X$ papke1996econometric. However, this approach requires parametric assumptions and the model coefficients lack a natural interpretation beyond their signs.} The underlying cause of these problems is that the rank of $W$ on the right hand side of the RRRX does not satisfy the properties of a rank after partialing out the effect of the covariates.

A variant of CRRR (CRRR') can be used in these situations. CRRR' regresses the marginal rank of $Y$ on the rank of $W$ conditional on $X$ and possibly $X$. Similar to CRRR, CRRR' bypasses the conceptual problems associated with RRRX. We show that the CRRR' slope corresponds to a correlation and therefore always lies between $-1$ and $1$. The square of this slope indeed corresponds to the partial R-squared of the conditional ranks of $W$. That is, it is the fraction of the variance of the marginal rank of $Y$ explained by the conditional rank of $W$, net of the impact of the covariates $X$. This measure is useful for performing variance decompositions. CRRR' is also suitable for subgroup analysis. The CRRR' slope indeed can be decomposed as the average of the CRRR' slopes in each group, weighted by the size of the group. Moreover, CRRR' predicts ranks in $[0,1]$. RRRX does not satisfy any of the previous properties. This is illustrated in the example presented in Section (ref). There are also more nuanced situations where we have two groups of covariates, $X=(X_1,X_2)$, and we are interested in the ranks of $Y$ conditional on $X_1$, but want to use $X_2$ to improve prediction or for some other purposes. For example, in an intergenerational mobility application, we might be interested in predicting regional child's income ranks using father's income and household characteristics such as father's immigration status. Our approach can also be applied to these settings. Here we would regress the ranks of $Y$ conditional on $X_1$ (region) on the ranks of $W$ conditional on $X$ (region and father's immigration status) and possibly $X$.

We provide an estimator of the CRRR coefficients based on distribution regression (DR). Here we do not distinguish between the different variations of CRRR because they all can be treated jointly by allowing different conditioning covariates or equivalently restricting some coefficients of the DR to be zero. Like the estimator of the RRR coefficients, our estimator consists of two steps. In the RRR, the first step estimates the marginal ranks of $Y$ and $W$ using the empirical distribution, and the second step runs the linear regression of the estimated ranks of $Y$ on the estimated ranks of $W$ or computes the sample correlation between these ranks. Both versions of the second step produce numerically identical results if there are no ties in the observed values of $Y$ and $W$. In CRRR, the first step estimates the conditional ranks by running logit or probit DRs of $Y$ on $X$ and $W$ on $X$ at multiple values of $Y$ and $W$ to trace the entire conditional distributions. The second steps are identical to RRR, but the linear regression and correlation versions are no longer numerically identical even if there are no ties. They are both, however, consistent estimators of the CRRR slope. The CRRR estimator is computationally tractable, albeit somewhat more demanding than RRR.

We derive the asymptotic distribution of the CRRR estimator and provide feasible inference theory. chetverikov2023inference noted that standard inference methods for linear regression do not apply to the RRR estimator because both the independent and dependent variables, namely the estimated ranks, are generated. The same problem applies to CRRR. The theory for the RRR estimator was derived using U-statistic theory hoeffding1948class or the delta method ren1995hadamard when $Y$ and $W$ are continuous. We employ the functional delta method approach to derive the theory because the CRRR estimator does not have a U-statistic structure. The application of the functional delta method to the CRRR estimator presents several differences with respect to RRR. The first ingredient of both approaches consists of writing the parameter of interest as a functional of inputs: the joint distribution of $Y$ and $W$ for RRR, or the conditional distributions of $Y$ and $W$ conditional on $X$ and the joint distribution of $Y$, $W$ and $X$, for CRRR. The CRRR functional is more complicated than the RRR functional and the inputs for CRRR live in more complex spaces than those for RRR. As a result, the Hadamard differentiability of the RRR functional established by ren1995hadamard does not cover the CRRR functional. We establish the Hadamard differentiabilty of the CRRR functional in the relevant spaces.

The second ingredient required for the application of the functional delta method is to establish functional central limit theorems for the estimators of the inputs. For RRR, this follows from the now classical large sample theory of the empirical distribution function. A challenge for CRRR is that existing theory for DR estimators of conditional distributions excludes the tails. In particular, the available functional central limit theorems only hold for compact strict subsets of the support. We deal with this problem by imposing assumptions on the tail behavior of the DR model that allows us to estimate the conditional distribution in the tails. We then obtain functional central limit theorems for DR estimators of conditional distributions over the entire support.

Combining these two ingredients we establish that the CRRR estimator follows a normal distribution around the CRRR slope in large samples via the delta method. The asymptotic variance has a complicated expression that might be difficult to estimate analytically. We develop the use of the exchangeable bootstrap to obtain standard errors and construct confidence intervals. Exchangeable bootstrap include the most common forms of bootstrap such as empirical, weighted and subsampling bootstrap as special cases. We establish its validity in large samples from the functional delta method for the bootstrap. Like hoeffding1948class and ren1995hadamard, our theory covers the case where $Y$ and $W$ are continuous. The theory can be extended to noncontinuous variables following the analysis of chetverikov2023inference for the RRR estimator. We leave this extension to future research.

We apply the CRRR estimator to analyze intergenerational income mobility in Switzerland using the Economic Well-Being of the Working and Retirement Age Population Data (WiSiER) from 1986 to 2016 for 11 cantons. This dataset contains rich information on socioeconomic, demographic and other variables merged from tax records, social insurance, unemployment records and surveys, and can be linked for fathers and children. We uncover a gender gap in intergenerational income mobility as the persistence between fathers and sons is stronger than between fathers and daughters, both with and without controlling for covariates. We also find that about $62\%$ and $52\%$ of the overall (unconditional) income persistence is explained by the within-group income persistence for sons and daughters, respectively; where groups are defined by child's and father's marital status, Swiss citizenship, high school graduation, experience, number of children and canton and year fixed effects. We also provide evidence supporting greater persistence for fathers with higher education and fathers with only one child. These results uncover the substantial role of both within-group and between-group persistence in explaining the intergenerational transmission of income.

Our methodology complements related methodological developments in the econometric and statistical literature. While chetverikov2023inference provides the inferential theory for RRR and RRRX with marginal ranks, their approach does not apply to conditional ranks employed in CRRR, as we explained earlier. kitagawa2018measurement studied measurement error in RRRs in the context of intergenerational income mobility. Another new development is lei2024causal which examines causal underpinnings of the marginal rank regressions. More closely related to our work, liu2018covariate introduced a covariate-adjusted Spearman coefficient based on the probability scale residuals of li2012new and shepherd2016probability, which we further discuss in Section (ref). They proposed a modelling strategy and an estimator based on a monotone transformation of a location-shift model, which is a special case of the distribution regression model chernozhukov+13inference.\footnote{liu2018covariate did not develop inference theory for their estimator in the case where $Y$ and $W$ are continuous, although they conjectured that it is possible; see Section 5 ibid.} The DR approach is more flexible and comprehensive, in that it can approximate the true conditional distribution function arbitrarily well by considering rich sets of basis functions with respect to the covariates. This is not generally possible with transformations of location models.\footnote{A transformation of a location-shift model takes the form $H(Y) = b(X)'\beta + \epsilon$, where $\epsilon$ is independent of $X$ and has a known distribution, $b(X)$ is a basis of functions of $X$, and $H$ is an unknown monotone transformation function. The model is more flexible than just a location model, but it does not allow the covariates to affect the distribution of $H(Y)$ other than through its location. No matter how rich the basis functions $b(X)$ are, the model is not guaranteed to cover the true conditional distribution, even in the limit where the dimension of $b(X)$ grows large.} gijbels2011conditional and veraverbeke2011estimation developed estimators of conditional measures of association using copulas, relying on the representation of these measures in terms of the conditional copula for continuous outcomes. They derived distribution theory for estimators of conditional versions of Kendall's tau and Blomqvist's beta blomqvist1950measure with scalar covariates. In contrast, our modelling and estimators are based on conditional distributions and apply to multivariate covariates.

\paragraph{Outline} The rest of the paper is organized as follows. Section (ref) contrasts CRRR with RRR via a simple conceptual example. Section (ref) introduces formally CRRR and derives its properties. Section (ref) describes the estimation procedure based on DR and a bootstrap algorithm to make inference. Section (ref) discusses an application examining the relationship between fathers' and their children's labor income using Swiss data, while Section (ref) provides asymptotic theory. Additional theoretical and numerical simulation results, and proofs are reported in the Appendix.

CRRR vs RRR: A Conceptual Example

To contrast the relative performances of CRRR and different forms of RRR in capturing the relationship between $Y$ and $W$ in the presence of covariates $X$, we examine a simple conceptual example in which $X$ is binary. This example is convenient because with a binary covariate there is no concern that the differences between the methods are driven by particular modeling strategies to specify various regression functions.

Let $Y$ be daughter's height (in cm), $W$ be father's height (in cm) and $X$ be a country indicator, say $X=0$ for the Netherlands and $X=1$ for Ireland. Conditional on $X$, $Y$ and $W$ follow a bivariate normal distribution with mean parameters that may depend on the value of $X$ and constant covariance matrix. More specifically,

equation[equation omitted — 276 chars of source]

where ${\mathrm{P}}(X=0) = {\mathrm{P}}(X=1) = 1/2$. For example, $\delta$ can be a negative country shock such as the Irish Famine that affects the father's height in Ireland, but not in the Netherlands. We consider two cases depending on the extent of the effect of the shock as measured by $\delta$:

itemize• No shock: $\delta = 0$. • Negative shock: $\delta = 12$.

Table (ref) compares measures of overall and within-country intergenerational height persistence based on rank correlation with the estimands of RRR, CRRR and two versions of RRRX. Thus, $\rho_{Y,W}$ and $\bar \rho_{Y,W \mid X}$ are the Spearman rank correlation coefficient between $Y$ and $W$ and the expected Spearman rank correlation between $Y$ and $W$ conditional on $X$; RRR is the RRR slope; CRRR is the CRRR slope; RRRX-A is the slope of the RRRX, that is, $\beta_1$ in $$ \tilde U = \beta_0 + \beta_1 \tilde V + \beta_2 X + \epsilon, \quad {\mathrm{E}}[\epsilon] = {\mathrm{E}}[\tilde V \epsilon] = {\mathrm{E}}[X \epsilon] = 0, $$ where $\tilde U$ and $\tilde V$ denote the marginal ranks of $Y$ and $W$, respectively; and RRRX-I is the average slope of the RRRs run separately by the values of $X$, that is, $\beta_1$ in $$ \tilde U = \beta_0 + \beta_1 \tilde V + \beta_2 [X-.5] + \beta_3 [X-.5] \tilde V + \epsilon, \quad {\mathrm{E}}[\epsilon] = {\mathrm{E}}[\tilde V \epsilon] = {\mathrm{E}}[X \epsilon] = {\mathrm{E}}[X \tilde V \epsilon] = 0 $$

table[table omitted — 700 chars of source]

We find that all the methods give the same answer when $\delta=0$, that is when the joint distribution of daughter's and father's heights is the same in both countries.\footnote{Up to numerical error, all the slopes are equal to the rank correlation of the bivariate normal with correlation $c=.6$, $\rho_S(Y,W)=6\arcsin{(c/2)}/\pi = .58$ cramer1999mathematical.} When $\delta=12$, RRR gives the overall mobility $\rho_{Y,W}$, whereas CRRR gives the within-country mobility $\bar \rho_{Y,W \mid X}$. Both forms of RRRX produce measures greater than one, which do not correspond to any rank correlation and might be difficult to interpret. Whether RRR or CRRR is the right measure depends on the application. In this case, CRRR measures average intergenerational mobility within each country whereas RRR measures intergenerational mobility pooling the two countries. They would lead to different conclusions about the effect of the Irish Famine. According to RRR, the famine reduces overall height persistence, whereas it does not have any effect within each country according to CRRR. Both versions of RRRX lead to the opposite conclusion that the famine increases height persistence. This conclusion does not correspond to a change in either overall or within-country mobility as measured by Spearman correlations.

figure[figure omitted — 284 chars of source]
figure[figure omitted — 288 chars of source]

While RRR and RRRX are frequently employed in intergenerational mobility analyses, the rank correlation is not always the object of interest. Many empirical studies in this literature focus on the so-called level of absolute upward mobility. This is defined as the expected marginal rank of a child with a father at a specified percentile, typically the 25th percentile chetty2014land. More generally, it is the conditional expectation function (CEF) of the child's rank given values for the father's rank and the covariates. Figures (ref) and (ref) report predicted daughter's height (marginal) ranks obtained from CRRR', the variant of CRRR that regresses marginal ranks of $Y$ on conditional ranks of $W$, and different versions of RRR for the scenarios $\delta=0$ and $\delta=12$, respectively. In this case CRRR' is the same as CRRR as the marginal and conditional ranks of $Y$ coincide because $Y$ is independent of $X$.\footnote{We also examined a variant of the data generating process of the model in which the conditional mean of $Y$ was affected by a different binary variable such that CRRR' was different from CRRR. It did not produce any qualitative differences from the conclusions that follow.} Accordingly, we shall refer to CRRR' as CRRR. RRR only uses the father's height $W$, whereas RRRX and CRRR also employ the value of the covariate $X$.\footnote{We can also include $X$ as an additional regressor in CRRR, but in this case does not change the results because $Y$ is independent of $X$ conditional on the conditional rank of $W$.} We do not distinguish between RRRX-A and RRRX-I because they produce very similar predictions. To make the methods comparable, CRRR evaluates the prediction at the father's conditional rank corresponding to the marginal quantile of the father's (marginal) rank.\footnote{More specifically, if the father's marginal rank is $v$ and $X=x$, then the corresponding father's conditional rank is $V_{w,x} = F_{W \mid X}(F_W^{-1}(v) \mid x)$, where $F_W$ and $F_{W \mid X}$ are the marginal and conditional distributions of $W$.} As a benchmark of comparison, we report the conditional expectation function (CEF) of the marginal rank of $Y$ given $X$ and $W$ which, using the properties of the multivariate normal, is:

equation[equation omitted — 189 chars of source]

where $\Phi$ is the distribution of the standard normal and $V_{w,x}$ is the conditional rank of $W$ at $W=w$ and $X=x$.\footnote{Note that it is equivalent conditioning on $W$ to conditioning on the marginal rank of $W$ because there is a one-to-one relationship between them. Indeed, $$ {\mathrm{E}}(\tilde U \mid \tilde V = \tilde v, X=x) = \Phi \left( \frac{.6 *\left( F^{-1}_W(\tilde v) - 180 + \delta x\right) }{4 \sqrt{2 - 0.6^2}} \right), $$ where $F_W$ is the distribution of $W$. } The true CEF can be computed in this example because we know the joint distribution of $(Y,W)$ conditional on $X$.

When $\delta =0$, $X$ does not help predict because $(Y,W)$ and $X$ are independent. All the methods yield the same prediction function in both countries, which is a linear approximation to the CEF. When $\delta = 12$, $X$ helps predict because the CEF is different in both countries. In this case, RRRX and CRRR give better approximations to the CEF than RRR, because they make use of the information in $X$. RRRX provides a linear approximation to the CEF in each country, but delivers predicted daughter's ranks lower than $0$ for father's ranks below the 28th percentile when $X=0$ (Netherlands), and greater than $1$ for father's ranks above the 72th percentile when $X=1$ (Ireland). RRRX therefore produces unreasonable predictions for many relevant values of the father's rank in half of the population. CRRR does not suffer from this problem. The implicit nonlinearity introduced by the use of the conditional father's rank, instead of the marginal rank, enforces that all the predicted daughter's ranks are between $0$ and $1$. In other words, while CRRR provides a linear approximation to the CEF with respect to the conditional father's rank, it provides a nonlinear approximation in terms of marginal father's rank. This approximation is bounded in $[0,1]$ because the absolute value of the slope is less than $1$. There is still approximation error, however, because the CEF is also a nonlinear function of the conditional father's rank in this case; see (ref).

When $\delta=12$, the CEF is convex when $X=0$ (Netherlands) and concave when $X=1$ (Ireland). Intuitively, a low marginal father's rank corresponds to a very low conditional father's rank in the Netherlands, which is associated with a very low conditional and marginal daughter's rank due to the high positive within-country correlation. Conversely, a high marginal father's rank corresponds to a very high conditional father's rank in Ireland, which is associated with a very high conditional and marginal daughter's rank. The CRRR predictions agree with these shape restrictions, whereas RRRX imposes linearity by construction. In both scenarios for $\delta$, the CRRR slope of $0.58$ indicates that about $34\% (=0.58^2)$ of the variability of the daughter's rank is explained by the father's rank net of covariates (country indicator), which is also the R-squared of CRRR. In this case, this explained fraction is the same in both countries. The RRRX slope of $1.07$ when $\delta=12$ does not have an interpretation as either a partial or overall R-squared.

table[table omitted — 765 chars of source]

Finally, we conduct a subgroup analysis using CRRR and RRR. Table (ref) compares the conditional Spearman rank correlation, CRRR slope and RRR slope for each value of $X$. CRRR produces measures that are invariant to both $X$ and $\delta$, which correspond to the rank correlations between $Y$ and $W$ conditional on $X$. If $Y$ and $X$ were not independent, then CRRR and CRRR' would give different results, but still both would produce slopes for each country that would average to the overall CRRR or CRRR' slope.\footnote{In the case of CRRR, the slopes for each country would also correspond to conditional Spearman rank correlations.} The RRR slopes are the same as the CRRR slopes when $X$ is irrelevant. RRR, however, delivers different slopes for the different values of $X$ when $\delta=12$, and also across the different values of $\delta$. The RRR slopes are greater than one when $\delta=12$, confirming that they do not correspond to correlations and making them hard to interpret. Moreover, the within-group RRR slopes do not average to the overall RRR slope in general hertz2008group.

While the above example is simple, it illustrates a number of important features which apply, or are likely to apply, in more complicated settings. First, in the presence of covariates, the RRRX slope does not correspond to any measure of rank correlation and can be difficult to interpret. In contrast, the CRRR slope corresponds to a specific form of rank correlation. Moreover, while neither CRRR or RRRX exactly predict the CEF for all values of the father's rank, the results here suggest that CRRR provides a more sensible and accurate approximation. Moreover, it directly produces a measure of the proportion of variability in the ranks of $Y$ explained by the ranks of $W$ net of the covariate $X$ and is suitable for subgroup analysis.

Conditional Rank-Rank Regression

Let $(Y,W)$ be a bivariate random variable with joint distribution $F_{Y,W}$ and marginal distributions $F_Y$ and $F_W$ for $Y$ and $W$, respectively. For example, $Y$ is child's income and $W$ is father's income. We assume that $Y$ and $W$ are continuous.

Canonical RRR

We start by reviewing the canonical rank-rank regression (RRR). Let $\tilde U := F_Y(Y)$ and $\tilde V:=F_W(W)$ denote the (marginal) ranks of $Y$ and $W$. By continuity of $Y$ and $W$, ranks are uniformly distributed, $\tilde U \sim U(0,1)$ and $\tilde V \sim U(0,1)$. The RRR of $Y$ on $W$ is defined as the correlation between $\tilde U$ and $\tilde V$ or the slope of the linear regression of $\tilde U$ on $\tilde V$ (or vice versa): $$ \rho := \mathrm{Cor}(\tilde U, \tilde V)= \frac{\operatorname{Cov}(\tilde U,\tilde V)}{\operatorname{Var}(\tilde U)} = \frac{\operatorname{Cov}(\tilde U,\tilde V)}{\operatorname{Var}(\tilde V)} = 12 \ {\mathrm{E}}[(\tilde U-.5)(\tilde V - .5)], $$ where all the equalities follow from the uniform distribution of $\tilde U$ and $\tilde V$. In statistics this correlation measure is the celebrated Spearman rank correlation between $Y$ and $W$, and is widely used to measure dependence between variables. It is invariant to rescaling and all increasing monotone transformations of the variables, and has gained prominence for that reason. The rank correlation has become popular in economics in studies of income and wealth mobility due to its interpretability as a measure of persistence and scale-free nature.

Conditional RRR

We introduce now the conditional rank-rank regression (CRRR). Let $X$ denote a vector of covariates related to $Y$ and $W$ including, for example, child's and father's education, age, marital status and nationality. Let $F_{Y\mid X}$ and $F_{W \mid X}$ denote the distributions of $Y$ and $W$ conditional on $X$. Then, $U := F_{Y \mid X}(Y \mid X)$ and $V:=F_{W \mid X}(W \mid X)$ are the conditional ranks of $Y$ and $W$, where conditioning is on $X$. For example, $U$ and $V$ would be child's and father's income ranks among families with the same composition in terms of covariates. By continuity of $Y$ and $W$, the conditional ranks follow the uniform distribution, conditional on $X$: $$U \mid X \sim U(0,1) \text{ and } V \mid X \sim U(0,1),$$ and also unconditionally. This implies the constant variance property, $$\operatorname{Var}(V) = \operatorname{Var}(U) =\operatorname{Var}(V \mid X) = \operatorname{Var}(U \mid X) = 1/12$$ and the constant mean property, $${\mathrm{E}} V = {\mathrm{E}} U = {\mathrm{E}}(V \mid X) = {\mathrm{E}} (U \mid X) =.5.$$ Note that both $U$ and $V$ are marginally independent of $X$ , but not necessarily jointly independent so the correlation between $U$ and $V$ can depend on $X$.\footnote{That is, $U \mathop{\perp\!\!\!\!\perp} X$ and $V \mathop{\perp\!\!\!\!\perp} X$, but generally $(U,V) \not\perp\!\!\!\!\perp X$, where $\mathop{\perp\!\!\!\!\perp}$ denotes stochastic independence.}

The CRRR of $Y$ on $W$ given $X$ is defined as either the correlation between $U$ and $V$ or the slope of the linear regression of $U$ on $V$ (or vice versa):

equation[equation omitted — 176 chars of source]

CRRR is the average conditional correlation between conditional ranks:

equation[equation omitted — 124 chars of source]

where $\rho_{Y,W\mid X}$ denotes the conditional Spearman rank correlation between $Y$ and $W$ conditional on $X$, which is equal to $\mathrm{Cor}(U,V\mid X)$ by definition. Equation (ref) follows from $\operatorname{Cov}(U,V) = {\mathrm{E}} [\operatorname{Cov}(U,V\mid X)]$ by the law of total covariance since $\operatorname{Cov}[{\mathrm{E}}(U\mid X), {\mathrm{E}}(V\mid X)] = 0$; moreover, the conditional variance of $U$ and $V$ is equal to the unconditional variance. In summary, CRRR is the average Spearman rank correlation between $Y$ and $W$ conditional on $X$, averaged over the distribution of $X$, which is a summary measure of within-group persistence.

By the properties of $U$ and $V$, the CRRR can also be represented as the rescaled covariance of conditional ranks:

equation[equation omitted — 87 chars of source]

a formula convenient for estimation. Moreover, in the regression version of the CRRR, the intercept, $\alpha_C$, is mechanically related to the slope, $\rho_C$, through

equation[equation omitted — 109 chars of source]

This relationship raises concerns about the interpretation of intercepts and slopes in the regression versions of the RRR used in intergenerational mobility studies as measures of absolute and relative mobility, repectively.\footnote{chetty2014land noted a relationship analogous to (ref) for the regression version of the canonical RRR.}

Finally, we note that correlation of conditional ranks is generally not equal to correlation of marginal (unconditional) ranks: $$ \rho_C \neq \rho $$ but the two agree under independence from $X$, namely $\rho_C = \rho$ if $Y \mathop{\perp\!\!\!\!\perp} X$ and $W \mathop{\perp\!\!\!\!\perp} X$, because in that case $U = \tilde U$ and $V = \tilde V$.

In the context of the income mobility application, $\rho_C$ measures within-group income persistence and $\rho$ measures overall income persistence, encompassing both within-group and between-group persistence. The between-group persistence can then be defined as the difference between the marginal rank and conditional rank correlations: $$\textrm{Between-group persistence} = \rho - \rho_C.$$ Assume, for example, that the covariates $X$ capture family characteristics such as size or parental education. The difference between the two measures can be explained as follows: The within-group or unexplained persistence $\rho_C$ captures the extent to which father's income rank facilitates child's income rank among families with the same observable characteristics. In other words, it measures the influence of father's income on child's income, where the variation in father's and child's incomes comes from unobserved characteristics such as family status, ability and the extent of social or professional networks. On the other hand, the between-group measure $\rho - \rho_C$ aims to capture the contribution of observed characteristics to income persistence.

We can further decompose the between-group persistence using the total law of covariance: $$ \rho - \rho_C = 12 \operatorname{Cov}[{\mathrm{E}}(\tilde U \mid X),{\mathrm{E}}(\tilde V \mid X)] + 12 {\mathrm{E}}[\operatorname{Cov}(\tilde U,\tilde V \mid X) - \operatorname{Cov}(U,V \mid X)], $$ where the first component is the covariance of conditional means of marginal ranks, and the second component is the average conditional covariance of marginal ranks net of the average within-group inequality.

Rank-rank regression with covariates (RRRX)

CRRR is different from RRR with covariates $X$ (RRRX) where $X$ is included additively (or non-additively) in the regression of marginal ranks, $\tilde U$ on $\tilde V$. We believe that our proposal is a more natural and adequate way to incorporate covariates. In fact, RRRX with additive covariates is no longer related to a rank correlation nor has to lie in the interval $[-1,1]$. RRRX is also more difficult to interpret as it does not correspond to a meaningful measure of within-group persistence. Making RRRX more flexible by including interactions between $X$ and $\tilde V$ does not mitigate any of these problems. In fact, making RRRX fully nonparametric also does not alleviate the problem. We show in the next section that even in the simplest case where $X$ is binary, the nonparametric RRRX does not capture meaningful economic quantities. When $X$ is discrete, the nonparametric approach (tabulating unconditional rank correlation by subgroups) does not either.

In what follows, we systematically explain the current approaches to RRRX and contrast these with the CRRR approach. We use the intergenerational income application to give context to the discussion.

example[RRR vs RRRX vs CRRR] Let $Y$ be child's income, $W$ be father's income and $X$ be an indicator for father's high school diploma. In this case, the marginal ranks $\tilde U$ and $\tilde V$ are relative to the distribution of income in the entire population that includes fathers with and without high school diploma, whereas the conditional ranks $U$ and $V$ are relative to the distribution of income of those with the same father's high school diploma status. RRR measures the correlation between the marginal ranks, whereas CRRR measures the average correlation between the conditional ranks, that is CRRR first obtains the rank correlation separately for fathers with and without high school diploma and then averages these correlations weighted by the proportions of each type in the population. CRRR therefore can be interpreted as a within-group or ceteris paribus effect, where the families are ranked and compared with families where the father's high school diploma status is held constant. The slope of the RRRX with covariates does not have a natural interpretation in terms of intergenerational mobility. It measures the coefficient in the regression of child's marginal rank on father's marginal rank, where the father's marginal rank is recentered to have the same mean for fathers with and without high school diploma. This slope does not have an interpretation as a within-group persistence. Moreover, it is not a rank correlation and can lie outside the interval $[-1,1]$, because the recentered father's marginal rank does not have the properties of a rank. In particular it no longer follows a uniform distribution. {\tiny {\ensuremath{\blacksquare}}}

Subgroup Analysis

When $X$ is discrete, it is common to run RRRs separately for each value of $X$ instead of including $X$ as an additive control. For example, abramitzky2021intergenerational run separate RRR of child's income on father's income by father's immigration status. The slopes of these regressions cannot be interpreted in terms of rank correlations or even as conditional correlations between the marginal ranks. To see this, note that the slope of the regression of $\tilde U$ on $\tilde V$ conditional on $X=x$, is not equal to conditional correlation of $\tilde U$ and $\tilde V$: $$ \frac{\operatorname{Cov}(\tilde U,\tilde V \mid X=x)}{\operatorname{Var}(\tilde V \mid X=x)} \neq \frac{\operatorname{Cov}(\tilde U,\tilde V \mid X=x)}{\sqrt{\operatorname{Var}(\tilde V \mid X=x) \operatorname{Var}(\tilde U \mid X=x)}}, $$ because marginal ranks have different conditional distributions, i.e. $\tilde U \overset{d}{\not\sim} \tilde V \mid X=x$, in general. The slope therefore does not generally correspond with the conditional correlation of the marginal ranks conditional nor the conditional rank correlation between $Y$ and $W$. We give an example in Section (ref) where this slope is greater than one.

Consider now the CRRR. Assume we are interested in conducting a subgroup analysis of intergenerational mobility with respect to father's high school diploma or immigration status. Let $X_1 \subseteq X$ be a set of variables that define the subpopulation of interest such as an indicator for high school diploma and/or Swiss nationality. Then, the CRRR slope conditional on $X_1=x_1$ is:\footnote{Indeed, by the law of total covariance with respect to $X$ and uniformity of $U$ and $V$ conditional on $X$, $$ \operatorname{Cov}(U,V \mid X_1) = {\mathrm{E}}[\operatorname{Cov}(U,V \mid X) \mid X_1] + \operatorname{Cov}[{\mathrm{E}}(U\mid X),{\mathrm{E}}(V\mid X) \mid X_1) = {\mathrm{E}}[\operatorname{Cov}(U,V \mid X) \mid X_1], $$ and $$ \operatorname{Var}(V \mid X_1) = \operatorname{Var}(V \mid X) = \operatorname{Var}(U \mid X), $$ almost surely.}

equation[equation omitted — 330 chars of source]

Hence, the CRRR slope for the subgroup defined by $X_1=x_1$ corresponds to the average conditional rank correlation between $Y$ and $W$, where the average is taken with respect to the distribution of $X$ conditional on $X_1=x_1$. This allow us, for example, to measure intergenerational mobility separately for families with fathers with and without high school diploma.\footnote{If $X_1 \not\subseteq X$, , the slope no longer has an interpretation as average conditional rank correlation because $V \not\sim U \mid X_1$ in general.}

Variations of CRRR (CRRR')

There are applications where the researcher might want to use different sets of covariates to obtain the conditional ranks $U$ and $V$. In the intergenerational mobility application, for example, we might not want to control for son's education to obtain the father's income rank. In this case the CRRR' slope still corresponds to an average correlation between the ranks. To see this, let $U=F_{Y \mid X_1}(Y \mid X_1)$ and $V=F_{W \mid X_2}(W \mid X_2)$ with $X_1 \neq X_2$ and $X= X_1 \cap X_2$, the set of covariates included in both $X_1$ and $X_2$, then: $$ \rho_C = \frac{\operatorname{Cov}(U,V)}{\operatorname{Var}(V)} = {\mathrm{E}}\left[\frac{\operatorname{Cov}(U,V \mid X) }{\sqrt{\operatorname{Var}(V \mid X)\operatorname{Var}(U \mid X)}}\right], $$ where we use the law of total covariance with respect to $X$, $U \mathop{\perp\!\!\!\!\perp} X$, $V \mathop{\perp\!\!\!\!\perp} X$ and iterated expectations. The CRRR' slope therefore corresponds to the correlation between the ranks $U$ and $V$ conditional on the common covariates $X$, averaged over the distribution of $X$. Note, however, that $\rho_C$ in this case does not correspond to an average conditional rank correlation between $Y$ and $W$. The source of the difference is that $U \neq F_{Y \mid X}(Y \mid X)$ and $V \neq F_{W \mid X}(W \mid X)$ in general.\footnote{This rank correlation can be obtained by constructing the conditional ranks as $U=F_{Y \mid X}(Y \mid X)$ and $V=F_{W \mid X}(W \mid X)$.} One exception occurs when $Y$ is independent of the components of $X_2$ not included in $X_1$ conditional on $X_1$, and $W$ is independent of the components of $X_1$ not included in $X_2$ conditional on $X_2$. In that case, $$ \rho_C = {\mathrm{E}}\left[\frac{\operatorname{Cov}(U,V \mid \bar X) }{\sqrt{\operatorname{Var}(V \mid \bar X)\operatorname{Var}(U \mid \bar X)}}\right] = {\mathrm{E}} [ \rho_{Y,W \mid \bar X} ], $$ where $\bar X = X_1 \cup X_2$. This result follows by the law of total covariance with respect to $\bar X$ and uniformity of $V$ and $U$ conditional on $\bar X$.\footnote{Note that if $V \mid X_1 \sim U(0,1)$ and $V \mathop{\perp\!\!\!\!\perp} X_2 \mid X_1$, then $V \mid X \sim U(0,1)$.}

Like CRRR, CRRR' is suitable for subgroup analysis in the following sense. Let $\rho_C(x_2)$ be the CRRR' slope in the group defined by $X_2=x_2$, that is $$ \rho_C(x_2) = \frac{\operatorname{Cov}(U,V \mid X_2 = x_2)}{\operatorname{Var}(V \mid X_2 = x_2)}. $$ Then, by the law of total covariance and $V \perp\!\!\!\perp X_2$, $$ \rho_C = \frac{{\mathrm{E}}\left[\operatorname{Cov}(U,V \mid X_2)\right]}{\operatorname{Var}(V)} = {\mathrm{E}}[\rho_C(X_2)], $$ that is, the CRRR' slope can be decomposed as the average of the CRRR' slopes in each group defined by $X_2$, weighted by the size of the group.

An interesting example occurs when $X_1 = \emptyset$ and $X_2 = X$. In this case, $U = \tilde U$ and $\rho_C$ is the correlation between the marginal ranks of $Y$ and the conditional ranks of $W$. Unlike the RRR, the inclusion of covariates in this CRRR does not affect the coefficient of $V$ because $V \mathop{\perp\!\!\!\!\perp} X$, and can be used to perform a variance decomposition of $\tilde U$. Let, $$ \tilde{U} = \rho_C V + X'\beta_C + \varepsilon, \quad {\mathrm{E}}[(V; X) \varepsilon] = 0, $$ be the extended CRRR with covariates, where the first term of $X$ is a constant. Then, $$ \text{Var}(\tilde U) = \rho_C^2 \text{Var}(V) + \text{Var}(X'\beta_C) + \text{Var}(\varepsilon), $$ where the first two terms of the right-hand-side correspond to the contribution of $V$ and $X$ to the variance of $\tilde U$, and the third terms to the unobserved component. Indeed, $\rho_C^2$ measures the fraction of the variance of $\tilde U$ explained by $V$ since $\text{Var}(\tilde U) = \text{Var}(V)$.

Properties of CRRR

We conclude this section by gathering the properties of the CRRR slope in the following lemma.

lemma[CRRR Properties] Assume that $Y$ and $W$ are continuous random variables, $X$ is a vector of covariates, $U = F_{Y \mid X}(Y \mid X)$ and $V = F_{W \mid X}(W \mid X)$. Then, (1) The CRRR slope, $\rho_C$, has the representations given in (ref) and (ref). (2) The slope $\rho_C$ is the expected conditional Spearman rank correlation between $Y$ and $W$: $ \rho_C = {\mathrm{E}}[\rho_{Y,W \mid X}]. $ (3) Subgroup analysis: if $X_1 \subseteq X$, then (ref) holds. Therefore, $ \rho_C(x_1)$ is the average conditional rank correlation between $Y$ and $W$ in the group defined by $X_1= x_1$. (4) Let $U_1=F_{Y \mid X_1}(Y \mid X_1)$ and $V_2=F_{W \mid X_2}(W \mid X_2)$ with $X_1 \neq X_2$ and $X= X_1 \cap X_2$, then $$ \rho_C = \operatorname{Cor}(U_1,V_2) = {\mathrm{E}}[\operatorname{Cor}(U_1,V_2 \mid X)], $$ that is $\rho_C$ is the average conditional correlation between $U_1$ and $V_2$ given the set of common covariates $X$.
remarkThis lemma simply records the observations given above. It is useful to connect here to liu2018covariate who introduced the covariate-adjusted Spearman correlation coefficient as the correlation between the probability scale residuals of $Y$ and $W$. These residuals are defined as $ r(Y,F_{Y\mid X})$ and $r(W,F_{W\mid X})$, where $r(r,F_{R\mid X}) = F_{R\mid X}(r \mid X) + F_{R\mid X}(r- \mid X) -1$ and $F_{R \mid X}(r- \mid x) = \lim_{u \nearrow r}F_{R \mid X}(u \mid x)$, for $R \in \{Y,W\}$. In the case where $Y$ and $W$ are continuous, the probability scale residuals are affine transformations of the conditional ranks because $F_{R \mid X}(r- \mid x) = F_{R \mid X}( r \mid x)$, e.g., $r(Y,F_{Y\mid X}) = 2 U -1$, and the covariate-adjusted Spearman correlation equals to the CRRR slope. The properties in Lemma (ref)(2) and (3) then follow from results in liu2018covariate when $Y$ and $W$ are continuous. The conceptual difference is that our definition and derivations are based on the characterization of the Spearman correlation as the correlation between ranks or grade correlation kruskal1958ordinal, whereas theirs are based on the characterization of the Spearman correlation in terms of concordance-discordance probabilities.

Distribution Regression Estimator of CRRR

DR Model for Conditional Distributions

For estimation purposes, it is convenient to model the conditional distributions $F_{Y\mid X}$ and $F_{W \mid X}$ using the distribution regression (DR) model: $$ F_{R \mid X}(r \mid x) = \Lambda(x'\beta_R(r)), \quad R \in \{Y,W\}, \quad r \in \mathcal{R}, $$ where $\Lambda$ is the standard normal or logistic distribution, $\mathcal{R}$ is the support of $R$ and the first component of $x$ is a constant. The specification can be made more flexible by replacing $x$ by a vector of transformations of $x$ with good approximating properties.

As the data to estimate the conditional distribution function at the tails are sparse, it is necessary to impose some structure. We assume that the conditional distribution far in the tails can be extrapolated from the conditional distribution not too far in the tails.\footnote{This is in line with approaches used in extreme value theory that impose restrictions on the tail behavior allowing similar extrapolations. For example, see embrechts:1997 for a broad reference on the theory of extremes and victor:annals or chernozhukov:2011 for similar approaches in the context of extremal quantile regression.} We formalize this approach by imposing restrictions on the coefficient of the DR model in the tails.

Let $\bar{\mathcal{R}}$ be a compact strict subset of $\mathcal{R}$, for $\mathcal{R} \in \{\mathcal{Y},\mathcal{W}\}$, where $\mathcal{Y}$ and $\mathcal{W}$ are the supports of $Y$ and $W$, respectively. Then, we assume: $$ F_{R\mid X}(r \mid x) = \Lambda((r-\bar r)\alpha_R(\bar r) + x'\beta_R(\bar{r})), \quad R \in \{Y,W\}, \quad r \in \mathcal{R}\setminus\bar{\mathcal{R}}, $$ where $\bar r := \arg \min_{r' \in \bar{\mathcal{R}}} |r-r'|$ and $\alpha_R(\bar r) > 0$. That is, we postulate that the random variable $R$ behaves in the tails like a random variable with distribution $\Lambda$, after subtracting the location shift $x'\beta_R(\bar{r})$ and dividing by the scale $\alpha_R(\bar r)$, which are different at the upper and lower tails. Thus, the DR coefficient is restricted in the tails by: $$\beta_{R,1}(r) = \beta_{R,1}(\bar r) + (r - \bar r)\alpha_R(\bar r), \quad \beta_{R,-1}(r) = \beta_{R,-1}(\bar r),\quad R \in \{Y,W\}, \quad r \in \mathcal{R}\setminus\bar{\mathcal{R}},$$ where $\beta_R(r)$ is partitioned into $(\beta_{R,1}(r),\beta_{R,-1}(r)')'$ where $\beta_{R,1}(r)$ is the intercept and $\beta_{R,-1}(r)$ are the slope components. That is, $r \mapsto \beta_{R,1}(r)$ is a linear function and $r \mapsto \beta_{R,-1}(r) $ is constant on $\mathcal{R}\setminus\bar{\mathcal{R}}$.

Under the DR model, the conditional ranks can be expressed as the following functionals of the parameters: $$ U = \Lambda(X'\beta_Y(Y)), \quad V = \Lambda(X'\beta_W(W)). $$

Estimation

We provide several estimators of the CRRR slope based on the different representations of $\rho_C$ in (ref) and (ref). This section presents correlation-based and fully-restricted estimators. Regression-based estimators are given in Appendix (ref). We recommend the use of at least the correlation-based and fully-restricted estimators. The fully-restricted estimator, based on (ref), uses all the information available and is the simplest to compute, but it might be sensitive to misspecification of the model for the conditional distributions. In particular, it can deliver estimates outside the interval $[-1,1]$ under misspecification. The correlation-based estimator is more robust in the sense that it is the only estimator that guarantees estimates in the interval $[-1,1]$ under misspecification.\footnote{While we impose correct specification of the DR model for the conditional distributions, the derivation of the theoretical results does not rely fundamentally on correct specification. We conjecture that the probability limit of the correlation-based estimator still has an interpretation as correlation of pseudo-ranks under misspecification, but leave the formal analysis to future research.} We show in Appendix (ref) that the correlation-based estimator is asymptotically equivalent to the average of the regression-based and reversed regression-based estimators.

Let $\{Z_i:=(Y_i,W_i,X_i)\}_{i=1}^n$ be a random sample of $Z:=(Y,W,X)$. The following algorithms describe the estimators of $\rho_{C} $. All of them are based on DR.

algorithm[algorithm omitted — 2,250 chars of source]
remark[Computation] If the set $\mathcal{R}_n$ contains many elements, in step (1) we can either replace it by a smaller fine mesh or use a computationally fast method similar to chernozhukov2022fast to speed-up computation.\footnote{By stochastic equicontinuity of the conditional distribution processes $(r,x) \mapsto \sqrt{n} \left[\Lambda(x'\widehat \beta_R(r)) - \Lambda(x' \beta_R(r))\right]$, $R \in \{Y,W\}$, in Lemma (ref), the meshwidth $\delta$ should be such that $\delta \sqrt{n} \to 0$ as $n \to \infty$.} Note that the optimization program to obtain $\widehat \alpha_R(\bar r)$ in step (2) only needs to be solved twice, once for $r_0$ in the upper tail and once for $r_0$ in the lower tail. Also, we require $m \geqslant 30$, which is thought to be the minimal sample size required to estimate one parameter.

Bootstrap Inference

Section (ref) shows that the estimators described in Algorithm (ref) follow normal distributions in large samples. The variances of these distributions, however, have complicated forms and are difficult to estimate. Section (ref) also shows that the asymptotic distributions can be consistently estimated using exchangeable bootstrap. Exchangeable bootstrap is a general resampling method that includes empirical, weighted, wild and subsampling bootstrap as special cases; see Comment (ref). The following algorithm describes how to obtain bootstrap draws of the estimators of $\rho_C$.

algorithm[algorithm omitted — 2,315 chars of source]
remark[Bootstrap Weights] van1996weak notes that by appropriately selecting the distribution of the weights, exchangeable bootstrap covers the most common bootstrap schemes as special cases. The empirical bootstrap corresponds to where $(w_{n1},...,w_{nn})$ is a multinomial vector with parameter $n$ and probabilities $(1/n,...,1/n)$. The weighted bootstrap corresponds to where $w_{n1},...,w_{nn}$ are i.i.d. nonnegative random variables with ${\mathrm{E}}(w_{n1})=\operatorname{Var}(w_{n1})=1$, e.g. standard exponential. The wild bootstrap corresponds to where $w_{n1}, ..., w_{nn}$ are i.i.d. vectors with ${\mathrm{E}}(w_{n1}^{2+\varepsilon}) < \infty$ for some $\varepsilon > 0$, and $\operatorname{Var}(w_{n1})=1$. The $m$ out of $n$ bootstrap corresponds to letting $(w_{n1},...,w_{nn})$ be equal to $\sqrt{n/m}$ times multinomial vectors with parameter $ m$ and probabilities $(1/n,...,1/n)$. The subsampling bootstrap corresponds to letting $(w_{n1},...,w_{nn})$ be a row in which the number $n(n-m)^{-1/2}m^{-1/2}$ appears $m$ times and $0$ appears $n-m$ times ordered at random, independent of the data.

We now show how to use the exchangeable bootstrap to obtain standard errors for the estimators of $\rho_C$ and construct asymptotic confidence intervals for $\rho_C$. Algorithm (ref) describes the procedure for $\widehat \rho_C$. A similar algorithm applies to $\breve \rho_C$. Let $B$ a prespecified number of bootstrap repetitions and $\alpha$ be the significance level for the confidence intervals. For example, $B=500$ and $\alpha=0.05$.

algorithm[algorithm omitted — 1,217 chars of source]

Empirical Application

We analyze intergenerational income mobility in Switzerland using the Economic Well-Being of the Working and Retirement Age Population Data (WiSiER).

Data

WiSiER data include Swiss individuals from 11 Cantons from 1982 to 2016. The Swiss Federal Statistical Office merged data from tax records, social insurance, unemployment data, and surveys, creating a unique opportunity to analyze mobility. An ID can match parents and children. While many approaches seem feasible, we compare fathers and children at the same age of 35. As a result, the observations stem from different periods, with most of our successful matches coming from 1982-1990 (fathers) and 2000-2016 (children). The primary outcome variable is yearly real insured labor income (AHV) in $1,000$ Swiss francs (CHF) at the age of 35. The following covariates are available for both fathers and children: months experience, indicators for high-education (12 or more years of schooling), Swiss citizenship, and being single, and number of own children. Further, we include the fathers age at child's birth, and year and canton fixed effects for the children. Finally, for the analysis we exclude the following observations: (i) children where there is no parent in the data, (ii) observations with no information on the child's or father's birth year, and (iii) whenever the father was younger than 15 at the birth of the child. We conduct separate analyses for the relationships with sons and daughters. Table (ref) reports descriptive statistics for the data used in the analysis. It shows that father's characteristics are similar in families with sons and daughters. This alleviates a potential concern about endogenous selection in the comparison between sons and daughters.

table[table omitted — 1,347 chars of source]

Rank-Rank Regressions

Table (ref) reports the results of RRR and CRRR. The CRRR results are obtained using Algorithms (ref) and (ref) for the correlation-based estimator with a logistic link function and a mesh of 200 points located at sample quantiles in a sequence of orders from $0.01$ to $0.99$ with increments of $0.98/199$. We use linear interpolation to obtain estimates of the conditional ranks corresponding to intermediate points in the mesh. The standard errors (SE) and 95% confidence intervals (95% CI) are computed by empirical bootstrap with 500 repetitions. Based on the results of numerical simulations reported in Appendix (ref), we do not impose tail restrictions. In results not reported, we find very similar estimates, standard errors and confidence intervals for regression-based and fully restricted estimators.\footnote{These results are available from the authors upon request.} We show the robustness of the results to the choice of link function in Section (ref).

We find significant positive income persistence in both father-son and father-daughter relationships, with and without covariates. However, the persistence is much stronger for sons than for daughters suggesting the presence of a gender gap in intergenerational transmission of income even after controlling for the father's and child's characteristics. Comparing RRR and CRRR, we find that within-group persistence accounts for approximately 62% of the overall income persistence for sons and about 52% for daughters. These results highlight the substantial role of both within-group and between-group differences in explaining intergenerational mobility.

A subgroup analysis reveals relatively more mobility in families with a larger number of children and with a low educated father. In particular, we find relatively less persistence for sons in large families and more for daughters of high educated fathers. This would be consistent with decreasing returns of intergenerational transfers with respect to family size and increasing with respect to father's education. This heterogeneity, however, is not statistically significant. We do not find differences in intergenerational mobility for families with immigrant fathers in Switzerland, unlike the results of abramitzky2021intergenerational for the U.S. This difference might be due to the small fraction of immigrant fathers in the sample, see Table (ref).

table[table omitted — 1,151 chars of source]

Transition Matrices

Figures (ref) and (ref) show heatmaps of transition matrices for father-son and father-daughter, respectively. These matrices are a parsimonious representation of the joint distribution of income for father and child discretized in cells defined by deciles. They are commonly used in intergenerational mobility studies to provide a more granular measure of persistence than the rank-rank regressions. We report all the entries in percent deviations from $0.1$ because all the entries should be equal to $0.1$ under perfect mobility, that is when the income of the child is independent of the income of the father. Panels (A) report transition matrices based on marginal ranks, similar to previous studies. Panels (B) report conditional transition matrices based on conditional ranks, which are new to this paper and capture within-group dependence. For father-son, we find that the highest values in panel (A) are concentrated on the diagonal, which is consistent with the positive RRR estimate in Table (ref). The results in panel (B) show a less clear pattern once we control for covariates, consistent with the lower CRRR estimates in Table (ref). The results for father-daughter show similar but weaker patterns as we expect from the smaller correlation estimates in Table (ref). Interestingly, for both sons and daughters the highest probability occurs at the bottom right corner of the very top deciles conditionally and unconditionally.

figure[figure omitted — 655 chars of source]
figure[figure omitted — 662 chars of source]

Rank-Rank Regressions Excluding Child's Covariates

One concern about the CRRR results in Table (ref) is that the child's covariates might be picking up indirect sources of intergenerational mobility of income. For example, fathers might invest in child's education to increase the child's income prospects. To deal with this concern, Table (ref) reports CRRR results where the child's covariates, other than year and canton fixed effects, are excluded from the covariate set $X$. These results are obtained using the correlation-based estimator with a logistic link function with the same parameter choices as in Table (ref).

As expected, not accounting for the child's covariates increases the importance of within-group persistence with the estimates increasing to about 80% for father-son and 69% for father-daughter. In both cases the increase is about 17-18%. The other conclusions remain unchanged. In particular, we still find a significant gender gap in intergenerational transmission of income, and relatively less persistence for sons in large families and more for daughters of high educated fathers.

table[table omitted — 1,145 chars of source]

Robustness to Link Function

Table (ref) reports the results of CRRR using the correlation-based estimator with a Gaussian or probit link function. The estimates, standard errors and confidence intervals are almost identical to Table (ref) showing the robustness of the results to the use of the logistic versus Gaussian link functions.

table[table omitted — 1,151 chars of source]

Asymptotic Theory

In this section we provide asymptotic theory for the estimators of the CRRR slope $\rho_C$. We focus on the correlation-based and fully-restricted estimators of Algorithm (ref). We derive their asymptotic distributions by the delta method. For example, we take the following steps for the correlation-based estimator:

enumerate• Express the parameter $\rho_C$ as a correlation-based functional of the function-valued inputs $F_{Y \mid X}$, $F_{W \mid X}$ and $F_Z$, where $Z = (Y,W,X')'$. That is: \begin{equation} \rho_C = \phi(F_{Y\mid X}, F_{Y\mid X}, F_Z) := \frac{\int [F_{Y \mid X}(y\mid x) - .5] [F_{W \mid X}(w\mid x) - .5] \mathrm{d} F_{Z}(z)}{\sqrt{\int [F_{W \mid X}(w\mid x) - .5]^2 \mathrm{d} F_{Z}(z) \int [F_{Y \mid X}(y\mid x) - .5]^2 \mathrm{d} F_{Z}(z)}}. \end{equation} • Show that the plug-in estimator of $\rho_C$ using $\phi$, $\tilde \rho_C$, is a restricted correlation-based estimator: $$ \tilde \rho_C = \phi(\widehat F_{Y\mid X}, \widehat F_{W\mid X}, \widehat F_Z) = \frac{\sum_{i=1}^n (\widehat U_i - .5)(\widehat V_i - .5) }{\sqrt{\sum_{i=1}^n (\widehat V_i - .5)^2 \sum_{i=1}^n (\widehat U_i - .5)^2}}, $$ where $\widehat F_{Y \mid X}$ and $\widehat F_{W \mid X}$ are the DR estimators of $F_{Y \mid X}$, $F_{W \mid X}$ and $\widehat F_Z$ is the empirical distribution function of $Z$. • Establish that the map $\phi$ is Hadamard differentiable in the relevant functional spaces at $(F_{Y\mid X}, F_{Y\mid X}, F_Z)$ with the affine and continuous derivative operator: $$ (z_Y,z_W,g_Z) \mapsto\phi'_{F_{Y\mid X}, F_{Y\mid X}, F_Z}(z_Y,z_W,g_Z), $$ where $z_Y$, $z_W$ and $g_Z$ are the limits of converging deviations from $F_{Y\mid X}$, $F_{Y\mid X}$ and $F_Z$. • Apply the functional delta method to obtain the limit of $\tilde \rho_C$ from the limit of the deviations of the estimators of the inputs $\sqrt{n}(\widehat F_{Y\mid X}-F_{Y\mid X})$, $\sqrt{n}(\widehat F_{W\mid X}-F_{W\mid X})$, and $\sqrt{n}(\widehat F_Z-F_Z)$, $$ \sqrt{n}(\tilde \rho_C - \rho_C) \rightsquigarrow \phi'_{F_{Y \mid X},F_{W \mid X},F_{Z}}(Z_Y,Z_W,G_{Z}), $$ where $\rightsquigarrow$ and $(Z_Y,Z_W,G_{Z})$ are defined below. • Show that $\widehat \rho_C$ has the same asymptotic distribution as $\tilde \rho_C$, $$ \sqrt{n}(\widehat \rho_C - \tilde \rho_C) \to_P 0. $$

The distribution of the fully-restricted estimator is derived following steps (1)-(4), and replacing the correlation-based functional in step (1) by the fully-restricted functional:

equation[equation omitted — 197 chars of source]
remark[Regression-Based Estimators] The limit distribution of the regression-based estimators is derived following analogous steps to the correlation-based estimator, and replacing the correlation-based representation of the functional in step (1) by the regression-based representation: \begin{equation} \rho_C = \varphi(F_{Y\mid X}, F_{Y\mid X}, F_Z) := \frac{\int [F_{Y \mid X}(y\mid x) - .5] [F_{W \mid X}(w\mid x) - .5] \mathrm{d} F_{Z}(z)}{\int [F_{W \mid X}(w\mid x) - .5]^2 \mathrm{d} F_{Z}(z) }. \end{equation} We provide the corresponding results in Appendix (ref).

Before stating formally the main results, we review the existing theory for the estimator of the RRR slope. The purpose of this review is to explain why the existing results do not cover the estimators of the CRRR slope. hoeffding1948class first derived the asymptotic distribution of the RRR slope estimator using the theory of U-statistics. We cannot follow the same approach because none of our estimators has a U-statistic representation. ren1995hadamard alternatively derived the asymptotic distribution of the RRR slope estimator using the delta method. ren1995hadamard used analogous steps to our procedure described above. The following remarks explain each step of our procedure and point out the challenges and differences with respect to ren1995hadamard.

remark[Functional Representation of $\rho_C$] The functional and the inputs of the functional representation of the RRR slope, $\rho$, are different from $\rho_C$. In particular, ren1995hadamard showed that: $$ \rho = \tilde \phi(F_{Y,W}) = 12 \int [F_{Y,W}(y,+\infty) - .5] [F_{Y,W}(+\infty,w) - .5] \mathrm{d} F_{Y,W}(y,w). $$ The functional $\tilde \phi$ is an special case of the fully-restricted functional $\phi_1$ in (ref) where there are no covariates $X$. In the case of RRR, depite being a regression-based estimator, the denominator simplifies because the sample variances of the estimated marginal ranks are deterministic when $Y$ and $W$ are continuous.\footnote{Indeed, these sample variances are equal to $(n^2-1)/(12n^2)$, see ren1995hadamard; and the regression-based and correlation-based versions of the RRR estimator are numerically identical if there are no ties in the observations of $W$ and $Y$.} This simplification is not available for the correlation-based and regression-based estimators of $\rho_C$ because the sample variances of the estimated conditional ranks are random.
remark[Plug-in Estimator] The plug-in estimator of $\rho_C$ using $\phi$ is: $$ \phi(\widehat F_{Y\mid X}, \widehat F_{W\mid X}, \widehat F_Z) = \frac{\int [\Lambda(x'\widehat \beta_Y(y)) - .5] [\Lambda(x'\widehat \beta_W(w)) - .5] \mathrm{d} \widehat F_{Z}(z)}{\int [\Lambda(x'\widehat \beta_W(w)) - .5]^2 \mathrm{d} \widehat F_{Z}(z)} = \tilde \rho_C, $$ where the second equality follows from the properties of the empirical distribution function $\widehat F_Z$. This proves Step (2) above.
remark[Hadamard Differentiability of $\phi$] The argument to establish differentiability of the RRR functional $\tilde \phi$ does not apply to the CRRR functional $\phi$ for several reasons. First, the expression of $\phi$ is different from $\tilde \phi$. Second, the inputs and their estimators are also different. Moreover, the estimators of the inputs of the CRRR functional live in more complicated spaces than the estimators of the inputs of the RRR functional. Thus, while the estimator of $F_Z$ lives in the space of Cadlag functions, the estimators of $F_{Y \mid X}$ and $F_{W \mid X}$ live in the space of bounded functions, but have limits in the space of continuous functions, once properly recentered and rescaled. Because of this difference, we need to establish Hadamard differentiability in the space of bounded functions, tangentially to the space of continuous functions.
remark[Limit Process of Input Estimators] To apply the delta method, we need to characterize the limit process for the estimator of the inputs. This characterization is much more challenging for CRRR than RRR. Thus, for example, ren1995hadamard can rely on existing functional central limit theorems for the empirical distribution to establish the limit process over the entire support of $Y$ and $W$. Unfortunately, the existing functional central limit theorems for DR estimators of conditional distributions have only been established on a compact strict subset of the support of $Y$ and $W$; see, for example, chernozhukov+13inference. We deal with this challenge by imposing restrictions on the DR model of the conditional distributions at the tails. These restrictions allow us to extend the central limit theorems to the entire support of $Y$ and $W$. In numerical simulations, however, we find that estimators with and without imposing the tail restrictions perform similarly in terms of bias, standard deviation and root mean squared error.

We formally state now the main results from the steps (4) and (5). The result from step (3) is relegated to Appendix (ref) because it is of more technical nature. We state all the results for the logistic link function because it produces analytically simpler expressions, but it can be readily extended to the Gaussian link at the cost of more cumbersome notation.

We start by imposing some conditions on the DR model.

assumption[DR Model] For $R \in \{Y,W\}$: (a) The conditional distribution function takes the form $F_{R \mid X}(r \mid x)=\Lambda(x^{\prime}\beta_{R}(r))$ for all $r \in \mathcal{R}$ and $x \in \mathcal{X}$, where $\Lambda(u) = (1+\exp(-u))^{-1}$, the standard logistic distribution. (b) The support $\mathcal{R}$ is an open interval in $\mathbb{R}$ and the conditional density function $ f_{R \mid X}(r \mid x)$ exists and is positive in $(r,x)$ on $(R,X)$; it is uniformly bounded and uniformly continuous in $(r,x)$ on $(R,X)$. (c) $E\|X \|^{2} < \infty$ and the minimum eigenvalue of: \begin{equation*} J_R(r) := {\mathrm{E}} \left[ \lambda(X^{\prime}\beta_R(r)) X X^{\prime}\right] , \end{equation*} is bounded away from zero uniformly over $r \in \mathcal{R}$, where $\lambda = \Lambda(1-\Lambda)$ is the derivative of $\Lambda$. (d) Let $\beta_R(y)$ be partitioned as $(\beta_{R,1}(r),\beta_{R,-1}(r)')'$ where $\beta_{R,1}(r)$ is the intercept and $\beta_{R,-1}(r)$ includes the rest of the components. Then, for $r \in \mathcal{R}\setminus\bar{\mathcal{R}}$, where $\bar{\mathcal{R}}$ is a closed subinterval of the interior of $\mathcal{R}$, $\beta_{R,1}(r) = \beta_{R,1}(\bar r) + (r - \bar r)\alpha_{R}(\bar r)$ for $\bar r := \arg \min_{r' \in \bar{\mathcal{R}}} |r-r'|$ and some $\alpha_R(\bar r) > 0$, and $\beta_{R,-1}(r) = \beta_{R,-1}(\bar r)$.
remark[DR Model] The conditions in Assumption (ref)(a)-(c) are the same as in chernozhukov+13inference. They are used to obtain a functional central limit for the DR estimator of the conditional distribution $F_{R \mid X}(r \mid x)$ on $\bar{\mathcal{R}}$. Assumption (ref)(d) imposes restrictions on the tails that allow us to extend the functional central limit theorem to $\mathcal{R}$.

In order to state the result about the limit process for the inputs, we define, for $R \in \{Y,W\}$,

multline*[multline* omitted — 722 chars of source]

Consider the empirical processes $ (r,x) \mapsto \widehat Z_R(r,x) := \sqrt{n}\left(\widehat F_{R \mid X}(r \mid x) - F_{R \mid X}(r \mid x) \right)$, $R \in \{Y,W\}$, and $f \mapsto \widehat G_Z(f) := \sqrt{n} \int f \mathrm{d} (\widehat F_Z - F_Z),$ where $\widehat F_{R \mid X}(r \mid x) := \Lambda(x'\widehat \beta_R(r))$, $\widehat F_Z$ is the empirical distribution function of $Z = (Y,W,X)$, and $\mathcal{F}$ is a class of measurable functions that (i) includes $F_{Y \mid X}$, $F_{W \mid X}$, $F_{Y \mid X}^2$, $F_{W \mid X}^2$, $F_{Y \mid X} F_{W \mid X}$ and the indicators of all the rectangles in $\bar{\mathbb{R}}^{d_x+2}$, where $\overline{\mathbb{R}} := \mathbb{R} \cup \{-\infty, \infty\}$ is the extended real line, and (ii) is totally bounded under the metric: $$\lambda(f,\tilde f) = \left[\int (f-\tilde f)^2 \mathrm{d} F_Z \right]^{1/2}, \quad f,\tilde f \in \mathcal{F}.$$ Let $Z_n\rightsquigarrow Z $ in $\mathbb{E}$ denote weak convergence of a stochastic process $Z_n$ to a random element $Z$ in a normed space $\mathbb{E }$, as defined in van1996weak.

lemma[Limit Processses for Inputs] Assume that Assumption (ref) holds, the support $\mathcal{X}$ is a compact subset of $\mathbb{R}^{d_x}$, and $\{Z_i=(Y_i,W_i,X_i)\}_{i=1}^n$ is a random sample of $Z=(Y,W,X)$. Then, in the metric space $\ell^{\infty}(\mathcal{Y}\mathcal{W}\mathcal{X}\mathcal{F})$, $$ (\widehat Z_Y(y,x), \widehat Z_W(w,x), \widehat G_Z(f)) \rightsquigarrow (Z_Y(y,x), Z_W(w,x), G_Z(f)) $$ as stochastic processes indexed by $(y,w,x,f)$. The limit process is a zero-mean tight Gaussian process such that: \begin{equation*} Z_{R}(r,x) = \mathbb{G}( \ell_{r,x}), \ \ R \in \{Y,W\}, and \ G_Z(f) = {\mathbb{G}}(f), \end{equation*} where ${\mathbb{G}}$ is a $P$-Brownian bridge.

The next result states a central limit theorem for $\tilde \rho_C$ and $\breve \rho_C$, and the asymptotic equivalence between $\tilde \rho_C$ and $\widehat \rho_C$.

theorem[Limit Distribution of $\widehat \rho_C$, $\tilde \rho_C$ and $\breve \rho_C$] Under the conditions of Lemma (ref): (1) in $\mathbb{R}$, \begin{equation*} \sqrt{n} \left( \tilde \rho_{C} - \rho_{C} \right) \rightsquigarrow Z_{\rho} := 12 \ [Z_{1,\rho} - \rho_{C} (Z_{2,\rho} + Z_{3,\rho})/2] \ \ and \ \ \sqrt{n} \left( \breve \rho_{C} - \rho_{C} \right) \rightsquigarrow 12 Z_{1,\rho}, \end{equation*} where $Z_{1,\rho}$, $Z_{2,\rho}$ and $Z_{3,\rho}$ are zero-mean Gaussian random variables defined by \begin{multline*} Z_{1,\rho} := \int \left\{ Z_{Y}(y,x)[F_{W\mid X}(w\mid x) - .5] + Z_{W}(w,x)[F_{Y\mid X}(y\mid x) - .5] \right\} \mathrm{d} F_{Z}(y,w,z) \\ + G_Z\left([F_{Y\mid X} - .5][F_{W\mid X} - .5]\right), \end{multline*} \begin{equation*} Z_{2,\rho} := 2 \int Z_{W}(w,x)[F_{W\mid X}(w \mid x) - .5] \mathrm{d} F_{Z}(y,w,z) + G_Z\left([F_{W\mid X} - .5]^2\right), \end{equation*} and \begin{equation*} Z_{3,\rho} := 2 \int Z_{Y}(y,x)[F_{Y\mid X}(y \mid x) - .5] \mathrm{d} F_{Z}(y,w,z) + G_Z\left([F_{Y\mid X} - .5]^2\right). \end{equation*} (2) $\widehat \rho_{C}$ has the same limit distribution as $\tilde \rho_{C}$ because $$ \sqrt{n} \left( \widehat \rho_{C} - \tilde \rho_{C} \right) \to_{{\mathrm{P}}} 0. $$

The variance of the limit processes $Z_{1,\rho}$ and $Z_{\rho}$ have complicated expressions that might be difficult to estimate analytically. To avoid this difficulty, we propose the use of bootstrap to make inference. We show that the exchangeable bootstrap draws of Algorithm (ref) have the same asymptotic distribution as the CRRR estimators under the following assumption on the weights:

assumption[Exchangeable Bootstrap] For each $n$, $(\omega_{n1}, ..., \omega_{nn})$ is an exchangeable,\footnote{A sequence of random variables $X_1, X_2, ..., X_n$ is exchangeable if for any finite permutation $\sigma$ of the indices $1,2, ..., n$ the joint distribution of the permuted sequence $X_{\sigma(1)}, X_{\sigma(2)}, ...,X_{\sigma(n)} $ is the same as the joint distribution of the original sequence.} nonnegative random vector, which is independent of the data, such that for some $\epsilon> 0$ \begin{equation} \begin{split} \sup_{n} {\mathrm{E}}[\omega_{n1}^{2+\epsilon}] < \infty, \ \ n^{-1}\sum _{i=1}^{n} \left( \omega_{ni} - \bar{\omega}_n \right)^{2} \to_{{\mathrm{P}}} 1, \ \ \bar \omega_n \to_{{\mathrm{P}}} 1, \end{split} \end{equation} where $\bar \omega_n = n^{-1} \sum_{i=1}^{n} \omega_{ni} $.

In order to state the results about bootstrap validity formally, we follow the notation and definitions in van1996weak. Let $D_{n}$ denote the data vector and $M_{n}$ be the vector of random variables used to generate bootstrap draws given $D_{n}$. Consider the random element $ \mathbb{Z}^{*}_{n} = \mathbb{Z}_{n}(D_{n}, M_{n})$ in a normed space $ \mathbb{E}$. We say that the bootstrap law of $\mathbb{Z}^{*}_{n}$ consistently estimates the law of some tight random element $\mathbb{Z}$ and write $\mathbb{Z}^{*}_{n} \rightsquigarrow_{{\mathrm{P}}} \mathbb{Z} $ in $\mathbb{E}$ if:

equation[equation omitted — 227 chars of source]

where $\text{BL}_{1}(\mathbb{E})$ denotes the space of functions with Lipschitz norm at most 1 and ${\mathrm{E}}_{M_{n}}$ denotes the conditional expectation with respect to $M_{n}$ given the data $D_{n}$; and $ \rightarrow_{{\mathrm{P}}}$ denotes convergence in (outer) probability.

We now provide a bootstrap central limit theorem for the estimators of the CRRR slope. This result follows from a functional central limit theorem for the input processes, which we establish in Lemma (ref) in Appendix (ref), and the functional delta method for the bootstrap.

theorem[Exchangeable Bootstrap Consistency] Under the conditions of Lemma (ref) and Assumption (ref): in $\mathbb{R}$, $$\sqrt{n} \left( \tilde \rho^{*}_{C} - \tilde \rho_{C} \right) \rightsquigarrow _{{\mathrm{P}} } Z_{\rho}, \ \ \ \ \sqrt{n} \left( \widehat \rho^{*}_{C} - \widehat \rho_{C} \right) \rightsquigarrow _{{\mathrm{P}} } Z_{\rho}, \ \ \text{ and } \ \ \sqrt{n} \left( \breve \rho^{*}_{C} - \widehat \rho_{C} \right) \rightsquigarrow _{{\mathrm{P}} } Z_{1,\rho}$$ that is, exchangeable bootstrap consistently estimates the law of the limit processes $Z_{\rho}$ and $Z_{1,\rho}$. In particular, $$ \widehat \sigma_{\rho} \to_{{\mathrm{P}}} \sigma_{\rho} \ \ \text{ and } \ \ {\mathrm{P}}\left\{ \rho_C \in \textrm{ACI}_{1-\alpha}(\rho_{C}) \right\} \to 1-\alpha \ \ \text{ as } \ \ n \to \infty, $$ where $\sigma_{\rho}$ is the standard deviation of the limit process $Z_{\rho}$, and $\widehat \sigma_{\rho}$ and $\textrm{ACI}_{1-\alpha}(\rho_{C})$ are defined in Algorithm (ref).

conclusion

This paper introduces the conditional rank-rank regression (CRRR) as an alternative to traditional rank-rank regressions with covariates (RRRX) for measuring within-group mobility and persistence. The CRRR uses conditional ranks of the variables of interest given covariates, in contrast to RRRX which uses marginal ranks net of covariate effects. We show that the CRRR slope preserves an intuitive interpretation as the average conditional rank correlation between the variables, similar to RRR without covariates. In contrast, the slope of RRRX loses the rank correlation interpretation and can take on values outside the interval $[-1, 1].$ The CRRR is also suitable for subgroup analysis, where the CRRR slopes maintain a rank correlation interpretation conditional on the groups.

We propose a distribution regression estimator for CRRR where the conditional distributions are modeled flexibly using parametric link functions. The estimator is easy to implement and computationally tractable. We derive asymptotic theory for the estimator based on the functional delta method. The analytic asymptotic variance is cumbersome, so we propose an exchangeable bootstrap procedure for inference. The bootstrap procedure is also used to construct confidence intervals. We illustrate the usefulness of CRRR in an empirical application to intergenerational income mobility in Switzerland. The application reveals stronger intergenerational persistence between fathers and sons than fathers and daughters, where the within-group persistence accounts for between $52\%$ and $79\%$ of the overall persistence. We also find some evidence of heterogeneity across groups defined by father's education and family size. The results are robust to the exclusion of child's covariates and the use of logistic or Gaussian link functions.

In summary, CRRR provides a well-grounded measure of within-group mobility and persistence. The distribution regression estimator, coupled with exchangeable bootstrap inference, provides a practical and flexible way to implement CRRR in empirical applications. We expect CRRR will be a useful addition to the toolkit of methods for studying mobility and persistence. A natural next step is to consider high dimensional settings where they might be many covariates. An extension where we derive orthogonal moment conditions for CRRR to apply double/debiased machine learning (DML) is underway.