Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
102,733 characters · 21 sections · 54 citation commands
Conditional Rank-Rank Regression$^*$
The linear regression of the rank of a variable $Y$ on the rank of another variable $W$ is commonly referred to as a rank-rank regression (RRR). It has become increasingly popular in empirical investigations in economics to analyze policy relevant issues such as, for example, mobility and sorting behavior beller2006intergenerational,dahl2008association,chetty2014land,adermon2018intergenerational,murphy2020top.\footnote{chetverikov2025inferencerankrankregressions have recently documented that 40 articles published between January 2013 and February 2024 in American Economic Review, Journal of Political Economy, Quarterly Journal of Economics, and Review of Economic Studies employ RRRs in their empirical analysis.} A desirable feature of RRR is that its slope coefficient corresponds to the Spearman rank correlation coefficient between $Y$ and $W$, a classical measure of association with desirable properties such as invariance to monotone transformations of the variables and robustness to outliers and heavy tails spearman1904proof,kendall1948rank.\footnote{maasoumi2022generalized questioned the economic interpretation of the RRR as a measure of mobility due to the use of linear regression and proposed alternative measures based on nonparametric regression.}
We propose a general method to control for covariates $X$ in RRRs. Covariates are commonly accounted for in RRRs either via their inclusion as additional regressors (RRRX), or by conducting RRR on residuals after partialling out their influence (RRR.res). We show that these two approaches can have undesirable properties. In particular, their resulting coefficients can be difficult to interpret and produce counterintuitive results. For example, chetverikov2025inferencerankrankregressions noted that the slope of RRRX is no longer the Spearman correlation and might lie outside the interval $[-1,1]$. Moreover, the ranks of the residuals in RRR.res can only be related to meaningful ranks of the original variables under restrictive assumptions on the distribution of the variables of interest conditional on the covariates.\footnote{For example, the ranks of the residuals of the linear regression of $Y$ on $X$ correspond to the ranks of $Y$ conditional on $X$ under the location-shift model $Y = X'\beta + \varepsilon$ where $\varepsilon$ is independent of $X$. However this relationship generally does not hold otherwise.} We call our proposal conditional rank-rank regression (CRRR) because it employs conditional ranks. Intuitively, we replace the linear partialing out of the covariates implicit or explicit in the existing approaches by a nonlinear partialing out. We show that our approach always delivers coefficients that are easy to interpret and which do not have the undesirable properties of those from the existing approaches.
In its canonical form, CRRR regresses the rank of $Y$ conditional on $X$ on the rank of $W$ conditional on $X$. It provides an alternative to RRRX and RRR.res and produces estimates that are easy to interpret. Indeed, we show that the CRRR slope is equal to the Spearman correlation of $Y$ and $W$ conditional on $X$, averaged over the distribution of $X$, which is a natural summary measure of within-group association with similar properties to the Spearman rank correlation. In Statistics, this type of measure of association is so-called covariate-adjusted and it has been frequently used in biostatistic applications gijbels2011conditional,liu2018covariate,eden2022nonparametric,wei2023partial. Another attractive feature of CRRR is that it is suitable for subgroup analysis, i.e. to run separate CRRR by groups defined from a categorical variable. If this variable is included in $X$, the CRRR slope for each group has the interpretation of the average conditional Spearman rank correlation for that group.
CRRR is also attractive for economic applications. For example, hagedorn2017identifying used ranks of residuals of wages on worker and establishment characteristics to analyze labor market sorting. However, these ranks based on residuals are statistical constructions that, in general, do not have a natural interpretation in terms of the original wages. Our proposal is to conduct the analysis with conditional wage ranks, which reflect wage ranks within groups of workers with the same characteristics, and are accordingly closely related to the observed wages. CRRR can also be applied to the analysis of sorting in other matching settings such as marriage and worker-firm markets haltiwanger2015cyclical,guiso2022assortative,haner2024marry.
The interpretation of the CRRR slope is different from the RRR slope. Assume, for example, in an investigation of sorting behavior of married couples that $Y$ is husband's wage, $W$ is wife's wage and $X$ is number of children. The RRR slope without covariates is the correlation between husband's and wife's wage ranks where the ranks are relative to the entire wage distribution. The CRRR slope is the rank correlation where the ranks are relative to the wage distribution of those households with the same number of children. The slope of RRRX is the regression slope of the husband's wage rank on wife's wage rank, where the ranks are relative to the entire wage distribution and the wife's rank is centered to have the same mean within households with the same number of children. The slope of this regression might be difficult to interpret as the centered rank is no longer a rank. We believe that in many settings the CRRR slope better reflects the relationship researchers are intending to capture when they control for covariates as it is closer to the ceteris paribus effect that economic models typically predict. Mathematically, the difference between CRRR and RRRX is the order in the application of the rank and covariate partialing out operators. RRRX obtain ranks first and partials out the covariates second, whereas CRRR reverses the order. The final outcome differs across procedures as the two operators do not commute due to the nonlinearity of the rank operator.
CRRR is different from RRR because they have different dependent variables, conditional versus unconditional (marginal) ranks. In some instances the explicit objective is to predict the marginal rank of $Y$ using $W$ and $X$. For example, in intergenerational mobility applications one might wish to predict the child's income rank using father's income and family household characteristics such as father's immigration status abramitzky2021intergenerational. Here, the marginal rank of $Y$ (child's rank in the entire distribution) might be more interesting than the rank of $Y$ conditional on $X$ (child's rank among those who have the same father's immigration status indicator). The commonly employed approach in these instances is to include $X$ as a covariate in the RRR or to perform subgroup analysis for different values of $X$. We refer to both approaches as RRRX as the subgroup analysis can also be implemented using RRR by including suitable interactions. As noted above, chetverikov2023inference observed that the coefficient of the rank of $W$ in RRRX might be difficult to interpret. In Section (ref) we show through an example that both forms of RRRX can deliver counterintuitive economic results in terms of mobility and predict ranks of $Y$ outside $[0,1]$.\footnote{Another way to avoid this problem is to use a fractional response regression to model the ranks of $Y$ conditional on the ranks of $W$ and $X$ papke1996econometric. However, this approach requires parametric assumptions and the model coefficients lack a natural interpretation beyond their signs.} The underlying cause of these problems is that the rank of $W$ on the right hand side of the RRRX does not satisfy the properties of a rank after partialing out the effect of the covariates.
A variant of CRRR (CRRR') can be used in these situations. CRRR' regresses the marginal rank of $Y$ on the rank of $W$ conditional on $X$ and possibly $X$. Similar to CRRR, CRRR' bypasses the conceptual problems associated with RRRX. We show that the CRRR' slope corresponds to a correlation and therefore always lies between $-1$ and $1$. The square of this slope indeed corresponds to the partial R-squared of the conditional ranks of $W$. That is, it is the fraction of the variance of the marginal rank of $Y$ explained by the conditional rank of $W$, net of the impact of the covariates $X$. This measure is useful for performing variance decompositions. CRRR' is also suitable for subgroup analysis. The CRRR' slope indeed can be decomposed as the average of the CRRR' slopes in each group, weighted by the size of the group. Moreover, CRRR' predicts ranks in $[0,1]$. RRRX does not satisfy any of the previous properties. This is illustrated in the example presented in Section (ref). There are also more nuanced situations where we have two groups of covariates, $X=(X_1,X_2)$, and we are interested in the ranks of $Y$ conditional on $X_1$, but want to use $X_2$ to improve prediction or for some other purposes. For example, in an intergenerational mobility application, we might be interested in predicting regional child's income ranks using father's income and household characteristics such as father's immigration status. Our approach can also be applied to these settings. Here we would regress the ranks of $Y$ conditional on $X_1$ (region) on the ranks of $W$ conditional on $X$ (region and father's immigration status) and possibly $X$.
We provide an estimator of the CRRR coefficients based on distribution regression (DR). Here we do not distinguish between the different variations of CRRR because they all can be treated jointly by allowing different conditioning covariates or equivalently restricting some coefficients of the DR to be zero. Like the estimator of the RRR coefficients, our estimator consists of two steps. In the RRR, the first step estimates the marginal ranks of $Y$ and $W$ using the empirical distribution, and the second step runs the linear regression of the estimated ranks of $Y$ on the estimated ranks of $W$ or computes the sample correlation between these ranks. Both versions of the second step produce numerically identical results if there are no ties in the observed values of $Y$ and $W$. In CRRR, the first step estimates the conditional ranks by running logit or probit DRs of $Y$ on $X$ and $W$ on $X$ at multiple values of $Y$ and $W$ to trace the entire conditional distributions. The second steps are identical to RRR, but the linear regression and correlation versions are no longer numerically identical even if there are no ties. They are both, however, consistent estimators of the CRRR slope. The CRRR estimator is computationally tractable, albeit somewhat more demanding than RRR.
We derive the asymptotic distribution of the CRRR estimator and provide feasible inference theory. chetverikov2023inference noted that standard inference methods for linear regression do not apply to the RRR estimator because both the independent and dependent variables, namely the estimated ranks, are generated. The same problem applies to CRRR. The theory for the RRR estimator was derived using U-statistic theory hoeffding1948class or the delta method ren1995hadamard when $Y$ and $W$ are continuous. We employ the functional delta method approach to derive the theory because the CRRR estimator does not have a U-statistic structure. The application of the functional delta method to the CRRR estimator presents several differences with respect to RRR. The first ingredient of both approaches consists of writing the parameter of interest as a functional of inputs: the joint distribution of $Y$ and $W$ for RRR, or the conditional distributions of $Y$ and $W$ conditional on $X$ and the joint distribution of $Y$, $W$ and $X$, for CRRR. The CRRR functional is more complicated than the RRR functional and the inputs for CRRR live in more complex spaces than those for RRR. As a result, the Hadamard differentiability of the RRR functional established by ren1995hadamard does not cover the CRRR functional. We establish the Hadamard differentiabilty of the CRRR functional in the relevant spaces.
The second ingredient required for the application of the functional delta method is to establish functional central limit theorems for the estimators of the inputs. For RRR, this follows from the now classical large sample theory of the empirical distribution function. A challenge for CRRR is that existing theory for DR estimators of conditional distributions excludes the tails. In particular, the available functional central limit theorems only hold for compact strict subsets of the support. We deal with this problem by imposing assumptions on the tail behavior of the DR model that allows us to estimate the conditional distribution in the tails. We then obtain functional central limit theorems for DR estimators of conditional distributions over the entire support.
Combining these two ingredients we establish that the CRRR estimator follows a normal distribution around the CRRR slope in large samples via the delta method. The asymptotic variance has a complicated expression that might be difficult to estimate analytically. We develop the use of the exchangeable bootstrap to obtain standard errors and construct confidence intervals. Exchangeable bootstrap include the most common forms of bootstrap such as empirical, weighted and subsampling bootstrap as special cases. We establish its validity in large samples from the functional delta method for the bootstrap. Like hoeffding1948class and ren1995hadamard, our theory covers the case where $Y$ and $W$ are continuous. The theory can be extended to noncontinuous variables following the analysis of chetverikov2023inference for the RRR estimator. We leave this extension to future research.
We apply the CRRR estimator to analyze intergenerational income mobility in Switzerland using the Economic Well-Being of the Working and Retirement Age Population Data (WiSiER) from 1986 to 2016 for 11 cantons. This dataset contains rich information on socioeconomic, demographic and other variables merged from tax records, social insurance, unemployment records and surveys, and can be linked for fathers and children. We uncover a gender gap in intergenerational income mobility as the persistence between fathers and sons is stronger than between fathers and daughters, both with and without controlling for covariates. We also find that about $62\%$ and $52\%$ of the overall (unconditional) income persistence is explained by the within-group income persistence for sons and daughters, respectively; where groups are defined by child's and father's marital status, Swiss citizenship, high school graduation, experience, number of children and canton and year fixed effects. We also provide evidence supporting greater persistence for fathers with higher education and fathers with only one child. These results uncover the substantial role of both within-group and between-group persistence in explaining the intergenerational transmission of income.
Our methodology complements related methodological developments in the econometric and statistical literature. While chetverikov2023inference provides the inferential theory for RRR and RRRX with marginal ranks, their approach does not apply to conditional ranks employed in CRRR, as we explained earlier. kitagawa2018measurement studied measurement error in RRRs in the context of intergenerational income mobility. Another new development is lei2024causal which examines causal underpinnings of the marginal rank regressions. More closely related to our work, liu2018covariate introduced a covariate-adjusted Spearman coefficient based on the probability scale residuals of li2012new and shepherd2016probability, which we further discuss in Section (ref). They proposed a modelling strategy and an estimator based on a monotone transformation of a location-shift model, which is a special case of the distribution regression model chernozhukov+13inference.\footnote{liu2018covariate did not develop inference theory for their estimator in the case where $Y$ and $W$ are continuous, although they conjectured that it is possible; see Section 5 ibid.} The DR approach is more flexible and comprehensive, in that it can approximate the true conditional distribution function arbitrarily well by considering rich sets of basis functions with respect to the covariates. This is not generally possible with transformations of location models.\footnote{A transformation of a location-shift model takes the form $H(Y) = b(X)'\beta + \epsilon$, where $\epsilon$ is independent of $X$ and has a known distribution, $b(X)$ is a basis of functions of $X$, and $H$ is an unknown monotone transformation function. The model is more flexible than just a location model, but it does not allow the covariates to affect the distribution of $H(Y)$ other than through its location. No matter how rich the basis functions $b(X)$ are, the model is not guaranteed to cover the true conditional distribution, even in the limit where the dimension of $b(X)$ grows large.} gijbels2011conditional and veraverbeke2011estimation developed estimators of conditional measures of association using copulas, relying on the representation of these measures in terms of the conditional copula for continuous outcomes. They derived distribution theory for estimators of conditional versions of Kendall's tau and Blomqvist's beta blomqvist1950measure with scalar covariates. In contrast, our modelling and estimators are based on conditional distributions and apply to multivariate covariates.
\paragraph{Outline} The rest of the paper is organized as follows. Section (ref) contrasts CRRR with RRR via a simple conceptual example. Section (ref) introduces formally CRRR and derives its properties. Section (ref) describes the estimation procedure based on DR and a bootstrap algorithm to make inference. Section (ref) discusses an application examining the relationship between fathers' and their children's labor income using Swiss data, while Section (ref) provides asymptotic theory. Additional theoretical and numerical simulation results, and proofs are reported in the Appendix.
To contrast the relative performances of CRRR and different forms of RRR in capturing the relationship between $Y$ and $W$ in the presence of covariates $X$, we examine a simple conceptual example in which $X$ is binary. This example is convenient because with a binary covariate there is no concern that the differences between the methods are driven by particular modeling strategies to specify various regression functions.
Let $Y$ be daughter's height (in cm), $W$ be father's height (in cm) and $X$ be a country indicator, say $X=0$ for the Netherlands and $X=1$ for Ireland. Conditional on $X$, $Y$ and $W$ follow a bivariate normal distribution with mean parameters that may depend on the value of $X$ and constant covariance matrix. More specifically,
where ${\mathrm{P}}(X=0) = {\mathrm{P}}(X=1) = 1/2$. For example, $\delta$ can be a negative country shock such as the Irish Famine that affects the father's height in Ireland, but not in the Netherlands. We consider two cases depending on the extent of the effect of the shock as measured by $\delta$:
Table (ref) compares measures of overall and within-country intergenerational height persistence based on rank correlation with the estimands of RRR, CRRR and two versions of RRRX. Thus, $\rho_{Y,W}$ and $\bar \rho_{Y,W \mid X}$ are the Spearman rank correlation coefficient between $Y$ and $W$ and the expected Spearman rank correlation between $Y$ and $W$ conditional on $X$; RRR is the RRR slope; CRRR is the CRRR slope; RRRX-A is the slope of the RRRX, that is, $\beta_1$ in $$ \tilde U = \beta_0 + \beta_1 \tilde V + \beta_2 X + \epsilon, \quad {\mathrm{E}}[\epsilon] = {\mathrm{E}}[\tilde V \epsilon] = {\mathrm{E}}[X \epsilon] = 0, $$ where $\tilde U$ and $\tilde V$ denote the marginal ranks of $Y$ and $W$, respectively; and RRRX-I is the average slope of the RRRs run separately by the values of $X$, that is, $\beta_1$ in $$ \tilde U = \beta_0 + \beta_1 \tilde V + \beta_2 [X-.5] + \beta_3 [X-.5] \tilde V + \epsilon, \quad {\mathrm{E}}[\epsilon] = {\mathrm{E}}[\tilde V \epsilon] = {\mathrm{E}}[X \epsilon] = {\mathrm{E}}[X \tilde V \epsilon] = 0 $$
We find that all the methods give the same answer when $\delta=0$, that is when the joint distribution of daughter's and father's heights is the same in both countries.\footnote{Up to numerical error, all the slopes are equal to the rank correlation of the bivariate normal with correlation $c=.6$, $\rho_S(Y,W)=6\arcsin{(c/2)}/\pi = .58$ cramer1999mathematical.} When $\delta=12$, RRR gives the overall mobility $\rho_{Y,W}$, whereas CRRR gives the within-country mobility $\bar \rho_{Y,W \mid X}$. Both forms of RRRX produce measures greater than one, which do not correspond to any rank correlation and might be difficult to interpret. Whether RRR or CRRR is the right measure depends on the application. In this case, CRRR measures average intergenerational mobility within each country whereas RRR measures intergenerational mobility pooling the two countries. They would lead to different conclusions about the effect of the Irish Famine. According to RRR, the famine reduces overall height persistence, whereas it does not have any effect within each country according to CRRR. Both versions of RRRX lead to the opposite conclusion that the famine increases height persistence. This conclusion does not correspond to a change in either overall or within-country mobility as measured by Spearman correlations.
While RRR and RRRX are frequently employed in intergenerational mobility analyses, the rank correlation is not always the object of interest. Many empirical studies in this literature focus on the so-called level of absolute upward mobility. This is defined as the expected marginal rank of a child with a father at a specified percentile, typically the 25th percentile chetty2014land. More generally, it is the conditional expectation function (CEF) of the child's rank given values for the father's rank and the covariates. Figures (ref) and (ref) report predicted daughter's height (marginal) ranks obtained from CRRR', the variant of CRRR that regresses marginal ranks of $Y$ on conditional ranks of $W$, and different versions of RRR for the scenarios $\delta=0$ and $\delta=12$, respectively. In this case CRRR' is the same as CRRR as the marginal and conditional ranks of $Y$ coincide because $Y$ is independent of $X$.\footnote{We also examined a variant of the data generating process of the model in which the conditional mean of $Y$ was affected by a different binary variable such that CRRR' was different from CRRR. It did not produce any qualitative differences from the conclusions that follow.} Accordingly, we shall refer to CRRR' as CRRR. RRR only uses the father's height $W$, whereas RRRX and CRRR also employ the value of the covariate $X$.\footnote{We can also include $X$ as an additional regressor in CRRR, but in this case does not change the results because $Y$ is independent of $X$ conditional on the conditional rank of $W$.} We do not distinguish between RRRX-A and RRRX-I because they produce very similar predictions. To make the methods comparable, CRRR evaluates the prediction at the father's conditional rank corresponding to the marginal quantile of the father's (marginal) rank.\footnote{More specifically, if the father's marginal rank is $v$ and $X=x$, then the corresponding father's conditional rank is $V_{w,x} = F_{W \mid X}(F_W^{-1}(v) \mid x)$, where $F_W$ and $F_{W \mid X}$ are the marginal and conditional distributions of $W$.} As a benchmark of comparison, we report the conditional expectation function (CEF) of the marginal rank of $Y$ given $X$ and $W$ which, using the properties of the multivariate normal, is:
where $\Phi$ is the distribution of the standard normal and $V_{w,x}$ is the conditional rank of $W$ at $W=w$ and $X=x$.\footnote{Note that it is equivalent conditioning on $W$ to conditioning on the marginal rank of $W$ because there is a one-to-one relationship between them. Indeed, $$ {\mathrm{E}}(\tilde U \mid \tilde V = \tilde v, X=x) = \Phi \left( \frac{.6 *\left( F^{-1}_W(\tilde v) - 180 + \delta x\right) }{4 \sqrt{2 - 0.6^2}} \right), $$ where $F_W$ is the distribution of $W$. } The true CEF can be computed in this example because we know the joint distribution of $(Y,W)$ conditional on $X$.
When $\delta =0$, $X$ does not help predict because $(Y,W)$ and $X$ are independent. All the methods yield the same prediction function in both countries, which is a linear approximation to the CEF. When $\delta = 12$, $X$ helps predict because the CEF is different in both countries. In this case, RRRX and CRRR give better approximations to the CEF than RRR, because they make use of the information in $X$. RRRX provides a linear approximation to the CEF in each country, but delivers predicted daughter's ranks lower than $0$ for father's ranks below the 28th percentile when $X=0$ (Netherlands), and greater than $1$ for father's ranks above the 72th percentile when $X=1$ (Ireland). RRRX therefore produces unreasonable predictions for many relevant values of the father's rank in half of the population. CRRR does not suffer from this problem. The implicit nonlinearity introduced by the use of the conditional father's rank, instead of the marginal rank, enforces that all the predicted daughter's ranks are between $0$ and $1$. In other words, while CRRR provides a linear approximation to the CEF with respect to the conditional father's rank, it provides a nonlinear approximation in terms of marginal father's rank. This approximation is bounded in $[0,1]$ because the absolute value of the slope is less than $1$. There is still approximation error, however, because the CEF is also a nonlinear function of the conditional father's rank in this case; see (ref).
When $\delta=12$, the CEF is convex when $X=0$ (Netherlands) and concave when $X=1$ (Ireland). Intuitively, a low marginal father's rank corresponds to a very low conditional father's rank in the Netherlands, which is associated with a very low conditional and marginal daughter's rank due to the high positive within-country correlation. Conversely, a high marginal father's rank corresponds to a very high conditional father's rank in Ireland, which is associated with a very high conditional and marginal daughter's rank. The CRRR predictions agree with these shape restrictions, whereas RRRX imposes linearity by construction. In both scenarios for $\delta$, the CRRR slope of $0.58$ indicates that about $34\% (=0.58^2)$ of the variability of the daughter's rank is explained by the father's rank net of covariates (country indicator), which is also the R-squared of CRRR. In this case, this explained fraction is the same in both countries. The RRRX slope of $1.07$ when $\delta=12$ does not have an interpretation as either a partial or overall R-squared.
Finally, we conduct a subgroup analysis using CRRR and RRR. Table (ref) compares the conditional Spearman rank correlation, CRRR slope and RRR slope for each value of $X$. CRRR produces measures that are invariant to both $X$ and $\delta$, which correspond to the rank correlations between $Y$ and $W$ conditional on $X$. If $Y$ and $X$ were not independent, then CRRR and CRRR' would give different results, but still both would produce slopes for each country that would average to the overall CRRR or CRRR' slope.\footnote{In the case of CRRR, the slopes for each country would also correspond to conditional Spearman rank correlations.} The RRR slopes are the same as the CRRR slopes when $X$ is irrelevant. RRR, however, delivers different slopes for the different values of $X$ when $\delta=12$, and also across the different values of $\delta$. The RRR slopes are greater than one when $\delta=12$, confirming that they do not correspond to correlations and making them hard to interpret. Moreover, the within-group RRR slopes do not average to the overall RRR slope in general hertz2008group.
While the above example is simple, it illustrates a number of important features which apply, or are likely to apply, in more complicated settings. First, in the presence of covariates, the RRRX slope does not correspond to any measure of rank correlation and can be difficult to interpret. In contrast, the CRRR slope corresponds to a specific form of rank correlation. Moreover, while neither CRRR or RRRX exactly predict the CEF for all values of the father's rank, the results here suggest that CRRR provides a more sensible and accurate approximation. Moreover, it directly produces a measure of the proportion of variability in the ranks of $Y$ explained by the ranks of $W$ net of the covariate $X$ and is suitable for subgroup analysis.
Let $(Y,W)$ be a bivariate random variable with joint distribution $F_{Y,W}$ and marginal distributions $F_Y$ and $F_W$ for $Y$ and $W$, respectively. For example, $Y$ is child's income and $W$ is father's income. We assume that $Y$ and $W$ are continuous.
We start by reviewing the canonical rank-rank regression (RRR). Let $\tilde U := F_Y(Y)$ and $\tilde V:=F_W(W)$ denote the (marginal) ranks of $Y$ and $W$. By continuity of $Y$ and $W$, ranks are uniformly distributed, $\tilde U \sim U(0,1)$ and $\tilde V \sim U(0,1)$. The RRR of $Y$ on $W$ is defined as the correlation between $\tilde U$ and $\tilde V$ or the slope of the linear regression of $\tilde U$ on $\tilde V$ (or vice versa): $$ \rho := \mathrm{Cor}(\tilde U, \tilde V)= \frac{\operatorname{Cov}(\tilde U,\tilde V)}{\operatorname{Var}(\tilde U)} = \frac{\operatorname{Cov}(\tilde U,\tilde V)}{\operatorname{Var}(\tilde V)} = 12 \ {\mathrm{E}}[(\tilde U-.5)(\tilde V - .5)], $$ where all the equalities follow from the uniform distribution of $\tilde U$ and $\tilde V$. In statistics this correlation measure is the celebrated Spearman rank correlation between $Y$ and $W$, and is widely used to measure dependence between variables. It is invariant to rescaling and all increasing monotone transformations of the variables, and has gained prominence for that reason. The rank correlation has become popular in economics in studies of income and wealth mobility due to its interpretability as a measure of persistence and scale-free nature.
We introduce now the conditional rank-rank regression (CRRR). Let $X$ denote a vector of covariates related to $Y$ and $W$ including, for example, child's and father's education, age, marital status and nationality. Let $F_{Y\mid X}$ and $F_{W \mid X}$ denote the distributions of $Y$ and $W$ conditional on $X$. Then, $U := F_{Y \mid X}(Y \mid X)$ and $V:=F_{W \mid X}(W \mid X)$ are the conditional ranks of $Y$ and $W$, where conditioning is on $X$. For example, $U$ and $V$ would be child's and father's income ranks among families with the same composition in terms of covariates. By continuity of $Y$ and $W$, the conditional ranks follow the uniform distribution, conditional on $X$: $$U \mid X \sim U(0,1) \text{ and } V \mid X \sim U(0,1),$$ and also unconditionally. This implies the constant variance property, $$\operatorname{Var}(V) = \operatorname{Var}(U) =\operatorname{Var}(V \mid X) = \operatorname{Var}(U \mid X) = 1/12$$ and the constant mean property, $${\mathrm{E}} V = {\mathrm{E}} U = {\mathrm{E}}(V \mid X) = {\mathrm{E}} (U \mid X) =.5.$$ Note that both $U$ and $V$ are marginally independent of $X$ , but not necessarily jointly independent so the correlation between $U$ and $V$ can depend on $X$.\footnote{That is, $U \mathop{\perp\!\!\!\!\perp} X$ and $V \mathop{\perp\!\!\!\!\perp} X$, but generally $(U,V) \not\perp\!\!\!\!\perp X$, where $\mathop{\perp\!\!\!\!\perp}$ denotes stochastic independence.}
The CRRR of $Y$ on $W$ given $X$ is defined as either the correlation between $U$ and $V$ or the slope of the linear regression of $U$ on $V$ (or vice versa):
CRRR is the average conditional correlation between conditional ranks:
where $\rho_{Y,W\mid X}$ denotes the conditional Spearman rank correlation between $Y$ and $W$ conditional on $X$, which is equal to $\mathrm{Cor}(U,V\mid X)$ by definition. Equation (ref) follows from $\operatorname{Cov}(U,V) = {\mathrm{E}} [\operatorname{Cov}(U,V\mid X)]$ by the law of total covariance since $\operatorname{Cov}[{\mathrm{E}}(U\mid X), {\mathrm{E}}(V\mid X)] = 0$; moreover, the conditional variance of $U$ and $V$ is equal to the unconditional variance. In summary, CRRR is the average Spearman rank correlation between $Y$ and $W$ conditional on $X$, averaged over the distribution of $X$, which is a summary measure of within-group persistence.
By the properties of $U$ and $V$, the CRRR can also be represented as the rescaled covariance of conditional ranks:
a formula convenient for estimation. Moreover, in the regression version of the CRRR, the intercept, $\alpha_C$, is mechanically related to the slope, $\rho_C$, through
This relationship raises concerns about the interpretation of intercepts and slopes in the regression versions of the RRR used in intergenerational mobility studies as measures of absolute and relative mobility, repectively.\footnote{chetty2014land noted a relationship analogous to (ref) for the regression version of the canonical RRR.}
Finally, we note that correlation of conditional ranks is generally not equal to correlation of marginal (unconditional) ranks: $$ \rho_C \neq \rho $$ but the two agree under independence from $X$, namely $\rho_C = \rho$ if $Y \mathop{\perp\!\!\!\!\perp} X$ and $W \mathop{\perp\!\!\!\!\perp} X$, because in that case $U = \tilde U$ and $V = \tilde V$.
In the context of the income mobility application, $\rho_C$ measures within-group income persistence and $\rho$ measures overall income persistence, encompassing both within-group and between-group persistence. The between-group persistence can then be defined as the difference between the marginal rank and conditional rank correlations: $$\textrm{Between-group persistence} = \rho - \rho_C.$$ Assume, for example, that the covariates $X$ capture family characteristics such as size or parental education. The difference between the two measures can be explained as follows: The within-group or unexplained persistence $\rho_C$ captures the extent to which father's income rank facilitates child's income rank among families with the same observable characteristics. In other words, it measures the influence of father's income on child's income, where the variation in father's and child's incomes comes from unobserved characteristics such as family status, ability and the extent of social or professional networks. On the other hand, the between-group measure $\rho - \rho_C$ aims to capture the contribution of observed characteristics to income persistence.
We can further decompose the between-group persistence using the total law of covariance: $$ \rho - \rho_C = 12 \operatorname{Cov}[{\mathrm{E}}(\tilde U \mid X),{\mathrm{E}}(\tilde V \mid X)] + 12 {\mathrm{E}}[\operatorname{Cov}(\tilde U,\tilde V \mid X) - \operatorname{Cov}(U,V \mid X)], $$ where the first component is the covariance of conditional means of marginal ranks, and the second component is the average conditional covariance of marginal ranks net of the average within-group inequality.
CRRR is different from RRR with covariates $X$ (RRRX) where $X$ is included additively (or non-additively) in the regression of marginal ranks, $\tilde U$ on $\tilde V$. We believe that our proposal is a more natural and adequate way to incorporate covariates. In fact, RRRX with additive covariates is no longer related to a rank correlation nor has to lie in the interval $[-1,1]$. RRRX is also more difficult to interpret as it does not correspond to a meaningful measure of within-group persistence. Making RRRX more flexible by including interactions between $X$ and $\tilde V$ does not mitigate any of these problems. In fact, making RRRX fully nonparametric also does not alleviate the problem. We show in the next section that even in the simplest case where $X$ is binary, the nonparametric RRRX does not capture meaningful economic quantities. When $X$ is discrete, the nonparametric approach (tabulating unconditional rank correlation by subgroups) does not either.
In what follows, we systematically explain the current approaches to RRRX and contrast these with the CRRR approach. We use the intergenerational income application to give context to the discussion.
When $X$ is discrete, it is common to run RRRs separately for each value of $X$ instead of including $X$ as an additive control. For example, abramitzky2021intergenerational run separate RRR of child's income on father's income by father's immigration status. The slopes of these regressions cannot be interpreted in terms of rank correlations or even as conditional correlations between the marginal ranks. To see this, note that the slope of the regression of $\tilde U$ on $\tilde V$ conditional on $X=x$, is not equal to conditional correlation of $\tilde U$ and $\tilde V$: $$ \frac{\operatorname{Cov}(\tilde U,\tilde V \mid X=x)}{\operatorname{Var}(\tilde V \mid X=x)} \neq \frac{\operatorname{Cov}(\tilde U,\tilde V \mid X=x)}{\sqrt{\operatorname{Var}(\tilde V \mid X=x) \operatorname{Var}(\tilde U \mid X=x)}}, $$ because marginal ranks have different conditional distributions, i.e. $\tilde U \overset{d}{\not\sim} \tilde V \mid X=x$, in general. The slope therefore does not generally correspond with the conditional correlation of the marginal ranks conditional nor the conditional rank correlation between $Y$ and $W$. We give an example in Section (ref) where this slope is greater than one.
Consider now the CRRR. Assume we are interested in conducting a subgroup analysis of intergenerational mobility with respect to father's high school diploma or immigration status. Let $X_1 \subseteq X$ be a set of variables that define the subpopulation of interest such as an indicator for high school diploma and/or Swiss nationality. Then, the CRRR slope conditional on $X_1=x_1$ is:\footnote{Indeed, by the law of total covariance with respect to $X$ and uniformity of $U$ and $V$ conditional on $X$, $$ \operatorname{Cov}(U,V \mid X_1) = {\mathrm{E}}[\operatorname{Cov}(U,V \mid X) \mid X_1] + \operatorname{Cov}[{\mathrm{E}}(U\mid X),{\mathrm{E}}(V\mid X) \mid X_1) = {\mathrm{E}}[\operatorname{Cov}(U,V \mid X) \mid X_1], $$ and $$ \operatorname{Var}(V \mid X_1) = \operatorname{Var}(V \mid X) = \operatorname{Var}(U \mid X), $$ almost surely.}
Hence, the CRRR slope for the subgroup defined by $X_1=x_1$ corresponds to the average conditional rank correlation between $Y$ and $W$, where the average is taken with respect to the distribution of $X$ conditional on $X_1=x_1$. This allow us, for example, to measure intergenerational mobility separately for families with fathers with and without high school diploma.\footnote{If $X_1 \not\subseteq X$, , the slope no longer has an interpretation as average conditional rank correlation because $V \not\sim U \mid X_1$ in general.}
There are applications where the researcher might want to use different sets of covariates to obtain the conditional ranks $U$ and $V$. In the intergenerational mobility application, for example, we might not want to control for son's education to obtain the father's income rank. In this case the CRRR' slope still corresponds to an average correlation between the ranks. To see this, let $U=F_{Y \mid X_1}(Y \mid X_1)$ and $V=F_{W \mid X_2}(W \mid X_2)$ with $X_1 \neq X_2$ and $X= X_1 \cap X_2$, the set of covariates included in both $X_1$ and $X_2$, then: $$ \rho_C = \frac{\operatorname{Cov}(U,V)}{\operatorname{Var}(V)} = {\mathrm{E}}\left[\frac{\operatorname{Cov}(U,V \mid X) }{\sqrt{\operatorname{Var}(V \mid X)\operatorname{Var}(U \mid X)}}\right], $$ where we use the law of total covariance with respect to $X$, $U \mathop{\perp\!\!\!\!\perp} X$, $V \mathop{\perp\!\!\!\!\perp} X$ and iterated expectations. The CRRR' slope therefore corresponds to the correlation between the ranks $U$ and $V$ conditional on the common covariates $X$, averaged over the distribution of $X$. Note, however, that $\rho_C$ in this case does not correspond to an average conditional rank correlation between $Y$ and $W$. The source of the difference is that $U \neq F_{Y \mid X}(Y \mid X)$ and $V \neq F_{W \mid X}(W \mid X)$ in general.\footnote{This rank correlation can be obtained by constructing the conditional ranks as $U=F_{Y \mid X}(Y \mid X)$ and $V=F_{W \mid X}(W \mid X)$.} One exception occurs when $Y$ is independent of the components of $X_2$ not included in $X_1$ conditional on $X_1$, and $W$ is independent of the components of $X_1$ not included in $X_2$ conditional on $X_2$. In that case, $$ \rho_C = {\mathrm{E}}\left[\frac{\operatorname{Cov}(U,V \mid \bar X) }{\sqrt{\operatorname{Var}(V \mid \bar X)\operatorname{Var}(U \mid \bar X)}}\right] = {\mathrm{E}} [ \rho_{Y,W \mid \bar X} ], $$ where $\bar X = X_1 \cup X_2$. This result follows by the law of total covariance with respect to $\bar X$ and uniformity of $V$ and $U$ conditional on $\bar X$.\footnote{Note that if $V \mid X_1 \sim U(0,1)$ and $V \mathop{\perp\!\!\!\!\perp} X_2 \mid X_1$, then $V \mid X \sim U(0,1)$.}
Like CRRR, CRRR' is suitable for subgroup analysis in the following sense. Let $\rho_C(x_2)$ be the CRRR' slope in the group defined by $X_2=x_2$, that is $$ \rho_C(x_2) = \frac{\operatorname{Cov}(U,V \mid X_2 = x_2)}{\operatorname{Var}(V \mid X_2 = x_2)}. $$ Then, by the law of total covariance and $V \perp\!\!\!\perp X_2$, $$ \rho_C = \frac{{\mathrm{E}}\left[\operatorname{Cov}(U,V \mid X_2)\right]}{\operatorname{Var}(V)} = {\mathrm{E}}[\rho_C(X_2)], $$ that is, the CRRR' slope can be decomposed as the average of the CRRR' slopes in each group defined by $X_2$, weighted by the size of the group.
An interesting example occurs when $X_1 = \emptyset$ and $X_2 = X$. In this case, $U = \tilde U$ and $\rho_C$ is the correlation between the marginal ranks of $Y$ and the conditional ranks of $W$. Unlike the RRR, the inclusion of covariates in this CRRR does not affect the coefficient of $V$ because $V \mathop{\perp\!\!\!\!\perp} X$, and can be used to perform a variance decomposition of $\tilde U$. Let, $$ \tilde{U} = \rho_C V + X'\beta_C + \varepsilon, \quad {\mathrm{E}}[(V; X) \varepsilon] = 0, $$ be the extended CRRR with covariates, where the first term of $X$ is a constant. Then, $$ \text{Var}(\tilde U) = \rho_C^2 \text{Var}(V) + \text{Var}(X'\beta_C) + \text{Var}(\varepsilon), $$ where the first two terms of the right-hand-side correspond to the contribution of $V$ and $X$ to the variance of $\tilde U$, and the third terms to the unobserved component. Indeed, $\rho_C^2$ measures the fraction of the variance of $\tilde U$ explained by $V$ since $\text{Var}(\tilde U) = \text{Var}(V)$.
We conclude this section by gathering the properties of the CRRR slope in the following lemma.
For estimation purposes, it is convenient to model the conditional distributions $F_{Y\mid X}$ and $F_{W \mid X}$ using the distribution regression (DR) model: $$ F_{R \mid X}(r \mid x) = \Lambda(x'\beta_R(r)), \quad R \in \{Y,W\}, \quad r \in \mathcal{R}, $$ where $\Lambda$ is the standard normal or logistic distribution, $\mathcal{R}$ is the support of $R$ and the first component of $x$ is a constant. The specification can be made more flexible by replacing $x$ by a vector of transformations of $x$ with good approximating properties.
As the data to estimate the conditional distribution function at the tails are sparse, it is necessary to impose some structure. We assume that the conditional distribution far in the tails can be extrapolated from the conditional distribution not too far in the tails.\footnote{This is in line with approaches used in extreme value theory that impose restrictions on the tail behavior allowing similar extrapolations. For example, see embrechts:1997 for a broad reference on the theory of extremes and victor:annals or chernozhukov:2011 for similar approaches in the context of extremal quantile regression.} We formalize this approach by imposing restrictions on the coefficient of the DR model in the tails.
Let $\bar{\mathcal{R}}$ be a compact strict subset of $\mathcal{R}$, for $\mathcal{R} \in \{\mathcal{Y},\mathcal{W}\}$, where $\mathcal{Y}$ and $\mathcal{W}$ are the supports of $Y$ and $W$, respectively. Then, we assume: $$ F_{R\mid X}(r \mid x) = \Lambda((r-\bar r)\alpha_R(\bar r) + x'\beta_R(\bar{r})), \quad R \in \{Y,W\}, \quad r \in \mathcal{R}\setminus\bar{\mathcal{R}}, $$ where $\bar r := \arg \min_{r' \in \bar{\mathcal{R}}} |r-r'|$ and $\alpha_R(\bar r) > 0$. That is, we postulate that the random variable $R$ behaves in the tails like a random variable with distribution $\Lambda$, after subtracting the location shift $x'\beta_R(\bar{r})$ and dividing by the scale $\alpha_R(\bar r)$, which are different at the upper and lower tails. Thus, the DR coefficient is restricted in the tails by: $$\beta_{R,1}(r) = \beta_{R,1}(\bar r) + (r - \bar r)\alpha_R(\bar r), \quad \beta_{R,-1}(r) = \beta_{R,-1}(\bar r),\quad R \in \{Y,W\}, \quad r \in \mathcal{R}\setminus\bar{\mathcal{R}},$$ where $\beta_R(r)$ is partitioned into $(\beta_{R,1}(r),\beta_{R,-1}(r)')'$ where $\beta_{R,1}(r)$ is the intercept and $\beta_{R,-1}(r)$ are the slope components. That is, $r \mapsto \beta_{R,1}(r)$ is a linear function and $r \mapsto \beta_{R,-1}(r) $ is constant on $\mathcal{R}\setminus\bar{\mathcal{R}}$.
Under the DR model, the conditional ranks can be expressed as the following functionals of the parameters: $$ U = \Lambda(X'\beta_Y(Y)), \quad V = \Lambda(X'\beta_W(W)). $$
We provide several estimators of the CRRR slope based on the different representations of $\rho_C$ in (ref) and (ref). This section presents correlation-based and fully-restricted estimators. Regression-based estimators are given in Appendix (ref). We recommend the use of at least the correlation-based and fully-restricted estimators. The fully-restricted estimator, based on (ref), uses all the information available and is the simplest to compute, but it might be sensitive to misspecification of the model for the conditional distributions. In particular, it can deliver estimates outside the interval $[-1,1]$ under misspecification. The correlation-based estimator is more robust in the sense that it is the only estimator that guarantees estimates in the interval $[-1,1]$ under misspecification.\footnote{While we impose correct specification of the DR model for the conditional distributions, the derivation of the theoretical results does not rely fundamentally on correct specification. We conjecture that the probability limit of the correlation-based estimator still has an interpretation as correlation of pseudo-ranks under misspecification, but leave the formal analysis to future research.} We show in Appendix (ref) that the correlation-based estimator is asymptotically equivalent to the average of the regression-based and reversed regression-based estimators.
Let $\{Z_i:=(Y_i,W_i,X_i)\}_{i=1}^n$ be a random sample of $Z:=(Y,W,X)$. The following algorithms describe the estimators of $\rho_{C} $. All of them are based on DR.
Section (ref) shows that the estimators described in Algorithm (ref) follow normal distributions in large samples. The variances of these distributions, however, have complicated forms and are difficult to estimate. Section (ref) also shows that the asymptotic distributions can be consistently estimated using exchangeable bootstrap. Exchangeable bootstrap is a general resampling method that includes empirical, weighted, wild and subsampling bootstrap as special cases; see Comment (ref). The following algorithm describes how to obtain bootstrap draws of the estimators of $\rho_C$.
We now show how to use the exchangeable bootstrap to obtain standard errors for the estimators of $\rho_C$ and construct asymptotic confidence intervals for $\rho_C$. Algorithm (ref) describes the procedure for $\widehat \rho_C$. A similar algorithm applies to $\breve \rho_C$. Let $B$ a prespecified number of bootstrap repetitions and $\alpha$ be the significance level for the confidence intervals. For example, $B=500$ and $\alpha=0.05$.
We analyze intergenerational income mobility in Switzerland using the Economic Well-Being of the Working and Retirement Age Population Data (WiSiER).
WiSiER data include Swiss individuals from 11 Cantons from 1982 to 2016. The Swiss Federal Statistical Office merged data from tax records, social insurance, unemployment data, and surveys, creating a unique opportunity to analyze mobility. An ID can match parents and children. While many approaches seem feasible, we compare fathers and children at the same age of 35. As a result, the observations stem from different periods, with most of our successful matches coming from 1982-1990 (fathers) and 2000-2016 (children). The primary outcome variable is yearly real insured labor income (AHV) in $1,000$ Swiss francs (CHF) at the age of 35. The following covariates are available for both fathers and children: months experience, indicators for high-education (12 or more years of schooling), Swiss citizenship, and being single, and number of own children. Further, we include the fathers age at child's birth, and year and canton fixed effects for the children. Finally, for the analysis we exclude the following observations: (i) children where there is no parent in the data, (ii) observations with no information on the child's or father's birth year, and (iii) whenever the father was younger than 15 at the birth of the child. We conduct separate analyses for the relationships with sons and daughters. Table (ref) reports descriptive statistics for the data used in the analysis. It shows that father's characteristics are similar in families with sons and daughters. This alleviates a potential concern about endogenous selection in the comparison between sons and daughters.
Table (ref) reports the results of RRR and CRRR. The CRRR results are obtained using Algorithms (ref) and (ref) for the correlation-based estimator with a logistic link function and a mesh of 200 points located at sample quantiles in a sequence of orders from $0.01$ to $0.99$ with increments of $0.98/199$. We use linear interpolation to obtain estimates of the conditional ranks corresponding to intermediate points in the mesh. The standard errors (SE) and 95% confidence intervals (95% CI) are computed by empirical bootstrap with 500 repetitions. Based on the results of numerical simulations reported in Appendix (ref), we do not impose tail restrictions. In results not reported, we find very similar estimates, standard errors and confidence intervals for regression-based and fully restricted estimators.\footnote{These results are available from the authors upon request.} We show the robustness of the results to the choice of link function in Section (ref).
We find significant positive income persistence in both father-son and father-daughter relationships, with and without covariates. However, the persistence is much stronger for sons than for daughters suggesting the presence of a gender gap in intergenerational transmission of income even after controlling for the father's and child's characteristics. Comparing RRR and CRRR, we find that within-group persistence accounts for approximately 62% of the overall income persistence for sons and about 52% for daughters. These results highlight the substantial role of both within-group and between-group differences in explaining intergenerational mobility.
A subgroup analysis reveals relatively more mobility in families with a larger number of children and with a low educated father. In particular, we find relatively less persistence for sons in large families and more for daughters of high educated fathers. This would be consistent with decreasing returns of intergenerational transfers with respect to family size and increasing with respect to father's education. This heterogeneity, however, is not statistically significant. We do not find differences in intergenerational mobility for families with immigrant fathers in Switzerland, unlike the results of abramitzky2021intergenerational for the U.S. This difference might be due to the small fraction of immigrant fathers in the sample, see Table (ref).
Figures (ref) and (ref) show heatmaps of transition matrices for father-son and father-daughter, respectively. These matrices are a parsimonious representation of the joint distribution of income for father and child discretized in cells defined by deciles. They are commonly used in intergenerational mobility studies to provide a more granular measure of persistence than the rank-rank regressions. We report all the entries in percent deviations from $0.1$ because all the entries should be equal to $0.1$ under perfect mobility, that is when the income of the child is independent of the income of the father. Panels (A) report transition matrices based on marginal ranks, similar to previous studies. Panels (B) report conditional transition matrices based on conditional ranks, which are new to this paper and capture within-group dependence. For father-son, we find that the highest values in panel (A) are concentrated on the diagonal, which is consistent with the positive RRR estimate in Table (ref). The results in panel (B) show a less clear pattern once we control for covariates, consistent with the lower CRRR estimates in Table (ref). The results for father-daughter show similar but weaker patterns as we expect from the smaller correlation estimates in Table (ref). Interestingly, for both sons and daughters the highest probability occurs at the bottom right corner of the very top deciles conditionally and unconditionally.
One concern about the CRRR results in Table (ref) is that the child's covariates might be picking up indirect sources of intergenerational mobility of income. For example, fathers might invest in child's education to increase the child's income prospects. To deal with this concern, Table (ref) reports CRRR results where the child's covariates, other than year and canton fixed effects, are excluded from the covariate set $X$. These results are obtained using the correlation-based estimator with a logistic link function with the same parameter choices as in Table (ref).
As expected, not accounting for the child's covariates increases the importance of within-group persistence with the estimates increasing to about 80% for father-son and 69% for father-daughter. In both cases the increase is about 17-18%. The other conclusions remain unchanged. In particular, we still find a significant gender gap in intergenerational transmission of income, and relatively less persistence for sons in large families and more for daughters of high educated fathers.
Table (ref) reports the results of CRRR using the correlation-based estimator with a Gaussian or probit link function. The estimates, standard errors and confidence intervals are almost identical to Table (ref) showing the robustness of the results to the use of the logistic versus Gaussian link functions.
In this section we provide asymptotic theory for the estimators of the CRRR slope $\rho_C$. We focus on the correlation-based and fully-restricted estimators of Algorithm (ref). We derive their asymptotic distributions by the delta method. For example, we take the following steps for the correlation-based estimator:
The distribution of the fully-restricted estimator is derived following steps (1)-(4), and replacing the correlation-based functional in step (1) by the fully-restricted functional:
Before stating formally the main results, we review the existing theory for the estimator of the RRR slope. The purpose of this review is to explain why the existing results do not cover the estimators of the CRRR slope. hoeffding1948class first derived the asymptotic distribution of the RRR slope estimator using the theory of U-statistics. We cannot follow the same approach because none of our estimators has a U-statistic representation. ren1995hadamard alternatively derived the asymptotic distribution of the RRR slope estimator using the delta method. ren1995hadamard used analogous steps to our procedure described above. The following remarks explain each step of our procedure and point out the challenges and differences with respect to ren1995hadamard.
We formally state now the main results from the steps (4) and (5). The result from step (3) is relegated to Appendix (ref) because it is of more technical nature. We state all the results for the logistic link function because it produces analytically simpler expressions, but it can be readily extended to the Gaussian link at the cost of more cumbersome notation.
We start by imposing some conditions on the DR model.
In order to state the result about the limit process for the inputs, we define, for $R \in \{Y,W\}$,
Consider the empirical processes $ (r,x) \mapsto \widehat Z_R(r,x) := \sqrt{n}\left(\widehat F_{R \mid X}(r \mid x) - F_{R \mid X}(r \mid x) \right)$, $R \in \{Y,W\}$, and $f \mapsto \widehat G_Z(f) := \sqrt{n} \int f \mathrm{d} (\widehat F_Z - F_Z),$ where $\widehat F_{R \mid X}(r \mid x) := \Lambda(x'\widehat \beta_R(r))$, $\widehat F_Z$ is the empirical distribution function of $Z = (Y,W,X)$, and $\mathcal{F}$ is a class of measurable functions that (i) includes $F_{Y \mid X}$, $F_{W \mid X}$, $F_{Y \mid X}^2$, $F_{W \mid X}^2$, $F_{Y \mid X} F_{W \mid X}$ and the indicators of all the rectangles in $\bar{\mathbb{R}}^{d_x+2}$, where $\overline{\mathbb{R}} := \mathbb{R} \cup \{-\infty, \infty\}$ is the extended real line, and (ii) is totally bounded under the metric: $$\lambda(f,\tilde f) = \left[\int (f-\tilde f)^2 \mathrm{d} F_Z \right]^{1/2}, \quad f,\tilde f \in \mathcal{F}.$$ Let $Z_n\rightsquigarrow Z $ in $\mathbb{E}$ denote weak convergence of a stochastic process $Z_n$ to a random element $Z$ in a normed space $\mathbb{E }$, as defined in van1996weak.
The next result states a central limit theorem for $\tilde \rho_C$ and $\breve \rho_C$, and the asymptotic equivalence between $\tilde \rho_C$ and $\widehat \rho_C$.
The variance of the limit processes $Z_{1,\rho}$ and $Z_{\rho}$ have complicated expressions that might be difficult to estimate analytically. To avoid this difficulty, we propose the use of bootstrap to make inference. We show that the exchangeable bootstrap draws of Algorithm (ref) have the same asymptotic distribution as the CRRR estimators under the following assumption on the weights:
In order to state the results about bootstrap validity formally, we follow the notation and definitions in van1996weak. Let $D_{n}$ denote the data vector and $M_{n}$ be the vector of random variables used to generate bootstrap draws given $D_{n}$. Consider the random element $ \mathbb{Z}^{*}_{n} = \mathbb{Z}_{n}(D_{n}, M_{n})$ in a normed space $ \mathbb{E}$. We say that the bootstrap law of $\mathbb{Z}^{*}_{n}$ consistently estimates the law of some tight random element $\mathbb{Z}$ and write $\mathbb{Z}^{*}_{n} \rightsquigarrow_{{\mathrm{P}}} \mathbb{Z} $ in $\mathbb{E}$ if:
where $\text{BL}_{1}(\mathbb{E})$ denotes the space of functions with Lipschitz norm at most 1 and ${\mathrm{E}}_{M_{n}}$ denotes the conditional expectation with respect to $M_{n}$ given the data $D_{n}$; and $ \rightarrow_{{\mathrm{P}}}$ denotes convergence in (outer) probability.
We now provide a bootstrap central limit theorem for the estimators of the CRRR slope. This result follows from a functional central limit theorem for the input processes, which we establish in Lemma (ref) in Appendix (ref), and the functional delta method for the bootstrap.
This paper introduces the conditional rank-rank regression (CRRR) as an alternative to traditional rank-rank regressions with covariates (RRRX) for measuring within-group mobility and persistence. The CRRR uses conditional ranks of the variables of interest given covariates, in contrast to RRRX which uses marginal ranks net of covariate effects. We show that the CRRR slope preserves an intuitive interpretation as the average conditional rank correlation between the variables, similar to RRR without covariates. In contrast, the slope of RRRX loses the rank correlation interpretation and can take on values outside the interval $[-1, 1].$ The CRRR is also suitable for subgroup analysis, where the CRRR slopes maintain a rank correlation interpretation conditional on the groups.
We propose a distribution regression estimator for CRRR where the conditional distributions are modeled flexibly using parametric link functions. The estimator is easy to implement and computationally tractable. We derive asymptotic theory for the estimator based on the functional delta method. The analytic asymptotic variance is cumbersome, so we propose an exchangeable bootstrap procedure for inference. The bootstrap procedure is also used to construct confidence intervals. We illustrate the usefulness of CRRR in an empirical application to intergenerational income mobility in Switzerland. The application reveals stronger intergenerational persistence between fathers and sons than fathers and daughters, where the within-group persistence accounts for between $52\%$ and $79\%$ of the overall persistence. We also find some evidence of heterogeneity across groups defined by father's education and family size. The results are robust to the exclusion of child's covariates and the use of logistic or Gaussian link functions.
In summary, CRRR provides a well-grounded measure of within-group mobility and persistence. The distribution regression estimator, coupled with exchangeable bootstrap inference, provides a practical and flexible way to implement CRRR in empirical applications. We expect CRRR will be a useful addition to the toolkit of methods for studying mobility and persistence. A natural next step is to consider high dimensional settings where they might be many covariates. An extension where we derive orthogonal moment conditions for CRRR to apply double/debiased machine learning (DML) is underway.