The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
128,053 characters
Identification and Estimation of Intergenerational Income Mobility Measures
\title{Identification and Estimation of Intergenerational Income Mobility Measures\footnote{I am grateful to Juan Carlos Escanciano and Jan Stuhler for
their guidance and support, and to Christophe Gaillac, Antonio Raiola, Stephen Jenkins, and Nazarii Salish for their insightful discussions and comments. I also thank participants of the International Association for Applied Econometrics 2025 Annual Conference, the Eleventh Italian Congress of Econometrics and Empirical Economics, the Econometrics Brown Bag Seminar at University College London, the PhD Workshops at Universidad Carlos III de Madrid, the ENTER Jamboree 2025 Conference, and the III at 10: New Directions in Inequality Research for their valuable feedback.}} \author{Alejandro Puerta-Cuartas\thanks{Department of Economics, Universidad Carlos III, Madrid, Spain, \href{E-mail: [email removed]}{E-mail: [email removed]} } \\\vspace{5mm}\today
} \pubMonth{}\pubYear{}\pubVolume{}\pubIssue{}\JEL{J62; D63; C13; C14 }\Keywords{Intergenerational mobility, inequality, semiparametric inference, missing data.}
\begin{abstract}
Measuring the intergenerational transmission of lifetime economic status is complicated by researchers often only observing snapshots of income at specific ages. Consequently, standard practice estimates intergenerational mobility using income averages, introducing life-cycle bias that compromises reliability and comparability across studies, time, and place. I develop a missing data framework that exploits available income data and observable characteristics to eliminate life-cycle bias. This method combines nonparametric identification with Neyman-orthogonal moments to construct debiased machine learning estimators for intergenerational income mobility measures under plausible missing-at-random and testable independence assumptions. I apply this framework to estimate the intergenerational elasticity for the U.S. using the Panel Study of Income Dynamics across birth cohorts from 1954 to 1977 with rolling 10-year windows. While existing approaches estimate values between 0.41 and 0.54, the proposed method yields substantially higher estimates ranging from 0.6 to 0.7, averaging 0.64. These results align closely with recent evidence using long time averages over mid-career periods, reinforcing high U.S. intergenerational persistence.
\vspace{2mm}
\end{abstract}\maketitle \clearpage \section{Introduction}\label{sec:1}
Many economic and causal parameters depend on lifetime outcomes such as income, earnings, or consumption. However, survey and administrative data typically cover only a limited segment of individuals’ working lives, posing significant challenges across fields such as household economics, education, and labor. For example, this limitation complicates measuring the degree to which income shocks transmit to consumption \citep{jappelli2010consumption}, distorts estimates of long-run earnings returns to education \citep{heckman2006earnings}, and introduces life-cycle bias when estimating intergenerational income persistence \citep{solon1992intergenerational}. This paper develops a missing data framework to construct consistent, locally robust estimators for intergenerational income mobility measures from incomplete income data and individual characteristics. The method combines nonparametric identification with Neyman-orthogonal moments to eliminate the life-cycle bias arising from incomplete income data. I illustrate the practical advantages of this approach by estimating the intergenerational elasticity (IGE) in the United States.\par
Existing approaches to measuring intergenerational income persistence rely on income averages, introducing life-cycle bias. This occurs because annual income snapshots only partially capture lifetime economic status, particularly when observed at early or late career stages \citep{haider2006life}. The problem is further exacerbated by parents and children typically being observed at different life stages \citep{jenkins1987snapshots}, and by individuals exhibiting heterogeneous income growth that varies with parental characteristics \citep{halvorsen2022earnings}.\par
While correcting individual sources of bias improves estimates, this piecemeal approach hinders comparability. Recent literature has made substantial progress by separately addressing life-cycle bias: \citet{mazumder2016estimating} use long-time averages for parents, while \citet{mello2022lifecycle} propose a life-cycle (LC) estimator that predicts children's income profiles accounting for income growth that varies with parental characteristics. Nevertheless, questions remain about the robustness of existing estimates and the reliability of comparisons across time and place \citep{mogstad2023family}. I formalize why comparability remains elusive: bias magnitudes vary with research design, income dynamics, and sampling rules, causing estimators relying on income averages to converge to context-specific parameters even after correcting specific biases. Estimates from different studies thus target fundamentally different parameters, compromising both the assessment of mobility levels and comparisons across studies and countries. This underscores the need to move beyond proxy refinement and focus on identification, which jointly addresses all sources of bias.\par
This paper shows that, despite lifetime income being unobserved, intergenerational income mobility measures can be nonparametrically identified from incomplete income data and observable characteristics. The key insight is that although individual income profiles are only partially observed, cross-sectional variation across ages and individual characteristics can be exploited to recover them. To illustrate this framework concretely, I establish nonparametric identification for the IGE, the most common measure of income persistence across generations \citep{nybom2017biases}. The covariance between children's and parents' lifetime incomes, the IGE numerator, is identified as the average conditional covariance of their annual incomes across all age pairs, and the denominator, the variance of parental lifetime income, is recovered from conditional autocovariances at nearby ages and predicted income profiles at distant ages.\par
Identification relies on two sets of assumptions. First, to recover income profiles and the autocovariance of parental income, a missing-at-random (MAR) condition is imposed, requiring that, conditional on individual characteristics, whether income is observed at a given age is unrelated to the income level itself. Additionally, a common support condition requires that in the population, the probability of observing income at a given age, conditional on observable characteristics, is bounded away from zero and one. Second, to eliminate life-cycle bias from both generations, two independence assumptions are imposed. For children, I assume that income prediction errors—the part of annual income unexplained by observable characteristics—are uncorrelated with parental lifetime income. This condition can be satisfied by modeling children's income profiles as functions of parental characteristics (such as income and education) interacted with age, capturing systematic differences in income growth across family backgrounds. For parents, I assume that prediction errors at distant periods (more than 10 years apart in the application) are uncorrelated. By conditioning on observable characteristics that capture persistent income components, these prediction errors reflect only transitory shocks that dissipate over time.\par
Building on nonparametric identification, I construct debiased machine learning estimators for intergenerational income mobility measures. For the IGE, the identification result shows that it can be recovered from conditional expectations of income profiles and parental income autocovariances. Estimating these nuisance parameters requires flexibly accommodating high-dimensional individual characteristics and heterogeneous, nonlinear income dynamics. While machine learning (ML) methods can adaptively learn these complex relationships, they introduce regularization and model selection bias that would propagate to IGE estimation. Following \cite{chernozhukov2018double} and more specifically \cite{chernozhukov2022locally}, I address this by constructing Neyman-orthogonal moments that incorporate the influence function of the first-stage ML estimates. This orthogonalization ensures local robustness: first-stage estimation errors affect the IGE estimate only at second order, thereby mitigating regularization and model selection bias. Additionally, I employ cross-fitting, which prevents overfitting and own-observation bias. This approach enables valid inference while leveraging the flexibility of modern ML methods.\par
Consistency and asymptotic normality are established for the resulting estimators, enabling valid inference that accounts for first-stage machine learning estimation of nuisance parameters. Simulations suggest the estimators exhibit sound finite-sample performance, with negligible bias that vanishes as sample size increases and coverage rates close to nominal levels. In contrast, a naive ML plug-in estimator uncorrected for first-step estimation errors and existing approaches relying on income averages exhibit sizable bias and severe undercoverage. Additionally, I develop a locally robust test for the assumption that children’s prediction errors are uncorrelated with parental lifetime income. I show that it attains correct asymptotic size, is consistent against fixed alternatives, and exhibits nontrivial power against local alternatives. To facilitate implementation, the proposed methods will soon be available in a companion user-friendly \texttt{R} package, \texttt{LRIGE}.\par
I apply the framework to estimate the IGE in the United States using the Panel Study of Income Dynamics (PSID). This dataset is particularly well-suited for this analysis: it has been widely used to study income mobility, provides long-term income histories, and includes rich individual characteristics relevant for income prediction. Importantly, the literature suggests that the MAR assumption for income missingness in the PSID is empirically plausible, provided relevant observables are accounted for \citep{fitzgerald2011attrition,schoeni2015implications}. The analysis focuses on birth cohorts spanning 1954 to 1977, using rolling 10-year windows and covering the lifetime period from ages 25 to 55. This sample design ensures both parents and children are observed during working life, guaranteeing both missing and non-missing observations at each age and satisfying the common support condition required for identification. Across cohort windows, the proposed test strongly supports that children's prediction errors are uncorrelated with parental permanent income: in 14 of 15 cases, the correlation is not statistically different from zero. Additionally, the results suggest that autocorrelation in parental income prediction errors becomes negligible beyond ten years in all 15 windows. Together with the empirical plausibility of the MAR assumption documented in prior work, these results validate the identifying assumptions underlying the framework. \par
The proposed method yields IGE estimates ranging from 0.6 to 0.7 across cohorts, with an average of 0.64. These results align closely with recent PSID-based studies that address life-cycle bias and find values exceeding 0.6 \citep{gouskova2010estimating,chau2012intergenerational,mazumder2016estimating}. In contrast, a conventional approach using three-year income averages during mid-life for both generations yields an average IGE of only 0.43 (ranging from 0.41 to 0.49). The LC estimator substantially improves upon this, producing an average IGE of 0.50 (ranging from 0.46 to 0.54). Decomposing estimates into their covariance and variance components reveals that the LC's covariance estimates closely match the locally robust benchmark, providing direct evidence that its explicit modeling of heterogeneous income profiles successfully addresses children's life-cycle bias. Finally, a plug-in ML estimator yields an average IGE of 0.60 (ranging from 0.46 to 0.71), with bias reaching 0.24 in one cohort and confidence intervals 41\% narrower on average than the debiased approach, underscoring the importance of orthogonalization for both bias correction and valid inference.\par
This paper makes two contributions. First, it develops a method for obtaining reliable and comparable intergenerational income mobility estimates across studies, time, and place. Although I focus on the IGE, the framework applies more broadly to other mobility measures. For example, the intergenerational correlation equals the IGE scaled by the ratio of parents’ to children’s income standard deviations. Extending the IGE result to the correlation thus requires only two additional steps: identifying children's income variance by the same argument used for parents, and constructing a Neyman-orthogonal moment for the correlation parameter rather than the elasticity. \par
Second, it introduces a missing data framework that enables locally robust inference when parameters depend on partially observed lifetime outcomes. Standard debiased machine learning assumes the parameter of interest is identified from observable data \citep{chernozhukov2022locally}. Satisfying this assumption is not straightforward when lifetime outcomes, such as permanent income, consumption, or educational returns, are only partially observed across the life cycle. I address this challenge by establishing that these parameters can be nonparametrically identified from incomplete income data and observable characteristics under MAR and orthogonality conditions, thereby enabling construction of Neyman-orthogonal moments for locally robust inference. The framework directly extends to estimating long-run educational returns and measuring intergenerational transmission of well-being, where lifetime outcomes for one or both generations are partially observed. It also applies to more complex settings, such as partial insurance models \citep{blundell2008consumption}. When rich longitudinal data are available, nonparametric identification makes it possible to exploit high-dimensional data through machine learning to recover the variance-covariance structure of consumption and income while maintaining locally robust inference.\par
The remainder of the paper is as follows: Section \ref{sec:2} formalizes that proxy-based estimators converge to context-specific limits, compromising comparability across studies. Section \ref{sec:3} presents nonparametric identification of the IGE, a locally robust estimator, and inference results including asymptotic normality and testing. Section \ref{sec:sims} reports simulations, and Section \ref{sec:app} applies the estimator to measure the IGE in the United States. Proofs are provided in the Appendix.
\section{Biases and Comparability in Intergenerational Elasticity Estimates}\label{sec:2}
The biases introduced by using income averages to assess intergenerational transmission of lifetime economic status affect all mobility measures. To derive explicit characterizations, however, I focus on a specific measure: the intergenerational elasticity of income, the most common measure of income persistence across generations \citep{nybom2017biases}. The IGE captures the degree to which income differences between parents are associated with income differences among their children. Formally, the IGE is defined by the regression
\begin{align}\label{eq:igemain} Y_c^P&=\alpha_0+\beta_0 Y_f^P+u, \quad \mathbb{E}\left[u \left(1, Y_f^P\right)'\right]=0,\end{align} where $Y_f^P$ and $Y_c^P$ denote the permanent component of log annual income for fathers and children \citep{solon1992intergenerational}, $u$ is an idiosyncratic error term uncorrelated with parental income, and $\beta_0$ is the intergenerational elasticity. Henceforth, I will refer to $Y^P$ as permanent income.\par
According to equation (\ref{eq:igemain}), the closed form solution for the IGE is given by \begin{align}\label{eq:beta} \beta_0&=\frac{\mathbb{E}\left[\left(Y_c^P-\mathbb{E}\left[Y_c^P\right]\right)\left(Y_f^P-\mathbb{E}\left[Y_f^P\right]\right)\right]}{\mathbb{E}\left[\left(Y_f^P-\mathbb{E}\left[Y_f^P\right]\right)^2\right]}. \end{align}
However, because permanent income data is rarely available in practice, researchers typically rely on short-term income snapshots \citep{mazumder2005fortunate}, using either a single-year observation or a multi-year average as a proxy for permanent income. This raises the question of what exactly the available estimators in the literature measure, and whether their estimates are comparable across settings.\par
To address this, I adopt the strategy that \cite{mogstad2024instrumental} refer to as \say{reverse engineering}. In the context of instrumental variables (IV), this approach begins with a practical problem: when treatment effects exhibit unobserved heterogeneity (UHTE), the classical linear IV model is misspecified. Yet a linear IV estimate can still be computed. The reverse engineering framework then seeks to determine what, if anything, this estimator measures. Thus, it proceeds by starting with the tool, and it attempts to reverse engineer an interpretation for it under suitable assumptions.\par
The idea of reverse engineering has already been applied in the intergenerational mobility literature to formally characterize the behavior of estimators based on proxy measures of permanent income \citep{solon1992intergenerational,nybom2016heterogeneous}. Specifically, it has been used to establish the probability limit of such estimators, thereby identifying the distinct sources of bias introduced by relying on imperfect proxies. \par
To set the grounds for the analysis, I follow the literature by first characterizing the probability limit of the estimator based on income proxies. I begin by defining the observed data typically available to researchers. Most empirical studies utilize longitudinal datasets such as the Panel Study of Income Dynamics (PSID), which contain only partial income trajectories for parents and children, along with additional individual and family characteristics. Formally, the observed data consist of an independent and identically distributed (i.i.d.) sample of $\bm W = \big(\bm Y_c \odot \bm D_c, \bm Y_f \odot \bm D_f, \bm D_c, \bm D_f, \bm X \big),$ where \(\bm Y_c\) and \(\bm Y_f\) are \(T\)-dimensional random vectors containing information on (log) annual child and parental income, respectively, the vectors \(\bm D_c\) and \(\bm D_f\) are \(T\)-dimensional indicator vectors, with elements \(D_{gt} = 1\) if \(Y_{gt}\) is observed and \(D_{gt} = 0\) otherwise, for \(g \in \{c,f\}\), \(\odot\) denotes the element-wise product, so that \(\bm Y_g \odot \bm D_g\) contains the observed entries of \(\bm Y_g\) and zeros elsewhere, and the vector $\bm X$ contains observed characteristics for both generations.\par
The standard approach to estimate the IGE, which I label the mid-life income (MI) estimator, consists of a two-step approach. First, it proxies permanent income as the average of $T_f$ and $T_c$ (log) annual income observations around mid-life for the fathers and the children, respectively. In the second step, it regresses the child's proxy measure on the parent's. This practice has a long history, dating back to early contributions that recognized the problem of measurement error and sought to approximate permanent income using multi-year averages \citep{de1973relation,hauser1975socioeconomic,freeman1978black,tsai1983sex}. Its limitations were also identified early on: \citet{creedy1977distribution} and \citet{jenkins1987snapshots} pointed out that this approach suffers from life-cycle bias \citep{creedy1977distribution,jenkins1987snapshots}, which arises because income \say{snapshots} fail to capture complete lifetime earnings profiles and because parents and children are often observed at different stages of their life cycles. Despite these limitations being identified decades ago, the MI estimator remains the standard practice in empirical studies of intergenerational mobility.
\par Formally, the MI estimand is defined as the slope coefficient in the projection:
\begin{align*} \tilde{Y}_c^P&=\alpha^{MI}+\beta^{MI} \tilde{Y}_f^P+u^{MI}, \quad \mathbb{E}\left[u^{MI} \left(1, \tilde{Y}_f^P\right)'\right]=0,\\\nonumber \tilde{Y}_g^P&\coloneqq\frac{1}{T_g}\sum_{j\in \mathcal{M}_g}Y_{gj}D_{gj}, \quad g\in \{c,f\},\end{align*}
where $D_{gj}=1$ when $Y_{gj}$ is observed and zero otherwise, $\mathcal{M}_g$ is a set of pre-defined mid-life years for generation $g$, and $T_{g}\coloneqq\sum_{j\in \mathcal{M}_g}D_{gj}$ is the number of years used for the average. \par
To establish the probability limit of the MI estimator using the reverse engineering approach \citep{mogstad2024instrumental}, I now impose standard assumptions used in the literature. A comprehensive discussion of the MI estimator’s definition, theoretical underpinnings, assumptions, and sources of bias can be found in Appendix \ref{sec:mi}.\par
\begin{assumption1}{1}{MI}(Annual Income Process)\label{as:1}
The relationship between annual and permanent income is governed by \begin{align*}
Y_{gt}&=\lambda_tY^P_g+v_{gt}, \quad \mathbb{E}\left[v_{gt} Y^P_g\right]=0, \quad g\in \{c,f\}, \quad t=1,...,T,\\ \lambda_t&=1, \forall t\in \mathcal{M}_g, \quad g\in \{c,f\}.\end{align*} where $\lambda_t$ captures that the persistence of permanent income may vary over the life-cycle period, and $v_{gt}$ is an age shock.
\end{assumption1} \begin{assumption1}{2}{MI}(Conditional Mean Independence) \label{as:2mi} The following conditional mean restrictions hold \begin{align*} \mathbb{E}\left[v_{ct}v_{fj}\big | D_{ct},D_{fj}\right]&=0, \quad t\in \mathcal{M}_c, \quad j\in \mathcal{M}_f,\\ \mathbb{E}\left[v_{fj}Y_c^P\big | D_{fj}\right]&=0, \quad j\in \mathcal{M}_f, \\ \mathbb{E}\left[v_{ft}Y_f^P | D_{ft}, D_{fj}\right]&=0, \quad tj \in \mathcal{M}_f, \\ \mathbb{E}\left[v_{gj} | D_{gj}\right]&=0, \quad g\in \{c,f\}, \quad j \in \mathcal{M}_j.\end{align*}
\end{assumption1}
The inconsistency of the MI estimator is well-documented in the literature. Proposition \ref{coro:1} restates this result to make explicit the four sources of bias that will be central to the discussion. While the main characterization of these biases is familiar, the proposition extends prior work by incorporating missing income data. Appendix \ref{sec:equi} shows that Proposition \ref{coro:1} reduces to the results of \cite{solon1992intergenerational} and coincides with \cite{nybom2016heterogeneous} under certain assumption variants. In addition, the proposition formalizes the empirical observation that IGE estimates are sensitive to sample inclusion criteria and missing income, by demonstrating how the observation probabilities $p_f\left({t,j}\in \mathcal{M}_f\right)$ and $p_c\left(t\in \mathcal{M}_c\right)$ shape the asymptotic bias.
\begin{proposition}
\label{coro:1} \fontsize{10}{12}\selectfont Under Assumptions \ref{as:1} and \ref{as:2mi}, the probability limit of the MI estimator is given by:\fontsize{10}{12}\selectfont
\begin{gather}\label{eq:mi_inconsistent}
\hat{\beta}_n^{MI}\overset{p}{\to}\frac{\beta_0\mathbb{E}\left[\left(Y_f^P-\mathbb{E}\left[Y_f^P\right]\right)^2\right]+\overbrace{\frac{1}{T_c}\sum_{t} \mathbb{E}\left[Y_f^Pv_{ct}\big | D_{ct}=1,t\in \mathcal{M}_c\right]}^\text{(c)}\times \overbrace{p_c\left(t\in \mathcal{M}_c\right)}^\text{(d)}}{\mathbb{E}\left[\left(Y_f^P-\mathbb{E}\left[Y_f^P\right]\right)^2\right]+\underbrace{\frac{1}{T_f^2}\sum_{t}\sum_{j}\mathbb{E}\left[v_{ft}v_{fj}\big | D_{ft}=1,D_{fj}=1, \left\{t,j\right\}\in \mathcal{M}_f\right]}_\text{(a)}\times \underbrace{p_f\left(\left\{t,j\right\}\in \mathcal{M}_f\right)}_\text{(b)}},
\end{gather} \fontsize{10}{12}\selectfont where $v_{ct}$ and $v_{ft}$ are children and parental age shocks to (log) annual income as defined in Assumption \ref{as:1}, and
$p_c\left(t\in \mathcal{M}_c\right)$ and $p_f\left(\left\{t,j\right\}\in \mathcal{M}_f\right)$ denote the probabilities of observing child income at mid-life year $t$, and parent income at mid-life years in years $t$ and $j$, respectively.
\end{proposition}
Proposition \ref{coro:1} presents a formal statement of the biases already familiar from prior work, including (a) the measurement error and life-cycle bias of parental income \citep{solon1992intergenerational,mazumder2005fortunate}, (b) the sensitivity of the IGE estimates to low, zero, and missing parental income observations \citep{couch1998sample,dahl2008association,chetty2014land,nybom2016heterogeneous}, (c) the measurement error and life-cycle bias of children's income \citep{nybom2016heterogeneous}, and (d) the sensitivity to the number of years and the selected year(s) to measure children's income \citep{mello2022lifecycle}. For ease of exposition, the proposition is stated under a classical errors-in-variables formulation, so that the age-shocks are treated as capturing both transitory fluctuations and the systematic life-cycle bias. In Appendix \ref{sec:equi} (equation \ref{eq:ns}) this restriction is relaxed by allowing for a generalized errors-in-variables structure that explicitly separates the life-cycle component. A brief discussion of each component that hinders the consistent estimation of the IGE by $\hat{\beta}_n^{MI}$ can be found in Appendix \ref{sec:mi}.\par
Identifying the distinct biases introduced by noisy measures has enabled the literature to refine estimation procedures by addressing specific sources of bias. For example, \cite{mazumder2005fortunate} addresses the measurement error and life-cycle bias of parental income (component (a) in equation (\ref{eq:mi_inconsistent})) by using long-term parental income averages centered at age 40. More recently, \cite{mello2022lifecycle} address life-cycle bias from using snapshots of children's income, captured by terms (c) and (d), by predicting children's income profiles from standard observables such as age and education, while allowing income growth to be steeper for children from more affluent families. \cite{lubotsky2006interpretation} show that including proxy variables separately in a regression and optimally weighting their coefficients yields less attenuated estimates than constructing a summary measure of the proxy variables.\par
Despite these advances, questions remain about the robustness of existing estimates and the reliability of comparisons across time and place \citep{mogstad2023family,mello2022lifecycle}. While the literature has rightly emphasized the sensitivity of IGE estimates to the biases in Proposition \ref{coro:1}, a more fundamental issue is that these biases systematically alter the target parameter itself. Crucially, the magnitude of each bias component varies with income dynamics, study design, and sample inclusion criteria. Due to heterogeneity in institutional contexts and study design choices, these factors differ across datasets, regions, countries, and time, rendering IGE estimates non-comparable. The following corollary makes this dependence explicit by formalizing that the target parameter is inherently dependent on the study design and the underlying income dynamics.\par
\begin{corollaryp}[Context-Dependence of the MI Estimand] \label{coro:context} Under the assumptions of Proposition \ref{coro:1}, the probability limit of the Mid-Life Income estimator is a context-dependent parameter \begin{align}\label{eq:power} \hat{\beta}_n^{MI} \overset{p}{\to} \beta^{MI}(\eta)\coloneqq\beta_0 \cdot \Delta(\eta), \end{align} where the distortion factor $\Delta(\eta) \neq 1$ captures departure from consistency. Formally, the research-design/context vector \[ \eta\coloneqq \big(\mathcal{M},T p,\Sigma\big),\]
collects all research-design choices and structural features of the income process: $\mathcal{M}\coloneqq\left(\mathcal{M}_f, \mathcal{M}_c\right)$ and $T\coloneqq\left(T_f, T_c\right)$ are the set of pre-defined mid-life years, and the number of years used for the average for generation $g\in \{c,f\}$, respectively; $p=(p_f,p_c)$ captures observation/availability and selection rules; and $\Sigma$ summarizes how permanent income and transitory fluctuations in parents’ and children’s earnings $\left(Y_f^P,v_{ft},v_{ct}\right)$ vary and relate to each other. Equation \eqref{eq:mi_inconsistent} provides the explicit form of $\Delta(\eta)$.\par
\end{corollaryp}
The central insight of Corollary \ref{coro:context} was anticipated by \cite{jenkins1987snapshots}, who demonstrated using a simple two-period life-cycle model that snapshot-based IGE estimates suffer from large life-cycle biases whose direction cannot be determined a priori. Because variance and covariance factors work in opposite directions in the bias calculation, life-cycle biases may be upwards or downwards depending on the specific parameters of the income process and the choice of estimator. This analysis revealed that same-stage-of-lifecycle estimates are not necessarily superior to contemporaneous ones, and that adjusting for within-generation age variation does not necessarily reduce bias.\par
The present analysis reaches similar conclusions through a complementary approach: reverse-engineering the standard mid-life income estimator rather than building from structural primitives. This reveals that $\hat{\beta}_n^{MI}$ converges to a context-dependent parameter $\beta^{MI}(\eta)$ where the distortion factor $\Delta(\eta)$ depends on midlife definitions, averaging windows, observation patterns, and income dynamics. While the structural model in \citep{jenkins1987snapshots} provided intuition for why biases arise and demonstrated their potential magnitude through calibrated examples, equation \eqref{eq:mi_inconsistent} shows how these biases manifest in standard practice, decomposing \(\Delta(\eta)\) into explicit components.\par
This formalization rationalizes empirical patterns that have raised concerns about the reliability of comparisons across studies, time, and place. Specifically, it shows that reliance on income proxies introduces systematic distortions, preventing identification of the true intergenerational elasticity. These distortions limit the reliability and comparability of resulting estimates across studies, as each converges to its own context-specific value $\beta^{MI}(\bm{\eta})$. A compelling example is the wide variation in recent U.S. estimates, which range from 0.35 to 0.65 \citep{mello2022lifecycle}. Corollary \ref{coro:context} makes explicit the potential drivers of this pattern. Even when components such as (a) and (c) in equation \eqref{eq:mi_inconsistent} are held constant, differences in the definition of mid-life income $\left(\mathcal{M}_g\right)$, the number of years averaged ($T_g$), or sample selection rules ($p_g$) alter the distortion factor $\Delta(\eta)$ in equation \eqref{eq:power}. As a result, each estimate converges to a different context-specific parameter $\beta^{MI}(\eta)$.\par
Distortions in $\Delta(\eta)$ also affect trend analyses, as both design choices and cohort-specific income dynamics can vary over time. In the U.S., the PSID’s transition from annual (1968–1997) to biennial interviews illustrates how survey design changes can alter the estimand: reduced income observations for recent cohorts modify $\Delta(\eta)$ through the observation probability $p_c$, potentially distorting mobility trends. Empirical evidence from Sweden further supports the theoretical distortions highlighted in Corollary \ref{coro:context}. \cite{mello2022lifecycle} show that MI-based estimates suggest a sharp decline in mobility for the 1950s–1970s cohorts, whereas their life-cycle estimator, which corrects for life-cycle bias in children’s income, indicates stable mobility across these cohorts, and a modest increase for those born in the 1980s.\par
Finally, Corollary \ref{coro:context} formalizes how differences in study design and income dynamics can undermine cross-country comparisons. Even under the same definitions of mid-life income, differences in transitory shock persistence (a), children’s income growth (c), and observation probabilities (b, d) alter the estimand $\beta^{MI}(\eta)$. This calls for caution in interpreting international patterns such as the Great Gatsby Curve: unlike scale-free measures like the Gini coefficient, MI-based IGE estimates reflect both underlying mobility and study-specific distortions captured by $\Delta(\eta)$.\par
Beyond its implications for common applications, Corollary \ref{coro:context} shows that eliminating individual biases alone does not guarantee comparability. Even after correcting for specific sources of bias, differences in study design, cohort composition, or the magnitude of residual distortions can still produce inconsistent estimates. This analysis highlights a fundamental shift in perspective: rather than addressing individual sources of bias, attention should be directed toward identifying the IGE, which simultaneously removes all biases and allows for reliable, comparable estimates.
\section{Identification, Estimation, and Inference for the IGE with Incomplete Data}\label{sec:3}
\subsection{Nonparametric Identification}
To establish identification of the intergenerational elasticity, I adopt the \say{forward-engineering} strategy proposed by \citet{mogstad2024instrumental}. In their terminology, this approach begins with a model and then constructs estimators under the assumption that the model is correctly specified. In the context of the IGE, my proposal precisely follows this logic: begin by defining permanent income and then derive the conditions under which the IGE is identified, given that definition.\par
The literature has generally characterized permanent income rather than attempting to provide a precise definition. For instance, in line with \cite{solon1992intergenerational}, \cite{mazumder2005fortunate} interprets it as the permanent component of log earnings, capturing true long-term earning capacity. In contrast, \cite{haider2006life} describes it as a long-run income variable, such as the log of the present discounted value of lifetime earnings. Other work \citep{black2011recent,corak2013income} refer more generally to log permanent earnings without elaborating further. Reflecting this theoretical heterogeneity, empirical research has proxied permanent income differently; while some studies compute it as the log of average annual income \citep{dahl2008association,mazumder2016estimating}, others take the average of log annual income \citep{zimmerman1992regression,bratberg2007trends}. \par
Our goal is not to define a unifying measure of permanent income or to identify a uniquely correct IGE. The intergenerational elasticity has never been directly observed; it has always been estimated under specific measurement choices and assumptions. I adopt a definition of permanent income that is theoretically grounded, empirically tractable, and consistent with standard practice in the literature. By settling on a workable definition, researchers can generate estimates of the IGE that are more reliable and comparable across studies and datasets. Thus, the forward-engineering approach provides a foundation for comparability and reliability. \par
\begin{definition}[Permanent income]\label{def1}
For an individual of generation $g$ (where $g \in \{c, f\}$ for child or father), permanent income $\left(Y^P_g\right)$ is defined as their average log annual income over a specific lifetime period from $t=1$ to T:
\begin{align}\label{eq:def_perm_inc} Y^P_g&\coloneqq\frac{1}{T}\sum_{t=1}^TY_{gt}, \quad g\in\{c,f\},\end{align}
where $Y_{gt}$ is log annual income in year $t$, with $t=1$ indicating the start age and $T$ the number of years covered.\end{definition}
Our definition of permanent income aligns with the literature that conceptualizes it as the permanent component of log earnings. This definition provides an empirically tractable measure that facilitates the identification of the intergenerational elasticity. The linearity of the sum-of-logs specification is crucial, as it permits the use of standard missing-at-random assumptions to recover permanent income from partially observed data. An alternative definition involving the log of the average introduces nonlinearities that preclude a similar identification strategy and require stronger assumptions about the joint distribution of income over lifetime. For a detailed discussion of these considerations, see Appendix~\ref{sec:lfinc}. When applied to the life-cycle estimator of \citet{mello2022lifecycle}, this definition yields results that are virtually identical to those obtained by defining permanent income as the log of the average, as shown in Table~\ref{tab:ige_cohort}.\par
The fundamental challenge in estimating (identifying) the IGE is its reliance on unobserved permanent income. Traditional approaches, such as the Generalized Error-in-Variables (GEIV) model \citep{haider2006life}, motivate the use of short-term averages of income—typically during mid-life—by assuming a parametric link between observed annual income and unobserved permanent income. In contrast, this definition underpins the nonparametric nature of the identification result, enabling us to recover the intergenerational elasticity without restrictive functional form assumptions.\par
While the IGE in equation (\ref{eq:beta}) depends on unobserved permanent income for both generations, the workable definition allows us to reformulate the target parameter in terms of partially observed (log) annual incomes:
\begin{align}\label{eq:beta_new1}
\beta_0&=\frac{\mathbb{E}\left[\left(Y_c^P-\mathbb{E}\left[Y_c^P\right]\right)\left(Y_f^P-\mathbb{E}\left[Y_f^P\right]\right)\right]}{\mathbb{E}\left[\left(Y_f^P-\mathbb{E}\left[Y_f^P\right]\right)^2\right]}= \frac{\sum_{t=1}^T \sum_{j=1}^T \mathbb{E}\left[\left(Y_{ct} - \mathbb{E}\left[Y_{c}^P\right]\right)\left(Y_{fj} - \mathbb{E}\left[Y_{f}^P\right]\right)\right]}{ \sum_{t=1}^T \sum_{j=1}^T \mathbb{E}\left[\left(Y_{ft} - \mathbb{E}\left[Y_{f}^P\right]\right)\left(Y_{fj} - \mathbb{E}\left[Y_{f}^P\right]\right)\right]},
\end{align} where the scaling component $1/T$ cancels out.
Although this is a crucial step, it is not sufficient for identification. To bridge this gap, I exploit observable characteristics by decomposing (log) annual income as
\begin{align}\label{eq:cond_mean_main}
Y_{gt} &= \mathbb{E}\left[Y_{gt} \mid \bm{X}_{gt}\right] + \epsilon_{gt}, \quad \mathbb{E}\left[\epsilon_{gt} \mid \bm{X}_{gt}\right] = 0, \quad g\in \{c,f\}, \quad t=1,\ldots,T,
\end{align} where $\bm{X}_{gt}$ denotes the elements in the observed characteristics $\bm X$ relevant for predicting (log) annual income of generation $g$ at time $t$, and \( \epsilon_{ct} \) is the nonparametric prediction error. Although \( \epsilon_{ct} \) can be interpreted as an age shock, it differs conceptually from the age shock \( v_{ct} \) in the GEIV model of Assumption \ref{as:1}. Specifically, \( \epsilon_{ct} \) captures the component of (log) annual income that is not explained by observed parental and own characteristics, that is, the residual from a predictive model based on observables. In contrast, \( v_{ct} \) reflects transitory deviations from an individual’s permanent income and arises within a latent factor structure that distinguishes between the permanent and transitory components of income.\par
Identification is established in two steps. Substituting the income decomposition into equation ~(\ref{eq:beta_new1}) and imposing conditional mean independence and orthogonality assumptions involving observables and prediction errors, thereby eliminating dependence on unobserved components. Then, I impose standard missing-at-random assumptions to recover the necessary conditional moments from the available data. For the income profiles, the MAR assumption enables identification of the conditional expectation $\mathbb{E}[Y_{gt} \mid \bm{X}_{gt}] = \mathbb{E}[Y_{gt} \mid \bm{X}_{gt}, D_{gt} = 1]$,
where $D_{gt}$ indicates income observability at time $t$. In a similar way, we can identify the conditional second moments arising in the denominator of equation (\ref{eq:beta_new1})
$\mathbb{E}[Y_{ft} Y_{fj} \mid \bm{X}_{ftj}] = \mathbb{E}[Y_{ft} Y_{fj} \mid \bm{X}_{ftj}, D_{ft} = 1, D_{fj}=1],$
where $\bm{X}_{ftj}$ comprises the elements in the observed characteristics $\bm X$ relevant for predicting the covariance between parental incomes at ages $t$ and $j$, and $\bm{X}_{ftj}$ is defined such that $\bm{X}_{ft} \subset\bm{X}_{ftj}$ for $t,j=1,...,T.$ With these foundations in place, I now formally state the complete set of identifying assumptions.\par
\begin{assumption1}{1}{NP}(Conditional Mean Independence and Orthogonality)\label{as:ortho_np}
\begin{enumerate}[i.] \item The observable characteristics satisfy: \begin{enumerate}[1.] \item $ \mathbb{E}\left[Y_{ct} \mid \bm{X}_{ct}, \bm{X}_{cj},\bm{X}_{fj}\right] = \mathbb{E}\left[Y_{ct} \mid \bm{X}_{ct}\right] \quad \text{ for } t,j=1,...T,$ \item $ \mathbb{E}\left[Y_{ft} \mid \bm{X}_{ft}, \bm{X}_{ftj}, \bm{X}_{cj}\right] = \mathbb{E}\left[Y_{ft} \mid \bm{X}_{ft}\right] \quad \text{ for } \bm{X}_{fj} \subset\bm{X}_{ftj},\quad t,j=1,...T.$ \end{enumerate}
\item The average covariance between children's prediction errors and parental permanent income across all observed years is zero \begin{align*} \frac{1}{T}\sum_{t=1}^T\mathbb{E}\left[\epsilon_{ct}Y_{f}^P\right]&=0,\quad \epsilon_{ct}\coloneqq Y_{ct}-\mathbb{E}\left[Y_{ct}\mid \bm X_{ct}\right], \quad t=1,\ldots,T. \end{align*} \item The average covariance of parental income prediction errors for $|t-j|>h$ is zero
\begin{align*} \frac{1}{T^2}\sum_{(t,j)\in \mathcal{H}}\mathbb{E}\left[ \epsilon_{ft} \epsilon_{fj}\right]=0,\quad \text{where } \mathcal{H}=\{(t,j): 1\leq t,j\leq T, |t-j|>h\}\end{align*} \end{enumerate}
\end{assumption1}
The first condition establishes that $\bm{X}_{gt}$ contains all relevant predictors for annual income for generation $g$ at time $t$, implying the remaining information in $\bm{X}$ provides no additional explanatory power. In Section \ref{sec:app} I illustrate that the specification of the characteristics predictive of income profiles and parental income covariance, namely, \(\bm{X}_{ct}\), \(\bm{X}_{ft}\), and \(\bm{X}_{ftj}\), can be designed to satisfy Assumption \ref{as:ortho_np}.$i$ by construction.\par
Children’s age shocks being correlated with parental permanent income constitutes a source of bias of the MI estimator (component (c) in equation (\ref{eq:mi_inconsistent})). One of the empirical patterns driving this dependence stems from children from affluent families exhibiting faster income growth, even after controlling for observables \citep{mello2022lifecycle}. The life-cycle estimator addresses this by projecting children's annual income into the space of observables. In particular, by including in $X_{ct}$ the interaction between average parental (log) annual income observations around mid-life $\left(\tilde{Y}_f^P=\frac{1}{T_f}\sum_{j\in \mathcal{M}_f}Y_{fj}D_{fj}\right)$ and children's age at time $t$, the prediction errors of children's income $\left(\epsilon_{gt}=Y_{gt}- \mathbb{E}\left[Y_{gt} \mid \bm{X}_{gt}\right]\right)$ become uncorrelated with parental permanent income $Y_f^P$. Accordingly, Assumption \ref{as:ortho_np}.$ii$ imposes that thee average covariance between children’s prediction errors and parental permanent income across all observed years is zero, once the relevant family characteristics are controlled for. In Section \ref{sec:test}, a test for Assumption \ref{as:ortho_np}$ii$ is proposed, and in the U.S. application, the test does not reject the validity of this assumption. \par
The requirement of Assumption \ref{as:ortho_np}.$iii$ arises from the fundamental mismatch between the complete income profiles required by equation (\ref{eq:beta_new1}) and the income snapshots typically available in practice. Specifically, joint observation of parental incomes \( (Y_{ft}, Y_{fj}) \) (i.e., \( D_{ft}=1, D_{fj}=1 \)) occurs only for relatively close time periods, such as incomes observed between ages 25 and 35 for a given individual. Consequently, income pairs for distant periods (\( |t-j| > h \)) are systematically absent in available data. Assumption \ref{as:ortho_np}.$iii$ addresses this empirical constraint by imposing that conditional on family characteristics
$\bm{X}_{ftj}$, parental income shocks (prediction errors $\epsilon_{ft}$ and $\epsilon_{fj}$) are uncorrelated for periods separated by more than $h$ years. The availability of rich family characteristics $\bm{X}$ makes this assumption empirically plausible, as it allows us to account for the persistent components of intertemporal dependence.\par
Income autocorrelation captures two distinct sources: a permanent component driven by family characteristics (e.g., wealth, neighborhood quality, and race), and a transitory component, driven by short-term shocks (e.g., unemployment spells, economic crises, or health events). Crucially, while the influence of transitory shocks decays as the time gap $(t - j)$ widens, the effect of family background characteristics remains over time. Assumption \ref{as:ortho_np}.$iii$ states that parental annual income from periods more than \(h\) years in the past influences current income solely through observed characteristics. This specification serves dual purposes: it realistically captures the (conditional) short-memory of transitory shocks while accommodating the limitations inherent in available longitudinal datasets. In the application, the autocorrelation in parental income prediction errors becomes negligible beyond ten years in all 15 windows, supporting the plausibility of this assumption. \par
The following assumption formalizes some necessary conditions for identifying the intergenerational elasticity using partial income data and family characteristics. First, it requires that income realizations, for both generations and across nearby ages for fathers, are independent of their observability conditional on family characteristics. This ensures that survey attrition or non-reporting is not systematically associated with unobserved income determinants, ruling out selection bias. Second, it imposes an overlap condition guaranteeing sufficient data coverage across individuals and age windows, preventing estimates from being driven by specific reporting patterns or missing subpopulations. Together, these conditions prevent two key threats to validity: estimates being distorted either by systematic missingness (e.g., concentrated among low-income families) or by over-reliance on narrow age clusters. When satisfied, they ensure that inference is driven by income dynamics rather than data availability.\par
According to equation (\ref{eq:beta_new1}), the IGE depends on two distinct components: the covariance between parent and child income and the covariance within parental income, which implies that identification requirements differ across generations. For children, unconfoundedness needs only to hold for single income observations since the IGE exploits contemporaneous parent-child pairs, whereas for fathers, stronger conditions on income tuples are required to capture the temporal structure of their income process.\par
\begin{assumption1}{2}{NP}(Missing At Random)\label{as:unc_np} \begin{enumerate}[i.] \item The missingness of children's annual income $Y_{ct}$ is as good as random once we control for $\bm X_{ct}$\begin{align*} Y_{ct}&\perp D_{ct}\mid \bm X_{ct}, \quad t=1,..., T. \end{align*}
\item Given family characteristics, there is both missing and non-missing children incomes for every age \begin{align*} 0<&p\left(D_{ct}=1\mid \bm X_{ct} \right)<1 \quad a.s, \quad t=1,..., T. \end{align*} \item The missingness of parental annual income pairs $\left(Y_{ft}, Y_{fj}\right) $ is as good as random once we control for $\bm X_{ftj}$\begin{align*} \left(Y_{ft}, Y_{fj}\right) \perp \left(D_{ft}, D_{fj}\right) \mid \bm{X}_{ftj}, \quad \text{ for all } t - j > h> 0,
\end{align*} where $\bm{X}_{ftj}$ are the family characteristics predictive of parental income covariance between years $t$ and $j$, and $\bm{X}_{ftj}\coloneqq \bm X_{ft}$ for $j=t$. \item Given family characteristics, there is both missing and non-missing parental incomes for every age and its neighboring ages \begin{align*} 0 < p\left(D_{ft}=1, D_{fj}=1 \mid \bm{X}_{ftj}\right) < 1 \quad \text{a.s.}, \quad \text{ for all } t - j > h> 0. \end{align*}\end{enumerate}\end{assumption1}
Assumption \ref{as:unc_np} imposes a missing-at-random structure for child and parental incomes and a positivity condition for identification, similar to the conditional independence assumptions in \citet{angrist1995identification}. The assumption that income missingness in the PSID is missing at random is supported by empirical evidence. \cite{fitzgerald1998analysis} finds that attrition in the PSID is selective, primarily affecting lower socioeconomic individuals and those with unstable earnings, marriage, and migration histories, but these factors explain little of the overall attrition, and regression-to-the-mean effects mitigate selection bias. This conclusion is reinforced by \cite{lillard1998panel}, who find that ignoring attrition induces only very mild biases in household income models.\par
\cite{fitzgerald2011attrition} examines attrition in intergenerational models of health, education, and earnings, finding that sibling correlations in outcomes are marginally higher among individuals who remain in the panel longer, though the differences are not statistically significant. Models of intergenerational links with covariates show negligible attrition bias for females. In contrast, the evidence for males is mixed but generally weak, suggesting that conditioning on observables largely mitigates selective attrition. The study finds little evidence of attrition bias, though analyses of educational and earnings outcomes for men appear to benefit from conditioning on observables.\par
\cite{schoeni2015implications} show that applying sample weights reduces differences in intergenerational income elasticity estimates between the full sample, the attriting sample, and the non-attriting sample, rendering these differences statistically insignificant. Their findings highlight that attrition, particularly higher among lower-income individuals, is influenced by the correlation between child and parental income outcomes, emphasizing the importance of incorporating both parental and child characteristics in analyses of intergenerational mobility.\par
Taken together, the literature suggests that the MAR assumption for income missingness in the PSID is empirically plausible, provided analyses carefully account for relevant observables. To address the concerns raised by \cite{schoeni2015implications}, particularly the influence of the correlation between parental and child income outcomes on observability, the analysis incorporates both parental and child characteristics in the conditioning set, thereby strengthening the plausibility of the MAR assumption.\par
The following Theorem establishes the nonparametric identification of the intergenerational elasticity in the presence of incomplete income data and family characteristics. This fundamental result ensures that estimates derived from the identification result are comparable across studies, providing a building block to analyze intergenerational mobility under valid inference.
\begin{theorem}\label{thm:2} Under assumptions \ref{as:ortho_np} and \ref{as:unc_np} and the definition of permanent income in equation (\ref{eq:def_perm_inc}), the IGE is nonparametrically identified as
\begin{gather}\label{eq:identification}
\beta_0=\frac{\mathbb{E}\left[\sum_{t=1}^T\left(\mu_{ct}\left(\bm{X}_{ct},1\right)-\mu_c^P\right)\sum_{j=1}^T\left(\mu_{fj}\left(\bm{X}_{fj},1\right)-\mu_f^P\right)\right]}{ \mathbb{E}\left[\sum_{|t-j| \leq h}\sigma_{tj}\left(\bm{X}_{ftj},1,1\right)+ \sum_{|t-j| > h}\left(\mu_{ft}\left(\bm{X}_{ft},1\right)-\mu_f^P\right)\left(\mu_{fj}\left(\bm{X}_{fj},1\right)-\mu_f^P\right)\right]},
\end{gather} where $\mu_{gt}(\bm{X}_{gt},1) \coloneqq \mathbb{E}\left[ Y_{gt} \mid \bm{X}_{gt}, D_{gt} = 1 \right]$, $\mu_g^P \coloneqq \mathbb{E}\left[ \sum_{t=1}^T \mu_{gt}(\bm{X}_{gt},1) \right]$, and $\sigma_{tj}\left(\bm{X}_{ftj},1,1\right) \coloneqq \mathbb{E}\left[\left(Y_{ft}-\mu_f^P\right)\left(Y_{fj}-\mu_f^P\right)\mid \bm X_{ftj}, D_{ft}=1, D_{fj}=1\right]$.
\end{theorem}
Theorem \ref{thm:2} establishes the identification of the intergenerational elasticity in the presence of incomplete income data. Specifically, it shows that, under the conditional mean independence and orthogonality conditions in Assumption \ref{as:ortho_np} and the standard missing-at-random assumptions in \ref{as:unc_np}, the IGE can be recovered from conditional expectations, including the conditional income profiles of parents and children, and the conditional covariance matrix of parental income.\par
To the best of my knowledge, the only existing identification result in this framework is that of \cite{an2022nonparametric}, who nonparametrically identify the mobility function relating children's to parents permanent income. While their more general framework nests the linear IGE as a special case, since they leave the relationship of parental and child incomes unspecified, my approach offers three important advantages. First, I relax their classical errors-in-variables model (Assumption \ref{as:1} with $\lambda_t=1$) for two measurement periods, by exploiting the definition of permanent income. Second, I relax the assumption that transitory shocks to children's income are uncorrelated with parental permanent income and parental transitory shocks. In contrast, I assume that the prediction error of the children's annual income is uncorrelated to parental permanent income conditional on family characteristics (Assumption \ref{as:ortho_np}). Finally, the proposed framework explicitly addresses the missing data structure inherent in real-world income observations, while incorporating all available information on both income dynamics and family characteristics.\par
Our identification result provides two valuable contributions to the study of intergenerational mobility. First, it resolves persistent methodological challenges by establishing sufficient conditions for identifying the intergenerational elasticity from incomplete income observations and family characteristics. Second, it provides the theoretical foundation for constructing a consistent estimator, enabling researchers to obtain valid and comparable estimates of the intergenerational elasticity.
\subsection{Locally Robust Estimation of the IGE}
I estimate the intergenerational elasticity $\beta_0$ using a Generalized Method of Moments (GMM) approach based on Theorem \ref{thm:2}. The analysis proceeds by rearranging equation (\ref{eq:identification}) to derive the moment condition used to identify $\beta_0$:
\fontsize{10}{12}\selectfont
\begin{align}\label{eq:ident} \nonumber \mathbb{E}\left[g_1\left(W,\gamma,\beta,\mu_c^P,\mu_f^P\right)\right]&=0, \\
\nonumber
g_1\left(W,\gamma,\beta,\mu_{c}^P,\mu_{f}^P\right)&=\beta\sum_{|t-j| \leq h} \sigma_{tj}\left(\bm X_{ftj}, 1, 1\right)+\beta\sum_{|t-j| > h}\left(\mu_{ft}\left(\bm X_{ft}, 1\right)-\mu_f^P\right)\left(\mu_{fj}\left(\bm X_{fj}, 1\right)-\mu_f^P\right)\\
&-\sum_{t=1}^T\left(\mu_{ct}\left(\bm X_{ct}, 1\right)-\mu_c^P\right)\sum_{j=1}^T\left(\mu_{fj}\left(\bm X_{fj}, 1\right)-\mu_f^P\right),
\end{align}
\normalsize where $\gamma\coloneqq (\sigma_{tj}, \mu_{ft}, \mu_{fj})$. Thus, the moment identifying the IGE depends on the income profiles and parental income covariance structure, captured by the nuisance parameter $\gamma$, as well as the population mean permanent incomes $\left(\mu_c^P, \mu_f^P\right)$. These means are themselves identified by the moment conditions (see equation (\ref{eq:cond_means})):
\begin{align*}\nonumber
\mathbb{E}\left[g_2\left(W,\gamma,\mu_c^P\right)\right]&=0,\quad
g_2\left(W,\gamma,\mu_c^P\right)=\sum_{t=1}^T\mu_{ct}\left(\bm X_{ct}, 1\right)-\mu_c^P,\\
\mathbb{E}\left[g_3\left(W,\gamma,\mu_f^P\right)\right]&=0,\quad
g_3\left(W,\gamma,\mu_f^P \right)=\sum_{t=1}^T\mu_{ft}\left(\bm X_{ft}, 1\right)-\mu_f^P.
\end{align*}
Finally, I define the augmented parameter $\theta\coloneqq \left(\beta,\mu_c^P,\mu_f^P\right)$, and combine the moment conditions into a single system for GMM estimation
\begin{align*}
g\left(W,\gamma,\theta\right)=\left(
g_1\left(W,\gamma,\theta\right) \
g_2\left(W,\gamma,\theta\right) \
g_3\left(W,\gamma,\theta\right)\right)'.
\end{align*} \par
Because the nuisance parameter $\gamma$ is unknown, a natural two-step estimation procedure is to first estimate the conditional expectations in $\gamma$, and then perform GMM estimation based on $g\left(W,\hat{\gamma},\theta\right)$. As in any two-step procedure, errors in the first step affect inference in the second. This issue is especially pronounced when machine learning (ML) is used, because regularization and model selection allow for bias to attain smaller variance. As a result, bias from the first step propagates to the second. This is formally captured by the sensitivity of the moment condition to small changes in the nuisance parameter:
\[
\frac{d}{d\tau} \mathbb{E}[g(W, \gamma_\tau, \theta)]\big |_{\tau=0} \neq 0,
\] indicating that the moment identifying $\theta$ is not locally robust to estimation error in $\gamma_0$. \par
\cite{chernozhukov2022locally} provide a general procedure for constructing orthogonal moment functions for GMM, where the moment conditions are locally insensitive to first-step estimation (see Appendix \ref{sec:illust} for an illustration). This property ensures that the resulting estimator is locally robust, meaning that estimation errors in the first step have no effect, locally, on the estimation of the parameter of interest. The authors show that an orthogonal (locally robust) moment function $\psi$ can be constructed by augmenting the identifying moment function $g$ with a correction term $\phi$:
\begin{align*}
\psi\left(W,\gamma,\alpha,\theta\right)=g\left(W,\gamma,\theta\right)+\phi\left(W,\gamma,\alpha,\theta\right),
\end{align*} where $\alpha$ encompasses additional nuisance parameters introduced by $\phi$.\par
\begin{proposition}\label{prop:1}
There exists a function $\phi$ such that the augmented moment condition
\begin{align*}
\psi\left(W,\gamma,\alpha,\theta\right)=g\left(W,\gamma,\theta\right)+\phi\left(W,\gamma,\alpha,\theta\right)
\end{align*} identifies the intergenerational elasticity, as well as the population mean of permanent incomes, and satisfies local robustness:
\[
\frac{d}{d\tau} \mathbb{E}[\psi(W, \gamma_\tau,\alpha, \theta)]\big |_{\tau=0} = 0,
\] where $\psi$ is given by equation (\ref{eq:phi_ref}). This latter property ensures that the estimator of the intergenerational elasticity is first-order insensitive to estimation error in the nuisance parameter $\gamma$.
\end{proposition}
The locally robust closed-form solution for the IGE follows directly from Proposition \ref{prop:1}. In particular, solving for $\beta$ in the orthogonal moment condition yields:
\begin{gather}\label{eq:lr_cf} \beta=\frac{\mathbb{E}\left[\sum_{t=1}^T\left(\mu_{ct}\left(\bm X_{ct}, 1\right)-\mu_c^P\right)\sum_{j=1}^T\left(\mu_{fj}\left(\bm X_{fj}, 1\right)-\mu_f^P\right)\right]+\mathbb{E}\left[\phi_4+\phi_5\right]}{\mathbb{E}\left[\sum_{|t-j| \leq h} \sigma_{tj}\left(\bm X_{ftj}, 1, 1\right)+\sum_{|t-j| > h}\left(\mu_{ft}\left(\bm X_{ft}, 1\right)-\mu_f^P\right)\left(\mu_{fj}\left(\bm X_{fj}, 1\right)-\mu_f^P\right)\right]+\mathbb{E}\left[\phi_1+\phi_2+\phi_3\right]},
\end{gather} where
\begin{align*}\nonumber \phi_1&=\beta \sum_{|t-j| \leq h}\frac{D_{ft}D_{fj}}{p\left(D_{ft}=1, D_{fj}=1 |\bm X_{ftj}\right)}\left(\left(Y_{ft}-\mu_f^P\right)\left(Y_{fj}-\mu_f^P\right) -\sigma_{tj}\left(\bm X_{ftj}, 1, 1\right)\right),\\
\phi_2&=\sum_{|t-j| > h}\left(\mu_{fj}\left(\bm X_{fj}, 1\right)-\mu_f^P\right)\frac{D_{ft}}{p\left(D_{ft}=1|\bm X_{ft}\right)}\left(Y_{ft}-\mu_{ft}\left(\bm X_{ft}, 1\right)\right),\\
\phi_3&=\sum_{|t-j| > h}\left(\mu_{ft}\left(\bm X_{ft}, 1\right)-\mu_f^P\right)\frac{D_{fj}}{p\left(D_{fj}=1|\bm X_{fj}\right)}\left(Y_{fj}-\mu_{fj}\left(\bm X_{fj}, 1\right)\right),\\
\phi_4&=\sum_{t=1}^T\sum_{t=j}^T\left(\mu_{ct}\left(\bm X_{ct}, 1\right)-\mu_c^P\right)\frac{D_{fj}}{p\left(D_{fj}=1|\bm X_{fj}\right)}\left(Y_{fj}-\mu_{fj}\left(\bm X_{fj}, 1\right)\right),\\
\phi_5&=\sum_{t=1}^T\sum_{t=j}^T\left(\mu_{fj}\left(\bm X_{fj}, 1\right)-\mu_f^P\right)\frac{D_{ct}}{p\left(D_{ct}=1|\bm X_{ct}\right)}\left(Y_{ct}-\mu_{ct}\left(\bm X_{ct}, 1\right)\right), \\\mu_c^P&=\mathbb{E}\left[\sum_{t=1}^T\mu_{ct}\left(\bm X_{ct}, 1\right)\right]+\mathbb{E}\left[\sum_{t=1}^T\frac{D_{ct}}{p\left(D_{ct}=1|\bm X_{ct}\right)}\left(Y_{ct}-\mu_{ct}\left(\bm X_{ct}, 1\right)\right)\right], \\
\mu_f^P&=\mathbb{E}\left[\sum_{t=1}^T\mu_{ft}\left(\bm X_{ft}, 1\right)\right]+\mathbb{E}\left[\sum_{t=1}^T\frac{D_{ft}}{p\left(D_{ft}=1|\bm X_{ft}\right)}\left(Y_{ft}-\mu_{ft}\left(\bm X_{ft}, 1\right)\right)\right].\\
\end{align*}
\par
The locally robust closed-form solution for $\beta$ in equation (\ref{eq:lr_cf}) corresponds to the expression given in Theorem \ref{thm:2}, augmented with correction terms that make it first-order insensitive to estimation errors in the nuisance parameters. The term $\phi_1$ corrects for errors in estimating the conditional covariance of parental income for closely spaced periods ($|t-j|\leq h$), while $\phi_2$ and $\phi_3$ address errors in estimating parental income profiles for more distant periods ($|t-j|>h$). Similarly, $\phi_4$ and $\phi_5$ correct for errors in estimating the conditional income profiles of children and parents, respectively. A critical feature of all correction terms ($\phi_1$ to $\phi_5$) is their inherent adjustment for non-random missingness by weighting prediction errors by the inverse propensity score. Finally, the closed-form expressions for the population means of permanent incomes ($\mu_c^P$ and $\mu_f^P$) also incorporate the corresponding prediction errors, ensuring that the estimator remains locally robust to first-step estimation mistakes.\par
Equation (\ref{eq:lr_cf}) motivates the definition of the augmented parameter
\(\theta \coloneqq (\beta, \mu_c^P, \mu_f^P)\) rather than including \(\mu_c^P\) and \(\mu_f^P\) in the nuisance parameter \(\gamma\). Each population mean \(\mu_g^P\) depends not only on the conditional income profiles \(\mu_{g,t}\) but also on the underlying population distribution. Consequently, small changes in the population distribution affect \(\mu_g^P\) both through the conditional profiles and through the expectation itself. Thus, including \(\mu_g^P\) in \(\gamma\) would therefore make the closed-form solution for \(\beta\) considerably more complex. By keeping \(\mu_c^P\) and \(\mu_f^P\) in \(\theta\), we separate the estimation of population permanent means from the first-step nuisance functions, which makes the locally robust solution more tractable.\par
To construct a debiased machine learning estimator for the IGE, I use the orthogonal moment condition in Proposition \ref{prop:1} (see equation (\ref{eq:phi_ref})) combined with cross-fitting to ensure robustness and mitigate overfitting. Following \citet{semenova2023inference}, cross-fitting in settings with dependence should be performed at the level of independent sampling units, in this case, families, rather than individual child--father pairs. Accordingly, let $f \in \{1, \dots, n_f\}$ index families, with $\mathcal{P}_f$ denoting the set of all child--father pairs in family $f$. The set of family indices is partitioned into $L$ mutually exclusive and exhaustive folds $\{\mathcal{F}_\ell\}_{\ell=1}^L$. For each fold $\ell = 1, \dots, L$, the nuisance parameters $\hat{\gamma}^{(\ell)}$ and $\hat{\alpha}^{(\ell)}$ are estimated using only data from families not in $\mathcal{F}_\ell$, thereby preserving independence between the samples used for first-stage estimation and those used for evaluation.
The debiased moment function is then computed as
\[
\hat{\psi}(\theta) = \frac{1}{n} \sum_{\ell=1}^L \sum_{f \in \mathcal{F}_\ell} \sum_{i \in \mathcal{P}_f} \sum_{(t,j) \in \mathcal{J}_i} \hat{\psi}_{i,tj}^{(\ell)},\quad \hat{\psi}_{i,tj}^{(\ell)} \coloneqq g\big(W_{i,tj}, \hat{\gamma}^{(\ell)}, \theta\big) + \phi\big(W_{i,tj}, \hat{\gamma}^{(\ell)}, \hat{\alpha}^{(\ell)}, \theta\big),
\]
where $\mathcal{J}_i$ denotes the set of all tuples $(t, j)$ observed for child--father pairs $i$, noting that a family may contribute multiple such pairs. Since the system is exactly identified, there is no need to compute fold-specific $\hat{\theta}^{(\ell)}$. The locally robust estimator of the IGE is thus obtained by solving
\[
\hat{\theta}^{LR}_n = \arg\min_{\theta \in \Theta \subset \mathbb{R}^3} \hat{\psi}(\theta)'\hat{\Upsilon}\hat{\psi}(\theta).
\]where $\hat{\Upsilon}$ is a positive semi-definite weighting matrix, and $\Theta$ denotes the set of parameter values. This objective function incorporates orthogonal moments and cross-fitting. While the influence function corrects for prediction errors in estimating the conditional expectations, cross-fitting eliminates overfitting in nuisance parameter estimation. Furthermore, by grouping folds at the family level, this approach aligns with the principle of leaving out dependent \say{neighbor} units in panel settings \citep{semenova2023inference}, ensuring that dependence within families does not bias the orthogonalization step.
\subsection{Asymptotic Theory and Inference}\subsubsection{Asymptotic Properties of the Locally Robust Estimator}
To provide rigorous justification for the empirical implementation of the proposed estimator, its large-sample behavior is examined. I begin by establishing consistency, which follows from standard M-estimation theory, adapted to the locally robust framework of \citet{chernozhukov2022locally}. While their main asymptotic results assume consistency, Theorem A3 provides primitive conditions under which it holds. The following Lemma adapts these conditions to the proposed setting.
\begin{lemma}[Consistency of the Locally Robust Estimator]\label{lem:consistency}
Let \( \hat{\theta}^{LR}_n \) be the solution to the cross-fitted orthogonal moment condition:
\begin{align*}
\hat{\theta}^{LR}_n &= \arg\min_{\theta \in \Theta\subset \mathbb{R}^3} \hat{\psi}(\theta)'\hat{\Upsilon}\hat{\psi}(\theta),\quad
\hat{\psi}(\theta) = \frac{1}{n} \sum_{\ell=1}^L \sum_{f \in \mathcal{F}_\ell} \sum_{i \in \mathcal{P}_f} \sum_{(t,j) \in \mathcal{J}_i} \hat{\psi}_{i,tj}^{(\ell)},\\ \hat{\psi}_{i,tj}^{(\ell)} &\coloneqq g\big(W_{i,tj}, \hat{\gamma}^{(\ell)}, \theta\big) + \phi\big(W_{i,tj}, \hat{\gamma}^{(\ell)}, \hat{\alpha}^{(\ell)}, \theta\big),
\end{align*} where $\hat{\Upsilon}$ is a positive semi-definite weighting matrix.
Then \( \hat{\theta}^{LR}_n \overset{p}{\to} \theta_0 \), by Theorem A3 in \citet{chernozhukov2022locally}, provided Assumptions \ref{as:ortho_np}, \ref{as:unc_np} and \ref{ass:clr} hold.
\end{lemma}
Lemma \ref{lem:consistency} shows that under mild regularity conditions \( \hat{\theta}^{LR}_n=\left(\hat{\beta}^{LR}_n, \hat{\mu}_{c,n}^P, \hat{\mu}_{F,n}^P\right) \) converges in probability to the true parameter \( \theta_0 \). The consistency of the locally robust estimator guarantees that, under the specified conditions, the estimated intergenerational elasticity \( \hat{\beta}^{LR}_n \) converges to the true value \( \beta_0 \) as the sample size increases. This ensures that the estimator remains stable even when machine learning methods are used to estimate nuisance parameters. As a result, the estimates of the intergenerational elasticity are both reliable and comparable across different studies.\par
Under the regularity conditions described in Appendix \ref{sec:as_p}, I establish the asymptotic normality of the proposed estimator, which explicitly accounts for uncertainty from the first-stage estimation of the nuisance parameters. This yields confidence intervals with valid coverage, a crucial requirement for drawing meaningful conclusions about intergenerational mobility patterns.\par
The following Lemma formalizes the validity of inference for the estimator \( \hat{\theta}^{LR}_n \), even when nuisance components are estimated using high-dimensional or nonparametric methods. This robustness is achieved through the use of orthogonal moment conditions, which ensure that estimation errors in the first stage enter the moment function only at second order. As a result, standard $\sqrt{n}$ asymptotic normality can be established under relatively weak conditions. Crucially, cross-fitting plays a central role in mitigating own-observation bias and avoids the need for stringent entropy or Donsker-type conditions, which are not known to hold for many machine learning first steps. Together, these features allow us to utilize flexible first-stage methods while maintaining valid inference.
\begin{lemma}[Asymptotic Normality of the Locally Robust Estimator]\label{lemma:asymptotic_normality}
Under Assumptions \ref{as:ortho_np}-\ref{ass:lr_new} and \ref{ass:lr8}, $\hat{\theta}^{LR}_n\xrightarrow{p}\theta_0$, and non-singularity of $G' \Upsilon G$, the asymptotic normality of the estimator \( \hat{\theta}^{LR}_n \) directly follows from Theorem 9 of \cite{chernozhukov2022locally}. Specifically, we have:
\[
\sqrt{n}(\hat{\theta}^{LR}_n - \theta_0) \xrightarrow{d} \mathcal{N}(0, V),
\]
where\( V = \left(G' \Upsilon G\right)^{-1} \), \( G = \mathbb{E}[\partial_\theta g(W, \gamma, \alpha, \theta)] \), and $\hat{\Upsilon}$ is the estimated efficient weighting matrix defined as $\hat{\Upsilon} = \hat{\Psi}^{-1}$ for $\hat{\Psi} = \frac{1}{n} \sum_{\ell=1}^L \sum_{i \in \mathcal{I}_\ell} \sum_{(tj) \in \mathcal{J}_i} \hat{\psi}_{i,tj}^{(\ell)}\hat{\psi}_{i,tj}^{(\ell)'}$.
In addition, if Assumption \ref{ass:lr_other} holds, then \( \hat{V} \xrightarrow{p} V \).
\end{lemma}
\par
Lemma \ref{lemma:asymptotic_normality} completes the theoretical framework by integrating the three contributions: (i) the nonparametric identification of the intergenerational elasticity in the presence of incomplete income data; (ii) a consistent, locally robust estimator that corrects for first-step prediction errors; and (iii) valid inference that accounts for uncertainty from the first-stage estimation of nuisance parameters. Appendix \ref{sec:as_p} characterizes the asymptotic variance
$V$ associated with this result. Together, these advances provide a theoretically grounded toolkit for studying income persistence through the lens of the intergenerational elasticity. \par
The asymptotic normality result in Lemma \ref{lemma:asymptotic_normality} is derived under the assumption of independently and i.i.d observations. In practice, however, datasets commonly include multiple children from the same family, introducing correlation within families. Accordingly, the asymptotic variance $V$ should be estimated using a cluster-robust approach that accounts for this dependence structure. While the i.i.d. assumption is adopted here for ease of exposition and to align with the general theoretical framework of \cite{chernozhukov2022locally}, the core identification and estimation strategy remains sound. The cluster-robust extension is a straightforward implementation detail for the variance estimation, where the moment functions $\hat{\psi}_{i,tj}$ is aggregated at the family level before constructing the variance-covariance matrix $\hat{\Psi}$, accounting for correlation within families in the standard errors.\par
\par
\subsection{Testing the Identification Assumptions}\label{sec:test} This section develops formal hypothesis tests for Assumption \ref{as:ortho_np}.$ii$ and discusses how to assess in practice \ref{as:ortho_np}.$iii.$. As illustrated In Section \ref{sec:app} the specification of the characteristics predictive of income profiles and parental income covariance, namely, \(\bm{X}_{ct}\), \(\bm{X}_{ft}\), and \(\bm{X}_{ftj}\), can be designed to satisfy Assumption \ref{as:ortho_np}.$i$ by construction. The MAR conditions in Assumptions \ref{as:unc_np}.$i$ and \ref{as:unc_np}.$iii$ are not directly testable from the observed data, as they involve unobserved missingness mechanisms. Nevertheless, as discussed above, the literature suggests that the MAR assumption for income missingness in the PSID is empirically plausible, provided that analyses carefully account for relevant observables. Consistent with the findings in \citet{schoeni2015implications}, I include both child and father characteristics in the conditioning set, thereby strengthening the plausibility of the MAR assumption in the analysis. Finally, the boundedness condition on the propensity score in Assumptions \ref{as:unc_np}.$ii$ and \ref{as:unc_np}.$iv$ can be assessed informally through visual inspection.\par
I start by considering a test for the orthogonality between children's prediction errors and parental permanent income \begin{align}\label{eq:test1} H_0: \frac{1}{T}\sum_{t=1}^T\mathbb{E}\left[\epsilon_{ct}Y_{f}^P\right]=0, \quad \quad \text{vs}\quad H_1: \frac{1}{T}\sum_{t=1}^T\mathbb{E}\left[\epsilon_{ct}Y_{f}^P\right] \neq 0,
\end{align}
where $\epsilon_{ct}\coloneqq Y_{ct}-\mathbb{E}\left[Y_{ct}\mid \bm X_{ct}\right]$ denotes the children's income prediction errors at time $t$ and $Y_f^P$ represents parental permanent income. The main challenge in testing this hypothesis is that both random variables are unobserved, and their machine learning estimation introduces regularization and model selection bias when testing $H_0$. To address these issues, a three-stage procedure is proposed. Establishing identification of the object of interest \(\theta_{cf} \coloneqq \frac{1}{T}\sum_{t=1}^T\mathbb{E}\left[\epsilon_{ct} Y_f^P\right]\). Next, constructing a locally robust estimator, and finally providing a $t-$test based on $\hat{\theta}_{cf}$.\par
A locally robust $t-$test for $H_0$ in (\ref{eq:test1}) is given by
\begin{align*} t_{cf,n} &= \frac{\hat{\theta}_{cf,n}}{\sqrt{\hat{V}_{cf,n}/n}}, \end{align*} where $\hat{\theta}_{cf,n}$ is the argument solving the cross-fitted locally robust moment in equation (\ref{eq:cf}), and $\hat{V}_{cf,n}$ is a consistent estimator of the asymptotic variance of $\hat{\theta}_{cf,n}$, that accounts for dependence within families (see Appendix \ref{sec:as_test} for the step-by-step derivation). Similar to the locally robust estimator for the IGE, this cluster-robust variance estimator is constructed by aggregating moment functions at the family level to allow for arbitrary correlation between observations from the same family, while maintaining independence across different families.\par
The following Theorem establishes the asymptotic properties of this locally robust $t-$test.
\begin{theorem}\label{th:3} (Size, Consistency, and Local Power of the Locally Robust $t-$Test I)\label{cor:wald_consistency_local} Under Assumptions \ref{as:ortho_np_n}, \ref{as:unc_np_new}, \ref{ass:4lr}-\ref{ass:lr8} and \ref{ass:clr}, the asymptotic properties of the locally robust $t$ statistic \begin{align*}
t_{cf,n} &= \frac{\hat{\theta}_{cf,n}}{\sqrt{\hat{V}_{cf,n}/n}}
\end{align*} are given by the following statements:
\begin{enumerate} \item (\emph{Asymptotic size}) Under $H_0:\theta_{cf0}=0$, \[
t_{cf,n} \xrightarrow{d} \mathcal{N}(0,1) \quad\text{and}\quad \lim_{n\to\infty}\Pr\left(|t_{cf,n}| > z_{1-\alpha/2}\right) = \alpha,
\]
where $z_{1-\alpha/2}$ is the $(1-\alpha/2)$-quantile of the standard normal distribution. \item (\emph{Consistency under fixed alternatives}) For any fixed alternative with $\theta_{cf0} \neq 0$,
\[
\lim_{n\to\infty}\Pr\!\left(|t_{cf,n}| > z_{1-\alpha/2}\right) = 1.
\] \item (\emph{Local alternatives}) Under $H_{1n}:\theta_{cf0} = \delta/\sqrt{n}$ with fixed $\delta \in \mathbb{R}$,
\[
t_{cf,n} \xrightarrow{d} \mathcal{N}\left(\frac{\delta}{\sqrt{V_{cf}}}, 1\right),
\]
so the limiting power is
\[
\lim_{n\to\infty}\Pr\!\left(|t_{cf,n}| > z_{1-\alpha/2}\right)
= 2\left[1 - \Phi\left(z_{1-\alpha/2} - \frac{|\delta|}{\sqrt{V_{cf}}}\right)\right]
> \alpha \quad \text{whenever } \delta \neq 0,
\]
where $\Phi(\cdot)$ denotes the standard normal cumulative distribution function.\end{enumerate}
\end{theorem}\par
While a similar locally robust test could be constructed for Assumption~\ref{as:ortho_np}.$iii$, implementing such a test faces a fundamental empirical constraint: jointly observed income pairs $(Y_{ft}, Y_{fj})$ for distant periods ($|t-j|>h$) are systematically sparse in longitudinal data, undermining the reliability and power of formal hypothesis testing. Instead, this assumption can be assessed in practice by analyzing the autocorrelation structure and variance decomposition of parental income prediction errors.
Since the variance of permanent income decomposes as $
\text{Var}(Y_f^P) = \text{Var}\big(\mathbb{E}[Y_{ft} \mid \mathbf{X}_{ft}]\big) + \text{Var}(\epsilon_{ft}),$ the contribution of distant-lag residual products $\sum_{|t-j|>h} \mathbb{E}[\epsilon_{ft}\epsilon_{fj}]$ to the IGE denominator depends on both their contribution to residual variance and the relative importance of residual variance itself. \par
I suggest an empirical assessment in two steps. first, examining whether the autocorrelation function $\text{Corr}(\epsilon_{ft},\epsilon_{fj})$ exhibits rapid decay. Second, quantifying the weighted contribution of distant lags to the residual variance structure. Rapid decay combined with small variance contribution provides transparent evidence that the orthogonality assumption is empirically plausible and has a negligible impact on IGE estimates.
Empirically, one should examine whether the autocorrelation function $\text{Corr}(\epsilon_{ft},\epsilon_{fj})$ decays to low levels beyond lag $h$, and what fraction of the residual variance structure is attributable to distant lags, computed as the weighted contribution $\sum_{k>h} (T-k) \mathbb{E}[\epsilon_{ft}\epsilon_{fj}]$ relative to the total. If both the autocorrelation decay is rapid and the variance contribution from distant lags is negligible, this provides evidence that imposing the orthogonality assumption has negligible impact on IGE estimates. \par While the proposal offers a straightforward approach to selecting $h$ based on observed autocorrelation patterns and variance contributions, a more data-driven procedure could be developed by adapting bandwidth selection methods from the time series literature to this setting. I leave such extensions for future research and implement the empirical assessment procedure in the application.
\section{Simulations}\label{sec:sims}
I consider the following data generating process (DGP) for generation $g$ at age $t$: \fontsize{10.5}{12.5}
\begin{align*}
Y_{gt}&=\gamma_{0,g}+\gamma_{1,g}X_{1,g}+\gamma_{2,g}X_{2,g}+\gamma_{3,g}t+\gamma_{4,g}t^2+\gamma_{5,g}X_{1,g}t+\epsilon_{gt},\quad t=1,...,T, \quad g\in\{c,f\},\\
\epsilon_{gt}&\sim \mathcal{N}\left(0,\sigma^2_\epsilon\right)\\
\begin{pmatrix}X_{j,c} \\ X_{j,f}
\end{pmatrix} &\sim \mathcal{N}\left(\begin{pmatrix}0 \\0
\end{pmatrix}, \begin{pmatrix}1 & \sigma_j \\\sigma_j & 1\end{pmatrix}\right), \quad j=1,2.
\end{align*}\normalsize
Using the definition of permanent income:
\begin{align}\label{eq:perm_sim}
Y_g^P&= \gamma_{0,g} + \gamma_{1,g} X_{1,g} + \gamma_{2,g} X_{2,g} + \gamma_{3,g} \bar{t} + \gamma_{4,g}\bar{t^2} + \gamma_{5,g} X_{1,g}\bar{t} + \bar{\epsilon}_{g}, \quad \quad g\in\{c,f\}, \\\nonumber
\bar{t}&=\frac{1}{T}
\sum_{t=1}^T t, \quad \bar{t^2}=\frac{1}{T}\sum_{t=1}^T t^2, \quad \bar{\epsilon}_{g}=\frac{1}{T}\sum_{t=1}^T\epsilon_{gt},
\end{align}
so the covariance between permanent incomes is given by \fontsize{10.5}{12.5}
\begin{align}\label{eq:covsim}
\text{Cov}\left(Y_c^P, Y_f^P\right) &=\text{Cov}\left(\gamma_{1,c} X_{1,c}, \gamma_{1,f} X_{1,f}\right) + \text{Cov}\left(\gamma_{2,c} X_{2,c}, \gamma_{2,f} X_{2,f}\right) + \text{Cov}\left(\gamma_{5,c} X_{1,c} \bar{t}, \gamma_{5,f} X_{1,f} \bar{t}\right)\\\nonumber&+ \text{Cov}\left(\gamma_{1,c} X_{1,c}, \gamma_{5,f} X_{1,f} \bar{t}\right) + \text{Cov}\left(\gamma_{5,c} X_{1,c} \bar{t}, \gamma_{1,f} X_{1,f}\right)\\\nonumber&= \gamma_{1,c} \gamma_{1,f} \sigma_1 + \gamma_{2,c} \gamma_{2,f} \sigma_2 + \gamma_{5,c} \gamma_{5,f} \bar{t^2} \sigma_1 + \gamma_{1,c} \gamma_{5,f} \bar{t} \sigma_1 + \gamma_{5,c} \gamma_{1,f} \bar{t} \sigma_1,\end{align}\normalsize where we have used that the covariates come from a bivariate normal distribution with zero mean and correlation $\sigma_j$ among generations.\par
According to equation (\ref{eq:perm_sim}), the variance of parental income corresponds to
\begin{align}\label{eq:varsim}
\text{Var}\left(Y_f^P\right) = \gamma_{1,f}^2 + \gamma_{2,f}^2 + \gamma_{5,f}^2 \bar{t^2} + 2 \gamma_{1,f} \gamma_{5,f} \bar{t} + \sigma_\epsilon^2/T.
\end{align}
Finally, by plugging equations (\ref{eq:covsim}) and (\ref{eq:varsim}) into (\ref{eq:beta}) yields
\begin{align}\label{eq:beta_sim}
\beta_0&=\frac{\mathbb{E}\left[\left(Y_c^P-\mathbb{E}\left[Y_c^P\right]\right)\left(Y_f^P-\mathbb{E}\left[Y_f^P\right]\right)\right]}{\mathbb{E}\left[\left(Y_f^P-\mathbb{E}\left[Y_f^P\right]\right)^2\right]}\nonumber\\
&=\frac{ \sigma_1 \left(\gamma_{1,c} \gamma_{1,f}+\gamma_{5,c} \gamma_{5,f} \bar{t^2}+\left(\gamma_{1,c} \gamma_{5,f}+\gamma_{5,c} \gamma_{1,f} \right)\bar{t}\right)+ \gamma_{2,c} \gamma_{2,f} \sigma_2}{\gamma_{1,f}^2 + \gamma_{2,f}^2 + \gamma_{5,f}^2 \bar{t^2} + 2 \gamma_{1,f} \gamma_{5,f} \bar{t} + \sigma_\epsilon^2/T}.
\end{align} Setting the parameter values to
\begin{align*}
\gamma_{0,c} &= 8.5, & \gamma_{0,f} &= 5, &
\gamma_{1,c} &= 0.275, & \gamma_{1,f} &= 0.4, \\
\gamma_{2,c} &= 0.2, & \gamma_{2,f} &= 0.25,&
\gamma_{3,c} &= 0.4, & \gamma_{3,f} &= 0.5, \\
\gamma_{4,c} &= -0.005, & \gamma_{4,f} &= -0.0045, &
\gamma_{5,c} &= 0.01, & \gamma_{5,f} &= 0.015, \\
\sigma_1 &= 0.75, & \sigma_2 &= 0.75, &
\sigma_\epsilon &= 1, & t &= 20, \dots, 60,
\end{align*} yields $T=41$, and $\beta_0=0.50$ according to equation (\ref{eq:beta_sim}).
\par
The simulation considers sample sizes $n$ = 100, 500, 1000, and 2000, with 20\%, 35\%, and 50\% of each sample randomly selected via income snapshots from parents and children. First, for each individual $i$, a contiguous observation period length is drawn from a right-censored Poisson distribution:
\begin{align*}
\ell_i = \min\left(\max(OW_i^P, 2), 41\right), \quad OW_i^P \sim \text{Pois}(\lambda),
\end{align*}
where $\lambda \approx 10$, $20$, andor $31$ years for 20\%, 35\%, and 50\% coverage respectively. Then, the observation window begins at a random age:
\begin{align*}
a_i \sim \mathcal{U}\left(20,\ 60-\ell_i+1\right)
\end{align*}
ensuring complete coverage within the 20-60 age range. Thus, only incomes satisfying $t \in [a_i, a_i+\ell_i)$ are observed, with other years being missing (completely at random), mimicking common data limitations in mobility studies. This creates contiguous observation blocks that mimic real-world data limitations where income histories may only be observed during certain life periods, mimicking realistic administrative or survey-based data constraints. The sampling is performed separately for children and parents.\par
I assess the performance of the Locally Robust (LR) estimator by examining its bias and coverage properties relative to three alternative approaches: (1) the plug-in machine learning estimator, (2) the mid-life income estimator, and (3) the life-cycle estimator. Income profiles are estimated for both generations using XGBoost Regression, which also allows us to compute the conditional covariance of parental income. Propensity scores are estimated via logistic regression. The core difference between the locally robust (LR) and plug-in machine learning estimators lies in their moment conditions: the LR estimator uses a Neyman-orthogonal moment that incorporates the influence function of the first-stage estimates, while the plug-in estimator relies on the uncorrected identifying moment. For the mid-life income estimator, fathers’ permanent income is proxied by averaging earnings from ages 30 to 40, and children’s income is based on a single mid-life earnings draw. In contrast, the life-cycle estimator uses the same paternal income proxy but estimates children’s permanent income as the average of predicted earnings over the life cycle from a correctly specified OLS regression. For the LR and ML estimators, I proceed in two steps: hyperparameter tuning using 5-fold cross-validation, followed by cross-fitting to prevent overfitting. In all simulations, 500 Monte Carlo replications are used. \par
Table \ref{tab:1} presents the finite-sample performance of four estimators for the intergenerational elasticity, evaluated through bias and coverage rates across 500 Monte Carlo replications. The true IGE is 0.5, with a nominal coverage rate of 0.95. The analysis spans three sample sizes (\(n = 100, 500, 1000, 2000\)) and three observation probabilities (\(\kappa = 0.20, 0.35, 0.50\)).\par
\begin{table}[ht]
\caption{Bias and coverage of different estimators for the IGE for different sample sizes and observation probability.}
\centering\resizebox{0.95\textwidth}{!}{
\begin{threeparttable}
\begin{tabular}{l cc cc cc cc}
\hline
& \multicolumn{2}{c}{Locally Robust} & \multicolumn{2}{c}{Plug-in Machine Learning} & \multicolumn{2}{c}{Life-cycle} & \multicolumn{2}{c}{Mid-life} \\ \cmidrule(lr){2-3} \cmidrule(lr){4-5} \cmidrule(lr){6-7} \cmidrule(lr){8-9}
$n$ & Bias & Coverage & Bias & Coverage & Bias & Coverage & Bias & Coverage \\
\hline \hline
$\kappa$ = 0.20 & & & & & & & & \\
\hline
100 & -0.00506 & 0.91 & -0.07358 & 0.56 & -0.17380 & 0.36 & -0.19890 & 0.86 \\
500 & -0.00347 & 0.95 & -0.03991 & 0.43 & -0.16430 & 0.00 & -0.18762 & 0.43 \\
1000 & -0.00129 & 0.92 & -0.02981 & 0.41 & -0.16310 & 0.00 & -0.18462 & 0.17 \\
2000 & -0.00132 & 0.93 & -0.03414 & 0.11 & -0.16420 & 0.00 & -0.18699 & 0.01 \\
\hline \hline
$\kappa$ = 0.35 & & & & & & & & \\
\hline
100 & -0.00685 & 0.91 & -0.05702 & 0.66 & -0.1540 & 0.27 & -0.18318 & 0.75 \\
500 & -0.00314 & 0.94 & -0.02800 & 0.66 & -0.14970 & 0.00 & -0.17540 & 0.17 \\
1000 & -0.00230 & 0.93 & -0.03077 & 0.35 & -0.14960 & 0.00 & -0.17099 & 0.01 \\
2000 & 0.00109 & 0.92 & -0.01376 & 0.67 & -0.14900 & 0.00 & -0.16894 & 0.00 \\
\hline \hline
$\kappa$ = 0.50 & & & & & & & & \\
\hline
100 & -0.00317 & 0.91 & -0.04778 & 0.72 & -0.16108 & 0.12 & -0.15682 & 0.73 \\
500 & 0.00561 & 0.93 & -0.01547 & 0.84 & -0.14557 & 0.00 & -0.15302 & 0.12 \\
1000 & 0.00460 & 0.93 & -0.01928 & 0.67 & -0.14741 & 0.00 & -0.15481 & 0.01 \\
2000 & 0.00306 & 0.92 & -0.00949 & 0.82 & -0.13861 & 0.00 & -0.15465 & 0.00 \\
\hline
\end{tabular}
\begin{tablenotes}
\footnotesize
\item Results based on 500 Monte Carlo replications with true IGE equal to 0.5 and the nominal coverage is 0.95.
\end{tablenotes}
\end{threeparttable}}
\label{tab:1}
\end{table}
The locally robust estimator exhibits superior performance, with bias decreasing as sample size increases (e.g., from $-0.0051$ at \(n=100\) to $-0.0015$ at \(n=2000\) for \(\kappa=0.20\)). This aligns with expected \(\sqrt{n}\)-consistency, reflecting its robustness to sample size variations. Coverage rates remain close to the nominal 0.95, ranging from 0.91 to 0.95 across all scenarios, with minor undercoverage at smaller sample sizes (\(n=100\)). Notably, both bias and coverage are largely insensitive to changes in \(\kappa\), indicating stability across varying observation probabilities.\par
In contrast, the plug-in machine learning estimator has substantially higher bias in absolute terms (e.g., $-0.0736$ at \(n=100\) vs. $-0.0095$ at \(n=2000\) for \(\kappa=0.50\)). Its coverage rates are consistently below the nominal 0.95, improving from 0.56 to 0.82 as sample size increases for \(\kappa=0.50\), but remaining inadequate. This poor performance underscores the limitations of the plug-in approach, particularly in smaller samples or lower observation probabilities, justifying the preference for the locally robust estimator.\par
The life-cycle and mid-life estimators exhibit substantial and persistent bias across all sample sizes and \(\kappa\) values, with LC bias ranging from $-0.174$ to $-0.139$, and MI bias from $-0.199$ to $-0.155$. Their coverage rates deteriorate to zero for larger samples due to miss-centered confidence intervals: as sampling variability decreases, the intervals narrow around biased point estimates, missing the true IGE. The similar performance of LC and MI estimators in these simulations reflects two factors. First, although the LC estimator addresses the children’s life-cycle bias, both estimators rely on the same mid-life proxies for parental permanent income, so neither fully accounts for parental life-cycle bias. Second, the simple DGP generates limited heterogeneity in children's income growth by family background, attenuating the LC estimator's designed advantage. In the PSID application, where such heterogeneity is pronounced, the LC estimator substantially outperforms the MI approach (see Table \ref{tab:ige_cohort} and Figure \ref{fig:three_plots}). These simulations thus primarily validate the locally robust estimator's properties rather than a comprehensive comparison of all methods.\par
Overall, the locally robust estimator emerges as the most reliable, offering low bias and near-nominal coverage across all settings. The plug-in machine learning estimator, while improving with larger samples, remains inferior due to higher bias and poor coverage. The life-cycle and mid-life estimators are consistently outperformed, highlighting the importance of consistent and locally robust estimation in analyzing the IGE.\par
\section{Consistent Estimation of the Intergenerational Elasticity in the United States}\label{sec:app}
In this section, I implement the locally robust estimator to measure the intergenerational elasticity of income in the United States. The analysis employs the Panel Study of Income Dynamics, the world’s longest-running longitudinal household survey. Launched in 1968 with a nationally representative sample of 5,000 U.S. families (over 18,000 individuals), the PSID has continuously tracked these families and their descendants, collecting rich data on income, wealth, employment, education, health, and other socioeconomic outcomes.\par
The analysis focuses on birth cohorts spanning 1954 to 1977, using rolling 10-year windows, yielding 15 overlapping cohort samples. This cohort selection ensures sufficient observations for both parents and children within the available PSID data span (1968-2023). Following \cite{lee2009trends}, I use the PSID core sample, corresponding to the Survey Research Center component, and define the income measure as family income, which allows us to include both male and female children. Individuals with only zero or missing income values are excluded. All dollar values are adjusted to 1968 dollars using the CPI. To handle nonpositive incomes, they are bottom-coded at the sample 1st percentile, which affects 0.12\% of the observations in the raw PSID data.\par
Following \cite{mazumder2016estimating}, I focus on the lifetime span from ages 25 to 55 (31 years), corresponding to the core working-life period. This age range is chosen to ensure adequate sample coverage for both generations within the PSID timeframe while capturing income during the stable working years. The 25-55 window allows us to observe sufficient years of income for parents (who are typically 25 years older than their children, see Table \ref{tab:sum}) while avoiding periods dominated by educational transitions at younger ages or retirement decisions at older ages. Sample sizes range from 1,038 to 1,103 child-father pairs across 574 to 757 families, with 11,747 to 19,694 child observations and 16,802 to 25,400 father observations (see Table \ref{tab:sample_sizes} for detailed counts by cohort window).\par
The family characteristics in the analysis are drawn from the rich data provided by the PSID and are organized into several domains. Education is measured by years of schooling completed and whether the household head reported having additional training beyond standard school or college. Regional location follows the PSID’s classification into Northeast, North Central, South, or West. Family structure includes the birth order of the children, the father’s age at first birth, and the PSID’s intergenerational mapping is used to incorporate the number of offspring per father. Assets are captured through indicators of housing and business ownership. Demographics include race (classified as White or Non-White), sex of the children (given the focus on fathers), religion, and age at the time of interview. While the PSID offers a broader set of variables, the analysis focuses on these selected characteristics to ensure consistency and availability across survey waves.\par
Based on this available data, I construct the characteristics predictive of income profiles and parental income covariance, namely, \(\bm{X}_{ct}\), \(\bm{X}_{ft}\), and \(\bm{X}_{ftj}\). This specification must account for three key requirements: handling missing data in observables, incorporating the dynamics of the income process, and satisfying Assumptions \ref{as:ortho_np} and \ref{as:unc_np}. To address the first, the variables are summarized over the lifetime span (ages 25–55) using averages for time-varying characteristics (excluding education), modes for religion and region, and the maximum value for education.\par
To accurately model yearly income as a function of covariates, it is essential to incorporate the dynamics of the income process and capture relevant empirical patterns. To this end, the covariate specification in \cite{mello2022lifecycle} is closely followed. Specifically, a quartic polynomial in age is included to capture concavities and nonlinearities in income profiles. For the child generation, a noisy proxy for parental permanent income is incorporated, defined as the three-year average of log family income when the child was aged 15–17. If family income data for this period are unavailable, the closest available three-year window within ages 15–17 is used. Additionally, interactions between this noisy proxy and parental education with a quadratic polynomial in age are included to capture the greater variability in income growth at younger ages and the typically faster income growth among children from high-income families. \par
Our covariate specification is designed to satisfy Assumption \ref{as:ortho_np}.$i$ by construction, while also making the remaining assumptions plausible in practice, although not guaranteed to hold. To satisfy Assumption \ref{as:ortho_np}.$i$, $\bm X_{ft}$ and $\bm X_{cj}$ are merged such that their non-overlapping components $\left(\bm X_{ft}\cap \bm X_{cj}\right)^c=\{age_{ft},age_{cj}\}$, are both deterministic. In doing so, I also include in \(\bm{X}_{ft}\) the noisy measure of parental permanent income along with its interaction with age. This time-invariant measure serves as a relevant predictor in the presence of missing income data, helping to compensate for the absence of income leads and lags. Moreover, it enhances the first-step estimation of income profiles (and parental income covariance), which is fundamentally a prediction task. Finally, to construct \(\bm{X}_{ftj}\), \(\bm{X}_{ft}\) and \(\bm{X}_{fj}\) are merged, ensuring $\left(\bm X_{ft}\cap \bm X_{ftj}\right)^c=\{age_{fj}\}$. The current specification of the covariates ensures that Assumption \ref{as:ortho_np}.$i$ is satisfied by construction; that is, the covariates used to predict children’s income satisfy:
$\mathbb{E}\left[Y_{ct} \mid \bm{X}_{ct}, \bm{X}_{cj}, \bm{X}_{fj}\right] = \mathbb{E}\left[Y_{ct} \mid \bm{X}_{ct}\right]$ for $t,j=1,\ldots,T,$
and those used to predict fathers’ income satisfy $
\mathbb{E}\left[Y_{ft} \mid \bm{X}_{ft}, \bm{X}_{ftj}, \bm{X}_{cj}\right] = \mathbb{E}\left[Y_{ft} \mid \bm{X}_{ft}\right]$ for $\bm{X}_{fj} \subset\bm{X}_{ftj}, t,j=1,...T.$ This follows from the complements of the intersections, $\left(\bm X_{ft}\cap \bm X_{ft}\right)^c=\{age_{ft},age_{cj}\}$ and $\left(\bm X_{ft}\cap \bm X_{ftj}\right)^c=\{age_{fj}\}$, consisting solely of age, which is deterministic and thus do not add stochastic variation beyond what is captured by $\bm X_{ct}$ and $\bm X_{ft}$.\par
Our covariate specification further enhances the plausibility of the remaining assumptions in practice. The inclusion of an interaction between parental permanent income and age, for example, strengthens the orthogonality condition between children's income prediction errors and parental permanent income required by Assumption \ref{as:ortho_np}.$ii$. In Section \ref{sec:test}, I develop a test to empirically evaluate Assumption \ref{as:ortho_np}.$ii$. As regards, Assumption \ref{as:ortho_np}.$iii$, I outline how to construct a similar test but assess this assumption in practice by examining the estimated autocorrelation function, which is more reliable given the sparse availability of income pairs over long time spans..\par
The dimensionality of the covariate set reflects the trade-off between the plausibility of the missing-at-random assumption and the boundedness of the propensity scores in Assumption \ref{as:unc_np}. While increasing the dimension of $\bm X_{gt}$, can make the conditional independence $Y_{ct}\perp D_{ct}|\bm X_{ct}$ more plausible, it may reduce the likelihood that the propensity score remains bounded away from zero. Nonetheless, it is not merely the dimensionality, but rather the informativeness of the covariates that determines whether missingness is conditionally at random. In other words, we aim to control for the relevant features such that, conditional on them, the missingness of annual income occurs conditionally at random. While the MAR conditions in Assumptions \ref{as:unc_np}.$i$ and \ref{as:unc_np}.$iii$ cannot be directly tested from observed data, the boundedness conditions in Assumptions \ref{as:unc_np}.$ii$ and \ref{as:unc_np}.$iv$ can be assessed informally through visual inspection.\par
\begin{table}[ht]
\centering
\caption{Summary Statistics for Children and Fathers}\label{tab:sum}
\resizebox{0.95\textwidth}{!}{
\begin{tabular}{lccccc|ccccc}
\toprule\hline &
\multicolumn{5}{c|}{\textbf{Children}} & \multicolumn{5}{c}{\textbf{Fathers}} \\
\cmidrule(r){1-6} \cmidrule(l){7-11}
Variable & Mean & Median & SD & Min & Max & Mean & Median & SD & Min & Max \\
\midrule
Annual Income & 9.22 & 9.31 & 0.90 & -1.42 & 13.8 & 9.30 & 9.36 & 0.87 & -1.53 & 12.9 \\
Proxy of Permanent Income & - & - & - & - & - & 9.33 & 9.41 & 0.72 & -1.06 & 12.1 \\
Education Level & 14.10 & 14.00 & 2.18 & 0.00 & 17.0 & 13.00 & 12.00 & 3.01 & 0.00 & 17.0 \\
House Ownership & 0.64 & 0.73 & 0.34 & 0.00 & 1.00 & 0.78 & 0.93 & 0.31 & 0.00 & 1.00 \\
Business Ownership & 0.16 & 0.00 & 0.25 & 0.00 & 1.00 & 0.18 & 0.00 & 0.28 & 0.00 & 1.00 \\
Additional Training & 0.22 & 0.00 & 0.42 & 0.00 & 1.00 & 0.15 & 0.00 & 0.36 & 0.00 & 1.00 \\
Religion & 0.86 & 1.00 & 0.35 & 0.00 & 1.00 & 0.88 & 1.00 & 0.32 & 0.00 & 1.00 \\
White & 0.92 & 1.00 & 0.28 & 0.00 & 1.00 & 0.91 & 1.00 & 0.29 & 0.00 & 1.00 \\
Sex & 0.51 & 1.00 & 0.50 & 0.00 & 1.00 & - & - & - & - & - \\
Birth Order & 2.11 & 2.00 & 1.40 & 1.00 & 12.0 & - & - & - & - & - \\
Northeast Region & 0.20 & 0.00 & 0.40 & 0.00 & 1.00 & 0.20 & 0.00 & 0.40 & 0.00 & 1.00 \\
South Region & 0.34 & 0.00 & 0.48 & 0.00 & 1.00 & 0.32 & 0.00 & 0.47 & 0.00 & 1.00 \\
West Region & 0.18 & 0.00 & 0.38 & 0.00 & 1.00 & 0.17 & 0.00 & 0.37 & 0.00 & 1.00 \\
Age at First Child & - & - & - & - & - & 25.40 & 25.00 & 4.88 & 13.00 & 50.0 \\
\hline\hline
\end{tabular}
}
\end{table} Table~\ref{tab:sum} provides summary statistics for socioeconomic characteristics of children and their fathers, revealing important intergenerational patterns. Fathers exhibit higher mean logged annual income (9.30 vs.\ 9.22) and home ownership rates (78\% vs.\ 64\%), while children show greater educational attainment (mean 14.1 vs. 13 years) and additional training participation (22\% vs.\ 15\%). The income measures exhibit slightly tighter dispersion for fathers, with smaller standard deviations (0.87 vs.\ 0.90 for annual income). Educational attainment shows greater variability among fathers (SD 3.01 vs.\ 2.18), potentially reflecting cohort differences in educational access. Both generations share nearly identical white composition (91\% vs.\ 92\%), religious affiliation (88\% vs.\ 86\%), and business ownership (18\% vs.\ 16\% for children). Half of the children are female, reflecting a balanced gender distribution. The median birth order indicates that most families in the data have two or more children, with relatively few only children. The mean and median values for the proxy of permanent income closely match those of annual income, but with less variation and a narrower range. Most fathers had their first child around age 25. Regional distributions show similar patterns across generations, with children slightly more concentrated in the South (34\% vs.\ 32\%) and both generations showing identical Northeast representation (20\%).\par
To assess the assumption that the children's prediction errors are uncorrelated with parental lifetime income, I implement the proposed locally robust test. This condition is crucial for unbiased estimation: if prediction errors systematically correlate with parental income, the IGE estimates will be biased. Results are shown in Figure \ref{fig:res_test} and Table \ref{tab:cohort_estimates}. In 14 of the 15 cohort windows, the LR test fails to reject the null hypothesis of zero covariance between children’s income prediction errors and parental permanent income, with an average $p$-value of 0.39. This finding provides empirical evidence that the life-cycle estimator of \cite{mello2022lifecycle} effectively addresses life-cycle bias from the children's side by properly accounting for how income trajectories vary with family background. Importantly, this empirical validation of the orthogonality condition strengthens confidence in both the identification strategy and the resulting IGE estimates.\par
\begin{figure}
\centering
\includegraphics[width=0.9\linewidth]{Figures/cohort_test_results_annotated.pdf}
\caption{Covariance of Children Prediction Errors and Parental Permanent Income Across Cohorts}
\label{fig:res_test}
\end{figure}
Our second orthogonality assumption requires that the average covariance of parental income prediction errors for distant periods is negligible. To assess this, I analyze the autocorrelation structure and variance decomposition of these residuals. Since the variance of permanent income (the denominator of the IGE) decomposes as $\text{Var}(Y_f^P) = \text{Var}\left(\mathbb{E}[Y_{ft} \mid \mathbf{X}_{ft}]\right) + \text{Var}(\epsilon_{ft})$,
the contribution of distant-lag residual products to the total variance depends on both their contribution to residual variance and the relative importance of residual variance itself. Figure \ref{fig:autocor} exhibits rapid autocorrelation decay across all 15 birth cohort windows: autocorrelations fall from 0.607 at lag 1 to 0.086 at lag 10 and stabilize at 0.072 for lags 11--13. Figure~\ref{fig:second_figure} shows that lags 0--10 account for 95.4\% of the residual variance, while lags 11--13 contribute only 4.6\%. The analysis examines autocorrelations through lag 13 because observation pairs become increasingly sparse at distant lags. Unmeasured lags beyond 13 would contribute even less due to both the observed monotonic decay and their mechanically lower weights in the variance structure. These findings strongly support the orthogonality assumption for $h=10$, which I adopt for implementing the proposed estimator.\par
\begin{figure}[htbp] \centering
\begin{subfigure}[b]{0.48\textwidth} \centering
\includegraphics[width=\linewidth]{Figures/autocorrelation_all_cohorts_overlay_viridis.pdf} \caption{Father Income Residual Autocorrelation by Birth Cohort Window} \label{fig:autocor}
\end{subfigure} \hfill \begin{subfigure}[b]{0.48\textwidth} \centering \includegraphics[width=\linewidth]{Figures/variance_contribution_grouped.pdf} \caption{Decomposition of Residual Income Variance by Lag Range} \label{fig:second_figure}
\end{subfigure}
\caption{Dependence of Parental Income Prediction Errors}
\end{figure}
Having validated the identifying assumptions empirically, I now turn to estimating the intergenerational elasticity using four alternative approaches: the locally robust (LR) estimator, the plug-in machine learning (ML) estimator, the life-cycle (LC) estimator, and the mid-life income (MI) estimator. Following \cite{mello2022lifecycle}, the LC estimator predicts children’s log income using an OLS specification that includes a quartic in age, interactions with education dummies, linear and quadratic interactions with parental income, individual fixed effects, and year fixed effects, with outliers removed. Lifetime income is then constructed as the mean of the exponentiated predicted values for each individual, multiplied by a smearing factor computed as the average of exponentiated residuals within parental income deciles, correcting for retransformation bias in the log-linear specification \citep{wooldridge2013introductory}. Fathers’ permanent income is proxied by the three-year average of log family income when the child was aged 15–17. To implement the mid-life income estimator,the same measure of parental permanent income is used, whereas children’s income is calculated as the three-year average of log income between ages 25 and 33, following \cite{solon1992intergenerational}.\par
The locally robust and plug-in machine learning estimators use the same covariates as the life-cycle estimator, augmented with the family characteristics listed in Table \ref{tab:sum}. Individual fixed effects are omitted as they cannot be accommodated under the cross-fitting procedure.
Propensity scores are estimated using a lasso-logit model, while income profiles for both generations and the conditional covariances of parental income are modeled using XGBoost regression. The key distinction between the LR and plug-in ML estimators lies in their moment conditions: the LR approach utilizes the orthogonal moment that accounts for the first-stage influence function, whereas the plug-in version relies solely on the identifying moment without such correction. For all predictive tasks in the LC, LR, and ML estimators, values are winsorized at the 1st and 99th percentiles.\par
Table \ref{tab:ige_cohort} presents intergenerational elasticity estimates for the United States using PSID core sample data across 15 birth cohorts (1954-1977). The locally robust estimator yields an average IGE of 0.643, systematically exceeding estimates from the plug-in machine learning approach (0.600), the life-cycle estimators (0.500 and 0.502 for the log-average and average-log specifications, respectively), and the standard mid-life income estimator (0.430). These economically meaningful differences underscore the importance of combining proper identification with local robustness for reliable mobility measurement. \par
\begin{table}[ht]
\centering
\caption{Intergenerational Elasticity Estimates by Cohort Using Alternative Estimators}
\label{tab:ige_cohort}
\begin{threeparttable}
\begin{adjustbox}{max width=0.8\textwidth,center}
\begin{tabular}{lccccc}
\toprule
Cohort & LR & ML & LC (log-avg) & LC (avg-log) & MI \\
\midrule
1954-1963 & 0.598 & 0.522 & 0.509 & 0.496 & 0.413 \\
(N = 1,099) & (0.414, 0.782) & (0.397, 0.647) & (0.420, 0.598) & (0.407, 0.585) & (0.317, 0.508) \\
1955-1964 & 0.596 & 0.557 & 0.502 & 0.494 & 0.415 \\
(N = 1,089) & (0.396, 0.796) & (0.415, 0.700) & (0.409, 0.596) & (0.401, 0.587) & (0.318, 0.513) \\
1956-1965 & 0.596 & 0.545 & 0.522 & 0.519 & 0.450 \\
(N = 1,083) & (0.441, 0.752) & (0.414, 0.676) & (0.428, 0.615) & (0.426, 0.612) & (0.351, 0.550) \\
1957-1966 & 0.624 & 0.616 & 0.513 & 0.517 & 0.419 \\
(N = 1,071) & (0.410, 0.837) & (0.482, 0.750) & (0.419, 0.608) & (0.422, 0.611) & (0.322, 0.517) \\
1958-1967 & 0.634 & 0.596 & 0.536 & 0.542 & 0.486 \\
(N = 1,097) & (0.453, 0.814) & (0.471, 0.721) & (0.446, 0.626) & (0.452, 0.631) & (0.394, 0.577) \\
1959-1968 & 0.659 & 0.619 & 0.519 & 0.526 & 0.432 \\
(N = 1,085) & (0.481, 0.836) & (0.520, 0.718) & (0.436, 0.603) & (0.444, 0.609) & (0.342, 0.521) \\
1960-1969 & 0.625 & 0.590 & 0.511 & 0.514 & 0.431 \\
(N = 1,062) & (0.428, 0.823) & (0.503, 0.677) & (0.429, 0.592) & (0.434, 0.595) & (0.347, 0.515) \\
1961-1970 & 0.615 & 0.579 & 0.496 & 0.495 & 0.421 \\
(N = 1,042) & (0.444, 0.785) & (0.499, 0.659) & (0.418, 0.573) & (0.419, 0.571) & (0.341, 0.501) \\
1962-1971 & 0.630 & 0.610 & 0.496 & 0.499 & 0.415 \\
(N = 1,045) & (0.468, 0.792) & (0.532, 0.688) & (0.418, 0.574) & (0.422, 0.576) & (0.333, 0.498) \\
1963-1972 & 0.672 & 0.676 & 0.504 & 0.507 & 0.428 \\
(N = 1,038) & (0.512, 0.831) & (0.589, 0.762) & (0.424, 0.585) & (0.427, 0.586) & (0.345, 0.512) \\
1964-1973 & 0.663 & 0.667 & 0.475 & 0.481 & 0.413 \\
(N = 1,060) & (0.508, 0.817) & (0.580, 0.754) & (0.396, 0.554) & (0.403, 0.559) & (0.330, 0.497) \\
1965-1974 & 0.702 & 0.711 & 0.504 & 0.511 & 0.451 \\
(N = 1,067) & (0.551, 0.854) & (0.623, 0.800) & (0.428, 0.580) & (0.436, 0.586) & (0.368, 0.535) \\
1966-1975 & 0.700 & 0.460 & 0.487 & 0.495 & 0.439 \\
(N = 1,074) & (0.587, 0.813) & (0.386, 0.535) & (0.407, 0.567) & (0.415, 0.574) & (0.361, 0.517) \\
1967-1976 & 0.673 & 0.612 & 0.474 & 0.477 & 0.428 \\
(N = 1,088) & (0.526, 0.819) & (0.541, 0.683) & (0.395, 0.552) & (0.399, 0.555) & (0.353, 0.504) \\
1968-1977 & 0.671 & 0.641 & 0.461 & 0.467 & 0.410 \\
(N = 1,103) & (0.504, 0.838) & (0.563, 0.720) & (0.381, 0.542) & (0.386, 0.547) & (0.336, 0.485) \\
\bottomrule
\end{tabular}
\end{adjustbox}
\begin{tablenotes}
\footnotesize
\item \hspace{12mm} 95\% confidence intervals clustered at the family level are reported in parentheses.
\end{tablenotes}
\end{threeparttable}
\end{table}
Our estimates of the IGE range from 0.6 to 0.7 across cohorts, with an average of 0.64. These results closely align with recent PSID-based studies reviewed by \cite{mazumder2018intergenerational} that attempt to mitigate life-cycle bias. Building on the theoretical framework developed by \cite{haider2006life}, \cite{gouskova2010estimating} produces a bias-corrected estimate of 0.63. Taking a different approach, \cite{chau2012intergenerational} embeds an earnings dynamics model with long time spans of income data and estimates the IGE to exceed 0.6. Most recently, \cite{mazumder2016estimating} leverages the full length of the PSID panel by using up to 15-year averages of fathers' income centered at age 40, estimating the IGE in family income to be greater than 0.6. Our locally robust estimates thus reinforce this convergent evidence from multiple PSID-based methodologies, all pointing to substantially higher intergenerational persistence than the conventional estimates of approximately 0.4. Moreover, we find that the IGE increases slightly across cohorts. The earliest three cohort windows average 0.597, the next six average 0.631, and the latest average 0.680. Despite wide confidence intervals, this 14\% progression provides suggestive evidence of a modest increase in persistence across the 1954-1977 birth cohorts.\par
The naive ML estimator yields an IGE of 0.60 (ranging from 0.46 to 0.71), systematically underestimating the true elasticity. On average, the plug-in ML estimator underestimates the IGE by 0.043 (7\%), though the bias varies considerably across cohorts. In the best case, the bias is only -0.008 (1.2\%). However, in the worst case (1966-1975 cohort), the plug-in approach underestimates the IGE by 34\%, yielding an estimate of 0.46 versus the locally robust estimate of 0.70. This represents a meaningful difference in the assessment of intergenerational mobility. In line with the simulation results (Table \ref{tab:1}), the plug-in ML exhibits considerable undercoverage: its confidence intervals are 41\% narrower on average than the locally robust estimator's, reflecting false precision from neglecting uncertainty in nuisance parameter estimation. \par The differences between the locally robust and plug-in ML estimates and confidence intervals underscore that local robustness corrections have real consequences for empirical conclusions. While conventional machine learning approaches excel at prediction, directly applying them to estimate causal or structural parameters without orthogonalization can lead to non-negligible bias that materially affects our understanding of mobility patterns. The 34\% underestimation observed in the 1966-1975 cohort illustrates that the effects of failing to debias can be substantial, even when the average bias across cohorts appears modest.\par
The standard mid-life income estimator produces an average intergenerational elasticity (IGE) of 0.430. Across the 15 cohort windows, the estimated IGE ranges from 0.410 to 0.486, reflecting a substantial downward bias. On average, the estimator underestimates the true IGE by 0.213, with the bias ranging from 23\% to 39\% across cohorts. This substantial underestimation stems from life-cycle bias in both generations: parental income measured at mid-life imperfectly captures permanent economic status, while children's three-year averages fail to account for differential income growth patterns across family backgrounds.\par
The life-cycle estimator, implemented using the log of average income as originally proposed by \cite{mello2022lifecycle}, represents a substantial improvement over the mid-life income approach, producing an average IGE of 0.500 with estimates ranging from 0.46 to 0.53. By predicting children's income profiles and accounting for heterogeneous income growth across family backgrounds, the LC estimator successfully addresses the bias children's life-cycle bias inherent in the MI approach. The improvement is considerable: while the MI estimator exhibits an average bias of 33\% (ranging from 23\% to 39\%), the LC estimator reduces this to 22\% (ranging from 12\% to 31\%). Nonetheless, the remaining underestimation suggests that the LC estimator, while correcting for children's bias, still relies on the same mid-life income proxies for parents as the MI estimator.\par
To assess the sensitivity of IGE estimates to the workable definition of permanent income, I implement the life-cycle estimator using the average of log income (our workable definition underlying the identification strategy in Theorem \ref{thm:2}). The estimates are nearly identical, with the average-log specification yielding an average IGE of 0.500 (range: 0.467-0.542) compared to 0.503 (range: 0.461-0.535) for the log-average specification. The close correspondence between these estimates—differing by less than 0.4\% on average validates the use of the average-log specification that enables nonparametric identification under the proposed missing data framework.\par
To shed light on the specific sources of bias for each estimator, I break down the IGE into its core components: the covariance of child and parent permanent income (numerator) and the variance of parental permanent income (denominator). Figure \ref{fig:three_plots} presents this decomposition across all estimators and birth cohorts, displaying both the point estimates and their deviations from the locally robust benchmark, which serve as the empirical measure of bias in this sample. Importantly, the nature of this bias differs fundamentally across estimators: the MI and LC estimators are affected by life-cycle bias and measurement error, due to their reliance on mid-life income averages as proxies for parental permanent income, whereas the plug-in ML estimator is susceptible to bias from regularization and model selection in the first-stage machine learning predictions.\par
\begin{figure}[htbp]
\centering
\captionsetup[subfigure]{justification=centering}
\begin{subfigure}[b]{0.8\textwidth}
\centering
\includegraphics[width=\textwidth]{Figures/plot_betas.pdf}
\caption{IGE Estimates by Birth Cohort and Estimator}
\label{fig:betas}
\end{subfigure}
\begin{subfigure}[b]{0.8\textwidth}
\centering
\includegraphics[width=\textwidth]{Figures/plot_covariances.pdf}
\caption{Covariance Estimates by Birth Cohort and Estimator}
\label{fig:covariances}
\end{subfigure}
\begin{subfigure}[b]{0.8\textwidth}
\centering
\includegraphics[width=\textwidth]{Figures/plot_variances.pdf}
\caption{Variance Estimates by Birth Cohort and Estimator}
\label{fig:variances}
\end{subfigure}
\caption{Intergenerational Elasticity, Covariances, and Variance Estimates and Bias Across Birth Cohorts and Estimators}
\label{fig:three_plots}
\end{figure}
The decomposition in Figure \ref{fig:three_plots} reveals a striking pattern that validates the proposed theoretical framework. Panel (a) reproduces the IGE estimates from Table \ref{tab:ige_cohort}, showing the systematic ordering of estimators across cohorts. Panels (b) and (c) expose the underlying mechanisms: the LC estimator performs remarkably well at estimating the covariance between child and parent permanent income, frequently outperforming even the plug-in ML estimator despite the latter's use of flexible machine learning methods. However, the LC estimator systematically overestimates the variance of parental permanent income, with deviations from the locally robust benchmark that are consistently positive across all cohorts.\par
Figure \ref{fig:covariances} provides strong empirical evidence that the LC estimator successfully addresses children's life-cycle bias. The covariance estimates closely track the locally robust benchmark across all cohorts, illustrating that the LC approach accurately captures the child-parent income relationship. This success stems from the estimator's explicit modeling of children's income profiles: by predicting how income evolves over the life cycle and accounting for heterogeneous income growth across family backgrounds, the LC estimator addresses the sensitivity to the age at which child income is measured. Remarkably, the LC estimator frequently outperforms even the plug-in ML estimator on this component, despite the latter's use of flexible machine learning methods. This superior performance is particularly noteworthy: it demonstrates that a well-specified parametric approach, informed by economic theory about income dynamics and family background effects, can outperform flexible machine learning methods in the absence of debiasing.\par
Panel (c) highlights the critical importance of addressing all sources of bias simultaneously. Both the LC and MI estimators systematically overestimate the variance of parental permanent income across cohorts. This arises from both relying on mid-life income averages that suffer from life-cycle bias and measurement error. Thus, the inflated denominator mechanically reduces the IGE estimates, leading to substantial underestimation even when the numerator is correctly measured. These results underscore the practical relevance of addressing all sources of bias: the LC estimator’s average IGE of 0.50 falls well short of the locally robust estimate of 0.64.\par
Overall, the empirical application provides compelling evidence that properly addressing both missing income data and estimation uncertainty is essential for credible inference on intergenerational mobility. The locally robust estimator consistently produces higher intergenerational elasticity estimates across all PSID birth cohorts, indicating lower mobility in the United States than suggested by conventional methods and aligning closely with recent evidence based on long-term income data. At the same time, our results confirm that the life-cycle estimator performs precisely as intended on the children’s side: it effectively corrects for children's life-cycle bias by modeling heterogeneous income growth across family background. However, its reliance on mid-life parental income proxies leaves residual bias from the parental side, leading to systematic underestimation of the IGE. The plug-in machine learning estimator, while flexible and powerful for prediction, fails to correct for first-stage estimation error, resulting in underestimation and undercoverage. Together, these findings illustrate the benefits of local robustness for combining the strengths of structured life-cycle modeling and flexible machine learning to produce reliable and measures of intergenerational mobility.
\section{Conclusions}\label{sec:conc}
This paper addresses the fundamental challenge of measuring the intergenerational transmission of lifetime economic status when researchers observe only snapshots of income at specific ages alongside individual characteristics. Constrained by these data limitations, standard practice estimates the intergenerational elasticity using income averages during mid-life, a procedure that introduces well-documented life-cycle bias. In response, the literature has made substantial progress in refining IGE estimates by separately addressing life-cycle bias from either the parent or child generation. I show that this piecemeal approach hinders comparability of estimates. This insight motivates our core contribution: moving beyond proxy refinement to establish identification conditions that jointly address all sources of bias, enabling consistent and comparable mobility estimates across studies, time, and place.\par
Building on these insights, I formalize that the magnitude of each bias source depends on study design, income dynamics, and sampling rules, causing proxy-based estimators to converge to context-specific parameters rather than the population IGE, which compromises their comparability across studies, time, and place.\par
I establish nonparametric identification of the IGE from incomplete income data and family characteristics under standard missing-at-random assumptions and testable orthogonality conditions. Moving beyond the conventional generalized errors-in-variables model, I adopt a workable definition of permanent income as the average of log annual earnings during working life—a specification that proves empirically robust, as IGE estimates remain virtually unchanged when permanent income is instead defined as the log of average lifetime income. \par
Building on this identification result, I develop a consistent and locally robust estimator by constructing an orthogonal moment function that ensures machine learning estimation of nuisance parameters—such as conditional income expectations and propensity scores—has no first-order effect on the IGE estimate. This debiasing procedure combines Neyman-orthogonal moments with cross-fitting at the family level to prevent overfitting while accommodating the dependent structure inherent in intergenerational data. I establish the estimator's asymptotic normality and construct valid confidence intervals that account for uncertainty introduced by first-stage estimation. Additionally, I develop locally robust tests for the orthogonality assumptions underlying identification, providing researchers with tools to assess the validity of our framework in different empirical contexts.\par
Our framework enables comparable IGE estimates across time and place in the presence of incomplete income data. By addressing the key methodological challenge of not observing lifetime income, our approach establishes a robust foundation for studying income persistence, enhancing the reliability and interpretability of mobility research across diverse economic settings.\par
Our simulation analysis illustrates the sound finite-sample performance of the locally robust estimator, which exhibits negligible bias that vanishes as sample size increases and coverage rates close to nominal levels across different scenarios. The estimator substantially outperforms alternative approaches: the plug-in machine learning estimator exhibits both higher bias and severe undercoverage due to failing to account for first-stage uncertainty, while traditional proxy-based methods show large and persistent biases that do not diminish with sample size.\par
Our locally robust estimates of the IGE range from 0.60 to 0.70 across cohorts, averaging 0.64. These substantially exceed conventional estimates relying on mid-life income averages, but align closely with recent evidence using long-time averages over mid-career periods, revealing considerably lower U.S. mobility than conventional estimates suggest. I also find a naive plug-in machine learning estimator exhibits considerable bias and severe undercoverage, underscoring that empirical relevance on constructing Neyman-orthogonal moments for both eliminating regularization bias and achieving valid inference.\par
Our study highlights three important directions for future research. First, our identification strategy requires longitudinal data that are often unavailable in developing countries, where understanding income persistence is most relevant. Future work should establish alternative identification results for data-scarce environments. Second, the shift in the literature toward rank-based measures has been partly motivated by concerns about the nonlinear relationship between log child income and log parent income. Accordingly, future research should study identification and locally robust estimation for a nonlinear version of the intergenerational elasticity. One promising direction involves estimating the regression of child permanent income on parent permanent income in levels. A quantile-specific elasticity can then be constructed by multiplying the marginal effect at each point in the parental income distribution by the ratio of average child to parent income at that quantile. This would provide a richer, distributional perspective on income persistence and allow researchers to quantify how mobility varies across the income ladder patterns while avoiding the limitations of log-linear specifications.\par
Finally, a central empirical challenge in economics is that many key parameters, from models of life-cycle income, savings, and consumption to measures of individual well-being, depend on lifetime outcomes, while only partial observations at certain ages are typically available. The methods proposed in this paper open the door for applications in other contexts with partially observed outcomes, providing a flexible framework for addressing similar empirical challenges.
\bibliography{cites}