EconBase
← Back to paper

Counterfactual Copula and Its Application to the Effects of College Education on Intergenerational Mobility

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

59,938 characters · 9 sections · 56 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Counterfactual Copula and Its Application to the Effects of College Education on Intergenerational Mobility

abstractThis paper proposes a nonparametric estimator of the counterfactual copula of two outcome variables that would be affected by a policy intervention. The proposed estimator allows policymakers to conduct ex-ante evaluations by comparing the estimated counterfactual and actual copulas as well as their corresponding measures of association. Asymptotic properties of the counterfactual copula estimator are established under regularity conditions. These conditions are also used to validate the nonparametric bootstrap for inference on counterfactual quantities. Simulation results indicate that our estimation and inference procedures perform well in moderately sized samples. Applying the proposed method to studying the effects of college education on intergenerational income mobility under two counterfactual scenarios, we find that while providing some college education to all children is unlikely to promote mobility, offering a college degree to children from less educated families can significantly reduce income persistence across generations. {\bf Keywords}: copula, counterfactual policy effect, intergenerational mobility {\bf JEL classifications}: C14, C31, J62

Introduction

Counterfactual analysis is an important method in policy evaluation, especially when neither a policy nor its quasi-experiment has yet been implemented. Recent studies, for example Rothe2010, ChernozhukovFernandez-ValEtAl2013a, and Hsu2022, focus on the counterfactual distribution of a scalar outcome variable affected by an exogenous policy intervention. However, a change in a policy variable may also cause a change in dependence structure between two outcome variables of interest. For example, in the intergenerational mobility literature, college education has been regarded as a “great equalizer” to reduce the correlation between incomes across generations by offering equal opportunities to children regardless of the advantages of origins (Torche2011, Torche2011 and Chetty2020, Chetty2020). Accordingly, policymakers would wish to evaluate ex ante to what extent income mobility could be improved if the higher education system were further expanded.

Motivated by such counterfactual policy evaluation, we are concerned with its effect on a bivariate copula function of the two outcome variables

align*[align* omitted — 82 chars of source]

where $(Y_1,Y_2)^\top$ is a two-dimensional vector of outcome variables that are continuously distributed, $X$ is a $d$-dimensional vector of covariates, $(\varepsilon_1,\varepsilon_2)^\top$ is individual unobserved heterogeneity in an arbitrary measurable space of unrestricted dimensionality, and $(g_1,g_2)^\top$ are functions that are unknown to researchers. Policymakers attempt to exogenously manipulate the covariate from $X$ to $X^*$, and this manipulation leads to the counterfactual outcome variables

align*[align* omitted — 98 chars of source]

but the functions $(g_1,g_2)^\top$ and unobserved heterogeneity $(\varepsilon_1,\varepsilon_2)^\top$ are unaffected by the exogenous manipulation. We call the copula $C_{Y_1^*Y_2^*}$ of $(Y^*_1,Y^*_2)^\top$ the counterfactual copula, in contrast to the actual copula $C_{Y_1Y_2}$, which is the copula of $(Y_1,Y_2)^\top$ in the status quo. The policy effects of interest could be the difference in copula, $C_{Y_1^*Y_2^*}-C_{Y_1Y_2}$, and the difference in some measure of association, $\nu(C_{Y_1^*Y_2^*})-\nu(C_{Y_1Y_2})$, where $\nu$ is a functional defined on the collection of all copulas. For example, the change in Spearman's rho between bivariate outcomes could be the main concern, as it corresponds to the rank-rank slope, a commonly used measure of intergenerational mobility in economics (Chetty2014, Chetty2014). In this case, the specific functional is $\nu_{\rho}:C\mapsto 12\int_{[0,1]^2} u_1u_2\d C(u_1,u_2)-3$.

We first show identification of the counterfactual copula under standard assumptions and adopt the principle of analogy to propose a nonparametric estimator $\widehat{C}_{Y_1^*Y_2^*}$ of the counterfactual copula in the presence of a random sample $\{(Y_{1i},Y_{2i},X_i,X_i^*)\}_{i=1}^n$. This proposed estimator, together with the empirical copula estimator $\widehat{C}_{Y_1Y_2}$ of the actual copula, allows researchers to evaluate the effects of a policy before its implementation. We then establish the joint weak convergence of $(\sqrt{n}(\widehat{C}_{Y_1^*Y_2^*}-C_{Y_1^*Y_2^*}),\sqrt{n}(\widehat{C}_{Y_1Y_2}-C_{Y_1Y_2}))^\top$ and $\sqrt{n}((\nu(\widehat{C}_{Y_1^*Y_2^*})-\nu(\widehat{C}_{Y_1Y_2}))-(\nu(C_{Y_1^*Y_2^*})-\nu(C_{Y_1Y_2})))$, provided that $\nu$ is Hadamard differentiable. Since the limiting processes would have complicated covariance structures with unknown nuisance functions, we also validate the nonparametric bootstrap to draw inference on counterfactual quantities of interest. Results of the simulation show that the counterfactual quantities of interest are $\sqrt{n}$-consistent, which is in line with our theoretical analysis.

Viewing our method as an essential complement to the existing toolkit for policy evaluation, we apply it to evaluating the effects of college education on intergenerational income mobility in the United States. We use data obtained from the Panel Study of Income Dynamics (PSID) and consider two counterfactual scenarios. In the first scenario, children with low educational attainment are assigned more years of schooling as if extra “compulsory” schooling years were required. We find that conferring a college degree on all children can substantially reduce income persistence by about one-third of the magnitude of the considered association measure, including Spearman's rho, Kendall's tau, Gini's gamma, and Blomqvist's beta. This finding echoes the sheepskin effects in literature, indicating that a college degree provides additional returns beyond educational attainment alone (Hungerford1987, Hungerford1987; Jaeger1996, Jaeger1996). From a pragmatic perspective, a policy typically covers some, rather than all, families; additionally, existing empirical results (Peters1992, Peters1992; Sikhova0, Sikhova0) suggest that parental education is a key determinant of offspring income. We thus consider the second scenario in which only children from less educated families are required to complete a four-year college degree. Our empirical results show that income persistence across generations is significantly lower if only children of parents with educational attainment less than, or equal to, ten years of schooling are required to complete a four-year college degree; in this case, families affected by this counterfactual policy account for about 15% of the sample. In contrast, expanding this policy to include families at the upper tail of the parental education distribution would result in only a negligible improvement in income mobility.

The rest of this paper is organized as follows. (ref) discusses the model and relevant measures of association, and provides identification and estimation of the counterfactual copula. In (ref), we derive the joint weak convergence of the counterfactual and empirical copula processes and validate the nonparametric bootstrap for counterfactual quantities. (ref) reports the simulation results, and (ref) presents the empirical study regarding the heterogeneous effects of college education on intergenerational mobility. Finally, (ref) concludes. Technical proofs of theorems and lemmas are deferred to (ref).

Identification and Estimation

As described in the introduction, we consider two continuous outcome variables

align*[align* omitted — 82 chars of source]

where $X$ is a $d$-dimensional vector of covariates, $\varepsilon_1$ and $\varepsilon_2$ are individual unobserved heterogeneity, and $g_1$ and $ g_2$ are structural functions.\footnote{The nonseparable model could be viewed as the reduced-form equations induced by some underlying structural equations. A simple example is the triangular nonparametric simultaneous equations model

align*[align* omitted — 80 chars of source]

Substituting the second equation into the first yields the reduced-form equations

align*[align* omitted — 138 chars of source]

where $\varepsilon_1=(e_1,e_2)^\top$ and $\varepsilon_2=e_2$. Unlike NeweyPowellEtAl1999, who are concerned with the structural function $\tilde{g}_1$, we are interested in the dependence structure between $Y_1$ and $Y_2$.} This bivariate setting is very flexible. First, the structural functions $g_1$ and $ g_2$ are unknown to researchers, so ad hoc parametric assumptions need not be imposed. In addition, the policy effect could depend on unobserved heterogeneity because both $\varepsilon_1$ and $\varepsilon_2$ are nonseparable. The dimensionality of $(\varepsilon_1,\varepsilon_2)^\top$ is also unspecified, so incorrect inference due to the exclusion of some important heterogeneity can be avoided. Finally, it is possible that some elements of $X$ might be excluded from the structural functions.\footnote{ To see this, denoting $X=(X_0,X_1,X_2)^\top$, we write $Y_1=g_1(X_0,X_1,\varepsilon_1)$ and $Y_2=g_2(X_0,X_2,\varepsilon_2)$ in the case where $X_0$ affects $(Y_1,Y_2)^\top$, $X_1$ affects only $Y_1$, and $X_2$ affects only $Y_2$.}

Suppose that a counterfactual policy, manipulating the covariate from $X$ to $X^*$, causes a change in bivariate outcomes from $(Y_1,Y_2)^\top$ to $(Y^*_1,Y^*_2)^\top$, where

align*[align* omitted — 90 chars of source]

This change in bivariate outcomes in turn causes a change in their copula, and the difference between the counterfactual copula and the actual copula \[ \Delta_{C}(u_1,u_2)\equiv C_{Y_1^*Y_2^*}(u_1,u_2)-C_{Y_1Y_2}(u_1,u_2) \] is called the policy effect on the copula. The change in copula may be accompanied by changes in measures of association; hence, policymakers would be interested in the corresponding policy effect on the association \[ \Delta_{\nu}\equiv \nu(C_{Y_1^*Y_2^*})-\nu(C_{Y_1Y_2}) \] for the following common measures:

enumerate[(i)] • Spearman's rho $\nu_{\rho}:C\mapsto 12\int_{[0,1]^2} u_1u_2\d C(u_1,u_2)-3$; • Kendall's tau $\nu_{\tau}:C\mapsto 4\int_{[0,1]^2} C(u_1,u_2)\d C(u_1,u_2)-1$; • Gini's gamma $\nu_{\gamma}:C\mapsto 2\int_{[0,1]^2} [|u_1+u_2-1|-|u_1-u_2|]\d C(u_1,u_2)$; • Blomqvist's beta $\nu_{\beta}: C\mapsto 4C(0.5,0.5)-1$.

We refer to Nelsen2006 for an excellent overview of these measures.

Since each aforementioned measure of association can be written in terms of some functional of the copula, the identification of those counterfactual policy effects on association is achieved if we can identify the counterfactual copula $C_{Y_1^*Y_2^*}$ and actual copula $C_{Y_1Y_2}$. Sklar1959's (Sklar1959) theorem implies that the actual copula is identified by

align*[align* omitted — 82 chars of source]

Given independent and identically distributed observations $\{(Y_{1i},Y_{2i},X_i,X_i^*)\}_{i=1}^n$, the actual copula of $(Y_1,Y_2)^\top$ is commonly estimated by the empirical copula proposed by Deheuvels1979; to be specific,

equation[equation omitted — 211 chars of source]

where $\widehat F^{-1}_{Y_j}(u)=\inf\{y:\widehat F_{Y_j}(y)\geq u\}$ is the inverse function of the empirical CDF estimator $\widehat{F}_{Y_j}(y)=n^{-1}\sum_{i=1}^n\operatorname{\operatorname{1}}\{Y_{ji}\leq y\}$ for $j\in\{1,2\}$.

To achieve the identification of counterfactual copula, we start by introducing assumptions as follows.\\[0.2cm] Assumption I (Identification)

enumerate[label=I\arabic*] • $(X,X^*)^\top$ and $(\varepsilon_1,\varepsilon_2)^\top$ are independent. • The support of $X^*$ is a subset of the support of $X$. • The counterfactual outcome variables $(Y_1^*,Y_2^*)^\top$ are continuously distributed.

Assumption (ref) requires exogeneity of $(X,X^*)^\top$, which may be strong in some empirical studies, and is sufficient for Assumption 2.3 in Hsu2022.\footnote{ Assumption (ref) could be relaxed by the control function approach but such consideration goes beyond the scope of the present paper.} Assumption (ref) excludes the extrapolation of covariates. Such extrapolation is inappropriate because none of parametric setups is imposed on the functions $g_1$ and $g_2$ outside the support of $X$. Assumptions (ref) and (ref), together with the availability of data on $(Y_1,Y_2,X,X^*)^\top$, allow us to identify the joint counterfactual cumulative distribution function (CDF) $F_{Y_1^*Y_2^*}$ and marginal counterfactual CDFs $(F_{Y_1^*},F_{Y_2^*})^\top$, as shown in Rothe2010 and ChernozhukovFernandez-ValEtAl2013a. Moreover, Assumption (ref) allows us to uniquely express the counterfactual copula $C_{Y_1^*Y_2^*}$ in terms of counterfactual CDFs $(F_{Y_1^*Y_2^*},F_{Y_1^*},F_{Y_2^*})^\top$ by Sklar's theorem. More details of Sklar's theorem can be found in Nelsen2006. We summarize the discussion in the following proposition.

proUnder Assumption {\normalfont I}, the counterfactual copula is identified by \begin{align*} C_{Y_1^*Y_2^*}(u_1,u_2) &=F_{Y_1^*Y_2^*}(F^{-1}_{Y_1^*}(u_1),F^{-1}_{Y_2^*}(u_2))\\ &=\int F_{Y_1Y_2|X}(F^{-1}_{Y_1^*}(u_1),F^{-1}_{Y_2^*}(u_2)|x)F_{X^*}(\d x), \end{align*} where $F^{-1}_{Y_j^*}(u)\equiv\inf\{y:F_{Y_j^*}(y)\geq u\}$ is the inverse function of \[ F_{Y_j^*}(y)=\int F_{Y_j|X}(y|x)F_{X^*}(\d x) \] for every $j\in\{1,2\}$, and \[ F_{Y_{1}^{*}Y_{2}^{*}}(y_{1},y_{2})=\int F_{Y_{1}Y_{2}|X}(y_{1},y_{2}|x)\; F_{X^{*}}(\d x). \]

(ref) shows that the counterfactual copula $C_{Y_1^*Y_2^*}$ is identified without any restriction on the correlation between $\varepsilon_1$ and $\varepsilon_2$.\footnote{If $\varepsilon_1$ and $\varepsilon_2$ are conditionally independent given $X$, then the counterfactual copula can be further simplified to \[ C_{Y_1^*Y_2^*}(u_1,u_2) =\int F_{Y_1|X}(F^{-1}_{Y_1^*}(u_1)|x)F_{Y_2|X}(F^{-1}_{Y_2^*}(u_2)|x)F_{X^*}(\d x). \] In addition, the formula of $F_{Y_j^*}$ can be simplified if not all elements of $X$ are factors of $Y_j$ for $j\in\{1,2\}$.} This proposition also suggests a nonparametric plug-in counterfactual copula estimator

align[align omitted — 190 chars of source]

where for a kernel function $K$ and a bandwidth $h=h_n$ shrinking with the sample size $n$, \[ \widehat{F}_{Y_1Y_2|X}(y_1,y_2|x)=\frac{\sum_{i=1}^n\operatorname{\operatorname{1}}\{Y_{1i}\leq y_1,Y_{2i}\leq y_2\}K\mathopen{}\mathclose\bgroup\originalleft(\frac{X_i-x}{h}\aftergroup\egroup\originalright)}{\sum_{i=1}^nK\mathopen{}\mathclose\bgroup\originalleft(\frac{X_i-x}{h}\aftergroup\egroup\originalright)}, \] and for $j\in\{1,2\}$, $\widehat F^{-1}_{Y_j^*}(u)=\inf\{y:\widehat F_{Y_j^*}(y)\geq u\}$ is the inverse function of \[ \widehat{F}_{Y_j^*}(y)=\frac{1}{n}\sum_{i=1}^n\widehat{F}_{Y_j|X}(y|X_i^*) \] with $\widehat{F}_{Y_1|X}(y|x)=\widehat{F}_{Y_1Y_2|X}(y,\infty|x)$ and $\widehat{F}_{Y_2|X}(y|x)=\widehat{F}_{Y_1Y_2|X}(\infty,y|x)$.

The proposed estimator in (ref) is easy to implement. To see this, we define the joint counterfactual CDF estimator to be

align[align omitted — 187 chars of source]

where the weights $\{W_{i,n}\}_{i=1}^{n}$ are given by

align*[align* omitted — 254 chars of source]

The joint counterfactual estimator in (ref) can be viewed as an extension of Rothe2010's (Rothe2010) marginal counterfactual CDF estimator to the two-dimensional setting; specifically, in this study, we have $\widehat{F}_{Y_1^*}(y_{1})=\widehat{F}_{Y_{1}^{*}Y_{2}^{*}}(y_{1},\infty)$ and $\widehat{F}_{Y_2^*}(y_{2})=\widehat{F}_{Y_{1}^{*}Y_{2}^{*}}(\infty,y_{2})$. The estimator $\widehat{F}_{Y_{1}^{*}Y_{2}^{*}}$, compared with $\widehat{F}_{Y_{1}Y_{2}}$ can be also applied to testing bivariate stochastic dominance, which is an important topic in decision making; see for example DenuitHuangEtAl2014. More importantly, we can rewrite the counterfactual copula estimator in (ref) as

align*[align* omitted — 196 chars of source]

This formula suggests that when evaluating $\widehat{C}_{Y_{1}^{*}Y_{2}^{*}}$ at multiple locations, we only require a one-off computation of weights $\{W_{i,n}\}_{i=1}^{n}$.\footnote{ Note that whenever every weight $W_{i,n}$ is nonnegative, both counterfactual CDF estimators $\widehat{F}_{Y^*_1}$ and $\widehat{F}_{Y^*_2}$ are nondecreasing; in this case, the counterfactual copula estimator $\widehat{C}_{Y_1^*Y_2^*}$ is $2$-increasing because for any $(u_1,u_2), (v_1,v_2)\in [0,1]^2$ with $u_1\leq v_1$ and $u_2\leq v_2$,

align*[align* omitted — 408 chars of source]

In the case where some weights are negative, we are unaware of any study dealing with necessary adjustments to satisfy the $2$-increasing property of a bivariate CDF. This task would be challenging and reserved for future work. } Although the counterfactual copula estimator in (ref) has the same weights $\{W_{i,n}\}_{i=1}^n$ as those in Rothe2010, it has summands with an indicator function evaluated at two estimated arguments $(\widehat{F}_{Y^*_1}^{-1}(u_1),\widehat{F}_{Y^*_2}^{-1}(u_2))^\top$, which introduces nonnegligible estimation effects, as we will see in the next section.

commentSimilarly, the identification of the actual copula is accomplished by Sklar's theorem, stating that $C_{Y_1Y_2}(u_1,u_2)=F_{Y_1Y_2}(F^{-1}_{Y_1}(u_1),F^{-1}_{Y_2}(u_2))$. In literature, the actual copula of $(Y_1,Y_2)^\top$ is commonly estimated by the empirical copula proposed by Deheuvels1979 \begin{equation} \widehat{C}_{Y_1Y_2}(u_1,u_2)\equiv\frac{1}{n}\sum_{i=1}^n\operatorname{\operatorname{1}}\{Y_{1i}\leq \widehat{F}_{Y_1}^{-1}(u_1),Y_{2i}\leq \widehat{F}_{Y_2}^{-1}(u_2)\}, \end{equation} where $\widehat F^{-1}_{Y_j}(u)=\inf\{y:\widehat F_{Y_j}(y)\geq u\}$ is the inverse function of the empirical CDF estimator $\widehat{F}_{Y_j}(y)=n^{-1}\sum_{i=1}^n\operatorname{\operatorname{1}}\{Y_{ji}\leq y\}$ for $j\in\{1,2\}$.

Combining estimators in (ref), we can estimate the policy effect on the copula by \[ \widehat{\Delta}_{C}(u_1,u_2)\equiv\widehat{C}_{Y_1^*Y_2^*}(u_1,u_2)-\widehat{C}_{Y_1Y_2}(u_1,u_2), \] and the policy effect on the association by \[ \widehat{\Delta}_{\nu}\equiv\nu(\widehat{C}_{Y_1^*Y_2^*})-\nu(\widehat{C}_{Y_1Y_2}). \]

Asymptotic Theory

The asymptotic properties of the empirical copula estimator in (ref) has been considerably studied over the past two decades. One recent advance is that Segers2012 proves the weak convergence of the empirical copula process $\widehat{\mathbb{C}}_{Y_1Y_2}\equiv\sqrt{n}(\widehat{C}_{Y_1Y_2}-C_{Y_1Y_2})$ in the space $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$ under the non-restrictive assumption on the copula in cases where data are independent and identically distributed.\footnote{ We write $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$ for the set of all bounded real-valued functions on $[0,1]^2$ equipped with the uniform norm $\|\cdot\|_{\infty}$.} We present this assumption on the actual copula as follows.\\[0.2cm] Assumption C (Non-restrictive Assumption on the Actual Copula)

enumerate[label=C\arabic*] • For each $j\in\{1,2\}$, the $j$-th partial derivative $\partial C_{Y_1Y_2}/\partial u_j$ exists and is continuous on the set $\{(u_1,u_2)^\top\in[0,1]^2:u_j\in(0,1)\}$.

To extend the partial derivatives to the whole unit square $[0,1]^2$, we follow Segers2012 and define \[ \frac{\partial C_{Y_1Y_2}(u_1,u_2)}{\partial u_1} \equiv

cases\limsup\limits_{u\downarrow0}\dfrac{C_{Y_1Y_2}(u,u_2)}{u} & if $u_1=0$, \\ \limsup\limits_{u\downarrow0}\dfrac{C_{Y_1Y_2}(1,u_2)-C_{Y_1Y_2}(1-u,u_2)}{u} & if $u_1=1$,

\] and \[ \frac{\partial C_{Y_1Y_2}(u_1,u_2)}{\partial u_2} \equiv

cases\limsup\limits_{u\downarrow0}\dfrac{C_{Y_1Y_2}(u_1,u)}{u} & if $u_2=0$, \\ \limsup\limits_{u\downarrow0}\dfrac{C_{Y_1Y_2}(u_1,1)-C_{Y_1Y_2}(u_1,1-u)}{u} & if $u_2=1$.

\] As discussed by Segers2012 and Buecher2014, Assumption C is necessary for the existence of the limiting process and continuity of its trajectories. Furthermore, Buecher2013 demonstrate that Assumption C implies Hadamard differentiability of the copula mapping; thus, weak convergence of the empirical copula process can be established as long as the empirical process $\sqrt{n}(\widetilde{G}_{Y_1Y_2}-C_{Y_1Y_2})$ converges weakly in the space $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$ to a tight, centered Gaussian process, where the empirical CDF \[ \widetilde{G}_{Y_1Y_2}(u_1,u_2)\equiv\frac{1}{n}\sum_{i=1}^n\operatorname{\operatorname{1}}\{F_{Y_1}(Y_{1i})\leq u_1,F_{Y_2}(Y_{2i})\leq u_2\} \] is constructed by the unobservable probability integral transforms $\{(F_{Y_1}(Y_{1i}),F_{Y_2}(Y_{2i}))\}_{i=1}^n$. For completeness, we summarize the discussion above and state the following proposition, which is obtained by Segers2012 for independent observations.\footnote{Buecher2013 and BuecherKojadinovic2016 obtain the same representation of $\widehat{\mathbb{C}}_{Y_1Y_2}$ for mixing observations.}

proUnder Assumption {\normalfont C}, the empirical copula process has the representation \begin{align*} &\widehat{\mathbb{C}}_{Y_1Y_2}(u_1,u_2)\\ =&\sqrt{n}(\widetilde{G}_{Y_1Y_2}(u_1,u_2)-C_{Y_1Y_2}(u_1,u_2)) -\sum_{j=1}^2\frac{\partial C_{Y_1Y_2}(u_1,u_2)}{\partial u_j}\sqrt{n}(\widetilde{G}_{Y_j}(u_j)-u_j)+R_n(u_1,u_2), \end{align*} where $\widetilde{G}_{Y_1}(u)=\widetilde{G}_{Y_1Y_2}(u,1)$, $\widetilde{G}_{Y_2}(u)=\widetilde{G}_{Y_1Y_2}(1,u)$, and the remainder term $R_n$ satisfies \[ \sup_{(u_1,u_2)\in[0,1]^2}|R_n(u_1,u_2)|=o_p(1). \] Consequently, the empirical copula process converges weakly to a Gaussian process in the space $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$.

We now show that the counterfactual copula process $\widehat{\mathbb{C}}_{Y^*_1Y^*_2}\equiv\sqrt{n}(\widehat{C}_{Y^*_1Y^*_2}-C_{Y^*_1Y^*_2})$ has a similar representation provided the following assumptions hold.\\[0.3cm] Assumption K (Kernel)\\ The kernel function $K:\mathbb{R}^{d}\rightarrow \mathbb{R}$ satisfies the following conditions.

enumerate[label=K\arabic*] • $K$ is of bounded variation, vanishes outside $[-1,1]^d$, and satisfies $\int K(u)\d u=1$. • There is a positive integer $r\geq 2$ such that $\int\prod_{\ell=1}^du_{\ell}^{\lambda_{\ell}}K(u)\d u=0$ for any $d$-dimensional vector $\lambda=(\lambda_1,\ldots,\lambda_{d})^\top$ of nonnegative integers with $\sum_{\ell=1}^d\lambda_{\ell}\leq r-1$. • For $u\in[-1,1]^d$, $K(u)$ is $r$-times differentiable with respect to $u$, and the derivatives are uniformly continuous and bounded. • For $u\in[-1,1]^d$, $K(u)=K(|u|)$.

Assumption B (Bandwidth)\\ The sequence $\{h_{n}\}_{n=1}^{\infty}$ of bandwidths satisfies the following conditions.

enumerate[label=B\arabic*] • $h_{n}\rightarrow 0$. • $\frac{\log n}{n^{1/2}h_{n}^{d}}\rightarrow 0$. • $n^{1/2}h_{n}^{r}\rightarrow 0$.

These assumptions on the kernel and bandwidth are mild and common in the literature on nonparametric and semiparametric econometrics. To make Assumptions (ref) and (ref) valid simultaneously, we require that the order of kernel $K$ be greater than the dimension of covariates, as is imposed in Rothe2010. We also need technical assumptions on the support and smoothness of density and distribution functions as follows.\\[0.3cm] Assumption S (Smoothness and Support)\\ Let $\mathcal{X}$, $\mathcal{Y}_1$, $\mathcal{Y}_2$, $\mathcal{Y}_1^*$, and $\mathcal{Y}_2^*$ be the support of $X$, $Y_1$, $Y_2$, $Y_1^*$, and $Y_2^*$, respectively. Also let $f_Z$ denote the density function of a generic random variable $Z$ with respect to the Lebesgue measure.

enumerate[label=S\arabic*] • The function $f_{X}(x)$ is $r$-times differentiable with respect to $x$ on the interior of $\mathcal{X}$, and its derivatives are bounded and uniformly continuous. In addition, the function $f_{X}(x)$ is bounded away from zero on $\mathcal{X}$. • For all $(y_1,y_2,x)\in \mathcal{Y}_1\mathcal{Y}_2\mathcal{X}$, the first $r$ partial derivatives with respect to $x$ of $F_{Y_1Y_2|X}(y_1,y_2|x)$ and $f_{Y_1Y_2|X}(y_1,y_2|x)$ are bounded. • The function $f_{X^*}(x)$ is $r$-times differentiable with respect to $x$ on the interior of $\mathcal{X}$, and its derivatives are bounded and uniformly continuous.\footnote{Let $f_{X^*}(x)=0$ if $x\notin\mathcal{X}^*$.} • Both $\operatorname{\operatorname{E}}[(\sup_{(y_1,y_2)\in\mathcal{Y}_1\mathcal{Y}_2}f_{Y_1Y_2|X}(y_1,y_2|X))^2]$ and $\operatorname{\operatorname{E}}[ (f_{X^*}(X)/f_{X}(X))^2]$ are finite. • For each $j\in\{1,2\}$, both $\mathcal{Y}_j$ and $\mathcal{Y}^*_j$ are bounded with respect to the Euclidean distance. • The function $F_{Y^*_1Y^*_2}$ has continuous partial derivatives of order up to 2 on $\mathcal{Y}_1^*\mathcal{Y}_2^*$. • The function $f_{Y^*_1Y^*_2}$ is bounded away from zero on $\mathcal{Y}_1^*\mathcal{Y}_2^*$.
proUnder Assumptions {\normalfont I}, {\normalfont K}, {\normalfont B}, and {\normalfont S}, the counterfactual copula process has the representation \begin{align*} &\widehat{\mathbb{C}}_{Y^*_1Y^*_2}(u_1,u_2)\\ =&\sqrt{n}(\widetilde{G}_{Y^*_1Y^*_2}(u_1,u_2)-C_{Y^*_1Y^*_2}(u_1,u_2))-\sum_{j=1}^2\frac{\partial C_{Y^*_1Y^*_2}(u_1,u_2)}{\partial u_j}\sqrt{n}(\widetilde{G}_{Y^*_j}(u_j)-u_j)+R^*_n(u_1,u_2), \end{align*} where $\widetilde{G}_{Y^*_1}(u)=\widetilde{G}_{Y^*_1Y^*_2}(u,1)$ and $\widetilde{G}_{Y^*_2}(u)=\widetilde{G}_{Y^*_1Y^*_2}(1,u)$ with \[ \widetilde{G}_{Y^*_1Y^*_2}(u_1,u_2)\equiv\frac{1}{n}\sum_{i=1}^nW_{i,n}\operatorname{\operatorname{1}}\{F_{Y^*_1}(Y_{1i})\leq u_1,F_{Y^*_2}(Y_{2i})\leq u_2\}, \] and $\{W_{i,n}\}_{i=1}^n$ are the weights given immediately after (ref). The extension to $[0,1]^2$ of $\partial C_{Y^*_1Y^*_2}/\partial u_j$ is similar to that of $\partial C_{Y_1Y_2}/\partial u_j$ for $j\in\{1,2\}$ , and the remainder term $R^*_n$ satisfies \[ \sup_{(u_1,u_2)\in[0,1]^2}|R^*_n(u_1,u_2)|=o_p(1). \]

According to (ref), the weak convergence of $\widehat{\mathbb{C}}_{Y^*_1Y^*_2}$ in $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$ is guaranteed by that of $\sqrt{n}(\widetilde{G}_{Y^*_1Y^*_2}-C_{Y^*_1Y^*_2})$ in the same space. Notice that $\widetilde{G}_{Y^*_1Y^*_2}(u_1,u_2) =\widehat{F}_{Y_1^*Y_2^*}(F_{Y_1^*}^{-1}(u_1),F_{Y_2^*}^{-1}(u_2))$ by construction, where $\widehat{F}_{Y_1^*Y_2^*}$ is the joint counterfactual CDF estimator defined in (ref). It follows that

align*[align* omitted — 223 chars of source]

Hence, the weak convergence of $\sqrt{n}(\widetilde{G}_{Y^*_1Y^*_2}-C_{Y^*_1Y^*_2})$ in $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$ is equivalent to that of $\sqrt{n}(\widehat{F}_{Y_1^*Y_2^*}-F_{Y_1^*Y_2^*})$ in $(\ell^\infty(\mathcal{Y}_1\mathcal{Y}_2),\|\cdot\|_\infty)$. To show the latter weak convergence, we extend the representation of the univariate counterfactual CDF estimator $\widehat{F}_{Y_j^*}$ obtained in Rothe2010 to that of joint counterfactual CDF estimator $\widehat{F}_{Y_1^*Y_2^*}$ in (ref) in (ref).

Since the policy effects on the copula and the association measures are based on comparing the actual and counterfactual copulas, we study the asymptotic behavior of the random map \[

bmatrix[bmatrix omitted — 24 chars of source]

\mapsto

bmatrix[bmatrix omitted — 144 chars of source]

=

bmatrix[bmatrix omitted — 98 chars of source]

. \] (ref) allow the random map to be approximated by the scaled sum of two-dimensional influence functions, which will be used to prove (ref) below.\footnote{The notation $\Rightarrow$ stands for weak convergence in a functional space equipped with an appropriate norm.}

thmSuppose that Assumptions {\normalfont I}, {\normalfont C}, {\normalfont K}, {\normalfont B}, and {\normalfont S} hold. Then in the space $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)\times(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$, \[ (\widehat{\mathbb{C}}_{Y_1Y_2},\widehat{\mathbb{C}}_{Y^*_1Y^*_2})^\top \Rightarrow (\mathbb{C}_{Y_1Y_2},\mathbb{C}_{Y^*_1Y^*_2})^\top\equiv\mathbb{C}, \] where $\mathbb{C}$ is a two-dimensional centered Gaussian process. Furthermore, let $\nu$ be a functional mapping from $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$ to some normed space $(\mathcal{V},\|\cdot\|_{\mathcal{V}})$. Suppose that $\nu$ is Hadamard differentiable at $C_{Y_1Y_2}$ and $C_{Y^*_1Y^*_2}$. Then in the space $(\mathcal{V},\|\cdot\|_{\mathcal{V}})\times(\mathcal{V},\|\cdot\|_{\mathcal{V}})$, \begin{align*} (\sqrt{n}(\nu(\widehat{C}_{Y_1Y_2})-\nu(C_{Y_1Y_2})), \sqrt{n}(\nu(\widehat{C}_{Y^*_1Y^*_2})-\nu(C_{Y^*_1Y^*_2})))^\top \Rightarrow(\mathbb{V}_{Y_1Y_2},\mathbb{V}_{Y^*_1Y^*_2})^\top\equiv\mathbb{V}, \end{align*} where $\mathbb{V}$ is a two-dimensional centered Gaussian process.

Applying (ref) and the continuous mapping theorem, we obtain $\sqrt{n}(\widehat{\Delta}_{C}-\Delta_{C})\Rightarrow (-1,1)\mathbb{C}$ in the space $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$ and $\sqrt{n}(\widehat{\Delta}_{\nu}-\Delta_{\nu})\Rightarrow (-1,1)\mathbb{V}$ in the space $(\mathcal{V},\|\cdot\|_{\mathcal{V}})$.

To find confidence bands for $C_{Y^*_1Y^*_2}$, $\Delta_{C}$, and $\Delta_{\nu}$, we apply a nonparametric bootstrap procedure. Let $\{(Y_{b,1i},Y_{b,2i},X_{b,i},X_{b,i}^*)\}_{i=1}^n$ be the bootstrap data with sample size $n$ drawn with replacement from the data $\{(Y_{1i},Y_{2i},X_i,X_i^*)\}_{i=1}^n$ at hand. By replacing the data at hand with the bootstrap data in (ref), we construct the bootstrap counterfactual copula estimator $\widehat{C}^{b}_{Y_1^*Y_2^*}$ and bootstrap empirical copula estimator $\widehat{C}^{b}_{Y_1Y_2}$, respectively. Additionally, we define the bootstrap counterfactual copula process to be $\widehat{\mathbb{C}}^{b}_{Y^*_1Y^*_2} \equiv\sqrt{n}(\widehat{C}^{b}_{Y^*_1Y^*_2}-\widehat{C}_{Y^*_1Y^*_2})$ and bootstrap empirical copula process $\widehat{\mathbb{C}}^{b}_{Y_1Y_2}\equiv\sqrt{n}(\widehat{C}^{b}_{Y_1Y_2}-\widehat{C}_{Y_1Y_2})$. To validate the nonparametric bootstrap for statistical inference, we prove the following theorem.

thmUnder the assumptions in (ref), in the space $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)\times(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$, \[ (\widehat{\mathbb{C}}^{b}_{Y_1Y_2},\widehat{\mathbb{C}}^{b}_{Y^*_1Y^*_2})^\top\Rightarrow(\mathbb{C}_{Y_1Y_2},\mathbb{C}_{Y^*_1Y^*_2})^\top=\mathbb{C}, \] conditional on the data $\{(Y_{1i},Y_{2i},X_i,X_i^*)\}_{i=1}^n$. Moreover, if the functional mapping $\nu$ from $(\ell^\infty([0,1]^2),\|\cdot\|_\infty)$ to some normed space $(\mathcal{V},\|\cdot\|_{\mathcal{V}})$ is Hadamard differentiable at $C_{Y_1Y_2}$ and $C_{Y^*_1Y^*_2}$, then in the space $(\mathcal{V},\|\cdot\|_{\mathcal{V}})\times (\mathcal{V},\|\cdot\|_{\mathcal{V}})$, \[ (\sqrt{n}(\nu(\widehat{C}^{b}_{Y_1Y_2})-\nu(\widehat{C}_{Y_1Y_2})), \sqrt{n}(\nu(\widehat{C}^{b}_{Y^*_1Y^*_2})-\nu(\widehat{C}_{Y^*_1Y^*_2})))^\top\Rightarrow(\mathbb{V}_{Y_1Y_2},\mathbb{V}_{Y^*_1Y^*_2})^\top=\mathbb{V}, \] conditional on the data $\{(Y_{1i},Y_{2i},X_i,X_i^*)\}_{i=1}^n$.

(ref) allows us to construct confidence bands for the counterfactual copula $C_{Y_1^*Y_2^*}$ and the actual copulas $C_{Y_1Y_2}$. It can also be applied to the construction of confidence intervals for an association measure $\nu(C_{Y^{*}_{1}Y^{*}_{2}})$ and a policy effect $\Delta_{\nu}= \nu(C_{Y_1^*Y_2^*})-\nu(C_{Y_1Y_2})$ for any functional mapping $\nu:(\ell^\infty([0,1]^2),\|\cdot\|_\infty)\to(\mathbb{R},|\cdot|)$ that is Hadamard differentiable at $C_{Y_1Y_2}$ and $C_{Y^*_1Y^*_2}$. Details about the construction are deferred to the next section.

Monte Carlo Simulation

We evaluate the finite-sample properties of the actual copula and counterfactual copula estimators. We also investigate bootstrap validity for inference on the association measures and corresponding policy effects. The data generating process under consideration is

align*[align* omitted — 79 chars of source]

where $(X,\varepsilon_1,\varepsilon_2)^\top$ follows a trivariate normal distribution with zero mean and covariance matrix \[

bmatrix[bmatrix omitted — 57 chars of source]

. \] Clearly, $(Y_1,Y_2)^\top$ is bivariate normally distributed and the actual copula is a Gaussian copula $C_{Y_1Y_2}(u_1,u_2)=\Phi_R(\Phi^{-1}(u_1),\Phi^{-1}(u_2))$, where $\Phi^{-1}$ is the inverse function of the standard normal CDF and $\Phi_R$ is the joint CDF of a bivariate normal with zero mean and correlation matrix \[ R=

bmatrix[bmatrix omitted — 70 chars of source]

. \] We consider an exogenous policy intervention \[ X^*=\pi(X)=0.5X, \] and this intervention does not affect the unobservables $(\varepsilon_1,\varepsilon_2)^\top$. Simple algebra shows that the counterfactual copula is also a Gaussian copula $C_{Y_1^*Y_2^*}(u_1,u_2)=\Phi_{R^*}(\Phi^{-1}(u_1),\Phi^{-1}(u_2))$ with correlation matrix \[ R^*=

bmatrix[bmatrix omitted — 68 chars of source]

. \] The data generating process and the counterfactual policy are not meant to mimic any data set in empirical studies; instead, they are only illustrative of the proposed method.

Given a random sample $\{(Y_{1i},Y_{2i},X_i)\}_{i=1}^n$ of size $n$, we estimate the actual copula by the rank-based empirical copula in Genest1995 for computational efficiency; specifically, \[ \widetilde{C}_{Y_1Y_2}(u_1,u_2)=\frac{1}{n}\sum_{i=1}^n\operatorname{\operatorname{1}}\{\widehat{F}_{Y_1}(Y_{1i})\leq u_1,\widehat{F}_{Y_2}(Y_{2i})\leq u_2\}. \] This rank-based empirical copula $\widetilde{C}_{Y_1Y_2}$ differs from $\widehat{C}_{Y_1Y_2}$ in (ref) but their difference is first-order asymptotically negligible; see FermanianRadulovicEtAl2004. Similarly, we modify the proposed counterfactual copula estimator $\widehat{C}_{Y_1^*Y_2^*}$ in (ref) as \[ \widetilde{C}_{Y_1^*Y_2^*}(u_1,u_2)=\frac{1}{n}\sum_{i=1}^nW_{i,n}\operatorname{\operatorname{1}}\{\widehat{F}_{Y_1^*}(Y_{1i})\leq u_1,\widehat{F}_{Y_2^*}(Y_{2i})\leq u_2\}. \] To construct the weights $W_{i,n}$, we use the Epanechnikov kernel with bandwidth $h=5.5s_Xn^{-1/3}$ where $s_X$ denotes the sample standard deviation of $X$. We have also tried other kernel functions and bandwidth constants, but the results are not too sensitive to these changes.\footnote{In particular, we have tried the equivalent kernel of the local linear regression for automatic boundary correction. See Fan1996 for more details. The simulation results are broadly similar and are available upon request.} The copula functions are estimated on the grids $\{(i/100, j/100): i, j= 0,1,\dots,100\}$. For every sample size $n\in\{100, 200, 400\}$, we conduct 1,000 simulation replications.

table[table omitted — 1,343 chars of source]

(ref) presents the simulation results regarding the actual and counterfactual copula estimators in terms of mean integrated absolute error (MIAE) and root mean integrated squared error (RMISE).\footnote{For an estimator $\widehat{C}$ of a copula $C$, the mean integrated absolute error of $\widehat{C}$ is \[ \text{MIAE}(\widehat{C})=\mathbb{E}\mathopen{}\mathclose\bgroup\originalleft[\int_0^1\int_0^1|\widehat{C}(u_1,u_2)-C(u_1,u_2)|\d u_1 \d u_2\aftergroup\egroup\originalright] \] and the root mean integrated squared error of $\widehat{C}$ is \[ \text{RMISE}(\widehat{C})=\sqrt{\mathbb{E}\mathopen{}\mathclose\bgroup\originalleft[\int_0^1\int_0^1|\widehat{C}(u_1,u_2)-C(u_1,u_2)|^2\d u_1\d u_2\aftergroup\egroup\originalright]}. \]} As expected, the rank-based empirical copula $\widetilde{C}_{Y_1Y_2}$ performs well in terms of MIAE and RMISE. It can also be seen that the MIAE and RMISE of $\widetilde{C}_{Y_1^*Y_2^*}$ shrink with increasing sample size and nearly halve as the size quadruples. This result suggests the $\sqrt{n}$-consistency of the proposed counterfactual estimator, which is in line with our theoretical analysis. For comparison, we also consider an oracle counterfactual copula estimator, which would be the empirical copula of $(Y_1^*,Y_2^*)$ if the counterfactual outcomes $\{(Y_{1i}^*,Y_{2i}^*)\}_{i=1}^n$ were observed. This oracle estimator, despite having smaller MIAE and RMISE, does not outweigh the proposed estimator considerably for moderate sample sizes.

We next turn to address the inferential questions about the measures of association and corresponding policy effects. To tackle the inferential tasks, we implement the nonparametric bootstrap. For $b=1,\dotsc,B$, let $(M^b_{1,n},\dotsc,M^b_{n,n})^{\top}$ be a multinomial random vector with success probabilities $(1/n,\dotsc,1/n)$. Let $\widetilde{C}^b_{Y_1Y_2}$ and $\widetilde{C}^b_{Y_1^*Y_2^*}$ denote the bootstrap estimators of actual and counterfactual copulas, respectively. To be specific,

align*[align* omitted — 380 chars of source]

where $\widehat{F}^b_{Y_j}(y)=n^{-1}\sum_{i=1}^nM^b_{i,n}\operatorname{\operatorname{1}}\{Y_{ji}\leq y\}$ and $\widehat{F}^b_{Y_j^*}(y)=n^{-1}\sum_{i=1}^nM^b_{i,n}W_{i,n}\operatorname{\operatorname{1}}\{Y_{ji}\leq y\}$ for $j\in\{1,2\}$. Also let $\widetilde{\Delta}_\nu=\nu(\widetilde{C}_{Y_1^*Y_2^*})-\nu(\widetilde{C}_{Y_1Y_2})$ and $\widetilde{\Delta}^{b}_\nu=\nu(\widetilde{C}^b_{Y_1^*Y_2^*})-\nu(\widetilde{C}^b_{Y_1Y_2})$ be the original and bootstrap estimators of policy effect $\Delta_\nu$, respectively. We construct the $(1-\alpha)$ confidence interval for the policy effect $\Delta_\nu$ as \[ \mathopen{}\mathclose\bgroup\originalleft[\widetilde{\Delta}_\nu-\frac{Q_\nu^{1-\alpha,B}}{\sqrt{n}},\quad\widetilde{\Delta}_\nu+\frac{Q_\nu^{1-\alpha,B}}{\sqrt{n}}\aftergroup\egroup\originalright], \] where $Q_\nu^{1-\alpha,B}$ is the $(1-\alpha)$-th quantile of \[ \mathopen{}\mathclose\bgroup\originalleft\{\mathopen{}\mathclose\bgroup\originalleft|\sqrt{n}(\widetilde{\Delta}^b_\nu-\widetilde{\Delta}_\nu)-\frac{\sqrt{n}}{B}\sum_{b'=1}^B(\widetilde{\Delta}^{b'}_\nu-\widetilde{\Delta}_\nu)\aftergroup\egroup\originalright|\aftergroup\egroup\originalright\}_{b=1}^B. \] We generate $B=1,000$ bootstrap replicates and set the nominal coverage level to be 95%. Note that following Kojadinovic2019, we consider the centered version of $\sqrt{n}(\widetilde{\Delta}^b_\nu-\widetilde{\Delta}_\nu)$. The rationale behind centering is that the bootstrap processes, whether for copulas or their corresponding functionals, converge weakly to centered Gaussian processes as stated in (ref). Hence, using centered replicates would always lead to better approximation in finite samples. It is also worth noting that rather than applying large-sample approximation to these measures, we can calculate them analytically. As shown in Meyer2013, the bivariate Gaussianity of $C_{Y_1Y_2}$ and $C_{Y^*_1Y^*_2}$ implies that the association measures, mentioned in (ref), can be expressed in terms of the Pearson correlation coefficient $r$, which is the off-diagonal element of the correlation matrix. Concretely, we have

align*[align* omitted — 513 chars of source]
table[table omitted — 2,267 chars of source]

(ref) reports the mean absolute errors (MAE), root mean squared errors (RMSE), and confidence interval coverage rates (CR) for the estimators of actual and counterfactual association measures and the corresponding policy effects. As expected from our theoretical analysis, both MAE and RMSE nearly halve as the sample size quadruples; furthermore, the empirical coverage rates are generally close to the pre-specified nominal level. These results suggest that the bootstrap method performs well in small samples.

Empirical Study

We apply the counterfactual copula method to studying the role of college education in intergenerational income mobility in the United States. Over the past decades, researchers have observed that the association between incomes across generations among college-educated children is weaker than that among less educated groups. That there is a “meritocratic power” of college education reducing the influence of family origin has been supported by a number of empirical studies, for example Hout1988, Torche2011, and Chetty2020. While most of the early literature hinges on the intergenerational income elasticity, this measure has its limitations as discussed in Mitnik2020. Alternatively, Chetty2014 advocate assessing positional mobility by intergenerational rank correlation, which is none other than Spearman's rho between parent and child income. Besides, it is well recognized that the observed high degree of mobility among children with college education reflects a combination of selection bias (the types of students admitted) and causal effects (the returns to college education). To address this confounding issue, Zhou2019 adopts a reweighting approach whereas Gregg2019 employ unconditional quantile regression with the inclusion of confounding factors. In view of these advances, we use the proposed copula method to investigate the heterogeneous effects of education on intergenerational income mobility. This approach allows us to simultaneously control for confounding factors and explore various aspects of mobility, including Spearman's rho and other measures of association discussed in Section (ref).

Data

The data we use are drawn from the Panel Study of Income Dynamics (PSID), one of the primary databases used in the literature on intergenerational mobility in the United States.\footnote{See Mazumder2018 for a comprehensive survey of the PSID.} Beginning in 1968, the PSID is a longitudinal study of a representative sample of US individuals and their family units. The sample has been reinterviewed annually through 1997 and biennially thereafter to collect information on the individuals and their descendants. As children from the original PSID families transition to adulthood and form their own households, the sample size grew from 4,800 family units in 1968 to almost double in 2019. However, family-level variables collected in each wave of the study are contained in separate single-year family files, which need to be merged by a cross-year individual file containing all individual-level variables. In this study, we use both family and individual files for the survey years 1968--2019. Additionally, we follow Lee2009 to exclude the Survey of Economic Opportunity (SEO) part of the PSID due to serious irregularities in the sampling of SEO respondents.

Like most of the literature, we construct permanent incomes of parents and children, $Y_1$ and $Y_2$, using at least 5-year averages of total family income across years.\footnote{See Table 2 of Mazumder2018 for a selected summary of related studies.} There are several advantages of using averages and using family income. First, averages of income likely provide a better measure of permanent income by reducing the impact of transitory shocks. Second, by using averaged data, we do not require individuals to report income every year. Third, using family income instead of individual one allows us to keep the unemployed in the analysis. Fourth, families with one high-income spouse will be properly treated as high-income families. To avoid heterogeneity in life cycle earnings profiles, we only average the total family income for parents and children aged 25 to 55 at the time of data collection. We also adjust the family income in all years to 2019 dollars using the CPI-U-RS prior to any calculations. Note that because of the use of family income for both parents and children, we cannot separate the income of children from that of their parents if they live together. Consequently, our sample consists of 3,895 households, with children being the head (or the spouse of the head) of another household in a subsequent reporting year.

In addition to income variables, we also consider several family characteristics $X$ that may affect income mobility. These covariates include gender, race, and years of schooling for the family head (parent), as well as gender, year of birth, and years of schooling for the child. (ref) presents the descriptive statistics for all variables under consideration.

table[table omitted — 1,012 chars of source]

Scenarios

We evaluate the effects of education on intergenerational income mobility in two counterfactual scenarios. In the first scenario, children of low educational attainment are given more years of schooling such that all children receive at least a certain amount of education. Specifically, the counterfactual covariates $X^*=(X^*_\text{cedu}, X^{*\top}_\text{other})^\top$ consist of the counterfactual children's years of schooling $X^*_\text{cedu}=\max\{X_\text{cedu},s\}$ with $X_\text{cedu}$ the actual value and $s$ the years of “compulsory” schooling varying from one experiment to another, and $X^*_\text{other}=X_\text{other}$ representing the other family characteristics, which are held constant. Incrementally increasing the parameter $s$ from 13 to 16 (corresponding to the case where a college degree is conferred on all children), we can examine the heterogeneous effects of college education on income mobility by comparing the counterfactual measures of association with the actual ones.

In the second scenario, children with less educated parents are required to complete a four-year college degree. To be specific, let $X^*_\text{cedu}=\max\{X_\text{cedu},16\}\cdot\operatorname{1}\{X_\text{pedu}\leq s'\}+X_\text{cedu}\cdot\operatorname{1}\{X_\text{pedu}>s'\}$, where $X_\text{pedu}$ is the parental years of schooling and $s'$ is the threshold. By shifting $s'$ from 6 to 17,\footnote{Note that only 1.34% (52 out of 3,895) of parents in the sample have less than 6 years of education.} we can investigate the trade-off between the effectiveness of a policy and the number of children affected by this policy. This trade-off may further provide policymakers with guidance on the public investments in education from the perspective on intergenerational income mobility.

Results

The results of our first counterfactual scenario are summarized in (ref), where the actual, counterfactual, and subgroup measures of association across children's compulsory schooling parameter $s$ are plotted.\footnote{ By a subgroup measure of association, we mean the association measure for a subgroup of families with children's actual years of schooling above or equal to $s$. } As can be seen in Panel (a), the actual Spearman's rho, depicted by the horizontal dashed line, falls well within the 95% confidence interval for the counterfactual Spearman's rho except for the case of $s=16$, in which all children complete a four-year college degree. This finding indicates that, other things being equal, providing some college education to all children is unlikely to boost intergenerational income mobility, while the receipt of a college degree can effectively reduce income persistence by about one-third of the magnitude. In addition, the subgroup Spearman's rho for children who attend college without receiving a college degree substantially differs from the corresponding (i.e., $s=13, 14, 15$) counterfactual Spearman's rho. Such differences suggest that the high mobility (low association) observed in each of these subgroups may be attributed to selection bias, rather than college education per se. These findings are robust to the use of the other measures of association, inclusive of Kendall's tau, Gini's gamma, and Blomqvist's beta, shown in Panels (b)-(d) respectively.

figure[figure omitted — 700 chars of source]

(ref) shows the counterfactual policy effects on the four association measures in the second scenario. As shown in Panel (a), the policy effect on Spearman's rho is significantly less than zero when the threshold parameter is set at $s'=10$. In this setting, children with parental eduction less than, or equal to, ten years of schooling are required to complete a four-year college degree, and those families affected by the counterfactual policy account for 14.87% (579 out of 3,895) of the sample. In addition, when the threshold parameter $s'$ is greater than 12, the relatively flat curve is suggestive of a negligible improvement in the policy effect on Spearman's rho. The robustness of these findings are confirmed by the similar curves plotted in Panels (b)-(d) where the other measures of association are adopted. Taken together, these results demonstrate that intergenerational income mobility can be effectively improved if children from less educated families (i.e., parental eduction less than or equal to ten years) are required to complete a four-year college degree.

figure[figure omitted — 735 chars of source]

Conclusion

In this paper, we propose a nonparametric estimator of the counterfactual copula in response to exogenous policy intervention and prove its weak convergence. We also establish bootstrap validity for inference on the counterfactual copula and the corresponding measures of association. Simulation results are in line with our theoretical findings. Finally, we illustrate the use of the proposed method with an empirical application investigating the role of college education in intergenerational income mobility. Our empirical results suggest that if children from less-educated families were given a college degree, then intergenerational income mobility would be substantially improved.