EconBase
← Back to paper

Nonparametric Causal Decomposition of Group Disparities

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

94,442 characters · 23 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric Causal Decomposition of Group Disparities

abstractWe introduce a new nonparametric causal decomposition approach that identifies the mechanisms by which a treatment variable contributes to a group-based outcome disparity. Our approach distinguishes three mechanisms: group differences in 1) treatment prevalence, 2) average treatment effects, and 3) selection into treatment based on individual-level treatment effects. Our approach reformulates classic Kitagawa-Blinder-Oaxaca decompositions in causal and nonparametric terms, complements causal mediation analysis by explaining group disparities instead of group effects, and isolates conceptually distinct mechanisms conflated in recent random equalization decompositions. In contrast to all prior approaches, our framework uniquely identifies differential selection into treatment as a novel disparity-generating mechanism. Our approach can be used for both the retrospective causal explanation of disparities and the prospective planning of interventions to change disparities. We present both an unconditional and a conditional decomposition, where the latter quantifies the contributions of the treatment within levels of certain covariates. We develop nonparametric estimators that are $\sqrt{n}$-consistent, asymptotically normal, semiparametrically efficient, and multiply robust. We apply our approach to analyze the mechanisms by which college graduation causally contributes to intergenerational income persistence (the disparity in adult income between the children of high- vs low-income parents). Empirically, we demonstrate a previously undiscovered role played by the new selection component in intergenerational income persistence.

Introduction

Social and health scientists often seek to decompose an outcome disparity between groups in terms of the contributions of an intermediate treatment variable. For example, how much, and in what ways, do racial differences in medical care contribute to racial disparities in health howe_african_2014? How does childbearing contribute to the gender wage gap cha_is_2023? And what are the roles of education in the association between parents' and their children's social class (social mobility) ishida_class_1995? The common structure of these questions is that they seek to quantify the mechanisms by which a treatment variable causally explains an observed descriptive group disparity.

Prior research has addressed such questions using three approaches, none of which is fully appropriate for the task of causally decomposing descriptive disparities. First, popular Kitagawa-Blinder-Oaxaca (KBO) decompositions kitagawa_components_1955, blinder_wage_1973, oaxaca_male-female_1973 are defined in terms of parametric regression coefficients and do not answer any causal question by design (fortin_decomposition_2011, fortin_decomposition_2011, p.13; lundberg_what_2021, lundberg_what_2021, p.542). Second, causal mediation analysis (CMA) vanderweele_explanation_2015, though formulated in causal terms, decomposes the causal effects of group membership rather than observed group disparities. Third, recently developed random-equalization decompositions vanderweele_causal_2014, jackson_decomposition_2018, jackson_meaningful_2021, lundberg_gap-closing_2024 conflate distinct disparity-generating mechanisms, which limits their interpretability, explanatory power, and usefulness for policy analysis. Importantly, all prior approaches neglect that outcome disparities can in part be explained by group-differential selection of individuals into treatment as a function of their individual treatment effects.

In this article, we propose a nonparametric causal decomposition approach that remedies these limitations. In contrast to KBO decompositions, our decompositions are formulated as model-free causal estimands with interventional interpretations. In contrast to CMA, we decompose observed group disparities rather than possibly ill-defined causal effects of group membership. In contrast to random equalization decompositions, our approach divides the outcome disparity into conceptually unambiguous components. In contrast to all prior approaches, we reveal a novel disparity-generating mechanism, termed the “selection component,’’ that has not previously been incorporated in decomposition analysis.

Our framework distinguishes three causal mechanisms through which an intermediate treatment variable can produce a group disparity in an outcome. First, mean outcomes may differ across groups because the groups receive the treatment at different rates (differential prevalence). Second, even if both groups have the same treatment prevalence, mean outcomes may differ across groups because the treatment has different average effects across groups (differential effects). Third, even if both groups have the same treatment prevalence and the same average treatment effects, mean outcomes will nonetheless differ across groups if members of one group select into treatment more strongly on their treatment effects than members of the other group (differential selection); for example, if the treatment is randomly distributed to members of one group, but specifically given to only those members who will benefit the most from the treatment in the other group. Conceptually, prior decompositions were limited to considering differential prevalence and differential effects, leading to the belief that these two are the only possible mechanisms ward_how_2019, diderichsen_differential_2019. Isolating differential selection into treatment as a source of group disparities is the central conceptual contribution of our approach.

We introduce an unconditional and a conditional decomposition. The unconditional decomposition is useful for investigating the overall (marginal) contributions of a treatment to an outcome disparity. The conditional decomposition is useful for investigating the contributions of the treatment within levels of one or more pre-treatment covariates. For example, in the social sciences, the unconditional decomposition might quantify the overall contributions of racial differences in incarceration rates to racial income disparities; and the conditional decomposition could identify the contributions of racial differences in mortgage receipt to racial disparities in home ownership conditional on credit scores, i.e., considering only the part of the racial differences in mortgage receipt that remains after accounting for racial differences in credit scores. In medical research, the unconditional decomposition might inform the overall contributions of vaccinations to racial health disparities; and the conditional decomposition can quantify the contributions of receiving a medical treatment to health disparities conditional on indications for that treatment.

We develop nonparametric estimators for our causal decompositions using efficient influence functions (EIF) under the assumption of conditional ignorability of the treatment. Our estimators can be implemented via data-adaptive methods such as machine learning (ML) and accommodate high-dimensional confounders. We derive the conditions under which the estimators are $\sqrt{n}$-consistent, asymptotically normal, and semiparametrically efficient. The estimators are also doubly or even quadruply robust in terms of consistency. The estimators are implemented in the R package cdgd yu_2023, available from CRAN.

As an empirical application, we study the causal contributions of college graduation to intergenerational income persistence (the complement to income mobility), defined as the disparity in income attainment across parental income groups. This application contributes to multiple literatures in sociology and economics. Policy-wise, it provides insights into how interventions on college education may alter intergenerational income persistence. R Code for the empirical application is available at \url{https://github.com/ang-yu/causal_decomposition_case_study}.

Our paper proceeds as follows. Section 2 introduces our causal decompositions and their interventional interpretations. We explicate the contributions of our framework by formally contrasting it with KBO decompositions, CMA, and random equalization decompositions. In section 3, we introduce the estimators and their asymptotic theory. Section 4 presents the empirical application. Section 5 concludes with extensions. All proofs are collected in the supplementary appendices.

Estimands

Unconditional decomposition

We consider a binary treatment variable $D_i \in \left\lbrace 0,1 \right\rbrace$ for each individual $i$. Let $Y_{i}^0$ and $Y_{i}^1$ be the potential outcomes rubin_estimating_1974 of $Y_i$ under the hypothetical intervention to set $D_i=0$ and $D_i=1$, respectively. Let $\tau_i \coloneqq Y_{i}^1 - Y_{i}^0$ denote the individual-level treatment effect. For expositional convenience, we assume that higher values of $Y_i$ are better. Suppose that the population contains two disjoint groups, $G_i=g \in \left\lbrace a,b \right\rbrace$, where $a$ denotes the advantaged group and $b$ denotes the disadvantaged group. We use subscript $g$ to indicate group-specific quantities, for example, $\operatorname{E}_g(Y_i) \coloneqq \operatorname{E}(Y_i \mid g)$. Henceforth, we suppress subscript $i$ to ease notation. In our empirical application, $G$ is parental income group, $Y$ is adult income, and $D$ is college graduation.

assu[SUTVA] $Y=(1-D) Y^0 + D Y^1$.

Assuming only the stable unit treatment value assumption (SUTVA) on the relationship between the treatment and the outcome rubin_randomization_1980,rubin_comment_1986, the observed outcome disparity between group $a$ and $b$ can be decomposed into four components:

align[align omitted — 552 chars of source]

First, the “baseline” component reflects the difference in mean baseline potential outcomes, $Y^0$, between groups, i.e., the outcome disparity in the complete absence of the treatment.\footnote{The baseline component is identical to the “counterfactual disparity measure” proposed by naimi_mediation_2016. In contrast to the two-way decomposition of naimi_mediation_2016, our four-way decomposition distinguishes more mechanisms.} In our application, the baseline component is the part of the income disparity in adulthood that is not attributable to college graduation in any way. Second, the “prevalence” component indicates how much of the group disparity is due to differential prevalence of the treatment. For example, it indicates the extent to which the difference in college graduation rates between parental income groups contributes to the outcome disparity. Third, the “effect” component reflects the difference in average treatment effects (ATE) between groups. Thus, it reveals the contribution of group-differential ATEs of college graduation to the adult income disparity.

Fourth, the “selection” component captures the contribution of differential selection into treatment based on the individual-level treatment effects. Selection into treatment within each group is captured by $\operatorname{Cov}_g(D,\tau)$. This covariance is positive if group members who would benefit more from the treatment are more likely to receive the treatment. In our example, differential selection will increase the income disparity in adulthood if selection into college graduation is more positive in the higher parental income group than in the lower parental income group.

Both the effect component and the selection component account for the contribution of effect heterogeneity to group disparities. Whereas the effect component captures the contribution of between-group effect heterogeneity, the selection component captures the contribution of within-group effect heterogeneity. To our knowledge, no prior decomposition has captured the role of differential selection into treatment in the generation of group disparities.

To further explicate our novel selection component, we provide two interpretations for the covariance between the treatment and treatment effect, $\operatorname{Cov}(D, \tau)$. First, when the receipt of the treatment is mainly based on self-selection (e.g, college graduation), the covariance may indicate the extent to which the choice to take up the treatment is rational with respect to returns to the treatment. This interpretation has been extensively studied in economics and sociology heckman_structural_2005, brand_who_2010, heckman_returns_2018, where a quantity closely related to $\operatorname{Cov}(D, \tau)$ has been labeled “sorting on gains”. We formally explicate the connection to “sorting on gains” in Appendix F. Second, if treatment assignment is mainly administered by external decision-makers (e.g., drug prescription), the covariance indicates how effectively the treatment is assigned to individuals (see Appendix B). Thus, the selection component could also be called the sorting component or the effectiveness component.

Finally, the sum of the prevalence, effect, and selection components can be thought of as the unconditional total contribution of the treatment. Thus, equation ((ref)) decomposes the total contribution of a treatment into three distinct causal mechanisms.

Interventional interpretation

We expect that most researchers will use our decomposition for retrospective causal explanation of existing disparities. However, since our decomposition is formulated in counterfactual terms, it is also prescriptive for future interventions to change disparities. Here, we explicate a three-step sequential intervention on the treatment, which successively eliminates the selection, prevalence, and effect components, in this order. The baseline component is the remaining disparity after the intervention.

We express this three-step sequential intervention using the randomized intervention notation didelez_direct_2006, geneletti_identifying_2007, where $R(D \mid g)$ represents a randomly drawn value of the treatment $D$ from group $g$. Then, $\operatorname{E}_g \left(Y^{R(D \mid g') } \right)$ denotes the post-intervention mean potential outcome for group $g$ after each member of group $g$ has received a random draw of the treatment from group $g'$. When $g=g'$, the intervention amounts to a random redistribution of the treatment within the group. Using its definition, we can rewrite the post-intervention mean potential outcome:

align[align omitted — 159 chars of source]

Then the components of the unconditional decomposition can be re-written as follows:

align*[align* omitted — 1,005 chars of source]

The first step of the intervention internally randomizes the treatment within each group. In this step, the pre-intervention disparity is $\operatorname{E}_a(Y) - \operatorname{E}_b(Y)$, and the post-intervention disparity is $\operatorname{E}_a \left(Y^{R(D \mid a)} \right) - \operatorname{E}_b \left(Y^{R(D \mid b)}\right)$. Since randomizing the treatment within each group sets $\operatorname{Cov}_a(D,\tau)=\operatorname{Cov}_b(D,\tau)=0$ and removes differential selection between groups, the change in disparity resulting from this step equals the selection component.

The second step of the intervention equalizes treatment prevalence by giving members of group $b$ random draws of the treatment from group $a$. In this step, the pre-intervention disparity is $\operatorname{E}_a \left(Y^{R(D \mid a)} \right) - \operatorname{E}_b \left(Y^{R(D \mid b)}\right)$, and the post-intervention disparity is $\operatorname{E}_a \left(Y^{R(D \mid a)} \right)-\operatorname{E}_b \left(Y^{R(D \mid a)} \right)$. Thus, the prevalence component is the change in disparity resulting from this equalization step.\footnote{This step also justifies the scaling factors in the prevalence and effect components. Intuitively, the prevalence component is scaled by $\operatorname{E}_b(\tau)$, because randomly changing treatment prevalence in group $b$ affects the outcome disparity only if the treatment has an effect in group $b$ on average. The scaling factor in the effect component follows algebraically to complete the decomposition. Different interventions would lead to different scaling factors. For example, if group $a$ is instead intervened to have the treatment prevalence of group $b$, the prevalence component would be scaled by $\operatorname{E}_a(\tau)$.}

The third step of the intervention sets the treatment to zero for all individuals in both groups. In the third step, the pre-intervention disparity is $\operatorname{E}_a \left(Y^{R(D \mid a)} \right)-\operatorname{E}_b \left(Y^{R(D \mid a)} \right)$, and the post-intervention disparity is $\operatorname{E}_a \left(Y^0 \right) - \operatorname{E}_b \left(Y^0 \right)$. When nobody receives the treatment, the influence of differential ATEs across groups is deactivated, which is why the change in disparity in this step equals the effect component.\footnote{The effect component also corresponds to the change in disparity brought about by giving group $a$ the ATE of group $b$. Although the notion of interventions on effects or structural relations often appears in the literature malinsky_intervening_2018, diderichsen_differential_2019, brady_rethinking_2017, these interventions cannot be expressed in the potential outcomes notation. In contrast, the third step of our sequential intervention fits in the potential outcomes framework, as it eliminates the effect component by intervening on the treatment variable.}

Finally, at the end of the three-step sequential intervention, the remaining disparity is the baseline component, which is the part of the disparity not attributable to the treatment and cannot be affected by the intervention on the treatment.\footnote{The three-step interventional interpretation also elucidates the role of the treatment coding scheme. If the labels of treatment ($D=1$) and control ($D=0$) are swapped, the selection and prevalence components will remain unchanged, as their corresponding intervention steps are agnostic to treatment labels. However, the effect and baseline components will change, because the third intervention step assigns everyone to what is labeled as control. Therefore, the choice of treatment labels is a substantive decision to the extent that the researcher seeks to interpret the baseline and effect components.} We present a visualization of the unconditional decomposition in Appendix B.

Importantly, each step of the sequential intervention removes one single component without affecting any other components. Conversely, each component of the decomposition isolates a distinct policy lever that can be sequentially intervened upon. Therefore, estimates of our unconditional decomposition directly enable policy makers to choose between implementing just the first step, the first two steps, or all three steps together. Such freedom of choice is useful, for example, when the first step is effective for reducing the disparity, but the later two steps would either empirically increase the disparity or be infeasible to implement (e.g., due to cost concerns). In our empirical example, if the selection component is positive, then a low-cost information intervention might ameliorate the income disparity by preventing suboptimal selection into college among lower-income students, even if the prevalence and effect components do not provide viable disparity-reducing policies.

Comparison with prior work

We compare our new approach to three prior decomposition frameworks. We highlight that no prior decomposition contains a selection component; furthermore, prior approaches require strong assumptions to recover any of our other components.

Comparison with the KBO decomposition

Disparities research in the social and health sciences traditionally employs KBO decompositions. The KBO decomposition that most closely resembles our approach decomposes the outcome disparity between groups into four components with respect to treatment $D$ and pre-treatment covariates ${\boldsymbol X}$:\footnote{In practice, many KBO decompositions are farther removed from our approach. In particular, research often does not separate $D$ from ${\boldsymbol X}$ or heed the temporal order of variables.}

align[align omitted — 576 chars of source]

where $\alpha_g$, $\beta_g$, and $\boldsymbol{\gamma}_g$ are coefficients from the group-specific linear regression that contains only main effects: $$Y=\alpha_g+\beta_g D + \boldsymbol{\gamma}_g^\intercal {\boldsymbol X} + \epsilon.$$

This decomposition attains a causal interpretation under (i) the causal assumption of conditional ignorability of the treatment, $Y^d \rotatebox[origin=c]{90}{$\models$} D \mid g, {\boldsymbol x}$, $\forall d, g, {\boldsymbol x}$, and (ii) the parametric assumption that the group-specific linear regressions are correctly specified. If and only if both assumptions are satisfied, the endowment and slope components in the KBO decomposition are equivalent to our prevalence and effect components, respectively; and the sum of the intercept and residual components equals our baseline component.\footnote{Prior work has offered alternative causal interpretations for KBO decompositions under various assumptions. For example, a prominent literature shows that KBO decompositions can estimate the average treatment effect on the treated fortin_decomposition_2011, kline_oaxaca-blinder_2011, yamaguchi_decomposition_2015. Similarly, chernozhukov_sorted_2018 show that a KBO decomposition can estimate the partial treatment effect. And huber_causal_2015 discusses using a KBO decomposition for estimating the natural indirect effect pearl_direct_2001. None of these interpretations accommodates a descriptive group variable, with one important exception: jackson_decomposition_2018 establish a connection between the KBO decomposition ((ref)) and the unconditional random equalization decomposition that is discussed in Section (ref).}

Our decomposition ((ref)) differs from the KBO decomposition in three respects. First, our decomposition is inherently causal, because it is directly formulated as estimands in potential outcomes notation. In contrast, the KBO decomposition requires additional assumptions to support a causal interpretation. Such assumptions are rarely stated in practice. Second, our decomposition is nonparametric, whereas the KBO decomposition is model-based and hence relies on a particular functional form. Third, the KBO decomposition does not contain a selection component, because the assumed functional form imposes effect homogeneity within each group. By contrast, our nonparametric decomposition does not impose any effect homogeneity and hence contains a selection component as a distinctive conceptual contribution.

Comparison with causal mediation analysis

CMA decomposes a total effect of an exposure on an outcome into components in terms of an intermediate mediator. When CMA is used to understand a group disparity, our group variable is treated as the exposure, and our treatment variable becomes the mediator. The CMA literature is vast and contains many different decompositions vanderweele_explanation_2015. We focus on a three-way CMA decomposition vanderweele_attributing_2014,vanderweele_unification_2014 that is useful for illustrating similarities and differences with our decomposition ((ref)):

align[align omitted — 594 chars of source]

where $Y^g$ and $D^g$ are, respectively, the potential outcomes of $Y$ and $D$ when assigned group $g$, and $Y^{g,d}$ is the potential outcome of $Y$ when jointly assigned both group $g$ and treatment $d$.

Certain equivalences between equation ((ref)) and our unconditional decomposition can be established under very strong assumptions: two unconditional ignorability assumptions for $G$, i.e., $Y^{g,d} \rotatebox[origin=c]{90}{$\models$} G$, $\forall d,g$, and $D^g \rotatebox[origin=c]{90}{$\models$} G$, $\forall g$; two SUTVA-type assumptions, $\operatorname{E}_g \left(Y^{g,d} \right)=\operatorname{E}_g \left(Y^d \right)$ and $\operatorname{E}_g \left(D^g \right)=\operatorname{E}_g (D), \forall d,g$; and a cross-world independence assumption pearl_direct_2001, $Y^{g,d} \rotatebox[origin=c]{90}{$\models$} D^{g'}, \forall d,g,g'$. Under these assumptions, the conditional direct effect (CDE) equals our baseline component; the pure indirect effect (PIE) equals our prevalence component; and the portion attributable to interaction (PAI) equals our effect component. These equivalences are intuitive: both the CDE and the baseline component capture a group-based outcome difference when the intermediate variable, $D$, is held at $0$; both the PIE and the prevalence component address the role of the prevalence of $D$ in the relationship between the group and the outcome; finally, the PAI and the effect component both reflect how the effect of $D$ interacts with group membership.

However, all existing CMA decompositions, to our knowledge, differ from our unconditional decomposition in three crucial respects. First, CMA decomposes a different quantity: a total effect of group membership, rather than the observed, descriptive, group disparity. Our focus on decomposing descriptive disparities is useful, because descriptive disparities between groups are often the object of interest in their own right in the social and health sciences and the focus of popular and policy concerns. For example, income disparities in adulthood between the children of rich and poor parents are often viewed as concerning, regardless of whether these disparities originate from the causal effect of parental income or from confounding factors such as parents’ education and race. Furthermore, some group variables, such as race and gender, may be immutable attributes on which it is hard to define an intervention rubin_estimating_1974, holland_statistics_1986. As a result, CMA estimands, when encoding a hypothetical intervention on group membership, may not even be well-defined.

Second, the identification assumptions of CMA are much stronger. Whereas, as we show in Section (ref), our approach requires only causal assumptions on $D$, CMA requires assumptions on both $D$ and $G$ vanderweele_explanation_2015. This is why, above, we need assumptions on $G$ to establish equivalences between equation ((ref)) and our unconditional decomposition.

Third, there is no selection component in CMA. Although it is possible to define an analogous selection component for CMA in terms of exposure-induced selection into the mediator based on the net effect of the mediator on the outcome, i.e., $\operatorname{Cov} \left(D^a, Y^{a, 1}-Y^{a, 0} \right)-\operatorname{Cov} \left(D^b, Y^{b, 1}-Y^{b, 0} \right)$, such a component, to our knowledge, has not appeared in existing CMA decompositions.\footnote{In a companion paper, we discuss in more detail the analogous selection component for CMA in relation to the difference between the total effect and its randomized interventional analogue yu2024naturalmediationeffectsdiffer. Also, zhou_attendance_2024 presents a decomposition that contains the first half of CMA's selection component, $\operatorname{Cov} (D^a, Y^{a, 1}-Y^{a, 0} )$.}

Comparison with the unconditional random equalization decomposition

The newest entry into decomposition methodology is the random equalization decompositions vanderweele_causal_2014, jackson_decomposition_2018, sudharsanan_educational_2021, lundberg_gap-closing_2024, which exist in unconditional (URED) and conditional versions (CRED).

The URED is defined in terms of a one-step intervention that equalizes treatment prevalence by assigning random draws of the treatment from the advantaged group, $R(D \mid a)$, to the disadvantaged group jackson_decomposition_2018:

gather[gather omitted — 312 chars of source]

Thus, the outcome disparity is decomposed into the change in disparity resulting from the intervention and the remaining disparity after the intervention.

The URED shares several similarities with our approach. It decomposes the observed disparity, and it is formulated as causal estimands with interventional interpretations. Furthermore, when there is no selection into treatment, the change in disparity of the URED equals our prevalence component, and the remaining disparity equals the sum of our baseline and effect components.

However, the URED also differs from our unconditional decomposition in important respects. First, as a two-way decomposition, the URED contains less information than our four-way decomposition. Correspondingly, the URED only informs a one-step intervention, which is less flexible than the three-step sequential intervention our decomposition provides.

Second, when selection is present, the URED does not isolate any of the mechanisms identified in our decomposition ((ref)). Notably, the URED does not have a component isolating the contribution of differential treatment prevalence across groups. Formally, the change in disparity of the URED equals $\operatorname{E}_b(\tau)[\operatorname{E}_a(D)-\operatorname{E}_b(D)]-\operatorname{Cov}_b(D, \tau)$, which mixes our prevalence component with selection into treatment in the disadvantaged group. Intuitively, by assigning the disadvantaged group random draws of the treatment, the URED's one-step intervention simultaneously equalizes treatment prevalence across groups and randomizes treatment assignment in the disadvantaged group. By contrast, as noted in Section (ref), the three-step sequential intervention corresponding to our unconditional decomposition neatly separates equalization from randomization. Consequently, unlike our prevalence component, the URED's change in disparity is generally non-zero even if there is no difference in treatment prevalence, $\operatorname{E}_a(D)=\operatorname{E}_b(D)$.

By the same token, the URED does not contain a selection component, because the contribution of differential selection to the disparity is split between the URED's two components. The change in disparity contains the negative of selection into treatment in the disadvantaged group, and the remaining disparity contains selection into treatment in the advantaged group alongside our baseline and effect components, $\operatorname{E}_a(Y^0)-\operatorname{E}_b(Y^0)+ \operatorname{E}_a(D)[\operatorname{E}_a(\tau)-\operatorname{E}_b(\tau)] + \operatorname{Cov}_a(D,\tau)$.\footnote{lundberg_gap-closing_2024 introduces a variant of the URED whose one-step intervention equalizes treatment prevalence across groups by assigning random draws from the marginal treatment distribution in the population, $R(D)$, to both groups. The resulting change in disparity is $\operatorname{E}_a(Y)-\operatorname{E}_b(Y)-\left[\operatorname{E}_a \left(Y^{R(D)} \right)-\operatorname{E}_b \left(Y^{R(D)} \right) \right]$. Rewriting this as $\operatorname{E}(\tau)[\operatorname{E}_a(D)-\operatorname{E}_b(D)] + \operatorname{Cov}_a(D, \tau) - \operatorname{Cov}_b(D, \tau) - [\operatorname{Pr}(G=a)-\operatorname{Pr}(G=b)][\operatorname{E}_a(D) - \operatorname{E}_b(D)][\operatorname{E}_a(\tau)-\operatorname{E}_b(\tau)] $ shows that this URED variant also fails to separate the contributions of differential prevalence and differential selection.}

Conditional decomposition

Our conditional decomposition evaluates the contributions of the treatment to the outcome disparity conditional on a vector of pre-treatment covariates ${\boldsymbol Q}$. Thus, the conditional decomposition quantifies how the disparity would change if differences in treatment selection, prevalence, and effect were removed only between members of the advantaged and disadvantaged groups who share the same values of ${\boldsymbol Q}$. The conditional decomposition is useful in practice because it informs interventions that may be normatively more desirable or practically more feasible than their unconditional counterparts jackson_meaningful_2021.\footnote{jackson_meaningful_2021 calls ${\boldsymbol Q}$ “treatment-allowable covariates” in the sense that the analyst allows group differences in treatment assignment that are associated with ${\boldsymbol Q}$ to remain and does not investigate their remediation.}

In our application, ${\boldsymbol Q}$ is academic achievement in high school. The conditional decomposition thus informs interventions on college graduation that are conducted only within levels of prior achievement, preserving the relationship between prior achievement and college degree attainment.

assu[Common support] $\text{supp}_a({\boldsymbol Q}) = \text{supp}_b({\boldsymbol Q})$.

In order to introduce the conditional decomposition, we assume common support on ${\boldsymbol Q}$, which rules out the scenario where certain ${\boldsymbol Q}$ values only exist in one group. Under Assumptions (ref) and (ref), we then obtain the following conditional decomposition:

align[align omitted — 1,143 chars of source]

The baseline component in the conditional decomposition is the same as the baseline component in the unconditional decomposition: it is the level of the disparity had nobody received the treatment. The conditional prevalence component represents the part of the disparity that is due to differences in treatment prevalence across groups within levels of ${\boldsymbol Q}$. In our application, it indicates the extent to which the difference in college graduation rates between equally achieving advantaged and disadvantaged students contributes to the income disparity in adulthood. The conditional effect component reflects the contribution of the difference in the group-specific conditional average treatment effect (CATE) given ${\boldsymbol Q}$. Thus, it reveals the contribution of group-differential CATEs of college graduation given prior academic achievement. The conditional selection component is the contribution of group-differential selection into treatment net of ${\boldsymbol Q}$, for example, the contribution of differential sorting into college graduation after accounting for achievement in high school.

Analogous to the unconditional total contribution of the treatment, the conditional total contribution of the treatment is thus the sum of the conditional prevalence, effect, and selection components. The ${\boldsymbol Q}$-distribution component, which appears only in the conditional decomposition, equals the difference between the unconditional and conditional total contributions of the treatment to the outcome disparity. Hence, the ${\boldsymbol Q}$-distribution component is the between-${\boldsymbol Q}$, as opposed to within-${\boldsymbol Q}$, part of the unconditional total contribution of the treatment.\footnote{Furthermore, the ${\boldsymbol Q}$-distribution component identifies the amount of the disparity associated with (i) the relationship between $G$ and ${\boldsymbol Q}$ and (ii) the within-group relationship between $\{D,\tau\}$ and ${\boldsymbol Q}$. Clearly, if $G \rotatebox[origin=c]{90}{$\models$} {\boldsymbol Q}$ or $\{D,\tau\} \rotatebox[origin=c]{90}{$\models$} {\boldsymbol Q} \mid g$, the ${\boldsymbol Q}$-distribution component will be zero. In our application, if there is no group difference in prior achievement or if prior achievement is not associated with college graduation and the effect of college graduation on adult income, then the ${\boldsymbol Q}$-distribution component will be zero. }

Interventional interpretation

Similar to the unconditional case, our conditional decomposition has an interventional interpretation that maps a three-step sequential intervention to the decomposition components. Let $\operatorname{E}_g \left(Y^{R(D \mid g',{\boldsymbol Q})} \right)$ be the mean potential outcome of group $g$ when its members with certain ${\boldsymbol Q}$ values are given treatments randomly drawn from those members of group $g'$ who have the same ${\boldsymbol Q}$ values. We can rewrite this mean potential outcome as:

align[align omitted — 275 chars of source]

This allows us to express the components of the conditional decomposition as:

align[align omitted — 1,252 chars of source]

Therefore, the first step of the intervention randomizes the treatment within groups and within ${\boldsymbol Q}$ levels. In this step, the pre-intervention disparity is $\operatorname{E}_a(Y)-\operatorname{E}_b(Y)$, the post-intervention disparity is $\operatorname{E}_a \left(Y^{R(D \mid a,{\boldsymbol Q})} \right)-\operatorname{E}_b \left(Y^{R(D \mid b,{\boldsymbol Q})} \right)$, and the resulting change in disparity equals the conditional selection component. The second step equalizes treatment prevalence across groups but within ${\boldsymbol Q}$ levels. In this step, the pre-intervention disparity is $\operatorname{E}_a \left(Y^{R(D \mid a,{\boldsymbol Q})} \right)-\operatorname{E}_b \left(Y^{R(D \mid b,{\boldsymbol Q})} \right)$, the post-intervention disparity is $\operatorname{E}_a \left(Y^{R(D \mid a,{\boldsymbol Q})} \right) - \operatorname{E}_b \left(Y^{R(D \mid a,{\boldsymbol Q})} \right)$, and the change in disparity equals the conditional prevalence component. The third step sets the treatment to zero for all individuals. Thus, this final step retains only the baseline component as the post-intervention disparity, while the change in disparity is the sum of the conditional effect and ${\boldsymbol Q}$-distribution components.\footnote{Different from the effect component in the unconditional decomposition, there does not exist an obvious interventional step that isolates the conditional effect component itself. Setting $D=0$ within each level of ${\boldsymbol Q}$ in analogy to the third step of the unconditional decomposition sets $D=0$ marginally and hence eliminates both the conditional effect and ${\boldsymbol Q}$-distribution components. Equalizing the CATE given ${\boldsymbol Q}$ across groups would eliminate the conditional effects component alone, but, as noted in Footnote (ref), interventions on effects are unconventional within the potential outcomes framework.} \footnote{For equation ((ref)) to hold, we only require $\text{supp}_g({\boldsymbol Q}) \subseteq \text{supp}_{g'}({\boldsymbol Q})$. This implies that the three-step intervention is well-defined as long as $\text{supp}_b({\boldsymbol Q}) \subseteq \text{supp}_{a}({\boldsymbol Q})$. Intuitively, at each level of ${\boldsymbol Q}$ in group $b$, we must be able to find members of group $a$ with the same ${\boldsymbol Q}$ values in order to conduct the equalization intervention (the second step). Note that $\text{supp}_b({\boldsymbol Q}) \subseteq \text{supp}_{a}({\boldsymbol Q})$ is a weaker condition than Assumption (ref). However, under the weaker condition, only the pre- and post-intervention disparities are well-defined, not all components in equation ((ref)).}

Comparison with path-specific effects in causal mediation analysis

Since there is no apparent counterpart of our conditional decomposition in the KBO framework, we proceed to comparison with CMA. As a subfield of CMA, path-specific effects provide fine-grained decompositions of the total effect of an exposure when there is more than one mediator avin_identiability_2005, vanderweele_explanation_2015. Treating the group variable $G$ as the exposure and all variables in $\{{\boldsymbol Q}, D\}$ as mediators, such that ${\boldsymbol Q}$ is temporally after $G$ but before $D$, we have the following path-specific decomposition of the total effect of the exposure:

align*[align* omitted — 781 chars of source]

where $Y^{g, {\boldsymbol Q}^{g'}, D^{g'', {\boldsymbol Q}^{g'''}}}$ is a potential outcome with “nested counterfactuals” indicating the assignment of $G$, ${\boldsymbol Q}$, and $D$. The total effect of $G$ is decomposed into the paths involving neither ${\boldsymbol Q}$ nor $D$ (i.e., $A \rightarrow Y$), the paths involving ${\boldsymbol Q}$ (i.e., the combination of $G \rightarrow {\boldsymbol Q} \rightarrow D \rightarrow Y$ and $G \rightarrow {\boldsymbol Q} \rightarrow Y$), and the paths involving $D$ but not ${\boldsymbol Q}$ (i.e., $G \rightarrow D \rightarrow Y$). Figure A1 in Appendix A.8. illustrates these paths.

The equivalence between the paths involving $D$ but not ${\boldsymbol Q}$ and the conditional prevalence component can be established under some very strong assumptions detailed in Appendix A.8. This equivalence is intuitive, as our conditional prevalence component captures the contribution of differential prevalence of the treatment $D$ after accounting for ${\boldsymbol Q}$. However, none of other components of our conditional decomposition is isolated in the path-specific decomposition under these assumptions. Moreover, unlike the path-specific decomposition, our conditional decomposition does not impose any temporal order between $G$ and ${\boldsymbol Q}$ or require ignorability assumptions on $G$ or ${\boldsymbol Q}$.

Comparison with the conditional random equalization decomposition

Analogous to the URED in Section (ref), the CRED jackson_meaningful_2021 is defined in terms of a one-step intervention that equalizes treatment prevalence within levels of ${\boldsymbol Q}$ by assigning members of the disadvantaged group who have certain ${\boldsymbol Q}$ values to random draws of the treatment from advantaged group members who have the same ${\boldsymbol Q}$ values. The CRED thus decomposes the outcome disparity into the change in disparity resulting from this intervention and the corresponding remaining disparity:

gather[gather omitted — 346 chars of source]

As before, the key difference between our conditional decomposition and the CRED is that the CRED is a two-way decomposition that does not isolate any of the mechanisms identified in our conditional decomposition. Notably, the CRED does not have a conditional prevalence component. Formally, the CRED's change in disparity equals

gather[gather omitted — 270 chars of source]

which mixes our conditional prevalence component with the disadvantaged group's conditional selection. Intuitively, within ${\boldsymbol Q}$ levels, the CRED's intervention simultaneously equalizes treatment across groups and randomizes treatment in the disadvantaged group. Consequently, unlike our conditional prevalence component, the URED’s change in disparity is generally non-zero even if there is no difference in treatment prevalence given any ${\boldsymbol Q}$ value, $\operatorname{E}_a(D \mid {\boldsymbol q}) = \operatorname{E}_b(D \mid {\boldsymbol q}), \forall {\boldsymbol q}$. The CRED does not contain a conditional selection component either, because the CRED splits the contribution of differential conditional selection between its two components.\footnote{lundberg_gap-closing_2024 also proposes a variant of the CRED, whose change in disparity can be written as $\operatorname{E}_a[\operatorname{Cov}_a(D,\tau \mid {\boldsymbol Q})] - \operatorname{E}_b[\operatorname{Cov}_b(D,\tau \mid {\boldsymbol Q})] + \int[\operatorname{E}_a(D \mid {\boldsymbol q}) - \operatorname{E}_b(D \mid {\boldsymbol q})][\operatorname{E}_a(\tau \mid {\boldsymbol q})f_a({\boldsymbol q}) p_b({\boldsymbol q}) + \operatorname{E}_b(\tau \mid {\boldsymbol q})f_b({\boldsymbol q}) p_a({\boldsymbol q})] \dd {\boldsymbol q}$, where $p_g({\boldsymbol q}) = \operatorname{Pr}(G=g \mid {\boldsymbol q})$. Hence, this CRED variant also does not separate the contributions of differential conditional prevalence and differential conditional selection.}

Identification, estimation, and inference

We identify our unconditional and conditional decompositions using the standard assumptions of conditional ignorability and overlap. Without loss of generality, let ${\boldsymbol Q} \subseteq {\boldsymbol X}$.

assu[Conditional ignorability] $Y^d \rotatebox[origin=c]{90}{$\models$} D \mid {\boldsymbol x}, g$, $\forall d, {\boldsymbol x}, g$.
assu[Overlap] $0 < \operatorname{E}(D \mid {\boldsymbol x}, g) <1$, $\forall {\boldsymbol x}, g$.

We develop nonparametric and efficient estimators for our decompositions. These estimators are “one-step” estimators based on the EIFs of the decomposition components, which remove the bias from naive substitution estimators bickel_efficient_1998, van_der_vaart_asymptotic_2000,hines_demystifying_2022. The estimators contain some nuisance functions, which can be estimated using flexible ML methods coupled with cross-fitting. Under conditions specified below, our estimators are $\sqrt{n}$-consistent, asymptotically normal, and semiparametrically efficient. Thus, we are able to construct asymptotically accurate Wald-type confidence intervals and hypothesis tests. Our estimators also have double or quadruple robustness properties.

To unburden notation, we define the following functions of the observed data: $\mu(d,{\boldsymbol X}, g)=\operatorname{E}(Y \mid d, {\boldsymbol X}, g)$, $\pi(d,{\boldsymbol X}, g) = \operatorname{Pr}(D=d \mid {\boldsymbol X}, g)$, and $\omega(d, {\boldsymbol Q}, g)=\operatorname{E}[\mu(d,{\boldsymbol X},g) \mid {\boldsymbol Q},g]$. Also recall that $p_g = \operatorname{Pr}(G=g)$, and $p_g({\boldsymbol Q}) = \operatorname{Pr}(G=g \mid {\boldsymbol Q})$. We use circumflexes to denote estimated quantities.

Unconditional decomposition

All components of the unconditional decomposition can be expressed as linear combinations of the total disparity and two generic functions of the potential outcomes evaluated at appropriate values of $d$, $g$, and $g'$: $\xi_{dg} \vcentcolon= \operatorname{E} \left(Y^d \mid g \right)$ and $\xi_{dgg'} \vcentcolon= \operatorname{E} \left(Y^d \mid g \right)\operatorname{E} \left(D \mid g' \right)$. The relationship between the components of the unconditional decomposition and the generic functions is as follows:

align*[align* omitted — 326 chars of source]

Hence, the EIFs and one-step estimators for the decomposition components directly follow from those for $\xi_{dg}$ and $\xi_{dgg'}$.\footnote{These generic functions also provide a basis for the estimation of Jackson and VanderWeele's (jackson_decomposition_2018) version of the URED, since its change in disparity can be represented as $\xi_{0b} + \xi_{1ba}-\xi_{0ba}-\operatorname{E}_b(Y)$.} Under Assumptions (ref), (ref), and (ref), $\xi_{dg}$ and $\xi_{dgg'}$ can be identified as the following nonparametric functionals:

align*[align* omitted — 218 chars of source]

These identification results then enable the derivation of the EIFs for $\xi_{dg}$ and $\xi_{dgg'}$.\footnote{When the conditional ignorability assumption ((ref)) is not credible, researchers may evaluate the robustness of their causal conclusions by performing a sensitivity analysis, as we do in Appendix H. Alternatively, they may abandon the causal interpretation in favor of a descriptive interpretation. The nonparametric descriptive quantity $\operatorname{E}[\mu(1,{\boldsymbol X},G)-\mu(0,{\boldsymbol X},G)]$ has been labeled the average controlled difference (ACD) between treatment and control given ${\boldsymbol X}$ and $G$ li_propensity_2013, li_balancing_2018. In the same spirit, one can also descriptively interpret the identified nonparametric functionals for our decomposition components. We define $\operatorname{E}[\mu(1,{\boldsymbol X},g)-\mu(0,{\boldsymbol X},g) \mid g]$ as the group specific ACD, $\mu(1,{\boldsymbol X},G)-\mu(0,{\boldsymbol X},G)$ as the controlled outcome difference (COD), and $\operatorname{E}(D \mid {\boldsymbol X},G)$ as the controlled treatment prevalence (CTP). Then, the functional for the prevalence component is the group difference in treatment prevalence scaled by group $b$'s ACD, the functional for the effect component is the group difference in ACD scaled by group $a$'s treatment prevalence, and the functional for the selection component is the group difference in the covariance between the COD and the CTP. Formally,

align*[align* omitted — 664 chars of source]

}

prop[EIF, unconditional decomposition] Under Assumptions (ref), (ref), and (ref), the EIF of $\xi_{dg}$ is $$\phi_{dg}(Y,D,{\boldsymbol X},G) \vcentcolon= \frac{\mathbb{1}(G=g)}{p_g} \left\{ \frac{\mathbb{1}(D=d)}{\pi(d, {\boldsymbol X},g)} [Y-\mu(d,{\boldsymbol X},g)] + \mu(d,{\boldsymbol X},g) - \xi_{dg} \right\},$$ and the EIF of $\xi_{dgg'}$ is \begin{align*} \phi_{dgg'}(Y,D,{\boldsymbol X},G) &\vcentcolon= \frac{\mathbb{1}(G=g)}{p_g} \left\{\frac{\mathbb{1}(D=d)}{\pi(d,{\boldsymbol X},g)}[Y-\mu(d,{\boldsymbol X},g)] + \mu(d,{\boldsymbol X},g) \right\} \operatorname{E} \left(D \mid g' \right) \\ &\phantom{=} + \frac{\mathbb{1}(G=g')}{p_{g'}} \xi_{dg} \left[D - \operatorname{E} \left(D \mid g' \right) \right] - \frac{\mathbb{1}(G=g)}{p_g} \xi_{dgg'}. \end{align*} In Appendix C, we also derive the general EIFs with survey weights.\footnote{Related EIFs have appeared in prior work. park_groupwise_2024 give the EIF for $\xi_{1g}-\xi_{0g}$. The EIF for $\xi_{0,g}+\xi_{1gg'}-\xi_{0gg'}$ coincides with the EIF for the quantity denoted as $\theta_c$ in diaz_nonparametric_2021, when pre-treatment confounders in the latter are omitted. Neither of these prior works accommodates survey weights.}

We use the EIFs as estimating equations, i.e., set their sample averages to zero and solve for $\xi_{dg}$ and $\xi_{dgg'}$. The one-step estimators of $\xi_{dg}$ and $\xi_{dgg'}$ thus are

align*[align* omitted — 513 chars of source]

Each estimator contains two nuisance functions, $\pi(d,{\boldsymbol X},g)$ and $\mu(d,{\boldsymbol X},g)$. The estimators are consistent as long as either one of the two nuisance functions is consistently estimated. Although $p_g$ and $\operatorname{E}(D \mid g)$ are technically also nuisance functions, their consistent estimation is left implicit hereafter.

prop[Double robustness in consistency, unconditional decomposition] Under Assumptions (ref), (ref), and (ref), either consistent estimation of $\mu(d,{\boldsymbol X},g)$ or of $\pi(d,{\boldsymbol X},g)$ is sufficient for the consistency of $\hat{\xi}_{dg}$ and $\hat{\xi}_{dgg'}$.\footnote{The estimator of our unconditional decomposition is doubly robust with respect to the same two nuisance functions as the classic augmented-inverse-probability-of-treatment-weighting (AIPW) estimator of the ATE robins_estimation_1994.}

The nuisance functions $\pi(d,{\boldsymbol X},g)$ and $\mu(d,{\boldsymbol X},g)$ can be estimated using various methods. We focus on a nonparametric approach where the nuisance functions are estimated using flexible ML models with cross-fitting. The use of cross-fitting allows for weaker conditions on the estimation of the nuisance functions chernozhukov_double/debiased_2018,kennedy_semiparametric_2022. In practice, we implement cross-fitting by randomly splitting the data into two disjoint subsamples, fitting the ML models in each subsample and plugging in values of ${\boldsymbol X}$ from the other sample to obtain the nuisance function estimates. To improve the finite-sample performance of the estimators, we stabilize the weight, $\mathbb{1}(D=d)/\hat{\pi}(d,{\boldsymbol X},g)$, by dividing it by its sample average.

To study the asymptotic distribution of the cross-fitted one-step estimators for the unconditional decomposition, we invoke three additional assumptions about the nuisance functions, which are the same as the assumptions required for the double ML estimator of the ATE chernozhukov_double/debiased_2018,kennedy_semiparametric_2022. We let $\| \cdot \|$ denote the $L_2$-norm. \refstepcounter{assu}

subassu[Boundedness] With probability 1, $\hat{\pi}(d,{\boldsymbol X},g) \geq \eta$, $\pi(d,{\boldsymbol X},g) \geq \eta$, and $|Y-\hat{\mu}(d,{\boldsymbol X},g)| \leq \zeta$, for some $\eta>0$ and some $\zeta < \infty$, $\forall d, g$.
subassu[Consistency] $\| \hat{\mu}(d,{\boldsymbol X},g) - \mu(d,{\boldsymbol X},g) \| =o_p(1)$, and $\| \hat{\pi}(d,{\boldsymbol X},g) - \pi(d,{\boldsymbol X},g) \| =o_p(1)$, $\forall d, g$.
subassu[Convergence rate] $\|\hat{\pi}(d,{\boldsymbol X},g)-\pi(d,{\boldsymbol X},g)\| \|\hat{\mu}(d,{\boldsymbol X},g)-\mu(d,{\boldsymbol X},g)\|=o_p(n^{-1/2})$, $\forall d, g$.
prop[Asymptotic distributions, unconditional decomposition] Under Assumptions (ref), (ref), (ref), (ref), (ref), and (ref), the cross-fitted one-step estimators for $\hat{\xi}_{dg}$ and $\hat{\xi}_{dgg'}$ are $\sqrt{n}$-consistent, asymptotically normal, and semiparametrically efficient, i.e., $\sqrt{n} \left(\hat{\xi}_{dg} - \xi_{dg} \right) \xrightarrow{d} \mathcal{N} \left(0, \sigma^2_{dg} \right)$, and $\sqrt{n}\left(\hat{\xi}_{dgg'} - \xi_{dgg'} \right) \xrightarrow{d} \mathcal{N} \left(0, \sigma^2_{dgg'} \right)$, where $\sigma^2_{dg} \vcentcolon= \operatorname{E}[\phi_{dg}(Y,D,{\boldsymbol X},G)^2]$ and $\sigma^2_{dgg'} \vcentcolon= \operatorname{E}[\phi_{dgg'}(Y,D,{\boldsymbol X},G)^2]$ are the respective semiparametric efficiency bounds.

We consistently estimate $\sigma^2_{dg}$ and $\sigma^2_{dgg'}$ using the averages of the squared estimated EIFs. The asymptotic distributions can then be used to construct hypothesis tests and confidence intervals. Since the unconditional decomposition components are simple additive functions of the observed disparity, $\xi_{dg}$, and $\xi_{dgg'}$, all properties established for the estimators of $\xi_{dg}$, and $\xi_{dgg'}$ (double robustness, $\sqrt{n}$-consistency, asymptotic normality, and semiparametric efficiency) carry over to the final estimators of the decomposition components.

Conditional decomposition

Relative to the unconditional decomposition, the estimation of the conditional decomposition requires consideration of one additional generic function: $$\xi_{dgg'g''} \vcentcolon= \operatorname{E} \left[\operatorname{E} \left(Y^d \mid {\boldsymbol Q}, g \right) \operatorname{E} \left(D \mid {\boldsymbol Q}, g' \right) \mid g'' \right],$$ where $(d, g, g', g'')$ denotes any one of eight combinations of treatment status and group memberships. The relationship between components of the conditional decomposition and the generic functions is as follows:

align*[align* omitted — 467 chars of source]

Since the estimation of $\xi_{dg}$ was discussed in the previous subsection, we now focus on $\xi_{dgg'g''}$. The EIFs, one-step estimators, and their asymptotic distributions for the components of the conditional decomposition will then follow.\footnote{We thereby also provide efficient and nonparametric estimation for the change in disparity in the CRED of jackson_meaningful_2021, which can be represented as $\xi_{0b}+\xi_{1bab}-\xi_{0bab}-\operatorname{E}_b(Y)$.} Under Assumptions (ref), (ref), (ref), and (ref), we identify $\xi_{dgg'g''}$ as $$\xi_{dgg'g''}=\operatorname{E} \left[\omega(d,{\boldsymbol Q},g) \operatorname{E} \left(D \mid {\boldsymbol Q}, g' \right) \mid g'' \right].$$

prop[EIF, conditional decomposition] Under Assumptions (ref), (ref), (ref), and (ref), the EIF of $\xi_{dgg'g''}$ is \begin{align*} &\phantom{=} \phi_{dgg'g”}(Y,D,{\boldsymbol X},G) \\ &=\frac{\mathbb{1}(G=g”)}{p_{g”}} \left[ \omega(d,{\boldsymbol Q},g) \operatorname{E} \left(D \mid {\boldsymbol Q}, g' \right) - \xi_{dgg'g”} \right] + \frac{\mathbb{1}(G=g')p_{g”}({\boldsymbol Q})}{p_{g'}({\boldsymbol Q})p_{g”}} \left[ D-\operatorname{E} \left(D \mid {\boldsymbol Q}, g' \right) \right] \omega(d,{\boldsymbol Q},g) \\ &\phantom{=} + \frac{\mathbb{1}(G=g) p_{g”}({\boldsymbol Q})}{p_g({\boldsymbol Q})p_{g”}} \left\{ \frac{\mathbb{1}(D=d)}{\pi(d,{\boldsymbol X},g)} [Y-\mu(d,{\boldsymbol X},g)]+\mu(d,{\boldsymbol X},g) - \omega(d,{\boldsymbol Q},g) \right\} \operatorname{E} \left(D \mid {\boldsymbol Q},g' \right). \end{align*} In Appendix C, we also derive the general EIF with survey weights.

We again construct the one-step estimator by using the EIF as an estimating equation.

align*[align* omitted — 799 chars of source]

This estimator contains five nuisance functions: $p_g({\boldsymbol Q})$, $\pi(d,{\boldsymbol X},g)$, $\mu(d,{\boldsymbol X},g)$, $\operatorname{E}(D \mid {\boldsymbol Q}, g)$, and $\omega(d,{\boldsymbol Q},g)$. As is the case for the unconditional decomposition, consistent estimation of the conditional decomposition does not require all nuisance functions to be consistently estimated. In particular, $\hat{\xi}_{dgg'g''}$ is quadruply robust to inconsistently estimated nuisance functions.

prop[Quadruple robustness in consistency, conditional decomposition] Under Assumptions (ref), (ref), (ref), and (ref), $\hat{\xi}_{dgg'g''}$ is consistent if one of four minimal conditions holds, as summarized in Table 1.
table[table omitted — 1,965 chars of source]

As before, we estimate the nuisance functions nonparametrically using ML with cross-fitting. To improve finite-sample performance, $\hat{\xi}_{dgg'g''}$ can be stabilized by dividing $\mathbb{1}(D=d)/\hat{\pi}(d,{\boldsymbol X},g)$, $\mathbb{1}(G=g')\hat{p}_{g''}({\boldsymbol Q})/\hat{p}_{g'}({\boldsymbol Q})\hat{p}_{g''}$, and $\mathbb{1}(G=g)\hat{p}_{g''}({\boldsymbol Q})/\hat{p}_g({\boldsymbol Q})p_{g''}$ by their respective sample averages.

The estimation of $\omega(d,{\boldsymbol Q},g)$ deserves particular attention, because its cross-fitting is nonstandard, and it can be doubly robust itself. Specifically, we adopt a pseudo-outcome approach van_der_laan_statistical_2006,semenova_debiased_2021, where the pseudo outcome for each $d$ is defined as $$ \delta_d \left(Y, D, {\boldsymbol X}, G \right) \vcentcolon= \frac{\mathbb{1}(D=d)}{\pi(d, {\boldsymbol X},G)} [Y-\mu(d,{\boldsymbol X},G)] + \mu(d,{\boldsymbol X},G),$$ which is motivated by the fact that $\omega(d,{\boldsymbol Q},g)=\operatorname{E} \left[\delta_d \left(Y, D, {\boldsymbol X}, G \right) \mid {\boldsymbol Q},g \right]$. We first randomly draw two disjoint subsamples from the data. Then we estimate $\pi(d, {\boldsymbol X},G)$ and $\mu(d,{\boldsymbol X},G)$ in each subsample without cross-fitting and obtain the estimated pseudo outcome $\hat{\delta}_d \left(Y, D, {\boldsymbol X}, G \right)$. Finally, we obtain estimates of $\omega(d,{\boldsymbol Q},g)$ using cross-fitting, i.e., we fit $\operatorname{E}\left[\hat{\delta}_d \left(Y, D, {\boldsymbol X}, G \right) \mid {\boldsymbol Q}, g \right]$ separately in each subsample and plug in values of ${\boldsymbol Q}$ from the respective other subsample. Using this procedure, we ensure that the fitting of $\omega(d,{\boldsymbol Q},g)$, which relies on estimating the pseudo outcome, is done separately in each subsample. Provided that $\operatorname{E}\left[\hat{\delta}_d \left(Y, D, {\boldsymbol X}, G \right) \mid {\boldsymbol Q}, g \right]$ can be consistently estimated, this approach enables consistent estimation of $\omega(d,{\boldsymbol Q},g)$ if either $\mu(d,{\boldsymbol X},g)$ or $\pi(d,{\boldsymbol X},g)$ is consistently estimated.

To establish the asymptotic distributions of the cross-fitted one-step estimators for the conditional decomposition, we invoke Assumptions (ref), which augments Assumptions (ref) with respect to the additional nuisance functions needed for the conditional decomposition. \refstepcounter{assu}

subassu[Boundedness] With probability 1, $\hat{\pi}(d,{\boldsymbol X},g) \geq \eta$, $\pi(d,{\boldsymbol X},g) \geq \eta$, $\hat{p}_g({\boldsymbol Q}) \geq \eta$, $p_g({\boldsymbol Q}) \geq \eta$, $|Y-\hat{\mu}(d,{\boldsymbol X},g)| \leq \zeta$, $|Y-\mu(d,{\boldsymbol X},g)| \leq \zeta$, $|\mu(d,{\boldsymbol X},g)| \leq \zeta$, $|\omega(d,{\boldsymbol Q},g)| \leq \zeta$, and $\left| \frac{\hat{p}_g({\boldsymbol Q})}{\hat{p}_{g'}({\boldsymbol Q})} \right| \leq \zeta$, for some $\eta>0$ and $\zeta < \infty$, $\forall d,g,g'$.
subassu[Consistency] $\| \hat{\mu}(d,{\boldsymbol X},g) - \mu(d,{\boldsymbol X},g) \| =o_p(1)$, $\| \hat{\pi}(d,{\boldsymbol X},g) - \pi(d,{\boldsymbol X},g) \| =o_p(1)$, $\left\| \hat{\omega}(d,{\boldsymbol Q},g) - \omega(d,{\boldsymbol Q},g) \right\| =o_p(1)$, $\left\| \hat{\operatorname{E}}(D \mid {\boldsymbol Q},g) - \operatorname{E}(D \mid {\boldsymbol Q},g) \right\| =o_p(1)$, and $\left\| \hat{p}_g({\boldsymbol Q}) -p_g({\boldsymbol Q}) \right\|=o_p(1)$, $\forall d,g$.
subassu[Convergence rate] First, we require $\left\|\hat{\pi}(d,{\boldsymbol X},g)-\pi(d,{\boldsymbol X},g) \right\| \|\hat{\mu}(d,{\boldsymbol X},g)-\mu(d,{\boldsymbol X},g)\|=o_p(n^{-1/2}), \forall d,g$. Second, depending on the specific combination of $g,g'$, and $g''$ in a $\xi_{dgg'g''}$, we require $\left\| \hat{\operatorname{E}}(D \mid {\boldsymbol Q}, g) - \operatorname{E}(D \mid {\boldsymbol Q}, g) \right\| = o_p(n^{-1/2}), \forall g$, when $g=g''$; $\left\| \hat{\omega}(d,{\boldsymbol Q},g) - \omega(d,{\boldsymbol Q},g) \right\|= o_p(n^{-1/2}), \forall d,g$, when $g'=g''$; and \\ $\left\| \hat{\omega}(d,{\boldsymbol Q},g) - \omega(d,{\boldsymbol Q},g) \right\| \left\| \hat{\operatorname{E}}(D \mid {\boldsymbol Q}, g) - \operatorname{E}(D \mid {\boldsymbol Q}, g) \right\| = o_p(n^{-1/2}), \forall d,g$, when $g=g'=g''$.
prop[Asymptotic distribution, conditional decomposition] Under Assumptions (ref), (ref), (ref), (ref), (ref), (ref), and (ref), the cross-fitted one-step estimator of $\xi_{dgg'g''}$ is $\sqrt{n}$-consistent, asymptotically normal, and semiparametrically efficient, i.e., $\sqrt{n} \left( \hat{\xi}_{dgg'g''} - \xi_{dgg'g''} \right) \xrightarrow{d} \mathcal{N} \left(0, \sigma^2_{dgg'g''} \right)$, where $\sigma^2_{dgg'g''} \vcentcolon= \operatorname{E} \left[\phi_{dgg'g''}(Y,D,{\boldsymbol X},G)^2 \right]$ is the semiparametric efficiency bound.

Since the conditional decomposition components are additive functions of the observed disparity, $\xi_{dg}$, and $\xi_{dgg'g''}$, double robustness, $\sqrt{n}$-consistency, asymptotic normality, and semiparametric efficiency all carry over to the final estimators of the decomposition components. We conduct hypothesis tests and construct confidence intervals analogously to the unconditional decomposition.

Finally, for ML-based estimation of the conditional decomposition, we note a tension between asymptotic normality and semiparametric efficiency on one hand, and consistency on the other. ML does not typically satisfy the convergence rate conditions for establishing asymptotic normality and semiparametric efficiency for three components of the conditional decomposition: conditional prevalence, conditional effect, and ${\boldsymbol Q}$-distribution. However, we still prefer ML for all components because it achieves consistent estimation without imposing parametric assumptions. Note that this tension does not exist for the baseline and conditional selection components, because their convergence rate conditions can reasonably be achieved by ML.

Application

Overview

Our application decomposes the contributions of college graduation to the perpetuation of income inequality across generations, which is also known as intergenerational income persistence, the complement to income mobility. Groups ($G$) are defined by parental income; the outcome ($Y$) is offspring's adult income; and the treatment ($D$) is college graduation. Although previous research has indirectly touched upon all four components of our unconditional decomposition to some extent, this analysis is the first to present a direct and unified decomposition. Our findings reveal that differential selection into college completion by income origin reduces intergenerational income persistence and increases mobility. This result has not previously been discovered and introduces a new mechanism into the classic Origin-Education-Destination triangle breen_social_2004 in social stratification research.

The baseline component of our decompositions represents the part of the disparity in adult income that is unaccounted for by college graduation. Prior research enumerates multiple sources of intergenerational income persistence that may operate independently of college graduation. For example, parental income is associated with a variety of pre-college characteristics that may directly influence income attainment, such as cognitive skills and noncognitive traits in adolescence heckman_effects_2006. Moreover, people from more privileged backgrounds likely benefit from their parents' human, social, and financial capitals regardless of their own formal educational attainment. Interventions on college graduation would not eliminate these channels of income persistence.

Conceptually speaking to the prevalence component in our unconditional decomposition, social scientists have long regarded education as a mediator in the intergenerational reproduction of socioeconomic inequalities (blau_american_1978; featherman_opportunity_1978; ishida_class_1995). Specifically, research documents large differences in college graduation rates across parental income groups ziol-guest_parent_2016, bailey_gains_2011, and simulations suggest that rising educational inequalities have strengthened intergenerational income persistence over time bloome_educational_2018. Nonetheless, prior work is descriptive in nature.

Related to our effect component, there is an active literature on heterogeneity in the effects of college graduation on adult income, although results vary. Brand et al. (brand_uncovering_2021) and Cheng et al. (cheng_heterogeneous_2021) find larger effects of college graduation on the income of people from more disadvantaged backgrounds. By contrast, zhou_equalization_2019, fiel_great_2020, and yu_leveraging_2021 find no statistically significant heterogeneity in the effects of college completion on income across parental income groups. However, none of these works evaluates the extent to which income disparities can be attributed to groupwise differential effects of college.

Finally, some prior work has addressed selection into college as a function of college effects on income, i.e. $\operatorname{Cov}(D, \tau)$. Here, too, results are mixed. brand_who_2010 and brand_uncovering_2021 find negative selection, i.e., those who are least likely to attend college would benefit most from it. By contrast, the instrumental variable analysis of heckman_returns_2018 finds positive selection into college. However, prior work estimates selection in the pooled population rather than within parental income groups, thereby missing the link between the difference in group-specific selection and the group-based outcome disparity that our approach identifies. (Appendix F clarifies and synthesizes related notions of “selection into college” in the social science literature.)

Data, variables and estimation

We analyze the National Longitudinal Survey of Youth 1979, a nationally representative U.S. cohort study of individuals born between 1957 and 1964. We restrict the sample to respondents who were between 14 to 17 years old at baseline in 1979 to ensure that income origin is measured prior to respondents' college graduation. We also limit the analysis to respondents who graduated from high school by age 29. The sample size of our complete-case analysis is $N=2,008$. Missingness mostly occurs in the outcome variable due to loss to follow-up (22%), with less missingness in other variables (<6%). Appendix Table A1 presents associations between outcome missingness and baseline covariates.

We contrast parental income-origin groups ($G$), defined as the top 40% and bottom 40% of family income averaged over the first three waves of the survey (1979, 1980, and 1981, when respondents were 14 to 20 years old) and divided by the square root of the family size to adjust for need zhou_equalization_2019. The treatment ($D$) is a binary indicator of whether the respondent graduated from college by age 29. The outcome is the percentile rank of the respondent's adult income, averaged over five survey waves between age 35 and 44, divided by the square root of family size. For the conditional decomposition, we define ${\boldsymbol Q}$ as the Armed Forces Qualification Test (AFQT) score, measured in 1980. The AFQT score is a widely used measure of academic achievement that predicts college completion.

We measure an extensive set of confounders (${\boldsymbol X}$) at baseline, including gender, race, parental income percentile, parental education, parental presence, the number of siblings, urban residence, educational expectation, friends' educational expectation, AFQT score, age at the baseline survey, the Rotter score of control locus, the Rosenberg self-esteem score, language spoken at home, Metropolitan Statistical Area category, separation from mother, school satisfaction, region of residence, and mother's working status.

We present four estimates each for our unconditional and conditional decompositions, using different models for the nuisance functions to assess robustness: three alternative ML methods (gradient boosting machine [GBM], neural networks, and random forests) and one set of parametric models. Depending on whether the left-hand-side variable is continuous or binary, the parametric models are linear or logit. Specifically, for $\mu(d,{\boldsymbol X},g)$, we use all two-way interactions between $D$ and $\{{\boldsymbol X}, G\}$, along with their main effects. For $\pi(d,{\boldsymbol X},g)$ and $p_g({\boldsymbol Q})$, the logit models contain only main effects. For $\operatorname{E}(D \mid {\boldsymbol Q},g)$ and $\operatorname{E}[\hat{\delta}_d \left(Y, D, {\boldsymbol X}, G \right) \mid {\boldsymbol Q}, g ]$, we include all main effects and two-way interactions between $G$ and ${\boldsymbol Q}$. Appendix H contains model diagnostics and a sensitivity analysis for unobserved confounding. Statistical significance is assessed at the 0.05 level.

Results

Figure (ref) presents our main results for the components of the unconditional and conditional decompositions across different models for the nuisance functions.\footnote{See Appendix Tables A3 and A4 for numerical details. To aid interpretation, Appendix Table A2 additionally reports estimated group-specific means of baseline potential outcome, treatment proportions, ATEs, and covariances between the treatment and the treatment effect, as well as group differences in these quantities.} Descriptively, we find that individuals from lower-income origins on average achieve income 21 percentiles lower in their 30s and 40s than individuals from higher-income origins. This confirms the existence of intergenerational income persistence and represents the total disparity that we decompose. Below, we present our results in the order of the three-step sequential interventions and start with the unconditional decomposition.

Our estimates for the selection component in the unconditional decomposition are consistently negative. These estimates are statistically significant for all three ML models but not for the parametric models of the nuisance functions. Consequently, randomizing college graduation within each income-origin group to remove selection (the first step of the three-step sequential intervention) would increase the total disparity in adult income by around 7%. Thus, this intervention would increase intergenerational income persistence and decrease income mobility. Differential selection into college graduation reduces intergenerational income persistence because selection is positive in the disadvantaged group and negative in the advantaged group (Appendix Table A2). This lends support to stipulations from earlier research that obtaining a college degree is more of a rational decision in pursuit of economic returns among disadvantaged individuals, and more of an adherence to the social norm of attending college among advantaged individuals mare_social_1980, hout_social_2012. The selection component is the central conceptual contribution of our approach and also the most novel finding of our empirical application.

The prevalence component in the unconditional decomposition is positive, substantively large, and statistically significant across all models for the nuisance functions, accounting for about 15% of the total disparity in adult income. Consequently, equalizing college graduation rates across income-origin groups as in the second step of the three-step sequential intervention would reduce the total disparity in adult income by about 15%. In other words, such an intervention would decease intergenerational income persistence and increase income mobility. Underlying the prevalence component is the striking gap in college graduation rates by parental income, as 34% of respondents in the higher-income group, but only 9% of the lower-income group, obtained a college degree by age 29 (see Appendix Table A2).

The effect component in the unconditional decomposition is statistically insignificant due to minimal between-group effect heterogeneity. In Appendix Table A2, we show that the group-specific ATEs of college graduation on adult income range from 11 to 15 percentiles across models, all of which are statistically significant. However, the group difference in ATEs is always statistically insignificant.

The baseline component constitutes nearly 90% of the total disparity in adult income across all models for the nuisance functions. This demonstrates that most of intergenerational income persistence is due to processes that do not involve college graduation or its effects.

In sum, our unconditional decomposition reveals that college graduation in the United States plays two contradictory roles in intergenerational income persistence. On the one hand, the higher college graduation rate among individuals from higher-income origins increases intergenerational income persistence. On the other hand, some part of the income persistence is offset by the positive selection into college among individuals from lower-income origins.

The conditional decomposition quantifies the contributions of college graduation within levels of the AFQT achievement score. Hence, it informs settings in which policy makers cannot, or do not want to, change the factual relationship between prior achievement and college completion, perhaps due to normative constraints or meritocratic preferences.

figure[figure omitted — 870 chars of source]

In the conditional decomposition, all estimates for the conditional selection component are negative and statistically insignificant. The conditional prevalence component is positive and statistically significant, although it is much smaller than its unconditional counterpart and only accounts for 3% to 6% of the total disparity. Hence, equalizing chances of graduating college within levels of prior achievement (in the manner of the second step of the three-step sequential intervention in Section (ref)) would still somewhat decrease intergenerational income persistence.

Estimates for the conditional effect component are negative and statistically insignificant. The ${\boldsymbol Q}$-distribution component is positive and statistically significant. It reflects the strong associations between the AFQT score ($Q$) on one hand and parental income ($G)$, college graduation ($D$), and the college effect on income ($\tau$) on the other (see Footnote (ref)). The positive ${\boldsymbol Q}$-distribution component also shows that, net of prior achievement, college graduation plays a much more limited role in producing intergenerational income persistence.\footnote{In Appendix G, as a robustness check, Table A5 presents estimates of the conditional decomposition based on an estimation procedure combining parametric and ML models for specific nuisance functions. As explained in the last paragraph of Section (ref) and Appendix G, this theoretically may make the asymptotic inference more exact for the conditional prevalence, conditional effect, and ${\boldsymbol Q}$-distribution components. In practice, however, the empirical findings are qualitatively unchanged.} Finally, by definition, the baseline component is the same in the unconditional and the conditional decompositions.

In Appendix H, we conduct a sensitivity analysis for the conditional ignorability assumption ((ref)) against unobserved confounding. Leveraging results in opacic_disparity_2023, we derive bias formulas for the unconditional and conditional prevalence components. We find that in order to reduce our point estimates to zero, there must be an unobserved confounder that is about three times as strong as educational expectation, which is shown to be an especially strong observed confounder. This suggests that our estimates are quite robust to unobserved confounding.

Finally, we also estimate the change-in-disparity components in the URED and CRED of jackson_decomposition_2018 and jackson_meaningful_2021. As explained above, these decompositions estimate the impacts of (marginally or conditionally) randomly equalizing the treatment and do not isolate differential prevalence of college graduation as a distinct mechanism. Consequently, both the URED and CRED underestimate the extent to which differential college graduation rates, independent of selection, contribute to intergenerational income persistence (see Appendix Tables A3 and A4).

Discussion

The goal of causal disparity analysis is to enumerate, quantify, and disambiguate the mechanisms by which a treatment variable contributes to an observed outcome disparity between groups. We introduced a new nonparametric decomposition approach that is more appropriate for the causal explanation of descriptive disparities and differentiates more mechanisms than prior approaches. In particular, we identify differential selection into treatment as a previously overlooked mechanism and novel policy lever. We developed ML-based estimators that are semiparametrically efficient, asymptotically normal, and multiply robust. Empirically, we demonstrate that our approach provides new insights by documenting that differential prevalence of, and selection into, college graduation play important but counteracting roles in the production of intergenerational income persistence.

In future work, we plan to extend our approach in multiple directions. First, we will develop analogous decompositions for non-binary treatment and group variables, multiple treatments, non-continuous outcomes, and time-to-event outcomes, under the conceptual umbrella of sequential interventions. Second, as an alternative to the conditional ignorability assumption, it will be valuable to develop an instrumental variable-based identification strategy for the decompositions, possibly via the marginal treatment effect framework heckman_structural_2005. Third, in social science applications, period or cohort can play the role of the group variable, so longitudinal changes can also be decomposed. Fourth, in terms of estimation, one could employ targeted learning van_der_laan_targeted_2011 for improved finite-sample performance.

Acknowledgments

We thank Paul Bauer, Xavier d'Haultfoeuille, Laurent Davezies, Eric Grodsky, Jim Heckman, Aleksei Opacic, Chan Park, Stephen Raudenbush, Ben Rosche, Michael Sobel, Jiwei Zhao, and Xiang Zhou for helpful suggestions. We are also grateful for the insightful comments of three AOAS reviewers. Earlier versions of this paper have been presented at GESIS, PAA, ACIC, RC28, JSM, ICLC, CREST, and the Universities of Wisconsin, Heidelberg, and Chicago.

Funding

The authors gratefully acknowledge core grants to the Center for Demography and Ecology (P2CHD047873) and to the Center for Demography of Health and Aging (P30AG017266) at UW-Madison, a Romnes Fellowship to Felix Elwert at UW-Madison, and a Wisconsin Partnership Program grant from the UW-Madison School of Medicine and Public Health to Christie Bartels and Felix Elwert.

Appendices