EconBase
← Back to paper

Advancing Distribution Decomposition Methods Beyond Common Supports: Applications to Racial Wealth Disparities

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

46,825 characters · 13 sections · 29 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Advancing Distribution Decomposition Methods Beyond Common Supports: Applications to Racial Wealth Disparities

abstractI generalize state-of-the-art approaches that decompose differences in the distribution of a variable of interest between two groups into a portion explained by covariates and a residual portion. The method that I propose relaxes the overlapping supports assumption, allowing the groups being compared to not necessarily share exactly the same covariate support. I illustrate my method revisiting the black-white wealth gap in the U.S. as a function of labor income and other variables. Traditionally used decomposition methods would trim (or assign zero weight to) observations that lie outside the common covariate support region. On the other hand, by allowing all observations to contribute to the existing wealth gap, I find that otherwise trimmed observations contribute from $3\%$ to $19\%$ to the overall wealth gap, at different portions of the wealth distribution.

\onehalfspacing

Introduction

Decomposition exercises have been developed and improved in the Economics literature for decades, dating to the seminal papers of oaxaca1973 and blinder1973. These tools decompose the difference between two groups with respect to a variable of interest $Y$ into a portion due to differences in observable group characteristics $X$ and another residual portion. If we consider $Y$ to be wealth, $X$ to be average lifetime labor income and the groups to be defined by blacks and whites, we would be able to compute the contribution of labor income in explaining the racial wealth gap with decomposition methods.

A central part of decompositions involves estimating a counterfactual quantity which is what values of $Y$ would a group have if it had the $X$ distribution of the other group; e.g. what would be the white's wealth if they had the black's distribution of average lifetime labor income. In order to build this counterfactual, decomposition methods find blacks who are similar to each white with respect to all observable characteristics other than race and use these blacks' wealth and a model to predict it. An issue with this approach is that it is not always possible to find blacks who are similar to whites with regards to $X$. For instance, the range of the average lifetime labor income distribution was not the same for these races in the 1980's and beginning of 1990's. There is a nontrivial set of blacks exhibiting average income of nearly zero, with no white households earning that low; and with a considerable amount of whites at the top of the lifetime earnings distribution, unmatched by blacks. These observations are not informative for building counterfactual wealth for the other race. Current decomposition methods either trim these observations, or assign virtually zero weight to them in their estimation process, even though they may contribute significantly to the overall observed gap in $Y$.

In fact, the process of trimming observations that do not share similar $X$ values in the other group is required to satisfy a theoretical assumption imposed by decomposition methods referred to as the “common support” or “overlapping support” assumption. In this paper, I build on the existing literature to propose a generic decomposition method which relaxes the overlapping support assumption, allowing all observations to account for the actual observed gap. Furthermore, the decomposition framework that I present in this paper allows for decompositions in functionals of $Y$ -- e.g. permitting the study of how $X$ influences the gap in each quantile or portion of the distribution of $Y$. Although the set up is generic, my estimation strategy is focused on decomposing the gap in the distribution of $Y$, and I illustrate my method analysing the black-white wealth gap in the United States as a function of average lifetime labor income.

In order to estimate the full impact of lifetime labor income on differences in wealth accumulation between blacks and whites, I make use of publicly available data from the Panel Study of Income Dynamics (PSID). Besides yearly household surveys dating from 1968, I also use the wealth supplement surveys from 1984, 1989 and 1994 to get information about the net wealth of American households, focusing on households where heads are either black or white with nation-wide representation. I follow the data pre-processing suggested by barskyboundcharleslupton2002 and I find that differences in the wealth distributions between blacks and whites are primarily due to differences in lifetime labor income, corroborating previous findings. However, by including observations outside of the common support of $X$, I find that: (i) the wealth distribution gap becomes bigger, especially for relatively higher values of wealth; and that (ii) observations outside of the common support contribute from $3\%$ to $19\%$ to the wealth gap distribution, non-trivial amounts that were previously ignored.

Literature: Decomposition methods in economics are still evolving, although little has been done with regards to relaxing the common support assumption (fortinlemieuxfirpo2011). barskyboundcharleslupton2002 are pioneers in thinking about this assumption. However, as shown in Section (ref), functional decompositions similar to theirs (e.g. dinardofortinlemieux1996) assign virtually zero weight to observations outside the common support region, which is almost equivalent to trimming these observations. Furthermore, these methods still impose that the support for $X$ of one of the groups has to be a subset of/overlap the support of $X$ for the other group, which restricts the decomposition exercise, allowing it to be done using only white counterfactuals, but not black counterfactuals\footnote{The composition methods are set so the counterfactual quantity of group $W$, while having $B$ observable characteristics, is only identified when the covariates support of group $W$ is a subset of the covariate support for group $B$. For the black-white case, the support of black households' labor income does not overlap the whites' support, limiting the construction of the black counterfactual wealth distribution while having whites labor income. The method I propose allows for both counterfactual quantities to be estimated, not requiring the deletion of observations. Section (ref) defines the counterfactual quantity more precisely.}. In comparison, besides fully accounting for observations outside the common support, my estimator allows the decomposition to be done with either group counterfactual.

Another approach was developed in a way to deal with lack of overlapping support nopo2008,garcianoposalardi2009,FogelModenesi2021wagegap, however this methodology applies only to decompositions of differences in the mean of $Y$. I build on nopo2008's set up to extend their approach to decompose differences in functionals of $Y$. I also propose a feasible estimation strategy to decompose differences between intergroup distributions of $Y$ leveraging established results for estimating conditional distributions developed by chernozhukovfernandezvalmelly2013 and implemented by chenchernozhukovfernandezvalmelly2016. Although chernozhukovfernandezvalmelly2013 use their counterfactual distributional approach to decompose gaps as well, they also impose the common support assumption. This assumption is equally imposed by a relatively recent functional decomposition method using recentered influence functions firpofortinlemieux2018.

Finally, in terms of application, I use the same dataset and data pre-processing decisions as barskyboundcharleslupton2002, with the difference that, in addition to labor income, I also control for other observable covariates of the head of the household.

Roadmap: This paper is organized as follows. Section (ref) is the main section, setting up the framework that allows for decomposition in functionals of $Y$ while relaxing the common support assumption. This Section also discusses how other methods handle observations outside the common support and it proposes an estimator for decompositions in the distribution of $Y$. Next, Section (ref) presents the PSID dataset used in this study, as well as details our sample cuts and descriptive statistics of the final sample. Section (ref) presents results of my decomposition applied to the black-wealth distributional wealth gap, contrasting results using my estimator with the status quo. Finally, Section (ref) presents final thoughts and potential next steps in this research area.

Methodology

A primitive form to study intergroup differences in a dependent variable of interest $Y$ can be done by studying differences in the distributions of $Y$ between groups. By computing counterfactual distributions and decomposing the intergroup gap in distributions, it is also possible to decompose a series of other gaps, namely gap in mean $Y$, or the gap in a certain quantile of $Y$. Therefore, in order to make my approach as flexible as possible, I choose to develop an identification and estimation strategy for decomposing differences in the distribution of $Y$.

A distributional framework for decompositions

A decomposition of difference in distributions allows the researcher to answer questions like “what is the contribution of differences in $X$ (e.g. labor income) in the proportion of agents from each group at a certain level $y$ of $Y$ (e.g. at average wealth in the U.S., or for households at the top/bottom of the wealth distribution)?” In order to perform this decomposition I first define my parameter of interest, i.e. the distributional gap on $Y$ between groups $W$ and $B$:

flalign\Delta(y) :=& H_W(y) - H_B(y) \nonumber \\ =& \int_{S_W} H_W(y|x) dF_W(x) - \int_{S_B} H_B(y|x) dF_B(x)

where $H_g(y)$ and $H_g(y|x)$ are the cumulative distribution function (c.d.f.) and the conditional c.d.f. of $Y$ for group $g \in \{B,W\}$; $F_g(x)$ is the c.d.f. of the group observable characteristics $X$ for group $g$; and $S_g$ is the support of $X$ for group $g$.

Decomposition methods typically decompose the gap into a portion due to differences in the composition of $X$, $F_g$, fixing $H_g(y|x)$, and another portion due to structural differences in the form of $H_g(y|x)$, while fixing the distribution of $X$ equal for both groups. In order compute these portions, I add and subtract the counterfactual distribution of $Y$, denoted by $H_0(y)$, to the gap function $\Delta(y)$ in equation (ref). This counterfactual quantity refers to the $Y$ distribution for members of group $W$, if they had group $B$'s distribution of $X$, more precisely defined as\footnote{Notice that the decomposition can also be performed by adding and subtracting the $B$ counterfactual distribution of $Y$, $\tilde H_0(y) := \int_{S_W} H_B(y|x) dF_W(x)$. This is the c.d.f. of $Y$ for the $B$ group, if group $B$ had group $W$ distribution of $X$.}:

equation[equation omitted — 82 chars of source]

which allows the decomposition to be rewritten as:

flalign\Delta(y) :=& H_W(y) - H_B(y) \pm H_0(y) \nonumber \\ =& \left[ H_W(y) - H_0(y) \right] + \left[ H_0(y) - H_B(y)\right] \nonumber \\ =& \underset{\Delta_X (y) := Composition}{ \underbrace{\left[ \int_{S_W} H_W(y|x) dF_W(x) - \int_{S_B} H_W(y|x) dF_B(x) \right]} } + \underset{\Delta_0 (y) := Structure}{ \underbrace{\left[ \int_{S_B} \left[H_A(y|x) - H_B(y|x)\right]dF_B(x) \right]} }

In the wealth gap example, the composition portion, $\Delta_X (y)$, corresponds to the change in the proportion of whites with wealth equal to or below $y$, if whites had blacks' distribution of labor income. On the other hand, the structural portion, $\Delta_0 (y)$, can be interpreted as the difference in the proportions of blacks with wealth equal to or below $y$ if the wealth accumulation process for blacks shifted to how whites accumulate wealth conditional on income.

Overlapping support assumption

Definition and purpose

The overlapping support assumption, sometimes referred to as the common support assumption, guarantees that only comparable individuals -- in terms of $X$ -- from group $B$ are used to build the $W$ counterfactual distribution $H_0(y)$. More formally, considering the decomposition set up above, this assumption can be stated as: \\

Definition: Overlapping Support (OS) Assumption \\ Denote by $S_g$ the support of the distribution of $X$ for group $g \in \{B,W\}$, the supports of $X$ overlap if $S_W \supset S_B$.\footnote{Alternatively, if the decomposition is performed using group $B$ counterfactual distribution of $Y$, $\tilde H_0(y) := \int_{S_W} H_B(y|x) dF_W(x)$, then the overlapping support assumption becomes $S_W \subset S_B$.} \\

This allows the counterfactual $Y$ distribution $H_0(y)$, in equation (ref), to be well defined as a c.d.f.. With $S_W$ being “wider” than than $S_B$, the integral in the counterfactual goes over regions of $x$ where the conditional distribution $H_W(y|x)$ also exists with positive measure, not allowing the integral over $S_B$ to go beyond 1.

How does the literature handle it?

This is a widely assumed assumption among dozens of decomposition methods, with the only exception of nopo2008 which only performs decomposition in mean $Y$. Little attention has been paid to this assumption, although the decomposition literature in economics is still active, especially when it comes to decompositions in functionals of $Y$ fortinlemieuxfirpo2011,firpofortinlemieux2018.

Most of the existing decomposition methods handle failure of this assumption by simply trimming observations that do not lie in a region of common support of $X$, defined by $S_B \cap S_W$ (e.g. oaxaca1973, blinder1973, MachadoMata2005, chernozhukovfernandezvalmelly2013, firpofortinlemieux2018). In certain scenarios, trimming observations outside of the common support is reasonable, as it removes extreme outliers and it disregards a nearly negligible amount of observations, which would not change results significantly. However, in situations like the black-white wealth gap, where approximately $10\%$ of the observations lie outside of this region, simply deleting them can change results significantly. In other situations where unbiased counterfactuals are needed, researchers might need to add a large number of controls in order to mitigate potential violations of the assumption of selection on observables / ignorability. This increases the dimensionality of $X$ and the complexity of the common support region, making it more likely that a bigger fraction of observations would assume $X$ values outside the common support (see FogelModenesi2021wagegap for a more detailed discussion).

An alternative to deleting observations outside of the common support can be done using the estimation strategy developed by dinardofortinlemieux1996, as pointed out by barskyboundcharleslupton2002. This consists of a re-weighting approach, which rewrites the counterfactual distribution in (ref) as:

equation[equation omitted — 141 chars of source]

where $\psi(x)$ is the reweighting factor\footnote{The weight $\psi(x)$ can be interpreted as the Radon-Nikodym derivative of $dF_B$ with respect to $dF_W$, changing the measure of integration.} used to shift the distribution of $W$ to a counterfactual distribution, allowing the integral to be taken with respect to the measure of $W$, over the support of $W$, $S_W$. The weight can defined and rewritten using Bayes rule as follows:

flalign\psi(x) = \frac{P(x|B)}{P(x|W)} = \frac{P(B|x)/P(B)}{P(W|x)/P(W)}

The literature uses the right hand side of equation (ref) to consistently estimate $\psi(x)$. Consider now an observation of from group $W$, assuming value $x^*$ that lies outside the region of common support, i.e. $x^* \notin S_B \cap S_W$ -- which is equivalent to $x^* \notin S_B$. This implies that $P(B|x^*) = 0$, making $\psi(x^*) = 0$, which is equivalent to trimming/deleting these observations outside the common support, as in the previous methods mentioned. In the black-white wealth gap, this corresponds to assigning zero weight to all whites at the top of the distribution of lifetime labor income.

Relaxing the overlapping support assumption

I propose a generalization in gap decompositions in functionals of $Y$, which relaxes the overlapping support assumption, allowing all observations to account for the actual observed gap in the decomposition exercise. The key insight to allow mismatching supports in the gap decomposition consists in rewriting the gap in equation (ref) having in mind 3 different portions of the support of $X$: (i) the common support region $S_W \cap S_B$, where observations in groups $B$ and $W$ can are comparable and can be used to construct counterfactual quantities of each other; (ii) the region outside the common support where only $W$ observations exist, $S_W \cap \bar S_B$; and (iii) the remaining region where only $B$ observations can be found $S_B \cap \bar S_W$, where $\bar S_g$ is the complement of $S_g$, $g \in \{ B,W\}$. I represent $H_g(y)$, $g \in \{ B,W\}$, as:

equation[equation omitted — 171 chars of source]

This representation allows me to rewrite the right hand side of equation $\ref{eq:ch3-dist_decomp}$ in terms that are either in the common support, or out of it. The step by step identification of these terms is in the appendix (ref). The final result is the following generic decomposition:

flalign\Delta(y) =: \Delta_X(y) + \Delta_0(y) + \Delta_W(y) + \Delta_B(y),

where

flalign*&\Delta_X(y) = \int_{S_W \cap S_B} H_W(y | x) \left[\frac{dF_W(x)}{\mu_W(S_B)} - \frac{dF_B(x)}{\mu_B(S_W)} \right] && \\ &\Delta_0(y) = \int_{S_W \cap S_B} \left[H_W(y | x) - H_B(y | x) \right] \frac{dF_B(x)}{\mu_B(S_W)} && \\ &\Delta_W(y) = \left[ \int_{\bar S_B} H_W(y | x) \frac{dF_W(x)}{\mu_W(\bar S_B)} - \int_{S_B} H_W(y | x) \frac{dF_W(x)}{\mu_W(S_B)} \right] \mu_W(\bar S_B) && \\ &\Delta_B(y) = \left[ \int_{S_W} H_B(y | x) \frac{dF_B(x)}{\mu_B(S_W)} - \int_{\bar S_W} H_B(y | x) \frac{dF_B(x)}{\mu_B(\bar S_W)}\right] \mu_B(\bar S_W) &&

where $\mu_g(\cdot)$ is the probability of finding observations of group $g$ as a function of different parts of the support of $X$. It worth mentioning that if the supports of $X$ for both groups fully coincide, i.e. $S_B = S_W$, then my proposed set up collapses to the conventional set up in (ref), used by decompositions in functionals of $Y$\footnote{Equation (ref) collapses to (ref), if $S_B = S_W$, because in this situation, $\mu_B(\bar S_W) = \mu_W(\bar S_B) =0$ and $\mu_W(S_B)=\mu_B(S_W) = 1$.}.

The interpretation of $\Delta_X(y)$ and $\Delta_0(y)$ in equation (ref) is exactly the same as in equation (ref), only over the region of common support. The new terms, $\Delta_g(y)$, $g \in \{ B,W\}$ allow for observations outside of the common support to account for the existing gap. Precisely, $\Delta_g(y)$ corresponds to the change in the proportion of observations with $Y$ below or equal to the value $y$ if observations from group $g$ inside the common support suddenly get the $X$ distribution of observations from the same group outside the common support. In other words, this term captures differences between observations inside and outside of the common support region within the same group.

Estimation strategy and plan for inference

In this section I propose an econometric approach to consistently estimate each of the terms in the generic estimator and I lay out a plan for inference, to be followed in the next iteration of this project.

In order to obtain the final estimator for the decomposition in this paper, I need to estimate two objects. Notice that all of the terms in the proposed estimator in equation (ref) are of the type

equation[equation omitted — 147 chars of source]

where $\tilde S$ is an arbitrary region in the support of $X$ and $\tilde F(x)$ is a well defined c.d.f. over that region. This makes this quantity equivalent to an expectation of the conditional distribution function $H_g(y | x)$ with respect to the measure defined by $d \tilde F$, hence I need to estimate these two objects.

First, the econometric literature has developed greatly in terms of estimation and inference of conditional distribution functions such as $H_g(y | x)$. By employing distribution regressions it possible to estimate $H_g(y | x)$ consistently for each group separately chernozhukovfernandezvalmelly2013,chenchernozhukovfernandezvalmelly2016\footnote{For further details in the estimation procedure for $H_g(y|x)$, see appendix (ref).}. Therefore the first step in my estimation procedure consist in fitting functional forms for $H_g(y | x)$, for each group.

For the second object, $d \tilde F(x)$, I can simply use the empirical c.d.f. to consistently estimate it. Take for example $\frac{dF_W(x)}{\mu_W(S_B)}$, which is equivalent to the measure of $x$ for a specific sub-population -- i.e. observation group $W$ that always lie in the common support. Assume that there are $N_W^{S_B}$ observations from this subpopulation, then each $x_i$ in this subpopulation is estimated to happen with empirical probability of $\frac{1}{N_W^{S_B}}$. Analogously, an estimator for $d \tilde F(x)$ is $\frac{1}{\tilde N}$, where $\tilde N$ corresponds to the number of observations of the subpopulation for which the measure $d \tilde F(x)$ is defined.

The two consistent estimators for $H_g(y | x)$ and for $d\tilde F(x)$ can be plugged into equation (ref), forming, by Slutsky theorem, a consistent estimator for $\theta_g^{\tilde S}(y)$:

equation[equation omitted — 139 chars of source]

After obtaining $\hat \theta_g^{\tilde S}(y)$ for each group $g$ and different region of the support of $X$, the only remaining quantities to obtain estimates for all of the parameters in the gap decomposition in (ref) are $\mu_W(\bar S_B)$ and $\mu_B(\bar S_W)$, which can be consistently estimated as the proportions of observations of each group outside the common support.

As for inference, the most complex object in the estimation regards $\hat H_g(y | x_i)$ for which computing standard errors is not trivial. However, chernozhukovfernandezvalmelly2013 showed that the distribution regression coefficients are asymptotically jointly normally distributed and that the whole distribution regression process converges to a Gaussian process. This result can be useful to compute standard errors for each of the parameters in my proposed decomposition in the next iteration of this project.

Data

In order to illustrate the proposed decomposition method, I reassess the black-white wealth gap using the same data source and data preparation as described in barskyboundcharleslupton2002. In particular, I make use of the publicly available yearly datasets from the Panel Study of Income Dynamics (PSID)\footnote{The public version of the yearly PSID datasets can be downloaded at \url{https://psidonline.isr.umich.edu/}.}. Since 1968, the PSID follows a nationally representative set of households in the United States, starting with nearly 5,000 households and asking questions mainly about family composition and income. In the years of 1984, 1989 and 1994, families were additionally surveyed about all assets and liability holdings that compose their net wealth, which is central information to this paper.

Although detailed wealth information was only available in the years of 1984, 1989 and 1994, I followed each household yearly since the beginning of the PSID, 1968, in order to obtain information regarding their labor income throughout their existence in the survey. This allowed me to calculate a proxy for lifetime labor earnings, by averaging each household's labor income (in prices of 1989), a quantity crucial to measure the impacts of labor income in the accumulated household wealth throughout decades. In addition to the wealth and labor income data, I also used the gender of the head of the household, education dummies and age as factors in the wealth decomposition -- contrasting with barskyboundcharleslupton2002, who only used labor income as an independent variable in their exercise.

In order to compute the total net wealth of a household, several variables are summed: net business worth; personal bank accounts balances (including certificates of deposit, treasury bills, savings bonds); real estate equity value (deducted by mortgage and loans taken); value of stocks, investment funds, mutual funds, retirement accounts, of vehicles; and any other current debts.

Following barskyboundcharleslupton2002, the sample is composed of households whose head is either black or white, from age 45 to 49, being head for at least 5 years prior each of the wealth survey years -- 1984, 1989 and 1994. The minimum of a 5 year window allows for more observations in order to estimate more precisely a proxy for lifetime labor income. The narrow age range restriction, on the other hand, has a few purposes. First, studying wealth differences as a function of labor income requires time for households to accumulate wealth, hence the relatively older age. Second, heads are significantly more likely to receive inheritances past the age of 50, which can be a confounding factor in this study, hence the upper threshold. Third, blacks were more likely to retire earlier, which can be another confounding factor that the age range aims to avoid. Fourth, the narrow age restriction is an attempt to control for other confounding factors related to age.

The remaining sample after all restrictions imposed consists in 382 households with black heads and 856 with white heads, as shown in table (ref). By construction, households from each race have similar age distribution, but they differ significantly with regards to the gender and education of their heads. Nearly half of the black households have female heads, while males are heads of $85.8\%$ of white households. Furthermore, $32\%$ of white heads have a college degree, in contrasting with only $7.6\%$ of black heads having one.

table[table omitted — 1,358 chars of source]

For total net wealth and average labor income, their values have been converted to 1989 constant dollars using the Consumer Price Index for all Urban consumers (CPI-U). Mean average labor earnings for black households is $\$21,800$, corresponding to $56\%$ of white's average earnings, as shown by table (ref). As pointed out by barskyboundcharleslupton2002, part of this discrepancy is due to black households having much lower marital rates, reducing their total labor income, relative to white households. The net wealth gap, on the other hand, is substantially bigger than the labor income gap between races. While whites have $\$208,600$ mean net wealth, black households have their mean wealth at only $17\%$ of white's wealth. Furthermore, the ratio of black-white labor income gap increases monotonically from $.01$ to $.76$ as we move from the 5th to the 95th percentiles of its distribution. Although the labor income ratio increases drastically by moving along the labor income distribution, the wealth ratio remains around $.20$ for the upper half of the wealth distribution for these groups.

The aforementioned descriptive statistics raise a few questions that we hope to address in the following section. First, how can different portions of the wealth gap be decomposed into a portion due to labor income and other demographics, and another residual portion? For example, would the labor income be more important to explain the wealth gap at the bottom of the wealth distribution, versus at the top? Second, how to decompose the entire observed wealth gap, without trimming observations that do not share common support with regards to observable characteristics? For instance, instead of dropping blacks at the bottom of the labor earnings distribution and whites from the top, where the labor earnings support is not matched by the other race, how can we incorporate these observations to the wealth gap decomposition, precisely measuring their contribution to different parts of the gap distribution?

table[table omitted — 2,135 chars of source]

Results

Common covariate support

As shown in table (ref) and graphically illustrated in figure (ref), the average lifetime labor income supports for black and whites do not fully overlap. There is an absence of black households at the top of the lifetime income distribution, as well as there are no whites at the very bottom of this distribution. Unmatched white households at the top of the lifetime labor income distribution correspond to $6.8\%$ of the total sample, while unmatched black households at the bottom represent $2.7\%$ of the total. Conventional decomposition methods would either trim or assign zero weight to these observations, resulting in not utilizing a total of $9.5\%$ of the sample.

figure[figure omitted — 355 chars of source]

Gap decomposition

The main ingredients of our parameter of interest $\Delta(y)$ are the wealth distribution functions $H_B(y)$, $H_W(y)$ and $H_0(y)$. We depict these functions for the black-white wealth gap decomposition in figure (ref). This figure conveys a similar message conveyed by table (ref), that the wealth distribution of whites 1st-order stochastically dominates the wealth distribution for blacks, i.e. $H_W(y) \leq H_B(y), \forall y$. In other words, a bigger fraction of blacks are less wealthy than whites at any net wealth level, with the difference between these distributions reaching its biggest values around $\$50$ thousand dollars of 1989, and not vanishing completely even at high levels of wealth. Another important function depicted in figure (ref) is the estimate for $H_0(y)$, corresponding to the counterfactual wealth distribution of whites households, if they had the average labor income distribution of black households. It is noticeable that the counterfactual white wealth distribution is relatively close to the blacks' wealth distribution up to their medians, however, counterfactual white households at the top of the wealth distribution are still wealthier than blacks, despite their income being similar.

figure[figure omitted — 535 chars of source]

Two versions of the gap in wealth distributions are plotted in figure (ref): one using all of the sample, denoted by $\Delta(y)$; and another deleting the observations outside of the common support, as done by traditional methods, denoted by $\Delta^{OS}(y)$. The solid black line is simply the difference between the c.d.f.s in figure (ref), conveying the same message that the gap is more accentuated on relatively lower portions of the wealth support, and that the gap persists even for higher wealth values. In comparison, by imposing the Overlapping Support (OS) assumption, the gap shrinks towards zero for wealth levels between $\$150k$ and $\$400k$ (see green dashed line). If one believes in a positive correlation between average labor income and wealth, the attenuation in the gap is expected, since black households with the lowest average labor income and white households with the highest average labor income were not accounted by $\Delta^{OS}(y)$, mechanically shrinking this gap. Furthermore, with this figure it is possible to know exactly where these deleted observations are located in the wealth distribution.

figure[figure omitted — 517 chars of source]

Now I analyze results regarding the decomposition of the gaps $\Delta(y)$ and $\Delta^{OS}(y)$. I add the decomposition of these two terms into composition portions ($\Delta_X(y)$ and $\Delta_X^{OS}(y)$) and structural portions ($\Delta_0(y)$ and $\Delta_0^{OS}(y)$) in figure (ref). For the decomposition of the gap relaxing the overlapping support assumption, I also compute and plot the contribution of the observations outside the common support, denoted by $\Delta_B(y)$ and $\Delta_W(y)$, in figure (ref). The terms regarding the new decomposition method were drawn with solid lines, and the terms representing the conventional decomposition methods imposing the OS assumption were drawn using dashed lines.

Among all the decomposition terms defined in equation (ref), the major contributor to the distributional wealth gap is by far the composition effect, either imposing the OS assumption or not, as denoted by the red lines in figure (ref). This effect peaks at around $\$50k$ and reduces as wealth values increase, but never vanishes. In practical terms, white households would be substantially less wealthy if they earned as black households do from their jobs throughout their lifetime. By relaxing the OS assumption, the composition effect of $X$ is magnified due to including more extreme observations with regards to their average labor income.

The structural effects are represented by blue lines in figure (ref) and their magnitudes do not change much whether imposing OS or not. These effects capture differences in how white households versus black households accumulate wealth for similar values of labor income, i.e. $H_W(y|x) - H_B(y|x)$, weighted by blacks average labor income distribution. The fact that the blue curves are similar indicates that deleting the extreme observations outside the common support does not alter significantly how these conditional distribution functions are fitted. In terms of the shape of $\Delta_0(y)$ and $\Delta_0^{OS}(y)$, they both indicate that if households accumulate wealth like whites, while earning like blacks, they would be poorer than black households until approximately $\$180k$ of wealth. After this point, i.e. for relatively wealthier households, accumulating wealth like white households makes them wealthier.

Finally I discuss the terms that indicate the direct impact of observations outside the common support into the distributional wealth gap, plotted with the solid green and yellow curves in figure (ref). Both of these $\Delta_B(y)$ and $\Delta_W(y)$ effects act in opposite directions as the observed gap, being positive, especially before $\$180k$ of wealth, approximating zero as wealth values increase. Due to a relatively larger portion of whites being outside of the common support area, with high labor income, the contribution of this group of observations is bigger in magnitude than the contribution of blacks outside of the common support. $\Delta_W(y)$ being positive indicates that for a given level $y$ of wealth, whites outside the common support accumulate wealth at a slower rate than whites within the common support. Alternatively, $\Delta_B(y)$ greater than zero indicates that blacks in the common support accumulate wealth at a slower rate than blacks outside the common support.

figure[figure omitted — 566 chars of source]

The participation in percentage terms of each of the decomposition factors $\Delta_X(y)$, $\Delta_0(y)$, $\Delta_B(y)$ and $\Delta_W(y)$ into the distributional gap $\Delta(y)$ for the new estimator is plotted in figure (ref). The contribution of each factor was computed as the fraction between the absolute value of the factor and the sum of the absolute values of all the factors, capturing the strength at which each factor shifts the gap. As expected, differences in the composition of the labor income $X$ accounts on average for approximately $75\%$ of the wealth gap (red solid line). Except for the composition effect, the contribution of the structural portion prevails among others for wealth values beyond $\$200k$ (blue solid line), ranging from $15\%$ to $26\%$ after this wealth value. After the composition effect, however, the contribution of the terms regarding the lack of common support, i.e. $\Delta_B(y) + \Delta_W(y)$, tend to be the second most important effect for wealth values until $\$175k$, where $75\%$ for the sample is located (black dashed line). This indicates the importance of including these observations in the gap decomposition analysis. Considering all wealth values, the contribution of these new terms ranges from $3\%$ to $19\%$ of the overall gap.

figure[figure omitted — 682 chars of source]

Conclusion

This paper expands the set of intergroup gap decomposition tools in Economics by proposing a new distributional decomposition method that relaxes the Overlapping Support (OS) assumption. Current decomposition methods impose the OS assumption, which results in deleting or assigning zero weight to observations from one group unmatched by observations in the other group in terms of observable characteristics. These observations, however, constitute part of the observed gap, hence deleting them changes the gap itself, as well as modifies each of the terms in the decomposition. On the other hand, the new decomposition estimator that I propose allows for observations outside the common covariate support to participate in the decomposition of the observed gap. The key insight to obtain this estimator regards rewriting the integral inside the gap, with respect to covariates , while splitting the covariates support into regions where groups share common support and where supports do not match. I show that the estimation procedure of each portion of the decomposition requires the estimation of (i) distributions of covariates $X$ over different regions of the support, and of (ii) the conditional c.d.f. of the variable of interest $Y$ given $X$. I estimate (i) using analogous empirical c.d.f. and (ii) using distribution regressions.

I illustrate the proposed method studying the black-white wealth gap in the United States, using the publicly available income and wealth data from the Panel Study of Income Dynamics (PSID) from 1968 to 1994. More specifically, I assess how differences in lifetime labor income between races contribute to the observed difference in their wealth accumulation, while also accounting for black and white households that have extreme values of labor income. I show that households with extreme values of income represent nearly $10\%$ of the sample, and current decomposition methods would disregard these observations, attenuating the existing gap and changing the portions it is decomposed to. Moreover, I show that, although intergroup differences in average labor income predominantly explains the wealth gap, the terms associated with the pure effect of observations outside of the common support in the decomposition contributes from $3\%$ to $19\%$ to the overall gap, reaching its highest values in parts of the wealth distribution where most of the sample is concentrated.