Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
118,081 characters · 17 sections · 70 citation commands
Decomposition of Differences in Distribution under Sample Selection and the Gender Wage Gap
Decomposition methods have been widely used to understand the sources of differences between the outcomes of two groups. Following the seminal works by Oaxaca1973 and Blinder1973, mean differences between two groups have been attributed to either differences in mean covariates (endowments effect/composition component) or differences in the slope parameters (coefficients effect/structural component). Subsequently, several methods have been proposed to account for differences between other features of the distribution, such as the unconditional quantiles. These decompositions typically rely on the ignorability assumption (selection on observables).
However, individuals present in the decomposition often constitute a self-selected sample. This requires accounting for it to estimate the structural parameters using sample selection methods. Moreover, the decomposition of a feature of the outcomes between two groups could depend on differences the unobservables that depend on the amount of self-selection and participation for each group. This is the case if one considers the distribution of actual outcomes rather than the distribution of potential outcomes that would be observed if everyone participated.
Additionally, sample selection raises the question of which population is the target of the analysis: participants or the entire population. In the first case, the comparison is between those individuals with non-zero outcomes. E.g., one could be interested in understanding the differences in labor earnings of two groups of employed workers. However, it can be argued that including non-participants provides a more comprehensive comparison. This would amount to include unemployed workers and those out of the labor force in the analysis.
This paper makes the following contributions. First, I propose several decompositions according to the target population and the distributional feature of interest. Specifically, I consider decompositions for participants and the entire population for actual outcomes. Second, I propose two ancillary decompositions of the propensity score and the average value of the unobservable variable that affects the outcome, which is interpreted as ability. Third, I estimate the functionals that are used in the decomposition using the Quantile Regression with Selection (QRS) estimator proposed by Arellano2017. I show the asymptotic properties of the estimated components of the decomposition, as well as how to carry uniformly valid inference. Fourth, I obtain new estimates of the evolution of the gender wage gap, finding that the unobserved ability of female workers has increased more than that of males. This, together with the fall in the gender participation gap have been two major factors in reducing the gender wage gap.
In the absence of self-selection, individual actual and potential outcomes coincide. In contrast, actual outcomes equal zero for non-participants when there is self-selection. Hence, there are two potentially relevant target populations to consider: just participants, who have a non-zero outcome, and the entire population, which also includes non-participants. The decomposition for both populations are related and, depending on the context, either could be more appropriate to analyze. I consider both of them, highlighting their differences and similarities in a general framework and under some simplifying assumptions. Furthermore, because differences in the unobservables can account for differences in the actual outcomes, the type of decompositions presented here differ from those that focus on the potential outcomes when every individual is assumed to participate.
Under ignorability, differences in outcomes can be split into differences in covariates and differences in the returns to the covariates. However, when there is sample selection on unobservables, there are two additional components that can explain the differences between the two groups. The first one is the selection component, which reflects how, caeteris paribus, the unobserved characteristics of a group may be more or less positively selected relative to those of the other group. The other one is the participation component, which reflects differences in outcomes that can be attributed to the differences in the participation rates between the two groups. This extends the decomposition into three components in a triangular model with a binary treatment considered in Pereda2022.
Most decompositions in a cross-sectional data setting have imposed no selection on unobservables.\footnote{There exist works that explicitly account for time-invariant unobservables in a panel data setting. See Fortin2011 and references therein for further details.} This problem has been highlighted, e.g., by Kunze2008 or Huber2015c. Some exceptions include Neuman2003,Neuman2004 and Mora2008, who used the Heckit correction to obtain their estimates corrected for selection, or Cukrowska2016, who used a multinomial correction model. More recently, Huber2020 used mediation analysis tools to decompose the mean gap with sample selection under the assumption that the effect of the unobservables is homogeneous (e.g., if the model is additively separable).
A common feature of these works is that they are only concerned with decompositions of mean differences. In contrast, the focus gradually expanded to analyze other features, including unconditional distributions. The latter are particularly relevant, as many statistics of interest can be expressed as a function of it. Regarding estimation, there exists a wide menu of available methods, including those based on quantile regression Machado2005,Melly2006,Chernozhukov2013, on distribution regression Chernozhukov2013, and on reweighting Dinardo1996,Firpo2018. A thorough review of the advantages and disadvantages of each method is found in Fortin2011.
A decomposition into four components has previously been considered by Chernozhukov2019c, who also used it in a sample selection framework with a distributional regression estimator. The main difference with respect to the one considered in this paper regards the structural functions employed to model the outcomes. In their paper, they use bivariate Gaussian distributions to model the sample selection locally, whereas the approach in this paper is to model the unobservables with a copula globally. This allows to interpret the unobservables as the conditional ranks in the distribution of potential outcomes. Moreover, it is straightforward to relate differences in the value of these unobservables between the two groups to differences in the copula, the propensity score and the distribution of covariates, which may be of independent interest for the researcher.
Other related works are Maasoumi2017 and Maasoumi2019. In the first one they estimate the differences between functions of the distributions of the two groups for the selected sample and they bound the differences for the entire population. In the second one they analyze the evolution of the gender wage gap in the US using the estimator proposed by Arellano2017, including a decomposition of the gap between the two distributions of potential wages for the entire population. There are some differences relative to these two papers. First, I consider differences between the actual distributions and relate them to the different primitives of the model. Second, I consider four different sources of variability, which can help explain which are the determinants of the gap. The latter can be a problem not just if sample selection is ignored, but also if one analyzes the distributions of potential outcomes ignoring participation.
For estimation purposes, I consider a nonseparable model with a univariate unobservable variable in the outcome equation, which naturally points at quantile regression methods. The QRS estimator can estimate the structural function that relates potential wages to covariates and unobservables and the copula that captures the amount of self-selection. This, together with an estimator of the propensity score and the sample distribution of covariates are all the ingredients needed to estimate the different components of the decomposition. Moreover, I discuss under which simplifying assumptions other alternative estimators can be used.
The decomposition methods have been applied to a wide variety of topics, including test scores gaps between genders Sohn2008, schools Krieg2006 or countries Mcewan2004, differences in students' enrollment Borooah2005, differences in health insurance coverage between different demographic groups Bustamante2009, or gender differences in smoking behavior Bauer2007. However, most decomposition studies revolved around wage gaps between genders Kunze2008, race Barsky2002, unionization status Card1996, or public-private employees Depalo2020.
I apply the methods presented in this paper to decompose the gender wage gap. I revisit the estimates by Maasoumi2019, which are based on the Current Population Survey for the period 1976-2013. I perform several decompositions, considering different population targets (actual earnings for participants and the full population, as well as potential outcomes for the full population) and several statistics (mean, unconditional distributions and generalized entropy indices). I find that considering just employed workers understates the wage gap relative to consider the entire population. However, the gender gap has substantially diminished over the period considered in both cases. The main contributing factors for this reduction are the increase in female participation, and the increase in the average value of both observed and unobserved characteristics of employed females.
The rest of the paper is organized as follows: Section (ref) presents the model and the main functions of interest. The decompositions of interest are presented in Section (ref), including particular cases of interest that are nested in the more general model. Section (ref) describes the estimation of the decompositions for the general model, describing its asymptotic properties and showing how to conduct inference. Section (ref) revisits the estimates of the evolution of the gender wage gap, and Section (ref) concludes.
Consider the following selection model:
where $Y$ denotes the continuous outcome of interest, $X$ a set of predetermined covariates, $Z\equiv\left(Z_{1},X'\right)$ is composed of the instrument $Z_{1}$ and the predetermined covariates, $S$ is a binary indicator for participation, $D$ is an observed variable that denotes the group to which the individuals of the population can belong, and $U$ and $V$ are two unobservable random variables. Equations (ref)-(ref) conform a sample selection model that allows for a broad class of differences between individuals of different groups. The set of groups is denoted by $\mathcal{D}$, and its dimension is assumed to be finite. For expositional purposes, the leading case is two groups, i.e., $D\in\mathcal{D}=\left\{0,1\right\}$.\footnote{This framework is reminiscent of those considered for the study of the marginal treatment effect (MTE; Bjorklund1987). Such framework can be rationalized by a generalized Roy model with imperfect information Pereda2022.}
This system can be used to model several economic phenomena, such as labor wages, denoted by $Y$. These could be different for people belonging to different demographic groups, like gender ($D$). Equation (ref) reflects the fact that only employed workers ($S=1$) have a wage. These are modeled by Equation (ref), and they depend on their observed and unobserved characteristics, respectively given by $X$ and $U$. The latter, together with $V$ are the unobservables of the model, which can be interpreted as the ability of the worker and the unobserved propensity to be unemployed, respectively.\footnote{One could also consider work intensity as an additional dependent variable in the model. E.g., one could consider weekly hours or worked weeks, as in Fernandez2022. Such extensions could also be considered by appropriately modeling the unobservable variables, possibly adding some extra ones. Such comprehensive models could be informative about features that are not captured by a simpler model, at the cost of making them computationally more complicated. Their study is left for future work.}
Following Heckman2005, the distribution of $V$ can be normalized to be uniformly distributed on the unit interval. This is convenient, as $\pi_{d}$ can now be interpreted as the propensity score.\footnote{Note that the propensity score is sensitive to the instrument used. Therefore, using different instruments would lead to different results. See, e.g., Mogstad2021 for further details on the interpretation of treatment effects in with multiple instruments, a framework related to the one considered in this paper.} Moreover, if the same normalization is applied to the distribution of $U$, Equation (ref) uses the Skorohod representation, allowing us to interpret $g_{d}\left(x,u\right)$ as the structural quantile function (SQF). These normalizations offer two advantages: the joint distribution of the unobservables conditional on $D=d$ and $X=x$, denoted by $C_{d,x}\left(u,v\right)\equiv\mathbb{P}\left(U\leq u,V\leq v|D=d,X=x\right)$, can be interpreted as a copula and, by taking the inverse of the SQF with respect to $U$, one obtains the conditional distribution of potential outcomes if all were participants.\footnote{This constitutes a conceptual difference relative to Chernozhukov2019c. They model the relation between the outcome variable and the participation decision using a bivariate model that is observationally equivalent. However, it is not possible to interpret the distribution of the unobservables as a copula with their representation. Hence, the interpretation of their unobservables is less clear, as they are locally defined.}
The copula also conveys some important information for the policy maker regarding some counterfactual scenarios. It captures the amount of self-selection on unobservables, linking the participation decision to the outcome. Negative amounts of correlation are associated with positive selection of individuals into participation. Consequently, the more negative the amount of correlation, the lower the potential outcomes of non-participants relative to participants.
To show that these normalizations are without loss of generality, consider an alternative data generating process determined by functions $\tilde{g}_{D}$ and $\tilde{\pi}_{D}$, as well as the distribution of the unobservables $\tilde{F}_{\tilde{U},\tilde{V}|D,X}$. Their observational equivalence is established by Lemma (ref):
The identification conditions are listed in Arellano2017.\footnote{Note that the identification is based on instrumental variables, rather than on a control function approach. A similar model was considered using control functions was considered by Fernandez2022. The two models are non-nested, and in particular Fernandez2022 works under the assumption that the unobservables of $Z$, including $X$. This contrasts with the assumptions in Arellano2017, which allows for non-trivial dependence between the unobservables and $X$. A comparison of the conditions required for identification with instrumental variables and control functions is provided, e.g., by Heckman2007b.} Apart from some standard assumptions (exclusion restriction, continuous outcomes, propensity score strictly within the unit interval and a well-defined continuous copula for the unobservables), they require either identification at infinity, or a continuous instrument and that the copula be real analytic with respect to its second argument.\footnote{The set of assumptions required for identification by Chernozhukov2019c is different, though they are related. They provide a comparison of both sets of assumptions in their paper.} One could relax the latter by imposing a parametric assumption, a possibility considered in Appendix (ref). Additionally, these assumptions can be relaxed if the copula is homogeneous with respect to the covariates and one uses variation in the covariates as a source of exogenous variation.
Decompositions typically vary according to the distributional feature that is decomposed, such as the difference of the means or of the unconditional quantiles. Sample selection introduces two additional margins of choice regarding which decomposition to make. First, because only a fraction of the population participates, one could consider a decomposition for participants or for the entire population, assigning a value of zero for non-participants. Second, the distribution of the unobservables for participants and non-participants differ, making the distributions of actual and potential outcomes differ, too. The focus here is on actual outcomes of both participants and the full population.
The decomposition of the outcomes for either participants or the entire population requires to account for the following primitive functions: the SQF, $g_{d}\left(x,u\right)$, the propensity score, $\pi_{d}\left(z\right)$, the marginal distribution of the covariates, $F_{Z}^{d}\left(z\right)$, and the conditional copula of the unobservables, $C_{d,x}\left(u,v\right)$, for $d=0,1$. This contrasts with decompositions of potential outcomes, which account for selection only to estimate the structural parameters, and then base the decomposition on the SQF and the propensity score.
The leading feature that is decomposed is the mean outcome. This requires assigning a value to the counterfactual mean outcome for the population of interest when a combination of the distribution of the covariates or the estimated structural functions are changed to match those of the other group. Following the previous discussion, the mean outcome for participants with the distribution of the observables of group $h$, the SQF of group $j$, the copula of group $k$ and the propensity score of group $l$ is given by
where $G_{d,x}\left(u,v\right)\equiv\mathbb{P}\left(U\leq u|D=d,X=x,V\leq v\right)=\frac{C_{d,x}\left(u,v\right)}{v}$ denotes the copula conditional on participation and $\mathcal{Z}$ denotes the support of $Z$. This is the channel through which the unobservables affect the mean outcome. Similarly, the counterfactual value for the entire population is given by
More generally, one can obtain equivalent expressions for the distributions for each group, that constitute the building blocks for other functionals of interest, notably unconditional quantiles. The unconditional cumulative distribution functions for participants and the entire population are respectively given by
Consequently, the unconditional quantile functions are given by\footnote{Note that if $Y$ is continuous and there is no bunching at any particular value, Equation (ref) equals the inverse of the cumulative distribution function, but the same is not true for Equation (ref), which has a mass point at $Y=0$.}
The decomposition in a sample selection model is more intricate than in the regular one. To see this, note that the covariates can have an impact on the final outcome through three different channels. First, they affect the distribution of potential outcomes through the SQF, the only channel present when there is no sample selection. Second, they affect the propensity to participate, making some individuals more likely to participate. Third, they affect the amount of self-selection through the distribution of the unobservables, which is indexed by $x$ in general. In other words, some characteristics may display some correlation with the unobservables, either reinforcing or mitigating the differences between the two groups.
Consequently, to better assess how the covariates affect the propensity score and the selection on unobservables channels, it is useful to report two ancillary decompositions, presented in Section (ref). Moreover, the decomposition is path dependent. I present one of the possible orders for the decomposition of mean differences for participants and describe each component in turn.\footnote{Specifically, there are a total of 24 decompositions, and the size of each component may vary in each of them. See Fortin2011 for further details on the limits of path dependent decompositions.}
Note that, by definition, $\mathbb{E}\left[Y|D=1,S=1\right]=\mathbb{E}\left[Y^{1111}|S=1\right]$ and $\mathbb{E}\left[Y|D=0,S=1\right]=\mathbb{E}\left[Y^{0000}|S=1\right]$. The difference between these two can be decomposed as:
Equation (ref) decomposes the mean difference between the two groups into four components. The first two are those present in decompositions under the ignorability assumption. The endowments component captures how differences in the covariates between the two groups lead to differences in mean outcomes. The other term present in this type of decompositions is the coefficients component. It reflects how differences in the SQF between the two groups are related to differences in mean outcomes, which may be partly driven by discrimination. For example, in the Oaxaca-Blinder decomposition, this term equals the difference in the OLS coefficients between the two groups, scaled by the average covariates of one of them.
The remaining two terms arise in a sample selection framework. The participation component has the easiest interpretation, as it relates differences in the probability of participating for both groups to differences in mean outcomes. To get some intuition, assume that more able individuals (i.e, those with high $U$) are also more prone to participate (i.e., low $V$). Then, as the propensity score increases from zero to one, the average ability of participants gets smaller towards its mean value, reducing the mean wage for participants.
The selection component links differences in the amount of self-selection into participation to differences in mean outcomes. The interpretation is slightly different from that of the participation component, as it affects the distribution of the unobservables without affecting the participation. Therefore, caeteris paribus, the higher the level of selection, the higher the average ability of participants and, consequently, the higher their mean outcome.
Note that because of the linearity of the expectation operator, each component can be written in terms of the difference of the primitive functions. For instance, the coefficients component can be written as
This convenient property does not generally hold, and in particular for the decomposition of the unconditional quantiles. Moreover, in some particular cases, the expression for some components can be simplified, as shown in Appendix (ref). The decomposition for the entire population is analogous to the one for participants:
The interpretation of the different components is similar, with one notable difference: because the outcome for non-participants equals zero by definition, the mean outcome for the entire population is the mean outcome for participants multiplied by the propensity score. Therefore, the endowments and participation components operate through two channels: first, by changing the proportion of participants; second, by affecting the average outcome of participants. In contrast, the coefficients and selection components operate exclusively through the second channel.
The last two considered decompositions are those of the unconditional quantiles of the outcome for both target populations.\footnote{Additional decompositions can be constructed analogously. E.g., if one is interested in differences in inequality, one could construct equivalent counterfactual values for the Gini index and apply the decomposition. As long as the statistic of interest can be expressed in terms of the same primitives of the model, and under some regularity conditions, they will be well-behaved. See, e.g., Chernozhukov2013 for further details.} They consist of the same four components as the mean decompositions, although there are some important differences. From a mathematical standpoint, the unconditional quantile function is not a linear operator. Therefore, their expressions are more convoluted even under some simplifying assumptions. From a policy perspective, they are more informative, as they allow to assess which segments of the population present the largest differences. This makes them particularly relevant in cases in which the differences are of opposite signs for different quantiles of the distribution.
Formally, the decomposition of the unconditional quantile distribution for participants at quantile $\tau$ equals
while the decomposition of the unconditional quantile distribution for the entire population at quantile $\tau$ is given by
The previous discussion highlighted the multiple channels through which the covariates can affect the final outcome in a sample selection framework. To better assess their role, researchers could perform two additional decompositions of the propensity score for the entire population and the average value of the unobservables for participants. These decompositions could be presented alongside the main decomposition, complementing the analysis.\footnote{It is worth stressing that, because the variable that is decomposed is not the main outcome, the components of the ancillary decompositions are not comparable to those of the main decompositions, although they are related to each other.}
The first one is the participation decomposition. This is a regular means decomposition without selection using the propensity score as the dependent variable. As such, differences in the propensity score would be split into the usual endowments and coefficients components.
The second one is the self-selection decomposition. This is more reminiscent to the decompositions of the mean outcomes. To see this, note that the average value of $U$ for participants with the distribution of the observables of group $h$, the copula of group $l$ and the propensity score of group $m$ is given by
Then, the difference of this statistic between the two groups can be decomposed as:
Similarly to the main decompositions, the endowments component captures differences in the copula and the propensity score due to differences in the covariates. On the other hand, the selection and participation components capture differences in the unobserved ability between the two groups due to differences in the copula and the propensity score, respectively.
For expositional convenience, I present the estimators of the decompositions for participants. The estimators of the decomposition for the entire population are similarly constructed using the analogy principle. Their exact expressions and their asymptotic properties are presented in Appendix (ref). Throughout the entire section, the following assumptions are maintained:
Assumption (ref) states the sampling process of the data. Assumption (ref) further restricts it to ensure that the number of individuals in each group converges to a fixed proportion with respect to the sample size. The propensity score is required to satisfy some mild regularity conditions given by Assumption (ref). It ensures that its asymptotic distribution is well-behaved, and it is satisfied by many estimation methods, such as maximum likelihood. Finally, Assumption (ref) ensures that the dependent variable has a finite conditional density for the entirety of its support.
Let $\upsilon_{m}\left(z,\tau,\pi,f\right)\equiv\left(g_{j}\left(x,\tau\right),c_{k,x}\left(\tau,\pi\right),\pi_{l}\left(z\right),\int_{\mathcal{Z}}fdF_{Z}^{h}\right)$ be the vector that contains the structural functions, where $m\equiv\left(h,j,k,l\right)\in\mathcal{M}\equiv\left\{\left(h,j,k,l\right)\in\mathcal{D}^{4}\right\}$ is used to keep notation compact.\footnote{Strictly speaking, the last component of this vector is not of the structural functions, which would be $F_{Z}^{h}$. However, the former is more convenient to derive the asymptotic theory, so I refer to $\upsilon_{\ell}$ as the vector of structural functions.} The counterfactual mean outcome is estimated as:
Each of the effects is computed as $\hat{\mathbb{E}}\left[Y^{m}|S=1\right]-\hat{\mathbb{E}}\left[Y^{m'}|S=1\right]$ for the appropriate choice of $m,m'$.\footnote{An alternative approach to the one presented in this paper would be to provide bounds of the quantities of interest. Maasoumi2017 do this in a similar setting, whereas Blundell2007c obtained estimates of changes in the distribution of wages by gender using bounds based a variety of assumptions. Unfortunately, as pointed out by the latter, the bounds on differentials tend to be much larger than the bounds on the quantities of interest. Therefore, while in principle it would be possible to bound the components presented in this paper, the width of these bounds could be too large to be informative in most empirical applications.} Similarly, the estimator of the unconditional distribution equals:
where $\hat{F}_{Y|S=1}^{m}\left(y\right)=\frac{1}{n_{h}}\sum_{i=1}^{n}\left[\varepsilon+\int_{\varepsilon}^{1-\varepsilon}\mathbf{1}\left(\hat{g}_{j}\left(X_{i},u\right)\leq y\right)d\hat{G}_{k,x}\left(u,\hat{\pi}_{l}\left(Z_{i}\right)\right)\right]\mathbf{1}\left(D_{i}=h\right)$.
Similarly, the effects of interest are computed as $\hat{Q}_{Y|S=1}^{m}\left(\tau\right)-\hat{Q}_{Y|S=1}^{m'}\left(\tau\right)$. To derive the asymptotic distribution of these estimators, let $\mathcal{U},\overline{\mathcal{P}}$ be closed subsets of $\left(0,1\right)$, and $\mathcal{F}$ denote the class of measurable functions that includes $\left\{F_{Y|Z}\left(y|z\right):y\in\mathcal{Y},z\in\mathcal{Z}\right\}$ as well as the indicators of all the rectangles in $\overline{\mathbb{R}}^{d_{z}}$, such that $\mathcal{F}$ is totally bounded under the metric $\sigma\left(f,\tilde{f}\right)=\left[\int\left(f-\tilde{f}\right)^{2}dF_{Z}\right]^{\sfrac{1}{2}}$, and denote the space of real-valued bounded functions defined on the index set by the supremum norm by $\ell^{\infty}$. The estimator $\hat{\upsilon}_{m}$ needs to satisfy the following condition:
To keep notation compact, denote the difference between two counterfactual values of a statistic indexed by $m$ and $m'$ by $\Delta^{m,m'}\left(\cdot\right)$. Moreover, denote the support of $Y$ by $\mathcal{Y}$. The asymptotic distribution of the components of the means and unconditional quantile decomposition are established in the following theorem:
Theorem (ref) establishes the asymptotic gaussianity of the counterfactuals of interest when the estimator of the structural functions satisfies Condition (ref). As a consequence, one needs to consider an estimator of the structural functions for which this condition holds. One such example is the estimator proposed by Arellano2017. This estimator is adapted to account for different groups, and it has to satisfy the following assumptions:
Assumption (ref) is made for convenience. Despite restricting the shape of the distribution of $Y$, it does not force conditional quantile curves to be parallel to each other for different values of the covariates. Moreover, it does not suffer from the curse of dimensionality as other more flexible alternatives, such as the partially linear model Lee2003. Assumption (ref) is also imposed for practical reasons, and it allows to use the most common parametric copulas, such as the Clayton, the Gaussian, or a Bernstein copula of a fixed order.\footnote{Note that Assumption (ref) does not provide a way to select the most appropriate copula given some data. This could be done, e.g., by choosing the copula that minimizes the criterion function given in Equation (ref) whenever the candidate copulas have the same number of parameters. Alternatively, it could be selected using cross validation as in Pereda2022.} Importantly, it allows the copula to be dependent on the covariates. Assumption (ref) is a regularity condition and Assumption (ref) ensures that the moments needed to derive the asymptotic distribution of the estimator have full rank. Lastly, Assumption (ref) ensures that the participation decision is not deterministic for any individual.
The following steps describe how to compute the estimator, and how to implement it to obtain an estimate of the decompositions:
Step 1 is standard and can be done, e.g., by logit or probit. Step 2 is a rotated quantile regression conditional on a particular value of the copula. Note that in practice one needs to set a grid of values of $\tau$, such as $\tau=\left\{0.01,...,0.99\right\}$. The third step is computationally expensive, as it involves the minimization over a non-convex space. The most common parametric copulas depend on few parameters, so that a grid search may be feasible. For Bernstein copulas, one can use the algorithm proposed in Pereda2022. The last two steps are immediate, and they yield the slope parameters, which are those estimated in step 2 for the value of the copula estimated in step 3, the SQF and the copula.
The following theorem establishes the uniform asymptotic distribution of the Arellano2017 estimator:
It remains to verify that it satisfies Condition (ref). This is done in the following corollary:
The expressions of the asymptotic variance of the different estimators are complex and depend on several density functions. Therefore, using resampling methods to obtain standard errors is preferable to obtaining closed-form expressions. In this paper I consider the weighted bootstrap Ma2005. The following assumption defines the weights used to obtain all the bootstrap estimates:
For the Arellano2017 estimator, the weighted bootstrap is implemented as follows:
This bootstrap estimator of the variance is based on the one presented in Chernozhukov2013. Even though it is possible to use the variance of the estimator across repetitions of the bootstrap to obtain the standard errors, it would require additional conditions for it to be valid Kato2011. This estimator only requires that the bootstrap converges in distribution to the asymptotic distribution of the sample estimator, which is established in the following theorem:
On top of providing uniform confidence bands for the functionals of interest and the intermediate functions, the weighted bootstrap can be used to carry out uniform inference using, e.g., a Kolmogorov-Smirnov test to any of the components of the decomposition of the unconditional quantiles. For example, one could test the null hypothesis that one of the components equals a specific value, $\Delta^{mm'}Q_{Y|S=1}\left(\tau\right)$. The test statistic would be given by
where $\hat{\Sigma}_{Q|S=1,mm'}\left(\tau\right)$ is an estimator of the asymptotic variance of $\Delta^{m,m'}\hat{Q}_{Y|S=1}\left(\tau\right)$, such as the one proposed in the weighted bootstrap algorithm. The critical value $c_{1-\alpha}$ would be $1-\alpha$ quantile of the distribution of the bootstrapped of the $KS_{n}$ statistic. Similar uniform confidence bands can be constructed for other functionals of interest.
The ancillary decompositions are also based on the vector $\upsilon_{\ell}\left(z,\tau,\pi,f\right)$. The counterfactual values of the mean propensity score and the mean value of the unobservables are given by
The asymptotic distribution of the components of the two ancillary distributions is established in the following theorem:
I study the evolution of the gender gap between earnings distributions using the Current Population Survey (CPS) dataset. I extend the analysis in Maasoumi2019 by decomposing several features of the distribution of actual earnings for employed workers and the entire population. In addition, I do the two ancillary decompositions regarding differences in participation and self-selection between men and women.
To preserve the comparability to Maasoumi2019, the sample covers the 1976-2013 period, and the regressions are done on a year-by-year basis, abstracting from any dynamics. I restrict the analysis to individuals between 18 and 64 years old, who work for wages and salary, do not live in group quarters and worked at least for 20 weeks and 25 hours per week in the previous year.\footnote{Note that this analysis excludes part-time workers, seasonal workers, self-employed, or the total number of hours worked (which affects the intensive margin). Despite this, the gaps considered in this paper are relevant for a large fraction of the working age population, and the elements analyzed could also be relevant for the analysis of a more comprehensive gender gap that accounts for the aforementioned factors. Such analysis is beyond the scope of this paper.} Like them, I estimate the propensity score using the probit estimator. I use the same regressors they used, i.e., a third degree polynomial of age, four levels of education, four regional dummies, marital status, an indicator for white race, and the interactions between age and the other listed covariates, plus another variable they did not use: the number of children.
The dependent variable is mean log hourly wages.\footnote{As in Maasoumi2019, it is computed as the logarithm of the total wage and salary income divided by the number of week and hours worked during the previous year and adjusted for inflation using the 1999 consumer price index adjustment factor.} The specification for the QRS estimator uses the same set of variables (except for the interactions between age and the remaining covariates), the same instrument (number of children below 5 years old) and two parametric copulas: the Frank and the Gaussian copulas.\footnote{A discussion of the validity of the exclusion restriction is provided in Maasoumi2019. They acknowledge that the instrument may be stronger for women than for men, as well as for earlier years than for the latter. In addition, it may be stronger for married workers than for those who are not.} However, there are some slight differences along several dimensions: I include the regressor number of children, the quantile grid, the propensity score, and the objective function used to estimate the copula.\footnote{For precision, the quantile grid I used for the estimation is $\left(0.01,0.02,...,0.99\right)$, while the one used by Maasoumi2019 was $\left(0.3,0.4,...,0.7\right)$; I use the propensity score as the instrument $\varphi\left(u,z\right)=\hat{\pi}\left(z\right)$ as suggested in Arellano2017, whereas Maasoumi2019 use $\varphi\left(u,z\right)=\sqrt{u\left(1-u\right)}\hat{\pi}\left(z\right)$, which puts less weight on values that are further away from the median; the objective function equals Equation (ref), while the objective function used by Maasoumi2019 is $\frac{1}{N}\sum_{i=1}^{N}\sum_{j=1}^{J}\left(\varphi\left(\tau_{j},z_{i}\right)\left(\mathbf{1}\left(y_{i}\leq x_{i}'\hat{\beta}\left(\tau_{j}\right)\right)-G\left(\tau_{j},\hat{\pi}\left(z_{i}\right);\hat{\theta}\right)\right)\right)$; the implemented quantile regression estimates from Stata in Maasoumi2019 were in some cases numerically slightly worse than the ones in Matlab, i.e., the value of the check function was smaller for the latter. See Appendix (ref) for further details.} Additionally, I also allow for more flexible specifications that separately estimate the main equation according to race (white vs non-white), level of education (college vs less than college) and marital status (married vs non-married), which are reported in Section (ref), and the estimates using the same specification as in Maasoumi2019, which are reported in Appendix (ref).\footnote{For completeness, I also report the results of the decompositions of potential outcomes and the estimates of the general entropy measures considered by Maasoumi2019 in Appendix (ref).}
To analyze the evolution of the gender earnings gap, I begin by analyzing the evolution of labor market participation for both genders. Table (ref) reports the average estimated propensity score by gender for the entire period. There has been a marked catch-up between female and male participation: in 1976, female participation was roughly one third, steadily increasing to over one half in the early 2000s, to fall slightly in the aftermath of the financial crisis. Meanwhile, male participation has been more stable: it has been equal to around two thirds until the financial crisis, falling to about 60% afterwards.
Consequently, the gender participation gap has more than halved during the period, from 33% to about 14% (Table (ref)). Its decomposition shows that almost the entirety of the gap is explained by the coefficients components. In words, there has been a structural increase in female participation into employment for women unrelated to gender differences in covariates. On the other hand, the endowments component has been either statistically not significant or slightly negative for most of the period. Indeed, only between 1987 and 1989 was this component positive and barely significant. Hence, the catch up in college education rates for female workers has not contributed to an increased participation in the labor force relative to men.
The second feature of interest is the evolution of differences in self selection, which depends mainly on the estimated copula. Because the values for different copulas are not directly comparable, I report the Kendall's $\tau$ correlation coefficients.\footnote{Unlike to the more common Spearman's $\rho$ correlation coefficient, Kendall's $\tau$ is invariant to the distribution of the marginals.} Table (ref) reports the baseline estimates for each year and gender, both with the Frank and the Gaussian copulas. These coefficients indicate that the amount of selection into employment has steadily increased for female workers: until the early 80s, there used to be negative selection that turned positive afterwards.\footnote{Recall that a negative (positive) coefficient implies positive (negative) selection into employment.} This result reinforces the dynamics on selection into employment found by Mulligan2008. On the other hand, the amount of selection for male workers has fluctuated more over time, being either above or below that of females depending on the year. The results are very similar with both copulas, which suggests that the choice of the parametric copula is of secondary importance. To assess the sensitivity of these results to the model used, I report in Table (ref) in Appendix (ref) a comparison of the correlation coefficients with those of the Heckman 2-stage estimator Heckman1979. The findings show that the copula estimated by both models are very similar throughout the entire period.
A more informative way to understand these estimates is to compare the average value of the unobservable $u$ across genders and periods. This is shown in Table (ref) and Figure (ref). For the entire period considered, the average value of $u$ for full-time employed females steadily increased from slightly above 40 to around 60. The average value for employed males has slightly increased over time. However, it displayed much more fluctuation over time, which may be partly related to the lack of strength of instrument for male workers. Still, there was a positive trend until the mid-nineties, steadily decreasing until the eve of the financial crisis, when it soared for a few years.\footnote{These trends can be seen more clearly if one smooths the copula estimates over time. In Figure (ref) in Appendix (ref) I compare the baseline estimates to a 5 year moving average of the copula.} In other words, along with the increase in participation, there has been an increase in the amount of self-selection into employment for women. Hence, potential earnings of non-employed women are lower than actual earnings of those employed, given the same observed characteristics.
The average selection difference between male and female workers has consequently become negative, decreasing by about 11.6 percentage points, although it also reflects the oscillation of the estimates for males. The participation components experienced a decrease of about 6 percentage points, roughly half of the difference. Hence, even if the amount of self-selection (copula) had been the same for both genders, the increase of female employment rates contributed to the increase in average unobserved ability, setting a gap relative to male workers. On the other hand, the selection component is less precisely estimated and it displays an unstable evolution. This is a direct consequence of the oscillating behavior of the male copula estimates. Regardless, the long term trend also points at an increase of the gap in favor of employed women. Lastly, because the baseline estimates are homogeneous with respect to the covariates, the endowments component is negligible.\footnote{Note that it is not exactly zero because there is variation in the copula because of differences in the propensity score, deriving themselves from differences in the distribution of $Z$ between genders.}
Next consider the distributions of actual earnings for participants and the entire population by gender. Specifically, I report their means and the value of their 10th, 25th, 50th, 75th and 90th percentiles in Tables (ref)-(ref). The actual earnings distribution for participants shows a small decrease in mean earnings for male workers and a slightly larger increase for mean female earnings.
This catch-up, however, masks an increase in the inequality within each gender, as interquantile ranges (IQR) increased both for male and female distributions: the 90-10 IQR increased from 132 to 172 percentage points for men, and from 112 to 158 for women; the 75-25 IQR similarly increased from 69 to 90 and 60 to 82 percentage points for male and female workers, respectively. Still, the evolution has been quite heterogeneous across the distributions of earnings. For male workers earnings followed a long term decrease for percentiles below the 75th, and a steady increase for those at the top of the distribution; for female workers there has been a gain for those above the 25th percentile, more pronounced at the top. Despite this catch-up, there is still a gap in favor of men at all quantiles.
By construction, earnings at any given quantile are smaller on the distribution for the full population than on that for participants, resulting in a smaller mean for both genders. Nonetheless, including non-participants increases the earnings gap. Following the increase in female labor participation, this distribution has steadily increased for females, reducing the gap relative to males by a bigger fraction than for the distribution of participants. On the other hand, mean earnings for males have oscillated across time following the changes in participation and average earnings for participants. Note that the fall in the male participation rate in the last years of the sample has been starker than that of females, prompting a decrease in mean earnings for both genders, along with a decrease of the gap.
Tables (ref)-(ref) report the decompositions of the mean earnings gap for the two populations considered.\footnote{The estimates of the decompositions of mean earnings with the Heckman 2-stage estimator are very similar to those found with the QRS estimator. These are shown in Figures (ref)-(ref) in Appendix (ref).} The mean gap for participants has more than halved during the period: from over 40% gender gap for workers, it was equal to 19%. Out of the four components, the largest one in every period has been the coefficients component. Its size displays some yearly variation driven by the fluctuation of the slope and copula parameters for male workers.\footnote{The erratic behavior of this and the selection components could be linked to the strength of the instrument, which is weaker for male workers. To see this, note that the autocorrelation of the total gap is 0.99. Out of the four components, it is also large for the endowments (0.99) and participation components (0.79), whereas for the coefficients and selection components are significantly smaller: 0.47 and 0.37, respectively. However, the autocorrelation of the sum of these two components is equal to 0.93. Therefore, it is plausible that the lack of strength of the instrument may be responsible for the large variations in the size of these two components even in consecutive years.} If one considers the moving average estimates of the copula parameters, it can be seen a decrease of this component until the mid-nineties, increasing afterwards.\footnote{See Figures (ref)-(ref) in Appendix (ref).} Analogously, the selection component also displays an erratic behavior, which is again a consequence of the estimates of the copula for males. The sign of this component is often negative, and when it is positive it is not significant at the 95% confidence level. Moreover, it displays a slightly downward trend, thus increasingly helping in the reduction of the mean gap (by 14 percentage points in the last year of analysis).
The dynamics of the remaining two components has been more stable: they were initially positive, and they eventually became negative, therefore reducing the mean gender gap. Moreover, the magnitude of these two components has been more modest than that of the other two: the endowments components changed from 1 to -4 percentage points, whereas the participation component experienced a more pronounced fall (from 5 to -4 percentage points).
One way to highlight the importance of accounting for selection is to compare the previous decomposition to the standard Oaxaca-Blinder decomposition. This is shown in Figure (ref). Both components of the latter display a different evolution over time. Specifically, the endowments component retains the downward trend, but its magnitude is much larger, whereas the coefficients component displays a steady downward trend. Hence, the latter appears to be a contributor to the decrease in the gender wage gap, in contrast with the analysis that accounts for sample selection.
The gap for the entire population has followed a similar trend, steadily decreasing to less than a half of the gap in 1976. Specifically, from a gap of over 100%, it fell to less than 50%. However, its magnitude has always been larger, owing to the gender participation gap. Indeed, the participation component constitutes the lion share of the gap, and its reduction has been responsible for the majority of the reduction of the gap: from an initial 81 percentage points, it fell to 33. The remaining three components display a similar behavior to the one found in the decomposition of actual earnings for participants, although their size is slightly scaled down.
The same decompositions are performed for the unconditional distributions. I present the estimates for a number of years (1976, 1984, 1992, 2000, 2007, 2013) in Figures (ref)-(ref). Additionally, I report the estimates for several quantiles in Tables (ref)-(ref) in Appendix (ref).
Figure (ref) shows the evolution of the gap for the entire quantile process. Several changes have taken place. First, the gap increases monotonically with the quantiles of the distribution in every year, with the exception of the extreme top quantiles. Second, there has been a generalized reduction at all quantiles and, the decrease in the gap in absolute value has also been larger for higher quantiles. Namely, the gap at the 10th percentile fell from 29% to 10%, whereas the gap at the 90th percentile decreased from 49% to 24%. As it was the case for the mean decomposition, the coefficients component has been the largest one for almost every quantile and every year. In contrast, the selection component has been relatively flat across quantiles, although it displayed great variation across years, switching sign several times.
Regarding the participation and endowments components, their behavior is quite similar: their magnitude is smaller than that of the other two components, they were initially positive for the majority of the distribution and had a mild upward slope, and they have progressively flattened out, becoming negative and therefore reducing the gender gap. Moreover, their magnitude is similar to that of the decomposition of the mean.
The distributional gap for the entire population has an unconventional shape, as it displays a thick spike for a large part of the distribution. Its width equals the difference in the participation rates between men and women, and its height equals earnings of male workers from the left tail of their gender distribution. Therefore, the width of this spike has progressively diminished over time with the reduction of the participation gap. However, it still remains the main factor of difference between the two distributions. Second, since the fraction of non-participants is positive for both men and women, the lower tail of the gap equals zero, as workers of both genders on that tail do not have any labor earnings.
Additionally, the participation component has a decreasing shape after the end of the spike, reflecting a shifting in the distribution between male and female worker. To see this, denote by $\tau_{f}$ the quantile at which women earnings becomes positive, i.e., $\tau_{f}=1-\mathbb{E}\left[\pi_{f}\left(Z\right)\right]$. Men above $\tau_{f}$ represent those above a certain level of earnings, which would be equivalent to comparing the distribution of earnings of employed women at a given quantile to the distribution of employed men of a higher quantile.
The other three components have a similar shape to the one found in the decomposition for participants, although their size is somewhat smaller. Consequently, the coefficients component is the second largest one, remaining an important determinant of the gap. Finally, the endowments component and selection components are flatter, relatively small, and they have become negative over time reducing the gap at all quantiles.
These results are related to the findings in Olivetti2008. In their cross-country comparison, they found that countries with the highest participation gaps tended to have lower wage gaps for participants. However, their estimated gaps accounting for self-selection showed that gaps increased much more in countries with larger participation gaps. This is captured by the participation component in the decomposition of the wage gap: for a given level of self-selection, narrowing the participation gap would make the participation component shrink to zero. However, this does not eliminate the impact of self-selection on unobservables, as differences in the amount of self-selection could be due to differences in the copula. Indeed, the estimates in this paper showed how, as participation gaps decreased over the considered period, the intensity of self-selection increased much more for women. Therefore, a more comprehensive cross-country analysis could account for this factor to explain how differences in self-selection intensity have determined different wage gaps in different countries.
One potentially strong assumption regards the fact that the baseline copulas used in the estimation are homogeneous across the covariates. This limits any selection differences across some of these characteristics to the propensity score channel. To address this issue, I repeat the estimation separately for three different categories: race (white vs non-white), education level (college graduates vs. less than college) and marital status (married vs unmarried).
Some of the estimates are slightly sensitive to the heterogeneous copulas. The mean value of unobserved ability is most similar when one obtains the estimates by race (Figure (ref), top panel). However, it has been in general larger for white males relative to the baseline estimates, although the estimates are more volatile. In contrast, the estimates for non-white males show a smaller value than the baseline for almost the entire period. The estimates for white females show a more stable evolution of the average value of unobserved ability, whereas for non-white female workers, the evolution has been positive, very closely to the baseline estimates.
A more evident difference relative to the baseline case appears in the estimates split by education level (Figure (ref), central panel), showing a great divide in the average level of unobserved ability between those with a college degree and those without. For the first group, this level has been higher and relatively stable, whereas for the latter it has been lower and decreasing steadily since the mid-nineties. The same divide can be observed for female workers although the average level of unobserved ability steadily increased for college-educated women, and has remained stable for lower-educated female workers.
Finally, if one allows the copula to be different for married and unmarried workers (Figure (ref), bottom panel), we find the largest differences for women: in particular, unmarried female workers tend to have a much higher level of unobserved ability, although the gap with married women has decreased in the last two decades. In contrast, the estimates for men are much similar to the baseline ones. The main difference corresponds to the period beginning in the late nineties, in which the average level of unobserved ability for married male workers is smaller.
Despite these differences in the amount of self-selection, the mean earnings gap for participants remains largely unaltered in these specifications (Figure (ref)). However, the decompositions do vary slightly. The most noticeable differences relative to the baseline estimates arise in the coefficients and selection components. In particular, the coefficients components is generally larger for the estimates with heterogeneous copulas by race and marital status, and smaller for the estimates that are heterogeneous by education level. The opposite is true for the selection component, as these two almost cancel each other out entirely. This reinforces the hypothesis that the instrument is weak for male workers. The only exception is the model with heterogeneous copula by marital status, for which the the participation components is more negative during the entire period, i.e., in favor of female workers.
In this paper I have introduced a new way to decompose differences in outcomes between two groups when there is self-selection into participation. In particular, rather than considering the decomposition of potential outcomes for the entire population, I decompose differences in actual outcomes for both participants and the entire population. These differences are decomposed into four components: endowments, coefficients, participation and selection. Moreover, I propose to provide two additional ancillary decompositions regarding differences in participation and self-selection.
I apply this methodology to analyze the labor earnings gap between males and females. I find that increases in female labor market participation and improvements in self-selection that led to an increase in unobserved ability for females have been responsible for a large share of the fall of the gap. Moreover, considering the gap for participants greatly underestimates the gap for the entire population, due to the still existing gender participation gap.