Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
102,908 characters · 32 sections · 64 citation commands
Decomposing Inequalities using Machine Learning and Overcoming Common Support Issues
\setcounter{tocdepth}{2}
The Kitagawa-Oaxaca-Blinder decomposition (kitagawa1955components; oaxaca1973male; blinder1973wage) is widely used in empirical studies to study differences in distributional statistics between two groups or two periods DiFoLe:96, FiFoLe:09, FiFoLe:18.\footnote{While the decomposition method is also referred to as the “Oaxaca-Blinder decomposition” in the economics literature, similar principles had been introduced earlier by Kitagawa in demographics and sociology. The links between these contributions are discussed in oaxaca2025oaxaca. Given the generality of our counterfactual framework in this paper, we do not attribute the method to any single author and instead cite all three contributions.} It separates an observed difference into an explained part due to individuals' observed characteristics and an unexplained part. This method is typically used to decompose average wage differences between men and women.
In classical decompositions, two major methodological choices are made. The first methodological choice consists of selecting a reference outcome. In the context of gender wage gap decompositions, the reference outcome is the wage an individual would receive in a counterfactual setting where men and women would be paid according to the same wage model. This means that each characteristic, such as education or experience, determines the reference outcome in the same way for men and women. In other words, the reference outcome is a counterfactual outcome that would be identical for an individual in one group and an individual in another group, sharing the same observable characteristics. Classical decompositions in the literature assume that the reference outcome is either the wage model of men or of women. Empirically, the final decompositions vary according to this methodological choice, with no consensus which group's potential outcome should be chosen as the reference FoLeFi:11. The second traditional methodological choice is to impose assumptions about outcome models, such as linearity, invertibility, and exogeneity, to perform ordinary least squares estimates. These standard assumptions can be relaxed by using reweighting techniques FoLeFi:11 and machine learning methods strittmatter2021gender. However, these techniques require an additional common support assumption.
In this paper, we propose a method to combine the advantages of both approaches, avoiding the restrictive assumptions of a standard linear model through machine learning and reweighting methods, without relying on the common support assumption. This is achieved by changing the reference outcome. Instead of choosing the outcome of one of the two groups as the reference outcome, we propose using a propensity score-weighted combination of the two potential outcomes as a reference, following neumark1988employers in the standard linear framework.
To understand why changing the reference outcome helps relaxing these standard assumptions simultaneously, we underline the idea that the common support hypothesis is needed whenever both reweighting techniques are used and the outcome of one of the two groups is taken as the reference outcome. Indeed, the combination of these two choices forces us to reweight the observed outcomes by the probability of belonging to the reference group, conditional on individual characteristics. This implies divisions by the conditional probability of belonging to the reference group, which is feasible only when these probabilities are bounded away from zero (hence the common support assumption). Intuitively, therefore, it is necessary for individuals to have comparable counterparts in the reference group, to avoid unstable weights driven by very small estimated probabilities. While this condition may be credible in situations close to randomized controlled trials, as in impact evaluation programs, it becomes more questionable in the context of inequality decomposition. For instance, using a large data set of 1.7 million employees in Switzerland, strittmatter2021gender find that with the 10 most important variables for explaining male wages, 55% of women in the private sector have no comparable men with respect to these variables. They show that the results are sensitive to the trimming used to have observationally comparable men and women, and thus that common support is a key issue. However, dividing by conditional probabilities is no longer required when choosing a reference outcome defined as a propensity score-weighted sum of the two potential outcomes. This choice of reference outcome was originally proposed by neumark1988employers within a linear framework, where it was limited to the use of only a few categorical variables. This reference outcome is sensitive to gender composition within sectors and is consistent with the standard concept of equal redistribution conditional on a set of characteristics. We extend this reference outcome to the case of a large set of characteristics, either categorical or continuous, and derive several estimators of the unexplained part of the mean difference. Unlike those based on standard reference outcomes, we show that the resulting reweighting estimators do not require the imposition of the common support hypothesis.
To tackle the assumption on the correct parametric specification of the reference outcome or the propensity score regressions, we consider Machine Learning (ML) techniques. These methods allow us to estimate potentially complex non-linear functions by mapping inputs to outputs and/or by performing variable selection when a large number of characteristics is available. Although it may be tempting to introduce ML estimates of unknown functions directly into standard estimators, this naive approach is not recommended because it provides biased estimators that converge at slower rates than the ${n}^{-1/2}$ rate obtained in the parametric framework. A more appropriate approach has been developed in recent literature on treatment effects. chernozhukov2018double provided a so-called double machine learning procedure to recover debiased estimators that achieve root-n convergence rates for the Average Treatment Effect (ATE) and Average Treatment Effect on the Treated (ATT). When the unexplained part in the Kitagawa-Oaxaca-Blinder decomposition can be interpreted as an ATT FoLeFi:11, double machine learning can thus be used to estimate it with good statistical properties strittmatter2021gender.
Our main contribution can be summarized as follows: drawing on Rubin's conceptual framework rubin1974estimating and reformulating the decomposition problem in terms of potential outcomes, we rewrite the unexplained part when the reference outcome is directly inspired by what is proposed by neumark1988employers. This means that we choose as a reference outcome a combination of the potential outcomes of the two groups, weighted by conditional probabilities of being in each group. While the unexplained part can be analogous to an Average Treatment effect on the Treated (ATT) or an Average Treatment effect on the Untreated (ATU) when choosing standard reference outcomes, this is no longer the case when we choose this reference outcome. Therefore, to apply the double machine learning method, we check that the orthogonality condition remains satisfied for this specific object. Furthermore, we demonstrate that this new construction of the unexplained part, based on this outcome, avoids any division by the propensity score. Thus, it becomes unnecessary to assume that the propensity scores remain far from 1 or 0, once the reference outcome is defined as a weighted average of the potential outcomes of the two groups. Nonetheless, adopting the reference outcome proposed by neumark1988employers has important implications for the choice of explanatory variables, as discussed in the paper.
The remainder of the article is structured as follows. In section (ref), we present the general framework and discuss the choice of the reference outcome. We present an equilibrium reference outcome based on neumark1988employers. Section (ref) is devoted to the identification strategy. The appropriate estimators are derived and it is shown that those based on the equilibrium reference outcome do not require the common support condition. Section (ref) presents the empirical strategy, distinguishing the standard parametric approach and the proposed non-parametric approach. Section (ref) shows two empirical illustrations. Finally, section (ref) concludes.
An individual may be either in group 1 or in group 0. We suppose groups to be two countries, two time periods or two subpopulations (e.g, females and males). Let $D$ be a random variable with support $\mathcal{X}_D=\{0,1\}$ denoting the group of an individual. We assume that the mean outcome is higher in group 1 than in group 0, so that group 1 is considered the advantaged group and group 0 the disadvantaged group. This is not an assumption, but it will be convenient for interpreting the models. Indeed, the labels of the two groups can always be interchanged to check it. One can express the problem in terms of potential outcomes, following the literature on treatment effects rubin1974estimating.
An individual in group 1 gets an outcome $Y(1)$ that is observed. In contrast, the outcome $Y(0)$ he would have received if he were in group 0 is not observed. For example, if the individual is a female, we only observe her wage as a female. $Y^\text{obs}$ is a random variable representing the outcome observed depending on the value of $D$ and is supposed to be continuous:\footnote{It is possible to extend the decomposition when $Y^\text{obs}$ is a limited dependent variable. In the case of a binary variable, one can assume an underlying probit model for example, and decompose the probability, as discussed in FoLeFi:11, section 3.5.}
Moreover, we assume that, in each group, the outcome may be explained by some individual characteristics $X \in \mathcal{X}$. They may be unevenly distributed between both categories. For example, men and women may not have the same average level of education. We want to decompose the observed difference in means between a part explained and a part not explained by the characteristics.
In general, the objective is to compare the outcome that individuals in the observed sample receive with an alternative outcome $Y(r)$ they could have received. In the simplest case, it can be the outcome received in the other group. Let $\Delta^\text{obs}$ be the expected observed difference in outcomes:
By adding and subtracting $\mathbb{E}\big[Y(r) \mid D= 0 \big]$ and $\mathbb{E}\big[Y(r) \mid D= 1 \big]$, we obtain
Thus, a difference in means between two groups can be decomposed in three parts which we denote by:
The decomposition of the observed difference in outcomes is sensitive to the choice of the reference outcome.
Thanks to the potential outcomes framework, we can bridge the gap between the inequality literature and the literature on treatment effects. When we choose as the reference outcome the potential outcome of one of the two groups, the unexplained component in decomposition analyses can formally resemble an ATT or an ATU. However, these analogies are purely notional. Binary variables such as gender or country of residence differ fundamentally from treatment assignments: they are not manipulable in an experimental sense. There is no experimental framework in which such variables would be randomly assigned across individuals. Drawing a parallel between a group dummy and a treatment variable in observational studies is also problematic. Unlike the baseline or pre-treatment covariates typically used to ensure comparability in treatment effect estimation, the covariates used for the decomposition are not necessarily determined before the group variable.
Taking $Y(r):=Y(0)$ implies that, in a world where individuals from both groups with the same characteristics receive identical outcomes, everyone obtains the potential outcome corresponding to the disadvantaged group $Y(0)$. In this case, $\Delta_S^{0,0}=0$. This leads to the decomposition:
Note that this is exactly the decomposition derived in angrist2009mostly, specifically in equation (3.2.1) of their book, in terms of Average Treatment effects on the Treated, ATT$:=\mathbb{E}[Y(1)-Y(0) \mid D=1]$, and selection bias.
Alternatively, one could consider that when people face no unexplained observed difference in means on the market, they all receive the outcome of the advantaged group, thus $Y(r) := Y(1)$ and $\Delta^{1,1}_S=0$. This leads to the following decomposition.
Here, the unexplained part of the difference in means is analogous to an Average Treatment effect on the Untreated, ATU$:=\mathbb{E}[Y(1)-Y(0) \mid D=0]$.
The choice of the reference outcome is arbitrary, and the decomposition can be viewed as arbitrary as well. The estimations of the unexplained part $\Delta_S^{0,1}$ (ATT) and $\Delta_S^{1,0}$ (ATU) may differ significantly in practice. There is no general guidance on this choice FoLeFi:11, and both estimations are often presented in empirical studies.
Building on the taste-based discrimination framework of becker1957economics and arrow1972some, neumark1988employers shows that the Kitagawa-Oaxaca-Blinder decomposition can be derived from a theoretical model of employers' discriminatory tastes. Setting $Y(r)=Y(1)$ reflects the idea of “pure” discrimination against women, while setting $Y(r)=Y(0)$ reflects the idea of “pure” nepotism toward men, that is, a willingness to hire men even at relatively higher wages. He also proposes a reference outcome where employers can be both discriminatory against females and nepotistic toward males. For an individual working in a given type of labor\footnote{For neumark1988employers, “types of labor” are job categories, illustrated by the distinction between skilled versus unskilled jobs. Alternative definitions could be based on other criteria, such as employment sectors.} $l$, the reference outcome is defined as the weighted average:
where $w_l$ and $(1-w_l)$ are the proportions of men and women working in labor type $l$. Here, the non-discriminatory wage is sensitive to the gender composition of each type of labor. This avoids hypotheses in employers' tastes such that they require more profits to compensate for hiring females at males wage, or they are not willing to accept lower profits to hire males at females wage. With a finite number of types of labor ($l=1, \dots, L$), much smaller than the number of observations ($L \ll N$), it is shown that the estimation requires only the estimation of the log wage regression for the full sample. Therefore, the implementation of this method is restricted to the use of only a few categorical variables, in order to have different types of labor with enough observations in each of them.
In the following, we extend the reference outcome in ($\ref{eq:y_neumark}$) to the case of a set of characteristics $X$, which may consist of many variables, either categorical or continuous. For an individual with a given set of characteristics $X$, the equilibrium reference outcome is defined as the weighted average:
where the weights are defined by the proportion of individuals who share the same set of characteristics in each group, that is,
In this more general setting, $Y(2)$ is nonlinear and the estimation of the equilibrium outcome requires non-parametric or machine learning methods (see the following subsection and section (ref)).
The equilibrium reference outcome aligns with the common notion of equal redistribution in economic inequality, which advocates for dividing a cake equally among a set of individuals. Indeed, when we consider several individuals who share the same set of characteristics $X=x$, the average equilibrium reference outcome is a weighted average of the outcome of individuals in each group. It corresponds to the total amount of outcome redistributed equally among individuals who share the same set of characteristics.
Choosing $Y(r)=Y(2)$, the decomposition in equation ((ref)) becomes:
Thus, the decomposition consists of three parts. Neither the advantage $\Delta_S^{r,1}$ nor the disadvantage $-\Delta_S^{r,0}$ disappear from the unexplained part when $r=2$ is chosen, unlike in the cases $r\in\{0,1\}$. The analogy between the unexplained part ($\Delta_S^{r,1}-\Delta_S^{r,0}$) and an object such as the ATT or ATU is no longer relevant.
By choosing $r \in \{0,1\}$, there is always a counterfactual quantity, never observed, to be estimated: either $\mathbb{E}\big[Y(0) \mid D=1\big]$ in the case $r=0$, or $\mathbb{E}\big[Y(1) \mid D=0\big]$ in the case $r=1$. If we take the example of the case $r=0$, we would like to determine what individuals in the group $D=1$ would receive if they were paid according to the model $Y(0)$. A common support problem arises when, for certain combinations of covariates $X$ observed in group $D=1$, there are no individuals in group $D=0$ sharing these characteristics $X$. In other words, $\mathbb{P}\big(D=0\mid X)=0$. In the absence of alter egos in the reference group, methodological choices are needed to estimate the counterfactual.
The first solution that can be proposed is the most standard one in classical decompositions: extrapolation. If several individuals in group $D=1$ are old, which is rarely or never observed in group $D=0$, then $Y(0)$ will be predicted outside the support of $X \mid D=0$ on the basis of a linear relationship between the age and the observed outcome in group $D=0$. Hence, the actual reference outcome would be:
$\tilde Y(0)$ corresponds to a counterfactual for which there is no observed alter egos in group $D=0$ since $\mathbb{P}(D=0\mid X)=0$. Thus, it cannot be estimated on actual observations, in contrary of $Y(0)$ that is observed for some individuals in group $D=0$. The only way to make a guess about this wage would be to extrapolate $\tilde Y(0)$ based on an estimation of $Y(0)$. The extrapolation $\tilde Y(0)$ requires additional assumptions, for instance the use of a parametric regression being stable outside the support of the sample used for the estimation. In a non-parametric regression framework, this approach is more problematic, as extrapolation outside the support is generally infeasible.
An alternative approach is to remove observations without (enough) pairs in the reference group. In practice, performing such trimming involves considering the following reference outcome:
where $\alpha$ is small and arbitrarily chosen. In the treatment effect literature, trimming is often performed for $\alpha=0.01$ or $\alpha=0.05$. However, there is no consensus on how to select $\alpha$. Putting ((ref)) in ((ref)) leads to replace $Y(0)$ by $Y(1)$ in the right component $\Delta_S^{0,1}$ for the men without (enough) alter egos only. In expectation, this amounts to removing those observations. Thus, trimming can be viewed as a specific extrapolation, $\tilde Y(0)=Y(1)$, where no discrimination is assumed outside the support of the reference group.
In this paper, we consider setting $r=2$, that is, we take the equilibrium outcome as the reference, as an alternative to extrapolation and trimming. An appealing feature of the equilibrium reference wage in ((ref))-((ref)) is that it is sensitive to the gender composition of each profile. Thus, the reference wage for a profile that is predominantly female will be closer to the current wage for women, and vice versa. The equilibrium reference outcome $Y(2)$ is smoothly designed and continuous. In contrast, the trimming reference outcome is discontinuous and changes abruptly from $Y(0)$ to $Y(1)$ at a given threshold defined by $\alpha$. In practice, the results are quite sensitive to the choice of $\alpha$ strittmatter2021gender. A nice feature of the equilibrium reference outcome is that such arbitrary preliminary choice is not required.
Figure (ref) illustrates the impact of the choice of the reference outcome on the difference in means decomposition with a lack of common support. Two sets of outcome values are drawn from linear regression models, for the advantaged group 1 and disadvantaged group 0, and the conditional probability model is a logit model.\footnote{$Y(1) = g(1,X) + \epsilon_1$ where $g(1,X)=0.3 + 0.42 \ X$ and $\epsilon_1 \sim \mathcal{N}(0,0.01)$ and $Y(0) = g(0,X) + \epsilon_0$ where $g(0,X)=0.2 + 0.2 \ X$ and $\epsilon_0 \sim \mathcal{N}(0,0.015)$ and $P(D=1|X)=1/(1+\exp(4-8 \ X))$} This design illustrates the presence of very few pairs for high and low values of $X$. The unexplained part of the decomposition in ((ref)) is based on $\mathbbm{E}[Y(r)|D=d]$, for $d\in (0,1)$. The latter is usually obtained from the expectation of the counterfactual $\mathbbm{E}[Y(r)|X]$, which is drawn with a black line in the figures (see next section for a discussion on the identification of the counterfactual).
Figure (ref) (left) shows the counterfactual with the outcome of the disadvantaged group taken as reference outcome, $\mathbbm{E}[Y(0)|X]$, in the case of extrapolation (black line). The counterfactual is extrapolated out of the support of the outcome values of group 0 in a continuous and smooth way, assuming that it is a linear function, $\mathbbm{E}[Y(0)|X]=\beta_0+\beta_1 X$. The unexplained part of the difference in means is given by the average of the deviations of group 1 observations from the counterfactual (orange vertical lines). It corresponds to the standard Kitagawa-Oaxaca-Blinder approach with $r=0$\footnote{Note that, in the classic approach, it is common to re-estimate both models using linear regression, as discussed in section (ref).}.
Figure (ref) (middle) shows the counterfactual with the outcome of the disadvantaged group taken as the reference outcome, in the case of trimming (black lines). Below the threshold $X=0.75$, the counterfactual is given by the conditional outcome of the disadvantaged group $\mathbbm{E}[Y(0)|X]$. Above the threshold, it is given by the conditional outcome of the advantaged group $\mathbbm{E}[Y(1)|X]$. This counterfactual is defined by a discontinuous line, with a jump at the selected threshold. The unexplained part of the difference in means is given by the average of the deviations of group 1 observations from the counterfactual (orange vertical lines). It is clear in this example that, by trimming observations with the largest deviations, the unexplained part becomes smaller than that obtained by extrapolation.
Figure (ref) (right) shows the counterfactual with the equilibrium reference outcome, $\mathbbm{E}[Y(2)|X]$ (black curve). It is a continuous and smooth function, moving from the outcome values of the disadvantaged group for low values of $X$ to the outcome values of the advantaged group for high values of $X$. This counterfactual takes into account the balance between the two groups, for given $X$. The unexplained part of the difference in means is given by the average of the deviations of group 1 observations from the counterfactual (orange vertical lines) minus the average of the deviations of group 0 observations from the counterfactual (green vertical lines). It is worth noting that the counterfactual is non-linear, even if $Y(0)$ and $Y(1)$ are drawn from two linear regression models. This is because $X$ is not categorical and does not define a limited number of profiles. Therefore, using the equilibrium reference outcome $(\ref{eq:y_newref})$ requires non-parametric or machine learning estimation methods to estimate the counterfactual.
Finally, interpreting the equilibrium reference wage in ((ref)) as “non-discriminatory" in gender gap analysis requires some caution. Indeed, the effect of discrimination is to redistribute wages within each profile, defined by a set of characteristics $X$. This implies that:
Each decomposition, either ((ref)), ((ref)), or ((ref)), requires to estimate observable and counterfactual quantities. The next session presents identification strategies for both.
We will focus here on the unexplained part of the observed difference in means defined in ((ref)):
The targeted parameters are the unexplained parts $\Delta^{r,1}_S$ and $\Delta^{r,0}_S$. Hence, we aim at estimating $\Delta^{r,d}_S$ with $d \in \{0,1\}$, which can be written as a difference between an observable mean and a counterfactual mean:
When $r=d$, it is equal to zero. Thus, we let $r \in \{0,1,2\} \setminus d$ herafter.
The observable part can be easily estimated by a sample mean. Indeed, since $Y(d)$ is the outcome always observed in group $D=d$, from the law of total probability we have:
The central problem in ((ref)) is the identification of the counterfactual mean, which is unobserved. The different strategies we will consider are:
The outcome regression approach is well-known in the literature on inequality; it is used in the standard Kitagawa-Oaxaca-Blinder decomposition method. The IPW and AIPW are often used in the treatment effects literature.
These strategies are detailed below and require more details on either the outcome or the probability model. We will see that the identification of the counterfactual in ((ref)) requires the following assumption:
This condition states that the potential outcome is independent of group membership $D$, conditional on covariates $X$. It implies mean independence:
Therefore, within the subpopulation of individuals with the same covariates, the difference in the distributions of the observed outcomes between individuals from the two groups fairly represents the difference in means in this subpopulation imbens_causal_2015. This is because within this subpopulation, the individuals in the two groups are both random samples from that subpopulation.
With the law of iterated expectations and assumption (ref) [ignorability] for $Y(r)$, the counterfactual in ((ref)) can be related to the observed outcome:
Let us define the conditional expectation by a regression function:
From ((ref)), ((ref)), ((ref)) and ((ref)), the unexplained advantage of group 1 ($\Delta_S^{1,0}$) and the unexplained disadvantage of group 0 ($-\Delta_S^{0,1}$) can be identified with:
We first show that the ignorability assumption allows us to link the potential equilibrium outcome to the observed outcome, while $Y(2)$ is never observed directly in the $D=0$ or $D=1$ group.
With ((ref)), ignorability and the law of iterated expectations, the counterfactual in ((ref)) can be related to the observed outcome as follows:
Let us define the conditional expectation by a regression function:
From ((ref)), ((ref)), ((ref)) and ((ref)), the unexplained advantage of group 1 ($\Delta_S^{2,1}$ ) and the unexplained disadvantage of group 0 ($-\Delta_S^{2,0}$) can be identified with:
The main idea of the Inverse Probability Weighting (IPW) is that an expectation computed from a subpopulation can be corrected to reflect the expectation for the whole population. $\text{Counterfactual}_{r,d}$ in ((ref)) is a conditional expectation. Therefore, IPW involves finding weights $\pi(X)$ such that:
When the outcome of the advantaged or disadvantaged group is taken as the reference outcome, an additional condition is required.
For example, to estimate $\mathbb{E}\big[Y(1)-Y(0) \mid D=1\big]$, this requires: $\mathbb{P}\big(D=0\mid X \big)> 0$. The consequence of this assumption is that there is no case where a specific set of individual characteristics forces the individual to be found in the group $D=1$ only. In other words, there are alter-egos in the group $D=0$ for every individual in group $D=1$.\footnote{In section (ref), we discuss in detail the links between this assumption in the context of inverse probability weighting, and the assumption of invertibility of $X'X$ generally made in the case of extrapolation.} In practice, IPW estimators require one to drop observations with extreme propensity scores, so that the resulting estimation does not depart from the truth due to a division by zero. This causes a lack of observations that could be avoided in the standard Kitagawa-Oaxaca-Blinder decomposition.
Under Assumptions (ref) [ignorability] and (ref) [common support], using Bayes' rule and the law of iterated expectations, one can show that the weights needed in ((ref)) for $(r, d) \in \{0,1\}^2$ are:\footnote{See for example kline2011oaxaca for a proof.}
Using ((ref)) in ((ref)) and the law of total probability, this leads to the following counterfactual:
Let us define the conditional probabilities by a regression model:
From ((ref)), ((ref)), ((ref)) and ((ref)), the unexplained advantage of group 1 ($\Delta_S^{0,1}$) and the unexplained disadvantage of group 0 ($-\Delta_S^{1,0}$) can be identified:
It is clear from this equation that $p(r,X)$ must be different from zero. The common support condition is then required to identify the unexplained part of the mean decomposition.
Similarly, under ignorability, we can derive $\pi(X)$ such that equation ((ref)) is satisfied in the case $r=2$, building on kline2011oaxaca:
From ((ref)), ((ref)), ((ref)) and ((ref)), the unexplained advantage of group 1 ($\Delta_S^{2,1}$) and the unexplained disadvantage of group 0 ($-\Delta_S^{2,0}$) can be identified:
Unlike in ((ref)), no restriction is required on $p(\cdot,X)$ as it does not appear at the denominator. Therefore, the assumption (ref) [commmon support] is not required to identify the unexplained part of the observed difference in outcomes.
A doubly-robust estimator is consistent if at least one of the models is correctly specified, either the outcome model or the probability model robins1995semiparametric. Thus, it combines both strategies. The Augmented Inverse Probability Weighting (AIPW) estimator is the most commonly known doubly robust estimator, and typically relies on parametric regressions to estimate the outcome model and the propensity score.\footnote{In fact, kline2011oaxaca shows that the standard Kitagawa-Oaxaca-Blinder estimator constitutes a propensity score re-weighting estimator based upon a linear model for the conditional odds of being treated. This includes assignment models with a latent log-logistic error structure and may yield negative weights. As such, it is a “doubly robust" estimator.} It provides an estimator for $\text{Counterfactual}_{r,d}$.
When the outcome of the advantaged or disadvantaged group is taken as reference outcome, $r \in \{0, 1\}$, providing that assumption (ref) [commmon support] holds, the estimator of the counterfactual mean is equal to:\footnote{((ref)) is equal to $\mathbb{E} \Big[ \frac{Y^\text{obs} \mathbbm{1}(D=r)}{\mathbb{P}(D=d)} \frac{p(d,X)}{p(r,X)} + \Big( \frac{\mathbbm{1}(D=d)}{\mathbb{P}(D=d)} - \frac{p(d,X)}{p(r,X)} \frac{\mathbbm{1}(D=r)}{\mathbb{P}(D=d)} \Big) g(r,X) \Big]$ from which we have ((ref)).}
When $g(r,X)$ is correctly specified, the first term in ((ref)) disappears and the counterfactual reduces to that of the outcome regression approach in ((ref)). When $p(r,X)$ is correctly specified, the second term in ((ref)) disappears and the counterfactual reduces to that of the IPW in ((ref)).
From ((ref)), ((ref)) and ((ref)), the unexplained advantage of groups 1 ($\Delta_S^{1,0}$) and the unexplained disadvantage of group 0 ($-\Delta_S^{0,1}$) can be identified:
It is clear from this equation that $p(r,X)$ must be different from zero and that the common support condition is required to identify the unexplained part of the mean decomposition.
When the equilibrium outcome is taken as reference outcome, $r=2$, the estimator of the couterfactual mean is equal to:
When $g(r,X)$ is correctly specified, the first term in ((ref)) disappears and the counterfactual reduces to that of the outcome regression approach in ((ref)). When $p(r,X)$ is correctly specified, the second term in ((ref)) disappears in average and the counterfactual reduces to that of the IPW in ((ref)).
From ((ref)), ((ref)) and ((ref)), the unexplained advantage of groups 1 ($\Delta_S^{2,0}$) and the unexplained disadvantage of group 0 ($-\Delta_S^{2,1}$) are identified:
Unlike in ((ref)), no restriction is required on $p(\cdot ,X)$ as it does not appear at the denominator. Therefore, assumption (ref) [commmon support] is not required to identify the unexplained part of the observed difference in outcomes.
In the previous section, the unexplained parts of each group are identified. Estimators of the unexplained part of the observed difference in ((ref)) can then be derived, replacing expectations and functions $g(r,\cdot)$ and $p(d,\cdot)$ by sample means and estimations $\hat g(r,\cdot)$ and $\hat p(d,\cdot)$. For the different reference outcomes, we obtain the estimators below (see details in Appendix (ref)).
Note that the weights from the IPW and AIPW estimators may not sum up to one in finite sample. One can rewrite the IPW and AIPW estimators as a function of weights that can be normalized to sum up to one in this case. BuDiMc:14 show that normalized IPW estimators with $r=0$ and $r=1$ perform often better than unnormalized IPW estimators in finite samples. We show normalized IPW and AIPW estimators for $r=\{0,1,2\}$ in Appendix (ref).
All these estimators require Assumption (ref) [ignorability]. Moreover, Assumption (ref) [common support] is required for $r=0$ and $r=1$. Therefore, trimming observations when $\hat p(1,X_i)$ is close to 0 when $r=1$ or 1 when $r=0$ is often performed in practice. It is not required for $r=2$.
Parametric regression models can be estimated efficiently under regular assumptions. However, the unbiasedness and consistency of the estimators depend on the correct specification of the model.
The standard decomposition approach is implemented with the estimators in ((ref)), ((ref)) and ((ref)), where the outcome regression $g(r,\cdot)$ is estimated from a parametric linear model:
With the exogeneity assumption, ((ref)) is equivalent to ((ref)) and ((ref)), and the regression function $g(r,.)$ can be estimated consistently from the linear outcome model above. The unknown coefficients $\beta_r$ are usually estimated by OLS, when $X'X$ is invertible.
In the case $r\in\{0,1\}$, under assumptions (ref) [linear outcome] and (ref) [exogeneity], the unexplained part of the observed difference can be estimated with $\hat\delta_1^{\text{reg}}$ in ((ref)) or $\hat\delta_0^{\text{reg}}$ in ((ref)), where $\hat g(1,X_i)$ and $\hat g(0,X_i)$ are replaced, respectively, by OLS estimators $X_i\hat\beta_1$ and $X_i\hat\beta_0$.
In the case $r=2$, even if the outcome regression in each group is linear, $g(2,\cdot)$ is not necessarily linear.\footnote{ Indeed, with assumption (ref) [linear outcome] for both $r=0$ and $r=1$ , we have: $g(2,X) = p(1,X) X \beta_1 + p(0,X) X \beta_0$ which is not linear, except in two cases: if $p(1,X)$ does not depend on $X$ ; or if $\beta_1=\beta_0$ except for the intercept.} This is illustrated in Figure (ref) with $X$ continuous (see section (ref)). Nonetheless, neumark1988employers shows that if $X$ defines a finite number of types of labor or workers, the equilibrium reference outcome can be estimated from a linear regression model, $g(2,X)=X\beta$, by regressing $Y^\text{obs}$ on $X$ from the full sample. Therefore, the unexplained part of the observed difference can be estimated with $\hat\delta_2^{\text{reg}}$ in ((ref)), where $\hat g(2,X_i)$ is replaced by $X_i\hat\beta$. However, the case where $X$ defines only a few categories is rather restrictive, and this approach is not often used in empirical studies.
{\bf Alternative approach}: The estimator of the unexplained part of the decomposition proposed by oaxaca1973male and blinder1973wage can be obtained by replacing $Y^\text{obs}_i$ by $X_i\hat\beta_0$ in ((ref)), or $Y^\text{obs}_i$ by $X_i\hat\beta_1$ in ((ref)). For instance, with $r=0$, the estimator of the unexplained part of the decomposition ((ref)) becomes
where $\hat\beta_1$ (resp. $\hat\beta_0$) is the OLS estimator of $\beta_1$ (resp. $\beta_0$) obtained from a regression of $Y^{\text{obs}}$ on $X$ in group 1 (resp. group 0). In this approach, outcome regressions are estimated for both groups. The first advantage in the linear case is that the resulting decomposition is directly interpretable. The second advantage is that, if the same mistake is made in both models, such as omitting a variable that linearly affects both outcomes in the same proportions, then the resulting difference will be robust to this misspecification. However, a different set of assumptions is required. Indeed, assumption (ref) [exogeneity] can be slightly relaxed, but assumptions (ref) [linear outcome] for {both} $r=0$ and $r=1$ are required.
{\bf Invertibility versus common support}: As indicated in section (ref), the common support assumption is made when using the IPW approach in the $r\in\{0,1\}$ case. In the case of extrapolation, this assumption is not necessary, but the invertibility of $X'X$ is assumed instead when using OLS to estimate the regression. Although these are two fundamentally different hypotheses, violations of both can arise from similar situations. Suppose we assume $r = 1$. In this case, the counterfactual to be predicted is $\mathbb{E}\big[Y(1) \mid D = 0\big]$. For extrapolation, a prediction model must be estimated on the $D=1$ subgroup, then predicted on the $D=0$ group. Suppose that there exists $X_k$, a dummy variable in the model which indicates that the individual has obtained diploma $k$ if $X_k = 1$, and suppose that all members of group $D=1$ have obtained diploma $k$, and that only some members of group $D=0$ have received it. Let's also assume that $X_k$ is a determinant of $Y(1)$. The effect of $X_k$, omitted from the linear regression estimated by OLS to invert $X'X$, will be captured in the intercept. This does not prevent accurate prediction of $Y(1)$ in group $D=0$ for individuals who have obtained diploma $k$. However, it fails to provide valid predictions of $Y(1)$ for individuals in group $D=0$ who do not hold the diploma. Indeed, the lack of variation in $X_k$ in the other group makes it impossible to estimate $Y(1)$ for this subgroup. At the same time, denoting $X_{-k}$ the set of covariates deprived of $X_k$, $\mathbb{P}(D=1 \mid X_k = 0, X_{-k})=0$, violating the common support assumption. The complete absence of alter egos in the reference group ($D=1$) for a subgroup in $D=0$ simultaneously undermines both strategies, even though the formal assumptions they rely on differ in general. In other words, the same data configuration can simultaneously violate both the invertibility and common support assumptions, although in other cases one assumption may hold while the other fails.
The previous standard decomposition approach requires the correct specification of the reference outcome regression model. Estimators based on the IPW or the AIPW strategies allow relaxing this hypothesis. In the parametric case, the estimation of the propensity score is often obtained by maximum likelihood from a probit or logit model:
{\bf The IPW} approach relies on the correct specification of a parametric model for the propensity score $p(1,X)$, defined in ((ref)).
In the case $r\in\{0,1 \}$, under assumptions (ref) [ignorability], (ref) [common support] and (ref) [logit/probit], the unexplained part of the observed difference can be estimated with $\hat\delta_1^{\text{ipw}}$ in ((ref)) and $\hat\delta_0^{\text{ipw}}$ in ((ref)), where $\hat p(1,X_i)$ is replaced by the maximum likelihood estimator $F(X_i\hat\beta)$. In practice, observations are trimmed when $F(X_i\hat\beta)$ is close to 0 or 1.
In the case $r=2$, under the same assumptions, except assumption (ref) [common support], the unexplained part of the observed difference can be estimated with ((ref)). It is worth noting that this estimator does not require any restriction on the shape of the outcome regression and remains valid even when the propensity score takes values equal to 0 or 1. Therefore, with a correct specification of the propensity score model, $\hat\delta_2^{\text{ipw}}$ in ((ref)) permits to obtain an estimation of the unexplained part of the decomposition based on the Neumark's reference outcome, without $X$ being restricted to a few number of categories and without trimming observations.
{\bf The AIPW} approach combines both strategies. The resulting estimators are doubly robust, meaning that they remain consistent if either the outcome model $g(r,\cdot)$ or the propensity score model $p(1, \cdot)$ is correctly specified.
In the case $r\in\{0,1 \}$, under assumptions (ref) [common support] and either (ref) [linear outcome] and (ref) [exogeneity] or either (ref) [ignorability] and (ref) [logit/probit], the unexplained part of the observed difference can be consistently estimated with $\hat\delta_1^{\text{aipw}}$ in ((ref)) and $\hat\delta_0^{\text{aipw}}$ in ((ref)) using a parametric approach. In practice, observations are trimmed when $\hat p(1,X_i)=F(X_i\hat\beta)$ is close to 0 or 1.
In the case $r=2$, the AIPW estimator remains consistent as long as either the outcome regression model on the pooled sample or the propensity score model is correctly specified. Thus, under the assumptions (ref) [linear outcome] and (ref) [exogeneity] with $r=2$, or under the assumptions (ref) [ignorability] and (ref) [logit/probit], the unexplained part of the observed difference can be estimated with $\hat\delta_2^{\text{aipw}}$ in ((ref)). Consequently, no trimming is required.
Supervised machine learning methods predict an outcome $Y^\text{obs}$ using covariates $X$. Here, we will use the term machine learning methods to refer to classical methods such as lasso, ridge, random forests, boosting, or neural networks (see hastie2005elements). Linear regression estimated by OLS and logit/probit models estimated by maximum likelihood may be presented as basic methods in standard textbooks. However, the term “machine learning” more generally refers to flexible methods capable of capturing nonlinear relationships between covariates and the outcome, interaction effects between variables on the outcome, and/or selecting relevant covariates for models. These methods can therefore be used to predict $\mathbb{E}\big[Y^\text{obs}\mid X\big]$, $\mathbb{E}\big[Y^\text{obs}\mid X, D=d \big]$ for any $d\in\{0,1\} $, or $\mathbb{P}\big(D=d \mid X\big)$ for any $d\in\{0,1\}$. This is why they can be considered as alternatives or extensions to the parametric methods presented in the previous sections.
However, these non-parametric methods are often based on a bias/variance tradeoff and therefore provide biased estimates with slower convergence rates than parametric models. This may not be an issue when the focus is on prediction, but it may be problematic when the quality of the estimation is key and inference is required, as in the estimation of a treatment effect. In our application, the parameter of interest is not directly the quantity predicted by supervised machine learning techniques. Rather, we aim to estimate it using plug-in strategies based on either an estimation for the outcome model, the propensity score model, or both. Since these naive plug-in approaches neither guarantee convergence to the target value at a sufficient rate nor allow for the construction of valid confidence intervals or the asymptotic normality of the estimators, alternative strategies have been proposed in the literature to adapt plug-in methods estimated via machine learning while preserving desirable properties for statistical inference. chernozhukov2018double developed a theory of inference with machine learning methods and proposed a double/debiased Machine Learning method providing estimators that can achieve $\sqrt{n}$-consistency and asymptotic Normality. Let $\theta_0$ be the true value of a low-dimensional parameter of interest and $\eta_0$ the true nuisance parameter. In our decomposition framework, $\theta_0$ would be $\Delta_S^{r,d}$ and $\eta_0$ would be $g(r, \cdot)$ and $p(d, \cdot)$, for $r \in \{0, 1, 2\}$ and $d \in \{0,1\}$. The double/debiased machine learning is based on two crucial ingredients: the Neyman orthogonality condition and cross-fitting.
{\bf Neyman orthogonality}: Let us consider scores $\psi(.)$ that satisfy the following identification condition:
and the Neyman orthogonality condition
where $W=(Y^{\mathrm{obs}}, D, X)$ and $c\in [0,1)$. An estimator of the parameter of interest $\theta_0$ can be obtained from the empirical analog of ((ref)) with $\eta_0$ replaced by a machine learning estimation $\hat\eta$. The orthogonality condition ((ref)) ensures the estimator of $\theta_0$ to be insensitive to small mistakes in the estimation of nuisance parameters. chernozhukov2018double show that, with scores defined by ((ref)) and ((ref)) and nuisance parameters quite well estimated by $\hat\eta$,\footnote{It is required that $\hat\eta$ converges at a rate faster than $n^{-1/4}$. This rate is achieved by many ML methods.} the estimator of $\theta_0$ obtained from
where $\sigma^2=\mathbb{E} [\psi^2(W_i;\theta_0,\eta_0)]$ and $n$ it the total number of observations in the sample.
{\bf Cross-fitting}: To avoid overfitting, often associated with machine learning methods, the parameter of interest and the nuisance parameters should be estimated from two different samples. Let us consider $K$ replications, where an estimator of $\theta_0$ is obtained on a random subsample $I_k$ with $n_k$ observations from
while the nuisance functions $\hat\eta_k$ are estimated from an auxiliary sample, which includes all observations that are not in $I_k$. This can be done for each replication. Finally, the parameter of interest $\theta_0$ can be estimated by averaging all $\hat\theta_k$
and the variance of $\hat\theta$ can be estimated with $\hat\sigma^2/n$ where
Hence, the standard error of $\hat\theta$ is
Estimation and inference of a parameter of interest can then be made with good properties in a quite general framework.
It is known from chernozhukov2018double that the score functions based on the AIPW estimators $\hat\delta_1^{\text{aipw}}$ in ((ref)), $\hat\delta_0^{\text{aipw}}$ in ((ref)) satisfy the Neyman orthogonality condition. Importantly, we extend this result by proving that the estimator $\hat\delta_2^{\text{aipw}}$ in ((ref)) which corresponds to the unexplained component of the decomposition when using neumark1988employers's equilibrium outcome as the reference, also satisfies the Neyman orthogonality condition. Thus, these estimators can be used with machine learning methods and cross-fitting to estimate the unexplained part of the observed difference in outcomes, as detailed in Algorithm (ref).
When the outcome of the disadvantaged group is taken as reference outcome, $r=0$, the score function is equal to
where $p(0,X)=1-p(1,X)$, and the nuisance functions are $\eta = \{g(r,\cdot), p(r,\cdot)\}$. chernozhukov2018double show that the score function $\psi_{0}$ satisfies the identification condition ((ref)) and the orthogonality condition ((ref)).\footnote{The score function $\psi_{0}$ in ((ref)) is similar to equation (5.4) in chernozhukov2018double with $\bar g(X)=g(0,X)$, $m(X)=p(0,X)$ and $p=\mathbbm{P}(D=d)$.} The estimator of $\delta$ is derived from ((ref)), by setting the expectation of ((ref)) equal to zero. Since $\mathbb{E} [\mathbbm{1}(D=d)] = \mathbb{P}(D=d)$, it is easy to see that the estimator of $\delta$ is equal to ((ref)) with $r=0$ and $d=1$.
When the outcome of the advantaged group is taken as reference outcome, $r=1$, the score function is equal to
The score function $\psi_{1}$ satisfies the identification condition ((ref)) and it respects the orthogonality condition ((ref)) as shown by chernozhukov2018double. Indeed, $\psi_{1}$ can be obtained by switching the role of the two groups in ((ref)) and by changing the sign of the second component.
The empirical scores are, respectively, equal to
They can be used to obtain variances of the estimators, by calculating the empirical variance of these scores, divided by $n$.
Therefore, appropriate estimators of the unexplained part of the difference in means with the reference outcome taken from group $r=0$ (ATT) or group $r=1$ (ATU) are obtained with the AIPW estimators $\hat\delta_1^{\text{aipw}}$ in ((ref)) or $\hat\delta_0^{\text{aipw}}$ in ((ref)), with variances obtained from ((ref))-((ref)).
When the equilibrium outcome is taken as reference outcome, $r=2$, let us consider the score function:
where $\eta = \{g(2,\cdot), p(1,\cdot)\}$. This score satisfies the identification condition ((ref)) and the orthogonality condition ((ref)) (see proof in Appendix (ref)).\footnote{Although the main paper focuses on the unexplained part, we also provide in Appendix (ref) an AIPW estimator for the explained part when $r=2$, its corresponding score function, and a proof that this score satisfies the Neyman-orthogonality condition to perform the full decomposition.}. The estimator of $\theta$ is derived from ((ref)), by setting the expectation of ((ref)) equal to zero. The empirical analog of $\theta$ is equal to the AIPW estimator in ((ref)) and then in ((ref)).
The empirical scores are equal to:
They can be used to obtain variances of the estimators, by calculating the empirical variance of these scores, divided by $n$. \
Therefore, an appropriate estimator of the unexplained part of the difference in means with the equilibrium reference outcome can be obtained with the AIPW estimator $\hat\delta_2^{\text{aipw}}$ in ((ref)), with variance obtained using the empirical scores ((ref)) in ((ref)), combined with cross-fitting.
We discuss several practical considerations: calibration, trimming and the use of pooled regression for the equilibrium reference outcome.
The use of machine learning models to estimate propensity scores requires some caution. Unlike parametric logit models, ML models for classification are generally not well-calibrated: estimated probabilities may not match empirical probabilities. In other words, an estimated probability of ‘70% of being treated' is not followed by ‘70% of units being treated' in samples. A poorly calibrated model can be problematic, since the double orthogonalization requires that outcome and propensity score regressions are quite well estimated. In practice, calibration plots, showing the empirical probability versus the predicted class probability, can be used to check if a model is not well-calibrated, and calibration correction methods can be used NiCa:05.
When the propensity score is close to zero or one, the variance of the IPW estimators can be huge when $r=0$ and $r=1$, because $\hat p(1,X_i)$ appears at the denominator in ((ref)) and ((ref)). In practice, it is often recommended to trim observations above/below a threshold to avoid very large weights. The standard estimators ((ref)) and ((ref)) may seem appealing, as they do not require propensity score estimation. In fact, these estimators are not transparent about the common support condition. Instead, they extrapolate to regions where there is no data. In linear regression models, it may seem reasonable to assume that the regression model is still the same outside of support, but this is often no longer the case in a non-parametric framework. Therefore, with propensity scores close to zero or one, an accurate estimation of the unexplained part is unlikely with none of the estimators. This problem does not arise with estimators based on the equilibrium reference outcome $r=2$.
From ((ref)), the equilibrium reference outcome can be obtained by regressing $Y^\text{obs}$ on $X$ from the pooled sample, as suggested by neumark1988employers. However, fortin2008gender and jann2008stata argue that this approach may transfer an inappropriate component in the explained part of the decomposition. Our potential outcomes approach clarifies the debate. We show that the correction proposed in the literature lead to a modification of the reference outcome.
Let us consider a linear framework, with the following regression model
jann2008stata suggests to estimate the regression from the pooled sample including $D$ as additional covariate:
and to use $\hat{\alpha}^* + X\hat\beta^*$ as reference outcome. It causes a change in the reference outcome, which becomes:
see Appendix (ref). It remains to replace the weights in ((ref)) or ((ref)) by the unconditional probability $\pi$.
To compare the implications of the reference outcome selection, it is useful to consider the explained part in a linear framework. For the standard reference outcome, $r=0,1$, and for the Jann's method, $r=3$, we have:
The explained parts are defined by the mean differences in $X$ between the two groups, multiplied by the coefficients of group 0, 1 or a weighted average between the two as $\beta+\pi\delta=(1-\pi)\beta+\pi(\beta+\delta)$. For the Neumark or equilibrium reference outcome, we have:
The explained part is still defined by the mean differences in $X$ between the two groups, but also by the mean differences in $p(1,X)$ and in $Xp(1,X)$. See Appendix (ref) for detailed calculations.
The new terms, multiplied by $\gamma$ and $\delta$ in ((ref)), come from the fact that $p(1,X)$ is used to define the equilibrium reference outcome $Y(2)$. At first sight, the presence of these two terms in the explained part seem consistent, since they differ from zero when the covariates $X$ are different between the two groups. However, they may be undesirable depending on the choice of $X$. To see why, let us consider the case where $X$ is correlated to $D$ but it does not explain $Y^\text{obs}$, that is:
For instance, if $X$ denotes “having long hair” and the difference in means refers to the gender wage gap, using $r=0,1,3$ would lead to the correct conclusion that mean wage differences are not explained by $X$, while using $r=2$ would lead to the wrong conclusion that it explains wage differences. fortin2008gender and jann2008stata suggest adding the variable $D$ in the regression model to correct this issue. In Appendix (ref), we demonstrate that this estimation method leads to change the reference outcome as it estimates on average a quantity similar to $\Delta_X^3$ and not to $\Delta_X^2$ in the linear case. We would therefore like to emphasize that Jann's strategy is not a “correction” of an error made when estimating the explained component in the case $r=2$. Rather, it is a change of reference outcome and target parameter that removes the influence of irrelevant covariates. It does not allow probabilities to be conditioned on the relevant covariables as the probabilities in equation ((ref)) are unconditional. This analysis highlights the importance of appropriate covariate selection in the case $r=2$, as using an inappropriate covariate may introduce undesirable component in the explained part.
To illustrate the findings of the previous sections, we use two datasets and estimate the unexplained part of the difference in means of log-wages between two groups. We consider several reference outcomes: of the disadvantaged group ($r=0$), advantaged group ($r=1$), and the proposed equilibrium reference ($r=2$). We calculate several estimators:
First, we consider a parametric approach, with an OLS estimation of the outcome linear regression and a logit estimation of the propensity score regression. The standard errors are computed by bootstrapping pairs, that is, by resampling randomly and with replacement lines in the original sample (Freedman:81). The number of bootstrap replications is equal to $B=999$.
Then, we consider a machine learning approach, with a gradient boosting estimation of the outcome and propensity score regressions (see section (ref)). Estimates of the unexplained part of the difference in means are obtained from Algorithm (ref), repeating the cross-fitting procedure $K=100$ times where the full sample is randomly split into two subsamples $I_k$ and $I_k^c$ of equal size. A statistical framework has been developed for AIPW estimators combined with cross-fitting (ML-AIPWu, ML-AIPWn). This is the so-called double machine learning method (chernozhukov2018double). Therefore, we show standard errors for AIPW estimators only. They are computed from ((ref))-((ref)) with score functions ((ref)), ((ref)) and ((ref)) for un-normalized weights and with ((ref)), ((ref)) and ((ref)) for normalized weights. Standard and IPW estimates are calculated using a machine learning approach for indicative purposes.
The data are obtained from the application in Hlav:18, on labor wages and demographic characteristics of 712 employed Hispanic workers in the Chicago metropolitan area.\footnote{The data can be obtained from the chicago data frame, included in the oaxaca R package.} The difference in means of log-wages between natives and foreign-born workers is equal to 0.1434. The covariates selected in the wage equation and in the propensity score regression are age, gender, and education.\footnote{We use the variables {\em age, female, LTHS, some.college, college} and {\em advanced.degree} from the chicago dataset.} The quantity of interest is the part of the mean difference unexplained by this set of characteristics.
Table (ref) shows estimates and standard errors of the unexplained part of the difference in means of log-wages between native and foreign-born workers. Several reference outcomes are considered: of foreign-born workers ($r=0$), native workers ($r=1$), and the equilibrium reference ($r=2$). The first and second columns show the results of the parametric approach ({\tt OLS}). The third and fourth columns show the results of the machine learning approach ({\tt ML}). We can see that:
Finally, the double machine learning method provides estimators robust to misspecification of regression models with valid inference (ML-AIPWu and ML-AIPWn). Using this method with several reference outcomes $r=0,1,2$, the results suggest that the part of the difference in means which is unexplained by the set of characteristics is quite large, much higher than that obtained by a parametric approach (OLS).
The data are obtained from CheHaSp:15, on labor wages and socio-economic characteristics of 29{,}217 employed men and women in the United States in 2012.\footnote{The data can be obtained from the cps2012 data frame, included in the hdm R package.} The difference in means of log-wages between men and women is equal to 0.2608. The covariates selected in the wage equation and in the propensity score regression are, among others, on marital status, education, and experience.\footnote{We use the variables {\em widowed, divorced, separated, nevermarried, hsd08, hsd911, hsg, cg, ad, mw, so, we, exp1, exp2, exp3} from the cps2012 dataset.} The quantity of interest is the part of the mean difference unexplained by this set of characteristics.
Table (ref) shows the unexplained part estimates of the difference in means of log-wages between men and women, that is, of unexplained part of the gender wage gap. Several reference outcomes are considered: of women workers ($r=0$), men workers ($r=1$), and the equilibrium reference ($r=2$). The first and second columns show the results of the parametric approach ({\tt OLS}). The third and fourth columns show the results of the machine learning approach ({\tt ML}). We can see that the results are very similar, whatever the estimator (Reg, IPWu, IPWn, AIPWu, AIPWn) or estimation method (OLS, ML) selected. They suggest that the parametric regressions for outcome and propensity score are correctly specified. This may be partly due to the fact that most explanatory variables are categorical (12 out of 13), and that interactions between covariates do not play a significant role. Overall, the results show that the gender wage gap is largely unexplained by the set of characteristics.
In this paper, we reformulate the problem of inequality decomposition using a potential outcomes framework. This approach clarifies the link between the unexplained component and objects like the Average Treatment Effect on the Treated (ATT) or Untreated (ATU). Changing the reference outcome - typically the potential outcome of one group - to an equilibrium outcome, as originally proposed by neumark1988employers, becomes straightforward within this framework.
Our main contribution is to analyze the unexplained component using this equilibrium reference outcome, demonstrating that it allows for a doubly robust estimator satisfying Neyman orthogonality, enabling estimation via double machine learning chernozhukov2018double. This provides an alternative to extrapolation or trimming strategies, relaxing two key assumptions: the common support condition and the need for correctly specified, low-dimensional parametric models. We present a simple algorithm to estimate this object, along with empirical applications and a broader methodological discussion aimed at applied researchers.
We highlight methodological issues that machine learning alone does not fully resolve. When choosing the potential outcome of one of the two groups as the reference outcome, which is the standard choice, adding many variables correlated with the group indicator to the models can lead to excessive trimming. This is equivalent to performing the decomposition on only a subgroup of the population. When using Neumark's reference outcome, trimming is not necessary, as the AIPW estimator no longer requires division by the propensity score. However, the key assumption becomes that the reference outcome is a weighting of the potential outcomes of the two groups by the propensity score. Thus, we implicitly assume that any variable correlated with the group indicator enters Neumark's outcome model, even when it is not a variable determining the potential outcomes models taken separately. This makes the decomposition based on Neumark's reference outcome sensitive to including irrelevant variables. This problem was already known in the linear case, and one solution proposed in the literature was to include the group indicator as a control in the estimation of the observed outcome model on the pooled sample. Thanks to our approach based on potential outcomes, we can demonstrate that this solution is not entirely satisfactory because it implicitly changes the reference outcome. We therefore conclude that applied researchers must carefully reflect on the relevance of the chosen variables, even when using these data-driven methods to analyze inequalities. Additionally, calibration issues should also receive special attention from empiricists, as they can significantly impact the robustness and interpretability of empirical findings.
\printbibliography[heading=bibliography]