EconBase
← Back to paper

Bounds on direct and indirect effects under treatment/mediator endogeneity and outcome attrition

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,761 characters · 5 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 329 chars of source]

Abstract: { Causal mediation analysis aims at disentangling a treatment effect into an indirect mechanism operating through an intermediate outcome or mediator, as well as the direct effect of the treatment on the outcome of interest. However, the evaluation of direct and indirect effects is frequently complicated by non-ignorable selection into the treatment and/or mediator, even after controlling for observables, as well as sample selection/outcome attrition. We propose a method for bounding direct and indirect effects in the presence of such complications using a method that is based on a sequence of linear programming problems. Considering inverse probability weighting by propensity scores, we compute the weights that would yield identification in the absence of complications and perturb them by an entropy parameter reflecting a specific amount of propensity score misspecification to set-identify the effects of interest. We apply our method to data from the National Longitudinal Survey of Youth 1979 to derive bounds on the explained and unexplained components of a gender wage gap decomposition that is likely prone to non-ignorable mediator selection and outcome attrition.}

{ Keywords: Causal mechanisms, direct effects, indirect effects, causal channels, mediation analysis, sample selection, bounds.}

{ JEL classification: C21. \quad }

{ {\scriptsize Addresses for correspondence: Martin Huber, University of Fribourg, Bd.\ de P\'{e}rolles 90, 1700 Fribourg, Switzerland; [email removed]. Luk\'{a}\v{s} Laff\'{e}rs, Department of Mathematics, Matej Bel University, Tajovsk\'{e}ho 40, SK97401 Bansk\'{a} Bystrica, Slovakia. Laff\'{e}rs acknowledges support provided by the Slovak Research and Development Agency under the contract No. APVV-17-0329 and VEGA-1/0692/20. }\thispagestyle{empty} }

{ \setcounter{footnote}{0} \setcounter{footnote}{0} \setcounter{page}{1} }

Introduction

Mediation analysis aims to decompose a treatment effect into an indirect causal mechanism operating through one or several intermediate variables, so-called mediators, as well as the direct effect,including any mechanisms not operating through the mediators of interest. For instance, early childhood interventions might affect labor market or health outcomes later in life through different mechanisms like the formation of cognitive or non-cognitive skills, see for instance \citeasnoun{HeckmanPintoSavelyev2013} and \citeasnoun{Keeleetal2015}. Furthermore, job seeker counseling may influence employment through assignment to training programs or other mechanisms in the counseling process, see \citeasnoun{HuberLechnerMellace2017}. Even with a randomly assigned treatment, direct and indirect effects are generally not identified by naively controlling for mediators, as this likely introduces selection bias, see \citeasnoun{RoGr92}. While much of the earlier work on mediation analysis assumed linear models and/or neglected selection issues, see \citeasnoun{Co57}, \citeasnoun{JuKe81}, and \citeasnoun{BaKe86}, more recent contributions discuss more general identification approaches and explicitly consider confounding. See for instance \citeasnoun{RoGr92}, \citeasnoun{Pearl01}, \citeasnoun{Robins2003}, \citeasnoun{PeSiva06}, \citeasnoun{VanderWeele09}, \citeasnoun{ImKeYa10}, \citeasnoun{Hong10}, \citeasnoun{AlbertNelson2011}, \citeasnoun{ImYa2011}, \citeasnoun{TchetgenTchetgenShpitser2011}, and \citeasnoun{VansteelandtBekaertLange2012}.

In most mediation studies, identification relies on a conditionally exogenous treatment and mediator given observed covariates and rules out non-ignorable outcome attrition or sample selection, i.e.\ that outcomes are only observed for a nonrandom subpopulation. This issue occurs for instance in wage regressions, where wages are only observed for the selective subgroup of employed individuals, see \citeasnoun{He76} and \citeasnoun{Heckman79}. For this reason, \citeasnoun{HuberSolovyeva2017} incorporate outcome attrition into mediation models, assuming conditional treatment and mediator exogeneity and tackling outcome attrition either by observed covariates (missing at random assumption, see e.g.\ \citeasnoun{Ru76b} and \citeasnoun{LittleRubin87}) or by instruments (if attrition is selective in unobservables). In many empirical problems, however, observed covariates might not be rich enough to convincingly control for treatment/mediator endogeneity and attrition bias while instruments that satisfy specific exclusion restrictions w.r.t.\ attrition (see for instance \citeasnoun{Daneve03} and \citeasnoun{Hu11b}) might not be available.

This paper provides a method for deriving bounds on direct and indirect effects when the treatment, the mediator and outcome attrition are likely selective even after controlling for observed covariates. Considering identification based on inverse probability weighting based on a combination of propensity scores, we compute the weights that would yield identification in the absence of complications (as provided in \citeasnoun{HuberSolovyeva2017}) and perturb them by an entropy parameter reflecting misspecification in the various propensity scores. Based on the framing the identification issue as an optimization problem to be solved by linear programming, we set-identify the mean potential outcomes and thus, the direct and indirect effects of interest.

Our contribution is related to further studies that used optimization, and in particular linear programming, to derive bounds on treatment effects under selection problems, see e.g.\ \citeasnoun{BaPe97}, \citeasnoun{manski2007partial}, \citeasnoun{honore2006bounds}, \citeasnoun{molinari2008partial}, \citeasnoun{freyberger2015identification}, \citeasnoun{laffers2017sensitivity}, \citeasnoun{laffers2019bounding}, among many others. Our paper is also related to the literature on sensitivity analysis in mediation analysis. \citeasnoun{VanderWeele2010}, for instance, provides a general formula for the bias of direct and indirect effects in the presence of an unobserved mediator-outcome confounder. By considering sensible values for differences in conditional mean outcomes across confounder values and for differences in the conditional mean of the confounder across treatment states, researches may investigate the sensitivity of the effects. \citeasnoun{ImKeYa10} propose a sensitivity check for parametric (both linear and nonlinear) mediation models based on specifying the correlation of unobserved terms in the mediation and outcome equations, assuming that the mediator-outcome confounders are not a function of the treatment. In contrast, \citeasnoun{TchetgenTchetgenShpitser2011} suggest a semiparametric procedure that allows for confounders of the mediator-outcome relation which are affected by the treatment based on specifying and calibrating the so-called selection bias function, which is agnostic about the dimension of unobserved confounders. See \citeasnoun{VanderWeeleChiba2014} and \citeasnoun{VansteelandtVanderWeele2012} for further selection bias functions.

As an alternative strategy, \citeasnoun{AlbertNelson2011} suggest considering the correlation of counterfactual values of post-treatment variables as sensitivity parameter. Finally, the paper that is the closest to our approach is \citeasnoun{HongQinYang2018}, which provides a method tailored to weighting estimators under the omission of both pre- and post-treatment confounders. The idea is that such confounders create a discrepancy between the correct weight an observation should obtain and the one actually used. The resulting bias can be represented by the covariance between the weight discrepancy and the outcome conditional on the treatment, which serves as base for conducting sensitivity analyses. Our approach is different to \citeasnoun{HongQinYang2018} in that we represent the discrepancy between the correct and observed weights using entropy parameters and, instead of deriving analytical formulas, rely on an optimization routine to obtain bounds on direct and indirect effects. This allows for a separate relaxation of the three main identification assumptions and thus may lead to a better understanding of the non-robustness of the results to violations of the various identification assumptions. We also note that the econometric setup of \citeasnoun{HuberSolovyeva2017} underlying our analysis invokes a different set of assumptions than \citeasnoun{HongQinYang2018} or any of the other previous methods, such that our approach permits investigating sensitivity also w.r.t.\ outcome attrition.\footnote{The studies mentioned and our own investigate the sensitivity of direct and indirect effects to prespecified deviations from the identifying assumptions. Alternatively, one may derive worst case bounds, which are based on the possibly most extreme forms of violations of specific assumptions, which typically implies a rather wide range of admissible effect values. See for instance \citeasnoun{Kaufmanetal05}, \citeasnoun{Caietal08}, \citeasnoun{Sj09}, and \citeasnoun{FlFl10}.}

We apply our method to data from the National Longitudinal Survey of Youth 1979, a panel study of young individuals in the U.S.\ aged 14 to 22 years in 1979. The specific sample considered has previously been analyzed by \citeasnoun{HuberSolovyeva2019} to decompose the gender gap in wages reported in the year 2000 into an indirect (or explained) component due to differences in mediators like education and occupation, as well as a direct (or unexplained) gender difference in wages not attributable to the observed mediators. While \citeasnoun{HuberSolovyeva2019} investigated the sensitivity of point estimation of explained and unexplained component under different identifying assumptions, our approach permits easing any of the conditional exogeneity assumptions on gender, the mediators, and selection into employment (as wages are only observed for working individuals) to derive bound son the parameters of interest. We find that the omission of confounders of the treatment and the mediators would potentially have the largest impact on the significance of the results. More specifically, the omission of a confounder that has the same predictive power as the first or second most important mediator entering the treatment propensity score would render all the effects insignificant. The results also show that in some specifications the choice of the link function in the estimation of probabilistic weights matter.

The remainder of this paper is organized as follows. Section (ref) introduces the variables as well as the direct and indirect effects of interest. Section (ref) restates the identifying assumptions of \citeasnoun{HuberSolovyeva2017}, under which the direct and indirect effects are point identified, and introduces the sensitivity analysis based on inverse probability weighting when relaxing these assumptions. Section (ref) presents an application to the decomposition of the U.S.\ gender wage gap using data from the National Longitudinal Survey of Youth 1979. Section (ref) concludes.

Variables and parameters of interest

Mediation analysis typically aims to disentangle the average treatment effect (ATE) of a binary treatment, denoted by $D$, on an outcome variable, denoted by $Y$, into a direct effect and an indirect effect operating through one or several mediators. We denote the latter by $M$, which is assumed to have bounded support and may be scalar or a vector of variables and contain discrete and/or continuous elements. For defining natural direct and indirect effects, we make use of the potential outcome framework, see for instance \citeasnoun{Rubin74}, which has been applied to causal mediation analysis for instance by \citeasnoun{TenHaveetal2007} and \citeasnoun{Albert2008}. Let to this end $M(d), Y(d,M(d'))$ denote the potential mediator state as a function of the treatment and potential outcome as a function of the treatment and the potential mediator, respectively, under treatments $d, d'$ $\in$ $\{0,1\}$. For each subject, only one potential outcome and mediator state, respectively, is observed, because the realized mediator and outcome values are $M=D\cdot M(1) + (1-D)\cdot M(0)$ and $Y=D\cdot Y(1,M(1)) + (1-D)\cdot Y(0,M(0))$.

The ATE, denoted by $\Delta$, is given by the total effect of the treatment operating through the direct or indirect mechanisms:

eqnarray[eqnarray omitted — 46 chars of source]

The (average) direct effect, denoted by $\theta (d)$, is characterized by the difference in mean potential outcomes under treatment and non-treatment when fixing the mediator at its potential value for $D=d$, which shuts down the indirect mechanism via $M$.

eqnarray[eqnarray omitted — 71 chars of source]

The (average) indirect effect, denoted by $\delta (d)$, is given by the difference in mean potential outcomes when exogenously varying the mediator to take its potential values under treatment and non-treatment, but keeping the treatment fixed at $D=d$ to shut down the direct effect.

eqnarray[eqnarray omitted — 71 chars of source]

\citeasnoun{RoGr92} and \citeasnoun{Robins2003} referred to these causal parameters as pure/total direct and indirect effects, \citeasnoun{FlFl09} as net and mechanism average treatment effects, and \citeasnoun{Pearl01} as natural direct and indirect effects, which is the denomination followed in the remained of this study.

The ATE is the sum of the natural direct and indirect effects defined upon opposite treatment states:

eqnarray[eqnarray omitted — 225 chars of source]

This follows from adding and subtracting either $E[Y(0,M(1))]$ or $E[Y(1,M(0))]$ in (ref).The notation $\theta(1),\theta(0)$ and $\delta(1),\delta(0)$ allows for effect heterogeneity as a function of the treatment state, i.e., the presence of interaction effects between the treatment and the mediator. For instance, the impact of a training ($M$) on employment ($Y$) might depend on whether a job seeker has received some form of counseling in the job search process ($D$). A different way to see this is that the direct effect of counseling ($D$) may depend on whether the job seeker attends a training ($M$).

Obviously, effects are not identified without invoking identifying assumptions. First, $Y(1,M(1))$ and $Y(0,M(0))$ are not observed for any subject at the same time, which constitutes the fundamental problem of causal inference. Second, neither $Y(1,M(0))$, nor $Y(0,M(1))$ is observed for any subject. Therefore, point identification of direct and indirect effects requires the treatment and the mediator to be exogenous at least conditional on observables, which appears, however, implausible in many empirical applications. Our sensitivity analysis outlined in Section (ref) relaxes such exogeneity conditions at the cost of giving up on point identification. This permits incorporating a vector of observed pre-treatment covariates, denoted by $X$, that may confound the causal relations between $D$ and $M$, $D$ and $Y$, and $M$ and $Y$. It is thus assumed that $X$ is insufficient to control for all sources of selection such that unobserved confounders render point identification of direct and indirect effects impossible, which appears plausible in many empirical contexts.

As a further complication to identification, our framework allows for considering outcome attrition/sample selection, implying that $Y$ is only observed for a non-random subpopulation. For instance, when investigating wage outcomes, as in \citeasnoun{Gr74}, the subpopulation of employed individuals for whom wages are observed might be positively selected in terms of unobservables like ability and motivation. As a further example, consider the effect of educational interventions on test scores, with scores being only observed for those participating in the test or reporting the results, see \citeasnoun{AnBeKr04}. We therefore introduce a binary selection indicator $S$, which indicates whether $Y$ is observed for a specific subject. While $S$ is allowed to be a function of $D$, $M$, and $X$, i.e.\ $S=S(D,M,X)$, it is assumed to neither be affected by nor to affect outcome $Y$. $S$ is therefore not a mediator, as selection per se does not causally influence the outcome, but might nevertheless create endogeneity bias when outcomes are only observed conditional on $S=1$.

Sensitivity analysis

The starting point for our sensitivity analysis is a set of assumptions provided in \citeasnoun{HuberSolovyeva2017}, which identifies direct and indirect effects by invoking conditional treatment and mediator exogeneity as well outcome attrition related to observed characteristics (known as missing at random assumption). Formally, the assumptions are as follows: \newline Assumption A1 (conditional independence of the treatment):\newline (a) $Y(d,m) \bot D | X=x$, (b) $M(d') \bot D | X=x$ for all $d,d' \in \{0,1\}$ and $m, x$ in the support of $M, X$.\newline Assumption A1 rules out unobservables jointly affecting the treatment on the one hand and the mediator and/or the outcome on the other hand conditional on $X$. In contrast, our sensitivity analysis permits that such unobserved confounders do exist. \newline Assumption A2 (conditional independence of the mediator):\newline $ Y(d,m) \bot M | D=d', X=x $ for all $d,d' \in \{0,1\}$ and $m,x$ in the support of $M,X$. \newline Assumption A2 rules out unobservables jointly affecting the mediator and the outcome conditional on $D$ and $X$. This only appears plausible if detailed information on possible confounders of the mediator-outcome relation is available in the data (even in experiments with random treatment assignment) and if post-treatment confounders of $M$ and $Y$ can be plausibly ruled out when controlling for $D$ and $X$. In contrast, our sensitivity analysis allows for unobserved confounders of the mediator-outcome relation. \newline Assumption A3 (conditional independence of selection):\newline $ Y \bot S | D=d, M=m, X=x $ for all $d \in \{0,1\}$ and $m,x$ in the support of $M,X$. \newline Assumption A3 rules out unobservables jointly affecting selection and the outcome conditional on $D,M,X$, such that outcomes are missing at random (MAR) in the denomination of \citeasnoun{Ru76b}, i.e.\ outcome attrition is selective w.r.t.\ observed characteristics only. In contrast, our sensitivity analysis permits outcome attrition to be selective w.r.t.\ unobservables. \newline Assumption A4 (common support):\newline (a) $\Pr(D=d| M=m, X=x)>0$ and (b) $\Pr(S=1| D=d, M=m, X=x)>0$ for all $d \in \{0,1\}$ and $m,x$ in the support of $M,X$.\newline Assumption A4 consists of two common support restrictions. The first requires the conditional probability to receive a specific treatment given $M,X$, henceforth referred to as propensity score, to be larger than zero for either treatment state. This also implies that $\Pr(D=d|X=x)>0$ and (by Bayes' theorem) that $\Pr(M=m| D=d, X=x)>0$, or in the case of $M$ being continuous, that the conditional density of $M$ given $D,X$ is larger than zero. Therefore, $M$ must not be deterministic in $D$ given $X$, as otherwise identification fails due to the lack of comparable units in terms of the mediator across treatment states. The second common support restriction requires that for any combination of $D,M,X$, the probability to be observed is larger than zero. Otherwise, the outcome is not observed for some specific combinations of these variables. Our sensitivity relies on the same set of common support assumptions, in order to make treated and non-treated subjects with observed and non-observed outcomes comparable in terms of observed characteristics.

As outlined in Theorem 1 of \citeasnoun{HuberSolovyeva2017}, Assumptions A1 to A4 permit identifying the mean potential outcomes based on inverse probability weighting (IPW) by

eqnarray[eqnarray omitted — 597 chars of source]

The direct and indirect effects of interest are obtained as differences between two out of the four mean potential outcomes. For notational ease, we henceforth denote the various propensity scores in (ref) by

eqnarray[eqnarray omitted — 111 chars of source]

This denomination is motivated by the fact that e.g.\ under A3, there are no confounders that jointly affect $S$ and $D$ conditional on $(M,X)$, $S$ and $M$ conditional on $(D,X)$ or $S$ and $X$ conditional on $(M,D).$ So under A3, $p^{A3}$ is the correct probability to be used for weighting in order to obtain mean potential outcomes. Conversely, if A3 does not hold, e.g. some important confounder is missing, then the correct weight differs from $p^{A3}$.

\tikzstyle{VertexStyle} = [shape = ellipse, minimum width = 6ex, draw]

\tikzstyle{EdgeStyle} = [->,>=stealth']

figure[figure omitted — 363 chars of source]

Figure (ref) illustrates our mediation framework with outcome attrition based on a directed acyclic graph, in which the arrows represent causal effects. It is worth noting that each of $D$, $M$, $S$, $X$, and $Y$ might be causally affected by further, unobserved variables not displayed in Figure (ref). As long as such unobservables do not jointly affect $D$ and $Y$, $D$ and $M$, $M$ and $Y$, or $S$ and $Y$ conditional on $X$, Assumptions A1 to A3 hold. In many empirical problems, however, some or all of A1 to A3 appear difficult defend. While, for instance, A1 holds by design in randomized experiments, it may appear less plausible in observational studies, in particular if the set of observed control variables is limited. Assumption A2 might seem unlikely in any identification design, in particular if there is a substantial time lag between $D$ and $M$ which makes mediator-outcome confounding more likely even when conditioning on pre-treatment covariates $X$. We therefore consider various relaxations of Assumptions A1 to A3.

Our approach modifies the IPW weights of the expressions in (ref) to investigate sensitivity and is thus related to \citeasnoun{HongQinYang2018}, who were the first to propose robustness checks in the context of IPW. However, our approach uses a different measure of discrepancies between the correct and observed weights than \citeasnoun{HongQinYang2018}, whose sensitivity check is based on the covariance between the weight discrepancy and the outcome conditional on the treatment. Also, our approach does not entail analytical formulae and thus relies on optimization routines. Even though this increases the computational burden, an advantage is that we are able to separately consider relaxations of the various identifying assumptions and thus gain insights on the sources of the potential non-robustness of the effects.

To see the intuition of our approach, suppose for the moment that there exist a scalar of a vector of unobserved confounders, denoted by $U$, that makes Assumption A3 fail. In this case, $p^{A3}$ is no longer the appropriate propensity score for identifying the mean potential outcomes. Our sensitivity analysis is based on specifying the magnitude by which $p^{A3}$ may deviate from the appropriate (but unidentified) propensity score that includes $U$ as conditioning variable. More formally, we consider the following entropy measure defined as the absolute difference between $p^{A3}$ and the appropriate propensity score for identification, $q^{A3} = \Pr(S=1|D,M,X,U)$:

eqnarray[eqnarray omitted — 97 chars of source]

This definition restricts the absolute error in the propensity score $p^{A3}$ due to omitting confounders $U$ to a multiple $ \epsilon^{A3}$ of the standard deviation of a random variable with a binomial distribution corresponding to that of $p^{A3}$. This definition also satisfies a symmetry property, so that the relaxation of the identifying assumption leads to the same set of values for $p^{A3}$ and for $(1-p^{A3}).$ The crucial task is to sensibly choose the value of $\epsilon^{A3}$, e.g.\ based on the richness of $X$, which determines the likely importance of omitted confounders $U$. In an analogous way, $q^{A1} = \Pr(D=1|X,U)$, $q^{A2} = \Pr(D=1|M,X,U)$ as well as $\epsilon^{A1}$ and $\epsilon^{A2}$ are to be defined. Suppose there were no unobserved confounders and that Assumptions A1, A2, A3 were satisfied. That would correspond to the situation with $\epsilon^{A1}=\epsilon^{A3}=\epsilon^{A3}=0.$ The greater is the particular entropy parameter, the larger is the permitted importance of unobserved confounders in the specific assumption.

Our sensitivity analysis provides bounds on estimates of the mean potential outcomes in ((ref)) for deviations of A1, A2, A3 when obeying the restrictions given by $\epsilon^{A1}$, $\epsilon^{A2}$, and $\epsilon^{A3}$ in (ref) as well as specific scaling constraints. The bounds on mean potential outcomes then translate into bounds on natural direct and indirect effects. Assume that we have available an i.i.d.\ sample of $(Y_i,D_i,M_i,X_i,S_i)$ for $n$ subjects, where $i$ $\in$ $\{1,2,....,n\}$ indexes a specific observation. For the sake of exposition, consider the estimation of $E[Y(1,M(1))]$. Under Assumptions A1, A3, and A4, this quantity can be estimated by $$\hat E[ Y(1,M(1))]= \frac{1}{c} \cdot \sum_{i = 1}^{n} Y_i \cdot D_i \cdot S_i \cdot \frac{1}{\hat p_i^{A1}}\cdot \frac{1}{ \hat p_i^{A3}},$$ where $\hat p_i^{A1}, \hat p_i^{A3}$ denote estimates of $p^{A1}, p^{A3}$ for observation $i$, which we obtain in our application presented below by logit regression, and $c$ denotes a normalizing constant, $c = \sum_{i = 1}^{n} \frac{D_i}{\hat p_i^{A1}}\cdot \frac{S_i}{ \hat p_i^{A3}}$.\footnote{Note that Assumption A2 is not required for the identification of $E[Y(d,M(d))]$ for $d \in \{0,1 \}$, but for $E[Y(d,M(1-d))].$} In the presence of confounders $U$ and a failure of assumptions A1 and/or A3, estimating $\hat E[ Y(1,M(1))]$ based on $\hat p_i^{A1}, \hat p_i^{A3}$ is generally inconsistent. To estimate the bounds on the mean potential outcome under such violations, the unknown population parameters $p^{A1}$ and $p^{A3}$ are replaced by their estimates $\hat p_i^{A1}$ and $\hat p_i^{A3}$ in (ref) to estimate the entropy measures $|q^{A1} - p^{A1}|$ and $|q^{A3} - p^{A3}|$, respectively.

Finding the bounds on $\hat E[ Y(1,M(1))]$ corresponds to solving the following optimization problem:

eqnarray[eqnarray omitted — 818 chars of source]

The equalities in (ref) are normalizing restrictions, while the expressions in (ref) relax the identifying assumptions and (ref) ensures that $q_i$ are proper probabilities. We note that as only those observations $i$ with $D_i = 1$ and $S_i = 1$ enter the calculations, we added a superscript $1$ to the entropy parameters $\epsilon$ (for $D_i=1$). This implies that one might pick different parameters for different treatment groups, if e.g.\ justified by contextual knowledge.

The optimization problem (ref) may be transformed into a computationally more convenient form. Let to this end $\omega_i^{A1} = 1/q_i^{A1}$ and $\omega_i^{A3} = 1/q_i^{A3}$. Then, an alternative representation is

eqnarray[eqnarray omitted — 1,095 chars of source]

For a fixed vector $\omega^{A1}$ or $\omega^{A3}$, this problem is a linear program. This suggests an algorithm that iteratively changes $\omega^{A1}$ and $\omega^{A2}$ until convergence using $\omega_i^{A1} = 1/\hat p_i^{A1}$ and $\omega_i^{A3} = 1/\hat p_i^{A3}$ as starting point. An important feature of our approach based on the optimization is that we impose no structure on the dependence of potentially omitted confounders and our outcome variable and exhaust all possibilities for the weights. This may in general lead to wider bounds and thus more prudent inference.

Bounds on $\hat E[Y(0,M(0))]$ (i.e.\ the estimate of $E[Y(0,M(0))]$) can be constructed in an analogous manner by using observations with $S_i=1$, $D_i=0$ and searching through $(1-q_i^{A1})$ instead. For bounds on $\hat E[Y(1,M(0))]$ and $\hat E[Y(0,M(1))]$, one has to optimize w.r.t.\ $q_i^{A1}$, $q_i^{A2}$ and $q_i^{A3}.$ All optimization problems are formally stated in Appendix (ref). After deriving upper and lower bounds on the mean potential outcomes, we may construct bounds on the effects of interest in the following way:

eqnarray[eqnarray omitted — 851 chars of source]

where subscripts $LB, UB$ stand for the lower and upper bounds of the respective estimated mean potential outcome and `$\hat{\ }$' implies that any of the causal effects are estimates rather than population parameters. The bounds on $\Delta$, $\theta(1)$, and $\theta(0)$ are sharp, because the observations that are used for calculating the two mean potential outcomes upon which the respective effect is defined are distinct. Consider for instance the lower bound on $\theta(1)$. In order to estimate $E[ Y(1,M(1))]_{LB}$, we use observations with $D_i=1$, while for $E[ Y(0,M(1))]_{UB}$ we only use observations with $D_i=0$. In contrast, the bounds for $\delta(1)$ and $\delta(0)$ are not necessarily sharp. As an example, consider $\delta(1)$. In order to estimate $E[ Y(1,M(1))]_{LB}$ and $E[ Y(1,M(0))]_{UB}$, we use observations with $D_i=1$ and the weights that entail $E[ Y(1,M(1))]_{LB}$ and $E[ Y(1,M(0))]_{UB}$ are not necessary the same.\footnote{ It is in principle possible to compute sharp bounds even for $\delta(1)$ and $\delta(0)$ at the cost of solving a more complex optimization problem. In such case, however, our heuristic algorithm of solving sequential linear programs cannot be used.}

An important question yet to be discussed is how to set the entropy parameters, which represent changes in the propensity scores due to omitting confounders (e.g.\ $\epsilon^{A1,1}$ and $\epsilon^{A3,1}$), in a meaningful way. Investigating the importance of observed confounders $X$ may provide some guidance for finding sensible values for the entropy parameters. Consider, for instance, a logistic regression for estimating $p_i^{A3}$:

eqnarray*[eqnarray* omitted — 167 chars of source]

where $\hat p_i^{A3}$ is obtained by maximum likelihood estimation. Now consider removing the most important predictor (of $S$) in $X$ and re-estimating the propensity score, denoted as $\hat p_{i,X1}^{A3}$, where $X1$ are the remaining covariates (without the most important predictor).

For each observation, one can then compute the entropy parameter when including and excluding the most important predictor in $X$, $\epsilon_{i,X1}^{A3} = \frac{|\hat p_{i,X1}^{A3} - \hat p_i^{A3}|}{ \sqrt{\hat p_i^{A3}(1-\hat p_i^{A3})}}$. One may ultimately pick the entropy parameter as average of $\epsilon_{i,X1}^{A3}$ in the subpopulation with $D_i=1$ and $S_i=1$:

eqnarray*[eqnarray* omitted — 129 chars of source]

This corresponds to the average change induced by omitting the most important predictor from $X$, which is used as a proxy for the importance of unobserved confounders $U$. There are different ways of measuring the importance of a predictor in a regression and natural choice seems to be the change in the deviance. The latter is the log-likelihood ratio statistic for testing the difference in the model fit when including and excluding a specific predictor, which is asymptotically chi-squared distributed. Similarly $\epsilon^{A3}_{Xj}$ and $\epsilon^{A3}_{Mj}$ would denote the average change in estimated probabilities that would omitting the $j$-th most important (measured the by deviance) from $X$ and $M$, respectively.\footnote{Another approach for setting the entropy parameter is to consider the change in estimated probabilities induced by a change of the link function, e.g.\ by picking the probit instead of logit function. Formally, $\epsilon^{A3,1}_{i,probit} = \frac{|\hat p_{i,probit}^{A3} - \hat p_i^{A3}|}{ \sqrt{\hat p_i^{A3}(1-\hat p_i^{A3})}}$, where $\hat p_{i,probit}^{A3}$ and $\hat p_i^{A3}$ correspond to the estimated probabilites under a probit and logit model, respectively. One may thus pick the entropy parameter as average $\epsilon^{A3,1}_{probit} = \sum_{i=1}^n \frac{D_i \cdot S_i \cdot \epsilon^{A3}_{i,probit} }{\sum_{i=1}^n D_i \cdot S_i }.$}

Application

We apply our method to data from the National Longitudinal Survey of Youth 1979 (NLSY79), a panel survey of young individuals who were aged 14 to 22 years in 1979, to decompose the gender wage gap in the year 2000 when respondents were 35 to 43 years old.\footnote{The NLSY79 data consists of three independent samples: a cross-sectional sample (6,111 subjects, or 48%) representing the non-institutionalized civilian youth; a supplemental sample (42%) oversampling civilian Hispanic, black, and economically disadvantaged nonblack/non-Hispanic young people; and a military sample (10%) comprised of youth serving in the military as of September 30, 1978 (\citeasnoun{NLSY792000}).} We use exactly the same sample definition as in \citeasnoun{HuberSolovyeva2019}, who consider five wage decomposition techniques with distinct identifying assumptions to investigate the sensitivity of point estimators of the indirect effect (or explained component) due to gender differences in mediators like education or occupation as well as the direct gender effect (or unexplained component). However, the consistency of these decomposition techniques relies on specific conditional exogeneity or instrumental variable assumptions w.r.t.\ to gender, the mediators, and the observability of wages, which are likely violated in practice. See also \citeasnoun{Huber2015} for a discussion of identification issues with in wage decompositions based on the causal mediation framework. In contrast, the approach suggested in this paper permits investigating the robustness of the direct and indirect effects of gender on wage under violations of conditional exogeneity.

The NLSY79 includes a rich set of individual characteristics, including socio-economic variables likes education, occupation, work experience, and a range of further labor market-relevant information. Our evaluation sample consists of 6,658 individuals (3,162 men and 3,496 women), after excluding 1,351 observations from the total NLSY79 sample in 2000 due to various data issues outlined in \citeasnoun{HuberSolovyeva2019}.\footnote{For instance, we excluded 502 persons reporting to work 1,000 hours or more in the past calendar year, but whose average hourly wages in the past calendar year were either missing or equal to zero. Furthermore, we dropped 54 working individuals with average hourly wages of less than \$1 in the past calendar year. We also excluded 608 observations with missing values in the mediators.} Treatment $D$ is a binary indicator for gender, taking the value zero for females and one for males. Outcome $Y$ corresponds to the log average hourly wage in the past calendar year reported in 2000. However, the wage outcome under full-time employment is not observed for everyone, as a non-negligible share (in particular among females) are in minor employment or not employed at all. This likely introduces sample selection issues, see \citeasnoun{Heckman79}. We therefore define the selection indicator $S$ to be one for individuals who indicated to have worked at least 1,000 hours in the past calendar year, as it is the case for 87% of males and 70% of females.

The vector of mediators $M$ includes individual variables reported in or constructed with reference to 1998 such that they arguably reflect decisions taken after birth and prior to the measurement of the outcome (i.e.\ on the causal path between $D$ and $Y$): marital status, years in marriage, the region and the number of years residing in that region, a dummy for living in an urban area (SMSA) and the number of years living in an urban area, education, dummies for the year of first employment, number of jobs ever had, tenure with the current employer (in weeks), industry and the number of years working in that industry, occupation and the number of years working in that occupation, whether employed in 1998 and total years of employment, a dummy for full-time employment and the share of full-time employment from 1994-98, total weeks of employment, the number of weeks unemployed and the number of weeks out of the labor force, and whether health problems prevented employment along with the history of health problems.

In the propensity scores, we control for a set of observed covariates $X$ arguably mostly determined at or prior to birth, namely race, religion, year of birth, birth order, parental place of birth (in the U.S. or abroad), and parental education. However, unobserved confounders causing non-ignorable selection into the treatment, mediators, and/or employment decision $S$ even after controlling for $X$ likely exist. For instance, risk preferences, attitudes towards competition and negotiations, and other socio-psychological factors, see e.g., \citeasnoun{Bertrand} and \citeasnoun{AzmatPet14}, are not available in our data. Such variables might possibly confound the mediator-outcome relationship. As a second example, selection into employment might be driven by innate ability or motivation, which likely also affect wages. Due to such endogeneity concerns, one should be cautious about deriving causal claims (e.g. about the amount of labor market discrimination) and policy recommendations based on wage gap decompositions, see the discussion in \citeasnoun{HuberSolovyeva2019}. In this context, our method is useful for assessing the sensitivity of the results to violations of some or all exogeneity conditions required for a causal interpretation of wage decompositions.

Table (ref) in Appendix (ref) provides descriptive statistics for the key variables of our analysis, namely means, mean differences across gender, and the respective p-values based on two-sample t-tests. The p-values suggest that females an males differ importantly in a range of variables like labor market experience, education, industry, and occupation. Our sensitivity analysis is based on assuming that omitted confounders in the various propensity score specifications behave similar to the first, second, or third important predictors among the pre-treatment covariates or mediators that enter the respective specification. For this reason, Table (ref) shows the three most important covariates and mediators in the different propensity score models presumably prone to confounding, where importance refers to the change in deviance as discussed at the end of Section (ref).

Tables \ (ref), (ref), (ref), and (ref) \ present the estimated effect bounds on $\theta(1)$, $\theta(0)$, $\delta(1)$, and $\delta(0)$, respectively, in squared brackets, when basing the entropy parameter on the respective first, second, or third most important predictors.\footnote{We note that in the decomposition literature, it is frequently the male wages that are considered as reference (or `fair') wages. This suggests considering $\theta(0)$ and $\delta(1)$ as unexplained and explained components, respectively. See \citeasnoun{Sloczynski2013} for an in-depth discussion of reference group choice in the potential outcome framework.} 95% confidence intervals are also reported in parentheses, based on subsampling with 500 replications and a subsample size of $\left\lfloor n^{0.7} \right\rfloor$.\footnote{We applied subsampling to the upper and lower bounds separately similar to \citeasnoun{laffers2017sensitivity} or \citeasnoun{demuynck2015bounding}. A computationally more expensive stepdown procedure of \citeasnoun{romano2010inference} may be used to control for the asymptotic coverage of the whole identified set. }We observe that the direct effects remain statistically significant at the 5% level under most relaxations and that potentially missing confounders would have the biggest impact via Assumption A2. Previous employment has considerable explanatory power for later labor market performance, as reflected in the relatively wide bounds in the column for the 2nd most important missing $M$ (employment in 1998) in Table (ref), where the lower 95% confidence bound goes below zero. Table (ref) displays the estimated bounds for the natural direct effect for $d=0$, when mediators are set to their potential values for women and the indirect mechanisms are shut down. For this effect, the choice of the link function seems important when considering Assumption A2 and the related regression for estimating $P(D=1|M,X).$ More precisely, if we would allow the difference between the correct and estimated weights to be of a similar magnitude as is the difference from using the probit instead of logit link function, then the upper bound on this effect may exceed $0.6$.

Concerning the natural indirect effects reported in Tables (ref) and (ref), the confidence interval on the effect estimate of $0.053$ includes the zero under most relaxations for males ($d=1$), while the lower confidence bound remains above zero under most relaxations for females ($d=0$). We also observe that confidence intervals are not necessarily symmetric around the estimated bounds and that the omission of the variable with the most predictive power (measured by the change in deviance) does not necessarily lead to the widest bounds. The latter is due to the fact that e.g.\ the strongest predictor of the treatment is not necessarily the strongest confounder, i.e.\ the predictor's association with the outcome is sufficiently weaker than that of other treatment predictors.

table[table omitted — 1,316 chars of source]
table[table omitted — 2,865 chars of source]
table[table omitted — 2,854 chars of source]
table[table omitted — 2,899 chars of source]
table[table omitted — 2,836 chars of source]

Conclusion

This paper proposed a sensitivity check for estimating natural direct and indirect effects in the presence of treatment and mediator endogeneity as well as selective attrition or missingness in outcomes. To this end, we considered identification based on inverse probability weighting using treatment and selection propensity scores and perturbed the respective propensity scores by an entropy parameter reflecting a specific amount of misspecification to set-identify the effects of interest. We demonstrated that this approach can be framed as a linear programming problem and discussed sensible choices of the entropy parameters based on the predictive power of the most important predictors in the propensity scores. Finally, we applied our method to data from the National Longitudinal Survey of Youth 1979 to derive bounds on the explained and unexplained components of a gender wage gap decomposition that is likely prone to non-ignorable mediator selection and sample selection in terms of the observability of the wage outcomes.