Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,438 characters · 8 sections · 81 citation commands
\sloppy
\thispagestyle{empty}
{\scriptsize We have benefited from comments of seminar participants in Neuch\^atel, D\"usseldorf, and Kassel, as well as conference participants at the European Causal Inference Meeting 2024 in Copenhagen. Addresses for correspondence: Martin Huber, University of Fribourg, Bd.\ de P\'{e}rolles 90, 1700 Fribourg, Switzerland; [email removed]. Kevin Kloiber, University of Munich, Ludwigstrasse 33, 80539 M\"unchen, Germany; [email removed]. Luk\'{a}\v{s} Laff\'{e}rs, Department of Mathematics, Matej Bel University, Tajovsk\'{e}ho 40, 974 01 Bansk\'{a} Bystrica, Slovakia; Department of Economics, Norwegian School of Economics, Helleveien 30, 5045 Bergen, Norway; [email removed]. Laff\'{e}rs acknowledges support provided by the Slovak Research and Development Agency under contracts VEGA 1/0398/23 and APVV-21-0360.}
\thispagestyle{empty} \setcounter{page}{1}
Causal mediation concerns the evaluation of the direct effect of a treatment on an outcome, as well as the indirect effect mediated by an intermediate outcome referred to as a mediator. For instance, when assessing the effect of two sequential training programs (such as job application training and an IT course) on employment, the direct effect corresponds to the effect of the first program, net of participation in the second program. Conversely, the evaluation of dynamic treatment effects pertains to assessing the effects of specific sequences of treatment and mediator values, such as the impact of participating in two sequential training programs versus not participating in any program. Effect identification in numerous studies within the fields of causal mediation and dynamic treatment evaluation relies on a sequential selection-on-observables or ignorability assumption, see for instance Ro86, RoHeBr00, Lech09, FlFl09, ImKeYa10, Hong10, TchetgenTchetgenShpitser2011, and Huber2012. This assumption postulates that, conditional on observed covariates, the treatment and the mediator are exogenous, such that their causal effects are not confounded by unobservables.
Related identification issues also arise when assessing causal effects in the context of sample selection or post-treatment outcome attrition, which implies that the outcome is only observed for a selectively non-random subpopulation. For example, wages are only observed for employed individuals. In such cases, sequential ignorability means that both the treatment and a selection indicator for the observability of the outcome (rather than a mediator) are exogenous conditional on observed characteristics. This allows for the identification of the total effect of the treatment on the outcome, as discussed, for instance, in Negi2020 and bia2023double. Whether the observed covariates plausibly satisfy the assumption of sequential ignorability is conventionally motivated based on theory, intuition, domain knowledge, or previous empirical findings. However, it is typically also subject to controversy, as unobserved confounders can rarely be completely ruled out in empirical applications relying on observational data.
This study introduces a test for conditions that imply sequential ignorability, facilitating the verification of identification in observational data. Testing relies on two types of observed variables for each of the treatment and mediator (or selection indicator): covariates to be controlled for and additional variables that are presumed to be separate instruments for the treatment and the mediator. The testable condition arises within a specific causal structure that rules out certain forms of reverse causality, such as those from the outcome to the mediator, treatment, covariates, or the suspected instruments and imposes that the suspected instruments have a (first stage) association with the treatment or the mediator on other observed variables. If, under these circumstances, the supposed instrument for the treatment is conditionally independent of the outcome, given the treatment and the covariates, and the supposed instrument for the mediator is conditionally independent of the outcome, given the treatment, mediator, and covariates, then the following conditions hold: (A) The instruments are valid, meaning they do not directly influence the outcome, except through their impact on the treatment or mediator, and are not associated with unobservable factors affecting the outcome, given the observed variables. (B) The treatment and the mediator are exogenous, implying that they are not associated with unobservable factors that affect the outcome, conditional on the observed variables.
Therefore, the testable conditional independence assumption implies that the treatment and the mediator satisfy a sequential ignorability assumption, which permits identifying dynamic treatment effects or the controlled direct effect discussed, for instance, in Pearl01. The controlled direct effect corresponds to the net effect of the treatment when holding the mediator fixed at a specific value (e.g., one). In sample selection models, this assumption allows for the evaluation of the total treatment effect. Furthermore, if it additionally holds that the supposed instrument for the treatment is conditionally independent of the mediator given the treatment and the covariates, the effect of the treatment on the mediator is identifiable. This is also a precondition for the evaluation of natural direct and indirect effects as discussed in RoGr92 and Pearl01, albeit their identification rests on somewhat stronger conditions than those implied by our testable implications. Natural effects are defined based on setting the mediator to its potential value under a specific treatment (rather than fixing it at a constant value like one, as is the case when evaluating the controlled direct effect).
Our study is, to the best of our knowledge, the first to propose an identification test in the contexts of causal mediation, dynamic treatment effects, or sample selection, building on the work of huberkueck2022 for testing the identification of the (total) effect of a single treatment variable. We extend the machine learning-based testing approach presented therein to enable testing conditional independence assumptions involving multiple variables, such as the treatment, mediator, and selection indicator. Additionally, our extension allows for the consideration of multivalued instruments for the treatment, mediator, or selection indicator, as opposed to binary instruments. The utilization of machine learning methods offers the advantage of data-driven control for important observed confounders during testing, which is particularly valuable in high-dimensional contexts where numerous potential control variables are available. More concisely, we consider a test that is based on the expectation of squared differences of regression functions including and excluding the respective instruments. The squared difference satisfies the Neyman1959-orthogonality condition, implying that we may account for covariates in the regression functions by machine learning without compromising on desirable asymptotic properties, given that specific regularity conditions hold, see doubleML.
We investigate the finite sample performance of our test in a simulation study and find that it performs very decently in terms of empirical size and power in our simulation designs when the sample size consists of several thousand observations (or more). Furthermore, we apply our method to large administrative data from Slovakia to test the identification of the effects of sequential labor market programs for jobseekers on employment. We use the availability in public employment service centers as instruments and also control for a rich set of jobseeker characteristics. Our findings indicate that our testable implications are not rejected for the specific sequences of programs tested, namely a graduate practice program followed by employment incentives (consisting of hiring incentives and subsidized employment).
Our paper contributes to a growing literature on testing identifying assumptions. For instance, deLunaJohansson2012 and BlackJooLaLondeSmithTaylor2015 consider the same testable implication as huberkueck2022 to assess the conditional exogeneity of a treatment based on a valid instrument, rather than jointly testing the treatment exogeneity and instrument validity.\ peters2015causal employ instruments to learn in a data-driven way which variables are treatments, in the sense that they directly influence an outcome of interest, under the assumption that they satisfy conditional exogeneity. The approach utilizes instruments in a way that enables the rejection of treatments violating conditional exogeneity, thereby providing the power to detect identification failures. angrist2015wanna test an analogous implication as in deLunaJohansson2012, BlackJooLaLondeSmithTaylor2015, and huberkueck2022, but when the treatment is unconfounded, in order to test the validity of an instrument within the framework of the regression discontinuity design (RDD). In a sharp RDD, the treatment is a deterministic function of a cutoff in a running variable and, therefore, satisfies conditional exogeneity by design. This allows for testing whether the running variable is a valid instrument, i.e., whether it is not associated with the outcome conditional on the treatment. If this holds true, causal effects can also be identified away from the cutoff, which is otherwise not feasible due to the lack of common support in the treatment across different values of the running variable.
The remainder of this study is organized as follows. Section (ref) presents the identifying assumptions and testable implications in scenarios involving only pre-treatment covariates as control variables. Section (ref) considers a modified causal framework that allows the first instrument (for the treatment) to directly influence the second instrument (for the mediator), which requires controlling for the second instrument in a specific way when testing. Section (ref) discusses a setup with dynamic confounding, which requires controlling for post-treatment covariates in addition to pre-treatment covariates when testing. Section (ref) outlines the machine learning-based test, which permits controlling for high-dimensional covariates in a data-driven manner. Section (ref) presents a simulation study that investigates the finite sample performance of our test. Section (ref) provides an empirical application to Slovak labor market data. Section (ref) concludes. The proofs of the identification results are provided in the Appendix.
In causal mediation analysis, the objective is to dissect the overall causal impact of a treatment variable \( D \) on an outcome variable \( Y \) into two distinct components: the direct effect of the treatment on the outcome and the indirect effect that operates through the mediator variable \( M \). A conceptually related framework is dynamic treatment evaluation, where \( D \) signifies an initial treatment and \( M \) a subsequent one, aiming to evaluate the effectiveness of different treatment sequences involving \( D \) and \( M \). In both mediation and dynamic treatment models, the variables \( D \), \( M \), and \( Y \) can exhibit either discrete or continuous distributions. However, in models featuring sample selection or post-treatment outcome attrition, \( M \) represents a binary selection indicator determining the observability of outcome \( Y \), while \( D \) represents the treatment. Despite this distinction, the identification challenges encountered in sample selection models are related to those in dynamic treatment models.
To formalize the assumptions required for identifying causal effects in mediation, dynamic treatment, and sample selection models, we make us of the potential outcome framework (e.g., Neyman23 and Rubin74) and denote random variables by capital letters and their specific values by lowercase letters. Specifically, we denote by $M(d)$ the potential mediator under treatment value $d \in \mathcal{D}$, where $\mathcal{D}$ is the support of the treatment. $Y(d,m)$ denotes the potential outcomes as a function of both the treatment and some mediator value $m \in \mathcal{M}$, where $\mathcal{M}$ is the support of the mediator. Our definition of $Y(d,m)$ and $M(d)$ as a function of an individual's treatment status $D=d$ and mediator value $M=m$ assumes that (i) an individual's potential outcomes are not affected by the treatment or mediator status of others, and (ii) there are no multiple versions of any treatment level $d$ and mediator level $d$ across individuals. This is known as Stable Unit Treatment Value Assumption (SUTVA), see e.g.\ Rubin80 and Cox58. Furthermore, we denote observed pre-treatment covariates as $X$, one or more observed instrumental variables for the treatment $D$ as $Z_1$, and one or more observed instrumental variables for the mediator $M$ as $Z_2$. The properties of these instrumental variables are yet to be established. Finally, let $\mathcal{X}, \mathcal{Z}_1, \mathcal{Z}_2$, and $\mathcal{Y}$ denote the support of $X, Z_1, Z_2$, and $Y$.
We subsequently discuss a range of identifying assumptions for causal mediation analysis, which permit identifying direct and indirect effects conditional on the covariates and testing identification based on specific assumptions on the instruments. Our first assumption establishes a specific causal structure between variables within our framework, positing that only certain variables exert a causal influence on others. We formalise this causal structure using the previously mentioned potential outcome notation, by applying the latter also to other variables than the outcome and the mediator. To this end, let \(A(b)\) and \(A(b,c,...)\) correspond to the potential value of variable \(A\) when setting variable \(B\) to \(b\), or variables \(B\), \(C\),... to \(b\), \(c\),..., respectively. Moreover, we assert that if there exists a causal relation between two variables (potentially conditional on other variables), then a statistical dependence necessarily exists between them, aligning with the principle of causal faithfulness.
The first line of Assumption (ref) rules out a reverse causal effect of outcome $Y$ on $D$, $X$, $M$, $Z_1$, or $Z_2$. In addition, the treatment $D$ must not casually affect $X$, $Z_1$, the mediator $M$ must not causally affect $X$, $D$, $Z_1$, $Z_2$, while $X$ might affect $D$, $M$, $Z_1$, $Z_2$ or $Y$, $Z_1$ might affect $D$. $Z_2$ might affect $M$. The second line of Assumption (ref) establishes causal faithfulness, which ensures that only variables that are d-separated, i.e., not linked via any causal paths, are statistically independent or conditionally independent, see e.g. Pearl00. To be more precise, we employ the d-separation criterion of pearl1988probabilistic, which relies on blocking causal paths between variables. A path between two (sets of) variables $A$ and $B$ is blocked when conditioning on a (set of) control variable(s) $C$ if
According to the d-separation criterion, $A$ and $B$ are d-separated when conditioning on a set of control variables $C$ if, and only if, $C$ blocks every path between $A$ and $B$.\ d-separation is sufficient for the (conditional) independence of two variables, providing a vital component for the proof of our Theorems. Causal faithfulness imposes that d-separation is also a necessary condition, such that two variables are statistically independent if and only if d-separation holds. A scenario in which faithfulness fails is that one variable affects another one via several causal paths (or mechanisms) which exactly cancel out such that the variables are independent, see e.g.\ the discussion in spirtes2000causation.
Next, we introduce two common support assumptions which are required for nonparametric identification and testing. To this end, let $f(A=a|B=b)$ denote the conditional density of variable $A$ given $B$ at values $A=a$ and $B=b$. We note that if $A$ is discrete rather than continuous, then $f(A=a|B=b)$ is a conditional probability rather than a density.
Assumption (ref) requires that conditional on $M$ and $X$, their exist observations with any values of $D$ and $Z_1$ that occur in the total population. Likewise, Assumption (ref) implies that conditional on $D$ and $X$, their exist observations with any values of $M$ and $Z_2$ that occur in the total population. In case of a violation of these common support conditions, effects can only be evaluated and/or tested over a subset of the support of $X$, $D$, and/or $M$.
Our next assumptions impose that the instruments are relevant in the sense that the instrument for treatment $Z_1$ is statistically associated with the treatment $D$ conditional covariates $X$, while the instrument for the mediator $Z_2$ is associated with the mediator $M$, conditional on the treatment $D$ and covariates $X$.
We note that $\not\!\perp\!\!\!\perp$ denotes statistical dependence. Assumptions (ref) and (ref) can be directly tested in the data by investigating the conditional associations of $D$ and $Z_1$ as well as $M$ and $Z_2$. If these assumptions do not hold, the testing approach lacks power in detecting violations of the identifying assumptions, as discussed in more detail in huberkueck2022.
The next assumptions invoke the conditional independence of the treatment on the one hand and the potential outcomes and potential mediators on the other hand conditional on the covariates, with ${\perp\!\!\!\perp}$ denoting statistical independence. These kind of assumptions are known as treatment exogeneity, selection-on-observables, or unconfoundedness, as for instance discussed in Im04 and ImWo08. The imply that given the covariates, the treatment is as good as random when aiming to assess its effect on the outcome or the mediator, respectively. \setcounter{assumption}{6}
Assumptions (ref) and (ref) require that conditional on covariates $X$, there exist no confounders jointly affecting the treatment $D$ on the one hand and the outcome $Y$ or the mediator $M$ on the other hand. This permits identifying the causal effect of $D$ on $M$ or $Y$ when controlling for $X$.
Next, we impose a related conditional independence assumption also w.r.t.\ to the mediator, requiring that the latter is conditionally independent of the potential outcomes conditional on the treatment and the pre-treatment covariates.
Assumption (ref) requires that conditional on $D$ and $X$, there exist no confounders jointly affecting the mediator $M$ and the outcome $Y$. This permits identifying the causal effect of $M$ on $Y$ when controlling for $D$ and $X$.
Our next assumptions concern the validity of the first instrument $Z_1$ and require that the latter is conditionally independent of the potential outcomes and the potential mediators given the covariates $X$.
\setcounter{assumption}{8}
Assumptions (ref) and (ref) rule out confounders jointly affecting $Z_1$ on the one hand and the outcome $Y$ or the mediator $M$ on the other hand when controlling for $X$, which is similar to Assumptions (ref) and (ref), which imposed such an exogeneity assumption w.r.t.\ the treatment. In addition, Assumptions (ref) and (ref) require that conditional on $X$, instrument $Z_1$ does not directly affect $M$ or $Y$ other than through $D$. This exclusion restriction implies that $M(d,z_1)=M(d)$ and $Y(d,m,z_1)=Y(d,m)$ for any value $z_1$ of instrument $Z_1$, otherwise the respective conditional independence assumptions are violated. Assumptions (ref) and (ref) are not sufficient for the nonparametric identification of the causal effects of $D$ on $Y$ or $M$ based on the instrument $Z_1$, which would require further assumptions like treatment monotonicity as discussed in Imbens+94 and Angrist+96. However, Assumptions (ref) and (ref) are useful for testing the identification of the causal effects of $D$ on $M$ and $Y$ as outlined further below.
Lastly, we also invoke an instrument validity assumption on the second instrument $Z_2$, requiring it to be conditionally independent of the potential outcomes conditional on $D$ and $X$.
Assumption (ref) rules out confounders jointly affecting $Z_2$ and $Y$ conditional on $D$ and $X$, and that $Z_2$ does not directly affect $Y$ (other than through $M$) such that the exclusion restriction $Y(d,m,z_2)=Y(d,m)$ holds for any value $z_2$ of $Z_2$. This assumption is useful for testing the identification of the causal effects of $M$ on $Y$.
If one is willing to impose the causal structure and faithfulness postulated in Assumption (ref) along with the (testable) conditional dependencies of the respective instruments and the treatment or the mediator as invoked in Assumptions (ref) and Assumption (ref), then the joint satisfaction of the conditional independence assumptions on the treatment, mediators, and instruments imposed in Assumptions (ref), (ref), (ref), (ref), (ref), and (ref) implies the following testable implications, which under the common support assumptions (ref) and (ref) are verifiable for all values in the support of the respective conditioning set:
Furthermore, conditional on Assumptions (ref), (ref), (ref), the joint satisfaction of the testable implications ((ref)), ((ref)), ((ref)) is not only implied by, but also imply the joint satisfaction of the conditional independence assumptions (ref), (ref), (ref), (ref), (ref), (ref), as formalized in the following theorem, which is proven in the appendix.
In practice, we may use the result in Theorem (ref) to test the satisfaction of the conditional independence assumptions for factual values of mediators and outcomes, i.e., for the potential outcomes $Y(d,m)$ and potential mediators $M(d)$ among subjects actually receiving $D=d$ and $M=d$, but not for counterfactual values $d'\neq d$ or $m' \neq m$ for those subjects with $D=d$ and $M=d$. If violations of the conditional independence assumptions (ref), (ref), (ref), (ref), (ref), (ref) exclusively concern counterfactual outcomes and mediators, but never (f)actual outcomes or mediators, testing cannot detect such violations. Therefore, our approach permits testing necessary, but not sufficient conditions for the identification of various causal effects discussed below. From a practical perspective, however, it seems unlikely that violations exclusively occur among counterfactual, but not among factual outcomes and mediators, because this would imply very particular, not to say implausible modeling constraints. Therefore, we would expect our test to have power in empirical applications, as violations of conditional independence should typically not affect only counterfactual outcomes or mediators, but also (at least some) factual ones.
We note that if the conditions ((ref)),((ref)),((ref)) hold for both factual and counterfactual mediators or outcomes, the following causal effects on the distributions of the mediator or the outcome are identified:
In causal mediation analysis, the controlled direct effect, which is based on forcing or prescribing the mediator to take a specific value $M=m$, is typically not the only causal parameter of interest. A large literature also focuses on natural direct and indirect effects defined in terms of potential (rather than prescribed) mediator states, as for instance discussed in RoGr92, Pearl01, ImKeYa10, and Huber2012. For instance, the average natural direct treatment effect of $d\neq d'$ on $Y$ conditional the potential mediator under treatment value $d$, $M(d)$, corresponds to $E[Y(d,M(d)) - Y(d',M(d))]$, while the average natural indirect effect of $D$ on $Y$ via $M$ when fixing the treatment at $D=d'$ is $E[Y(d,M(d)) - Y(d,M(d'))]$.
It is worth noting that the identification of natural effects not only hinges on the satisfaction of Assumptions (ref), (ref), and (ref) for both factual and counterfactual outcomes, but even on additional requirements. One additional condition that yields identification and has been suggested in Pearl01 is invoking conditional independence of potential mediators and potential outcomes across treatment states $d\neq d'$:
It is obvious that this conditional independence is an inherently counterfactual assumption, as either only $d$ or $d'$ can be observed for any subject in the sample. For this reason, we cannot test this assumption in the data. Furthermore, it is worth noting that the conditional independence in (ref) as well as Assumptions (ref) and (ref) are implied by the following, joint conditional independence assumption of ImKeYa10, which they impose in addition to Assumption (ref) for the identification of natural direct and indirect effects:
The conditional independence assumption in expression (ref) implies that the joint distribution of potential outcomes and potential mediators is conditionally independent of treatment assignment. In contrast, Assumptions (ref) and (ref) only impose conditional independence w.r.t.\ the marginal distributions of the potential outcomes and potential mediators. Yet, we argue that testing Assumptions (ref) and (ref) for actual outcomes typically also has nontrivial power against violations of the counterfactual condition (ref). The reason is that only under quite specific mediator and outcome models, it can be the case that Assumptions (ref) and (ref) always hold for factual mediators and outcomes, while we have at the same time that the conditional independence (ref) is violated for certain counterfactual outcomes and mediators. For instance, if the treatment is randomly assigned given $X$, then both the joint conditional independence assumption (ref) as well as Assumptions (ref) and (ref) on the marginal distributions are satisfied. However, if conditional randomization of $D$ fails, then we would suspect violations of all assumptions (ref), (ref), (ref), and even so for at least some factual outcomes and mediators.
Figure (ref) presents a causal model in which the identifying assumptions of the various causal effects discussed above as well as the testable implications given in Theorem (ref) are satisfied, based on a causal graph in which causal relations between variables are represented by arrows2, see e.g.\ Pearl00. Treatment $D$ affects outcome $Y$ both directly and via the mediator $M$ and is conditionally independent of potential mediators and outcomes, because no unobserved variables, whose effects are depicted by dashed lines (to indicate their non-observability), jointly affect $D$ and the post-treatment variables $M$ and $Y$ conditional on covariates $X$. Analogously, there are no unobservables jointly affecting the mediator $M$ and the outcome $Y$ conditional on $X$ and $D$. Furthermore, there are no confounders of the instrument for the treatment, $Z_1$, and the $M$ or $Y$ conditional on $X$, and $Z_1$ does not directly affect $M$ or $Y$ other than through $D$. Moreover, there are no confounders of the instrument for the mediator, $Z_2$, and $Y$ conditional on $X$ and $D$, and $Z_2$ does not directly affect $Y$ other than through $M$.
It is worth noting that if $Z_2$ has a direct causal effect on $M$ as in Figure (ref), then the satisfaction of the identifying assumptions underlying Theorem (ref) requires that $Z_1$ and $Z_2$ are conditionally independent of each other given $D$ and $X$. To see this, note that if e.g.\ $Z_1$ has a direct effect on $Z_2$ conditional on $D$ and $X$ and $Z_2$ has a direct effect on $M$, $Z_1$ is a confounder that jointly affects treatment $D$ (by Assumption (ref)) on the one hand and $M$ and $Y$ on the other hand. Related issues also occur if unobserved confounders affect both $Z_1$ and $Z_2$ and $Z_2$ directly affects $M$. The left graph of Figure (ref) provides such a causal model in which Assumptions (ref) and (ref) are violated.
If, on the contrary, $Z_2$ is associated with $M$ solely through unobserved confounders of both $Z_2$ and $M$ (rather than through a direct effect), then all identifying assumptions in Theorem (ref) hold even if $Z_1$ has a direct causal effect on $Z_2$ given $D$ and $X$, and/or unobserved confounders affect both $Z_1$ and $Z_2$, as illustrated in the right graph of Figure (ref). The rationale behind this is that $Z_2$ does not lie on any causal pathway through which the treatment $D$ influences the mediator $M$, as $Z_2$ does not affect $M$. Here, adjusting for $Z_2$ would even be harmful as it would introduce so-called M-bias, a specific form of collider or selection bias, see e.g. Pearl00. The reason is that by conditioning on $Z_2$, one introduces a spurious association between the treatment $D$ and the unobservable $U_2$ which also affects $M$ (and $Y$ via $M$). Such a spurious association comes from the fact that $D$ either directly affects the collider variable $Z_2$ or is associated with the latter through $U_1$ (which both affects $Z_2$ and $D$ via $Z_1$).
In causal models in which the first instrument $Z_1$ affects the second instrument $Z_2$ and $Z_2$ affects the mediator $M$, controlling for $Z_2$ is required to block any causal effects of $Z_1$ on $M$ or $Y$ that do not operate via the treatment $D$, conditional on covariates $X$. This scenario is illustrated in the left graph of Figure (ref), where controlling for $Z_2$ does not introduce collider bias, very much in contrast to the scenario depicted in the right graph of Figure (ref). Consequently, there may exist an alternative testing approach for identification that includes $Z_2$ in some of the conditioning sets when verifying the conditional independence of $Z_1$ and $Y$. For this reason, we subsequently modify the conditional independence assumptions imposed on the treatment and the first instrument to hold when controlling for both $X$ and $Z_2$, rather than $X$ alone. To this end, we replace Assumptions (ref), (ref), (ref), and (ref) by the following assumptions:
\setcounter{assumption}{6}
\setcounter{assumption}{8}
This adjustment in the identifying assumptions yields the following testable implications instead of the previous ones (ref) and (ref), while the previous implication (ref) remains unaffected:
We note that there exists even a further testable implication, which follows from the fact that any causal association of $Z_1$ on $Y$ via $Z_2$ necessarily operates through $M$:
This testable implication ((ref)), however, does not add any new information and thus, no additional power in testing when compared to the previous testable implications. This is formalized in the following lemma, which states that given Assumption (ref) and implications ((ref)), ((ref)), ((ref)) implies ((ref)) and vice versa:
The following theorem formalizes that our modified set of identifying assumptions, which involve conditioning on $Z_2$, implies and is implied by the respective testable implications:
The proofs of Lemma (ref) and Theorem (ref) are provided in the appendix.
It is important to note even though the testable implications involve conditioning on $Z_2$, this does not automatically imply that controlling for $Z_2$ is also appropriate for the identification of causal effects. In fact, if the testable implications in Theorem (ref) hold and $Z_2$ is a confounder of $D$ and $M$, or unobserved confounders affect both $D$ and $Z_2$ while $Z_2$ affects $M$, then controlling for $Z_2$ (along with $X$) is required when evaluating the causal effect of $D$ on $M$ or $Y$. However, if $D$ affects $Z_2$ and no unobserved confounders affect $D$ and $Z_2$, then controlling for $Z_2$ is only required for testing, but not for evaluating the causal effects of $D$. In fact, conditioning on $Z_2$ would block any impact of $D$ on $M$ and $Y$ that operates via $Z_2$, such that only a partial effect could be identified.
In many empirical applications, it may not hold that pre-treatment covariates are sufficiently informative to fully control for confounders jointly affecting the mediator $M$ and outcome $Y$, implying a violation of the conditional independence of the mediator as imposed by Assumption (ref). This issue appears particularly relevant if the mediator is measured at a substantially later point in time than the treatment. Just as the control variables of the treatment are typically measured shortly prior to treatment assignment, it then seems reasonable to also control for possible confounders of the mediator-outcome relation shortly prior to selection into the mediator. In such a scenario, several control variables for the mediator might be affected by the treatment. For instance, when assessing the treatment effect of education ($D$) on earnings ($Y$) when considering work experience as mediator ($M$), post-treatment covariates like health or mental well-being ($W$) might affect both $M$ and $Y$. For this reason, the subsequent discussion considers the case that the treatment can influence observed post-treatment confounders of the mediator-outcome relation to be controlled for, which are denoted by $W$, while $\mathcal{W}$ denotes their support. Figure (ref) depicts a causal graph in which both pre-treatment covariates $X$ and post-treatment covariates $W$ are present. $W$ may be affected by treatment $D$, pre-treatment covariates $X$, and the first instrument $Z_1$, and may itself affect mediator $M$, outcome $Y$, and the second instrument $Z_2$.
For our testing approach, this implies that the post-treatment covariates $W$ now are required to enter the conditioning set when imposing the conditional independence assumptions with respect to the second instrument and the mediator. To ensure that the test has power, there must be a first-stage association between instrument $Z_2$ and treatment $M$ when also controlling for $W$ in addition to $D$ and $X$. We also rule out that $M$ has a direct effect on $W$, based on the fact that $W$ are pre-mediator confounders and thus cannot be causally affected by the mediator. For these reasons, we replace Assumptions (ref), (ref), (ref), and (ref) in Section (ref) by the following four assumptions.
\setcounter{assumption}{1}
\setcounter{assumption}{5}
\setcounter{assumption}{7}
\setcounter{assumption}{9}
Under these modifications of also including $W$ in the conditioning set, the testable implication (ref) is substituted by the following implication, which states that the second instrument is conditionally independent of the outcome given the treatment, the mediator, as well as the pre- and post-treatment covariates:
Similar to the previously analyzed causal scenarios, our revised set of identifying assumptions, which includes conditioning on $W$, implies and is implied by a corresponding set of testable implications. This result is formally stated in the following theorem.
The proof is provided in the appendix.
We note that if the conditional independence assumptions (ref), (ref), (ref), (ref), (ref), and (ref) hold for both factual and counterfactual mediators or outcomes, the effect of treatment $D$ on outcome $Y$ and mediator $M$ (by Assumptions (ref) and (ref)), the effect of the mediator $M$ on outcome $Y$ (by Assumption (ref)), and the dynamic effect of specific sequences of $D$ and $M$, including the controlled direct effect, (by Assumptions (ref) and (ref)) are identified, as well as the effect of treatment $D$ on outcome $Y$ in sample selection models, where $M$ represents a binary indicator for the observability of $Y$ (by Assumptions (ref) and (ref)). See, for instance, LechnerMiquel2010, bia2023double, and huber2023causal for discussions of identification in dynamic treatment and sample selection models with post-treatment control variables. Concerning natural direct and indirect effects, identification is not attained under post-treatment confounders even when imposing the joint conditional independence assumption in expression (ref) along with Assumption (ref) and analogous conditional independence assumption on $W$. In contrast to the case when all confounders are pre-treatment, identification under post-treatment confounders is infeasible without specific parametric conditions. See, for instance, the discussions in AvinShpitserPearl2005, Robins2003, who imposes the absence of treatment-mediator interaction effects on the outcome to obtain identification, and ImYa2011, who assume that any treatment-mediator interaction effects are homogeneous across subjects.
This section introduces our testing approach, focusing primarily on the assumptions and testable implications discussed in Section (ref). However, we note that the causal scenarios outlined in Sections (ref) and (ref) can also be accommodated by appropriately adjusting the conditioning sets during testing. Moreover, instead of verifying complete statistical independence as stipulated in the testable implications (ref), (ref), and (ref), we test the conditional mean independence of the outcome or mediator and the instruments, which is sufficient for identifying (conditional or unconditional) average causal effects or the treatment and the mediator. To formalize the testable implications in terms of conditional mean independence, we denote by \( \mu_B(a) = E(B|A=a) \) the conditional mean of random variable \( B \), where \( A \) represents a set of conditioning variables taking values \( a \). Then, testing whether implications (ref), (ref), (ref) hold on average involves verifying the following null hypothesis.
The intuition behind hypothesis (ref) is that if \( Y \) and \( M \) are conditionally mean independent of the instruments, then whether the instruments are included or excluded from the respective conditioning sets should not alter the conditional means of the outcome and mediator, respectively. Following huberkueck2022, our test relies on a quadratic formulation of the null hypothesis, which is a common approach in specification tests for nonparametric regression, see e.g. Ra97, RaHaLi06, hong1995consistent and wooldridge1992test. To this end, let us denote by
which implies the null hypothesis
We note that testing based on $\theta$ can accommodate both discrete and continuous variables for the instruments, treatment, mediator, outcome, and covariates. In contrast, huberkueck2022 focused on a binary instrument, testing the squared difference in conditional mean outcomes with instrument values of zero versus one, rather than comparing the inclusion versus exclusion of (possibly non-binary) instruments, as proposed here. However, akin to huberkueck2022, we test the null hypothesis (ref) using a moment condition which makes use of the following score function to verify whether it is mean zero:
where $V=(Y,D,M,X,Z_1,Z_2)$ represents the random variables, while $\eta_1(V)=(\mu_Y(D,X)$, $\mu_M(D,X)$, $\mu_Y(D,M,X))'$, $\eta_2(V)=(\mu_Y(D,X,Z_1),\mu_M(D,X,Z_1), \mu_Y(D,M,X,Z_2))'$ are the respective conditional means. Additionally, $\zeta$ is an independent mean-zero random variable with $\|\zeta\|_{P,q}<C$, where $C$ is a positive constant, for $q>2$ and variance $\sigma_\zeta^2>0$.\footnote{For any random vector $R = (R_1,...,R_l),$ we have $||R||_q = \max_{1\leq j \leq l}||R_l||_q, $ where $||R_l||_q = (E[|R_l|^q])^{\frac{1}{q}}.$}
Introducing the random perturbation \( \zeta \) serves to prevent our testing approach from having a degenerate variance under the null hypothesis \( H_0 \). The variance \( \sigma_\zeta^2 \) acts as a tuning parameter that should be adjusted based on the sample size \( n \). This adjustment involves a trade-off: a higher variance brings the testing based on \( \phi \) closer to the nominal size in smaller samples, but it may diminish the test's power by potentially obscuring genuine deviations from the null hypothesis. As shown in huberkueck2022, score functions of this quadratic type satisfy Neyman orthogonality under the null hypothesis, which implies that estimators and tests based on such score functions are relatively robust to misspecifications of the conditional mean outcomes, see doubleML. This robustness is particularly advantageous when estimating the conditional mean outcomes using machine learning techniques, which often entail regularization bias in estimation.
For estimating the target parameter \( \theta \), we employ cross-fitting as described in doubleML. This approach involves estimating the models for the conditional mean outcomes and the score function (ref) underlying the test using distinct observations from the sample to mitigate overfitting. We assume an i.i.d. sample with \( n \) subjects, where \( i \) ranges from 1 to \( n \), representing the index of subjects in the sample. The subjects are randomly divided into \( K \) subsamples or folds of size \( n/K \), which is assumed to be an integer for simplicity. Let \( I_{k_i} \) with \( k_i \) ranging from 1 to \( K \) denote the fold containing subject \( i \), and \( I_{k_i}^c \) denote its complement, consisting of all remaining folds excluding subject \( i \). For any subject \( i \) in fold \( k_i \), we obtain predictions \( \hat{\eta}_1^{k_i}(V_i) \) and \( \hat{\eta}_2^{k_i}(V_i) \) of the true conditional means \( \eta_1(V_i) \) and \( \eta_2(V_i) \) respectively, based on estimating the model parameters (e.g., coefficients in a lasso regression) of the conditional means in the complement set \( I_{k_i}^c \). The cross-fitted estimator of \( \theta \) then corresponds to
with $(\hat{\mu}_Y^{k_i}(D_i,X_i), \hat{\mu}_M^{k_i}(D_i,X_i), \hat{\mu}_Y^{k_i}(D_i,M_i,X_i))'=\hat{\eta}_1^{k_i}(V_i)$\\ and $(\hat{\mu}_Y^{k_i}(D_i,X_i,Z_{1,i}), \hat{\mu}_M^{k_i}(D_i,X_i,,Z_{1,i}), \hat{\mu}_Y^{k_i}(D_i,M_i,X_i,Z_{2,i}))'=\hat{\eta}_2^{k_i}(V_i)$.
The results in huberkueck2022, although focusing on a binary instrument, imply that under the null hypothesis \( H_0 \), the estimator \( \hat{\theta} \) follows an asymptotically normal distribution, provided certain regularity conditions are met, such as a convergence rate of the estimators \( \hat{\eta}_1^{k_i}(V_i) \) and \( \hat{\eta}_2^{k_i}(V_i) \) satisfying \( o(n^{-1/4}) \). Appendix (ref) provides a formal proof for the Neyman orthogonality of score function $\phi(V,\theta,\eta)$ and the asymptotic normality of testing under the null hypothesis, along with the required regularity conditions. However, we point out that under the alternative hypothesis, the estimator is non-normal. Regarding the variance of \( \hat{\theta} \), we note that for any subject \( i \), the three squared differences adjusted by \( \zeta_i \) within the brackets of expression (ref) are not independent of each other. Consequently, the conventional variance formula of the score function (ref), \( E[(\eta_1(V)-\eta_2(V))^4]+\sigma_\zeta^2 \), is inappropriate for inference under the null hypothesis. For this reason, we apply cluster-robust variance estimation by clustering at the subject-level to address this issue.
This section presents a simulation study aimed at evaluating the finite sample performance of our testing approach suggested in the previous section. Our first simulation design is based on the following model, which is related to the the causal framework considered in Section (ref):
with $X$, $Z_1$, $Z_2$, $U_1$, $U_2$, and $U_3$ being independent of each other. The treatment variable $D$ is a binary variable modeled by an indicator function $I\{...\}$ whose value is determined by pre-treatment covariates $X$, given that the coefficient vector $\beta$ is nonzero, the first instrument $Z_1$, and the unobservable $U_1$. The mediator $M$ is a linear function of $X$ (if coefficient vector $\beta \neq 0$), treatment $D$, instrument $Z_2$, unobservable $U_2$, and unobservable $U_1$ (if coefficient $\delta\neq 0$). The outcome variable $Y$ is a linear function of treatment $D$, mediator $M$, covariates $X$ (if $\beta \neq 0)$, instruments $Z_1$ and $Z_2$ (if coefficient $\gamma\neq0$), unobservable $U_3$, and unobservable $U_1$ (if $\delta\neq 0$). The causal model implies that the (total) treatment effect is $1+0.5\cdot0.5=1.25$, the direct effect (not operating via the mediator $M$) is $1$, and the indirect (of $D$ on $Y$ via $M$) is $0.5\cdot0.5=0.25$. The unobserved terms $U_1,U_2,U_3$ and the instruments $Z_1,Z_2$ are standard normally distributed and independent of each other and of $X$. $X$ is a vector of normally distributed covariates with zero means and a covariance matrix $\sigma^2_X$, where the covariance of the $i$th and $j$th covariate in $X$ corresponds to $0.5^{|i-j|}$. The coefficient vector $\beta$ gauges the effects of the covariates on $Y$, $M$, and $D$, quantifying the degree of confounding due to observables. The $i$th element of $\beta$ is set to $0.5/i^2$ for $i=1,\ldots,p$, implying a quadratic decay in the relevance of any additional covariate $i$ for confounding.
Our study comprises 1000 simulations for each of two distinct sample sizes \( n \) consisting of 1000 and 4000 observations respectively, with the number of covariates \( X \) set to 200. To test the null hypothesis (ref), we estimate the conditional means involved in the score function (ref) using lasso regression and 2-fold cross-fitting. We generate the random variable \( \zeta \) from a mean-zero normal distribution, \( \mathcal{N}(0,\sigma_\zeta^2) \), with the standard deviation \( \sigma_\zeta \) being inversely proportional to the sample size, \( \sigma_\zeta = 500/n \). This choice ensures that the influence of \( \zeta \) on the estimated violations becomes asymptotically negligible. Furthermore, we estimate the total as well as the natural direct and indirect treatment effects using Double Machine Learning (DML) with cross-fitting, employing the default options of the medDML() command in the R package causalweight by BodoryHuber2018. This method applies lasso regression to estimate the outcome, mediator, and treatment equations, and requires that selection-on-observables holds with respect to the treatment and the mediator (given \( X \) and given \( D,X \) respectively), as imposed in expression (ref) and Assumption (ref). The procedure drops observations with conditional treatment probabilities, known as propensity scores, close to zero or one (i.e., smaller than a threshold of 0.01 or 1% or larger than 0.99 or 99%) from the estimation to prevent an inflation of the propensity score-based weights, which could lead to increased variance in effect estimation.
We note that setting the coefficients $\delta=0$ and $\gamma=0$ in the simulations implies the satisfaction of any conditional independence assumptions imposed on the treatment, the mediator, or the instruments, such that the testable implications in Theorem (ref) all hold. In contrast, when \( \delta \neq 0 \), the variable \( U_1 \) becomes an unobserved confounder jointly affecting the treatment, mediator, and outcome variables. This violates the selection-on-observables assumptions on the treatment and the mediator. Furthermore, when \( \gamma \neq 0 \), the instruments exhibit direct effects on the outcome, violating instrument validity. Table (ref) presents the results of our simulations under various choices of \( \delta \) and \( \gamma \). It includes the test's rejection rate (rej.\ rate) at the 5% level of statistical significance and its average p-value (mean pval). Additionally, we report the absolute biases (bias) and root mean squared errors (RMSE) of the DML estimator of the total, direct, and indirect effects to assess how violations of the testable implications impact effect estimation performance. These statistics are provided for the direct effects under both treatment and control (dir.(1), dir.(0)), defined as \( E[Y(1,M(1))-Y(0,M(1))] \) and \( E[Y(1,M(0))-Y(0,M(0))] \), as well as for the indirect effects under both treatment and control (indir.(1), indir.(0)), defined as \( E[Y(1,M(1))-Y(1,M(0))] \) and \( E[Y(0,M(1))-Y(0,M(0))] \). However, it is worth noting that in our simulation design without treatment-mediator interaction effects, the respective effects are equivalent under treatment and control.
The top panel of Table (ref) shows the results for \( \delta = 0 \) and \( \gamma = 0 \), ensuring satisfaction of all testable implications. The empirical rejection rate of the test is close to the nominal rate of 5%, with values of 0.044 (or 4%) for the smaller sample size \( n = 1000 \) and 0.047 (or 4.7%) for \( n = 4000 \). Correspondingly, the average p-value across the 1000 simulations is relatively high, slightly exceeding 50%. Regarding effect estimation, we find that the absolute biases of the total, direct, and indirect effects are generally close to zero and decrease with increasing sample size. Additionally, the RMSE decreases as the sample size increases from \( n = 1000 \) to \( n = 4000 \), decaying roughly at a rate proportional to \( \sqrt{n} \). The intermediate panel reports the results for \( \delta = 1 \) and \( \gamma = 0 \), indicating a violation of the selection-on-observables assumptions. Accordingly, the rejection rate of our test increases to 68.8% under the smaller sample size and reaches 100% under \( n = 4000 \), while the average p-value decreases from 12.2% to 0%, which points to a very decent power of our test in the scenario considered. As expected, the DML estimator of the total, direct, and indirect effects performs poorly under both sample sizes, with absolute biases and RMSEs substantially different from zero.
In the bottom panel, we consider \( \delta = 0 \) and \( \gamma = 0.2 \), implying a violation of instrument validity due to moderate direct effects of the instruments on the outcome. We observe that the test's statistical power increases sharply with sample size, from 8.6% under \( n = 1000 \) to 100% under \( n = 4000 \). Regarding the estimated effects, we note that the absolute biases are generally not substantial but not negligible either, and they only diminish slightly as the sample size increases. This phenomenon occurs because the instruments confound the treatment-outcome and mediator-outcome associations due to their direct effects on the outcome, which is not accounted for in the effect estimations that only consider \( X \) as covariates to be controlled for. Overall, the simulation results suggest that the test exhibits satisfactory size and power to detect violations of the testable implications, which increase with sample size.
Our second simulation design explores a scenario where the first instrument, $Z_1$, exerts a direct influence on the second instrument, $Z_2$. This structural adjustment is reflected by redefining $Z_2$ in model (ref) as follows, while the remaining variable definitions are the same as before:
Table (ref) presents the results of 1000 simulations when testing the implications of Theorem (ref), which now partly contain $Z_2$ in the conditioning set. In the scenario where all testable implications are met ($\delta$ = 0 & $\gamma$ = 0), as depicted in the top panel, the empirical size of the test closely mirrors the nominal size of 5%, while the average p-value is larger than 50%. Notably, the absolute biases of effect estimates remain close to zero across both sample sizes of $n = 1,000$ and $n = 4,000$ and the RMSE follows a decay rate roughly proportional to $\sqrt{n}$. Transitioning to the intermediate panel with $\delta = 1$ and $\gamma = 0$, implying a violation of the selection-on-observables assumptions, testing power quickly increases in the sample size, while the effect estimators are substantially biased. In the bottom panel with $\delta = 0$ and $\gamma = 0.2$, implying that the instruments directly affect the outcome, the test promptly gains power in the sample size, while the effect estimators are moderately biased. These findings are qualitatively similar to those of the initial simulation design discussed above.
This section presents an empirical application of our test to rich administrative data from Slovakia previously analyzed by almp to evaluate a Slovak active labor market intervention known as the `Youth Guarantee'. The dataset contains comprehensive information from administrative records maintained by the Slovak public employment service centers, which are part of the Central Office for Labour, Social Affairs, and Family of the Slovak Republic (COLSAF), including the unemployment history, socio-economic characteristics, participation in labor market programs, and regional characteristics.
We focus on the dynamic effects of a specific sequential labor market intervention aimed at unemployed youth for job seekers registering as unemployed at the beginning of 2016. The program sequence starts with a three to six-month training called graduate practice (GP), designed to foster early-career workplace experience, albeit with less substantial monetary contributions than the subsequent interventions. The time lag between unemployment registration and the start of GP must be at least one month by law. After completing GP, participants could advance to a second program that is typically twelve months long and provides employment incentives (EI) that combine hiring incentives with subsidized employment, by offsetting up to 75% of the costs for employing unemployed youth over twelve months, which is then followed by a compulsory employment term of six months. svabova_kramarova_2021 and stefanik_can_2020 find these interventions to have moderate but statistically significant positive effects on employment, but negative effects on earnings. In our analysis, the first intervention, $D$, is defined as receiving GP within the first 6 months after the unemployment registration, while the second intervention, $M$, is defined as receiving employment incentives within months seven to twelve since unemployment registration. The timeline of the interventions is depicted in Figure (ref).
Central to our testing approach are two instrumental variables, $Z_1$ and $Z_2$, designed to capture the local availability of programs $D$ and $M$, respectively, at public employment service (PES) centers. The first instrument $Z_1$ is computed based on the ratio of jobseekers enrolled in intervention $D$ in the previous year (2015) to the total influx of new jobseekers at the respective PES office during the previous two years (2014 and 2015). An analogous methodology is applied to compute the second instrument $Z_2$ related to intervention $M$. Our approach is in line with a range of studies that leverage regional variations in the availability of interventions across employment agencies, particularly to construct instruments for participation in specific labor market interventions. Examples include froelich_lechner_2010, lechner_2013, boockmann_2014, markussen_2014, FrolichLechner2015, and, in particular, caliendo_2017, dauth_2020, and lang_2022, who use a similar approach as ours by comparing the number of participants in a program to the overall number of eligible jobseekers. By applying this approach in the Slovak context, our instruments aim to quantify the local propensity for receiving a specific labor market intervention. It is worth noting that the application to a particular intervention needs to be submitted by the jobseeker themselves, while it is the regional PES management that makes the final decision of acceptance, as a function of budgetary restrictions and caseworker recommendations. As such, program enrollment variability is significantly shaped by regional factors, such as unemployment rates, proximity to PES offices, and demographic considerations, including the proportion of the Roma population in the area, which we can observe and control for.
In addition to regional control variables, our data set contains rich individual-level information. The pre-treatment covariates ($X$) are measured prior to the GP intervention (\(D\)) and consist of 264 variables that include besides the regional information for instance a jobseeker's marital status, the presence of dependents, comprehensive measures of education and skills, employment histories, prior claims of unemployment benefits, willingness to relocate for work, health information, and caseworker assessments of employability prospects. We also include five post-treatment covariates ($W$) in our analysis, which are measured after (and may be influenced by) \(D\) but prior to the start of the EI intervention (\(M\)) and might potentially affect both EI participation and the outcome variable. The variables \(W\) include information on a jobseeker's participation in other labor market programs than GP ($D$) during the first intervention period and whether the jobseeker was absent from the unemployment register in months four to six (e.g., due to being employed). Finally, our outcome variable (\(Y\)) is a binary indicator for employment three years after the initial unemployment registration. The causal framework considered for testing relies on the assumptions and testable implications discussed in Section (ref), related to the illustration in Figure (ref).
Our evaluation sample comprises 12,436 individuals aged 15 to 29 who registered as unemployed in 2016, remained jobless for a minimum of three months, and were tracked for up to 48 months after registration. Among them, 2,491 participated solely in GP ($D$), 1,597 solely in EI ($M$), 2,025 in both programs, and 6,323 in neither. In line with suggestions by doubleML, Table (ref) displays the median p-value (pval) from 21 runs of the testing procedure outlined in Section (ref), which applies lasso regression and 10-fold cross-fitting to estimate conditional means as simulated in Section (ref). This procedure tests the implications of Theorem (ref) detailed in Section (ref). The median p-value for these runs is 24.19%, the related test statistics 0.000421, and the standard error 0.000360, such that we cannot reject the null hypothesis of the joint satisfaction of the testable implications at conventional significance levels. Furthermore, we note that the instruments $Z_1$ and $Z_2$ exhibit highly statistically significant first-stage associations with $D$ and $M$, respectively. When utilizing the {\tt DoubleMLData()} command (with default options) in the {\tt DoubleML} package by Bachetal2024 to estimate the first-stage associations of $D$ and $Z_1$ conditional on $X$ and of $M$ and $Z_2$ conditional on $(D,X,W)$ based on DML with lasso regression models for the instruments and interventions, the p-values are below 1%. This suggests the satisfaction of Assumptions (ref) and (ref), which are important for the power of the test.
Due to the insignificant test result, we suspect that the selection-on-observables assumptions regarding $D$ and $M$ are (close to being) satisfied. For this reason, Table (ref) also presents an estimate of the ATE (effect) of participating in both GP and EI versus not participating in any program, $E[Y(1,1)-Y(0,0)]$, along with the standard error of the effect estimate (effect_se) and its p-value (effect_pval). Estimation is based on the {\tt dyntreatDML} command (with default options) of the {\tt causalweight} package, which employs cross-fitted DML and lasso regression for the estimation of the models for $D$, $M$, and $Y$, as discussed in BodoryHuberLaffers. The ATE estimate suggests that participating in both programs increases the employment probability by 8.55 percentage points, and this effect is highly statistically significant, with a p-value of less than 0.1%.
Table (ref) also reports the number of discarded observations due to extreme propensity scores related to $D$ and/or $M$ (effect_ntrimmed), indicating that the product of the propensities to receive $D$ and $M$ is either smaller than 1% or larger than 99%. Specifically, 6,280 observations, or 50.5% of the original sample, were discarded, implying that only roughly half of the individuals in the sample are sufficiently similar in terms of their covariates across different interventions to be considered for effect estimation. It is noteworthy that various propensity score thresholds for discarding observations, such as smaller than 0.5% and larger than 99.5%, or smaller than 5% and larger than 95%, yield very similar effect estimates.
A further analysis was conducted to examine the robustness of our test results to changes in the set of control variables. Specifically, the number of covariates included in the model was reduced by excluding post-treatment variables $W$ and substantially reducing $X$ to only include age and four binary indicators representing the highest level of educational attainment. This revised specification yielded a median test statistic of 0.00066, with a corresponding standard error of 0.00036 and a p-value of 6.92%. Therefore, our test rejects the null hypothesis at the 10% significance level, which is in line with our impression that the assumption of sequential ignorability does not appear plausible under such a limited set of control variables.
This paper introduces a new method for testing the identification of causal effects in mediation, dynamic treatment, and sample selection models, by making use of covariates to be controlled for and suspected instruments. It establishes sufficient testable conditions for identifying such effects in observational data. These conditions jointly imply the exogeneity of the treatment and the mediator conditional on observed variables, as well as the validity of the instruments for the treatment and mediator (or sample selection indicator). We explore various causal frameworks and testable implications, contingent upon whether the covariates to be controlled for are solely pre-treatment or also include post-treatment variables, and depending on the causal association between the instruments for the treatment and the mediator. We suggest a machine learning-based testing approach that handles high-dimensional covariates in a data-adaptive manner and demonstrates a decent performance in our simulation study. Furthermore, we apply the test to evaluate a sequence of active labor market programs in Slovakia, incorporating both pre- and post-treatment control variables, and do not reject the testable implications for the identification of dynamic treatments when employing the local availability of labor market programs as instruments.
\setlength\baselineskip{14.0pt}
{ \setcounter{equation}{0}