Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,055 characters · 8 sections · 86 citation commands
Difference-in-differences for mediation analysis using double machine learning
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 {
} \fi
{\it Keywords:} Difference-in-differences, mediation, causal inference, machine learning, parallel trends, controlled direct effect, dynamic treatment effects, natural direct effect, natural indirect effect.
\thispagestyle{empty} \setcounter{page}{1}
\spacingset{1.45}
Difference-in-differences (DiD) Snow1855, Ashenfelter78 is among the most popular methods for treatment evaluation, that is, for assessing the impact of a treatment on an outcome of interest in observational studies, provided that outcomes can be observed both before and after the introduction of the treatment. In such studies, the outcomes of treated individuals observed before and after treatment typically do not allow researchers to directly infer the treatment effect, due to confounding time trends in the treated individuals’ counterfactual outcomes — those that would have occurred in the absence of treatment. The DiD approach addresses this issue by using a control group that does not receive the treatment and by invoking a parallel trend assumption, which requires that the treated group’s unobserved counterfactual outcomes under non-treatment follow the same mean outcome trend as observed in the control group. This permits identifying the average treatment effect on the treated (ATET). While the standard DiD setup focuses on the evaluation of the total effect of a treatment, researchers are often also interested in the causal mechanisms through which the effect operates, such as indirect effects transmitted via intermediate variables (so-called mediators) or the treatment's direct effect, net of effects transmitted through mediators. Furthermore, one may wish to study the joint effects of specific treatment and mediator values, which corresponds to dynamic treatment effects obtained by setting the treatment and mediator to specific levels.
Building on this motivation, our study proposes a new DiD framework for causal mediation and dynamic treatment effect analysis that accommodates (possibly multivalued discrete or continuous) treatments and mediators. Depending on the form of the imposed parallel trends assumption, the framework allows the assessment of average direct and indirect treatment effects, as well as joint effects of the treatment and the mediator. Identification of dynamic treatment effects --- such as the joint effect of treatment and mediator, or the (controlled) direct effect of the treatment when fixing the mediator at a specific value --- relies on a conditional parallel trends assumption imposed on mean potential outcomes under no treatment and no mediation, or under specific nonzero treatment or mediator values, conditional on observed covariates. The assumption requires that, conditional on covariates, the mean potential outcomes of the treatment group (characterized by specific values of the treatment and the mediator) would have followed the same average trend as those of the control group not receiving the treatment and mediator (or exposed to a different treatment-mediator combination than the treatment group), had the treatment group been assigned the same treatment and mediator values as the control group.
A further class of causal parameters studied in causal mediation analysis are natural direct and indirect effects RoGr92, Pearl01, which are defined in terms of potential mediator values as functions of the treatment, rather than fixed mediator values. For example, the natural indirect effect captures the effect of the treatment on the outcome operating through treatment-induced changes in the mediator. We show that additional assumptions permit identification of such natural effects, by imposing conditional parallel trends either on mean potential outcomes across treatment groups (in addition to across treatment–mediator combinations), or on the distribution of the mediator across treatment groups.
We propose flexible DiD estimators for both repeated cross sections, where subjects differ across time periods, and panel data, where the same subjects are repeatedly observed over time. The estimators build on the double/debiased machine learning (DML) framework of Chetal2018 to control for covariates in a data-driven manner, which is particularly advantageous when the number of potential control variables is large relative to the sample size. To this end, we estimate the conditional mean outcome and treatment models using machine learning methods and employ these nuisance estimates as plug-ins in doubly robust (DR) score functions Robins+94, RoRo95, which we adapt to the DiD setting for estimating direct and indirect effects.
We show that our estimators satisfy the Neyman1959-orthogonality condition, implying that they are relatively robust—i.e., first-order insensitive—to estimation errors in the nuisance parameters under suitable regularity conditions, such as approximate sparsity when lasso regression is used for nuisance estimation Bellonietal2014. Following Chetal2018, we further employ cross-fitting to ensure that nuisance parameters and score functions are not estimated on the same subsamples, thereby mitigating overfitting bias. Furthermore, we present simulation evidence demonstrating favorable finite-sample performance of our estimators for sample sizes in the order of several thousand observations. Finally, we provide an empirical application that revisits the setting of farbmacher2022causal, in which the direct effect of health care coverage on general health is disentangled from the indirect effect operating through routine checkups, using data from the National Longitudinal Survey of Youth 1997 (NLSY97). Even though the point estimates of both the total and direct effects suggest a health-improving effect, we find no statistically significant evidence of a short-term effect of health care coverage on general health among individuals who obtain coverage, either through routine checkups or through other causal mechanisms.
Our study contributes to a growing literature on the semi- and nonparametric DiD-based identification of causal effects. This framework avoids potential misspecification errors of classical linear two-way fixed effects models --- see, for instance, the discussion in GoodmanBacon2018 and de2020two, among others. In contrast to much of the existing work, we not only consider binary treatments but also multivalued (and possibly continuous) treatments, as in callaway2024difference and dechaisemartin2023differenceindifferences, who, however, do not focus on mediation analysis or on controlling for observed covariates. In this paper, we invoke a conditional parallel trends assumption within a semiparametric DiD framework, implying that parallel trends are assumed to hold conditional on observed covariates, as considered in abadie2005 for binary treatments (for ATET estimation rather than mediation analysis).
We study identification in both panel data and repeated cross-section settings. Our approach for repeated cross-sections can accommodate covariate distributions that vary over time within treatment groups, whereas many existing methods require these distributions to remain time-invariant within groups, as pointed out by Hong2013. This limitation affects, for example, DiD estimators for binary or discrete treatments proposed by abadie2005 (inverse probability weighting), SantAnnaZhao2018 (DR estimation), and Chang2020 (DML). A notable exception is the DML approach by Zimmert2018, which allows for time-varying covariate distributions within treatment groups but applies it to ATET estimation for binary treatments rather than to mediation analysis (with possibly non-binary treatments and mediators), as we do. Further related work includes zhang2025, who consider DML-based DiD estimation of the ATET for continuous versus no treatment, and haddad2024, who extend the framework to the evaluation of time-varying treatments and comparisons of strictly positive treatment doses (as also considered by dechaisemartin2023differenceindifferences, though without covariate adjustment). In contrast, our paper considers (potentially non-binary) treatments and mediators to assess direct and indirect effects on the treated, rather than the ATET alone.
Our paper is also related to recent work on DiD methods for mediation analysis. Deuchertetal2019 consider a binary, random treatment and a binary, nonrandom mediator to identify direct and indirect effects in specific subpopulations defined by potential mediator states as functions of treatment, based on parallel trend assumptions within those subpopulations. Their results demonstrate that natural direct and indirect effects defined in terms of potential mediators RoGr92, Pearl01 for the total or treated population can be obtained only if specific parallel trend assumptions hold across multiple subpopulations. In this paper, we impose parallel trend assumptions that are strong enough to identify natural direct and indirect effects among the treated, but assume them to hold conditional on observed control variables, for which we adjust in a data-adaptive manner using machine learning. As a further distinction, we also consider non-binary treatments and mediators.
Similar to our setting, Schenk2024 assumes both the treatment and mediator to be nonrandom and discusses the identification of natural direct and indirect effects among the treated based on specific parallel trend assumptions. The study highlights interesting trade-offs between different sets of assumptions: for a binary treatment and mediator, for instance, one may impose parallel trend assumptions that only hold across some, but not all subpopulations defined by potential mediators, but must then make additional assumptions such as monotonicity of the mediator in treatment. In contrast, our paper imposes stronger parallel trend assumptions across groups defined by treatment and potential mediator values. As a result, monotonicity is not required, and identification is also feasible for non-binary treatments and mediators. Furthermore, unlike Schenk2024, we employ machine learning methods to control for covariates in a data-adaptive way when imposing parallel trends.
blackwell2022difference considers the identification of controlled direct effects under a random treatment and a nonrandom mediator - that is, when fixing the mediator at a specific value to assess the direct treatment effect, rather than at its potential value as in the natural direct effect. Also in this paper, we derive identification results for the controlled direct effect. However, our framework allows both the treatment and the mediator to be nonrandom and furthermore, accommodates data-adaptive adjustment for covariates using machine learning. Our paper is further related to the literature on dynamic treatment effects (of which the identification of controlled direct effects is a special case); see, for instance, Ro86 and RoHeBr00. Within this literature, DR identification has been studied for instance in Tranetal2019, as well as DML estimation; see LewisSyrgkanis2020, bodory2022evaluating, and bradic2024high. However, these studies rely on selection-on-observables assumptions, while our paper focuses DiD methods that invoke parallel trend assumptions.
The remainder of this paper is organized as follows. Section (ref) discusses the DiD-based identification of dynamic treatment effects and controlled direct effects in repeated cross-sectional data. Section (ref) focuses on the identification of natural direct and indirect effects in causal mediation analysis in repeated cross-sectional data. Section (ref) considers identification in panel data. Section (ref) proposes estimators based on DML with cross-fitting and also provides asymptotic results for the proposed estimators. Section (ref) provides a simulation study for the repeated cross-section and panel data cases. Section (ref) presents an empirical application that decomposes the effect of health care coverage on general health into a direct effect and an indirect effect operating through routine checkups. Section (ref) concludes.
This section discusses DiD-based identification of dynamic treatment effects as well as direct and indirect effects in repeated cross-sectional data. We consider a discretely distributed treatment and mediator, denoted by $D$ and $M$, respectively. Furthermore, let $T$ denote the time period, where $T=0$ is the baseline period prior to treatment and mediator assignment, while $T=1$ is a post-treatment and post-mediator period in which the effect on the outcome is measured. We use capital letters to refer to random variables and lowercase letters for specific realizations of these variables. $Y_t$ denotes the observed outcome in period $T=t$, while $Y_t(d,m)$ represents the potential outcome in period $t$ when hypothetically setting the treatment to $D=d$ and the mediator to $M=m$; see e.g.\ Neyman23 and Rubin74 for more discussion on the potential outcome framework. Throughout the paper, we impose the stable unit treatment value assumption (SUTVA), implying that a subject's potential outcomes are not affected by the treatment or mediators of others; see, for instance, the discussion in Rubin80 and Cox58. Finally, we denote by $X$ observed covariates that may serve as control variables.
In our discussion, we refer to the treatment group as those subjects receiving a specific treatment value $D=d$ and mediator value $M=m$. We consider the identification of the average treatment effect on the treated (ATET) for this treatment–mediator combination, which is defined by comparing the observed average outcomes in the treatment group to the average potential outcomes for the same group under counterfactual treatment and mediator values $d'$ and $m'$. Formally, the ATET of treatment and mediator states $D=d, M=m$ versus $D=d', M=m'$ in the treatment group with $D=d, M=m$ in the post treatment period $T=1$ is given by
For instance, for a binary treatment and mediator and for $d=m=1$, $d'=m'=0$, the ATET $\Delta_{1,1}(1,1,0,0) = E[Y_1(1,1) - Y_1(0,0) \mid D=1, M=1, T=1]$ corresponds to the average effect of receiving both the treatment and the mediator versus receiving neither, among the group that actually receives the treatment and the mediator and is observed in the post-treatment period. More generally, for $d \neq d'$ and $m \neq m'$, the effect in (ref) corresponds to the joint effect of the treatment and mediator on the outcome among those with $D=d$ and $M=m$. This framework permits the consideration of dynamic treatment effects, i.e., the effects of different sequences of treatments and mediators, as, for example, in Ro86 and RoHeBr00 (which rely on sequential selection-on-observables assumptions, whereas our study relies on parallel trends assumptions). In contrast, if we fix the mediator value at $m=m'$, the resulting ATET corresponds to the controlled direct effect; see, e.g., Pearl01. For instance, $\Delta_{1,1}(1,0,0,0) = E[Y_1(1,0) - Y_1(0,0) \mid D=1, M=1, T=1]$ represents the controlled direct effect when fixing the mediator at $m=m'=0$.
For identification of the ATET, we impose the following assumptions related to Lechner2010 for binary treatments, but now imposed w.r.t.\ to a combination of discrete treatment and mediator variables:
Assumption (ref) imposes parallel trends in the mean potential outcome under the treatment–mediator combination $D=d', M=m'$ across the treatment group with $D=d, M=m$ on the one hand and the control group with $D=d', M=m'$ on the other hand, conditional on covariates $X$. For binary treatment and mediator variables with $d=m=1$ and $d'=m'=0$, this implies $E[Y_1(0,0)-Y_0(0,0)|D=1,M(1)=1,X]=E[Y_1(0,0)-Y_0(0,0)|D=0,M(0)=0,X]$, where we have used the fact that, by the observational rule, $M = M(1)$ conditional on $D=1$, while $M = M(0)$ conditional on $D=0$. Therefore, the parallel trend assumption is imposed not only across groups with treatment values $D=1$ and $D=0$, but also across groups defined by potential mediators under treatment ($M(1)=1$) and under nontreatment ($M(0)=0$). It is worth emphasizing that these are generally distinct subgroups in terms of how the mediator responds to treatment, an issue that fits into the causal framework of principal stratification, see frangakis2002principal, and is closely related to compliance in instrumental variable contexts, see Angrist+96: The subpopulation with $M(1)=1$ consists of subjects complying with the treatment in the sense that they take the mediator as a result of being treated, characterized by $M(0)=0, M(1)=1$, as well as subjects who always take the mediator, characterized by $M(0)=M(1)=1$. In contrast, the subpopulation with $M(0)=0$ includes complying subjects ($M(0)=0, M(1)=1$) and subjects who never take the mediator ($M(0)=M(1)=0$).
We note that this specific type of parallel trends assumption, $E[Y_1(0,0)-Y_0(0,0)|D=1,M(1)=1,X]=E[Y_1(0,0)-Y_0(0,0)|D=0,M(0)=0,X]$, is commonly imposed in the DiD literature on multiple treatment periods, including staggered treatment adoption across groups. See, for instance, the discussions in borusyak2024revisiting, CallawaySantAnna2018, GoodmanBacon2018, cha19, and sun2021estimating. Indeed, the multiple treatment period setup can be interpreted through the lens of our mediation framework by defining $D$ as treatment adoption in an earlier period and $M$ as treatment adoption in a later period, while typically allowing for additional treatment periods beyond these two.
For this reason, Assumption (ref) implicitly imposes parallel trends across (mediator) always- and never-takers. Depending on the context and the values of $d, d', m, m'$ considered, our parallel trend assumptions are more stringent than those considered for causal mediation in Deuchertetal2019 under random treatment and nonrandom mediator, or in Schenk2024 under nonrandom treatment and mediator, who tailor parallel trend assumptions to specific subpopulations defined by potential mediator states as a function of treatment. Consequently, identification of particular direct and indirect effects in Deuchertetal2019 and Schenk2024 may require additional assumptions, such as monotonicity of the mediator in treatment, e.g., $\Pr(M(1)-M(0) \geq 0 \mid D = 1) = 1$. Such monotonicity is not imposed here, due to the stronger parallel trend assumption across subgroups defined by potential mediators (and treatments). This parallel trend assumption imposes restrictions on the selection bias arising from unobserved confounders that jointly affect the treatment/mediator and the outcome, which can cause treatment and control groups defined by different treatment–mediator combinations, $D=d, M=m$ and $D=d', M=m'$, to have different average potential outcomes under control, $Y_t(d', m')$. Specifically, this selection bias must be constant across time periods $t=0$ and $t=1$, conditional on $X$.|D=d,
Assumption (ref) requires that, conditional on $X$, the ATET is zero in the baseline period. This rules out, on average, any anticipation effects among treated units in period $T=0$ with respect to the treatment and mediator to come. In other words, since the treatment and mediator have not yet been realized, they cannot, on average, induce behavioral adjustments in the treated group that affect pretreatment outcomes in anticipation of future treatment or mediator values.
Assumption (ref) requires that, for any subject with $D=d,M=m$ in the post-treatment period $T=1$, there exist, in terms of $X$, comparable subjects with $D=d,M=m$ in the baseline period, with $D=d',M=m'$ in the post-treatment period, and with $D=d',M=m'$ in the baseline period. In other words, for every combination of covariates observed among treated and mediated units in the post-treatment period, there must exist comparable observations in all other three treatment/mediator–time cells. This common support condition ensures that the counterfactual outcome distributions for the target population with $D=d,M=m$ in the post-treatment period can be recovered from observable data in the other cells.
Assumption (ref) imposes that the covariates are not causally affected by the treatment or the mediator. This can be a delicate restriction if the covariates include post-treatment variables (as is often the case in repeated cross sections); see, for instance, the discussion in caetano2022difference. Controlling for post-treatment $X$ generally introduces bias if $X$ is affected by $D$ and/or $M$, and $X$ either affects $Y$, is correlated with unobservables affecting $Y$, or both. Assumption (ref) rules out such bias by requiring that potential covariates are not a function of the treatment or mediator, but this assumption should be carefully scrutinized in practice whenever one conditions on post-treatment covariates.
Under these assumptions, the mean potential outcome under non-treatment $D=d'$ and mediator state $M=m'$ is identified in the post-treatment period among the group receiving treatment doses $D=d$ and $M=m$, conditional on $X$:
The first equality in (ref) follows from subtracting and adding $E[Y_0(d',m')|D=d,M=m,X]$ and the second from Assumption (ref). The third equality follows from Assumption (ref). The fourth equality follows from the fact that $Y_t=Y_t(d,m)$ conditional on $D=d$ and $M=m$, which is known as the observational rule. Finally, note that by Assumption 4, $X$ is not causally affected by either $D$ or $M$, implying that conditioning on $X$ does not block any causal effect of $D$ or $M$ on $Y$.
It follows that $E[Y_1(d',m')|D=d,M=m]$ can be obtained by averaging the expression in the last line of (ref) over $X$ conditional on $D=d, M=m$ in the post-treatment period $T=1$. This step requires satisfaction of Assumption (ref) (common support):
Furthermore, we have by the observational rule that
For notational convenience, we henceforth denote the conditional mean outcome by $\mu_{d,m}(t,X)=E[Y_t \mid D=d,M=m,X]$. Furthermore, let $\Pi_{d,m,t}=\Pr(D=d,M=m,T=t)$ denote the joint probability of treatment, mediator, and period, and $\rho_{d,m,t}(X)=\Pr(D=d,M=m,T=t\mid X)$ the corresponding conditional probability given $X$, also known as propensity score. Let $I{\cdot}$ denote the indicator function, which equals one if its argument is satisfied and zero otherwise. For the moment, assume that the treatment $D$ and mediator $M$ are discrete. Taking the difference between (ref) and the average of (ref) then identifies the ATET of treatment and mediator states $D=d, M=m$ versus $D=d', M=m'$ in the group with $D=d, M=m$ in the post-treatment period, as given in equation (ref):
As a methodological contribution, we propose an alternative, doubly robust (DR) expression for the ATET which is numerically equivalent to (ref) under our identifying assumptions and, in contrast to (ref), satisfies Neyman1959-orthogonality. This property makes it particularly attractive for combination with machine learning methods for the estimation of the conditional mean outcomes and propensity scores. The DR identification result is provided in Proposition (ref).
Appendix (ref) provides formal proofs that equation (ref) identifies the ATET and satisfies Neyman orthogonality.
It is worth noting that, in addition to the ATET as defined in (ref), strengthening Assumptions 1–4 to not only hold for the population with $D=d, M=m$, but also in analogous manner for the population with $D=d',M=m'$ permits identification of the average treatment effect (ATE) in the total population in the post-treatment period, defined as
This result follows because, under the stated assumptions applied symmetrically to both $(D=d,M=m)$ and $(D=d',M=m')$ cells, the conditional mean potential outcomes $E[Y_1(d,m)\mid X]$ and $E[Y_1(d',m')\mid X]$ are identified for all vlues of $X$ in the population.
Assuming for instance that both $D$ and $M$ are binary, imposing Assumption (ref) symmetrically to $(D=1,M=1)$ and $(D=0,M=0)$ imposes parallel trends both without and with treatment and mediation, that is, $E[Y_1(0,0)-Y_0(0,0)|D=1,M=1,X]=E[Y_1(0,0)-Y_0(0,0)|D=0,M=0,X]$ and $E[Y_1(1,1)-Y_0(1,1)|D=1,M=1,X]=E[Y_1(1,1)-Y_0(1,1)|D=0,M=0,X]$, respectively. It is worth noting that these two assumptions jointly imply homogeneity in average effects. To see this, note that by subtracting the two parallel trend conditions, we obtain
where the second and third equalities follow from Assumption (ref). Under this stronger set of assumptions, the ATE is identified by
A corresponding DR expression for identifying the ATE that is Neyman orthogonal (as formally shown in Appendix (ref)) is given in Proposition (ref), where $p_t(X) = \Pr(T = t \mid X)$ denotes the conditional period probability given covariates.
While the last section focused on the identification of directed controlled effects as well as dynamic treatment effects, this section discusses the disentangling of the total ATET into natural direct and indirect effects defined in terms of potential mediators as a function of the treatment RoGr92, Pearl01, as aimed for in causal mediation analysis; see, for instance, Huber2019 for a survey. The motivation is that while canonical DiD has been developed for assessing the average effect of some treatment $D$ on the treated population, we are frequently interested in decomposing this ATET into various causal mechanisms, such as the average indirect effect operating through a mediator $M$. For this reason, we now consider the (total) ATET of treatment $D$ in the post-treatment period, defined as
where the second equality follows from the conventional definition of a potential outcome as a function of $D$ only, such that $Y_1(d,M(d))=Y_1(d)$, as typically considered in the canonical DiD literature.
By adding and subtracting $Y_1(d',M(d))$ within the expectation in (ref), it can be easily seen that the ATET can be decomposed into a natural direct effect, defined as
and a natural indirect effect, defined as
While $E[Y_1(d,M(d))|D=d,T=1]=E[Y_1|D=d]$ is directly observed in the data, identification of the ATET and the natural effects hinges on the identification of the counterfactuals
where $\mathcal{M}$ denotes the support of $M$, which is here assumed to be discrete.
The identification of the counterfactual in equation (ref) requires that Assumption (ref) holds for potential outcomes $Y_1(d',m)$ across treatment groups $D=d$ and $D=d'$ when fixing the mediator at any value $m$ in the support of $M(d)$ among the group with $D=d, T=1$: $E[Y_1(d',m)-Y_0(d',m)|D=d,M=m,X]=E[Y_1(d',m)-Y_0(d',m)|D=d',M=m,X]$. This condition permits identification of $E[Y_1(d',m)|D=d,M(d)=m]$, while $\Pr(M(d)=m \mid D=d, T=1)$ is directly identified from the observed data. Under this assumption, $E[Y_1(d',M(d)) \mid D=d]$ is given by
where the second equality follows from the observational rule, the third from the law of iterated expectations, and the fourth from Assumption (ref) if it holds for all values $m$ in the support of $M(d)$.
A DR expression for the counterfactual, which under the identifying assumptions is numerically equivalent to the fourth line of equation (ref) and satisfies Neyman orthogonality (as formally shown in Appendix (ref)), is given in Proposition (ref):
Rather than conditioning on each specific mediator value $m$ and averaging over mediator values among the treated in the post-treatment period using the weights $\Pr(M=m \mid D=d, T=1)$, we may instead average directly over the distribution of the realized mediator $M$ among the treated in the post-treatment period. That is, instead of computing for any conditional mean outcome
we may consider the equivalent expression
which integrates over the observed distribution of $M$ given $D=d$ and $T=1$. Analogously, for any debiasing term based on inverse weighting by the propensity score, instead of \[ E\!\left[ \frac{(Y_T - \mu_{D,M}(T,X)) \cdot \rho_{d,m,1}(X)}{\Pi_{d,m,1}} \cdot \frac{I\{D=d'\} \cdot I\{M=m\} \cdot I\{T=t\}}{\rho_{d',m,t}(X)} \right]\cdot \Pr(M=m \mid D=d, T=1), \] we may consider \[ E\!\left[ \frac{(Y_T - \mu_{D,M}(T,X)) \cdot \pi_{d,1}(M,X)}{P_{d,1}} \cdot \frac{I\{D=d'\} \cdot I\{T=t\}}{\pi_{d',t}(M,X)} \right], \] where $\pi_{d,t}(M,X) = \Pr(D=d, T=t \mid M, X)$ is the joint propensity score of treatment $d$ and time $t$ given $M$ and $X$. Therefore, the identification result in equation (ref) is equivalent to the following expression, which avoids conditional mediator probabilities:
The identification of the mean potential outcome in (ref), that is, the counterfactual under treatment $d'$, requires a different parallel trends assumption than Assumption (ref) (for identifying the counterfactual under treatment $d'$ and mediator $m'$). The assumption is now imposed with respect to treatment only (rather than both treatment and mediator), and is formalized in the following assumption.
It is worth noting that, in contrast to Assumption (ref), the validity of Assumption (ref) implies that those confounders affecting treatment and outcome, which are differenced out by a DiD approach, do not directly affect the mediator $M$ other than through the treatment $D$. If there were unobserved variables affecting both $D$ and $M$, the assumption would fail because differencing mean outcomes over time cannot account for these confounders.
To see this, consider the following example. Suppose the outcome in some time period $T$ is a (possibly unknown) function, denoted by $\mathcal{F}_1$, of the observed variables $D$, $M$, and $X$, which may interact with time $T$, and an additively separable, time-constant function, $\mathcal{F}_2$, of time-invariant unobservables: \[ Y_T = \mathcal{F}_1(D, M, X, T) + \mathcal{F}_2(U). \] Furthermore, assume that the treatment is a function of $X$ and $U$, \[ D = \mathcal{F}_D(X, U), \] while the mediator is a function of $D$ and $X$, \[ M = \mathcal{F}_M(D, X). \] Random, time-varying, and additively separable error terms could be added to the models for $Y_T$, $D$, and $M$, but they are omitted for simplicity.
The outcome model implies that differencing mean potential outcomes under treatment value $d'$ across periods conditional on $X=x$ eliminates the time-constant component $\mathcal{F}_2(U)$: \[ E[Y_1(d',M(d')) - Y_0(d',M(d')) \mid X=x] = E[\mathcal{F}_1(d', M(d'), x, 1) - \mathcal{F}_1(d', M(d'), x, 0)]. \] For this reason, Assumption (ref) holds because $U$, while affecting $D$, does not affect $E[\mathcal{F}_1(d', M(d'), x, 1) - \mathcal{F}_1(d', M(d'), x, 0)]$, as the potential mediator $M(d') = \mathcal{F}_M(d', X)$ is not a function of $U$. However, changing the mediator model to $M = \mathcal{F}_M(D, X, U)$ would violate Assumption (ref), as the potential mediator becomes $M(d') = \mathcal{F}_M(d', X, U)$. In this case, $U$ influences both the mean potential outcome difference \[ E[Y_1(d',M(d')) - Y_0(d',M(d')) \mid X=x] = E[\mathcal{F}_1(d', \mathcal{F}_M(d', x, U), x, 1) - \mathcal{F}_1(d', \mathcal{F}_M(d', x, U), x, 0)], \] through the mediator, and also the treatment, such that it acts as a confounder of the treatment and the difference in mean potential outcomes.
When representing the potential outcome only as a function of $d'$, i.e.\ $Y_t(d')=Y_t(d',M(d'))$, we see that Assumption (ref) corresponds to the parallel trend assumption in the canonical DiD literature on ATET identification, $E[Y_1(d')-Y_0(d')|D=d,X]=E[Y_1(d')-Y_0(d')|D=d',X]$.The identification of $E[Y_1(d',M(d'))|D=d,X]$ can be shown by following analogous steps as in equation (ref). Denoting by $\mu_{d}(0,X)=E[Y_0|D=d,X]$, we have that under Assumptions (ref)--(ref) and (ref) satisfied for the distribution of $M(d')$ among $D=d, T=1$,
Denoting by $P_{d,t}=\Pr(D=d,T=t)$ and $\pi_{d,t}(X)=\Pr(D=d,T=t|X)$, a DR robust expression of (ref) is given in Proposition (ref). For $d=1$ and $d'=0$, the expression for counterfactual $E[Y_1(d',M(d'))|D=d, T=1]$ in the proposition is equivalent to that derived in Zimmert2018, who shows that it satisfies Neyman orthogonality (and for completeness, identification based on expression (ref) and Neyman orthogonality is also demonstrated in Appendix (ref)).
As an alternative to Assumption (ref), we may impose Assumption (ref), which permits identifying $E[Y_1(d,m)| D=d, M=m,T=1]$, for all mediator values occurring among the treated and an additional parallel trend assumption on the mediator that permits identifying $\Pr(M(d')=m|D=d,T=1)$. This assumption pertains not only to the conditional mean of the mediator but to its entire distribution for all $m$ in the support of $M(d')$ given $D=d,T=1$, see CALLAWAY2018395 and CallawayLi for related distributional parallel trend conditions. The approach requires that the past value of the mediator can be observed in the pretreatment period. For this reason, we introduce additional notation and denote by $M_0$ the mediator in the pretreatment period ($T=0$), while $M$ continues to denote the post-treatment mediator through which $D$ may affect the post-treatment outcome $Y_t$. The parallel trend assumption on the conditional distribution of the mediator is formalized as follows:
It is worth noting that if the pre-treatment value of the mediator is zero (or more generally, the same) for everyone, Assumption (ref) collapses to a standard selection-on-observables assumption with respect to the mediator, as discussed in Im04. Specifically, in that case, we have $\Pr(M_0(D) = 0| D, X) = 1$ and $\Pr(M_0(D) \neq 0 | D, X) = 0$, so that Assumption (ref) reduces to $\Pr(M(d') = m|D=d,X)=\Pr(M(d') = m|D=d',X)$. This, together with Assumption (ref), in turn implies Assumption (ref), since conditional on $X$ the distribution of $M(d')$ cannot depend on unobserved variables that jointly affect the treatment and mediator. Hence, scenarios such as the one discussed further above with an unobserved term $U$ entering both $D=\mathcal{F}_D(X,U)$ and $M=\mathcal{F}_M(D,X,U)$ are ruled out. Consequently, when the pre-treatment mediator is degenerate (i.e., has the same value for everyone), jointly imposing Assumptions (ref) and (ref) provides no different identifying content beyond imposing Assumption (ref) alone.
In addition to parallel trends with respect to the mediator, we also need to impose a no-anticipation assumption on the mediator values in the pretreatment period, analogous to Assumption (ref) for the outcomes:
The newly introduced assumptions permit the identification of the mean counterfactual outcome based on the following approach:
where the first equality follows from the law of iterated expectations, the third from Assumptions (ref) and (ref) (parallel trends for mean potential outcomes and potential mediators, respectively), as well as from the satisfaction of Assumptions (ref) and (ref) (no anticipation) across mediator values $m$ in the support of $M(d')$ among the treated in the post-treatment period. The last equality follows from the observational rule.
As shorthand notation, we henceforth denote by $\nu_{d,m}(0,X)=\Pr(M_0=m\mid D=d,X)$ and $\nu_{d,m}(1,X)=\Pr(M=m\mid D=d,X)=E[I\{M=m\}\mid D=d,X]$ the conditional mediator probabilities in the pre- and post-treatment periods, respectively. Furthermore, we note that $M_T=M_0$ if $T=0$ and $M$ if $T=1$. The DR expression corresponding to (ref) for identifying $E[Y_1(d',M(d'))\mid D=d, T=1]$ is given in Proposition (ref), which is also shown to satisfy Neyman orthogonality in Appendix (ref).
As a final remark in this section, we note that all previous results derived for discrete treatments and mediators can be extended to the case of continuous treatments and mediators. To this end, each indicator function for treatment values $d$ and $d'$ as well as mediator values $m$ and $m'$ is replaced by a kernel weighting function with a bandwidth approaching zero to achieve identification. Consider, for instance, a continuous treatment $D$ taking value $d$, while maintaining a discretely distributed mediator. We denote by $\omega(D;d,h)$ a weighting function that depends on the distance between $D$ and the reference value $d$, as well as on a non-negative bandwidth parameter $h$: $\omega(D;d,h)=\frac{1}{h}\mathcal{K}\left(\frac{D-d}{h}\right)$, where $\mathcal{K}$ denotes a well-behaved kernel function. The closer $h$ is to zero, the less weight is assigned to observations with larger discrepancies between $D$ and $d$.
Furthermore, note that the previously defined (conditional) probabilities, such as $\Pi_{d,m,t}$ and $\rho_{d,m,t}(X)$, correspond to (conditional) density functions when the treatment is continuous. As discussed in fan1996estimation, the conditional treatment density --- also referred to as the generalized propensity score in ImaivanDyk2004 and HiranoImbens2005 --- can be expressed as $\rho_{d,m,t}(X)=\lim_{h \rightarrow 0} E[\omega(D;d,h) \cdot I\{M=m\} \cdot I\{T=t\}|X]$. Furthermore, $\Pi_{d,m,t}=\lim_{h \rightarrow 0} E[\omega(D;d,h) \cdot I\{M=m\} \cdot I\{T=t\}]$.
Applying these modifications, for instance, to the DR expression in equation (ref) yields
If the mediator is continuous as well, analogous modifications apply to $I\{M=m\}$ and $I\{M=m'\}$, which are then replaced by kernel weighting functions.
This section discusses identification in panel data under the identifying assumptions proposed in the previous section. Panel data permit taking outcome differences within subjects over time, such that we now consider conditional means of within-subject differences in outcomes over time, defined as $\mu_{d',m'}(X)=E[Y_1-Y_0|D=d',M=m',X]$, instead of differences in time-specific conditional means, $\mu_{d',m'}(1,X)-\mu_{d',m'}(0,X)=E[Y_1|D=d',M=m',X]-E[Y_0|D=d',M=m',X]$, as in the previous section. Furthermore, as the same individuals are observed in both time periods $T=1$ and $T=0$, it follows that the treatment and mediator distributions are independent of $T$, both unconditionally and conditional on covariates $X$, the distribution of which does not change over time either. For this reason, $\Pi_{d,m,1}=\Pr(D=d,M=m,T=1)=\Pr(D=d,M=m)=\Pi_{d,m}$ and $\rho_{d,m,1}(X)=\Pr(D=d,M=m,T=1|X)=\Pr(D=d,M=m|X)=\rho_{d,m}(X)$.
We first reconsider the identification of the treatment effect in equation (ref), which (as conditioning on $T=1$ is unnecessary) simplifies to $\Delta_{d,m}(d,m,d',m')=E[Y(d,m)-Y(d',m')|D=d,M=m]$. Not conditioning on $T$ and taking within-subject differences allows simplifying the previous identification result provided in equation (ref) to the DR expression in Proposition (ref), with the proof of Neyman orthogonality given in Appendix (ref). We note that this proposition is closely related to the conventional DR expression for ATET evaluation of discrete treatments under the selection-on-observables assumption (see, for example, Section 4.6 in huber2023causal), with the difference that here, outcome differences (rather than outcome levels) and indicators for both the treatment and the mediator (rather than the treatment alone) enter the expression.
In analogous manner, the DR expression in Proposition (ref) for the ATE defined in equation (ref), now denoted by $\Delta(d,m,d',m')=E[Y(d,m)-Y(d',m')]$, can be simplified when compared to to repeated cross sections. Notably, in panel data we have $\Pr(T=1)=\Pr(T=0)$ and $p_1(X)=\Pr(T=1\mid X)=\Pr(T=0\mid X)$, as each subject is observed in both time periods. Hence, reweighting based on $T$, $\Pr(T=1)$, and $p_1(X)$ is no longer required, which yields the equation stated in Proposition (ref), with the proof of Neyman orthogonality provided in Appendix (ref). The equation is closely related to the standard DR expression for ATE identification under the selection-on-observables assumption, with the difference that here, outcome differences (rather than outcome levels) and indicators for both the treatment and the mediator (rather than the treatment alone) enter the expression.
Likewise, the DR identification result of Proposition (ref), which identifies the counterfactual $E[Y_1(d', M(d)) \mid D=d]$ (where conditioning on $T=1$ is now omitted), simplifies in panel data to the expression in Proposition (ref).
In analogy to (ref) for repeated cross sections, an alternative (but numerically equivalent) DR expression to that in Proposition (ref), which, however, avoids conditional mediator probabilities, is given by:
Furthermore, the DR expression in Proposition (ref) for identifying the counterfactual $E[Y_1(d',M(d')) \mid D=d]$ (where conditioning $T=1$ is now omitted) simplifies, because $P_{d,t}=\Pr(D=d, T=t)=\Pr(D=d)$ and $\pi_{d,t}(X)=\Pr(D=d, T=t \mid X)=\Pr(D=d \mid X)=\pi_{d}(X)$. Denoting $\mu_{d'}(X)=E[Y_1 - Y_0 \mid D=d', X]$, we obtain the expression in Proposition (ref), with the proof of Neyman orthogonality provided in Appendix (ref). This result is analogous to the DR expression, for instance, provided in SantAnnaZhao2018, although in that case it pertains to the ATET rather than to the counterfactual outcome, as considered here.
Finally, we consider the panel data version of the result in Proposition (ref) for the identification of $E[Y_1(d',M(d'))\mid D=d]$ based on an alternative set of assumptions (including a parallel trends assumption on the mediator). To this end, we denote by $\nu_{d',m}(X)=E[ I\{M=m\}-I\{M_0=m\}\mid D=d',X]$ the difference in the conditional probability that the mediator takes value $m$ between post- and pre-treatment periods, given treatment and covariates. The result is given in Proposition (ref), for which identification and Neyman orthogonality are shown in Appendix (ref).
In this section, we outline DiD estimation within the DML framework using cross-fitting, following Chetal2018. The approach is based on the sample analogs of the DR identification results in Sections (ref) to (ref). We focus on estimating the ATET $\Delta_{d,m}(d,m,d',m')$ defined in equation (ref), using the sample analog of equation (ref) from Proposition (ref). Estimation for other DR results proceeds analogously using the corresponding normalized sample analogs; hence, we omit a detailed discussion.
Let $\mathcal{W} = {W_i \mid 1 \leq i \leq n}$, with $W_i = (Y_{i,T}, D_i, M_i, X_i, T_i)$ for $i=1,\ldots,n$, denote an i.i.d.\ sample of size $n$. The nuisance parameters, comprising conditional mean outcomes and propensity scores, are
with corresponding estimates
We estimate $\Delta_{d,m}(d,m,d',m')$ using the following DML algorithm with cross-fitting to avoid overfitting, where within-group weights --- defined by treatment state and time period --- are normalized to sum to one:\newline DML algorithm:
Under specific regularity conditions --- in particular, that the nuisance parameter estimators converge at rate $o(n^{-1/4})$ to the true models --- DML-based effect or counterfactual estimators achieve $\sqrt{n}$-consistency and asymptotic normality for discrete treatments and mediators, analogous to Chetal2018 for ATET and ATE with binary treatments. For example, lasso regression can be used to estimate the nuisance functions, and the required convergence rate is attainable under approximate sparsity, as shown in Bellonietal2014. When treatments and/or mediators are continuous and kernel smoothing is applied as in equation (ref), convergence rates are slower than $\sqrt{n}$ due to the local nature of the estimation. Nonetheless, asymptotic normality can still be achieved under appropriate regularity conditions; see, for instance, zhang2025 and haddad2024 for asymptotic results on DiD-based DML with continuous treatments.
This section presents a simulation study to assess the finite-sample behavior of the proposed methods for repeated cross-sections. We first focus on natural direct and indirect effects and consider the following data-generating process (DGP):
$X$ is a vector of $p=100$ covariates. Each covariate $X_j$, $j\in\{1,...,p\}$, depends on a random error term $Q_j$ and a binary time index $T$, where $T=0$ denotes the pre-treatment period and $T=1$ denotes the post-treatment period, each occurring with equal probability. The binary treatment $D$ is determined by the observed covariates $X$ and two unobserved terms, $U$ and $V_d$. The mediator $M$ is continuous and is a function of $X$, $D$, and the unobserved term $V_m$. The outcome $Y_T$ depends on $X$, $T$, and unobserved terms $U$ and $W$; in the post-treatment period, it additionally depends on $D$, $M$, and their interaction. The impact of $X$ on $D$, $M$, and $Y_T$ is captured by the coefficient vector $\beta$. The $j$th element of $\beta$ is set to $0.4/j^2$, implying a quadratic decay in covariate importance such that covariates with lower indices $j$ exert stronger confounding. The unobserved variable $U$ enters both the treatment and outcome equations and therefore acts as a confounder of the treatment-outcome relationship; this confounding is addressed by our proposed approach. Importantly, $M$ does not depend on $U$, ensuring satisfaction of Assumption (ref). All error terms $Q_j$, $U$, $V_d$, $V_m$, and $W$ are independently drawn from a standard normal distribution.
We consider two sample sizes, $n=2000$ and $n=8000$, and generate 1000 independent simulation replications. The parameters of interest are the natural direct effect, the natural indirect effect, and the total effect, as defined in equations (ref), (ref), and (ref), respectively. These effects depend on counterfactual outcomes identified in Propositions (ref) and (ref). Specifically, the counterfactuals are estimated using normalized sample analogs of equations (ref) and (ref).
Our estimator is implemented in the statistical software R Rcore2025 and uses four-fold cross-fitting. All nuisance parameters are estimated by cross-validated lasso regression, with logit specifications for the propensity scores $\pi_{d,t}(X)$ and $\pi_{d,t}(M,X)$ and linear specifications for the conditional mean outcomes $\mu_d(t,X)$ and $\mu_{d,M}(t,X)$. To safeguard against too influential observations in IPW, we impose a trimming rule based on the estimated propensity scores. Observations are dropped within each of the subgroups $(d,0)$, $(d',1)$, $(d',0)$ whenever either $\pi_{d,t}(X)$ or $\pi_{d,t}(M,X)$ falls below a pre-specified threshold. The trimming threshold is set to 0.05 in all simulations. Inference is conducted using an asymptotic approximation. Standard errors are computed based on the estimated variance of the score function associated with the estimator.
The upper panel of Table (ref) provides the simulation results for estimating the (total) ATET and decomposing it into the natural direct and indirect effects. The true ATET equals $\Delta_1 (1,0,M(1),M(0))=E[Y_1(1)-Y_1(0)|D=1, T=1]=1.5+M(1)=2.419$, with a natural direct effect of $\Delta_1 (1,0,M(1),M(1))=E[Y_1(1,M(1))-Y_1(0,M(1))|D=1, T=1]=1+M(1)=1.919$ and a natural indirect effect of $\Delta_1 (0,0,M(1),M(0))=E[Y_1(0,M(1))-Y_1(0,M(0))|D=1, T=1]=M(1)-M(0)=0.5$. For each estimator and sample size, the table reports the bias (`bias'), standard deviation (`std'), root mean squared error (`rmse'), average standard error (`avse'), and the coverage rate of the 95% confidence intervals (`cover').
For $n=2000$, all estimators exhibit only small finite-sample bias. A direct comparison in the table suggests that the natural indirect effect is least biased; however, because its true effect size is smaller than those of the natural direct and total effects, its bias is more pronounced relative to the magnitude of the true effect. Root mean squared errors are moderate across all estimators. Regarding statistical inference, the average standard errors tend to be slightly smaller than the corresponding empirical standard deviations, resulting in mild undercoverage of the 95% confidence intervals. This indicates that statistical inference is slightly optimistic in finite samples. When the sample size is increased to $n=8000$, the standard deviations are approximately halved, which suggests that the estimators converge to the true effects at an $n^{-1/2}$ rate.
Next, we focus on dynamic treatment effects and controlled direct effects. To this end, we slightly modify the DGP. In contrast to the previous setting with a continuous mediator, we now assume that the mediator is binary. Specifically, the mediator equation in the DGP is replaced by \[ M = I\{X\beta+0.5\cdot D+V_m>0\}. \] This modification ensures that the mediator levels $M=m$ and $M=m'$ used to define the dynamic treatment effect and the controlled direct effect occur with positive probability. All other components of the DGP remain unchanged. As before, we consider sample sizes of $n=2000$ and $n=8000$ and generate 1000 independent simulation replications. The effect of interest is the ATET defined in equation (ref). Estimation is based on the normalized sample analog of equation (ref) in Proposition (ref). In the simulation, we examine three scenarios that differ in the values at which $d$, $d'$, $m$, and $m'$ are fixed. Specifically, we estimate (i) the ATET of receiving both the treatment and the mediator versus receiving neither, (ii) the controlled direct effect under $m=m'=0$, and (iii) the controlled direct effect under $m=m'=1$.
The estimation procedure mirrors that used for the natural direct and indirect effects. We again employ four-fold cross-fitting with nuisance parameters estimated via cross-validated lasso, using logit specifications for the propensity scores and linear specifications for the conditional mean outcomes. The trimming rule and statistical inference are implemented as before. The lower panel of Table (ref) provides the simulation results for estimating the ATET at different values of $d$, $d'$, $m$, and $m'$. The true effects equal $\Delta_{1,1}(1,1,0,0) = E[Y_1(1,1) - Y_1(0,0) \mid D=1, M=1, T=1]=4-1=3$, $\Delta_{1,1}(1,0,0,0) = E[Y_1(1,0) - Y_1(0,0) \mid D=1, M=1, T=1]=2-1=1$ and $\Delta_{1,1}(1,1,0,1) = E[Y_1(1,1) - Y_1(0,1) \mid D=1, M=1, T=1]=3-1=2$ in the three scenarios considered.
The results are broadly similar to those for the natural direct and indirect effects in the upper panel. For $n=2000$, all estimators again exhibit small finite-sample bias, although their standard deviations and RMSEs are somewhat larger in absolute terms than those of the natural effects. Coverage rates remain slightly below 95%, reflecting that the average standard errors tend to be slightly smaller than the standard deviations. For $n=8000$, the standard deviations are approximately half as large as those at $n=2000$, as expected for estimators that converge to the true effects at an $n^{-1/2}$ rate. In terms of nearly all the performance metrics considered, the results for the controlled direct effect under $m=m'=0$ are somewhat worse than in the other two scenarios. This likely reflects that, under our DGP, the treatment–mediator combination $d=1$ and $m=0$ occurs least frequently, so support is weakest in this case.
In this section, we illustrate our approach by revisiting the empirical application of farbmacher2022causal using data from the National Longitudinal Survey of Youth 1997 (NLSY97), a nationally representative U.S. sample of 8984 respondents who were ages 12--17 at the first interview in 1997\footnote{For additional information on the NLSY97 sample, see U.S. Bureau of Labor Statistics nlsy97}. farbmacher2022causal examine the causal effect of health care coverage ($D$) on general health ($Y$) and decompose this effect into an indirect effect operating through the incidence of routine checkups ($M$) and a direct effect capturing all remaining causal mechanisms. While we apply our method to this same research question, our analysis differs in two important respects. First, instead of the average treatment effect (ATE), we estimate the average treatment effect on the treated (ATET). Second, our identification strategy differs. The approach of farbmacher2022causal relies on a selection-on-observables assumption, whereas our DiD framework instead invokes the conditional common trend assumptions stated in Assumptions (ref) and (ref).
We largely follow farbmacher2022causal in the definition of the main variables. The treatment is a binary indicator equal to one if the respondent had any kind of health care coverage in 2006 and zero otherwise. The outcome variable is self-reported general health, recorded on a five-point ordinal scale with categories excellent, very good, good, fair, and poor, where higher values correspond to poorer health status. We observe the pre-treatment outcome in 2005 and the post-treatment outcome in 2008. The mediator captures whether the respondent visited a doctor for a routine checkup within the past 12 months and is measured in 2007, i.e., after treatment but before the post-treatment outcome. farbmacher2022causal control for a rich set of demographic, socioeconomic, household, and health-related variables. We use the same set of control variables, including interactions and higher-order terms, to ensure comparability. One exception arises due to our different research design: we do not include the lagged outcome variables used in farbmacher2022causal among the control variables. All interactions involving these variables are excluded as well.
Our DiD framework affects sample construction in two important ways. First, we require respondents to report their general health in both 2005 and 2008. Second, we exclude respondents who had any form of health care coverage prior to 2006. As a result, both the size and composition of our final sample differ markedly from those in farbmacher2022causal. Our sample is considerably smaller, comprising 1020 observations compared to the 7486 observations in their study. To illustrate the differences in sample composition, we report descriptive statistics using the same selection of control variables as in farbmacher2022causal. The variables displayed in Table (ref) represent only a subset of the nearly 1000 control variables included in our dataset. While farbmacher2022causal find that treated and untreated respondents differ substantially with respect to the variables reported in the table, we find no statistically significant differences between the two groups for the majority of these variables in our sample. However, there are some notable exceptions. For instance, women, respondents living in urban areas, and more highly educated individuals are more likely to have health care coverage. A similar pattern emerges for the mediator. farbmacher2022causal again report substantial differences between respondents who underwent a routine checkup and those who did not, whereas in our sample the two groups are largely comparable across the variables shown in Table (ref). Statistically significant differences arise only for a limited number of characteristics, such as gender, highest completed grade, certain ethnicity categories (`Black' and `White or Other'), and the number of household members under age 18.
Given that the NLSY97 is a panel dataset, we implement the panel-data version of our DR estimator based on equations (ref) and (ref) of Propositions (ref) and (ref) to assess the total, natural direct, and natural indirect effects of health care coverage on self-reported general health. As in the simulations, our estimators are based on four-fold cross-fitting. The nuisance parameters are estimated by cross-validated lasso regression. Following farbmacher2022causal, we set the trimming threshold to 0.02 (2%). However, no observations are dropped due to this trimming rule.
Table (ref) reports the estimated effects, standard errors, and $p$-values. The ATET is not statistically significant at any conventional level, nor are the natural direct and natural indirect effects. For individuals who obtain health care coverage, we find no statistically significant evidence of a short-term effect of coverage on general health, either through routine checkups or via other causal pathways. Nevertheless, the point estimates of both the total and direct effects are negative, indicating a health-improving effect, since lower values of the outcome correspond to better health. The results are broadly comparable to those reported by farbmacher2022causal. In their application and depending on the estimation approach, most of the natural direct and indirect effects are not statistically significant. However, in contrast to our ATET, their ATE is statistically significant at the 10% or 5% level, depending on the estimation approach. Regarding effect magnitudes, the total effect and the natural direct effect are of comparable size in both applications. Our point estimates are slightly larger in absolute value than those obtained by farbmacher2022causal; however, the corresponding standard errors are also larger, indicating lower precision. Although the natural indirect effects are of similar magnitude and close to zero in both applications, the sign of the estimated effect differs, being slightly positive in our analysis and negative in farbmacher2022causal. Overall, despite relying on a different identification strategy and focusing on the ATET rather than the ATE, our findings point in the same substantive direction as those of farbmacher2022causal.
This paper suggests a flexible difference-in-differences (DiD) framework for causal mediation and dynamic treatment effect analysis. By extending the DiD design beyond total treatment effects, our approach allows evaluating controlled direct effects, natural direct and indirect effects operating through mediators, and joint treatment–mediator (or dynamic treatment) effects. Identification of controlled direct effects and dynamic treatment effects is achieved under conditional parallel trends assumptions on mean potential outcomes across treatment-mediator states. For the identification of natural direct and indirect effects, we impose additional parallel trend assumptions on mean potential outcomes or mediator distributions across treatment states.
On the estimation side, we propose doubly robust DiD estimators within the double machine learning framework for both repeated cross sections and panel data. By combining machine learning–based nuisance estimation with Neyman-orthogonal score functions and cross-fitting, our estimators allow for data-driven covariate adjustment while retaining valid large-sample inference under specific regularity conditions. We establish asymptotic normality of our methods and demonstrate their favorable finite-sample performance in a simulation study. We also provide an empirical application using data from the National Longitudinal Survey of Youth 1997 to disentangle the direct effect of health care coverage on general health from the indirect effect operating through routine checkups. Despite point estimates of both the total and direct effects suggesting health improvements, we find no statistically significant short-term effect of health care coverage on general health among those who obtain coverage, either via routine checkups or other causal pathways.
{ \setcounter{equation}{0}