Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
99,778 characters · 22 sections · 63 citation commands
Balancing Weights for Causal Mediation Analysis
Keywords: causal mediation analysis; natural direct and indirect effects; propensity score instability; balancing weights; finite--sample covariate imbalance.
Causal mediation analysis has become an essential framework for researchers who seek to understand the mechanisms of a causal effect. Many traditional studies focus on estimating the average treatment effect (ATE). However, researchers are now increasingly attempting to uncover the underlying causal pathways by decomposing the ATE into its mediated and unmediated components. This is important because improving policies or interventions often requires knowing why an effect exists, not simply confirming its existence. Reflecting this practical importance, this framework has become widespread across various disciplines, with both theoretical and empirical studies actively pursued (Imai-GeneralApproachAnalysis-2010o, Imai-IdentificationInferenceEffects-2010a, Imai-UnpackingBlackStudies-2011p, VanderWeele-ExplanationCausalInteraction-2015c, Huber-MediationAnalysis-2020y, Celli-CausalMediationModels-2022i, Nguyen-ClarifyingCausalLearn-2020r, Nguyen-ClarifyingCausalOutcomes-2022p). This paper contributes to this field by developing a novel weighting estimator for causal mediation analysis that enhances the finite--sample performance.
We introduce notation and define the estimands. Let $\mathcal{X}$, $\mathcal{M}$, and $\mathcal{Y}$ denote the supports of the covariates, mediator, and outcome, respectively. The observed data are $(Y,M,D,X)\in\mathcal{Y}\times\mathcal{M}\times\{0,1\}\times\mathcal{X}$. We use potential outcomes to formalize causal mechanisms. For $d\in\{0,1\}$, let $M_d$ denote the mediator that would be observed if $D$ were set to $d$, and let $Y_{dm}$ denote the outcome that would be observed if $(D,M)$ were set to $(d,m)$. We write $Y_d:=Y_{dM_d}$ and, critically for causal mediation analysis, $Y_{dM_{d'}}$ for the cross-world outcome where the treatment is $d$ while the mediator follows the distribution it would have under $d'$.
We consider two key estimands in causal mediation analysis: the natural direct effect (NDE) and the natural indirect effect (NIE). The NDE, defined as $\text{NDE}(d):=\mathbb{E}[Y_{1M_d} - Y_{0M_d}]$, measures the causal effect of changing the treatment from $0$ to $1$ while fixing the mediator to the distribution it would take under $D=d$. In this way, the NDE captures the part of the treatment effect that does not operate through the mediator and reflects only the direct path from treatment to outcome. The NIE, defined as $\text{NIE}(d):= \mathbb{E}[Y_{dM_1} - Y_{dM_0}]$, measures the effect of the treatment through the mediator on the outcome. The $\text{ATE} \, (:= \mathbb{E}[Y_1 - Y_0])$, can be decomposed as $\text{ATE} = \text{NDE}(1) + \text{NIE}(0) = \text{NIE}(1) + \text{NDE}(0)$, which are depicted in Figure (ref). Accordingly, we are interested in the targets $\theta_{d,1-d} := \mathbb{E}[Y_{dM_{1-d}}], \theta_d := \mathbb{E}[Y_d], d\in\{0,1\}$.
We focus on two main weighting estimators for causal mediation analysis: the inverse probability weighting (IPW) estimator and its augmented counterpart, the efficient influence function (EIF)–based estimator. The IPW estimator of Huber-IdentifyingCausalMechanisms-2014x for $\theta_{1,0}$ is given by
where $\widehat{\xi}_{d}(M,X)$ and $\widehat{\pi}_{d}(X)$ are estimators of $P(D=d\mid M,X)$ and $P(D=d\mid X)$, respectively. Since this weight also plays a role in the EIF–based estimator discussed below, we denote it by $ \widehat{w}^{\mathrm{EIF}}_2(M_i,X_i) = \widehat{\xi}_{0}(M_i,X_i)/\widehat{\pi}_{0}(X_i)\widehat{\xi}_{1}(M_i,X_i). $ Notably, these propensity score estimators appear in the denominator, which gives rise to the issues of interest. We also consider an augmented version, the EIF–based estimator of TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y, as introduced in (ref), which combines weighting with outcome regression (regression imputation). Specifically, the EIF–based estimator for $\theta_{1,0}$ employs not only $\widehat{w}^{\mathrm{EIF}}_2(M_i, X_i)$ but also $ \widehat{w}^{\mathrm{EIF}}_1(X_i) = 1/\widehat{\pi}_0(X_i), $ which also depends on the inverse of the estimated propensity score.\footnote{The subscripts 1 and 2 in $\widehat{w}^{\mathrm{EIF}}_1(X_i)$ and $\widehat{w}^{\mathrm{EIF}}_2(M_i,X_i)$ correspond to the discussion of the proposed two-step algorithm in Section (ref).}
Both the IPW estimator and the EIF--based estimator rely on the weights based on the inverse of the estimated propensity scores and are therefore susceptible to two key issues noted in previous studies: (i) instability and (ii) finite-sample covariate imbalance (Hirshberg-ApproachesWeightingInference-2017l, Chattopadhyay-BalancingVsPractice-2020m, Ben-Michael-BalancingActInference-2021x, Cohn-BalancingWeightsInference-2023j).
To directly address these two concerns, we propose alternative weighting methods by extending balancing weights to causal mediation analysis. In particular, we build on the “minimal--dispersion approximate balancing weights” (hereafter, minimal weights) of Wang-MinimalDispersionConsiderations-2019m, as this framework encompasses several existing methods as special cases, including entropy balancing Hainmueller-EntropyBalancingStudies-2012c, stable balancing weights Zubizarreta-StableWeightsData-2015v, and empirical balancing calibration weights Chan-GloballyEfficientWeighting-2016v. The method’s name describes the two simultaneous objectives of the underlying optimization problem: “minimal dispersion”, i.e.\ minimizing a measure of weight dispersion such as a squared loss or an entropy‐based loss, and “approximate balancing”, i.e.\ enforcing covariate balance up to pre-specified tolerances. In this way, the method addresses both of the issues discussed above—(i) instability and (ii) finite-sample covariate imbalance. At the same time, it removes the need for ad hoc post–estimation covariate balance diagnostics or trimming of extreme weights, both of which are commonly applied in practice.
We propose a “two-step minimal-dispersion approximate balancing weights ”method (hereafter, two-step minimal weights) for causal mediation. The name “two-step” reflects our strategy to first balance the marginal distribution of the covariates, and then the joint distribution of the covariates and the mediator. This sequential approach is motivated by our analysis in Section (ref), which reveals that imbalances in these specific distributions generate the bias. While a similar decomposition of the bias source has been recently discussed by liu2024two, our study leverages this insight to construct the different weighting algorithm. For clarity, we denote the resulting two-step minimal weights by $\widehat{w}^{\mathrm{MW}}_1(X_i)$ and $\widehat{w}^{\mathrm{MW}}_2(M_i,X_i)$.
As a theoretical contribution, we establish the consistency of the weights $\widehat{w}^{\mathrm{MW}}_1(X_i)$ and $\widehat{w}^{\mathrm{MW}}_2(M_i, X_i)$ and the asymptotic normality and semiparametric efficiency of the resulting estimators: IPW estimator and EIF--based estimator with weights $\widehat{w}^{\mathrm{MW}}_1(X_i)$ and $\widehat{w}^{\mathrm{MW}}_2(M_i, X_i)$. Firstly, by considering a nonparametric setting in which the number of balancing constraints increases with the sample size, we show that the optimization-based weights $\widehat{w}^{\mathrm{MW}}_1(X_i)$ and $\widehat{w}^{\mathrm{MW}}_2(M_i, X_i)$ converge, respectively, to $\xi_{0}(M_i, X_i) / \pi_{0}(X_i) \xi_{1}(M_i, X_i)$ and $1/\pi_0(X_i)$. The proof of this result relies on the dual formulation and is derived under standard M--estimation assumptions. We also impose conditions that are standard in sieve estimation (see, e.g., Newey-ConvergenceRatesEstimators-1997k; Chen-Chapter76Models-2007p), and our asymptotic arguments parallel those in the balancing-weight literature (e.g., Wang-MinimalDispersionConsiderations-2019m; Fan-OptimalCovariateEstimation-2023s). Secondly, under additional assumptions such as the complexity of the function class, we establish that the estimators are asymptotically normally distributed with variance achieving the semiparametric efficiency bound. These results imply that, under the stated conditions, the proposed estimators are asymptotically equivalent to those based on $\widehat{w}^{\mathrm{EIF}}_1(X_i)$ and $\widehat{w}^{\mathrm{EIF}}_2(M_i, X_i)$. We therefore conclude that the proposed method enhances finite-sample performance while preserving its desirable asymptotic properties.
The proofs of these asymptotic properties are provided in the Appendix. While we employ techniques similar to those in Wang-MinimalDispersionConsiderations-2019m for the ATE, extending them to causal mediation analysis is non-trivial. Specifically, a key challenge lies in controlling the influence of the first-step weight estimation on the second step. We address this issue by exploiting the structure of the efficient influence function, building on insights from recent work such as Farbmacher-CausalMediationLearning-2022c.
To evaluate the finite–sample performance of our method, we conduct extensive simulations based on two settings from TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y and Wong-Kernel-basedCovariateStudies-2018k, under both (A) correct specification and (B) misspecification. We compare our approach with three baseline estimators: the IPW estimator, the EIF–based estimator, and the regression imputation estimator. A concise summary of the simulation results is reported in Table (ref). Overall, the method performs as well as or better than the baseline estimators. In addition, we implement other competing weighting schemes, including covariate balancing propensity scores (CBPS)\footnote{The CBPS method was originally proposed by Imai-CovariateBalancingScore-2014v. Its extension to causal mediation analysis is provided in Appendix (ref) of this paper.} and an oracle variant based on the true propensity scores. The proposed method further outperforms these alternatives, confirming its broad advantages.
We illustrate the practical value of our approach using the media–framing data of brader2008triggers. To focus squarely on the role of weighting quality, we consider two generic estimators: an IPW–type estimator obtained by plugging arbitrary weights into estimator ((ref)), and an EIF–type estimator defined as its augmented counterpart with outcome regression. In particular, we show that using our proposed minimal weights can reduce standard errors in both the EIF–type and IPW–type estimators, thereby providing clearer results.
This paper contributes to both the literature on causal mediation analysis and the literature on balancing weights. The key references we build upon for the estimation of the NDE and NIE are TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y and Huber-IdentifyingCausalMechanisms-2014x. For a review of estimators, including weighting estimators and regression estimators, see Nguyen-CausalMediationEffects-2023j. While Nguyen-CausalMediationEffects-2023j discusses the general idea of achieving balance between covariates and the mediator, it does not explore concrete balancing-weight algorithms or their theoretical properties, which are the focus of the present paper.
While the literature on balancing weights is extensive for the ATE Hainmueller-EntropyBalancingStudies-2012c, Zubizarreta-StableWeightsData-2015v, Chan-GloballyEfficientWeighting-2016v, Hirshberg-ApproachesWeightingInference-2017l, Chattopadhyay-BalancingVsPractice-2020m, Ben-Michael-BalancingActInference-2021x, Cohn-BalancingWeightsInference-2023j, its extension to causal mediation analysis remains limited. Although several optimization-based weighting approaches have been proposed recently, our method remains distinct. For instance, Huang-NonparametricEstimationTreatment-2024x, building on Ai-UnifiedFrameworkModels-2021c, proposed estimating nonparametric propensity scores by minimizing entropy loss. However, their approach does not explicitly address the issue of (ii) finite-sample covariate imbalance. liu2024two introduced a minimax estimation framework following dikkala2020minimax that accounts for covariate imbalance, but their algorithmic approach differs fundamentally from ours. Our method is specifically designed to simultaneously tackle two key issues: (i) instability and (ii) finite-sample covariate imbalance. In this regard, the work most closely related to ours is Chan-EfficientNonparametricEffects-2016g, which also seeks to determine weights by minimizing a distance measure. However, our algorithm diverges by incorporating approximate balancing constraints to explicitly optimize the bias--variance trade-off. Furthermore, we establish our asymptotic properties under weaker assumptions than those required in Chan-EfficientNonparametricEffects-2016g.
The remainder of this paper is organized as follows. Section (ref) introduces the basic setup and the identification of the key estimands. Section (ref) presents the EIF--based and IPW estimators and provides a formal bias‐decomposition analysis. Section (ref) proposes the two‐-step minimal‐-weights algorithm. Section (ref) establishes the asymptotic properties of the proposed estimators. Section (ref) presents the results of simulation studies. Section (ref) then reports an empirical application. Finally, Section (ref) concludes the paper and outlines directions for future research. Implementation notes and proofs of the main theorems are provided in the Appendix.
This section introduced notations and assumptions for identification. We state the standard consistency, sequential ignorability, and positivity assumptions, and present identification formulas for $\theta_{d,1-d}$ and $\theta_d$.
To facilitate the subsequent discussion, we introduce the following notations.
$\theta_{d, 1-d}$ and $\theta_{d}$ for $d \in \{0,1\}$ are identified by the following theorem.
For the proof of Theorem (ref), see TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y. The identification can be established through several complementary results:
Theorem (ref) can be viewed as a combination of these results. Intuitively, identification can be achieved either through weighting alone or through outcome expectations alone, while their combination naturally yields an alternative representation. Indeed, the expression
constitutes an efficient influence function for $\theta_{d,1-d}$. The derivation of the efficient influence function for $\mathbb{E}[Y_{dM_{1-d}}]$ can be found in TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y.\footnote{In fact, the estimator in (ref) is an alternative expression of that in TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y, where the original form is given by
In this work, we adopt expression (ref) to avoid estimating the mediator's density and to facilitate the discussion of the covariate balancing propensity score method in Appendix (ref). The estimator (ref) is derived by a straightforward application of Bayes' rule to (ref) and thus retains the properties described in TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y.} The estimator based on this representation is therefore referred to as the EIF-based estimator.
We now provide a more detailed explanation of the assumptions for identification. Consistency assumes that the observed mediator and outcome are equal to the potential mediator and outcome under the observed treatment and mediator assignments. The sequential ignorability assumption consists of the following four conditional independences. First, $ Y_{dm} \mathrel{\perp\!\!\!\perp} D \mid X, $ which ensures that the effect of treatment \( D \) on the outcome \( Y \) is unconfounded given \( X \). Second, $ M_{d} \mathrel{\perp\!\!\!\perp} D \mid X, $ implying that the effect of treatment \( D \) on the mediator \( M \) is unconfounded given \( X \). Third, $ Y_{dm} \mathrel{\perp\!\!\!\perp} M \mid D = d, X, $ which states that the effect of the mediator \( M \) on the outcome is unconfounded given \( X \). Finally, the so-called cross-world independence assumption $ Y_{dm} \mathrel{\perp\!\!\!\perp} M_{1-d} \mid X $ is crucial. It ensures the absence of confounding between the mediator and the outcome that is induced by the treatment. Such a confounder, often called a treatment-induced confounder (denoted by $L$ in Figure (ref)), is a variable affected by $D$ that subsequently affects both $M$ and $Y$. For further details on the sequential ignorability assumptions, see, for example, Pearl-DirectIndirectEffects-2001g, Imai-IdentificationInferenceEffects-2010a, Vanderweele-EffectDecompositionConfounder-2014z. Positivity assumption requires that each treatment and mediator level has a nonzero probability for every given value of the covariates.
We now proceed from identification to estimation. Estimating the NDE and NIE requires estimating the quantity $\mathbb{E}[Y_{dM_{1-d}}]$. This estimand is unique to causal mediation because it relies on the potential outcome under two treatment statuses. For notational simplicity, we will henceforth focus on estimating $\theta_{1,0} := \mathbb{E}[Y_{1M_0}]$.\footnote{All the following analysis applies similarly to $ \theta_{0,1} := \mathbb{E}[Y_{0M_1}].$}
This section treats the two canonical weighting estimators for $\theta_{1,0}$: the EIF-based estimator and the IPW estimator introduced in Section (ref). In addition, we introduce the IPW--type estimator and the EIF--type estimator, which replace the weights based on the inverse of the estimated propensity scores with a set of general weights. For both the EIF--type and IPW--type estimators, we provide a finite--sample bias decomposition. This decomposition formally demonstrates the importance of addressing (ii) finite-sample covariate imbalance, one of the two concerns highlighted in Section (ref), alongside (i) instability.
The EIF–based estimator is given by
where
We denote $\widehat{\pi}_0(X_i)$, $\widehat{\xi}_d(M_i,X_i)$, $\widehat{\mu}_1(M_i,X_i)$, and $\widehat{\eta}_{1,0}(X_i)$ as estimators of $\pi_0(X_i)$, $\xi_d(M_i,X_i)$ for $d\in\{0,1\}$, $\mu_1(M_i,X_i)$, and $\eta_{1,0}(X_i)$, respectively.
One of the key properties of the EIF--based estimator is a semiparametric efficiency.
For the proof, see TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y. In Section (ref), we show that our proposed estimator also attains this semiparametric efficiency bound under certain regularity conditions \footnote{Another key property of the EIF--based estimator is its multiple robustness.
If one of the following conditions hold:
then the estimator \(\widehat{\theta}^{\mathrm{EIF}}_{1,0}\) is consistent.
This result implies that the estimator remains consistent even when some nuisance parameters are misspecified. Although multiple robustness is not the main focus of this paper, we can show, similar to TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y, that when the balancing constraints for the weights are correctly specified, the estimator remains consistent even if the outcome models are misspecified. In fact, in Section (ref), we confirm this consistency under outcome model misspecification.}.
Although the EIF--based estimator is multiple robust and semiparametrically efficient, its reliance on the inverse of propensity scores leads to the aforementioned issues of instability and covariate imbalance. To formally investigate the latter, we consider a general EIF--type weighting estimator, which replaces the specific EIF weights \( \widehat{w}^{\text{EIF}}_2(M_i,X_i) \) and \( \widehat{w}^{\text{EIF}}_1(X_i) \) with generic weights \( \widehat{w}_2(M_i,X_i) \) and \( \widehat{w}_1(X_i) \):
Decomposing the bias of this estimator requires defining two sets of terms. First, we define the estimation errors in the nuisance functions as \[ \widetilde{\mu}_1(M_i,X_i) := \mu_1(M_i,X_i) - \widehat{\mu}_1(M_i,X_i), \quad \widetilde{\eta}_{1,0}(X_i) := \eta_{1,0}(X_i) - \widehat{\eta}_{1,0}(X_i). \] Second, we define the random errors as \[ u_i := Y_i - \mu_1(M_i,X_i), \quad v_i := \mu_1(M_i,X_i) - \eta_{1,0}(X_i). \] With these components defined, we can formally state the sources of the estimator's bias in the following proposition.
We prove Proposition (ref). First, the difference between the EIF--type weighting estimator and the target parameter can be decomposed as:
The first two terms reflect imbalances in \( \widetilde{\mu}_1(M_i,X_i) \) and \( \widetilde{\eta}_{1,0}(X_i) \). Imbalance in \( \widetilde{\mu}_1(M_i,X_i) \) arises when the weights \( \widehat{w}_2(M_i,X_i) \) do not balance the distributions of the mediator and covariates between the treatment group and the control group weighted by \( \widehat{w}_1(X_i) \). Similarly, imbalance in \( \widetilde{\eta}_{1,0}(X_i) \) arises when the weights \( \widehat{w}_1(X_i) \) fail to balance the covariate distributions between the control group and the full sample. The third and fourth terms capture the errors \( u_i \) and \( v_i \), while the last term represents the sampling variation of \( \eta_{1,0}(X_i) \).
Next, by taking expectations on both sides, we show that the last three terms are equal to $0$. First, note that $\mathbb{E}[\eta_{1,0}(X)] = \theta_{1,0}$. For the first noise term, using the law of total expectation, we have:
Similarly, for the second noise term,
This completes the proof of Proposition (ref). The key insight from this decomposition is that the finite--sample bias stems solely from imbalances in the estimation errors, $\widetilde{\mu}_{1}(M,X)$ and $\widetilde{\eta}_{1,0}(X)$. Our proposed weighting scheme is therefore designed to directly target this source of bias.
Furthermore, this bias decomposition sheds light on another property of the EIF--based estimator (ref): it is asymptotically unbiased due to its population covariate balance.
The proof is deferred to Appendix (ref). Proposition (ref) shows that the moment equalities hold for arbitrary measurable functions \( f(M,X) \) and \(g(X)\). Consequently, as long as \(\xi_d(M,X)\) and \( \pi_d(X) \) are consistently estimated, the EIF--based estimator (ref) remains asymptotically unbiased, regardless of the specification of \( \widetilde{\mu}(M,X) \) and \( \widetilde{\eta}(X) \). However, the EIF weights \( \widehat{w}^{\text{EIF}}_1(X_i) \) and \( \widehat{w}^{\text{EIF}}_2(M_i,X_i) \) are not constructed to ensure finite--sample covariate balance, which introduces bias as Proposition (ref) implies.
To complement the EIF--type estimator, which leverages an outcome model, we introduce an alternative that relies exclusively on weights to estimate the NDE and NIE. This estimator is particularly useful for isolating and evaluating the performance of the weighting scheme itself. We define this estimator here and show that its finite--sample bias can be decomposed in a manner analogous to that of the EIF--type estimator. The IPW–type estimator is defined as: $$ \widehat{\theta}_{1,0}^{\mathrm{IPW-type}} = \frac{1}{n}\sum_{i=1}^n D_i \widehat{w}_2(M_i,X_i) Y_i. $$ Following a similar process as in the proof of Proposition (ref), the IPW–type estimator has the following bias decomposition.
The difference between this bias decompositon and that of the EIF--type estimator (Proposition (ref)) is that the bias of the IPW--type estimator stems from the imbalance of the true $\mu_1(M_i,X_i)$ and $\eta_{1,0}(X_i)$, rather than its estimation error $\widetilde{\mu}_1(M_i,X_i)$ and $\widetilde{\eta}_{1,0}(X_i)$.
Therefore, for both the IPW–type and EIF–type estimators, it is crucial that the weights balance mediators and covariates, such as those included in the specifications of $\mu_1(M_i,X_i)$ and $\eta_{1,0}(X_i)$.
To address the issues of (i) instability and (ii) finite-sample covariate balance in weighting estimators discussed previously, this section introduces an optimization-based approach. Applying the “minimal weight” method Wang-MinimalDispersionConsiderations-2019m to causal mediation analysis requires an adaptation to account for the two distinct sources of bias identified in Section (ref). To this end, we propose the “two--step minimal weights” method. This algorithm constructs weights by solving an optimization problem, specifically designed to sequentially correct the two types of imbalances inherent in causal mediation analysis.
In what follows, we first formally present the two--step algorithm. We then introduce a data-driven procedure for tuning its key tolerance hyperparameters which govern the trade-off between bias and variance.
We propose the following two--step algorithm to derive weights for the weighting estimators of \(\theta_{1,0}\).
To provide an overview, Algorithm (ref) is tailored to causal mediation by coupling a mediation-aware balance design with an explicit stabilization mechanism. It adopts a sequential scheme that first balances $X$ in Step 1 (control versus full sample) and then balances $(M,X)$ in Step 2 (treated versus reweighted control), reflecting the two imbalances that arise in our bias decomposition in Proposition (ref) and Proposition (ref). Each step minimizes a convex dispersion penalty $f(\cdot)$ subject to approximate balance constraints governed by tolerances $\{\varepsilon_j\}_{j=1}^K$ and $\{\delta_j\}_{j=1}^L$.
To begin, we discuss the choice of the loss function. Typical choices for $f$ include entropy balancing ($f(w)=w\log w$) and stable balancing weights ($f(w)=(w-1/n_0)^2$, with $n_0$ the number of controls), both of which tame the dispersion of inverse-propensity-type weights while preserving the targeted balance moments.
Next, we turn to the balancing constraints. In Step 1, the constraints are applied to the target functions $\{c_j(X)\}_{j=1}^K$, reflecting the assumption that $\eta_{1,0}(X)$ (or its estimation error $\widetilde{\eta}_{1,0}(X)$) can be well approximated by a linear combination of these bases. Setting $\varepsilon_j = 0$ corresponds to an exact sample balance.
In practice, to normalize the weights, one may add $c_{K+1}(X_i) = 1$ with $\varepsilon_{K+1} = 0$. The non-negativity condition of the weights is typically achieved either by directly including it as a constraint or by an appropriate choice of loss function, which we explain when presenting Theorem (ref) that derives the dual form of the Algorithm (ref). Step 2 applies the same procedure to $\{b_j(M,X)\}_{j=1}^L$, aligning the treated group with the reweighted control group on the joint $(M,X)$ distribution.
Finally, the choice of hyperparameters $\{\varepsilon_j\}_{j=1}^K$ and $\{\delta_j\}_{j=1}^L$ which determine the bias--variance trade-off, is discussed in the next subsection.
Then, we can construct the EIF‐type weighting estimator and IPW–type weighting estimator with normalized two--step minimal weights:
The properties of these estimators are examined in Section (ref). In our subsequent simulation and empirical application sections, we will employ both estimators to evaluate the properties of the proposed weights.
One of the key features of our proposed method is the use of tolerances $\{\varepsilon_j\}_{j=1}^K$ and $\{\delta_j\}_{j=1}^L$, which allow for approximate rather than exact covariate balance. Since the weights are constructed without using outcome data, standard cross-validation is not applicable. Therefore, we adapt the procedure from Wang-MinimalDispersionConsiderations-2019m, which selects the tolerances by evaluating the stability of covariate balance across bootstrap samples, as detailed in Algorithm (ref).
The core idea of this algorithm is to select the tolerances that yield the most stable covariate balance across multiple bootstrap resamples of the data. By averaging the imbalance measure over these resamples, the procedure aims to find hyperparameters that perform well not only on the original sample but also under bootstrap resampling.
However, there are a couple of downsides to this algorithm. First, the sequential optimization can sometimes get stuck in a local optimum. For example, even if the first stage selects an $\varepsilon^*$ that provides good balance for the $\{c_k(X)\}_{k=1}^K$ basis functions, this does not necessarily ensure good balance for the $\{b_l(M,X)\}_{l=1}^L$ basis functions in the second stage. Second, the algorithm applies the same $\varepsilon^*$ and $\delta^*$ to all basis functions. Ideally, we would prefer to balance covariates that are more strongly related to the outcome, but the fundamental difficulty is that doing so would make the weights outcome-dependent.
In this section, we establish the key asymptotic properties of our proposed two--step minimal weights. We begin by deriving the convergence rates of the weights. Building on this result, we then prove the main theoretical claim of the paper: that the proposed estimator is consistent, asymptotically normal, and attains the semiparametric efficiency bound. We note that both the EIF--type and IPW--type estimators exhibit the same asymptotic properties under nearly identical assumptions. Thus, we conclude that the proposed method improves finite-sample behavior without compromising asymptotic guarantees.
First, we investigate the convergence rates of the two--step minimal weights derived in Algorithm (ref). For the weights derived in the first step optimization problem (ref) and (ref), Theorem 2 of Wang-MinimalDispersionConsiderations-2019m directly yields the convergence rates. Specifically, it has been established that $n \widehat{w}^{\mathrm{MW}}_1(x)$ consistently estimates $1/\pi_0(x)$. For completeness, we describe the assumption and the result in the Appendix (ref). In parallel to their theorem, we obtain the following results for the weights derived in the second step optimization problem (ref) and (ref).
We begin by deriving the dual of the optimization problem.
The proof is provided in Appendix (ref). First, note that problem (ref) can be viewed as a shrinkage estimation problem. The weights are estimated by a generalized linear model on the basis functions $B(M_i,X_i)$ with link function $\rho'$. The dual variables in $\lambda$ can be interpreted as the coefficients of the basis functions. From an optimization perspective, they represent the shadow prices of the covariate balance constraints: a larger dual variable indicates that balancing the associated covariate is more important or more difficult, and that relaxing the constraint would lead to a larger reduction in the objective function. The inclusion of an $\ell_1$ penalty further mitigates the influence of covariates that are particularly hard to balance, thereby preventing the weights from becoming overly dependent on them. In addition, this formulation helps ensure the nonnegativity of the weights. For example, if we set $f(w) = w \log w$, then $\rho'(t) = \exp(t - 1)$, which guarantees that the resulting weights are nonnegative.
By utilizing Theorem (ref), we establish the convergence rates of the weights. To this end, the following conditions are assumed.
These assumptions are analogous to Assumption 1 of Wang-MinimalDispersionConsiderations-2019m. In Assumption (ref), conditions (ref) represent standard regularity assumptions that ensure the consistency of M-estimators. Given the result that $\widehat{m}_1^{\mathrm{MW}}(X)$ is consistent for $1/\pi_0(X)$, and using the first-order condition together with Proposition (ref), we can deduce that there exists a $\lambda^\circ$ such that $w^*_2(m,x) = n \rho'(B(m,x)^\top \lambda^\circ)$. Condition (ref) enables the consistency of $\lambda^\dagger$ to imply the consistency of the estimated weights. Noting that $\rho'(t) = (f')^{-1}(t)$ by a simple calculation, it is clear that this condition is met by standard choices of $f$ used in the definition of two--step minimal weights (ref), such as the squared loss $f(w) = w^2$ and the entropy-based loss $f(w) = w \log w$. Condition (ref) imposes a technical restriction on the magnitude of the basis functions. Condition (ref) controls the growth rate of the number of basis functions relative to the sample size. Condition (ref) ensures that the weights $w^*_2(m,x)$ can be uniformly approximated, which typically requires that the basis $B(m,x)$ is sufficiently rich. Conditions (ref)--(ref) are satisfied by a wide range of basis functions, including regression splines, trigonometric polynomials, and wavelets (see, e.g., Newey-ConvergenceRatesEstimators-1997k; Chen-Chapter76Models-2007p). Lastly, condition (ref) characterizes the degree to which the equality constraints for exact covariate balance may be relaxed by hyperparameters without compromising the asymptotic properties of the estimated weights.
Under these assumptions, we establish the convergence rate of \(\widehat{w}_2(M_i,X_i)\) to \(w_2^*(M_i,X_i)\). Since normalization makes $\widehat{w}^{\mathrm{MW}}_2(m,x)$ tend to zero as $n \to \infty$, we rescale by $n$ to obtain a nontrivial convergence rate.
The first component of the convergence rate, \(O_p\bigl( \sqrt{L^2 \log L/n} + L^{1 - r_2}\bigr)\), parallels Theorem 2 of Wang-MinimalDispersionConsiderations-2019m, while the second component, \(O_p(L K^{-r_1})\), reflects the influence of the first-step estimation.
Next, we establish the asymptotic normality of the EIF--type estimator and IPW--type estimator based on our proposed weights.
We require the following Assumption (ref).
Assumption (ref) is analogous to that of Wang-MinimalDispersionConsiderations-2019m. Condition (ref) states the standard regularity requirements, ensuring the estimators have finite moments. Condition (ref) then requires a uniform approximation for $\mu$ and $\eta$, similar to Assumption (ref) (ref). Condition (ref) introduces a complexity constraint, requiring that the metric entropy of the classes $\mathcal{W}_1$, $\mathcal{W}_2$, $\mathcal{U}$, and $\mathcal{T}$ does not grow too rapidly as $\varepsilon \to 0$. This condition is met by H\"older classes of smoothness $s$ on any bounded convex subset of $\mathbb{R}^d$ when $ s/d > 1$ (see, e.g., vanDerVaart-WeakConvergenceStatistics-1996k). Consistent with Wang-MinimalDispersionConsiderations-2019m and Fan-OptimalCovariateEstimation-2023s, these uniform approximation rate requirements are regarded as some of the least stringent in the literature (e.g. Hirano-EfficientEstimationScore-2003o; Chan-GloballyEfficientWeighting-2016v; Chan-EfficientNonparametricEffects-2016g). Additionally, condition (ref) restricts how quickly the number of basis functions, $K$ and $L$, may increase with the sample size $n$. The permissible rate for this increase is governed by the sum of the approximation errors ($r_{2}+r_{\mu}$, $r_1 + r_{\mu}$, and $r_1 + r_{\eta}$) of the propensity score and outcome estimators. This feature, that the product structures become crucial in the asymptotics, resembles the findings of Farbmacher-CausalMediationLearning-2022c on debiased machine learning for causal mediation.
Finally, we can establish the asymptotic distribution of the EIF--type estimator and show that it achieves the semiparametric efficiency bound.
For the IPW--type estimator, we obtain the same asymptotic result.
In this section, we conduct two Monte Carlo simulation studies to evaluate the performance of our proposed two–step minimal weights against several competing estimators. In addition, we implement the hyperparameter tuning of the balancing tolerances (Algorithm (ref)).
First, we consider a simulation setting adapted from TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y. The data are generated as follows:
We denote $\text{Bernoulli}(p)$ as a Bernoulli distribution where the variable takes the value 1 with probability \( p \), and $\mathcal{N}(0,1)$ as a normal distribution with mean zero and variance one. In this data-generating process, the covariates $ \{ X_1, \dots, X_4 \}$ are complex nonlinear transformations of the underlying covariates $ \{ Z_1, \dots, Z_4 \}$. The estimands are $\text{NDE}(0)$ and $\text{NDE}(1)$. The true values of both direct effects are $1$.
We compare the performance of several weighting methods when applied to the EIF--type and IPW--type estimators for NDE. This allows us to assess the performance of the weights both in conjunction with an outcome model and independently. The weighting methods are as follows:
All weights are normalized to sum to one. To specify the estimators, we introduce the following notation:
Then, the EIF--type estimators for $\text{NDE}(0)$ and $\text{NDE}(1)$ are given by the following two equations, respectively:
where the terms $\sum_{i=1}^n (1 - D_i) \widetilde{w}_1 (Y_i - \widehat{m}_0(X_i)) - 1/n \sum_{i=1}^n \widehat{m}_0(X_i)$ and $\sum_{i=1}^n D_i \widehat{w}_1 (Y_i - \widehat{m}_1(X_i)) + 1/n \sum_{i=1}^n \widehat{m}_1(X_i)$ are the estimators for $\mathbb{E}[Y_0]$ and $\mathbb{E}[Y_1]$. The IPW--type weighting estimators for $\text{NDE}(0)$ and $\text{NDE}(1)$ are given by the following two equations, respectively:
where the terms $\sum_{i = 1}^n (1 - D_i) \widetilde{w}_1(X_i) Y_i$ and $\sum_{i = 1}^n D_i \widehat{w}_1(X_i) Y_i$ are the estimators for $\mathbb{E}[Y_0]$ and $\mathbb{E}[Y_1]$.
To evaluate the balance of covariates and the mediator achieved by the weights, we adopt the target absolute standardized mean difference (TASMD) from Chattopadhyay-BalancingVsPractice-2020m. The TASMDs for the first and second weighting steps are defined as follows: $$
$$ Here, $\bar{X}_{w_1,c}$ and $\bar{X}_{w_2,t}$ are the weighted means of a covariate $X$ in the control group (using weights $\widehat{w}_1$) and the treatment group (using weights $\widehat{w}_2$), respectively. $\bar{X}_{full}$ is the unweighted mean of $X$ in the full population. The denominators $s_c$ and $s_t$ are the unweighted sample standard deviations of $X$ in the control and treated groups, respectively.
We implement the Monte Carlo simulation with $2000$ iterations in two settings described below.
Figures (ref) and (ref) display the Monte Carlo means and variances for each estimator. Specifically, Figure (ref) presents the results for $\text{NDE}(1)$, while Figure (ref) corresponds to $\text{NDE}(0)$. We first examine Setting (A), which is depicted in the first two columns of the plots in both figures. In this setting, for both $\text{NDE}(1)$ and $\text{NDE}(0)$, the well-specified outcome model masks the underlying differences in weight quality for the EIF-type estimators, resulting in similarly strong performance across all methods. The IPW--type estimators, however, isolate the direct impact of the weighting schemes. In this context, our proposed MW is the only method that remains unbiased with substantially low variance. Notably, both standard EIF weights and even the True PS-based weights result in unignorable bias and large variance. While trimming the EIF weights reduces variance, it fails to correct the bias. This result highlights that even a perfectly specified propensity score model is insufficient when sample-level covariate balance is not enforced. It also reveals the inherent instability of inverting propensity scores.
Next, we examine Setting (B), which is displayed in the remaining columns of Figures (ref) and (ref). These results demonstrate the proposed method's robustness to outcome model misspecification. For the EIF--type estimators, our proposed MW method continues to yield unbiased estimates with low variance, performing as well as it did in setting (A). In contrast, the estimators using standard EIF weights and True PS weights are now substantially biased, as their weights fail to correct for the misspecified outcome model. A similar pattern of MW's superior performance is observed for the IPW--type estimators.
The covariate balances of each weighting method are illustrated in Figure (ref) and Figure (ref) for setting (B) with $n=1000$. Figure (ref) presents the TASMD for the first-step weights, comparing the weighted control group to the full population. Figure (ref) shows the TASMD for the second-step weights, comparing the weighted treated group to the weighted control group. In both steps, our proposed MW achieves a TASMD of virtually zero across all specified covariates $\{ Z_1, \dots, Z_4, X_1, \dots, X_4 \}$ and the mediator $M$. This demonstrates that MW successfully enforces perfect sample-level balance as designed. In contrast, all competing methods, including EIF, EIF trimmed, CBPS, and even weights based on True PS, exhibit imbalances for several covariates, as indicated by their non-zero TASMD values.
\FloatBarrier
In light of Remark (ref), we now demonstrate the utility of MW in a different setting. To this end, we adapt the simulation design of Wong-Kernel-basedCovariateStudies-2018k to a causal mediation framework. The data are generated as follows:
Note that $\text{NDE}(1) = \text{NDE}(0) = 0$ because the treatment effect arises only through interactions with covariates. In this setting, a regression imputation model that omits the interaction terms between the treatment and the covariates is misspecified and thus expected to perform poorly. We consider the following two settings:
The performance of all weighting estimators in these simulation settings is presented in Figures (ref) and (ref). In setting (A), the performance of all estimators is degraded by the outcome model misspecification. While our proposed MW estimator consistently achieves the lowest variance, it exhibits a non-trivial bias when estimating $\text{NDE}(0)$.
Conversely, the advantages of the MW estimator are demonstrated in setting (B). In this setting, the propensity score models and the balancing constraints are enriched with the true covariates \{$Z_1, \dots, Z_4$\} and their interactions with the mediator \{$M \times Z_1, \dots, M \times Z_4$\}, even though the main outcome models remain misspecified. As a result, MW substantially outperforms all competing estimators. It is the only method to simultaneously achieve low bias and variance for both $\text{NDE}(1)$ and $\text{NDE}(0)$, as indicated by the bolded values in the table. The covariate balance diagnostics for setting (B) with $n = 1000$ in Figures (ref) and (ref) support these results. The figures clearly show that only MW achieves near-perfect balance in both weighting steps, whereas all other methods exhibit imbalance.
For comparison, we consider a regression imputation (RI) estimator; the results are shown in Figures (ref) and (ref). As expected, the RI estimator is severely biased in setting (A) due to its inability to capture the treatment-covariate interactions. Notably, this poor performance occurs even though the RI model utilizes the same set of covariates as the MW estimator, highlighting the weakness of the simple linear imputation approach in this setting. In setting (B), when the RI model is correctly specified to include these interactions, its performance becomes comparable to that of our MW estimator.
In summary, these results provide a compelling case for the proposed MW estimator. In setting (A), where the simple RI estimator fails due to model misspecification, MW performs comparably to the EIF--type estimator and offers improvement over RI. Conversely, in setting (B), where a correctly specified RI model performs well, MW matches its high level of accuracy and efficiency, with both methods outperforming the EIF--type estimator. This demonstrates that the estimator with MW matches or exceeds the performance of the best alternative method in both scenarios.
\FloatBarrier
Finally, we conduct a simulation study to evaluate the hyperparameter tuning procedure detailed in Algorithm (ref). We focus on setting (B) from the Wong and Chan simulation. The Monte Carlo simulation is performed with a sample size of $n = 500$ and $100$ iterations. For the tuning algorithm, we construct a grid of $500$ points for the hyperparameters $\varepsilon$ and $\delta$, and we employ bootstrap resampling with $R = 50$ repetitions.
The results, presented in Table (ref), compare the performance of the MW estimator under exact balancing (i.e., $\varepsilon=0, \delta=0$) against the estimator with optimally chosen hyperparameters. The tuning algorithm leads to a reduction in bias across all four estimator configurations. This reduction is a consequence of the algorithm preferring approximate balance over exact balance, as illustrated by the TASMD plots in Figures (ref) and (ref), which show that the selected hyperparameters yield small but non-zero balance errors. For the EIF--type estimators, the reduction in bias is accompanied by a slight increase in variance, resulting in a marginally higher mean squared error (MSE). In contrast, for the IPW--type estimators, the tuning algorithm successfully reduces both bias and variance, leading to an improved overall MSE.
In this section, we analyze the framing dataset using our proposed two--step minimal weights estimator. The dataset originates from brader2008triggers, who investigated how media framing of immigration news influences public opinion and political behavior. It is available in the R package mediation (R-mediation). In their experiment, participants were randomly assigned to read an article about immigration and subsequently answered a series of follow-up questions from which the outcome, mediators, and covariates were measured. Our objective is to assess the causal pathways through which framing influences information-seeking behavior: to what extent the effect arises directly from exposure to the media frame and ethnic cues, and to what extent it is mediated by shifts in perceived harm and negative emotional responses.
The sample consists of 265 individuals, with the main variables defined as follows:
We specify all nuisance models—the outcome regression models, the propensity score models, and the balancing constraints, using the same covariate set. We estimate ATE, NDE, and NIE using two types of estimators, EIF--type estimator and IPW--type estimator, across four different weighting schemes: EIF, EIF (trimmed), our proposed MW, and CBPS.
First, we assess the finite--sample balance of each weighting method by examining the TASMD. Figure (ref) displays the balance for the first-step weights (weighting the control group to resemble the full population), while Figure (ref) shows the balance for the second-step weights (weighting the treated group to resemble the re-weighted control group). In both figures, our proposed MW estimator achieves a TASMD of virtually zero for all 20 covariates included in the model.
We next examine the estimates reported in Figure (ref). The ATE is negative in all specifications. In other words, participants exposed to a positively framed story with a European (Russian) immigrant were more likely to request information from anti-immigration organizations than those exposed to a negatively framed story with a Latino (Mexican) immigrant. This aligns with the findings of brader2008triggers, who suggest that European cues, especially under a positive frame, may spark curiosity and information seeking. Indeed, consistent with this explanation, both $\mathrm{NDE}(1)$ and $\mathrm{NDE}(0)$ are negative, indicating that the media cue itself shifts the outcome downward. By contrast, $\mathrm{NIE}(1)$ and $\mathrm{NIE}(0)$ are positive, suggesting that changes in mediators induced by the treatment move behavior upward. This can be explained by the observation that exposure to a negatively framed story featuring a Latino (Mexican) immigrant changes perceived harm and emotions, leading individuals to seek information from anti-immigration organizations. Taken together, this implies that the positive indirect pathway partly offsets the negative direct effect.
Most estimates are not statistically significant. The important exception is $\mathrm{NDE}(0)$: under the EIF-type specification, all weighting methods yield significant results ($p \approx 0.01$-$0.03$), with MW producing the smallest variance and the narrowest interval. Under the IPW/HT-type specification, significance appears only with MW ($p \approx 0.013$), again due to a notable reduction in variance.
This paper addressed the challenge of estimating natural direct and indirect effects in the presence of finite--sample instability and covariate imbalance, which are two critical limitations of the EIF--based estimator and IPW estimator. To this end, we introduced a novel two--step minimal weights estimator specifically tailored for the causal mediation setting.
Our theoretical contribution began with a formal decomposition of the EIF--type estimator's bias. This analysis pinpointed imbalances in the marginal covariate distribution and the joint distribution of covariates and the mediator as the primary sources of bias, thereby motivating our two--step weight estimation algorithm, which sequentially targets these specific imbalances. We established the convergence rates for the proposed weights and proved that the resulting estimator is asymptotically normal and achieves the semiparametric efficiency bound.
Our simulation studies and empirical application provided empirical support for this approach. The two--step proposed minimal weights consistently reduced both bias and variance compared to competitors, particularly in challenging scenarios with model misspecification where standard methods failed. This underscores a key implication for applied researchers: directly targeting sample-level balance and weight stability can yield substantially more reliable estimates of causal mechanisms.
As a direction for future research, we note that causal mediation analysis is closely connected to the study of long-term treatment effects and statistical surrogacy, since mediators are often interpreted as short-term outcomes or surrogate endpoints (Chen-SemiparametricEstimationEffects-2023s, Imbens-Long-termCausalCombination-2024k, athey2024estimatingtreatmenteffectsusing). This topic has attracted increasing attention in econometrics, and we expect that estimators analogous to those proposed in this paper can be developed in this setting as well.