EconBase
← Back to paper

Balancing Weights for Causal Mediation Analysis

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

99,778 characters · 22 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Balancing Weights for Causal Mediation Analysis

abstractThis paper develops methods for estimating the natural direct and indirect effects in causal mediation analysis. The efficient influence function–based estimator (EIF--based estimator) and the inverse probability weighting estimator (IPW estimator), which are standard in causal mediation analysis, both rely on the inverse of the estimated propensity scores, and thus they are vulnerable to two key issues: (i) instability and (ii) finite-sample covariate imbalance. We propose estimators based on the weights obtained by an algorithm that directly penalizes weight dispersion while enforcing approximate covariate and mediator balance, thereby improving stability and mitigating bias in finite samples. We establish the convergence rates of the proposed weights and show that the resulting estimators are asymptotically normal and achieve the semiparametric efficiency bound. Monte Carlo simulations demonstrate that the proposed estimator outperforms not only the EIF--based estimator and the IPW estimator but also the regression imputation estimator in challenging scenarios with model misspecification. Furthermore, the proposed method is applied to a real dataset from a study examining the effects of media framing on immigration attitudes.

Keywords: causal mediation analysis; natural direct and indirect effects; propensity score instability; balancing weights; finite--sample covariate imbalance.

Introduction

Causal mediation analysis has become an essential framework for researchers who seek to understand the mechanisms of a causal effect. Many traditional studies focus on estimating the average treatment effect (ATE). However, researchers are now increasingly attempting to uncover the underlying causal pathways by decomposing the ATE into its mediated and unmediated components. This is important because improving policies or interventions often requires knowing why an effect exists, not simply confirming its existence. Reflecting this practical importance, this framework has become widespread across various disciplines, with both theoretical and empirical studies actively pursued (Imai-GeneralApproachAnalysis-2010o, Imai-IdentificationInferenceEffects-2010a, Imai-UnpackingBlackStudies-2011p, VanderWeele-ExplanationCausalInteraction-2015c, Huber-MediationAnalysis-2020y, Celli-CausalMediationModels-2022i, Nguyen-ClarifyingCausalLearn-2020r, Nguyen-ClarifyingCausalOutcomes-2022p). This paper contributes to this field by developing a novel weighting estimator for causal mediation analysis that enhances the finite--sample performance.

We introduce notation and define the estimands. Let $\mathcal{X}$, $\mathcal{M}$, and $\mathcal{Y}$ denote the supports of the covariates, mediator, and outcome, respectively. The observed data are $(Y,M,D,X)\in\mathcal{Y}\times\mathcal{M}\times\{0,1\}\times\mathcal{X}$. We use potential outcomes to formalize causal mechanisms. For $d\in\{0,1\}$, let $M_d$ denote the mediator that would be observed if $D$ were set to $d$, and let $Y_{dm}$ denote the outcome that would be observed if $(D,M)$ were set to $(d,m)$. We write $Y_d:=Y_{dM_d}$ and, critically for causal mediation analysis, $Y_{dM_{d'}}$ for the cross-world outcome where the treatment is $d$ while the mediator follows the distribution it would have under $d'$.

We consider two key estimands in causal mediation analysis: the natural direct effect (NDE) and the natural indirect effect (NIE). The NDE, defined as $\text{NDE}(d):=\mathbb{E}[Y_{1M_d} - Y_{0M_d}]$, measures the causal effect of changing the treatment from $0$ to $1$ while fixing the mediator to the distribution it would take under $D=d$. In this way, the NDE captures the part of the treatment effect that does not operate through the mediator and reflects only the direct path from treatment to outcome. The NIE, defined as $\text{NIE}(d):= \mathbb{E}[Y_{dM_1} - Y_{dM_0}]$, measures the effect of the treatment through the mediator on the outcome. The $\text{ATE} \, (:= \mathbb{E}[Y_1 - Y_0])$, can be decomposed as $\text{ATE} = \text{NDE}(1) + \text{NIE}(0) = \text{NIE}(1) + \text{NDE}(0)$, which are depicted in Figure (ref). Accordingly, we are interested in the targets $\theta_{d,1-d} := \mathbb{E}[Y_{dM_{1-d}}], \theta_d := \mathbb{E}[Y_d], d\in\{0,1\}$.

figure[figure omitted — 809 chars of source]

We focus on two main weighting estimators for causal mediation analysis: the inverse probability weighting (IPW) estimator and its augmented counterpart, the efficient influence function (EIF)–based estimator. The IPW estimator of Huber-IdentifyingCausalMechanisms-2014x for $\theta_{1,0}$ is given by

equation[equation omitted — 157 chars of source]

where $\widehat{\xi}_{d}(M,X)$ and $\widehat{\pi}_{d}(X)$ are estimators of $P(D=d\mid M,X)$ and $P(D=d\mid X)$, respectively. Since this weight also plays a role in the EIF–based estimator discussed below, we denote it by $ \widehat{w}^{\mathrm{EIF}}_2(M_i,X_i) = \widehat{\xi}_{0}(M_i,X_i)/\widehat{\pi}_{0}(X_i)\widehat{\xi}_{1}(M_i,X_i). $ Notably, these propensity score estimators appear in the denominator, which gives rise to the issues of interest. We also consider an augmented version, the EIF–based estimator of TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y, as introduced in (ref), which combines weighting with outcome regression (regression imputation). Specifically, the EIF–based estimator for $\theta_{1,0}$ employs not only $\widehat{w}^{\mathrm{EIF}}_2(M_i, X_i)$ but also $ \widehat{w}^{\mathrm{EIF}}_1(X_i) = 1/\widehat{\pi}_0(X_i), $ which also depends on the inverse of the estimated propensity score.\footnote{The subscripts 1 and 2 in $\widehat{w}^{\mathrm{EIF}}_1(X_i)$ and $\widehat{w}^{\mathrm{EIF}}_2(M_i,X_i)$ correspond to the discussion of the proposed two-step algorithm in Section (ref).}

Both the IPW estimator and the EIF--based estimator rely on the weights based on the inverse of the estimated propensity scores and are therefore susceptible to two key issues noted in previous studies: (i) instability and (ii) finite-sample covariate imbalance (Hirshberg-ApproachesWeightingInference-2017l, Chattopadhyay-BalancingVsPractice-2020m, Ben-Michael-BalancingActInference-2021x, Cohn-BalancingWeightsInference-2023j).

itemize• Instability. The weights $\widehat{w}^{\mathrm{EIF}}_1(X_i)$ and $\widehat{w}^{\mathrm{EIF}}_2(M_i,X_i)$ can become highly unstable when the estimated propensity scores approach zero or one. Not only that, but standard propensity score estimation methods do not explicitly penalize the instability of these inverse weights. For example, consider an individual with a true propensity score of $0.05$, which corresponds to an inverse weight of $20$. If the estimated propensity score is $0.09$, then its inverse is approximately $11$. However, if the estimate is $0.01$, its inverse rises to $100$, and there is no penalty for this type of explosion in the weights.\footnote{This example is taken from Ben-Michael-BalancingActInference-2021x.} In a causal mediation setting, this issue becomes particularly severe because, as seen in the estimator ((ref)), the denominator involves a product of inverse probability terms. Therefore, it is necessary to develop an algorithm that directly penalizes the dispersion of the weights. • Finite-sample covariate imbalance. The weights $\widehat{w}^{\mathrm{EIF}}_1(X_i)$ and $\widehat{w}^{\mathrm{EIF}}_2(M_i,X_i)$ asymptotically balance the distributions of covariates and mediators (we will see this in Proposition (ref), Section (ref)). However, this property does not necessarily hold in finite samples. Specifically, for certain functions $f(M_i,X_i)$ and $g(X_i)$, we may observe that \begin{align*} & \frac{1}{n}\sum_{i=1}^n D_i \widehat{w}^{EIF}_2(M_i,X_i) f(M_i,X_i) \neq \frac{1}{n}\sum_{i=1}^n (1-D_i) \widehat{w}^{EIF}_1(X_i) f(M_i,X_i), \\ & \frac{1}{n}\sum_{i=1}^n (1-D_i) \widehat{w}^{EIF}_1(X_i) g(X_i) \neq \frac{1}{n}\sum_{i=1}^n g(X_i). \end{align*} As we will show in Section (ref), if $f(M_i, X_i)$ and $g(X_i)$ are important predictors of the outcome, then such imbalances in finite samples lead to finite-sample bias. Therefore, it is necessary to develop an algorithm that explicitly enforces covariate and mediator balance in finite samples.\footnote{Although it is possible to assess balance after weight estimation and re-estimate if needed, such an ad hoc iterative procedure can invalidate subsequent inference.}

To directly address these two concerns, we propose alternative weighting methods by extending balancing weights to causal mediation analysis. In particular, we build on the “minimal--dispersion approximate balancing weights” (hereafter, minimal weights) of Wang-MinimalDispersionConsiderations-2019m, as this framework encompasses several existing methods as special cases, including entropy balancing Hainmueller-EntropyBalancingStudies-2012c, stable balancing weights Zubizarreta-StableWeightsData-2015v, and empirical balancing calibration weights Chan-GloballyEfficientWeighting-2016v. The method’s name describes the two simultaneous objectives of the underlying optimization problem: “minimal dispersion”, i.e.\ minimizing a measure of weight dispersion such as a squared loss or an entropy‐based loss, and “approximate balancing”, i.e.\ enforcing covariate balance up to pre-specified tolerances. In this way, the method addresses both of the issues discussed above—(i) instability and (ii) finite-sample covariate imbalance. At the same time, it removes the need for ad hoc post–estimation covariate balance diagnostics or trimming of extreme weights, both of which are commonly applied in practice.

We propose a “two-step minimal-dispersion approximate balancing weights ”method (hereafter, two-step minimal weights) for causal mediation. The name “two-step” reflects our strategy to first balance the marginal distribution of the covariates, and then the joint distribution of the covariates and the mediator. This sequential approach is motivated by our analysis in Section (ref), which reveals that imbalances in these specific distributions generate the bias. While a similar decomposition of the bias source has been recently discussed by liu2024two, our study leverages this insight to construct the different weighting algorithm. For clarity, we denote the resulting two-step minimal weights by $\widehat{w}^{\mathrm{MW}}_1(X_i)$ and $\widehat{w}^{\mathrm{MW}}_2(M_i,X_i)$.

As a theoretical contribution, we establish the consistency of the weights $\widehat{w}^{\mathrm{MW}}_1(X_i)$ and $\widehat{w}^{\mathrm{MW}}_2(M_i, X_i)$ and the asymptotic normality and semiparametric efficiency of the resulting estimators: IPW estimator and EIF--based estimator with weights $\widehat{w}^{\mathrm{MW}}_1(X_i)$ and $\widehat{w}^{\mathrm{MW}}_2(M_i, X_i)$. Firstly, by considering a nonparametric setting in which the number of balancing constraints increases with the sample size, we show that the optimization-based weights $\widehat{w}^{\mathrm{MW}}_1(X_i)$ and $\widehat{w}^{\mathrm{MW}}_2(M_i, X_i)$ converge, respectively, to $\xi_{0}(M_i, X_i) / \pi_{0}(X_i) \xi_{1}(M_i, X_i)$ and $1/\pi_0(X_i)$. The proof of this result relies on the dual formulation and is derived under standard M--estimation assumptions. We also impose conditions that are standard in sieve estimation (see, e.g., Newey-ConvergenceRatesEstimators-1997k; Chen-Chapter76Models-2007p), and our asymptotic arguments parallel those in the balancing-weight literature (e.g., Wang-MinimalDispersionConsiderations-2019m; Fan-OptimalCovariateEstimation-2023s). Secondly, under additional assumptions such as the complexity of the function class, we establish that the estimators are asymptotically normally distributed with variance achieving the semiparametric efficiency bound. These results imply that, under the stated conditions, the proposed estimators are asymptotically equivalent to those based on $\widehat{w}^{\mathrm{EIF}}_1(X_i)$ and $\widehat{w}^{\mathrm{EIF}}_2(M_i, X_i)$. We therefore conclude that the proposed method enhances finite-sample performance while preserving its desirable asymptotic properties.

The proofs of these asymptotic properties are provided in the Appendix. While we employ techniques similar to those in Wang-MinimalDispersionConsiderations-2019m for the ATE, extending them to causal mediation analysis is non-trivial. Specifically, a key challenge lies in controlling the influence of the first-step weight estimation on the second step. We address this issue by exploiting the structure of the efficient influence function, building on insights from recent work such as Farbmacher-CausalMediationLearning-2022c.

To evaluate the finite–sample performance of our method, we conduct extensive simulations based on two settings from TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y and Wong-Kernel-basedCovariateStudies-2018k, under both (A) correct specification and (B) misspecification. We compare our approach with three baseline estimators: the IPW estimator, the EIF–based estimator, and the regression imputation estimator. A concise summary of the simulation results is reported in Table (ref). Overall, the method performs as well as or better than the baseline estimators. In addition, we implement other competing weighting schemes, including covariate balancing propensity scores (CBPS)\footnote{The CBPS method was originally proposed by Imai-CovariateBalancingScore-2014v. Its extension to causal mediation analysis is provided in Appendix (ref) of this paper.} and an oracle variant based on the true propensity scores. The proposed method further outperforms these alternatives, confirming its broad advantages.

table[table omitted — 672 chars of source]

We illustrate the practical value of our approach using the media–framing data of brader2008triggers. To focus squarely on the role of weighting quality, we consider two generic estimators: an IPW–type estimator obtained by plugging arbitrary weights into estimator ((ref)), and an EIF–type estimator defined as its augmented counterpart with outcome regression. In particular, we show that using our proposed minimal weights can reduce standard errors in both the EIF–type and IPW–type estimators, thereby providing clearer results.

Related literature

This paper contributes to both the literature on causal mediation analysis and the literature on balancing weights. The key references we build upon for the estimation of the NDE and NIE are TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y and Huber-IdentifyingCausalMechanisms-2014x. For a review of estimators, including weighting estimators and regression estimators, see Nguyen-CausalMediationEffects-2023j. While Nguyen-CausalMediationEffects-2023j discusses the general idea of achieving balance between covariates and the mediator, it does not explore concrete balancing-weight algorithms or their theoretical properties, which are the focus of the present paper.

While the literature on balancing weights is extensive for the ATE Hainmueller-EntropyBalancingStudies-2012c, Zubizarreta-StableWeightsData-2015v, Chan-GloballyEfficientWeighting-2016v, Hirshberg-ApproachesWeightingInference-2017l, Chattopadhyay-BalancingVsPractice-2020m, Ben-Michael-BalancingActInference-2021x, Cohn-BalancingWeightsInference-2023j, its extension to causal mediation analysis remains limited. Although several optimization-based weighting approaches have been proposed recently, our method remains distinct. For instance, Huang-NonparametricEstimationTreatment-2024x, building on Ai-UnifiedFrameworkModels-2021c, proposed estimating nonparametric propensity scores by minimizing entropy loss. However, their approach does not explicitly address the issue of (ii) finite-sample covariate imbalance. liu2024two introduced a minimax estimation framework following dikkala2020minimax that accounts for covariate imbalance, but their algorithmic approach differs fundamentally from ours. Our method is specifically designed to simultaneously tackle two key issues: (i) instability and (ii) finite-sample covariate imbalance. In this regard, the work most closely related to ours is Chan-EfficientNonparametricEffects-2016g, which also seeks to determine weights by minimizing a distance measure. However, our algorithm diverges by incorporating approximate balancing constraints to explicitly optimize the bias--variance trade-off. Furthermore, we establish our asymptotic properties under weaker assumptions than those required in Chan-EfficientNonparametricEffects-2016g.

Outline

The remainder of this paper is organized as follows. Section (ref) introduces the basic setup and the identification of the key estimands. Section (ref) presents the EIF--based and IPW estimators and provides a formal bias‐decomposition analysis. Section (ref) proposes the two‐-step minimal‐-weights algorithm. Section (ref) establishes the asymptotic properties of the proposed estimators. Section (ref) presents the results of simulation studies. Section (ref) then reports an empirical application. Finally, Section (ref) concludes the paper and outlines directions for future research. Implementation notes and proofs of the main theorems are provided in the Appendix.

Causal Mediation Analysis: Setup and Identification

This section introduced notations and assumptions for identification. We state the standard consistency, sequential ignorability, and positivity assumptions, and present identification formulas for $\theta_{d,1-d}$ and $\theta_d$.

To facilitate the subsequent discussion, we introduce the following notations.

itemize\( f_{Y|M, D, X}(y) \), \( f_{M \mid D, X}(m) \), and \( f_X(x) \) denote the (conditional) densities of the outcome, mediator, and covariates, respectively. • Conditional outcome expectation ($m$): The function $m_d(X) := \mathbb{E}[Y\mid D=d,X]$ denotes the conditional mean of the outcome under treatment status \(d\) given covariates \(X\). • Conditional outcome expectation with mediator ($\mu$): The function $\mu_{d}(M,X) := \mathbb{E}[Y|M, D = d, X]$ denotes the conditional mean of the outcome under treatment status \(d\) given the mediator, treatment, and covariates. • Counterfactual outcome expectation ($\eta$): The function $\eta_{d, 1-d}(X) := \mathbb{E}[\mathbb{E}[Y|M, D = d, X] | D = 1 - d, X] = \int_{m \in \mathcal{M}}\mu_{d}(m,X)f_{M \mid D = 1-d, X}(m)dm$ denotes the expected value of the outcome under treatment status $d$ and the mediator distribution that would occur under the opposite treatment status $1-d$. • Propensity score ($\pi$): The function $\pi_{d}(X) := P(D=d|X)$ denotes the conditional probability of receiving the treatment $d$ given the covariates. • Propensity score given mediator ($\xi$): The function $\xi_{d}(M,X) := P(D=d|M,X)$ denotes the conditional probability of receiving the treatment $d$ given both the mediator and the covariates.

$\theta_{d, 1-d}$ and $\theta_{d}$ for $d \in \{0,1\}$ are identified by the following theorem.

theoremIf all of these conditions hold: \begin{enumerate} • Consistency: If $D = d$ then $M_d = M$ (w.p.1), and if $D = d$ and $M = m$, then $Y_{dm} = Y$ (w.p.1). • Sequential ignorability: For each \( d \in \{0, 1\} \) and $m \in \mathcal{M}$, \[ Y_{dm} \mathrel{\perp\!\!\!\perp} D \mid X, \quad M_{d} \mathrel{\perp\!\!\!\perp} D \mid X, \quad Y_{dm} \mathrel{\perp\!\!\!\perp} M \mid D = d, X, \text{ and } Y_{dm} \mathrel{\perp\!\!\!\perp} M_{1-d} \mid X. \] • Positivity: \[ 0 < f_{M|D,X}(m) < 1 \text{ for each } m \in \mathcal{M}, \text{and } 0 < \pi_{d}(X) < 1 \text{ for each } d \in \{0, 1\}. \] \end{enumerate} Then, \( \theta_{d, 1-d} \) and $\theta_{d}$ are identified as follows: \begin{align*} \theta_{d, 1-d} &= \mathbb{E}\Bigl[\frac{(dD + (1-d)(1-D))\xi_{1-d}(M,X)}{\pi_{1-d}(X)\xi_d(M,X)} (Y - \mu_d(M,X)) \\ &\quad + \frac{(1-d)D + d(1-D)}{\pi_{1-d}(X)}(\mu_d(M,X) - \eta_{d,1-d}(X)) + \eta_{d,1-d}(X) \Bigr] \\ \theta_{d} &= \mathbb{E}\left[ \frac{(dD + (1-d)(1-D))}{\pi_d(X)}(Y - m_d(X)) + m_d(X) \right]. \end{align*}

For the proof of Theorem (ref), see TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y. The identification can be established through several complementary results:

align*[align* omitted — 262 chars of source]

Theorem (ref) can be viewed as a combination of these results. Intuitively, identification can be achieved either through weighting alone or through outcome expectations alone, while their combination naturally yields an alternative representation. Indeed, the expression

align*[align* omitted — 202 chars of source]

constitutes an efficient influence function for $\theta_{d,1-d}$. The derivation of the efficient influence function for $\mathbb{E}[Y_{dM_{1-d}}]$ can be found in TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y.\footnote{In fact, the estimator in (ref) is an alternative expression of that in TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y, where the original form is given by

align[align omitted — 394 chars of source]

In this work, we adopt expression (ref) to avoid estimating the mediator's density and to facilitate the discussion of the covariate balancing propensity score method in Appendix (ref). The estimator (ref) is derived by a straightforward application of Bayes' rule to (ref) and thus retains the properties described in TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y.} The estimator based on this representation is therefore referred to as the EIF-based estimator.

We now provide a more detailed explanation of the assumptions for identification. Consistency assumes that the observed mediator and outcome are equal to the potential mediator and outcome under the observed treatment and mediator assignments. The sequential ignorability assumption consists of the following four conditional independences. First, $ Y_{dm} \mathrel{\perp\!\!\!\perp} D \mid X, $ which ensures that the effect of treatment \( D \) on the outcome \( Y \) is unconfounded given \( X \). Second, $ M_{d} \mathrel{\perp\!\!\!\perp} D \mid X, $ implying that the effect of treatment \( D \) on the mediator \( M \) is unconfounded given \( X \). Third, $ Y_{dm} \mathrel{\perp\!\!\!\perp} M \mid D = d, X, $ which states that the effect of the mediator \( M \) on the outcome is unconfounded given \( X \). Finally, the so-called cross-world independence assumption $ Y_{dm} \mathrel{\perp\!\!\!\perp} M_{1-d} \mid X $ is crucial. It ensures the absence of confounding between the mediator and the outcome that is induced by the treatment. Such a confounder, often called a treatment-induced confounder (denoted by $L$ in Figure (ref)), is a variable affected by $D$ that subsequently affects both $M$ and $Y$. For further details on the sequential ignorability assumptions, see, for example, Pearl-DirectIndirectEffects-2001g, Imai-IdentificationInferenceEffects-2010a, Vanderweele-EffectDecompositionConfounder-2014z. Positivity assumption requires that each treatment and mediator level has a nonzero probability for every given value of the covariates.

figure[figure omitted — 1,257 chars of source]

Weighting Estimators and Their Finite-Sample Bias

We now proceed from identification to estimation. Estimating the NDE and NIE requires estimating the quantity $\mathbb{E}[Y_{dM_{1-d}}]$. This estimand is unique to causal mediation because it relies on the potential outcome under two treatment statuses. For notational simplicity, we will henceforth focus on estimating $\theta_{1,0} := \mathbb{E}[Y_{1M_0}]$.\footnote{All the following analysis applies similarly to $ \theta_{0,1} := \mathbb{E}[Y_{0M_1}].$}

This section treats the two canonical weighting estimators for $\theta_{1,0}$: the EIF-based estimator and the IPW estimator introduced in Section (ref). In addition, we introduce the IPW--type estimator and the EIF--type estimator, which replace the weights based on the inverse of the estimated propensity scores with a set of general weights. For both the EIF--type and IPW--type estimators, we provide a finite--sample bias decomposition. This decomposition formally demonstrates the importance of addressing (ii) finite-sample covariate imbalance, one of the two concerns highlighted in Section (ref), alongside (i) instability.

Efficient influence function-based estimator

The EIF–based estimator is given by

align[align omitted — 385 chars of source]

where

equation[equation omitted — 244 chars of source]

We denote $\widehat{\pi}_0(X_i)$, $\widehat{\xi}_d(M_i,X_i)$, $\widehat{\mu}_1(M_i,X_i)$, and $\widehat{\eta}_{1,0}(X_i)$ as estimators of $\pi_0(X_i)$, $\xi_d(M_i,X_i)$ for $d\in\{0,1\}$, $\mu_1(M_i,X_i)$, and $\eta_{1,0}(X_i)$, respectively.

One of the key properties of the EIF--based estimator is a semiparametric efficiency.

theorem[Semiparametric efficiency, TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y] If all nuisance parameters are correctly specified, then the EIF--based estimator \( \widehat{\theta}^{\mathrm{EIF}}_{1,0} \) achieves the semiparametric efficiency bound.

For the proof, see TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y. In Section (ref), we show that our proposed estimator also attains this semiparametric efficiency bound under certain regularity conditions \footnote{Another key property of the EIF--based estimator is its multiple robustness.

If one of the following conditions hold:

itemize• Both \(\widehat{\pi}_0(X_i)\) and \(\widehat{\xi}_1(M_i,X_i)\) are consistent; • Both \(\widehat{\pi}_0(X_i)\) and \(\widehat{\mu}_1(M_i,X_i)\) are consistent; • Both \(\widehat{\mu}_1(M_i,X_i)\) and \(\widehat{\eta}_{1,0}(X_i)\) are consistent;

then the estimator \(\widehat{\theta}^{\mathrm{EIF}}_{1,0}\) is consistent.

This result implies that the estimator remains consistent even when some nuisance parameters are misspecified. Although multiple robustness is not the main focus of this paper, we can show, similar to TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y, that when the balancing constraints for the weights are correctly specified, the estimator remains consistent even if the outcome models are misspecified. In fact, in Section (ref), we confirm this consistency under outcome model misspecification.}.

EIF–type estimator and its finite--sample bias

Although the EIF--based estimator is multiple robust and semiparametrically efficient, its reliance on the inverse of propensity scores leads to the aforementioned issues of instability and covariate imbalance. To formally investigate the latter, we consider a general EIF--type weighting estimator, which replaces the specific EIF weights \( \widehat{w}^{\text{EIF}}_2(M_i,X_i) \) and \( \widehat{w}^{\text{EIF}}_1(X_i) \) with generic weights \( \widehat{w}_2(M_i,X_i) \) and \( \widehat{w}_1(X_i) \):

align[align omitted — 374 chars of source]

Decomposing the bias of this estimator requires defining two sets of terms. First, we define the estimation errors in the nuisance functions as \[ \widetilde{\mu}_1(M_i,X_i) := \mu_1(M_i,X_i) - \widehat{\mu}_1(M_i,X_i), \quad \widetilde{\eta}_{1,0}(X_i) := \eta_{1,0}(X_i) - \widehat{\eta}_{1,0}(X_i). \] Second, we define the random errors as \[ u_i := Y_i - \mu_1(M_i,X_i), \quad v_i := \mu_1(M_i,X_i) - \eta_{1,0}(X_i). \] With these components defined, we can formally state the sources of the estimator's bias in the following proposition.

proposition[The bias of the EIF--type weighting estimator in finite samples] For the EIF--type weighting estimator (ref), the bias in finite samples is given by \begin{align*} \mathbb{E}\left[ \widehat{\theta}^{\mathrm{EIF-type}}_{1,0} - \theta_{1,0} \right] &= \mathbb{E} \Biggl[ \frac{1}{n}\sum_{i=1}^n D_i \widehat{w}_2(M_i,X_i) \widetilde{\mu}_1(M_i,X_i) - \frac{1}{n}\sum_{i=1}^n (1-D_i) \widehat{w}_1(X_i) \widetilde{\mu}_1(M_i,X_i)\\ &\quad + \frac{1}{n}\sum_{i=1}^n (1-D_i) \widehat{w}_1(X_i) \widetilde{\eta}_{1,0}(X_i) - \frac{1}{n}\sum_{i=1}^n \widetilde{\eta}_{1,0}(X_i) \Biggr]. \end{align*}

We prove Proposition (ref). First, the difference between the EIF--type weighting estimator and the target parameter can be decomposed as:

align*[align* omitted — 787 chars of source]

The first two terms reflect imbalances in \( \widetilde{\mu}_1(M_i,X_i) \) and \( \widetilde{\eta}_{1,0}(X_i) \). Imbalance in \( \widetilde{\mu}_1(M_i,X_i) \) arises when the weights \( \widehat{w}_2(M_i,X_i) \) do not balance the distributions of the mediator and covariates between the treatment group and the control group weighted by \( \widehat{w}_1(X_i) \). Similarly, imbalance in \( \widetilde{\eta}_{1,0}(X_i) \) arises when the weights \( \widehat{w}_1(X_i) \) fail to balance the covariate distributions between the control group and the full sample. The third and fourth terms capture the errors \( u_i \) and \( v_i \), while the last term represents the sampling variation of \( \eta_{1,0}(X_i) \).

Next, by taking expectations on both sides, we show that the last three terms are equal to $0$. First, note that $\mathbb{E}[\eta_{1,0}(X)] = \theta_{1,0}$. For the first noise term, using the law of total expectation, we have:

align*[align* omitted — 240 chars of source]

Similarly, for the second noise term,

align*[align* omitted — 141 chars of source]

This completes the proof of Proposition (ref). The key insight from this decomposition is that the finite--sample bias stems solely from imbalances in the estimation errors, $\widetilde{\mu}_{1}(M,X)$ and $\widetilde{\eta}_{1,0}(X)$. Our proposed weighting scheme is therefore designed to directly target this source of bias.

Furthermore, this bias decomposition sheds light on another property of the EIF--based estimator (ref): it is asymptotically unbiased due to its population covariate balance.

proposition[Population covariate balance of EIF weights (ref)] Let \( f(M,X) \) and \( g(X) \) be any bounded functions. We denote $w^*_1(X) = 1/\pi_0(X)$ and $w^*_2(M, X) = \xi_0(M,X)/\pi_0(X) \xi_1(M,X)$. Then, \begin{align*} \mathbb{E}[D w^*_2(M, X) f(M,X)] &= \mathbb{E}[(1-D) w^*_1(X) f(M,X)], \\ \mathbb{E}[(1-D) w^*_1(X)g(X)] &= \mathbb{E}[g(X)]. \end{align*}

The proof is deferred to Appendix (ref). Proposition (ref) shows that the moment equalities hold for arbitrary measurable functions \( f(M,X) \) and \(g(X)\). Consequently, as long as \(\xi_d(M,X)\) and \( \pi_d(X) \) are consistently estimated, the EIF--based estimator (ref) remains asymptotically unbiased, regardless of the specification of \( \widetilde{\mu}(M,X) \) and \( \widetilde{\eta}(X) \). However, the EIF weights \( \widehat{w}^{\text{EIF}}_1(X_i) \) and \( \widehat{w}^{\text{EIF}}_2(M_i,X_i) \) are not constructed to ensure finite--sample covariate balance, which introduces bias as Proposition (ref) implies.

IPW–type estimator and its finite--sample bias

To complement the EIF--type estimator, which leverages an outcome model, we introduce an alternative that relies exclusively on weights to estimate the NDE and NIE. This estimator is particularly useful for isolating and evaluating the performance of the weighting scheme itself. We define this estimator here and show that its finite--sample bias can be decomposed in a manner analogous to that of the EIF--type estimator. The IPW–type estimator is defined as: $$ \widehat{\theta}_{1,0}^{\mathrm{IPW-type}} = \frac{1}{n}\sum_{i=1}^n D_i \widehat{w}_2(M_i,X_i) Y_i. $$ Following a similar process as in the proof of Proposition (ref), the IPW–type estimator has the following bias decomposition.

proposition[The bias of the IPW--type weighting estimator in finite samples] For the IPW--type weighting estimator, the bias in finite samples is given by \begin{align} \mathbb{E}\left[ \widehat{\theta}^{\mathrm{IPW-type}}_{1,0} - \theta_{1,0} \right] &= \mathbb{E} \Biggl[ \frac{1}{n}\sum_{i=1}^n D_i \widehat{w}_2(M_i,X_i) \mu_1(M_i,X_i) - \frac{1}{n}\sum_{i=1}^n (1-D_i) \widehat{w}_1(X_i) \mu_1(M_i,X_i)\\ &\quad + \frac{1}{n}\sum_{i=1}^n (1-D_i) \widehat{w}_1(X_i) \eta_{1,0}(X_i) - \frac{1}{n}\sum_{i=1}^n \eta_{1,0}(X_i) \Biggr]. \end{align}

The difference between this bias decompositon and that of the EIF--type estimator (Proposition (ref)) is that the bias of the IPW--type estimator stems from the imbalance of the true $\mu_1(M_i,X_i)$ and $\eta_{1,0}(X_i)$, rather than its estimation error $\widetilde{\mu}_1(M_i,X_i)$ and $\widetilde{\eta}_{1,0}(X_i)$.

Therefore, for both the IPW–type and EIF–type estimators, it is crucial that the weights balance mediators and covariates, such as those included in the specifications of $\mu_1(M_i,X_i)$ and $\eta_{1,0}(X_i)$.

Proposed Method: Two-Step Minimal Weights

To address the issues of (i) instability and (ii) finite-sample covariate balance in weighting estimators discussed previously, this section introduces an optimization-based approach. Applying the “minimal weight” method Wang-MinimalDispersionConsiderations-2019m to causal mediation analysis requires an adaptation to account for the two distinct sources of bias identified in Section (ref). To this end, we propose the “two--step minimal weights” method. This algorithm constructs weights by solving an optimization problem, specifically designed to sequentially correct the two types of imbalances inherent in causal mediation analysis.

In what follows, we first formally present the two--step algorithm. We then introduce a data-driven procedure for tuning its key tolerance hyperparameters which govern the trade-off between bias and variance.

Two-step minimal weights

We propose the following two--step algorithm to derive weights for the weighting estimators of \(\theta_{1,0}\).

Algorithm[Two-step minimal weights] We define the weights derived by the following two--step optimization as “two--step minimal weights”. Here, $f(w)$ denotes a weight function specified by the researcher, subject to the condition that $f''(w) > 0$. \begin{itemize} • Obtain \(\{\widehat{w}^{\text{MW}}_1(X_i)\}_{i=1}^{n}\) by solving the constrained optimization problem: \begin{align} \min_{ \{w_{1,i}\}_{i=1}^n } \quad &\sum_{i=1}^n (1-D_i) f(w_{1i}), \\ subject to \quad &\left| \sum_{i=1}^n (1-D_i) w_{1i} c_j(X_i) - \frac{1}{n}\sum_{i=1}^n c_j(X_i) \right| \leq \varepsilon_j, \quad j=1,\ldots,K, \end{align} where \(\{c_j(X_i)\}_{j=1}^{K}\) are smooth basis functions of the covariates. • Using the weights \(\{\widehat{w}^{\text{MW}}_1(X_i)\}_{i=1}^{n}\) from Step 1, obtain \(\{\widehat{w}^{\text{MW}}_2(X_i, M_i)\}_{i=1}^{n}\) by solving: \begin{align} \min_{ \{w_{2,i}\}_{i=1}^n } \quad &\sum_{i = 1}^n D_i f(w_{2i}), \\ subject to \quad &\left| \sum_{i=1}^n D_i w_{2i} b_j(X_i, M_i) - \sum_{i=1}^n (1-D_i) \widehat{w}^{MW}_1(X_i) b_j(M_i,X_i) \right| \leq \delta_j, \quad j=1,\ldots,L, \end{align} where \(\{b_j(X_i, M_i)\}_{j=1}^{L}\) are smooth basis functions of the covariates and mediator. \end{itemize}

To provide an overview, Algorithm (ref) is tailored to causal mediation by coupling a mediation-aware balance design with an explicit stabilization mechanism. It adopts a sequential scheme that first balances $X$ in Step 1 (control versus full sample) and then balances $(M,X)$ in Step 2 (treated versus reweighted control), reflecting the two imbalances that arise in our bias decomposition in Proposition (ref) and Proposition (ref). Each step minimizes a convex dispersion penalty $f(\cdot)$ subject to approximate balance constraints governed by tolerances $\{\varepsilon_j\}_{j=1}^K$ and $\{\delta_j\}_{j=1}^L$.

To begin, we discuss the choice of the loss function. Typical choices for $f$ include entropy balancing ($f(w)=w\log w$) and stable balancing weights ($f(w)=(w-1/n_0)^2$, with $n_0$ the number of controls), both of which tame the dispersion of inverse-propensity-type weights while preserving the targeted balance moments.

Next, we turn to the balancing constraints. In Step 1, the constraints are applied to the target functions $\{c_j(X)\}_{j=1}^K$, reflecting the assumption that $\eta_{1,0}(X)$ (or its estimation error $\widetilde{\eta}_{1,0}(X)$) can be well approximated by a linear combination of these bases. Setting $\varepsilon_j = 0$ corresponds to an exact sample balance.

In practice, to normalize the weights, one may add $c_{K+1}(X_i) = 1$ with $\varepsilon_{K+1} = 0$. The non-negativity condition of the weights is typically achieved either by directly including it as a constraint or by an appropriate choice of loss function, which we explain when presenting Theorem (ref) that derives the dual form of the Algorithm (ref). Step 2 applies the same procedure to $\{b_j(M,X)\}_{j=1}^L$, aligning the treated group with the reweighted control group on the joint $(M,X)$ distribution.

Finally, the choice of hyperparameters $\{\varepsilon_j\}_{j=1}^K$ and $\{\delta_j\}_{j=1}^L$ which determine the bias--variance trade-off, is discussed in the next subsection.

Then, we can construct the EIF‐type weighting estimator and IPW–type weighting estimator with normalized two--step minimal weights:

align*[align* omitted — 424 chars of source]

The properties of these estimators are examined in Section (ref). In our subsequent simulation and empirical application sections, we will employ both estimators to evaluate the properties of the proposed weights.

remarkAn extension to the case of multiple mediators is straightforward by including all mediators in the second step of our procedure, but the interpretation requires caution. For instance, suppose there are two mediators, $M_{\mathrm{I}}$ and $M_{\mathrm{II}}$. Without additional assumptions, it is not possible to separately identify the pathways $D \to M_{\mathrm{I}} \to Y$, $D \to M_{\mathrm{II}} \to Y$, and $D \to M_{\mathrm{I}} \to M_{\mathrm{II}} \to Y$. Under the current identification assumptions and estimation strategy, we can identify and estimate only the NIE through the joint mediators, $D \to \{M_{\mathrm{I}}, M_{\mathrm{II}}\} \to Y$, along with its corresponding NDE. In the empirical application presented in Section (ref), we estimate the NDE and NIE considering joint mediators. Extending balancing methods to path-specific effects (see Miles-SemiparametricEstimationConfounding-2020l) is left for future research.

Hyperparameter tuning for tolerances

One of the key features of our proposed method is the use of tolerances $\{\varepsilon_j\}_{j=1}^K$ and $\{\delta_j\}_{j=1}^L$, which allow for approximate rather than exact covariate balance. Since the weights are constructed without using outcome data, standard cross-validation is not applicable. Therefore, we adapt the procedure from Wang-MinimalDispersionConsiderations-2019m, which selects the tolerances by evaluating the stability of covariate balance across bootstrap samples, as detailed in Algorithm (ref).

AlgorithmLet \( \mathcal{E} \subset [0, (nK)^{-1/2}] \) denote a finite grid of candidate values for $\{\varepsilon_j\}_{j=1}^K$ and let \( \mathcal{D} \subset [0, (nL)^{-1/2}] \) denote a finite grid of candidate values for $\{\delta_j\}_{j=1}^L$. These grids are based on Assumption (ref) and Assumption (ref). Step 1: For each \( \varepsilon' \in \mathcal{E} \): \begin{enumerate} • Apply the Step 1 of Algorithm (ref) with tolerances \( \varepsilon_j = \varepsilon'\) for $j = 1, \dots, K$ to obtain the weights \( \{ \widehat{w}_{1}(X_i) \}_{i=1}^n \) and normalize them. • For \( r = 1, \dots, R \), draw a bootstrap sample \( S_r \) (with replacement) from the original data and compute the covariate balance criterion \[ C_r := \sum_{j=1}^{K} \left\| \left( \sum_{i \in S_r} \widehat{w}_1(X_i)(1 - D_i)c_j(X_i) - \frac{1}{n} \sum_{i=1}^{n} c_j(X_i) \right) \Big/ \mathrm{sd}(c_j(X)) \right\|_2. \] • Compute the average imbalance $ \bar{C}(\varepsilon') := \frac{1}{R} \sum_{r=1}^{R} C_r. $ \end{enumerate} Select the optimal imbalance level $ \varepsilon^* := \arg\min_{\varepsilon' \in \mathcal{E}} \bar{C}(\varepsilon'). $ Step 2: For each \( \delta' \in \mathcal{D} \): \begin{enumerate} • Using the optimal weights \( \{ \widehat{w}_{1}(X_i) \}_{i=1}^n \) obtained with \( \varepsilon^* \), apply the Step 2 of Algorithm (ref) with tolerances \( \delta_j = \delta'\) for $j = 1, \dots, L$ to obtain the weights \( \{ \widehat{w}_{2}(X_i, M_i) \}_{i=1}^n \) and normalize them. • For \( r = 1,\dots,R \), draw a bootstrap sample \( S_r \) (with replacement) from the original data and compute the covariate balance criterion \[ B_r := \sum_{j=1}^{L} \left\| \left( \sum_{i \in S_r} \widehat{w}_2(M_i,X_i)D_ib_j(M_i,X_i) - \sum_{i=1}^{n} \widehat{w}_{1}(X_i)(1-D_i)b_j(M_i,X_i) \right) \Big/ \mathrm{sd}(b_j(M,X)) \right\|_2. \] • Compute the average imbalance $ \bar{B}(\delta') := \frac{1}{R} \sum_{j=1}^{R} B_r. $ \end{enumerate} Select the optimal imbalance level $ \delta^* := \arg\min_{\delta' \in \mathcal{D}} \bar{B}(\delta'). $

The core idea of this algorithm is to select the tolerances that yield the most stable covariate balance across multiple bootstrap resamples of the data. By averaging the imbalance measure over these resamples, the procedure aims to find hyperparameters that perform well not only on the original sample but also under bootstrap resampling.

However, there are a couple of downsides to this algorithm. First, the sequential optimization can sometimes get stuck in a local optimum. For example, even if the first stage selects an $\varepsilon^*$ that provides good balance for the $\{c_k(X)\}_{k=1}^K$ basis functions, this does not necessarily ensure good balance for the $\{b_l(M,X)\}_{l=1}^L$ basis functions in the second stage. Second, the algorithm applies the same $\varepsilon^*$ and $\delta^*$ to all basis functions. Ideally, we would prefer to balance covariates that are more strongly related to the outcome, but the fundamental difficulty is that doing so would make the weights outcome-dependent.

Asymptotic properties of the two--step minimal weights

In this section, we establish the key asymptotic properties of our proposed two--step minimal weights. We begin by deriving the convergence rates of the weights. Building on this result, we then prove the main theoretical claim of the paper: that the proposed estimator is consistent, asymptotically normal, and attains the semiparametric efficiency bound. We note that both the EIF--type and IPW--type estimators exhibit the same asymptotic properties under nearly identical assumptions. Thus, we conclude that the proposed method improves finite-sample behavior without compromising asymptotic guarantees.

Convergence rates of the two-step minimal weights

First, we investigate the convergence rates of the two--step minimal weights derived in Algorithm (ref). For the weights derived in the first step optimization problem (ref) and (ref), Theorem 2 of Wang-MinimalDispersionConsiderations-2019m directly yields the convergence rates. Specifically, it has been established that $n \widehat{w}^{\mathrm{MW}}_1(x)$ consistently estimates $1/\pi_0(x)$. For completeness, we describe the assumption and the result in the Appendix (ref). In parallel to their theorem, we obtain the following results for the weights derived in the second step optimization problem (ref) and (ref).

We begin by deriving the dual of the optimization problem.

theoremThe dual problem of (ref) and (ref) is the following unconstrained optimization problem: \begin{align} \min_{\lambda} &\frac{1}{n} \sum_{i=1}^{n} \Bigl( D_i n \rho\bigl(B(M_i,X_i)^\top \lambda\bigr) - (1 - D_i) n \widehat{w}^{\mathrm{MW}}_1(X_i) B(M_i,X_i)^\top \lambda \Bigr) + |\lambda|^\top \delta, \end{align} where \(\lambda = (\lambda_1, \dots, \lambda_L)^\top\) is the vector of dual variables associated with the \(L\) balancing constraints $B(M_i,X_i) = \bigl(b_1(M_i,X_i), b_2(M_i,X_i), \dots, b_L(M_i,X_i)\bigr)^\top$, \(\delta = (\delta_1, \dots, \delta_L)^\top\) are tolerances, and \(\rho(t) = (f')^{-1}(t)t - f((f')^{-1}(t))\). Moreover, the primal solution \(\widehat{w}^{\mathrm{MW}}_2(M_i,X_i)\) satisfies \[ \widehat{w}^{\mathrm{MW}}_2(M_i,X_i) = \rho'\Bigl( B(M_i,X_i)^\top \lambda^\dagger \Bigr) \quad (i = 1, \dots, n), \] where \(\lambda^\dagger\) is the solution to the dual optimization problem.

The proof is provided in Appendix (ref). First, note that problem (ref) can be viewed as a shrinkage estimation problem. The weights are estimated by a generalized linear model on the basis functions $B(M_i,X_i)$ with link function $\rho'$. The dual variables in $\lambda$ can be interpreted as the coefficients of the basis functions. From an optimization perspective, they represent the shadow prices of the covariate balance constraints: a larger dual variable indicates that balancing the associated covariate is more important or more difficult, and that relaxing the constraint would lead to a larger reduction in the objective function. The inclusion of an $\ell_1$ penalty further mitigates the influence of covariates that are particularly hard to balance, thereby preventing the weights from becoming overly dependent on them. In addition, this formulation helps ensure the nonnegativity of the weights. For example, if we set $f(w) = w \log w$, then $\rho'(t) = \exp(t - 1)$, which guarantees that the resulting weights are nonnegative.

By utilizing Theorem (ref), we establish the convergence rates of the weights. To this end, the following conditions are assumed.

assumptionThe following conditions hold: \begin{enumerate}[label=(\roman*),ref=(\roman*)] • The minimizer \[ \lambda^\circ = \arg\min_{\lambda \in \Lambda} \mathbb{E} \Bigl[D n \rho\bigl( B(M,X)^\top \lambda \bigr) - (1 - D) n \widehat{w}^{\text{MW}}_1(X) B(M,X)^\top \lambda \Bigr] \] is unique, where \(\Lambda\) is the compact parameter space for \(\lambda\). \(\lambda^\circ \in \text{int}(\Lambda)\), where $\text{int}(\cdot)$ stands for the interior of a set; • There exists a constant \(0 < c_0 < 1/2\) such that \[ c_0 \leq n \rho'(v) \leq 1 - c_0 \] for any \(v = B(M,X)^\top\lambda\) with \(\lambda \in \text{int}(\Lambda)\); also, there exist constants \(0 < c_1 < c_2 \) such that \[ 0 < c_1 \leq n \rho''(v) \leq c_2 \] for any \(v = B(M,X)^\top\lambda\) with \(\lambda \in \text{int}(\Lambda)\); • There exist constants $c>0$ and $C<\infty$ such that \[ \sup_{(m,x) \in \mathcal{M} \times \mathcal{X}} \|B(m,x)\|_2 \leq C L^{1/2}, \quad \|B(M,X)\|_{P,2} \leq C L^{1/2}, \] and \[ c \le \lambda_{\min}\left(\mathbb{E}[B(M,X) B(M,X)^\top]\right) \le \lambda_{\max}\left(\mathbb{E}[B(M,X) B(M,X)^\top]\right) \le C, \] with $\mathbb{E}[B(M,X) B(M,X)^\top] \preceq C I.$ Here, for a symmetric matrix $A$, $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$ denote its minimum and maximum eigenvalues, respectively, and $A \preceq C I$ means that $C I - A$ is positive semidefinite. $I$ denotes the identity matrix. Also, we denote \(\| f(Z) \|_{P,2} = (\mathbb{E}[\|f(Z)\|_2^2])^{1/2} = \left(\int \|f(z)\|_2^2 dF_Z(z) \right)^{1/2} \). • The number of basis functions \(L\) satisfies $L \to \infty$ as $n \to \infty$, \(L^2 \log L = o(n)\) and $L K^{- r_1} = o(1)$; • There exist constants \(r_2 > 1\) and \(\lambda^*\) such that the true propensity score function satisfies \[ \sup_{(m,x) \in \mathcal{M} \times \mathcal{X}} \Bigl| w^*_2(m,x) - n\rho'(B(m,x)^\top \lambda^*) \Bigr| = O\bigl(L^{-r_2}\bigr), \] where \(w^*_2(m,x) = \xi_0(m,x) / \pi_0(x) \xi_1(m,x) \); • \( \|\delta\|_\infty = o_p((nL)^{-1/2}) \) and so $\|\delta\|_2 = o_p(n^{-1/2})$. \end{enumerate}

These assumptions are analogous to Assumption 1 of Wang-MinimalDispersionConsiderations-2019m. In Assumption (ref), conditions (ref) represent standard regularity assumptions that ensure the consistency of M-estimators. Given the result that $\widehat{m}_1^{\mathrm{MW}}(X)$ is consistent for $1/\pi_0(X)$, and using the first-order condition together with Proposition (ref), we can deduce that there exists a $\lambda^\circ$ such that $w^*_2(m,x) = n \rho'(B(m,x)^\top \lambda^\circ)$. Condition (ref) enables the consistency of $\lambda^\dagger$ to imply the consistency of the estimated weights. Noting that $\rho'(t) = (f')^{-1}(t)$ by a simple calculation, it is clear that this condition is met by standard choices of $f$ used in the definition of two--step minimal weights (ref), such as the squared loss $f(w) = w^2$ and the entropy-based loss $f(w) = w \log w$. Condition (ref) imposes a technical restriction on the magnitude of the basis functions. Condition (ref) controls the growth rate of the number of basis functions relative to the sample size. Condition (ref) ensures that the weights $w^*_2(m,x)$ can be uniformly approximated, which typically requires that the basis $B(m,x)$ is sufficiently rich. Conditions (ref)--(ref) are satisfied by a wide range of basis functions, including regression splines, trigonometric polynomials, and wavelets (see, e.g., Newey-ConvergenceRatesEstimators-1997k; Chen-Chapter76Models-2007p). Lastly, condition (ref) characterizes the degree to which the equality constraints for exact covariate balance may be relaxed by hyperparameters without compromising the asymptotic properties of the estimated weights.

Under these assumptions, we establish the convergence rate of \(\widehat{w}_2(M_i,X_i)\) to \(w_2^*(M_i,X_i)\). Since normalization makes $\widehat{w}^{\mathrm{MW}}_2(m,x)$ tend to zero as $n \to \infty$, we rescale by $n$ to obtain a nontrivial convergence rate.

theoremUnder the conditions in Assumption (ref), we have: \begin{align*} \sup_{(m,x) \in \mathcal{M} \times \mathcal{X}} \left| n \widehat{w}^{\mathrm{MW}}_2(m,x) - w_2^*(m,x) \right| &= O_p\left( \sqrt{\frac{L^2\log L}{n}} + L^{1 - r_2} + L K^{-r_1} \right) = o_p(1), \\ \left \| n \widehat{w}^{\mathrm{MW}}_2(M,X) - w_2^*(M,X) \right\|_{P,2} &= O_p \left( \sqrt{\frac{L^2 \log L}{n}} + L^{1 - r_2} + L K^{- r_1} \right) = o_p(1), \end{align*} where \(r_1\) is defined in Assumption (ref) in Appendix (ref).\footnote{When considering the norm of an estimator $\widehat{f}$ with the notation $\| \widehat{f}(Z) \|_{P,2}$, the expectation is not taken with respect to the same sample of $Z$ that was used to construct $\widehat{f}$. Instead, we conceptually introduce an independent copy $Z'$ of $Z$ and evaluate \[ \| \widehat{f}(Z') \|_{P,2} = \left( \int \|\widehat{f}(z')\|_2^2 \, dF_{Z'}(z') \right)^{1/2}. \] }

The first component of the convergence rate, \(O_p\bigl( \sqrt{L^2 \log L/n} + L^{1 - r_2}\bigr)\), parallels Theorem 2 of Wang-MinimalDispersionConsiderations-2019m, while the second component, \(O_p(L K^{-r_1})\), reflects the influence of the first-step estimation.

Asymptotic normality and semiparametric efficiency

Next, we establish the asymptotic normality of the EIF--type estimator and IPW--type estimator based on our proposed weights.

EIF--type weighting estimator

We require the following Assumption (ref).

assumptionThe following conditions hold: \begin{enumerate}[label=(\roman*),ref=(\roman*)] • $\mathbb{E}[|Y- \mu_1(M,X)|] < \infty$, $\mathbb{E}[|\mu_1(M,X) - \eta_{1,0}(X)|] < \infty$, • There exist $r_\mu, r_\eta > 1$ and $\beta, \gamma$ such that $\mu_1(m,x)$ and $\eta_{1,0}(x)$ satisfies \begin{align*} \sup_{(m,x) \in \mathcal{M} \times \mathcal{X}} |\mu_1(m,x) - B(m,x)^\top \beta| &= O(L^{-r_\mu}); \\ \sup_{x \in \mathcal{X}} |\eta_{1,0}(x) - C(x)^\top \gamma| &= O(K^{-r_\eta}). \end{align*} • For the sets of smooth functions $\mathcal{W}_1$, $\mathcal{W}_2$, $\mathcal{U}$ and $\mathcal{T}$ such that $w^*_1 \in \mathcal{W}_1$, $w^*_2 \in \mathcal{W}_2$, $\mu \in \mathcal{U}$ and $\eta \in \mathcal{T}$, \begin{align*} & \log n_{[]}(\varepsilon, \mathcal{W}_1, L_2(P)) \leq C(1/\varepsilon)^{1/k_1}, \quad \log n_{[]}(\varepsilon, \mathcal{W}_2, L_2(P)) \leq C(1/\varepsilon)^{1/k_2} \\ & \log n_{[]}(\varepsilon, \mathcal{U}, L_2(P)) \leq C(1/\varepsilon)^{1/k_\mu}, \quad \log n_{[]}(\varepsilon, \mathcal{T}, L_2(P)) \leq C(1/\varepsilon)^{1/k_\eta} \end{align*} for a positive constant $C$ and $k_1, k_2, k_\mu, k_\eta > 1/2$, with $n_{[]}(\varepsilon, S, L_2(P))$ denoting the bracketing number of the set $S$ by $\varepsilon$-brackets; Moreover, we assume that the function classes $\mathcal{W}_1, \mathcal{W}_2, \mathcal{U},$ and $\mathcal{T}$ admit envelope functions belonging to $L_2(P)$; • $n^{1/2}L^{1 - r_2 - r_\mu} = o(1)$, $n^{1/2}L^{1 -r_\mu}K^{-r_1} = o(1)$, $n^{1/2} K^{1 - r_1} L^{-r_\mu} = o(1)$ and $n^{1/2} K^{1 -r_1 - r_\eta} = o(1)$. \end{enumerate}

Assumption (ref) is analogous to that of Wang-MinimalDispersionConsiderations-2019m. Condition (ref) states the standard regularity requirements, ensuring the estimators have finite moments. Condition (ref) then requires a uniform approximation for $\mu$ and $\eta$, similar to Assumption (ref) (ref). Condition (ref) introduces a complexity constraint, requiring that the metric entropy of the classes $\mathcal{W}_1$, $\mathcal{W}_2$, $\mathcal{U}$, and $\mathcal{T}$ does not grow too rapidly as $\varepsilon \to 0$. This condition is met by H\"older classes of smoothness $s$ on any bounded convex subset of $\mathbb{R}^d$ when $ s/d > 1$ (see, e.g., vanDerVaart-WeakConvergenceStatistics-1996k). Consistent with Wang-MinimalDispersionConsiderations-2019m and Fan-OptimalCovariateEstimation-2023s, these uniform approximation rate requirements are regarded as some of the least stringent in the literature (e.g. Hirano-EfficientEstimationScore-2003o; Chan-GloballyEfficientWeighting-2016v; Chan-EfficientNonparametricEffects-2016g). Additionally, condition (ref) restricts how quickly the number of basis functions, $K$ and $L$, may increase with the sample size $n$. The permissible rate for this increase is governed by the sum of the approximation errors ($r_{2}+r_{\mu}$, $r_1 + r_{\mu}$, and $r_1 + r_{\eta}$) of the propensity score and outcome estimators. This feature, that the product structures become crucial in the asymptotics, resembles the findings of Farbmacher-CausalMediationLearning-2022c on debiased machine learning for causal mediation.

Finally, we can establish the asymptotic distribution of the EIF--type estimator and show that it achieves the semiparametric efficiency bound.

theoremSuppose that Assumptions (ref) and (ref) hold. Furthermore, we assume that $\widehat{\mu}_1(M_i,X_i)$ and $\widehat{\eta}_{1,0}(X_i)$ are linear combinations of $B(M_i,X_i)$ and $C(X_i)$. Then, $$ \sqrt{n}\left(\widehat{\theta}^{\mathrm{EIF-MW}}_{1,0} - \theta_{1,0} \right) \to \mathcal{N}(0,V_{opt}),$$ where $V_{opt}$ is the semiparametric efficiency bound given by the efficient influence function: $$ \frac{D_i \xi_0(M_i,X_i)}{\pi_0(X_i)\xi_1(M_i,X_i)} \left(Y_i- \mu_1(M_i,X_i)\right) + \frac{1 - D_i}{\pi_0(X_i)} \left(\mu_1(M_i,X_i) - \eta_{1,0}(X_i) \right) + \eta_{1,0}(X_i) - \theta_{1,0}. $$ The variance can be estimated by the sample variance of the efficient influence function, in which the weights, which consist of propensity scores, are simply replaced by two--step minimal weights. Specifically, \begin{align*} \widehat{V}_{opt}^{\mathrm{EIF-MW}} &= \frac{1}{n}\sum_{i=1}^n \Biggl[ D_i n \widehat{w}_2^{\mathrm{MW}}(M_i,X_i)\Bigl(Y_i - \widehat{\mu}_1(M_i,X_i)\Bigr) \\ &\quad + (1 - D_i) n \widehat{w}_1^{\mathrm{MW}}(X_i)\Bigl(\widehat{\mu}_1(M_i,X_i) - \widehat{\eta}_{1,0}(X_i)\Bigr) + \widehat{\eta}_{1,0}(X_i) - \widehat{\theta}^{\mathrm{EIF-MW}}_{1,0} \Biggr]^2 , \end{align*}

IPW--type weighting estimator

For the IPW--type estimator, we obtain the same asymptotic result.

theoremSuppose that Assumptions (ref) and (ref) hold. Then, $$ \sqrt{n}\left( \widehat{\theta}_{1,0}^{\mathrm{IPW-MW}} - \theta_{1,0} \right) \to \mathcal{N}(0,V_{opt}), $$ where $V_{opt}$ is the semiparametric efficiency bound given by the efficient influence function. The variance can be estimated consistently by the following estimator: \begin{align*} \widehat{V}_{opt}^{\mathrm{IPW-MW}} &= \frac{1}{n}\sum_{i=1}^n \Biggl[ D_i n \widehat{w}_2^{\mathrm{MW}}(M_i,X_i)\Bigl(Y_i - \widehat{\mu}_1(M_i,X_i)\Bigr) \\ &\quad + (1 - D_i) n \widehat{w}_1^{\mathrm{MW}}(X_i)\Bigl(\widehat{\mu}_1(M_i,X_i) - \widehat{\eta}_{1,0}(X_i)\Bigr) + \widehat{\eta}_{1,0}(X_i) - \widehat{\theta}^{\mathrm{IPW-MW}}_{1,0} \Biggr]^2 , \end{align*} where \begin{align*} \widehat{\mu}_1(M_i,X_i) &= B(M_i,X_i)^\top \left( \frac{1}{n}\sum_{i=1}^n D_i B(M_i,X_i) B(M_i,X_i)^\top \right)^{-1} \frac{1}{n}\sum_{i=1}^n D_i B(M_i,X_i) Y_i \\ \widehat{\eta}_{1,0}(X_i) &= C(X_i)^\top \left( \frac{1}{n}\sum_{i=1}^n (1-D_i)C(X_i) C(X_i)^\top \right)^{-1} \frac{1}{n}\sum_{i=1}^n (1-D_i)C(X_i) \widehat{\mu}_1(M_i,X_i) \end{align*}

Simulation Studies

In this section, we conduct two Monte Carlo simulation studies to evaluate the performance of our proposed two–step minimal weights against several competing estimators. In addition, we implement the hyperparameter tuning of the balancing tolerances (Algorithm (ref)).

Tchetgen Tchetgen and Shpitser (2012) setting

First, we consider a simulation setting adapted from TchetgenTchetgen-SemiparametricTheoryAnalysis-2012y. The data are generated as follows:

align*[align* omitted — 603 chars of source]

We denote $\text{Bernoulli}(p)$ as a Bernoulli distribution where the variable takes the value 1 with probability \( p \), and $\mathcal{N}(0,1)$ as a normal distribution with mean zero and variance one. In this data-generating process, the covariates $ \{ X_1, \dots, X_4 \}$ are complex nonlinear transformations of the underlying covariates $ \{ Z_1, \dots, Z_4 \}$. The estimands are $\text{NDE}(0)$ and $\text{NDE}(1)$. The true values of both direct effects are $1$.

We compare the performance of several weighting methods when applied to the EIF--type and IPW--type estimators for NDE. This allows us to assess the performance of the weights both in conjunction with an outcome model and independently. The weighting methods are as follows:

itemize• EIF weights: These are the EIF--based weights defined in (ref), where the nuisance parameters are estimated using parametric methods such as logistic regression and linear regression. • EIF trimmed: A variant of the EIF--based weights that trims the propensity scores (or their products) to lie within the interval \([0.01, 0.99]\). • MW: Two--step minimal weights obtained via Algorithm (ref). As a dispersion penalty, we choose the entropy penalty $f(w) = w \log w$. • CBPS: Covariate balancing propensity scores obtained via Algorithm (ref) in Appendix (ref). • True PS: EIF--based weights computed by substituting the estimated propensity scores with the true propensity scores.

All weights are normalized to sum to one. To specify the estimators, we introduce the following notation:

itemize• Weights with exchanged treatment status ($\widetilde{w}_1, \widetilde{w}_2$): In addition to the standard weights $\widehat{w}_1$ and $\widehat{w}_2$ described above, our estimators also require a second set of weights defined under exchanged treatment status. We denote these by $\widetilde{w}_1$ and $\widetilde{w}_2$, obtained by swapping the roles of the treatment ($D=1$) and control ($D=0$) groups in the original weight definitions. For EIF weights, we define $\widetilde{w}_1(X_i) = 1/\widehat{\pi}_1(X_i)$ and $\widetilde{w}_2(M_i, X_i) = \widehat{\xi}_1(M_i, X_i) / (\widehat{\pi}_1(X_i)\widehat{\xi}_0(M_i, X_i))$. For the two--step minimal weights, these are constructed by applying Algorithm (ref) with treatment and control reversed: Step 1 is applied to the treatment group to balance covariates against the full population, followed by Step 2 applied to the control group.

Then, the EIF--type estimators for $\text{NDE}(0)$ and $\text{NDE}(1)$ are given by the following two equations, respectively:

align*[align* omitted — 840 chars of source]

where the terms $\sum_{i=1}^n (1 - D_i) \widetilde{w}_1 (Y_i - \widehat{m}_0(X_i)) - 1/n \sum_{i=1}^n \widehat{m}_0(X_i)$ and $\sum_{i=1}^n D_i \widehat{w}_1 (Y_i - \widehat{m}_1(X_i)) + 1/n \sum_{i=1}^n \widehat{m}_1(X_i)$ are the estimators for $\mathbb{E}[Y_0]$ and $\mathbb{E}[Y_1]$. The IPW--type weighting estimators for $\text{NDE}(0)$ and $\text{NDE}(1)$ are given by the following two equations, respectively:

align*[align* omitted — 291 chars of source]

where the terms $\sum_{i = 1}^n (1 - D_i) \widetilde{w}_1(X_i) Y_i$ and $\sum_{i = 1}^n D_i \widehat{w}_1(X_i) Y_i$ are the estimators for $\mathbb{E}[Y_0]$ and $\mathbb{E}[Y_1]$.

To evaluate the balance of covariates and the mediator achieved by the weights, we adopt the target absolute standardized mean difference (TASMD) from Chattopadhyay-BalancingVsPractice-2020m. The TASMDs for the first and second weighting steps are defined as follows: $$

alignedTASMD_cp(X) = \frac{\lvert\bar{X}_{\widehat{w}_1,c} - \bar{X}_{full}\rvert}{s_c}, \quad TASMD_tc(X) = \frac{\lvert\bar{X}_{\widehat{w}_2,t} - \bar{X}_{\widehat{w}_1,c}\rvert}{s_t}.

$$ Here, $\bar{X}_{w_1,c}$ and $\bar{X}_{w_2,t}$ are the weighted means of a covariate $X$ in the control group (using weights $\widehat{w}_1$) and the treatment group (using weights $\widehat{w}_2$), respectively. $\bar{X}_{full}$ is the unweighted mean of $X$ in the full population. The denominators $s_c$ and $s_t$ are the unweighted sample standard deviations of $X$ in the control and treated groups, respectively.

We implement the Monte Carlo simulation with $2000$ iterations in two settings described below.

itemize• All models and the basis functions for balancing constraints use $\{Z_1, Z_2, Z_3, Z_4\}$. • The outcome regression models for $\mu$ and $\eta$ use $\{X_1, X_2, X_3, X_4\}$, whereas the propensity score models for $\pi$ and $\eta$ as well as the basis functions for balancing constraints use $\{X_1, X_2, X_3, X_4, Z_1, Z_2, Z_3, Z_4\}$.

Figures (ref) and (ref) display the Monte Carlo means and variances for each estimator. Specifically, Figure (ref) presents the results for $\text{NDE}(1)$, while Figure (ref) corresponds to $\text{NDE}(0)$. We first examine Setting (A), which is depicted in the first two columns of the plots in both figures. In this setting, for both $\text{NDE}(1)$ and $\text{NDE}(0)$, the well-specified outcome model masks the underlying differences in weight quality for the EIF-type estimators, resulting in similarly strong performance across all methods. The IPW--type estimators, however, isolate the direct impact of the weighting schemes. In this context, our proposed MW is the only method that remains unbiased with substantially low variance. Notably, both standard EIF weights and even the True PS-based weights result in unignorable bias and large variance. While trimming the EIF weights reduces variance, it fails to correct the bias. This result highlights that even a perfectly specified propensity score model is insufficient when sample-level covariate balance is not enforced. It also reveals the inherent instability of inverting propensity scores.

Next, we examine Setting (B), which is displayed in the remaining columns of Figures (ref) and (ref). These results demonstrate the proposed method's robustness to outcome model misspecification. For the EIF--type estimators, our proposed MW method continues to yield unbiased estimates with low variance, performing as well as it did in setting (A). In contrast, the estimators using standard EIF weights and True PS weights are now substantially biased, as their weights fail to correct for the misspecified outcome model. A similar pattern of MW's superior performance is observed for the IPW--type estimators.

figure[figure omitted — 627 chars of source]
figure[figure omitted — 286 chars of source]

The covariate balances of each weighting method are illustrated in Figure (ref) and Figure (ref) for setting (B) with $n=1000$. Figure (ref) presents the TASMD for the first-step weights, comparing the weighted control group to the full population. Figure (ref) shows the TASMD for the second-step weights, comparing the weighted treated group to the weighted control group. In both steps, our proposed MW achieves a TASMD of virtually zero across all specified covariates $\{ Z_1, \dots, Z_4, X_1, \dots, X_4 \}$ and the mediator $M$. This demonstrates that MW successfully enforces perfect sample-level balance as designed. In contrast, all competing methods, including EIF, EIF trimmed, CBPS, and even weights based on True PS, exhibit imbalances for several covariates, as indicated by their non-zero TASMD values.

figure[figure omitted — 725 chars of source]
remarkAs Table (ref) shows, in setting (A), the results of MW closely resemble those from regression imputation (RI) using $\{Z_1, Z_2, Z_3, Z_4\}$ as regressors. In setting (B), the results are similarly close to those from regression imputation with $\{X_1, X_2, X_3, X_4, Z_1, Z_2, Z_3, Z_4\}$ as regressors. These patterns are consistent with the findings of BrunsSmithDukesFellerOgburn2025, which show that a balancing weights estimator with exact zero constraints can be expressed as the sum of a regression imputation term and an approximation error. In our simulation, the approximation error is possibly small, making the two approaches appear nearly equivalent.
remarkWe include both $\{Z_1, Z_2, Z_3, Z_4\}$ and $\{ X_1, X_2, X_3, X_4 \}$ in the weight models due to the bias decomposition discussed in Section (ref), which suggests that the covariates used in the regression models should also be incorporated into the weighting models. For instance, if the balancing constraints of MW and CBPS target $\{ Z_1, Z_2, Z_3, Z_4 \}$, while the regression estimators are linear combinations of $\{ X_1, X_2, X_3, X_4 \}$, this mismatch introduces a bias that does not vanish asymptotically (see Proposition (ref), Proposition (ref) and Appendix (ref)).

\FloatBarrier

Wong and Chan (2018) setting

In light of Remark (ref), we now demonstrate the utility of MW in a different setting. To this end, we adapt the simulation design of Wong-Kernel-basedCovariateStudies-2018k to a causal mediation framework. The data are generated as follows:

align*[align* omitted — 633 chars of source]

Note that $\text{NDE}(1) = \text{NDE}(0) = 0$ because the treatment effect arises only through interactions with covariates. In this setting, a regression imputation model that omits the interaction terms between the treatment and the covariates is misspecified and thus expected to perform poorly. We consider the following two settings:

itemize• All models and the balancing constraints use covariates $\{Z_1, Z_2, Z_3, Z_4\}$ without interaction terms. • Outcome regression models use $\{X_1, X_2, ..., X_{10}\}$, while propensity score $\pi$ and the balancing constraint in the first step use $\{ X_1, X_2, ..., X_{10}, Z_1, ..., Z_4 \}$ and propensity score $\xi$ and the balancing constraint in the second step use $\{X_1, X_2, ..., X_{10}, Z_1, ..., Z_4, M \times Z_1, M \times Z_2, M \times Z_3, M \times Z_4\}$.

The performance of all weighting estimators in these simulation settings is presented in Figures (ref) and (ref). In setting (A), the performance of all estimators is degraded by the outcome model misspecification. While our proposed MW estimator consistently achieves the lowest variance, it exhibits a non-trivial bias when estimating $\text{NDE}(0)$.

Conversely, the advantages of the MW estimator are demonstrated in setting (B). In this setting, the propensity score models and the balancing constraints are enriched with the true covariates \{$Z_1, \dots, Z_4$\} and their interactions with the mediator \{$M \times Z_1, \dots, M \times Z_4$\}, even though the main outcome models remain misspecified. As a result, MW substantially outperforms all competing estimators. It is the only method to simultaneously achieve low bias and variance for both $\text{NDE}(1)$ and $\text{NDE}(0)$, as indicated by the bolded values in the table. The covariate balance diagnostics for setting (B) with $n = 1000$ in Figures (ref) and (ref) support these results. The figures clearly show that only MW achieves near-perfect balance in both weighting steps, whereas all other methods exhibit imbalance.

For comparison, we consider a regression imputation (RI) estimator; the results are shown in Figures (ref) and (ref). As expected, the RI estimator is severely biased in setting (A) due to its inability to capture the treatment-covariate interactions. Notably, this poor performance occurs even though the RI model utilizes the same set of covariates as the MW estimator, highlighting the weakness of the simple linear imputation approach in this setting. In setting (B), when the RI model is correctly specified to include these interactions, its performance becomes comparable to that of our MW estimator.

In summary, these results provide a compelling case for the proposed MW estimator. In setting (A), where the simple RI estimator fails due to model misspecification, MW performs comparably to the EIF--type estimator and offers improvement over RI. Conversely, in setting (B), where a correctly specified RI model performs well, MW matches its high level of accuracy and efficiency, with both methods outperforming the EIF--type estimator. This demonstrates that the estimator with MW matches or exceeds the performance of the best alternative method in both scenarios.

figure[figure omitted — 269 chars of source]
figure[figure omitted — 269 chars of source]
figure[figure omitted — 704 chars of source]

\FloatBarrier

Hyperparameter tuning

Finally, we conduct a simulation study to evaluate the hyperparameter tuning procedure detailed in Algorithm (ref). We focus on setting (B) from the Wong and Chan simulation. The Monte Carlo simulation is performed with a sample size of $n = 500$ and $100$ iterations. For the tuning algorithm, we construct a grid of $500$ points for the hyperparameters $\varepsilon$ and $\delta$, and we employ bootstrap resampling with $R = 50$ repetitions.

The results, presented in Table (ref), compare the performance of the MW estimator under exact balancing (i.e., $\varepsilon=0, \delta=0$) against the estimator with optimally chosen hyperparameters. The tuning algorithm leads to a reduction in bias across all four estimator configurations. This reduction is a consequence of the algorithm preferring approximate balance over exact balance, as illustrated by the TASMD plots in Figures (ref) and (ref), which show that the selected hyperparameters yield small but non-zero balance errors. For the EIF--type estimators, the reduction in bias is accompanied by a slight increase in variance, resulting in a marginally higher mean squared error (MSE). In contrast, for the IPW--type estimators, the tuning algorithm successfully reduces both bias and variance, leading to an improved overall MSE.

table[table omitted — 901 chars of source]
figure[figure omitted — 858 chars of source]

Empirical Application: The Effect of Media Framing on Immigration Attitudes

In this section, we analyze the framing dataset using our proposed two--step minimal weights estimator. The dataset originates from brader2008triggers, who investigated how media framing of immigration news influences public opinion and political behavior. It is available in the R package mediation (R-mediation). In their experiment, participants were randomly assigned to read an article about immigration and subsequently answered a series of follow-up questions from which the outcome, mediators, and covariates were measured. Our objective is to assess the causal pathways through which framing influences information-seeking behavior: to what extent the effect arises directly from exposure to the media frame and ethnic cues, and to what extent it is mediated by shifts in perceived harm and negative emotional responses.

The sample consists of 265 individuals, with the main variables defined as follows:

itemize• Treatment ($D$): We constructed an indicator for whether participants were exposed to a negatively framed article featuring the photo of a Latino (Mexican) immigrant or to a positively framed article featuring the photo of a European (Russian) immigrant.\footnote{The positive versus negative framing differed in whether the article emphasized the beneficial consequences of immigration (e.g., strengthening the economy, increasing tax revenues, enriching American culture) or its harmful consequences (e.g., lowering wages, consuming public resources, eroding American values). The framing was further reinforced by portraying state governors as either welcoming or concerned about immigration and by depicting citizens as having either positive or negative experiences with immigrants. With respect to immigrant identity, the Mexican versus Russian cue was conveyed through the photograph and its caption, which read: “[Jose Sanchez/Nikolai Vandinsky] is one of thousands of new immigrants who arrived in the United States during the first half of this year.”} • Mediator 1 ($M_{\mathrm{I}}$): Perceived harm from increased immigration, measured on a 2--8 scale.\footnote{After reading the article, participants answered questions such as: (i) “How likely is it that immigration will have a negative financial impact on many Americans?” and (ii) “How likely is it that immigration will have a negative financial impact on you or your family?” The two items were summed to form the perceived-harm index. Higher values indicate greater perceived harm.} • Mediator 2 ($M_{\mathrm{II}}$): Anxiety about increased immigration, measured with a single four-point item (Very / Somewhat / A little / Not at all). • Mediator 3 ($M_{\mathrm{III}}$): Negative affect during the experiment, based on three items (anxious, worried, angry), summed to a 3--12 scale.\footnote{Both $M_2$ and $M_3$ were derived from the same emotion battery: “How [anxious/proud/angry/hopeful/worried/excited] does it make you feel when thinking about high levels of immigration?” (Very, somewhat, a little, or not at all).} • Outcome ($Y$): Whether subjects requested information from anti-immigration organizations (1 = receive, 0 = not receive).\footnote{They asked if they would like more information about immigration from a variety of sources, including nonpartisan research centers, the U.S.government, pro-immigrant groups, and anti-immigrant groups.} • Covariates ($X$): Pre‑treatment demographics (age, education level, gender, and income), along with their quadratic and cubic terms, as well as all two-way and three-way interactions (excluding higher-order terms for dummy variables), yielding 20 covariates in total.

We specify all nuisance models—the outcome regression models, the propensity score models, and the balancing constraints, using the same covariate set. We estimate ATE, NDE, and NIE using two types of estimators, EIF--type estimator and IPW--type estimator, across four different weighting schemes: EIF, EIF (trimmed), our proposed MW, and CBPS.

First, we assess the finite--sample balance of each weighting method by examining the TASMD. Figure (ref) displays the balance for the first-step weights (weighting the control group to resemble the full population), while Figure (ref) shows the balance for the second-step weights (weighting the treated group to resemble the re-weighted control group). In both figures, our proposed MW estimator achieves a TASMD of virtually zero for all 20 covariates included in the model.

figure[figure omitted — 647 chars of source]

We next examine the estimates reported in Figure (ref). The ATE is negative in all specifications. In other words, participants exposed to a positively framed story with a European (Russian) immigrant were more likely to request information from anti-immigration organizations than those exposed to a negatively framed story with a Latino (Mexican) immigrant. This aligns with the findings of brader2008triggers, who suggest that European cues, especially under a positive frame, may spark curiosity and information seeking. Indeed, consistent with this explanation, both $\mathrm{NDE}(1)$ and $\mathrm{NDE}(0)$ are negative, indicating that the media cue itself shifts the outcome downward. By contrast, $\mathrm{NIE}(1)$ and $\mathrm{NIE}(0)$ are positive, suggesting that changes in mediators induced by the treatment move behavior upward. This can be explained by the observation that exposure to a negatively framed story featuring a Latino (Mexican) immigrant changes perceived harm and emotions, leading individuals to seek information from anti-immigration organizations. Taken together, this implies that the positive indirect pathway partly offsets the negative direct effect.

Most estimates are not statistically significant. The important exception is $\mathrm{NDE}(0)$: under the EIF-type specification, all weighting methods yield significant results ($p \approx 0.01$-$0.03$), with MW producing the smallest variance and the narrowest interval. Under the IPW/HT-type specification, significance appears only with MW ($p \approx 0.013$), again due to a notable reduction in variance.

figure[figure omitted — 489 chars of source]

Conclusion

This paper addressed the challenge of estimating natural direct and indirect effects in the presence of finite--sample instability and covariate imbalance, which are two critical limitations of the EIF--based estimator and IPW estimator. To this end, we introduced a novel two--step minimal weights estimator specifically tailored for the causal mediation setting.

Our theoretical contribution began with a formal decomposition of the EIF--type estimator's bias. This analysis pinpointed imbalances in the marginal covariate distribution and the joint distribution of covariates and the mediator as the primary sources of bias, thereby motivating our two--step weight estimation algorithm, which sequentially targets these specific imbalances. We established the convergence rates for the proposed weights and proved that the resulting estimator is asymptotically normal and achieves the semiparametric efficiency bound.

Our simulation studies and empirical application provided empirical support for this approach. The two--step proposed minimal weights consistently reduced both bias and variance compared to competitors, particularly in challenging scenarios with model misspecification where standard methods failed. This underscores a key implication for applied researchers: directly targeting sample-level balance and weight stability can yield substantially more reliable estimates of causal mechanisms.

As a direction for future research, we note that causal mediation analysis is closely connected to the study of long-term treatment effects and statistical surrogacy, since mediators are often interpreted as short-term outcomes or surrogate endpoints (Chen-SemiparametricEstimationEffects-2023s, Imbens-Long-termCausalCombination-2024k, athey2024estimatingtreatmenteffectsusing). This topic has attracted increasing attention in econometrics, and we expect that estimators analogous to those proposed in this paper can be developed in this setting as well.